
A first-party Blackfrost-AI BF16 model for security professionals conducting authorized research, assessment, engineering, and response work.
The release contains the complete BF16 SafeTensors checkpoint, weight index, configuration, tokenizer and processor assets, the packaged chat template, upstream license, this model card, and a documented deployment kit for the validated serving profile. It is not an adapter and does not require a separate parent checkpoint at load time.
| Field | Released artifact |
|---|---|
| Clean model name | CYBER-FROST-3.8-BF16 |
| Former name | BLACKFROST-3.8-ICED-BF16 |
| Architecture | Qwen4ExpForConditionalGeneration |
| Precision | BF16 |
| Weight layout | 131 SafeTensors shards |
| Total parameters reported by Hub metadata | approximately 180B |
| Configured context | 262,144 tokens |
| Context exercised in the published performance trial | 32,768 tokens |
| Native speculative head | one MTP layer |
| Validated modality | text |
The configuration includes a vision tower, but this release has not received a multimodal quality evaluation. Do not infer validated image or video capability from the presence of processor files.
Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope.
Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective, not a claim that the model has no refusal behavior or that every answer is safe or correct.
Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.
Cyber-Frost was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published.
Domain coverage includes:
Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are therefore operator provenance statements, not independent benchmark findings.
Training-data provenance and licensing review for the mixed-source corpus remains in progress. That review is one reason access remains manually gated.
The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding.
The configured context ceiling is not a blanket quality guarantee. The published BF16 trial exercised 32,768 tokens. Longer contexts, high concurrency, multimodal requests, and tool-heavy agent loops require separate validation.
Qwen/Qwen3.8-Flash-Next at immutable revision de4b8e4d43b917e7706784d8bb445c9af86a3540.BLACKFROST-3.8-FLASH-BF16.BLACKFROST-3.8-ICED-BF16 and is now named CYBER-FROST-3.8-BF16. The rename does not represent another training run.Tokenizer and processor lineage comes from the pinned Qwen foundation checkpoint. The packaged Blackfrost chat template is release-specific.
The released weight payload is tied to Hub revision 5321904427c4ef54df8a667edcbc2d1184e4286e. On 2026-09-15, every weight shard was checked against its corresponding Hub LFS object identifier with no mismatch. The old ICED label and the Cyber-Frost repository resolve to this same payload.
| Artifact | SHA-256 |
|---|---|
Sanitized config.json prepared for this card update | c77cfd62621660b25e4c811bf44bf2ad4f63af98f23a874952a99e4e85f189f2 |
model.safetensors.index.json | 99e815241ef03325536b0aaa4441deea45174c17fae31e10f0bb456410c590de |
| Qwen Community License file | a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08 |
The sanitized configuration removes legacy local-path metadata only; it does not alter architecture or inference behavior.
There were no API errors. The evaluation used an earlier internal suite revision that was subsequently corrected, so these results are development evidence rather than a standardized leaderboard score. The zero safe over-refusal result supports the intended low-friction behavior on that slice; the uneven harmful-request refusal rates also show why this checkpoint must not be treated as a safety control.
No standardized cyber-capability benchmark has yet been qualified for this exact BF16 checkpoint. In particular, this card does not claim a CyberMetric, SecBench, MMLU, HumanEval, or IFEval score. Refusal behavior is not a substitute for measuring security competence.
The measured profile used four NVIDIA B300 SXM6 GPUs, tensor parallelism 4, vLLM 0.29.1rc1.dev13+g1cfd97281, 32,768-token serving context, sequential requests, and thinking enabled. The frozen suite covered reasoning, code, prose, and tool-result synthesis with temperature 1.0, top-p 0.95, top-k 20, seed 38421, and at most 256 generated tokens.
| Mode | Completion tokens | Median decode | Mean decode | Median TTFT | Draft acceptance | Relative median |
|---|---|---|---|---|---|---|
| No draft | 1,760 | 190.10 tok/s | 190.51 tok/s | 0.696 s | n/a | baseline |
| Native MTP, k=2 | 1,820 | 265.00 tok/s | 264.27 tok/s | 0.660 s | 55.0% | +39.4% |
| Native MTP, k=3 | 1,799 | 263.10 tok/s | 288.43 tok/s | 0.684 s | 49.6% | +38.4% |
These are single-pass development measurements, not production throughput guarantees. Workload, concurrency, context, runtime build, cache precision, and sampling can materially change the result.
The embedded MTP tensors predate the final trunk-weight behavioral stage. Treat the MTP head as a provisional acceleration baseline rather than a freshly adapted draft head. With the tested vLLM build, native MTP plus the model's Mamba groups also disables cross-request prefix-cache reuse, and fused multi-step draft decode is unavailable for the experimental attention backend.
The repository includes a Blackfrost chat template that supplies a default operating prompt, supports caller-provided system context, exposes Qwen-style reasoning controls, and serializes tool calls. Its default is thinking enabled; supported reasoning-effort values are xhigh, medium, and low. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20.
The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.
Changing the template, system message, reasoning mode, or sampling can materially change refusal behavior and output quality. Record those settings when reporting results.
Use the included DEPLOYMENT/ kit for the validated four-GPU BF16 serving profile, environment variables, launch command, and smoke tests. The clean API model identifier is CYBER-FROST-3.8-BF16.
Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.
It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.
The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.
Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.
Report reproducible model or packaging issues through the repository's Discussions page without including secrets, client data, live targets, or sensitive exploit details.

A first-party Blackfrost-AI BF16 model for security professionals conducting authorized research, assessment, engineering, and response work.
The release contains the complete BF16 SafeTensors checkpoint, weight index, configuration, tokenizer and processor assets, the packaged chat template, upstream license, this model card, and a documented deployment kit for the validated serving profile. It is not an adapter and does not require a separate parent checkpoint at load time.
| Field | Released artifact |
|---|---|
| Clean model name | CYBER-FROST-3.8-BF16 |
| Former name | BLACKFROST-3.8-ICED-BF16 |
| Architecture | Qwen4ExpForConditionalGeneration |
| Precision | BF16 |
| Weight layout | 131 SafeTensors shards |
| Total parameters reported by Hub metadata | approximately 180B |
| Configured context | 262,144 tokens |
| Context exercised in the published performance trial | 32,768 tokens |
| Native speculative head | one MTP layer |
| Validated modality | text |
The configuration includes a vision tower, but this release has not received a multimodal quality evaluation. Do not infer validated image or video capability from the presence of processor files.
Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope.
Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective, not a claim that the model has no refusal behavior or that every answer is safe or correct.
Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.
Cyber-Frost was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published.
Domain coverage includes:
Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are therefore operator provenance statements, not independent benchmark findings.
Training-data provenance and licensing review for the mixed-source corpus remains in progress. That review is one reason access remains manually gated.
The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding.
The configured context ceiling is not a blanket quality guarantee. The published BF16 trial exercised 32,768 tokens. Longer contexts, high concurrency, multimodal requests, and tool-heavy agent loops require separate validation.
Qwen/Qwen3.8-Flash-Next at immutable revision de4b8e4d43b917e7706784d8bb445c9af86a3540.BLACKFROST-3.8-FLASH-BF16.BLACKFROST-3.8-ICED-BF16 and is now named CYBER-FROST-3.8-BF16. The rename does not represent another training run.Tokenizer and processor lineage comes from the pinned Qwen foundation checkpoint. The packaged Blackfrost chat template is release-specific.
The released weight payload is tied to Hub revision 5321904427c4ef54df8a667edcbc2d1184e4286e. On 2026-09-15, every weight shard was checked against its corresponding Hub LFS object identifier with no mismatch. The old ICED label and the Cyber-Frost repository resolve to this same payload.
| Artifact | SHA-256 |
|---|---|
Sanitized config.json prepared for this card update | c77cfd62621660b25e4c811bf44bf2ad4f63af98f23a874952a99e4e85f189f2 |
model.safetensors.index.json | 99e815241ef03325536b0aaa4441deea45174c17fae31e10f0bb456410c590de |
| Qwen Community License file | a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08 |
The sanitized configuration removes legacy local-path metadata only; it does not alter architecture or inference behavior.
There were no API errors. The evaluation used an earlier internal suite revision that was subsequently corrected, so these results are development evidence rather than a standardized leaderboard score. The zero safe over-refusal result supports the intended low-friction behavior on that slice; the uneven harmful-request refusal rates also show why this checkpoint must not be treated as a safety control.
No standardized cyber-capability benchmark has yet been qualified for this exact BF16 checkpoint. In particular, this card does not claim a CyberMetric, SecBench, MMLU, HumanEval, or IFEval score. Refusal behavior is not a substitute for measuring security competence.
The measured profile used four NVIDIA B300 SXM6 GPUs, tensor parallelism 4, vLLM 0.29.1rc1.dev13+g1cfd97281, 32,768-token serving context, sequential requests, and thinking enabled. The frozen suite covered reasoning, code, prose, and tool-result synthesis with temperature 1.0, top-p 0.95, top-k 20, seed 38421, and at most 256 generated tokens.
| Mode | Completion tokens | Median decode | Mean decode | Median TTFT | Draft acceptance | Relative median |
|---|---|---|---|---|---|---|
| No draft | 1,760 | 190.10 tok/s | 190.51 tok/s | 0.696 s | n/a | baseline |
| Native MTP, k=2 | 1,820 | 265.00 tok/s | 264.27 tok/s | 0.660 s | 55.0% | +39.4% |
| Native MTP, k=3 | 1,799 | 263.10 tok/s | 288.43 tok/s | 0.684 s | 49.6% | +38.4% |
These are single-pass development measurements, not production throughput guarantees. Workload, concurrency, context, runtime build, cache precision, and sampling can materially change the result.
The embedded MTP tensors predate the final trunk-weight behavioral stage. Treat the MTP head as a provisional acceleration baseline rather than a freshly adapted draft head. With the tested vLLM build, native MTP plus the model's Mamba groups also disables cross-request prefix-cache reuse, and fused multi-step draft decode is unavailable for the experimental attention backend.
The repository includes a Blackfrost chat template that supplies a default operating prompt, supports caller-provided system context, exposes Qwen-style reasoning controls, and serializes tool calls. Its default is thinking enabled; supported reasoning-effort values are xhigh, medium, and low. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20.
The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.
Changing the template, system message, reasoning mode, or sampling can materially change refusal behavior and output quality. Record those settings when reporting results.
Use the included DEPLOYMENT/ kit for the validated four-GPU BF16 serving profile, environment variables, launch command, and smoke tests. The clean API model identifier is CYBER-FROST-3.8-BF16.
Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.
It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.
The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.
Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.
Report reproducible model or packaging issues through the repository's Discussions page without including secrets, client data, live targets, or sensitive exploit details.