Scion-35B-A3B: ternary MoE experts + trained corrections
8
13 commits
6 linked in READMEs
updated Sep 30, 2026
GitHub: Build-scripts and docs | Forensics study | Runtime fork | Discussions
A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.
2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size
The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.
--lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing.Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14–1.37× faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).TAIL-EXPERIMENT-PLAN.md).sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, the ffn_moe_out virtual target and embedded-adapter support.| Item | Specification |
|---|---|
| Base model | empero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0) |
| Parameters | 34.9B total, about 3B active per token (256 experts, top-8 plus shared) |
| Architecture | qwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward |
| Context length | 262,144 tokens (inherited from the base model) |
| Weight format | Ternary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512) |
| Low-bit coverage | Expert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32 |
| Deployed size | 10.558 GiB / 11.337 GB (text only, single file) |
| Backends | llama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested |
| License | Apache-2.0 (inherited from the base) |
Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per
group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the
model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output).
The corrections are rank-512 low-rank branches on the attention output and the
MoE block output, plus exact router deltas. They ship in the compact legacy
q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are
merged into the file. Effective overall: 2.61 bpw.
| Format | bpw | Size | vs BF16 |
|---|---|---|---|
| BF16 (reference) | 16.38 | 71.07 GB | 1.0x |
| Q8_0 | 8.72 | 37.80 GB | 1.9x |
| Q4_K_M | 5.01 | 21.71 GB | 3.3x |
| IQ2_M | 2.90 | 12.56 GB | 5.7x |
| Scion-35B-A3B | 2.61 | 11.34 GB | 6.3x |
Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).
| Component | Pack | Size | Residency |
|---|---|---|---|
| Language model (this repo) | PQ2_0 experts, Q8_0 rest, embedded corrections | 11.34 GB | resident; the whole model |
| k=1 MTP drafter (separate repo) | fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding | 50 MB | transient; ~1 ms/eval on GPU |
| Uncorrected body (not uploaded) | PQ2_0 experts and Q8_0 rest | 10.46 GiB | for swap tests |
| Corrections (not uploaded) | rank-512 branches and router deltas (q1_0_g128) | 98 MiB | for swap tests |
The language model is the single released body file; the drafter is a separate
50 MB file in its own repo
(Scion-35B-A3B-mtp-drafter).
The two-file variant (body plus separate adapter) exists for reproducing the
merge and swapping corrections at runtime. Ask in Discussions if you want it.
About the Hub's quant chip. The Hub parses file names and labels this file
Q2_0; the same happens on Prism ML's ownPQ2_0releases. The container is Prism'sPQ2_0(legacy nameQ1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64Q2_0(type 42). No single quant name fits the file anyway, because it is a mix:PQ2_0expert banks, the embedded corrections in the legacyq1_0_g128container, andQ8_0for the rest (norms and routers in F32).
Recommended values, from the base model card:
temperature=0.6,top_p=0.95,top_k=20
This is a reasoning distill, and answers open with a long thinking segment.
Allow generous max_new_tokens (for example -n 16384); a small cap ends
generation mid-thought, before any answer.
A simple prompt works, for example You are a helpful assistant. The base is a
reasoning SFT distill and does not require a special system prompt.
-ngl 99 on a single card. When VRAM is tight, offload experts to CPU with
-ncmoe, which is the VRAM-budget dial; the ternary container roughly halves
the CPU-tail penalty compared to an f16 expert bank.-t 8 on an 8C/16T CPU; SMT siblings
collapse CPU expert throughput).The runtime is the fork, and the fork is the source of truth for running these files:
sky-is-green/prism-ml-llama.cpp, branchmoe-corr-runtime.
The PQ2_0 container, the legacy Q1_0_g128 import, the ffn_moe_out virtual
target and embedded adapters all live in the fork. Stock llama.cpp will not
run this file: it treats PQ2_0 and Q1_0_g128 as unknown tensor types.
Upstream's own Q2_0 (type 42, g64) is a different container and is not a
substitute.
The fork's default branch is the one you want, so a plain clone is enough.
# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON CPU-only: no flag
If the model will not load:
tensor '...' has invalid ggml type 142. should be in [0, 43)means the binary you ran is a build of upstreammaster, which knows 43 types and has never heard ofPQ2_0(type id 142). The branch above definesGGML_TYPE_PQ2_0 = 142andGGML_TYPE_COUNT = 144, and./verify-container-support.shchecks exactly that in a second. The same error also comes from a stalebuild/directory or from an olderllama-cliearlier on yourPATH; the reliable fix for all three is to delete the checkout and the build directory and start from the clone above. See also the fork README.
# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-ngl 99 -c 4096 -t 8 \
--temp 0.6 --top-p 0.95 --top-k 20 \
-p "Explain quantum computing in simple terms." -n 16384
# server
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080
-ngl 99 offloads every layer (0 is CPU-only), -c sets context up to
262144, and -t should be the physical core count. Verified on CPU and ROCm
(gfx1100, RX 7900 XT, where the release was built and measured). A CUDA build is
expected to work; reports are welcome in Discussions.
The drafter lives in its own repo:
Scion-35B-A3B-mtp-drafter
(50 MB). It is a small frozen-body head that predicts the model's own next
token from its post-norm hidden state plus the next token's embedding; it reuses
the model's own output_norm/output/token_embd, so it adds no second
vocabulary projection. It is not a stock llama.cpp drafter: the fork's
draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and
runs it entirely against the target context, with no draft model.
hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .
# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080
# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8
Measured on one RX 7900 XT (greedy):
| Workload | Baseline | + drafter | Speedup | Draft acceptance |
|---|---|---|---|---|
| Bench prompt, 400 tokens | 12.41 ms/tok | 9.07 ms/tok | 1.37× | 0.738 |
| Open-ended CLI generation | 49.0 tok/s | 57.2 tok/s | 1.17× | 0.46–0.74 |
| Server confirmation | — | — | 1.14–1.21× | 0.44–0.70 |
Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent — treat ~1.2× as the typical figure and 1.37× as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.
Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks);
KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and
Winogrande 400 zero-shot (about ±2% CI). BF16 and all quants were measured in one
H100 session; this release was measured on the local card and cross-checked
against the pod (IQ2_M KLD local vs pod: 0.2%).
| Variant | Size | bpw | HellaSwag | Winogrande |
|---|---|---|---|---|
| BF16 (reference) | 71.07 GB | 16.38 | 81.25 | 76.00 |
| Q4_K_M | 21.71 GB | 5.01 | 80.00 | 76.00 |
| IQ2_M | 12.56 GB | 2.90 | 79.00 | 75.75 |
| Q2_K | 13.84 GB | 3.19 | 76.50 | 73.25 |
| Scion-35B-A3B | 11.34 GB | 2.61 | 79.00 | 76.25 |
Within ±2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.
The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added
the remaining standard multiple-choice tasks on the same release and the same
fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):
| Task | Scion-35B-A3B |
|---|---|
| ARC-Challenge (299) | 57.19 ±2.87 |
| ARC-Easy (570) | 80.00 ±1.68 |
| MMLU (2000, seed 1234) | 40.15 ±1.10 |
| TruthfulQA (817) | 32.56 ±1.64 |
| wikitext-2 PPL (c512, 580 chunks) | 8.2731 ±0.054 |
The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the
grid's 8.3539; the model is unchanged.
| Variant | PPL (lower better) | KLD mean vs BF16 (lower better) | KLD 99.9% |
|---|---|---|---|
| BF16 | 7.1595 | — | — |
| Q4_K_M | 7.2354 | 0.0314 | 1.141 |
| IQ2_M | 8.4133 | 0.1636 | 3.963 |
| Q2_K | 8.4729 | 0.1493 | 3.276 |
| Scion-35B-A3B | 8.3539 | 0.2694 | 4.758 |
The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35
after) more than they improved the full-distribution tail. PPL ranks this build
first of the 2-bit class and KLD ranks it last, and that disagreement is the
research result: the training matches the teacher's top-50 logits, and the rest
of the distribution is unconstrained. Tail-aware training was attempted at
full scale (a tail-conditional KD term plus router bias plus a hard-window
curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the
teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are
nearly inert, and the retrain ties this release on every task while PPL regresses
~2%. The gap stands, and the negative is recorded in the project's register.

All ten community quants plus BF16; up is better, left is smaller. Full table
and method notes:
QUANT-RETENTION-35B.md.
| model | size GB | bpw | PPL | KLD mean | KLD 99.9% | HellaSwag | Winogrande |
|---|---|---|---|---|---|---|---|
| Scion-35B-A3B | 11.34 | 2.61 | 8.354 | 0.269 | 4.758 | 79.00 | 76.25 |
| IQ2_M | 12.56 | 2.90 | 8.413 | 0.164 | 3.963 | 79.00 | 75.75 |
| Q2_K | 13.84 | 3.19 | 8.473 | 0.149 | 3.276 | 76.50 | 73.25 |
| IQ3_M | 16.34 | 3.77 | 7.520 | 0.057 | 1.225 | 80.00 | 74.25 |
| Q3_K_M | 17.66 | 4.07 | 7.432 | 0.058 | 1.611 | 79.75 | 76.00 |
| IQ4_XS | 19.63 | 4.53 | 7.264 | 0.022 | 0.636 | 80.75 | 75.75 |
| Q4_K_M | 21.71 | 5.01 | 7.235 | 0.031 | 1.141 | 80.00 | 76.00 |
| Q5_K_M | 25.35 | 5.84 | 7.273 | 0.015 | 0.717 | 80.50 | 76.25 |
| Q6_K | 29.21 | 6.73 | 7.153 | 0.008 | 0.360 | 80.75 | 74.75 |
| Q8_0 | 37.80 | 8.72 | 7.160 | 0.004 | 0.209 | 80.25 | 75.50 |
| BF16 | 71.07 | 16.38 | 7.160 | — | — | 81.25 | 76.00 |
max_new_tokens
accordingly.Q2_0: it is filename-derived; see the note above, and
use the fork.@misc{scion35b2026,
title = {Scion-35B-A3B: ternary MoE experts with trained corrections},
author = {SkyIsNotGreen},
year = {2026},
month = {September},
url = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}
Apache-2.0, inherited from the base model; see LICENSE and
NOTICE.
Base weights: empero-ai and
Qwen / Alibaba Cloud (Apache-2.0);
BF16 conversion: MrFuzzihead;
container and kernels: Prism ML (MIT) with
TAARDIS conventions (MIT);
engine: llama.cpp (MIT).
Scion-35B-A3B: ternary MoE experts + trained corrections
8
13 commits
6 linked in READMEs
updated Sep 30, 2026
GitHub: Build-scripts and docs | Forensics study | Runtime fork | Discussions
A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.
2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size
The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.
--lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing.Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14–1.37× faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).TAIL-EXPERIMENT-PLAN.md).sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, the ffn_moe_out virtual target and embedded-adapter support.| Item | Specification |
|---|---|
| Base model | empero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0) |
| Parameters | 34.9B total, about 3B active per token (256 experts, top-8 plus shared) |
| Architecture | qwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward |
| Context length | 262,144 tokens (inherited from the base model) |
| Weight format | Ternary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512) |
| Low-bit coverage | Expert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32 |
| Deployed size | 10.558 GiB / 11.337 GB (text only, single file) |
| Backends | llama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested |
| License | Apache-2.0 (inherited from the base) |
Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per
group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the
model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output).
The corrections are rank-512 low-rank branches on the attention output and the
MoE block output, plus exact router deltas. They ship in the compact legacy
q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are
merged into the file. Effective overall: 2.61 bpw.
| Format | bpw | Size | vs BF16 |
|---|---|---|---|
| BF16 (reference) | 16.38 | 71.07 GB | 1.0x |
| Q8_0 | 8.72 | 37.80 GB | 1.9x |
| Q4_K_M | 5.01 | 21.71 GB | 3.3x |
| IQ2_M | 2.90 | 12.56 GB | 5.7x |
| Scion-35B-A3B | 2.61 | 11.34 GB | 6.3x |
Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).
| Component | Pack | Size | Residency |
|---|---|---|---|
| Language model (this repo) | PQ2_0 experts, Q8_0 rest, embedded corrections | 11.34 GB | resident; the whole model |
| k=1 MTP drafter (separate repo) | fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding | 50 MB | transient; ~1 ms/eval on GPU |
| Uncorrected body (not uploaded) | PQ2_0 experts and Q8_0 rest | 10.46 GiB | for swap tests |
| Corrections (not uploaded) | rank-512 branches and router deltas (q1_0_g128) | 98 MiB | for swap tests |
The language model is the single released body file; the drafter is a separate
50 MB file in its own repo
(Scion-35B-A3B-mtp-drafter).
The two-file variant (body plus separate adapter) exists for reproducing the
merge and swapping corrections at runtime. Ask in Discussions if you want it.
About the Hub's quant chip. The Hub parses file names and labels this file
Q2_0; the same happens on Prism ML's ownPQ2_0releases. The container is Prism'sPQ2_0(legacy nameQ1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64Q2_0(type 42). No single quant name fits the file anyway, because it is a mix:PQ2_0expert banks, the embedded corrections in the legacyq1_0_g128container, andQ8_0for the rest (norms and routers in F32).
Recommended values, from the base model card:
temperature=0.6,top_p=0.95,top_k=20
This is a reasoning distill, and answers open with a long thinking segment.
Allow generous max_new_tokens (for example -n 16384); a small cap ends
generation mid-thought, before any answer.
A simple prompt works, for example You are a helpful assistant. The base is a
reasoning SFT distill and does not require a special system prompt.
-ngl 99 on a single card. When VRAM is tight, offload experts to CPU with
-ncmoe, which is the VRAM-budget dial; the ternary container roughly halves
the CPU-tail penalty compared to an f16 expert bank.-t 8 on an 8C/16T CPU; SMT siblings
collapse CPU expert throughput).The runtime is the fork, and the fork is the source of truth for running these files:
sky-is-green/prism-ml-llama.cpp, branchmoe-corr-runtime.
The PQ2_0 container, the legacy Q1_0_g128 import, the ffn_moe_out virtual
target and embedded adapters all live in the fork. Stock llama.cpp will not
run this file: it treats PQ2_0 and Q1_0_g128 as unknown tensor types.
Upstream's own Q2_0 (type 42, g64) is a different container and is not a
substitute.
The fork's default branch is the one you want, so a plain clone is enough.
# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON CPU-only: no flag
If the model will not load:
tensor '...' has invalid ggml type 142. should be in [0, 43)means the binary you ran is a build of upstreammaster, which knows 43 types and has never heard ofPQ2_0(type id 142). The branch above definesGGML_TYPE_PQ2_0 = 142andGGML_TYPE_COUNT = 144, and./verify-container-support.shchecks exactly that in a second. The same error also comes from a stalebuild/directory or from an olderllama-cliearlier on yourPATH; the reliable fix for all three is to delete the checkout and the build directory and start from the clone above. See also the fork README.
# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-ngl 99 -c 4096 -t 8 \
--temp 0.6 --top-p 0.95 --top-k 20 \
-p "Explain quantum computing in simple terms." -n 16384
# server
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080
-ngl 99 offloads every layer (0 is CPU-only), -c sets context up to
262144, and -t should be the physical core count. Verified on CPU and ROCm
(gfx1100, RX 7900 XT, where the release was built and measured). A CUDA build is
expected to work; reports are welcome in Discussions.
The drafter lives in its own repo:
Scion-35B-A3B-mtp-drafter
(50 MB). It is a small frozen-body head that predicts the model's own next
token from its post-norm hidden state plus the next token's embedding; it reuses
the model's own output_norm/output/token_embd, so it adds no second
vocabulary projection. It is not a stock llama.cpp drafter: the fork's
draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and
runs it entirely against the target context, with no draft model.
hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .
# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080
# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
-md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8
Measured on one RX 7900 XT (greedy):
| Workload | Baseline | + drafter | Speedup | Draft acceptance |
|---|---|---|---|---|
| Bench prompt, 400 tokens | 12.41 ms/tok | 9.07 ms/tok | 1.37× | 0.738 |
| Open-ended CLI generation | 49.0 tok/s | 57.2 tok/s | 1.17× | 0.46–0.74 |
| Server confirmation | — | — | 1.14–1.21× | 0.44–0.70 |
Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent — treat ~1.2× as the typical figure and 1.37× as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.
Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks);
KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and
Winogrande 400 zero-shot (about ±2% CI). BF16 and all quants were measured in one
H100 session; this release was measured on the local card and cross-checked
against the pod (IQ2_M KLD local vs pod: 0.2%).
| Variant | Size | bpw | HellaSwag | Winogrande |
|---|---|---|---|---|
| BF16 (reference) | 71.07 GB | 16.38 | 81.25 | 76.00 |
| Q4_K_M | 21.71 GB | 5.01 | 80.00 | 76.00 |
| IQ2_M | 12.56 GB | 2.90 | 79.00 | 75.75 |
| Q2_K | 13.84 GB | 3.19 | 76.50 | 73.25 |
| Scion-35B-A3B | 11.34 GB | 2.61 | 79.00 | 76.25 |
Within ±2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.
The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added
the remaining standard multiple-choice tasks on the same release and the same
fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):
| Task | Scion-35B-A3B |
|---|---|
| ARC-Challenge (299) | 57.19 ±2.87 |
| ARC-Easy (570) | 80.00 ±1.68 |
| MMLU (2000, seed 1234) | 40.15 ±1.10 |
| TruthfulQA (817) | 32.56 ±1.64 |
| wikitext-2 PPL (c512, 580 chunks) | 8.2731 ±0.054 |
The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the
grid's 8.3539; the model is unchanged.
| Variant | PPL (lower better) | KLD mean vs BF16 (lower better) | KLD 99.9% |
|---|---|---|---|
| BF16 | 7.1595 | — | — |
| Q4_K_M | 7.2354 | 0.0314 | 1.141 |
| IQ2_M | 8.4133 | 0.1636 | 3.963 |
| Q2_K | 8.4729 | 0.1493 | 3.276 |
| Scion-35B-A3B | 8.3539 | 0.2694 | 4.758 |
The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35
after) more than they improved the full-distribution tail. PPL ranks this build
first of the 2-bit class and KLD ranks it last, and that disagreement is the
research result: the training matches the teacher's top-50 logits, and the rest
of the distribution is unconstrained. Tail-aware training was attempted at
full scale (a tail-conditional KD term plus router bias plus a hard-window
curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the
teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are
nearly inert, and the retrain ties this release on every task while PPL regresses
~2%. The gap stands, and the negative is recorded in the project's register.

All ten community quants plus BF16; up is better, left is smaller. Full table
and method notes:
QUANT-RETENTION-35B.md.
| model | size GB | bpw | PPL | KLD mean | KLD 99.9% | HellaSwag | Winogrande |
|---|---|---|---|---|---|---|---|
| Scion-35B-A3B | 11.34 | 2.61 | 8.354 | 0.269 | 4.758 | 79.00 | 76.25 |
| IQ2_M | 12.56 | 2.90 | 8.413 | 0.164 | 3.963 | 79.00 | 75.75 |
| Q2_K | 13.84 | 3.19 | 8.473 | 0.149 | 3.276 | 76.50 | 73.25 |
| IQ3_M | 16.34 | 3.77 | 7.520 | 0.057 | 1.225 | 80.00 | 74.25 |
| Q3_K_M | 17.66 | 4.07 | 7.432 | 0.058 | 1.611 | 79.75 | 76.00 |
| IQ4_XS | 19.63 | 4.53 | 7.264 | 0.022 | 0.636 | 80.75 | 75.75 |
| Q4_K_M | 21.71 | 5.01 | 7.235 | 0.031 | 1.141 | 80.00 | 76.00 |
| Q5_K_M | 25.35 | 5.84 | 7.273 | 0.015 | 0.717 | 80.50 | 76.25 |
| Q6_K | 29.21 | 6.73 | 7.153 | 0.008 | 0.360 | 80.75 | 74.75 |
| Q8_0 | 37.80 | 8.72 | 7.160 | 0.004 | 0.209 | 80.25 | 75.50 |
| BF16 | 71.07 | 16.38 | 7.160 | — | — | 81.25 | 76.00 |
max_new_tokens
accordingly.Q2_0: it is filename-derived; see the note above, and
use the fork.@misc{scion35b2026,
title = {Scion-35B-A3B: ternary MoE experts with trained corrections},
author = {SkyIsNotGreen},
year = {2026},
month = {September},
url = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}
Apache-2.0, inherited from the base model; see LICENSE and
NOTICE.
Base weights: empero-ai and
Qwen / Alibaba Cloud (Apache-2.0);
BF16 conversion: MrFuzzihead;
container and kernels: Prism ML (MIT) with
TAARDIS conventions (MIT);
engine: llama.cpp (MIT).