SkyIsNotGreen/Scion-35B-A3B

Model

Scion-35B-A3B: ternary MoE experts + trained corrections

8

13 commits

6 linked in READMEs

updated Sep 30, 2026

See the code

README

GitHub: Build-scripts and docs  |  Forensics study  |  Runtime fork  |  Discussions

Scion-35B-A3B: ternary MoE experts + trained corrections

A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.

2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size

The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.

Highlights

  • One 11.34 GB file at 2.61 bpw (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4_K_M is 21.71 GB and IQ2_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
  • Task retention inside the Q4 and BF16 noise band: HellaSwag 400 gives 79.00% (BF16 81.25, Q4_K_M 80.00) and Winogrande gives 76.25% (BF16 76.00, Q4_K_M 76.00). It has the joint-best Winogrande row in the table and the best PPL of the 2-bit class (8.354 against IQ2_M 8.413 and Q2_K 8.473).
  • Trained, not calibrated: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
  • One file, no --lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing.
  • A k=1 speculative drafter ships alongside (Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14–1.37× faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).
  • The gap is stated, not hidden: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4_K_M at 0.031. Tail-aware training was attempted at full scale and did not transfer (TAIL-EXPERIMENT-PLAN.md).

Resources

  • GitHub sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.
  • Bonsai 2 ternary forensics: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
  • Runtime: sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, the ffn_moe_out virtual target and embedded-adapter support.
  • Retention grid: this release against every community quant of the same model, same protocol.
  • Discussions: questions, test reports and failures are all welcome.

Model Overview

ItemSpecification
Base modelempero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0)
Parameters34.9B total, about 3B active per token (256 experts, top-8 plus shared)
Architectureqwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward
Context length262,144 tokens (inherited from the base model)
Weight formatTernary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512)
Low-bit coverageExpert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32
Deployed size10.558 GiB / 11.337 GB (text only, single file)
Backendsllama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested
LicenseApache-2.0 (inherited from the base)

Weight Representation: ternary PQ2_0 + trained corrections

Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output). The corrections are rank-512 low-rank branches on the attention output and the MoE block output, plus exact router deltas. They ship in the compact legacy q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are merged into the file. Effective overall: 2.61 bpw.

Memory Requirement

FormatbpwSizevs BF16
BF16 (reference)16.3871.07 GB1.0x
Q8_08.7237.80 GB1.9x
Q4_K_M5.0121.71 GB3.3x
IQ2_M2.9012.56 GB5.7x
Scion-35B-A3B2.6111.34 GB6.3x

Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).

Shipped Components

ComponentPackSizeResidency
Language model (this repo)PQ2_0 experts, Q8_0 rest, embedded corrections11.34 GBresident; the whole model
k=1 MTP drafter (separate repo)fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding50 MBtransient; ~1 ms/eval on GPU
Uncorrected body (not uploaded)PQ2_0 experts and Q8_0 rest10.46 GiBfor swap tests
Corrections (not uploaded)rank-512 branches and router deltas (q1_0_g128)98 MiBfor swap tests

The language model is the single released body file; the drafter is a separate 50 MB file in its own repo (Scion-35B-A3B-mtp-drafter). The two-file variant (body plus separate adapter) exists for reproducing the merge and swapping corrections at runtime. Ask in Discussions if you want it.

About the Hub's quant chip. The Hub parses file names and labels this file Q2_0; the same happens on Prism ML's own PQ2_0 releases. The container is Prism's PQ2_0 (legacy name Q1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64 Q2_0 (type 42). No single quant name fits the file anyway, because it is a mix: PQ2_0 expert banks, the embedded corrections in the legacy q1_0_g128 container, and Q8_0 for the rest (norms and routers in F32).

Best Practices

Generation Parameters

Recommended values, from the base model card:

  • temperature=0.6, top_p=0.95, top_k=20

This is a reasoning distill, and answers open with a long thinking segment. Allow generous max_new_tokens (for example -n 16384); a small cap ends generation mid-thought, before any answer.

System Prompt

A simple prompt works, for example You are a helpful assistant. The base is a reasoning SFT distill and does not require a special system prompt.

Choosing Context and Offload

  • The weights fit a 12 GB card; about 16 GB is comfortable with context.
  • -ngl 99 on a single card. When VRAM is tight, offload experts to CPU with -ncmoe, which is the VRAM-budget dial; the ternary container roughly halves the CPU-tail penalty compared to an f16 expert bank.
  • Threads should equal physical cores (-t 8 on an 8C/16T CPU; SMT siblings collapse CPU expert throughput).
  • Do not layer-split across two cards when one card fits; the proxy measurements showed a 40% generation loss.

Quickstart

The runtime is the fork, and the fork is the source of truth for running these files: sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime.

These files need the fork build

The PQ2_0 container, the legacy Q1_0_g128 import, the ffn_moe_out virtual target and embedded adapters all live in the fork. Stock llama.cpp will not run this file: it treats PQ2_0 and Q1_0_g128 as unknown tensor types. Upstream's own Q2_0 (type 42, g64) is a different container and is not a substitute. The fork's default branch is the one you want, so a plain clone is enough.

# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh          # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON     CPU-only: no flag

If the model will not load: tensor '...' has invalid ggml type 142. should be in [0, 43) means the binary you ran is a build of upstream master, which knows 43 types and has never heard of PQ2_0 (type id 142). The branch above defines GGML_TYPE_PQ2_0 = 142 and GGML_TYPE_COUNT = 144, and ./verify-container-support.sh checks exactly that in a second. The same error also comes from a stale build/ directory or from an older llama-cli earlier on your PATH; the reliable fix for all three is to delete the checkout and the build directory and start from the clone above. See also the fork README.

# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -ngl 99 -c 4096 -t 8 \
    --temp 0.6 --top-p 0.95 --top-k 20 \
    -p "Explain quantum computing in simple terms." -n 16384

# server
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080

-ngl 99 offloads every layer (0 is CPU-only), -c sets context up to 262144, and -t should be the physical core count. Verified on CPU and ROCm (gfx1100, RX 7900 XT, where the release was built and measured). A CUDA build is expected to work; reports are welcome in Discussions.

Speculative decoding (k=1 drafter)

The drafter lives in its own repo: Scion-35B-A3B-mtp-drafter (50 MB). It is a small frozen-body head that predicts the model's own next token from its post-norm hidden state plus the next token's embedding; it reuses the model's own output_norm/output/token_embd, so it adds no second vocabulary projection. It is not a stock llama.cpp drafter: the fork's draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and runs it entirely against the target context, with no draft model.

hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .

# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080

# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8

Measured on one RX 7900 XT (greedy):

WorkloadBaseline+ drafterSpeedupDraft acceptance
Bench prompt, 400 tokens12.41 ms/tok9.07 ms/tok1.37×0.738
Open-ended CLI generation49.0 tok/s57.2 tok/s1.17×0.46–0.74
Server confirmation——1.14–1.21×0.44–0.70

Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent — treat ~1.2× as the typical figure and 1.37× as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.

Benchmarks

Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks); KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and Winogrande 400 zero-shot (about ±2% CI). BF16 and all quants were measured in one H100 session; this release was measured on the local card and cross-checked against the pod (IQ2_M KLD local vs pod: 0.2%).

Task Retention (400 tasks each)

VariantSizebpwHellaSwagWinogrande
BF16 (reference)71.07 GB16.3881.2576.00
Q4_K_M21.71 GB5.0180.0076.00
IQ2_M12.56 GB2.9079.0075.75
Q2_K13.84 GB3.1976.5073.25
Scion-35B-A3B11.34 GB2.6179.0076.25

Within ±2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.

Additional tasks (re-measured 2026-09-30)

The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added the remaining standard multiple-choice tasks on the same release and the same fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):

TaskScion-35B-A3B
ARC-Challenge (299)57.19 ±2.87
ARC-Easy (570)80.00 ±1.68
MMLU (2000, seed 1234)40.15 ±1.10
TruthfulQA (817)32.56 ±1.64
wikitext-2 PPL (c512, 580 chunks)8.2731 ±0.054

The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the grid's 8.3539; the model is unchanged.

Distributional Fidelity: the known gap

VariantPPL (lower better)KLD mean vs BF16 (lower better)KLD 99.9%
BF167.1595——
Q4_K_M7.23540.03141.141
IQ2_M8.41330.16363.963
Q2_K8.47290.14933.276
Scion-35B-A3B8.35390.26944.758

The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35 after) more than they improved the full-distribution tail. PPL ranks this build first of the 2-bit class and KLD ranks it last, and that disagreement is the research result: the training matches the teacher's top-50 logits, and the rest of the distribution is unconstrained. Tail-aware training was attempted at full scale (a tail-conditional KD term plus router bias plus a hard-window curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are nearly inert, and the retrain ties this release on every task while PPL regresses ~2%. The gap stands, and the negative is recorded in the project's register.

Full Grid

Scion-35B-A3B against every community quant of the same model, plus BF16

All ten community quants plus BF16; up is better, left is smaller. Full table and method notes: QUANT-RETENTION-35B.md.

Full per-variant table
modelsize GBbpwPPLKLD meanKLD 99.9%HellaSwagWinogrande
Scion-35B-A3B11.342.618.3540.2694.75879.0076.25
IQ2_M12.562.908.4130.1643.96379.0075.75
Q2_K13.843.198.4730.1493.27676.5073.25
IQ3_M16.343.777.5200.0571.22580.0074.25
Q3_K_M17.664.077.4320.0581.61179.7576.00
IQ4_XS19.634.537.2640.0220.63680.7575.75
Q4_K_M21.715.017.2350.0311.14180.0076.00
Q5_K_M25.355.847.2730.0150.71780.5076.25
Q6_K29.216.737.1530.0080.36080.7574.75
Q8_037.808.727.1600.0040.20980.2575.50
BF1671.0716.387.160——81.2576.00

Use Cases

  • A 35B-A3B on one consumer GPU: 11.3 GB of weights fit a 20 GB card with room for context, and can be tiered further with CPU expert offload.
  • Low-bit research and testing: a reference point for "ternary experts plus trained corrections, no full-precision masters, no full-model QAT". The build-scripts and every negative result are public.
  • Local-first serving: the model was built for an offline assistant stack, so a single file with embedded corrections keeps deployment to one download and one set of flags.

Limitations

  • KLD tail (stated above): distributional fidelity is strong-2-bit, not Q4-class. PPL, HellaSwag and Winogrande look Q4-class; KLD is where the gap lives. A full-scale tail-aware retrain was attempted and did not transfer.
  • Text only: the base model has a vision tower, and this file carries no vision tensors (Q8_0 language path only).
  • Reasoning distill: long thinking traces; budget max_new_tokens accordingly.
  • Protocol caveats: the KLD figure is 50 chunks; task numbers are 400-task runs (about ±2% CI); the 2-bit and 3-bit competitors are imatrix-calibrated on Wikipedia-like data, which flatters their wikitext KLD, and the corrections were trained on fineweb.
  • Platform coverage: CPU and ROCm (gfx1100) verified. CUDA is expected to work but has not been tested, and Metal is untested.
  • The Hub chip says Q2_0: it is filename-derived; see the note above, and use the fork.
  • Not affiliated with Prism ML, empero-ai, or Alibaba Cloud. It builds on Prism ML's engine work (fork and PQ2_0 container) and community GGUF conversions.

Citation

@misc{scion35b2026,
    title  = {Scion-35B-A3B: ternary MoE experts with trained corrections},
    author = {SkyIsNotGreen},
    year   = {2026},
    month  = {September},
    url    = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}

License and Attribution

Apache-2.0, inherited from the base model; see LICENSE and NOTICE. Base weights: empero-ai and Qwen / Alibaba Cloud (Apache-2.0); BF16 conversion: MrFuzzihead; container and kernels: Prism ML (MIT) with TAARDIS conventions (MIT); engine: llama.cpp (MIT).

2-bit
conversational
draft-model
endpoints_compatible
gguf
llama-cpp
llama.cpp
pq2_0
quantized
qwen3.8
scion
speculative-decoding
ternary
text-generation

SkyIsNotGreen/Scion-35B-A3B

Model

Scion-35B-A3B: ternary MoE experts + trained corrections

8

13 commits

6 linked in READMEs

updated Sep 30, 2026

See the code

README

GitHub: Build-scripts and docs  |  Forensics study  |  Runtime fork  |  Discussions

Scion-35B-A3B: ternary MoE experts + trained corrections

A full 35B-A3B MoE in one 11.3 GB GGUF for llama.cpp. Ternary expert banks plus small trained corrections, with no full-precision masters and no full-model QAT.

2.61 bpw | 11.34 GB (6.3x smaller than the BF16 reference) | best PPL of the 2-bit class | task retention at Q4-class level, at about half of Q4_K_M's size

The name comes from grafting. A scion is the shoot grafted onto a rootstock, and here the trained corrections are grafted onto a 2.125 bpw ternary body.

Highlights

  • One 11.34 GB file at 2.61 bpw (expert banks at 2.125 bpw). The BF16 reference of the same model is 71.07 GB; Q4_K_M is 21.71 GB and IQ2_M is 12.56 GB. It is the smallest published build of this model I have seen at this quality level.
  • Task retention inside the Q4 and BF16 noise band: HellaSwag 400 gives 79.00% (BF16 81.25, Q4_K_M 80.00) and Winogrande gives 76.25% (BF16 76.00, Q4_K_M 76.00). It has the joint-best Winogrande row in the table and the best PPL of the 2-bit class (8.354 against IQ2_M 8.413 and Q2_K 8.473).
  • Trained, not calibrated: rank-512 correction branches on the attention output and the MoE block output, plus router deltas, trained by output-KD against the BF16 teacher with the deployed quantizer in the loop (ternary Lloyd g128). No imatrix and no calibration corpus, which is what separates this build from the imatrix-calibrated quants on the chart.
  • One file, no --lora: the corrections are embedded (adapter.embedded=true) and attached at load. There is no adapter plumbing.
  • A k=1 speculative drafter ships alongside (Scion-35B-A3B-mtp-drafter, 50 MB): it drafts the next token from the model's own hidden state and gives 1.14–1.37× faster generation on llama.cpp's speculative path, with unchanged outputs. It needs the fork runtime (it is not a stock llama.cpp drafter).
  • The gap is stated, not hidden: full-vocabulary KLD against BF16 is 0.269 mean, a strong 2-bit-class result but still behind Q4_K_M at 0.031. Tail-aware training was attempted at full scale and did not transfer (TAIL-EXPERIMENT-PLAN.md).

Resources

  • GitHub sky-is-green/scion: the source of truth for this work. It carries the harness that produced the file, the full MoE write-up, the port decisions and the negative register.
  • Bonsai 2 ternary forensics: the dense-model study this method grew out of (format recovery, the trained-weight residual, the calibration-artifact result).
  • Runtime: sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime, a fork of Prism ML's llama.cpp with the PQ2_0 container, the ffn_moe_out virtual target and embedded-adapter support.
  • Retention grid: this release against every community quant of the same model, same protocol.
  • Discussions: questions, test reports and failures are all welcome.

Model Overview

ItemSpecification
Base modelempero-ai/Qwen3.8-35B-A3B-Distill, a Qwen3.8-line reasoning distill built on Qwen/Qwen3.6-35B-A3B (Apache-2.0)
Parameters34.9B total, about 3B active per token (256 experts, top-8 plus shared)
Architectureqwen3_5_moe (qwen35moe in llama.cpp): 40 layers, hybrid linear and full attention, MoE feed-forward
Context length262,144 tokens (inherited from the base model)
Weight formatTernary PQ2_0 g128 expert banks (codes in {-1, 0, +1} plus one fp16 group scale per 128 weights), Q8_0 for the rest, embedded corrections in the legacy q1_0_g128 container (rank-512)
Low-bit coverageExpert banks only; attention, embeddings and the head stay Q8_0, and norms, routers and the output stay F32
Deployed size10.558 GiB / 11.337 GB (text only, single file)
Backendsllama.cpp fork; verified on CPU and ROCm/gfx1100, CUDA build expected to work but untested
LicenseApache-2.0 (inherited from the base)

Weight Representation: ternary PQ2_0 + trained corrections

Each expert weight takes a value from {-1, 0, +1} with one shared FP16 scale per group of 128 weights: 2-bit slots at 2.125 bits per weight. The rest of the model is Q8_0 (attention, embeddings, LM head) and F32 (norms, routers, output). The corrections are rank-512 low-rank branches on the attention output and the MoE block output, plus exact router deltas. They ship in the compact legacy q1_0_g128 container (2-bit codes plus one fp16 group scale per 128) and are merged into the file. Effective overall: 2.61 bpw.

Memory Requirement

FormatbpwSizevs BF16
BF16 (reference)16.3871.07 GB1.0x
Q8_08.7237.80 GB1.9x
Q4_K_M5.0121.71 GB3.3x
IQ2_M2.9012.56 GB5.7x
Scion-35B-A3B2.6111.34 GB6.3x

Sizes are the published on-disk files of the same model, measured under one protocol (wikitext-2 PPL, KLD against BF16, HellaSwag and Winogrande 400; see Benchmarks).

Shipped Components

ComponentPackSizeResidency
Language model (this repo)PQ2_0 experts, Q8_0 rest, embedded corrections11.34 GBresident; the whole model
k=1 MTP drafter (separate repo)fp16 head (fc1/gelu/fc2) over the model's own hidden + next-token embedding50 MBtransient; ~1 ms/eval on GPU
Uncorrected body (not uploaded)PQ2_0 experts and Q8_0 rest10.46 GiBfor swap tests
Corrections (not uploaded)rank-512 branches and router deltas (q1_0_g128)98 MiBfor swap tests

The language model is the single released body file; the drafter is a separate 50 MB file in its own repo (Scion-35B-A3B-mtp-drafter). The two-file variant (body plus separate adapter) exists for reproducing the merge and swapping corrections at runtime. Ask in Discussions if you want it.

About the Hub's quant chip. The Hub parses file names and labels this file Q2_0; the same happens on Prism ML's own PQ2_0 releases. The container is Prism's PQ2_0 (legacy name Q1_0_g128, type 142/43, identical byte layout): 2-bit codes with one fp16 group scale per 128 weights. It is not upstream llama.cpp's g64 Q2_0 (type 42). No single quant name fits the file anyway, because it is a mix: PQ2_0 expert banks, the embedded corrections in the legacy q1_0_g128 container, and Q8_0 for the rest (norms and routers in F32).

Best Practices

Generation Parameters

Recommended values, from the base model card:

  • temperature=0.6, top_p=0.95, top_k=20

This is a reasoning distill, and answers open with a long thinking segment. Allow generous max_new_tokens (for example -n 16384); a small cap ends generation mid-thought, before any answer.

System Prompt

A simple prompt works, for example You are a helpful assistant. The base is a reasoning SFT distill and does not require a special system prompt.

Choosing Context and Offload

  • The weights fit a 12 GB card; about 16 GB is comfortable with context.
  • -ngl 99 on a single card. When VRAM is tight, offload experts to CPU with -ncmoe, which is the VRAM-budget dial; the ternary container roughly halves the CPU-tail penalty compared to an f16 expert bank.
  • Threads should equal physical cores (-t 8 on an 8C/16T CPU; SMT siblings collapse CPU expert throughput).
  • Do not layer-split across two cards when one card fits; the proxy measurements showed a 40% generation loss.

Quickstart

The runtime is the fork, and the fork is the source of truth for running these files: sky-is-green/prism-ml-llama.cpp, branch moe-corr-runtime.

These files need the fork build

The PQ2_0 container, the legacy Q1_0_g128 import, the ffn_moe_out virtual target and embedded adapters all live in the fork. Stock llama.cpp will not run this file: it treats PQ2_0 and Q1_0_g128 as unknown tensor types. Upstream's own Q2_0 (type 42, g64) is a different container and is not a substitute. The fork's default branch is the one you want, so a plain clone is enough.

# build the runtime
git clone https://github.com/sky-is-green/prism-ml-llama.cpp
cd prism-ml-llama.cpp
./verify-container-support.sh          # must print "RESULT: OK"
rm -rf build
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-cli llama-server
# ROCm: -DGGML_HIP=ON     CPU-only: no flag

If the model will not load: tensor '...' has invalid ggml type 142. should be in [0, 43) means the binary you ran is a build of upstream master, which knows 43 types and has never heard of PQ2_0 (type id 142). The branch above defines GGML_TYPE_PQ2_0 = 142 and GGML_TYPE_COUNT = 144, and ./verify-container-support.sh checks exactly that in a second. The same error also comes from a stale build/ directory or from an older llama-cli earlier on your PATH; the reliable fix for all three is to delete the checkout and the build directory and start from the clone above. See also the fork README.

# fetch the weights
hf download SkyIsNotGreen/Scion-35B-A3B Scion-35B-A3B-PQ2_0-corr.gguf --local-dir .
# chat; the model thinks by default, so leave room for the trace
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -ngl 99 -c 4096 -t 8 \
    --temp 0.6 --top-p 0.95 --top-k 20 \
    -p "Explain quantum computing in simple terms." -n 16384

# server
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf -ngl 99 -c 4096 -t 8 --port 8080

-ngl 99 offloads every layer (0 is CPU-only), -c sets context up to 262144, and -t should be the physical core count. Verified on CPU and ROCm (gfx1100, RX 7900 XT, where the release was built and measured). A CUDA build is expected to work; reports are welcome in Discussions.

Speculative decoding (k=1 drafter)

The drafter lives in its own repo: Scion-35B-A3B-mtp-drafter (50 MB). It is a small frozen-body head that predicts the model's own next token from its post-norm hidden state plus the next token's embedding; it reuses the model's own output_norm/output/token_embd, so it adds no second vocabulary projection. It is not a stock llama.cpp drafter: the fork's draft-mtp-sidecar implementation loads it (auto-detected from the GGUF) and runs it entirely against the target context, with no draft model.

hf download SkyIsNotGreen/Scion-35B-A3B-mtp-drafter Scion-35B-A3B-mtp-drafter.gguf --local-dir .

# server (no extra flags; the sidecar is auto-detected)
./build/bin/llama-server -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8 --port 8080

# CLI
./build/bin/llama-cli -m Scion-35B-A3B-PQ2_0-corr.gguf \
    -md Scion-35B-A3B-mtp-drafter.gguf -ngl 99 -c 4096 -t 8

Measured on one RX 7900 XT (greedy):

WorkloadBaseline+ drafterSpeedupDraft acceptance
Bench prompt, 400 tokens12.41 ms/tok9.07 ms/tok1.37×0.738
Open-ended CLI generation49.0 tok/s57.2 tok/s1.17×0.46–0.74
Server confirmation——1.14–1.21×0.44–0.70

Teacher-forced acceptance: fineweb 0.483 / wikitext 0.360. Speedup is prompt-dependent — treat ~1.2× as the typical figure and 1.37× as the benchmark-prompt best. Drafting leaves the model's outputs unchanged apart from rare near-tie flips on batched verification.

Benchmarks

Community protocol, the same for every row: wikitext-2 PPL (c512, 580 chunks); KLD against BF16 logits over 50 chunks (25.5k tokens); HellaSwag 400 and Winogrande 400 zero-shot (about ±2% CI). BF16 and all quants were measured in one H100 session; this release was measured on the local card and cross-checked against the pod (IQ2_M KLD local vs pod: 0.2%).

Task Retention (400 tasks each)

VariantSizebpwHellaSwagWinogrande
BF16 (reference)71.07 GB16.3881.2576.00
Q4_K_M21.71 GB5.0180.0076.00
IQ2_M12.56 GB2.9079.0075.75
Q2_K13.84 GB3.1976.5073.25
Scion-35B-A3B11.34 GB2.6179.0076.25

Within ±2% noise of Q4_K_M and BF16; tied with IQ2_M on HellaSwag and ahead of it on Winogrande at 1.2 GB less. Treat sub-1% differences as ties.

Additional tasks (re-measured 2026-09-30)

The 2026-09-27 grid covered HellaSwag and Winogrande. A later local run added the remaining standard multiple-choice tasks on the same release and the same fork build (one RX 7900 XT, llama-perplexity --multiple-choice, same seed):

TaskScion-35B-A3B
ARC-Challenge (299)57.19 ±2.87
ARC-Easy (570)80.00 ±1.68
MMLU (2000, seed 1234)40.15 ±1.10
TruthfulQA (817)32.56 ±1.64
wikitext-2 PPL (c512, 580 chunks)8.2731 ±0.054

The fresh PPL uses --no-warmup and -np 8, so its level sits ~1% below the grid's 8.3539; the model is unchanged.

Distributional Fidelity: the known gap

VariantPPL (lower better)KLD mean vs BF16 (lower better)KLD 99.9%
BF167.1595——
Q4_K_M7.23540.03141.141
IQ2_M8.41330.16363.963
Q2_K8.47290.14933.276
Scion-35B-A3B8.35390.26944.758

The corrections improved mean token likelihood (PPL 11.60 uncorrected, 8.35 after) more than they improved the full-distribution tail. PPL ranks this build first of the 2-bit class and KLD ranks it last, and that disagreement is the research result: the training matches the teacher's top-50 logits, and the rest of the distribution is unconstrained. Tail-aware training was attempted at full scale (a tail-conditional KD term plus router bias plus a hard-window curriculum, the cur05 recipe) and did not transfer: on a 40-layer body the teacher puts ~98% of its mass inside the top-512 cache, so the tail terms are nearly inert, and the retrain ties this release on every task while PPL regresses ~2%. The gap stands, and the negative is recorded in the project's register.

Full Grid

Scion-35B-A3B against every community quant of the same model, plus BF16

All ten community quants plus BF16; up is better, left is smaller. Full table and method notes: QUANT-RETENTION-35B.md.

Full per-variant table
modelsize GBbpwPPLKLD meanKLD 99.9%HellaSwagWinogrande
Scion-35B-A3B11.342.618.3540.2694.75879.0076.25
IQ2_M12.562.908.4130.1643.96379.0075.75
Q2_K13.843.198.4730.1493.27676.5073.25
IQ3_M16.343.777.5200.0571.22580.0074.25
Q3_K_M17.664.077.4320.0581.61179.7576.00
IQ4_XS19.634.537.2640.0220.63680.7575.75
Q4_K_M21.715.017.2350.0311.14180.0076.00
Q5_K_M25.355.847.2730.0150.71780.5076.25
Q6_K29.216.737.1530.0080.36080.7574.75
Q8_037.808.727.1600.0040.20980.2575.50
BF1671.0716.387.160——81.2576.00

Use Cases

  • A 35B-A3B on one consumer GPU: 11.3 GB of weights fit a 20 GB card with room for context, and can be tiered further with CPU expert offload.
  • Low-bit research and testing: a reference point for "ternary experts plus trained corrections, no full-precision masters, no full-model QAT". The build-scripts and every negative result are public.
  • Local-first serving: the model was built for an offline assistant stack, so a single file with embedded corrections keeps deployment to one download and one set of flags.

Limitations

  • KLD tail (stated above): distributional fidelity is strong-2-bit, not Q4-class. PPL, HellaSwag and Winogrande look Q4-class; KLD is where the gap lives. A full-scale tail-aware retrain was attempted and did not transfer.
  • Text only: the base model has a vision tower, and this file carries no vision tensors (Q8_0 language path only).
  • Reasoning distill: long thinking traces; budget max_new_tokens accordingly.
  • Protocol caveats: the KLD figure is 50 chunks; task numbers are 400-task runs (about ±2% CI); the 2-bit and 3-bit competitors are imatrix-calibrated on Wikipedia-like data, which flatters their wikitext KLD, and the corrections were trained on fineweb.
  • Platform coverage: CPU and ROCm (gfx1100) verified. CUDA is expected to work but has not been tested, and Metal is untested.
  • The Hub chip says Q2_0: it is filename-derived; see the note above, and use the fork.
  • Not affiliated with Prism ML, empero-ai, or Alibaba Cloud. It builds on Prism ML's engine work (fork and PQ2_0 container) and community GGUF conversions.

Citation

@misc{scion35b2026,
    title  = {Scion-35B-A3B: ternary MoE experts with trained corrections},
    author = {SkyIsNotGreen},
    year   = {2026},
    month  = {September},
    url    = {https://huggingface.co/SkyIsNotGreen/Scion-35B-A3B}
}

License and Attribution

Apache-2.0, inherited from the base model; see LICENSE and NOTICE. Base weights: empero-ai and Qwen / Alibaba Cloud (Apache-2.0); BF16 conversion: MrFuzzihead; container and kernels: Prism ML (MIT) with TAARDIS conventions (MIT); engine: llama.cpp (MIT).

2-bit
conversational
draft-model
endpoints_compatible
gguf
llama-cpp
llama.cpp
pq2_0
quantized
qwen3.8
scion
speculative-decoding
ternary
text-generation