Independent forensics of Bonsai 2 27B ternary quantization: format and basis recovery, the trained-weights residual, and the public PTQ calibration artifact.
Python
0
96 commits
updated Sep 27, 2026
Independent forensics on Bonsai 2 27B, Prism ML's ~2 bpw ternary model built
from Qwen/Qwen3.8-27B. We recovered the storage format and rotation basis,
localized the quality gap to trained weights, and showed that a widely cited
public PTQ baseline is a calibration-passage artifact.
This is a study. It is not affiliated with Prism ML, Alibaba Cloud, or ThakiCloud, and it redistributes no model weights. The public write-up and discussion thread: Prism model page, discussion #62.
q
Mirror-Descent / RMD recipe is the lineage, not what Bonsai 2 runs: a direct
mirror-map arm is falsified inside the ternary loop, the zero share (0.328) is
the quantizer's central band rather than a formed attractor, and the ~8% of
trits RTN cannot reach is boundary placement from training, not an optimizer
signature. Rotation-in-the-loop + STE/KD + a managed schedule accounts for the
gap. See docs/FORENSIC-ARCHIVE.md.This repository is the dense-model forensics; the programme it started kept going. The paths it produced, in order:
docs/DSPARK-TRACK.md) and the 4B scale rung
(docs/SCALING-PROTOCOL.md).docs/WHITEPAPER.md — the full dense-forensics write-up.docs/FORENSIC-ARCHIVE.md — the settled Mirror-Descent question.docs/FAILURES.md — every falsified route, with evidence and revisit cost.sky-is-green/scion — harness, port decisions, serving experiments, release docs and its own negative register. Reference release: Scion-35B-A3B (11.34 GB / 2.61 bpw). The write-up split with it (scion/docs/MOE-EXTENSION.md).The full documentation index is docs/README.md; the
Gate-1/2/3 reports are in research/.
bonsai_forensics/ core library (rotation, quant, GPTQ, PQ2_0 codec,
oracle, gates, recovery) + spec.md
scripts/gate3/ SCR noise test, 27B prefix Hessians, GPTQ sweep,
H-objective diagnostics
scripts/pilot/ block-wise KD and data-selection pilots; the DSpark
drafter track (train / parity / export / benchmark)
scripts/eval_llama_server.py standalone smoke / perplexity evaluation
tests/ternary/ offline test suite (synthetic tensors; no downloads)
docs/ write-ups and decision records; start at docs/README.md
research/ Gate-1/2/3 write-ups
configs/ run configs, pinned model registry
No weights or large artifacts are committed. The scripts expect data under
artifacts/ (gitignored).
The original ladder is a legacy Qwen3 canary, not a complete Qwen3.8/Bonsai-2
replication. Qwen3.8-27B is a hybrid Qwen3_5ForConditionalGeneration model
with a different projection inventory. Before any cross-architecture run, use
the config-only preflight and read the audit:
configs/model_registry.yaml — pinned model
inventory, architecture profiles, and local feasibility.docs/MODEL-REGISTRY.md — candidate ranking and
cross-architecture pilot order.docs/REPRODUCIBILITY-AUDIT.md — prioritized
gaps, acceptance criteria, and the cross-architecture pilot protocol.scripts/pilot/inspect_model_targets.py — no-weight/no-GPU target census.scripts/pilot/run_cross_arch_pilot.sh — explicit, preflighted pilot runner.The current Qwen3 ladder should be reported as adaptive validation retention. Do not call an in-memory STE result PQ2_0/Bonsai-2 parity until the packed quantizer/export and persistent-basis checks in the audit are complete.
python -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
pip install -e .
The offline suite needs only numpy/pyyaml/pytest. The gate and pilot scripts
need torch/transformers/datasets; on ROCm install torch from the ROCm index.
Qwen3.5/Qwen3.8 cross-architecture runs require the qwen35 extra (the
recorded node uses Transformers 5.5.0).
Download the released model and (for forensics) the public base. Both are Apache-2.0:
huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf \
Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir artifacts/oracle/bonsai27
huggingface-cli download Qwen/Qwen3.8-27B \
--local-dir artifacts/base27
Then, from the repo root:
# Gate 1: does rotate+RTN in their basis reproduce their trits?
python -m bonsai_forensics.gate1_forensics \
--gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--layers 0,3,62,63 --out artifacts/gate1/gate1-report.json
# Gate 2: swap all PQ2_0 payloads for rotate+RTN and evaluate
python -m bonsai_forensics.gate2_rtn_artifact \
--source artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--dest artifacts/gate2/patched.gguf \
--manifest artifacts/oracle/bonsai27/hadamard-manifest.json \
--base-dir artifacts/base27 --out artifacts/gate2/patch-report.json
# Gate 3 diagnostics (see scripts/gate3/*.py for their arguments)
python scripts/gate3/scr_test.py
python scripts/gate3/prefix27.py --layers 0,3 --windows 8
python scripts/gate3/sweep27.py
python scripts/gate3/hobj27.py
# Evaluate a GGUF (one heavy ROCm process at a time)
python scripts/eval_llama_server.py smoke \
--server-bin /path/to/llama-server \
--gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--ngl 99 --ctx 4096 --no-thinking --out artifacts/eval/smoke.json
The scripts take their artifact locations from the repo root; set
HIP_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES to pin the device.
pytest tests -q
python -m bonsai_forensics.spec_hash --check
CI (.github/workflows/ci.yml) runs pytest tests -q on Python 3.13 with CPU
torch. The suite is offline and uses synthetic tensors only. It pins the frozen
wire contract (bonsai_forensics/spec.md, canonical hash
0d2c008b4aee726351f9b90e44ec003c18b579d8690db24c77a089d9e1fc652b).
One corpus (about 1 MB of Shakespeare), 8x512 context, single runs each. The effect sizes are far outside the noise and the payload-swap experiment is causal, but this is not a benchmark suite. We did not run 27B QAT, so we cannot claim to reproduce Prism's numbers; we localized the gap and identified the only remaining route.
This repository was extracted from a private research project. Task identifiers
(T30, T31, T32, ...), round numbers, and references to "project records" /
"project handoff" in the docs and research write-ups refer to that project's
internal coordination log, which is not published. Commit hashes and artifact
paths cited as evidence in docs/FAILURES.md are provenance from the original
working repository; the reports in research/ and the numbers reproduced here
are the authoritative record. No model weights are redistributed.
Apache-2.0 (see LICENSE and NOTICE). The analyzed
model, Bonsai 2 27B, is by Prism ML (Apache-2.0), derived from Qwen3.8-27B
by Alibaba Cloud (Apache-2.0); "Created using Bonsai by Prism ML." The
public PTQ reference reproduced by bonsai_forensics/reference.py is
ThakiCloud/bonsai-1bit-repro (Apache-2.0). Runtime evaluation uses
Prism's llama.cpp fork (MIT) on ggml (MIT).
Independent forensics of Bonsai 2 27B ternary quantization: format and basis recovery, the trained-weights residual, and the public PTQ calibration artifact.
Python
0
96 commits
updated Sep 27, 2026
Independent forensics on Bonsai 2 27B, Prism ML's ~2 bpw ternary model built
from Qwen/Qwen3.8-27B. We recovered the storage format and rotation basis,
localized the quality gap to trained weights, and showed that a widely cited
public PTQ baseline is a calibration-passage artifact.
This is a study. It is not affiliated with Prism ML, Alibaba Cloud, or ThakiCloud, and it redistributes no model weights. The public write-up and discussion thread: Prism model page, discussion #62.
q
Mirror-Descent / RMD recipe is the lineage, not what Bonsai 2 runs: a direct
mirror-map arm is falsified inside the ternary loop, the zero share (0.328) is
the quantizer's central band rather than a formed attractor, and the ~8% of
trits RTN cannot reach is boundary placement from training, not an optimizer
signature. Rotation-in-the-loop + STE/KD + a managed schedule accounts for the
gap. See docs/FORENSIC-ARCHIVE.md.This repository is the dense-model forensics; the programme it started kept going. The paths it produced, in order:
docs/DSPARK-TRACK.md) and the 4B scale rung
(docs/SCALING-PROTOCOL.md).docs/WHITEPAPER.md — the full dense-forensics write-up.docs/FORENSIC-ARCHIVE.md — the settled Mirror-Descent question.docs/FAILURES.md — every falsified route, with evidence and revisit cost.sky-is-green/scion — harness, port decisions, serving experiments, release docs and its own negative register. Reference release: Scion-35B-A3B (11.34 GB / 2.61 bpw). The write-up split with it (scion/docs/MOE-EXTENSION.md).The full documentation index is docs/README.md; the
Gate-1/2/3 reports are in research/.
bonsai_forensics/ core library (rotation, quant, GPTQ, PQ2_0 codec,
oracle, gates, recovery) + spec.md
scripts/gate3/ SCR noise test, 27B prefix Hessians, GPTQ sweep,
H-objective diagnostics
scripts/pilot/ block-wise KD and data-selection pilots; the DSpark
drafter track (train / parity / export / benchmark)
scripts/eval_llama_server.py standalone smoke / perplexity evaluation
tests/ternary/ offline test suite (synthetic tensors; no downloads)
docs/ write-ups and decision records; start at docs/README.md
research/ Gate-1/2/3 write-ups
configs/ run configs, pinned model registry
No weights or large artifacts are committed. The scripts expect data under
artifacts/ (gitignored).
The original ladder is a legacy Qwen3 canary, not a complete Qwen3.8/Bonsai-2
replication. Qwen3.8-27B is a hybrid Qwen3_5ForConditionalGeneration model
with a different projection inventory. Before any cross-architecture run, use
the config-only preflight and read the audit:
configs/model_registry.yaml — pinned model
inventory, architecture profiles, and local feasibility.docs/MODEL-REGISTRY.md — candidate ranking and
cross-architecture pilot order.docs/REPRODUCIBILITY-AUDIT.md — prioritized
gaps, acceptance criteria, and the cross-architecture pilot protocol.scripts/pilot/inspect_model_targets.py — no-weight/no-GPU target census.scripts/pilot/run_cross_arch_pilot.sh — explicit, preflighted pilot runner.The current Qwen3 ladder should be reported as adaptive validation retention. Do not call an in-memory STE result PQ2_0/Bonsai-2 parity until the packed quantizer/export and persistent-basis checks in the audit are complete.
python -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
pip install -e .
The offline suite needs only numpy/pyyaml/pytest. The gate and pilot scripts
need torch/transformers/datasets; on ROCm install torch from the ROCm index.
Qwen3.5/Qwen3.8 cross-architecture runs require the qwen35 extra (the
recorded node uses Transformers 5.5.0).
Download the released model and (for forensics) the public base. Both are Apache-2.0:
huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf \
Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir artifacts/oracle/bonsai27
huggingface-cli download Qwen/Qwen3.8-27B \
--local-dir artifacts/base27
Then, from the repo root:
# Gate 1: does rotate+RTN in their basis reproduce their trits?
python -m bonsai_forensics.gate1_forensics \
--gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--layers 0,3,62,63 --out artifacts/gate1/gate1-report.json
# Gate 2: swap all PQ2_0 payloads for rotate+RTN and evaluate
python -m bonsai_forensics.gate2_rtn_artifact \
--source artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--dest artifacts/gate2/patched.gguf \
--manifest artifacts/oracle/bonsai27/hadamard-manifest.json \
--base-dir artifacts/base27 --out artifacts/gate2/patch-report.json
# Gate 3 diagnostics (see scripts/gate3/*.py for their arguments)
python scripts/gate3/scr_test.py
python scripts/gate3/prefix27.py --layers 0,3 --windows 8
python scripts/gate3/sweep27.py
python scripts/gate3/hobj27.py
# Evaluate a GGUF (one heavy ROCm process at a time)
python scripts/eval_llama_server.py smoke \
--server-bin /path/to/llama-server \
--gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
--ngl 99 --ctx 4096 --no-thinking --out artifacts/eval/smoke.json
The scripts take their artifact locations from the repo root; set
HIP_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES to pin the device.
pytest tests -q
python -m bonsai_forensics.spec_hash --check
CI (.github/workflows/ci.yml) runs pytest tests -q on Python 3.13 with CPU
torch. The suite is offline and uses synthetic tensors only. It pins the frozen
wire contract (bonsai_forensics/spec.md, canonical hash
0d2c008b4aee726351f9b90e44ec003c18b579d8690db24c77a089d9e1fc652b).
One corpus (about 1 MB of Shakespeare), 8x512 context, single runs each. The effect sizes are far outside the noise and the payload-swap experiment is causal, but this is not a benchmark suite. We did not run 27B QAT, so we cannot claim to reproduce Prism's numbers; we localized the gap and identified the only remaining route.
This repository was extracted from a private research project. Task identifiers
(T30, T31, T32, ...), round numbers, and references to "project records" /
"project handoff" in the docs and research write-ups refer to that project's
internal coordination log, which is not published. Commit hashes and artifact
paths cited as evidence in docs/FAILURES.md are provenance from the original
working repository; the reports in research/ and the numbers reproduced here
are the authoritative record. No model weights are redistributed.
Apache-2.0 (see LICENSE and NOTICE). The analyzed
model, Bonsai 2 27B, is by Prism ML (Apache-2.0), derived from Qwen3.8-27B
by Alibaba Cloud (Apache-2.0); "Created using Bonsai by Prism ML." The
public PTQ reference reproduced by bonsai_forensics/reference.py is
ThakiCloud/bonsai-1bit-repro (Apache-2.0). Runtime evaluation uses
Prism's llama.cpp fork (MIT) on ggml (MIT).