sky-is-green/bonsai2-ternary-forensics

Independent forensics of Bonsai 2 27B ternary quantization: format and basis recovery, the trained-weights residual, and the public PTQ calibration artifact.

See the code

README

Bonsai 2 ternary forensics

License Python PyTorch NumPy pytest Last commit Repo size

Independent forensics on Bonsai 2 27B, Prism ML's ~2 bpw ternary model built from Qwen/Qwen3.8-27B. We recovered the storage format and rotation basis, localized the quality gap to trained weights, and showed that a widely cited public PTQ baseline is a calibration-passage artifact.

This is a study. It is not affiliated with Prism ML, Alibaba Cloud, or ThakiCloud, and it redistributes no model weights. The public write-up and discussion thread: Prism model page, discussion #62.

Findings at a glance

  • The format is not the secret. The GGUF declares its own blockwise Walsh-Hadamard rotation (block 1024, explicit signs, input axis). Once the tensor layouts are lined up, plain absmean round-to-nearest on the rotated base agrees with Prism's trits 0.920 mean / 0.885 min (chance: 0.33).
  • The residual carries the quality. Replacing all 402 ternary payloads with our own rotate+RTN weights, keeping metadata byte-identical, collapses perplexity from 18.5851 to 23,606 (1,270x) in the same binary, corpus, and settings.
  • It is not a quantizer trick. GPTQ with real Hessians, damping sweeps, act-order, rotation-aware Hessians, and group scale search all move the codes away from Prism's; their codes fit our activation Hessians worse than plain RTN. The residual is trained weight movement.
  • The public 2.1x is a calibration artifact. Reproducing the public QuIP reference gives the published 2.1x on a fixed passage; on held-out windows with real calibration data the same configurations collapse (1,178x to ~35,000x). Public PTQ at ~2 bpw does not survive off-calibration here.
  • The mechanism is ordinary ternary QAT in the rotated basis. The large-q Mirror-Descent / RMD recipe is the lineage, not what Bonsai 2 runs: a direct mirror-map arm is falsified inside the ternary loop, the zero share (0.328) is the quantizer's central band rather than a formed attractor, and the ~8% of trits RTN cannot reach is boundary placement from training, not an optimizer signature. Rotation-in-the-loop + STE/KD + a managed schedule accounts for the gap. See docs/FORENSIC-ARCHIVE.md.

Where this led (research paths)

This repository is the dense-model forensics; the programme it started kept going. The paths it produced, in order:

  1. Dense forensics → settled (findings above): container recovered, the residual is trained placement, public PTQ is a calibration artifact (F1–F4); Track B ships the capability from Prism's released weights.
  2. Local QAT/KD canary → priced, deliberately not funded (F5–F8, F11): 1.7B STE+KD was retracted to ~48% retention on a clean holdout; the 27B proof-run is mapped and unfunded.
  3. MoE extension → its own repository (sky-is-green/scion): in-place ternary MoE + trained corrections; the placement rule, the AUTOGRID noise-floor map, the deployable Lloyd quantizer; OLMoE 566.7 → 14.48, then the 35B release Scion-35B-A3B (11.34 GB / 2.61 bpw, task-level Q4-class).
  4. Serving (measured there): ternary halves the CPU-offload penalty vs f16; threads = physical cores; don't layer-split when one card fits; the hot-expert cache is a measured no-gain — sparse dispatch is the real project.
  5. Open path (current) — the KLD tail beyond the teacher's top-50: scion/docs/TAIL-EXPERIMENT-PLAN.md (queued; calibration-free constraint).
  6. Shelved — the DSpark drafter track (docs/DSPARK-TRACK.md) and the 4B scale rung (docs/SCALING-PROTOCOL.md).

Start here

  1. docs/WHITEPAPER.md — the full dense-forensics write-up.
  2. docs/FORENSIC-ARCHIVE.md — the settled Mirror-Descent question.
  3. docs/FAILURES.md — every falsified route, with evidence and revisit cost.
  4. The MoE extension moved to its own repository: sky-is-green/scion — harness, port decisions, serving experiments, release docs and its own negative register. Reference release: Scion-35B-A3B (11.34 GB / 2.61 bpw). The write-up split with it (scion/docs/MOE-EXTENSION.md).

The full documentation index is docs/README.md; the Gate-1/2/3 reports are in research/.

What is here

bonsai_forensics/        core library (rotation, quant, GPTQ, PQ2_0 codec,
                         oracle, gates, recovery) + spec.md
scripts/gate3/           SCR noise test, 27B prefix Hessians, GPTQ sweep,
                         H-objective diagnostics
scripts/pilot/           block-wise KD and data-selection pilots; the DSpark
                         drafter track (train / parity / export / benchmark)
scripts/eval_llama_server.py  standalone smoke / perplexity evaluation
tests/ternary/           offline test suite (synthetic tensors; no downloads)
docs/                    write-ups and decision records; start at docs/README.md
research/                Gate-1/2/3 write-ups
configs/                 run configs, pinned model registry

No weights or large artifacts are committed. The scripts expect data under artifacts/ (gitignored).

Architecture robustness and reproducibility

The original ladder is a legacy Qwen3 canary, not a complete Qwen3.8/Bonsai-2 replication. Qwen3.8-27B is a hybrid Qwen3_5ForConditionalGeneration model with a different projection inventory. Before any cross-architecture run, use the config-only preflight and read the audit:

  • configs/model_registry.yaml — pinned model inventory, architecture profiles, and local feasibility.
  • docs/MODEL-REGISTRY.md — candidate ranking and cross-architecture pilot order.
  • docs/REPRODUCIBILITY-AUDIT.md — prioritized gaps, acceptance criteria, and the cross-architecture pilot protocol.
  • scripts/pilot/inspect_model_targets.py — no-weight/no-GPU target census.
  • scripts/pilot/run_cross_arch_pilot.sh — explicit, preflighted pilot runner.

The current Qwen3 ladder should be reported as adaptive validation retention. Do not call an in-memory STE result PQ2_0/Bonsai-2 parity until the packed quantizer/export and persistent-basis checks in the audit are complete.

Install

python -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
pip install -e .

The offline suite needs only numpy/pyyaml/pytest. The gate and pilot scripts need torch/transformers/datasets; on ROCm install torch from the ROCm index. Qwen3.5/Qwen3.8 cross-architecture runs require the qwen35 extra (the recorded node uses Transformers 5.5.0).

Reproduce

Download the released model and (for forensics) the public base. Both are Apache-2.0:

huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf \
    Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir artifacts/oracle/bonsai27
huggingface-cli download Qwen/Qwen3.8-27B \
    --local-dir artifacts/base27

Then, from the repo root:

# Gate 1: does rotate+RTN in their basis reproduce their trits?
python -m bonsai_forensics.gate1_forensics \
    --gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --layers 0,3,62,63 --out artifacts/gate1/gate1-report.json

# Gate 2: swap all PQ2_0 payloads for rotate+RTN and evaluate
python -m bonsai_forensics.gate2_rtn_artifact \
    --source artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --dest artifacts/gate2/patched.gguf \
    --manifest artifacts/oracle/bonsai27/hadamard-manifest.json \
    --base-dir artifacts/base27 --out artifacts/gate2/patch-report.json

# Gate 3 diagnostics (see scripts/gate3/*.py for their arguments)
python scripts/gate3/scr_test.py
python scripts/gate3/prefix27.py --layers 0,3 --windows 8
python scripts/gate3/sweep27.py
python scripts/gate3/hobj27.py

# Evaluate a GGUF (one heavy ROCm process at a time)
python scripts/eval_llama_server.py smoke \
    --server-bin /path/to/llama-server \
    --gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --ngl 99 --ctx 4096 --no-thinking --out artifacts/eval/smoke.json

The scripts take their artifact locations from the repo root; set HIP_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES to pin the device.

Tests

pytest tests -q
python -m bonsai_forensics.spec_hash --check

CI (.github/workflows/ci.yml) runs pytest tests -q on Python 3.13 with CPU torch. The suite is offline and uses synthetic tensors only. It pins the frozen wire contract (bonsai_forensics/spec.md, canonical hash 0d2c008b4aee726351f9b90e44ec003c18b579d8690db24c77a089d9e1fc652b).

Caveats

One corpus (about 1 MB of Shakespeare), 8x512 context, single runs each. The effect sizes are far outside the noise and the payload-swap experiment is causal, but this is not a benchmark suite. We did not run 27B QAT, so we cannot claim to reproduce Prism's numbers; we localized the gap and identified the only remaining route.

Provenance

This repository was extracted from a private research project. Task identifiers (T30, T31, T32, ...), round numbers, and references to "project records" / "project handoff" in the docs and research write-ups refer to that project's internal coordination log, which is not published. Commit hashes and artifact paths cited as evidence in docs/FAILURES.md are provenance from the original working repository; the reports in research/ and the numbers reproduced here are the authoritative record. No model weights are redistributed.

License and attribution

Apache-2.0 (see LICENSE and NOTICE). The analyzed model, Bonsai 2 27B, is by Prism ML (Apache-2.0), derived from Qwen3.8-27B by Alibaba Cloud (Apache-2.0); "Created using Bonsai by Prism ML." The public PTQ reference reproduced by bonsai_forensics/reference.py is ThakiCloud/bonsai-1bit-repro (Apache-2.0). Runtime evaluation uses Prism's llama.cpp fork (MIT) on ggml (MIT).

bonsai
gguf
llm
mixture-of-experts
model-compression
quantization
reproducibility
ternary-quantization

sky-is-green/bonsai2-ternary-forensics

Independent forensics of Bonsai 2 27B ternary quantization: format and basis recovery, the trained-weights residual, and the public PTQ calibration artifact.

See the code

README

Bonsai 2 ternary forensics

License Python PyTorch NumPy pytest Last commit Repo size

Independent forensics on Bonsai 2 27B, Prism ML's ~2 bpw ternary model built from Qwen/Qwen3.8-27B. We recovered the storage format and rotation basis, localized the quality gap to trained weights, and showed that a widely cited public PTQ baseline is a calibration-passage artifact.

This is a study. It is not affiliated with Prism ML, Alibaba Cloud, or ThakiCloud, and it redistributes no model weights. The public write-up and discussion thread: Prism model page, discussion #62.

Findings at a glance

  • The format is not the secret. The GGUF declares its own blockwise Walsh-Hadamard rotation (block 1024, explicit signs, input axis). Once the tensor layouts are lined up, plain absmean round-to-nearest on the rotated base agrees with Prism's trits 0.920 mean / 0.885 min (chance: 0.33).
  • The residual carries the quality. Replacing all 402 ternary payloads with our own rotate+RTN weights, keeping metadata byte-identical, collapses perplexity from 18.5851 to 23,606 (1,270x) in the same binary, corpus, and settings.
  • It is not a quantizer trick. GPTQ with real Hessians, damping sweeps, act-order, rotation-aware Hessians, and group scale search all move the codes away from Prism's; their codes fit our activation Hessians worse than plain RTN. The residual is trained weight movement.
  • The public 2.1x is a calibration artifact. Reproducing the public QuIP reference gives the published 2.1x on a fixed passage; on held-out windows with real calibration data the same configurations collapse (1,178x to ~35,000x). Public PTQ at ~2 bpw does not survive off-calibration here.
  • The mechanism is ordinary ternary QAT in the rotated basis. The large-q Mirror-Descent / RMD recipe is the lineage, not what Bonsai 2 runs: a direct mirror-map arm is falsified inside the ternary loop, the zero share (0.328) is the quantizer's central band rather than a formed attractor, and the ~8% of trits RTN cannot reach is boundary placement from training, not an optimizer signature. Rotation-in-the-loop + STE/KD + a managed schedule accounts for the gap. See docs/FORENSIC-ARCHIVE.md.

Where this led (research paths)

This repository is the dense-model forensics; the programme it started kept going. The paths it produced, in order:

  1. Dense forensics → settled (findings above): container recovered, the residual is trained placement, public PTQ is a calibration artifact (F1–F4); Track B ships the capability from Prism's released weights.
  2. Local QAT/KD canary → priced, deliberately not funded (F5–F8, F11): 1.7B STE+KD was retracted to ~48% retention on a clean holdout; the 27B proof-run is mapped and unfunded.
  3. MoE extension → its own repository (sky-is-green/scion): in-place ternary MoE + trained corrections; the placement rule, the AUTOGRID noise-floor map, the deployable Lloyd quantizer; OLMoE 566.7 → 14.48, then the 35B release Scion-35B-A3B (11.34 GB / 2.61 bpw, task-level Q4-class).
  4. Serving (measured there): ternary halves the CPU-offload penalty vs f16; threads = physical cores; don't layer-split when one card fits; the hot-expert cache is a measured no-gain — sparse dispatch is the real project.
  5. Open path (current) — the KLD tail beyond the teacher's top-50: scion/docs/TAIL-EXPERIMENT-PLAN.md (queued; calibration-free constraint).
  6. Shelved — the DSpark drafter track (docs/DSPARK-TRACK.md) and the 4B scale rung (docs/SCALING-PROTOCOL.md).

Start here

  1. docs/WHITEPAPER.md — the full dense-forensics write-up.
  2. docs/FORENSIC-ARCHIVE.md — the settled Mirror-Descent question.
  3. docs/FAILURES.md — every falsified route, with evidence and revisit cost.
  4. The MoE extension moved to its own repository: sky-is-green/scion — harness, port decisions, serving experiments, release docs and its own negative register. Reference release: Scion-35B-A3B (11.34 GB / 2.61 bpw). The write-up split with it (scion/docs/MOE-EXTENSION.md).

The full documentation index is docs/README.md; the Gate-1/2/3 reports are in research/.

What is here

bonsai_forensics/        core library (rotation, quant, GPTQ, PQ2_0 codec,
                         oracle, gates, recovery) + spec.md
scripts/gate3/           SCR noise test, 27B prefix Hessians, GPTQ sweep,
                         H-objective diagnostics
scripts/pilot/           block-wise KD and data-selection pilots; the DSpark
                         drafter track (train / parity / export / benchmark)
scripts/eval_llama_server.py  standalone smoke / perplexity evaluation
tests/ternary/           offline test suite (synthetic tensors; no downloads)
docs/                    write-ups and decision records; start at docs/README.md
research/                Gate-1/2/3 write-ups
configs/                 run configs, pinned model registry

No weights or large artifacts are committed. The scripts expect data under artifacts/ (gitignored).

Architecture robustness and reproducibility

The original ladder is a legacy Qwen3 canary, not a complete Qwen3.8/Bonsai-2 replication. Qwen3.8-27B is a hybrid Qwen3_5ForConditionalGeneration model with a different projection inventory. Before any cross-architecture run, use the config-only preflight and read the audit:

  • configs/model_registry.yaml — pinned model inventory, architecture profiles, and local feasibility.
  • docs/MODEL-REGISTRY.md — candidate ranking and cross-architecture pilot order.
  • docs/REPRODUCIBILITY-AUDIT.md — prioritized gaps, acceptance criteria, and the cross-architecture pilot protocol.
  • scripts/pilot/inspect_model_targets.py — no-weight/no-GPU target census.
  • scripts/pilot/run_cross_arch_pilot.sh — explicit, preflighted pilot runner.

The current Qwen3 ladder should be reported as adaptive validation retention. Do not call an in-memory STE result PQ2_0/Bonsai-2 parity until the packed quantizer/export and persistent-basis checks in the audit are complete.

Install

python -m venv .venv && . .venv/bin/activate
pip install -r requirements.txt
pip install -e .

The offline suite needs only numpy/pyyaml/pytest. The gate and pilot scripts need torch/transformers/datasets; on ROCm install torch from the ROCm index. Qwen3.5/Qwen3.8 cross-architecture runs require the qwen35 extra (the recorded node uses Transformers 5.5.0).

Reproduce

Download the released model and (for forensics) the public base. Both are Apache-2.0:

huggingface-cli download prism-ml/Ternary-Bonsai-2-27B-gguf \
    Ternary-Bonsai-2-27B-PQ2_0.gguf --local-dir artifacts/oracle/bonsai27
huggingface-cli download Qwen/Qwen3.8-27B \
    --local-dir artifacts/base27

Then, from the repo root:

# Gate 1: does rotate+RTN in their basis reproduce their trits?
python -m bonsai_forensics.gate1_forensics \
    --gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --layers 0,3,62,63 --out artifacts/gate1/gate1-report.json

# Gate 2: swap all PQ2_0 payloads for rotate+RTN and evaluate
python -m bonsai_forensics.gate2_rtn_artifact \
    --source artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --dest artifacts/gate2/patched.gguf \
    --manifest artifacts/oracle/bonsai27/hadamard-manifest.json \
    --base-dir artifacts/base27 --out artifacts/gate2/patch-report.json

# Gate 3 diagnostics (see scripts/gate3/*.py for their arguments)
python scripts/gate3/scr_test.py
python scripts/gate3/prefix27.py --layers 0,3 --windows 8
python scripts/gate3/sweep27.py
python scripts/gate3/hobj27.py

# Evaluate a GGUF (one heavy ROCm process at a time)
python scripts/eval_llama_server.py smoke \
    --server-bin /path/to/llama-server \
    --gguf artifacts/oracle/bonsai27/Ternary-Bonsai-2-27B-PQ2_0.gguf \
    --ngl 99 --ctx 4096 --no-thinking --out artifacts/eval/smoke.json

The scripts take their artifact locations from the repo root; set HIP_VISIBLE_DEVICES / CUDA_VISIBLE_DEVICES to pin the device.

Tests

pytest tests -q
python -m bonsai_forensics.spec_hash --check

CI (.github/workflows/ci.yml) runs pytest tests -q on Python 3.13 with CPU torch. The suite is offline and uses synthetic tensors only. It pins the frozen wire contract (bonsai_forensics/spec.md, canonical hash 0d2c008b4aee726351f9b90e44ec003c18b579d8690db24c77a089d9e1fc652b).

Caveats

One corpus (about 1 MB of Shakespeare), 8x512 context, single runs each. The effect sizes are far outside the noise and the payload-swap experiment is causal, but this is not a benchmark suite. We did not run 27B QAT, so we cannot claim to reproduce Prism's numbers; we localized the gap and identified the only remaining route.

Provenance

This repository was extracted from a private research project. Task identifiers (T30, T31, T32, ...), round numbers, and references to "project records" / "project handoff" in the docs and research write-ups refer to that project's internal coordination log, which is not published. Commit hashes and artifact paths cited as evidence in docs/FAILURES.md are provenance from the original working repository; the reports in research/ and the numbers reproduced here are the authoritative record. No model weights are redistributed.

License and attribution

Apache-2.0 (see LICENSE and NOTICE). The analyzed model, Bonsai 2 27B, is by Prism ML (Apache-2.0), derived from Qwen3.8-27B by Alibaba Cloud (Apache-2.0); "Created using Bonsai by Prism ML." The public PTQ reference reproduced by bonsai_forensics/reference.py is ThakiCloud/bonsai-1bit-repro (Apache-2.0). Runtime evaluation uses Prism's llama.cpp fork (MIT) on ggml (MIT).

bonsai
gguf
llm
mixture-of-experts
model-compression
quantization
reproducibility
ternary-quantization