prism-ml/Ternary-Bonsai-4B-unpacked

Model

Ternary-Bonsai-4B — Unpacked FP16 Safetensors

7

1 commits

1 linked in READMEs

updated Apr 16, 2026

See the code

README

Ternary-Bonsai-4B — Unpacked FP16 Safetensors

FP16 safetensors (HuggingFace format) of the ternary Bonsai-4B model. This repo exists for users who want to run Ternary Bonsai with stock HuggingFace tooling or frameworks that don't yet support any of the packed ternary format. The MLX 2-bit format is currently the only packed format available; more formats for other backends are coming soon.

We strongly recommend using the natively packed models instead. The packed format is where all the benefits of Bonsai come from — up to 9x memory reduction, 5x faster inference, and lower energy per token. This unpacked FP16 version is full-size and does not provide any of those advantages.

For the optimized ternary release model (recommended):

bonsai
prismml
qwen3
safetensors
ternary

prism-ml/Ternary-Bonsai-4B-unpacked

Model

Ternary-Bonsai-4B — Unpacked FP16 Safetensors

7

1 commits

1 linked in READMEs

updated Apr 16, 2026

See the code

README

Ternary-Bonsai-4B — Unpacked FP16 Safetensors

FP16 safetensors (HuggingFace format) of the ternary Bonsai-4B model. This repo exists for users who want to run Ternary Bonsai with stock HuggingFace tooling or frameworks that don't yet support any of the packed ternary format. The MLX 2-bit format is currently the only packed format available; more formats for other backends are coming soon.

We strongly recommend using the natively packed models instead. The packed format is where all the benefits of Bonsai come from — up to 9x memory reduction, 5x faster inference, and lower energy per token. This unpacked FP16 version is full-size and does not provide any of those advantages.

For the optimized ternary release model (recommended):

bonsai
prismml
qwen3
safetensors
ternary