senna-lang/bonsai-4b-system-one

Lightweight System One model for edge devices: tree-masked multi-question inference on frozen Ternary-Bonsai-4B (MLX/PyTorch)

Python

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a lightweight, local Jev-like System One with Ternary-Bonsai-4B — and used it as a coding-agent judge (r/LocalLLaMA)

I built a local Jev-like System One on top of Ternary-Bonsai-4B. It takes one record and answers multiple choice, rating, or yes/no questions about it in one forward pass. I adapted the inference path to share the record prefix across questions. A tree attention mask lets each question attend to…

0

Oct 6, 2026

README

bonsai-4b-system-one

A lightweight System One model for edge devices, built on Ternary-Bonsai-4B. In the style of Jev, it answers many choice, rating, and yes/no questions about one record in a single forward pass, returning a probability per option instead of generating text.

The model itself is unchanged: it stays frozen, with no adapter and no fine-tuning. The System One behavior comes from the inference code: every question becomes a branch after the shared record, and a tree attention mask lets each branch see the record and its own tokens but never another branch. The record is encoded once, and each answer is read from the native LM head.

On Apple Silicon it runs from the 2-bit packed weights (about 1.1 GB) with MLX; a PyTorch backend uses the FP16 weights. A version with a LoRA adapter trained for this request format is planned separately.

Experimental research artifact. English-oriented; not for high-stakes decisions. Not affiliated with PrismML, Qwen/Alibaba Cloud, or TypeSafe.

Accuracy

Zero-shot accuracy on public datasets that were not used to design the prompt or readout (500 rows each, sampled with a fixed seed; protocol in docs/BENCHMARKS.md):

DatasetTaskTypeAccuracy95% CI
AG Newsnews topicchoice, 4 options0.8680.838–0.896
MASSIVE (en-US)voice-assistant scenariochoice, 18 options0.5620.518–0.604
MNLI (matched)entailment / neutral / contradictionchoice, 3 options0.8240.790–0.856
BoolQyes/no question about a passagenoul0.8360.802–0.868
SST-55-level sentimentscore, 5 levels0.4020.360–0.444

SST-5 mean absolute error: 0.77 levels. One canonical option order, single forward pass per request, MLX on an Apple M2. Row-level predictions: results/accuracy/.

Score-type questions (ratings) are the weakest point. See Limitations.

Speed

Apple M2 (24 GB), public fixture (16 records × 3 repeats), median request latency, model loading excluded:

Questions per record124816
MLX, one tree-masked pass0.53 s0.69 s0.98 s1.77 s3.43 s
MLX, one pass per question0.53 s1.05 s2.00 s4.18 s8.66 s
Speedup, MLX1.00×1.51×2.03×2.37×2.52×
Speedup, PyTorch MPS0.97×1.60×2.12×2.46×2.63×

On the Mac, time grows with token count, so the gain comes from encoding the record once (581 instead of 1,324 tokens at N = 16). CUDA latency has not been measured. Details in docs/RESULTS.md and results/latency/.

How it works

flowchart LR
    R[Shared prefix: record] --> B1[Branch 1: question + options]
    R --> B2[Branch 2: question + options]
    R --> B3[Branch N: ...]
    B1 & B2 & B3 --> F[One tree-masked forward pass]
    F --> P[Answer-letter probabilities per branch]
  • Each question is rendered as a chat prompt. The prompts' common token prefix (the record) is stored once; each question's suffix becomes a branch.
  • A causal tree attention mask lets a branch attend to the shared prefix and its own tokens, never to another branch. Position IDs restart after the prefix for every branch.
  • The native LM head reads the answer-letter logits (A, B, …) at the end of each branch. Nothing is generated.

The model weights are not modified; everything above is in this code. See docs/METHOD.md.

Install

PyTorch and MLX pin different transformers versions; use separate environments.

git clone https://github.com/senna-lang/bonsai-4b-system-one && cd bonsai-4b-system-one
python -m venv .venv && . .venv/bin/activate

pip install -e '.[mlx]'     # Apple Silicon: 2-bit packed weights, ~1.1 GB
# or
pip install -e '.[torch]'   # CUDA, Apple MPS, or CPU: FP16 weights, ~8 GB

Run the examples

python examples/run_example.py --backend mlx   # or --backend torch

The script downloads the pinned Bonsai weights from the Hugging Face Hub on first use, then scores the fictional requests in examples/requests.jsonl.

Python API

from system_one_bonsai import Branch, PackedRequest
from system_one_bonsai.inference_mlx import load_model, predict_request   # or system_one_bonsai.inference

model, tokenizer = load_model()
request = PackedRequest(
    state="A customer received a replacement device after the first one stopped charging.",
    branches=[
        Branch("choice", "Which issue is described?", ["Delivery delay", "Charging failure", "Billing error"]),
        Branch("noul", "Does the record mention a replacement?"),
    ],
)
probabilities = predict_request(model, tokenizer, request)
# probabilities[0]: [P(Delivery delay), P(Charging failure), P(Billing error)]
# probabilities[1]: [P(No), P(Yes)]
Branch kindoptionsOutput
choice2–20 optionsone probability per option
score2–20 ordered rating levelsone probability per level; compute sum(level * p) with your own level values for an expected score
noulempty (fixed [No, Yes])[P(No), P(Yes)]

Probabilities are conditional on the listed options and are not guaranteed to be calibrated. This code uses one canonical option order; it does not average over option orders.

Local server (System One API format)

Loading the model takes a few seconds, so for repeated use keep it resident behind a local HTTP endpoint:

python -m system_one_bonsai.serve            # MLX on 127.0.0.1:8765; --backend torch, --host, --port

POST /v1/systemone accepts and returns the request and answer shapes that System One API clients use, so a client of that API can point its base URL here:

curl -s localhost:8765/v1/systemone -d '{
  "state": {"ticket": "I was charged twice for my order and want my money back."},
  "questions": {
    "team":    {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "Payments and refunds", "shipping": "Delivery", "tech": "Product issues"}},
    "urgency": {"type": "score",  "instructions": "How urgent is this?", "criteria": ["low", "medium", "high"]},
    "refund":  {"type": "noul",   "instructions": "Does the customer ask for a refund?"}
  }}'
# {"answers": {"team": {"type": "choice", "choice": "billing", "probabilities": {...}, "confidence": ...},
#              "urgency": {"type": "score", "score": <expected level index>, "confidence": ...},
#              "refund": {"type": "noul", "noul": <P(yes)>}}, "usage": {"input_tokens": ..., "output_tokens": 0}}

How a request maps onto this model: state is rendered as indented JSON and becomes the record; each choice option reads label: description; score levels are the criteria, lowest first; a yes/no question (noul, or bool) gets its true/false criteria appended to the instructions. confidence is the highest option probability. The format was derived from a client implementation, not from a published specification, so compatibility is not guaranteed. The server has no authentication, ignores API keys, binds to localhost by default, and answers one request at a time; do not expose it to a network.

For example, the pi coding agent sends its Jev classifier calls here with this entry in ~/.pi/agent/models.json (verified with pi's classify()):

{ "providers": { "typesafe": { "baseUrl": "http://127.0.0.1:8765/v1", "apiKey": "local" } } }

Tests

pip install -e '.[test]' && pytest

The tests use a stand-in tokenizer and need no model weights. They check request validation, prefix sharing, branch spans, position IDs, and the local server's request translation and answer shapes.

Verification

fixtures/behavior.jsonl is a small fictional fixture (16 records, 42 branches) for checking behavior with the real model:

  • scripts/check_branch_isolation.py — a branch's probabilities do not change when other branches are added, reordered, or removed (max diff 0.0034 MLX, 0.0039 MPS).
  • scripts/compare_backends.py — PyTorch (MPS) and MLX agree (42/42 argmax, max diff 0.0035).
  • scripts/bench_latency.py — request latency of tree-masked packing vs one forward per question.
  • scripts/bench_accuracy.py — the public accuracy benchmark above.

See docs/RESULTS.md for commands and numbers.

Limitations

The numbers below were measured in the development project on data that is not distributed (500 held-out branches, compared with the soft answer distributions of a DeepSeek Flash teacher, not human labels, option orders averaged, MLX). They are reported for context and cannot be re-run from this repository.

Branch kindnThis model (no adapter)Bonsai + LoRA trained for this format
choice1630.9080.890
noul1360.8530.890
score2010.7010.771
all5000.8100.842

Teacher top-1 agreement. Without an adapter, choice questions are as good as with one; score questions lose 7 points (paired 95% CI 1.5–12.9 points), and noul probabilities are less calibrated (ECE 0.091 vs 0.030). Use this model mainly for choice and yes/no questions; treat score probabilities with care.

License

Apache-2.0 (see LICENSE). The Bonsai model is obtained separately under its own license (Apache-2.0). See THIRD_PARTY.md.

apple-silicon
bonsai
classification
decision-model
huggingface
jev
mlx
python
pytorch
ternary
tree-attention
typed-decisions

senna-lang/bonsai-4b-system-one

Lightweight System One model for edge devices: tree-masked multi-question inference on frozen Ternary-Bonsai-4B (MLX/PyTorch)

Python

0

5 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a lightweight, local Jev-like System One with Ternary-Bonsai-4B — and used it as a coding-agent judge (r/LocalLLaMA)

I built a local Jev-like System One on top of Ternary-Bonsai-4B. It takes one record and answers multiple choice, rating, or yes/no questions about it in one forward pass. I adapted the inference path to share the record prefix across questions. A tree attention mask lets each question attend to…

0

Oct 6, 2026

README

bonsai-4b-system-one

A lightweight System One model for edge devices, built on Ternary-Bonsai-4B. In the style of Jev, it answers many choice, rating, and yes/no questions about one record in a single forward pass, returning a probability per option instead of generating text.

The model itself is unchanged: it stays frozen, with no adapter and no fine-tuning. The System One behavior comes from the inference code: every question becomes a branch after the shared record, and a tree attention mask lets each branch see the record and its own tokens but never another branch. The record is encoded once, and each answer is read from the native LM head.

On Apple Silicon it runs from the 2-bit packed weights (about 1.1 GB) with MLX; a PyTorch backend uses the FP16 weights. A version with a LoRA adapter trained for this request format is planned separately.

Experimental research artifact. English-oriented; not for high-stakes decisions. Not affiliated with PrismML, Qwen/Alibaba Cloud, or TypeSafe.

Accuracy

Zero-shot accuracy on public datasets that were not used to design the prompt or readout (500 rows each, sampled with a fixed seed; protocol in docs/BENCHMARKS.md):

DatasetTaskTypeAccuracy95% CI
AG Newsnews topicchoice, 4 options0.8680.838–0.896
MASSIVE (en-US)voice-assistant scenariochoice, 18 options0.5620.518–0.604
MNLI (matched)entailment / neutral / contradictionchoice, 3 options0.8240.790–0.856
BoolQyes/no question about a passagenoul0.8360.802–0.868
SST-55-level sentimentscore, 5 levels0.4020.360–0.444

SST-5 mean absolute error: 0.77 levels. One canonical option order, single forward pass per request, MLX on an Apple M2. Row-level predictions: results/accuracy/.

Score-type questions (ratings) are the weakest point. See Limitations.

Speed

Apple M2 (24 GB), public fixture (16 records × 3 repeats), median request latency, model loading excluded:

Questions per record124816
MLX, one tree-masked pass0.53 s0.69 s0.98 s1.77 s3.43 s
MLX, one pass per question0.53 s1.05 s2.00 s4.18 s8.66 s
Speedup, MLX1.00×1.51×2.03×2.37×2.52×
Speedup, PyTorch MPS0.97×1.60×2.12×2.46×2.63×

On the Mac, time grows with token count, so the gain comes from encoding the record once (581 instead of 1,324 tokens at N = 16). CUDA latency has not been measured. Details in docs/RESULTS.md and results/latency/.

How it works

flowchart LR
    R[Shared prefix: record] --> B1[Branch 1: question + options]
    R --> B2[Branch 2: question + options]
    R --> B3[Branch N: ...]
    B1 & B2 & B3 --> F[One tree-masked forward pass]
    F --> P[Answer-letter probabilities per branch]
  • Each question is rendered as a chat prompt. The prompts' common token prefix (the record) is stored once; each question's suffix becomes a branch.
  • A causal tree attention mask lets a branch attend to the shared prefix and its own tokens, never to another branch. Position IDs restart after the prefix for every branch.
  • The native LM head reads the answer-letter logits (A, B, …) at the end of each branch. Nothing is generated.

The model weights are not modified; everything above is in this code. See docs/METHOD.md.

Install

PyTorch and MLX pin different transformers versions; use separate environments.

git clone https://github.com/senna-lang/bonsai-4b-system-one && cd bonsai-4b-system-one
python -m venv .venv && . .venv/bin/activate

pip install -e '.[mlx]'     # Apple Silicon: 2-bit packed weights, ~1.1 GB
# or
pip install -e '.[torch]'   # CUDA, Apple MPS, or CPU: FP16 weights, ~8 GB

Run the examples

python examples/run_example.py --backend mlx   # or --backend torch

The script downloads the pinned Bonsai weights from the Hugging Face Hub on first use, then scores the fictional requests in examples/requests.jsonl.

Python API

from system_one_bonsai import Branch, PackedRequest
from system_one_bonsai.inference_mlx import load_model, predict_request   # or system_one_bonsai.inference

model, tokenizer = load_model()
request = PackedRequest(
    state="A customer received a replacement device after the first one stopped charging.",
    branches=[
        Branch("choice", "Which issue is described?", ["Delivery delay", "Charging failure", "Billing error"]),
        Branch("noul", "Does the record mention a replacement?"),
    ],
)
probabilities = predict_request(model, tokenizer, request)
# probabilities[0]: [P(Delivery delay), P(Charging failure), P(Billing error)]
# probabilities[1]: [P(No), P(Yes)]
Branch kindoptionsOutput
choice2–20 optionsone probability per option
score2–20 ordered rating levelsone probability per level; compute sum(level * p) with your own level values for an expected score
noulempty (fixed [No, Yes])[P(No), P(Yes)]

Probabilities are conditional on the listed options and are not guaranteed to be calibrated. This code uses one canonical option order; it does not average over option orders.

Local server (System One API format)

Loading the model takes a few seconds, so for repeated use keep it resident behind a local HTTP endpoint:

python -m system_one_bonsai.serve            # MLX on 127.0.0.1:8765; --backend torch, --host, --port

POST /v1/systemone accepts and returns the request and answer shapes that System One API clients use, so a client of that API can point its base URL here:

curl -s localhost:8765/v1/systemone -d '{
  "state": {"ticket": "I was charged twice for my order and want my money back."},
  "questions": {
    "team":    {"type": "choice", "instructions": "Which team should handle this?", "criteria": {"billing": "Payments and refunds", "shipping": "Delivery", "tech": "Product issues"}},
    "urgency": {"type": "score",  "instructions": "How urgent is this?", "criteria": ["low", "medium", "high"]},
    "refund":  {"type": "noul",   "instructions": "Does the customer ask for a refund?"}
  }}'
# {"answers": {"team": {"type": "choice", "choice": "billing", "probabilities": {...}, "confidence": ...},
#              "urgency": {"type": "score", "score": <expected level index>, "confidence": ...},
#              "refund": {"type": "noul", "noul": <P(yes)>}}, "usage": {"input_tokens": ..., "output_tokens": 0}}

How a request maps onto this model: state is rendered as indented JSON and becomes the record; each choice option reads label: description; score levels are the criteria, lowest first; a yes/no question (noul, or bool) gets its true/false criteria appended to the instructions. confidence is the highest option probability. The format was derived from a client implementation, not from a published specification, so compatibility is not guaranteed. The server has no authentication, ignores API keys, binds to localhost by default, and answers one request at a time; do not expose it to a network.

For example, the pi coding agent sends its Jev classifier calls here with this entry in ~/.pi/agent/models.json (verified with pi's classify()):

{ "providers": { "typesafe": { "baseUrl": "http://127.0.0.1:8765/v1", "apiKey": "local" } } }

Tests

pip install -e '.[test]' && pytest

The tests use a stand-in tokenizer and need no model weights. They check request validation, prefix sharing, branch spans, position IDs, and the local server's request translation and answer shapes.

Verification

fixtures/behavior.jsonl is a small fictional fixture (16 records, 42 branches) for checking behavior with the real model:

  • scripts/check_branch_isolation.py — a branch's probabilities do not change when other branches are added, reordered, or removed (max diff 0.0034 MLX, 0.0039 MPS).
  • scripts/compare_backends.py — PyTorch (MPS) and MLX agree (42/42 argmax, max diff 0.0035).
  • scripts/bench_latency.py — request latency of tree-masked packing vs one forward per question.
  • scripts/bench_accuracy.py — the public accuracy benchmark above.

See docs/RESULTS.md for commands and numbers.

Limitations

The numbers below were measured in the development project on data that is not distributed (500 held-out branches, compared with the soft answer distributions of a DeepSeek Flash teacher, not human labels, option orders averaged, MLX). They are reported for context and cannot be re-run from this repository.

Branch kindnThis model (no adapter)Bonsai + LoRA trained for this format
choice1630.9080.890
noul1360.8530.890
score2010.7010.771
all5000.8100.842

Teacher top-1 agreement. Without an adapter, choice questions are as good as with one; score questions lose 7 points (paired 95% CI 1.5–12.9 points), and noul probabilities are less calibrated (ECE 0.091 vs 0.030). Use this model mainly for choice and yes/no questions; treat score probabilities with care.

License

Apache-2.0 (see LICENSE). The Bonsai model is obtained separately under its own license (Apache-2.0). See THIRD_PARTY.md.

apple-silicon
bonsai
classification
decision-model
huggingface
jev
mlx
python
pytorch
ternary
tree-attention
typed-decisions