7dollarbooks/bonsai2-logit-bias-test

PowerShell

0

3 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens (r/LocalLLaMA)

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other…

2

Oct 2, 2026

README

−2 logit bias on Bonsai 2 27B (RTX 5060 Laptop)

Run by Joseph Murray Adams.

Date: 2026-10-02
Result: on a fixed 50-question MATH-500 sample, a −2 logit bias on “wait”, “maybe”, and “perhaps” changed accuracy from 44/50 to 43/50 and average length from 845.2 to 872.3 tokens. Tokens per second stayed 29.3. This does not reproduce the Qwen3.5-4B report of higher accuracy and fewer tokens.

Suggested post title: −2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

Reproduction record

FieldValue
Date2026-10-02
Model fileTernary-Bonsai-2-27B-PTQ1_0.gguf
Model size5.53 GiB, 26.90 B params, 1.75 bpw ternary, group 128
Model repohttps://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
Forkhttps://github.com/PrismML-Eng/llama.cpp
Releaseprism-b10743-adfffbe (2026-09-25)
Buildadfffbe41 (10743). Server fingerprint b10743-adfffbe41
Windows assetsllama-prism-b10743-adfffbe-bin-win-cuda-13.3-x64.zip and cudart-llama-bin-win-cuda-13.3-x64.zip
GPUNVIDIA GeForce RTX 5060 Laptop, Blackwell, compute capability 12.0, 8123 MiB
DriverGame Ready 617.14 notebook, CUDA user-mode 13.4. Power max limit 85 W, NitroSense Turbo, on AC. Load temperature about 76°C
Context4096 (-c 4096)
GPU layers-ngl 99, flash attention -fa on. Weights resident at 6508 MiB
Sampling, both arms--temp 0 --top-k 40 --top-p 0.95 --min-p 0 --seed 42
Reasoning--reasoning-budget 2048 (Prism’s Medium cap). Generation cap -n 3072. Parallel 1
BenchmarkHuggingFaceH4/MATH-500, 50 questions, one shared order
ScoreLast \boxed{} span, then strip whitespace, \left, \right, and backticks. Exact match against the dataset answer. A truncated reply scores 0

Run A is the server above with no logit bias. Run B is the same server plus:

--logit-bias 11158-2 --logit-bias 3655-2 --logit-bias 13784-2 --logit-bias 35542-2 --logit-bias 6970-2 --logit-bias 20734-2 --logit-bias 63068-2 --logit-bias 8106-2 --logit-bias 30442-2
FormToken id
wait11158
wait3655
Wait13784
maybe35542
maybe6970
Maybe20734
perhaps63068
perhaps8106
Perhaps30442

Each form was one token (/tokenize, add_special=false, with_pieces=true). The bias is −2, not a ban.

Prompt, both arms:

Solve this problem. Put the final answer in \boxed{}.

<problem>

Question order: .NET Random(42) shuffle of the 500-row test split, first 50, saved as questions-50.json and reused for run B. Runner: run_math500.ps1 in this folder. Set $root (line 10) to your own folder before running it.

Prism’s published thinking recipe is temperature 1.0, top-k 20, min-p 0.05. Their benchmark tables used min-p 0.0. This test holds temperature 0 so the bias is the only difference. These accuracy numbers are not a Prism leaderboard row.

Results

Runncorrectaccuracyavg completion tokensavg tokens/struncations
A baseline504488%845.229.272
B −2 hedge bias504386%872.329.282

Accuracy −2 points. Length +3.2% (845.2 → 872.3). Speed unchanged.

Forty-seven outcomes matched. Three changed:

QuestionAB
precalculus/1199wrong, truncated at 3072correct, 2221 tokens
prealgebra/1733correct, 2285 tokenswrong, truncated at 3072
precalculus/441correct, 2308 tokenswrong, 1852 tokens

Misses that stayed misses: number_theory/598, algebra/2253, prealgebra/1742, intermediate_algebra/662, intermediate_algebra/1454.

Per question, completion tokens. A t marks a reply that hit the 3072 cap.

#idAB
1test/precalculus/1201.json16312413
2test/algebra/849.json8585
3test/precalculus/801.json23942381
4test/number_theory/598.json729729
5test/intermediate_algebra/623.json12601276
6test/geometry/967.json156162
7test/algebra/2264.json494498
8test/algebra/2253.json279279
9test/number_theory/1257.json485485
10test/precalculus/1313.json17651988
11test/counting_and_probability/1114.json194194
12test/intermediate_algebra/1247.json15221978
13test/geometry/483.json379378
14test/prealgebra/1742.json204204
15test/intermediate_algebra/662.json3072 t2886
16test/number_theory/427.json473473
17test/prealgebra/1733.json22853072 t
18test/intermediate_algebra/1454.json23953072 t
19test/algebra/1332.json192192
20test/precalculus/986.json817813
21test/prealgebra/1640.json446446
22test/algebra/346.json127127
23test/prealgebra/192.json160160
24test/precalculus/1199.json3072 t2221
25test/algebra/2232.json210210
26test/algebra/1457.json314314
27test/algebra/1072.json449449
28test/counting_and_probability/430.json11191119
29test/number_theory/691.json147147
30test/algebra/2476.json261261
31test/counting_and_probability/525.json14161496
32test/geometry/353.json263262
33test/precalculus/1105.json108106
34test/algebra/170.json277272
35test/algebra/769.json160162
36test/prealgebra/1558.json254291
37test/precalculus/441.json23081852
38test/intermediate_algebra/834.json615615
39test/intermediate_algebra/1837.json387386
40test/number_theory/1002.json929929
41test/counting_and_probability/14.json21382130
42test/algebra/1004.json129129
43test/precalculus/1146.json25052243
44test/algebra/2046.json344344
45test/prealgebra/1995.json282282
46test/precalculus/541.json316316
47test/prealgebra/1572.json200184
48test/algebra/1837.json10111011
49test/geometry/1140.json989989
50test/geometry/795.json513604

Files in this repository

FileContents
run_math500.ps1The script that sent the 50 questions
questions-50.jsonThe 50 problems, in order
token-ids.csvHedge-word token ids
results-A.csv, results-B.csvPer-question scores
summary-A.txt, summary-B.txtOne-line summaries
raw-A/, raw-B/score row followed by the server response, two JSON objects back to back
server-A.log, server-B.logllama-server logs

What this is not

A LocalLLaMA report on Qwen3.5-4B saw higher MATH-500 accuracy and fewer tokens from a −2 bias on these words. The paper “Wait, We Don’t Need to ‘Wait’!” shortened traces and held accuracy, and it did not use a flat −2 bias. A −2 bias only makes the tokens less likely.

Simon Willison’s 20–44 tokens/s figures are Apple Silicon. This run is a Windows laptop at about 30 tokens/s while answering, and 31.69 tokens/s on llama-bench TG128.

Separate speed row

llama-bench on the same file, same GPU, same day, AC power, 85 W cap:

.\llama-bench.exe -m C:\Qu_models\bonsai\Ternary-Bonsai-2-27B-PTQ1_0.gguf -p 512 -n 128 -ngl 99 -fa 1
Testtokens/s
pp512711.71 ± 15.69
tg12831.69 ± 0.13

The bench line names the architecture inside the GGUF as qwen35 27B PTQ1_0. The file is Ternary-Bonsai-2-27B-PTQ1_0.gguf. PQ2_0 was not tested. This row is for Prism’s community-benchmarks page, not for the logit-bias thread. There is no RTX 5060 Laptop entry in that table. Template: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/community-benchmarks/bonsai2/TEMPLATE-llama-cpp.md

7dollarbooks/bonsai2-logit-bias-test

PowerShell

0

3 commits

updated Oct 2, 2026

See the code

See what people are saying

SourceMessageScoreDate

−2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens (r/LocalLLaMA)

A recent post here reported that a −2 logit bias on "wait", "maybe", and "perhaps" made Qwen3.5-4B more accurate and shorter on 50 MATH-500 questions. I tried it on Ternary Bonsai 2 27B (PTQ1_0, 5.53 GiB) on an RTX 5060 Laptop 8 GB under Windows, Prism llama.cpp build adfffbe41. It went the other…

2

Oct 2, 2026

README

−2 logit bias on Bonsai 2 27B (RTX 5060 Laptop)

Run by Joseph Murray Adams.

Date: 2026-10-02
Result: on a fixed 50-question MATH-500 sample, a −2 logit bias on “wait”, “maybe”, and “perhaps” changed accuracy from 44/50 to 43/50 and average length from 845.2 to 872.3 tokens. Tokens per second stayed 29.3. This does not reproduce the Qwen3.5-4B report of higher accuracy and fewer tokens.

Suggested post title: −2 logit bias on Bonsai 2 27B: 44/50 → 43/50 on MATH-500, +3% tokens

Reproduction record

FieldValue
Date2026-10-02
Model fileTernary-Bonsai-2-27B-PTQ1_0.gguf
Model size5.53 GiB, 26.90 B params, 1.75 bpw ternary, group 128
Model repohttps://huggingface.co/prism-ml/Ternary-Bonsai-2-27B-gguf
Forkhttps://github.com/PrismML-Eng/llama.cpp
Releaseprism-b10743-adfffbe (2026-09-25)
Buildadfffbe41 (10743). Server fingerprint b10743-adfffbe41
Windows assetsllama-prism-b10743-adfffbe-bin-win-cuda-13.3-x64.zip and cudart-llama-bin-win-cuda-13.3-x64.zip
GPUNVIDIA GeForce RTX 5060 Laptop, Blackwell, compute capability 12.0, 8123 MiB
DriverGame Ready 617.14 notebook, CUDA user-mode 13.4. Power max limit 85 W, NitroSense Turbo, on AC. Load temperature about 76°C
Context4096 (-c 4096)
GPU layers-ngl 99, flash attention -fa on. Weights resident at 6508 MiB
Sampling, both arms--temp 0 --top-k 40 --top-p 0.95 --min-p 0 --seed 42
Reasoning--reasoning-budget 2048 (Prism’s Medium cap). Generation cap -n 3072. Parallel 1
BenchmarkHuggingFaceH4/MATH-500, 50 questions, one shared order
ScoreLast \boxed{} span, then strip whitespace, \left, \right, and backticks. Exact match against the dataset answer. A truncated reply scores 0

Run A is the server above with no logit bias. Run B is the same server plus:

--logit-bias 11158-2 --logit-bias 3655-2 --logit-bias 13784-2 --logit-bias 35542-2 --logit-bias 6970-2 --logit-bias 20734-2 --logit-bias 63068-2 --logit-bias 8106-2 --logit-bias 30442-2
FormToken id
wait11158
wait3655
Wait13784
maybe35542
maybe6970
Maybe20734
perhaps63068
perhaps8106
Perhaps30442

Each form was one token (/tokenize, add_special=false, with_pieces=true). The bias is −2, not a ban.

Prompt, both arms:

Solve this problem. Put the final answer in \boxed{}.

<problem>

Question order: .NET Random(42) shuffle of the 500-row test split, first 50, saved as questions-50.json and reused for run B. Runner: run_math500.ps1 in this folder. Set $root (line 10) to your own folder before running it.

Prism’s published thinking recipe is temperature 1.0, top-k 20, min-p 0.05. Their benchmark tables used min-p 0.0. This test holds temperature 0 so the bias is the only difference. These accuracy numbers are not a Prism leaderboard row.

Results

Runncorrectaccuracyavg completion tokensavg tokens/struncations
A baseline504488%845.229.272
B −2 hedge bias504386%872.329.282

Accuracy −2 points. Length +3.2% (845.2 → 872.3). Speed unchanged.

Forty-seven outcomes matched. Three changed:

QuestionAB
precalculus/1199wrong, truncated at 3072correct, 2221 tokens
prealgebra/1733correct, 2285 tokenswrong, truncated at 3072
precalculus/441correct, 2308 tokenswrong, 1852 tokens

Misses that stayed misses: number_theory/598, algebra/2253, prealgebra/1742, intermediate_algebra/662, intermediate_algebra/1454.

Per question, completion tokens. A t marks a reply that hit the 3072 cap.

#idAB
1test/precalculus/1201.json16312413
2test/algebra/849.json8585
3test/precalculus/801.json23942381
4test/number_theory/598.json729729
5test/intermediate_algebra/623.json12601276
6test/geometry/967.json156162
7test/algebra/2264.json494498
8test/algebra/2253.json279279
9test/number_theory/1257.json485485
10test/precalculus/1313.json17651988
11test/counting_and_probability/1114.json194194
12test/intermediate_algebra/1247.json15221978
13test/geometry/483.json379378
14test/prealgebra/1742.json204204
15test/intermediate_algebra/662.json3072 t2886
16test/number_theory/427.json473473
17test/prealgebra/1733.json22853072 t
18test/intermediate_algebra/1454.json23953072 t
19test/algebra/1332.json192192
20test/precalculus/986.json817813
21test/prealgebra/1640.json446446
22test/algebra/346.json127127
23test/prealgebra/192.json160160
24test/precalculus/1199.json3072 t2221
25test/algebra/2232.json210210
26test/algebra/1457.json314314
27test/algebra/1072.json449449
28test/counting_and_probability/430.json11191119
29test/number_theory/691.json147147
30test/algebra/2476.json261261
31test/counting_and_probability/525.json14161496
32test/geometry/353.json263262
33test/precalculus/1105.json108106
34test/algebra/170.json277272
35test/algebra/769.json160162
36test/prealgebra/1558.json254291
37test/precalculus/441.json23081852
38test/intermediate_algebra/834.json615615
39test/intermediate_algebra/1837.json387386
40test/number_theory/1002.json929929
41test/counting_and_probability/14.json21382130
42test/algebra/1004.json129129
43test/precalculus/1146.json25052243
44test/algebra/2046.json344344
45test/prealgebra/1995.json282282
46test/precalculus/541.json316316
47test/prealgebra/1572.json200184
48test/algebra/1837.json10111011
49test/geometry/1140.json989989
50test/geometry/795.json513604

Files in this repository

FileContents
run_math500.ps1The script that sent the 50 questions
questions-50.jsonThe 50 problems, in order
token-ids.csvHedge-word token ids
results-A.csv, results-B.csvPer-question scores
summary-A.txt, summary-B.txtOne-line summaries
raw-A/, raw-B/score row followed by the server response, two JSON objects back to back
server-A.log, server-B.logllama-server logs

What this is not

A LocalLLaMA report on Qwen3.5-4B saw higher MATH-500 accuracy and fewer tokens from a −2 bias on these words. The paper “Wait, We Don’t Need to ‘Wait’!” shortened traces and held accuracy, and it did not use a flat −2 bias. A −2 bias only makes the tokens less likely.

Simon Willison’s 20–44 tokens/s figures are Apple Silicon. This run is a Windows laptop at about 30 tokens/s while answering, and 31.69 tokens/s on llama-bench TG128.

Separate speed row

llama-bench on the same file, same GPU, same day, AC power, 85 W cap:

.\llama-bench.exe -m C:\Qu_models\bonsai\Ternary-Bonsai-2-27B-PTQ1_0.gguf -p 512 -n 128 -ngl 99 -fa 1
Testtokens/s
pp512711.71 ± 15.69
tg12831.69 ± 0.13

The bench line names the architecture inside the GGUF as qwen35 27B PTQ1_0. The file is Ternary-Bonsai-2-27B-PTQ1_0.gguf. PQ2_0 was not tested. This row is for Prism’s community-benchmarks page, not for the logit-bias thread. There is no RTX 5060 Laptop entry in that table. Template: https://github.com/PrismML-Eng/Bonsai-demo/blob/main/community-benchmarks/bonsai2/TEMPLATE-llama-cpp.md

Languages

PowerShell

100.0%