playtest-graded benchmark for local LLM stacks: a 10-file game build, 19 static checks, a runtime soak, and a human playing it
See the codeDoes your local LLM setup actually work for coding agents, or does it just benchmark well?
The best build this harness ever measured, 4,204 lines of polished arcade game, 17/19 on the static scorer, had a dead keyboard. The model exported its keymap as a Node module and never wired it into the browser. Syntax checks passed.
The two-minute runtime soak ran at 60fps with 2.5M draw calls. Then a human pressed Enter and nothing happened, because the first keydown threw on an undefined global and took the entire input system with it.
That's failure class #14. There are fourteen others, and every one was found the same way: by a person playing the game. Zero were found by static analysis.
That gap is the reason this repo exists.
Not the model. Your whole inference stack: engine, quant, drafter, speculative decoding, chat template, contract wording, thinking budget. Several of these move the result more than the model choice does, and standard benchmarks can't see any of them.
One cell:
A coding agent gets a fixed contract: build "Neon Overdrive", a 10-file HTML5 canvas breakout game. Exact file list, physics rules, serve behavior, pacing constants, all specified, all static-checkable.
Static scorer (19 checks) grades the code: structure, collision correctness, contract compliance, two different velocity-multiplication bug patterns. 3.
Runtime soak (120s, headless Chromium via raw CDP): boot errors are fatal, frames and draw calls must advance, reloads are detected. 4. Behavior probe (behavior-probe.sh): a scripted real-time pass for serve gating (the ball must not move before Space) and brick reflection (the ball must bounce, not pass through).
Catches collision-wiring classes the soak and the static scorer cannot see. 5. You play it. Two minutes with the arrow keys.
This is the grade that matters. The fifteen failure classes in the ledger all live here.
Wall clock: 5-40 minutes depending on model and effort tier. The result line is one comment's worth of data, generated for you.
The subject is validated on llama.cpp (Vulkan) + AMD Strix Halo, Qwen3.8-27B (four quants) and Qwen3.8-Flash-Next 125B-A6B, with DFlash2 and native-MTP speculative decoding. It will work on any local stack that can run a coding agent; the receipts below are from this rig.
Companion repo with the runnable pi-agent stack (server scripts, wiring, extensions): qwen38-strix-halo-harness.
| benchmark | what it measures |
|---|---|
| WebGen-Bench | model capability: multi-file website generation, browser tests |
| LMGame / BALROG | agents playing games |
| neon-ladder | your local configuration, through a real build workload |
Model-capability benchmarks can't catch a chat template that silently ignores your effort flags, or a drafter whose acceptance collapses on reasoning-heavy turns. Those are configuration failures, and they're what this harness is built to surface.
Behind every number published here:
| layer | what I run |
|---|---|
| Hardware | AMD Strix Halo APU, Ryzen AI MAX+ 395, Radeon 8060S iGPU (gfx1151), 109GB usable unified memory |
| OS / driver | CachyOS Linux, Mesa RADV 26.2.1 system driver, boot args amdgpu.gttsize=126976 ttm.pages_limit=32505856 |
| Engine | llama.cpp Vulkan, Nathan's strix-halo-vulkan releases (validated v0.6.11 through v0.7.4.1; 0.7.4+ = throughput parity + greedy repeatability fixes) + my adaptive draft-sizing port |
| Models | Qwen3.8-27B: Unsloth UD-Q4_K_XL-v3 (the 27B pick), Q5/Q6/Q8 for the tier map · Qwen3.8-Flash-Next: Unsloth UD-IQ4_XS (87.25GB, sha-pinned) |
| Drafters | DFlash2 Q4_K_M sidecar (27B; adaptive n3-7 or fixed n4) · native MTP head Q8_0 (Flash-Next; fixed n4) |
| Vision | mmproj-F16.GGUF from each model repo, passed as -mm ... --mmproj-offload on both servers (pi sessions have vision on GPU) |
| KV cache | 27B: f16 target / q8_0 drafter · Flash-Next: q8_0/q8_0 · contexts: 262144 (27B) / 32768-65536 (Flash-Next) |
| Chat template | Sharp v22.4.0 on the 27B (+19% sustained vs stock; flag order matters: see the run recipe) · embedded template on Flash-Next |
| Agent harness | pi coding agent (0.84.x) + pi-llama-cpp provider, thinkingTokenBudgetField: "thinking_budget_tokens", Qwen card sampling (temp 1.0, top_p 0.95, top_k 20) |
| Gates (this bundle) | static scorer · 120s runtime smoke gate · runner with repair-first retry and server-down abort |
Two engine-driver findings that cost real time to learn (three-way A/B, same source commit): a native build beats the portable payload by 6-24% on decode (JSON-class 28.3 → 40.3 t/s), and the system Mesa stable beat the bundled devel snapshot on agent-relevant classes (deep decode +9.5%, emission +13%). Budget GPU memory against mem_info_vram_total + mem_info_gtt_total; both heaps count; the Vulkan allocator uses both. This nominal-128GB machine with a 16GB BIOS UMA carve exposes 16 + 109.7 ≈ 125.7 GiB total, split between a pinned VRAM heap and an evictable GTT heap.
One adaptive-drafting boundary worth knowing before you tune: the acceptance controller wins on DFlash2 block drafts and loses to fixed n4 on MTP chained drafts, measured both directions on the same harness. Use whichever wins on your stack, not whichever is newer.
| file | role |
|---|---|
contract.txt | the game contract: file structure, physics, MENU/START + SERVE rules, workflow |
prompt-multi-v2_12.txt | the bench prompt (current): contract + workflow + attempt marker slot. v2_12 adds MENU/START RULE, SERVE RULE (ball glued to paddle until Space) and the input-wiring verify clause. Published cells to date (FN band, halogen 0.6.1 own-format, halogen 0.7.0 GGUF) ran v2_9 - same contract minus those clauses - so cross-engine cells stay comparable; new cells should use this file. |
game-score.py | static scorer, 19 checks, versioned + md5-frozen |
smoke-gate.sh | 120s runtime soak: boot errors fatal, frame/draw-ops advancement, reload detection |
wsmin.js | raw-WebSocket CDP client the gate polls through |
cdp.js | single-eval CDP client for diagnostics (the gate does not need it) |
run.sh | runner: success = files + SMOKE-OK; boot errors trigger a repair pass, else wipe and reroll |
soak-probe.sh | standalone diagnostic soak for post-mortems |
report.sh | assembles the one-comment result line from run artifacts (auto-invoked by run.sh on success) |
Use the bundled scorer with the bundled contract. They are calibrated as a set. Scores from modified contracts or different scorers are not comparable.
Three dependencies, then clone:
# 1. System deps
sudo pacman -S nodejs chromium # or your distro's equivalents
# 2. The pi coding agent (or build from source: github.com/earendil-works/pi)
yay -S pi-coding-agent-bin
# 3. This bundle
git clone https://github.com/aic0d3r/neon-ladder && cd neon-ladder
The agent harness is a separate repo: qwen38-strix-halo-harness holds the pi extensions, setup.sh (one command: providers, sidecar, halogen launch), start.sh (daily use), crash recovery, and the NPU retrieval toolkit. If you want the full daily-driver stack, install that too and run its ./setup.sh. This repo is the benchmark: game builds, prompts, scorers, results.
For the ladder you only need pi wired to your server. Minimal ~/.pi/agent/models.json entry (the llama.cpp path):
{
"providers": {
"llamacpp": {
"baseUrl": "http://127.0.0.1:8080/v1",
"api": "openai-completions",
"apiKey": "dummy",
"models": [{
"id": "my-model",
"reasoning": true,
"contextWindow": 262144,
"maxTokens": 32768,
"compat": {
"thinkingFormat": "chat-template",
"chatTemplateKwargs": { "reasoning_effort": {"$var": "thinking.effort"}, "enable_thinking": {"$var": "thinking.enabled"} },
"thinkingTokenBudgetField": "thinking_budget_tokens"
}
}]
}
}
}
Start your server, run the cell, play the game. Reference server command:
llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
-md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
-mm mmproj-F16.gguf --mmproj-offload \
--spec-type draft-dflash --spec-draft-n-max 4 \
-ngl all -fa on -ctk f16 -ctv f16 -ctkd q8_0 -ctvd q8_0 \
-c 262144 -b 4096 -ub 4096 --jinja \
--chat-template-file sharp-v22.4.0.jinja \
--host 127.0.0.1 --port 8080 --metrics
# KEY: --jinja must come BEFORE --chat-template-file, and never use
# --chat-template with a file path: it fails silently.
The cell itself:
export LLADDER=$HOME/neon-ladder-work && mkdir -p $LLADDER
bash run.sh build-run1 game-run1 medium "$(cat contract.txt)"
# ... typically 15-45 min later, on success the runner prints your result line
Open build-run1/index.html, press Enter, play two minutes. Then:
QUANT="Q4_K_XL" DRAFTER="DFlash2-Q4_M n4" PLAYTEST="Y - plays well" \
RIG="Strix Halo 395+8060S, Nathan v0.7.4.1" bash report.sh $LLADDER/build-run1 game-run1
That prints (and saves) a line like:
Q8_K_XL / DFlash2-Q4_M n4 / medium / 23min / static 17/19 / SMOKE-OK / Y - clean and polished / Strix Halo 395+8060S, 109GB, Nathan v0.7.4.1
Static score, soak verdict, effort, and wall clock come from the run artifacts automatically. You supply quant, drafter, playtest verdict, and rig, the four things only you know.
Divergent results are data, not contradiction. If you know the rules:
By available GPU memory (check mem_info_vram_total + mem_info_gtt_total, both heaps count):
| budget | stack | what you get |
|---|---|---|
| ~24 GB | Qwen3.8-27B UD-Q4_K_XL + DFlash2 Q4_K_M, fixed n4, ctx ≤64k | the reliable-build config, every 27B receipt above |
| 24–48 GB | same, full 256k ctx | plus IDE/browser headroom, the reference rig |
| 48–91 GB | add the Q8-27B profile for deep-context reading | fastest prefill tier (253 avg / 330 peak t/s) |
| ≥91 GB | Qwen3.8-Flash-Next UD-IQ4_XS + MTP Q8_0, Nathan v0.7.3+ built natively (v0.7.4.1 current), --reasoning-effort medium --reasoning-budget 2048 (server default: clients that send a per-request budget, e.g. pi, override it), this contract | the speed lane and the daily driver on 128GB: 12-min clean medium builds, 40 t/s sustained decode on low-effort sessions, 353 t/s prefill |
By task:
--spec-draft-adaptive n3-7): +25% on tool-call classes. On MTP stacks, stick to fixed n4 (adaptive loses there).Three knobs that matter more than they look:
--chat-template: it fails silently (use --chat-template-file, --jinja first).Every row is a receipt from this harness. Playtest verdicts are the user's, recorded as given.
| stack | contract era | verdict |
|---|---|---|
| Qwen3.8-27B (Q4/Q5/Q6/Q8) + DFlash2 | early revisions | 6/6 playable, PPL plateau 7.079–7.089 |
| 27B + adaptive draft sizing (N=5 cells) | current | 4/5 clean playtests, decode 18.0 vs 17.7 fixed-n4, acceptance 65.3% vs 60.4% |
| 27B low, clean-env control | current | 17/19 (1412 LOC), 39 min, 4.3× the reasoning of the MoE below, quality a wash, 8× the wall time |
| 27B low / medium / high, N=2 | v0.7.4.1 | 16/19 (43 min) · 13/19 but plays well · 15/19, 38 min artifact-complete: the "84-minute deep build" was verify tail; no 27B tier produces rich builds |
| Flash-Next + native MTP, first try | v0.7.3 | 17/19, 16 min, try-1 SUCCESS |
| Flash-Next + pad clause | current | fully clean playtest: 12 min, 30.7 t/s decode, 353 t/s prefill (best measured) |
| Flash-Next low, clean env | v0.7.3 | 5-min build, 40 t/s decode (48.7 peak), 16/19, zero human edits |
| Flash-Next low, 8× thinking budget | v0.7.3 | 14/19, 9 min. Budget raise scored worse, cap never engaged: budget is burn-out protection, not a quality knob |
| Flash-Next medium, N=3 | v0.7.4.1 | 15-18/19. The 18 is the best score measured on this machine; 12-17 min |
| Flash-Next high, N=2 | v0.7.3 + v0.7.4.1 | 16-17/19, 2855-3056 LOC, ~40 min to artifacts, richest builds, best playtests |
| Flash-Next minimal (256 budget), N=2 | v0.7.4.1 | 14-15/19, 7-9 min, builds a working level 1; level-progression froze in playtest: the scaffolding lane, not the game lane |
| Flash-Next xhigh (16384 budget) | v0.7.4.1 | 17/19, 4204 LOC, richest ever, ~35 min; playtest: keyboard dead on first keydown (class #14) |
| gate runtime v2.2.1 | 2026-09-06 | leak-guard: CDP probes reap their detached chromium group on signal exit; scorer untouched, all scores comparable |
Final board (2026-09-06): every effort tier at N≥2 on both models. Static bands fully overlap (13-18). The best score is Flash-Next medium at 18/19.
Fifteen playtest-found failure classes; the last three (gameover-exit, dead keyboard, level-2 freeze) were runtime wiring that passed static score and the soak.
Spec-decode controllers are mechanism-dependent. Acceptance-adaptive draft sizing wins on block drafters and loses on chained MTP drafts, measured both ways on the same harness. 3. Fast and healthy ≠ agent-ready. A stack can benchmark beautifully and still fail a build contract.
That gap is the reason this repo exists. 4. Per-model blind spots are usually spec gaps. Two models each failed the same subsystem on every roll.
One explicit contract sentence cured each, 2-for-2. 5. On explicit contracts, model size buys speed, not quality. A 125B MoE and a 27B, matched effort and environment: one check apart, both playable, the MoE in 7.8× less wall time with 4.3× less reasoning.
Pick by your token budget, not by a quality assumption.
node --check), chromium (headless)pi config) or your numbers carry a variable the recipe does not account forBoth moved to qwen38-strix-halo-harness:
rag-index.py / rag-query.py (NPU semantic search, ~100ms end-to-end), the codebase_search
Pi extension, progress-tracker crash recovery, sidecar compaction, and the one-command setup. Measured against this repo on 2026-10-01: 8/10 top-3 retrieval hits at ~100ms.
playtest-graded benchmark for local LLM stacks: a 10-file game build, 19 static checks, a runtime soak, and a human playing it
See the codeDoes your local LLM setup actually work for coding agents, or does it just benchmark well?
The best build this harness ever measured, 4,204 lines of polished arcade game, 17/19 on the static scorer, had a dead keyboard. The model exported its keymap as a Node module and never wired it into the browser. Syntax checks passed.
The two-minute runtime soak ran at 60fps with 2.5M draw calls. Then a human pressed Enter and nothing happened, because the first keydown threw on an undefined global and took the entire input system with it.
That's failure class #14. There are fourteen others, and every one was found the same way: by a person playing the game. Zero were found by static analysis.
That gap is the reason this repo exists.
Not the model. Your whole inference stack: engine, quant, drafter, speculative decoding, chat template, contract wording, thinking budget. Several of these move the result more than the model choice does, and standard benchmarks can't see any of them.
One cell:
A coding agent gets a fixed contract: build "Neon Overdrive", a 10-file HTML5 canvas breakout game. Exact file list, physics rules, serve behavior, pacing constants, all specified, all static-checkable.
Static scorer (19 checks) grades the code: structure, collision correctness, contract compliance, two different velocity-multiplication bug patterns. 3.
Runtime soak (120s, headless Chromium via raw CDP): boot errors are fatal, frames and draw calls must advance, reloads are detected. 4. Behavior probe (behavior-probe.sh): a scripted real-time pass for serve gating (the ball must not move before Space) and brick reflection (the ball must bounce, not pass through).
Catches collision-wiring classes the soak and the static scorer cannot see. 5. You play it. Two minutes with the arrow keys.
This is the grade that matters. The fifteen failure classes in the ledger all live here.
Wall clock: 5-40 minutes depending on model and effort tier. The result line is one comment's worth of data, generated for you.
The subject is validated on llama.cpp (Vulkan) + AMD Strix Halo, Qwen3.8-27B (four quants) and Qwen3.8-Flash-Next 125B-A6B, with DFlash2 and native-MTP speculative decoding. It will work on any local stack that can run a coding agent; the receipts below are from this rig.
Companion repo with the runnable pi-agent stack (server scripts, wiring, extensions): qwen38-strix-halo-harness.
| benchmark | what it measures |
|---|---|
| WebGen-Bench | model capability: multi-file website generation, browser tests |
| LMGame / BALROG | agents playing games |
| neon-ladder | your local configuration, through a real build workload |
Model-capability benchmarks can't catch a chat template that silently ignores your effort flags, or a drafter whose acceptance collapses on reasoning-heavy turns. Those are configuration failures, and they're what this harness is built to surface.
Behind every number published here:
| layer | what I run |
|---|---|
| Hardware | AMD Strix Halo APU, Ryzen AI MAX+ 395, Radeon 8060S iGPU (gfx1151), 109GB usable unified memory |
| OS / driver | CachyOS Linux, Mesa RADV 26.2.1 system driver, boot args amdgpu.gttsize=126976 ttm.pages_limit=32505856 |
| Engine | llama.cpp Vulkan, Nathan's strix-halo-vulkan releases (validated v0.6.11 through v0.7.4.1; 0.7.4+ = throughput parity + greedy repeatability fixes) + my adaptive draft-sizing port |
| Models | Qwen3.8-27B: Unsloth UD-Q4_K_XL-v3 (the 27B pick), Q5/Q6/Q8 for the tier map · Qwen3.8-Flash-Next: Unsloth UD-IQ4_XS (87.25GB, sha-pinned) |
| Drafters | DFlash2 Q4_K_M sidecar (27B; adaptive n3-7 or fixed n4) · native MTP head Q8_0 (Flash-Next; fixed n4) |
| Vision | mmproj-F16.GGUF from each model repo, passed as -mm ... --mmproj-offload on both servers (pi sessions have vision on GPU) |
| KV cache | 27B: f16 target / q8_0 drafter · Flash-Next: q8_0/q8_0 · contexts: 262144 (27B) / 32768-65536 (Flash-Next) |
| Chat template | Sharp v22.4.0 on the 27B (+19% sustained vs stock; flag order matters: see the run recipe) · embedded template on Flash-Next |
| Agent harness | pi coding agent (0.84.x) + pi-llama-cpp provider, thinkingTokenBudgetField: "thinking_budget_tokens", Qwen card sampling (temp 1.0, top_p 0.95, top_k 20) |
| Gates (this bundle) | static scorer · 120s runtime smoke gate · runner with repair-first retry and server-down abort |
Two engine-driver findings that cost real time to learn (three-way A/B, same source commit): a native build beats the portable payload by 6-24% on decode (JSON-class 28.3 → 40.3 t/s), and the system Mesa stable beat the bundled devel snapshot on agent-relevant classes (deep decode +9.5%, emission +13%). Budget GPU memory against mem_info_vram_total + mem_info_gtt_total; both heaps count; the Vulkan allocator uses both. This nominal-128GB machine with a 16GB BIOS UMA carve exposes 16 + 109.7 ≈ 125.7 GiB total, split between a pinned VRAM heap and an evictable GTT heap.
One adaptive-drafting boundary worth knowing before you tune: the acceptance controller wins on DFlash2 block drafts and loses to fixed n4 on MTP chained drafts, measured both directions on the same harness. Use whichever wins on your stack, not whichever is newer.
| file | role |
|---|---|
contract.txt | the game contract: file structure, physics, MENU/START + SERVE rules, workflow |
prompt-multi-v2_12.txt | the bench prompt (current): contract + workflow + attempt marker slot. v2_12 adds MENU/START RULE, SERVE RULE (ball glued to paddle until Space) and the input-wiring verify clause. Published cells to date (FN band, halogen 0.6.1 own-format, halogen 0.7.0 GGUF) ran v2_9 - same contract minus those clauses - so cross-engine cells stay comparable; new cells should use this file. |
game-score.py | static scorer, 19 checks, versioned + md5-frozen |
smoke-gate.sh | 120s runtime soak: boot errors fatal, frame/draw-ops advancement, reload detection |
wsmin.js | raw-WebSocket CDP client the gate polls through |
cdp.js | single-eval CDP client for diagnostics (the gate does not need it) |
run.sh | runner: success = files + SMOKE-OK; boot errors trigger a repair pass, else wipe and reroll |
soak-probe.sh | standalone diagnostic soak for post-mortems |
report.sh | assembles the one-comment result line from run artifacts (auto-invoked by run.sh on success) |
Use the bundled scorer with the bundled contract. They are calibrated as a set. Scores from modified contracts or different scorers are not comparable.
Three dependencies, then clone:
# 1. System deps
sudo pacman -S nodejs chromium # or your distro's equivalents
# 2. The pi coding agent (or build from source: github.com/earendil-works/pi)
yay -S pi-coding-agent-bin
# 3. This bundle
git clone https://github.com/aic0d3r/neon-ladder && cd neon-ladder
The agent harness is a separate repo: qwen38-strix-halo-harness holds the pi extensions, setup.sh (one command: providers, sidecar, halogen launch), start.sh (daily use), crash recovery, and the NPU retrieval toolkit. If you want the full daily-driver stack, install that too and run its ./setup.sh. This repo is the benchmark: game builds, prompts, scorers, results.
For the ladder you only need pi wired to your server. Minimal ~/.pi/agent/models.json entry (the llama.cpp path):
{
"providers": {
"llamacpp": {
"baseUrl": "http://127.0.0.1:8080/v1",
"api": "openai-completions",
"apiKey": "dummy",
"models": [{
"id": "my-model",
"reasoning": true,
"contextWindow": 262144,
"maxTokens": 32768,
"compat": {
"thinkingFormat": "chat-template",
"chatTemplateKwargs": { "reasoning_effort": {"$var": "thinking.effort"}, "enable_thinking": {"$var": "thinking.enabled"} },
"thinkingTokenBudgetField": "thinking_budget_tokens"
}
}]
}
}
}
Start your server, run the cell, play the game. Reference server command:
llama-server -m Qwen3.8-27B-UD-Q4_K_XL.gguf \
-md Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
-mm mmproj-F16.gguf --mmproj-offload \
--spec-type draft-dflash --spec-draft-n-max 4 \
-ngl all -fa on -ctk f16 -ctv f16 -ctkd q8_0 -ctvd q8_0 \
-c 262144 -b 4096 -ub 4096 --jinja \
--chat-template-file sharp-v22.4.0.jinja \
--host 127.0.0.1 --port 8080 --metrics
# KEY: --jinja must come BEFORE --chat-template-file, and never use
# --chat-template with a file path: it fails silently.
The cell itself:
export LLADDER=$HOME/neon-ladder-work && mkdir -p $LLADDER
bash run.sh build-run1 game-run1 medium "$(cat contract.txt)"
# ... typically 15-45 min later, on success the runner prints your result line
Open build-run1/index.html, press Enter, play two minutes. Then:
QUANT="Q4_K_XL" DRAFTER="DFlash2-Q4_M n4" PLAYTEST="Y - plays well" \
RIG="Strix Halo 395+8060S, Nathan v0.7.4.1" bash report.sh $LLADDER/build-run1 game-run1
That prints (and saves) a line like:
Q8_K_XL / DFlash2-Q4_M n4 / medium / 23min / static 17/19 / SMOKE-OK / Y - clean and polished / Strix Halo 395+8060S, 109GB, Nathan v0.7.4.1
Static score, soak verdict, effort, and wall clock come from the run artifacts automatically. You supply quant, drafter, playtest verdict, and rig, the four things only you know.
Divergent results are data, not contradiction. If you know the rules:
By available GPU memory (check mem_info_vram_total + mem_info_gtt_total, both heaps count):
| budget | stack | what you get |
|---|---|---|
| ~24 GB | Qwen3.8-27B UD-Q4_K_XL + DFlash2 Q4_K_M, fixed n4, ctx ≤64k | the reliable-build config, every 27B receipt above |
| 24–48 GB | same, full 256k ctx | plus IDE/browser headroom, the reference rig |
| 48–91 GB | add the Q8-27B profile for deep-context reading | fastest prefill tier (253 avg / 330 peak t/s) |
| ≥91 GB | Qwen3.8-Flash-Next UD-IQ4_XS + MTP Q8_0, Nathan v0.7.3+ built natively (v0.7.4.1 current), --reasoning-effort medium --reasoning-budget 2048 (server default: clients that send a per-request budget, e.g. pi, override it), this contract | the speed lane and the daily driver on 128GB: 12-min clean medium builds, 40 t/s sustained decode on low-effort sessions, 353 t/s prefill |
By task:
--spec-draft-adaptive n3-7): +25% on tool-call classes. On MTP stacks, stick to fixed n4 (adaptive loses there).Three knobs that matter more than they look:
--chat-template: it fails silently (use --chat-template-file, --jinja first).Every row is a receipt from this harness. Playtest verdicts are the user's, recorded as given.
| stack | contract era | verdict |
|---|---|---|
| Qwen3.8-27B (Q4/Q5/Q6/Q8) + DFlash2 | early revisions | 6/6 playable, PPL plateau 7.079–7.089 |
| 27B + adaptive draft sizing (N=5 cells) | current | 4/5 clean playtests, decode 18.0 vs 17.7 fixed-n4, acceptance 65.3% vs 60.4% |
| 27B low, clean-env control | current | 17/19 (1412 LOC), 39 min, 4.3× the reasoning of the MoE below, quality a wash, 8× the wall time |
| 27B low / medium / high, N=2 | v0.7.4.1 | 16/19 (43 min) · 13/19 but plays well · 15/19, 38 min artifact-complete: the "84-minute deep build" was verify tail; no 27B tier produces rich builds |
| Flash-Next + native MTP, first try | v0.7.3 | 17/19, 16 min, try-1 SUCCESS |
| Flash-Next + pad clause | current | fully clean playtest: 12 min, 30.7 t/s decode, 353 t/s prefill (best measured) |
| Flash-Next low, clean env | v0.7.3 | 5-min build, 40 t/s decode (48.7 peak), 16/19, zero human edits |
| Flash-Next low, 8× thinking budget | v0.7.3 | 14/19, 9 min. Budget raise scored worse, cap never engaged: budget is burn-out protection, not a quality knob |
| Flash-Next medium, N=3 | v0.7.4.1 | 15-18/19. The 18 is the best score measured on this machine; 12-17 min |
| Flash-Next high, N=2 | v0.7.3 + v0.7.4.1 | 16-17/19, 2855-3056 LOC, ~40 min to artifacts, richest builds, best playtests |
| Flash-Next minimal (256 budget), N=2 | v0.7.4.1 | 14-15/19, 7-9 min, builds a working level 1; level-progression froze in playtest: the scaffolding lane, not the game lane |
| Flash-Next xhigh (16384 budget) | v0.7.4.1 | 17/19, 4204 LOC, richest ever, ~35 min; playtest: keyboard dead on first keydown (class #14) |
| gate runtime v2.2.1 | 2026-09-06 | leak-guard: CDP probes reap their detached chromium group on signal exit; scorer untouched, all scores comparable |
Final board (2026-09-06): every effort tier at N≥2 on both models. Static bands fully overlap (13-18). The best score is Flash-Next medium at 18/19.
Fifteen playtest-found failure classes; the last three (gameover-exit, dead keyboard, level-2 freeze) were runtime wiring that passed static score and the soak.
Spec-decode controllers are mechanism-dependent. Acceptance-adaptive draft sizing wins on block drafters and loses on chained MTP drafts, measured both ways on the same harness. 3. Fast and healthy ≠ agent-ready. A stack can benchmark beautifully and still fail a build contract.
That gap is the reason this repo exists. 4. Per-model blind spots are usually spec gaps. Two models each failed the same subsystem on every roll.
One explicit contract sentence cured each, 2-for-2. 5. On explicit contracts, model size buys speed, not quality. A 125B MoE and a 27B, matched effort and environment: one check apart, both playable, the MoE in 7.8× less wall time with 4.3× less reasoning.
Pick by your token budget, not by a quality assumption.
node --check), chromium (headless)pi config) or your numbers carry a variable the recipe does not account forBoth moved to qwen38-strix-halo-harness:
rag-index.py / rag-query.py (NPU semantic search, ~100ms end-to-end), the codebase_search
Pi extension, progress-tracker crash recovery, sidecar compaction, and the one-command setup. Measured against this repo on 2026-10-01: 8/10 top-3 retrieval hits at ~100ms.