slee-persis/GVS5H

GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard | Fable 5 Level Coding for a Fifth the Price - or on a Single GPU

271

stars

16

commits

Python

primary language

Sep 11, 2026

updated

README

GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard

Results

Manager vs single call, four models — LCB-100, 5 passes, 128k max tokens, reasoning ON

What one pass costs — LCB-100, 5 passes, single call vs manager, against Fable 5

Abstract. Frontier coding performance is typically bought with larger proprietary models at high cost. We introduce ledger-based zero-shot self-orchestration, a training-free method in which fresh instances of one model decompose problems and coordinate through a shared filesystem holding a plan, notes and current solution. Across nine open and closed-weight models on the 100 latest hard LiveCodeBench problems, the method yields gains of up to 23.2 percentage points on pinned backends and offers two routes to frontier-level accuracy. Orchestrated GPT-5.6-Terra reaches 88.0% pass@1 against Fable 5's 90.4% at 19% of the cost, and locally served, open-weight Qwen3.8-27B rises from 69.2% to 92.4%, slightly exceeding Fable 5. Gains are not universal: some models are unchanged or worse. Transcript analysis attributes the gain to decomposition and persistent context. Inference-time organization can approach frontier coding accuracy at a fraction of the cost, or slightly exceed it on self-hostable weights.

the paper

Running the code

Needs uv and an API key for the model you want to test.

cd codebase/v2-current
export OPENAI_API_KEY=...

LCB_RELEASE=release_v6 \
ESCALATION_CLOUD_MAX_TOKENS=128000 \
ESCALATION_CLOUD_TIMEOUT=7200 \
MULTIAGENT_MODEL=openai:gpt-5.6-terra \
uv run --no-project --python 3.12 --with 'datasets<4' --with numpy --with anthropic \
  python escalation/run_bench.py --engine multiagent --only lcb --lcb 100 --parallel 8
  • --engine multiagent runs the manager; --engine single is the one-call baseline.
  • Other models: anthropic:<model>, dashscope:<model>, openrouter:<model>, each with its own *_API_KEY.
  • The pass@1 score prints at the end. Results are written to runs/results.json, workspaces to runs/ws/.

License

Code is under the MIT License. The paper, figures and run data are under CC BY 4.0. The LiveCodeBench fork, the benchmark problem statements and the LaTeX template files keep their own licenses. See NOTICE.md for which license covers which path.

Contributors

slee-persis

9 commits

persis-hsun

7 commits

slee-persis/GVS5H

GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard | Fable 5 Level Coding for a Fifth the Price - or on a Single GPU

271

stars

16

commits

Python

primary language

Sep 11, 2026

updated

README

GVS5H: Five Qwen3.8-27B Models Match Claude Fable 5 on LiveCodeBench Hard

Results

Manager vs single call, four models — LCB-100, 5 passes, 128k max tokens, reasoning ON

What one pass costs — LCB-100, 5 passes, single call vs manager, against Fable 5

Abstract. Frontier coding performance is typically bought with larger proprietary models at high cost. We introduce ledger-based zero-shot self-orchestration, a training-free method in which fresh instances of one model decompose problems and coordinate through a shared filesystem holding a plan, notes and current solution. Across nine open and closed-weight models on the 100 latest hard LiveCodeBench problems, the method yields gains of up to 23.2 percentage points on pinned backends and offers two routes to frontier-level accuracy. Orchestrated GPT-5.6-Terra reaches 88.0% pass@1 against Fable 5's 90.4% at 19% of the cost, and locally served, open-weight Qwen3.8-27B rises from 69.2% to 92.4%, slightly exceeding Fable 5. Gains are not universal: some models are unchanged or worse. Transcript analysis attributes the gain to decomposition and persistent context. Inference-time organization can approach frontier coding accuracy at a fraction of the cost, or slightly exceed it on self-hostable weights.

the paper

Running the code

Needs uv and an API key for the model you want to test.

cd codebase/v2-current
export OPENAI_API_KEY=...

LCB_RELEASE=release_v6 \
ESCALATION_CLOUD_MAX_TOKENS=128000 \
ESCALATION_CLOUD_TIMEOUT=7200 \
MULTIAGENT_MODEL=openai:gpt-5.6-terra \
uv run --no-project --python 3.12 --with 'datasets<4' --with numpy --with anthropic \
  python escalation/run_bench.py --engine multiagent --only lcb --lcb 100 --parallel 8
  • --engine multiagent runs the manager; --engine single is the one-call baseline.
  • Other models: anthropic:<model>, dashscope:<model>, openrouter:<model>, each with its own *_API_KEY.
  • The pass@1 score prints at the end. Results are written to runs/results.json, workspaces to runs/ws/.

License

Code is under the MIT License. The paper, figures and run data are under CC BY 4.0. The LiveCodeBench fork, the benchmark problem statements and the LaTeX template files keep their own licenses. See NOTICE.md for which license covers which path.

Contributors

slee-persis

9 commits

persis-hsun

7 commits

Languages

Python

98.6%

TeX

1.1%