Open benchmark runtime for document-grounded decision models
Python
1
50 commits
updated Sep 27, 2026
pip install git+https://github.com/Hanno-Labs/decision-bench.git
uv add git+https://github.com/Hanno-Labs/decision-bench.git
Smoke-test a supported decision model on the pinned benchmark. This example uses
Bosun v3.1 0.6B, whose
native decision-token readout is supported directly by run-hf.
hf download Hanno-Labs/bosun-v3.1-0.6b \
--revision aaa9dd06d4d6501b33df61942472fed9284bc5e6 \
--local-dir models/bosun-v3.1-0.6b
decision-bench run-hf task_specs/decisionbench-dev.toml \
models/bosun-v3.1-0.6b results/bosun-v3.1-0.6b --smoke
Before running anything else, see the complete supported adapters and models. If your model is listed, use its runner; only add an adapter when its native decision readout is not already supported.
| 📈 Leaderboard | Compare reviewed results and filter by task, family, domain, or primitive |
| 🏃 Get Started | Install DecisionBench and run the frozen suite |
| 📋 Tasks and Views | Understand the 23,900 rows, nine families, three primitives, and reasoning track |
| 🤖 Models | See supported adapters and models, or add a new native readout contract |
| 📊 Results | Load, inspect, and submit reproducible results |
| 🧪 Evaluation | Learn the metrics, artifacts, and comparability rules |
| 🤝 Contributing | Add models, tasks, benchmarks, and result records |
Report bugs and request features for any DecisionBench component in the central issue tracker. Send code changes to the repository that owns that component.
Choose the path that matches what you want to bring to DecisionBench:
🤖 Add a Model →Add a compatibility adapter so DecisionBench can evaluate a new decision model. |
📊 Submit Results →Run a supported model and submit its reviewed, reproducible scores. |
🧩 Add a Task →Contribute one dataset-backed decision problem with labels, provenance, and tests. |
🗂️ Add a Benchmark →Curate existing tasks into a named evaluation for a domain or purpose. |
DecisionBench is under active development. Until the benchmark paper is published, cite the
repository and the individual datasets listed in the
task catalog. Machine-readable
citation metadata lives in CITATION.cff.
Python
96.0%
Shell
3.9%
Open benchmark runtime for document-grounded decision models
Python
1
50 commits
updated Sep 27, 2026
pip install git+https://github.com/Hanno-Labs/decision-bench.git
uv add git+https://github.com/Hanno-Labs/decision-bench.git
Smoke-test a supported decision model on the pinned benchmark. This example uses
Bosun v3.1 0.6B, whose
native decision-token readout is supported directly by run-hf.
hf download Hanno-Labs/bosun-v3.1-0.6b \
--revision aaa9dd06d4d6501b33df61942472fed9284bc5e6 \
--local-dir models/bosun-v3.1-0.6b
decision-bench run-hf task_specs/decisionbench-dev.toml \
models/bosun-v3.1-0.6b results/bosun-v3.1-0.6b --smoke
Before running anything else, see the complete supported adapters and models. If your model is listed, use its runner; only add an adapter when its native decision readout is not already supported.
| 📈 Leaderboard | Compare reviewed results and filter by task, family, domain, or primitive |
| 🏃 Get Started | Install DecisionBench and run the frozen suite |
| 📋 Tasks and Views | Understand the 23,900 rows, nine families, three primitives, and reasoning track |
| 🤖 Models | See supported adapters and models, or add a new native readout contract |
| 📊 Results | Load, inspect, and submit reproducible results |
| 🧪 Evaluation | Learn the metrics, artifacts, and comparability rules |
| 🤝 Contributing | Add models, tasks, benchmarks, and result records |
Report bugs and request features for any DecisionBench component in the central issue tracker. Send code changes to the repository that owns that component.
Choose the path that matches what you want to bring to DecisionBench:
🤖 Add a Model →Add a compatibility adapter so DecisionBench can evaluate a new decision model. |
📊 Submit Results →Run a supported model and submit its reviewed, reproducible scores. |
🧩 Add a Task →Contribute one dataset-backed decision problem with labels, provenance, and tests. |
🗂️ Add a Benchmark →Curate existing tasks into a named evaluation for a domain or purpose. |
DecisionBench is under active development. Until the benchmark paper is published, cite the
repository and the individual datasets listed in the
task catalog. Machine-readable
citation metadata lives in CITATION.cff.
Python
96.0%
Shell
3.9%