Hanno-Labs/decision-bench

Open benchmark runtime for document-grounded decision models

Python

1

50 commits

updated Sep 27, 2026

See the code

README

DecisionBench


The evaluation ecosystem for decision models

License CI

Installation · Documentation · Leaderboard · Results · Issues · Citing

Installation

pip install git+https://github.com/Hanno-Labs/decision-bench.git
uv add git+https://github.com/Hanno-Labs/decision-bench.git

Example Usage

Smoke-test a supported decision model on the pinned benchmark. This example uses Bosun v3.1 0.6B, whose native decision-token readout is supported directly by run-hf.

hf download Hanno-Labs/bosun-v3.1-0.6b \
  --revision aaa9dd06d4d6501b33df61942472fed9284bc5e6 \
  --local-dir models/bosun-v3.1-0.6b

decision-bench run-hf task_specs/decisionbench-dev.toml \
  models/bosun-v3.1-0.6b results/bosun-v3.1-0.6b --smoke

Before running anything else, see the complete supported adapters and models. If your model is listed, use its runner; only add an adapter when its native decision readout is not already supported.

Overview

📈 LeaderboardCompare reviewed results and filter by task, family, domain, or primitive
🏃 Get StartedInstall DecisionBench and run the frozen suite
📋 Tasks and ViewsUnderstand the 23,900 rows, nine families, three primitives, and reasoning track
🤖 ModelsSee supported adapters and models, or add a new native readout contract
📊 ResultsLoad, inspect, and submit reproducible results
🧪 EvaluationLearn the metrics, artifacts, and comparability rules
🤝 ContributingAdd models, tasks, benchmarks, and result records

Contribute

Report bugs and request features for any DecisionBench component in the central issue tracker. Send code changes to the repository that owns that component.

Choose the path that matches what you want to bring to DecisionBench:

🤖 Add a Model →

Add a compatibility adapter so DecisionBench can evaluate a new decision model.

📊 Submit Results →

Run a supported model and submit its reviewed, reproducible scores.

🧩 Add a Task →

Contribute one dataset-backed decision problem with labels, provenance, and tests.

🗂️ Add a Benchmark →

Curate existing tasks into a named evaluation for a domain or purpose.

Citing

DecisionBench is under active development. Until the benchmark paper is published, cite the repository and the individual datasets listed in the task catalog. Machine-readable citation metadata lives in CITATION.cff.

benchmark
decision-model
jev
llm-evaluation
model-evaluation
typesafe-ai

Hanno-Labs/decision-bench

Open benchmark runtime for document-grounded decision models

Python

1

50 commits

updated Sep 27, 2026

See the code

README

DecisionBench


The evaluation ecosystem for decision models

License CI

Installation · Documentation · Leaderboard · Results · Issues · Citing

Installation

pip install git+https://github.com/Hanno-Labs/decision-bench.git
uv add git+https://github.com/Hanno-Labs/decision-bench.git

Example Usage

Smoke-test a supported decision model on the pinned benchmark. This example uses Bosun v3.1 0.6B, whose native decision-token readout is supported directly by run-hf.

hf download Hanno-Labs/bosun-v3.1-0.6b \
  --revision aaa9dd06d4d6501b33df61942472fed9284bc5e6 \
  --local-dir models/bosun-v3.1-0.6b

decision-bench run-hf task_specs/decisionbench-dev.toml \
  models/bosun-v3.1-0.6b results/bosun-v3.1-0.6b --smoke

Before running anything else, see the complete supported adapters and models. If your model is listed, use its runner; only add an adapter when its native decision readout is not already supported.

Overview

📈 LeaderboardCompare reviewed results and filter by task, family, domain, or primitive
🏃 Get StartedInstall DecisionBench and run the frozen suite
📋 Tasks and ViewsUnderstand the 23,900 rows, nine families, three primitives, and reasoning track
🤖 ModelsSee supported adapters and models, or add a new native readout contract
📊 ResultsLoad, inspect, and submit reproducible results
🧪 EvaluationLearn the metrics, artifacts, and comparability rules
🤝 ContributingAdd models, tasks, benchmarks, and result records

Contribute

Report bugs and request features for any DecisionBench component in the central issue tracker. Send code changes to the repository that owns that component.

Choose the path that matches what you want to bring to DecisionBench:

🤖 Add a Model →

Add a compatibility adapter so DecisionBench can evaluate a new decision model.

📊 Submit Results →

Run a supported model and submit its reviewed, reproducible scores.

🧩 Add a Task →

Contribute one dataset-backed decision problem with labels, provenance, and tests.

🗂️ Add a Benchmark →

Curate existing tasks into a named evaluation for a domain or purpose.

Citing

DecisionBench is under active development. Until the benchmark paper is published, cite the repository and the individual datasets listed in the task catalog. Machine-readable citation metadata lives in CITATION.cff.

benchmark
decision-model
jev
llm-evaluation
model-evaluation
typesafe-ai

Languages

Python

96.0%

Shell

3.9%