The first System One benchmark scored against human uncertainty, not just a single right answer.
Does your decision model hesitate where humans disagree?
System One models like Jev return a probability with every decision, so you can set a threshold: automate the confident cases and send the uncertain ones to a human. But other benchmarks score these models against a single "right" answer, so nobody checks whether the probabilities actually mean anything.
NVIDIA's HelpSteer2 (CC BY 4.0) makes that check possible. Every AI response in it was rated by several human annotators, and their individual ratings are public. If three people split 2 to 1 on a response, a good model should be uncertain about it too.
DoubtBench turns that data into about 7,500 Score, Choice and Noul questions. It scores models on two things: getting the answer right, and whether their probabilities line up with human disagreement.
Validation split: 7,455 questions. Two reference rows frame every model. Human (ceiling) replays the annotators' own votes. Uniform (floor) spreads its probability evenly and knows nothing.
| Rank | Model | Open | Params | DoubtBench | Accuracy | ECE / floor | JS div. ↓ | Disagreement corr. | $ / 1k q | p50 latency | Date | Version |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| — | human (ceiling) | — | — | 90.4 | 82.6% | 13.76× | 0.000 | 1.00 | — | — | 2026-10-03 | 1.0 |
| 1 | Claude Haiku 4.5 (verbalized) | closed | undisclosed | 69.2 | 45.8% | 7.27× | 0.292 | 0.15 | $0.6689 | 2166 ms | 2026-10-03 | 1.0 |
| 2 | Jev 1.13 | closed | undisclosed | 67.8 | 48.2% | 12.80× | 0.284 | 0.14 | $0.0134 | 98 ms | 2026-10-03 | 1.0 |
| 3 | Claude Opus 5.5 (verbalized) | closed | undisclosed | 67.8 | 44.0% | 7.97× | 0.306 | 0.06 | $3.1368 | 3735 ms | 2026-10-04 | 1.0 |
| 4 | Claude Sonnet 5.5 (verbalized) | closed | undisclosed | 67.4 | 42.5% | 8.16× | 0.301 | 0.12 | $1.5522 | 2420 ms | 2026-10-04 | 1.0 |
| — | uniform (floor) | — | 0 | 54.8 | 18.4% | 21.39× | 0.396 | — | — | — | 2026-10-03 | 1.0 |
Four models so far: Jev 1.13, a purpose-built System One model, and three Claude models asked to state a probability for every option.
| DoubtBench score | Accuracy | Disagreement corr. | Cost of a full run | Median latency | |
|---|---|---|---|---|---|
| Human annotators (ceiling) | 90.4 | 82.6% | 1.00 | ||
| Claude Haiku 4.5 | 69.2 | 45.8% | 0.15 | $5.00 | 2.2 s |
| Jev 1.13 | 67.8 | 48.2% | 0.14 | $0.10 | 0.1 s |
| Claude Opus 5.5 | 67.8 | 44.0% | 0.06 | $23.30 | 3.7 s |
| Claude Sonnet 5.5 | 67.4 | 42.5% | 0.12 | $11.50 | 2.4 s |
| Uniform guessing (floor) | 54.8 | 18.4% | — |
No model hesitates where humans disagree. The disagreement correlation measures how much a model's uncertainty rises on responses where the annotators split. The ceiling is 1.0. All four models land between 0.06 and 0.15. They are barely more uncertain on the responses humans argued about than on the ones everyone agreed on. If you plan to send uncertain cases to a human, that matters.
Jev is the most accurate, by far the cheapest and the fastest. It gets more answers right than any Claude model, and the gap holds up statistically (paired McNemar test, p < 0.001 against each). A full run costs about 10 cents, against $5 to $23 for Claude, at 98 ms median latency.
The Claude models score well on calibration because they're less confident. Their average confidence is about 0.54, against Jev's 0.65. Hedging brings stated confidence closer to actual accuracy, which is what keeps them level with Jev on the composite score.
Bigger isn't better here. Haiku 4.5 beats both Sonnet 5.5 and Opus 5.5 on accuracy. The larger models struggle most with the yes/no "is this response acceptable?" question, at 55% against Haiku's 72% and Jev's 75%. Opus is the least swayed by response order, though: it keeps the same answer 90% of the time when the two responses swap places, against 80% for Jev and 64% for Haiku.
How these were run: Claude models answer through structured JSON output, with the questions about one
response batched into a single call, as with Jev. Sonnet 5.5 and Opus 5.5 ran at low effort, and
higher effort may do better. They also refused 30 and 36 questions (jailbreak-style prompts, flagged
cyber), which count as wrong. Haiku and Jev refused none.
Laya, a SemIf/Kev model and an open-weight instruct model are next. Submit yours.
The human row isn't a model. It replays what the annotators actually voted, so read it as the realistic best case.
Its accuracy is 82.6%, not 100%, because HelpSteer2's official label is a rounded average of the votes. Sometimes nobody picked that value: one vote for 2 and one for 4 give a label of 3. That happens for 8% of attribute scores and about a quarter of the preference items. For the same reason, its ECE and temperature look poor even though it matches the humans perfectly.
So compare models to the human row, not to 100. The cleanest comparison is accuracy on questions where every annotator agreed. There the human row scores 100%, Jev 1.13 scores 62%, and the Claude models score 53% to 55%.
pip install doubtbench # or: pip install -e ".[dev]" from a clone
doubtbench build # download HelpSteer2 @ pinned revision -> data/doubtbench.jsonl
TYPESAFE_API_KEY=... doubtbench run --adapter jev --model jev-1.13.0
A full Jev run is about 1,900 calls and ~2M input tokens, so roughly $0.09 at $0.042 per million (an
estimate from request size). All six questions about one response go in a single call. No TypeSafe key? Use --adapter jev_vercel with AI_GATEWAY_API_KEY.
Claude models (any of them) with an Anthropic API key:
pip install "doubtbench[claude]"
ANTHROPIC_API_KEY=... doubtbench run --adapter claude --model claude-haiku-4-5
Open models through any OpenAI-compatible server (vLLM, llama.cpp, MLX, SGLang) via logprob readout:
vllm serve Qwen/Qwen3.5-4B-Instruct &
doubtbench run --adapter logprob_local --base-url http://localhost:8000/v1 --model Qwen/Qwen3.5-4B-Instruct
Other commands: doubtbench score results/<run_id>, doubtbench compare results/a results/b,
doubtbench leaderboard, doubtbench validate. Add --limit 200 for a quick smoke run.
Take item hs2-val-000088-helpfulness. The user wrote "Could you" and the assistant replied "Sure,
I can help with that. What would you like to know?" Three annotators rated helpfulness: two gave
3 (mostly helpful) and one gave 4 (fully helpful).
| P(2) | P(3) | P(4) | Correct? | JS divergence to humans | |
|---|---|---|---|---|---|
| Humans (2 of 3 votes, 1 of 3 votes) | 0.67 | 0.33 | 0 | ||
| Model A, overconfident | 0.98 | 0.02 | ✓ | 0.143 | |
| Model B, hesitates like humans | 0.05 | 0.62 | 0.33 | ✓ | 0.026 |
Illustrative probabilities; the human votes are real. Both models pick 3, so a single-answer benchmark scores them the same. But model A claims 98% certainty on a question the humans split on, and model B's 62% is about right. DoubtBench tells them apart.
| Task | Type | Questions (val) | What it asks |
|---|---|---|---|
hs2_attribute_score | Score 0–4 | 5,110 | Helpfulness, correctness, coherence, complexity, verbosity of one response |
hs2_is_acceptable | Noul | 1,022 | Is helpfulness ≥ 3? |
hs2_pairwise_pref | Choice | 882 | Response 1, response 2, or tie. Each pair appears twice, swapped, to measure position bias |
hs2_pref_strength | Score −3…+3 | 441 | How strongly one response is preferred |
Attribute questions (verbatim):
| Attribute | Question |
|---|---|
| helpfulness | How helpful is the response to the prompt overall? |
| correctness | Does the response include all pertinent facts without errors? |
| coherence | How consistent and clear is the response? |
| complexity | How much expertise would it take to write this response? |
| verbosity | How much detail does the response include relative to what was asked? |
Level descriptions are in doubtbench/constants.py. One JSON object per line,
compatible with existing System One harnesses (state, type, question, options/levels,
label), plus soft_label (annotator vote shares) and source (CC BY attribution).
Accuracy (plus within-1 for scores), score MAE, NLL, Brier, ECE with a resampled noise floor, a refit temperature (T > 1 overconfident), soft-label KL / Jensen-Shannon, disagreement correlation (Spearman between model entropy and human entropy), accuracy on unanimous vs split items, selective accuracy at 0.7/0.8/0.9 confidence, position consistency, p50/p95 latency and cost per 1k questions. Every metric is broken down by task, attribute, question type and single vs multi-turn.
DoubtBench score = mean of Accuracy, Calibration (1 − ECE) and Human agreement (1 − JS), each 0–100. The components are always shown next to it. Details: docs/METHODOLOGY.md.
logprob_local / systemone_compat); see doubtbench/adapters/.doubtbench run --adapter ... --model ... on the default (validation) data. No --limit.submissions/<model>.yaml and your results/<run_id>/ folder (run.json, raw.jsonl).
CI re-scores from your raw answers and comments the numbers.Full guide: docs/SUBMITTING.md. Don't have the hardware? Open an issue and we'll run it for you.
nvidia/HelpSteer2@990b2711, with SHA-256 of every input and output in data/MANIFEST.json. doubtbench build is deterministic.data/dropped_harmful.json.data/overlap_report.json.--full adds HelpSteer2 train. Those results are reported separately and flagged as possibly contaminated, because HelpSteer2 train is widely used to train reward models.Code: MIT (LICENSE). Data: derived from HelpSteer2 under CC BY 4.0. See DATA_LICENSE.md for attribution and the list of modifications.
If you use DoubtBench, please cite it (CITATION.cff) and the HelpSteer2 papers:
@misc{wang2024helpsteer2,
title={HelpSteer2: Open-source dataset for training top-performing reward models},
author={Zhilin Wang and Yi Dong and Olivier Delalleau and Jiaqi Zeng and Gerald Shen and Daniel Egert and Jimmy J. Zhang and Makesh Narsimhan Sreedhar and Oleksii Kuchaiev},
year={2024}, eprint={2406.08673}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
@misc{wang2024helpsteer2preference,
title={HelpSteer2-Preference: Complementing Ratings with Preferences},
author={Zhilin Wang and Alexander Bukharin and Olivier Delalleau and Daniel Egert and Gerald Shen and Jiaqi Zeng and Oleksii Kuchaiev and Yi Dong},
year={2024}, eprint={2410.01257}, archivePrefix={arXiv}, primaryClass={cs.LG}
}
Python
100.0%
The first System One benchmark scored against human uncertainty, not just a single right answer.
Does your decision model hesitate where humans disagree?
System One models like Jev return a probability with every decision, so you can set a threshold: automate the confident cases and send the uncertain ones to a human. But other benchmarks score these models against a single "right" answer, so nobody checks whether the probabilities actually mean anything.
NVIDIA's HelpSteer2 (CC BY 4.0) makes that check possible. Every AI response in it was rated by several human annotators, and their individual ratings are public. If three people split 2 to 1 on a response, a good model should be uncertain about it too.
DoubtBench turns that data into about 7,500 Score, Choice and Noul questions. It scores models on two things: getting the answer right, and whether their probabilities line up with human disagreement.
Validation split: 7,455 questions. Two reference rows frame every model. Human (ceiling) replays the annotators' own votes. Uniform (floor) spreads its probability evenly and knows nothing.
| Rank | Model | Open | Params | DoubtBench | Accuracy | ECE / floor | JS div. ↓ | Disagreement corr. | $ / 1k q | p50 latency | Date | Version |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| — | human (ceiling) | — | — | 90.4 | 82.6% | 13.76× | 0.000 | 1.00 | — | — | 2026-10-03 | 1.0 |
| 1 | Claude Haiku 4.5 (verbalized) | closed | undisclosed | 69.2 | 45.8% | 7.27× | 0.292 | 0.15 | $0.6689 | 2166 ms | 2026-10-03 | 1.0 |
| 2 | Jev 1.13 | closed | undisclosed | 67.8 | 48.2% | 12.80× | 0.284 | 0.14 | $0.0134 | 98 ms | 2026-10-03 | 1.0 |
| 3 | Claude Opus 5.5 (verbalized) | closed | undisclosed | 67.8 | 44.0% | 7.97× | 0.306 | 0.06 | $3.1368 | 3735 ms | 2026-10-04 | 1.0 |
| 4 | Claude Sonnet 5.5 (verbalized) | closed | undisclosed | 67.4 | 42.5% | 8.16× | 0.301 | 0.12 | $1.5522 | 2420 ms | 2026-10-04 | 1.0 |
| — | uniform (floor) | — | 0 | 54.8 | 18.4% | 21.39× | 0.396 | — | — | — | 2026-10-03 | 1.0 |
Four models so far: Jev 1.13, a purpose-built System One model, and three Claude models asked to state a probability for every option.
| DoubtBench score | Accuracy | Disagreement corr. | Cost of a full run | Median latency | |
|---|---|---|---|---|---|
| Human annotators (ceiling) | 90.4 | 82.6% | 1.00 | ||
| Claude Haiku 4.5 | 69.2 | 45.8% | 0.15 | $5.00 | 2.2 s |
| Jev 1.13 | 67.8 | 48.2% | 0.14 | $0.10 | 0.1 s |
| Claude Opus 5.5 | 67.8 | 44.0% | 0.06 | $23.30 | 3.7 s |
| Claude Sonnet 5.5 | 67.4 | 42.5% | 0.12 | $11.50 | 2.4 s |
| Uniform guessing (floor) | 54.8 | 18.4% | — |
No model hesitates where humans disagree. The disagreement correlation measures how much a model's uncertainty rises on responses where the annotators split. The ceiling is 1.0. All four models land between 0.06 and 0.15. They are barely more uncertain on the responses humans argued about than on the ones everyone agreed on. If you plan to send uncertain cases to a human, that matters.
Jev is the most accurate, by far the cheapest and the fastest. It gets more answers right than any Claude model, and the gap holds up statistically (paired McNemar test, p < 0.001 against each). A full run costs about 10 cents, against $5 to $23 for Claude, at 98 ms median latency.
The Claude models score well on calibration because they're less confident. Their average confidence is about 0.54, against Jev's 0.65. Hedging brings stated confidence closer to actual accuracy, which is what keeps them level with Jev on the composite score.
Bigger isn't better here. Haiku 4.5 beats both Sonnet 5.5 and Opus 5.5 on accuracy. The larger models struggle most with the yes/no "is this response acceptable?" question, at 55% against Haiku's 72% and Jev's 75%. Opus is the least swayed by response order, though: it keeps the same answer 90% of the time when the two responses swap places, against 80% for Jev and 64% for Haiku.
How these were run: Claude models answer through structured JSON output, with the questions about one
response batched into a single call, as with Jev. Sonnet 5.5 and Opus 5.5 ran at low effort, and
higher effort may do better. They also refused 30 and 36 questions (jailbreak-style prompts, flagged
cyber), which count as wrong. Haiku and Jev refused none.
Laya, a SemIf/Kev model and an open-weight instruct model are next. Submit yours.
The human row isn't a model. It replays what the annotators actually voted, so read it as the realistic best case.
Its accuracy is 82.6%, not 100%, because HelpSteer2's official label is a rounded average of the votes. Sometimes nobody picked that value: one vote for 2 and one for 4 give a label of 3. That happens for 8% of attribute scores and about a quarter of the preference items. For the same reason, its ECE and temperature look poor even though it matches the humans perfectly.
So compare models to the human row, not to 100. The cleanest comparison is accuracy on questions where every annotator agreed. There the human row scores 100%, Jev 1.13 scores 62%, and the Claude models score 53% to 55%.
pip install doubtbench # or: pip install -e ".[dev]" from a clone
doubtbench build # download HelpSteer2 @ pinned revision -> data/doubtbench.jsonl
TYPESAFE_API_KEY=... doubtbench run --adapter jev --model jev-1.13.0
A full Jev run is about 1,900 calls and ~2M input tokens, so roughly $0.09 at $0.042 per million (an
estimate from request size). All six questions about one response go in a single call. No TypeSafe key? Use --adapter jev_vercel with AI_GATEWAY_API_KEY.
Claude models (any of them) with an Anthropic API key:
pip install "doubtbench[claude]"
ANTHROPIC_API_KEY=... doubtbench run --adapter claude --model claude-haiku-4-5
Open models through any OpenAI-compatible server (vLLM, llama.cpp, MLX, SGLang) via logprob readout:
vllm serve Qwen/Qwen3.5-4B-Instruct &
doubtbench run --adapter logprob_local --base-url http://localhost:8000/v1 --model Qwen/Qwen3.5-4B-Instruct
Other commands: doubtbench score results/<run_id>, doubtbench compare results/a results/b,
doubtbench leaderboard, doubtbench validate. Add --limit 200 for a quick smoke run.
Take item hs2-val-000088-helpfulness. The user wrote "Could you" and the assistant replied "Sure,
I can help with that. What would you like to know?" Three annotators rated helpfulness: two gave
3 (mostly helpful) and one gave 4 (fully helpful).
| P(2) | P(3) | P(4) | Correct? | JS divergence to humans | |
|---|---|---|---|---|---|
| Humans (2 of 3 votes, 1 of 3 votes) | 0.67 | 0.33 | 0 | ||
| Model A, overconfident | 0.98 | 0.02 | ✓ | 0.143 | |
| Model B, hesitates like humans | 0.05 | 0.62 | 0.33 | ✓ | 0.026 |
Illustrative probabilities; the human votes are real. Both models pick 3, so a single-answer benchmark scores them the same. But model A claims 98% certainty on a question the humans split on, and model B's 62% is about right. DoubtBench tells them apart.
| Task | Type | Questions (val) | What it asks |
|---|---|---|---|
hs2_attribute_score | Score 0–4 | 5,110 | Helpfulness, correctness, coherence, complexity, verbosity of one response |
hs2_is_acceptable | Noul | 1,022 | Is helpfulness ≥ 3? |
hs2_pairwise_pref | Choice | 882 | Response 1, response 2, or tie. Each pair appears twice, swapped, to measure position bias |
hs2_pref_strength | Score −3…+3 | 441 | How strongly one response is preferred |
Attribute questions (verbatim):
| Attribute | Question |
|---|---|
| helpfulness | How helpful is the response to the prompt overall? |
| correctness | Does the response include all pertinent facts without errors? |
| coherence | How consistent and clear is the response? |
| complexity | How much expertise would it take to write this response? |
| verbosity | How much detail does the response include relative to what was asked? |
Level descriptions are in doubtbench/constants.py. One JSON object per line,
compatible with existing System One harnesses (state, type, question, options/levels,
label), plus soft_label (annotator vote shares) and source (CC BY attribution).
Accuracy (plus within-1 for scores), score MAE, NLL, Brier, ECE with a resampled noise floor, a refit temperature (T > 1 overconfident), soft-label KL / Jensen-Shannon, disagreement correlation (Spearman between model entropy and human entropy), accuracy on unanimous vs split items, selective accuracy at 0.7/0.8/0.9 confidence, position consistency, p50/p95 latency and cost per 1k questions. Every metric is broken down by task, attribute, question type and single vs multi-turn.
DoubtBench score = mean of Accuracy, Calibration (1 − ECE) and Human agreement (1 − JS), each 0–100. The components are always shown next to it. Details: docs/METHODOLOGY.md.
logprob_local / systemone_compat); see doubtbench/adapters/.doubtbench run --adapter ... --model ... on the default (validation) data. No --limit.submissions/<model>.yaml and your results/<run_id>/ folder (run.json, raw.jsonl).
CI re-scores from your raw answers and comments the numbers.Full guide: docs/SUBMITTING.md. Don't have the hardware? Open an issue and we'll run it for you.
nvidia/HelpSteer2@990b2711, with SHA-256 of every input and output in data/MANIFEST.json. doubtbench build is deterministic.data/dropped_harmful.json.data/overlap_report.json.--full adds HelpSteer2 train. Those results are reported separately and flagged as possibly contaminated, because HelpSteer2 train is widely used to train reward models.Code: MIT (LICENSE). Data: derived from HelpSteer2 under CC BY 4.0. See DATA_LICENSE.md for attribution and the list of modifications.
If you use DoubtBench, please cite it (CITATION.cff) and the HelpSteer2 papers:
@misc{wang2024helpsteer2,
title={HelpSteer2: Open-source dataset for training top-performing reward models},
author={Zhilin Wang and Yi Dong and Olivier Delalleau and Jiaqi Zeng and Gerald Shen and Daniel Egert and Jimmy J. Zhang and Makesh Narsimhan Sreedhar and Oleksii Kuchaiev},
year={2024}, eprint={2406.08673}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
@misc{wang2024helpsteer2preference,
title={HelpSteer2-Preference: Complementing Ratings with Preferences},
author={Zhilin Wang and Alexander Bukharin and Olivier Delalleau and Daniel Egert and Gerald Shen and Jiaqi Zeng and Oleksii Kuchaiev and Yi Dong},
year={2024}, eprint={2410.01257}, archivePrefix={arXiv}, primaryClass={cs.LG}
}
Python
100.0%