trust123500/doubtbench

Python

2

11 commits

updated Oct 4, 2026

See the code

See what people are saying

README

DoubtBench

The first System One benchmark scored against human uncertainty, not just a single right answer.

Does your decision model hesitate where humans disagree?

System One models like Jev return a probability with every decision, so you can set a threshold: automate the confident cases and send the uncertain ones to a human. But other benchmarks score these models against a single "right" answer, so nobody checks whether the probabilities actually mean anything.

NVIDIA's HelpSteer2 (CC BY 4.0) makes that check possible. Every AI response in it was rated by several human annotators, and their individual ratings are public. If three people split 2 to 1 on a response, a good model should be uncertain about it too.

DoubtBench turns that data into about 7,500 Score, Choice and Noul questions. It scores models on two things: getting the answer right, and whether their probabilities line up with human disagreement.

DoubtBench score by model Model entropy vs human disagreement

Leaderboard

Validation split: 7,455 questions. Two reference rows frame every model. Human (ceiling) replays the annotators' own votes. Uniform (floor) spreads its probability evenly and knows nothing.

RankModelOpenParamsDoubtBenchAccuracyECE / floorJS div. ↓Disagreement corr.$ / 1k qp50 latencyDateVersion
—human (ceiling)——90.482.6%13.76×0.0001.00——2026-10-031.0
1Claude Haiku 4.5 (verbalized)closedundisclosed69.245.8%7.27×0.2920.15$0.66892166 ms2026-10-031.0
2Jev 1.13closedundisclosed67.848.2%12.80×0.2840.14$0.013498 ms2026-10-031.0
3Claude Opus 5.5 (verbalized)closedundisclosed67.844.0%7.97×0.3060.06$3.13683735 ms2026-10-041.0
4Claude Sonnet 5.5 (verbalized)closedundisclosed67.442.5%8.16×0.3010.12$1.55222420 ms2026-10-041.0
—uniform (floor)—054.818.4%21.39×0.396———2026-10-031.0

First results

Four models so far: Jev 1.13, a purpose-built System One model, and three Claude models asked to state a probability for every option.

DoubtBench scoreAccuracyDisagreement corr.Cost of a full runMedian latency
Human annotators (ceiling)90.482.6%1.00
Claude Haiku 4.569.245.8%0.15$5.002.2 s
Jev 1.1367.848.2%0.14$0.100.1 s
Claude Opus 5.567.844.0%0.06$23.303.7 s
Claude Sonnet 5.567.442.5%0.12$11.502.4 s
Uniform guessing (floor)54.818.4%—

No model hesitates where humans disagree. The disagreement correlation measures how much a model's uncertainty rises on responses where the annotators split. The ceiling is 1.0. All four models land between 0.06 and 0.15. They are barely more uncertain on the responses humans argued about than on the ones everyone agreed on. If you plan to send uncertain cases to a human, that matters.

Jev is the most accurate, by far the cheapest and the fastest. It gets more answers right than any Claude model, and the gap holds up statistically (paired McNemar test, p < 0.001 against each). A full run costs about 10 cents, against $5 to $23 for Claude, at 98 ms median latency.

The Claude models score well on calibration because they're less confident. Their average confidence is about 0.54, against Jev's 0.65. Hedging brings stated confidence closer to actual accuracy, which is what keeps them level with Jev on the composite score.

Bigger isn't better here. Haiku 4.5 beats both Sonnet 5.5 and Opus 5.5 on accuracy. The larger models struggle most with the yes/no "is this response acceptable?" question, at 55% against Haiku's 72% and Jev's 75%. Opus is the least swayed by response order, though: it keeps the same answer 90% of the time when the two responses swap places, against 80% for Jev and 64% for Haiku.

How these were run: Claude models answer through structured JSON output, with the questions about one response batched into a single call, as with Jev. Sonnet 5.5 and Opus 5.5 ran at low effort, and higher effort may do better. They also refused 30 and 36 questions (jailbreak-style prompts, flagged cyber), which count as wrong. Haiku and Jev refused none.

Laya, a SemIf/Kev model and an open-weight instruct model are next. Submit yours.

Reading the table

  • DoubtBench score: the average of accuracy, calibration and human agreement, each out of 100.
  • Accuracy: how often the top answer matches the official label.
  • ECE / floor: how far confidence is from actual accuracy, divided by what random noise alone would give. Near 1× is ideal.
  • JS divergence: how different the model's probabilities are from the annotators' votes. 0 is a perfect match.
  • Disagreement correlation: does the model get less sure exactly where humans disagree? 1.0 is perfect, 0 means no relationship.

Why the human row isn't 100

The human row isn't a model. It replays what the annotators actually voted, so read it as the realistic best case.

Its accuracy is 82.6%, not 100%, because HelpSteer2's official label is a rounded average of the votes. Sometimes nobody picked that value: one vote for 2 and one for 4 give a label of 3. That happens for 8% of attribute scores and about a quarter of the preference items. For the same reason, its ECE and temperature look poor even though it matches the humans perfectly.

So compare models to the human row, not to 100. The cleanest comparison is accuracy on questions where every annotator agreed. There the human row scores 100%, Jev 1.13 scores 62%, and the Claude models score 53% to 55%.

Run it in 3 commands

pip install doubtbench                       # or: pip install -e ".[dev]" from a clone
doubtbench build                             # download HelpSteer2 @ pinned revision -> data/doubtbench.jsonl
TYPESAFE_API_KEY=... doubtbench run --adapter jev --model jev-1.13.0

A full Jev run is about 1,900 calls and ~2M input tokens, so roughly $0.09 at $0.042 per million (an estimate from request size). All six questions about one response go in a single call. No TypeSafe key? Use --adapter jev_vercel with AI_GATEWAY_API_KEY.

Claude models (any of them) with an Anthropic API key:

pip install "doubtbench[claude]"
ANTHROPIC_API_KEY=... doubtbench run --adapter claude --model claude-haiku-4-5

Open models through any OpenAI-compatible server (vLLM, llama.cpp, MLX, SGLang) via logprob readout:

vllm serve Qwen/Qwen3.5-4B-Instruct &
doubtbench run --adapter logprob_local --base-url http://localhost:8000/v1 --model Qwen/Qwen3.5-4B-Instruct

Other commands: doubtbench score results/<run_id>, doubtbench compare results/a results/b, doubtbench leaderboard, doubtbench validate. Add --limit 200 for a quick smoke run.

An example

Take item hs2-val-000088-helpfulness. The user wrote "Could you" and the assistant replied "Sure, I can help with that. What would you like to know?" Three annotators rated helpfulness: two gave 3 (mostly helpful) and one gave 4 (fully helpful).

P(2)P(3)P(4)Correct?JS divergence to humans
Humans (2 of 3 votes, 1 of 3 votes)0.670.330
Model A, overconfident0.980.02✓0.143
Model B, hesitates like humans0.050.620.33✓0.026

Illustrative probabilities; the human votes are real. Both models pick 3, so a single-answer benchmark scores them the same. But model A claims 98% certainty on a question the humans split on, and model B's 62% is about right. DoubtBench tells them apart.

Tasks

TaskTypeQuestions (val)What it asks
hs2_attribute_scoreScore 0–45,110Helpfulness, correctness, coherence, complexity, verbosity of one response
hs2_is_acceptableNoul1,022Is helpfulness ≥ 3?
hs2_pairwise_prefChoice882Response 1, response 2, or tie. Each pair appears twice, swapped, to measure position bias
hs2_pref_strengthScore −3…+3441How strongly one response is preferred

Attribute questions (verbatim):

AttributeQuestion
helpfulnessHow helpful is the response to the prompt overall?
correctnessDoes the response include all pertinent facts without errors?
coherenceHow consistent and clear is the response?
complexityHow much expertise would it take to write this response?
verbosityHow much detail does the response include relative to what was asked?

Level descriptions are in doubtbench/constants.py. One JSON object per line, compatible with existing System One harnesses (state, type, question, options/levels, label), plus soft_label (annotator vote shares) and source (CC BY attribution).

Metrics

Accuracy (plus within-1 for scores), score MAE, NLL, Brier, ECE with a resampled noise floor, a refit temperature (T > 1 overconfident), soft-label KL / Jensen-Shannon, disagreement correlation (Spearman between model entropy and human entropy), accuracy on unanimous vs split items, selective accuracy at 0.7/0.8/0.9 confidence, position consistency, p50/p95 latency and cost per 1k questions. Every metric is broken down by task, attribute, question type and single vs multi-turn.

DoubtBench score = mean of Accuracy, Calibration (1 − ECE) and Human agreement (1 − JS), each 0–100. The components are always shown next to it. Details: docs/METHODOLOGY.md.

Submit your model

  1. Write an adapter (or reuse logprob_local / systemone_compat); see doubtbench/adapters/.
  2. doubtbench run --adapter ... --model ... on the default (validation) data. No --limit.
  3. Open a PR with submissions/<model>.yaml and your results/<run_id>/ folder (run.json, raw.jsonl). CI re-scores from your raw answers and comments the numbers.

Full guide: docs/SUBMITTING.md. Don't have the hardware? Open an issue and we'll run it for you.

Data hygiene

  • Pinned source: nvidia/HelpSteer2@990b2711, with SHA-256 of every input and output in data/MANIFEST.json. doubtbench build is deterministic.
  • Harmful content: a keyword filter (code) drops jailbreak prompts and operational harm; dropped indices are in data/dropped_harmful.json.
  • Contamination: items sharing a normalized text or any 8-word sequence with JevBench public items are dropped; see data/overlap_report.json.
  • Train split: --full adds HelpSteer2 train. Those results are reported separately and flagged as possibly contaminated, because HelpSteer2 train is widely used to train reward models.

License and citation

Code: MIT (LICENSE). Data: derived from HelpSteer2 under CC BY 4.0. See DATA_LICENSE.md for attribution and the list of modifications.

If you use DoubtBench, please cite it (CITATION.cff) and the HelpSteer2 papers:

@misc{wang2024helpsteer2,
  title={HelpSteer2: Open-source dataset for training top-performing reward models},
  author={Zhilin Wang and Yi Dong and Olivier Delalleau and Jiaqi Zeng and Gerald Shen and Daniel Egert and Jimmy J. Zhang and Makesh Narsimhan Sreedhar and Oleksii Kuchaiev},
  year={2024}, eprint={2406.08673}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
@misc{wang2024helpsteer2preference,
  title={HelpSteer2-Preference: Complementing Ratings with Preferences},
  author={Zhilin Wang and Alexander Bukharin and Olivier Delalleau and Daniel Egert and Gerald Shen and Jiaqi Zeng and Oleksii Kuchaiev and Yi Dong},
  year={2024}, eprint={2410.01257}, archivePrefix={arXiv}, primaryClass={cs.LG}
}

trust123500/doubtbench

Python

2

11 commits

updated Oct 4, 2026

See the code

See what people are saying

README

DoubtBench

The first System One benchmark scored against human uncertainty, not just a single right answer.

Does your decision model hesitate where humans disagree?

System One models like Jev return a probability with every decision, so you can set a threshold: automate the confident cases and send the uncertain ones to a human. But other benchmarks score these models against a single "right" answer, so nobody checks whether the probabilities actually mean anything.

NVIDIA's HelpSteer2 (CC BY 4.0) makes that check possible. Every AI response in it was rated by several human annotators, and their individual ratings are public. If three people split 2 to 1 on a response, a good model should be uncertain about it too.

DoubtBench turns that data into about 7,500 Score, Choice and Noul questions. It scores models on two things: getting the answer right, and whether their probabilities line up with human disagreement.

DoubtBench score by model Model entropy vs human disagreement

Leaderboard

Validation split: 7,455 questions. Two reference rows frame every model. Human (ceiling) replays the annotators' own votes. Uniform (floor) spreads its probability evenly and knows nothing.

RankModelOpenParamsDoubtBenchAccuracyECE / floorJS div. ↓Disagreement corr.$ / 1k qp50 latencyDateVersion
—human (ceiling)——90.482.6%13.76×0.0001.00——2026-10-031.0
1Claude Haiku 4.5 (verbalized)closedundisclosed69.245.8%7.27×0.2920.15$0.66892166 ms2026-10-031.0
2Jev 1.13closedundisclosed67.848.2%12.80×0.2840.14$0.013498 ms2026-10-031.0
3Claude Opus 5.5 (verbalized)closedundisclosed67.844.0%7.97×0.3060.06$3.13683735 ms2026-10-041.0
4Claude Sonnet 5.5 (verbalized)closedundisclosed67.442.5%8.16×0.3010.12$1.55222420 ms2026-10-041.0
—uniform (floor)—054.818.4%21.39×0.396———2026-10-031.0

First results

Four models so far: Jev 1.13, a purpose-built System One model, and three Claude models asked to state a probability for every option.

DoubtBench scoreAccuracyDisagreement corr.Cost of a full runMedian latency
Human annotators (ceiling)90.482.6%1.00
Claude Haiku 4.569.245.8%0.15$5.002.2 s
Jev 1.1367.848.2%0.14$0.100.1 s
Claude Opus 5.567.844.0%0.06$23.303.7 s
Claude Sonnet 5.567.442.5%0.12$11.502.4 s
Uniform guessing (floor)54.818.4%—

No model hesitates where humans disagree. The disagreement correlation measures how much a model's uncertainty rises on responses where the annotators split. The ceiling is 1.0. All four models land between 0.06 and 0.15. They are barely more uncertain on the responses humans argued about than on the ones everyone agreed on. If you plan to send uncertain cases to a human, that matters.

Jev is the most accurate, by far the cheapest and the fastest. It gets more answers right than any Claude model, and the gap holds up statistically (paired McNemar test, p < 0.001 against each). A full run costs about 10 cents, against $5 to $23 for Claude, at 98 ms median latency.

The Claude models score well on calibration because they're less confident. Their average confidence is about 0.54, against Jev's 0.65. Hedging brings stated confidence closer to actual accuracy, which is what keeps them level with Jev on the composite score.

Bigger isn't better here. Haiku 4.5 beats both Sonnet 5.5 and Opus 5.5 on accuracy. The larger models struggle most with the yes/no "is this response acceptable?" question, at 55% against Haiku's 72% and Jev's 75%. Opus is the least swayed by response order, though: it keeps the same answer 90% of the time when the two responses swap places, against 80% for Jev and 64% for Haiku.

How these were run: Claude models answer through structured JSON output, with the questions about one response batched into a single call, as with Jev. Sonnet 5.5 and Opus 5.5 ran at low effort, and higher effort may do better. They also refused 30 and 36 questions (jailbreak-style prompts, flagged cyber), which count as wrong. Haiku and Jev refused none.

Laya, a SemIf/Kev model and an open-weight instruct model are next. Submit yours.

Reading the table

  • DoubtBench score: the average of accuracy, calibration and human agreement, each out of 100.
  • Accuracy: how often the top answer matches the official label.
  • ECE / floor: how far confidence is from actual accuracy, divided by what random noise alone would give. Near 1× is ideal.
  • JS divergence: how different the model's probabilities are from the annotators' votes. 0 is a perfect match.
  • Disagreement correlation: does the model get less sure exactly where humans disagree? 1.0 is perfect, 0 means no relationship.

Why the human row isn't 100

The human row isn't a model. It replays what the annotators actually voted, so read it as the realistic best case.

Its accuracy is 82.6%, not 100%, because HelpSteer2's official label is a rounded average of the votes. Sometimes nobody picked that value: one vote for 2 and one for 4 give a label of 3. That happens for 8% of attribute scores and about a quarter of the preference items. For the same reason, its ECE and temperature look poor even though it matches the humans perfectly.

So compare models to the human row, not to 100. The cleanest comparison is accuracy on questions where every annotator agreed. There the human row scores 100%, Jev 1.13 scores 62%, and the Claude models score 53% to 55%.

Run it in 3 commands

pip install doubtbench                       # or: pip install -e ".[dev]" from a clone
doubtbench build                             # download HelpSteer2 @ pinned revision -> data/doubtbench.jsonl
TYPESAFE_API_KEY=... doubtbench run --adapter jev --model jev-1.13.0

A full Jev run is about 1,900 calls and ~2M input tokens, so roughly $0.09 at $0.042 per million (an estimate from request size). All six questions about one response go in a single call. No TypeSafe key? Use --adapter jev_vercel with AI_GATEWAY_API_KEY.

Claude models (any of them) with an Anthropic API key:

pip install "doubtbench[claude]"
ANTHROPIC_API_KEY=... doubtbench run --adapter claude --model claude-haiku-4-5

Open models through any OpenAI-compatible server (vLLM, llama.cpp, MLX, SGLang) via logprob readout:

vllm serve Qwen/Qwen3.5-4B-Instruct &
doubtbench run --adapter logprob_local --base-url http://localhost:8000/v1 --model Qwen/Qwen3.5-4B-Instruct

Other commands: doubtbench score results/<run_id>, doubtbench compare results/a results/b, doubtbench leaderboard, doubtbench validate. Add --limit 200 for a quick smoke run.

An example

Take item hs2-val-000088-helpfulness. The user wrote "Could you" and the assistant replied "Sure, I can help with that. What would you like to know?" Three annotators rated helpfulness: two gave 3 (mostly helpful) and one gave 4 (fully helpful).

P(2)P(3)P(4)Correct?JS divergence to humans
Humans (2 of 3 votes, 1 of 3 votes)0.670.330
Model A, overconfident0.980.02✓0.143
Model B, hesitates like humans0.050.620.33✓0.026

Illustrative probabilities; the human votes are real. Both models pick 3, so a single-answer benchmark scores them the same. But model A claims 98% certainty on a question the humans split on, and model B's 62% is about right. DoubtBench tells them apart.

Tasks

TaskTypeQuestions (val)What it asks
hs2_attribute_scoreScore 0–45,110Helpfulness, correctness, coherence, complexity, verbosity of one response
hs2_is_acceptableNoul1,022Is helpfulness ≥ 3?
hs2_pairwise_prefChoice882Response 1, response 2, or tie. Each pair appears twice, swapped, to measure position bias
hs2_pref_strengthScore −3…+3441How strongly one response is preferred

Attribute questions (verbatim):

AttributeQuestion
helpfulnessHow helpful is the response to the prompt overall?
correctnessDoes the response include all pertinent facts without errors?
coherenceHow consistent and clear is the response?
complexityHow much expertise would it take to write this response?
verbosityHow much detail does the response include relative to what was asked?

Level descriptions are in doubtbench/constants.py. One JSON object per line, compatible with existing System One harnesses (state, type, question, options/levels, label), plus soft_label (annotator vote shares) and source (CC BY attribution).

Metrics

Accuracy (plus within-1 for scores), score MAE, NLL, Brier, ECE with a resampled noise floor, a refit temperature (T > 1 overconfident), soft-label KL / Jensen-Shannon, disagreement correlation (Spearman between model entropy and human entropy), accuracy on unanimous vs split items, selective accuracy at 0.7/0.8/0.9 confidence, position consistency, p50/p95 latency and cost per 1k questions. Every metric is broken down by task, attribute, question type and single vs multi-turn.

DoubtBench score = mean of Accuracy, Calibration (1 − ECE) and Human agreement (1 − JS), each 0–100. The components are always shown next to it. Details: docs/METHODOLOGY.md.

Submit your model

  1. Write an adapter (or reuse logprob_local / systemone_compat); see doubtbench/adapters/.
  2. doubtbench run --adapter ... --model ... on the default (validation) data. No --limit.
  3. Open a PR with submissions/<model>.yaml and your results/<run_id>/ folder (run.json, raw.jsonl). CI re-scores from your raw answers and comments the numbers.

Full guide: docs/SUBMITTING.md. Don't have the hardware? Open an issue and we'll run it for you.

Data hygiene

  • Pinned source: nvidia/HelpSteer2@990b2711, with SHA-256 of every input and output in data/MANIFEST.json. doubtbench build is deterministic.
  • Harmful content: a keyword filter (code) drops jailbreak prompts and operational harm; dropped indices are in data/dropped_harmful.json.
  • Contamination: items sharing a normalized text or any 8-word sequence with JevBench public items are dropped; see data/overlap_report.json.
  • Train split: --full adds HelpSteer2 train. Those results are reported separately and flagged as possibly contaminated, because HelpSteer2 train is widely used to train reward models.

License and citation

Code: MIT (LICENSE). Data: derived from HelpSteer2 under CC BY 4.0. See DATA_LICENSE.md for attribution and the list of modifications.

If you use DoubtBench, please cite it (CITATION.cff) and the HelpSteer2 papers:

@misc{wang2024helpsteer2,
  title={HelpSteer2: Open-source dataset for training top-performing reward models},
  author={Zhilin Wang and Yi Dong and Olivier Delalleau and Jiaqi Zeng and Gerald Shen and Daniel Egert and Jimmy J. Zhang and Makesh Narsimhan Sreedhar and Oleksii Kuchaiev},
  year={2024}, eprint={2406.08673}, archivePrefix={arXiv}, primaryClass={cs.CL}
}
@misc{wang2024helpsteer2preference,
  title={HelpSteer2-Preference: Complementing Ratings with Preferences},
  author={Zhilin Wang and Alexander Bukharin and Olivier Delalleau and Daniel Egert and Gerald Shen and Jiaqi Zeng and Oleksii Kuchaiev and Yi Dong},
  year={2024}, eprint={2410.01257}, archivePrefix={arXiv}, primaryClass={cs.LG}
}

Languages

Python

100.0%