[NeurIPS 2025] Reasoning Models Better Express Their Confidence"
See the code[tweet (breif overview of the paper)]
🙁 LLMs are overconfident even when they are dead wrong.
🧐 What about reasoning models? Can they actually tell us “My answer is only 60% likely to be correct”?
❗Our paper suggests that they can! Through extensive analysis, we investigate what enables this emergent ability.
# clone the repository
pip install -e lm-eval-harness
pip install -e evalchemy
pip install vllm
bash evalchemy/scripts/reasoning_no_force.sh
bash evalchemy/scripts/reasoning_force.sh
bash evalchemy/scripts/non_reasoning.sh
results/calculate_metrics.ipynb to calculate ECE, Brier Score, and AUROC for the outputs.bash evalchemy/reasoning_slope.sh
bash evalchemy/non_reasoning_slope.sh
results/linear_regression.ipynb to run linear regression on the calibration metrics.Change the dataset path and the model name appropriately referring to the list below.
Reasoning Models
Non-Reasoning Models
bash evalchemy/reasoning_ablations.sh
The code used to create the ablated CoTs are available in ablation_data/.
bash evalchemy/non_reasoning_slow_think.sh
The few-shot slow thinking examples are available in evalchemy/eval/chat_benchmarks/non_reasoning_slow_think/few_shot_prompt.py
@inproceedings{
yoon2025reasoning,
title={Reasoning Models Better Express Their Confidence},
author={Dongkeun Yoon and Seungone Kim and Sohee Yang and Sunkyoung Kim and Soyeon Kim and Yongil Kim and Eunbi Choi and Yireun Kim and Minjoon Seo},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=rbBtoVnduo}
}
Python
92.4%
Jupyter Notebook
6.3%
Shell
1.0%
[NeurIPS 2025] Reasoning Models Better Express Their Confidence"
See the code[tweet (breif overview of the paper)]
🙁 LLMs are overconfident even when they are dead wrong.
🧐 What about reasoning models? Can they actually tell us “My answer is only 60% likely to be correct”?
❗Our paper suggests that they can! Through extensive analysis, we investigate what enables this emergent ability.
# clone the repository
pip install -e lm-eval-harness
pip install -e evalchemy
pip install vllm
bash evalchemy/scripts/reasoning_no_force.sh
bash evalchemy/scripts/reasoning_force.sh
bash evalchemy/scripts/non_reasoning.sh
results/calculate_metrics.ipynb to calculate ECE, Brier Score, and AUROC for the outputs.bash evalchemy/reasoning_slope.sh
bash evalchemy/non_reasoning_slope.sh
results/linear_regression.ipynb to run linear regression on the calibration metrics.Change the dataset path and the model name appropriately referring to the list below.
Reasoning Models
Non-Reasoning Models
bash evalchemy/reasoning_ablations.sh
The code used to create the ablated CoTs are available in ablation_data/.
bash evalchemy/non_reasoning_slow_think.sh
The few-shot slow thinking examples are available in evalchemy/eval/chat_benchmarks/non_reasoning_slow_think/few_shot_prompt.py
@inproceedings{
yoon2025reasoning,
title={Reasoning Models Better Express Their Confidence},
author={Dongkeun Yoon and Seungone Kim and Sohee Yang and Sunkyoung Kim and Soyeon Kim and Yongil Kim and Eunbi Choi and Yireun Kim and Minjoon Seo},
booktitle={The Thirty-ninth Annual Conference on Neural Information Processing Systems},
year={2025},
url={https://openreview.net/forum?id=rbBtoVnduo}
}
Python
92.4%
Jupyter Notebook
6.3%
Shell
1.0%