Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw detection, falsification, calibration, and critique quality into one 0 to 100 score.
Live leaderboard · Methods · Use the dataset
Grok-4.2 leads at 74.5. Claude-Opus-4.7 and Claude-Opus-4.6 follow within 1.3 points. Those small gaps are not reliable rank differences, so the more useful question is which models sit on the efficient frontier.
The frontier shows the best available tradeoffs between two things people often confuse:
Claude-Opus-4.6, Claude-Opus-4.7, Grok-4.2, and Grok-4.3 form the frontier in this panel. A model below that line is beaten by another model on both measures.
Gemma, gpt-oss, and Llama are omitted from this one scatter plot because they stretch the axes and make the central comparison harder to read. They remain in the full ranking above and the table below.
A model that gives the right answer for the wrong reason is brittle. A model that identifies flaws but stays overconfident is hard to trust. REFUTE separates those failure modes instead of rewarding polished prose alone.
That produces useful diagnoses, not just a ranking:
Grok-4.1-Fast writes strong critiques but is weak on the underlying facts. gpt-oss-120b often knows the paper but struggles to critique and calibrate. Llama-3.3-70B is weak across most components. Different model families fail in different ways.
| Rank | Model | Truth | Know | Flaws | Falsify | Scope | Skill | Brier↓ |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok-4.2 | 74.5 | 85% | 81% | 98.8% | 98.3% | 7.59 | 0.173 |
| 2 | Claude-Opus-4.7 | 73.4 | 78.8% | 84% | 97.5% | 98.3% | 6.95 | 0.166 |
| 3 | Claude-Opus-4.6 | 73.2 | 78.8% | 83% | 93.8% | 95% | 6.61 | 0.149 |
| 4 | Grok-3-Mini | 72.2 | 80% | 83% | 96.2% | 98.3% | 7.46 | 0.189 |
| 5 | Grok-4.3 | 71.5 | 80% | 81% | 97.5% | 100% | 7.61 | 0.198 |
| 6 | Gemini-3.1-Pro | 71.2 | 88.8% | 85% | 98.8% | 100% | 6.42 | 0.216 |
| 7 | Kimi-K2.6 | 71.0 | 77.5% | 80% | 92.5% | 100% | 6.42 | 0.163 |
| 8 | GPT-5.2 | 69.2 | 82.5% | 79% | 85% | 98.3% | 7.04 | 0.191 |
| 9 | GPT-5.6 | 67.8 | 82.5% | 86% | 95% | 100% | 7.03 | 0.287 |
| 10 | Gemma-4-31B | 67.2 | 82.5% | 77% | 97.5% | 96.7% | 5.62 | 0.205 |
| 11 | GPT-5.4 | 64.3 | 76.2% | 75% | 95% | 98.3% | 6.96 | 0.242 |
| 12 | DeepSeek-V4-Pro | 62.9 | 77.5% | 80% | 83.8% | 100% | 6.32 | 0.246 |
| 13 | Grok-4.1-Fast | 59.4 | 66.2% | 68% | 80% | 95% | 7.04 | 0.228 |
| 14 | gpt-oss-120b | 57.5 | 80% | 74% | 76.2% | 96.7% | 4.47 | 0.494 |
| 15 | Llama-3.3-70B | 46.3 | 61.3% | 54% | 60% | 76.7% | 4.47 | 0.238 |
Question columns are accuracy. Skill is scored out of 10. Brier is calibration error, so lower is better. Blank API answers count as wrong.
Five additional models have partial results but no Truth Score: GLM-5.1, Qwen3-235B, Qwen3.5-397B, GLM-5, and Cogito-v2.1.
320 auto-graded questions ask four things: What did the study find? What would overturn the claim? How far does the evidence reach? Which summary is honest?
The scope question is nearly solved in this panel, so it receives only 5% of Truth Score. Knowledge and honest-summary questions are tied at 78%, leaving meaningful room for improvement on both.
Replicate returns succeeded but produces an empty answer on REFUTE biology and medical prompts, usually after only 2 to 7 output tokens. Ordinary non-science prompts work. This appears to be route-level safety filtering, not a token-budget failure. We will not turn refusals into a score. A valid run needs direct Anthropic API access or another route that does not blank scientific prompts.
Interactive leaderboard · Full methods
from datasets import load_dataset
ds = load_dataset("BGPT-OFFICIAL/refute", "refute_discrimination_hard", split="train")
print(ds[0]["prompt"][:400])
# Grade the final line: ANSWER=A/B/C/D
Integrators: INTEGRATORS.md · Eval protocol
Public rows strip DOI, title, and source hash.
Full methods: TECHNICAL_REPORT.md · release summary
Why skill isn't truth · Changelog · Cite
@misc{bgpt_refute_2026,
title = {REFUTE: Reasoning Over Evidence Benchmark},
author = {{BGPT Team}},
year = {2026},
url = {https://huggingface.co/datasets/BGPT-OFFICIAL/refute}
}
Built by BGPT, structured evidence from full-text science papers.
127 commits
12 commits
Models can sound convincing while misreading a result or expressing more confidence than the evidence deserves. That matters when people use them to summarize papers, compare studies, or decide what to investigate next.
REFUTE tests whether a model knows the finding, spots quiet flaws, names what would overturn a claim, and matches its confidence to the evidence.
Truth Score is the main result. It combines factual accuracy, flaw detection, falsification, calibration, and critique quality into one 0 to 100 score.
Live leaderboard · Methods · Use the dataset
Grok-4.2 leads at 74.5. Claude-Opus-4.7 and Claude-Opus-4.6 follow within 1.3 points. Those small gaps are not reliable rank differences, so the more useful question is which models sit on the efficient frontier.
The frontier shows the best available tradeoffs between two things people often confuse:
Claude-Opus-4.6, Claude-Opus-4.7, Grok-4.2, and Grok-4.3 form the frontier in this panel. A model below that line is beaten by another model on both measures.
Gemma, gpt-oss, and Llama are omitted from this one scatter plot because they stretch the axes and make the central comparison harder to read. They remain in the full ranking above and the table below.
A model that gives the right answer for the wrong reason is brittle. A model that identifies flaws but stays overconfident is hard to trust. REFUTE separates those failure modes instead of rewarding polished prose alone.
That produces useful diagnoses, not just a ranking:
Grok-4.1-Fast writes strong critiques but is weak on the underlying facts. gpt-oss-120b often knows the paper but struggles to critique and calibrate. Llama-3.3-70B is weak across most components. Different model families fail in different ways.
| Rank | Model | Truth | Know | Flaws | Falsify | Scope | Skill | Brier↓ |
|---|---|---|---|---|---|---|---|---|
| 1 | Grok-4.2 | 74.5 | 85% | 81% | 98.8% | 98.3% | 7.59 | 0.173 |
| 2 | Claude-Opus-4.7 | 73.4 | 78.8% | 84% | 97.5% | 98.3% | 6.95 | 0.166 |
| 3 | Claude-Opus-4.6 | 73.2 | 78.8% | 83% | 93.8% | 95% | 6.61 | 0.149 |
| 4 | Grok-3-Mini | 72.2 | 80% | 83% | 96.2% | 98.3% | 7.46 | 0.189 |
| 5 | Grok-4.3 | 71.5 | 80% | 81% | 97.5% | 100% | 7.61 | 0.198 |
| 6 | Gemini-3.1-Pro | 71.2 | 88.8% | 85% | 98.8% | 100% | 6.42 | 0.216 |
| 7 | Kimi-K2.6 | 71.0 | 77.5% | 80% | 92.5% | 100% | 6.42 | 0.163 |
| 8 | GPT-5.2 | 69.2 | 82.5% | 79% | 85% | 98.3% | 7.04 | 0.191 |
| 9 | GPT-5.6 | 67.8 | 82.5% | 86% | 95% | 100% | 7.03 | 0.287 |
| 10 | Gemma-4-31B | 67.2 | 82.5% | 77% | 97.5% | 96.7% | 5.62 | 0.205 |
| 11 | GPT-5.4 | 64.3 | 76.2% | 75% | 95% | 98.3% | 6.96 | 0.242 |
| 12 | DeepSeek-V4-Pro | 62.9 | 77.5% | 80% | 83.8% | 100% | 6.32 | 0.246 |
| 13 | Grok-4.1-Fast | 59.4 | 66.2% | 68% | 80% | 95% | 7.04 | 0.228 |
| 14 | gpt-oss-120b | 57.5 | 80% | 74% | 76.2% | 96.7% | 4.47 | 0.494 |
| 15 | Llama-3.3-70B | 46.3 | 61.3% | 54% | 60% | 76.7% | 4.47 | 0.238 |
Question columns are accuracy. Skill is scored out of 10. Brier is calibration error, so lower is better. Blank API answers count as wrong.
Five additional models have partial results but no Truth Score: GLM-5.1, Qwen3-235B, Qwen3.5-397B, GLM-5, and Cogito-v2.1.
320 auto-graded questions ask four things: What did the study find? What would overturn the claim? How far does the evidence reach? Which summary is honest?
The scope question is nearly solved in this panel, so it receives only 5% of Truth Score. Knowledge and honest-summary questions are tied at 78%, leaving meaningful room for improvement on both.
Replicate returns succeeded but produces an empty answer on REFUTE biology and medical prompts, usually after only 2 to 7 output tokens. Ordinary non-science prompts work. This appears to be route-level safety filtering, not a token-budget failure. We will not turn refusals into a score. A valid run needs direct Anthropic API access or another route that does not blank scientific prompts.
Interactive leaderboard · Full methods
from datasets import load_dataset
ds = load_dataset("BGPT-OFFICIAL/refute", "refute_discrimination_hard", split="train")
print(ds[0]["prompt"][:400])
# Grade the final line: ANSWER=A/B/C/D
Integrators: INTEGRATORS.md · Eval protocol
Public rows strip DOI, title, and source hash.
Full methods: TECHNICAL_REPORT.md · release summary
Why skill isn't truth · Changelog · Cite
@misc{bgpt_refute_2026,
title = {REFUTE: Reasoning Over Evidence Benchmark},
author = {{BGPT Team}},
year = {2026},
url = {https://huggingface.co/datasets/BGPT-OFFICIAL/refute}
}
Built by BGPT, structured evidence from full-text science papers.
127 commits
12 commits