ReportBench is a comprehensive benchmark for evaluating the factual quality and citation behavior of Deep Research agents. Leveraging expert-authored survey papers as ground truth, ReportBench reverse-engineers domain-specific prompts and provides automated tools to assess both cited and non-cited content. [GitHub] / [Paper]
ReportBench addresses this need by:
The dataset construction pipeline consists of four phases:
Survey Paper Identification
Fine-Grained Reference Extraction
Prompt Generation
Application Domain Distribution
ReportBench’s evaluation workflow consists of two complementary validation procedures:
We evaluated two Deep Research products alongside their corresponding base LLMs using the ReportBench benchmark. Table 1 summarizes precision, recall, average references per report, citation match rate, cited statement count, non-cited factual accuracy, and non-cited statement count.
| Test Model | Precision | Recall | Avg Refs | Cit. Match Rate | Cit. Stmt Count | Non-Cit Acc | Non-Cit Stmt Count |
|---|---|---|---|---|---|---|---|
| OpenAI Deep Research | 0.385 | 0.033 | 9.89 | 78.87% | 88.2 | 95.83% | 38.9 |
| Gemini Deep Research | 0.145 | 0.036 | 32.42 | 72.94% | 96.2 | 92.21% | 49.6 |
| gemini-2.5-flash | 0.237 | 0.012 | 5.47 | 44.88% | 12.1 | 98.52% | 11.5 |
| gemini-2.5-pro | 0.269 | 0.010 | 4.27 | 59.24% | 6.58 | 96.08% | 9.35 |
| o3 | 0.299 | 0.031 | 12.26 | 31.43% | 16.16 | 82.22% | 11.51 |
| claude4-sonnet | 0.337 | 0.021 | 6.74 | 73.67% | 14.93 | 92.64% | 17.07 |
Table 1. Performance metrics of Deep Research products and their base models.
Format: JSON Lines (.jsonl) — one JSON object per line. Each object describes a task and includes a ground_truth array containing reference objects (bibs).
{"arxiv_id":"2312.04861","submitter":"Shanliang Yao","authors":"Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding, Eng Gee Lim, Yong Yue, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Xiaohui Zhu, Yutao Yue","title":"Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review","comments":"Accepted by TITS","journal-ref":"IEEE Transactions on Intelligent Transportation Systems 2025","doi":"10.1109\/TITS.2025.3554781","report-no":null,"categories":"cs.CV cs.AI","license":"http:\/\/arxiv.org\/licenses\/nonexclusive-distrib\/1.0\/","abstract":"With the rapid ... radar.","versions":[{"version":"v1","created":"Fri, 8 Dec 2023 06:31:19 GMT"},{"version":"v2","created":"Fri, 19 Apr 2024 08:55:34 GMT"},{"version":"v3","created":"Mon, 21 Apr 2025 08:37:24 GMT"}],"update_date":"2025-04-22","authors_parsed":[["Yao","Shanliang"],["Guan","Runwei"],["Peng","Zitian"],["Xu","Chenhang"],["Shi","Yilu"],["Ding","Weiping"],["Lim","Eng Gee"],["Yue","Yong"],["Seo","Hyungjoon"],["Man","Ka Lok"],["Ma","Jieming"],["Zhu","Xiaohui"],["Yue","Yutao"]],"application_domain":"Transportation and Intelligent Mobility","prompt":"Please help me research the academic advancements in different radar data representation methods in the field of autonomous driving, and ensure only papers published before April 2025 are referenced.","ground_truth":[{"bib_id":"zheng2022tj4dradset","title":"TJ4DRadSet: A 4D radar dataset for autonomous driving","author":"Zheng, Lianqing and Ma, Zhixiong and Zhu, Xichan and Tan, Bin and Li, Sen and Long, Kai and Sun, Weiqi and Chen, Sihan and Zhang, Lu and Wan, Mengyue and others","meta_info":{"organization":"IEEE","year":"2022","pages":"493--498","booktitle":"2022 IEEE 25th international conference on intelligent transportation systems (ITSC)"}},...,]}
{"arxiv_id":"2308.06419","submitter":"Mahsa Golchoubian","authors":"Mahsa Golchoubian, Moojan Ghafurian, Kerstin Dautenhahn, Nasser Lashgarian Azad","title":"Pedestrian Trajectory Prediction in Pedestrian-Vehicle Mixed Environments: A Systematic Review","comments":"Published in IEEE Transactions on Intelligent Transportation Systems","journal-ref":null,"doi":"10.1109\/TITS.2023.3291196","report-no":null,"categories":"cs.RO cs.AI cs.HC cs.LG cs.SY eess.SY","license":"http:\/\/arxiv.org\/licenses\/nonexclusive-distrib\/1.0\/","abstract":"Planning an ... environments are discussed.","versions":[{"version":"v1","created":"Fri, 11 Aug 2023 23:58:51 GMT"}],"update_date":"2023-08-15","authors_parsed":[["Golchoubian","Mahsa"],["Ghafurian","Moojan"],["Dautenhahn","Kerstin"],["Azad","Nasser Lashgarian"]],"application_domain":"Transportation and Intelligent Mobility","prompt":"Please help me summarize the research status in the field of pedestrian trajectory prediction in unstructured environments with human-vehicle interactions prior to August 2023.","ground_truth":[{"bib_id":"helbing1995social","title":"Social force model for pedestrian dynamics","author":"Helbing, Dirk and Molnar, Peter","meta_info":{"publisher":"APS","year":"1995","pages":"4282--4286","number":"5","volume":"51","journal":"Physical review E"}},{"bib_id":"yang2018crowd","title":"Crowd motion detection and prediction for transportation efficiency in shared spaces","author":"Yang, Dongfang and Maroli, John M and Li, Linhui and El-Shaer, Menna and Jabr, Bander A and Redmill, Keith and \u00d6zguner, F\u00fcsun and \u00d6zguner, \u00dcmit","meta_info":{"organization":"IEEE","year":"2018","pages":"1--6","booktitle":"2018 IEEE International Science of Smart City Operations and Platforms Engineering in Partnership with Global City Teams Challenge (SCOPE-GCTC)"}}, ...,]}
@misc{li2025reportbenchevaluatingdeepresearch,
title={ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks},
author={Minghao Li and Ying Zeng and Zhihao Cheng and Cong Ma and Kai Jia},
year={2025},
eprint={2508.15804},
archivePrefix={arXiv},
5 commits
ReportBench is a comprehensive benchmark for evaluating the factual quality and citation behavior of Deep Research agents. Leveraging expert-authored survey papers as ground truth, ReportBench reverse-engineers domain-specific prompts and provides automated tools to assess both cited and non-cited content. [GitHub] / [Paper]
ReportBench addresses this need by:
The dataset construction pipeline consists of four phases:
Survey Paper Identification
Fine-Grained Reference Extraction
Prompt Generation
Application Domain Distribution
ReportBench’s evaluation workflow consists of two complementary validation procedures:
We evaluated two Deep Research products alongside their corresponding base LLMs using the ReportBench benchmark. Table 1 summarizes precision, recall, average references per report, citation match rate, cited statement count, non-cited factual accuracy, and non-cited statement count.
| Test Model | Precision | Recall | Avg Refs | Cit. Match Rate | Cit. Stmt Count | Non-Cit Acc | Non-Cit Stmt Count |
|---|---|---|---|---|---|---|---|
| OpenAI Deep Research | 0.385 | 0.033 | 9.89 | 78.87% | 88.2 | 95.83% | 38.9 |
| Gemini Deep Research | 0.145 | 0.036 | 32.42 | 72.94% | 96.2 | 92.21% | 49.6 |
| gemini-2.5-flash | 0.237 | 0.012 | 5.47 | 44.88% | 12.1 | 98.52% | 11.5 |
| gemini-2.5-pro | 0.269 | 0.010 | 4.27 | 59.24% | 6.58 | 96.08% | 9.35 |
| o3 | 0.299 | 0.031 | 12.26 | 31.43% | 16.16 | 82.22% | 11.51 |
| claude4-sonnet | 0.337 | 0.021 | 6.74 | 73.67% | 14.93 | 92.64% | 17.07 |
Table 1. Performance metrics of Deep Research products and their base models.
Format: JSON Lines (.jsonl) — one JSON object per line. Each object describes a task and includes a ground_truth array containing reference objects (bibs).
{"arxiv_id":"2312.04861","submitter":"Shanliang Yao","authors":"Shanliang Yao, Runwei Guan, Zitian Peng, Chenhang Xu, Yilu Shi, Weiping Ding, Eng Gee Lim, Yong Yue, Hyungjoon Seo, Ka Lok Man, Jieming Ma, Xiaohui Zhu, Yutao Yue","title":"Exploring Radar Data Representations in Autonomous Driving: A Comprehensive Review","comments":"Accepted by TITS","journal-ref":"IEEE Transactions on Intelligent Transportation Systems 2025","doi":"10.1109\/TITS.2025.3554781","report-no":null,"categories":"cs.CV cs.AI","license":"http:\/\/arxiv.org\/licenses\/nonexclusive-distrib\/1.0\/","abstract":"With the rapid ... radar.","versions":[{"version":"v1","created":"Fri, 8 Dec 2023 06:31:19 GMT"},{"version":"v2","created":"Fri, 19 Apr 2024 08:55:34 GMT"},{"version":"v3","created":"Mon, 21 Apr 2025 08:37:24 GMT"}],"update_date":"2025-04-22","authors_parsed":[["Yao","Shanliang"],["Guan","Runwei"],["Peng","Zitian"],["Xu","Chenhang"],["Shi","Yilu"],["Ding","Weiping"],["Lim","Eng Gee"],["Yue","Yong"],["Seo","Hyungjoon"],["Man","Ka Lok"],["Ma","Jieming"],["Zhu","Xiaohui"],["Yue","Yutao"]],"application_domain":"Transportation and Intelligent Mobility","prompt":"Please help me research the academic advancements in different radar data representation methods in the field of autonomous driving, and ensure only papers published before April 2025 are referenced.","ground_truth":[{"bib_id":"zheng2022tj4dradset","title":"TJ4DRadSet: A 4D radar dataset for autonomous driving","author":"Zheng, Lianqing and Ma, Zhixiong and Zhu, Xichan and Tan, Bin and Li, Sen and Long, Kai and Sun, Weiqi and Chen, Sihan and Zhang, Lu and Wan, Mengyue and others","meta_info":{"organization":"IEEE","year":"2022","pages":"493--498","booktitle":"2022 IEEE 25th international conference on intelligent transportation systems (ITSC)"}},...,]}
{"arxiv_id":"2308.06419","submitter":"Mahsa Golchoubian","authors":"Mahsa Golchoubian, Moojan Ghafurian, Kerstin Dautenhahn, Nasser Lashgarian Azad","title":"Pedestrian Trajectory Prediction in Pedestrian-Vehicle Mixed Environments: A Systematic Review","comments":"Published in IEEE Transactions on Intelligent Transportation Systems","journal-ref":null,"doi":"10.1109\/TITS.2023.3291196","report-no":null,"categories":"cs.RO cs.AI cs.HC cs.LG cs.SY eess.SY","license":"http:\/\/arxiv.org\/licenses\/nonexclusive-distrib\/1.0\/","abstract":"Planning an ... environments are discussed.","versions":[{"version":"v1","created":"Fri, 11 Aug 2023 23:58:51 GMT"}],"update_date":"2023-08-15","authors_parsed":[["Golchoubian","Mahsa"],["Ghafurian","Moojan"],["Dautenhahn","Kerstin"],["Azad","Nasser Lashgarian"]],"application_domain":"Transportation and Intelligent Mobility","prompt":"Please help me summarize the research status in the field of pedestrian trajectory prediction in unstructured environments with human-vehicle interactions prior to August 2023.","ground_truth":[{"bib_id":"helbing1995social","title":"Social force model for pedestrian dynamics","author":"Helbing, Dirk and Molnar, Peter","meta_info":{"publisher":"APS","year":"1995","pages":"4282--4286","number":"5","volume":"51","journal":"Physical review E"}},{"bib_id":"yang2018crowd","title":"Crowd motion detection and prediction for transportation efficiency in shared spaces","author":"Yang, Dongfang and Maroli, John M and Li, Linhui and El-Shaer, Menna and Jabr, Bander A and Redmill, Keith and \u00d6zguner, F\u00fcsun and \u00d6zguner, \u00dcmit","meta_info":{"organization":"IEEE","year":"2018","pages":"1--6","booktitle":"2018 IEEE International Science of Smart City Operations and Platforms Engineering in Partnership with Global City Teams Challenge (SCOPE-GCTC)"}}, ...,]}
@misc{li2025reportbenchevaluatingdeepresearch,
title={ReportBench: Evaluating Deep Research Agents via Academic Survey Tasks},
author={Minghao Li and Ying Zeng and Zhihao Cheng and Cong Ma and Kai Jia},
year={2025},
eprint={2508.15804},
archivePrefix={arXiv},
5 commits