Hothan/OlympiadBench

Dataset

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024]

47

8 commits

2 linked in READMEs

updated Jun 8, 2025

See the code

README

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024]

📖 arXiv | GitHub

Note: We have made adjustments to the image content in the multimodal portion of the dataset and fixed previous issues where some images in the English physics subset were not displayed properly. If your usage involves images, please re-download the dataset (we recommend all users to download the latest version).

Additionally, some entries in the solution field may also include images. However, due to image display limitations on Hugging Face, we did not include them in this update. If you need the images embedded in the solution field, please download the full dataset from the whole data link. This version contains all the original image content.

Thank you for your support of OlympiadBench — we hope it is helpful to your work.

Dataset Description

OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning.

More details are at our GitHub.

Contact

Citation

If you do find our code helpful or use our benchmark dataset, please citing our paper.

BibTeX:

@article{he2024olympiadbench,
  title={Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems},
  author={He, Chaoqun and Luo, Renjie and Bai, Yuzhuo and Hu, Shengding and Thai, Zhen Leng and Shen, Junhao and Hu, Jinyi and Han, Xu and Huang, Yujie and Zhang, Yuxiang and others},
  journal={arXiv preprint arXiv:2402.14008},
  year={2024}
}
math
physics

Contributors

Hothan

8 commits

Hothan/OlympiadBench

Dataset

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024]

47

8 commits

2 linked in READMEs

updated Jun 8, 2025

See the code

README

OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems[ACL 2024]

📖 arXiv | GitHub

Note: We have made adjustments to the image content in the multimodal portion of the dataset and fixed previous issues where some images in the English physics subset were not displayed properly. If your usage involves images, please re-download the dataset (we recommend all users to download the latest version).

Additionally, some entries in the solution field may also include images. However, due to image display limitations on Hugging Face, we did not include them in this update. If you need the images embedded in the solution field, please download the full dataset from the whole data link. This version contains all the original image content.

Thank you for your support of OlympiadBench — we hope it is helpful to your work.

Dataset Description

OlympiadBench is an Olympiad-level bilingual multimodal scientific benchmark, featuring 8,476 problems from Olympiad-level mathematics and physics competitions, including the Chinese college entrance exam. Each problem is detailed with expert-level annotations for step-by-step reasoning. Notably, the best-performing model, GPT-4V, attains an average score of 17.97% on OlympiadBench, with a mere 10.74% in physics, highlighting the benchmark rigor and the intricacy of physical reasoning.

More details are at our GitHub.

Contact

Citation

If you do find our code helpful or use our benchmark dataset, please citing our paper.

BibTeX:

@article{he2024olympiadbench,
  title={Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems},
  author={He, Chaoqun and Luo, Renjie and Bai, Yuzhuo and Hu, Shengding and Thai, Zhen Leng and Shen, Junhao and Hu, Jinyi and Han, Xu and Huang, Yujie and Zhang, Yuxiang and others},
  journal={arXiv preprint arXiv:2402.14008},
  year={2024}
}
math
physics

Contributors

Hothan

8 commits