Astrobench is a specialized benchmarking dataset for evaluating the performance of Large Language Models in astronomy and astrophysics knowledge recall. This dataset consists of multiple-choice questions derived from the Annual Review of Astronomy and Astrophysics, designed to test models' comprehension and knowledge of astronomical research.
Here's how to load and use the dataset:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("AstroMLab/Astrobench_MCQ_v1_Public")
# Print the first question and its MCQ options
index = 0
example = dataset['train'][index]
print("Question:", example['question'])
print("\nOptions:")
print("A:", example['A'])
print("B:", example['B'])
print("C:", example['C'])
print("D:", example['D'])
This public release intentionally omits the correct answers (keys) to the multiple-choice questions. This design choice is deliberate to:
We encourage researchers and developers to evaluate their models using this benchmark. To receive an evaluation score:
Questions were generated following specific criteria:
If you use this dataset in your research, please cite:
@ARTICLE{2024arXiv240711194T,
author = {{Ting}, Yuan-Sen and {Dung Nguyen}, Tuan and {Ghosal}, Tirthankar and {Pan}, Rui and {Arora}, Hardik and {Sun}, Zechang and {de Haan}, Tijmen and {Ramachandra}, Nesar and {Wells}, Azton and {Madireddy}, Sandeep and {Accomazzi}, Alberto},
title = "{AstroMLab 1: Who Wins Astronomy Jeopardy!?}",
journal = {arXiv e-prints},
keywords = {Astrophysics - Instrumentation and Methods for Astrophysics, Astrophysics - Earth and Planetary Astrophysics, Astrophysics - Astrophysics of Galaxies, Astrophysics - Solar and Stellar Astrophysics, Computer Science - Artificial Intelligence, Computer Science - Computation and Language},
year = 2024,
month = jul,
eid = {arXiv:2407.11194},
pages = {arXiv:2407.11194},
doi = {10.48550/arXiv.2407.11194},
archivePrefix = {arXiv},
eprint = {2407.11194},
primaryClass = {astro-ph.IM},
adsurl = {https://ui.adsabs.harvard.edu/abs/2024arXiv240711194T},
adsnote = {Provided by the SAO/NASA Astrophysics Data System}
}
Astrobench is a specialized benchmarking dataset for evaluating the performance of Large Language Models in astronomy and astrophysics knowledge recall. This dataset consists of multiple-choice questions derived from the Annual Review of Astronomy and Astrophysics, designed to test models' comprehension and knowledge of astronomical research.
Here's how to load and use the dataset:
from datasets import load_dataset
# Load the dataset
dataset = load_dataset("AstroMLab/Astrobench_MCQ_v1_Public")
# Print the first question and its MCQ options
index = 0
example = dataset['train'][index]
print("Question:", example['question'])
print("\nOptions:")
print("A:", example['A'])
print("B:", example['B'])
print("C:", example['C'])
print("D:", example['D'])
This public release intentionally omits the correct answers (keys) to the multiple-choice questions. This design choice is deliberate to:
We encourage researchers and developers to evaluate their models using this benchmark. To receive an evaluation score:
Questions were generated following specific criteria:
If you use this dataset in your research, please cite:
@ARTICLE{2024arXiv240711194T,
author = {{Ting}, Yuan-Sen and {Dung Nguyen}, Tuan and {Ghosal}, Tirthankar and {Pan}, Rui and {Arora}, Hardik and {Sun}, Zechang and {de Haan}, Tijmen and {Ramachandra}, Nesar and {Wells}, Azton and {Madireddy}, Sandeep and {Accomazzi}, Alberto},
title = "{AstroMLab 1: Who Wins Astronomy Jeopardy!?}",
journal = {arXiv e-prints},
keywords = {Astrophysics - Instrumentation and Methods for Astrophysics, Astrophysics - Earth and Planetary Astrophysics, Astrophysics - Astrophysics of Galaxies, Astrophysics - Solar and Stellar Astrophysics, Computer Science - Artificial Intelligence, Computer Science - Computation and Language},
year = 2024,
month = jul,
eid = {arXiv:2407.11194},
pages = {arXiv:2407.11194},
doi = {10.48550/arXiv.2407.11194},
archivePrefix = {arXiv},
eprint = {2407.11194},
primaryClass = {astro-ph.IM},
adsurl = {https://ui.adsabs.harvard.edu/abs/2024arXiv240711194T},
adsnote = {Provided by the SAO/NASA Astrophysics Data System}
}