CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
BibTeX:
@misc{lin2024criticbench,
title={CriticBench: Benchmarking LLMs for Critique-Correct Reasoning},
author={Zicheng Lin and Zhibin Gou and Tian Liang and Ruilin Luo and Haowei Liu and Yujiu Yang},
year={2024},
eprint={2402.14809},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
APA:
Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., & Yang, Y. (2024). CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
5 commits
CriticBench is a comprehensive benchmark designed to assess LLMs' abilities to generate, critique/discriminate and correct reasoning across a variety of tasks. CriticBench encompasses five reasoning domains: mathematical, commonsense, symbolic, coding, and algorithmic. It compiles 15 datasets and incorporates responses from three LLM families.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
Users should be made aware of the risks, biases and limitations of the dataset. More information needed for further recommendations.
BibTeX:
@misc{lin2024criticbench,
title={CriticBench: Benchmarking LLMs for Critique-Correct Reasoning},
author={Zicheng Lin and Zhibin Gou and Tian Liang and Ruilin Luo and Haowei Liu and Yujiu Yang},
year={2024},
eprint={2402.14809},
archivePrefix={arXiv},
primaryClass={cs.CL}
}
APA:
Lin, Z., Gou, Z., Liang, T., Luo, R., Liu, H., & Yang, Y. (2024). CriticBench: Benchmarking LLMs for Critique-Correct Reasoning.
[More Information Needed]
[More Information Needed]
[More Information Needed]
[More Information Needed]
5 commits