Dataset Card for "livebench/model_judgment"
2
13 commits
1 linked in READMEs
updated Apr 7, 2025
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
This dataset contains all model judgments (scores) currently used to create the leaderboard. Our github readme contains instructions for downloading the model judgments (specifically see the section for download_leaderboard.py).
For more information, see our paper.
Dataset Card for "livebench/model_judgment"
2
13 commits
1 linked in READMEs
updated Apr 7, 2025
LiveBench is a benchmark for LLMs designed with test set contamination and objective evaluation in mind. It has the following properties:
This dataset contains all model judgments (scores) currently used to create the leaderboard. Our github readme contains instructions for downloading the model judgments (specifically see the section for download_leaderboard.py).
For more information, see our paper.