4 repos
lmsys/mt_bench_human_judgments
Content
147
11 commits
rwq-elo/rwq-battle-records
RWQ battle records dataset
0
8 commits
lmsys/lmsys-arena-human-preference-55k
Dataset for [Kaggle competition](https://www.kaggle.com/competitions/lmsys-chatbot-arena/overview)…
159
5 commits
ScalerLab/JudgeBench
JudgeBench: A Benchmark for Evaluating LLM-Based Judges
12
1 commits