arXiv:2310.07641 · 2 repos reference this paper in their README
princeton-nlp/LLMBar
141
·
[ICLR 2024] Evaluating Large Language Models at Evaluating Instruction Following
allenai/reward-bench
109
No description