arXiv:2403.13787 · 3 repos reference this paper in their README
guijinSON/MM-Eval
20
·
Official implementation for "MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge…
prometheus-eval/prometheus-eval
1,111
Evaluate your LLM's response with Prometheus and GPT4 💯
allenai/reward-bench
109
No description