arXiv:2410.15522 · 3 repos reference this paper in their README
guijinSON/MM-Eval
20
·
Official implementation for "MM-Eval: A Multilingual Meta-Evaluation Benchmark for LLM-as-a-Judge…
prometheus-eval/prometheus-eval
1,111
Evaluate your LLM's response with Prometheus and GPT4 💯
Cohere-Labs-Community/m-rewardbench
44
Evaluating Reward Models in Multilingual Settings (ACL Main '25)