Official data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
TLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)
At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool for meta-evaluation research.
@article{son2024llm,
title={LLM-as-a-Judge \& Reward Model: What They Can and Cannot Do},
author={Son, Guijin and Ko, Hyunwoo and Lee, Hoyoung and Kim, Yewon and Hong, Seunghyeok},
journal={arXiv preprint arXiv:2409.11239},
year={2024}
}
spthsrbwls123@yonsei.ac.kr
9 commits
Official data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
TLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)
At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool for meta-evaluation research.
@article{son2024llm,
title={LLM-as-a-Judge \& Reward Model: What They Can and Cannot Do},
author={Son, Guijin and Ko, Hyunwoo and Lee, Hoyoung and Kim, Yewon and Hong, Seunghyeok},
journal={arXiv preprint arXiv:2409.11239},
year={2024}
}
spthsrbwls123@yonsei.ac.kr
9 commits