HAERAE-HUB/KUDGE

Dataset

7

stars

9

commits

1

linked in READMEs

Sep 20, 2024

updated

README

Official data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
TLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)

Dataset Description

At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool for meta-evaluation research.

Subsets

  • Pointwise/Pairwise: The pointwise, and pairwise subset of Kudge. You may directly input the 'judge_query' column to a LLM to use it as an LLM-as-a-Judge.
  • Pointwise/Pairwise-False: A manually created subset with responses corrupted with false information, may be used to test the robustness of automated evaluators against factual hallucinations.
  • Human Annotations: Raw human annotation dataset collected. 5,638 Instances (Note: Expected 5,760, but some are missing due to system errors)

How to Cite.

@article{son2024llm,
  title={LLM-as-a-Judge \& Reward Model: What They Can and Cannot Do},
  author={Son, Guijin and Ko, Hyunwoo and Lee, Hoyoung and Kim, Yewon and Hong, Seunghyeok},
  journal={arXiv preprint arXiv:2409.11239},
  year={2024}
}

Point of Context

spthsrbwls123@yonsei.ac.kr

Contributors

amphora

9 commits

HAERAE-HUB/KUDGE

Dataset

7

stars

9

commits

1

linked in READMEs

Sep 20, 2024

updated

README

Official data repository for LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
TLDR; Automated Evaluators (LLM-as-a-Judge, Reward Models) can be transferred to non-English settings without additional training. (most of the times)

Dataset Description

At the best of our knowledge, KUDGE is the only, non-English, human-annotated meta-evaluation dataset at this point. Consisted of 5,012 human annotation from native Korean speakers, we expect KUDGE to be widely used as a tool for meta-evaluation research.

Subsets

  • Pointwise/Pairwise: The pointwise, and pairwise subset of Kudge. You may directly input the 'judge_query' column to a LLM to use it as an LLM-as-a-Judge.
  • Pointwise/Pairwise-False: A manually created subset with responses corrupted with false information, may be used to test the robustness of automated evaluators against factual hallucinations.
  • Human Annotations: Raw human annotation dataset collected. 5,638 Instances (Note: Expected 5,760, but some are missing due to system errors)

How to Cite.

@article{son2024llm,
  title={LLM-as-a-Judge \& Reward Model: What They Can and Cannot Do},
  author={Son, Guijin and Ko, Hyunwoo and Lee, Hoyoung and Kim, Yewon and Hong, Seunghyeok},
  journal={arXiv preprint arXiv:2409.11239},
  year={2024}
}

Point of Context

spthsrbwls123@yonsei.ac.kr

Contributors

amphora

9 commits