LGAI-EXAONE/KoMT-Bench

Dataset

42

stars

3

commits

2

linked in READMEs

Aug 8, 2024

updated

evaluation
instruction-following
language model
LLM-as-a-judge

README

KoMT-Bench

Introduction

We present KoMT-Bench, a benchmark designed to evaluate the capability of language models in following instructions in Korean. KoMT-Bench is an in-house dataset created by translating MT-Bench [1] dataset into Korean and modifying some questions to reflect the characteristics and cultural nuances of the Korean language. After the initial translation and modification, we requested expert linguists to conduct a thorough review of our benchmark dataset.

To conduct evaluations on KoMT-Bench, please visit the official KoMT-Bench GitHub repository in which the evaluation scripts are provided.

Here are examples from KoMT-Bench:

CategoryMT-BenchKoMT-Bench
Writing
1st TurnImagine you are writing a blog post comparing two popular smartphone models. Develop an outline for the blog post, including key points and subheadings to effectively compare and contrast the features, performance, and user experience of the two models. Please answer in fewer than 200 words.두 개의 인기 스마트폰 모델을 비교하는 블로그 게시물을 작성한다고 가정합니다. 두 모델의 기능, 성능, 사용자 경험을 효과적으로 비교하고 대조할 수 있도록 핵심 사항과 소제목을 포함하여 블로그 게시물의 개요를 작성하세요. 200자 이내로 답하십시오.
2nd TurnTake your previous response and rephrase it as a limerick.이전 답변을 충청도 사투리로 재작성하십시오.
Math
1st TurnWhen a number is divided by 10, the remainder is 4. What is the remainder when twice the number is divided by 4?어떤 숫자를 10으로 나눈 나머지는 4입니다. 그 숫자의 두 배를 4로 나눈 나머지를 구하세요.
2nd TurnWhat about when twice the number is divided by 5?그 숫자의 두 배를 5로 나누면 어떨까요?
Humanities
1st TurnProvide insights into the correlation between economic indicators such as GDP, inflation, and unemployment rates. Explain how fiscal and monetary policies affect those indicators.GDP, 인플레이션, 실업률과 같은 경제 지표 간의 상관관계에 대한 통찰을 제시하세요. 이러한 지표들에 재정 및 통화 정책이 어떤 영향을 미치는지 설명하세요.
2nd TurnNow, explain them again like I'm five.이제 제가 5살이라 생각하고 다시 설명해 주세요.

Models Results

Here are the evaluation results of various language models including EXAONE 3.0 7.8B instruction-tuned model on KoMT-Bench. Please refer to EXAONE 3.0 technical report for details.

EXAONE 3.0 7.8B Inst.Llama 3.1 8B Inst.Gemma 2 9B Inst.QWEN 2 7B Inst.Phi 3 7B Inst.Mistral 7B Inst.
KoMT-Bench8.926.067.927.694.875.20

References

[1] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 46595–46623. Curran Associates, Inc., 2023.


Citation

@misc{KoMT-Bench,
  author = {LG AI Research},
  title = {KoMT-Bench},
  year = {2024},
  publisher = {Hugging Face},
  journal = {Hugging Face repository},
  howpublished = {\url{https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench}}
}

Contributors

LG-AI-EXAONE

2 commits

nuxlear

1 commits

LGAI-EXAONE/KoMT-Bench

Dataset

42

stars

3

commits

2

linked in READMEs

Aug 8, 2024

updated

evaluation
instruction-following
language model
LLM-as-a-judge

README

KoMT-Bench

Introduction

We present KoMT-Bench, a benchmark designed to evaluate the capability of language models in following instructions in Korean. KoMT-Bench is an in-house dataset created by translating MT-Bench [1] dataset into Korean and modifying some questions to reflect the characteristics and cultural nuances of the Korean language. After the initial translation and modification, we requested expert linguists to conduct a thorough review of our benchmark dataset.

To conduct evaluations on KoMT-Bench, please visit the official KoMT-Bench GitHub repository in which the evaluation scripts are provided.

Here are examples from KoMT-Bench:

CategoryMT-BenchKoMT-Bench
Writing
1st TurnImagine you are writing a blog post comparing two popular smartphone models. Develop an outline for the blog post, including key points and subheadings to effectively compare and contrast the features, performance, and user experience of the two models. Please answer in fewer than 200 words.두 개의 인기 스마트폰 모델을 비교하는 블로그 게시물을 작성한다고 가정합니다. 두 모델의 기능, 성능, 사용자 경험을 효과적으로 비교하고 대조할 수 있도록 핵심 사항과 소제목을 포함하여 블로그 게시물의 개요를 작성하세요. 200자 이내로 답하십시오.
2nd TurnTake your previous response and rephrase it as a limerick.이전 답변을 충청도 사투리로 재작성하십시오.
Math
1st TurnWhen a number is divided by 10, the remainder is 4. What is the remainder when twice the number is divided by 4?어떤 숫자를 10으로 나눈 나머지는 4입니다. 그 숫자의 두 배를 4로 나눈 나머지를 구하세요.
2nd TurnWhat about when twice the number is divided by 5?그 숫자의 두 배를 5로 나누면 어떨까요?
Humanities
1st TurnProvide insights into the correlation between economic indicators such as GDP, inflation, and unemployment rates. Explain how fiscal and monetary policies affect those indicators.GDP, 인플레이션, 실업률과 같은 경제 지표 간의 상관관계에 대한 통찰을 제시하세요. 이러한 지표들에 재정 및 통화 정책이 어떤 영향을 미치는지 설명하세요.
2nd TurnNow, explain them again like I'm five.이제 제가 5살이라 생각하고 다시 설명해 주세요.

Models Results

Here are the evaluation results of various language models including EXAONE 3.0 7.8B instruction-tuned model on KoMT-Bench. Please refer to EXAONE 3.0 technical report for details.

EXAONE 3.0 7.8B Inst.Llama 3.1 8B Inst.Gemma 2 9B Inst.QWEN 2 7B Inst.Phi 3 7B Inst.Mistral 7B Inst.
KoMT-Bench8.926.067.927.694.875.20

References

[1] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, Hao Zhang, Joseph E Gonzalez, and Ion Stoica. Judging llm-as-a-judge with mt-bench and chatbot arena. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Information Processing Systems, volume 36, pages 46595–46623. Curran Associates, Inc., 2023.


Citation

@misc{KoMT-Bench,
  author = {LG AI Research},
  title = {KoMT-Bench},
  year = {2024},
  publisher = {Hugging Face},
  journal = {Hugging Face repository},
  howpublished = {\url{https://huggingface.co/datasets/LGAI-EXAONE/KoMT-Bench}}
}

Contributors

LG-AI-EXAONE

2 commits

nuxlear

1 commits