Research Paper coming soon!
\(K^{2} Eval\) is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion.
The design principle behind \(K^{2} Eval\) centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating benchmarks like MT-Bench or Vicuna-Instructions-80, which would produce Korean-language instructions devoid of cultural relevance. In addition, \(K^{2} Eval\) comprised of question, scoring rubric, evaluation criteria, gold reference answer for the standardized assessment.
The following figure shows the differences between MT-Bench, Vicuna-Instructions-80, and LogicKor.

The following table shows the distribution of subjects and abilities in \(K^{2} Eval\).
| Knowledge Type | Reasoning Type | # of Instance |
|---|---|---|
| Art | Empathetic Reasoning | 5 |
| Culinary | Brainstorming | 5 |
| Culinary | Cause & Effect Analysis | 5 |
| Culture & Traditions | Comparative Analysis | 5 |
| Geography | Cause & Effect Analysis | 5 |
| Geography | Comparative Analysis | 5 |
| Geography | Numerical Estimation | 5 |
| History | Creative Writing | 5 |
| History | Numerical Estimation | 10 |
| Linguistics | Cause & Effect Analysis | 5 |
| Linguistics | Empathetic Reasoning | 5 |
| Literaure | Comparative Analysis | 5 |
| Literature | Creative Writing | 10 |
| Politivs & Economy | Proposing Solutions | 5 |
| Social Issues | Proposing Solutions | 10 |
We assess the benchmark's separability introduced by Arena-Hard to check that the benchmark can effectively differentiate between models. The separability refers to the percentage of model pairs with non-overlapping confidence intervals of benchmark scores, determined via bootstrapping.
The \(K^{2} Eval\) demonstrates high separability at 73.76%, which exceeds that of MT-Bench and LogicKor. Although it is lower than Arena-Hard-v0.1, we suspect this is primarily due to the dataset size. The following table show the result of separability analysis.
| Dataset | Separability | # Models | # Instances |
|---|---|---|---|
| K2-Eval | 73.76% | 31 | 90 |
| LogicKor | 52.94% | 34 | 40 |
| MT-Bench | 22.60% | 20 | 80 |
| Arena-Hard-v0.1 | 87.40% | 20 | 500 |
In our research, we employ 15 human judges for annotation. The judges are provided with instructions, reference answers, rubrics and model reponses and tasked to score between 1 to 5. All responses are scored a minimum of two times to ensure quality.
We observe HyperCLOVA X to show the highest performance on the benchmark. We also discover the Importance of targeted instruction tuning using Korean data. Specifically, models such as EEVE-Korean-Instruct-10.8B and KULLM3 exhibit human preference comparable to much larger models like Command-R-Plus-104B and Mixtral-8x22B-Instruct. This indicates that localized tuning that addresses linguistic and cultural nuances is necessary beyond raw computational budget or size to improve human preference.

Guijin Son, Ko Hyun Woo, Hoyoung Lee, Seunghyeok Hong, Yewon Kim, Jungwoo Kim
Special thanks to our annotators.
Hagyun Gill, Jiyeon Kim, Chaejun Seo, Hayoung Eun, Dahyun Lee, Seonging Cho, Inhae Cho, Yeonjun Choi ,Sujin Jang, Hyejung Choi
For any questions contact us via the following email :)
spthsrbwls123@yonsei.ac.kr
17 commits
5 commits
Research Paper coming soon!
\(K^{2} Eval\) is a novel benchmark featuring 90 handwritten instructions that require in-depth knowledge of Korean language and culture for accurate completion.
The design principle behind \(K^{2} Eval\) centers on collecting instructions that necessitate knowledge specific to Korean culture and context in order to solve. This approach distinguishes our work from simply translating benchmarks like MT-Bench or Vicuna-Instructions-80, which would produce Korean-language instructions devoid of cultural relevance. In addition, \(K^{2} Eval\) comprised of question, scoring rubric, evaluation criteria, gold reference answer for the standardized assessment.
The following figure shows the differences between MT-Bench, Vicuna-Instructions-80, and LogicKor.

The following table shows the distribution of subjects and abilities in \(K^{2} Eval\).
| Knowledge Type | Reasoning Type | # of Instance |
|---|---|---|
| Art | Empathetic Reasoning | 5 |
| Culinary | Brainstorming | 5 |
| Culinary | Cause & Effect Analysis | 5 |
| Culture & Traditions | Comparative Analysis | 5 |
| Geography | Cause & Effect Analysis | 5 |
| Geography | Comparative Analysis | 5 |
| Geography | Numerical Estimation | 5 |
| History | Creative Writing | 5 |
| History | Numerical Estimation | 10 |
| Linguistics | Cause & Effect Analysis | 5 |
| Linguistics | Empathetic Reasoning | 5 |
| Literaure | Comparative Analysis | 5 |
| Literature | Creative Writing | 10 |
| Politivs & Economy | Proposing Solutions | 5 |
| Social Issues | Proposing Solutions | 10 |
We assess the benchmark's separability introduced by Arena-Hard to check that the benchmark can effectively differentiate between models. The separability refers to the percentage of model pairs with non-overlapping confidence intervals of benchmark scores, determined via bootstrapping.
The \(K^{2} Eval\) demonstrates high separability at 73.76%, which exceeds that of MT-Bench and LogicKor. Although it is lower than Arena-Hard-v0.1, we suspect this is primarily due to the dataset size. The following table show the result of separability analysis.
| Dataset | Separability | # Models | # Instances |
|---|---|---|---|
| K2-Eval | 73.76% | 31 | 90 |
| LogicKor | 52.94% | 34 | 40 |
| MT-Bench | 22.60% | 20 | 80 |
| Arena-Hard-v0.1 | 87.40% | 20 | 500 |
In our research, we employ 15 human judges for annotation. The judges are provided with instructions, reference answers, rubrics and model reponses and tasked to score between 1 to 5. All responses are scored a minimum of two times to ensure quality.
We observe HyperCLOVA X to show the highest performance on the benchmark. We also discover the Importance of targeted instruction tuning using Korean data. Specifically, models such as EEVE-Korean-Instruct-10.8B and KULLM3 exhibit human preference comparable to much larger models like Command-R-Plus-104B and Mixtral-8x22B-Instruct. This indicates that localized tuning that addresses linguistic and cultural nuances is necessary beyond raw computational budget or size to improve human preference.

Guijin Son, Ko Hyun Woo, Hoyoung Lee, Seunghyeok Hong, Yewon Kim, Jungwoo Kim
Special thanks to our annotators.
Hagyun Gill, Jiyeon Kim, Chaejun Seo, Hayoung Eun, Dahyun Lee, Seonging Cho, Inhae Cho, Yeonjun Choi ,Sujin Jang, Hyejung Choi
For any questions contact us via the following email :)
spthsrbwls123@yonsei.ac.kr
17 commits
5 commits