MediaTek-Research/TCEval-v2

Dataset

28

stars

98

commits

2

linked in READMEs

Apr 2, 2024

updated

README

TCEval v2

TCEval-v2 is a Traditional Chinese evaluation suite for foundation models derived from TCEval-v1. It covers 5 capabilities, including contextual QA, knowledge, classification, and table understanding.

Benchmark

  • Contextual QA
    • drcd : DRCD is a Traditional Chinese machine reading comprehension dataset containing 10,014 paragraphs from 2,108 Wikipedia articles and over 30,000 questions.
  • Knowledge
    • tmmluplus (provided by MediaTek Research and iKala): Taiwan Massive Multitask Language Understanding + (TMMLU+) is curated from examinations in Taiwan, consisting of 67 subjects spanning across multiple disciplines, from vocational to academic fields, and covering elementary to professional proficiency levels. It is designed to identify a model’s knowledge and problem-solving blind spots similar to human evaluations. It is categorized into STEM, humanties, social sciences and other (similar to MMLU), for a higher level overview of the model capabilities.
  • Table Understanding
    • penguin_table (translate from a subset of BIG-Bench): The “penguins in a table” task contained in BIG-bench asks a language model to answer questions about the animals contained in a table, or multiple tables, described in the context.
  • Chat and instruction following
    • mt_bench_tw (translated from MT Bench): MT-Bench-TW is a Traditional Chinese version of MT-bench, which is a series of open-ended questions that evaluate a chatbot’s multi-turn conversational and instruction-following ability. MT-Bench-TW inherits the categorization of MT-Bench, which includes a wide variety of core capabilities, such as reasoning and writing.

If you find the dataset useful in your work, please cite:

@misc{hsu2023advancing,
    title={Advancing the Evaluation of Traditional Chinese Language Models: Towards a Comprehensive Benchmark Suite}, 
    author={Chan-Jan Hsu and Chang-Le Liu and Feng-Ting Liao and Po-Chun Hsu and Yi-Chang Chen and Da-shan Shiu},
    year={2023},
    eprint={2309.08448},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

Contributors

YC-Chen

93 commits

cllatMTK

2 commits

Splend1dchan

2 commits

pochunhsu

1 commits

MediaTek-Research/TCEval-v2

Dataset

28

stars

98

commits

2

linked in READMEs

Apr 2, 2024

updated

README

TCEval v2

TCEval-v2 is a Traditional Chinese evaluation suite for foundation models derived from TCEval-v1. It covers 5 capabilities, including contextual QA, knowledge, classification, and table understanding.

Benchmark

  • Contextual QA
    • drcd : DRCD is a Traditional Chinese machine reading comprehension dataset containing 10,014 paragraphs from 2,108 Wikipedia articles and over 30,000 questions.
  • Knowledge
    • tmmluplus (provided by MediaTek Research and iKala): Taiwan Massive Multitask Language Understanding + (TMMLU+) is curated from examinations in Taiwan, consisting of 67 subjects spanning across multiple disciplines, from vocational to academic fields, and covering elementary to professional proficiency levels. It is designed to identify a model’s knowledge and problem-solving blind spots similar to human evaluations. It is categorized into STEM, humanties, social sciences and other (similar to MMLU), for a higher level overview of the model capabilities.
  • Table Understanding
    • penguin_table (translate from a subset of BIG-Bench): The “penguins in a table” task contained in BIG-bench asks a language model to answer questions about the animals contained in a table, or multiple tables, described in the context.
  • Chat and instruction following
    • mt_bench_tw (translated from MT Bench): MT-Bench-TW is a Traditional Chinese version of MT-bench, which is a series of open-ended questions that evaluate a chatbot’s multi-turn conversational and instruction-following ability. MT-Bench-TW inherits the categorization of MT-Bench, which includes a wide variety of core capabilities, such as reasoning and writing.

If you find the dataset useful in your work, please cite:

@misc{hsu2023advancing,
    title={Advancing the Evaluation of Traditional Chinese Language Models: Towards a Comprehensive Benchmark Suite}, 
    author={Chan-Jan Hsu and Chang-Le Liu and Feng-Ting Liao and Po-Chun Hsu and Yi-Chang Chen and Da-shan Shiu},
    year={2023},
    eprint={2309.08448},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}

Contributors

YC-Chen

93 commits

cllatMTK

2 commits

Splend1dchan

2 commits

pochunhsu

1 commits