Dataflex-selection Dataset (Alpaca-format)
0
8 commits
1 linked in READMEs
updated Nov 25, 2025
This dataset contains two splits:
A high-quality instruction–response dataset in Alpaca format, designed for training and fine-tuning large language models on instruction-following tasks. The dataset includes structured triplets:
The validation split is derived from the MMLU benchmark and has been further enhanced with GPT-5–generated Chain-of-Thought (CoT) reasoning. Each sample contains the full question, the correct answer, and a model-generated prediction, enabling detailed evaluation and error analysis.
This split can be used for automated evaluation, benchmarking, data selection experiments (e.g., TSDS / NEAR), and studying reasoning failure cases.
Each validation sample contains the following fields:
"A. ...") "A")Dataflex-selection Dataset (Alpaca-format)
0
8 commits
1 linked in READMEs
updated Nov 25, 2025
This dataset contains two splits:
A high-quality instruction–response dataset in Alpaca format, designed for training and fine-tuning large language models on instruction-following tasks. The dataset includes structured triplets:
The validation split is derived from the MMLU benchmark and has been further enhanced with GPT-5–generated Chain-of-Thought (CoT) reasoning. Each sample contains the full question, the correct answer, and a model-generated prediction, enabling detailed evaluation and error analysis.
This split can be used for automated evaluation, benchmarking, data selection experiments (e.g., TSDS / NEAR), and studying reasoning failure cases.
Each validation sample contains the following fields:
"A. ...") "A")