This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated commonsense question-answering data proved highly beneficial.

4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent improvement that enable non-instruction-tuned models to outperform or match general instruction-tuned LLMs.
To use only our human-written few-shot examples, XS(8) or XS(32), filter Column 4 is_few_shot == 1.
We release our LoRA adapters that are fine-tuned on the XL dataset version for the Mistral 7B v0.2 architecture here.
The dataset is a collection of multiple-choice questions with corresponding options and answers. There are always 2 answer options provided (yes or no), of which a single option is correct. Each sample in the dataset is represented as a single row in a table, with four columns:
Column 1: question
Column 2: options
Column 3: answer
Column 4: is_few_shot
Example: A sample has the following layout:
"question": "Does exposure to blue lights from computers and phones help promote sleep?"
"options": ["A. Yes", "B. No"]
"answer": "B"
"is_few_shot": 0
If you use our code, datasets, or model checkpoints in your research, please cite the following paper:
@article{ziegler2025craft,
author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
pages = {1693-1721},
year = {2025},
month = {12},
issn = {2307-387X},
doi = {10.1162/TACL.a.56},
url = {https://doi.org/10.1162/TACL.a.56},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}
This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated commonsense question-answering data proved highly beneficial.

4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent improvement that enable non-instruction-tuned models to outperform or match general instruction-tuned LLMs.
To use only our human-written few-shot examples, XS(8) or XS(32), filter Column 4 is_few_shot == 1.
We release our LoRA adapters that are fine-tuned on the XL dataset version for the Mistral 7B v0.2 architecture here.
The dataset is a collection of multiple-choice questions with corresponding options and answers. There are always 2 answer options provided (yes or no), of which a single option is correct. Each sample in the dataset is represented as a single row in a table, with four columns:
Column 1: question
Column 2: options
Column 3: answer
Column 4: is_few_shot
Example: A sample has the following layout:
"question": "Does exposure to blue lights from computers and phones help promote sleep?"
"options": ["A. Yes", "B. No"]
"answer": "B"
"is_few_shot": 0
If you use our code, datasets, or model checkpoints in your research, please cite the following paper:
@article{ziegler2025craft,
author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
pages = {1693-1721},
year = {2025},
month = {12},
issn = {2307-387X},
doi = {10.1162/TACL.a.56},
url = {https://doi.org/10.1162/TACL.a.56},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}