[TACL, EMNLP 2025 Oral] Code, datasets, and checkpoints for the paper "CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation"
See the code
This repository contains the code, datasets, and additionally required files for the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
We make all size variations of our crafted datasets available on Hugging Face:
To use our human-written few-shots, simply filter the dataset for is_few_shot == 1, or load the .jsonl from assets/{task}/few-shot/corpus-task-32.jsonl.
Our 8 few-shot-based experiments simply use the first 8 lines of each file.
Models trained on our synthetic datasets match the performance of general instruction-tuned LLMs and can even outperform training on human-curated data for tasks like summarization.

Our synthesized data is also more robust against distribution shifts because we do not generate data for a specific test set, but for an overall task. This can be seen from the 5-gram overlap between our crafted datasets and the test sets, and comparing it to the in-domin train sets. At maximum, our synthetic datasets have a 0.4% overlap, while in-domain train sets range up to 17.9% 5-gram overlap (see table below).
| BioQA | CSQA | MedQA | Summarization | |
|---|---|---|---|---|
| CRAFTXS | 0.0% | 0.0% | 0.0% | 0.0% |
| CRAFTS | 0.0% | 0.1% | 0.1% | 0.0% |
| CRAFTM | 0.0% | 0.2% | 0.1% | 0.1% |
| CRAFTL | 0.0% | 0.4% | 0.3% | 0.2% |
| CRAFTXL | 0.0% | 0.2% | 0.2% | 0.2% |
| Baseline (In-domain Train Set) | 17.9% | 4.4% | 1.1% | 0.3% |
Consequently, CRAFT's performance is stronger and more consistent across other test sets, e.g., MMLU.
| Dataset | Baseline | CRAFTXL |
|---|---|---|
| In-domain | 89.9 | 78.1 |
| MMLUMedical Genetics | 60.0 | 69.0 |
| MMLUAnatomy | 55.6 | 57.0 |
| MMLUHigh School Biology | 69.3 | 67.4 |
| MMLUCollege Biology | 66.7 | 74.3 |
| MMLU-Avg | 62.9 | 66.9 |
Here, we provide the download links for the adapter checkpoints resulting from our fine-tuning on the CRAFT-XL versions.
Our experiments are based around Python 3.10.9, Pytorch 2.2.1, and vllm 0.4.1. For details, check requirements.txt.
The pipeline consists of 5 steps that have to be performed sequentially.
In general, you have to follow the same steps as if you were reproducing our experiments as described below.
You can either use our embedding database and corpora mentioned under Step 0 below, or you can extend our embedding database with your public/private corpora, or you can create your own specialized embedding database with corresponding corpora using your private or other public datasets.
Currently, our code only features the experiments from our paper ready, so you will need to adapt the scripts and run configs a bit.
You can still use large parts from our run configs from code/run_configs/ as the baseline, but you will need to change the paths pointing to our databases and corpora files.
Additionally, when running finetuning and evaluation, you will need to write code for your task sample design, as well as provide your evaluation dataset and format it accordingly.
Nonetheless, the general structure stays the same.
To create an embedding database, we provide the files we used to embed our corpora and create the database.
python3 code/create_embeddings.py $(code/run_configs/embed/stackexchange.cfg).h5 database with 16-bit precision NumPy arrays using multi-qa-MiniLM-L6-cos-v1 from the SentenceTransformer suite as the embedding model.assets/{task}/few-shot/corpus-task-32.jsonl
If you have any questions, feel free to open a GitHub issue.
Have a look at code/utils/args.py for all available runtime arguments.
We provide our pre-filled argparse run configs for all experiments under code/run_configs/.
The few-shots for all tasks are also available in assets/{task}/few-shot/corpus-task-32.jsonl, so you can start running/reproducing our experiments.
datasets/embeddings.h5
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumen version of C4 from Hugging Face. We used the Git download version. It is not mentioned there that you have to run git lfs checkout after everything is downloaded so that the lazy files are actually linked to the downloaded files.datasets/wikipedia/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumdatasets/wikihow/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumdatasets/stackexchange/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumassets/ has the following subfolder available: assets/{task}/corpus_samples/, assets/{task}/outputs/, assets/{task}/results/, assets/{task}/task_samples/model_ckpts directory. All LoRA adapters will be saved heremodels/hf_models directory and place the model you want to use for task sample creation in there (e.g. Mistral 7B Instruct v0.2), as well as the model you want to fine-tune (e.g. Mistral 7B v0.2), and the model you want to evaluate again (e.g. Mistral 7B Instruct v0.2, too). If you evaluate a generative task, such as summarization, you should also load LLaMA 3 70B Instruct into this directory.hf-llama.privkey and place it in the root directorymodel_ckpts/{taks}/mistral-7b-v0.2-32-mixed-25k-seed_1234/ (check the specific seed in the Hugging Face repository description)python3 code/retrieve_docs.py $(cat code/run_configs/{task}/retrieve/32-mixed-50k.cfg) or any other size of documents you want to retrieveassets/{task}/corpus_samples/32-mixed-50.jsonlpython3 code/create_task_samples.py $(cat code/run_configs/{task}/create_tasks/mistral-7b-instruct-v0.2-32-mixed-50k.cfg)assets/{task}/task_samples/32-mixed-25000/ as .arrow files
.arrow) versions of task samples under assets/{task}/task_samples/32-mixed-50k-raw.jsonl, assets/{task}/task_samples/32-mixed-50k-clean.jsonl, assets/{task}/task_samples/32-mixed-50k-error_msgs.csv.stderr outputs.python3 code/run_finetune.py $(cat code/run_configs/{task}/finetune/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) (or seed 2024, 9999)model_ckpts/{taks}/mistral-7b-v0.2-32-mixed-25k-seed_1234/
python3 code/run_eval.py $(cat code/run_configs/{task}/evaluate/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) (or seed 2024, 9999)'bioqa' was chosen as the task).assets/{task}/outputs/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonlassets/{task}/results/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonpython3 code/run_eval_winrate.py $(cat code/run_configs/{task}/evaluate/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) to calculate the win rate of the model outputs (saved under assets/{task}/outputs/) and the evaluation dataset answers with LLaMA 3 70B Instruct as the judgeassets/{task}/results/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonYour are done! Should you have any questions, feel free to open a GitHub issue.
If you use our code, datasets, or model checkpoints in your research, please cite the following paper:
@article{ziegler2025craft,
author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
pages = {1693-1721},
year = {2025},
month = {12},
issn = {2307-387X},
doi = {10.1162/TACL.a.56},
url = {https://doi.org/10.1162/TACL.a.56},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}
378 followers · starred Sep 2024
415 followers · starred Sep 2024
Python
100.0%
[TACL, EMNLP 2025 Oral] Code, datasets, and checkpoints for the paper "CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation"
See the code
This repository contains the code, datasets, and additionally required files for the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.
We make all size variations of our crafted datasets available on Hugging Face:
To use our human-written few-shots, simply filter the dataset for is_few_shot == 1, or load the .jsonl from assets/{task}/few-shot/corpus-task-32.jsonl.
Our 8 few-shot-based experiments simply use the first 8 lines of each file.
Models trained on our synthetic datasets match the performance of general instruction-tuned LLMs and can even outperform training on human-curated data for tasks like summarization.

Our synthesized data is also more robust against distribution shifts because we do not generate data for a specific test set, but for an overall task. This can be seen from the 5-gram overlap between our crafted datasets and the test sets, and comparing it to the in-domin train sets. At maximum, our synthetic datasets have a 0.4% overlap, while in-domain train sets range up to 17.9% 5-gram overlap (see table below).
| BioQA | CSQA | MedQA | Summarization | |
|---|---|---|---|---|
| CRAFTXS | 0.0% | 0.0% | 0.0% | 0.0% |
| CRAFTS | 0.0% | 0.1% | 0.1% | 0.0% |
| CRAFTM | 0.0% | 0.2% | 0.1% | 0.1% |
| CRAFTL | 0.0% | 0.4% | 0.3% | 0.2% |
| CRAFTXL | 0.0% | 0.2% | 0.2% | 0.2% |
| Baseline (In-domain Train Set) | 17.9% | 4.4% | 1.1% | 0.3% |
Consequently, CRAFT's performance is stronger and more consistent across other test sets, e.g., MMLU.
| Dataset | Baseline | CRAFTXL |
|---|---|---|
| In-domain | 89.9 | 78.1 |
| MMLUMedical Genetics | 60.0 | 69.0 |
| MMLUAnatomy | 55.6 | 57.0 |
| MMLUHigh School Biology | 69.3 | 67.4 |
| MMLUCollege Biology | 66.7 | 74.3 |
| MMLU-Avg | 62.9 | 66.9 |
Here, we provide the download links for the adapter checkpoints resulting from our fine-tuning on the CRAFT-XL versions.
Our experiments are based around Python 3.10.9, Pytorch 2.2.1, and vllm 0.4.1. For details, check requirements.txt.
The pipeline consists of 5 steps that have to be performed sequentially.
In general, you have to follow the same steps as if you were reproducing our experiments as described below.
You can either use our embedding database and corpora mentioned under Step 0 below, or you can extend our embedding database with your public/private corpora, or you can create your own specialized embedding database with corresponding corpora using your private or other public datasets.
Currently, our code only features the experiments from our paper ready, so you will need to adapt the scripts and run configs a bit.
You can still use large parts from our run configs from code/run_configs/ as the baseline, but you will need to change the paths pointing to our databases and corpora files.
Additionally, when running finetuning and evaluation, you will need to write code for your task sample design, as well as provide your evaluation dataset and format it accordingly.
Nonetheless, the general structure stays the same.
To create an embedding database, we provide the files we used to embed our corpora and create the database.
python3 code/create_embeddings.py $(code/run_configs/embed/stackexchange.cfg).h5 database with 16-bit precision NumPy arrays using multi-qa-MiniLM-L6-cos-v1 from the SentenceTransformer suite as the embedding model.assets/{task}/few-shot/corpus-task-32.jsonl
If you have any questions, feel free to open a GitHub issue.
Have a look at code/utils/args.py for all available runtime arguments.
We provide our pre-filled argparse run configs for all experiments under code/run_configs/.
The few-shots for all tasks are also available in assets/{task}/few-shot/corpus-task-32.jsonl, so you can start running/reproducing our experiments.
datasets/embeddings.h5
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumen version of C4 from Hugging Face. We used the Git download version. It is not mentioned there that you have to run git lfs checkout after everything is downloaded so that the lazy files are actually linked to the downloaded files.datasets/wikipedia/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumdatasets/wikihow/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumdatasets/stackexchange/cleaned/
http address. If your browser autocompletes to https, you may need to manually adjust the link. Sometimes, you may also have to paste the address into the search bar directly.checksumassets/ has the following subfolder available: assets/{task}/corpus_samples/, assets/{task}/outputs/, assets/{task}/results/, assets/{task}/task_samples/model_ckpts directory. All LoRA adapters will be saved heremodels/hf_models directory and place the model you want to use for task sample creation in there (e.g. Mistral 7B Instruct v0.2), as well as the model you want to fine-tune (e.g. Mistral 7B v0.2), and the model you want to evaluate again (e.g. Mistral 7B Instruct v0.2, too). If you evaluate a generative task, such as summarization, you should also load LLaMA 3 70B Instruct into this directory.hf-llama.privkey and place it in the root directorymodel_ckpts/{taks}/mistral-7b-v0.2-32-mixed-25k-seed_1234/ (check the specific seed in the Hugging Face repository description)python3 code/retrieve_docs.py $(cat code/run_configs/{task}/retrieve/32-mixed-50k.cfg) or any other size of documents you want to retrieveassets/{task}/corpus_samples/32-mixed-50.jsonlpython3 code/create_task_samples.py $(cat code/run_configs/{task}/create_tasks/mistral-7b-instruct-v0.2-32-mixed-50k.cfg)assets/{task}/task_samples/32-mixed-25000/ as .arrow files
.arrow) versions of task samples under assets/{task}/task_samples/32-mixed-50k-raw.jsonl, assets/{task}/task_samples/32-mixed-50k-clean.jsonl, assets/{task}/task_samples/32-mixed-50k-error_msgs.csv.stderr outputs.python3 code/run_finetune.py $(cat code/run_configs/{task}/finetune/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) (or seed 2024, 9999)model_ckpts/{taks}/mistral-7b-v0.2-32-mixed-25k-seed_1234/
python3 code/run_eval.py $(cat code/run_configs/{task}/evaluate/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) (or seed 2024, 9999)'bioqa' was chosen as the task).assets/{task}/outputs/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonlassets/{task}/results/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonpython3 code/run_eval_winrate.py $(cat code/run_configs/{task}/evaluate/lora/mistral-7b-v0.2-32-mixed-25k-seed_1234.cfg) to calculate the win rate of the model outputs (saved under assets/{task}/outputs/) and the evaluation dataset answers with LLaMA 3 70B Instruct as the judgeassets/{task}/results/lora/mistral-7b-v0.2-lora-32-mixed-25k-seed_1234.jsonYour are done! Should you have any questions, feel free to open a GitHub issue.
If you use our code, datasets, or model checkpoints in your research, please cite the following paper:
@article{ziegler2025craft,
author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
journal = {Transactions of the Association for Computational Linguistics},
volume = {13},
pages = {1693-1721},
year = {2025},
month = {12},
issn = {2307-387X},
doi = {10.1162/TACL.a.56},
url = {https://doi.org/10.1162/TACL.a.56},
eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}
378 followers · starred Sep 2024
415 followers · starred Sep 2024
Python
100.0%