ingoziegler/CRAFT-MedQA

Dataset

CRAFT-MedQA

1

11 commits

1 linked in READMEs

updated Dec 8, 2025

See the code

README

CRAFT-MedQA

This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.

The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated medicine question-answering data proved highly beneficial.

MedQA Performance

4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent improvement that enable non-instruction-tuned models to outperform general instruction-tuned LLMs.

To use only our human-written few-shot examples, XS(8) or XS(32), filter Column 4 is_few_shot == 1.

We release our LoRA adapters that are fine-tuned on the XL dataset version for the Mistral 7B v0.2 architecture here.

Dataset Format

The dataset is a collection of multiple-choice questions with corresponding options and answers. There are always 4 answer options provided, of which a single option is correct. Each sample in the dataset is represented as a single row in a table, with four columns:

Column 1: question

  • Data Type: String
  • Description: The question being asked. This column contains the text of the question.

Column 2: options

  • Data Type: List of Strings
  • Description: The possible answer options for the question. This column contains a list of strings, where each string represents a possible answer choice.

Column 3: answer

  • Data Type: String
  • Description: The correct answer to the question. This column contains a single letter string, which corresponds to one of the options listed in Column 2.

Column 4: is_few_shot

  • Data Type: Integer
  • Description: A flag indicating whether the question is a human-written few-shot example. This column contains a binary value (0 or 1), where 0 indicates that the question is not a few-shot example, and 1 indicates that it is.

Example: A sample has the following layout:

"question": "During a laparoscopic appendectomy, what is inserted into the abdomen through an incision to allow the introduction of the laparoscope?"
"options": ["A. A trocar and harmless gas", "B. A tube for draining an abscess", "C. A surgical instrument for removing the appendix", "D. The laparoscope itself"]
"answer": "A"
"is_few_shot": 0

Citation

If you use our code, datasets, or model checkpoints in your research, please cite the following paper:

@article{ziegler2025craft,
    author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
    title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
    journal = {Transactions of the Association for Computational Linguistics},
    volume = {13},
    pages = {1693-1721},
    year = {2025},
    month = {12},
    issn = {2307-387X},
    doi = {10.1162/TACL.a.56},
    url = {https://doi.org/10.1162/TACL.a.56},
    eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}
biology
medical
synthetic

ingoziegler/CRAFT-MedQA

Dataset

CRAFT-MedQA

1

11 commits

1 linked in READMEs

updated Dec 8, 2025

See the code

README

CRAFT-MedQA

This is a synthetic dataset generated with the CRAFT framework proposed in the paper CRAFT Your Dataset: Task-Specific Synthetic Data Generation Through Corpus Retrieval and Augmentation.

The correctness of the data has not been verified in detail, but training on this data and evaluating on human-curated medicine question-answering data proved highly beneficial.

MedQA Performance

4 synthetic dataset sizes (S, M, L, XL) are available, and training on them yields consistent improvement that enable non-instruction-tuned models to outperform general instruction-tuned LLMs.

To use only our human-written few-shot examples, XS(8) or XS(32), filter Column 4 is_few_shot == 1.

We release our LoRA adapters that are fine-tuned on the XL dataset version for the Mistral 7B v0.2 architecture here.

Dataset Format

The dataset is a collection of multiple-choice questions with corresponding options and answers. There are always 4 answer options provided, of which a single option is correct. Each sample in the dataset is represented as a single row in a table, with four columns:

Column 1: question

  • Data Type: String
  • Description: The question being asked. This column contains the text of the question.

Column 2: options

  • Data Type: List of Strings
  • Description: The possible answer options for the question. This column contains a list of strings, where each string represents a possible answer choice.

Column 3: answer

  • Data Type: String
  • Description: The correct answer to the question. This column contains a single letter string, which corresponds to one of the options listed in Column 2.

Column 4: is_few_shot

  • Data Type: Integer
  • Description: A flag indicating whether the question is a human-written few-shot example. This column contains a binary value (0 or 1), where 0 indicates that the question is not a few-shot example, and 1 indicates that it is.

Example: A sample has the following layout:

"question": "During a laparoscopic appendectomy, what is inserted into the abdomen through an incision to allow the introduction of the laparoscope?"
"options": ["A. A trocar and harmless gas", "B. A tube for draining an abscess", "C. A surgical instrument for removing the appendix", "D. The laparoscope itself"]
"answer": "A"
"is_few_shot": 0

Citation

If you use our code, datasets, or model checkpoints in your research, please cite the following paper:

@article{ziegler2025craft,
    author={Ziegler, Ingo and K{\"o}ksal, Abdullatif and Elliott, Desmond and Sch{\"u}tze, Hinrich},
    title = {CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation},
    journal = {Transactions of the Association for Computational Linguistics},
    volume = {13},
    pages = {1693-1721},
    year = {2025},
    month = {12},
    issn = {2307-387X},
    doi = {10.1162/TACL.a.56},
    url = {https://doi.org/10.1162/TACL.a.56},
    eprint = {https://direct.mit.edu/tacl/article-pdf/doi/10.1162/TACL.a.56/2568491/tacl.a.56.pdf},
}
biology
medical
synthetic