RealMedQA is a biomedical question answering dataset consisting of realistic question and answer pairs. The questions were created by medical students and a large language model (LLM), while the answers are guideline recommendations provided by the UK's National Institute for Health and Care Excellence (NICE). The full paper describing the dataset and the experiments has been accepted to the American Medical Informatics Association (AMIA) Annual Symposium and is available here.
Initially, 12,543 guidelines were retrieved using the NICE syndication API. As we were interested in only the guidelines that pertain to clinical practice, we only used the guidelines that came under 'Conditions and diseases' which reduced the number to 7,385.
We created an instruction sheet with examples which we provided to both the humans (medical students) and the LLM to generate the several questions for each guideline recommendation. The instruction sheet was fed as a prompt along with each recommendation to the LLM, while the humans created the questions using Google forms.
Both the QA pairs generated by the LLM and those generated by human annotators were verified by humans for quality. The verifiers were asked whether each question:
A total of 800 human QA pairs and 400 LLM QA pairs were verified.
The dataset is structured according to the following columns:
If you use the dataset, please cite our work using the following reference:
@misc{kell2024realmedqapilotbiomedicalquestion,
title={RealMedQA: A pilot biomedical question answering dataset containing realistic clinical questions},
author={Gregory Kell and Angus Roberts and Serge Umansky and Yuti Khare and Najma Ahmed and Nikhil Patel and Chloe Simela and Jack Coumbe and Julian Rozario and Ryan-Rhys Griffiths and Iain J. Marshall},
year={2024},
eprint={2408.08624},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2408.08624},
}
RealMedQA is a biomedical question answering dataset consisting of realistic question and answer pairs. The questions were created by medical students and a large language model (LLM), while the answers are guideline recommendations provided by the UK's National Institute for Health and Care Excellence (NICE). The full paper describing the dataset and the experiments has been accepted to the American Medical Informatics Association (AMIA) Annual Symposium and is available here.
Initially, 12,543 guidelines were retrieved using the NICE syndication API. As we were interested in only the guidelines that pertain to clinical practice, we only used the guidelines that came under 'Conditions and diseases' which reduced the number to 7,385.
We created an instruction sheet with examples which we provided to both the humans (medical students) and the LLM to generate the several questions for each guideline recommendation. The instruction sheet was fed as a prompt along with each recommendation to the LLM, while the humans created the questions using Google forms.
Both the QA pairs generated by the LLM and those generated by human annotators were verified by humans for quality. The verifiers were asked whether each question:
A total of 800 human QA pairs and 400 LLM QA pairs were verified.
The dataset is structured according to the following columns:
If you use the dataset, please cite our work using the following reference:
@misc{kell2024realmedqapilotbiomedicalquestion,
title={RealMedQA: A pilot biomedical question answering dataset containing realistic clinical questions},
author={Gregory Kell and Angus Roberts and Serge Umansky and Yuti Khare and Najma Ahmed and Nikhil Patel and Chloe Simela and Jack Coumbe and Julian Rozario and Ryan-Rhys Griffiths and Iain J. Marshall},
year={2024},
eprint={2408.08624},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2408.08624},
}