This repository contains the data and annotations described in the papers:
The dataset consists of 57 mock medical primary care consultations held over 5 days by 7 Babylon clinicians and 57 Babylon employees acting as patients, using case cards with presenting complaints, symptoms, medical & general history etc. The data in this repository includes:
audio folder);transcripts folder);notes folder);human_eval_data folder).The scripts folder includes some data transformation scripts
(utterance extraction, transcript collation etc.)
More detailed descriptions are found in each folder's README.md files.
Due to their size, the audio files are stored using Git Large File Storage (https://git-lfs.github.com/). To clone the repository:
brew install git-lfsgit lfs installgit clone https://github.com/babylonhealth/primock57.git@inproceedings{korfiatis2022primock57,
title={(in press): PriMock57: A Dataset Of Primary Care Mock Consultations},
author={Papadopoulos Korfiatis, Alex and Moramarco, Francesco and Sarac, Radmila and Savkov, Aleksandar},
booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics},
year={2022}
}
@inproceedings{moramarco2022human,
title={(In press): Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation},
author={Moramarco, Francesco and Papadopoulos Korfiatis, Alex and Perera, Mark and Juric, Damir and Flann, Jack and Reiter, Ehud and Belz, Anya and Savkov, Aleksandar},
booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics},
year={2022}
}
5 commits
This repository contains the data and annotations described in the papers:
The dataset consists of 57 mock medical primary care consultations held over 5 days by 7 Babylon clinicians and 57 Babylon employees acting as patients, using case cards with presenting complaints, symptoms, medical & general history etc. The data in this repository includes:
audio folder);transcripts folder);notes folder);human_eval_data folder).The scripts folder includes some data transformation scripts
(utterance extraction, transcript collation etc.)
More detailed descriptions are found in each folder's README.md files.
Due to their size, the audio files are stored using Git Large File Storage (https://git-lfs.github.com/). To clone the repository:
brew install git-lfsgit lfs installgit clone https://github.com/babylonhealth/primock57.git@inproceedings{korfiatis2022primock57,
title={(in press): PriMock57: A Dataset Of Primary Care Mock Consultations},
author={Papadopoulos Korfiatis, Alex and Moramarco, Francesco and Sarac, Radmila and Savkov, Aleksandar},
booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics},
year={2022}
}
@inproceedings{moramarco2022human,
title={(In press): Human Evaluation and Correlation with Automatic Metrics in Consultation Note Generation},
author={Moramarco, Francesco and Papadopoulos Korfiatis, Alex and Perera, Mark and Juric, Damir and Flann, Jack and Reiter, Ehud and Belz, Anya and Savkov, Aleksandar},
booktitle={Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics},
year={2022}
}
5 commits
Python
93.7%
Shell
6.3%