Arxiv: Arxiv | Code: Open-PMC Github | Model Checkpoint: Hugging Face
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
This dataset is designed for research in:
The dataset primarily contains text in English.
Each record in the dataset contains:
The dataset does not contain predefined splits. Users can split the data as needed for training, validation, and testing.
The dataset does not contain additional manual annotations.
This dataset is designed for research purposes only and should not be used for:
If you find the code useful for your research, please consider citing
@article{baghbanzadeh2025open,
title={Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning},
author={Baghbanzadeh, Negin and Ashkezari, Sajad and Dolatabadi, Elham and Afkanpour, Arash},
journal={arXiv preprint arXiv:2506.02738},
year={2025}
}
This dataset is licensed under CC-BY-4.0, meaning it can be used for research purposes with appropriate attribution.
500 commits
Arxiv: Arxiv | Code: Open-PMC Github | Model Checkpoint: Hugging Face
This dataset consists of image-text pairs extracted from medical papers available on PubMed Central. It has been curated to support research in medical image understanding, particularly in natural language processing (NLP) and computer vision tasks related to medical imagery. The dataset includes:
This dataset is designed for research in:
The dataset primarily contains text in English.
Each record in the dataset contains:
The dataset does not contain predefined splits. Users can split the data as needed for training, validation, and testing.
The dataset does not contain additional manual annotations.
This dataset is designed for research purposes only and should not be used for:
If you find the code useful for your research, please consider citing
@article{baghbanzadeh2025open,
title={Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning},
author={Baghbanzadeh, Negin and Ashkezari, Sajad and Dolatabadi, Elham and Afkanpour, Arash},
journal={arXiv preprint arXiv:2506.02738},
year={2025}
}
This dataset is licensed under CC-BY-4.0, meaning it can be used for research purposes with appropriate attribution.
500 commits