Yunkun-Zhang/Data-Centric-FM-Healthcare

A survey on data-centric foundation models in healthcare.

78

28 commits

updated Feb 14, 2026

See the code

README

Data-Centric Foundation Models in Computational Healthcare

:fire::fire::fire: A survey on data-centric foundation models in computational healthcare

Project Page | Paper [arXiv]

Last updated: 2026/02/05

:pencil: If you find this repo helps, please kindly cite our survey, thanks!

@article{zhang2024data,
  title={Data-Centric Foundation Models in Computational Healthcare: A Survey},
  author={Zhang, Yunkun and Gao, Jin and Tan, Zheling and Zhou, Lingfeng and Ding, Kexin and Zhou, Mu and Zhang, Shaoting and Wang, Dequan},
  journal={arXiv},
  year={2024},
  eprint={2401.02458},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  doi={10.48550/arXiv.2401.02458},
  url={https://arxiv.org/abs/2401.02458}
}

In this repository, we provide an up-to-date list of healthcare-related foundation models and datasets, which are also mentioned in our survey paper.

:book: Contents


Healthcare and Medical Foundation Models

A star (*) after the pre-training data shows that the authors constructed the data with more than three sources.

Language Models

ModelSubfieldPaperCodeBasePre-Training Data
Baichuan-M2MedicineBaichuan-M2: Scaling Medical Capability with Large Verifier SystemGithubQwen2.5*
Baichuan-M1MedicineBaichuan-M1: Pushing the Medical Capability of Large Language Models-Transformer20T tokens*
EHRMambaClinicEHRMamba: Towards Generalizable and Scalable Foundation Models for Electronic Health RecordsGithubMambaMIMIC-IV
MMedLM 2MedicineTowards building multilingual language model for medicineGithubInternLM 2MMedC*
BiMediXMedicineBiMediX: Bilingual Medical Mixture of Experts LLMGithubMixtralBiMed1.3M*
Me LLaMAMedicineMe LLaMA: Foundation Large Language Models for Medical ApplicationsGithubLLaMA 2*
BioMistralBiomedicineBioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains-MistralPubMed Central
PULSEMedicine-GithubInternLM*
MeditronMedicineMEDITRON-70B: Scaling Medical Pretraining for Large Language ModelsGithubLLaMA 2GAP-Replay*
TaiyiBiomedicineTaiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical TasksGithubQwen-7B / GLM4-9BBigBio + CBLUE
BioMedGPTBiomedicineBioMedGPT: An Open Multimodal Large Language Model for BioMedicineGithubLLaMA 2S2ORC
Clinical LLaMA-LoRAClinicParameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain-LLaMAMIMIC-IV
Med-PaLM 2ClinicToward expert-level medical question answering with large language modelsGooglePaLM 2MedQA
PMC-LLaMAMedicinePMC-LLaMA: toward building open-source language models for medicineGithubLLaMAMedC
MedAlpacaMedicineMedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training DataGithubLLaMAMedical Meadow
BenTsao (HuaTuo)BiomedicineHuaTuo: Tuning LLaMA Model with Chinese Medical KnowledgeGithubLLaMACMeKG
ChatDoctorMedicineChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain KnowledgeGithubLLaMAHealthCareMagic*
Clinical-T5ClinicClinical-T5: Large Language Models Built Using Mimic Clinical TextPhysioNetT5MIMIC-III + MIMIC-IV
Med-PaLMClinicLarge Language Models Encode Clinical KnowledgeGooglePaLMMedQA
BioGPTBiomedicineBioGPT: Generative Pre-Trained Transformer for Biomedical Text Generation and MiningGithubGPT-2PubMed
BioLinkBERTBiomedicineLinkBERT: Pretraining Language Models with Document LinksGithubBERTPubMed (citation links)
PubMedBERTBiomedicineDomain-Specific Language Model Pretraining for Biomedical Natural Language ProcessingMicrosoftBERTPubMed
BioBERTBiomedicineBioBERT: A Pre-Trained Biomedical Language Representation Model for Biomedical Text MiningGithubBERTPubMed + PMC
BlueBERTBiomedicineAn Empirical Study of Multi-Task Learning on BERT for Biomedical Text MiningGithubBERTPubMed + MIMIC-III
Clinical BERTClinicPublicly Available Clinical BERT EmbeddingsGithubBERTMIMIC-III
SciBERTBiomedicineSciBERT: A Pretrained Language Model for Scientific TextGithubBERTSemantic Scholar

Vision Models

ModelSubfieldPaperCodeBasePre-Training Data
FastGliomaPathologyFoundation models for fast, label-free detection of glioma infiltration--Label-free optical microscopy (4M images)*
MedLSAMRadiologyMedLSAM: Localize and Segment Anything Model for 3D CT ImagesGithubSAM*
BiomedParseBiomedicineA Foundation Model for Joint Segmentation, Detection and Recognition of Biomedical Objects across Nine ModalitiesGithubSEEMBiomedParseData*
Universal ModelRadiologyUniversal and Extensible Language-Vision Models for Organ Segmentation and Tumor Detection from Abdominal Computed TomographyGithub-*
CHIEFPathologyA pathology foundation model for cancer diagnosis and prognosis predictionGithubCTransPath*
USFMSonographyUSFM: A Universal Ultrasound Foundation Model Generalized to Tasks and Organs towards Label Efficient Image Analysis-MIM3M-US*
BrainSegFounderRadiologyBrainSegFounder: towards 3D foundation models for neuroimage segmentationGithubSwinUNETRUK Biobank + BraTS + ATLAS
MedSAMMedicineSegment Anything in Medical ImagesGithubSAM*
Prov-GigaPathPathologyA Whole-Slide Foundation Model for Digital Pathology from Real-World DataGithub-Prov-Path*
BEPHPathologyA Foundation Model for Generalizable Cancer Diagnosis and Survival Prediction from Histopathological ImagesGithubBEiTv2*
Pai et al.RadiologyFoundation Model for Cancer Imaging BiomarkersGithubSimCLR*
VIS-MAERadiologyVIS-MAE: An Efficient Self-supervised Learning Approach on Medical Image Segmentation and Classification-MAE*
SegmentAnyBoneRadiologySegmentAnyBone: A Universal Model that Segments Any Bone at Any Location on MRIGithubSAM*
RudolfVPathologyRudolfV: A Foundation Model by Pathologists for Pathologists-DINOv2*
PathoDuetPathologyPathoDuet: Foundation Models for Pathological Slide Analysis of H&E and IHC StainsGithubMoCo v3TCGA + HyReCo + BCI
UNIPathologyTowards a general-purpose foundation model for computational pathology-DINOv2Mass-100K
REMEDISRadiologyRobust and Data-Efficient Generalization of Self-Supervised Machine Learning for Diagnostic ImagingGithubSimCLRMIMIC-IV + CheXpert
VirchowPathologyA foundation model for clinical-grade computational pathology and rare cancers detection-DINOv2*
RETFoundRetinopathyA Foundation Model for Generalizable Disease Detection from Retinal ImagesGithubMAE*
CTransPathPathologyTransformer-Based Unsupervised Contrastive Learning for Histopathological Image ClassificationGithub-TCGA + PAIP
HIPTPathologyScaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningGithubDINOTCGA

Vision-Language Models

ModelSubfieldPaperCodeBasePre-Training Data
TITANPathologyA multimodal whole-slide foundation model for pathology-CONCH335k WSIs + reports + synthetic captions*
DentVLMDentistryDentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice-Qwen2-VL2.4M oral images + 88.6K dental VQA*
LingshuMedicineLingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and ReasoningProjectQwen2.5-VL*
MUSKPathologyA Vision–Language Foundation Model for Precision OncologyGithubBEiT*
MONETDermatologyTransparent Medical Image AI via an Image–Text Foundation Model Grounded in Medical LiteratureGithubCLIPPubMed + textbooks
Uni-MedMedicineUni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE-CLIP + LLaMA 2*
MaCoRadiologyEnhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive LearningGithubCLIP + MAEMIMIC-CXR
RadFoundRadiologyExpert-Level Vision-Language Foundation Model for Real-World Radiology and Comprehensive Evaluation--RadVLCorpus*
BiomedGPTBiomedicineA Generalist Vision–Language Foundation Model for Diverse Biomedical TasksGithub-*
PRISMPathologyPRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology-CoCa*
Med-GeminiMedicineCapabilities of Gemini Models in Medicine-Gemini*
EchoCLIPSonographyVision-Language Foundation Model for Echocardiogram InterpretationGithubCLIP*
Med-PaLM MBiomedicineTowards Generalist Biomedical AI-PaLMMultiMedBench*
ChemDFMChemistryDeveloping ChemDFM as a large language foundation model for chemistry-LLaMAPubMed + USPTO
CheXagentRadiologyA Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray InterpretationGithubBLIP-2CheXinstruct*
SATRadiologyLarge-vocabulary segmentation for medical images with text promptsGithub-SAT-DS*
PathChatPathologyVision–language AI assistance in human pathology-LLaVAPathChatInstruct*
Qilin-Med-VLRadiologyQilin-Med-VL: Towards Chinese Large Vision-Language Model for General HealthcareGithubLLaVAChi-Med-VL*
CXR-CLIPRadiologyCXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-trainingGithubCLIPMIMIC-CXR + CheXpert + ChestX-ray14
PathLDMPathologyPathLDM: Text conditioned Latent Diffusion Model for HistopathologyGithubLatent DiffusionTCGA-BRCA + GPT-3.5
RadFMRadiologyTowards generalist foundation model for radiology by leveraging web-scale 2D&3D medical dataGithub-MedMD*
KADRadiologyKnowledge-Enhanced Visual-Language Pre-Training on Chest Radiology ImagesGithubCLIPMIMIC-CXR + UMLS
Med-FlamingoMedicineMed-Flamingo: A Multimodal Medical Few-Shot LearnerGithubFlamingoMTB + PMC-OA
CONCHPathologyA Visual-Language Foundation Model for Computational PathologyGithubCoCaPubMed + PMC
QuiltNetPathologyQuilt-1M: One Million Image-Text Pairs for HistopathologyGithubCLIPQuilt-1M*
PathAsstPathologyPathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyGithubCLIPPathCap + PathInstruct*
PLIPPathologyA Visual-Language Foundation Model for Pathology Image Analysis Using Medical TwitterHuggingfaceCLIPOpenPath*
MI-ZeroPathologyVisual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology ImagesGithubCLIPARCH
LLaVA-MedBiomedicineLLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One DayGithubLLaVAPMC-15M + GPT-4
MedVInTBiomedicinePMC-VQA: Visual Instruction Tuning for Medical Visual Question AnsweringGithub-PMC-VQA*
PMC-CLIPBiomedicinePMC-CLIP: Contrastive Language-Image Pre-Training Using Biomedical DocumentsGithubCLIPPMC-OA*
BiomedCLIPBiomedicineA Multimodal Biomedical Foundation Model Trained from Fifteen Million Image–Text PairsHuggingfaceCLIPPMC-15M*
MedKLIPRadiologyMedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisGithubCLIPMIMIC-CXR
MedCLIPMedicineMedCLIP: Contrastive Learning from Unpaired Medical Images and TextGithubCLIPCheXpert + MIMIC-CXR
CheXzeroRadiologyExpert-Level Detection of Pathologies from Unannotated Chest X-ray Images via Self-Supervised LearningGithubCLIPMIMIC-CXR
PubMedCLIPRadiologyPubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain?GithubCLIPROCO

Protein and Molecule Models

Other Models

Datasets for Foundation Model

Text

Dataset (Paper)DescriptionLink
MedBench (DOI)A Chinese medical LLM benchmark with 300,901 Chinese questions covering 43 clinical specialties, combined with an automatic evaluation systemOfficial site
MMedBench (DOI)A multilingual medical QA benchmark, where questions are categorized into 21 topicsGithub
MMedC (DOI)A multilingual medical corpus containing over 25.5B tokensGithub
BiMed1.3M (DOI)An English and Arabic bilingual dataset of 1.3M samples of medical QA and chatGithub
GAP-Replay (arXiv)48.1B tokens from 4 medical corpora including guidelines, abstracts, papers, and replayGithub
Huatuo-26M (DOI)26M Chinese medical QA pairsGithub
Medical Meadow (arXiv)16M medical QA pairs collected from 9 sourcesGithub
MultiMedQA (Nature)6 existing and 1 online-collected medical QA datasetNature
BigBio (NeurIPS)126+ biomedical NLP datasets covering 13 task categories and 10+ languagesGithub
MedMCQA (MLR)194K multiple-choice questions covering 2.4K healthcare topicsOfficial site
MedQA-USMLE (MDPI)61,097 multiple choice questions based on USMLE in three languagesGithub
CBLUE (DOI)A Chinese biomedical language understanding evaluation benchmark with 18 datasetsOfficial site
BLURB (DOI)13 biomedical NLP datasets in 6 tasksOfficial site
PubMedQA (DOI)1K expert-annotated, 61.2K unlabeled, and 211.3K artificially generated biomedical QA instancesOfficial site
BLUE (DOI)5 language tasks with 10 biomedical and clinical text datasetsGithub
webMedQA (BMC)63,284 real-world Chinese medical questions with over 300K answersGithub
MedMentions (DOI)4,392 papers annotated by experts with mentions of UMLS entitiesGithub
MIMIC-III (Nature)Critical care data for over 40,000 patientsOfficial site
ClinicalTrials.govAn online database of clinical research studies, including clinical trials and observational studiesOfficial site

Imaging

Dataset (Paper)DescriptionLink
3M-US2,187,915 ultrasound images of 12 common organs-
AbdomenAtlas (arXiv)20,460 3D CT volumes from 112 hospitals, with 673K masks of anatomical structuresGithub
BiomedParseData (Nature)1.1M images, 3.4M image-mask-label triples, and 6.8M image-mask-description triplesGithub
Mass-100K (DOI)100M tissue patches from 100,426 diagnostic H&E WSIs across 20 major tissue types-
RETFound (Nature)Unannotated retinal images, containing 904,170 CFPs and 736,442 OCT scansNature
AbdomenAtlas-8K (NeurIPS)8,448 CT volumes with per-voxel annotated eight abdominal organsGithub
Med-MNIST v2 (Nature)12 2D and 6 3D datasets for biomedical image classificationOfficial site
EchoNet-Dynamic (DOI)10,030 expert-annotated echocardiogram videosOfficial site
CheXpert (DOI)224,316 chest radiographs of 65,240 patientsOfficial site
Kather Colon Dataset (PMC)100K histological images of human colorectal cancer and healthy tissueZenodo
DeepLesion (PMC)32K CT scans with annotations and semantic labels from radiological reportsNIH
ChestXray-NIHCC (DOI)100K radiographs with labels from more than 30,000 patientsNIH
ISICAn archive containing 23K skin lesion images with labels & ImagingOfficial site

Genomics

Dataset (Paper)DescriptionLink
1000 Genomes Project (Nature)A comprehensive catalog of human genetic variationsOfficial site
ENCODE (Nature)A platform of genomics data and encyclopedia with integrative-level and ground-level annotationsNIH
dbSNP (NIH)A collection of human single nucleotide variations, microsatellites, and small-scale insertions and deletionsNIH

Drug

Dataset (Paper)DescriptionLink
DrugChat (arXiv)143,517 question-answer pairs covering 10,834 drug compounds, collected from PubChem and ChEMBL-
PubChem (NIH)A collection of 900+ sources of chemical information dataNIH
DrugBank (NIH)A web-enabled structured database of molecular information about drugsOfficial site
ChEMBL (NIH)20M bioactivity measurements for 2.4M distinct compounds and 15K protein targetsOfficial site

Multi-Modal

Dataset (Paper)DescriptionLink
MultiMedBench (NEJM AI)A multi-modal benchmark comprising 12 data sources and 14 tasks-
RadGenome-Chest CT (arXiv)A dataset of 3D chest CT, including 197 organ-level segmentation masks, 665K multi-granularity grounded reports, and 1.3M grounded VQA pairs-
OmniMedVQA (DOI)131,813 question-answering items with 120,530 images from 12 modalities and 26 human anatomical regions, collected from 75 medical datasets-
SAT-DS (DOI)11,462 scans with 142,254 segmentation annotations spanning 8 human body regions from 31 medical image segmentation datasets, together with domain knowledge from e-Anatomy and UMLSGithub
PathChatInstruct (DOI)250K+ diverse disease-agnostic visual-language instructions with image and text-
Chi-Med-VL (arXiv)580,014 image-text pairs and 469,441 question-answer pairs for general healthcare in ChineseGithub
MedMD (DOI)15.5M 2D scans and 180k 3D radiology scans with textual descriptionsGithub
OpenPath (Nature)208,414 pathology images paired with natural language descriptionsHuggingface
Quilt-1M (NeurIPS)1M image-text pairs for histopathologyGithub
Med-MMHL (arXiv)Human- and LLM-generated misinformation detection datasetGithub
Mol-Instructions (OpenReview)148K molecule-oriented, 505K protein-oriented, and biomolecular text instructionsHuggingface
PathInstruct (DOI)180K samples of LLM-generated instruction-following dataGithub
PMC-VQA (arXiv)227K VQA pairs of 149K images of various modalities or diseasesGithub
PMC-OA (DOI)1.6M fine-grained biomedical image-text pairsGithub
PathCap (DOI)142K pathology image-caption pairs from various sourcesGithub
SwissProtCLAP (DOI)441K text-protein sequence pairsGithub
MIMIC-IV (Nature)Clinical information for hospital stays of over 60,000 patientsOfficial site
MIMIC-CXR (Nature)227,835 chest imaging studies with free-text reports for 65,379 patientsPhysioNet
TCGAA landmark cancer genomics program, molecularly characterized over 20,000 primary cancer and matched normal samples spanning 33 cancer typesOfficial site

Contributors

Yunkun-Zhang

26 commits

shiyegao

2 commits

Yunkun-Zhang/Data-Centric-FM-Healthcare

A survey on data-centric foundation models in healthcare.

78

28 commits

updated Feb 14, 2026

See the code

README

Data-Centric Foundation Models in Computational Healthcare

:fire::fire::fire: A survey on data-centric foundation models in computational healthcare

Project Page | Paper [arXiv]

Last updated: 2026/02/05

:pencil: If you find this repo helps, please kindly cite our survey, thanks!

@article{zhang2024data,
  title={Data-Centric Foundation Models in Computational Healthcare: A Survey},
  author={Zhang, Yunkun and Gao, Jin and Tan, Zheling and Zhou, Lingfeng and Ding, Kexin and Zhou, Mu and Zhang, Shaoting and Wang, Dequan},
  journal={arXiv},
  year={2024},
  eprint={2401.02458},
  archivePrefix={arXiv},
  primaryClass={cs.LG},
  doi={10.48550/arXiv.2401.02458},
  url={https://arxiv.org/abs/2401.02458}
}

In this repository, we provide an up-to-date list of healthcare-related foundation models and datasets, which are also mentioned in our survey paper.

:book: Contents


Healthcare and Medical Foundation Models

A star (*) after the pre-training data shows that the authors constructed the data with more than three sources.

Language Models

ModelSubfieldPaperCodeBasePre-Training Data
Baichuan-M2MedicineBaichuan-M2: Scaling Medical Capability with Large Verifier SystemGithubQwen2.5*
Baichuan-M1MedicineBaichuan-M1: Pushing the Medical Capability of Large Language Models-Transformer20T tokens*
EHRMambaClinicEHRMamba: Towards Generalizable and Scalable Foundation Models for Electronic Health RecordsGithubMambaMIMIC-IV
MMedLM 2MedicineTowards building multilingual language model for medicineGithubInternLM 2MMedC*
BiMediXMedicineBiMediX: Bilingual Medical Mixture of Experts LLMGithubMixtralBiMed1.3M*
Me LLaMAMedicineMe LLaMA: Foundation Large Language Models for Medical ApplicationsGithubLLaMA 2*
BioMistralBiomedicineBioMistral: A Collection of Open-Source Pretrained Large Language Models for Medical Domains-MistralPubMed Central
PULSEMedicine-GithubInternLM*
MeditronMedicineMEDITRON-70B: Scaling Medical Pretraining for Large Language ModelsGithubLLaMA 2GAP-Replay*
TaiyiBiomedicineTaiyi: A Bilingual Fine-Tuned Large Language Model for Diverse Biomedical TasksGithubQwen-7B / GLM4-9BBigBio + CBLUE
BioMedGPTBiomedicineBioMedGPT: An Open Multimodal Large Language Model for BioMedicineGithubLLaMA 2S2ORC
Clinical LLaMA-LoRAClinicParameter-Efficient Fine-Tuning of LLaMA for the Clinical Domain-LLaMAMIMIC-IV
Med-PaLM 2ClinicToward expert-level medical question answering with large language modelsGooglePaLM 2MedQA
PMC-LLaMAMedicinePMC-LLaMA: toward building open-source language models for medicineGithubLLaMAMedC
MedAlpacaMedicineMedAlpaca -- An Open-Source Collection of Medical Conversational AI Models and Training DataGithubLLaMAMedical Meadow
BenTsao (HuaTuo)BiomedicineHuaTuo: Tuning LLaMA Model with Chinese Medical KnowledgeGithubLLaMACMeKG
ChatDoctorMedicineChatDoctor: A Medical Chat Model Fine-Tuned on a Large Language Model Meta-AI (LLaMA) Using Medical Domain KnowledgeGithubLLaMAHealthCareMagic*
Clinical-T5ClinicClinical-T5: Large Language Models Built Using Mimic Clinical TextPhysioNetT5MIMIC-III + MIMIC-IV
Med-PaLMClinicLarge Language Models Encode Clinical KnowledgeGooglePaLMMedQA
BioGPTBiomedicineBioGPT: Generative Pre-Trained Transformer for Biomedical Text Generation and MiningGithubGPT-2PubMed
BioLinkBERTBiomedicineLinkBERT: Pretraining Language Models with Document LinksGithubBERTPubMed (citation links)
PubMedBERTBiomedicineDomain-Specific Language Model Pretraining for Biomedical Natural Language ProcessingMicrosoftBERTPubMed
BioBERTBiomedicineBioBERT: A Pre-Trained Biomedical Language Representation Model for Biomedical Text MiningGithubBERTPubMed + PMC
BlueBERTBiomedicineAn Empirical Study of Multi-Task Learning on BERT for Biomedical Text MiningGithubBERTPubMed + MIMIC-III
Clinical BERTClinicPublicly Available Clinical BERT EmbeddingsGithubBERTMIMIC-III
SciBERTBiomedicineSciBERT: A Pretrained Language Model for Scientific TextGithubBERTSemantic Scholar

Vision Models

ModelSubfieldPaperCodeBasePre-Training Data
FastGliomaPathologyFoundation models for fast, label-free detection of glioma infiltration--Label-free optical microscopy (4M images)*
MedLSAMRadiologyMedLSAM: Localize and Segment Anything Model for 3D CT ImagesGithubSAM*
BiomedParseBiomedicineA Foundation Model for Joint Segmentation, Detection and Recognition of Biomedical Objects across Nine ModalitiesGithubSEEMBiomedParseData*
Universal ModelRadiologyUniversal and Extensible Language-Vision Models for Organ Segmentation and Tumor Detection from Abdominal Computed TomographyGithub-*
CHIEFPathologyA pathology foundation model for cancer diagnosis and prognosis predictionGithubCTransPath*
USFMSonographyUSFM: A Universal Ultrasound Foundation Model Generalized to Tasks and Organs towards Label Efficient Image Analysis-MIM3M-US*
BrainSegFounderRadiologyBrainSegFounder: towards 3D foundation models for neuroimage segmentationGithubSwinUNETRUK Biobank + BraTS + ATLAS
MedSAMMedicineSegment Anything in Medical ImagesGithubSAM*
Prov-GigaPathPathologyA Whole-Slide Foundation Model for Digital Pathology from Real-World DataGithub-Prov-Path*
BEPHPathologyA Foundation Model for Generalizable Cancer Diagnosis and Survival Prediction from Histopathological ImagesGithubBEiTv2*
Pai et al.RadiologyFoundation Model for Cancer Imaging BiomarkersGithubSimCLR*
VIS-MAERadiologyVIS-MAE: An Efficient Self-supervised Learning Approach on Medical Image Segmentation and Classification-MAE*
SegmentAnyBoneRadiologySegmentAnyBone: A Universal Model that Segments Any Bone at Any Location on MRIGithubSAM*
RudolfVPathologyRudolfV: A Foundation Model by Pathologists for Pathologists-DINOv2*
PathoDuetPathologyPathoDuet: Foundation Models for Pathological Slide Analysis of H&E and IHC StainsGithubMoCo v3TCGA + HyReCo + BCI
UNIPathologyTowards a general-purpose foundation model for computational pathology-DINOv2Mass-100K
REMEDISRadiologyRobust and Data-Efficient Generalization of Self-Supervised Machine Learning for Diagnostic ImagingGithubSimCLRMIMIC-IV + CheXpert
VirchowPathologyA foundation model for clinical-grade computational pathology and rare cancers detection-DINOv2*
RETFoundRetinopathyA Foundation Model for Generalizable Disease Detection from Retinal ImagesGithubMAE*
CTransPathPathologyTransformer-Based Unsupervised Contrastive Learning for Histopathological Image ClassificationGithub-TCGA + PAIP
HIPTPathologyScaling Vision Transformers to Gigapixel Images via Hierarchical Self-Supervised LearningGithubDINOTCGA

Vision-Language Models

ModelSubfieldPaperCodeBasePre-Training Data
TITANPathologyA multimodal whole-slide foundation model for pathology-CONCH335k WSIs + reports + synthetic captions*
DentVLMDentistryDentVLM: A Multimodal Vision-Language Model for Comprehensive Dental Diagnosis and Enhanced Clinical Practice-Qwen2-VL2.4M oral images + 88.6K dental VQA*
LingshuMedicineLingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and ReasoningProjectQwen2.5-VL*
MUSKPathologyA Vision–Language Foundation Model for Precision OncologyGithubBEiT*
MONETDermatologyTransparent Medical Image AI via an Image–Text Foundation Model Grounded in Medical LiteratureGithubCLIPPubMed + textbooks
Uni-MedMedicineUni-Med: A Unified Medical Generalist Foundation Model For Multi-Task Learning Via Connector-MoE-CLIP + LLaMA 2*
MaCoRadiologyEnhancing Representation in Radiography-Reports Foundation Model: A Granular Alignment Algorithm Using Masked Contrastive LearningGithubCLIP + MAEMIMIC-CXR
RadFoundRadiologyExpert-Level Vision-Language Foundation Model for Real-World Radiology and Comprehensive Evaluation--RadVLCorpus*
BiomedGPTBiomedicineA Generalist Vision–Language Foundation Model for Diverse Biomedical TasksGithub-*
PRISMPathologyPRISM: A Multi-Modal Generative Foundation Model for Slide-Level Histopathology-CoCa*
Med-GeminiMedicineCapabilities of Gemini Models in Medicine-Gemini*
EchoCLIPSonographyVision-Language Foundation Model for Echocardiogram InterpretationGithubCLIP*
Med-PaLM MBiomedicineTowards Generalist Biomedical AI-PaLMMultiMedBench*
ChemDFMChemistryDeveloping ChemDFM as a large language foundation model for chemistry-LLaMAPubMed + USPTO
CheXagentRadiologyA Vision-Language Foundation Model to Enhance Efficiency of Chest X-ray InterpretationGithubBLIP-2CheXinstruct*
SATRadiologyLarge-vocabulary segmentation for medical images with text promptsGithub-SAT-DS*
PathChatPathologyVision–language AI assistance in human pathology-LLaVAPathChatInstruct*
Qilin-Med-VLRadiologyQilin-Med-VL: Towards Chinese Large Vision-Language Model for General HealthcareGithubLLaVAChi-Med-VL*
CXR-CLIPRadiologyCXR-CLIP: Toward Large Scale Chest X-ray Language-Image Pre-trainingGithubCLIPMIMIC-CXR + CheXpert + ChestX-ray14
PathLDMPathologyPathLDM: Text conditioned Latent Diffusion Model for HistopathologyGithubLatent DiffusionTCGA-BRCA + GPT-3.5
RadFMRadiologyTowards generalist foundation model for radiology by leveraging web-scale 2D&3D medical dataGithub-MedMD*
KADRadiologyKnowledge-Enhanced Visual-Language Pre-Training on Chest Radiology ImagesGithubCLIPMIMIC-CXR + UMLS
Med-FlamingoMedicineMed-Flamingo: A Multimodal Medical Few-Shot LearnerGithubFlamingoMTB + PMC-OA
CONCHPathologyA Visual-Language Foundation Model for Computational PathologyGithubCoCaPubMed + PMC
QuiltNetPathologyQuilt-1M: One Million Image-Text Pairs for HistopathologyGithubCLIPQuilt-1M*
PathAsstPathologyPathAsst: A Generative Foundation AI Assistant towards Artificial General Intelligence of PathologyGithubCLIPPathCap + PathInstruct*
PLIPPathologyA Visual-Language Foundation Model for Pathology Image Analysis Using Medical TwitterHuggingfaceCLIPOpenPath*
MI-ZeroPathologyVisual Language Pretrained Multiple Instance Zero-Shot Transfer for Histopathology ImagesGithubCLIPARCH
LLaVA-MedBiomedicineLLaVA-Med: Training a Large Language-and-Vision Assistant for Biomedicine in One DayGithubLLaVAPMC-15M + GPT-4
MedVInTBiomedicinePMC-VQA: Visual Instruction Tuning for Medical Visual Question AnsweringGithub-PMC-VQA*
PMC-CLIPBiomedicinePMC-CLIP: Contrastive Language-Image Pre-Training Using Biomedical DocumentsGithubCLIPPMC-OA*
BiomedCLIPBiomedicineA Multimodal Biomedical Foundation Model Trained from Fifteen Million Image–Text PairsHuggingfaceCLIPPMC-15M*
MedKLIPRadiologyMedKLIP: Medical Knowledge Enhanced Language-Image Pre-Training for X-ray DiagnosisGithubCLIPMIMIC-CXR
MedCLIPMedicineMedCLIP: Contrastive Learning from Unpaired Medical Images and TextGithubCLIPCheXpert + MIMIC-CXR
CheXzeroRadiologyExpert-Level Detection of Pathologies from Unannotated Chest X-ray Images via Self-Supervised LearningGithubCLIPMIMIC-CXR
PubMedCLIPRadiologyPubMedCLIP: How Much Does CLIP Benefit Visual Question Answering in the Medical Domain?GithubCLIPROCO

Protein and Molecule Models

Other Models

Datasets for Foundation Model

Text

Dataset (Paper)DescriptionLink
MedBench (DOI)A Chinese medical LLM benchmark with 300,901 Chinese questions covering 43 clinical specialties, combined with an automatic evaluation systemOfficial site
MMedBench (DOI)A multilingual medical QA benchmark, where questions are categorized into 21 topicsGithub
MMedC (DOI)A multilingual medical corpus containing over 25.5B tokensGithub
BiMed1.3M (DOI)An English and Arabic bilingual dataset of 1.3M samples of medical QA and chatGithub
GAP-Replay (arXiv)48.1B tokens from 4 medical corpora including guidelines, abstracts, papers, and replayGithub
Huatuo-26M (DOI)26M Chinese medical QA pairsGithub
Medical Meadow (arXiv)16M medical QA pairs collected from 9 sourcesGithub
MultiMedQA (Nature)6 existing and 1 online-collected medical QA datasetNature
BigBio (NeurIPS)126+ biomedical NLP datasets covering 13 task categories and 10+ languagesGithub
MedMCQA (MLR)194K multiple-choice questions covering 2.4K healthcare topicsOfficial site
MedQA-USMLE (MDPI)61,097 multiple choice questions based on USMLE in three languagesGithub
CBLUE (DOI)A Chinese biomedical language understanding evaluation benchmark with 18 datasetsOfficial site
BLURB (DOI)13 biomedical NLP datasets in 6 tasksOfficial site
PubMedQA (DOI)1K expert-annotated, 61.2K unlabeled, and 211.3K artificially generated biomedical QA instancesOfficial site
BLUE (DOI)5 language tasks with 10 biomedical and clinical text datasetsGithub
webMedQA (BMC)63,284 real-world Chinese medical questions with over 300K answersGithub
MedMentions (DOI)4,392 papers annotated by experts with mentions of UMLS entitiesGithub
MIMIC-III (Nature)Critical care data for over 40,000 patientsOfficial site
ClinicalTrials.govAn online database of clinical research studies, including clinical trials and observational studiesOfficial site

Imaging

Dataset (Paper)DescriptionLink
3M-US2,187,915 ultrasound images of 12 common organs-
AbdomenAtlas (arXiv)20,460 3D CT volumes from 112 hospitals, with 673K masks of anatomical structuresGithub
BiomedParseData (Nature)1.1M images, 3.4M image-mask-label triples, and 6.8M image-mask-description triplesGithub
Mass-100K (DOI)100M tissue patches from 100,426 diagnostic H&E WSIs across 20 major tissue types-
RETFound (Nature)Unannotated retinal images, containing 904,170 CFPs and 736,442 OCT scansNature
AbdomenAtlas-8K (NeurIPS)8,448 CT volumes with per-voxel annotated eight abdominal organsGithub
Med-MNIST v2 (Nature)12 2D and 6 3D datasets for biomedical image classificationOfficial site
EchoNet-Dynamic (DOI)10,030 expert-annotated echocardiogram videosOfficial site
CheXpert (DOI)224,316 chest radiographs of 65,240 patientsOfficial site
Kather Colon Dataset (PMC)100K histological images of human colorectal cancer and healthy tissueZenodo
DeepLesion (PMC)32K CT scans with annotations and semantic labels from radiological reportsNIH
ChestXray-NIHCC (DOI)100K radiographs with labels from more than 30,000 patientsNIH
ISICAn archive containing 23K skin lesion images with labels & ImagingOfficial site

Genomics

Dataset (Paper)DescriptionLink
1000 Genomes Project (Nature)A comprehensive catalog of human genetic variationsOfficial site
ENCODE (Nature)A platform of genomics data and encyclopedia with integrative-level and ground-level annotationsNIH
dbSNP (NIH)A collection of human single nucleotide variations, microsatellites, and small-scale insertions and deletionsNIH

Drug

Dataset (Paper)DescriptionLink
DrugChat (arXiv)143,517 question-answer pairs covering 10,834 drug compounds, collected from PubChem and ChEMBL-
PubChem (NIH)A collection of 900+ sources of chemical information dataNIH
DrugBank (NIH)A web-enabled structured database of molecular information about drugsOfficial site
ChEMBL (NIH)20M bioactivity measurements for 2.4M distinct compounds and 15K protein targetsOfficial site

Multi-Modal

Dataset (Paper)DescriptionLink
MultiMedBench (NEJM AI)A multi-modal benchmark comprising 12 data sources and 14 tasks-
RadGenome-Chest CT (arXiv)A dataset of 3D chest CT, including 197 organ-level segmentation masks, 665K multi-granularity grounded reports, and 1.3M grounded VQA pairs-
OmniMedVQA (DOI)131,813 question-answering items with 120,530 images from 12 modalities and 26 human anatomical regions, collected from 75 medical datasets-
SAT-DS (DOI)11,462 scans with 142,254 segmentation annotations spanning 8 human body regions from 31 medical image segmentation datasets, together with domain knowledge from e-Anatomy and UMLSGithub
PathChatInstruct (DOI)250K+ diverse disease-agnostic visual-language instructions with image and text-
Chi-Med-VL (arXiv)580,014 image-text pairs and 469,441 question-answer pairs for general healthcare in ChineseGithub
MedMD (DOI)15.5M 2D scans and 180k 3D radiology scans with textual descriptionsGithub
OpenPath (Nature)208,414 pathology images paired with natural language descriptionsHuggingface
Quilt-1M (NeurIPS)1M image-text pairs for histopathologyGithub
Med-MMHL (arXiv)Human- and LLM-generated misinformation detection datasetGithub
Mol-Instructions (OpenReview)148K molecule-oriented, 505K protein-oriented, and biomolecular text instructionsHuggingface
PathInstruct (DOI)180K samples of LLM-generated instruction-following dataGithub
PMC-VQA (arXiv)227K VQA pairs of 149K images of various modalities or diseasesGithub
PMC-OA (DOI)1.6M fine-grained biomedical image-text pairsGithub
PathCap (DOI)142K pathology image-caption pairs from various sourcesGithub
SwissProtCLAP (DOI)441K text-protein sequence pairsGithub
MIMIC-IV (Nature)Clinical information for hospital stays of over 60,000 patientsOfficial site
MIMIC-CXR (Nature)227,835 chest imaging studies with free-text reports for 65,379 patientsPhysioNet
TCGAA landmark cancer genomics program, molecularly characterized over 20,000 primary cancer and matched normal samples spanning 33 cancer typesOfficial site

Contributors

Yunkun-Zhang

26 commits

shiyegao

2 commits