MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos.
Python
33
6 commits
updated Apr 18, 2026

[π Paper] ο½ [π€ Dataset] ο½ [π€ Model (1.3B)] ο½ [π€ Model (14B)] ο½ [π Blog]
Apr 18, 2026: We released our training code.Jul 8, 2025: We released our paper, data and project. Models are coming soon. Please stay tuned!Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation.
[!NOTE] We open-sourced our models, data, and code here.
You can β¬οΈdownload our full MedVideoCap-55K from HuggingFace. Our dataset has several features:
# Install
pip install -r requirements.txt
# Train LoRA on 1.3B
bash scripts/train_lora_1.3B.sh
# Train full on 14B with DeepSpeed
bash scripts/train_full_14B.sh
# Generate single video
python inference.py --model_name "Wan-AI/Wan2.1-T2V-1.3B" \
--checkpoint_dir "./models/train/wan_lora_1.3B" \
--is_lora --prompt "your prompt"
# Batch generation
bash scripts/infer_batch.sh --metadata_path data/metadata.json
For more details, please refer to quick start.
Please refer to evaluation.
Our works are inspired by the following works.
@misc{wang2025medgenunlockingmedicalvideo,
title={MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos},
author={Rongsheng Wang and Junying Chen and Ke Ji and Zhenyang Cai and Shunian Chen and Yunjin Yang and Benyou Wang},
year={2025},
eprint={2507.05675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05675},
}
Python
62.7%
Shell
37.3%
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos.
Python
33
6 commits
updated Apr 18, 2026

[π Paper] ο½ [π€ Dataset] ο½ [π€ Model (1.3B)] ο½ [π€ Model (14B)] ο½ [π Blog]
Apr 18, 2026: We released our training code.Jul 8, 2025: We released our paper, data and project. Models are coming soon. Please stay tuned!Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and simulation, requiring not only high visual fidelity but also strict medical accuracy. However, current models often produce unrealistic or erroneous content when applied to medical prompts, largely due to the lack of large-scale, high-quality datasets tailored to the medical domain. To address this gap, we introduce MedVideoCap-55K, the first large-scale, diverse, and caption-rich dataset for medical video generation. It comprises over 55,000 curated clips spanning real-world medical scenarios, providing a strong foundation for training generalist medical video generation models. Built upon this dataset, we develop MedGen, which achieves leading performance among open-source models and rivals commercial systems across multiple benchmarks in both visual quality and medical accuracy. We hope our dataset and model can serve as a valuable resource and help catalyze further research in medical video generation.
[!NOTE] We open-sourced our models, data, and code here.
You can β¬οΈdownload our full MedVideoCap-55K from HuggingFace. Our dataset has several features:
# Install
pip install -r requirements.txt
# Train LoRA on 1.3B
bash scripts/train_lora_1.3B.sh
# Train full on 14B with DeepSpeed
bash scripts/train_full_14B.sh
# Generate single video
python inference.py --model_name "Wan-AI/Wan2.1-T2V-1.3B" \
--checkpoint_dir "./models/train/wan_lora_1.3B" \
--is_lora --prompt "your prompt"
# Batch generation
bash scripts/infer_batch.sh --metadata_path data/metadata.json
For more details, please refer to quick start.
Please refer to evaluation.
Our works are inspired by the following works.
@misc{wang2025medgenunlockingmedicalvideo,
title={MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos},
author={Rongsheng Wang and Junying Chen and Ke Ji and Zhenyang Cai and Shunian Chen and Yunjin Yang and Benyou Wang},
year={2025},
eprint={2507.05675},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2507.05675},
}
Python
62.7%
Shell
37.3%