15 repos
Synthetic dataset generation and curation for training vision-language models in medical imaging and healthcare contexts. The cluster centers on MedVLSynther, a system for creating paired medical image-text datasets at scale, with multiple dataset versions (1K through 13K samples) and format variants designed for different model architectures and training paradigms. These resources address the challenge of obtaining large, well-annotated medical multimodal datasets needed to develop and evaluate models that combine visual understanding of medical images with natural language reasoning.