12 repos
Datasets and annotation frameworks for training vision-language models that combine image and text data. The cluster centers on large-scale image captioning datasets (LLaVA-ReCap variants covering different source distributions, CC3M, and general web data) along with instruction-tuning and evaluation datasets designed to improve multimodal AI systems. These resources support the development of models that can understand and reason about visual content paired with natural language.