10 repos
Datasets and tools for collecting, annotating, and preparing multimodal image-text data for training vision-language models. The cluster centers on projects like LLaVA-ReCap variants and GPT-4V datasets that generate or curate large-scale captioned image collections, enabling development of models that combine visual understanding with natural language processing. Repositories here focus on data pipeline construction, caption generation at scale, and dataset standardization rather than model architecture or training inference.