Multimodal Vision-Language Models

43 repos across 4 sub-areas

Large-scale transformer-based models that combine vision and language capabilities for tasks like image understanding, visual question answering, and cross-modal feature extraction. The cluster centers on the InternVL family of models—ranging from compact 2B parameters to large 76B variants—which demonstrate how to build efficient multimodal systems across different scale points. Repositories here focus on model architectures, safetensor checkpoints, multilingual support, and practical implementations of vision-language integration using modern transformer frameworks.