43 repos across 4 sub-areas
Large-scale transformer-based models that combine vision and language capabilities for tasks like image understanding, visual question answering, and cross-modal feature extraction. The cluster centers on the InternVL family of models—ranging from compact 2B parameters to large 76B variants—which demonstrate how to build efficient multimodal systems across different scale points. Repositories here focus on model architectures, safetensor checkpoints, multilingual support, and practical implementations of vision-language integration using modern transformer frameworks.
Vision-Language Models and Multimodal AI
40 repos
Large language models augmented with vision capabilities for understanding and reasoning over images and text together. The cluster centers on multimodal transformer architectures that process both visual and textual inputs, enabling applications like image captioning, visual question answering, and cross-modal retrieval. The InternVL family of models forms the core, with variants ranging from 2B to 76B parameters optimized for different inference budgets and multilingual understanding.
Multimodal Vision-Language Models
1 repos
Models and frameworks for processing and generating text from images, combining visual and language understanding in conversational systems. The cluster centers on InternVL3.5 variants—a family of open vision-language models spanning multiple scales (1B to 240B parameters)—alongside supporting infrastructure for model distribution via safetensors, transformer-based architectures, and integration with conversational interfaces. Repositories here enable image-to-text reasoning, visual question answering, and multimodal dialogue applications.
Cluster 462331
1 repos
Cluster 462332
1 repos