43 repos across 4 sub-areas
Large-scale transformer-based models that combine vision and language capabilities for tasks like image understanding, visual question answering, and cross-modal feature extraction. The cluster centers on the InternVL family of models—ranging from compact 2B parameters to large 76B variants—which demonstrate how to build efficient multimodal systems across different scale points. Repositories here focus on model architectures, safetensor checkpoints, multilingual support, and practical implementations of vision-language integration using modern transformer frameworks.
Cluster 462330
40 repos
Multimodal Vision-Language Models
1 repos
Models and frameworks for processing and generating text from images, combining visual and language understanding in conversational systems. The cluster centers on InternVL3.5 variants—a family of open vision-language models spanning multiple scales (1B to 240B parameters)—alongside supporting infrastructure for model distribution via safetensors, transformer-based architectures, and integration with conversational interfaces. Repositories here enable image-to-text reasoning, visual question answering, and multimodal dialogue applications.
Cluster 462331
1 repos
Cluster 462332
1 repos