43 repos across 3 sub-areas
Multimodal AI systems that combine visual and textual understanding across multiple languages, enabling models to process images alongside conversational input. The cluster centers on the InternVL family of vision-language transformers (ranging from 1B to 241B parameters), which integrate computer vision capabilities with natural language processing. Repositories here represent implementations, variants, and applications of these architectures for tasks like visual question answering, image captioning, and cross-lingual visual reasoning.
Cluster 650509
34 repos
Multilingual Vision-Language Models
8 repos
Large language models trained to understand and generate text about images, with support for multiple languages. The cluster centers on the InternVL family of chat models—compact variants like Mini-InternVL-Chat-2B and full-scale versions up to InternVL-Chat-V1-5—which combine vision transformers with language modeling to enable image captioning, visual question answering, and multimodal dialogue across language boundaries. These repositories provide model weights, inference code, and fine-tuning examples for building production systems that merge computer vision with natural language understanding.
Cluster 650507
1 repos