8 repos
Large language models trained to understand and generate text about images, with support for multiple languages. The cluster centers on the InternVL family of chat models—compact variants like Mini-InternVL-Chat-2B and full-scale versions up to InternVL-Chat-V1-5—which combine vision transformers with language modeling to enable image captioning, visual question answering, and multimodal dialogue across language boundaries. These repositories provide model weights, inference code, and fine-tuning examples for building production systems that merge computer vision with natural language understanding.