40 repos
Large language models augmented with vision capabilities for understanding and reasoning over images and text together. The cluster centers on multimodal transformer architectures that process both visual and textual inputs, enabling applications like image captioning, visual question answering, and cross-modal retrieval. The InternVL family of models forms the core, with variants ranging from 2B to 76B parameters optimized for different inference budgets and multilingual understanding.