Vision-Language Models and Multimodal AI

40 repos

Large language models augmented with vision capabilities for understanding and reasoning over images and text together. The cluster centers on multimodal transformer architectures that process both visual and textual inputs, enabling applications like image captioning, visual question answering, and cross-modal retrieval. The InternVL family of models forms the core, with variants ranging from 2B to 76B parameters optimized for different inference budgets and multilingual understanding.

image-text-to-text ·2,570
multilingual ·2,570
internvl ·2,570
transformers ·2,570
custom_code ·2,559
conversational ·2,557
safetensors ·2,557
feature-extraction ·2,506
internvl_chat ·2,506
tensorboard ·966