1 repo
Models and frameworks for processing and generating text from images, combining visual and language understanding in conversational systems. The cluster centers on InternVL3.5 variants—a family of open vision-language models spanning multiple scales (1B to 240B parameters)—alongside supporting infrastructure for model distribution via safetensors, transformer-based architectures, and integration with conversational interfaces. Repositories here enable image-to-text reasoning, visual question answering, and multimodal dialogue applications.