8 repos
This cluster focuses on pretrained models that combine visual and language understanding, enabling AI systems to process and reason over both images and text simultaneously. The repositories center on multimodal architectures and cross-modality learning, with CogVLM and its variants as core implementations. Researchers exploring this area will find model implementations, training approaches, and applications of vision-language systems for tasks requiring integrated visual and linguistic reasoning.