Vision-Language Models and Multimodal AI

8 repos

This cluster focuses on pretrained models that combine visual and language understanding, enabling AI systems to process and reason over both images and text simultaneously. The repositories center on multimodal architectures and cross-modality learning, with CogVLM and its variants as core implementations. Researchers exploring this area will find model implementations, training approaches, and applications of vision-language systems for tasks requiring integrated visual and linguistic reasoning.

safetensors ·567
transformers ·567
text-generation ·567
custom_code ·567
en ·543
chat ·220
conversational ·220
cogvlm2 ·220