Vision-Language Models

12 repos

Multimodal AI systems that combine visual and textual understanding, enabling models to process and reason over both images and language together. This cluster centers on pretrained vision-language models like CogVLM, which integrate image encoders with large language models to support tasks such as visual question answering, image captioning, and cross-modal reasoning. Repositories here include model implementations, pretrained checkpoints, and related multimodal learning frameworks.

Python · 3
pretrained-models ·15,917
language-model ·15,917
multi-modal ·15,917
visual-language-models ·13,486
cross-modality ·13,486
cogvlm ·2,431
custom_code ·567
text-generation ·567
transformers ·567
safetensors ·567