12 repos
Multimodal AI systems that combine visual and textual understanding, enabling models to process and reason over both images and language together. This cluster centers on pretrained vision-language models like CogVLM, which integrate image encoders with large language models to support tasks such as visual question answering, image captioning, and cross-modal reasoning. Repositories here include model implementations, pretrained checkpoints, and related multimodal learning frameworks.