Multimodal Vision-Language Models

10 repos

Large-scale pretrained models that combine vision and language understanding to process and generate content across both modalities. The cluster centers on practical implementations of vision-language architectures (CogVLM, VisualGLM, InternVL), along with supporting tools, training frameworks, and applications built on these foundations. Repositories here span model implementations, fine-tuning infrastructure, and downstream applications that leverage multimodal reasoning.

Python · 9
TypeScript · 1
multi-modal ·41,567
gpt ·25,654
pretrained-models ·25,521
language-model ·24,400
visual-language-models ·13,480
cross-modality ·13,480
llm ·10,156
image-text-retrieval ·10,156
gpt-4v ·10,156
gpt-4o ·10,156