10 repos
Large-scale pretrained models that combine vision and language understanding to process and generate content across both modalities. The cluster centers on practical implementations of vision-language architectures (CogVLM, VisualGLM, InternVL), along with supporting tools, training frameworks, and applications built on these foundations. Repositories here span model implementations, fine-tuning infrastructure, and downstream applications that leverage multimodal reasoning.