Multimodal Large Language Models

11 repos

Research and implementation of vision-language models that combine large language models with visual understanding capabilities. The cluster centers on mixture-of-experts (MoE) architectures for efficient multimodal processing, with repositories implementing lightweight variants using different base models (Phi, StableLM, Qwen). This area covers model training, optimization, and inference techniques for building efficient multimodal AI systems that can understand and reason about both text and images.

Python · 2
custom_code ·104
text-generation ·104
transformers ·104
endpoints_compatible ·103
safetensors ·88
moe_llava_phi ·72
moe_llava_stablelm ·16
moe_llava_qwen ·15
pytorch ·15
llava_qwen ·1