9 repos
Practical implementations and variants of multimodal language models that combine vision encoders (CLIP, SigLIP, Fuyu, Idefics2) with large language models (LLaMA 3, others) to enable vision-language understanding. The cluster centers on the Mantis family of models and related instruction-tuned variants that adapt these multimodal architectures for tasks like visual question answering and image captioning. Developers exploring this area will find model implementations, training code, and adapter architectures for binding vision and language components.