Multimodal LLM Vision Architectures

9 repos

Practical implementations and variants of multimodal language models that combine vision encoders (CLIP, SigLIP, Fuyu, Idefics2) with large language models (LLaMA 3, others) to enable vision-language understanding. The cluster centers on the Mantis family of models and related instruction-tuned variants that adapt these multimodal architectures for tasks like visual question answering and image captioning. Developers exploring this area will find model implementations, training code, and adapter architectures for binding vision and language components.

Python · 3
mllm ·4,409
llava ·4,177
llm ·4,128
lmm ·3,873
llama3 ·3,588
gpt4 ·3,540
eagle ·3,540
demo ·3,540
large-language-models ·3,540
huggingface ·3,540