Multimodal vision-language models

7 repos

This cluster focuses on vision-language models that combine image and text understanding, particularly around the LLaVA architecture and its variants. Repositories here contain implementations, pretrained models, and extensions for image-to-text generation and multimodal reasoning tasks using transformer-based approaches. Learners will find practical resources for building and fine-tuning models that process both visual and textual inputs.

text-generation ·58
transformers ·58
safetensors ·58
endpoints_compatible ·53
video LLM ·53
llava ·53
conversational ·33
llava_llama ·5
image-text-to-text ·2