7 repos
This cluster focuses on vision-language models that combine image and text understanding, particularly around the LLaVA architecture and its variants. Repositories here contain implementations, pretrained models, and extensions for image-to-text generation and multimodal reasoning tasks using transformer-based approaches. Learners will find practical resources for building and fine-tuning models that process both visual and textual inputs.