Multimodal Vision-Language Models

15 repos

Libraries and model implementations for connecting vision and language understanding, enabling AI systems to process and generate text based on images. This cluster centers on LLaVA (Large Language and Vision Assistant) and related variants, which combine image encoders with language models using transformer architectures. Builders here will find pretrained model weights, fine-tuning approaches via LoRA adapters, and PyTorch-based implementations for vision-language tasks like image captioning and visual question-answering.

transformers ·1,763
text-generation ·1,763
image-text-to-text ·1,680
llava ·1,397
pytorch ·1,232
safetensors ·451
llava_mistral ·246
conversational ·246
llamafile ·188
GGUF ·188