15 repos
Libraries and model implementations for connecting vision and language understanding, enabling AI systems to process and generate text based on images. This cluster centers on LLaVA (Large Language and Vision Assistant) and related variants, which combine image encoders with language models using transformer architectures. Builders here will find pretrained model weights, fine-tuning approaches via LoRA adapters, and PyTorch-based implementations for vision-language tasks like image captioning and visual question-answering.