Vision-Language Models and Image-to-Text

44 repos across 2 sub-areas

Multimodal AI systems that combine vision and language understanding to perform tasks like image captioning, visual question answering, and image-to-text generation. This cluster centers on transformer-based architectures (particularly variants of LLaVA and BLIP2) that bridge computer vision and natural language processing, enabling models to understand and describe visual content. Repositories here include pretrained model weights, fine-tuning implementations, and inference code built primarily with PyTorch.