28 repos
Libraries, models, and implementations for connecting vision and language through transformer-based architectures, enabling tasks like image captioning, visual question answering, and image-to-text generation. The cluster centers on PyTorch-based multimodal frameworks that combine image encoders with language models, with several repositories featuring Japanese-language variants and instruction-tuned chat models built on foundations like LLaMA and Stable LM.