Multimodal Vision-Language Models

28 repos

Libraries, models, and implementations for connecting vision and language through transformer-based architectures, enabling tasks like image captioning, visual question answering, and image-to-text generation. The cluster centers on PyTorch-based multimodal frameworks that combine image encoders with language models, with several repositories featuring Japanese-language variants and instruction-tuned chat models built on foundations like LLaMA and Stable LM.

Python · 1
transformers ·2,368
image-captioning ·2,331
image-to-text ·2,269
image-text-to-text ·2,225
pytorch ·2,170
endpoints_compatible ·2,057
vision-encoder-decoder ·937
tf ·889
blip ·889
safetensors ·501