Multimodal Vision-Language Models

42 repos across 2 sub-areas

Libraries and model implementations for vision-language tasks, particularly image captioning and image-to-text generation using transformer architectures. The cluster centers on Japanese language variants of LLaVA models and related multimodal AI systems that combine visual understanding with text generation. While the language composition appears sparse in the metadata, the consistent presence of image-captioning, transformers, and image-text-to-text topics across repositories indicates a focused area around building and deploying models that process both images and text.