Multimodal Vision-Language Models

10 repos

Open-source implementations and variants of multimodal models that process both images and text, translating visual information into textual descriptions or answers. The cluster centers on the Molmo family of models—compact vision-language transformers designed for efficient image-to-text tasks—alongside related multimodal architectures. Repositories here include model weights, training code, and web-optimized implementations for inference and fine-tuning.

multimodal ·1,034
en ·1,034
image-text-to-text ·1,034
molmo ·1,034
transformers ·1,034
olmo ·1,018
custom_code ·1,010
text-generation ·888
conversational ·853
safetensors ·853