10 repos
Open-source implementations and variants of multimodal models that process both images and text, translating visual information into textual descriptions or answers. The cluster centers on the Molmo family of models—compact vision-language transformers designed for efficient image-to-text tasks—alongside related multimodal architectures. Repositories here include model weights, training code, and web-optimized implementations for inference and fine-tuning.