11 repos
Libraries and model implementations for multimodal AI systems that generate text from visual and textual inputs. The cluster centers on the MGM (Multimodal Generation Model) family of vision-language models at various scales (7B–13B parameters), alongside supporting infrastructure for text generation, model endpoints, and safetensor format compatibility. Repositories here focus on enabling developers to work with and deploy these integrated vision-and-language capabilities.