Vision-Language Model Generation

11 repos

Libraries and model implementations for multimodal AI systems that generate text from visual and textual inputs. The cluster centers on the MGM (Multimodal Generation Model) family of vision-language models at various scales (7B–13B parameters), alongside supporting infrastructure for text generation, model endpoints, and safetensor format compatibility. Repositories here focus on enabling developers to work with and deploy these integrated vision-and-language capabilities.

generation ·126
endpoints_compatible ·126
vision-language model ·126
text-generation ·126
safetensors ·126
transformers ·126
llama ·89
conversational ·74
gemma ·21
mixtral ·16