Multimodal Foundation Models

10 repos

Large language and vision models designed for any-to-any generation across text, image, and potentially other modalities. The cluster centers on Lumina-mGPT variants and Chameleon-based architectures that use safetensors for model serialization, representing a family of multimodal foundation models at different scales (7B to 34B parameters) and context window sizes. These are research and production implementations of unified models capable of understanding and generating multiple data types within a single architecture.

Python · 2
safetensors ·168
chameleon ·168
text-to-image ·84
any-to-any ·74
Any2Any ·65
transformers ·57
endpoints_compatible ·57
text-generation ·48
image-to-image ·10
image-text-to-text ·9