2 repos
Large-scale multimodal models that combine vision and language capabilities for image understanding and generation tasks. The cluster centers on model variants across different parameter scales (2B through 8x7B) optimized for inference deployment, with emphasis on safe tensor serialization and endpoint compatibility for production use. Repositories here represent both the core model architectures and infrastructure for serving vision-language generation at scale.