Multimodal Vision-Language Model Generation

2 repos

Large-scale multimodal models that combine vision and language capabilities for image understanding and generation tasks. The cluster centers on model variants across different parameter scales (2B through 8x7B) optimized for inference deployment, with emphasis on safe tensor serialization and endpoint compatibility for production use. Repositories here represent both the core model architectures and infrastructure for serving vision-language generation at scale.

Python · 2
large-language-models ·3,513
vision-language-model ·3,513
generation ·3,324
large-multimodal-models ·189
llava ·189
llm ·189
multimodal ·189
multi-modality ·189
large-context ·189