NVIDIA Cosmos multimodal AI models

13 repos

Foundation models and transfer learning implementations from NVIDIA's Cosmos suite, encompassing video generation, reasoning, and image upscaling capabilities. These repositories provide pre-trained model checkpoints, fine-tuning examples, and specialized variants (single-to-multi-view synthesis, 4K upscaling, safety guardrails) built on diffusers and safetensors infrastructure. Intended for researchers and practitioners integrating state-of-the-art multimodal perception and generation into downstream applications.

cosmos ·1,211
nvidia ·1,211
safetensors ·996
diffusers ·597
conversational ·434
image-text-to-text ·434
qwen3_vl ·434
cosmos3 ·387
cosmos3_omni ·387
omnimodel ·353