OpenMOSS-Team/MOVA-720p

Model

132

stars

13

commits

1

linked in READMEs

Feb 11, 2026

updated

any-to-any
diffusers
image-text-to-audio-video
image-text-to-video
image-to-audio-video
image-to-video
MOSI
MOVA
OpenMOSS
safetensors
sglang-diffusion
SII

README

MOVA: Towards Scalable and Synchronized Video–Audio Generation

We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the "silent era" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.

🌟Key Highlights

  • Native Bimodal Generation: Moves beyond clunky cascaded pipelines. MOVA generates high-fidelity video and synchronized audio in a single inference pass, eliminating error accumulation.
  • Precise Lip-Sync & Sound FX: Achieves state-of-the-art performance in multilingual lip-synchronization and environment-aware sound effects.
  • Fully Open-Source: In a field dominated by closed-source models (Sora 2, Veo 3, Kling), we are releasing model weights, inference code, training pipelines, and LoRA fine-tuning scripts.
  • Asymmetric Dual-Tower Architecture: Leverages the power of pre-trained video and audio towers, fused via a bidirectional cross-attention mechanism for rich modality interaction.

Demo

Model Details

Model Description

MOVA addresses the limitations of proprietary systems like Sora 2 and Veo 3 by offering a fully open-source framework for Image-to-Video-Audio (IT2VA) and Text-to-Video-Audio (T2VA) tasks. The model employs an asymmetric dual-tower architecture fused via a bidirectional cross-attention mechanism, leveraging a Mixture-of-Experts (MoE) design with 32B total parameters (18B active during inference) to ensure high-quality synthesis with efficient deployment. Alongside the model weights, we provide a fine-grained bimodal data pipeline and support for LoRA fine-tuning, empowering the community to advance research in synchronized cinematic synthesis.

Model Sources

Model Usage

Please refer to the Quick Start section on the GitHub page for model usage and inference scripts.

Evaluation

We evaluate our model through both objective benchmarks and subjective human evaluations. Below are the Elo scores and win rates comparing MOVA to existing open-source models.

Citation

@article{yu2026mova,
  title={MOVA: Towards Scalable and Synchronized Video-Audio Generation},
  author={Yu, Donghua and Chen, Mingshu and Chen, Qi and Luo, Qi and Wu, Qianyi and Cheng, Qinyuan and Li, Ruixiao and Liang, Tianyi and Zhang, Wenbo and Tu, Wenming and others},
  journal={arXiv preprint arXiv:2602.08794},
  year={2026}
}

Contributors

yhzx233

9 commits

Cqy2019

2 commits

nielsr

1 commits

pxyy

1 commits

OpenMOSS-Team/MOVA-720p

Model

132

stars

13

commits

1

linked in READMEs

Feb 11, 2026

updated

any-to-any
diffusers
image-text-to-audio-video
image-text-to-video
image-to-audio-video
image-to-video
MOSI
MOVA
OpenMOSS
safetensors
sglang-diffusion
SII

README

MOVA: Towards Scalable and Synchronized Video–Audio Generation

We introduce MOVA (MOSS Video and Audio), a foundation model designed to break the "silent era" of open-source video generation. Unlike cascaded pipelines that generate sound as an afterthought, MOVA synthesizes video and audio simultaneously for perfect alignment.

🌟Key Highlights

  • Native Bimodal Generation: Moves beyond clunky cascaded pipelines. MOVA generates high-fidelity video and synchronized audio in a single inference pass, eliminating error accumulation.
  • Precise Lip-Sync & Sound FX: Achieves state-of-the-art performance in multilingual lip-synchronization and environment-aware sound effects.
  • Fully Open-Source: In a field dominated by closed-source models (Sora 2, Veo 3, Kling), we are releasing model weights, inference code, training pipelines, and LoRA fine-tuning scripts.
  • Asymmetric Dual-Tower Architecture: Leverages the power of pre-trained video and audio towers, fused via a bidirectional cross-attention mechanism for rich modality interaction.

Demo

Model Details

Model Description

MOVA addresses the limitations of proprietary systems like Sora 2 and Veo 3 by offering a fully open-source framework for Image-to-Video-Audio (IT2VA) and Text-to-Video-Audio (T2VA) tasks. The model employs an asymmetric dual-tower architecture fused via a bidirectional cross-attention mechanism, leveraging a Mixture-of-Experts (MoE) design with 32B total parameters (18B active during inference) to ensure high-quality synthesis with efficient deployment. Alongside the model weights, we provide a fine-grained bimodal data pipeline and support for LoRA fine-tuning, empowering the community to advance research in synchronized cinematic synthesis.

Model Sources

Model Usage

Please refer to the Quick Start section on the GitHub page for model usage and inference scripts.

Evaluation

We evaluate our model through both objective benchmarks and subjective human evaluations. Below are the Elo scores and win rates comparing MOVA to existing open-source models.

Citation

@article{yu2026mova,
  title={MOVA: Towards Scalable and Synchronized Video-Audio Generation},
  author={Yu, Donghua and Chen, Mingshu and Chen, Qi and Luo, Qi and Wu, Qianyi and Cheng, Qinyuan and Li, Ruixiao and Liang, Tianyi and Zhang, Wenbo and Tu, Wenming and others},
  journal={arXiv preprint arXiv:2602.08794},
  year={2026}
}

Contributors

yhzx233

9 commits

Cqy2019

2 commits

nielsr

1 commits

pxyy

1 commits