10 repos
Methods and models for generating synchronized audio and video content, with emphasis on techniques like diffusion transformers (DiT) and instruction-based control for multimodal synthesis. The cluster spans generation frameworks, evaluation methodologies, and fine-tuning approaches that enable systems to produce coherent audio-visual outputs where sound and motion are jointly modeled rather than generated independently.