Eyeline-Labs/Survey-Video-Diffusion

The official paper summary of TMLR'25 paper "Survey of Video Diffusion Models: Foundations, Implementations, and Applications"

44

68 commits

updated Feb 2, 2026

See the code

README

Survey of Video Diffusion Models: Foundations, Implementations, and Applications

Paper

1University of Waterloo, 2Duke University, 3Netflix Eyeline Studios
*Contributed Equally, †Corresponding Author

Abstract

In this survey and github repository, we provide a comprehensive overview of the recent advances in video diffusion models. We cover the foundations of video generative models, including GANs, auto-regressive models, and diffusion models. We also discuss the learning foundations, including classic denoising diffusion models, flow matching, and training-free methods. Additionally, we explore various architectures, including UNet and diffusion transformers. We discuss the applications of video diffusion models, including video generation, enhancement, personalization, and 3D-aware video generation. Finally, we highlight the benefits of video diffusion models to other domains, such as video representation learning and video retrieval.

Moreover, to facilitate the understanding of video diffusion models, we provide a cheatsheet including commonly used training datasets, training engineering techniques, and evaluation metrics. We also provide a list of video diffusion models in academia and industry.

image

Table of Contents

Foundations

Video generative paradigms

GAN video models

Papers are listed generally in reverse order of their publication timestamps.

Auto-regressive video models

Papers are listed generally in reverse order of their publication timestamps.

Video diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Auto-regressive video diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Learning foundations

Classic denoising diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Flow matching and rectified flow

Papers are listed generally in reverse order of their publication timestamps.

Learning from feedback and reward models

Papers are listed generally in reverse order of their publication timestamps.

One-shot and few-shot learning

Papers are listed generally in reverse order of their publication timestamps.

Training-free methods

Papers are listed generally in reverse order of their publication timestamps.

Token learning

Papers are listed generally in reverse order of their publication timestamps.

Guidances

Classifier guidance

Papers are listed generally in reverse order of their publication timestamps.

Classifier-free guidance

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Classifier-Free Diffusion GuidancearXivStarWebsite2022

Diffusion model frameworks

Pixel diffusion and latent diffusion

Papers are listed generally in reverse order of their publication timestamps.

Optical-flow-based diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Noise scheduling

Papers are listed generally in reverse order of their publication timestamps.

Agent-based diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Architectures

UNet

Papers are listed generally in reverse order of their publication timestamps.

Diffusion transformers

Papers are listed generally in reverse order of their publication timestamps.

VAE for latent space compression

Papers are listed generally in reverse order of their publication timestamps.

Text encoders

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Magic 1-for-1: Generating one minute video clips within one minutearXivStarWebsitearXiv 2025
SkyReels v1: Human-centric video foundation modelarXivStar-arXiv 2025
Hunyuanvideo: A systematic framework for large video generative modelsarXiv-WebsitearXiv 2025
Step-video-t2v technical report: The practice, challenges, and future of video foundation modelarXiv-WebsitearXiv 2025
Identity-Preserving Text-to-Video Generation by Frequency DecompositionarXivStarWebsitearXiv 2024
An empirical study and analysis of text-to-image generation using large language model-powered textual representationarXiv--arXiv 2024
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understandingarXiv-WebsitearXiv 2024
FIT: Flexible Vision Transformer for Diffusion ModelarXivStarWebsitearXiv 2024
CogVideoX: Text-to-Video Diffusion Models with an Expert TransformerarXivStarWebsitearXiv 2024
Scaling Rectified Flow Transformers for High-Resolution Image SynthesisarXivStarWebsiteICML 2024
SIT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant TransformersarXivStarWebsitearXiv 2024
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesisarXiv-WebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Open-Sora-PlanarXivStarWebsitearXiv 2024
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
Latte: Latent Diffusion Transformer for Video GenerationarXivStarWebsitearXiv 2024
FluxarXivStarWebsitearXiv 2023
Scalable Diffusion Models with TransformersarXivStarWebsiteICCV 2023
All are Worth Words: A ViT Backbone for Diffusion ModelsarXivStarWebsiteCVPR 2023
Baichuan 2: Open Large-Scale Language ModelsarXivStarWebsitearXiv 2023
LLaMA 2: Open Foundation and Fine-Tuned Chat ModelsarXivStarWebsitearXiv 2023
LLaMA: Open and Efficient Foundation Language ModelsarXivStarWebsitearXiv 2023
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte ModelsarXivStar-TACL 2022
Imagen Video: High Definition Video Generation with Diffusion ModelsarXiv-WebsitearXiv 2022
Hierarchical Text-Conditional Image Generation with CLIP LatentsarXiv-WebsitearXiv 2022
High-Resolution Image Synthesis with Latent Diffusion ModelsarXivStarWebsiteCVPR 2022
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsarXivStar-arXiv 2021
GLM: General Language Model Pretraining with Autoregressive Blank InfillingarXivStarWebsitearXiv 2021
Learning Transferable Visual Models From Natural Language SupervisionarXivStarWebsiteICML 2021
Zero-Shot Text-to-Image GenerationarXiv--ICML 2021
Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerarXivStar-JMLR 2020
BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingarXivStarWebsiteNAACL 2019

Implementation

Datasets

More datasets could be found on Pixabay, Mixkit, Pond5, Adobe Stock, Shutterstock, Getty, Coverr, Videvo, Depositphotos, Storyblocks, Dissolve, Freepik, Vimeo, and Envato. Also, there are some datasets at Midjourney V5.1 Cleaned Data, Unsplash-lite, AnimateBench, Pexels-400k, and LAION-AESTHETICS.

Title
arXiv
GitHub
Website
Conference & Year
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality TeachersarXivStarWebsiteCVPR 2024
VBench: Comprehensive Benchmark Suite for Video Generative ModelsarXivStarWebsiteCVPR 2024
InternVid: Learning Text-to-Video Generation from Web-scale Video-Text DataarXivStarWebsiteICLR 2024
MiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsarXivStarWebsiteNeurIPS 2024
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion ModelsarXivStarWebsiteNeurIPS 2024
Vript: A Video Is Worth Thousands of WordsarXivStar-NeurIPS 2024
VideoCrafter2arXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorarXivStarWebsiteICCV 2023
Temporally Consistent Transformers for Video GenerationarXivStarWebsiteICML 2023
Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and MethodarXivStar-NeurIPS 2023
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video GenerationarXivStar-NeurIPS 2023
AIGCBench: Comprehensive evaluation of image-to-video content generated by AIarXivStarWebsiteTBench 2023
AdaPool: Exponential Adaptive Pooling for Information-Retaining DownsamplingarXivStar-TIP 2023
Swap Attention in Spatiotemporal Diffusions for Text-to-Video GenerationarXivStar-arXiv 2023
Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsarXivStar-CVPR 2022
The DEVIL is in the Details: A Diagnostic Evaluation Benchmark for Video InpaintingarXivStar-CVPR 2022
VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-ResolutionarXiv-WebsiteCVPR 2022
Learning Audio-Video Modalities from Image CaptionsarXiv-WebsiteECCV 2022
The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video EditingarXivStar-ECCV 2022
Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingarXivStarWebsiteNeurIPS 2022
Scaling Autoregressive Models for Content-Rich Text-to-Image GenerationarXiv-WebsiteTMLR 2022
ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningarXivStarWebsiteICCV 2021
Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalarXivStarWebsiteICCV 2021
MERLOT: Multimodal Neural Script Knowledge ModelsarXivStarWebsiteNeurIPS 2021
Learning Video Representations from Textual Web SupervisionarXiv--arXiv 2020
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsarXiv-WebsiteICCV 2019
VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearcharXiv-WebsiteICCV 2019
Towards Automatic Learning of Procedures from Web Instructional VideosarXiv-WebsiteAAAI 2018
How2: A Large-scale Dataset for Multimodal Language UnderstandingarXivStarWebsitearXiv 2018
Quo Vadis, Action Recognition? A New Model and the Kinetics DatasetarXivStar-CVPR 2017
Localizing Moments in Video with Natural LanguagearXivStar-ICCV 2017
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language--WebsiteCVPR 2016
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding--WebsiteCVPR 2015
A Dataset for Movie DescriptionarXiv--CVPR 2015
UCF101: A Dataset of 101 Human Actions Classes From Videos in The WildarXiv-WebsitearXiv 2012

Training engineering

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsarXivStarWebsiteICLR 2025
SAM 2: Segment Anything in Images and VideosarXiv-WebsiteICLR 2025
Motion Prompting: Controlling Video Generation with Motion TrajectoriesarXiv-WebsiteCVPR 2025
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
DynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsarXivStarWebsiteECCV 2024
CogVLM2: Visual Language Models for Image and Video UnderstandingarXivStar-arXiv 2024
CogVideoX: Text-to-Video Diffusion Models with an Expert TransformerarXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understandingarXiv-WebsitearXiv 2024
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense CaptioningarXivStarWebsitearXiv 2024
Cogvideo: Large-scale pretraining for text-to-video generation via transformersarXivStar-ICLR 2023
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsarXivStar-arXiv 2023
LLaMA: Open and Efficient Foundation Language ModelsarXivStarWebsitearXiv 2023
LLaMA 2: Open Foundation and Fine-Tuned Chat ModelsarXivStarWebsitearXiv 2023
ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningarXivStar-NeurIPS 2022
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessarXivStar-NeurIPS 2022
Visual Prompt TuningarXivStar-ECCV 2022
CoCa: Contrastive Captioners are Image-Text Foundation ModelsarXivStar-TMLR 2022
ZeRO: memory optimizations toward training trillion parameter modelsarXiv--Supercomputing 2020

Evaluation metrics and benchmarking findings

Papers are listed generally in reverse order of their publication timestamps.

Industry models

Title
arXiv
GitHub
Website
Conference & Year
Magic 1-for-1: Generating one minute video clips within one minutearXivStarWebsitearXiv 2025
SkyReels v1: Human-centric video foundation modelarXivStar-arXiv 2025
Step-Video-T2VarXiv-WebsitearXiv 2024
HunyuanVideoarXiv-WebsitearXiv 2024
Sora--Website2024
STIVarXivStarWebsitearXiv 2024
LTX-VideoarXivStarWebsitearXiv 2024
AllegroarXivStarWebsitearXiv 2024
JimengarXiv-WebsitearXiv 2024
Mochi 1arXiv-WebsitearXiv 2024
EasyAnimatearXivStarWebsitearXiv 2024
Vidu--Website2024
VideoCrafter2arXivStarWebsitearXiv 2024
VideoCrafter1arXivStarWebsitearXiv 2023
MiraarXiv-WebsitearXiv 2024
Hailuo AI--Website2024
LumierearXiv-WebsitearXiv 2024
VideoPoetarXiv-WebsitearXiv 2023
LumaAI Ray 2--Website2024
LumaAI Dream Machine--Website2023
Veo-2--Website2024
Veo-1--Website2023
Nova Real--Website2024
Wanx 2.1--Website2024
Kling--Website2024
Show-1arXivStarWebsiteNeurIPS 2023
MovieGenarXiv-WebsitearXiv 2024
Pika--Website2023
Vchitect-2.0--Website2024
OptisarXivStarWebsiteNeurIPS 2023
VLoggerarXivStarWebsiteICCV 2023
SeinearXivStarWebsiteCVPR 2023
LaviearXivStarWebsiteICCV 2023
MiracleVision--Website2023
PhenakiarXivStarWebsiteICLR 2024
W.A.L.TarXiv-WebsitearXiv 2024
Imagen videoarXiv-Website2022
GEN-3 Alpha--Website2024
GEN-2--Website2023
GEN-1--Website2022

Academia models

Title
arXiv
GitHub
Website
Conference & Year
RepVideo: Rethinking Cross-Layer Representation for Video GenerationarXivStarWebsitearXiv 2025
CausVid: Causality-Aware Video Generation with Slow-Fast Diffusion ModelsarXivStarWebsiteCVPR 2025
Open-Sora Plan: Open-Source Large Video Generation ModelarXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video SynthesisarXiv-WebsitearXiv 2024
SiT: Exploring Flow and Diffusion-Based Generative Models with Scalable Interpolant TransformersarXivStarWebsitearXiv 2024
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided PlanningarXivStarWebsiteCOLM 2024
AnimateLCM: Accelerating the Animation of Personalized Diffusion Models and Adapters with Decoupled Consistency LearningarXivStarWebsitearXiv 2024
I4VGEN: Interactive Video Generation via Integrated Dynamic ControlarXivStarWebsitearXiv 2024
SimDA: Simple Diffusion Adapter for Efficient Text-to-Video GenerationarXivStarWebsitearXiv 2023
AnimateDiff-v2arXivStarWebsiteICLR 2024
Animate-A-Story: Storytelling with Retrieval-Augmented Video GenerationarXivStarWebsitearXiv 2023
VideoGen: A Reference-Guided Latent Diffusion Approach for High-Definition Text-to-Video GenerationarXiv-WebsitearXiv 2023
Dysen-VDM: Diffusion Model with Dynamic Spatio-Temporal Fusion for Video GenerationarXivStar-arXiv 2023
HiGen: Hierarchical 3D Feature Generation for 3D-Aware Image Synthesis and ManipulationarXivStar-arXiv 2023
ModelScope Text-to-Video Technical ReportarXiv-WebsitearXiv 2023
InstructVideo: Instructing Video Diffusion Models with Human FeedbackarXiv-WebsiteCVPR 2024
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNeurIPS 2023
VideoFusion: Decomposed Diffusion Models for High-Quality Video GenerationarXiv--CVPR 2023
MagViT-v2: Masked Generative Video TransformerarXivStar-arXiv 2023
MagViT: Masked Generative Video TransformerarXivStar-arXiv 2022
Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video GenerationarXiv-WebsitearXiv 2023
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion ModelsarXivStarWebsiteCVPR 2023
Video Diffusion ModelsarXivStarWebsitearXiv 2022
Make-A-Video: Text-to-Video Generation without Text-Video DataarXivStarWebsiteICLR 2023
MagicVideo: Efficient Video Generation With Latent Diffusion ModelsarXiv-WebsitearXiv 2022
CogVideoX: Enhancing Video Understanding in the Era of Large Language ModelsarXivStarWebsitearXiv 2024
CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersarXivStarWebsiteICLR 2023
VideoGPT: Video Generation using VQ-VAE and TransformersarXivStarWebsitearXiv 2021

Applications

Conditions

Image condition

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerarXivStarWebsiteICLR 2025
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion ControlarXiv-WebsiteICLR 2025
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text GuidancearXivStarWebsiteICASSP 2025
EMO: Emote Portrait Alive-Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak ConditionsarXiv-WebsiteECCV 2024
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion ModelsarXivStarWebsiteCVPR 2025
MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and EditingarXivStarWebsiteECCV 2024
ConsistI2V: Enhancing Visual Consistency for Image-to-Video GenerationarXivStarWebsiteTMLR
I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsarXivStarWebsiteSIGGRAPH 2024
ID-Animator: Zero-Shot Identity-Preserving Human Video GenerationarXivStarWebsitearXiv 2024
CamCo: Camera-Controllable 3D-Consistent Image-to-Video GenerationarXiv-WebsitearXiv 2024
Generative Image DynamicsarXiv-WebsiteCVPR 2024
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image ModelsarXivStarWebsiteCVPR 2024
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion ModelsarXiv-WebsiteCVPR 2024
AtomoVideo: High Fidelity Image-to-Video GenerationarXiv-WebsitearXiv 2024
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modelingarXivStarWebsiteSIGGRAPH 2024
Seer: Language Instructed Video Prediction with Latent Diffusion ModelsarXivStarWebsiteICLR 2024
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion GuidancearXivStarWebsitearXiv 2023
VideoBooth: Diffusion-based Video Generation with Image PromptsarXivStarWebsiteCVPR 2024
Sparsectrl: Adding sparse controls to text-to-video diffusion modelsarXivStarWebsiteECCV 2024
DynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsarXivStarWebsiteECCV 2024
Adding Conditional Control to Text-to-Image Diffusion ModelsarXivStarWebsiteICCV 2023
Stable video diffusion: Scaling latent video diffusion models to large datasetsarXivStarWebsiteArxix 2023
Make pixels dance: High-dynamic video generationarXiv-WebsiteCVPR 2024
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsarXivStarWebsitearXiv 2023
Videocrafter1: Open diffusion models for high-quality video generationarXivStarWebsitearXiv 2023
VDT: General-purpose Video Diffusion Transformers via Mask ModelingarXivStarWebsiteICLR 2024
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNIPS 2023
Conditional Image-to-Video Generation with Latent Flow Diffusion ModelsarXivStarWebsiteCVPR 2023

Spatial condition

Papers are listed generally in reverse order of their publication timestamps.

Camera parameter condition

Papers are listed generally in reverse order of their publication timestamps.

Audio condition

Papers are listed generally in reverse order of their publication timestamps.

High-level video condition

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationarXivStarWebsiteICLR 2024
MotionClone: Training-Free Motion Cloning for Controllable Video GenerationarXivStarWebsiteICLR 2025
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion ModelsarXivStarWebsiteSIGGRAPH Asia 2024
ReVideo: Remake a Video with Motion and Content ControlarXivStarWebsiteNeurIPS 2024
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingarXivStarWebsiteACM MM 2024
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing TasksarXivStarWebsiteTMLR 2024
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingarXivStarWebsitearXiv 2024
VidToMe: Video Token Merging for Zero-Shot Video EditingarXivStarWebsiteCVPR 2024
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisarXiv-WebsiteCVPR 2024
SAVE: Protagonist Diversification with Structure Agnostic Video EditingarXivStarWebsiteECCV 2024
RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion ModelsarXivStarWebsiteCVPR 2024
DiffusionAtlas: High-Fidelity Consistent Diffusion Video EditingarXiv-WebsitearXiv 2023
DragVideo: Interactive Drag-style Video EditingarXivStarWebsiteECCV 2024
Drag-A-Video: Non-rigid Video Editing with Point-based InteractionarXivStarWebsitearXiv 2023
VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point CorrespondencearXivStarWebsiteCVPR 2024
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video EditingarXivStarWebsiteCVPR 2024
Motion-Conditioned Image Animation for Video EditingarXivStarWebsitearXiv 2023
MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionarXivStarWebsiteICML 2024
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character AnimationarXivStarWebsiteCVPR 2024
Consistent Video-to-Video Transfer Using Synthetic DatasetarXivStarWebsiteICLR 2024
MotionDirector: Motion Customization of Text-to-Video Diffusion ModelsarXivStarWebsiteECCV 2024
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
MagicEdit: High-Fidelity and Temporally Coherent Video EditingarXivStarWebsitearXiv 2023
CoDeF: Content Deformation Fields for Temporally Consistent Video ProcessingarXivStarWebsiteCVPR 2024
StableVideo: Text-driven Consistency-aware Diffusion Video EditingarXivStarWebsiteICCV 2023
VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNetarXivStarWebsitearXiv 2023
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNeurIPS 2023
Rerender A Video: Zero-Shot Text-Guided Video-to-Video TranslationarXivStarWebsiteSIGGRAPH Asia 2023
Video Colorization with Pre-trained Text-to-Image Diffusion ModelsarXivStarWebsitearXiv 2023
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video EditingarXiv-WebsiteTMLR 2024
DisCo: Disentangled Control for Realistic Human Dance GenerationarXivStarWebsiteCVPR 2024
Towards Consistent Video Editing with Text-to-Image Diffusion ModelsarXiv--NeurIPS 2023
Video ControlNet: Towards Temporally Consistent Synthetic-to-Real Video Translation Using Conditional Image Diffusion ModelsarXiv---
ControlVideo: Training-free Controllable Text-to-Video GenerationarXivStarWebsiteICLR 2024
InstructVid2Vid: Controllable Video Editing with Natural Language InstructionsarXiv--arXiv 2023
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free VideosarXivStarWebsiteAAAI 2024
DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionarXivStarWebsiteICCV 2023
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion ModelsarXivStar-IEEE Trans On Multimedia, 2023
Pix2Video: Video Editing using Image DiffusionarXivStarWebsiteICCV 2023
Structure and Content-Guided Video Synthesis with Diffusion ModelsarXiv-WebsiteICCV 2023
Shape-aware Text-driven Layered Video EditingarXivStarWebsiteCVPR 2023
DPE: Disentanglement of Pose and Expression for General Video Portrait EditingarXivStarWebsiteCVPR 2023
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationarXivStarWebsiteICCV 2023
Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video EncodingarXivStarWebsiteCVPR 2023
Layered Neural Atlases for Consistent Video EditingarXivStarWebsiteSIGGRAPH Asia 2021

Other conditions

Papers are listed generally in reverse order of their publication timestamps.

Enhancement

Video denoising and deblurring

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Video restoration based on deep learning: a comprehensive survey---2022

Video inpainting

Papers are listed generally in reverse order of their publication timestamps.

Video interpolation and extrapolation/prediction

Papers are listed generally in reverse order of their publication timestamps.

Video super-resolution

Papers are listed generally in reverse order of their publication timestamps.

Combining multiple video enhancement tasks

Papers are listed generally in reverse order of their publication timestamps.

Personalization

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Dynamic Concepts Personalization from Single VideosarXiv-WebsiteSIGGRAPH 2025
VideoAlchemy: Open-set Personalization in Video GenerationarXivStarWebsiteCVPR 2025
[PersonalVideo: High ID-Fidelity V

Truncated — view the full README on GitHub.

Contributors

yimuwangcs

29 commits

weipang142857

18 commits

ningyu1991

12 commits

limacv

5 commits

Eyeline-Labs/Survey-Video-Diffusion

The official paper summary of TMLR'25 paper "Survey of Video Diffusion Models: Foundations, Implementations, and Applications"

44

68 commits

updated Feb 2, 2026

See the code

README

Survey of Video Diffusion Models: Foundations, Implementations, and Applications

Paper

1University of Waterloo, 2Duke University, 3Netflix Eyeline Studios
*Contributed Equally, †Corresponding Author

Abstract

In this survey and github repository, we provide a comprehensive overview of the recent advances in video diffusion models. We cover the foundations of video generative models, including GANs, auto-regressive models, and diffusion models. We also discuss the learning foundations, including classic denoising diffusion models, flow matching, and training-free methods. Additionally, we explore various architectures, including UNet and diffusion transformers. We discuss the applications of video diffusion models, including video generation, enhancement, personalization, and 3D-aware video generation. Finally, we highlight the benefits of video diffusion models to other domains, such as video representation learning and video retrieval.

Moreover, to facilitate the understanding of video diffusion models, we provide a cheatsheet including commonly used training datasets, training engineering techniques, and evaluation metrics. We also provide a list of video diffusion models in academia and industry.

image

Table of Contents

Foundations

Video generative paradigms

GAN video models

Papers are listed generally in reverse order of their publication timestamps.

Auto-regressive video models

Papers are listed generally in reverse order of their publication timestamps.

Video diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Auto-regressive video diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Learning foundations

Classic denoising diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Flow matching and rectified flow

Papers are listed generally in reverse order of their publication timestamps.

Learning from feedback and reward models

Papers are listed generally in reverse order of their publication timestamps.

One-shot and few-shot learning

Papers are listed generally in reverse order of their publication timestamps.

Training-free methods

Papers are listed generally in reverse order of their publication timestamps.

Token learning

Papers are listed generally in reverse order of their publication timestamps.

Guidances

Classifier guidance

Papers are listed generally in reverse order of their publication timestamps.

Classifier-free guidance

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Classifier-Free Diffusion GuidancearXivStarWebsite2022

Diffusion model frameworks

Pixel diffusion and latent diffusion

Papers are listed generally in reverse order of their publication timestamps.

Optical-flow-based diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Noise scheduling

Papers are listed generally in reverse order of their publication timestamps.

Agent-based diffusion models

Papers are listed generally in reverse order of their publication timestamps.

Architectures

UNet

Papers are listed generally in reverse order of their publication timestamps.

Diffusion transformers

Papers are listed generally in reverse order of their publication timestamps.

VAE for latent space compression

Papers are listed generally in reverse order of their publication timestamps.

Text encoders

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Magic 1-for-1: Generating one minute video clips within one minutearXivStarWebsitearXiv 2025
SkyReels v1: Human-centric video foundation modelarXivStar-arXiv 2025
Hunyuanvideo: A systematic framework for large video generative modelsarXiv-WebsitearXiv 2025
Step-video-t2v technical report: The practice, challenges, and future of video foundation modelarXiv-WebsitearXiv 2025
Identity-Preserving Text-to-Video Generation by Frequency DecompositionarXivStarWebsitearXiv 2024
An empirical study and analysis of text-to-image generation using large language model-powered textual representationarXiv--arXiv 2024
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understandingarXiv-WebsitearXiv 2024
FIT: Flexible Vision Transformer for Diffusion ModelarXivStarWebsitearXiv 2024
CogVideoX: Text-to-Video Diffusion Models with an Expert TransformerarXivStarWebsitearXiv 2024
Scaling Rectified Flow Transformers for High-Resolution Image SynthesisarXivStarWebsiteICML 2024
SIT: Exploring Flow and Diffusion-based Generative Models with Scalable Interpolant TransformersarXivStarWebsitearXiv 2024
Kolors: Effective training of diffusion model for photorealistic text-to-image synthesisarXiv-WebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Open-Sora-PlanarXivStarWebsitearXiv 2024
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
Latte: Latent Diffusion Transformer for Video GenerationarXivStarWebsitearXiv 2024
FluxarXivStarWebsitearXiv 2023
Scalable Diffusion Models with TransformersarXivStarWebsiteICCV 2023
All are Worth Words: A ViT Backbone for Diffusion ModelsarXivStarWebsiteCVPR 2023
Baichuan 2: Open Large-Scale Language ModelsarXivStarWebsitearXiv 2023
LLaMA 2: Open Foundation and Fine-Tuned Chat ModelsarXivStarWebsitearXiv 2023
LLaMA: Open and Efficient Foundation Language ModelsarXivStarWebsitearXiv 2023
ByT5: Towards a Token-Free Future with Pre-trained Byte-to-Byte ModelsarXivStar-TACL 2022
Imagen Video: High Definition Video Generation with Diffusion ModelsarXiv-WebsitearXiv 2022
Hierarchical Text-Conditional Image Generation with CLIP LatentsarXiv-WebsitearXiv 2022
High-Resolution Image Synthesis with Latent Diffusion ModelsarXivStarWebsiteCVPR 2022
GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsarXivStar-arXiv 2021
GLM: General Language Model Pretraining with Autoregressive Blank InfillingarXivStarWebsitearXiv 2021
Learning Transferable Visual Models From Natural Language SupervisionarXivStarWebsiteICML 2021
Zero-Shot Text-to-Image GenerationarXiv--ICML 2021
Exploring the Limits of Transfer Learning with a Unified Text-to-Text TransformerarXivStar-JMLR 2020
BERT: Pre-training of Deep Bidirectional Transformers for Language UnderstandingarXivStarWebsiteNAACL 2019

Implementation

Datasets

More datasets could be found on Pixabay, Mixkit, Pond5, Adobe Stock, Shutterstock, Getty, Coverr, Videvo, Depositphotos, Storyblocks, Dissolve, Freepik, Vimeo, and Envato. Also, there are some datasets at Midjourney V5.1 Cleaned Data, Unsplash-lite, AnimateBench, Pexels-400k, and LAION-AESTHETICS.

Title
arXiv
GitHub
Website
Conference & Year
Panda-70M: Captioning 70M Videos with Multiple Cross-Modality TeachersarXivStarWebsiteCVPR 2024
VBench: Comprehensive Benchmark Suite for Video Generative ModelsarXivStarWebsiteCVPR 2024
InternVid: Learning Text-to-Video Generation from Web-scale Video-Text DataarXivStarWebsiteICLR 2024
MiraData: A Large-Scale Video Dataset with Long Durations and Structured CaptionsarXivStarWebsiteNeurIPS 2024
VidProM: A Million-scale Real Prompt-Gallery Dataset for Text-to-Video Diffusion ModelsarXivStarWebsiteNeurIPS 2024
Vript: A Video Is Worth Thousands of WordsarXivStar-NeurIPS 2024
VideoCrafter2arXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorarXivStarWebsiteICCV 2023
Temporally Consistent Transformers for Video GenerationarXivStarWebsiteICML 2023
Bitstream-Corrupted Video Recovery: A Novel Benchmark Dataset and MethodarXivStar-NeurIPS 2023
FETV: A Benchmark for Fine-Grained Evaluation of Open-Domain Text-to-Video GenerationarXivStar-NeurIPS 2023
AIGCBench: Comprehensive evaluation of image-to-video content generated by AIarXivStarWebsiteTBench 2023
AdaPool: Exponential Adaptive Pooling for Information-Retaining DownsamplingarXivStar-TIP 2023
Swap Attention in Spatiotemporal Diffusions for Text-to-Video GenerationarXivStar-arXiv 2023
Advancing High-Resolution Video-Language Representation with Large-Scale Video TranscriptionsarXivStar-CVPR 2022
The DEVIL is in the Details: A Diagnostic Evaluation Benchmark for Video InpaintingarXivStar-CVPR 2022
VFHQ: A High-Quality Dataset and Benchmark for Video Face Super-ResolutionarXiv-WebsiteCVPR 2022
Learning Audio-Video Modalities from Image CaptionsarXiv-WebsiteECCV 2022
The Anatomy of Video Editing: A Dataset and Benchmark Suite for AI-Assisted Video EditingarXivStar-ECCV 2022
Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingarXivStarWebsiteNeurIPS 2022
Scaling Autoregressive Models for Content-Rich Text-to-Image GenerationarXiv-WebsiteTMLR 2022
ACAV100M: Automatic Curation of Large-Scale Datasets for Audio-Visual Video Representation LearningarXivStarWebsiteICCV 2021
Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalarXivStarWebsiteICCV 2021
MERLOT: Multimodal Neural Script Knowledge ModelsarXivStarWebsiteNeurIPS 2021
Learning Video Representations from Textual Web SupervisionarXiv--arXiv 2020
HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video ClipsarXiv-WebsiteICCV 2019
VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language ResearcharXiv-WebsiteICCV 2019
Towards Automatic Learning of Procedures from Web Instructional VideosarXiv-WebsiteAAAI 2018
How2: A Large-scale Dataset for Multimodal Language UnderstandingarXivStarWebsitearXiv 2018
Quo Vadis, Action Recognition? A New Model and the Kinetics DatasetarXivStar-CVPR 2017
Localizing Moments in Video with Natural LanguagearXivStar-ICCV 2017
MSR-VTT: A Large Video Description Dataset for Bridging Video and Language--WebsiteCVPR 2016
ActivityNet: A Large-Scale Video Benchmark for Human Activity Understanding--WebsiteCVPR 2015
A Dataset for Movie DescriptionarXiv--CVPR 2015
UCF101: A Dataset of 101 Human Actions Classes From Videos in The WildarXiv-WebsitearXiv 2012

Training engineering

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
LLaVA-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal ModelsarXivStarWebsiteICLR 2025
SAM 2: Segment Anything in Images and VideosarXiv-WebsiteICLR 2025
Motion Prompting: Controlling Video Generation with Motion TrajectoriesarXiv-WebsiteCVPR 2025
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
DynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsarXivStarWebsiteECCV 2024
CogVLM2: Visual Language Models for Image and Video UnderstandingarXivStar-arXiv 2024
CogVideoX: Text-to-Video Diffusion Models with an Expert TransformerarXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Hunyuan-dit: A powerful multi-resolution diffusion transformer with fine-grained chinese understandingarXiv-WebsitearXiv 2024
PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense CaptioningarXivStarWebsitearXiv 2024
Cogvideo: Large-scale pretraining for text-to-video generation via transformersarXivStar-ICLR 2023
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large DatasetsarXivStar-arXiv 2023
LLaMA: Open and Efficient Foundation Language ModelsarXivStarWebsitearXiv 2023
LLaMA 2: Open Foundation and Fine-Tuned Chat ModelsarXivStarWebsitearXiv 2023
ST-Adapter: Parameter-Efficient Image-to-Video Transfer LearningarXivStar-NeurIPS 2022
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessarXivStar-NeurIPS 2022
Visual Prompt TuningarXivStar-ECCV 2022
CoCa: Contrastive Captioners are Image-Text Foundation ModelsarXivStar-TMLR 2022
ZeRO: memory optimizations toward training trillion parameter modelsarXiv--Supercomputing 2020

Evaluation metrics and benchmarking findings

Papers are listed generally in reverse order of their publication timestamps.

Industry models

Title
arXiv
GitHub
Website
Conference & Year
Magic 1-for-1: Generating one minute video clips within one minutearXivStarWebsitearXiv 2025
SkyReels v1: Human-centric video foundation modelarXivStar-arXiv 2025
Step-Video-T2VarXiv-WebsitearXiv 2024
HunyuanVideoarXiv-WebsitearXiv 2024
Sora--Website2024
STIVarXivStarWebsitearXiv 2024
LTX-VideoarXivStarWebsitearXiv 2024
AllegroarXivStarWebsitearXiv 2024
JimengarXiv-WebsitearXiv 2024
Mochi 1arXiv-WebsitearXiv 2024
EasyAnimatearXivStarWebsitearXiv 2024
Vidu--Website2024
VideoCrafter2arXivStarWebsitearXiv 2024
VideoCrafter1arXivStarWebsitearXiv 2023
MiraarXiv-WebsitearXiv 2024
Hailuo AI--Website2024
LumierearXiv-WebsitearXiv 2024
VideoPoetarXiv-WebsitearXiv 2023
LumaAI Ray 2--Website2024
LumaAI Dream Machine--Website2023
Veo-2--Website2024
Veo-1--Website2023
Nova Real--Website2024
Wanx 2.1--Website2024
Kling--Website2024
Show-1arXivStarWebsiteNeurIPS 2023
MovieGenarXiv-WebsitearXiv 2024
Pika--Website2023
Vchitect-2.0--Website2024
OptisarXivStarWebsiteNeurIPS 2023
VLoggerarXivStarWebsiteICCV 2023
SeinearXivStarWebsiteCVPR 2023
LaviearXivStarWebsiteICCV 2023
MiracleVision--Website2023
PhenakiarXivStarWebsiteICLR 2024
W.A.L.TarXiv-WebsitearXiv 2024
Imagen videoarXiv-Website2022
GEN-3 Alpha--Website2024
GEN-2--Website2023
GEN-1--Website2022

Academia models

Title
arXiv
GitHub
Website
Conference & Year
RepVideo: Rethinking Cross-Layer Representation for Video GenerationarXivStarWebsitearXiv 2025
CausVid: Causality-Aware Video Generation with Slow-Fast Diffusion ModelsarXivStarWebsiteCVPR 2025
Open-Sora Plan: Open-Source Large Video Generation ModelarXivStarWebsitearXiv 2024
Open-Sora: Democratizing Efficient Video Production for AllarXivStarWebsitearXiv 2024
Snap Video: Scaled Spatiotemporal Transformers for Text-to-Video SynthesisarXiv-WebsitearXiv 2024
SiT: Exploring Flow and Diffusion-Based Generative Models with Scalable Interpolant TransformersarXivStarWebsitearXiv 2024
VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-Guided PlanningarXivStarWebsiteCOLM 2024
AnimateLCM: Accelerating the Animation of Personalized Diffusion Models and Adapters with Decoupled Consistency LearningarXivStarWebsitearXiv 2024
I4VGEN: Interactive Video Generation via Integrated Dynamic ControlarXivStarWebsitearXiv 2024
SimDA: Simple Diffusion Adapter for Efficient Text-to-Video GenerationarXivStarWebsitearXiv 2023
AnimateDiff-v2arXivStarWebsiteICLR 2024
Animate-A-Story: Storytelling with Retrieval-Augmented Video GenerationarXivStarWebsitearXiv 2023
VideoGen: A Reference-Guided Latent Diffusion Approach for High-Definition Text-to-Video GenerationarXiv-WebsitearXiv 2023
Dysen-VDM: Diffusion Model with Dynamic Spatio-Temporal Fusion for Video GenerationarXivStar-arXiv 2023
HiGen: Hierarchical 3D Feature Generation for 3D-Aware Image Synthesis and ManipulationarXivStar-arXiv 2023
ModelScope Text-to-Video Technical ReportarXiv-WebsitearXiv 2023
InstructVideo: Instructing Video Diffusion Models with Human FeedbackarXiv-WebsiteCVPR 2024
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNeurIPS 2023
VideoFusion: Decomposed Diffusion Models for High-Quality Video GenerationarXiv--CVPR 2023
MagViT-v2: Masked Generative Video TransformerarXivStar-arXiv 2023
MagViT: Masked Generative Video TransformerarXivStar-arXiv 2022
Latent-Shift: Latent Diffusion with Temporal Shift for Efficient Text-to-Video GenerationarXiv-WebsitearXiv 2023
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion ModelsarXivStarWebsiteCVPR 2023
Video Diffusion ModelsarXivStarWebsitearXiv 2022
Make-A-Video: Text-to-Video Generation without Text-Video DataarXivStarWebsiteICLR 2023
MagicVideo: Efficient Video Generation With Latent Diffusion ModelsarXiv-WebsitearXiv 2022
CogVideoX: Enhancing Video Understanding in the Era of Large Language ModelsarXivStarWebsitearXiv 2024
CogVideo: Large-scale Pretraining for Text-to-Video Generation via TransformersarXivStarWebsiteICLR 2023
VideoGPT: Video Generation using VQ-VAE and TransformersarXivStarWebsitearXiv 2021

Applications

Conditions

Image condition

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerarXivStarWebsiteICLR 2025
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion ControlarXiv-WebsiteICLR 2025
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text GuidancearXivStarWebsiteICASSP 2025
EMO: Emote Portrait Alive-Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak ConditionsarXiv-WebsiteECCV 2024
Cinemo: Consistent and Controllable Image Animation with Motion Diffusion ModelsarXivStarWebsiteCVPR 2025
MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and EditingarXivStarWebsiteECCV 2024
ConsistI2V: Enhancing Visual Consistency for Image-to-Video GenerationarXivStarWebsiteTMLR
I2V-Adapter: A General Image-to-Video Adapter for Diffusion ModelsarXivStarWebsiteSIGGRAPH 2024
ID-Animator: Zero-Shot Identity-Preserving Human Video GenerationarXivStarWebsitearXiv 2024
CamCo: Camera-Controllable 3D-Consistent Image-to-Video GenerationarXiv-WebsitearXiv 2024
Generative Image DynamicsarXiv-WebsiteCVPR 2024
PIA: Your Personalized Image Animator via Plug-and-Play Modules in Text-to-Image ModelsarXivStarWebsiteCVPR 2024
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion ModelsarXiv-WebsiteCVPR 2024
AtomoVideo: High Fidelity Image-to-Video GenerationarXiv-WebsitearXiv 2024
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modelingarXivStarWebsiteSIGGRAPH 2024
Seer: Language Instructed Video Prediction with Latent Diffusion ModelsarXivStarWebsiteICLR 2024
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion GuidancearXivStarWebsitearXiv 2023
VideoBooth: Diffusion-based Video Generation with Image PromptsarXivStarWebsiteCVPR 2024
Sparsectrl: Adding sparse controls to text-to-video diffusion modelsarXivStarWebsiteECCV 2024
DynamiCrafter: Animating Open-domain Images with Video Diffusion PriorsarXivStarWebsiteECCV 2024
Adding Conditional Control to Text-to-Image Diffusion ModelsarXivStarWebsiteICCV 2023
Stable video diffusion: Scaling latent video diffusion models to large datasetsarXivStarWebsiteArxix 2023
Make pixels dance: High-dynamic video generationarXiv-WebsiteCVPR 2024
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion ModelsarXivStarWebsitearXiv 2023
Videocrafter1: Open diffusion models for high-quality video generationarXivStarWebsitearXiv 2023
VDT: General-purpose Video Diffusion Transformers via Mask ModelingarXivStarWebsiteICLR 2024
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNIPS 2023
Conditional Image-to-Video Generation with Latent Flow Diffusion ModelsarXivStarWebsiteCVPR 2023

Spatial condition

Papers are listed generally in reverse order of their publication timestamps.

Camera parameter condition

Papers are listed generally in reverse order of their publication timestamps.

Audio condition

Papers are listed generally in reverse order of their publication timestamps.

High-level video condition

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
TokenFlow: Unified Image Tokenizer for Multimodal Understanding and GenerationarXivStarWebsiteICLR 2024
MotionClone: Training-Free Motion Cloning for Controllable Video GenerationarXivStarWebsiteICLR 2025
I2VEdit: First-Frame-Guided Video Editing via Image-to-Video Diffusion ModelsarXivStarWebsiteSIGGRAPH Asia 2024
ReVideo: Remake a Video with Motion and Content ControlarXivStarWebsiteNeurIPS 2024
AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingarXivStarWebsiteACM MM 2024
AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing TasksarXivStarWebsiteTMLR 2024
UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingarXivStarWebsitearXiv 2024
VidToMe: Video Token Merging for Zero-Shot Video EditingarXivStarWebsiteCVPR 2024
FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisarXiv-WebsiteCVPR 2024
SAVE: Protagonist Diversification with Structure Agnostic Video EditingarXivStarWebsiteECCV 2024
RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion ModelsarXivStarWebsiteCVPR 2024
DiffusionAtlas: High-Fidelity Consistent Diffusion Video EditingarXiv-WebsitearXiv 2023
DragVideo: Interactive Drag-style Video EditingarXivStarWebsiteECCV 2024
Drag-A-Video: Non-rigid Video Editing with Point-based InteractionarXivStarWebsitearXiv 2023
VideoSwap: Customized Video Subject Swapping with Interactive Semantic Point CorrespondencearXivStarWebsiteCVPR 2024
A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video EditingarXivStarWebsiteCVPR 2024
Motion-Conditioned Image Animation for Video EditingarXivStarWebsitearXiv 2023
MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware DiffusionarXivStarWebsiteICML 2024
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character AnimationarXivStarWebsiteCVPR 2024
Consistent Video-to-Video Transfer Using Synthetic DatasetarXivStarWebsiteICLR 2024
MotionDirector: Motion Customization of Text-to-Video Diffusion ModelsarXivStarWebsiteECCV 2024
SimDA: Simple Diffusion Adapter for Efficient Video GenerationarXivStarWebsiteCVPR 2024
MagicEdit: High-Fidelity and Temporally Coherent Video EditingarXivStarWebsitearXiv 2023
CoDeF: Content Deformation Fields for Temporally Consistent Video ProcessingarXivStarWebsiteCVPR 2024
StableVideo: Text-driven Consistency-aware Diffusion Video EditingarXivStarWebsiteICCV 2023
VideoControlNet: A Motion-Guided Video-to-Video Translation Framework by Using Diffusion Model with ControlNetarXivStarWebsitearXiv 2023
VideoComposer: Compositional Video Synthesis with Motion ControllabilityarXivStarWebsiteNeurIPS 2023
Rerender A Video: Zero-Shot Text-Guided Video-to-Video TranslationarXivStarWebsiteSIGGRAPH Asia 2023
Video Colorization with Pre-trained Text-to-Image Diffusion ModelsarXivStarWebsitearXiv 2023
VidEdit: Zero-Shot and Spatially Aware Text-Driven Video EditingarXiv-WebsiteTMLR 2024
DisCo: Disentangled Control for Realistic Human Dance GenerationarXivStarWebsiteCVPR 2024
Towards Consistent Video Editing with Text-to-Image Diffusion ModelsarXiv--NeurIPS 2023
Video ControlNet: Towards Temporally Consistent Synthetic-to-Real Video Translation Using Conditional Image Diffusion ModelsarXiv---
ControlVideo: Training-free Controllable Text-to-Video GenerationarXivStarWebsiteICLR 2024
InstructVid2Vid: Controllable Video Editing with Natural Language InstructionsarXiv--arXiv 2023
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free VideosarXivStarWebsiteAAAI 2024
DreamPose: Fashion Image-to-Video Synthesis via Stable DiffusionarXivStarWebsiteICCV 2023
Zero-Shot Video Editing Using Off-The-Shelf Image Diffusion ModelsarXivStar-IEEE Trans On Multimedia, 2023
Pix2Video: Video Editing using Image DiffusionarXivStarWebsiteICCV 2023
Structure and Content-Guided Video Synthesis with Diffusion ModelsarXiv-WebsiteICCV 2023
Shape-aware Text-driven Layered Video EditingarXivStarWebsiteCVPR 2023
DPE: Disentanglement of Pose and Expression for General Video Portrait EditingarXivStarWebsiteCVPR 2023
Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video GenerationarXivStarWebsiteICCV 2023
Diffusion Video Autoencoders: Toward Temporally Consistent Face Video Editing via Disentangled Video EncodingarXivStarWebsiteCVPR 2023
Layered Neural Atlases for Consistent Video EditingarXivStarWebsiteSIGGRAPH Asia 2021

Other conditions

Papers are listed generally in reverse order of their publication timestamps.

Enhancement

Video denoising and deblurring

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Video restoration based on deep learning: a comprehensive survey---2022

Video inpainting

Papers are listed generally in reverse order of their publication timestamps.

Video interpolation and extrapolation/prediction

Papers are listed generally in reverse order of their publication timestamps.

Video super-resolution

Papers are listed generally in reverse order of their publication timestamps.

Combining multiple video enhancement tasks

Papers are listed generally in reverse order of their publication timestamps.

Personalization

Papers are listed generally in reverse order of their publication timestamps.

Title
arXiv
GitHub
Website
Conference & Year
Dynamic Concepts Personalization from Single VideosarXiv-WebsiteSIGGRAPH 2025
VideoAlchemy: Open-set Personalization in Video GenerationarXivStarWebsiteCVPR 2025
[PersonalVideo: High ID-Fidelity V

Truncated — view the full README on GitHub.

Contributors

yimuwangcs

29 commits

weipang142857

18 commits

ningyu1991

12 commits

limacv

5 commits