[ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
774
174 commits
updated Jul 31, 2026
πππA curated list of papers on controllable video generation. Please join us for more comprehensive summary. If you have any additions to the list, please raise them in the issue section. ζ¬’θΏθ‘₯ε π
If you find this repo is useful for your research, welcome to π this repo and cite our work using the following BibTeX:
@article{ma2025controllable,
title={Controllable video generation: A survey},
author={Ma, Yue and Feng, Kunyu and Hu, Zhongyuan and Wang, Xinyu and Wang, Yucheng and Zheng, Mingzhe and He, Xuanhua and Zhu, Chenyang and Liu, Hongyu and He, Yingqing and others},
journal={arXiv preprint arXiv:2507.16869},
year={2025}
}
ctrl + F and then type the author name. The dropdown list of authors will automatically expand when searching.DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes (23 Oct 2024)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures (10 Feb 2026)
MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation (30 May 2025)
HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions (29 May 2025)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios (6 May 2025)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation (14 Mar 2025)
Point-to-Point Video Generation (14 Mar 2025)
Pose Guided Human Video Generation (14 Mar 2025)
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers (15 Feb 2025)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance (12 Feb 2025)
Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance (10 Feb 2025)
DirectorLLM for Human-Centric Video Generation (19 Dec 2024)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation (16 Dec 2024)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation (12 Dec 2024)
StableAnimator: High-Quality Identity-Preserving Human Image Animation (26 Nov 2024)
AnimateAnywhere: Context-Controllable Human Video Generation with ID-Consistent One-shot Learning (28 Oct 2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation (12 Aug 2024)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models (12 Jul 2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance (28 Jun 2024)
Follow-Your-Pose v2: Multiple-Condition Guided Character Image Animation for Stable Pose Control (05 Jun 2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation (3 Jun 2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation (28 May 2024)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation (26 May 2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control (23 May 2024)
Zero-shot High-fidelity and Pose-controllable Character Animation (21 Apr 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (21 Mar 2024)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons (24 Jan 2024)
DreaMoving: A Human Video Generation Framework based on Diffusion Models (8 Dec 2023)
Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models (3 Dec 2023)
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation (28 Nov 2023)
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model (27 Nov 2023)
MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion (18 Nov 2023)
Dancing Avatar: Pose and Text-Guided Human Motion Videos Synthesis with Image Diffusion Model (15 Aug 2023)
DISCO: Disentangled Control for Realistic Human Dance Generation (30 Jun 2023)
DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion (12 Apr 2023)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos (3 Apr 2023)
vid2vid:Video-to-Video Synthesis (3 Dec 2018)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses (30 Nov 2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation (12 Aug 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (21 Mar 2024)
MoonShot: Towards Controllable Video Generation and Editing with Multimodal Conditions (3 Jan 2024)
SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models (28 Nov 2023)
GD-VDM: Generated Depth for better Diffusion-based Video Generation (19 Jun 2023)
Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance (1 Jun 2023)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning (23 May 2023)
ControlVideo: Training-free Controllable Text-to-Video Generation (22 May 2023)
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation (24 Mar 2024)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation (14 Mar 2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition (26 Nov 2024)
EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation (15 Nov 2024)
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization (18 Oct 2024)
EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions (12 Jul 2024)
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control (3 Jul 2024)
Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation (4 Jun 2024)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos (3 Apr 2023)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
SketchVideo: Sketch-based Video Generation and Editing (30 Mar 2025)
LayerAnimate: Layer-level Control for Animation (22 Mar 2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models (13 Mar 2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation (13 Mar 2025)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control (17 Feb 2025)
AniDoc: Animation Creation Made Easier (30 Jan 2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training (8 Dec 2024)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model (4 Dec 2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories(frame+track+text) (3 Dec 2024)
MOTIONFLOW: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation (Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
Trajectory Attention: Enhancing Video Generation with Fine-Grained Motion Control (28 Nov 2024)
Open-Sora Plan: Open-Source Large Video Generation Model (28 Nov 2024)
ToonCrafter: Generative Cartoon Interpolation (19 Nov 2024)
MagicStick: Controllable Video Editing via Control Handle Transformations (18 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion (7 Nov 2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation (22 Oct 2024)
Boosting Camera Motion Control for Video Diffusion Transformers (14 Oct 2024)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention (14 Oct 2024)
LVCD: Reference-based Lineart Video Colorization with Diffusion Models (19 Sep 2024)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation (16 Sep 2024)
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis (3 Sep 2024)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (17 Jul 2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation (16 Jul 2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model (24 May 2024)
A Recipe for Scaling up Text-to-Video Generation with Text-free Videos (25 Dec 2023)
VideoLCM: Video Latent Consistency Model (14 Dec 2023)
SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models (28 Nov 2023)
Make Pixels Dance: High-Dynamic Video Generation (18 Nov 2023)
TaleCrafter: Interactive Story Visualization with Multiple Characters (30 May 2023)
Sketching the Future(STF): Applying Conditional Control Techniques to Text-to-Video Models (10 May 2023)
SketchBetween: Video-to-Video Synthesis for Sprite Animation via Sketches (1 Sep 2022)
Sketch Me A Video (10 Oct 2021)
vid2vid:Video-to-Video Synthesis (3 Dec 2018)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving (02 Dec 2024)
Zehuan Wu, Jingcheng Ni, Xiaodong Wang, Yuxin Guo, Rui Chen, Lewei Lu, Jifeng Dai, Yuwen Xiong
Higher fidelity autonomous vehicle video generation with bounding-box controlled object motion (8 Dec 2024 )
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control (21 Nov 2024)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes (06 Sep 2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation (09 Sep 2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation (09 Sep 2024)
DiVE: DiT-based Video Generation with Enhanced Control (03 Sep 2024)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation (3 Jun 2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis (2 Feb 2024)
Panacea: Panoramic and Controllable Video Generation for Autonomous Driving (28 Nov 2023)
MagicDrive: Street View Generation with Diverse 3D Geometry Control (04 Oct 2023)
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving (18 Sep 2023)
LLM-grounded Video Diffusion Models (29 Sep 2023)
Multi-object Video Generation from Single Frame Layouts (06 May 2023)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation(30 May 2025)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation (27 May 2025)
SkyReels-A2: Compose Anything in Video Diffusion Transformers (3 Apr 2025)
AnimateAnywhere: Rouse the Background in Human Image Animation (28 Apr 2025)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis (19 Apr 2025)
PERSONALVIDEO: High ID-Fidelity Video Customization with Static Images (16 Mar 2025)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance (12 Feb 2025)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts (4 Feb 2025)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion (23 Jan 2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention (03 Dec 2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition (25 Nov 2024)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling (24 Sep 2024)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation (25 Jun 2024)
Magic-Me: Identity-Specific Video Customized Diffusion (14 Feb 2024)
Vlogger: Make Your Dream A Vlog (17 Jan 2024)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis (19 Apr 2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models (27 Mar 2025)
DreamRelation: Relation-Centric Video Customization (10 Mar 2025)
Get In Video: Add Anything You Want to the Video (8 Mar 2025)
Phantom: Subject-consistent Video Generation via Cross-modal Alignment (16 Feb 2025)
Multi-subject Open-set Personalization in Video Generation (10 Jan 2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning (8 Jan 2025)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models (27 Dec 2024)
Customcrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities (27 Dec 2024)
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training (20 Dec 2024)
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner (13 Dec 2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation (25 Nov 2024)
VideoAlchemy: Open-set Personalization in Video Generation (15 Nov 2024)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration (7 Nov 2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation (29 Oct 2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation (19 Aug 2024)
Still-Moving: Customized Video Generation Without Customized Video Data (11 Jul 2024)
VIMI: Grounding Video Generation through Multi-modal Instruction (8 Jul 2024)
Customvideo: Customizing Text-to-Video Generation with Multiple Subjects (22 May 2024)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control (21 May 2024)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (9 May 2024)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control (28 Mar 2024)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production (12 Mar 2024)
VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM (2 Jan 2024)
Dreamvideo: Composing Your Dream Videos with Customized Subject and Motion (7 Dec 2023)
VideoBooth: Diffusion-based Video Generation with Image Prompts (1 Dec 2023)
Videodreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning (2 Nov 2023)
Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation (13 Jul 2023)
TaleCrafter: Interactive Story Visualization with Multiple Characters (30 May 2023)
Dreamix: Video Diffusion Models are General Video Editors (2 Feb 2023)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
HunyuanVideo: A Systematic Framework For Large Video Generative Models (11 Mar 2025)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think (2 Mar 2025)
Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation (17 Feb 2025)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation (12 Feb 2025)
Autoregressive Video Generation without Vector Quantization (9 Jan 2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation (6 Jan 2025)
STIV: Scalable Text and Image Conditioned Video Generation (10 Dec 2024)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation (8 Dec 2024)
Lumiere: A Space-Time Diffusion Model for Video Generation (3 Dec 2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model (6 Nov 2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models (20 Oct 2024)
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation (3 Oct 2024)
DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors (1 Oct 2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (27 Sep 2024)
Structure and Content-Guided Video Synthesis with Diffusion Models (27 Sep 2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance (16 Sep 2024)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning (2 Aug 2024)
MoVideo: Motion-Aware Video Generation with Diffusion Model (29 Jul 2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models (13 Jul 2024)
EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture (5 Jul 2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation (1 Jul 2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation (13 Jun 2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction (10 Jun 2024)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models (25 Apr 2024)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models (25 Mar 2024)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts (13 Mar 2024)
AtomoVideo: High Fidelity Image-to-Video Generation (5 Mar 2024)
Tuning-Free Noise Rectification for High Fidelity Image-to-Video Generation (5 Mar 2024)
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling (31 Jan 2024)
UniVG: Towards UNIfied-modal Video Generation (17 Jan 2024)
Decouple Content and Motion for Conditional Image-to-Video Generation (14 Dec 2023)
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance (4 Dec 2023)
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (25 Nov 2023)
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models (7 Nov 2023)
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation (30 Oct 2023)
Synthesizing Videos from Images for Image-to-Video Adaptation (27 Oct 2023)
VideoDoodles: Hand-Drawn Animations on Videos with Scene-Aware Canvases (26 Jul 2023)
LaMD: Latent Motion Diffusion for Video Generation (23 Apr 2023)
Prompt Image to Life: Training-Free Text-Driven Image-to-Video Generation (2023)
Make It Move: Controllable Image-to-Video Generation With Text Descriptions (31 Mar 2022)
VideoGPT: Video Generation using VQ-VAE and Transformers (14 Sep 2021)
ImaGINator: Conditional Spatio-Temporal GAN for Video Generation (2020)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model (30 May 2024)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling (29 Jan 2024)
Motion-Conditioned Diffusion Model for Controllable Video Synthesis (27 Apr 2023)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (26 Nov 2024)
TrajectoryMover: Generative Movement of Object Trajectories in Videos (31 Mar 2026)
DiTraj: Training-free Trajectory Control For Video Diffusion Transformer (29 Sep 2025)
ATI: Any Trajectory Instruction for Controllable Video Generation (10 Jun 2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation (29 May 2025)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation (27 May 2025)
AnimateAnywhere: Rouse the Background in Human Image Animation (28 Apr 2025)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation (21 Apr 2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation (2 Apr 2025)
LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis (28 Mar 2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance (20 Mar 2025)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation (14 Mar 2025)
LayerAnimate: Layer-level Control for Animation (22 Mar 2025)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (22 Mar 2025)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (10 Mar 2025)
C-Drag: Chain-of-Thought Driven Motion Controller for Video Generation (27 Feb 2025)
SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation (25 Feb 2025)
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation (6 Feb 2025)
MotionBridge: Dynamic Video Inbetweening with Flexible Controls (7 Jan 2025)
TrackGo: A Flexible and Efficient Method for Controllable Video Generation (5 Jan 2025)
Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object (8 Jan 2025)
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation (12 Dec 2024)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses (10 Dec 2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories (3 Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
InTraGen: Trajectory-controlled Video Generation for Object Interactions (25 Nov 2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation (29 Oct 2024)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships (14 Oct 2024)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model (11 Jul 2024)
Freetraj: Tuning-free Trajectory Control in Video Diffusion Models (T2V) (24 Jun 2024)
Image Conductor: Precision Control for Interactive Video Synthesis (21 Jun 2024)
ReVideo: Remake a Video with Motion and Content (22 May 2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control (27 May 2024)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion (6 May 2024)
Peekaboo: Interactive Video Generation via Masked-Diffusion (T2V) (19 Apr 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (19 Apr 2024)
TrailBlazer: Trajectory Control for Diffusion-Based Video Generation (T2V) (8 Apr 2024)
Draganything: Motion Control for Anything Using Entity Representation (15 Mar 2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis (2 Feb 2024)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling (31 Jan 2024)
DragNuwa (Image+Text_Traj) (16 Aug 2023)
MCDiff Motion-Conditioned Diffusion Model for Controllable Video Synthesis (27 Apr 2023)
Controllable Video Generation With Sparse Trajectories (16 Dec 2018)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting (5 Jun 2025)
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation (4 Jun 2025)
AKiRa: Augmentation Kit on Rays for optical video generation (Jun 2025)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation (21 Apr 2025)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography (10 Apr 2025)
OmniCam: Unified Multimodal Video Generation via Camera Control (3 Apr 2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation (2 Apr 2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories (27 Mar 2025)
Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis (25 Mar 2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers (22 Mar 2025)
Aether: Geometric-Aware Unified World Modeling (18 Mar 2025)
EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning (16 Mar 2025)
I2V3D: Controllable Image-to-Video Generation with 3D Guidance (12 Mar 2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models (13 Mar 2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation (13 Mar 2025)
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video (14 Mar 2025)
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control (5 Mar 2025)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (10 Mar 2025)
I2VCONTROL-CAMERA: Precise Video Camera Control with Adjustable Motion Strength (28 Feb 2025)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation (12 Feb 2025)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion (12 Feb 2025)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control (14 Feb 2025)
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation (7 Feb 2025)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale (10 Dec 2024)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training (8 Dec 2024)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model (4 Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
Trajectory Attention: Enhancing Video Generation with Fine-Grained Motion Control (28 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion (7 Nov 2024)
Boosting Camera Motion Control for Video Diffusion Transformers (14 Oct 2024)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention (14 Oct 2024)
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis (3 Sep 2024)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (17 Jul 2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation (16 Jul 2024)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation (4 Jun 2024)
MotionMaster: Training-free Camera Motion Transfer For Video Generation (1 May 2024)
MOTIONFLOW: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation (Dec 2024)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios (6 May 2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning (5 Jun 2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation (29 May 2025)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer (15 Apr 2025)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models(hand mask sequence as control signal) (4 Apr 2025)
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance (3 Apr 2025)
Video Motion Transfer with Diffusion Transformers (27 Mar 2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories (27 Mar 2025)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models (25 Mar 2025)
DreamRelation: Relation-Centric Video Customization (10 Mar 2025)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching (18 Feb 2025)
MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent (5 Feb 2025)
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss (13 Jan 2025)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing (13 Jan 2025)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control (9 Jan 2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models (19 Dec 2024)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models (2 Dec 2024)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion ModelsOnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models (15 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
AnimateAnything: Consistent and Controllable Animation for Video Generation (16 Nov 2024)
Motionbooth: Motion-aware customized text-to-video generation (29 Oct 2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation (22 Oct 2024)
Motion Inversion for Video Customization (16 Oct 2024)
Zero-Shot Controllable Image-to-Video Animation via Motion Decomposition (21 Jul 2024)
DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing (15 Jul 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (1 Jun 2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models (28 Aug 2024)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (7 Dec 2023)
DreamVideo: Composing Your Dream Videos with Customized Subject and Motion (7 Dec 2023)
VMC: Video Motion Customization with Pre-trained Diffusion Models (1 Dec 2023)
Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer (3 Dec 2023)
MotionDirector: Motion Customization for Text-to-Video Diffusion Models (12 Oct 2023)
Truncated β view the full README on GitHub.
[ArXiv 2025] A survey about controllable video generation: This repo is the official awesome of "Controllable video generation: A survey"
774
174 commits
updated Jul 31, 2026
πππA curated list of papers on controllable video generation. Please join us for more comprehensive summary. If you have any additions to the list, please raise them in the issue section. ζ¬’θΏθ‘₯ε π
If you find this repo is useful for your research, welcome to π this repo and cite our work using the following BibTeX:
@article{ma2025controllable,
title={Controllable video generation: A survey},
author={Ma, Yue and Feng, Kunyu and Hu, Zhongyuan and Wang, Xinyu and Wang, Yucheng and Zheng, Mingzhe and He, Xuanhua and Zhu, Chenyang and Liu, Hongyu and He, Yingqing and others},
journal={arXiv preprint arXiv:2507.16869},
year={2025}
}
ctrl + F and then type the author name. The dropdown list of authors will automatically expand when searching.DynamicCity: Large-Scale 4D Occupancy Generation from Dynamic Scenes (23 Oct 2024)
Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures (10 Feb 2026)
MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation (30 May 2025)
HyperMotion: DiT-Based Pose-Guided Human Image Animation of Complex Motions (29 May 2025)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios (6 May 2025)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation (14 Mar 2025)
Point-to-Point Video Generation (14 Mar 2025)
Pose Guided Human Video Generation (14 Mar 2025)
SkyReels-A1: Expressive Portrait Animation in Video Diffusion Transformers (15 Feb 2025)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance (12 Feb 2025)
Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance (10 Feb 2025)
DirectorLLM for Human-Centric Video Generation (19 Dec 2024)
Generative Inbetweening through Frame-wise Conditions-Driven Video Generation (16 Dec 2024)
DisPose: Disentangling Pose Guidance for Controllable Human Image Animation (12 Dec 2024)
StableAnimator: High-Quality Identity-Preserving Human Image Animation (26 Nov 2024)
AnimateAnywhere: Context-Controllable Human Video Generation with ID-Consistent One-shot Learning (28 Oct 2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation (12 Aug 2024)
TCAN: Animating Human Images with Temporally Consistent Pose Guidance using Diffusion Models (12 Jul 2024)
MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance (28 Jun 2024)
Follow-Your-Pose v2: Multiple-Condition Guided Character Image Animation for Stable Pose Control (05 Jun 2024)
UniAnimate: Taming Unified Video Diffusion Models for Consistent Human Image Animation (3 Jun 2024)
VividPose: Advancing Stable Video Diffusion for Realistic Human Image Animation (28 May 2024)
Disentangling Foreground and Background Motion for Enhanced Realism in Human Video Generation (26 May 2024)
PoseCrafter: One-Shot Personalized Video Synthesis Following Flexible Pose Control (23 May 2024)
Zero-shot High-fidelity and Pose-controllable Character Animation (21 Apr 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (21 Mar 2024)
Do You Guys Want to Dance: Zero-Shot Compositional Human Dance Generation with Multiple Persons (24 Jan 2024)
DreaMoving: A Human Video Generation Framework based on Diffusion Models (8 Dec 2023)
Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models (3 Dec 2023)
Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation (28 Nov 2023)
MagicAnimate: Temporally Consistent Human Image Animation using Diffusion Model (27 Nov 2023)
MagicPose: Realistic Human Poses and Facial Expressions Retargeting with Identity-aware Diffusion (18 Nov 2023)
Dancing Avatar: Pose and Text-Guided Human Motion Videos Synthesis with Image Diffusion Model (15 Aug 2023)
DISCO: Disentangled Control for Realistic Human Dance Generation (30 Jun 2023)
DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion (12 Apr 2023)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos (3 Apr 2023)
vid2vid:Video-to-Video Synthesis (3 Dec 2018)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
DreamDance: Animating Human Images by Enriching 3D Geometry Cues from 2D Poses (30 Nov 2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation (12 Aug 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (21 Mar 2024)
MoonShot: Towards Controllable Video Generation and Editing with Multimodal Conditions (3 Jan 2024)
SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models (28 Nov 2023)
GD-VDM: Generated Depth for better Diffusion-based Video Generation (19 Jun 2023)
Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance (1 Jun 2023)
Control-A-Video: Controllable Text-to-Video Diffusion Models with Motion Prior and Reward Feedback Learning (23 May 2023)
ControlVideo: Training-free Controllable Text-to-Video Generation (22 May 2023)
HunyuanPortrait: Implicit Condition Control for Enhanced Portrait Animation (24 Mar 2024)
TASTE-Rob: Advancing Video Generation of Task-Oriented Hand-Object Interaction for Generalizable Robotic Manipulation (14 Mar 2025)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition (26 Nov 2024)
EchoMimicV2: Towards Striking, Simplified, and Semi-Body Human Animation (15 Nov 2024)
Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss Optimization (18 Oct 2024)
EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions (12 Jul 2024)
LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control (3 Jul 2024)
Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation (4 Jun 2024)
Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos (3 Apr 2023)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
SketchVideo: Sketch-based Video Generation and Editing (30 Mar 2025)
LayerAnimate: Layer-level Control for Animation (22 Mar 2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models (13 Mar 2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation (13 Mar 2025)
VidSketch: Hand-drawn Sketch-Driven Video Generation with Diffusion Control (17 Feb 2025)
AniDoc: Animation Creation Made Easier (30 Jan 2025)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training (8 Dec 2024)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model (4 Dec 2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories(frame+track+text) (3 Dec 2024)
MOTIONFLOW: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation (Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
Trajectory Attention: Enhancing Video Generation with Fine-Grained Motion Control (28 Nov 2024)
Open-Sora Plan: Open-Source Large Video Generation Model (28 Nov 2024)
ToonCrafter: Generative Cartoon Interpolation (19 Nov 2024)
MagicStick: Controllable Video Editing via Control Handle Transformations (18 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion (7 Nov 2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation (22 Oct 2024)
Boosting Camera Motion Control for Video Diffusion Transformers (14 Oct 2024)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention (14 Oct 2024)
LVCD: Reference-based Lineart Video Colorization with Diffusion Models (19 Sep 2024)
EasyControl: Transfer ControlNet to Video Diffusion for Controllable Generation and Interpolation (16 Sep 2024)
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis (3 Sep 2024)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (17 Jul 2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation (16 Jul 2024)
Ctrl-Adapter: An Efficient and Versatile Framework for Adapting Diverse Controls to Any Diffusion Model (24 May 2024)
A Recipe for Scaling up Text-to-Video Generation with Text-free Videos (25 Dec 2023)
VideoLCM: Video Latent Consistency Model (14 Dec 2023)
SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models (28 Nov 2023)
Make Pixels Dance: High-Dynamic Video Generation (18 Nov 2023)
TaleCrafter: Interactive Story Visualization with Multiple Characters (30 May 2023)
Sketching the Future(STF): Applying Conditional Control Techniques to Text-to-Video Models (10 May 2023)
SketchBetween: Video-to-Video Synthesis for Sprite Animation via Sketches (1 Sep 2022)
Sketch Me A Video (10 Oct 2021)
vid2vid:Video-to-Video Synthesis (3 Dec 2018)
HoloDrive: Holistic 2D-3D Multi-Modal Street Scene Generation for Autonomous Driving (02 Dec 2024)
Zehuan Wu, Jingcheng Ni, Xiaodong Wang, Yuxin Guo, Rui Chen, Lewei Lu, Jifeng Dai, Yuwen Xiong
Higher fidelity autonomous vehicle video generation with bounding-box controlled object motion (8 Dec 2024 )
MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control (21 Nov 2024)
DreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Scenes (06 Sep 2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation (09 Sep 2024)
DriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generation (09 Sep 2024)
DiVE: DiT-based Video Generation with Enhanced Control (03 Sep 2024)
Unleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video Generation (3 Jun 2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis (2 Feb 2024)
Panacea: Panoramic and Controllable Video Generation for Autonomous Driving (28 Nov 2023)
MagicDrive: Street View Generation with Diverse 3D Geometry Control (04 Oct 2023)
DriveDreamer: Towards Real-world-driven World Models for Autonomous Driving (18 Sep 2023)
LLM-grounded Video Diffusion Models (29 Sep 2023)
Multi-object Video Generation from Single Frame Layouts (06 May 2023)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
MTVCrafter: 4D Motion Tokenization for Open-World Human Image Animation(30 May 2025)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation (27 May 2025)
SkyReels-A2: Compose Anything in Video Diffusion Transformers (3 Apr 2025)
AnimateAnywhere: Rouse the Background in Human Image Animation (28 Apr 2025)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis (19 Apr 2025)
PERSONALVIDEO: High ID-Fidelity Video Customization with Static Images (16 Mar 2025)
AnyCharV: Bootstrap Controllable Character Video Generation with Fine-to-Coarse Guidance (12 Feb 2025)
Movie Weaver: Tuning-Free Multi-Concept Video Personalization with Anchored Prompts (4 Feb 2025)
EchoVideo: Identity-Preserving Human Video Generation by Multimodal Feature Fusion (23 Jan 2025)
VideoGen-of-Thought: Step-by-step generating multi-shot video with minimal manual intervention (03 Dec 2024)
Identity-Preserving Text-to-Video Generation by Frequency Decomposition (25 Nov 2024)
MIMO: Controllable Character Video Synthesis with Spatial Decomposed Modeling (24 Sep 2024)
ID-Animator: Zero-Shot Identity-Preserving Human Video Generation (25 Jun 2024)
Magic-Me: Identity-Specific Video Customized Diffusion (14 Feb 2024)
Vlogger: Make Your Dream A Vlog (17 Jan 2024)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
Concat-ID: Towards Universal Identity-Preserving Video Synthesis (19 Apr 2025)
VideoMage: Multi-Subject and Motion Customization of Text-to-Video Diffusion Models (27 Mar 2025)
DreamRelation: Relation-Centric Video Customization (10 Mar 2025)
Get In Video: Add Anything You Want to the Video (8 Mar 2025)
Phantom: Subject-consistent Video Generation via Cross-modal Alignment (16 Feb 2025)
Multi-subject Open-set Personalization in Video Generation (10 Jan 2025)
ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning (8 Jan 2025)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models (27 Dec 2024)
Customcrafter: Customized Video Generation with Preserving Motion and Concept Composition Abilities (27 Dec 2024)
CustomTTT: Motion and Appearance Customized Video Generation via Test-Time Training (20 Dec 2024)
SUGAR: Subject-Driven Video Customization in a Zero-Shot Manner (13 Dec 2024)
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation (25 Nov 2024)
VideoAlchemy: Open-set Personalization in Video Generation (15 Nov 2024)
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration (7 Nov 2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation (29 Oct 2024)
Anim-Director: A Large Multimodal Model Powered Agent for Controllable Animation Video Generation (19 Aug 2024)
Still-Moving: Customized Video Generation Without Customized Video Data (11 Jul 2024)
VIMI: Grounding Video Generation through Multi-modal Instruction (8 Jul 2024)
Customvideo: Customizing Text-to-Video Generation with Multiple Subjects (22 May 2024)
DisenStudio: Customized Multi-subject Text-to-Video Generation with Disentangled Spatial Control (21 May 2024)
Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers (9 May 2024)
SubjectDrive: Scaling Generative Data in Autonomous Driving via Subject Control (28 Mar 2024)
AesopAgent: Agent-driven Evolutionary System on Story-to-Video Production (12 Mar 2024)
VideoDrafter: Content-Consistent Multi-Scene Video Generation with LLM (2 Jan 2024)
Dreamvideo: Composing Your Dream Videos with Customized Subject and Motion (7 Dec 2023)
VideoBooth: Diffusion-based Video Generation with Image Prompts (1 Dec 2023)
Videodreamer: Customized Multi-Subject Text-to-Video Generation with Disen-Mix Finetuning (2 Nov 2023)
Animate-A-Story: Storytelling with Retrieval-Augmented Video Generation (13 Jul 2023)
TaleCrafter: Interactive Story Visualization with Multiple Characters (30 May 2023)
Dreamix: Video Diffusion Models are General Video Editors (2 Feb 2023)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models (8 Jun 2025)
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation (23 May 2025)
OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding (15 Apr 2025)
HunyuanVideo: A Systematic Framework For Large Video Generative Models (11 Mar 2025)
Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think (2 Mar 2025)
Adapting Image-to-Video Diffusion Models for Large-Motion Frame Interpolation (17 Feb 2025)
Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation (12 Feb 2025)
Autoregressive Video Generation without Vector Quantization (9 Jan 2025)
Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation (6 Jan 2025)
STIV: Scalable Text and Image Conditioned Video Generation (10 Dec 2024)
MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation (8 Dec 2024)
Lumiere: A Space-Time Diffusion Model for Video Generation (3 Dec 2024)
Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model (6 Nov 2024)
FrameBridge: Improving Image-to-Video Generation with Bridge Models (20 Oct 2024)
I4VGen: Image as Free Stepping Stone for Text-to-Video Generation (3 Oct 2024)
DynamiCrafter: Animating Open-Domain Images with Video Diffusion Priors (1 Oct 2024)
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation (27 Sep 2024)
Structure and Content-Guided Video Synthesis with Diffusion Models (27 Sep 2024)
DreamVideo: High-Fidelity Image-to-Video Generation with Image Retention and Text Guidance (16 Sep 2024)
Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning (2 Aug 2024)
MoVideo: Motion-Aware Video Generation with Diffusion Model (29 Jul 2024)
I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models (13 Jul 2024)
EasyAnimate: A High-Performance Long Video Generation Method based on Transformer Architecture (5 Jul 2024)
ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation (1 Jul 2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation (13 Jun 2024)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction (10 Jun 2024)
TI2V-Zero: Zero-Shot Image Conditioning for Text-to-Video Diffusion Models (25 Apr 2024)
TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models (25 Mar 2024)
Follow-Your-Click: Open-domain Regional Image Animation via Short Prompts (13 Mar 2024)
AtomoVideo: High Fidelity Image-to-Video Generation (5 Mar 2024)
Tuning-Free Noise Rectification for High Fidelity Image-to-Video Generation (5 Mar 2024)
Motion-i2v: Consistent and controllable image-to-video generation with explicit motion modeling (31 Jan 2024)
UniVG: Towards UNIfied-modal Video Generation (17 Jan 2024)
Decouple Content and Motion for Conditional Image-to-Video Generation (14 Dec 2023)
AnimateAnything: Fine-Grained Open Domain Image Animation with Motion Guidance (4 Dec 2023)
Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets (25 Nov 2023)
I2VGen-XL: High-Quality Image-to-Video Synthesis via Cascaded Diffusion Models (7 Nov 2023)
VideoCrafter1: Open Diffusion Models for High-Quality Video Generation (30 Oct 2023)
Synthesizing Videos from Images for Image-to-Video Adaptation (27 Oct 2023)
VideoDoodles: Hand-Drawn Animations on Videos with Scene-Aware Canvases (26 Jul 2023)
LaMD: Latent Motion Diffusion for Video Generation (23 Apr 2023)
Prompt Image to Life: Training-Free Text-Driven Image-to-Video Generation (2023)
Make It Move: Controllable Image-to-Video Generation With Text Descriptions (31 Mar 2022)
VideoGPT: Video Generation using VQ-VAE and Transformers (14 Sep 2021)
ImaGINator: Conditional Spatio-Temporal GAN for Video Generation (2020)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model (30 May 2024)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling (29 Jan 2024)
Motion-Conditioned Diffusion Model for Controllable Video Synthesis (27 Apr 2023)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (26 Nov 2024)
TrajectoryMover: Generative Movement of Object Trajectories in Videos (31 Mar 2026)
DiTraj: Training-free Trajectory Control For Video Diffusion Transformer (29 Sep 2025)
ATI: Any Trajectory Instruction for Controllable Video Generation (10 Jun 2025)
MOVi: Training-free Text-conditioned Multi-Object Video Generation (29 May 2025)
Frame In-N-Out: Unbounded Controllable Image-to-Video Generation (27 May 2025)
AnimateAnywhere: Rouse the Background in Human Image Animation (28 Apr 2025)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation (21 Apr 2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation (2 Apr 2025)
LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis (28 Mar 2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance (20 Mar 2025)
Tora: Trajectory-oriented Diffusion Transformer for Video Generation (14 Mar 2025)
LayerAnimate: Layer-level Control for Animation (22 Mar 2025)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (22 Mar 2025)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (10 Mar 2025)
C-Drag: Chain-of-Thought Driven Motion Controller for Video Generation (27 Feb 2025)
SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation (25 Feb 2025)
MotionCanvas: Cinematic Shot Design with Controllable Image-to-Video Generation (6 Feb 2025)
MotionBridge: Dynamic Video Inbetweening with Flexible Controls (7 Jan 2025)
TrackGo: A Flexible and Efficient Method for Controllable Video Generation (5 Jan 2025)
Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object (8 Jan 2025)
OmniDrag: Enabling Motion Control for Omnidirectional Image-to-Video Generation (12 Dec 2024)
ObjCtrl-2.5D: Training-free Object Control with Camera Poses (10 Dec 2024)
Motion Prompting: Controlling Video Generation with Motion Trajectories (3 Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
InTraGen: Trajectory-controlled Video Generation for Object Interactions (25 Nov 2024)
MotionBooth: Motion-Aware Customized Text-to-Video Generation (29 Oct 2024)
DragEntity: Trajectory Guided Video Generation using Entity and Positional Relationships (14 Oct 2024)
MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model (11 Jul 2024)
Freetraj: Tuning-free Trajectory Control in Video Diffusion Models (T2V) (24 Jun 2024)
Image Conductor: Precision Control for Interactive Video Synthesis (21 Jun 2024)
ReVideo: Remake a Video with Motion and Content (22 May 2024)
Collaborative Video Diffusion: Consistent Multi-video Generation with Camera Control (27 May 2024)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion (6 May 2024)
Peekaboo: Interactive Video Generation via Masked-Diffusion (T2V) (19 Apr 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (19 Apr 2024)
TrailBlazer: Trajectory Control for Diffusion-Based Video Generation (T2V) (8 Apr 2024)
Draganything: Motion Control for Anything Using Entity Representation (15 Mar 2024)
Boximator: Generating Rich and Controllable Motions for Video Synthesis (2 Feb 2024)
Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling (31 Jan 2024)
DragNuwa (Image+Text_Traj) (16 Aug 2023)
MCDiff Motion-Conditioned Diffusion Model for Controllable Video Synthesis (27 Apr 2023)
Controllable Video Generation With Sparse Trajectories (16 Dec 2018)
Follow-Your-Creation: Empowering 4D Creation through Video Inpainting (5 Jun 2025)
Voyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation (4 Jun 2025)
AKiRa: Augmentation Kit on Rays for optical video generation (Jun 2025)
Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation (21 Apr 2025)
GenDoP: Auto-regressive Camera Trajectory Generation as a Director of Photography (10 Apr 2025)
OmniCam: Unified Multimodal Video Generation via Camera Control (3 Apr 2025)
VidCRAFT3: Camera, Object, and Lighting Control for Image-to-Video Generation (2 Apr 2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories (27 Mar 2025)
Optical Flow Meets Video Diffusion Model for Enhanced Camera-Controlled Video Synthesis (25 Mar 2025)
AC3D: Analyzing and Improving 3D Camera Control in Video Diffusion Transformers (22 Mar 2025)
Aether: Geometric-Aware Unified World Modeling (18 Mar 2025)
EgoSim: Egocentric Exploration in Virtual Worlds with Multi-modal Conditioning (16 Mar 2025)
I2V3D: Controllable Image-to-Video Generation with 3D Guidance (12 Mar 2025)
CameraCtrl II: Dynamic Scene Exploration via Camera-controlled Video Diffusion Models (13 Mar 2025)
CameraCtrl: Enabling Camera Control for Text-to-Video Generation (13 Mar 2025)
ReCamMaster: Camera-Controlled Generative Rendering from A Single Video (14 Mar 2025)
GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control (5 Mar 2025)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (10 Mar 2025)
I2VCONTROL-CAMERA: Precise Video Camera Control with Adjustable Motion Strength (28 Feb 2025)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation (12 Feb 2025)
Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion (12 Feb 2025)
RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control (14 Feb 2025)
3DTrajMaster: Mastering 3D Trajectory for Multi-Entity Motion in Video Generation (7 Feb 2025)
You See it, You Got it: Learning 3D Creation on Pose-Free Videos at Scale (10 Dec 2024)
Latent-Reframe: Enabling Camera Control for Video Diffusion Model without Training (8 Dec 2024)
CamI2V: Camera-Controlled Image-to-Video Diffusion Model (4 Dec 2024)
I2VControl: Disentangled and Unified Video Motion Synthesis Control (30 Nov 2024)
Trajectory Attention: Enhancing Video Generation with Fine-Grained Motion Control (28 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
DimensionX: Create Any 3D and 4D Scenes from a Single Image with Controllable Video Diffusion (7 Nov 2024)
Boosting Camera Motion Control for Video Diffusion Transformers (14 Oct 2024)
Cavia: Camera-controllable Multi-view Video Diffusion with View-Integrated Attention (14 Oct 2024)
ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis (3 Sep 2024)
VD3D: Taming Large Video Diffusion Transformers for 3D Camera Control (17 Jul 2024)
MotionCtrl: A Unified and Flexible Motion Controller for Video Generation (16 Jul 2024)
CamCo: Camera-Controllable 3D-Consistent Image-to-Video Generation (4 Jun 2024)
MotionMaster: Training-free Camera Motion Transfer For Video Generation (1 May 2024)
MOTIONFLOW: Learning Implicit Motion Flow for Complex Camera Trajectory Control in Video Generation (Dec 2024)
FlexiAct: Towards Flexible Action Control in Heterogeneous Scenarios (6 May 2025)
Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning (5 Jun 2025)
MotionPro: A Precise Motion Controller for Image-to-Video Generation (29 May 2025)
DualReal: Joint Training for Lossless Identity-Motion Fusion in Video Customization (4 May 2025)
UniAnimate-DiT: Human Image Animation with Large-Scale Video Diffusion Transformer (15 Apr 2025)
InterDyn: Controllable Interactive Dynamics with Video Diffusion Models(hand mask sequence as control signal) (4 Apr 2025)
DreamActor-M1: Holistic, Expressive and Robust Human Image Animation with Hybrid Guidance (3 Apr 2025)
Video Motion Transfer with Diffusion Transformers (27 Mar 2025)
Motion Prompting: Controlling Video Generation with Motion Trajectories (27 Mar 2025)
EfficientMT: Efficient Temporal Adaptation for Motion Transfer in Text-to-Video Diffusion Models (25 Mar 2025)
DreamRelation: Relation-Centric Video Customization (10 Mar 2025)
MotionMatcher: Motion Customization of Text-to-Video Diffusion Models via Motion Feature Matching (18 Feb 2025)
MotionAgent: Fine-grained Controllable Video Generation via Motion Field Agent (5 Feb 2025)
Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss (13 Jan 2025)
SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing (13 Jan 2025)
Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control (9 Jan 2025)
Spectral Motion Alignment for Video Motion Transfer using Diffusion Models (19 Dec 2024)
MoTrans: Customized Motion Transfer with Text-driven Video Diffusion Models (2 Dec 2024)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion ModelsOnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models (15 Nov 2024)
Video Diffusion Models are Training-free Motion Interpreter and Controller (12 Nov 2024)
AnimateAnything: Consistent and Controllable Animation for Video Generation (16 Nov 2024)
Motionbooth: Motion-aware customized text-to-video generation (29 Oct 2024)
MotionClone: Training-Free Motion Cloning for Controllable Video Generation (22 Oct 2024)
Motion Inversion for Video Customization (16 Oct 2024)
Zero-Shot Controllable Image-to-Video Animation via Motion Decomposition (21 Jul 2024)
DreamMotion: Space-Time Self-Similar Score Distillation for Zero-Shot Video Editing (15 Jul 2024)
Champ: Controllable and Consistent Human Image Animation with 3D Parametric Guidance (1 Jun 2024)
Customize-A-Video: One-Shot Motion Customization of Text-to-Video Diffusion Models (28 Aug 2024)
Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation (7 Dec 2023)
DreamVideo: Composing Your Dream Videos with Customized Subject and Motion (7 Dec 2023)
VMC: Video Motion Customization with Pre-trained Diffusion Models (1 Dec 2023)
Space-Time Diffusion Features for Zero-Shot Text-Driven Motion Transfer (3 Dec 2023)
MotionDirector: Motion Customization for Text-to-Video Diffusion Models (12 Oct 2023)
Truncated β view the full README on GitHub.