QuenithAI/Video-Generation-Paper-List

Tracking the latest and greatest research papers on video generation.

171

59 commits

updated Mar 28, 2026

See the code

README

Awesome Video Generation by QuenithAI

A curated collection of papers, models, and resources for the field of Video Generation.

Awesome   PRs Welcome   Issues Welcome

[!NOTE] This repository is proudly maintained by the frontline research mentors at QuenithAI (应达学术). It aims to provide the most comprehensive and cutting-edge map of papers and technologies in the field of video generation.

Your contributions are also vital—feel free to open an issue or submit a pull request to become a collaborator of this repository. We expect your participation!

If you require expert 1-on-1 guidance on your submissions to top-tier conferences and journals, we invite you to contact us via WeChat or E-mail.


本仓库由 「应达学术」(QuenithAI) 的一线科研导师团队倾力打造并持续维护,旨在为您呈现视频生成领域最全面、最前沿的视频生成领域的论文。

您的贡献对我们和社区来说至关重要——我们诚邀有志之士通过 open an issuesubmit a pull request 来成为这个项目的合作者之一,期待您的加入!

如果您在冲刺科研顶会的道路上需要专业的1V1指导,欢迎通过微信邮件联系我们

⚡ Latest Updates
  • (Mar 14th, 2026): We have updated all accepted papers in ICLR 2026.
  • (Mar 9th, 2026): We have updated all accepted papers in CVPR 2026.
  • (Nov 19th, 2025): We have updated all accepted papers in AAAI 2026.
  • (Sep 13th, 2025): Add a new direction: 🎯 Reinforcement Learning for Video Generation.
  • (Aug 21th, 2025): Add a new direction: 🗣️ Audio-Driven Video Generation.
  • (Aug 20th, 2025): Initial commit and repository structure established.

📚 Table of Contents


📜 Papers & Models

✍️ Survey Papers

🎥 Text-to-Video (T2V) Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] Minute-Long Videos with Dual Parallelisms
    Paper Project Page GitHub

  • [AAAI 2026] FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
    Paper Project Page GitHub Hugging Face

  • [AAAI 2026] DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
    Paper Project Page GitHub

  • [AAAI 2026] GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
    Paper Project Page GitHub

  • [AAAI 2026] EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
    Paper Project Page

  • [CVPR 2026] ID-Composer: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
    ArXiv GitHub Project Page

  • [CVPR 2026] StreamDiT: Real-Time Streaming Text-to-Video Generation
    Paper Project Page

  • [CVPR 2026] Wan-Alpha: Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
    Paper Project Page

  • [CVPR 2026] HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
    Paper GitHub

  • [CVPR 2026] VISTA: A Test-Time Self-Improving Video Generation Agent
    Paper Project Page

  • [CVPR 2026] Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
    Paper Project Page GitHub

  • [CVPR 2026] LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
    Paper

  • [CVPR 2026] CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
    Paper Project Page

  • [CVPR 2026] Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
    Paper

  • [CVPR 2026 Findings] BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
    Paper

  • [ICLR 2026] Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
    Paper

  • [ICLR 2026] VideoRepair: Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
    Paper Project Page GitHub

  • [ICLR 2026] M4V: Multimodal Mamba for Efficient Text-to-Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] Video-As-Prompt: Unified Semantic Control for Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Generating Human Motion Videos using a Cascaded Text-to-Video Framework
    Paper

  • [ICLR 2026] TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
    Paper Project Page

  • [ICLR 2026] NoisEasier: Boosting Text-to-Video Generation with Direct Noise Optimization
    Paper Project Page

  • [ICLR 2026] Video-MSG: Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
    Paper Project Page GitHub

  • [ICLR 2026] FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
    Paper

  • [ICLR 2026] DiffuPhyGS: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors
    Paper

  • [ICLR 2026] RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
    Paper

  • [ICLR 2026] JDM: Joint Distribution Modeling for Fine-Grained Text-to-Video Generation
    Paper

  • [ICLR 2026] Towards One-step Causal Video Generation via Adversarial Self-Distillation
    Paper GitHub

  • [ICLR 2026] TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
    Paper

  • [ICLR 2026] Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
    Paper Project Page GitHub

  • [ICLR 2026] CCC: Prompt Evolution for Video Generation via Structured MLLM Feedback
    Paper

  • [ICLR 2026] Ask-A-Video: Controlling Video Generation with Vision Language Models
    Paper

  • [ICLR 2026] BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Subject-driven Video Generation Emerges from Experience Replays
    Paper

  • [ICLR 2026] Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
    Paper GitHub Hugging Face

  • [ICLR 2026] Cosmos-Eval: Towards Explainable Evaluation of Physics and Semantics in Text-to-Video Models
    Paper

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
    ArXiv GitHub

  • [CVPR 2025] Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
    ArXiv GitHub

  • [CVPR 2025] Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2025] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
    ArXiv Project Page GitHub

  • [CVPR 2025] Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2025] TransPixeler: Advancing Text-to-Video Generation with Transparency
    ArXiv Project Page GitHub

  • [CVPR 2025] LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
    ArXiv GitHub

  • [CVPR 2025] Improving Text-to-Video Generation via Instance-aware Structured Caption
    ArXiv GitHub

  • [CVPR 2025] Compositional Text-to-Video Generation with Blob Video Representations
    ArXiv Project Page

  • [CVPR 2025] Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
    ArXiv Project Page

  • [ICCV 2025] T2Bs: Text‑to‑Character Blendshapes via Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] Animate Your Word: Bringing Text to Life via Video Diffusion Prior
    Paper Project Page GitHub

  • [NeurIPS 2025] Safe‑Sora: Safe Text‑to‑Video Generation via Graphical Watermarking
    Paper Project Page GitHub Hugging Face

  • [ICCV 2025] Prompt‑A‑Video: Prompt Your Video Diffusion Model via Preference‑Aligned LLM
    Paper GitHub Hugging Face

  • [ICCV 2025] MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text‑to‑Video Generation
    Paper Project Page

  • [ICCV 2025] TITAN‑Guide: Taming Inference‑Time Alignment for Guided Text‑to‑Video Diffusion Models
    Paper Project Page

  • [ICCV 2025] Video‑T1: Test‑Time Scaling for Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] AnimateYourMesh: Feed‑Forward 4D Foundation Model for Text‑Driven Mesh Animation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2025] CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
    Paper

  • [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
    Paper Project Page GitHub Hugging Face

  • [NeurIPS 2025] Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
    Paper Project Page

  • [NeurIPS 2025] ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
    Paper GitHub

  • [NeurIPS 2025] PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-Aware Mechanisms
    Paper Project Page GitHub

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] Vlogger: Make Your Dream A Vlog
    ArXiv GitHub

  • [CVPR 2024] Make Pixels Dance: High-Dynamic Video Generation
    ArXiv Project Page Demo

  • [CVPR 2024] VGen: Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2024] GenTron: Delving Deep into Diffusion Transformers for Image and Video Generation
    ArXiv Project Page

  • [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2024] MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
    ArXiv Project Page Demo

  • [CVPR 2024] Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models
    ArXiv Project Page

  • [CVPR 2024] PEEKABOO: Interactive Video Generation via Masked-Diffusion
    ArXiv Project Page GitHub Demo

  • [CVPR 2024] EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
    ArXiv Project Page GitHub

  • [CVPR 2024] A Recipe for Scaling up Text-to-Video Generation with Text-free Videos
    ArXiv Project Page GitHub

  • [CVPR 2024] BIVDiff: A Training-free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
    ArXiv Project Page

  • [CVPR 2024] Mind the Time: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
    ArXiv Project Page

  • [CVPR 2024] MotionDirector: Motion Customization of Text-to-Video Diffusion Models
    ArXiv GitHub

  • [CVPR 2024] Hierarchical Patch-wise Diffusion Models for High-Resolution Video Generation
    Paper Project Page

  • [CVPR 2024] DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation
    Paper GitHub

  • [CVPR 2024] Grid Diffusion Models for Text-to-Video Generation
    ArXiv GitHub Demo

  • [ECCV 2024] Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
    ArXiv Project Page

  • [ECCV 2024] W.A.L.T.: Photorealistic Video Generation with Diffusion Models
    Paper Project Page

  • [ECCV 2024] MoVideo: Motion-Aware Video Generation with Diffusion Models
    Paper

  • [ECCV 2024] DrivingDiffusion: Layout-Guided Multi-View Driving Scenarios Video Generation with Latent Diffusion Model
    Paper Project Page GitHub

  • [ECCV 2024] MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
    Paper

  • [ECCV 2024] HARIVO: Harnessing Text-to-Image Models for Video Generation
    Paper Project Page

  • [ECCV 2024] MEVG: Multi-event Video Generation with Text-to-Video Models
    Paper Project Page

  • [NeurIPS 2024] DEMO: Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
    Paper GitHub

  • [ICML 2024] Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
    Paper Project Page GitHub Hugging Face

  • [ICLR 2024] VDT: General-purpose Video Diffusion Transformers via Mask Modeling
    ArXiv Project Page GitHub

  • [ICLR 2024] VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video Generation
    Paper

  • [AAAI 2024] Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
    ArXiv Project Page GitHub

  • [AAAI 2024] E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
    ArXiv

  • [AAAI 2024] ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
    ArXiv Project Page GitHub

  • [AAAI 2024] F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text to-Video Synthesis
    ArXiv

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [CVPR 2023] Align your Latents: High-resolution Video Synthesis with Latent Diffusion Models
    ArXiv Project Page GitHub

  • [CVPR 2023] Text2Video-Zero: Text-to-image Diffusion Models are Zero-shot Video Generators
    Paper Project Page GitHub Demo

  • [CVPR 2023] Video Probabilistic Diffusion Models in Projected Latent Space
    Paper GitHub

  • [ICCV 2023] PYOCO: Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
    Paper Project Page

  • [ICCV 2023] Gen-1: Structure and Content-guided Video Synthesis with Diffusion Models
    Paper Project Page

  • [NeurIPS 2023] Video Diffusion Models
    ArXiv Project Page

  • [NeurIPS 2023] UniPi: Learning Universal Policies via Text-Guided Video Generation
    Paper Project Page GitHub

  • [NeurIPS 2023] VideoComposer: Compositional Video Synthesis with Motion Controllability
    ArXiv Project Page GitHub

  • [ICLR 2023] CogVideo: Large-scale Pretraining for Text-to-video Generation via Transformers
    Paper GitHub Demo

  • [ICLR 2023] Make-A-Video: Text-to-video Generation without Text-video Data
    ArXiv Project Page GitHub

  • [ICLR 2023] Phenaki: Variable Length Video Generation From Open Domain Textual Description
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

🖼️ Image-to-Video (I2V) Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] IPRO: Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
    Paper Project Page GitHub

  • [CVPR 2026] ALG: Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
    Paper Project Page GitHub

  • [CVPR 2026] ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
    Paper Project Page

  • [CVPR 2026] FlexiMMT: Let Your Image Move with Your Motion! - Implicit Multi-Object Multi-Motion Transfer
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
    Paper Project Page GitHub

  • [ICLR 2026] MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Anchor Frame Bridging for Coherent First-Last Frame Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
    Paper Project Page GitHub Hugging Face

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
    ArXiv

  • [CVPR 2025] MotionPro: A Precise Motion Controller for Image-to-Video Generation
    ArXiv GitHub

  • [CVPR 2025] Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
    ArXiv

  • [CVPR 2025] Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
    ArXiv

  • [CVPR 2025] I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models
    Paper

  • [CVPR 2025] LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
    Paper GitHub

  • [ICCV 2025] AnyI2V: Animating Any Conditional Image with Motion Control
    Paper GitHub

  • [ICCV 2025] Versatile Transition Generation with Image-to-Video Diffusion
    Paper

  • [ICCV 2025] TIP‑I2V: A Million‑Scale Real Text and Image Prompt Dataset for Image‑to‑Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICCV 2025] Unified Video Generation via Next‑Set Prediction in Continuous Domain
    Paper

  • [NeurIPS 2025] GenRec: Unifying Video Generation and Recognition with Diffusion Models
    Paper GitHub Hugging Face

  • [ICCV 2025] Precise Action‑to‑Video Generation Through Visual Action Prompts
    Paper Project Page

  • [ICCV 2025] STIV: Scalable Text and Image Conditioned Video Generation
    Paper Project Page Hugging Face

  • [ICLR 2025] FrameBridge: Improving Image‑to‑Video Generation with Bridge Models
    Paper GitHub Project Page

  • [ICLR 2025] SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
    Paper Project Page GitHub

  • [ICLR 2025] Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
    Paper

  • [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
    Paper Project Page GitHub Hugging Face

  • [NeurIPS 2025] MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
    Paper GitHub Hugging Face

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] Animate Anyone: Consistent and Controllable Image-to-video Synthesis for Character Animation
    ArXiv Project Page GitHub

  • [CVPR 2024] Your Image Is My Video: Reshaping the Receptive Field via Image-to-Video Differentiable AutoAugmentation and Fusion
    Paper

  • [CVPR 2024] TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
    Paper

  • [CVPR 2024] Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning
    Paper

  • [CVPR 2024] Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
    Paper GitHub

  • [ECCV 2024] MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
    Paper GitHub

  • [ECCV 2024] $\mathrm R2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
    Paper GitHub

  • [ECCV 2024] PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
    Paper GitHub

  • [ECCV 2024] Rethinking Image-to-Video Adaptation: An Object-Centric Perspective
    Paper

  • [NeurIPS 2024] TPC: Test-time Procrustes Calibration for Diffusion-based Human Image Animation
    Paper

  • [NeurIPS 2024] Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
    Paper GitHub

  • [ICML 2024] Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
    Paper Project Page GitHub Hugging Face

  • [SIGGRAPH 2024] I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
    Paper GitHub

  • [SIGGRAPH 2024] Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
    Paper GitHub

  • [AAAI 2024] Continuous Piecewise-Affine Based Motion Model for Image Animation
    Paper GitHub

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [ICCV 2023] DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
    Paper GitHub

  • [ICCV 2023] Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

✂️ Video-to-Video (V2V) Editing

✨ 2026

✅ Published Papers

  • [AAAI 2026] Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
    Paper Project Page GitHub

  • [AAAI 2026] FAME: Fairness-aware Attention-modulated Video Editing
    Paper

  • [CVPR 2026] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] VideoCoF: Unified Video Editing with Temporal Reasoner
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
    Paper Project Page GitHub

  • [CVPR 2026] EasyV2V: A High-quality Instruction-based Video Editing Framework
    Paper Project Page

  • [CVPR 2026] VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Generative Video Motion Editing with 3D Point Tracks
    Paper Project Page

  • [CVPR 2026] NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
    Paper GitHub Hugging Face

  • [CVPR 2026] EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
    Paper Project Page GitHub

  • [CVPR 2026] PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
    Paper Project Page

  • [CVPR 2026] FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
    Paper Project Page

  • [ICLR 2026] Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
    Paper Project Page GitHub

  • [ICLR 2026] UNIC: Unified In-Context Video Editing
    Paper Project Page

  • [ICLR 2026] LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
    Paper Project Page GitHub

  • [ICLR 2026] DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
    Paper GitHub Hugging Face

  • [ICLR 2026] Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
    Paper Project Page GitHub

  • [ICLR 2026] FlowGuide: Precision-Guided Enhancement for Face Image and Video Editing
    Paper GitHub

  • [ICLR 2026] DragStream: Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
    Paper Project Page GitHub

  • [ICLR 2026] Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
    Paper Project Page

  • [ICLR 2026] FastVMT: Eliminating Redundancy in Video Motion Transfer
    Paper Project Page GitHub

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
    Paper

  • [CVPR 2025] VideoDirector: Precise Video Editing via Text-to-Video Models
    Paper GitHub

  • [CVPR 2025] VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing
    Paper

  • [CVPR 2025] Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing
    Paper

  • [CVPR 2025] Unity in Diversity: Video Editing via Gradient-Latent Purification
    Paper

  • [CVPR 2025] VEU-Bench: Towards Comprehensive Understanding of Video Editing
    Paper GitHub

  • [CVPR 2025] SketchVideo: Sketch-based Video Generation and Editing
    Paper GitHub

  • [CVPR 2025] FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video
    Paper GitHub

  • [CVPR 2025] Visual Prompting for One-shot Controllable Video Editing without Inversion
    Paper

  • [CVPR 2025] FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
    Paper GitHub

  • [ICCV 2025] VACE: All-in-One Video Creation and Editing
    Paper Project Page GitHub

  • [ICCV 2025] Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
    Paper Project Page

  • [ICCV 2025] DIVE: Taming DINO for Subject-Driven Video Editing
    Paper Project Page

  • [ICCV 2025] DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors
    Paper Project Page

  • [ICCV 2025] QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video Editing
    Paper

  • [ICCV 2025] Teleportraits: Training-Free People Insertion into Any Scene
    Paper

  • [ICLR 2025] VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
    Paper GitHub

  • [NeurIPS 2025] REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
    Paper Project Page

  • [AAAI 2025] FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
    Paper GitHub

  • [AAAI 2025] EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models
    Paper GitHub

  • [AAAI 2025] VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
    Paper GitHub

  • [AAAI 2025] Re-Attentional Controllable Video Diffusion Editing
    Paper GitHub

  • [WACV 2025] IP-FaceDiff: Identity-Preserving Facial Video Editing with Diffusion
    Paper

  • [WACV 2025] SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
    Paper GitHub

  • [WACV 2025] MagicStick: Controllable Video Editing via Control Handle Transformations
    Paper

  • [WACV 2025] Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
    Paper

  • [WACV 2025] FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
    Paper GitHub

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
    Paper GitHub

  • [CVPR 2024] VidToMe: Video Token Merging for Zero-Shot Video Editing
    Paper GitHub

  • [CVPR 2024] Video-P2P: Video Editing with Cross-Attention Control
    Paper GitHub

  • [CVPR 2024] CCEdit: Creative and Controllable Video Editing via Diffusion Models
    Paper

  • [CVPR 2024] RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models
    Paper GitHub

  • [CVPR 2024] DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
    Paper

  • [CVPR 2024] MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
    Paper

  • [CVPR 2024] MotionEditor: Editing Video Motion via Content-Aware Diffusion
    Paper GitHub

  • [CVPR 2024] CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-Driven Video Editing
    Paper GitHub

  • [ICLR 2024] Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
    Paper GitHub

  • [ICLR 2024] Video Decomposition Prior: Editing Videos Layer by Layer
    Paper

  • [ICLR 2024] FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
    Paper

  • [ICLR 2024] TokenFlow: Consistent Diffusion Features for Consistent Video Editing
    Paper GitHub

  • [ECCV 2024] VIDEOSHOP: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
    Paper GitHub

  • [ECCV 2024] DragVideo: Interactive Drag-Style Video Editing
    Paper GitHub

  • [ECCV 2024] WAVE: Warping DDIM Inversion Features for Zero-Shot Text-to-Video Editing
    Paper

  • [ECCV 2024] DreamMotion: Space-Time Self-similar Score Distillation for Zero-Shot Video Editing
    Paper

  • [ECCV 2024] Object-Centric Diffusion for Efficient Video Editing
    Paper GitHub

  • [ECCV 2024] Video Editing via Factorized Diffusion Distillation
    Paper

  • [ECCV 2024] SAVE: Protagonist Diversification with Structure Agnostic Video Editing
    Paper

  • [ECCV 2024] DNI: Dilutional Noise Initialization for Diffusion Video Editing
    Paper

  • [ECCV 2024] MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and Editing
    Paper

  • [ECCV 2024] DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency
    Paper

💡 Pre-Print Papers

⇧ Back to ToC

🕹️ Controllable Video Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
    Paper Project Page GitHub

  • [AAAI 2026] MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
    Paper Project Page

  • [CVPR 2026] Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] First Frame Is the Place to Go for Video Content Customization
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] UCPE: Unified Camera Positional Encoding for Controlled Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
    Paper

  • [CVPR 2026] ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
    Paper Project Page GitHub

  • [CVPR 2026] VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
    Paper Project Page GitHub

  • [CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
    Paper Project Page GitHub

  • [CVPR 2026] ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
    Paper Project Page GitHub

  • [CVPR 2026] BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] RealWonder: Real-Time Physical Action-Conditioned Video Generation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] LAMP: Language-Assisted Motion Planning for Controllable Video Generation
    Paper GitHub Hugging Face

  • [CVPR 2026] WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
    Paper GitHub

  • [CVPR 2026] StableWorld: Towards Stable and Consistent Long Interactive Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] Controllable Video Generation with Provable Disentanglement
    Paper

  • [ICLR 2026] 3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation
    Paper

  • [ICLR 2026] MotionStream: Real-Time Video Generation with Interactive Motion Controls
    Paper Project Page GitHub

  • [ICLR 2026] Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock Denoising
    Paper Project Page GitHub

  • [ICLR 2026] Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
    Paper Project Page GitHub

  • [ICLR 2026] Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control
    Paper

  • [ICLR 2026] Target-Aware Video Diffusion Models
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
    Paper Project Page GitHub

  • [ICLR 2026] MATRIX: Mask Track Alignment for Interaction-aware Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
    Paper Project Page

  • [ICLR 2026] Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models
    Paper Project Page

  • [ICLR 2026] ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
    Paper Project Page GitHub

  • [ICLR 2026] Light-X: Generative 4D Video Rendering with Camera and Illumination Control
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Generative View Stitching
    Paper Project Page GitHub

  • [ICLR 2026] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
    Paper Project Page GitHub

  • [ICLR 2026] LightCtrl: Training-free Controllable Video Relighting
    Paper GitHub

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner
    Paper

  • [CVPR 2025] AnimateAnything: Consistent and Controllable Animation for Video Generation
    Paper

  • [CVPR 2025] Customized Condition Controllable Generation for Video Soundtrack
    Paper GitHub

  • [CVPR 2025] StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
    Paper

  • [ICCV 2025] Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
    Paper Project Page GitHub

  • [ICCV 2025] MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
    Paper GitHub

  • [ICCV 2025] MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
    Paper Project Page GitHub

  • [ICCV 2025] InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
    Paper Project Page GitHub

  • [ICCV 2025] Free-Form Motion Control (SynFMC): Controlling the 6D Poses of Camera and Objects in Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
    Paper Project Page GitHub

  • [ICCV 2025] MagicMotion: Video Generation with a Smart Director
    Paper Project Page GitHub

  • [ICCV 2025] UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving
    Paper Project Page GitHub

  • [ICLR 2025] MotionClone: Training-Free Motion Cloning for Controllable Video Generation
    Paper GitHub

  • [AAAI 2025] CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
    Paper GitHub

  • [AAAI 2025] TrackGo: A Flexible and Efficient Method for Controllable Video Generation
    Paper

  • [WACV 2025] Fine-grained Controllable Video Generation via Object Appearance and Context
    Paper

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
    Paper GitHub

  • [CVPR 2024] Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
    Paper GitHub

  • [AAAI 2024] Decouple Content and Motion for Conditional Image-to-Video Generation
    Paper

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [CVPR 2023] Conditional Image-to-Video Generation with Latent Flow Diffusion Models
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

🗣️ Audio-Driven Video Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] EchoMimicV3: EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
    Paper Project Page GitHub Hugging Face

  • [AAAI 2026] FantasyTalking2: FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026 Findings] UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
    Paper

  • [CVPR 2026] Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
    Paper Project Page GitHub

  • [CVPR 2026] ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
    Paper Project Page

  • [CVPR 2026] StreamAvatar: *Streaming Diffusion Models for

Truncated — view the full README on GitHub.

aaai
cvpr
deep-learning
diffusion-models
eccv
iccv
iclr
icml
image-to-video
image-to-video-generation
neurips
nips
siggraph
text-to-video
text-to-video-generation
video-editing
video-generation
video-to-video
video-to-video-translation

Contributors

QuenithAI

48 commits

SHYuanBest

4 commits

Haodong-Yan

2 commits

feifeiobama

2 commits

QuenithAI/Video-Generation-Paper-List

Tracking the latest and greatest research papers on video generation.

171

59 commits

updated Mar 28, 2026

See the code

README

Awesome Video Generation by QuenithAI

A curated collection of papers, models, and resources for the field of Video Generation.

Awesome   PRs Welcome   Issues Welcome

[!NOTE] This repository is proudly maintained by the frontline research mentors at QuenithAI (应达学术). It aims to provide the most comprehensive and cutting-edge map of papers and technologies in the field of video generation.

Your contributions are also vital—feel free to open an issue or submit a pull request to become a collaborator of this repository. We expect your participation!

If you require expert 1-on-1 guidance on your submissions to top-tier conferences and journals, we invite you to contact us via WeChat or E-mail.


本仓库由 「应达学术」(QuenithAI) 的一线科研导师团队倾力打造并持续维护,旨在为您呈现视频生成领域最全面、最前沿的视频生成领域的论文。

您的贡献对我们和社区来说至关重要——我们诚邀有志之士通过 open an issuesubmit a pull request 来成为这个项目的合作者之一,期待您的加入!

如果您在冲刺科研顶会的道路上需要专业的1V1指导,欢迎通过微信邮件联系我们

⚡ Latest Updates
  • (Mar 14th, 2026): We have updated all accepted papers in ICLR 2026.
  • (Mar 9th, 2026): We have updated all accepted papers in CVPR 2026.
  • (Nov 19th, 2025): We have updated all accepted papers in AAAI 2026.
  • (Sep 13th, 2025): Add a new direction: 🎯 Reinforcement Learning for Video Generation.
  • (Aug 21th, 2025): Add a new direction: 🗣️ Audio-Driven Video Generation.
  • (Aug 20th, 2025): Initial commit and repository structure established.

📚 Table of Contents


📜 Papers & Models

✍️ Survey Papers

🎥 Text-to-Video (T2V) Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] Minute-Long Videos with Dual Parallelisms
    Paper Project Page GitHub

  • [AAAI 2026] FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
    Paper Project Page GitHub Hugging Face

  • [AAAI 2026] DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
    Paper Project Page GitHub

  • [AAAI 2026] GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
    Paper Project Page GitHub

  • [AAAI 2026] EmoVid: A Multimodal Emotion Video Dataset for Emotion-Centric Video Understanding and Generation
    Paper Project Page

  • [CVPR 2026] ID-Composer: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
    ArXiv GitHub Project Page

  • [CVPR 2026] StreamDiT: Real-Time Streaming Text-to-Video Generation
    Paper Project Page

  • [CVPR 2026] Wan-Alpha: Video Generation with Stable Transparency via Shiftable RGB-A Distribution Learner
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] MultiShotMaster: A Controllable Multi-Shot Video Generation Framework
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] OneStory: Coherent Multi-Shot Video Generation with Adaptive Memory
    Paper Project Page

  • [CVPR 2026] HoloCine: Holistic Generation of Cinematic Multi-Shot Long Video Narratives
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] SwitchCraft: Training-Free Multi-Event Video Generation with Attention Controls
    Paper GitHub

  • [CVPR 2026] VISTA: A Test-Time Self-Improving Video Generation Agent
    Paper Project Page

  • [CVPR 2026] Infinity-RoPE: Action-Controllable Infinite Video Generation Emerges From Autoregressive Self-Rollout
    Paper Project Page GitHub

  • [CVPR 2026] LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video Generation
    Paper

  • [CVPR 2026] CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
    Paper Project Page

  • [CVPR 2026] Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
    Paper

  • [CVPR 2026 Findings] BrandFusion: A Multi-Agent Framework for Seamless Brand Integration in Text-to-Video Generation
    Paper

  • [ICLR 2026] Improving Dynamic Object Interactions in Text-to-Video Generation with AI Feedback
    Paper

  • [ICLR 2026] VideoRepair: Self-Correcting Text-to-Video Generation with Misalignment Detection and Localized Refinement
    Paper Project Page GitHub

  • [ICLR 2026] M4V: Multimodal Mamba for Efficient Text-to-Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] Video-As-Prompt: Unified Semantic Control for Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Generating Human Motion Videos using a Cascaded Text-to-Video Framework
    Paper

  • [ICLR 2026] TGT: Text-Grounded Trajectories for Locally Controlled Video Generation
    Paper Project Page

  • [ICLR 2026] NoisEasier: Boosting Text-to-Video Generation with Direct Noise Optimization
    Paper Project Page

  • [ICLR 2026] Video-MSG: Training-free Guidance in Text-to-Video Generation via Multimodal Planning and Structured Noise Initialization
    Paper Project Page GitHub

  • [ICLR 2026] FaithfulFaces: Pose-Faithful Facial Identity Preservation for Text-to-Video Generation
    Paper

  • [ICLR 2026] DiffuPhyGS: Text-to-Video Generation with 3D Gaussians and Learnable Physical Properties via Diffusion Priors
    Paper

  • [ICLR 2026] RichSpace: Enriching Text-to-Video Prompt Space via Text Embedding Interpolation
    Paper

  • [ICLR 2026] JDM: Joint Distribution Modeling for Fine-Grained Text-to-Video Generation
    Paper

  • [ICLR 2026] Towards One-step Causal Video Generation via Adversarial Self-Distillation
    Paper GitHub

  • [ICLR 2026] TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation
    Paper

  • [ICLR 2026] Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
    Paper Project Page GitHub

  • [ICLR 2026] CCC: Prompt Evolution for Video Generation via Structured MLLM Feedback
    Paper

  • [ICLR 2026] Ask-A-Video: Controlling Video Generation with Vision Language Models
    Paper

  • [ICLR 2026] BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Subject-driven Video Generation Emerges from Experience Replays
    Paper

  • [ICLR 2026] Kaleido: Open-Sourced Multi-Subject Reference Video Generation Model
    Paper GitHub Hugging Face

  • [ICLR 2026] Cosmos-Eval: Towards Explainable Evaluation of Physics and Semantics in Text-to-Video Models
    Paper

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM
    ArXiv GitHub

  • [CVPR 2025] Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
    ArXiv GitHub

  • [CVPR 2025] Retrieval-Augmented Prompt Optimization for Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2025] Identity-Preserving Text-to-Video Generation by Frequency Decomposition
    ArXiv Project Page GitHub

  • [CVPR 2025] Exploiting Intersections in Diffusion Trajectories for Model-Agnostic, Zero-Shot, Training-Free Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2025] TransPixeler: Advancing Text-to-Video Generation with Transparency
    ArXiv Project Page GitHub

  • [CVPR 2025] LLM-Guided Iterative Self-Refinement for Physics-Grounded Text-to-Video Generation
    ArXiv GitHub

  • [CVPR 2025] Improving Text-to-Video Generation via Instance-aware Structured Caption
    ArXiv GitHub

  • [CVPR 2025] Compositional Text-to-Video Generation with Blob Video Representations
    ArXiv Project Page

  • [CVPR 2025] Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
    ArXiv Project Page

  • [ICCV 2025] T2Bs: Text‑to‑Character Blendshapes via Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] Animate Your Word: Bringing Text to Life via Video Diffusion Prior
    Paper Project Page GitHub

  • [NeurIPS 2025] Safe‑Sora: Safe Text‑to‑Video Generation via Graphical Watermarking
    Paper Project Page GitHub Hugging Face

  • [ICCV 2025] Prompt‑A‑Video: Prompt Your Video Diffusion Model via Preference‑Aligned LLM
    Paper GitHub Hugging Face

  • [ICCV 2025] MotionShot: Adaptive Motion Transfer across Arbitrary Objects for Text‑to‑Video Generation
    Paper Project Page

  • [ICCV 2025] TITAN‑Guide: Taming Inference‑Time Alignment for Guided Text‑to‑Video Diffusion Models
    Paper Project Page

  • [ICCV 2025] Video‑T1: Test‑Time Scaling for Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] AnimateYourMesh: Feed‑Forward 4D Foundation Model for Text‑Driven Mesh Animation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2025] OpenVid-1M: A Large-Scale High-Quality Dataset for Text-to-Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2025] CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
    Paper

  • [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
    Paper Project Page GitHub Hugging Face

  • [NeurIPS 2025] Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
    Paper Project Page

  • [NeurIPS 2025] ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
    Paper GitHub

  • [NeurIPS 2025] PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-Aware Mechanisms
    Paper Project Page GitHub

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] Vlogger: Make Your Dream A Vlog
    ArXiv GitHub

  • [CVPR 2024] Make Pixels Dance: High-Dynamic Video Generation
    ArXiv Project Page Demo

  • [CVPR 2024] VGen: Hierarchical Spatio-temporal Decoupling for Text-to-Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2024] GenTron: Delving Deep into Diffusion Transformers for Image and Video Generation
    ArXiv Project Page

  • [CVPR 2024] SimDA: Simple Diffusion Adapter for Efficient Video Generation
    ArXiv Project Page GitHub

  • [CVPR 2024] MicroCinema: A Divide-and-Conquer Approach for Text-to-Video Generation
    ArXiv Project Page Demo

  • [CVPR 2024] Generative Rendering: Controllable 4D-Guided Video Generation with 2D Diffusion Models
    ArXiv Project Page

  • [CVPR 2024] PEEKABOO: Interactive Video Generation via Masked-Diffusion
    ArXiv Project Page GitHub Demo

  • [CVPR 2024] EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
    ArXiv Project Page GitHub

  • [CVPR 2024] A Recipe for Scaling up Text-to-Video Generation with Text-free Videos
    ArXiv Project Page GitHub

  • [CVPR 2024] BIVDiff: A Training-free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion Models
    ArXiv Project Page

  • [CVPR 2024] Mind the Time: Scaled Spatiotemporal Transformers for Text-to-Video Synthesis
    ArXiv Project Page

  • [CVPR 2024] MotionDirector: Motion Customization of Text-to-Video Diffusion Models
    ArXiv GitHub

  • [CVPR 2024] Hierarchical Patch-wise Diffusion Models for High-Resolution Video Generation
    Paper Project Page

  • [CVPR 2024] DiffPerformer: Iterative Learning of Consistent Latent Guidance for Diffusion-based Human Video Generation
    Paper GitHub

  • [CVPR 2024] Grid Diffusion Models for Text-to-Video Generation
    ArXiv GitHub Demo

  • [ECCV 2024] Emu Video: Factorizing Text-to-Video Generation by Explicit Image Conditioning
    ArXiv Project Page

  • [ECCV 2024] W.A.L.T.: Photorealistic Video Generation with Diffusion Models
    Paper Project Page

  • [ECCV 2024] MoVideo: Motion-Aware Video Generation with Diffusion Models
    Paper

  • [ECCV 2024] DrivingDiffusion: Layout-Guided Multi-View Driving Scenarios Video Generation with Latent Diffusion Model
    Paper Project Page GitHub

  • [ECCV 2024] MagDiff: Multi-Alignment Diffusion for High-Fidelity Video Generation and Editing
    Paper

  • [ECCV 2024] HARIVO: Harnessing Text-to-Image Models for Video Generation
    Paper Project Page

  • [ECCV 2024] MEVG: Multi-event Video Generation with Text-to-Video Models
    Paper Project Page

  • [NeurIPS 2024] DEMO: Enhancing Motion in Text-to-Video Generation with Decomposed Encoding and Conditioning
    Paper GitHub

  • [ICML 2024] Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
    Paper Project Page GitHub Hugging Face

  • [ICLR 2024] VDT: General-purpose Video Diffusion Transformers via Mask Modeling
    ArXiv Project Page GitHub

  • [ICLR 2024] VersVideo: Leveraging Enhanced Temporal Diffusion Models for Versatile Video Generation
    Paper

  • [AAAI 2024] Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos
    ArXiv Project Page GitHub

  • [AAAI 2024] E2HQV: High-Quality Video Generation from Event Camera via Theory-Inspired Model-Aided Deep Learning
    ArXiv

  • [AAAI 2024] ConditionVideo: Training-Free Condition-Guided Text-to-Video Generation
    ArXiv Project Page GitHub

  • [AAAI 2024] F3-Pruning: A Training-Free and Generalized Pruning Strategy towards Faster and Finer Text to-Video Synthesis
    ArXiv

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [CVPR 2023] Align your Latents: High-resolution Video Synthesis with Latent Diffusion Models
    ArXiv Project Page GitHub

  • [CVPR 2023] Text2Video-Zero: Text-to-image Diffusion Models are Zero-shot Video Generators
    Paper Project Page GitHub Demo

  • [CVPR 2023] Video Probabilistic Diffusion Models in Projected Latent Space
    Paper GitHub

  • [ICCV 2023] PYOCO: Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models
    Paper Project Page

  • [ICCV 2023] Gen-1: Structure and Content-guided Video Synthesis with Diffusion Models
    Paper Project Page

  • [NeurIPS 2023] Video Diffusion Models
    ArXiv Project Page

  • [NeurIPS 2023] UniPi: Learning Universal Policies via Text-Guided Video Generation
    Paper Project Page GitHub

  • [NeurIPS 2023] VideoComposer: Compositional Video Synthesis with Motion Controllability
    ArXiv Project Page GitHub

  • [ICLR 2023] CogVideo: Large-scale Pretraining for Text-to-video Generation via Transformers
    Paper GitHub Demo

  • [ICLR 2023] Make-A-Video: Text-to-video Generation without Text-video Data
    ArXiv Project Page GitHub

  • [ICLR 2023] Phenaki: Variable Length Video Generation From Open Domain Textual Description
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

🖼️ Image-to-Video (I2V) Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] IPRO: Identity-Preserving Image-to-Video Generation via Reward-Guided Optimization
    Paper Project Page GitHub

  • [CVPR 2026] ALG: Improving Motion in Image-to-Video Models via Adaptive Low-Pass Guidance
    Paper Project Page GitHub

  • [CVPR 2026] ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation
    Paper Project Page

  • [CVPR 2026] FlexiMMT: Let Your Image Move with Your Motion! - Implicit Multi-Object Multi-Motion Transfer
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
    Paper Project Page GitHub

  • [ICLR 2026] MAGREF: Masked Guidance for Any-Reference Video Generation with Subject Disentanglement
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Anchor Frame Bridging for Coherent First-Last Frame Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] LumosX: Relate Any Identities with Their Attributes for Personalized Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] ReactID: Synchronizing Realistic Actions and Identity in Personalized Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Beyond Skeletons: Learning Animation Directly from Driving Videos with Same2X Training Strategy
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Unleashing Guidance Without Classifiers for Human-Object Interaction Animation
    Paper Project Page GitHub Hugging Face

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] MotionStone: Decoupled Motion Intensity Modulation with Diffusion Transformer for Image-to-Video Generation
    ArXiv

  • [CVPR 2025] MotionPro: A Precise Motion Controller for Image-to-Video Generation
    ArXiv GitHub

  • [CVPR 2025] Through-The-Mask: Mask-based Motion Trajectories for Image-to-Video Generation
    ArXiv

  • [CVPR 2025] Extrapolating and Decoupling Image-to-Video Generation Models: Motion Modeling is Easier Than You Think
    ArXiv

  • [CVPR 2025] I2VGuard: Safeguarding Images against Misuse in Diffusion-based Image-to-Video Models
    Paper

  • [CVPR 2025] LeviTor: 3D Trajectory Oriented Image-to-Video Synthesis
    Paper GitHub

  • [ICCV 2025] AnyI2V: Animating Any Conditional Image with Motion Control
    Paper GitHub

  • [ICCV 2025] Versatile Transition Generation with Image-to-Video Diffusion
    Paper

  • [ICCV 2025] TIP‑I2V: A Million‑Scale Real Text and Image Prompt Dataset for Image‑to‑Video Generation
    Paper Project Page GitHub Hugging Face

  • [ICCV 2025] Unified Video Generation via Next‑Set Prediction in Continuous Domain
    Paper

  • [NeurIPS 2025] GenRec: Unifying Video Generation and Recognition with Diffusion Models
    Paper GitHub Hugging Face

  • [ICCV 2025] Precise Action‑to‑Video Generation Through Visual Action Prompts
    Paper Project Page

  • [ICCV 2025] STIV: Scalable Text and Image Conditioned Video Generation
    Paper Project Page Hugging Face

  • [ICLR 2025] FrameBridge: Improving Image‑to‑Video Generation with Bridge Models
    Paper GitHub Project Page

  • [ICLR 2025] SG-I2V: Self-Guided Trajectory Control in Image-to-Video Generation
    Paper Project Page GitHub

  • [ICLR 2025] Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
    Paper

  • [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
    Paper Project Page GitHub Hugging Face

  • [NeurIPS 2025] MotionRAG: Motion Retrieval-Augmented Image-to-Video Generation
    Paper GitHub Hugging Face

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] Animate Anyone: Consistent and Controllable Image-to-video Synthesis for Character Animation
    ArXiv Project Page GitHub

  • [CVPR 2024] Your Image Is My Video: Reshaping the Receptive Field via Image-to-Video Differentiable AutoAugmentation and Fusion
    Paper

  • [CVPR 2024] TRIP: Temporal Residual Learning with Image Noise Prior for Image-to-Video Diffusion Models
    Paper

  • [CVPR 2024] Enhanced Motion-Text Alignment for Image-to-Video Transfer Learning
    Paper

  • [CVPR 2024] Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
    Paper GitHub

  • [ECCV 2024] MOFA-Video: Controllable Image Animation via Generative Motion Field Adaptions in Frozen Image-to-Video Diffusion Model
    Paper GitHub

  • [ECCV 2024] $\mathrm R2$-Tuning: Efficient Image-to-Video Transfer Learning for Video Temporal Grounding
    Paper GitHub

  • [ECCV 2024] PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
    Paper GitHub

  • [ECCV 2024] Rethinking Image-to-Video Adaptation: An Object-Centric Perspective
    Paper

  • [NeurIPS 2024] TPC: Test-time Procrustes Calibration for Diffusion-based Human Image Animation
    Paper

  • [NeurIPS 2024] Identifying and Solving Conditional Image Leakage in Image-to-Video Diffusion Model
    Paper GitHub

  • [ICML 2024] Video-LaVIT: Unified Video-Language Pre-training with Decoupled Visual-Motional Tokenization
    Paper Project Page GitHub Hugging Face

  • [SIGGRAPH 2024] I2V-Adapter: A General Image-to-Video Adapter for Diffusion Models
    Paper GitHub

  • [SIGGRAPH 2024] Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling
    Paper GitHub

  • [AAAI 2024] Continuous Piecewise-Affine Based Motion Model for Image Animation
    Paper GitHub

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [ICCV 2023] DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion
    Paper GitHub

  • [ICCV 2023] Disentangling Spatial and Temporal Learning for Efficient Image-to-Video Transfer Learning
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

✂️ Video-to-Video (V2V) Editing

✨ 2026

✅ Published Papers

  • [AAAI 2026] Vid-CamEdit: Video Camera Trajectory Editing with Generative Rendering from Estimated Geometry
    Paper Project Page GitHub

  • [AAAI 2026] FAME: Fairness-aware Attention-modulated Video Editing
    Paper

  • [CVPR 2026] Ditto: Scaling Instruction-Based Video Editing with a High-Quality Synthetic Dataset
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] VideoCoF: Unified Video Editing with Temporal Reasoner
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] V-RGBX: Video Editing with Accurate Controls over Intrinsic Properties
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
    Paper Project Page GitHub

  • [CVPR 2026] EasyV2V: A High-quality Instruction-based Video Editing Framework
    Paper Project Page

  • [CVPR 2026] VIVA: VLM-Guided Instruction-Based Video Editing with Reward Optimization
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Generative Video Motion Editing with 3D Point Tracks
    Paper Project Page

  • [CVPR 2026] NOVA: Sparse Control, Dense Synthesis for Pair-Free Video Editing
    Paper GitHub Hugging Face

  • [CVPR 2026] EditCtrl: Disentangled Local and Global Control for Real-Time Generative Video Editing
    Paper Project Page GitHub

  • [CVPR 2026] PropFly: Learning to Propagate via On-the-Fly Supervision from Pre-trained Video Diffusion Models
    Paper Project Page

  • [CVPR 2026] FFP-300K: Scaling First-Frame Propagation for Generalizable Video Editing
    Paper Project Page

  • [ICLR 2026] Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
    Paper Project Page GitHub

  • [ICLR 2026] UNIC: Unified In-Context Video Editing
    Paper Project Page

  • [ICLR 2026] LoRA-Edit: Controllable First-Frame-Guided Video Editing via Mask-Aware LoRA Fine-Tuning
    Paper Project Page GitHub

  • [ICLR 2026] DreamSwapV: Mask-guided Subject Swapping for Any Customized Video Editing
    Paper GitHub Hugging Face

  • [ICLR 2026] Any-to-Bokeh: Arbitrary-Subject Video Refocusing with Video Diffusion Model
    Paper Project Page GitHub

  • [ICLR 2026] FlowGuide: Precision-Guided Enhancement for Face Image and Video Editing
    Paper GitHub

  • [ICLR 2026] DragStream: Streaming Drag-Oriented Interactive Video Manipulation: Drag Anything, Anytime!
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Follow-Your-Creation: Empowering 4D Creation through Video Inpainting
    Paper Project Page GitHub

  • [ICLR 2026] Follow-Your-Motion: Video Motion Transfer via Efficient Spatial-Temporal Decoupled Finetuning
    Paper Project Page

  • [ICLR 2026] FastVMT: Eliminating Redundancy in Video Motion Transfer
    Paper Project Page GitHub

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
    Paper

  • [CVPR 2025] VideoDirector: Precise Video Editing via Text-to-Video Models
    Paper GitHub

  • [CVPR 2025] VideoSPatS: Video SPatiotemporal Splines for Disentangled Occlusion, Appearance and Motion Modeling and Editing
    Paper

  • [CVPR 2025] Align-A-Video: Deterministic Reward Tuning of Image Diffusion Models for Consistent Video Editing
    Paper

  • [CVPR 2025] Unity in Diversity: Video Editing via Gradient-Latent Purification
    Paper

  • [CVPR 2025] VEU-Bench: Towards Comprehensive Understanding of Video Editing
    Paper GitHub

  • [CVPR 2025] SketchVideo: Sketch-based Video Generation and Editing
    Paper GitHub

  • [CVPR 2025] FATE: Full-head Gaussian Avatar with Textural Editing from Monocular Video
    Paper GitHub

  • [CVPR 2025] Visual Prompting for One-shot Controllable Video Editing without Inversion
    Paper

  • [CVPR 2025] FADE: Frequency-Aware Diffusion Model Factorization for Video Editing
    Paper GitHub

  • [ICCV 2025] VACE: All-in-One Video Creation and Editing
    Paper Project Page GitHub

  • [ICCV 2025] Reangle-A-Video: 4D Video Generation as Video-to-Video Translation
    Paper Project Page

  • [ICCV 2025] DIVE: Taming DINO for Subject-Driven Video Editing
    Paper Project Page

  • [ICCV 2025] DynamicFace: High-Quality and Consistent Face Swapping for Image and Video using Composable 3D Facial Priors
    Paper Project Page

  • [ICCV 2025] QK-Edit: Revisiting Attention-based Injection in MM-DiT for Image and Video Editing
    Paper

  • [ICCV 2025] Teleportraits: Training-Free People Insertion into Any Scene
    Paper

  • [ICLR 2025] VideoGrain: Modulating Space-Time Attention for Multi-Grained Video Editing
    Paper GitHub

  • [NeurIPS 2025] REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
    Paper Project Page

  • [AAAI 2025] FreeMask: Rethinking the Importance of Attention Masks for Zero-Shot Video Editing
    Paper GitHub

  • [AAAI 2025] EditBoard: Towards a Comprehensive Evaluation Benchmark for Text-Based Video Editing Models
    Paper GitHub

  • [AAAI 2025] VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment
    Paper GitHub

  • [AAAI 2025] Re-Attentional Controllable Video Diffusion Editing
    Paper GitHub

  • [WACV 2025] IP-FaceDiff: Identity-Preserving Facial Video Editing with Diffusion
    Paper

  • [WACV 2025] SST-EM: Advanced Metrics for Evaluating Semantic, Spatial and Temporal Aspects in Video Editing
    Paper GitHub

  • [WACV 2025] MagicStick: Controllable Video Editing via Control Handle Transformations
    Paper

  • [WACV 2025] Ada-VE: Training-Free Consistent Video Editing Using Adaptive Motion Prior
    Paper

  • [WACV 2025] FastVideoEdit: Leveraging Consistency Models for Efficient Text-to-Video Editing
    Paper GitHub

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] A Video is Worth 256 Bases: Spatial-Temporal Expectation-Maximization Inversion for Zero-Shot Video Editing
    Paper GitHub

  • [CVPR 2024] VidToMe: Video Token Merging for Zero-Shot Video Editing
    Paper GitHub

  • [CVPR 2024] Video-P2P: Video Editing with Cross-Attention Control
    Paper GitHub

  • [CVPR 2024] CCEdit: Creative and Controllable Video Editing via Diffusion Models
    Paper

  • [CVPR 2024] RAVE: Randomized Noise Shuffling for Fast and Consistent Video Editing with Diffusion Models
    Paper GitHub

  • [CVPR 2024] DynVideo-E: Harnessing Dynamic NeRF for Large-Scale Motion- and View-Change Human-Centric Video Editing
    Paper

  • [CVPR 2024] MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
    Paper

  • [CVPR 2024] MotionEditor: Editing Video Motion via Content-Aware Diffusion
    Paper GitHub

  • [CVPR 2024] CAMEL: CAusal Motion Enhancement Tailored for Lifting Text-Driven Video Editing
    Paper GitHub

  • [ICLR 2024] Ground-A-Video: Zero-shot Grounded Video Editing using Text-to-image Diffusion Models
    Paper GitHub

  • [ICLR 2024] Video Decomposition Prior: Editing Videos Layer by Layer
    Paper

  • [ICLR 2024] FLATTEN: optical FLow-guided ATTENtion for consistent text-to-video editing
    Paper

  • [ICLR 2024] TokenFlow: Consistent Diffusion Features for Consistent Video Editing
    Paper GitHub

  • [ECCV 2024] VIDEOSHOP: Localized Semantic Video Editing with Noise-Extrapolated Diffusion Inversion
    Paper GitHub

  • [ECCV 2024] DragVideo: Interactive Drag-Style Video Editing
    Paper GitHub

  • [ECCV 2024] WAVE: Warping DDIM Inversion Features for Zero-Shot Text-to-Video Editing
    Paper

  • [ECCV 2024] DreamMotion: Space-Time Self-similar Score Distillation for Zero-Shot Video Editing
    Paper

  • [ECCV 2024] Object-Centric Diffusion for Efficient Video Editing
    Paper GitHub

  • [ECCV 2024] Video Editing via Factorized Diffusion Distillation
    Paper

  • [ECCV 2024] SAVE: Protagonist Diversification with Structure Agnostic Video Editing
    Paper

  • [ECCV 2024] DNI: Dilutional Noise Initialization for Diffusion Video Editing
    Paper

  • [ECCV 2024] MagDiff: Multi-alignment Diffusion for High-Fidelity Video Generation and Editing
    Paper

  • [ECCV 2024] DeCo: Decoupled Human-Centered Diffusion Video Editing with Motion Consistency
    Paper

💡 Pre-Print Papers

⇧ Back to ToC

🕹️ Controllable Video Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] OmniVDiff: Omni Controllable Video Diffusion for Generation and Understanding
    Paper Project Page GitHub

  • [AAAI 2026] MotionFlow: Attention-Driven Motion Transfer in Video Diffusion Models
    Paper Project Page

  • [CVPR 2026] Stand-In: A Lightweight and Plug-and-Play Identity Control for Video Generation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] First Frame Is the Place to Go for Video Content Customization
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] UCPE: Unified Camera Positional Encoding for Controlled Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
    Paper

  • [CVPR 2026] ReDirector: Creating Any-Length Video Retakes with Rotary Camera Encoding
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] NeoVerse: Enhancing 4D World Model with in-the-wild Monocular Videos
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] SeeU: Seeing the Unseen World via 4D Dynamics-aware Generation
    Paper Project Page GitHub

  • [CVPR 2026] VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] Goal Force: Teaching Video Models To Accomplish Physics-Conditioned Goals
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
    Paper Project Page GitHub

  • [CVPR 2026] Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models
    Paper Project Page GitHub

  • [CVPR 2026] ID-Crafter: VLM-Grounded Online RL for Compositional Multi-Subject Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] PerpetualWonder: Long-Horizon Action-Conditioned 4D Scene Generation
    Paper Project Page GitHub

  • [CVPR 2026] BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] RealWonder: Real-Time Physical Action-Conditioned Video Generation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026] LAMP: Language-Assisted Motion Planning for Controllable Video Generation
    Paper GitHub Hugging Face

  • [CVPR 2026] WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
    Paper GitHub

  • [CVPR 2026] StableWorld: Towards Stable and Consistent Long Interactive Video Generation
    Paper Project Page GitHub

  • [CVPR 2026] SMRABooth: Subject and Motion Representation Alignment for Customized Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] Controllable Video Generation with Provable Disentanglement
    Paper

  • [ICLR 2026] 3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation
    Paper

  • [ICLR 2026] MotionStream: Real-Time Video Generation with Interactive Motion Controls
    Paper Project Page GitHub

  • [ICLR 2026] Time-to-Move: Training-Free Motion-Controlled Video Generation via Dual-Clock Denoising
    Paper Project Page GitHub

  • [ICLR 2026] Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
    Paper Project Page GitHub

  • [ICLR 2026] Pusa V1.0: Unlocking Temporal Control in Pretrained Video Diffusion Models via Vectorized Timestep Adaptation
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] TTOM: Test-Time Optimization and Memorization for Compositional Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] MIMIC: Mask-Injected Manipulation Video Generation with Interaction Control
    Paper

  • [ICLR 2026] Target-Aware Video Diffusion Models
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] NewtonGen: Physics-consistent and Controllable Text-to-Video Generation via Neural Newtonian Dynamics
    Paper Project Page GitHub

  • [ICLR 2026] MATRIX: Mask Track Alignment for Interaction-aware Video Generation
    Paper Project Page GitHub

  • [ICLR 2026] ConsisDrive: Identity-Preserving Driving World Models for Video Generation by Instance Mask
    Paper Project Page

  • [ICLR 2026] Animating the Uncaptured: Humanoid Mesh Animation with Video Diffusion Models
    Paper Project Page

  • [ICLR 2026] ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
    Paper Project Page GitHub

  • [ICLR 2026] Light-X: Generative 4D Video Rendering with Camera and Illumination Control
    Paper Project Page GitHub Hugging Face

  • [ICLR 2026] Generative View Stitching
    Paper Project Page GitHub

  • [ICLR 2026] Learning Video Generation for Robotic Manipulation with Collaborative Trajectory Control
    Paper Project Page GitHub

  • [ICLR 2026] LightCtrl: Training-free Controllable Video Relighting
    Paper GitHub

💡 Pre-Print Papers

✨ 2025

✅ Published Papers

  • [CVPR 2025] IM-Zero: Instance-level Motion Controllable Video Generation in a Zero-shot Manner
    Paper

  • [CVPR 2025] AnimateAnything: Consistent and Controllable Animation for Video Generation
    Paper

  • [CVPR 2025] Customized Condition Controllable Generation for Video Soundtrack
    Paper GitHub

  • [CVPR 2025] StarGen: A Spatiotemporal Autoregression Framework with Video Diffusion Model for Scalable and Controllable Scene Generation
    Paper

  • [ICCV 2025] Perception-as-Control: Fine-grained Controllable Image Animation with 3D-aware Motion Representation
    Paper Project Page GitHub

  • [ICCV 2025] MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
    Paper GitHub

  • [ICCV 2025] MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control
    Paper Project Page GitHub

  • [ICCV 2025] InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models
    Paper Project Page GitHub

  • [ICCV 2025] Free-Form Motion Control (SynFMC): Controlling the 6D Poses of Camera and Objects in Video Generation
    Paper Project Page GitHub

  • [ICCV 2025] RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera Control
    Paper Project Page GitHub

  • [ICCV 2025] MagicMotion: Video Generation with a Smart Director
    Paper Project Page GitHub

  • [ICCV 2025] UniMLVG: Unified Framework for Multi-view Long Video Generation with Comprehensive Control Capabilities for Autonomous Driving
    Paper Project Page GitHub

  • [ICLR 2025] MotionClone: Training-Free Motion Cloning for Controllable Video Generation
    Paper GitHub

  • [AAAI 2025] CAGE: Unsupervised Visual Composition and Animation for Controllable Video Generation
    Paper GitHub

  • [AAAI 2025] TrackGo: A Flexible and Efficient Method for Controllable Video Generation
    Paper

  • [WACV 2025] Fine-grained Controllable Video Generation via Object Appearance and Context
    Paper

💡 Pre-Print Papers

✨ 2024

✅ Published Papers

  • [CVPR 2024] 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
    Paper GitHub

  • [CVPR 2024] Panacea: Panoramic and Controllable Video Generation for Autonomous Driving
    Paper GitHub

  • [AAAI 2024] Decouple Content and Motion for Conditional Image-to-Video Generation
    Paper

💡 Pre-Print Papers

✨ 2023

✅ Published Papers

  • [CVPR 2023] Conditional Image-to-Video Generation with Latent Flow Diffusion Models
    Paper GitHub

💡 Pre-Print Papers

⇧ Back to ToC

🗣️ Audio-Driven Video Generation

✨ 2026

✅ Published Papers

  • [AAAI 2026] EchoMimicV3: EchoMimicV3: 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation
    Paper Project Page GitHub Hugging Face

  • [AAAI 2026] FantasyTalking2: FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation
    Paper Project Page GitHub Hugging Face

  • [CVPR 2026 Findings] UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
    Paper

  • [CVPR 2026] Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural Conversation
    Paper Project Page GitHub

  • [CVPR 2026] ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
    Paper Project Page

  • [CVPR 2026] StreamAvatar: *Streaming Diffusion Models for

Truncated — view the full README on GitHub.

aaai
cvpr
deep-learning
diffusion-models
eccv
iccv
iclr
icml
image-to-video
image-to-video-generation
neurips
nips
siggraph
text-to-video
text-to-video-generation
video-editing
video-generation
video-to-video
video-to-video-translation

Contributors

QuenithAI

48 commits

SHYuanBest

4 commits

Haodong-Yan

2 commits

feifeiobama

2 commits