ealicesora/Awesome-Autoregressive-Video-Diffusion

Collection of forcing related autoregressive video Gen

101

11 commits

updated Mar 31, 2026

See the code

README

Awesome Autoregressive Video Diffusion Awesome

A curated list of papers, code, and resources about Self-Forcing-style autoregressive video diffusion.

This repo focuses on the research line around:

  • train-test gap in autoregressive video generation
  • long-horizon stability
  • streaming / real-time inference
  • distillation and ODE initialization strategies

Table of Contents


What is this list

Self-Forcing family methods aim to reduce the mismatch between:

  • training-time inputs (often teacher-forced or clean context), and
  • inference-time rollout (model-generated context with accumulated errors).

This repository tracks papers and implementations that directly address this issue for autoregressive video diffusion.


Papers

YearTitlePaperCode / ProjectNotes
2025Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionNeurIPS 2025 Spotlight · arXiv:2506.08009ProjectStarting point of the self-forcing line
2025Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationNeurIPS 2025 Poster · arXiv:2506.09350ProjectAdversarial student-forcing post-training to convert bidirectional video diffusion into frame-level autoregressive streaming generation (1NFE)
2025LongLive: Real-time Interactive Long Video GenerationarXiv:2509.22622Code · ProjectFrame-level causal AR video generation for real-time interactive long videos via KV-recache, streaming long tuning, and short-window attention with a frame sink
2025Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeICLR 2026 · arXiv:2509.25161Code · ProjectRolling/chunk-style long-horizon generation
2025Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationarXiv:2512.04678CodeReward-weighted distillation for streaming quality
2025Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationICLR 2026 · arXiv:2510.02283-Minute-scale long video extension
2025End-to-End Training for Autoregressive Video Diffusion via Self-ResamplingarXiv:2512.15702-End-to-end AR training and stabilization
2025From Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsCVPR 2025 · arXiv:2412.07772CodeBidirectional→causal + Video-DMD distillation for low-latency streaming generation
2025Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative CompressionarXiv:2512.05081Code · ProjectTraining-free KV management (Deep Sink + Participative Compression) for long-horizon video extrapolation
2025Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait AnimationarXiv:2512.21734ProjectReal-time infinite interactive portrait animation via chunk-wise generation + temporal knot overlap + running-ahead
2026Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationarXiv:2602.02214Code · ProjectAR teacher + ODE init + causal distillation
2026Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse AttentionarXiv:2602.01801ProjectTraining-free attention acceleration with Temporal Cache (TempCache) + sparse attention (AnnSA/AnnCA) for near-constant throughput long rollouts
2026LIVE: Long-horizon Interactive Video World ModelingarXiv:2602.03747Code · ProjectCycle-consistency (forward rollout + reverse recovery) to control long-horizon error accumulation
2026Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse AttentionarXiv:2602.04789CodeSparse attention for autoregressive video diffusion via Chunk-Aware Growth + Hierarchical Sparse Attention
2026Context Forcing: Consistent Autoregressive Video Generation with Long ContextarXiv:2602.06028Code · ProjectAligning student context length to a long-context teacher + Slow-Fast Memory for improved long-video consistency
2026Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video DiffusionarXiv:2601.21896-Salience-guided KV cache management (PaFu-KV) for AR video diffusion; distills bidirectional teacher signals via a lightweight SEH to retain top-k informative tokens for better long-horizon quality/efficiency trade-off
2026Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video DiffusionarXiv:2602.07775Code · ProjectTraining-free cache maintenance on top of Self Forcing via attention sink, sliding indices, and rolling/sliding semantics for open-ended minute-scale generation
2026Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand GesturesarXiv:2602.09600ProjectDistilled bidirectional-to-causal AR diffusion for egocentric hand-object interaction video generation with occlusion-invariant hand conditioning and Plücker-ray camera embeddings
2025Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on MinecraftarXiv:2510.03198ProjectHybrid Training + Chained Forward Training + geometry-indexed spatial memory for balancing novel-scene exploration and revisit consistency

Contributing

Contributions are welcome.

Please open a PR and add entries with the following format:

  • Title
  • Year
  • Paper link
  • Code link (optional)
  • Project page (optional)
  • 2-5 keywords
  • One-line relevance to self-forcing

Citation

If this repository helps your research, please consider citing:

@misc{awesome_self_forcing_video_diffusion,
  title={Awesome Self-Forcing for Video Diffusion},
  author={Community Contributors},
  year={2026},
  howpublished={\url{https://github.com/ealicesora/awesome-video-forcing}}
}

ealicesora/Awesome-Autoregressive-Video-Diffusion

Collection of forcing related autoregressive video Gen

101

11 commits

updated Mar 31, 2026

See the code

README

Awesome Autoregressive Video Diffusion Awesome

A curated list of papers, code, and resources about Self-Forcing-style autoregressive video diffusion.

This repo focuses on the research line around:

  • train-test gap in autoregressive video generation
  • long-horizon stability
  • streaming / real-time inference
  • distillation and ODE initialization strategies

Table of Contents


What is this list

Self-Forcing family methods aim to reduce the mismatch between:

  • training-time inputs (often teacher-forced or clean context), and
  • inference-time rollout (model-generated context with accumulated errors).

This repository tracks papers and implementations that directly address this issue for autoregressive video diffusion.


Papers

YearTitlePaperCode / ProjectNotes
2025Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionNeurIPS 2025 Spotlight · arXiv:2506.08009ProjectStarting point of the self-forcing line
2025Autoregressive Adversarial Post-Training for Real-Time Interactive Video GenerationNeurIPS 2025 Poster · arXiv:2506.09350ProjectAdversarial student-forcing post-training to convert bidirectional video diffusion into frame-level autoregressive streaming generation (1NFE)
2025LongLive: Real-time Interactive Long Video GenerationarXiv:2509.22622Code · ProjectFrame-level causal AR video generation for real-time interactive long videos via KV-recache, streaming long tuning, and short-window attention with a frame sink
2025Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeICLR 2026 · arXiv:2509.25161Code · ProjectRolling/chunk-style long-horizon generation
2025Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationarXiv:2512.04678CodeReward-weighted distillation for streaming quality
2025Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationICLR 2026 · arXiv:2510.02283-Minute-scale long video extension
2025End-to-End Training for Autoregressive Video Diffusion via Self-ResamplingarXiv:2512.15702-End-to-end AR training and stabilization
2025From Slow Bidirectional to Fast Autoregressive Video Diffusion ModelsCVPR 2025 · arXiv:2412.07772CodeBidirectional→causal + Video-DMD distillation for low-latency streaming generation
2025Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative CompressionarXiv:2512.05081Code · ProjectTraining-free KV management (Deep Sink + Participative Compression) for long-horizon video extrapolation
2025Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait AnimationarXiv:2512.21734ProjectReal-time infinite interactive portrait animation via chunk-wise generation + temporal knot overlap + running-ahead
2026Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationarXiv:2602.02214Code · ProjectAR teacher + ODE init + causal distillation
2026Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse AttentionarXiv:2602.01801ProjectTraining-free attention acceleration with Temporal Cache (TempCache) + sparse attention (AnnSA/AnnCA) for near-constant throughput long rollouts
2026LIVE: Long-horizon Interactive Video World ModelingarXiv:2602.03747Code · ProjectCycle-consistency (forward rollout + reverse recovery) to control long-horizon error accumulation
2026Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse AttentionarXiv:2602.04789CodeSparse attention for autoregressive video diffusion via Chunk-Aware Growth + Hierarchical Sparse Attention
2026Context Forcing: Consistent Autoregressive Video Generation with Long ContextarXiv:2602.06028Code · ProjectAligning student context length to a long-context teacher + Slow-Fast Memory for improved long-video consistency
2026Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video DiffusionarXiv:2601.21896-Salience-guided KV cache management (PaFu-KV) for AR video diffusion; distills bidirectional teacher signals via a lightweight SEH to retain top-k informative tokens for better long-horizon quality/efficiency trade-off
2026Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video DiffusionarXiv:2602.07775Code · ProjectTraining-free cache maintenance on top of Self Forcing via attention sink, sliding indices, and rolling/sliding semantics for open-ended minute-scale generation
2026Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand GesturesarXiv:2602.09600ProjectDistilled bidirectional-to-causal AR diffusion for egocentric hand-object interaction video generation with occlusion-invariant hand conditioning and Plücker-ray camera embeddings
2025Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on MinecraftarXiv:2510.03198ProjectHybrid Training + Chained Forward Training + geometry-indexed spatial memory for balancing novel-scene exploration and revisit consistency

Contributing

Contributions are welcome.

Please open a PR and add entries with the following format:

  • Title
  • Year
  • Paper link
  • Code link (optional)
  • Project page (optional)
  • 2-5 keywords
  • One-line relevance to self-forcing

Citation

If this repository helps your research, please consider citing:

@misc{awesome_self_forcing_video_diffusion,
  title={Awesome Self-Forcing for Video Diffusion},
  author={Community Contributors},
  year={2026},
  howpublished={\url{https://github.com/ealicesora/awesome-video-forcing}}
}