Collection of forcing related autoregressive video Gen
101
11 commits
updated Mar 31, 2026
A curated list of papers, code, and resources about Self-Forcing-style autoregressive video diffusion.
This repo focuses on the research line around:
Self-Forcing family methods aim to reduce the mismatch between:
This repository tracks papers and implementations that directly address this issue for autoregressive video diffusion.
| Year | Title | Paper | Code / Project | Notes |
|---|---|---|---|---|
| 2025 | Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion | NeurIPS 2025 Spotlight · arXiv:2506.08009 | Project | Starting point of the self-forcing line |
| 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation | NeurIPS 2025 Poster · arXiv:2506.09350 | Project | Adversarial student-forcing post-training to convert bidirectional video diffusion into frame-level autoregressive streaming generation (1NFE) |
| 2025 | LongLive: Real-time Interactive Long Video Generation | arXiv:2509.22622 | Code · Project | Frame-level causal AR video generation for real-time interactive long videos via KV-recache, streaming long tuning, and short-window attention with a frame sink |
| 2025 | Rolling Forcing: Autoregressive Long Video Diffusion in Real Time | ICLR 2026 · arXiv:2509.25161 | Code · Project | Rolling/chunk-style long-horizon generation |
| 2025 | Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation | arXiv:2512.04678 | Code | Reward-weighted distillation for streaming quality |
| 2025 | Self-Forcing++: Towards Minute-Scale High-Quality Video Generation | ICLR 2026 · arXiv:2510.02283 | - | Minute-scale long video extension |
| 2025 | End-to-End Training for Autoregressive Video Diffusion via Self-Resampling | arXiv:2512.15702 | - | End-to-end AR training and stabilization |
| 2025 | From Slow Bidirectional to Fast Autoregressive Video Diffusion Models | CVPR 2025 · arXiv:2412.07772 | Code | Bidirectional→causal + Video-DMD distillation for low-latency streaming generation |
| 2025 | Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression | arXiv:2512.05081 | Code · Project | Training-free KV management (Deep Sink + Participative Compression) for long-horizon video extrapolation |
| 2025 | Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation | arXiv:2512.21734 | Project | Real-time infinite interactive portrait animation via chunk-wise generation + temporal knot overlap + running-ahead |
| 2026 | Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation | arXiv:2602.02214 | Code · Project | AR teacher + ODE init + causal distillation |
| 2026 | Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention | arXiv:2602.01801 | Project | Training-free attention acceleration with Temporal Cache (TempCache) + sparse attention (AnnSA/AnnCA) for near-constant throughput long rollouts |
| 2026 | LIVE: Long-horizon Interactive Video World Modeling | arXiv:2602.03747 | Code · Project | Cycle-consistency (forward rollout + reverse recovery) to control long-horizon error accumulation |
| 2026 | Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention | arXiv:2602.04789 | Code | Sparse attention for autoregressive video diffusion via Chunk-Aware Growth + Hierarchical Sparse Attention |
| 2026 | Context Forcing: Consistent Autoregressive Video Generation with Long Context | arXiv:2602.06028 | Code · Project | Aligning student context length to a long-context teacher + Slow-Fast Memory for improved long-video consistency |
| 2026 | Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion | arXiv:2601.21896 | - | Salience-guided KV cache management (PaFu-KV) for AR video diffusion; distills bidirectional teacher signals via a lightweight SEH to retain top-k informative tokens for better long-horizon quality/efficiency trade-off |
| 2026 | Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion | arXiv:2602.07775 | Code · Project | Training-free cache maintenance on top of Self Forcing via attention sink, sliding indices, and rolling/sliding semantics for open-ended minute-scale generation |
| 2026 | Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures | arXiv:2602.09600 | Project | Distilled bidirectional-to-causal AR diffusion for egocentric hand-object interaction video generation with occlusion-invariant hand conditioning and Plücker-ray camera embeddings |
| 2025 | Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft | arXiv:2510.03198 | Project | Hybrid Training + Chained Forward Training + geometry-indexed spatial memory for balancing novel-scene exploration and revisit consistency |
Contributions are welcome.
Please open a PR and add entries with the following format:
If this repository helps your research, please consider citing:
@misc{awesome_self_forcing_video_diffusion,
title={Awesome Self-Forcing for Video Diffusion},
author={Community Contributors},
year={2026},
howpublished={\url{https://github.com/ealicesora/awesome-video-forcing}}
}
Collection of forcing related autoregressive video Gen
101
11 commits
updated Mar 31, 2026
A curated list of papers, code, and resources about Self-Forcing-style autoregressive video diffusion.
This repo focuses on the research line around:
Self-Forcing family methods aim to reduce the mismatch between:
This repository tracks papers and implementations that directly address this issue for autoregressive video diffusion.
| Year | Title | Paper | Code / Project | Notes |
|---|---|---|---|---|
| 2025 | Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion | NeurIPS 2025 Spotlight · arXiv:2506.08009 | Project | Starting point of the self-forcing line |
| 2025 | Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation | NeurIPS 2025 Poster · arXiv:2506.09350 | Project | Adversarial student-forcing post-training to convert bidirectional video diffusion into frame-level autoregressive streaming generation (1NFE) |
| 2025 | LongLive: Real-time Interactive Long Video Generation | arXiv:2509.22622 | Code · Project | Frame-level causal AR video generation for real-time interactive long videos via KV-recache, streaming long tuning, and short-window attention with a frame sink |
| 2025 | Rolling Forcing: Autoregressive Long Video Diffusion in Real Time | ICLR 2026 · arXiv:2509.25161 | Code · Project | Rolling/chunk-style long-horizon generation |
| 2025 | Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation | arXiv:2512.04678 | Code | Reward-weighted distillation for streaming quality |
| 2025 | Self-Forcing++: Towards Minute-Scale High-Quality Video Generation | ICLR 2026 · arXiv:2510.02283 | - | Minute-scale long video extension |
| 2025 | End-to-End Training for Autoregressive Video Diffusion via Self-Resampling | arXiv:2512.15702 | - | End-to-end AR training and stabilization |
| 2025 | From Slow Bidirectional to Fast Autoregressive Video Diffusion Models | CVPR 2025 · arXiv:2412.07772 | Code | Bidirectional→causal + Video-DMD distillation for low-latency streaming generation |
| 2025 | Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression | arXiv:2512.05081 | Code · Project | Training-free KV management (Deep Sink + Participative Compression) for long-horizon video extrapolation |
| 2025 | Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait Animation | arXiv:2512.21734 | Project | Real-time infinite interactive portrait animation via chunk-wise generation + temporal knot overlap + running-ahead |
| 2026 | Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation | arXiv:2602.02214 | Code · Project | AR teacher + ODE init + causal distillation |
| 2026 | Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse Attention | arXiv:2602.01801 | Project | Training-free attention acceleration with Temporal Cache (TempCache) + sparse attention (AnnSA/AnnCA) for near-constant throughput long rollouts |
| 2026 | LIVE: Long-horizon Interactive Video World Modeling | arXiv:2602.03747 | Code · Project | Cycle-consistency (forward rollout + reverse recovery) to control long-horizon error accumulation |
| 2026 | Light Forcing: Accelerating Autoregressive Video Diffusion via Sparse Attention | arXiv:2602.04789 | Code | Sparse attention for autoregressive video diffusion via Chunk-Aware Growth + Hierarchical Sparse Attention |
| 2026 | Context Forcing: Consistent Autoregressive Video Generation with Long Context | arXiv:2602.06028 | Code · Project | Aligning student context length to a long-context teacher + Slow-Fast Memory for improved long-video consistency |
| 2026 | Past- and Future-Informed KV Cache Policy with Salience Estimation in Autoregressive Video Diffusion | arXiv:2601.21896 | - | Salience-guided KV cache management (PaFu-KV) for AR video diffusion; distills bidirectional teacher signals via a lightweight SEH to retain top-k informative tokens for better long-horizon quality/efficiency trade-off |
| 2026 | Rolling Sink: Bridging Limited-Horizon Training and Open-Ended Testing in Autoregressive Video Diffusion | arXiv:2602.07775 | Code · Project | Training-free cache maintenance on top of Self Forcing via attention sink, sliding indices, and rolling/sliding semantics for open-ended minute-scale generation |
| 2026 | Hand2World: Autoregressive Egocentric Interaction Generation via Free-Space Hand Gestures | arXiv:2602.09600 | Project | Distilled bidirectional-to-causal AR diffusion for egocentric hand-object interaction video generation with occlusion-invariant hand conditioning and Plücker-ray camera embeddings |
| 2025 | Memory Forcing: Spatio-Temporal Memory for Consistent Scene Generation on Minecraft | arXiv:2510.03198 | Project | Hybrid Training + Chained Forward Training + geometry-indexed spatial memory for balancing novel-scene exploration and revisit consistency |
Contributions are welcome.
Please open a PR and add entries with the following format:
If this repository helps your research, please consider citing:
@misc{awesome_self_forcing_video_diffusion,
title={Awesome Self-Forcing for Video Diffusion},
author={Community Contributors},
year={2026},
howpublished={\url{https://github.com/ealicesora/awesome-video-forcing}}
}