Recent Advances in On-Policy Distillation for Multimodal LLMs
HTML
21
110 commits
updated Sep 25, 2026
A curated, auto-refreshed list of multimodal On-Policy Distillation (OPD / OPSD) papers β organized by Image QA Β· Video QA Β· Audio QA (plus generation, speculative decoding, and embodied/VLA).
π Live interactive reader β searchable, filterable, bilingual (EN / δΈζ), one click, no install Β Β·Β instant mirror (no Pages needed)
What is OPD? C1: the student samples its own trajectories y ~ Ο_student(Β·|x) during training; C2: a teacher provides per-token / sequence-level supervision on those student-generated samples. OPSD is the special case where the teacher is the same model conditioned on privileged information.
Each paper is tagged with arXiv link Β· date Β· first-author affiliation Β· code Β· β stars Β· citations. β Stars and citations are refreshed daily by a GitHub Action (β via GitHub API; citations via Semantic Scholar). For four-point summaries per paper, open the interactive reader.
π Stats last updated: 2026-09-25 08:49 UTC
| Subfield | # |
|---|---|
| πΌοΈ Image QA / VQA / Visual Reasoning | 14 |
| π¬ Video QA / Video Reasoning / Temporal Grounding | 7 |
| π Audio QA / Speech | 7 |
| π¨ Image / Video Generation (Diffusion Β· Flow) | 12 |
| β‘ Multimodal Speculative-Decoding Distillation | 4 |
| π€ Embodied / VLA / GUI Visual Agents | 8 |
| Total | 52 |
On-policy distillation that transfers reasoning into vision-language models and trains on VQA / visual-reasoning rollouts.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts | link | 2026-06-09 | Microsoft | β | β | 0 |
| Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization | link | 2026-06-08 | OPPO AI Center | GitHub | 2 | 2 |
| Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation | link | 2026-06-07 | Peking University | β | β | 3 |
| Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization | link | 2026-06-05 | Tianjin University | GitHub | 7 | 0 |
| ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation | link | 2026-06-04 | Shanghai AI Laboratory | GitHub | 22 | 8 |
| Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation | link | 2026-06-04 | Tsinghua University | β | β | 0 |
| Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding | link | 2026-05-30 | KAIST | β | β | 7 |
| Visual-Advantage On-Policy Distillation for Vision-Language Models | link | 2026-05-21 | Institute of Automation, CAS | β | β | 12 |
| Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation | link | 2026-05-18 | ISCAS | GitHub | 320 | 49 |
| DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation | link | 2026-05-15 | NVIDIA Research | β | β | 0 |
| Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe | link | 2026-05-05 | Zhejiang University | GitHub | 57 | 36 |
| Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL | link | 2026-04-30 | HKUST (GZ) | GitHub | 101 | 7 |
| KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA | link | 2026-01-30 | Chapman University | GitHub | 2 | 0 |
| VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation | link | 2025-10-27 | University of Tuebingen | β | β | 21 |
OPD / self-distillation for video question answering, video reasoning and temporal grounding (incl. closely-related AoTD, VITAL).
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning | link | 2026-06-10 | Shanghai Innovation Institute | β | β | 2 |
| World Model Self-Distillation: Training World Models to Solve General Tasks | link | 2026-06-10 | University of Bern | β | β | 0 |
| World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning | link | 2026-06-02 | University of Macau | β | β | 1 |
| VISD: Enhancing Video Reasoning via Structured Self-Distillation | link | 2026-05-07 | HUST | β | β | 10 |
| Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation | link | 2026-02-03 | Xiaomi | β | β | 28 |
| π Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning | link | 2025-08-06 | Tsinghua University | β | β | 86 |
| π Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation | link | 2024-12-02 | Shanghai Jiao Tong University | GitHub | 61 | 43 |
Cross-modal transfer of text reasoning into audio/speech, and OPD for audio understanding / ASR.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing | link | 2026-06-14 | Shenzhen University | β | β | 4 |
| Data-Efficient On-Policy Distillation for Automatic Speech Recognition | link | 2026-05-27 | AutoArk-AI | β | β | 1 |
| Qwen3.5-Omni Technical Report | link | 2026-04-17 | Alibaba | β | β | 137 |
| X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs | link | 2026-03-06 | Tencent Hunyuan | β | β | 12 |
| CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation | link | 2026-01-23 | Baidu | β | β | 9 |
| Step-Audio-R1 Technical Report | link | 2025-11-19 | StepFun | GitHub | 699 | 46 |
| Qwen3-Omni Technical Report | link | 2025-09-22 | Alibaba | β | β | 502 |
OPD / self-distillation for diffusion and flow-matching generative models (few-step generation, trajectory self-distillation, adversarial distillation).
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| Knowledge Distillation for Visual Autoregressive Models | link | 2026-06-04 | Qualcomm AI Research | β | β | 1 |
| GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models | link | 2026-05-28 | UCL | β | β | 3 |
| Adversarial Dual On-Policy Distillation from Expressive Teacher | link | 2026-05-26 | NTU | β | β | 0 |
| CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation | link | 2026-05-25 | Zhejiang University | β | β | 3 |
| DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models | link | 2026-05-14 | Fudan University | β | β | 23 |
| AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation | link | 2026-05-13 | NUS | β | β | 15 |
| TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM | link | 2026-05-10 | Renmin University of China | GitHub | 3 | 1 |
| Flow-OPD: On-Policy Distillation for Flow Matching Models | link | 2026-05-08 | USTC | β | β | 21 |
| D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models | link | 2026-05-06 | HKUST | β | β | 15 |
| LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation | link | 2025-12-29 | SII / SJTU | β | β | 7 |
| pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation | link | 2025-10-16 | Stanford University | GitHub | 468 | 25 |
| Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator | link | 2025-03-19 | Γcole Polytechnique | β | β | 6 |
Training on-policy draft models for vision-language models to speed up inference.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding | link | 2025-09-17 | Peking University | β | β | 22 |
| SpecVLM: Fast Speculative Decoding in Vision-Language Models | link | 2025-09-15 | Xi'an Jiaotong University | β | β | 6 |
| Speculative Decoding Reimagined for Multimodal Large Language Models | link | 2025-05-20 | Xiamen University | β | β | 7 |
| MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models | link | 2025-05-15 | Cerebras | β | β | 3 |
The student is a visual agent or VLA policy supervised on its own visual trajectories.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving | link | 2026-06-01 | Univ. of Wisconsin-Madison | β | β | 0 |
| HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents | link | 2026-05-08 | Xiaohongshu | GitHub | 76 | 9 |
| LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning | link | 2026-05-08 | Moore Threads | β | β | 3 |
| Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding | link | 2026-05-01 | IIE, CAS | β | β | 9 |
| Co-Evolving Policy Distillation | link | 2026-04-29 | IIE, CAS | β | β | 4 |
| HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents | link | 2026-04-08 | Tencent | GitHub | 872 | 17 |
| VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation | link | 2026-03-27 | HKUST (GZ) | β | β | 7 |
| Refined Policy Distillation: From VLA Generalists to RL Experts | link | 2025-03-06 | Univ. of Tech. Nuremberg | GitHub | 23 | 28 |
This list is compiled and de-duplicated from three awesome repositories, plus web search for a few multimodal entries missing from them. Full credit to the maintainers of:
Summaries are paraphrased from the papers' arXiv abstracts and may contain errors β please refer to the original papers. To add a paper, edit papers.json; the tables and the interactive reader regenerate automatically. β stars and citations are snapshots that change over time.
101 commits
9 commits
HTML
94.2%
Python
5.8%
Recent Advances in On-Policy Distillation for Multimodal LLMs
HTML
21
110 commits
updated Sep 25, 2026
A curated, auto-refreshed list of multimodal On-Policy Distillation (OPD / OPSD) papers β organized by Image QA Β· Video QA Β· Audio QA (plus generation, speculative decoding, and embodied/VLA).
π Live interactive reader β searchable, filterable, bilingual (EN / δΈζ), one click, no install Β Β·Β instant mirror (no Pages needed)
What is OPD? C1: the student samples its own trajectories y ~ Ο_student(Β·|x) during training; C2: a teacher provides per-token / sequence-level supervision on those student-generated samples. OPSD is the special case where the teacher is the same model conditioned on privileged information.
Each paper is tagged with arXiv link Β· date Β· first-author affiliation Β· code Β· β stars Β· citations. β Stars and citations are refreshed daily by a GitHub Action (β via GitHub API; citations via Semantic Scholar). For four-point summaries per paper, open the interactive reader.
π Stats last updated: 2026-09-25 08:49 UTC
| Subfield | # |
|---|---|
| πΌοΈ Image QA / VQA / Visual Reasoning | 14 |
| π¬ Video QA / Video Reasoning / Temporal Grounding | 7 |
| π Audio QA / Speech | 7 |
| π¨ Image / Video Generation (Diffusion Β· Flow) | 12 |
| β‘ Multimodal Speculative-Decoding Distillation | 4 |
| π€ Embodied / VLA / GUI Visual Agents | 8 |
| Total | 52 |
On-policy distillation that transfers reasoning into vision-language models and trains on VQA / visual-reasoning rollouts.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| Self-Distillation Policy Optimization via Visual Feedback: Bridging Code and Visual Artifacts | link | 2026-06-09 | Microsoft | β | β | 0 |
| Stabilizing On-Policy Distillation for MLLM Reasoning with Global Normalization | link | 2026-06-08 | OPPO AI Center | GitHub | 2 | 2 |
| Thinking Without Images: Internalizing Visual Manipulation with On-Policy Self-Distillation | link | 2026-06-07 | Peking University | β | β | 3 |
| Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization | link | 2026-06-05 | Tianjin University | GitHub | 7 | 0 |
| ViCuR: Visual Cues as Recoverable Privilege for Multimodal On-Policy Distillation | link | 2026-06-04 | Shanghai AI Laboratory | GitHub | 22 | 8 |
| Learning Visual Spatial Planning from Symbolic State via Modality-Gap-Aware Self-Distillation | link | 2026-06-04 | Tsinghua University | β | β | 0 |
| Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding | link | 2026-05-30 | KAIST | β | β | 7 |
| Visual-Advantage On-Policy Distillation for Vision-Language Models | link | 2026-05-21 | Institute of Automation, CAS | β | β | 12 |
| Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation | link | 2026-05-18 | ISCAS | GitHub | 320 | 49 |
| DeltaPrompts: Escaping the Zero-Delta Trap in Multimodal Distillation | link | 2026-05-15 | NVIDIA Research | β | β | 0 |
| Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe | link | 2026-05-05 | Zhejiang University | GitHub | 57 | 36 |
| Beyond SFT-to-RL: Pre-alignment via Black-Box On-Policy Distillation for Multimodal RL | link | 2026-04-30 | HKUST (GZ) | GitHub | 101 | 7 |
| KEPO: Knowledge-Enhanced Preference Optimization for Multimodal Reasoning with Applications to Medical VQA | link | 2026-01-30 | Chapman University | GitHub | 2 | 0 |
| VOLD: Reasoning Transfer from LLMs to Vision-Language Models via On-Policy Distillation | link | 2025-10-27 | University of Tuebingen | β | β | 21 |
OPD / self-distillation for video question answering, video reasoning and temporal grounding (incl. closely-related AoTD, VITAL).
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning | link | 2026-06-10 | Shanghai Innovation Institute | β | β | 2 |
| World Model Self-Distillation: Training World Models to Solve General Tasks | link | 2026-06-10 | University of Bern | β | β | 0 |
| World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning | link | 2026-06-02 | University of Macau | β | β | 1 |
| VISD: Enhancing Video Reasoning via Structured Self-Distillation | link | 2026-05-07 | HUST | β | β | 10 |
| Video-OPD: Efficient Post-Training of Multimodal Large Language Models for Temporal Video Grounding via On-Policy Distillation | link | 2026-02-03 | Xiaomi | β | β | 28 |
| π Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning | link | 2025-08-06 | Tsinghua University | β | β | 86 |
| π Enhancing Video-LLM Reasoning via Agent-of-Thoughts Distillation | link | 2024-12-02 | Shanghai Jiao Tong University | GitHub | 61 | 43 |
Cross-modal transfer of text reasoning into audio/speech, and OPD for audio understanding / ASR.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| OmniOPSD: Rationale-Privileged On-Policy Self-Distillation for Affective Computing | link | 2026-06-14 | Shenzhen University | β | β | 4 |
| Data-Efficient On-Policy Distillation for Automatic Speech Recognition | link | 2026-05-27 | AutoArk-AI | β | β | 1 |
| Qwen3.5-Omni Technical Report | link | 2026-04-17 | Alibaba | β | β | 137 |
| X-OPD: Cross-Modal On-Policy Distillation for Capability Alignment in Speech LLMs | link | 2026-03-06 | Tencent Hunyuan | β | β | 12 |
| CORD: Bridging the Audio-Text Reasoning Gap via Weighted On-policy Cross-modal Distillation | link | 2026-01-23 | Baidu | β | β | 9 |
| Step-Audio-R1 Technical Report | link | 2025-11-19 | StepFun | GitHub | 699 | 46 |
| Qwen3-Omni Technical Report | link | 2025-09-22 | Alibaba | β | β | 502 |
OPD / self-distillation for diffusion and flow-matching generative models (few-step generation, trajectory self-distillation, adversarial distillation).
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| Knowledge Distillation for Visual Autoregressive Models | link | 2026-06-04 | Qualcomm AI Research | β | β | 1 |
| GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models | link | 2026-05-28 | UCL | β | β | 3 |
| Adversarial Dual On-Policy Distillation from Expressive Teacher | link | 2026-05-26 | NTU | β | β | 0 |
| CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation | link | 2026-05-25 | Zhejiang University | β | β | 3 |
| DiffusionOPD: A Unified Perspective of On-Policy Distillation in Diffusion Models | link | 2026-05-14 | Fudan University | β | β | 23 |
| AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation | link | 2026-05-13 | NUS | β | β | 15 |
| TAD: Temporal-Aware Trajectory Self-Distillation for Fast and Accurate Diffusion LLM | link | 2026-05-10 | Renmin University of China | GitHub | 3 | 1 |
| Flow-OPD: On-Policy Distillation for Flow Matching Models | link | 2026-05-08 | USTC | β | β | 21 |
| D-OPSD: On-Policy Self-Distillation for Continuously Tuning Step-Distilled Diffusion Models | link | 2026-05-06 | HKUST | β | β | 15 |
| LiveTalk: Real-Time Multimodal Interactive Video Diffusion via Improved On-Policy Distillation | link | 2025-12-29 | SII / SJTU | β | β | 7 |
| pi-Flow: Policy-Based Few-Step Generation via Imitation Distillation | link | 2025-10-16 | Stanford University | GitHub | 468 | 25 |
| Di$\mathtt{[M]}$O: Distilling Masked Diffusion Models into One-step Generator | link | 2025-03-19 | Γcole Polytechnique | β | β | 6 |
Training on-policy draft models for vision-language models to speed up inference.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| ViSpec: Accelerating Vision-Language Models with Vision-Aware Speculative Decoding | link | 2025-09-17 | Peking University | β | β | 22 |
| SpecVLM: Fast Speculative Decoding in Vision-Language Models | link | 2025-09-15 | Xi'an Jiaotong University | β | β | 6 |
| Speculative Decoding Reimagined for Multimodal Large Language Models | link | 2025-05-20 | Xiamen University | β | β | 7 |
| MASSV: Multimodal Adaptation and Self-Data Distillation for Speculative Decoding of Vision-Language Models | link | 2025-05-15 | Cerebras | β | β | 3 |
The student is a visual agent or VLA policy supervised on its own visual trajectories.
| Paper | arXiv | Date | First-author affiliation | Code | β Stars | Citations |
|---|---|---|---|---|---|---|
| GeoDrive-Bench: Benchmarking Region-Specific Multimodal Reasoning in Autonomous Driving | link | 2026-06-01 | Univ. of Wisconsin-Madison | β | β | 0 |
| HyperEyes: Dual-Grained Efficiency-Aware Reinforcement Learning for Parallel Multimodal Search Agents | link | 2026-05-08 | Xiaohongshu | GitHub | 76 | 9 |
| LiteGUI: Distilling Compact GUI Agents with Reinforcement Learning | link | 2026-05-08 | Moore Threads | β | β | 3 |
| Learn where to Click from Yourself: On-Policy Self-Distillation for GUI Grounding | link | 2026-05-01 | IIE, CAS | β | β | 9 |
| Co-Evolving Policy Distillation | link | 2026-04-29 | IIE, CAS | β | β | 4 |
| HY-Embodied-0.5: Embodied Foundation Models for Real-World Agents | link | 2026-04-08 | Tencent | GitHub | 872 | 17 |
| VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation | link | 2026-03-27 | HKUST (GZ) | β | β | 7 |
| Refined Policy Distillation: From VLA Generalists to RL Experts | link | 2025-03-06 | Univ. of Tech. Nuremberg | GitHub | 23 | 28 |
This list is compiled and de-duplicated from three awesome repositories, plus web search for a few multimodal entries missing from them. Full credit to the maintainers of:
Summaries are paraphrased from the papers' arXiv abstracts and may contain errors β please refer to the original papers. To add a paper, edit papers.json; the tables and the interactive reader regenerate automatically. β stars and citations are snapshots that change over time.
101 commits
9 commits
HTML
94.2%
Python
5.8%