A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.
208
52 commits
updated Sep 8, 2026
A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.
Multi-Token Prediction (MTP) is an emerging paradigm that enhances the efficiency and capability of language and multimodal models by allowing them to predict multiple tokens simultaneously. This repository collects recent research and implementations in this exciting direction.
Venue information is included only when it is confirmed by official proceedings, an official conference program, or the paper's current metadata. Unpublished works are labeled as arXiv preprints or technical reports. Year sections follow the formal publication year when available; otherwise, they use the first public release year.
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| LoopMTP: A looped transformer guided by latent multi-token prediction | Lamarr Institute / University of Bonn / Fraunhofer IAIS | arXiv 2026 | arXiv | - |
| AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction | CUHK | arXiv 2026 | arXiv | - |
| AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding | Tencent | arXiv 2026 | arXiv | Code |
| Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context | NVIDIA | arXiv 2026 | arXiv | Code |
| K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling | Alibaba DAMO Academy / Hupan Lab / Zhejiang University / HKUST | arXiv 2026 | arXiv | Code |
| Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs | Renmin University of China | arXiv 2026 | arXiv | Code |
| BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion | SJTU / MMLab CUHK / CAS | arXiv 2026 | arXiv | - |
| How Transformers Learn to Plan via Multi-Token Prediction | UCLA / SJTU / UPenn / RIKEN AIP | COLM 2026 | arXiv | - |
| Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement | Shenzhen University / Microsoft Research Asia | ACL 2026 | ACL Anthology | Code |
| Self-Distillation for Multi-Token Prediction | Tencent | arXiv 2026 | arXiv | - |
| Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing | Qualcomm AI Research | ICML 2026 | arXiv | - |
| Efficient Document Parsing via Parallel Token Prediction | Tencent / Renmin University of China | CVPR 2026 Findings | arXiv | - |
| Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models | HIT / Baidu | arXiv 2026 | arXiv | - |
| DFlash: Block Diffusion for Flash Speculative Decoding | z-lab | ICML 2026 | arXiv | Code |
| Multi-Token Prediction via Self-Distillation | University of Maryland | ICML 2026 | arXiv | Code |
| Temporal Guidance for Large Language Models | NUAA | arXiv 2026 | arXiv | - |
| Parallel Token Prediction for Language Models | UC Irvine / CZI / Pyramidal AI | ICLR 2026 | arXiv | Code |
| Peeking Into The Future For Contextual Biasing | Samsung Research America | ICASSP 2026 | arXiv | - |
| Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits | University of Edinburgh | ICML 2026 | arXiv | - |
| Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries | Mila / CMU / FAIR at Meta | ICLR 2026 | arXiv | - |
| MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction | NEU | ICASSP 2026 | arXiv | - |
| Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction | Meta / UNC Chapel Hill | WACV 2026 | CVF | Code |
| What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study | Fudan University | AAAI 2026 | AAAI | Code |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| Next-Latent Prediction Transformers Learn Compact World Models | Microsoft Research | arXiv 2025 | arXiv | Code |
| MiMo-V2-Flash Technical Report | Xiaomi | Technical Report 2025 | Report | Code |
| FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction | Tencent | arXiv 2025 | arXiv | Code |
| Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential | Apple | arXiv 2025 | arXiv | - |
| Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation | ByteDance Seed | NeurIPS 2025 | arXiv | Project Page |
| Improving Large Language Models with Concept-Aware Fine-Tuning | NTU | arXiv 2025 | arXiv | Code |
| DONUT: A Decoder-Only Model for Trajectory Prediction | RWTH Aachen University | ICCV 2025 | CVF | Project Page |
| Generating Long Semantic IDs in Parallel for Recommendation | UC San Diego / Meta AI | KDD 2025 | arXiv | Code |
| Pre-Training Curriculum for Multi-Token Prediction in Language Models | Humboldt-Universität zu Berlin | ACL 2025 | ACL Anthology | Code |
| L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models | NUS | NeurIPS 2025 | arXiv | Code |
| Multi-Token Prediction Needs Registers | Athena Research Center | NeurIPS 2025 | NeurIPS | Code |
| MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining | Xiaomi | Technical Report 2025 | arXiv | Code |
| Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (Outstanding Paper) | Google Research / CMU | ICML 2025 | ICML | - |
| GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM | TeleAI | arXiv 2025 | arXiv | - |
| VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation | SJTU | EMNLP 2025 | ACL Anthology | - |
| On multi-token prediction for efficient LLM inference | Sony | ICLR 2025 Workshop (SLLM) | ICLR | - |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| DeepSeek-V3 Technical Report | DeepSeek AI | Technical Report 2024 | arXiv | - |
| Better & Faster Large Language Models via Multi-token Prediction | Meta | ICML 2024 | PMLR | - |
| ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training | USTC | Findings of EMNLP 2020 | ACL Anthology | - |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding | Tencent | arXiv 2026 | arXiv | Code |
| Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context | NVIDIA | arXiv 2026 | arXiv | Code |
| DFlash: Block Diffusion for Flash Speculative Decoding | z-lab | ICML 2026 | arXiv | Code |
| EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test | Peking University | NeurIPS 2025 | NeurIPS | Code |
| Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding | KAIST | ICASSP 2025 | arXiv | Project Page |
| EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty | Peking University | ICML 2024 | PMLR | Code |
| Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding | MIT | COLM 2024 | OpenReview | Code |
| Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads | Princeton University | ICML 2024 | PMLR | Code |
| Blockwise Parallel Decoding for Deep Autoregressive Models | UC Berkeley | NeurIPS 2018 | NeurIPS | - |
We welcome contributions! Please feel free to submit a PR or open an issue if you'd like to add new papers, tools, or correct any mistakes.
A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.
208
52 commits
updated Sep 8, 2026
A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.
Multi-Token Prediction (MTP) is an emerging paradigm that enhances the efficiency and capability of language and multimodal models by allowing them to predict multiple tokens simultaneously. This repository collects recent research and implementations in this exciting direction.
Venue information is included only when it is confirmed by official proceedings, an official conference program, or the paper's current metadata. Unpublished works are labeled as arXiv preprints or technical reports. Year sections follow the formal publication year when available; otherwise, they use the first public release year.
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| LoopMTP: A looped transformer guided by latent multi-token prediction | Lamarr Institute / University of Bonn / Fraunhofer IAIS | arXiv 2026 | arXiv | - |
| AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction | CUHK | arXiv 2026 | arXiv | - |
| AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding | Tencent | arXiv 2026 | arXiv | Code |
| Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context | NVIDIA | arXiv 2026 | arXiv | Code |
| K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language Modeling | Alibaba DAMO Academy / Hupan Lab / Zhejiang University / HKUST | arXiv 2026 | arXiv | Code |
| Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMs | Renmin University of China | arXiv 2026 | arXiv | Code |
| BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous Diffusion | SJTU / MMLab CUHK / CAS | arXiv 2026 | arXiv | - |
| How Transformers Learn to Plan via Multi-Token Prediction | UCLA / SJTU / UPenn / RIKEN AIP | COLM 2026 | arXiv | - |
| Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement | Shenzhen University / Microsoft Research Asia | ACL 2026 | ACL Anthology | Code |
| Self-Distillation for Multi-Token Prediction | Tencent | arXiv 2026 | arXiv | - |
| Efficient Training-Free Multi-Token Prediction via Embedding-Space Probing | Qualcomm AI Research | ICML 2026 | arXiv | - |
| Efficient Document Parsing via Parallel Token Prediction | Tencent / Renmin University of China | CVPR 2026 Findings | arXiv | - |
| Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language Models | HIT / Baidu | arXiv 2026 | arXiv | - |
| DFlash: Block Diffusion for Flash Speculative Decoding | z-lab | ICML 2026 | arXiv | Code |
| Multi-Token Prediction via Self-Distillation | University of Maryland | ICML 2026 | arXiv | Code |
| Temporal Guidance for Large Language Models | NUAA | arXiv 2026 | arXiv | - |
| Parallel Token Prediction for Language Models | UC Irvine / CZI / Pyramidal AI | ICLR 2026 | arXiv | Code |
| Peeking Into The Future For Contextual Biasing | Samsung Research America | ICASSP 2026 | arXiv | - |
| Fast and Expressive Multi-Byte Prediction with Probabilistic Circuits | University of Edinburgh | ICML 2026 | arXiv | - |
| Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries | Mila / CMU / FAIR at Meta | ICLR 2026 | arXiv | - |
| MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token Prediction | NEU | ICASSP 2026 | arXiv | - |
| Enhancing Visual Planning with Auxiliary Tasks and Multi-token Prediction | Meta / UNC Chapel Hill | WACV 2026 | CVF | Code |
| What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic Study | Fudan University | AAAI 2026 | AAAI | Code |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| Next-Latent Prediction Transformers Learn Compact World Models | Microsoft Research | arXiv 2025 | arXiv | Code |
| MiMo-V2-Flash Technical Report | Xiaomi | Technical Report 2025 | Report | Code |
| FastMTP: Accelerating LLM Inference with Enhanced Multi-Token Prediction | Tencent | arXiv 2025 | arXiv | Code |
| Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential | Apple | arXiv 2025 | arXiv | - |
| Chain-of-Action: Trajectory Autoregressive Modeling for Robotic Manipulation | ByteDance Seed | NeurIPS 2025 | arXiv | Project Page |
| Improving Large Language Models with Concept-Aware Fine-Tuning | NTU | arXiv 2025 | arXiv | Code |
| DONUT: A Decoder-Only Model for Trajectory Prediction | RWTH Aachen University | ICCV 2025 | CVF | Project Page |
| Generating Long Semantic IDs in Parallel for Recommendation | UC San Diego / Meta AI | KDD 2025 | arXiv | Code |
| Pre-Training Curriculum for Multi-Token Prediction in Language Models | Humboldt-Universität zu Berlin | ACL 2025 | ACL Anthology | Code |
| L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language Models | NUS | NeurIPS 2025 | arXiv | Code |
| Multi-Token Prediction Needs Registers | Athena Research Center | NeurIPS 2025 | NeurIPS | Code |
| MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to Posttraining | Xiaomi | Technical Report 2025 | arXiv | Code |
| Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (Outstanding Paper) | Google Research / CMU | ICML 2025 | ICML | - |
| GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLM | TeleAI | arXiv 2025 | arXiv | - |
| VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality Generation | SJTU | EMNLP 2025 | ACL Anthology | - |
| On multi-token prediction for efficient LLM inference | Sony | ICLR 2025 Workshop (SLLM) | ICLR | - |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| DeepSeek-V3 Technical Report | DeepSeek AI | Technical Report 2024 | arXiv | - |
| Better & Faster Large Language Models via Multi-token Prediction | Meta | ICML 2024 | PMLR | - |
| ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training | USTC | Findings of EMNLP 2020 | ACL Anthology | - |
| Title | Institution | Venue | Paper | Code |
|---|---|---|---|---|
| AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding | Tencent | arXiv 2026 | arXiv | Code |
| Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token Context | NVIDIA | arXiv 2026 | arXiv | Code |
| DFlash: Block Diffusion for Flash Speculative Decoding | z-lab | ICML 2026 | arXiv | Code |
| EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time Test | Peking University | NeurIPS 2025 | NeurIPS | Code |
| Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding | KAIST | ICASSP 2025 | arXiv | Project Page |
| EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty | Peking University | ICML 2024 | PMLR | Code |
| Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding | MIT | COLM 2024 | OpenReview | Code |
| Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads | Princeton University | ICML 2024 | PMLR | Code |
| Blockwise Parallel Decoding for Deep Autoregressive Models | UC Berkeley | NeurIPS 2018 | NeurIPS | - |
We welcome contributions! Please feel free to submit a PR or open an issue if you'd like to add new papers, tools, or correct any mistakes.