Xiaohao-Liu/Awesome-Multi-Token-Prediction

A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.

208

52 commits

updated Sep 8, 2026

See the code

README

Awesome Multi-Token Prediction (MTP!)

A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.

Multi-Token Prediction (MTP) is an emerging paradigm that enhances the efficiency and capability of language and multimodal models by allowing them to predict multiple tokens simultaneously. This repository collects recent research and implementations in this exciting direction.


Venue information is included only when it is confirmed by official proceedings, an official conference program, or the paper's current metadata. Unpublished works are labeled as arXiv preprints or technical reports. Year sections follow the formal publication year when available; otherwise, they use the first public release year.

🔬 Recent Papers (2026)

TitleInstitutionVenuePaperCode
LoopMTP: A looped transformer guided by latent multi-token predictionLamarr Institute / University of Bonn / Fraunhofer IAISarXiv 2026arXiv-
AdaMTP: An Adaptive Training Paradigm for Multi-Token PredictionCUHKarXiv 2026arXiv-
AngelSpec: Towards Real-World High Performance Inference with Speculative DecodingTencentarXiv 2026arXivCode
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextNVIDIAarXiv 2026arXivCode
K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language ModelingAlibaba DAMO Academy / Hupan Lab / Zhejiang University / HKUSTarXiv 2026arXivCode
Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMsRenmin University of ChinaarXiv 2026arXivCode
BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous DiffusionSJTU / MMLab CUHK / CASarXiv 2026arXiv-
How Transformers Learn to Plan via Multi-Token PredictionUCLA / SJTU / UPenn / RIKEN AIPCOLM 2026arXiv-
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic EnhancementShenzhen University / Microsoft Research AsiaACL 2026ACL AnthologyCode
Self-Distillation for Multi-Token PredictionTencentarXiv 2026arXiv-
Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingQualcomm AI ResearchICML 2026arXiv-
Efficient Document Parsing via Parallel Token PredictionTencent / Renmin University of ChinaCVPR 2026 FindingsarXiv-
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language ModelsHIT / BaiduarXiv 2026arXiv-
DFlash: Block Diffusion for Flash Speculative Decodingz-labICML 2026arXivCode
Multi-Token Prediction via Self-DistillationUniversity of MarylandICML 2026arXivCode
Temporal Guidance for Large Language ModelsNUAAarXiv 2026arXiv-
Parallel Token Prediction for Language ModelsUC Irvine / CZI / Pyramidal AIICLR 2026arXivCode
Peeking Into The Future For Contextual BiasingSamsung Research AmericaICASSP 2026arXiv-
Fast and Expressive Multi-Byte Prediction with Probabilistic CircuitsUniversity of EdinburghICML 2026arXiv-
Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesMila / CMU / FAIR at MetaICLR 2026arXiv-
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token PredictionNEUICASSP 2026arXiv-
Enhancing Visual Planning with Auxiliary Tasks and Multi-token PredictionMeta / UNC Chapel HillWACV 2026CVFCode
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyFudan UniversityAAAI 2026AAAICode

🔬 Recent Papers (2025)

TitleInstitutionVenuePaperCode
Next-Latent Prediction Transformers Learn Compact World ModelsMicrosoft ResearcharXiv 2025arXivCode
MiMo-V2-Flash Technical ReportXiaomiTechnical Report 2025ReportCode
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token PredictionTencentarXiv 2025arXivCode
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction PotentialApplearXiv 2025arXiv-
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationByteDance SeedNeurIPS 2025arXivProject Page
Improving Large Language Models with Concept-Aware Fine-TuningNTUarXiv 2025arXivCode
DONUT: A Decoder-Only Model for Trajectory PredictionRWTH Aachen UniversityICCV 2025CVFProject Page
Generating Long Semantic IDs in Parallel for RecommendationUC San Diego / Meta AIKDD 2025arXivCode
Pre-Training Curriculum for Multi-Token Prediction in Language ModelsHumboldt-Universität zu BerlinACL 2025ACL AnthologyCode
L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsNUSNeurIPS 2025arXivCode
Multi-Token Prediction Needs RegistersAthena Research CenterNeurIPS 2025NeurIPSCode
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to PosttrainingXiaomiTechnical Report 2025arXivCode
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (Outstanding Paper)Google Research / CMUICML 2025ICML-
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLMTeleAIarXiv 2025arXiv-
VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationSJTUEMNLP 2025ACL Anthology-
On multi-token prediction for efficient LLM inferenceSonyICLR 2025 Workshop (SLLM)ICLR-

📚 Earlier Works & Foundations

TitleInstitutionVenuePaperCode
DeepSeek-V3 Technical ReportDeepSeek AITechnical Report 2024arXiv-
Better & Faster Large Language Models via Multi-token PredictionMetaICML 2024PMLR-
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-trainingUSTCFindings of EMNLP 2020ACL Anthology-

🧠 Speculative Decoding + MTP

TitleInstitutionVenuePaperCode
AngelSpec: Towards Real-World High Performance Inference with Speculative DecodingTencentarXiv 2026arXivCode
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextNVIDIAarXiv 2026arXivCode
DFlash: Block Diffusion for Flash Speculative Decodingz-labICML 2026arXivCode
EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestPeking UniversityNeurIPS 2025NeurIPSCode
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative DecodingKAISTICASSP 2025arXivProject Page
EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyPeking UniversityICML 2024PMLRCode
Hydra: Sequentially-Dependent Draft Heads for Medusa DecodingMITCOLM 2024OpenReviewCode
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsPrinceton UniversityICML 2024PMLRCode
Blockwise Parallel Decoding for Deep Autoregressive ModelsUC BerkeleyNeurIPS 2018NeurIPS-

🧩 Want to Contribute?

We welcome contributions! Please feel free to submit a PR or open an issue if you'd like to add new papers, tools, or correct any mistakes.

✅ Guidelines:

  • Add relevant papers or projects related to MTP.
  • Use consistent formatting.
  • Include links where available.
llm
llm-infer
llm-training

Contributors

Xiaohao-Liu

45 commits

tuteng0915

2 commits

gavin-richie

1 commits

Xiaohao-Liu/Awesome-Multi-Token-Prediction

A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.

208

52 commits

updated Sep 8, 2026

See the code

README

Awesome Multi-Token Prediction (MTP!)

A curated list of papers, tools, and resources on Multi-Token Prediction (MTP) and related techniques in Large Language Models (LLMs), Speech-Language Models (SLMs), and more.

Multi-Token Prediction (MTP) is an emerging paradigm that enhances the efficiency and capability of language and multimodal models by allowing them to predict multiple tokens simultaneously. This repository collects recent research and implementations in this exciting direction.


Venue information is included only when it is confirmed by official proceedings, an official conference program, or the paper's current metadata. Unpublished works are labeled as arXiv preprints or technical reports. Year sections follow the formal publication year when available; otherwise, they use the first public release year.

🔬 Recent Papers (2026)

TitleInstitutionVenuePaperCode
LoopMTP: A looped transformer guided by latent multi-token predictionLamarr Institute / University of Bonn / Fraunhofer IAISarXiv 2026arXiv-
AdaMTP: An Adaptive Training Paradigm for Multi-Token PredictionCUHKarXiv 2026arXiv-
AngelSpec: Towards Real-World High Performance Inference with Speculative DecodingTencentarXiv 2026arXivCode
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextNVIDIAarXiv 2026arXivCode
K-Forcing: Joint Next-K-Token Decoding via Push-Forward Language ModelingAlibaba DAMO Academy / Hupan Lab / Zhejiang University / HKUSTarXiv 2026arXivCode
Pair-In, Pair-Out: Latent Multi-Token Prediction for Efficient LLMsRenmin University of ChinaarXiv 2026arXivCode
BitLM: Unlocking Multi-Token Language Generation with Bitwise Continuous DiffusionSJTU / MMLab CUHK / CASarXiv 2026arXiv-
How Transformers Learn to Plan via Multi-Token PredictionUCLA / SJTU / UPenn / RIKEN AIPCOLM 2026arXiv-
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic EnhancementShenzhen University / Microsoft Research AsiaACL 2026ACL AnthologyCode
Self-Distillation for Multi-Token PredictionTencentarXiv 2026arXiv-
Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingQualcomm AI ResearchICML 2026arXiv-
Efficient Document Parsing via Parallel Token PredictionTencent / Renmin University of ChinaCVPR 2026 FindingsarXiv-
Beyond Token-Level Policy Gradients for Complex Reasoning with Large Language ModelsHIT / BaiduarXiv 2026arXiv-
DFlash: Block Diffusion for Flash Speculative Decodingz-labICML 2026arXivCode
Multi-Token Prediction via Self-DistillationUniversity of MarylandICML 2026arXivCode
Temporal Guidance for Large Language ModelsNUAAarXiv 2026arXiv-
Parallel Token Prediction for Language ModelsUC Irvine / CZI / Pyramidal AIICLR 2026arXivCode
Peeking Into The Future For Contextual BiasingSamsung Research AmericaICASSP 2026arXiv-
Fast and Expressive Multi-Byte Prediction with Probabilistic CircuitsUniversity of EdinburghICML 2026arXiv-
Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesMila / CMU / FAIR at MetaICLR 2026arXiv-
MTP-S2UT: Enhancing Speech-to-Speech Translation Quality with Multi-token PredictionNEUICASSP 2026arXiv-
Enhancing Visual Planning with Auxiliary Tasks and Multi-token PredictionMeta / UNC Chapel HillWACV 2026CVFCode
What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyFudan UniversityAAAI 2026AAAICode

🔬 Recent Papers (2025)

TitleInstitutionVenuePaperCode
Next-Latent Prediction Transformers Learn Compact World ModelsMicrosoft ResearcharXiv 2025arXivCode
MiMo-V2-Flash Technical ReportXiaomiTechnical Report 2025ReportCode
FastMTP: Accelerating LLM Inference with Enhanced Multi-Token PredictionTencentarXiv 2025arXivCode
Your LLM Knows the Future: Uncovering Its Multi-Token Prediction PotentialApplearXiv 2025arXiv-
Chain-of-Action: Trajectory Autoregressive Modeling for Robotic ManipulationByteDance SeedNeurIPS 2025arXivProject Page
Improving Large Language Models with Concept-Aware Fine-TuningNTUarXiv 2025arXivCode
DONUT: A Decoder-Only Model for Trajectory PredictionRWTH Aachen UniversityICCV 2025CVFProject Page
Generating Long Semantic IDs in Parallel for RecommendationUC San Diego / Meta AIKDD 2025arXivCode
Pre-Training Curriculum for Multi-Token Prediction in Language ModelsHumboldt-Universität zu BerlinACL 2025ACL AnthologyCode
L-MTP: Leap Multi-Token Prediction Beyond Adjacent Context for Large Language ModelsNUSNeurIPS 2025arXivCode
Multi-Token Prediction Needs RegistersAthena Research CenterNeurIPS 2025NeurIPSCode
MiMo: Unlocking the Reasoning Potential of Language Model – From Pretraining to PosttrainingXiaomiTechnical Report 2025arXivCode
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction (Outstanding Paper)Google Research / CMUICML 2025ICML-
GOAT-TTS: Expressive and Realistic Speech Generation via A Dual-Branch LLMTeleAIarXiv 2025arXiv-
VocalNet: Speech LLMs with Multi-Token Prediction for Faster and High-Quality GenerationSJTUEMNLP 2025ACL Anthology-
On multi-token prediction for efficient LLM inferenceSonyICLR 2025 Workshop (SLLM)ICLR-

📚 Earlier Works & Foundations

TitleInstitutionVenuePaperCode
DeepSeek-V3 Technical ReportDeepSeek AITechnical Report 2024arXiv-
Better & Faster Large Language Models via Multi-token PredictionMetaICML 2024PMLR-
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-trainingUSTCFindings of EMNLP 2020ACL Anthology-

🧠 Speculative Decoding + MTP

TitleInstitutionVenuePaperCode
AngelSpec: Towards Real-World High Performance Inference with Speculative DecodingTencentarXiv 2026arXivCode
Windowed-MTP: Removing the Full-Context Draft-KV Tax at Million-Token ContextNVIDIAarXiv 2026arXivCode
DFlash: Block Diffusion for Flash Speculative Decodingz-labICML 2026arXivCode
EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestPeking UniversityNeurIPS 2025NeurIPSCode
Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative DecodingKAISTICASSP 2025arXivProject Page
EAGLE: Speculative Sampling Requires Rethinking Feature UncertaintyPeking UniversityICML 2024PMLRCode
Hydra: Sequentially-Dependent Draft Heads for Medusa DecodingMITCOLM 2024OpenReviewCode
Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsPrinceton UniversityICML 2024PMLRCode
Blockwise Parallel Decoding for Deep Autoregressive ModelsUC BerkeleyNeurIPS 2018NeurIPS-

🧩 Want to Contribute?

We welcome contributions! Please feel free to submit a PR or open an issue if you'd like to add new papers, tools, or correct any mistakes.

✅ Guidelines:

  • Add relevant papers or projects related to MTP.
  • Use consistent formatting.
  • Include links where available.
llm
llm-infer
llm-training

Contributors

Xiaohao-Liu

45 commits

tuteng0915

2 commits

gavin-richie

1 commits