Curated collection of research on the limitations of next-token prediction and methods that go beyond it.
33
27 commits
updated Jul 10, 2026
A curated list of research studying the limitations and learning dynamics of next-token prediction, as well as methods that move beyond it.
Next-token prediction has driven many of the breakthroughs in modern language modeling. Yet a growing body of work highlights its limitations such as its struggle with long-range planning and data inefficiency. Is next-token prediction truly the end-all objective for language modeling? What expressive shortcomings arise when models are trained only to predict the next token? Can we extract richer gradients from existing datasets beyond the next token?
If you know of a paper/blog related to this repository that we missed out on, feel free to open a pull request. Please follow this format when suggesting a new paper:
* **Paper Title** <br>
*Author(s)* <br>
Conference, Year. [[Paper]](link) [[Code]](link) [[Website]](link)
The Pitfalls of Next-Token Prediction
Gregor Bachmann, Vaishnavh Nagarajan
ICML, 2024. [Paper] [Code]
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, Mark Ibrahim
NeurIPS, 2024. [Paper]
Reasoning Bias of Next Token Prediction Training
Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu
arXiv, 2025. [Paper]
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
Vaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi Raghunathan
ICML, 2025. [Paper] [Code]
Alternatives To Next Token Prediction In Text Generation — A Survey
Charlie Wyatt, Aditya Joshi, Flora Salim
arXiv, 2025. [Paper]
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
ICML, 2026. [Paper]
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
Mark Rofin, Jalal Naghiyev, Michael Hahn
ICLR, 2026. [Paper] [Code] [Website]
How Transformers Learn to Plan via Multi-Token Prediction
Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
arXiv, 2026. [Paper]
Learn from your own latents and not from tokens: A sample-complexity theory
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
arXiv, 2026. [Paper]
Clarification of scope: We are primarily focused on foundational training methods that augment the learning objective and are broadly applicable during pretraining. We exclude methods that operate atop of a pretrained next-token prediction backbone, e.g., speculative decoding or finetuning-only methods.
This section collates methods that predict several future tokens using a shared model trunk. For generality, we include any approach that augments the training objective with token-level losses beyond the next token, for e.g., future n-gram or forward-reverse predictions.
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, Ming Zhou
EMNLP, 2020. [Paper] [Code]
Image Captioners Are Scalable Vision Learners Too
Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, Lucas Beyer
NeurIPS, 2023. [Paper] [Code]
PaSS: Parallel Speculative Sampling
Giovanni Monea, Armand Joulin, Edouard Grave
arXiv, 2023. [Paper]
Better & Faster Large Language Models via Multi-token Prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
ICML, 2024. [Paper]
DeepSeek-V3 Technical Report
DeepSeek-AI
arXiv, 2024. [Paper] [Code]
Efficient Joint Prediction of Multiple Future Tokens
Kwangjun Ahn, Alex Lamb, John Langford
arXiv, 2025. [Paper]
Improving Large Language Models with Concept-Aware Fine-Tuning
Michael K. Chen, Xikun Zhang, Jiaxing Huang, Dacheng Tao
arXiv, 2025. [Paper] [Code]
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
Chenze Shao, Fandong Meng, Jie Zhou
ICLR, 2025. [Paper] [Code]
The Belief State Transformer
Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Jayden Teoh, Bryon Xu, David Yan, Dinesh Jayaraman, Alex Lamb, John Langford
ICLR, 2025. [Paper] [Code] [Website]
Pre-Training Curriculum for Multi-Token Prediction in Language Models
Ansar Aynetdinov, Alan Akbik
ACL, 2025. [Paper] [Code]
Predicting the Order of Upcoming Tokens Improves Language Modeling
Zayd M. K. Zuhri, Erland Hilman Fuadi, Alham Fikri Aji
ICML, 2026. [Paper] [Code]
Multi-Token Prediction Needs Registers
Anastasios Gerontopoulos, Spyros Gidaris, Nikos Komodakis
NeurIPS, 2025. [Paper] [Code]
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA
arXiv, 2026. [Paper] [Code] [Website]
Efficient Pre-Training with Token Superposition
Bowen Peng, Théo Gigant, Jeffrey Quesnelle
arXiv, 2026. [Paper] [Website]
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
Qimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou, Rui Mao, Wei Chen, Naipeng Chao
ACL, 2026. [Paper] [Code]
We intentionally use “latent” in a broad sense here to keep the categorization simple. It may refer to final-layer hidden states, learned summaries of future tokens, conceptual embeddings of text, or other intermediate representations. As a rule of thumb, methods in this category incorporate auxiliary objectives to reconstruct some latent representation of text during training, rather than relying solely on token-level predictions.
Semformer: Transformer Language Models with Semantic Planning
Yongjing Yin, Junran Ding, Kai Song, Yue Zhang
EMNLP, 2024. [Paper] [Code]
Large Concept Models: Language Modeling in a Sentence Representation Space
Large Concept Models Team (FAIR at Meta)
arXiv, 2024. [Paper] [Code]
LLM Pretraining with Continuous Concepts
Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, Xian Li
arXiv, 2025. [Paper] [Code]
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki, Ioannis Mitliagkas, David Lopez-Paz, Kartik Ahuja
ICLR, 2026. [Paper]
Continuous Autoregressive Language Models
Chenze Shao, Darren Li, Fandong Meng, Jie Zhou
arXiv, 2025. [Paper] [Code] [Website]
Next-Latent Prediction Transformers Learn Compact World Models
Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Pratyusha Sharma, Riashat Islam, Alex Lamb, John Langford
arXiv, 2025. [Paper] [Code] [Website]
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
Xingwei Qu, Shaowen Wang, Zihao Huang, Kai Hua, Fan Yin, Rui-Jie Zhu, Jundong Zhou, Qiyang Min, Zihao Wang, Yizhi Li, Tianyu Zhang, He Xing, Zheng Zhang, Yuxuan Song, Tianyu Zheng, Zhiyuan Zeng, Chenghua Lin, Ge Zhang, Wenhao Huang
arXiv, 2025. [Paper]
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
Yuliang Liu, Yunchong Song, Yixuan Wang, Kewen Ge, Alex Lamb, Qipeng Guo, Kai Chen, Bowen Zhou, Zhouhan Lin
arXiv, 2026. [Paper] [Code]
This section covers methods that augment training sequences in ways that help overcome the myopic biases of standard next-token prediction.
Efficient Training of Language Models to Fill in the Middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen
arXiv, 2022. [Paper]
σ-GPTs: A New Approach to Autoregressive Models
Arnaud Pannatier, Evann Courdier, François Fleuret
ECML, 2024. [Paper] [Code]
Looking beyond the next token
Abitha Thankaraj, Yiding Jiang, J. Zico Kolter, Yonatan Bisk
arXiv, 2025. [Paper]
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
arXiv, 2026. [Paper] [Code]
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
Yinhan Lu, Eric Elmoznino, Léo Gagnon, Sarthak Mittal, Tejas Kasetty, Guillaume Lajoie
arXiv, 2026. [Paper]
This section covers continuous generation methods, i.e., approaches that generate text by iteratively refining a global continuous representation of the entire output, such as through diffusion, flow matching, or energy minimization, rather than emitting each token directly in a single pass. We focus on notable continuous generation methods and papers demonstrating how they address limitations of next-token prediction.
Diffusion Language Models (DLM) are a rapidly growing area of research. We direct readers to this repository for awesome curated list of DLM papers instead.
Discrete Flow Matching
Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman
NeurIPS, 2024. [Paper]
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann
NeurIPS, 2024. [Paper] [Code] [Website]
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, Lingpeng Kong
ICLR, 2025. [Paper] [Code]
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov
ICLR, 2025. [Paper] [Code] [Website]
Large Language Diffusion Models
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li
NeurIPS, 2025. [Paper] [Code] [Website]
Energy-Based Transformers are Scalable Learners and Thinkers
Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han, Hyeonjeong Ha, Aman Chadha, Yilun Du, Heng Ji, Jundong Li, Tariq Iqbal
arXiv, 2025. [Paper] [Code] [Website]
Curated collection of research on the limitations of next-token prediction and methods that go beyond it.
33
27 commits
updated Jul 10, 2026
A curated list of research studying the limitations and learning dynamics of next-token prediction, as well as methods that move beyond it.
Next-token prediction has driven many of the breakthroughs in modern language modeling. Yet a growing body of work highlights its limitations such as its struggle with long-range planning and data inefficiency. Is next-token prediction truly the end-all objective for language modeling? What expressive shortcomings arise when models are trained only to predict the next token? Can we extract richer gradients from existing datasets beyond the next token?
If you know of a paper/blog related to this repository that we missed out on, feel free to open a pull request. Please follow this format when suggesting a new paper:
* **Paper Title** <br>
*Author(s)* <br>
Conference, Year. [[Paper]](link) [[Code]](link) [[Website]](link)
The Pitfalls of Next-Token Prediction
Gregor Bachmann, Vaishnavh Nagarajan
ICML, 2024. [Paper] [Code]
The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, Mark Ibrahim
NeurIPS, 2024. [Paper]
Reasoning Bias of Next Token Prediction Training
Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu
arXiv, 2025. [Paper]
Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
Vaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi Raghunathan
ICML, 2025. [Paper] [Code]
Alternatives To Next Token Prediction In Text Generation — A Survey
Charlie Wyatt, Aditya Joshi, Flora Salim
arXiv, 2025. [Paper]
Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
ICML, 2026. [Paper]
Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
Mark Rofin, Jalal Naghiyev, Michael Hahn
ICLR, 2026. [Paper] [Code] [Website]
How Transformers Learn to Plan via Multi-Token Prediction
Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
arXiv, 2026. [Paper]
Learn from your own latents and not from tokens: A sample-complexity theory
Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
arXiv, 2026. [Paper]
Clarification of scope: We are primarily focused on foundational training methods that augment the learning objective and are broadly applicable during pretraining. We exclude methods that operate atop of a pretrained next-token prediction backbone, e.g., speculative decoding or finetuning-only methods.
This section collates methods that predict several future tokens using a shared model trunk. For generality, we include any approach that augments the training objective with token-level losses beyond the next token, for e.g., future n-gram or forward-reverse predictions.
ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, Ming Zhou
EMNLP, 2020. [Paper] [Code]
Image Captioners Are Scalable Vision Learners Too
Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, Lucas Beyer
NeurIPS, 2023. [Paper] [Code]
PaSS: Parallel Speculative Sampling
Giovanni Monea, Armand Joulin, Edouard Grave
arXiv, 2023. [Paper]
Better & Faster Large Language Models via Multi-token Prediction
Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
ICML, 2024. [Paper]
DeepSeek-V3 Technical Report
DeepSeek-AI
arXiv, 2024. [Paper] [Code]
Efficient Joint Prediction of Multiple Future Tokens
Kwangjun Ahn, Alex Lamb, John Langford
arXiv, 2025. [Paper]
Improving Large Language Models with Concept-Aware Fine-Tuning
Michael K. Chen, Xikun Zhang, Jiaxing Huang, Dacheng Tao
arXiv, 2025. [Paper] [Code]
Beyond Next Token Prediction: Patch-Level Training for Large Language Models
Chenze Shao, Fandong Meng, Jie Zhou
ICLR, 2025. [Paper] [Code]
The Belief State Transformer
Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Jayden Teoh, Bryon Xu, David Yan, Dinesh Jayaraman, Alex Lamb, John Langford
ICLR, 2025. [Paper] [Code] [Website]
Pre-Training Curriculum for Multi-Token Prediction in Language Models
Ansar Aynetdinov, Alan Akbik
ACL, 2025. [Paper] [Code]
Predicting the Order of Upcoming Tokens Improves Language Modeling
Zayd M. K. Zuhri, Erland Hilman Fuadi, Alham Fikri Aji
ICML, 2026. [Paper] [Code]
Multi-Token Prediction Needs Registers
Anastasios Gerontopoulos, Spyros Gidaris, Nikos Komodakis
NeurIPS, 2025. [Paper] [Code]
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
NVIDIA
arXiv, 2026. [Paper] [Code] [Website]
Efficient Pre-Training with Token Superposition
Bowen Peng, Théo Gigant, Jeffrey Quesnelle
arXiv, 2026. [Paper] [Website]
Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
Qimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou, Rui Mao, Wei Chen, Naipeng Chao
ACL, 2026. [Paper] [Code]
We intentionally use “latent” in a broad sense here to keep the categorization simple. It may refer to final-layer hidden states, learned summaries of future tokens, conceptual embeddings of text, or other intermediate representations. As a rule of thumb, methods in this category incorporate auxiliary objectives to reconstruct some latent representation of text during training, rather than relying solely on token-level predictions.
Semformer: Transformer Language Models with Semantic Planning
Yongjing Yin, Junran Ding, Kai Song, Yue Zhang
EMNLP, 2024. [Paper] [Code]
Large Concept Models: Language Modeling in a Sentence Representation Space
Large Concept Models Team (FAIR at Meta)
arXiv, 2024. [Paper] [Code]
LLM Pretraining with Continuous Concepts
Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, Xian Li
arXiv, 2025. [Paper] [Code]
Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki, Ioannis Mitliagkas, David Lopez-Paz, Kartik Ahuja
ICLR, 2026. [Paper]
Continuous Autoregressive Language Models
Chenze Shao, Darren Li, Fandong Meng, Jie Zhou
arXiv, 2025. [Paper] [Code] [Website]
Next-Latent Prediction Transformers Learn Compact World Models
Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Pratyusha Sharma, Riashat Islam, Alex Lamb, John Langford
arXiv, 2025. [Paper] [Code] [Website]
Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
Xingwei Qu, Shaowen Wang, Zihao Huang, Kai Hua, Fan Yin, Rui-Jie Zhu, Jundong Zhou, Qiyang Min, Zihao Wang, Yizhi Li, Tianyu Zhang, He Xing, Zheng Zhang, Yuxuan Song, Tianyu Zheng, Zhiyuan Zeng, Chenghua Lin, Ge Zhang, Wenhao Huang
arXiv, 2025. [Paper]
Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
Yuliang Liu, Yunchong Song, Yixuan Wang, Kewen Ge, Alex Lamb, Qipeng Guo, Kai Chen, Bowen Zhou, Zhouhan Lin
arXiv, 2026. [Paper] [Code]
This section covers methods that augment training sequences in ways that help overcome the myopic biases of standard next-token prediction.
Efficient Training of Language Models to Fill in the Middle
Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen
arXiv, 2022. [Paper]
σ-GPTs: A New Approach to Autoregressive Models
Arnaud Pannatier, Evann Courdier, François Fleuret
ECML, 2024. [Paper] [Code]
Looking beyond the next token
Abitha Thankaraj, Yiding Jiang, J. Zico Kolter, Yonatan Bisk
arXiv, 2025. [Paper]
Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
arXiv, 2026. [Paper] [Code]
Simplifying the Modeling of Arbitrary Conditionals in Natural Language
Yinhan Lu, Eric Elmoznino, Léo Gagnon, Sarthak Mittal, Tejas Kasetty, Guillaume Lajoie
arXiv, 2026. [Paper]
This section covers continuous generation methods, i.e., approaches that generate text by iteratively refining a global continuous representation of the entire output, such as through diffusion, flow matching, or energy minimization, rather than emitting each token directly in a single pass. We focus on notable continuous generation methods and papers demonstrating how they address limitations of next-token prediction.
Diffusion Language Models (DLM) are a rapidly growing area of research. We direct readers to this repository for awesome curated list of DLM papers instead.
Discrete Flow Matching
Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman
NeurIPS, 2024. [Paper]
Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann
NeurIPS, 2024. [Paper] [Code] [Website]
Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, Lingpeng Kong
ICLR, 2025. [Paper] [Code]
Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov
ICLR, 2025. [Paper] [Code] [Website]
Large Language Diffusion Models
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li
NeurIPS, 2025. [Paper] [Code] [Website]
Energy-Based Transformers are Scalable Learners and Thinkers
Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han, Hyeonjeong Ha, Aman Chadha, Yilun Du, Heng Ji, Jundong Li, Tariq Iqbal
arXiv, 2025. [Paper] [Code] [Website]