JaydenTeoh/beyond-next-token-prediction

Curated collection of research on the limitations of next-token prediction and methods that go beyond it.

33

27 commits

updated Jul 10, 2026

See the code

README

Awesome Beyond Next-Token Prediction Papers Awesome

A curated list of research studying the limitations and learning dynamics of next-token prediction, as well as methods that move beyond it.

Next-token prediction has driven many of the breakthroughs in modern language modeling. Yet a growing body of work highlights its limitations such as its struggle with long-range planning and data inefficiency. Is next-token prediction truly the end-all objective for language modeling? What expressive shortcomings arise when models are trained only to predict the next token? Can we extract richer gradients from existing datasets beyond the next token?

Table of Contents

Contributing

If you know of a paper/blog related to this repository that we missed out on, feel free to open a pull request. Please follow this format when suggesting a new paper:

* **Paper Title** <br>
*Author(s)* <br>
Conference, Year. [[Paper]](link) [[Code]](link) [[Website]](link)

Studies/Benchmarks

  • The Pitfalls of Next-Token Prediction
    Gregor Bachmann, Vaishnavh Nagarajan
    ICML, 2024. [Paper] [Code]

  • The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
    Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, Mark Ibrahim
    NeurIPS, 2024. [Paper]

  • Reasoning Bias of Next Token Prediction Training
    Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu
    arXiv, 2025. [Paper]

  • Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
    Vaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi Raghunathan
    ICML, 2025. [Paper] [Code]

  • Alternatives To Next Token Prediction In Text Generation — A Survey
    Charlie Wyatt, Aditya Joshi, Flora Salim
    arXiv, 2025. [Paper]

  • Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
    Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
    ICML, 2026. [Paper]

  • Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
    Mark Rofin, Jalal Naghiyev, Michael Hahn
    ICLR, 2026. [Paper] [Code] [Website]

  • How Transformers Learn to Plan via Multi-Token Prediction
    Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
    arXiv, 2026. [Paper]

  • Learn from your own latents and not from tokens: A sample-complexity theory
    Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
    arXiv, 2026. [Paper]

Methods

Clarification of scope: We are primarily focused on foundational training methods that augment the learning objective and are broadly applicable during pretraining. We exclude methods that operate atop of a pretrained next-token prediction backbone, e.g., speculative decoding or finetuning-only methods.

Multi-Token Prediction Methods

This section collates methods that predict several future tokens using a shared model trunk. For generality, we include any approach that augments the training objective with token-level losses beyond the next token, for e.g., future n-gram or forward-reverse predictions.

  • ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
    Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, Ming Zhou
    EMNLP, 2020. [Paper] [Code]

  • Image Captioners Are Scalable Vision Learners Too
    Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, Lucas Beyer
    NeurIPS, 2023. [Paper] [Code]

  • PaSS: Parallel Speculative Sampling
    Giovanni Monea, Armand Joulin, Edouard Grave
    arXiv, 2023. [Paper]

  • Better & Faster Large Language Models via Multi-token Prediction
    Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
    ICML, 2024. [Paper]

  • DeepSeek-V3 Technical Report
    DeepSeek-AI
    arXiv, 2024. [Paper] [Code]

  • Efficient Joint Prediction of Multiple Future Tokens
    Kwangjun Ahn, Alex Lamb, John Langford
    arXiv, 2025. [Paper]

  • Improving Large Language Models with Concept-Aware Fine-Tuning
    Michael K. Chen, Xikun Zhang, Jiaxing Huang, Dacheng Tao
    arXiv, 2025. [Paper] [Code]

  • Beyond Next Token Prediction: Patch-Level Training for Large Language Models
    Chenze Shao, Fandong Meng, Jie Zhou
    ICLR, 2025. [Paper] [Code]

  • The Belief State Transformer
    Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Jayden Teoh, Bryon Xu, David Yan, Dinesh Jayaraman, Alex Lamb, John Langford
    ICLR, 2025. [Paper] [Code] [Website]

  • Pre-Training Curriculum for Multi-Token Prediction in Language Models
    Ansar Aynetdinov, Alan Akbik
    ACL, 2025. [Paper] [Code]

  • Predicting the Order of Upcoming Tokens Improves Language Modeling
    Zayd M. K. Zuhri, Erland Hilman Fuadi, Alham Fikri Aji
    ICML, 2026. [Paper] [Code]

  • Multi-Token Prediction Needs Registers
    Anastasios Gerontopoulos, Spyros Gidaris, Nikos Komodakis
    NeurIPS, 2025. [Paper] [Code]

  • Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
    NVIDIA
    arXiv, 2026. [Paper] [Code] [Website]

  • Efficient Pre-Training with Token Superposition
    Bowen Peng, Théo Gigant, Jeffrey Quesnelle
    arXiv, 2026. [Paper] [Website]

  • Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
    Qimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou, Rui Mao, Wei Chen, Naipeng Chao
    ACL, 2026. [Paper] [Code]

Latent Prediction Methods

We intentionally use “latent” in a broad sense here to keep the categorization simple. It may refer to final-layer hidden states, learned summaries of future tokens, conceptual embeddings of text, or other intermediate representations. As a rule of thumb, methods in this category incorporate auxiliary objectives to reconstruct some latent representation of text during training, rather than relying solely on token-level predictions.

  • Semformer: Transformer Language Models with Semantic Planning
    Yongjing Yin, Junran Ding, Kai Song, Yue Zhang
    EMNLP, 2024. [Paper] [Code]

  • Large Concept Models: Language Modeling in a Sentence Representation Space
    Large Concept Models Team (FAIR at Meta)
    arXiv, 2024. [Paper] [Code]

  • LLM Pretraining with Continuous Concepts
    Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, Xian Li
    arXiv, 2025. [Paper] [Code]

  • Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
    Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki, Ioannis Mitliagkas, David Lopez-Paz, Kartik Ahuja
    ICLR, 2026. [Paper]

  • Continuous Autoregressive Language Models
    Chenze Shao, Darren Li, Fandong Meng, Jie Zhou
    arXiv, 2025. [Paper] [Code] [Website]

  • Next-Latent Prediction Transformers Learn Compact World Models
    Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Pratyusha Sharma, Riashat Islam, Alex Lamb, John Langford
    arXiv, 2025. [Paper] [Code] [Website]

  • Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
    Xingwei Qu, Shaowen Wang, Zihao Huang, Kai Hua, Fan Yin, Rui-Jie Zhu, Jundong Zhou, Qiyang Min, Zihao Wang, Yizhi Li, Tianyu Zhang, He Xing, Zheng Zhang, Yuxuan Song, Tianyu Zheng, Zhiyuan Zeng, Chenghua Lin, Ge Zhang, Wenhao Huang
    arXiv, 2025. [Paper]

  • Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
    Yuliang Liu, Yunchong Song, Yixuan Wang, Kewen Ge, Alex Lamb, Qipeng Guo, Kai Chen, Bowen Zhou, Zhouhan Lin
    arXiv, 2026. [Paper] [Code]

Sequence Augmentation Techniques

This section covers methods that augment training sequences in ways that help overcome the myopic biases of standard next-token prediction.

  • Efficient Training of Language Models to Fill in the Middle
    Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen
    arXiv, 2022. [Paper]

  • σ-GPTs: A New Approach to Autoregressive Models
    Arnaud Pannatier, Evann Courdier, François Fleuret
    ECML, 2024. [Paper] [Code]

  • Looking beyond the next token
    Abitha Thankaraj, Yiding Jiang, J. Zico Kolter, Yonatan Bisk
    arXiv, 2025. [Paper]

  • Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
    Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
    arXiv, 2026. [Paper] [Code]

  • Simplifying the Modeling of Arbitrary Conditionals in Natural Language
    Yinhan Lu, Eric Elmoznino, Léo Gagnon, Sarthak Mittal, Tejas Kasetty, Guillaume Lajoie
    arXiv, 2026. [Paper]

Continuous Generation Methods

This section covers continuous generation methods, i.e., approaches that generate text by iteratively refining a global continuous representation of the entire output, such as through diffusion, flow matching, or energy minimization, rather than emitting each token directly in a single pass. We focus on notable continuous generation methods and papers demonstrating how they address limitations of next-token prediction.

Diffusion Language Models (DLM) are a rapidly growing area of research. We direct readers to this repository for awesome curated list of DLM papers instead.

  • Discrete Flow Matching
    Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman
    NeurIPS, 2024. [Paper]

  • Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
    Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann
    NeurIPS, 2024. [Paper] [Code] [Website]

  • Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
    Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, Lingpeng Kong
    ICLR, 2025. [Paper] [Code]

  • Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
    Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov
    ICLR, 2025. [Paper] [Code] [Website]

  • Large Language Diffusion Models
    Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li
    NeurIPS, 2025. [Paper] [Code] [Website]

  • Energy-Based Transformers are Scalable Learners and Thinkers
    Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han, Hyeonjeong Ha, Aman Chadha, Yilun Du, Heng Ji, Jundong Li, Tariq Iqbal
    arXiv, 2025. [Paper] [Code] [Website]

language-model
machine-learning
multi-token-prediction
next-token-prediction
sequence-to-sequence
transformers

JaydenTeoh/beyond-next-token-prediction

Curated collection of research on the limitations of next-token prediction and methods that go beyond it.

33

27 commits

updated Jul 10, 2026

See the code

README

Awesome Beyond Next-Token Prediction Papers Awesome

A curated list of research studying the limitations and learning dynamics of next-token prediction, as well as methods that move beyond it.

Next-token prediction has driven many of the breakthroughs in modern language modeling. Yet a growing body of work highlights its limitations such as its struggle with long-range planning and data inefficiency. Is next-token prediction truly the end-all objective for language modeling? What expressive shortcomings arise when models are trained only to predict the next token? Can we extract richer gradients from existing datasets beyond the next token?

Table of Contents

Contributing

If you know of a paper/blog related to this repository that we missed out on, feel free to open a pull request. Please follow this format when suggesting a new paper:

* **Paper Title** <br>
*Author(s)* <br>
Conference, Year. [[Paper]](link) [[Code]](link) [[Website]](link)

Studies/Benchmarks

  • The Pitfalls of Next-Token Prediction
    Gregor Bachmann, Vaishnavh Nagarajan
    ICML, 2024. [Paper] [Code]

  • The Factorization Curse: Which Tokens You Predict Underlie the Reversal Curse and More
    Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, Mark Ibrahim
    NeurIPS, 2024. [Paper]

  • Reasoning Bias of Next Token Prediction Training
    Pengxiao Lin, Zhongwang Zhang, Zhi-Qin John Xu
    arXiv, 2025. [Paper]

  • Roll the dice & look before you leap: Going beyond the creative limits of next-token prediction
    Vaishnavh Nagarajan, Chen Henry Wu, Charles Ding, Aditi Raghunathan
    ICML, 2025. [Paper] [Code]

  • Alternatives To Next Token Prediction In Text Generation — A Survey
    Charlie Wyatt, Aditya Joshi, Flora Salim
    arXiv, 2025. [Paper]

  • Autoregressive Language Models are Secretly Energy-Based Models: Insights into the Lookahead Capabilities of Next-Token Prediction
    Mathieu Blondel, Michael E. Sander, Germain Vivier-Ardisson, Tianlin Liu, Vincent Roulet
    ICML, 2026. [Paper]

  • Understanding the Emergence of Seemingly Useless Features in Next-Token Predictors
    Mark Rofin, Jalal Naghiyev, Michael Hahn
    ICLR, 2026. [Paper] [Code] [Website]

  • How Transformers Learn to Plan via Multi-Token Prediction
    Jianhao Huang, Zhanpeng Zhou, Renqiu Xia, Baharan Mirzasoleiman, Weijie Su, Wei Huang
    arXiv, 2026. [Paper]

  • Learn from your own latents and not from tokens: A sample-complexity theory
    Daniel J. Korchinski, Alessandro Favero, Matthieu Wyart
    arXiv, 2026. [Paper]

Methods

Clarification of scope: We are primarily focused on foundational training methods that augment the learning objective and are broadly applicable during pretraining. We exclude methods that operate atop of a pretrained next-token prediction backbone, e.g., speculative decoding or finetuning-only methods.

Multi-Token Prediction Methods

This section collates methods that predict several future tokens using a shared model trunk. For generality, we include any approach that augments the training objective with token-level losses beyond the next token, for e.g., future n-gram or forward-reverse predictions.

  • ProphetNet: Predicting Future N-gram for Sequence-to-Sequence Pre-training
    Weizhen Qi, Yu Yan, Yeyun Gong, Dayiheng Liu, Nan Duan, Jiusheng Chen, Ruofei Zhang, Ming Zhou
    EMNLP, 2020. [Paper] [Code]

  • Image Captioners Are Scalable Vision Learners Too
    Michael Tschannen, Manoj Kumar, Andreas Steiner, Xiaohua Zhai, Neil Houlsby, Lucas Beyer
    NeurIPS, 2023. [Paper] [Code]

  • PaSS: Parallel Speculative Sampling
    Giovanni Monea, Armand Joulin, Edouard Grave
    arXiv, 2023. [Paper]

  • Better & Faster Large Language Models via Multi-token Prediction
    Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière, David Lopez-Paz, Gabriel Synnaeve
    ICML, 2024. [Paper]

  • DeepSeek-V3 Technical Report
    DeepSeek-AI
    arXiv, 2024. [Paper] [Code]

  • Efficient Joint Prediction of Multiple Future Tokens
    Kwangjun Ahn, Alex Lamb, John Langford
    arXiv, 2025. [Paper]

  • Improving Large Language Models with Concept-Aware Fine-Tuning
    Michael K. Chen, Xikun Zhang, Jiaxing Huang, Dacheng Tao
    arXiv, 2025. [Paper] [Code]

  • Beyond Next Token Prediction: Patch-Level Training for Large Language Models
    Chenze Shao, Fandong Meng, Jie Zhou
    ICLR, 2025. [Paper] [Code]

  • The Belief State Transformer
    Edward S. Hu, Kwangjun Ahn, Qinghua Liu, Haoran Xu, Manan Tomar, Ada Langford, Jayden Teoh, Bryon Xu, David Yan, Dinesh Jayaraman, Alex Lamb, John Langford
    ICLR, 2025. [Paper] [Code] [Website]

  • Pre-Training Curriculum for Multi-Token Prediction in Language Models
    Ansar Aynetdinov, Alan Akbik
    ACL, 2025. [Paper] [Code]

  • Predicting the Order of Upcoming Tokens Improves Language Modeling
    Zayd M. K. Zuhri, Erland Hilman Fuadi, Alham Fikri Aji
    ICML, 2026. [Paper] [Code]

  • Multi-Token Prediction Needs Registers
    Anastasios Gerontopoulos, Spyros Gidaris, Nikos Komodakis
    NeurIPS, 2025. [Paper] [Code]

  • Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
    NVIDIA
    arXiv, 2026. [Paper] [Code] [Website]

  • Efficient Pre-Training with Token Superposition
    Bowen Peng, Théo Gigant, Jeffrey Quesnelle
    arXiv, 2026. [Paper] [Website]

  • Toward Consistent World Models with Multi-Token Prediction and Latent Semantic Enhancement
    Qimin Zhong, Hao Liao, Haiming Qin, Mingyang Zhou, Rui Mao, Wei Chen, Naipeng Chao
    ACL, 2026. [Paper] [Code]

Latent Prediction Methods

We intentionally use “latent” in a broad sense here to keep the categorization simple. It may refer to final-layer hidden states, learned summaries of future tokens, conceptual embeddings of text, or other intermediate representations. As a rule of thumb, methods in this category incorporate auxiliary objectives to reconstruct some latent representation of text during training, rather than relying solely on token-level predictions.

  • Semformer: Transformer Language Models with Semantic Planning
    Yongjing Yin, Junran Ding, Kai Song, Yue Zhang
    EMNLP, 2024. [Paper] [Code]

  • Large Concept Models: Language Modeling in a Sentence Representation Space
    Large Concept Models Team (FAIR at Meta)
    arXiv, 2024. [Paper] [Code]

  • LLM Pretraining with Continuous Concepts
    Jihoon Tack, Jack Lanchantin, Jane Yu, Andrew Cohen, Ilia Kulikov, Janice Lan, Shibo Hao, Yuandong Tian, Jason Weston, Xian Li
    arXiv, 2025. [Paper] [Code]

  • Beyond Multi-Token Prediction: Pretraining LLMs with Future Summaries
    Divyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki, Ioannis Mitliagkas, David Lopez-Paz, Kartik Ahuja
    ICLR, 2026. [Paper]

  • Continuous Autoregressive Language Models
    Chenze Shao, Darren Li, Fandong Meng, Jie Zhou
    arXiv, 2025. [Paper] [Code] [Website]

  • Next-Latent Prediction Transformers Learn Compact World Models
    Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Pratyusha Sharma, Riashat Islam, Alex Lamb, John Langford
    arXiv, 2025. [Paper] [Code] [Website]

  • Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
    Xingwei Qu, Shaowen Wang, Zihao Huang, Kai Hua, Fan Yin, Rui-Jie Zhu, Jundong Zhou, Qiyang Min, Zihao Wang, Yizhi Li, Tianyu Zhang, He Xing, Zheng Zhang, Yuxuan Song, Tianyu Zheng, Zhiyuan Zeng, Chenghua Lin, Ge Zhang, Wenhao Huang
    arXiv, 2025. [Paper]

  • Next Concept Prediction in Discrete Latent Space Leads to Stronger Language Models
    Yuliang Liu, Yunchong Song, Yixuan Wang, Kewen Ge, Alex Lamb, Qipeng Guo, Kai Chen, Bowen Zhou, Zhouhan Lin
    arXiv, 2026. [Paper] [Code]

Sequence Augmentation Techniques

This section covers methods that augment training sequences in ways that help overcome the myopic biases of standard next-token prediction.

  • Efficient Training of Language Models to Fill in the Middle
    Mohammad Bavarian, Heewoo Jun, Nikolas Tezak, John Schulman, Christine McLeavey, Jerry Tworek, Mark Chen
    arXiv, 2022. [Paper]

  • σ-GPTs: A New Approach to Autoregressive Models
    Arnaud Pannatier, Evann Courdier, François Fleuret
    ECML, 2024. [Paper] [Code]

  • Looking beyond the next token
    Abitha Thankaraj, Yiding Jiang, J. Zico Kolter, Yonatan Bisk
    arXiv, 2025. [Paper]

  • Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining
    Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
    arXiv, 2026. [Paper] [Code]

  • Simplifying the Modeling of Arbitrary Conditionals in Natural Language
    Yinhan Lu, Eric Elmoznino, Léo Gagnon, Sarthak Mittal, Tejas Kasetty, Guillaume Lajoie
    arXiv, 2026. [Paper]

Continuous Generation Methods

This section covers continuous generation methods, i.e., approaches that generate text by iteratively refining a global continuous representation of the entire output, such as through diffusion, flow matching, or energy minimization, rather than emitting each token directly in a single pass. We focus on notable continuous generation methods and papers demonstrating how they address limitations of next-token prediction.

Diffusion Language Models (DLM) are a rapidly growing area of research. We direct readers to this repository for awesome curated list of DLM papers instead.

  • Discrete Flow Matching
    Itai Gat, Tal Remez, Neta Shaul, Felix Kreuk, Ricky T. Q. Chen, Gabriel Synnaeve, Yossi Adi, Yaron Lipman
    NeurIPS, 2024. [Paper]

  • Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
    Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann
    NeurIPS, 2024. [Paper] [Code] [Website]

  • Beyond Autoregression: Discrete Diffusion for Complex Reasoning and Planning
    Jiacheng Ye, Jiahui Gao, Shansan Gong, Lin Zheng, Xin Jiang, Zhenguo Li, Lingpeng Kong
    ICLR, 2025. [Paper] [Code]

  • Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models
    Marianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang, Zhixuan Qi, Jiaqi Han, Subham Sekhar Sahoo, Volodymyr Kuleshov
    ICLR, 2025. [Paper] [Code] [Website]

  • Large Language Diffusion Models
    Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, Chongxuan Li
    NeurIPS, 2025. [Paper] [Code] [Website]

  • Energy-Based Transformers are Scalable Learners and Thinkers
    Alexi Gladstone, Ganesh Nanduru, Md Mofijul Islam, Peixuan Han, Hyeonjeong Ha, Aman Chadha, Yilun Du, Heng Ji, Jundong Li, Tariq Iqbal
    arXiv, 2025. [Paper] [Code] [Website]

language-model
machine-learning
multi-token-prediction
next-token-prediction
sequence-to-sequence
transformers