[ICML 2026] Scaling Beyond Masked Diffusion Language Models
See the codeBy Subham Sekhar Sahoo, Jean-Marie Lamercier, Justin Deschenaux, Zhihan Yang, Jingyu Liu, John Thickstun, Ante Jukic
In this repo, we release the state-of-the-art diffusion language models:
Sahoo et al., "Simple and Effective Masked Diffusion Language Model", NeurIPS 2024.
We pre-train on SlimPajama.
For scaling-law experiments, set:
ALGO = ar / mdlm / esolm / duoMODEL = 6M / 19M / ... / 2121M (Full list)x1e18):
FLOPS = 6 / 10 / 30 / 60 / 100in the following command:
./auto_resubmit.sh -n 5 -m <MODEL> -f <FLOPS> -b 32 -N 1 -t chinchilla-mdlm scripts/<ALGO>/train_slim_mdlm.sh
We use Nvidia's Nemotron-Pre-Training-Dataset for pre-training the models which is now available on HuggingFace.
To train the 1.7B (non-embedding parameters) model, set:
ALGO = ar / mdlm / esolm / duoPHASE = 1 / 2
in the following command:./auto_resubmit.sh -n 10 -m 2121M -b 2 -N 16 -D nvidia -p <PHASE> -t ar scripts/<ALGO>/train.sh
1.7B Checkpoints will be released on March 1st, 2026.
This repository was built off of MDLM, DUO, and Eso-LMs.
@inproceedings{
sahoo2026scaling,
title={Scaling Beyond Masked Diffusion Language Models},
author={Subham Sekhar Sahoo and Jean-Marie Lemercier and Zhihan Yang and Justin Deschenaux and Jingyu Liu and John Thickstun and Ante Juki{\'c}},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=xKJG1CNtQM}
}
19 commits
Python
76.3%
Shell
23.7%
[ICML 2026] Scaling Beyond Masked Diffusion Language Models
See the codeBy Subham Sekhar Sahoo, Jean-Marie Lamercier, Justin Deschenaux, Zhihan Yang, Jingyu Liu, John Thickstun, Ante Jukic
In this repo, we release the state-of-the-art diffusion language models:
Sahoo et al., "Simple and Effective Masked Diffusion Language Model", NeurIPS 2024.
We pre-train on SlimPajama.
For scaling-law experiments, set:
ALGO = ar / mdlm / esolm / duoMODEL = 6M / 19M / ... / 2121M (Full list)x1e18):
FLOPS = 6 / 10 / 30 / 60 / 100in the following command:
./auto_resubmit.sh -n 5 -m <MODEL> -f <FLOPS> -b 32 -N 1 -t chinchilla-mdlm scripts/<ALGO>/train_slim_mdlm.sh
We use Nvidia's Nemotron-Pre-Training-Dataset for pre-training the models which is now available on HuggingFace.
To train the 1.7B (non-embedding parameters) model, set:
ALGO = ar / mdlm / esolm / duoPHASE = 1 / 2
in the following command:./auto_resubmit.sh -n 10 -m 2121M -b 2 -N 16 -D nvidia -p <PHASE> -t ar scripts/<ALGO>/train.sh
1.7B Checkpoints will be released on March 1st, 2026.
This repository was built off of MDLM, DUO, and Eso-LMs.
@inproceedings{
sahoo2026scaling,
title={Scaling Beyond Masked Diffusion Language Models},
author={Subham Sekhar Sahoo and Jean-Marie Lemercier and Zhihan Yang and Justin Deschenaux and Jingyu Liu and John Thickstun and Ante Juki{\'c}},
booktitle={Forty-third International Conference on Machine Learning},
year={2026},
url={https://openreview.net/forum?id=xKJG1CNtQM}
}
19 commits
Python
76.3%
Shell
23.7%