yifanyu/I-DLM-8B

Model

13

stars

60

commits

1

linked in READMEs

Apr 15, 2026

updated

conversational
custom_code
diffusion-lm
feature-extraction
introspective-decoding
safetensors
sdar
text-generation
transformers
Browse cluster: Large Language Model Fine-tuning and Deployment

README

I-DLM-8B

Introspective Diffusion Language Model (8B) — a diffusion language model converted from Qwen3-8B that matches AR quality while enabling parallel token generation.

[Project Page] [Paper] [Code]

Highlights

  • First DLM to match same-scale AR quality across 15 benchmarks
  • Introspective Strided Decoding (ISD): single-pass generation + verification with p/q acceptance criterion
  • AR-compatible serving via SGLang (paged KV cache, continuous batching, CUDA graphs)
  • 2.9–4.1× higher throughput than prior DLMs at high concurrency

Results

Quality (I-DLM-8B vs baselines)

BenchmarkI-DLM-8BQwen3-8B (AR)LLaDA-2.1-mini (16B)SDAR (8B)
ARC-C95.895.890.291.9
MMLU82.483.574.578.6
MMLU-Pro73.175.164.856.9
GPQA-D55.658.946.040.2
GPQA54.955.453.3---
GSM8K95.096.089.091.7
MATH-50096.895.885.078.6
MathBench89.193.184.276.9
AIME-2469.673.143.310.0
AIME-2560.865.443.310.0
HumanEval93.395.186.078.7
MBPP92.293.482.172.0
LiveCodeBench-v645.750.330.416.6
IFEval84.784.783.261.4

Usage

Note: This model checkpoint is hosted on HuggingFace for weight distribution. For inference, please use our SGLang-based ISD pipeline which implements the Introspective Strided Decoding algorithm described in the paper. Direct loading via transformers is not currently supported for reproducing paper results.

# Install
git clone https://github.com/Introspective-Diffusion/I-DLM.git
cd I-DLM/inference && bash install.sh

# Launch server
python -m sglang.launch_server \
    --model-path yifanyu/I-DLM-8B \
    --trust-remote-code --tp-size 1 --dtype bfloat16 \
    --mem-fraction-static 0.85 --max-running-requests 32 \
    --attention-backend flashinfer --dllm-algorithm IDLMBlockN \
    --dllm-algorithm-config inference/configs/idlm_blockN4_config.yaml \
    --port 30000

# Generate
curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Prove sqrt(2) is irrational."}],"max_tokens":4096}'

See the inference README for detailed setup, evaluation, and benchmarking.

Method

I-DLM recovers introspective consistency (AR models' inherent self-agreement) through:

  1. Strict causal masking across both masked and clean tokens
  2. Logit shift (Dream shift): hidden state at position i predicts token i+1
  3. All-masked training with auto-balanced loss: CE loss on both noisy and clean token positions, dynamically balanced
ModelHuggingFaceDescription
I-DLM-8Byifanyu/I-DLM-8BConverted from Qwen3-8B
I-DLM-32Byifanyu/I-DLM-32BConverted from Qwen3-32B
I-DLM-8B-LoRAyifanyu/I-DLM-8B-lora-r128Gated LoRA adapter (rank=128) for lossless R-ISD

Citation

@article{yu2026introspective,
  title={Introspective Diffusion Language Models},
  author={Yu, Yifan and Jian, Yuqing and Wang, Junxiong and Zhou, Zhongzhu
          and Zhuang, Donglin and Fang, Xinyu and Yanamandra, Sri
          and Wu, Xiaoxia and Wu, Qingyang and Song, Shuaiwen Leon
          and Dao, Tri and Athiwaratkun, Ben and Zou, James
          and Lai, Fan and Xu, Chenfeng},
  journal={arXiv preprint arXiv:2604.11035},
  year={2026}
}

Contributors

yifanyu

59 commits

nielsr

1 commits

yifanyu/I-DLM-8B

Model

13

stars

60

commits

1

linked in READMEs

Apr 15, 2026

updated

conversational
custom_code
diffusion-lm
feature-extraction
introspective-decoding
safetensors
sdar
text-generation
transformers
Browse cluster: Large Language Model Fine-tuning and Deployment

README

I-DLM-8B

Introspective Diffusion Language Model (8B) — a diffusion language model converted from Qwen3-8B that matches AR quality while enabling parallel token generation.

[Project Page] [Paper] [Code]

Highlights

  • First DLM to match same-scale AR quality across 15 benchmarks
  • Introspective Strided Decoding (ISD): single-pass generation + verification with p/q acceptance criterion
  • AR-compatible serving via SGLang (paged KV cache, continuous batching, CUDA graphs)
  • 2.9–4.1× higher throughput than prior DLMs at high concurrency

Results

Quality (I-DLM-8B vs baselines)

BenchmarkI-DLM-8BQwen3-8B (AR)LLaDA-2.1-mini (16B)SDAR (8B)
ARC-C95.895.890.291.9
MMLU82.483.574.578.6
MMLU-Pro73.175.164.856.9
GPQA-D55.658.946.040.2
GPQA54.955.453.3---
GSM8K95.096.089.091.7
MATH-50096.895.885.078.6
MathBench89.193.184.276.9
AIME-2469.673.143.310.0
AIME-2560.865.443.310.0
HumanEval93.395.186.078.7
MBPP92.293.482.172.0
LiveCodeBench-v645.750.330.416.6
IFEval84.784.783.261.4

Usage

Note: This model checkpoint is hosted on HuggingFace for weight distribution. For inference, please use our SGLang-based ISD pipeline which implements the Introspective Strided Decoding algorithm described in the paper. Direct loading via transformers is not currently supported for reproducing paper results.

# Install
git clone https://github.com/Introspective-Diffusion/I-DLM.git
cd I-DLM/inference && bash install.sh

# Launch server
python -m sglang.launch_server \
    --model-path yifanyu/I-DLM-8B \
    --trust-remote-code --tp-size 1 --dtype bfloat16 \
    --mem-fraction-static 0.85 --max-running-requests 32 \
    --attention-backend flashinfer --dllm-algorithm IDLMBlockN \
    --dllm-algorithm-config inference/configs/idlm_blockN4_config.yaml \
    --port 30000

# Generate
curl http://localhost:30000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model":"default","messages":[{"role":"user","content":"Prove sqrt(2) is irrational."}],"max_tokens":4096}'

See the inference README for detailed setup, evaluation, and benchmarking.

Method

I-DLM recovers introspective consistency (AR models' inherent self-agreement) through:

  1. Strict causal masking across both masked and clean tokens
  2. Logit shift (Dream shift): hidden state at position i predicts token i+1
  3. All-masked training with auto-balanced loss: CE loss on both noisy and clean token positions, dynamically balanced
ModelHuggingFaceDescription
I-DLM-8Byifanyu/I-DLM-8BConverted from Qwen3-8B
I-DLM-32Byifanyu/I-DLM-32BConverted from Qwen3-32B
I-DLM-8B-LoRAyifanyu/I-DLM-8B-lora-r128Gated LoRA adapter (rank=128) for lossless R-ISD

Citation

@article{yu2026introspective,
  title={Introspective Diffusion Language Models},
  author={Yu, Yifan and Jian, Yuqing and Wang, Junxiong and Zhou, Zhongzhu
          and Zhuang, Donglin and Fang, Xinyu and Yanamandra, Sri
          and Wu, Xiaoxia and Wu, Qingyang and Song, Shuaiwen Leon
          and Dao, Tri and Athiwaratkun, Ben and Zou, James
          and Lai, Fan and Xu, Chenfeng},
  journal={arXiv preprint arXiv:2604.11035},
  year={2026}
}

Contributors

yifanyu

59 commits

nielsr

1 commits