yaof20/DenseMixer

Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient

Python

68

36 commits

updated Aug 3, 2025

See the code

README

🎨 DenseMixer 🎨

Improving MoE Post-Training with Precise Router Gradients (Blog)

What is DenseMixer? • Key Features • Experiments • Quick Start • Efficiency • Citation

DenseMixer is a novel MoE post-training technique that empowers MoE training with more precise router gradient estimation, consistently outperforming conventional MoE training in downstream tasks.

What is DenseMixer?

DenseMixer addresses the non-differentiable Top-K routing problem in MoE training via straight-through estimator (STE). This enables more precise router gradients by computing all experts's output during forward pass for better gradient estimation during backward pass. For more technical details, please refer to our blog.

🚀 Key Features

  • Plug-and-play: Zero code changes required
  • Universal compatibility: Works with any MoE using Top-K routing
  • Performance gains: Consistently outperforms conventional MoE training
  • Parameter-efficient: Compatible with LoRA and other PEFT methods
  • No inference overhead: Zero impact on model inference speed

📈 Experiments

DenseMixer consistently outperforms conventional MoE training across:

  • Model scales: 7B, 14B, 30B parameters
  • Architectures: With/without shared experts
  • Training methods: From scratch and up-cycling
  • Data types: Instruction tuning and long reasoning data

DenseMixer Performance Gains

Reproducible Experiments: For detailed training scripts, configurations, and evaluation code, please refer to the experiments folder.

📊 Qwen1.5-MoE-A2.7B (14B): +2.2% average improvement across 7 tasks

Full Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Base Model38.6938.8432.3116.8318.2028.2916.5327.10
Frozen Router53.3735.2037.1082.2033.0138.2932.7544.56
Conventional53.4234.6036.4381.8029.2537.8033.0243.76
DenseMixer55.1635.4039.6883.4033.8340.5633.9045.99
Gain+1.74+0.80+3.25+1.60+4.58+2.76+0.88+2.23

LoRA Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Frozen Router -lora46.7731.4036.5871.0030.3030.1928.0839.19
Conventional -lora43.8934.0038.4164.8028.8037.9926.1439.15
DenseMixer -lora47.2435.4038.4171.8031.8040.2029.2542.01
Gain+3.35+1.40+0.00+7.00+3.00+2.21+3.11+2.86
📊 OLMoE-1B-7B: +2.9% average improvement across 7 tasks

Full Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Base Model15.8519.8010.970.205.707.4011.0910.14
Frozen Router44.8817.87.2372.8022.5036.0528.2932.79
Conventional45.9423.418.9274.6022.3535.9926.8935.44
DenseMixer49.0025.1220.7377.4023.0240.6432.5538.35
Gain+3.06+1.72+1.81+2.80+0.67+4.65+5.66+2.91

LoRA Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Frozen Router -lora45.0324.217.0755.8021.3037.7028.1932.76
Conventional -lora44.5824.215.8560.2021.6037.3026.2232.85
DenseMixer -lora45.3826.216.4866.6024.7040.8029.4335.66
Gain+0.80+2.00+0.63+6.40+3.10+3.50+3.21+2.81
📊 Qwen3-30B-A3B: +3.7% improvement on GPQA-Diamond

Nemotron-Code Dataset (35K samples):

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Base Model65.2460.0653.6016.8548.94
Conventional92.2386.8980.8032.2667.21
DenseMixer93.5989.0282.0034.3168.80
Gain+1.36+2.13+1.20+2.05+1.59

Stanford S1 Dataset (1K samples):

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Base Model38.8820.637.7134.8172.8034.97
Conventional54.8061.5645.6357.3393.4062.54
DenseMixer58.5263.8545.8358.5193.6064.06
Gain+3.72+2.29+0.20+1.18+0.20+1.52

Results shown for temperature=0.6, top_p=0.95 decoding parameters

Additional Decoding Parameters:

Temperature=0.7, top_p=0.8

Nemotron-Code:

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Conventional91.0185.3776.8029.3965.59
DenseMixer91.9286.8980.8031.8967.32

Stanford S1:

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Conventional54.2361.6744.2755.4192.2061.56
DenseMixer55.8063.1345.3157.1893.0062.88
Temperature=1.0, top_p=0.7

Nemotron-Code:

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Conventional90.8586.5979.0033.4267.59
DenseMixer93.2988.8784.3934.4068.99

Stanford S1:

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Conventional56.5563.6546.1559.1193.0063.69
DenseMixer58.1462.7147.5057.7793.8063.98

⚡ Quick Start

1. Installation

pip install densemixer

2. Setup (One-time)

densemixer setup

3. Enable DenseMixer

export DENSEMIXER_ENABLED=1

4. Use Your MoE Models

from transformers import Qwen3MoeForCausalLM

# DenseMixer automatically patches the model
model = Qwen3MoeForCausalLM.from_pretrained("Qwen/Qwen3-MoE-30B-A3B")

# Train as usual - no code changes needed!

🔧 Configuration

DenseMixer currently supports the following models.

DenseMixer uses environment variables for configuration:

VariableDescriptionDefault
DENSEMIXER_ENABLEDMaster switch (set to 1 to enable)0
DENSEMIXER_QWEN3Enable for Qwen3-MoE models1
DENSEMIXER_QWEN2Enable for Qwen1.5-MoE models1
DENSEMIXER_OLMOEEnable for OLMoE models1

Usage Examples

Enable for all models:

export DENSEMIXER_ENABLED=1
python your_training_script.py

Enable only for specific models:

export DENSEMIXER_ENABLED=1
export DENSEMIXER_QWEN3=1
export DENSEMIXER_QWEN2=0
export DENSEMIXER_OLMOE=0
python your_training_script.py

Disable (default behavior):

# No environment variables needed
python your_training_script.py

📊 Logging

DenseMixer provides intelligent logging to track when custom forward methods are used:

INFO - densemixer - DenseMixer: Using custom forward method for Qwen3-MoE
INFO - densemixer - DenseMixer: Using custom forward method for OLMoE

You can also customize the logging as below.

import logging

# Set logging level
logging.getLogger("densemixer").setLevel(logging.INFO)

# Or disable logging entirely
logging.getLogger("densemixer").setLevel(logging.WARNING)

⚡ Efficiency Analysis

FLOPs: 1.46x overhead vs conventional training (theoretical analysis on Qwen3-30B-A3B)

📊 Detailed FLOPs Analysis
Model Training Cost Analysis Results --- Conventional Training for Qwen3-30B-A3B ---
Number of parameters: 30,431,444,992
Number of Forward TFLOPs per layer: 16.85
Number of Backward TFLOPs per layer: 33.70
Number of TFLOPs per layer: 50.54
Peak memory cost: 157.93 GBs


Model Training Cost Analysis Results --- DenseMixer Training for Qwen3-30B-A3B ---
Number of parameters: 30,431,444,992
Number of Forward TFLOPs per layer: 40.04
Number of Backward TFLOPs per layer: 33.70 # we assume DenseMixer doesn't change backward significantly
Number of TFLOPs per layer: 73.74
Peak memory cost: 164.96 GBs

FLOPs: DenseMixer / Conventional = 1.46x

Detailed FLOPs analysis available in efficiency_analysis/flops_compute.py

Memory: Negligible overhead - model weights are already loaded on GPU

Time: Negligible when training with small scale of data

Detailed FLOPs analysis available in efficiency_analysis/flops_compute.py

ModelDatasetConventionalDenseMixerOverhead
Qwen1.5-MoEIntent (7K)22 min24 min+9%
Qwen3-MoES1 (1K)2.8h3.6h+29%

🚧 Roadmap & Future Improvements

While DenseMixer already delivers significant performance gains, we're working on several improvements to make it more efficient and production-ready:

  • Optimize backward FLOPs: Currently backward FLOPs increase alongside forward FLOPs, but this is unnecessary. We'll optimize the backward pass to reduce training overhead while maintaining performance gains
  • Integrate modern MoE kernels: Replace transformers' native for-loop MoE implementation with efficient kernels (e.g., grouped_gemm) for better large-scale training performance
  • Industry integration: Extend support beyond open-instruct/llama-factory to MoE-optimized frameworks like Megatron for seamless production adoption

📚 Citation

If you find our work useful, please cite us:

@misc{yao2025densemixer,
  title = {DenseMixer: Improving MoE Post-Training with Precise Router Gradient},
  url = {https://fengyao.notion.site/moe-posttraining},
  author = {Yao, Feng and Cui, Junxia and Zhang, Ruohan and Liu, Liyuan and Hao, Shibo and Zhang, Li and Dong, Chengyu and Wang, Shuohang and Shen, Yelong and Gao, Jianfeng and Shang, Jingbo},
  journal = {Feng Yao's Notion},
  year = {2025},
  month = jun
}

Questions?

If you have any questions related to the code or the blog, feel free to reach out to us at fengyao@ucsd.edu.

yaof20/DenseMixer

Official implementation for DenseMixer: Improving MoE Post-Training with Precise Router Gradient

Python

68

36 commits

updated Aug 3, 2025

See the code

README

🎨 DenseMixer 🎨

Improving MoE Post-Training with Precise Router Gradients (Blog)

What is DenseMixer? • Key Features • Experiments • Quick Start • Efficiency • Citation

DenseMixer is a novel MoE post-training technique that empowers MoE training with more precise router gradient estimation, consistently outperforming conventional MoE training in downstream tasks.

What is DenseMixer?

DenseMixer addresses the non-differentiable Top-K routing problem in MoE training via straight-through estimator (STE). This enables more precise router gradients by computing all experts's output during forward pass for better gradient estimation during backward pass. For more technical details, please refer to our blog.

🚀 Key Features

  • Plug-and-play: Zero code changes required
  • Universal compatibility: Works with any MoE using Top-K routing
  • Performance gains: Consistently outperforms conventional MoE training
  • Parameter-efficient: Compatible with LoRA and other PEFT methods
  • No inference overhead: Zero impact on model inference speed

📈 Experiments

DenseMixer consistently outperforms conventional MoE training across:

  • Model scales: 7B, 14B, 30B parameters
  • Architectures: With/without shared experts
  • Training methods: From scratch and up-cycling
  • Data types: Instruction tuning and long reasoning data

DenseMixer Performance Gains

Reproducible Experiments: For detailed training scripts, configurations, and evaluation code, please refer to the experiments folder.

📊 Qwen1.5-MoE-A2.7B (14B): +2.2% average improvement across 7 tasks

Full Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Base Model38.6938.8432.3116.8318.2028.2916.5327.10
Frozen Router53.3735.2037.1082.2033.0138.2932.7544.56
Conventional53.4234.6036.4381.8029.2537.8033.0243.76
DenseMixer55.1635.4039.6883.4033.8340.5633.9045.99
Gain+1.74+0.80+3.25+1.60+4.58+2.76+0.88+2.23

LoRA Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Frozen Router -lora46.7731.4036.5871.0030.3030.1928.0839.19
Conventional -lora43.8934.0038.4164.8028.8037.9926.1439.15
DenseMixer -lora47.2435.4038.4171.8031.8040.2029.2542.01
Gain+3.35+1.40+0.00+7.00+3.00+2.21+3.11+2.86
📊 OLMoE-1B-7B: +2.9% average improvement across 7 tasks

Full Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Base Model15.8519.8010.970.205.707.4011.0910.14
Frozen Router44.8817.87.2372.8022.5036.0528.2932.79
Conventional45.9423.418.9274.6022.3535.9926.8935.44
DenseMixer49.0025.1220.7377.4023.0240.6432.5538.35
Gain+3.06+1.72+1.81+2.80+0.67+4.65+5.66+2.91

LoRA Fine-tuning Results:

MethodGSMMBPPHumanEvalIntentLawSummaryTranslationAvg
Frozen Router -lora45.0324.217.0755.8021.3037.7028.1932.76
Conventional -lora44.5824.215.8560.2021.6037.3026.2232.85
DenseMixer -lora45.3826.216.4866.6024.7040.8029.4335.66
Gain+0.80+2.00+0.63+6.40+3.10+3.50+3.21+2.81
📊 Qwen3-30B-A3B: +3.7% improvement on GPQA-Diamond

Nemotron-Code Dataset (35K samples):

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Base Model65.2460.0653.6016.8548.94
Conventional92.2386.8980.8032.2667.21
DenseMixer93.5989.0282.0034.3168.80
Gain+1.36+2.13+1.20+2.05+1.59

Stanford S1 Dataset (1K samples):

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Base Model38.8820.637.7134.8172.8034.97
Conventional54.8061.5645.6357.3393.4062.54
DenseMixer58.5263.8545.8358.5193.6064.06
Gain+3.72+2.29+0.20+1.18+0.20+1.52

Results shown for temperature=0.6, top_p=0.95 decoding parameters

Additional Decoding Parameters:

Temperature=0.7, top_p=0.8

Nemotron-Code:

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Conventional91.0185.3776.8029.3965.59
DenseMixer91.9286.8980.8031.8967.32

Stanford S1:

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Conventional54.2361.6744.2755.4192.2061.56
DenseMixer55.8063.1345.3157.1893.0062.88
Temperature=1.0, top_p=0.7

Nemotron-Code:

MethodHumanEval (avg@4)HumanEval+ (avg@4)MBPP (avg@1)LiveCodeBench (avg@4)Avg
Conventional90.8586.5979.0033.4267.59
DenseMixer93.2988.8784.3934.4068.99

Stanford S1:

MethodGPQA Diamond (avg@8)AIME 2024 (avg@32)AIME 2025 (avg@32)Olympiad Bench (avg@1)MATH-500 (avg@1)Avg
Conventional56.5563.6546.1559.1193.0063.69
DenseMixer58.1462.7147.5057.7793.8063.98

⚡ Quick Start

1. Installation

pip install densemixer

2. Setup (One-time)

densemixer setup

3. Enable DenseMixer

export DENSEMIXER_ENABLED=1

4. Use Your MoE Models

from transformers import Qwen3MoeForCausalLM

# DenseMixer automatically patches the model
model = Qwen3MoeForCausalLM.from_pretrained("Qwen/Qwen3-MoE-30B-A3B")

# Train as usual - no code changes needed!

🔧 Configuration

DenseMixer currently supports the following models.

DenseMixer uses environment variables for configuration:

VariableDescriptionDefault
DENSEMIXER_ENABLEDMaster switch (set to 1 to enable)0
DENSEMIXER_QWEN3Enable for Qwen3-MoE models1
DENSEMIXER_QWEN2Enable for Qwen1.5-MoE models1
DENSEMIXER_OLMOEEnable for OLMoE models1

Usage Examples

Enable for all models:

export DENSEMIXER_ENABLED=1
python your_training_script.py

Enable only for specific models:

export DENSEMIXER_ENABLED=1
export DENSEMIXER_QWEN3=1
export DENSEMIXER_QWEN2=0
export DENSEMIXER_OLMOE=0
python your_training_script.py

Disable (default behavior):

# No environment variables needed
python your_training_script.py

📊 Logging

DenseMixer provides intelligent logging to track when custom forward methods are used:

INFO - densemixer - DenseMixer: Using custom forward method for Qwen3-MoE
INFO - densemixer - DenseMixer: Using custom forward method for OLMoE

You can also customize the logging as below.

import logging

# Set logging level
logging.getLogger("densemixer").setLevel(logging.INFO)

# Or disable logging entirely
logging.getLogger("densemixer").setLevel(logging.WARNING)

⚡ Efficiency Analysis

FLOPs: 1.46x overhead vs conventional training (theoretical analysis on Qwen3-30B-A3B)

📊 Detailed FLOPs Analysis
Model Training Cost Analysis Results --- Conventional Training for Qwen3-30B-A3B ---
Number of parameters: 30,431,444,992
Number of Forward TFLOPs per layer: 16.85
Number of Backward TFLOPs per layer: 33.70
Number of TFLOPs per layer: 50.54
Peak memory cost: 157.93 GBs


Model Training Cost Analysis Results --- DenseMixer Training for Qwen3-30B-A3B ---
Number of parameters: 30,431,444,992
Number of Forward TFLOPs per layer: 40.04
Number of Backward TFLOPs per layer: 33.70 # we assume DenseMixer doesn't change backward significantly
Number of TFLOPs per layer: 73.74
Peak memory cost: 164.96 GBs

FLOPs: DenseMixer / Conventional = 1.46x

Detailed FLOPs analysis available in efficiency_analysis/flops_compute.py

Memory: Negligible overhead - model weights are already loaded on GPU

Time: Negligible when training with small scale of data

Detailed FLOPs analysis available in efficiency_analysis/flops_compute.py

ModelDatasetConventionalDenseMixerOverhead
Qwen1.5-MoEIntent (7K)22 min24 min+9%
Qwen3-MoES1 (1K)2.8h3.6h+29%

🚧 Roadmap & Future Improvements

While DenseMixer already delivers significant performance gains, we're working on several improvements to make it more efficient and production-ready:

  • Optimize backward FLOPs: Currently backward FLOPs increase alongside forward FLOPs, but this is unnecessary. We'll optimize the backward pass to reduce training overhead while maintaining performance gains
  • Integrate modern MoE kernels: Replace transformers' native for-loop MoE implementation with efficient kernels (e.g., grouped_gemm) for better large-scale training performance
  • Industry integration: Extend support beyond open-instruct/llama-factory to MoE-optimized frameworks like Megatron for seamless production adoption

📚 Citation

If you find our work useful, please cite us:

@misc{yao2025densemixer,
  title = {DenseMixer: Improving MoE Post-Training with Precise Router Gradient},
  url = {https://fengyao.notion.site/moe-posttraining},
  author = {Yao, Feng and Cui, Junxia and Zhang, Ruohan and Liu, Liyuan and Hao, Shibo and Zhang, Li and Dong, Chengyu and Wang, Shuohang and Shen, Yelong and Gao, Jianfeng and Shang, Jingbo},
  journal = {Feng Yao's Notion},
  year = {2025},
  month = jun
}

Questions?

If you have any questions related to the code or the blog, feel free to reach out to us at fengyao@ucsd.edu.