AngelSlim/Qwen3-8B-DFly-Block8

Model

A more accessible, comprehensive, and efficient toolkit for large model compression.

1

8 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

AngelSlim

A more accessible, comprehensive, and efficient toolkit for large model compression.

✒️ TechnicalReport   |    📖 Documentation   |   🤗 Hugging Face   |   🤖 ModelScope

💬 WeChat |   🫨 Discord

Qwen3-8B No-Thinking Drafters: MTP and DFly

Model Summary

This repository contains speculative-decoding drafter checkpoints for Qwen3-8B, including an MTP drafter and a DFly drafter. The drafters propose candidate tokens for the target model to verify, improving inference throughput while preserving the target model's output distribution.

The models were trained with AngelSpec and are intended to be deployed in vLLM together with the Qwen3-8B target model.

  • Target model: Qwen3-8B
  • Inference mode: No-thinking
  • Drafter types: MTP and DFly
  • Training framework: AngelSpec
  • Recommended serving backend: vLLM speculative decoding
  • MTP speculative tokens: 3
  • DFly speculative tokens: 8

Training Data

For Qwen3-8B, prompts are taken from Open-PerfectBlend (Xu et al., 2024). Responses are regenerated with the Qwen3-8B target model in non-thinking mode, and the resulting target-generated responses are used to train the drafter.

Regenerating responses with the same target model helps the drafter better match the target distribution, which is critical for speculative decoding acceptance rate and end-to-end acceleration.

Training Hyperparameters

These Qwen3-8B draft models use the following training configuration:

HyperparameterValue
Batch size512
Loss scheduleLK loss cold start for 300 steps, then end-to-end TV loss
Learning rate6.0e-4
Minimum learning rate6.0e-5
LR decay styleCosine
Warmup ratio0.04
Weight decay0.0
Epochs10
Maximum sequence length4096

Evaluation

The table reports acceleration-related metrics across math, code, and chat benchmarks.

Target ModelDrafterMath500GSM8KHumanEvalMBPPLiveCodeBenchMT-BenchAvg.
Qwen3-8BMTP3.533.563.333.223.252.573.24
Qwen3-8BDFlash4.975.544.774.504.463.164.57
Qwen3-8BDSpark5.876.255.565.255.203.775.32
Qwen3-8BDFly6.066.425.605.345.363.675.41

DFly achieves the highest average result among the listed Qwen3-8B drafters in this evaluation, with an average score of 5.41.

Deployment

Set TARGET to the Qwen3-8B target model and DRAFT to the corresponding drafter repository or local checkpoint path. The relevant inference code can be found in PR https://github.com/vllm-project/vllm/pull/50246.

MTP-3

export TARGET="Qwen/Qwen3-8B"
export DRAFT="<path-or-hf-repo-to-this-MTP-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 1 \
  --speculative-config "{"method":"mtp","model":"$DRAFT","num_speculative_tokens":3}"

DFly-7

export TARGET="Qwen/Qwen3-8B"
export DRAFT="<path-or-hf-repo-to-this-DFly-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 1 \
  --speculative-config "{"method":"dspark","model":"$DRAFT","num_speculative_tokens":7}"

Technical Report

For more details, see the technical report:

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
arXiv: https://arxiv.org/abs/2607.25852

pytorch
qwen3

AngelSlim/Qwen3-8B-DFly-Block8

Model

A more accessible, comprehensive, and efficient toolkit for large model compression.

1

8 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

AngelSlim

A more accessible, comprehensive, and efficient toolkit for large model compression.

✒️ TechnicalReport   |    📖 Documentation   |   🤗 Hugging Face   |   🤖 ModelScope

💬 WeChat |   🫨 Discord

Qwen3-8B No-Thinking Drafters: MTP and DFly

Model Summary

This repository contains speculative-decoding drafter checkpoints for Qwen3-8B, including an MTP drafter and a DFly drafter. The drafters propose candidate tokens for the target model to verify, improving inference throughput while preserving the target model's output distribution.

The models were trained with AngelSpec and are intended to be deployed in vLLM together with the Qwen3-8B target model.

  • Target model: Qwen3-8B
  • Inference mode: No-thinking
  • Drafter types: MTP and DFly
  • Training framework: AngelSpec
  • Recommended serving backend: vLLM speculative decoding
  • MTP speculative tokens: 3
  • DFly speculative tokens: 8

Training Data

For Qwen3-8B, prompts are taken from Open-PerfectBlend (Xu et al., 2024). Responses are regenerated with the Qwen3-8B target model in non-thinking mode, and the resulting target-generated responses are used to train the drafter.

Regenerating responses with the same target model helps the drafter better match the target distribution, which is critical for speculative decoding acceptance rate and end-to-end acceleration.

Training Hyperparameters

These Qwen3-8B draft models use the following training configuration:

HyperparameterValue
Batch size512
Loss scheduleLK loss cold start for 300 steps, then end-to-end TV loss
Learning rate6.0e-4
Minimum learning rate6.0e-5
LR decay styleCosine
Warmup ratio0.04
Weight decay0.0
Epochs10
Maximum sequence length4096

Evaluation

The table reports acceleration-related metrics across math, code, and chat benchmarks.

Target ModelDrafterMath500GSM8KHumanEvalMBPPLiveCodeBenchMT-BenchAvg.
Qwen3-8BMTP3.533.563.333.223.252.573.24
Qwen3-8BDFlash4.975.544.774.504.463.164.57
Qwen3-8BDSpark5.876.255.565.255.203.775.32
Qwen3-8BDFly6.066.425.605.345.363.675.41

DFly achieves the highest average result among the listed Qwen3-8B drafters in this evaluation, with an average score of 5.41.

Deployment

Set TARGET to the Qwen3-8B target model and DRAFT to the corresponding drafter repository or local checkpoint path. The relevant inference code can be found in PR https://github.com/vllm-project/vllm/pull/50246.

MTP-3

export TARGET="Qwen/Qwen3-8B"
export DRAFT="<path-or-hf-repo-to-this-MTP-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 1 \
  --speculative-config "{"method":"mtp","model":"$DRAFT","num_speculative_tokens":3}"

DFly-7

export TARGET="Qwen/Qwen3-8B"
export DRAFT="<path-or-hf-repo-to-this-DFly-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 1 \
  --speculative-config "{"method":"dspark","model":"$DRAFT","num_speculative_tokens":7}"

Technical Report

For more details, see the technical report:

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
arXiv: https://arxiv.org/abs/2607.25852

pytorch
qwen3