AngelSlim/Hy3-DFly-Block8

Model

A more accessible, comprehensive, and efficient toolkit for large model compression.

2

8 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

AngelSlim

A more accessible, comprehensive, and efficient toolkit for large model compression.

✒️ TechnicalReport   |    📖 Documentation   |   🤗 Hugging Face   |   🤖 ModelScope

💬 WeChat |   🫨 Discord

Hy3-A21B No-Thinking Drafters: MTP and DFly

Model Summary

This repository contains no-thinking speculative-decoding drafter checkpoints for Hy3-A21B, including an MTP drafter and a DFly drafter. These draft models are trained with AngelSpec and are intended to accelerate Hy3-A21B inference by proposing candidate tokens that are verified by the target model.

They should be deployed together with the corresponding Hy3-A21B target model under the no-thinking inference setting.

  • Target model: Hy3-A21B
  • Inference mode: No-thinking
  • Drafter types: MTP and DFly
  • Training / inference framework: AngelSpec
  • Recommended serving backend: vLLM speculative decoding
  • Recommended MTP speculative tokens: 3
  • Recommended DFly speculative tokens: 8

Training Data

The no-thinking Hy3 drafter training starts from Open-PerfectBlend-style instruction prompts. Target responses are regenerated with Hy3-A21B under the no-thinking setting so that the drafter learns the target-side token distribution used during deployment.

Domain-Specific Data Expansion for DFly

To further strengthen code generation and mathematical reasoning, the DFly training corpus augments Open-PerfectBlend with 700K domain-specific prompts:

  • 500K code prompts from OpenCodeInstruct and OpenCodeReasoning.
    • OpenCodeInstruct provides broad code-instruction coverage.
    • OpenCodeReasoning emphasizes distilled reasoning traces for competitive programming.
    • Together, they improve both general code completion and multi-step program synthesis.
  • 200K math prompts from Big-Math, adding coverage of diverse and challenging reasoning problems.

Before response generation, prompts that share a contiguous sequence of at least 16 tokens with an evaluation example are removed. Responses for the filtered prompts are then regenerated with Hy3-A21B, mixed with Open-PerfectBlend, and used to train the combined corpus for six epochs with the same objective.

This expansion mainly improves math/code acceptance length. The gain is consistent across structured benchmarks, while chat-oriented performance may not improve uniformly because the final checkpoint is specialized toward code and math.

Evaluation: Accepted Length

The following table reports the mean accepted length on Hy3-A21B at temperature 1 under the no-thinking setting. The average is the arithmetic mean over the six benchmarks.

Target ModelDrafterMath500GSM8KHumanEvalMBPPLiveCodeBenchMT-BenchAvg.
Hy3-A21BMTP3.303.303.133.042.842.403.00
Hy3-A21BDFlash4.014.234.364.053.102.383.69
Hy3-A21BDFly5.235.535.525.414.072.964.79

Evaluation: Throughput Benefit

The following results are measured on HY3 295B-A21B with TP=8 at temperature 1 across concurrency levels. Tok/s denotes output-token throughput. Spd. denotes throughput relative to autoregressive decoding (AR). Each cell uses 3 x 120s measurement windows, and Avg. is the arithmetic mean across the six datasets.

Conc.MethodGSM8K Tok/sGSM8K Spd.Math500 Tok/sMath500 Spd.HumanEval Tok/sHumanEval Spd.MBPP Tok/sMBPP Spd.LiveCodeBench Tok/sLiveCodeBench Spd.MT-Bench Tok/sMT-Bench Spd.Avg. Tok/sAvg. Spd.
c4AR287.91.00x293.91.00x279.71.00x294.41.00x284.51.00x290.01.00x288.41.00x
c4MTP-3495.51.72x517.01.76x473.61.69x476.71.62x414.01.46x384.51.33x460.21.60x
c4DFlash-8569.51.98x587.92.00x588.12.10x575.71.96x408.41.44x371.61.28x516.91.79x
c4DFly-8635.42.21x643.92.19x647.22.31x661.52.25x455.21.60x384.01.32x571.21.98x
c8AR426.31.00x435.71.00x400.31.00x433.71.00x413.41.00x421.81.00x421.81.00x
c8MTP-3756.91.78x791.51.82x717.41.79x729.11.68x620.81.50x590.11.40x701.01.66x
c8DFlash-8857.02.01x860.11.97x866.22.16x866.22.00x609.11.47x545.51.29x767.41.82x
c8DFly-8964.82.26x959.52.20x974.12.43x1001.52.31x641.31.55x565.81.34x851.22.02x
c16AR650.71.00x670.61.00x594.81.00x667.91.00x617.21.00x628.11.00x638.21.00x
c16MTP-31146.61.76x1174.01.75x1076.11.81x1103.31.65x893.51.45x876.61.40x1045.01.64x
c16DFlash-81446.12.22x1430.42.13x1433.32.41x1440.42.16x970.51.57x901.01.43x1270.31.99x
c16DFly-81623.82.50x1607.52.40x1595.42.68x1677.82.51x1046.71.70x961.51.53x1418.82.22x
c32AR918.61.00x1000.81.00x850.81.00x1018.71.00x893.21.00x921.11.00x933.91.00x
c32MTP-31923.82.09x1989.81.99x1780.72.09x1863.61.83x1394.21.56x1469.41.60x1736.91.86x
c32DFlash-82261.82.46x2334.32.33x2170.62.55x2339.32.30x1489.91.67x1453.41.58x2008.22.15x
c32DFly-82527.42.75x2608.92.61x2429.12.86x2741.32.69x1602.61.79x1549.41.68x2243.12.40x
c64AR1156.91.00x1381.31.00x1197.61.00x1513.91.00x1229.61.00x1306.91.00x1297.71.00x
c64MTP-32932.02.53x3170.02.29x2639.62.20x3011.31.99x2050.41.67x2350.01.80x2692.22.08x
c64DFlash-82523.42.18x2947.12.13x2655.22.22x2671.81.76x2015.81.64x1815.81.39x2438.21.89x
c64DFly-82827.82.44x3301.82.39x2965.02.48x3130.92.07x2195.01.79x1936.61.48x2726.22.11x

Deployment

Set TARGET to the Hy3-A21B target model path or repository and DRAFT to the corresponding drafter checkpoint. The relevant inference code can be found in PR https://github.com/vllm-project/vllm/pull/50246.

MTP-3

export TARGET="<path-or-hf-repo-to-hy3-a21b-target>"
export DRAFT="<path-or-hf-repo-to-hy3-mtp-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 8 \
  --speculative-config "{"method":"mtp", "model":"$DRAFT", "num_speculative_tokens":3}"

DFly-8

export TARGET="<path-or-hf-repo-to-hy3-a21b-target>"
export DRAFT="<path-or-hf-repo-to-hy3-dfly-no-thinking-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 8 \
  --speculative-config "{"method":"dspark", "model":"$DRAFT","num_speculative_tokens":7}"

Technical Report

For more details, see the technical report:

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
arXiv: https://arxiv.org/abs/2607.25852

pytorch
qwen3

AngelSlim/Hy3-DFly-Block8

Model

A more accessible, comprehensive, and efficient toolkit for large model compression.

2

8 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

AngelSlim

A more accessible, comprehensive, and efficient toolkit for large model compression.

✒️ TechnicalReport   |    📖 Documentation   |   🤗 Hugging Face   |   🤖 ModelScope

💬 WeChat |   🫨 Discord

Hy3-A21B No-Thinking Drafters: MTP and DFly

Model Summary

This repository contains no-thinking speculative-decoding drafter checkpoints for Hy3-A21B, including an MTP drafter and a DFly drafter. These draft models are trained with AngelSpec and are intended to accelerate Hy3-A21B inference by proposing candidate tokens that are verified by the target model.

They should be deployed together with the corresponding Hy3-A21B target model under the no-thinking inference setting.

  • Target model: Hy3-A21B
  • Inference mode: No-thinking
  • Drafter types: MTP and DFly
  • Training / inference framework: AngelSpec
  • Recommended serving backend: vLLM speculative decoding
  • Recommended MTP speculative tokens: 3
  • Recommended DFly speculative tokens: 8

Training Data

The no-thinking Hy3 drafter training starts from Open-PerfectBlend-style instruction prompts. Target responses are regenerated with Hy3-A21B under the no-thinking setting so that the drafter learns the target-side token distribution used during deployment.

Domain-Specific Data Expansion for DFly

To further strengthen code generation and mathematical reasoning, the DFly training corpus augments Open-PerfectBlend with 700K domain-specific prompts:

  • 500K code prompts from OpenCodeInstruct and OpenCodeReasoning.
    • OpenCodeInstruct provides broad code-instruction coverage.
    • OpenCodeReasoning emphasizes distilled reasoning traces for competitive programming.
    • Together, they improve both general code completion and multi-step program synthesis.
  • 200K math prompts from Big-Math, adding coverage of diverse and challenging reasoning problems.

Before response generation, prompts that share a contiguous sequence of at least 16 tokens with an evaluation example are removed. Responses for the filtered prompts are then regenerated with Hy3-A21B, mixed with Open-PerfectBlend, and used to train the combined corpus for six epochs with the same objective.

This expansion mainly improves math/code acceptance length. The gain is consistent across structured benchmarks, while chat-oriented performance may not improve uniformly because the final checkpoint is specialized toward code and math.

Evaluation: Accepted Length

The following table reports the mean accepted length on Hy3-A21B at temperature 1 under the no-thinking setting. The average is the arithmetic mean over the six benchmarks.

Target ModelDrafterMath500GSM8KHumanEvalMBPPLiveCodeBenchMT-BenchAvg.
Hy3-A21BMTP3.303.303.133.042.842.403.00
Hy3-A21BDFlash4.014.234.364.053.102.383.69
Hy3-A21BDFly5.235.535.525.414.072.964.79

Evaluation: Throughput Benefit

The following results are measured on HY3 295B-A21B with TP=8 at temperature 1 across concurrency levels. Tok/s denotes output-token throughput. Spd. denotes throughput relative to autoregressive decoding (AR). Each cell uses 3 x 120s measurement windows, and Avg. is the arithmetic mean across the six datasets.

Conc.MethodGSM8K Tok/sGSM8K Spd.Math500 Tok/sMath500 Spd.HumanEval Tok/sHumanEval Spd.MBPP Tok/sMBPP Spd.LiveCodeBench Tok/sLiveCodeBench Spd.MT-Bench Tok/sMT-Bench Spd.Avg. Tok/sAvg. Spd.
c4AR287.91.00x293.91.00x279.71.00x294.41.00x284.51.00x290.01.00x288.41.00x
c4MTP-3495.51.72x517.01.76x473.61.69x476.71.62x414.01.46x384.51.33x460.21.60x
c4DFlash-8569.51.98x587.92.00x588.12.10x575.71.96x408.41.44x371.61.28x516.91.79x
c4DFly-8635.42.21x643.92.19x647.22.31x661.52.25x455.21.60x384.01.32x571.21.98x
c8AR426.31.00x435.71.00x400.31.00x433.71.00x413.41.00x421.81.00x421.81.00x
c8MTP-3756.91.78x791.51.82x717.41.79x729.11.68x620.81.50x590.11.40x701.01.66x
c8DFlash-8857.02.01x860.11.97x866.22.16x866.22.00x609.11.47x545.51.29x767.41.82x
c8DFly-8964.82.26x959.52.20x974.12.43x1001.52.31x641.31.55x565.81.34x851.22.02x
c16AR650.71.00x670.61.00x594.81.00x667.91.00x617.21.00x628.11.00x638.21.00x
c16MTP-31146.61.76x1174.01.75x1076.11.81x1103.31.65x893.51.45x876.61.40x1045.01.64x
c16DFlash-81446.12.22x1430.42.13x1433.32.41x1440.42.16x970.51.57x901.01.43x1270.31.99x
c16DFly-81623.82.50x1607.52.40x1595.42.68x1677.82.51x1046.71.70x961.51.53x1418.82.22x
c32AR918.61.00x1000.81.00x850.81.00x1018.71.00x893.21.00x921.11.00x933.91.00x
c32MTP-31923.82.09x1989.81.99x1780.72.09x1863.61.83x1394.21.56x1469.41.60x1736.91.86x
c32DFlash-82261.82.46x2334.32.33x2170.62.55x2339.32.30x1489.91.67x1453.41.58x2008.22.15x
c32DFly-82527.42.75x2608.92.61x2429.12.86x2741.32.69x1602.61.79x1549.41.68x2243.12.40x
c64AR1156.91.00x1381.31.00x1197.61.00x1513.91.00x1229.61.00x1306.91.00x1297.71.00x
c64MTP-32932.02.53x3170.02.29x2639.62.20x3011.31.99x2050.41.67x2350.01.80x2692.22.08x
c64DFlash-82523.42.18x2947.12.13x2655.22.22x2671.81.76x2015.81.64x1815.81.39x2438.21.89x
c64DFly-82827.82.44x3301.82.39x2965.02.48x3130.92.07x2195.01.79x1936.61.48x2726.22.11x

Deployment

Set TARGET to the Hy3-A21B target model path or repository and DRAFT to the corresponding drafter checkpoint. The relevant inference code can be found in PR https://github.com/vllm-project/vllm/pull/50246.

MTP-3

export TARGET="<path-or-hf-repo-to-hy3-a21b-target>"
export DRAFT="<path-or-hf-repo-to-hy3-mtp-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 8 \
  --speculative-config "{"method":"mtp", "model":"$DRAFT", "num_speculative_tokens":3}"

DFly-8

export TARGET="<path-or-hf-repo-to-hy3-a21b-target>"
export DRAFT="<path-or-hf-repo-to-hy3-dfly-no-thinking-drafter>"

vllm serve "$TARGET" \
  --tensor-parallel-size 8 \
  --speculative-config "{"method":"dspark", "model":"$DRAFT","num_speculative_tokens":7}"

Technical Report

For more details, see the technical report:

AngelSpec: Towards Real-World High Performance Inference with Speculative Decoding
arXiv: https://arxiv.org/abs/2607.25852

pytorch
qwen3