This is a block-size-16 Domino draft model for the target model
Qwen/Qwen3.6-27B. It is intended
for speculative decoding with the Domino-enabled SGLang branch.
Install the matching SGLang branch:
uv pip install "git+https://github.com/jianuo-huang/sglang.git@feat/domino-tensor-parallel#subdirectory=python"
Launch the regular b16 configuration:
python -m sglang.launch_server \
--model-path Qwen/Qwen3.6-27B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path huang2020/Qwen3.6-27B-Domino \
--speculative-dflash-block-size 16 \
--speculative-num-draft-tokens 16 \
--tp-size 2 \
--attention-backend flashinfer \
--mamba-scheduler-strategy extra_buffer \
--trust-remote-code
For b8, set --speculative-num-draft-tokens 8 while keeping
--speculative-dflash-block-size 16.
SGLang implementation: PR #32018.
Methods. AR is target-only decoding. MTP-S3/S7/S15 use the built-in Qwen3.6 MTP heads with 3/7/15 steps, 4/8/16 draft tokens, and top-k 1. DFlash uses the official z-lab/Qwen3.6-27B-DFlash checkpoint at revision 0919688.
For DFlash and Domino, b16 is the regular block-16 run and the target verifies all 16 positions. b8 keeps the same block-16 draft backbone but the target verifies only the first 8 positions. The draft backbone itself is not shortened.
Setting. Qwen/Qwen3.6-27B, TP2/BF16 on 2×A100 80GB, FlashInfer, O4096, thinking enabled, greedy sampling (temperature=0, top_p=1, top_k=1), C1/C8/C32, and three fresh-server repeats per cell. Workloads are GSM8K-128, MATH500-128, HumanEval-164, MBPP-128, MT-Bench-80, and Alpaca-128.

Each bar is mean output tok/s over three runs. The black outline marks the fastest speculative configuration for each workload.
Each cell is output tok/s (speedup versus AR). Bold marks the fastest speculative configuration in each row.
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 47.2 (1.00x) | 126.8 (2.68x) | 151.9 (3.22x) | 133.9 (2.84x) | 179.2 (3.79x) | 200.9 (4.25x) | 206.0 (4.36x) | 248.0 (5.25x) |
| MATH500 | 47.3 (1.00x) | 132.3 (2.80x) | 168.1 (3.55x) | 151.6 (3.20x) | 203.2 (4.29x) | 240.1 (5.07x) | 217.7 (4.60x) | 270.6 (5.72x) |
| HumanEval | 47.2 (1.00x) | 125.3 (2.65x) | 149.9 (3.18x) | 129.9 (2.75x) | 188.0 (3.98x) | 211.1 (4.47x) | 198.5 (4.20x) | 235.1 (4.98x) |
| MBPP | 47.6 (1.00x) | 122.6 (2.57x) | 141.8 (2.98x) | 117.0 (2.46x) | 177.5 (3.73x) | 186.2 (3.91x) | 189.0 (3.97x) | 214.0 (4.49x) |
| MT-Bench | 47.1 (1.00x) | 115.1 (2.44x) | 125.0 (2.65x) | 100.7 (2.14x) | 141.5 (3.00x) | 143.8 (3.05x) | 155.7 (3.31x) | 161.9 (3.44x) |
| Alpaca | 47.2 (1.00x) | 112.1 (2.38x) | 119.8 (2.54x) | 96.0 (2.03x) | 135.7 (2.87x) | 133.9 (2.84x) | 150.1 (3.18x) | 157.7 (3.34x) |
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 324.4 (1.00x) | 778.7 (2.40x) | 913.0 (2.81x) | 673.3 (2.08x) | 1016.1 (3.13x) | 899.0 (2.77x) | 1153.2 (3.55x) | 1131.2 (3.49x) |
| MATH500 | 331.6 (1.00x) | 846.3 (2.55x) | 1056.6 (3.19x) | 805.6 (2.43x) | 1191.4 (3.59x) | 1107.4 (3.34x) | 1255.2 (3.78x) | 1252.5 (3.78x) |
| HumanEval | 331.9 (1.00x) | 797.5 (2.40x) | 939.0 (2.83x) | 684.4 (2.06x) | 1104.2 (3.33x) | 985.8 (2.97x) | 1155.9 (3.48x) | 1078.5 (3.25x) |
| MBPP | 326.7 (1.00x) | 756.2 (2.32x) | 872.3 (2.67x) | 564.5 (1.73x) | 995.4 (3.05x) | 850.4 (2.60x) | 1053.7 (3.23x) | 950.0 (2.91x) |
| MT-Bench | 330.6 (1.00x) | 713.7 (2.16x) | 776.3 (2.35x) | 528.6 (1.60x) | 808.0 (2.44x) | 655.2 (1.98x) | 881.5 (2.67x) | 752.8 (2.28x) |
| Alpaca | 332.6 (1.00x) | 719.8 (2.16x) | 740.9 (2.23x) | 510.9 (1.54x) | 786.8 (2.37x) | 616.7 (1.85x) | 862.9 (2.59x) | 722.7 (2.17x) |
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 862.9 (1.00x) | 1555.9 (1.80x) | 1495.6 (1.73x) | 1034.2 (1.20x) | 1601.7 (1.86x) | 1214.0 (1.41x) | 1817.5 (2.11x) | 1574.5 (1.82x) |
| MATH500 | 963.3 (1.00x) | 1834.1 (1.90x) | 1795.0 (1.86x) | 1250.0 (1.30x) | 1920.0 (1.99x) | 1478.7 (1.54x) | 2021.3 (2.10x) | 1716.2 (1.78x) |
| HumanEval | 953.4 (1.00x) | 1701.1 (1.78x) | 1607.3 (1.69x) | 1099.3 (1.15x) | 1795.8 (1.88x) | 1351.4 (1.42x) | 1879.7 (1.97x) | 1541.3 (1.62x) |
| MBPP | 919.5 (1.00x) | 1577.2 (1.72x) | 1443.2 (1.57x) | 948.5 (1.03x) | 1595.0 (1.73x) | 1218.6 (1.33x) | 1547.2 (1.68x) | 1342.4 (1.46x) |
| MT-Bench | 866.8 (1.00x) | 1352.0 (1.56x) | 1144.4 (1.32x) | 721.8 (0.83x) | 1168.1 (1.35x) | 817.3 (0.94x) | 1286.9 (1.48x) | 920.3 (1.06x) |
| Alpaca | 788.1 (1.00x) | 1387.6 (1.76x) | 1130.1 (1.43x) | 718.4 (0.91x) | 1128.9 (1.43x) | 790.0 (1.00x) | 1315.0 (1.67x) | 894.1 (1.13x) |
Arithmetic mean of the per-workload TPS ratios; Overall is the arithmetic mean of all 18 workload-by-concurrency ratios.
| C | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| 1 | 2.59x | 3.02x | 2.57x | 3.61x | 3.93x | 3.94x | 4.54x |
| 8 | 2.33x | 2.68x | 1.90x | 2.98x | 2.59x | 3.22x | 2.98x |
| 32 | 1.75x | 1.60x | 1.07x | 1.71x | 1.27x | 1.84x | 1.48x |
| Overall | 2.22x | 2.43x | 1.85x | 2.77x | 2.60x | 3.00x | 3.00x |
Mean output tokens per target verification step, including the target bonus token. The maxima are 4/8/16 for MTP-S3/S7/S15, 8 for b8, and 16 for b16. Bold marks the highest value within the directly comparable max-8 and max-16 groups.
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.556 | 5.503 | 6.855 | 5.449 | 6.878 | 6.335 | 9.163 |
| MATH500 | 3.608 | 5.658 | 7.055 | 5.700 | 7.401 | 6.201 | 8.760 |
| HumanEval | 3.422 | 5.056 | 6.048 | 5.264 | 6.469 | 5.655 | 7.445 |
| MBPP | 3.349 | 4.782 | 5.439 | 4.972 | 5.732 | 5.405 | 6.803 |
| MT-Bench | 3.206 | 4.487 | 5.188 | 4.278 | 4.947 | 4.758 | 5.851 |
| Alpaca | 3.182 | 4.428 | 5.032 | 4.176 | 4.698 | 4.707 | 5.937 |
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.544 | 5.532 | 6.819 | 5.429 | 6.796 | 6.309 | 9.310 |
| MATH500 | 3.604 | 5.660 | 7.001 | 5.663 | 7.386 | 6.194 | 8.813 |
| HumanEval | 3.429 | 5.055 | 6.024 | 5.267 | 6.440 | 5.659 | 7.418 |
| MBPP | 3.338 | 4.787 | 5.442 | 4.954 | 5.770 | 5.404 | 6.787 |
| MT-Bench | 3.198 | 4.496 | 5.201 | 4.294 | 4.967 | 4.749 | 5.908 |
| Alpaca | 3.203 | 4.415 | 5.022 | 4.146 | 4.669 | 4.743 | 5.835 |
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.554 | 5.510 | 6.889 | 5.442 | 6.873 | 6.288 | 9.313 |
| MATH500 | 3.608 | 5.646 | 7.032 | 5.697 | 7.321 | 6.206 | 8.789 |
| HumanEval | 3.423 | 5.061 | 6.034 | 5.293 | 6.466 | 5.677 | 7.449 |
| MBPP | 3.346 | 4.755 | 5.497 | 4.942 | 5.752 | 5.373 | 6.842 |
| MT-Bench | 3.197 | 4.505 | 5.224 | 4.302 | 4.983 | 4.768 | 5.923 |
| Alpaca | 3.185 | 4.426 | 5.006 | 4.171 | 4.692 | 4.721 | 5.812 |
All 144 displayed method/workload/concurrency cells completed three measured runs.
These are observed serving results. BF16 greedy trajectories can differ across methods, so the TPS differences are not pure kernel-attribution claims.
@article{huang2026domino,
title={Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding},
author={Huang, Jianuo and Zhang, Yaojie and Zhang, Qituan and Lin, Hao and Xu, Hanlin and Zhang, Linfeng},
journal={arXiv preprint arXiv:2605.29707},
year={2026}
}
Domino builds on DFlash and its block-diffusion drafting formulation.
This is a block-size-16 Domino draft model for the target model
Qwen/Qwen3.6-27B. It is intended
for speculative decoding with the Domino-enabled SGLang branch.
Install the matching SGLang branch:
uv pip install "git+https://github.com/jianuo-huang/sglang.git@feat/domino-tensor-parallel#subdirectory=python"
Launch the regular b16 configuration:
python -m sglang.launch_server \
--model-path Qwen/Qwen3.6-27B \
--speculative-algorithm DFLASH \
--speculative-draft-model-path huang2020/Qwen3.6-27B-Domino \
--speculative-dflash-block-size 16 \
--speculative-num-draft-tokens 16 \
--tp-size 2 \
--attention-backend flashinfer \
--mamba-scheduler-strategy extra_buffer \
--trust-remote-code
For b8, set --speculative-num-draft-tokens 8 while keeping
--speculative-dflash-block-size 16.
SGLang implementation: PR #32018.
Methods. AR is target-only decoding. MTP-S3/S7/S15 use the built-in Qwen3.6 MTP heads with 3/7/15 steps, 4/8/16 draft tokens, and top-k 1. DFlash uses the official z-lab/Qwen3.6-27B-DFlash checkpoint at revision 0919688.
For DFlash and Domino, b16 is the regular block-16 run and the target verifies all 16 positions. b8 keeps the same block-16 draft backbone but the target verifies only the first 8 positions. The draft backbone itself is not shortened.
Setting. Qwen/Qwen3.6-27B, TP2/BF16 on 2×A100 80GB, FlashInfer, O4096, thinking enabled, greedy sampling (temperature=0, top_p=1, top_k=1), C1/C8/C32, and three fresh-server repeats per cell. Workloads are GSM8K-128, MATH500-128, HumanEval-164, MBPP-128, MT-Bench-80, and Alpaca-128.

Each bar is mean output tok/s over three runs. The black outline marks the fastest speculative configuration for each workload.
Each cell is output tok/s (speedup versus AR). Bold marks the fastest speculative configuration in each row.
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 47.2 (1.00x) | 126.8 (2.68x) | 151.9 (3.22x) | 133.9 (2.84x) | 179.2 (3.79x) | 200.9 (4.25x) | 206.0 (4.36x) | 248.0 (5.25x) |
| MATH500 | 47.3 (1.00x) | 132.3 (2.80x) | 168.1 (3.55x) | 151.6 (3.20x) | 203.2 (4.29x) | 240.1 (5.07x) | 217.7 (4.60x) | 270.6 (5.72x) |
| HumanEval | 47.2 (1.00x) | 125.3 (2.65x) | 149.9 (3.18x) | 129.9 (2.75x) | 188.0 (3.98x) | 211.1 (4.47x) | 198.5 (4.20x) | 235.1 (4.98x) |
| MBPP | 47.6 (1.00x) | 122.6 (2.57x) | 141.8 (2.98x) | 117.0 (2.46x) | 177.5 (3.73x) | 186.2 (3.91x) | 189.0 (3.97x) | 214.0 (4.49x) |
| MT-Bench | 47.1 (1.00x) | 115.1 (2.44x) | 125.0 (2.65x) | 100.7 (2.14x) | 141.5 (3.00x) | 143.8 (3.05x) | 155.7 (3.31x) | 161.9 (3.44x) |
| Alpaca | 47.2 (1.00x) | 112.1 (2.38x) | 119.8 (2.54x) | 96.0 (2.03x) | 135.7 (2.87x) | 133.9 (2.84x) | 150.1 (3.18x) | 157.7 (3.34x) |
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 324.4 (1.00x) | 778.7 (2.40x) | 913.0 (2.81x) | 673.3 (2.08x) | 1016.1 (3.13x) | 899.0 (2.77x) | 1153.2 (3.55x) | 1131.2 (3.49x) |
| MATH500 | 331.6 (1.00x) | 846.3 (2.55x) | 1056.6 (3.19x) | 805.6 (2.43x) | 1191.4 (3.59x) | 1107.4 (3.34x) | 1255.2 (3.78x) | 1252.5 (3.78x) |
| HumanEval | 331.9 (1.00x) | 797.5 (2.40x) | 939.0 (2.83x) | 684.4 (2.06x) | 1104.2 (3.33x) | 985.8 (2.97x) | 1155.9 (3.48x) | 1078.5 (3.25x) |
| MBPP | 326.7 (1.00x) | 756.2 (2.32x) | 872.3 (2.67x) | 564.5 (1.73x) | 995.4 (3.05x) | 850.4 (2.60x) | 1053.7 (3.23x) | 950.0 (2.91x) |
| MT-Bench | 330.6 (1.00x) | 713.7 (2.16x) | 776.3 (2.35x) | 528.6 (1.60x) | 808.0 (2.44x) | 655.2 (1.98x) | 881.5 (2.67x) | 752.8 (2.28x) |
| Alpaca | 332.6 (1.00x) | 719.8 (2.16x) | 740.9 (2.23x) | 510.9 (1.54x) | 786.8 (2.37x) | 616.7 (1.85x) | 862.9 (2.59x) | 722.7 (2.17x) |
| Workload | AR | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|---|
| GSM8K | 862.9 (1.00x) | 1555.9 (1.80x) | 1495.6 (1.73x) | 1034.2 (1.20x) | 1601.7 (1.86x) | 1214.0 (1.41x) | 1817.5 (2.11x) | 1574.5 (1.82x) |
| MATH500 | 963.3 (1.00x) | 1834.1 (1.90x) | 1795.0 (1.86x) | 1250.0 (1.30x) | 1920.0 (1.99x) | 1478.7 (1.54x) | 2021.3 (2.10x) | 1716.2 (1.78x) |
| HumanEval | 953.4 (1.00x) | 1701.1 (1.78x) | 1607.3 (1.69x) | 1099.3 (1.15x) | 1795.8 (1.88x) | 1351.4 (1.42x) | 1879.7 (1.97x) | 1541.3 (1.62x) |
| MBPP | 919.5 (1.00x) | 1577.2 (1.72x) | 1443.2 (1.57x) | 948.5 (1.03x) | 1595.0 (1.73x) | 1218.6 (1.33x) | 1547.2 (1.68x) | 1342.4 (1.46x) |
| MT-Bench | 866.8 (1.00x) | 1352.0 (1.56x) | 1144.4 (1.32x) | 721.8 (0.83x) | 1168.1 (1.35x) | 817.3 (0.94x) | 1286.9 (1.48x) | 920.3 (1.06x) |
| Alpaca | 788.1 (1.00x) | 1387.6 (1.76x) | 1130.1 (1.43x) | 718.4 (0.91x) | 1128.9 (1.43x) | 790.0 (1.00x) | 1315.0 (1.67x) | 894.1 (1.13x) |
Arithmetic mean of the per-workload TPS ratios; Overall is the arithmetic mean of all 18 workload-by-concurrency ratios.
| C | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| 1 | 2.59x | 3.02x | 2.57x | 3.61x | 3.93x | 3.94x | 4.54x |
| 8 | 2.33x | 2.68x | 1.90x | 2.98x | 2.59x | 3.22x | 2.98x |
| 32 | 1.75x | 1.60x | 1.07x | 1.71x | 1.27x | 1.84x | 1.48x |
| Overall | 2.22x | 2.43x | 1.85x | 2.77x | 2.60x | 3.00x | 3.00x |
Mean output tokens per target verification step, including the target bonus token. The maxima are 4/8/16 for MTP-S3/S7/S15, 8 for b8, and 16 for b16. Bold marks the highest value within the directly comparable max-8 and max-16 groups.
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.556 | 5.503 | 6.855 | 5.449 | 6.878 | 6.335 | 9.163 |
| MATH500 | 3.608 | 5.658 | 7.055 | 5.700 | 7.401 | 6.201 | 8.760 |
| HumanEval | 3.422 | 5.056 | 6.048 | 5.264 | 6.469 | 5.655 | 7.445 |
| MBPP | 3.349 | 4.782 | 5.439 | 4.972 | 5.732 | 5.405 | 6.803 |
| MT-Bench | 3.206 | 4.487 | 5.188 | 4.278 | 4.947 | 4.758 | 5.851 |
| Alpaca | 3.182 | 4.428 | 5.032 | 4.176 | 4.698 | 4.707 | 5.937 |
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.544 | 5.532 | 6.819 | 5.429 | 6.796 | 6.309 | 9.310 |
| MATH500 | 3.604 | 5.660 | 7.001 | 5.663 | 7.386 | 6.194 | 8.813 |
| HumanEval | 3.429 | 5.055 | 6.024 | 5.267 | 6.440 | 5.659 | 7.418 |
| MBPP | 3.338 | 4.787 | 5.442 | 4.954 | 5.770 | 5.404 | 6.787 |
| MT-Bench | 3.198 | 4.496 | 5.201 | 4.294 | 4.967 | 4.749 | 5.908 |
| Alpaca | 3.203 | 4.415 | 5.022 | 4.146 | 4.669 | 4.743 | 5.835 |
| Workload | MTP-S3 | MTP-S7 | MTP-S15 | DFlash b8 | DFlash b16 | Domino b8 | Domino b16 |
|---|---|---|---|---|---|---|---|
| GSM8K | 3.554 | 5.510 | 6.889 | 5.442 | 6.873 | 6.288 | 9.313 |
| MATH500 | 3.608 | 5.646 | 7.032 | 5.697 | 7.321 | 6.206 | 8.789 |
| HumanEval | 3.423 | 5.061 | 6.034 | 5.293 | 6.466 | 5.677 | 7.449 |
| MBPP | 3.346 | 4.755 | 5.497 | 4.942 | 5.752 | 5.373 | 6.842 |
| MT-Bench | 3.197 | 4.505 | 5.224 | 4.302 | 4.983 | 4.768 | 5.923 |
| Alpaca | 3.185 | 4.426 | 5.006 | 4.171 | 4.692 | 4.721 | 5.812 |
All 144 displayed method/workload/concurrency cells completed three measured runs.
These are observed serving results. BF16 greedy trajectories can differ across methods, so the TPS differences are not pure kernel-attribution claims.
@article{huang2026domino,
title={Domino: Decoupling Causal Modeling from Autoregressive Drafting in Speculative Decoding},
author={Huang, Jianuo and Zhang, Yaojie and Zhang, Qituan and Lin, Hao and Xu, Hanlin and Zhang, Linfeng},
journal={arXiv preprint arXiv:2605.29707},
year={2026}
}
Domino builds on DFlash and its block-diffusion drafting formulation.