Agens Volundr 32B Preview — DFlash2 drafter
1
9 commits
2 linked in READMEs
updated Oct 5, 2026
A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.
The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.
Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:
| Output | Speed-up |
|---|---|
| JSON | 3.6× |
| Code | 2.0× |
| Replies with thinking on (code, Cantonese) | 1.6× |
| Cantonese chat | 1.3× |
On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.
Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at
github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the
model card, and run the container with
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True:
--speculative-algorithm DFLASH \
--speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
--speculative-num-draft-tokens 8 \
--speculative-draft-model-quantization unquant \
--mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
--max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4
Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out
--enable-multimodal).
Apache-2.0. The licence files and NOTICE ship with the weights.
Agens Volundr 32B Preview — DFlash2 drafter
1
9 commits
2 linked in READMEs
updated Oct 5, 2026
A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.
The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.
Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:
| Output | Speed-up |
|---|---|
| JSON | 3.6× |
| Code | 2.0× |
| Replies with thinking on (code, Cantonese) | 1.6× |
| Cantonese chat | 1.3× |
On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.
Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at
github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the
model card, and run the container with
-e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True:
--speculative-algorithm DFLASH \
--speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
--speculative-num-draft-tokens 8 \
--speculative-draft-model-quantization unquant \
--mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
--max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4
Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out
--enable-multimodal).
Apache-2.0. The licence files and NOTICE ship with the weights.