Blockway/Agens-Volundr-32B-Preview-DFlash2

Model

Agens Volundr 32B Preview — DFlash2 drafter

1

9 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

Agens Volundr 32B Preview — DFlash2 drafter

A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.

The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.

Speed-up

Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:

OutputSpeed-up
JSON3.6×
Code2.0×
Replies with thinking on (code, Cantonese)1.6×
Cantonese chat1.3×

On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.

Use

Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the model card, and run the container with -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True:

    --speculative-algorithm DFLASH \
    --speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --speculative-draft-model-quantization unquant \
    --mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
    --max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4

Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out --enable-multimodal).

Licence

Apache-2.0. The licence files and NOTICE ship with the weights.

agens
blockway
dflash2
draft-model
qwen3
safetensors
speculative-decoding
text-generation
volundr

Blockway/Agens-Volundr-32B-Preview-DFlash2

Model

Agens Volundr 32B Preview — DFlash2 drafter

1

9 commits

2 linked in READMEs

updated Oct 5, 2026

See the code

README

Agens Volundr 32B Preview — DFlash2 drafter

A DFlash2 block drafter for speculative decoding with Agens Volundr 32B Preview. At each step it drafts 8 tokens at once from the target model's hidden states (taps at layers 8, 21, 35, 53 and 67), and the target verifies them in one pass. The target checks every token, so answers keep the target model's quality; with greedy decoding the output can differ from non-speculative decoding only where two tokens are numerically tied.

The drafter has 5 layers (about 1.9B parameters, 3.8 GB in bf16) and shares the target's embeddings and output head, which are not stored here.

Speed-up

Single user on two 48 GB GPUs, bf16, the Blockway sglang build, same server with and without the drafter:

OutputSpeed-up
JSON3.6×
Code2.0×
Replies with thinking on (code, Cantonese)1.6×
Cantonese chat1.3×

On average 2.56 drafted tokens are accepted per step. The drafter is for single-user generation: the drafter server runs 4 requests at a time, so with 8 or more concurrent users plain decoding gives more total throughput.

Use

Requires the Blockway sglang build (ghcr.io/blockwayz/agens-sglang, source at github.com/BlockWayz/agens-sglang). Add these flags to the two-GPU bf16 command from the model card, and run the container with -e PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True:

    --speculative-algorithm DFLASH \
    --speculative-draft-model-path Blockway/Agens-Volundr-32B-Preview-DFlash2 \
    --speculative-num-draft-tokens 8 \
    --speculative-draft-model-quantization unquant \
    --mem-fraction-static 0.90 --context-length 16384 --chunked-prefill-size 2048 \
    --max-mamba-cache-size 4 --max-running-requests 4 --cuda-graph-max-bs 4

Speculative decoding needs --attention-backend flashinfer and serves text requests (leave out --enable-multimodal).

Licence

Apache-2.0. The licence files and NOTICE ship with the weights.

agens
blockway
dflash2
draft-model
qwen3
safetensors
speculative-decoding
text-generation
volundr