Blockway/Agens-Volundr-32B-Preview

Model

Agens Volundr 32B — Preview

3

33 commits

2 linked in READMEs

updated Oct 4, 2026

See the code

README

Agens Volundr 32B — Preview

Agens Volundr 32B is an open-weights model from Blockway, built for reasoning, coding and agent work. This is a Preview release: the architecture, format and post-training pipeline are final; continued pre-training is still running and the full v1 will replace this checkpoint. In this Preview, long agentic sessions are the weak spot (see Known limitations).

It runs on Blockway's own mixer stack: 72 layers — 54 of KDA (linear attention with channel-wise gating: fixed-size state, so a long session costs no extra memory per turn), 17 of BCSA (Blockway Compressed-Sparse Attention: a 4,096-token dense window plus a compressed far field selected by a learned indexer) and one dense attention layer — with Engram (n-gram conditional memory) and mHC (manifold hyper-connections, four residual streams). About 32 billion parameters.

Volundr speaks English, Chinese (Simplified and Traditional) and Cantonese; handles text and images; calls tools; and ships with Blockway's control tokens and chat template (see Format and serving).

Results (same harness for every row)

Benchmark comparison

BenchmarkAgens Volundr 32B PreviewQwen3.8-27BAgens Pilot
Coding
Competitive coding — LiveCodeBench v6 (2025-02 to 2025-05), mean of 2 runs63.4 ▲ +4.259.258.8
Code completion — HumanEval81.7 ▲ +4.377.480.5
Agentic coding — SWE-bench Verified, 50 tasks (Codeway)44.058.064.0
General
Competition maths — AIME 2025, avg@4, mean of 2 runs74.6 ▲ +2.971.766.7
Maths — MATH-50098.2 ▲ +1.696.697.2
Chinese knowledge — CMMLU86.5 ▲ +0.186.486.1
Instruction following — IFEval (prompt strict)87.688.485.6
Instruction following — IFBench (prompt strict)66.369.363.0
Scientific reasoning — GPQA Diamond81.783.881.8
Knowledge — MMLU-Pro79.980.480.3
Agent · focus of the full v1
Multi-turn tool use — τ²-bench (airline, retail, telecom)74.279.280.0
Parallel function calls — BFCL v4 parallel92.094.092.0
Knowing when not to call — BFCL v4 irrelevance80.881.785.8

▲ Volundr ahead of Qwen3.8-27B (difference in points). Rows without ▲ are where this Preview trails; closing them is the focus of the full v1. Thinking on, temperature 0.6 and the same output-token limit for every model unless noted; answers cut off by the limit count as wrong. HumanEval: greedy code completion. GPQA Diamond: 6K-token thinking budget, then the answer is forced, for every model (Volundr: mean of 4 runs). τ²-bench: agent at temperature 0, Qwen3.8-27B as the user simulator for every model; the 13 retail tasks that need an LLM judge are scored as failed for every model (no LLM judges are used anywhere). BFCL v4: non-live categories, official harness at temperature 0.001. AIME 2025 and LiveCodeBench: mean of 2 runs. SWE-bench: 50-task subset of SWE-bench Verified ("Verified mini") run through Codeway, Blockway's coding-agent harness, with each model's default sampling and identical limits. MMLU-Pro and CMMLU: fixed stratified subsets (1,400 and 2,010 questions). Qwen3.8-27B and Agens Pilot were run by us on the same harness; these are not their publishers' figures.

Reading the table: Volundr is ahead of Qwen3.8-27B on competitive coding (LiveCodeBench v6, +4.2), code completion (HumanEval, +4.3) and competition maths (AIME 2025, +2.9; MATH-500, +1.6), and level on Chinese knowledge (CMMLU). It trails on agentic work, most clearly on SWE-bench through Codeway, where it often fell into repetition loops in long sessions, and by a few points on instruction following and GPQA. Closing those gaps is the focus of the full v1: it continues pre-training to about 10B tokens and adds training on long agentic sessions.

Format and serving

Volundr's chat format uses Blockway's control tokens: <|agens_start|> / <|agens_end|> for turns, <|think|> … <|/think|> for reasoning, <|call|> … <|/call|> for tool calls and <|result|> … <|/result|> for tool results. Each control token is a single reserved id, so turn boundaries, reasoning and tool calls never depend on how ordinary text tokenises. When no system message is supplied, the template inserts You are Agens, an AI assistant developed by Blockway.

Thinking is on by default (enable_thinking: false turns it off); reasoning_effort accepts low / medium / xhigh. Tool calls use the <function=…><parameter=…> form; the agens parsers in our sglang build expose them as OpenAI-style tool_calls.

Reasoning budget. Our sglang build adds a per-request reasoning_budget (int) that closes the thinking block at N generated tokens, so a hard constraint prompt always ends in an answer. Inference only, weights untouched. The IFBench rows use reasoning_budget: 6000.

Serving. Volundr runs on the Blockway sglang build (model class, agens parsers, reasoning budget). bf16 on two 48 GB GPUs:

docker run --rm --gpus '"device=0,1"' --ipc=host --network host --shm-size 32g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/blockwayz/agens-sglang:preview-sm89 \
  python3 -m sglang.launch_server \
    --model-path Blockway/Agens-Volundr-32B-Preview \
    --tp-size 2 --attention-backend flashinfer --page-size 1 \
    --disable-radix-cache --disable-prefill-cuda-graph --disable-custom-all-reduce \
    --mem-fraction-static 0.87 --context-length 32768 \
    --max-mamba-cache-size 16 --max-running-requests 16 --cuda-graph-max-bs 16 \
    --reasoning-parser agens --tool-call-parser agens \
    --host 127.0.0.1 --port 30000

Images: ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) and :preview-sm90 (H100/H200). The source patch is at github.com/BlockWayz/agens-sglang; its README has the commands for INT4 on one 48 GB GPU, bf16 on one H200 and the drafter, and explains each flag. Add --enable-multimodal --mm-feature-transport cpu to accept images. A vLLM plugin is not yet available.

Speed (measured on the Blockway sglang build, no drafter, one user unless noted; prompt of the given length plus 256 generated tokens):

SetupContextPrefill (tok/s)Decode (tok/s)
bf16, two 48 GB GPUs1K2,12225.1
bf16, two 48 GB GPUs8K2,18024.1
bf16, two 48 GB GPUs32K1,91624.1
bf16, two 48 GB GPUs64K1,67924.0
bf16, two 48 GB GPUs128K1,29723.9
INT4, one 48 GB GPU1K2,23631.0
INT4, one 48 GB GPU8K2,20029.3
INT4, one 48 GB GPU32K1,76429.1

Decode speed stays flat from 1K to 128K context: the 54 KDA layers carry a fixed-size state and the 17 BCSA layers read a bounded window plus 512 selected blocks, so the per-token cost barely grows with the conversation. With concurrent users (1K-token prompts, 256 new tokens each), total throughput is 127 tok/s at 8 users and 130 tok/s at 16 on two 48 GB GPUs (bf16), and 117 tok/s at 8 users on one 48 GB GPU (INT4, with 9 request slots; see the serving README).

Speculative decoding. A DFlash2 drafter for the Blockway sglang build is published as Blockway/Agens-Volundr-32B-Preview-DFlash2. Single user on two 48 GB GPUs, it makes JSON output up to 3.6× faster, code 2.0×, and replies with thinking on about 1.6×.

Architecture

Three quarters of the stack keep a fixed-size state instead of a growing KV cache, so long sessions stay cheap in memory and in decode time.

ComponentLayersHow it worksWhat it buys
KDA (Kimi Delta Attention)54Linear attention: a gated delta-rule update of a fixed-size state per head, with per-channel decayMemory and per-token decode cost independent of context length; no KV cache in these layers
BCSA (Blockway Compressed-Sparse Attention)17Exact attention over a 4,096-token sliding window, plus the far field pooled into 4-token blocks of which a learned indexer selects the top 512; both halves share one softmaxAttention cost per token bounded by the window plus 512 blocks, while keeping precise local detail and long-range recall
Dense attention1Standard full attentionOne exact global read of the whole context
Engram2Hashed 2-gram and 3-gram lookup into a 1M-row × 512 memory table, gated into the residual stream at layers 2 and 22Extra parametric memory read by a table lookup instead of matrix multiplies, at almost no compute cost
mHCallFour residual streams mixed by learned matrices kept doubly stochastic (Sinkhorn projection)Stable signal propagation through a 72-layer stack

What this means in practice: only 18 of 72 layers keep a KV cache. On two 48 GB GPUs (bf16) the server holds a 220K-token cache pool; the INT4 build (31.7 GiB) runs on a single 48 GB GPU.

Training

The Preview has had about 1.6B tokens of continued pre-training: repository-level code (each repository packed in dependency order, tests next to their sources), code and maths corpora, web text in English, Chinese and Cantonese, and GitHub issues, with chat-format replay under logit distillation so reasoning and turn-taking stay intact. Post-training covered identity, tool use (parallel calls, and declining when no tool fits) and preference training on pairs mined from the model's own samples.

Teacher outputs in the training data come only from open-weight models under permissive licences (Apache-2.0, MIT). There are no outputs from Anthropic, OpenAI or Google models.

Known limitations (Preview)

  • Long agentic coding sessions: in our SWE-bench runs through Codeway, Volundr fell into repetition loops in 36 of 50 tasks (output repeating a line, or tool calls written as plain text), against none for Qwen3.8-27B and Agens Pilot. Prefer short sessions or a supervising harness until v1.
  • Retrieval at the full 262K window has not been re-validated on this checkpoint.
  • Cantonese is a supported language, not a specialty.
  • The drafter is for single-user generation: with 8 or more concurrent requests, plain decoding gives more total throughput.

Acknowledgements

The mixer stack builds on published research: Kimi Delta Attention (Moonshot AI), and DeepSeek's sparse attention, Engram and manifold hyper-connections. The combination, the BCSA design and the training programme are Blockway's.

Licence

Apache-2.0. The licence files and MODIFICATIONS.md ship with the weights.

agens
agent
blockway
cantonese
code
conversational
custom_code
image-text-to-text
linear-attention
long-context
safetensors
sparse-attention
text-generation
tool-use
transformers
volundr

Blockway/Agens-Volundr-32B-Preview

Model

Agens Volundr 32B — Preview

3

33 commits

2 linked in READMEs

updated Oct 4, 2026

See the code

README

Agens Volundr 32B — Preview

Agens Volundr 32B is an open-weights model from Blockway, built for reasoning, coding and agent work. This is a Preview release: the architecture, format and post-training pipeline are final; continued pre-training is still running and the full v1 will replace this checkpoint. In this Preview, long agentic sessions are the weak spot (see Known limitations).

It runs on Blockway's own mixer stack: 72 layers — 54 of KDA (linear attention with channel-wise gating: fixed-size state, so a long session costs no extra memory per turn), 17 of BCSA (Blockway Compressed-Sparse Attention: a 4,096-token dense window plus a compressed far field selected by a learned indexer) and one dense attention layer — with Engram (n-gram conditional memory) and mHC (manifold hyper-connections, four residual streams). About 32 billion parameters.

Volundr speaks English, Chinese (Simplified and Traditional) and Cantonese; handles text and images; calls tools; and ships with Blockway's control tokens and chat template (see Format and serving).

Results (same harness for every row)

Benchmark comparison

BenchmarkAgens Volundr 32B PreviewQwen3.8-27BAgens Pilot
Coding
Competitive coding — LiveCodeBench v6 (2025-02 to 2025-05), mean of 2 runs63.4 ▲ +4.259.258.8
Code completion — HumanEval81.7 ▲ +4.377.480.5
Agentic coding — SWE-bench Verified, 50 tasks (Codeway)44.058.064.0
General
Competition maths — AIME 2025, avg@4, mean of 2 runs74.6 ▲ +2.971.766.7
Maths — MATH-50098.2 ▲ +1.696.697.2
Chinese knowledge — CMMLU86.5 ▲ +0.186.486.1
Instruction following — IFEval (prompt strict)87.688.485.6
Instruction following — IFBench (prompt strict)66.369.363.0
Scientific reasoning — GPQA Diamond81.783.881.8
Knowledge — MMLU-Pro79.980.480.3
Agent · focus of the full v1
Multi-turn tool use — τ²-bench (airline, retail, telecom)74.279.280.0
Parallel function calls — BFCL v4 parallel92.094.092.0
Knowing when not to call — BFCL v4 irrelevance80.881.785.8

▲ Volundr ahead of Qwen3.8-27B (difference in points). Rows without ▲ are where this Preview trails; closing them is the focus of the full v1. Thinking on, temperature 0.6 and the same output-token limit for every model unless noted; answers cut off by the limit count as wrong. HumanEval: greedy code completion. GPQA Diamond: 6K-token thinking budget, then the answer is forced, for every model (Volundr: mean of 4 runs). τ²-bench: agent at temperature 0, Qwen3.8-27B as the user simulator for every model; the 13 retail tasks that need an LLM judge are scored as failed for every model (no LLM judges are used anywhere). BFCL v4: non-live categories, official harness at temperature 0.001. AIME 2025 and LiveCodeBench: mean of 2 runs. SWE-bench: 50-task subset of SWE-bench Verified ("Verified mini") run through Codeway, Blockway's coding-agent harness, with each model's default sampling and identical limits. MMLU-Pro and CMMLU: fixed stratified subsets (1,400 and 2,010 questions). Qwen3.8-27B and Agens Pilot were run by us on the same harness; these are not their publishers' figures.

Reading the table: Volundr is ahead of Qwen3.8-27B on competitive coding (LiveCodeBench v6, +4.2), code completion (HumanEval, +4.3) and competition maths (AIME 2025, +2.9; MATH-500, +1.6), and level on Chinese knowledge (CMMLU). It trails on agentic work, most clearly on SWE-bench through Codeway, where it often fell into repetition loops in long sessions, and by a few points on instruction following and GPQA. Closing those gaps is the focus of the full v1: it continues pre-training to about 10B tokens and adds training on long agentic sessions.

Format and serving

Volundr's chat format uses Blockway's control tokens: <|agens_start|> / <|agens_end|> for turns, <|think|> … <|/think|> for reasoning, <|call|> … <|/call|> for tool calls and <|result|> … <|/result|> for tool results. Each control token is a single reserved id, so turn boundaries, reasoning and tool calls never depend on how ordinary text tokenises. When no system message is supplied, the template inserts You are Agens, an AI assistant developed by Blockway.

Thinking is on by default (enable_thinking: false turns it off); reasoning_effort accepts low / medium / xhigh. Tool calls use the <function=…><parameter=…> form; the agens parsers in our sglang build expose them as OpenAI-style tool_calls.

Reasoning budget. Our sglang build adds a per-request reasoning_budget (int) that closes the thinking block at N generated tokens, so a hard constraint prompt always ends in an answer. Inference only, weights untouched. The IFBench rows use reasoning_budget: 6000.

Serving. Volundr runs on the Blockway sglang build (model class, agens parsers, reasoning budget). bf16 on two 48 GB GPUs:

docker run --rm --gpus '"device=0,1"' --ipc=host --network host --shm-size 32g \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  ghcr.io/blockwayz/agens-sglang:preview-sm89 \
  python3 -m sglang.launch_server \
    --model-path Blockway/Agens-Volundr-32B-Preview \
    --tp-size 2 --attention-backend flashinfer --page-size 1 \
    --disable-radix-cache --disable-prefill-cuda-graph --disable-custom-all-reduce \
    --mem-fraction-static 0.87 --context-length 32768 \
    --max-mamba-cache-size 16 --max-running-requests 16 --cuda-graph-max-bs 16 \
    --reasoning-parser agens --tool-call-parser agens \
    --host 127.0.0.1 --port 30000

Images: ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) and :preview-sm90 (H100/H200). The source patch is at github.com/BlockWayz/agens-sglang; its README has the commands for INT4 on one 48 GB GPU, bf16 on one H200 and the drafter, and explains each flag. Add --enable-multimodal --mm-feature-transport cpu to accept images. A vLLM plugin is not yet available.

Speed (measured on the Blockway sglang build, no drafter, one user unless noted; prompt of the given length plus 256 generated tokens):

SetupContextPrefill (tok/s)Decode (tok/s)
bf16, two 48 GB GPUs1K2,12225.1
bf16, two 48 GB GPUs8K2,18024.1
bf16, two 48 GB GPUs32K1,91624.1
bf16, two 48 GB GPUs64K1,67924.0
bf16, two 48 GB GPUs128K1,29723.9
INT4, one 48 GB GPU1K2,23631.0
INT4, one 48 GB GPU8K2,20029.3
INT4, one 48 GB GPU32K1,76429.1

Decode speed stays flat from 1K to 128K context: the 54 KDA layers carry a fixed-size state and the 17 BCSA layers read a bounded window plus 512 selected blocks, so the per-token cost barely grows with the conversation. With concurrent users (1K-token prompts, 256 new tokens each), total throughput is 127 tok/s at 8 users and 130 tok/s at 16 on two 48 GB GPUs (bf16), and 117 tok/s at 8 users on one 48 GB GPU (INT4, with 9 request slots; see the serving README).

Speculative decoding. A DFlash2 drafter for the Blockway sglang build is published as Blockway/Agens-Volundr-32B-Preview-DFlash2. Single user on two 48 GB GPUs, it makes JSON output up to 3.6× faster, code 2.0×, and replies with thinking on about 1.6×.

Architecture

Three quarters of the stack keep a fixed-size state instead of a growing KV cache, so long sessions stay cheap in memory and in decode time.

ComponentLayersHow it worksWhat it buys
KDA (Kimi Delta Attention)54Linear attention: a gated delta-rule update of a fixed-size state per head, with per-channel decayMemory and per-token decode cost independent of context length; no KV cache in these layers
BCSA (Blockway Compressed-Sparse Attention)17Exact attention over a 4,096-token sliding window, plus the far field pooled into 4-token blocks of which a learned indexer selects the top 512; both halves share one softmaxAttention cost per token bounded by the window plus 512 blocks, while keeping precise local detail and long-range recall
Dense attention1Standard full attentionOne exact global read of the whole context
Engram2Hashed 2-gram and 3-gram lookup into a 1M-row × 512 memory table, gated into the residual stream at layers 2 and 22Extra parametric memory read by a table lookup instead of matrix multiplies, at almost no compute cost
mHCallFour residual streams mixed by learned matrices kept doubly stochastic (Sinkhorn projection)Stable signal propagation through a 72-layer stack

What this means in practice: only 18 of 72 layers keep a KV cache. On two 48 GB GPUs (bf16) the server holds a 220K-token cache pool; the INT4 build (31.7 GiB) runs on a single 48 GB GPU.

Training

The Preview has had about 1.6B tokens of continued pre-training: repository-level code (each repository packed in dependency order, tests next to their sources), code and maths corpora, web text in English, Chinese and Cantonese, and GitHub issues, with chat-format replay under logit distillation so reasoning and turn-taking stay intact. Post-training covered identity, tool use (parallel calls, and declining when no tool fits) and preference training on pairs mined from the model's own samples.

Teacher outputs in the training data come only from open-weight models under permissive licences (Apache-2.0, MIT). There are no outputs from Anthropic, OpenAI or Google models.

Known limitations (Preview)

  • Long agentic coding sessions: in our SWE-bench runs through Codeway, Volundr fell into repetition loops in 36 of 50 tasks (output repeating a line, or tool calls written as plain text), against none for Qwen3.8-27B and Agens Pilot. Prefer short sessions or a supervising harness until v1.
  • Retrieval at the full 262K window has not been re-validated on this checkpoint.
  • Cantonese is a supported language, not a specialty.
  • The drafter is for single-user generation: with 8 or more concurrent requests, plain decoding gives more total throughput.

Acknowledgements

The mixer stack builds on published research: Kimi Delta Attention (Moonshot AI), and DeepSeek's sparse attention, Engram and manifold hyper-connections. The combination, the BCSA design and the training programme are Blockway's.

Licence

Apache-2.0. The licence files and MODIFICATIONS.md ship with the weights.

agens
agent
blockway
cantonese
code
conversational
custom_code
image-text-to-text
linear-attention
long-context
safetensors
sparse-attention
text-generation
tool-use
transformers
volundr