INT4 build (4-bit weights, group size 128) of Agens Volundr 32B Preview. It runs on a single 48 GB GPU with the Blockway sglang build. The results section below is the bf16 model; this INT4 checkpoint, measured on the same harness:
| Benchmark | INT4 (this repo) | bf16 |
|---|---|---|
| AIME 2025, avg@4 | 73.3 (1 run) | 74.6 (mean of 2) |
| LiveCodeBench v6 | 56.5 (1 run) | 63.4 (mean of 2) |
| MATH-500 | 97.0 | 98.2 |
| IFEval (prompt strict) | 87.4 | 87.6 |
| MMLU-Pro | 80.2 | 79.9 |
| CMMLU | 85.8 | 86.5 |
Agens Volundr 32B is an open-weights model from Blockway, built for reasoning, coding and agent work. This is a Preview release: the architecture, format and post-training pipeline are final; continued pre-training is still running and the full v1 will replace this checkpoint. In this Preview, long agentic sessions are the weak spot (see Known limitations).
It runs on Blockway's own mixer stack: 72 layers — 54 of KDA (linear attention with channel-wise gating: fixed-size state, so a long session costs no extra memory per turn), 17 of BCSA (Blockway Compressed-Sparse Attention: a 4,096-token dense window plus a compressed far field selected by a learned indexer) and one dense attention layer — with Engram (n-gram conditional memory) and mHC (manifold hyper-connections, four residual streams). About 32 billion parameters.
Volundr speaks English, Chinese (Simplified and Traditional) and Cantonese; handles text and images; calls tools; and ships with Blockway's control tokens and chat template (see Format and serving).

| Benchmark | Agens Volundr 32B Preview | Qwen3.8-27B | Agens Pilot |
|---|---|---|---|
| Coding | |||
| Competitive coding — LiveCodeBench v6 (2025-02 to 2025-05), mean of 2 runs | 63.4 ▲ +4.2 | 59.2 | 58.8 |
| Code completion — HumanEval | 81.7 ▲ +4.3 | 77.4 | 80.5 |
| Agentic coding — SWE-bench Verified, 50 tasks (Codeway) | 44.0 | 58.0 | 64.0 |
| General | |||
| Competition maths — AIME 2025, avg@4, mean of 2 runs | 74.6 ▲ +2.9 | 71.7 | 66.7 |
| Maths — MATH-500 | 98.2 ▲ +1.6 | 96.6 | 97.2 |
| Chinese knowledge — CMMLU | 86.5 ▲ +0.1 | 86.4 | 86.1 |
| Instruction following — IFEval (prompt strict) | 87.6 | 88.4 | 85.6 |
| Instruction following — IFBench (prompt strict) | 66.3 | 69.3 | 63.0 |
| Scientific reasoning — GPQA Diamond | 81.7 | 83.8 | 81.8 |
| Knowledge — MMLU-Pro | 79.9 | 80.4 | 80.3 |
| Agent · focus of the full v1 | |||
| Multi-turn tool use — τ²-bench (airline, retail, telecom) | 74.2 | 79.2 | 80.0 |
| Parallel function calls — BFCL v4 parallel | 92.0 | 94.0 | 92.0 |
| Knowing when not to call — BFCL v4 irrelevance | 80.8 | 81.7 | 85.8 |
▲ Volundr ahead of Qwen3.8-27B (difference in points). Rows without ▲ are where this Preview trails; closing them is the focus of the full v1. Thinking on, temperature 0.6 and the same output-token limit for every model unless noted; answers cut off by the limit count as wrong. HumanEval: greedy code completion. GPQA Diamond: 6K-token thinking budget, then the answer is forced, for every model (Volundr: mean of 4 runs). τ²-bench: agent at temperature 0, Qwen3.8-27B as the user simulator for every model; the 13 retail tasks that need an LLM judge are scored as failed for every model (no LLM judges are used anywhere). BFCL v4: non-live categories, official harness at temperature 0.001. AIME 2025 and LiveCodeBench: mean of 2 runs. SWE-bench: 50-task subset of SWE-bench Verified ("Verified mini") run through Codeway, Blockway's coding-agent harness, with each model's default sampling and identical limits. MMLU-Pro and CMMLU: fixed stratified subsets (1,400 and 2,010 questions). Qwen3.8-27B and Agens Pilot were run by us on the same harness; these are not their publishers' figures.
Reading the table: Volundr is ahead of Qwen3.8-27B on competitive coding (LiveCodeBench v6, +4.2), code completion (HumanEval, +4.3) and competition maths (AIME 2025, +2.9; MATH-500, +1.6), and level on Chinese knowledge (CMMLU). It trails on agentic work, most clearly on SWE-bench through Codeway, where it often fell into repetition loops in long sessions, and by a few points on instruction following and GPQA. Closing those gaps is the focus of the full v1: it continues pre-training to about 10B tokens and adds training on long agentic sessions.
Volundr's chat format uses Blockway's control tokens: <|agens_start|> / <|agens_end|> for turns, <|think|> … <|/think|> for reasoning, <|call|> … <|/call|> for tool calls and <|result|> … <|/result|> for tool results. Each control token is a single reserved id, so turn boundaries, reasoning and tool calls never depend on how ordinary text tokenises. When no system message is supplied, the template inserts You are Agens, an AI assistant developed by Blockway.
Thinking is on by default (enable_thinking: false turns it off); reasoning_effort accepts low / medium / xhigh. Tool calls use the <function=…><parameter=…> form; the agens parsers in our sglang build expose them as OpenAI-style tool_calls.
Reasoning budget. Our sglang build adds a per-request reasoning_budget (int) that closes the thinking block at N generated tokens, so a hard constraint prompt always ends in an answer. Inference only, weights untouched. The IFBench rows use reasoning_budget: 6000.
Serving. Volundr runs on the Blockway sglang build (model class, agens parsers, reasoning budget). INT4 on one 48 GB GPU:
docker run --rm --gpus '"device=0"' --ipc=host --network host --shm-size 32g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/blockwayz/agens-sglang:preview-sm89 \
python3 -m sglang.launch_server \
--model-path Blockway/Agens-Volundr-32B-Preview-INT4 \
--tp-size 1 --attention-backend flashinfer --page-size 1 \
--disable-radix-cache --disable-prefill-cuda-graph \
--mem-fraction-static 0.86 --context-length 16384 \
--max-mamba-cache-size 8 --max-running-requests 8 --cuda-graph-max-bs 8 \
--reasoning-parser agens --tool-call-parser agens \
--host 127.0.0.1 --port 30000
Images: ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) and :preview-sm90 (H100/H200). The source patch is at github.com/BlockWayz/agens-sglang; its README has the commands for INT4 on one 48 GB GPU, bf16 on one H200 and the drafter, and explains each flag. Add --enable-multimodal --mm-feature-transport cpu to accept images. A vLLM plugin is not yet available.
Speed (measured on the Blockway sglang build, no drafter, one user unless noted; prompt of the given length plus 256 generated tokens):
| Setup | Context | Prefill (tok/s) | Decode (tok/s) |
|---|---|---|---|
| bf16, two 48 GB GPUs | 1K | 2,122 | 25.1 |
| bf16, two 48 GB GPUs | 8K | 2,180 | 24.1 |
| bf16, two 48 GB GPUs | 32K | 1,916 | 24.1 |
| bf16, two 48 GB GPUs | 64K | 1,679 | 24.0 |
| bf16, two 48 GB GPUs | 128K | 1,297 | 23.9 |
| INT4, one 48 GB GPU | 1K | 2,236 | 31.0 |
| INT4, one 48 GB GPU | 8K | 2,200 | 29.3 |
| INT4, one 48 GB GPU | 32K | 1,764 | 29.1 |
Decode speed stays flat from 1K to 128K context: the 54 KDA layers carry a fixed-size state and the 17 BCSA layers read a bounded window plus 512 selected blocks, so the per-token cost barely grows with the conversation. With concurrent users (1K-token prompts, 256 new tokens each), total throughput is 127 tok/s at 8 users and 130 tok/s at 16 on two 48 GB GPUs (bf16), and 117 tok/s at 8 users on one 48 GB GPU (INT4, with 9 request slots; see the serving README).
Speculative decoding. A DFlash2 drafter for the Blockway sglang build is published as Blockway/Agens-Volundr-32B-Preview-DFlash2. Single user on two 48 GB GPUs, it makes JSON output up to 3.6× faster, code 2.0×, and replies with thinking on about 1.6×.
Three quarters of the stack keep a fixed-size state instead of a growing KV cache, so long sessions stay cheap in memory and in decode time.
| Component | Layers | How it works | What it buys |
|---|---|---|---|
| KDA (Kimi Delta Attention) | 54 | Linear attention: a gated delta-rule update of a fixed-size state per head, with per-channel decay | Memory and per-token decode cost independent of context length; no KV cache in these layers |
| BCSA (Blockway Compressed-Sparse Attention) | 17 | Exact attention over a 4,096-token sliding window, plus the far field pooled into 4-token blocks of which a learned indexer selects the top 512; both halves share one softmax | Attention cost per token bounded by the window plus 512 blocks, while keeping precise local detail and long-range recall |
| Dense attention | 1 | Standard full attention | One exact global read of the whole context |
| Engram | 2 | Hashed 2-gram and 3-gram lookup into a 1M-row × 512 memory table, gated into the residual stream at layers 2 and 22 | Extra parametric memory read by a table lookup instead of matrix multiplies, at almost no compute cost |
| mHC | all | Four residual streams mixed by learned matrices kept doubly stochastic (Sinkhorn projection) | Stable signal propagation through a 72-layer stack |
What this means in practice: only 18 of 72 layers keep a KV cache. On two 48 GB GPUs (bf16) the server holds a 220K-token cache pool; the INT4 build (31.7 GiB) runs on a single 48 GB GPU.
The Preview has had about 1.6B tokens of continued pre-training: repository-level code (each repository packed in dependency order, tests next to their sources), code and maths corpora, web text in English, Chinese and Cantonese, and GitHub issues, with chat-format replay under logit distillation so reasoning and turn-taking stay intact. Post-training covered identity, tool use (parallel calls, and declining when no tool fits) and preference training on pairs mined from the model's own samples.
Teacher outputs in the training data come only from open-weight models under permissive licences (Apache-2.0, MIT). There are no outputs from Anthropic, OpenAI or Google models.
The mixer stack builds on published research: Kimi Delta Attention (Moonshot AI), and DeepSeek's sparse attention, Engram and manifold hyper-connections. The combination, the BCSA design and the training programme are Blockway's.
Apache-2.0. The licence files and MODIFICATIONS.md ship with the weights.
INT4 build (4-bit weights, group size 128) of Agens Volundr 32B Preview. It runs on a single 48 GB GPU with the Blockway sglang build. The results section below is the bf16 model; this INT4 checkpoint, measured on the same harness:
| Benchmark | INT4 (this repo) | bf16 |
|---|---|---|
| AIME 2025, avg@4 | 73.3 (1 run) | 74.6 (mean of 2) |
| LiveCodeBench v6 | 56.5 (1 run) | 63.4 (mean of 2) |
| MATH-500 | 97.0 | 98.2 |
| IFEval (prompt strict) | 87.4 | 87.6 |
| MMLU-Pro | 80.2 | 79.9 |
| CMMLU | 85.8 | 86.5 |
Agens Volundr 32B is an open-weights model from Blockway, built for reasoning, coding and agent work. This is a Preview release: the architecture, format and post-training pipeline are final; continued pre-training is still running and the full v1 will replace this checkpoint. In this Preview, long agentic sessions are the weak spot (see Known limitations).
It runs on Blockway's own mixer stack: 72 layers — 54 of KDA (linear attention with channel-wise gating: fixed-size state, so a long session costs no extra memory per turn), 17 of BCSA (Blockway Compressed-Sparse Attention: a 4,096-token dense window plus a compressed far field selected by a learned indexer) and one dense attention layer — with Engram (n-gram conditional memory) and mHC (manifold hyper-connections, four residual streams). About 32 billion parameters.
Volundr speaks English, Chinese (Simplified and Traditional) and Cantonese; handles text and images; calls tools; and ships with Blockway's control tokens and chat template (see Format and serving).

| Benchmark | Agens Volundr 32B Preview | Qwen3.8-27B | Agens Pilot |
|---|---|---|---|
| Coding | |||
| Competitive coding — LiveCodeBench v6 (2025-02 to 2025-05), mean of 2 runs | 63.4 ▲ +4.2 | 59.2 | 58.8 |
| Code completion — HumanEval | 81.7 ▲ +4.3 | 77.4 | 80.5 |
| Agentic coding — SWE-bench Verified, 50 tasks (Codeway) | 44.0 | 58.0 | 64.0 |
| General | |||
| Competition maths — AIME 2025, avg@4, mean of 2 runs | 74.6 ▲ +2.9 | 71.7 | 66.7 |
| Maths — MATH-500 | 98.2 ▲ +1.6 | 96.6 | 97.2 |
| Chinese knowledge — CMMLU | 86.5 ▲ +0.1 | 86.4 | 86.1 |
| Instruction following — IFEval (prompt strict) | 87.6 | 88.4 | 85.6 |
| Instruction following — IFBench (prompt strict) | 66.3 | 69.3 | 63.0 |
| Scientific reasoning — GPQA Diamond | 81.7 | 83.8 | 81.8 |
| Knowledge — MMLU-Pro | 79.9 | 80.4 | 80.3 |
| Agent · focus of the full v1 | |||
| Multi-turn tool use — τ²-bench (airline, retail, telecom) | 74.2 | 79.2 | 80.0 |
| Parallel function calls — BFCL v4 parallel | 92.0 | 94.0 | 92.0 |
| Knowing when not to call — BFCL v4 irrelevance | 80.8 | 81.7 | 85.8 |
▲ Volundr ahead of Qwen3.8-27B (difference in points). Rows without ▲ are where this Preview trails; closing them is the focus of the full v1. Thinking on, temperature 0.6 and the same output-token limit for every model unless noted; answers cut off by the limit count as wrong. HumanEval: greedy code completion. GPQA Diamond: 6K-token thinking budget, then the answer is forced, for every model (Volundr: mean of 4 runs). τ²-bench: agent at temperature 0, Qwen3.8-27B as the user simulator for every model; the 13 retail tasks that need an LLM judge are scored as failed for every model (no LLM judges are used anywhere). BFCL v4: non-live categories, official harness at temperature 0.001. AIME 2025 and LiveCodeBench: mean of 2 runs. SWE-bench: 50-task subset of SWE-bench Verified ("Verified mini") run through Codeway, Blockway's coding-agent harness, with each model's default sampling and identical limits. MMLU-Pro and CMMLU: fixed stratified subsets (1,400 and 2,010 questions). Qwen3.8-27B and Agens Pilot were run by us on the same harness; these are not their publishers' figures.
Reading the table: Volundr is ahead of Qwen3.8-27B on competitive coding (LiveCodeBench v6, +4.2), code completion (HumanEval, +4.3) and competition maths (AIME 2025, +2.9; MATH-500, +1.6), and level on Chinese knowledge (CMMLU). It trails on agentic work, most clearly on SWE-bench through Codeway, where it often fell into repetition loops in long sessions, and by a few points on instruction following and GPQA. Closing those gaps is the focus of the full v1: it continues pre-training to about 10B tokens and adds training on long agentic sessions.
Volundr's chat format uses Blockway's control tokens: <|agens_start|> / <|agens_end|> for turns, <|think|> … <|/think|> for reasoning, <|call|> … <|/call|> for tool calls and <|result|> … <|/result|> for tool results. Each control token is a single reserved id, so turn boundaries, reasoning and tool calls never depend on how ordinary text tokenises. When no system message is supplied, the template inserts You are Agens, an AI assistant developed by Blockway.
Thinking is on by default (enable_thinking: false turns it off); reasoning_effort accepts low / medium / xhigh. Tool calls use the <function=…><parameter=…> form; the agens parsers in our sglang build expose them as OpenAI-style tool_calls.
Reasoning budget. Our sglang build adds a per-request reasoning_budget (int) that closes the thinking block at N generated tokens, so a hard constraint prompt always ends in an answer. Inference only, weights untouched. The IFBench rows use reasoning_budget: 6000.
Serving. Volundr runs on the Blockway sglang build (model class, agens parsers, reasoning budget). INT4 on one 48 GB GPU:
docker run --rm --gpus '"device=0"' --ipc=host --network host --shm-size 32g \
-v ~/.cache/huggingface:/root/.cache/huggingface \
ghcr.io/blockwayz/agens-sglang:preview-sm89 \
python3 -m sglang.launch_server \
--model-path Blockway/Agens-Volundr-32B-Preview-INT4 \
--tp-size 1 --attention-backend flashinfer --page-size 1 \
--disable-radix-cache --disable-prefill-cuda-graph \
--mem-fraction-static 0.86 --context-length 16384 \
--max-mamba-cache-size 8 --max-running-requests 8 --cuda-graph-max-bs 8 \
--reasoning-parser agens --tool-call-parser agens \
--host 127.0.0.1 --port 30000
Images: ghcr.io/blockwayz/agens-sglang:preview-sm89 (48 GB Ada GPUs) and :preview-sm90 (H100/H200). The source patch is at github.com/BlockWayz/agens-sglang; its README has the commands for INT4 on one 48 GB GPU, bf16 on one H200 and the drafter, and explains each flag. Add --enable-multimodal --mm-feature-transport cpu to accept images. A vLLM plugin is not yet available.
Speed (measured on the Blockway sglang build, no drafter, one user unless noted; prompt of the given length plus 256 generated tokens):
| Setup | Context | Prefill (tok/s) | Decode (tok/s) |
|---|---|---|---|
| bf16, two 48 GB GPUs | 1K | 2,122 | 25.1 |
| bf16, two 48 GB GPUs | 8K | 2,180 | 24.1 |
| bf16, two 48 GB GPUs | 32K | 1,916 | 24.1 |
| bf16, two 48 GB GPUs | 64K | 1,679 | 24.0 |
| bf16, two 48 GB GPUs | 128K | 1,297 | 23.9 |
| INT4, one 48 GB GPU | 1K | 2,236 | 31.0 |
| INT4, one 48 GB GPU | 8K | 2,200 | 29.3 |
| INT4, one 48 GB GPU | 32K | 1,764 | 29.1 |
Decode speed stays flat from 1K to 128K context: the 54 KDA layers carry a fixed-size state and the 17 BCSA layers read a bounded window plus 512 selected blocks, so the per-token cost barely grows with the conversation. With concurrent users (1K-token prompts, 256 new tokens each), total throughput is 127 tok/s at 8 users and 130 tok/s at 16 on two 48 GB GPUs (bf16), and 117 tok/s at 8 users on one 48 GB GPU (INT4, with 9 request slots; see the serving README).
Speculative decoding. A DFlash2 drafter for the Blockway sglang build is published as Blockway/Agens-Volundr-32B-Preview-DFlash2. Single user on two 48 GB GPUs, it makes JSON output up to 3.6× faster, code 2.0×, and replies with thinking on about 1.6×.
Three quarters of the stack keep a fixed-size state instead of a growing KV cache, so long sessions stay cheap in memory and in decode time.
| Component | Layers | How it works | What it buys |
|---|---|---|---|
| KDA (Kimi Delta Attention) | 54 | Linear attention: a gated delta-rule update of a fixed-size state per head, with per-channel decay | Memory and per-token decode cost independent of context length; no KV cache in these layers |
| BCSA (Blockway Compressed-Sparse Attention) | 17 | Exact attention over a 4,096-token sliding window, plus the far field pooled into 4-token blocks of which a learned indexer selects the top 512; both halves share one softmax | Attention cost per token bounded by the window plus 512 blocks, while keeping precise local detail and long-range recall |
| Dense attention | 1 | Standard full attention | One exact global read of the whole context |
| Engram | 2 | Hashed 2-gram and 3-gram lookup into a 1M-row × 512 memory table, gated into the residual stream at layers 2 and 22 | Extra parametric memory read by a table lookup instead of matrix multiplies, at almost no compute cost |
| mHC | all | Four residual streams mixed by learned matrices kept doubly stochastic (Sinkhorn projection) | Stable signal propagation through a 72-layer stack |
What this means in practice: only 18 of 72 layers keep a KV cache. On two 48 GB GPUs (bf16) the server holds a 220K-token cache pool; the INT4 build (31.7 GiB) runs on a single 48 GB GPU.
The Preview has had about 1.6B tokens of continued pre-training: repository-level code (each repository packed in dependency order, tests next to their sources), code and maths corpora, web text in English, Chinese and Cantonese, and GitHub issues, with chat-format replay under logit distillation so reasoning and turn-taking stay intact. Post-training covered identity, tool use (parallel calls, and declining when no tool fits) and preference training on pairs mined from the model's own samples.
Teacher outputs in the training data come only from open-weight models under permissive licences (Apache-2.0, MIT). There are no outputs from Anthropic, OpenAI or Google models.
The mixer stack builds on published research: Kimi Delta Attention (Moonshot AI), and DeepSeek's sparse attention, Engram and manifold hyper-connections. The combination, the BCSA design and the training programme are Blockway's.
Apache-2.0. The licence files and MODIFICATIONS.md ship with the weights.