Castagna Veloce is an inference engine built and tuned for a single GPU: the AMD Instinct MI50
(Vega 20, gfx906), with 32 GB of HBM2 per card. It runs current large models across 2 to
8 cards: the DeepSeek V4 Flash, GLM-5.3-Flash and Qwen3.8 Flash-Next mixtures of experts and
the dense Qwen3.8 27B. The cards are joined by tensor parallelism over plain PCIe, and
speculative decoding (MTP, DSpark) makes single-user generation fast.
It is a fork of llama.cpp. It uses the same GGUF
weights and the same llama-server with its OpenAI-compatible API and Web UI. On top of that
it adds gfx906 kernels, fused operations for MoE-era architectures, and multi-GPU plumbing
designed for cards without matrix cores or fast interconnects.
Issues and contributions are welcome, especially from other MI50 and MI60 owners.
| Model | Cards | Inference modes | Hugging Face weights | Single-user speed |
|---|---|---|---|---|
| DeepSeek V4 Flash | 4 (TP4) | AR, DSpark | antirez IQ2XXS 0731 · DSpark 0731 | 1,161 tok/s pp; 81.3 tok/s tg with DSpark (63.5 AR) |
| GLM-5.3-Flash | 4 (TP4) | AR, MTP, images | Unsloth UD-Q2_K_XL · mmproj F16 | 978 tok/s pp; 79.5 tok/s tg with MTP (70.2 AR) |
| Qwen3.8 27B | 2 (TP2) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q4_0 · mmproj F16 | 702 tok/s pp; 69.5 tok/s tg with MTP (43.9 AR) |
| Qwen3.8 Flash-Next, decode profile | 4 (two TP2 pairs) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q8_0 · mmproj F16 | 1,939 tok/s pp; 119.0 tok/s tg with MTP (79.5 AR) |
| Qwen3.8 Flash-Next, prefill profile | 4 (layer split) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q8_0 · mmproj F16 | 2,615 tok/s pp; 79.1 tok/s tg with MTP (52.3 AR) |
| GLM-5.3-Flash Q4 | 8 (TP4×2) | AR, MTP, images | Unsloth UD-Q4_K_XL | 1,299 tok/s pp; 78.9 tok/s tg with MTP (63.7 AR) |
| DeepSeek V4.1 Flash | 8 (TP4×2) | AR, images | vcruz305 Q2_K · smalinin mmproj BF16 | 1,316 tok/s pp; 85.7 tok/s tg plain (86.4 AR) |
pp is llama-bench prompt processing of a 2,048-token prompt. AR is plain autoregressive
decoding (llama-bench, 128 tokens). The bold tg numbers are the llama-server decode
rate over 20 varied prompts of 200
greedy tokens each, with the model's drafter where it helps. Speculative speed depends on the
text: code, lists and translations draft far better than stories. All numbers are for a
single user, measured 2026-10-05/06 on the current code. Each model guide lists the exact
command and settings, and BENCHMARKS.md has the full method.
Everything below was written for Castagna Veloce, and none of it is in upstream llama.cpp (checked against v0.6.0). Most of it can be switched off with an environment variable for A/B tests. The speed figures were measured on 4 cards (Qwen3.8 27B on 2) when each feature landed.
DSV4_COMP_POOL, a new op that runs DeepSeek V4's compressor gather, pooling and
norm as one kernel: decode +8%. Skipping its padding blocks gave another +7%.LLAMA_CKPT_LAST_VERIFY): one graph rebuild fewer per
request, +1–2% throughput and a faster first token.LLAMA_MTP_DRAFT_VOCAB): the drafter scores only the first
96K token ids, split across the cards, while the target still checks every draft against
the full vocabulary. This is the idea of
FR-Spec applied to an MTP head.Castagna Veloce can run tensor parallelism in pairs of cards. Each pair splits every
matrix of its layers across its two cards, the layers are divided between the pairs, and the
pairs run as a pipeline. Qwen3.8 Flash-Next's decode profile uses this layout on 4 cards,
while DeepSeek V4, GLM-5.3 and DeepSeek V4.1 run as groups of four. Upstream llama.cpp's
-sm tensor always puts all cards into one group.
flowchart LR
subgraph P1["Pair 1: first half of the layers"]
G0["GPU 0"] <-->|all-reduce| G1["GPU 1"]
end
subgraph P2["Pair 2: second half of the layers"]
G2["GPU 2"] <-->|all-reduce| G3["GPU 3"]
end
P1 -->|activations| P2
Why pairs suit the MI50:
Choosing the layout:
LLAMA_TP_GROUP | Layout | Used by |
|---|---|---|
2 | pairs: 2 + 2 on 4 cards, 2 + 2 + 2 + 2 on 8 | Qwen3.8 Flash-Next, decode profile |
4 | groups of four: one group on 4 cards, two on 8 | DeepSeek V4, GLM-5.3, DeepSeek V4.1 |
0 | one group over all cards, as in upstream llama.cpp |
Left unset with -sm tensor on 4 or more cards, it also gives pairs, so the model guides
always set it. LLAMA_TP_LAYER_SPLIT=a,b,... sets the share of layers each group gets; by
default it follows each group's free memory.
gfx906). Every kernel choice is measured on these
cards. Other GPUs still build, but they are untested.test-backend-ops case checked
against the CPU backend, and every model's perplexity is re-checked after each change.
Changes meant to be exact are verified bit-identical. Changes that alter rounding say so
and are judged over many prompts, not one.GGML_CUDA_* environment
variable, to turn it off for A/B tests and bug reports.llama-server, the usual tools and flags.Supported setup: Ubuntu 26.04 LTS with its own ROCm 7.1 packages, which still ship gfx906
code. Kernel 7.0, HIP 7.1, LLVM/Clang 21.
sudo apt install build-essential cmake git hipcc libamdhip64-dev libhipblas-dev librocblas-dev \
rocm-device-libs-21 rocminfo rocm-smi
git clone https://github.com/benpeterson40/castagna-veloce
cd castagna-veloce
cmake -S . -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx906 -DGPU_TARGETS=gfx906 \
-DCMAKE_BUILD_TYPE=Release -DGGML_HIP_NO_VMM=ON -DGGML_HIP_GRAPHS=ON -DGGML_CUDA_FA=ON \
-DGGML_SCHED_MAX_COPIES=8 -DLLAMA_CURL=OFF
cmake --build build -j --target llama-server llama-cli llama-bench llama-perplexity
rocminfo | grep -c gfx906 # one line per card
Your user needs access to /dev/kfd and /dev/dri/renderD*, usually through the render
and video groups. Log out and back in after changing membership. If CMake picks the wrong
compiler, pass -DCMAKE_HIP_COMPILER=/usr/lib/llvm-21/bin/clang++.
For example GLM-5.3-Flash at 2 bits (about 109 GB, 4 cards):
pip install -U huggingface_hub # provides the hf command
hf download unsloth/GLM-5.3-Flash-GGUF --revision 621d456e93e926e4b52f85cff5f634358c1828f9 \
--include "UD-Q2_K_XL/*" --local-dir models/GLM-5.3-Flash-GGUF
HIP_VISIBLE_DEVICES=0,1,2,3 LLAMA_TP_GROUP=4 LLAMA_CKPT_LAST_VERIFY=1 \
./build/bin/llama-server \
-m models/GLM-5.3-Flash-GGUF/UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00001-of-00004.gguf \
-ngl 999 -sm tensor -fa on -c 65536 -b 4096 -ub 1024 --jinja -np 1 \
--spec-type draft-mtp --spec-draft-n-max 1 \
--host 0.0.0.0 --port 8080
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Say something"}]}'
LLAMA_TP_GROUP=4 makes the 4 cards one tensor-parallel group, and --spec-type draft-mtp
turns on the model's own MTP drafter. The GLM-5.3-Flash guide
explains every setting.
The Web UI is at http://localhost:8080. For Open WebUI, the OpenAI SDK and similar
clients, set the base URL to http://localhost:8080/v1. Each model guide
has the download and serve commands for that model.
<think> and has no
enable_thinking switch. Set the depth per request with
"chat_template_kwargs": {"reasoning_effort": "low"} (or "high"; the default is
"max"), or turn reasoning off with --reasoning-budget 0.--mmproj with the model's projector file (see its guide). The encoder
runs on the first card, and a very large photo can need over a gigabyte there, so cap
images with --image-max-tokens 2048."chat_template_kwargs": {"enable_thinking": false}, especially for images.-np 1) is the tuned configuration. The kernels for 1–6-token
batches are what make speculative decoding fast.gfx906
chip and should work, but it is untested. The 16 GB MI50 needs more cards for the same
models.echo high | sudo tee /sys/class/drm/card*/device/power_dpm_force_performance_level)
avoids clock ramps between drafting and verifying. With speculative decoding we measured
+0.4% on GLM-5.3 and about +5% on Flash-Next (the latter on an earlier build). The
published numbers use the default auto on all cards but the first.Castagna Veloce is a fork of llama.cpp and keeps its MIT license. The MI50-specific changes are MIT licensed as well. Model weights are not included and keep their publishers' licenses.
Castagna Veloce is not affiliated with or endorsed by AMD. AMD, Instinct and Radeon are trademarks of Advanced Micro Devices, Inc.
deepseek41 model and its vision encoder), ported here from an
earlier revision.
Castagna Veloce is an inference engine built and tuned for a single GPU: the AMD Instinct MI50
(Vega 20, gfx906), with 32 GB of HBM2 per card. It runs current large models across 2 to
8 cards: the DeepSeek V4 Flash, GLM-5.3-Flash and Qwen3.8 Flash-Next mixtures of experts and
the dense Qwen3.8 27B. The cards are joined by tensor parallelism over plain PCIe, and
speculative decoding (MTP, DSpark) makes single-user generation fast.
It is a fork of llama.cpp. It uses the same GGUF
weights and the same llama-server with its OpenAI-compatible API and Web UI. On top of that
it adds gfx906 kernels, fused operations for MoE-era architectures, and multi-GPU plumbing
designed for cards without matrix cores or fast interconnects.
Issues and contributions are welcome, especially from other MI50 and MI60 owners.
| Model | Cards | Inference modes | Hugging Face weights | Single-user speed |
|---|---|---|---|---|
| DeepSeek V4 Flash | 4 (TP4) | AR, DSpark | antirez IQ2XXS 0731 · DSpark 0731 | 1,161 tok/s pp; 81.3 tok/s tg with DSpark (63.5 AR) |
| GLM-5.3-Flash | 4 (TP4) | AR, MTP, images | Unsloth UD-Q2_K_XL · mmproj F16 | 978 tok/s pp; 79.5 tok/s tg with MTP (70.2 AR) |
| Qwen3.8 27B | 2 (TP2) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q4_0 · mmproj F16 | 702 tok/s pp; 69.5 tok/s tg with MTP (43.9 AR) |
| Qwen3.8 Flash-Next, decode profile | 4 (two TP2 pairs) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q8_0 · mmproj F16 | 1,939 tok/s pp; 119.0 tok/s tg with MTP (79.5 AR) |
| Qwen3.8 Flash-Next, prefill profile | 4 (layer split) | AR, MTP, images | Unsloth UD-Q4_K_XL · MTP Q8_0 · mmproj F16 | 2,615 tok/s pp; 79.1 tok/s tg with MTP (52.3 AR) |
| GLM-5.3-Flash Q4 | 8 (TP4×2) | AR, MTP, images | Unsloth UD-Q4_K_XL | 1,299 tok/s pp; 78.9 tok/s tg with MTP (63.7 AR) |
| DeepSeek V4.1 Flash | 8 (TP4×2) | AR, images | vcruz305 Q2_K · smalinin mmproj BF16 | 1,316 tok/s pp; 85.7 tok/s tg plain (86.4 AR) |
pp is llama-bench prompt processing of a 2,048-token prompt. AR is plain autoregressive
decoding (llama-bench, 128 tokens). The bold tg numbers are the llama-server decode
rate over 20 varied prompts of 200
greedy tokens each, with the model's drafter where it helps. Speculative speed depends on the
text: code, lists and translations draft far better than stories. All numbers are for a
single user, measured 2026-10-05/06 on the current code. Each model guide lists the exact
command and settings, and BENCHMARKS.md has the full method.
Everything below was written for Castagna Veloce, and none of it is in upstream llama.cpp (checked against v0.6.0). Most of it can be switched off with an environment variable for A/B tests. The speed figures were measured on 4 cards (Qwen3.8 27B on 2) when each feature landed.
DSV4_COMP_POOL, a new op that runs DeepSeek V4's compressor gather, pooling and
norm as one kernel: decode +8%. Skipping its padding blocks gave another +7%.LLAMA_CKPT_LAST_VERIFY): one graph rebuild fewer per
request, +1–2% throughput and a faster first token.LLAMA_MTP_DRAFT_VOCAB): the drafter scores only the first
96K token ids, split across the cards, while the target still checks every draft against
the full vocabulary. This is the idea of
FR-Spec applied to an MTP head.Castagna Veloce can run tensor parallelism in pairs of cards. Each pair splits every
matrix of its layers across its two cards, the layers are divided between the pairs, and the
pairs run as a pipeline. Qwen3.8 Flash-Next's decode profile uses this layout on 4 cards,
while DeepSeek V4, GLM-5.3 and DeepSeek V4.1 run as groups of four. Upstream llama.cpp's
-sm tensor always puts all cards into one group.
flowchart LR
subgraph P1["Pair 1: first half of the layers"]
G0["GPU 0"] <-->|all-reduce| G1["GPU 1"]
end
subgraph P2["Pair 2: second half of the layers"]
G2["GPU 2"] <-->|all-reduce| G3["GPU 3"]
end
P1 -->|activations| P2
Why pairs suit the MI50:
Choosing the layout:
LLAMA_TP_GROUP | Layout | Used by |
|---|---|---|
2 | pairs: 2 + 2 on 4 cards, 2 + 2 + 2 + 2 on 8 | Qwen3.8 Flash-Next, decode profile |
4 | groups of four: one group on 4 cards, two on 8 | DeepSeek V4, GLM-5.3, DeepSeek V4.1 |
0 | one group over all cards, as in upstream llama.cpp |
Left unset with -sm tensor on 4 or more cards, it also gives pairs, so the model guides
always set it. LLAMA_TP_LAYER_SPLIT=a,b,... sets the share of layers each group gets; by
default it follows each group's free memory.
gfx906). Every kernel choice is measured on these
cards. Other GPUs still build, but they are untested.test-backend-ops case checked
against the CPU backend, and every model's perplexity is re-checked after each change.
Changes meant to be exact are verified bit-identical. Changes that alter rounding say so
and are judged over many prompts, not one.GGML_CUDA_* environment
variable, to turn it off for A/B tests and bug reports.llama-server, the usual tools and flags.Supported setup: Ubuntu 26.04 LTS with its own ROCm 7.1 packages, which still ship gfx906
code. Kernel 7.0, HIP 7.1, LLVM/Clang 21.
sudo apt install build-essential cmake git hipcc libamdhip64-dev libhipblas-dev librocblas-dev \
rocm-device-libs-21 rocminfo rocm-smi
git clone https://github.com/benpeterson40/castagna-veloce
cd castagna-veloce
cmake -S . -B build -DGGML_HIP=ON -DAMDGPU_TARGETS=gfx906 -DGPU_TARGETS=gfx906 \
-DCMAKE_BUILD_TYPE=Release -DGGML_HIP_NO_VMM=ON -DGGML_HIP_GRAPHS=ON -DGGML_CUDA_FA=ON \
-DGGML_SCHED_MAX_COPIES=8 -DLLAMA_CURL=OFF
cmake --build build -j --target llama-server llama-cli llama-bench llama-perplexity
rocminfo | grep -c gfx906 # one line per card
Your user needs access to /dev/kfd and /dev/dri/renderD*, usually through the render
and video groups. Log out and back in after changing membership. If CMake picks the wrong
compiler, pass -DCMAKE_HIP_COMPILER=/usr/lib/llvm-21/bin/clang++.
For example GLM-5.3-Flash at 2 bits (about 109 GB, 4 cards):
pip install -U huggingface_hub # provides the hf command
hf download unsloth/GLM-5.3-Flash-GGUF --revision 621d456e93e926e4b52f85cff5f634358c1828f9 \
--include "UD-Q2_K_XL/*" --local-dir models/GLM-5.3-Flash-GGUF
HIP_VISIBLE_DEVICES=0,1,2,3 LLAMA_TP_GROUP=4 LLAMA_CKPT_LAST_VERIFY=1 \
./build/bin/llama-server \
-m models/GLM-5.3-Flash-GGUF/UD-Q2_K_XL/GLM-5.3-Flash-UD-Q2_K_XL-00001-of-00004.gguf \
-ngl 999 -sm tensor -fa on -c 65536 -b 4096 -ub 1024 --jinja -np 1 \
--spec-type draft-mtp --spec-draft-n-max 1 \
--host 0.0.0.0 --port 8080
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"messages": [{"role": "user", "content": "Say something"}]}'
LLAMA_TP_GROUP=4 makes the 4 cards one tensor-parallel group, and --spec-type draft-mtp
turns on the model's own MTP drafter. The GLM-5.3-Flash guide
explains every setting.
The Web UI is at http://localhost:8080. For Open WebUI, the OpenAI SDK and similar
clients, set the base URL to http://localhost:8080/v1. Each model guide
has the download and serve commands for that model.
<think> and has no
enable_thinking switch. Set the depth per request with
"chat_template_kwargs": {"reasoning_effort": "low"} (or "high"; the default is
"max"), or turn reasoning off with --reasoning-budget 0.--mmproj with the model's projector file (see its guide). The encoder
runs on the first card, and a very large photo can need over a gigabyte there, so cap
images with --image-max-tokens 2048."chat_template_kwargs": {"enable_thinking": false}, especially for images.-np 1) is the tuned configuration. The kernels for 1–6-token
batches are what make speculative decoding fast.gfx906
chip and should work, but it is untested. The 16 GB MI50 needs more cards for the same
models.echo high | sudo tee /sys/class/drm/card*/device/power_dpm_force_performance_level)
avoids clock ramps between drafting and verifying. With speculative decoding we measured
+0.4% on GLM-5.3 and about +5% on Flash-Next (the latter on an earlier build). The
published numbers use the default auto on all cards but the first.Castagna Veloce is a fork of llama.cpp and keeps its MIT license. The MI50-specific changes are MIT licensed as well. Model weights are not included and keep their publishers' licenses.
Castagna Veloce is not affiliated with or endorsed by AMD. AMD, Instinct and Radeon are trademarks of Advanced Micro Devices, Inc.
deepseek41 model and its vision encoder), ported here from an
earlier revision.