Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2
C++
351
438 commits
updated Sep 28, 2026
Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware:
Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory.
Contrinutions are welcome!
See the changelog and GitHub Releases for user-facing changes and release history.
All model documentation lives under docs/models:
| Model | Inference modes | Hugging Face weights | Benchmarks | Quality |
|---|---|---|---|---|
| Qwen3.8 27B | Q4/Q8, images, AR, DFlash2 | Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_M | Q4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · Benchmarks | Quality |
| Qwen3.8 Flash-Next | Q4, images, AR, MTP | Unsloth Q4_K_XL · MTP Q8_0 | 1,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · Benchmarks | Quality |
| DeepSeek V4 Flash | AR, DSpark | antirez Flash 0731 IQ2XXS · DSpark | 484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · Benchmarks | Quality |
| Qwen3-ASR 1.7B | Speech recognition | BF16 | 15.27× realtime · Benchmarks | Quality |
| Qwen3-TTS 1.7B | Speech synthesis and voice cloning | BF16 CustomVoice / VoiceDesign / Base | Up to 2.54× realtime; 201 ms to first audio (CustomVoice) · Benchmarks | Quality |
| Qwen-Image-2.1 | BF16 image generation and editing | Complete pipeline | In progress · Benchmarks | Quality |
| MiniMax H3 | BF16 text to video/audio | FL2VA pipeline | In progress · Benchmarks | Quality |
Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.
hf download unsloth/Qwen3.8-27B-GGUF \
Qwen3.8-27B-UD-Q8_K_XL.gguf \
--revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
--repo-type model \
--local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
--repo-type model \
--local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
--userns=keep-id:uid=1000,gid=1000 \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--ulimit memlock=-1 \
-p 8080:8080 \
-v ./models:/models:ro \
ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
gufo serve --host 0.0.0.0 --port 8080 llm \
--model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
--speculative dflash2 \
--dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.gguf
Rootless Podman needs crun for --group-add keep-groups. Your host user must
have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the
render and video groups; log out and back in after changing membership.
Container groups named video/render do not preserve host supplementary
groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and
podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device.
See Podman's rootless group-access guidance.
On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while
SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading
“Memory critical” error. Check the host audit log:
sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'
If it shows a denied map for the container, Podman documents this fix:
sudo setsebool -P container_use_devices true
This persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B-UD-Q8_K_XL",
"messages": [{"role": "user", "content": "Say something"}]
}'
For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to
http://localhost:8080/v1. Chat Completions supports text, images, tools and
streaming; Responses supports text and streaming. See the API contract.
The text server uses the model's native context by default and generates until
EOS or the context is full. --context N sets context capacity per session;
--max-tokens N sets a default response limit that clients can override.
Reasoning tokens count toward that response limit.
Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns
one production configuration for both Nix and ordinary Linux builds. Tests,
profilers, tuning executables and Python reference runners are not installed
with the production package. No build.sh wrapper is needed.
nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.gguf
flake.lock pins the dependencies. nix develop adds profiling,
model-download and independent evaluation tools; these are not runtime
requirements. Optional benchmark baselines are selected separately with
nix shell .#ds4-reference, .#llama-cpp-reference or
.#llama-cpp-mtp-reference; see benchmarking.
See testing for the small hosted CI suite and
explicit local quality checks.
Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.
| Dependency | Used for |
|---|---|
| ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLAS | GPU execution and matrix multiplication |
| hipCUB, rocPRIM, rocWMMA headers | Compiled GPU kernels |
| ICU, libcurl, OpenSSL, libpng, libjpeg | Tokenization, HTTPS, hashing and images |
| FFmpeg and ffprobe | Video/audio output; invoked as separate executables |
Install ROCm using AMD's Linux instructions.
Use the development packages for the libraries above. ROCm normally installs
under /opt/rocm.
For example, on Debian/Ubuntu the ordinary system libraries are:
sudo apt install build-essential cmake ninja-build pkg-config \
libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg
# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
hipcub-dev rocprim-dev rocwmma-dev
cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.gguf
Configuring fails at find_package(hipblas) when those ROCm packages are
missing. Other distributions name them -devel instead of -dev.
For nonstandard installations, pass ordinary CMake paths, for example
cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm".
If compiler discovery picks a system Clang, also pass
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++.
Use cmake --install build/release to install Gufo,
its runtime data and license notices. The GPU driver must allow your user to
access /dev/kfd and /dev/dri; model weights are acquired separately.
The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.
Gufo's original code is MIT licensed. Adapted code and dependencies
retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md
and licenses/, installed under share/licenses/gufo. Model weights are not
bundled and retain their publishers' terms.
The initial design is informed by the following open source projects:
C++
75.0%
HIP
15.7%
Python
7.8%
Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2
C++
351
438 commits
updated Sep 28, 2026
Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware:
Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory.
Contrinutions are welcome!
See the changelog and GitHub Releases for user-facing changes and release history.
All model documentation lives under docs/models:
| Model | Inference modes | Hugging Face weights | Benchmarks | Quality |
|---|---|---|---|---|
| Qwen3.8 27B | Q4/Q8, images, AR, DFlash2 | Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_M | Q4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · Benchmarks | Quality |
| Qwen3.8 Flash-Next | Q4, images, AR, MTP | Unsloth Q4_K_XL · MTP Q8_0 | 1,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · Benchmarks | Quality |
| DeepSeek V4 Flash | AR, DSpark | antirez Flash 0731 IQ2XXS · DSpark | 484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · Benchmarks | Quality |
| Qwen3-ASR 1.7B | Speech recognition | BF16 | 15.27× realtime · Benchmarks | Quality |
| Qwen3-TTS 1.7B | Speech synthesis and voice cloning | BF16 CustomVoice / VoiceDesign / Base | Up to 2.54× realtime; 201 ms to first audio (CustomVoice) · Benchmarks | Quality |
| Qwen-Image-2.1 | BF16 image generation and editing | Complete pipeline | In progress · Benchmarks | Quality |
| MiniMax H3 | BF16 text to video/audio | FL2VA pipeline | In progress · Benchmarks | Quality |
Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.
hf download unsloth/Qwen3.8-27B-GGUF \
Qwen3.8-27B-UD-Q8_K_XL.gguf \
--revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
--repo-type model \
--local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
--revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
--repo-type model \
--local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
--userns=keep-id:uid=1000,gid=1000 \
--device /dev/kfd \
--device /dev/dri \
--group-add keep-groups \
--ulimit memlock=-1 \
-p 8080:8080 \
-v ./models:/models:ro \
ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
gufo serve --host 0.0.0.0 --port 8080 llm \
--model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
--speculative dflash2 \
--dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.gguf
Rootless Podman needs crun for --group-add keep-groups. Your host user must
have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the
render and video groups; log out and back in after changing membership.
Container groups named video/render do not preserve host supplementary
groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and
podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device.
See Podman's rootless group-access guidance.
On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while
SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading
“Memory critical” error. Check the host audit log:
sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'
If it shows a denied map for the container, Podman documents this fix:
sudo setsebool -P container_use_devices true
This persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.
Then, from another terminal, ask it something through the OpenAI-compatible API:
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen3.8-27B-UD-Q8_K_XL",
"messages": [{"role": "user", "content": "Say something"}]
}'
For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to
http://localhost:8080/v1. Chat Completions supports text, images, tools and
streaming; Responses supports text and streaming. See the API contract.
The text server uses the model's native context by default and generates until
EOS or the context is full. --context N sets context capacity per session;
--max-tokens N sets a default response limit that clients can override.
Reasoning tokens count toward that response limit.
Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns
one production configuration for both Nix and ordinary Linux builds. Tests,
profilers, tuning executables and Python reference runners are not installed
with the production package. No build.sh wrapper is needed.
nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.gguf
flake.lock pins the dependencies. nix develop adds profiling,
model-download and independent evaluation tools; these are not runtime
requirements. Optional benchmark baselines are selected separately with
nix shell .#ds4-reference, .#llama-cpp-reference or
.#llama-cpp-mtp-reference; see benchmarking.
See testing for the small hosted CI suite and
explicit local quality checks.
Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.
| Dependency | Used for |
|---|---|
| ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLAS | GPU execution and matrix multiplication |
| hipCUB, rocPRIM, rocWMMA headers | Compiled GPU kernels |
| ICU, libcurl, OpenSSL, libpng, libjpeg | Tokenization, HTTPS, hashing and images |
| FFmpeg and ffprobe | Video/audio output; invoked as separate executables |
Install ROCm using AMD's Linux instructions.
Use the development packages for the libraries above. ROCm normally installs
under /opt/rocm.
For example, on Debian/Ubuntu the ordinary system libraries are:
sudo apt install build-essential cmake ninja-build pkg-config \
libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg
# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
hipcub-dev rocprim-dev rocwmma-dev
cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.gguf
Configuring fails at find_package(hipblas) when those ROCm packages are
missing. Other distributions name them -devel instead of -dev.
For nonstandard installations, pass ordinary CMake paths, for example
cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm".
If compiler discovery picks a system Clang, also pass
-DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++.
Use cmake --install build/release to install Gufo,
its runtime data and license notices. The GPU driver must allow your user to
access /dev/kfd and /dev/dri; model weights are acquired separately.
The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.
Gufo's original code is MIT licensed. Adapted code and dependencies
retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md
and licenses/, installed under share/licenses/gufo. Model weights are not
bundled and retain their publishers' terms.
The initial design is informed by the following open source projects:
C++
75.0%
HIP
15.7%
Python
7.8%