gufo-org/gufo

Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2

C++

351

438 commits

updated Sep 28, 2026

See the code

See what people are saying

SourceMessageScoreDate

If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context. (r/LocalLLaMA)

Here's the project. I have nothing to do with it. I'm just an amazed user. https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BENCHMARKS.md Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat. "[6204 chunks in 119.0 s | encode: 1239…

20

Sep 28, 2026

README

Gufo: the Strix Halo inference engine

Gufo logo

Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware: Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory.

Contrinutions are welcome!

See the changelog and GitHub Releases for user-facing changes and release history.

Models and benchmarks

All model documentation lives under docs/models:

ModelInference modesHugging Face weightsBenchmarksQuality
Qwen3.8 27BQ4/Q8, images, AR, DFlash2Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_MQ4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · BenchmarksQuality
Qwen3.8 Flash-NextQ4, images, AR, MTPUnsloth Q4_K_XL · MTP Q8_01,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · BenchmarksQuality
DeepSeek V4 FlashAR, DSparkantirez Flash 0731 IQ2XXS · DSpark484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · BenchmarksQuality
Qwen3-ASR 1.7BSpeech recognitionBF1615.27× realtime · BenchmarksQuality
Qwen3-TTS 1.7BSpeech synthesis and voice cloningBF16 CustomVoice / VoiceDesign / BaseUp to 2.54× realtime; 201 ms to first audio (CustomVoice) · BenchmarksQuality
Qwen-Image-2.1BF16 image generation and editingComplete pipelineIn progress · BenchmarksQuality
MiniMax H3BF16 text to video/audioFL2VA pipelineIn progress · BenchmarksQuality

Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.

Philosophy

  • Contributions are welcome! We need the help of Strix Halo community to keep improving gufo!
  • We would like this to be the one-stop shop for Strix Halo Local AI enthusiasts: batteries included for text, audio, image, and video models.
  • Build and optimize specifically for the Strix Halo 128 GiB hardware. Smaller memory configurations should still work and preserve the speed benefits for models that can fit on memory.
  • Support only the best available models for their size that can run on this hardware: less code to maintain, more focused optimization and testing work.
  • Preserve quality when optimizing. Each model's quality report records independent numerical checks, execution consistency and unresolved gaps. Don't reuse kernels across different models to limit blast radius of a code change.
  • Treat concurrent requests, cancellation and conversation caching as first-class workloads.
  • Keep production dependencies small and development tools separate.

Quickstart

hf download unsloth/Qwen3.8-27B-GGUF \
  Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
  Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
  --revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
  --userns=keep-id:uid=1000,gid=1000 \
  --device /dev/kfd \
  --device /dev/dri \
  --group-add keep-groups \
  --ulimit memlock=-1 \
  -p 8080:8080 \
  -v ./models:/models:ro \
  ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
  gufo serve --host 0.0.0.0 --port 8080 llm \
  --model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --speculative dflash2 \
  --dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.gguf

Rootless Podman needs crun for --group-add keep-groups. Your host user must have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the render and video groups; log out and back in after changing membership. Container groups named video/render do not preserve host supplementary groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device. See Podman's rootless group-access guidance.

On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading “Memory critical” error. Check the host audit log:

sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'

If it shows a denied map for the container, Podman documents this fix:

sudo setsebool -P container_use_devices true

This persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.

Then, from another terminal, ask it something through the OpenAI-compatible API:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-UD-Q8_K_XL",
    "messages": [{"role": "user", "content": "Say something"}]
  }'

For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to http://localhost:8080/v1. Chat Completions supports text, images, tools and streaming; Responses supports text and streaming. See the API contract.

The text server uses the model's native context by default and generates until EOS or the context is full. --context N sets context capacity per session; --max-tokens N sets a default response limit that clients can override. Reasoning tokens count toward that response limit.

Build from source

Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns one production configuration for both Nix and ordinary Linux builds. Tests, profilers, tuning executables and Python reference runners are not installed with the production package. No build.sh wrapper is needed.

With Nix

nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.gguf

flake.lock pins the dependencies. nix develop adds profiling, model-download and independent evaluation tools; these are not runtime requirements. Optional benchmark baselines are selected separately with nix shell .#ds4-reference, .#llama-cpp-reference or .#llama-cpp-mtp-reference; see benchmarking. See testing for the small hosted CI suite and explicit local quality checks.

Without Nix

Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.

DependencyUsed for
ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLASGPU execution and matrix multiplication
hipCUB, rocPRIM, rocWMMA headersCompiled GPU kernels
ICU, libcurl, OpenSSL, libpng, libjpegTokenization, HTTPS, hashing and images
FFmpeg and ffprobeVideo/audio output; invoked as separate executables

Install ROCm using AMD's Linux instructions. Use the development packages for the libraries above. ROCm normally installs under /opt/rocm.

For example, on Debian/Ubuntu the ordinary system libraries are:

sudo apt install build-essential cmake ninja-build pkg-config \
  libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg

# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
  hipcub-dev rocprim-dev rocwmma-dev

cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.gguf

Configuring fails at find_package(hipblas) when those ROCm packages are missing. Other distributions name them -devel instead of -dev. For nonstandard installations, pass ordinary CMake paths, for example cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm". If compiler discovery picks a system Clang, also pass -DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++. Use cmake --install build/release to install Gufo, its runtime data and license notices. The GPU driver must allow your user to access /dev/kfd and /dev/dri; model weights are acquired separately.

The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.

License

Gufo's original code is MIT licensed. Adapted code and dependencies retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md and licenses/, installed under share/licenses/gufo. Model weights are not bundled and retain their publishers' terms.

Reference Projects

The initial design is informed by the following open source projects:

amd
gfx1151
llm
ryzen-ai
strix-halo

gufo-org/gufo

Strix Halo inference engine. Qwen Flash Next Q4_K_XL: 1,628.52pp, 59.41tg single user, 157.22 tok/s 8 users; Qwen27B Q4_K_XL: 656.33pp, 70.56tg tok/s single user with DFlash2

C++

351

438 commits

updated Sep 28, 2026

See the code

See what people are saying

SourceMessageScoreDate

If you are running Qwen 3.8 Flash Next on Strix Halo, use this software for inference. It's so much faster than llama.cpp especially at high context. (r/LocalLLaMA)

Here's the project. I have nothing to do with it. I'm just an amazed user. https://github.com/gufo-org/gufo/blob/main/docs/models/qwen3.8-flash-next/BENCHMARKS.md Those benchmark numbers hold up on real work loads. Here are some numbers I got during a chat. "[6204 chunks in 119.0 s | encode: 1239…

20

Sep 28, 2026

README

Gufo: the Strix Halo inference engine

Gufo logo

Gufo is a vertical local inference engine specifically built and optimized for the AMD Strix Halo hardware: Ryzen AI MAX+ 395 systems with Radeon 8060S (gfx1151), up to 128 GiB of unified memory.

Contrinutions are welcome!

See the changelog and GitHub Releases for user-facing changes and release history.

Models and benchmarks

All model documentation lives under docs/models:

ModelInference modesHugging Face weightsBenchmarksQuality
Qwen3.8 27BQ4/Q8, images, AR, DFlash2Unsloth Q4_K_XL / Q8_K_XL · DFlash2 Q4_K_MQ4: 656.33 tok/s pp; up to 70.56 tok/s tg single user and 123.00 aggregated tok/s on 8 concurrent requests with DFlash2 · BenchmarksQuality
Qwen3.8 Flash-NextQ4, images, AR, MTPUnsloth Q4_K_XL · MTP Q8_01,628.52 tok/s pp; up to 59.41 tok/s tg single user and 157.22 aggregated tok/s on 8 concurrent requests with MTP · BenchmarksQuality
DeepSeek V4 FlashAR, DSparkantirez Flash 0731 IQ2XXS · DSpark484.62 tok/s pp; up to 26.62 tok/s tg single user and 54.74 aggregated tok/s on 8 concurrent requests with DSpark · BenchmarksQuality
Qwen3-ASR 1.7BSpeech recognitionBF1615.27× realtime · BenchmarksQuality
Qwen3-TTS 1.7BSpeech synthesis and voice cloningBF16 CustomVoice / VoiceDesign / BaseUp to 2.54× realtime; 201 ms to first audio (CustomVoice) · BenchmarksQuality
Qwen-Image-2.1BF16 image generation and editingComplete pipelineIn progress · BenchmarksQuality
MiniMax H3BF16 text to video/audioFL2VA pipelineIn progress · BenchmarksQuality

Peak measured workloads; text pp is autoregressive (AR), while tg uses the named speculative mode. Peaks include repetitive output; aggregate tg sums individual request decode rates. Qwen27B's single-user peak uses the short-prompt C1 workload. Audio excludes loading. Each model guide lists the required files and complete benchmark settings.

Philosophy

  • Contributions are welcome! We need the help of Strix Halo community to keep improving gufo!
  • We would like this to be the one-stop shop for Strix Halo Local AI enthusiasts: batteries included for text, audio, image, and video models.
  • Build and optimize specifically for the Strix Halo 128 GiB hardware. Smaller memory configurations should still work and preserve the speed benefits for models that can fit on memory.
  • Support only the best available models for their size that can run on this hardware: less code to maintain, more focused optimization and testing work.
  • Preserve quality when optimizing. Each model's quality report records independent numerical checks, execution consistency and unresolved gaps. Don't reuse kernels across different models to limit blast radius of a code change.
  • Treat concurrent requests, cancellation and conversation caching as first-class workloads.
  • Keep production dependencies small and development tools separate.

Quickstart

hf download unsloth/Qwen3.8-27B-GGUF \
  Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --revision 4ca720788d1e01f1bff70c033e0d0028fd02e502 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-GGUF
hf download z-lab/Qwen3.8-27B-DFlash2-GGUF \
  Qwen3.8-27B-DFlash2-Q4_K_M.gguf \
  --revision 2d9571f8ce46e151f61c6499c99dee6079e1d610 \
  --repo-type model \
  --local-dir models/Qwen3.8-27B-DFlash2-GGUF
podman pull ghcr.io/gufo-org/toolboxes/gufo-runtime:latest
podman run --rm \
  --userns=keep-id:uid=1000,gid=1000 \
  --device /dev/kfd \
  --device /dev/dri \
  --group-add keep-groups \
  --ulimit memlock=-1 \
  -p 8080:8080 \
  -v ./models:/models:ro \
  ghcr.io/gufo-org/toolboxes/gufo-runtime:latest \
  gufo serve --host 0.0.0.0 --port 8080 llm \
  --model /models/Qwen3.8-27B-GGUF/Qwen3.8-27B-UD-Q8_K_XL.gguf \
  --speculative dflash2 \
  --dflash-model /models/Qwen3.8-27B-DFlash2-GGUF/Qwen3.8-27B-DFlash2-Q4_K_M.gguf

Rootless Podman needs crun for --group-add keep-groups. Your host user must have read/write access to /dev/kfd and /dev/dri/renderD*, usually through the render and video groups; log out and back in after changing membership. Container groups named video/render do not preserve host supplementary groups. Check id, ls -l /dev/kfd /dev/dri/renderD*, and podman info --format '{{.Host.OCIRuntime.Name}}' if ROCm reports no device. See Podman's rootless group-access guidance.

On Fedora or another SELinux-enforcing host, GPU enumeration can succeed while SELinux blocks mapping /dev/kfd, causing ROCr to report a misleading “Memory critical” error. Check the host audit log:

sudo ausearch -m avc -ts recent | grep -E '/dev/kfd|hsa_device_t'

If it shows a denied map for the container, Podman documents this fix:

sudo setsebool -P container_use_devices true

This persistently allows containers to access device labels for devices passed into them; it affects all containers on that host. Review that policy scope before enabling it. See Podman's device documentation and the SELinux container policy.

Then, from another terminal, ask it something through the OpenAI-compatible API:

curl http://localhost:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen3.8-27B-UD-Q8_K_XL",
    "messages": [{"role": "user", "content": "Say something"}]
  }'

For Open WebUI, VS Code and OpenAI SDK clients, set the API base URL to http://localhost:8080/v1. Chat Completions supports text, images, tools and streaming; Responses supports text and streaming. See the API contract.

The text server uses the model's native context by default and generates until EOS or the context is full. --context N sets context capacity per session; --max-tokens N sets a default response limit that clients can override. Reasoning tokens count toward that response limit.

Build from source

Linux x86-64 on AMD Strix Halo (gfx1151) is the supported target. CMake owns one production configuration for both Nix and ordinary Linux builds. Tests, profilers, tuning executables and Python reference runners are not installed with the production package. No build.sh wrapper is needed.

With Nix

nix build
./result/bin/gufo diagnose
./result/bin/gufo serve llm --model /path/to/model.gguf

flake.lock pins the dependencies. nix develop adds profiling, model-download and independent evaluation tools; these are not runtime requirements. Optional benchmark baselines are selected separately with nix shell .#ds4-reference, .#llama-cpp-reference or .#llama-cpp-mtp-reference; see benchmarking. See testing for the small hosted CI suite and explicit local quality checks.

Without Nix

Install a C++20 compiler, CMake 3.21+, Ninja, pkg-config and the following development libraries. The currently qualified toolchain is GCC 15.3 and ROCm 7.2.3. Attention and audio convolution kernels are compiled directly from HIP. Python, Triton/AOTriton, Composable Kernel and MIOpen are not production build or runtime requirements.

DependencyUsed for
ROCm HIP compiler/runtime, hipBLAS, hipBLASLt, rocBLASGPU execution and matrix multiplication
hipCUB, rocPRIM, rocWMMA headersCompiled GPU kernels
ICU, libcurl, OpenSSL, libpng, libjpegTokenization, HTTPS, hashing and images
FFmpeg and ffprobeVideo/audio output; invoked as separate executables

Install ROCm using AMD's Linux instructions. Use the development packages for the libraries above. ROCm normally installs under /opt/rocm.

For example, on Debian/Ubuntu the ordinary system libraries are:

sudo apt install build-essential cmake ninja-build pkg-config \
  libicu-dev libcurl4-openssl-dev libssl-dev libpng-dev libjpeg-dev ffmpeg

# ROCm libraries from the table, named as AMD's repository ships them.
sudo apt install hipblas-dev hipblaslt-dev rocblas-dev \
  hipcub-dev rocprim-dev rocwmma-dev

cmake --preset release -DCMAKE_INSTALL_PREFIX="$HOME/.local"
cmake --build --preset release --parallel 4
./build/release/gufo diagnose
./build/release/gufo serve llm --model /path/to/model.gguf

Configuring fails at find_package(hipblas) when those ROCm packages are missing. Other distributions name them -devel instead of -dev. For nonstandard installations, pass ordinary CMake paths, for example cmake --preset release -DCMAKE_PREFIX_PATH="/opt/rocm". If compiler discovery picks a system Clang, also pass -DCMAKE_HIP_COMPILER=/opt/rocm/llvm/bin/clang++. Use cmake --install build/release to install Gufo, its runtime data and license notices. The GPU driver must allow your user to access /dev/kfd and /dev/dri; model weights are acquired separately.

The same source, compiler flags and install rules serve both builds. Nix pins the complete toolchain for reproducible comparisons; changing the compiler or math libraries requires the affected model's quality checks.

License

Gufo's original code is MIT licensed. Adapted code and dependencies retain their own notices in NOTICE, THIRD_PARTY_NOTICES.md and licenses/, installed under share/licenses/gufo. Model weights are not bundled and retain their publishers' terms.

Reference Projects

The initial design is informed by the following open source projects:

amd
gfx1151
llm
ryzen-ai
strix-halo

Languages

C++

75.0%

HIP

15.7%

Python

7.8%