Yamz-Labs/kyojin

Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.

Python

0

5 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine (r/LocalLLaMA)

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395. |Model|GLM-5.3-Flash|MiMo-V2.6-Flash-MOPD| |:-|:-|:-| |Size|99.7 GB|105 GB| |Prefill|580 tok/s at…

10

Oct 3, 2026

README

Yamz

Kyojin

Kyojin is the Yamz inference engine for AMD Strix Halo, built on ExLlamaV3 by turboderp. It adds a ROCm decode and prefill path for the AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, unified memory) and serving for two large MoE models with multi-token prediction:

  • GLM-5.3-Flash (glm_moe_dsa: MLA attention, sparse indexer, MTP)
  • MiMo-V2.6-Flash (mimo_v2, DFlash drafter)

The CUDA paths of upstream are kept. AMD additions sit behind USE_ROCM and architecture guards. This repository holds the engine, the serving scripts and the benchmark harnesses. The quantisation pipeline that produced the model packs is not part of it.

Measured on one Strix Halo machine: GLM-5.3-Flash prefills at 546 to 584 tok/s (3.5K to 64K context) and decodes at 26 to 30 tok/s with MTP; MiMo-V2.6-Flash decodes at 32 to 44 tok/s with speculative decoding on (29 tok/s plain) and prefills at about 650 tok/s (numbers and sources below). Both models run in 128 GB.

Weights: Yamz on Hugging Face - yamz-labs/GLM-5.3-Flash-EXL3-Yamz, yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz.

Quickstart (Strix Halo, ROCm)

Requirements: Linux (Ubuntu/Debian tested), a gfx1151 machine with 128 GB, Python 3.12, gcc, a ROCm 7 libhsa-runtime64.so.1 (the PyTorch copy segfaults on gfx1151 and Ubuntu's own libhsa-runtime64-1 is ROCm 5.7, too old; tools/strix_halo/env.sh finds the one inside the SDK wheel below first, then /opt/rocm, or take EXL3_HSA_LIB=<path>), ROCm 7.0 or newer (ROCm 6.4 has no gfx1151 code), a ROCm build of PyTorch for gfx1151, and a ROCm SDK devel tree with the hipsparse/ and thrust/ headers (the rocm-sdk-devel wheel).

git clone https://github.com/Yamz-Labs/kyojin && cd kyojin
python3 -m venv .venv && source .venv/bin/activate
# 1. ROCm torch + SDK first (AMD gfx1151 wheels; plain `pip install torch` gives a CUDA/CPU build that cannot run here):
pip install --pre torch rocm-sdk-devel --index-url https://rocm.nightlies.amd.com/v2/gfx1151/
pip install -r requirements.txt
rocm-sdk init                              # expands the devel headers (about 12 GB on disk)
export EXL3_ROCM_SDK=$(rocm-sdk path --root)   # or the path of your own devel tree
source tools/strix_halo/env.sh             # run from the repository root; sets LD_PRELOAD and PYTHONPATH (torch needs it to import)
./build.sh                                 # compiles the extension for gfx1151 into the repo root (exllamav3_ext*.so)
bash tools/strix_halo/env.sh --check       # prints versions, runs a small GPU matmul
hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz --local-dir ./glm-pack
python tools/glm/serve.py --model ./glm-pack --port 8000 -c 131072 --num-draft 2
curl http://localhost:8000/v1/models

Run source tools/strix_halo/env.sh again in every new shell before serving. Status: build and MiMo serving were verified from a fresh clone on a second Strix Halo machine (build in about 8 minutes). Both published packs were checked there against SHA256SUMS and served (one chat request each); env.sh --check has not been run there yet. The first request after a build is slow while the kernels warm up.

MiMo: hf download yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz --local-dir ./mimo-pack, then python tools/mimo/serve.py --model ./mimo-pack --port 8000 -c 131072. Speculative decoding (DFlash, 4 bpw drafter, confidence-truncated drafts) is on by default: the server uses the pack's drafter/ directory (or $MIMO_DRAFTER, or --drafter <dir>). Without a drafter it logs one line and decodes plain. Set MIMO_SPEC=0 in the lane script (tools/lanes/serve_mimo.sh) or pass --no-dflash to serve.py for plain decode. Greedy output under speculation is not token-identical to plain decode: near-tied logits can flip under the batched verify. A loaded drafter costs 2 to 4 % prefill. Details: tools/mimo/SERVE.md.

Measured numbers

One machine: Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X, ROCm. Other GPUs are untested.

Model / packContextPrefill tok/sDecode tok/sSource
GLM-5.3-Flash, 82 GB pack, MTP 24K / 16K / 64K / 128K644.6 / 620.9 / 617.7 / 608.232.1 prose, 33.8 chat, 35.8 code; 38.9 chat at 128K contextrun
GLM-5.3-Flash (99.7 GB pack), MTP 2, -c 524288, raw server3.7K / 14.2K611.9 / 585.027.1 default sampling, 28.1 greedyrun, single runs
same, through an OpenAI-style proxy3.7K / 14.4K590.0 / 589.829.1 default sampling, 28.1 greedyrun, single runs
GLM-5.3-Flash, 99 GB pack, MTP 2, -c 13107224K609.028.1 (28.12, 28.09)run, same harness as the next row
GLM-5.3-Flash, 82 GB pack, hook off, today's engine, -c 131072, client temperature 0, mean of 33.5K / 14K661 / 64532.0 prose, 33.6 chat, 35.7 coderun, second machine (same CPU)
GLM-5.3-Flash (99.7 GB), hook on + agent-lane flags, -c 98304, prefill mean of 3, decode mean of 63.5K / 14K / 64K580 / 584 / 54629.0 / 30.3 / 27.6 at temperature 0; 26.0 / 28.5 / 26.4 at temperature 1.0, top-p 0.95run
MiMo-V2.6-Flash-MOPD, 105 GB pack, speculative decoding on (default), 4 bpw drafter (702 MB), -c 32768, client temperature 0, decode medians of 6 runs over two loads, 128 tokens-32.1 prose, 34.8 chat, 44.3 code (plain: 28.9 on all three, 1.11x / 1.21x / 1.53x)run, public tree
MiMo-V2.6-Flash-MOPD, 105 GB pack, plain decode (no draft), 32K window, client temperature 04.1K / 23.7K594 / 61325.2 to 27.8run, single runs

First launch. The engine tunes its dense GEMM kernels on the first requests and keeps the result in a cache. On a fresh install the first GLM prefills run at 200 to 240 tok/s; speed reaches the figures above within a few requests and stays there on later launches.

llama.cpp (ROCm, UD-IQ1_S 1.56 bpw, -fa 1 -ub 2048, no MTP), same machine: pp4096 197.9, pp16384 158.3, tg128 16.74, tg at 64K 7.19 tok/s. Coarser quant: engine and format are compared together.

Quality against the official FP8 weights (129 held-out rows): see the model cards.

Share your numbers

Start a server from the quickstart, then run tools/bench.sh (standard library only, --base and --model select the server). It measures prefill on a prompt of about 3.5K tokens and decode on prose, chat and code with the prompts behind the table above, and prints one Markdown block with your hardware and versions. Paste it into a benchmark report. Results from other gfx1151 machines and other ROCm GPUs are the most useful contribution. See CONTRIBUTING.md.

Optional refusal hook

The engine can project one fixed direction out of the residual stream at run time. It edits no weights and no quantised data. Off unless EXL3_ABLIT_RUNTIME=/path/to/spec.json is set or the model folder contains uncensor_spec.json.

  • Bundled spec: if <model_dir>/uncensor_spec.json exists (with uncensor_direction.st next to it, or uncensor_spec.safetensors), the engine applies it at load and logs -- ablit runtime: spec <path> active (bundled in the model directory). An explicit EXL3_ABLIT_RUNTIME=<path> wins. EXL3_ABLIT_RUNTIME=off or --no-uncensor (GLM and MiMo servers) disables it. A missing or malformed spec stops the load with an error that names the file.
  • The spec is a JSON file (hidden, n_layers, per-layer weights attn_w[L], mlp_w[L]) plus a unit vector r in spec.safetensors.
  • After each attention and MLP block of layer L the output becomes y - w_L * r * (r . y).
  • EXL3_ABLIT_TORCH=1 forces the PyTorch path instead of the Triton kernel.
  • A direction fitted on refusals makes the model refuse less. Whoever uses a spec is responsible for it. This repository contains no spec file. Steering presets are published separately in the Yamz presets repository (https://github.com/yamz-labs/yamz-presets).
  • Code: exllamav3/modules/ablit_runtime.py. Test: tests/test_ablit_runtime_cpu.py, tests/test_uncensor_bundled_cpu.py.

Tests

pip install pytest
for t in tests/test_ablit_runtime_cpu.py tests/test_uncensor_bundled_cpu.py tools/glm/test_serve.py tools/mimo/test_serve.py tools/mimo/test_toolcalls.py; do PYTHONPATH=. pytest -q $t; done

Run each file separately because two files share a name (test_serve.py). These tests need no GPU and no built extension. Last run (CPU only, clean clone and venv): 3, 16, 7 of 8, 13 and 15 passed (tools/glm/test_serve.py takes a few minutes; its one failure, test_dense_tune_path_follows_the_cpp_tuner, needs the built extension). GPU tests need a built extension and a free GPU.

Build on CUDA

The CUDA build is the upstream one and is unchanged. Install a CUDA 12.4 or newer build of PyTorch, then pip install -r requirements.txt && pip install .. README.upstream.md has the full upstream guide (wheels, PyPI, uv, Windows, architecture list, conversion tool, examples). README.strix-halo.md has the kernel notes and benchmark harnesses.

Credits and licence

ExLlamaV3 by turboderp (MIT, LICENSE unchanged). ROCm decode path for gfx12 from sdougbrown/exllamav3; first gfx1151 port from vcruz305/exllamav3-amd. exllamav3/vendor/fla is flash-linear-attention (MIT). GLM-5.3-Flash is by Z.ai, MiMo-V2.6-Flash by Xiaomi; check each base licence before redistributing weights. Additions: MIT. This project is not affiliated with Z.ai, Xiaomi or turboderp.

amd
exl3
inference
llm
quantization
rocm
strix-halo

Yamz-Labs/kyojin

Kyojin: the Yamz inference engine for AMD Strix Halo (ROCm, gfx1151), built on ExLlamaV3. Runs 300B-class MoE models on one 128 GB mini PC.

Python

0

5 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

Two ~300B MoE models, each on ONE 128 GB mini PC (AMD Strix Halo): GLM-5.3-Flash at ~580 tok/s prefill, MiMo-V2.6-Flash up to 44 tok/s decode. EXL3 weights + open ROCm engine (r/LocalLLaMA)

We built an engine, Kyojin, on top of ExLlamaV3 for Strix Halo (gfx1151, ROCm), and packed two 300B-class MoE models so each fits one 128 GB machine. First release, all measured on Ryzen AI Max+ 395. |Model|GLM-5.3-Flash|MiMo-V2.6-Flash-MOPD| |:-|:-|:-| |Size|99.7 GB|105 GB| |Prefill|580 tok/s at…

10

Oct 3, 2026

README

Yamz

Kyojin

Kyojin is the Yamz inference engine for AMD Strix Halo, built on ExLlamaV3 by turboderp. It adds a ROCm decode and prefill path for the AMD Ryzen AI Max+ 395 (Radeon 8060S, gfx1151, unified memory) and serving for two large MoE models with multi-token prediction:

  • GLM-5.3-Flash (glm_moe_dsa: MLA attention, sparse indexer, MTP)
  • MiMo-V2.6-Flash (mimo_v2, DFlash drafter)

The CUDA paths of upstream are kept. AMD additions sit behind USE_ROCM and architecture guards. This repository holds the engine, the serving scripts and the benchmark harnesses. The quantisation pipeline that produced the model packs is not part of it.

Measured on one Strix Halo machine: GLM-5.3-Flash prefills at 546 to 584 tok/s (3.5K to 64K context) and decodes at 26 to 30 tok/s with MTP; MiMo-V2.6-Flash decodes at 32 to 44 tok/s with speculative decoding on (29 tok/s plain) and prefills at about 650 tok/s (numbers and sources below). Both models run in 128 GB.

Weights: Yamz on Hugging Face - yamz-labs/GLM-5.3-Flash-EXL3-Yamz, yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz.

Quickstart (Strix Halo, ROCm)

Requirements: Linux (Ubuntu/Debian tested), a gfx1151 machine with 128 GB, Python 3.12, gcc, a ROCm 7 libhsa-runtime64.so.1 (the PyTorch copy segfaults on gfx1151 and Ubuntu's own libhsa-runtime64-1 is ROCm 5.7, too old; tools/strix_halo/env.sh finds the one inside the SDK wheel below first, then /opt/rocm, or take EXL3_HSA_LIB=<path>), ROCm 7.0 or newer (ROCm 6.4 has no gfx1151 code), a ROCm build of PyTorch for gfx1151, and a ROCm SDK devel tree with the hipsparse/ and thrust/ headers (the rocm-sdk-devel wheel).

git clone https://github.com/Yamz-Labs/kyojin && cd kyojin
python3 -m venv .venv && source .venv/bin/activate
# 1. ROCm torch + SDK first (AMD gfx1151 wheels; plain `pip install torch` gives a CUDA/CPU build that cannot run here):
pip install --pre torch rocm-sdk-devel --index-url https://rocm.nightlies.amd.com/v2/gfx1151/
pip install -r requirements.txt
rocm-sdk init                              # expands the devel headers (about 12 GB on disk)
export EXL3_ROCM_SDK=$(rocm-sdk path --root)   # or the path of your own devel tree
source tools/strix_halo/env.sh             # run from the repository root; sets LD_PRELOAD and PYTHONPATH (torch needs it to import)
./build.sh                                 # compiles the extension for gfx1151 into the repo root (exllamav3_ext*.so)
bash tools/strix_halo/env.sh --check       # prints versions, runs a small GPU matmul
hf download yamz-labs/GLM-5.3-Flash-EXL3-Yamz --local-dir ./glm-pack
python tools/glm/serve.py --model ./glm-pack --port 8000 -c 131072 --num-draft 2
curl http://localhost:8000/v1/models

Run source tools/strix_halo/env.sh again in every new shell before serving. Status: build and MiMo serving were verified from a fresh clone on a second Strix Halo machine (build in about 8 minutes). Both published packs were checked there against SHA256SUMS and served (one chat request each); env.sh --check has not been run there yet. The first request after a build is slow while the kernels warm up.

MiMo: hf download yamz-labs/MiMo-V2.6-Flash-MOPD-EXL3-Yamz --local-dir ./mimo-pack, then python tools/mimo/serve.py --model ./mimo-pack --port 8000 -c 131072. Speculative decoding (DFlash, 4 bpw drafter, confidence-truncated drafts) is on by default: the server uses the pack's drafter/ directory (or $MIMO_DRAFTER, or --drafter <dir>). Without a drafter it logs one line and decodes plain. Set MIMO_SPEC=0 in the lane script (tools/lanes/serve_mimo.sh) or pass --no-dflash to serve.py for plain decode. Greedy output under speculation is not token-identical to plain decode: near-tied logits can flip under the batched verify. A loaded drafter costs 2 to 4 % prefill. Details: tools/mimo/SERVE.md.

Measured numbers

One machine: Ryzen AI Max+ 395, Radeon 8060S (gfx1151), 128 GB LPDDR5X, ROCm. Other GPUs are untested.

Model / packContextPrefill tok/sDecode tok/sSource
GLM-5.3-Flash, 82 GB pack, MTP 24K / 16K / 64K / 128K644.6 / 620.9 / 617.7 / 608.232.1 prose, 33.8 chat, 35.8 code; 38.9 chat at 128K contextrun
GLM-5.3-Flash (99.7 GB pack), MTP 2, -c 524288, raw server3.7K / 14.2K611.9 / 585.027.1 default sampling, 28.1 greedyrun, single runs
same, through an OpenAI-style proxy3.7K / 14.4K590.0 / 589.829.1 default sampling, 28.1 greedyrun, single runs
GLM-5.3-Flash, 99 GB pack, MTP 2, -c 13107224K609.028.1 (28.12, 28.09)run, same harness as the next row
GLM-5.3-Flash, 82 GB pack, hook off, today's engine, -c 131072, client temperature 0, mean of 33.5K / 14K661 / 64532.0 prose, 33.6 chat, 35.7 coderun, second machine (same CPU)
GLM-5.3-Flash (99.7 GB), hook on + agent-lane flags, -c 98304, prefill mean of 3, decode mean of 63.5K / 14K / 64K580 / 584 / 54629.0 / 30.3 / 27.6 at temperature 0; 26.0 / 28.5 / 26.4 at temperature 1.0, top-p 0.95run
MiMo-V2.6-Flash-MOPD, 105 GB pack, speculative decoding on (default), 4 bpw drafter (702 MB), -c 32768, client temperature 0, decode medians of 6 runs over two loads, 128 tokens-32.1 prose, 34.8 chat, 44.3 code (plain: 28.9 on all three, 1.11x / 1.21x / 1.53x)run, public tree
MiMo-V2.6-Flash-MOPD, 105 GB pack, plain decode (no draft), 32K window, client temperature 04.1K / 23.7K594 / 61325.2 to 27.8run, single runs

First launch. The engine tunes its dense GEMM kernels on the first requests and keeps the result in a cache. On a fresh install the first GLM prefills run at 200 to 240 tok/s; speed reaches the figures above within a few requests and stays there on later launches.

llama.cpp (ROCm, UD-IQ1_S 1.56 bpw, -fa 1 -ub 2048, no MTP), same machine: pp4096 197.9, pp16384 158.3, tg128 16.74, tg at 64K 7.19 tok/s. Coarser quant: engine and format are compared together.

Quality against the official FP8 weights (129 held-out rows): see the model cards.

Share your numbers

Start a server from the quickstart, then run tools/bench.sh (standard library only, --base and --model select the server). It measures prefill on a prompt of about 3.5K tokens and decode on prose, chat and code with the prompts behind the table above, and prints one Markdown block with your hardware and versions. Paste it into a benchmark report. Results from other gfx1151 machines and other ROCm GPUs are the most useful contribution. See CONTRIBUTING.md.

Optional refusal hook

The engine can project one fixed direction out of the residual stream at run time. It edits no weights and no quantised data. Off unless EXL3_ABLIT_RUNTIME=/path/to/spec.json is set or the model folder contains uncensor_spec.json.

  • Bundled spec: if <model_dir>/uncensor_spec.json exists (with uncensor_direction.st next to it, or uncensor_spec.safetensors), the engine applies it at load and logs -- ablit runtime: spec <path> active (bundled in the model directory). An explicit EXL3_ABLIT_RUNTIME=<path> wins. EXL3_ABLIT_RUNTIME=off or --no-uncensor (GLM and MiMo servers) disables it. A missing or malformed spec stops the load with an error that names the file.
  • The spec is a JSON file (hidden, n_layers, per-layer weights attn_w[L], mlp_w[L]) plus a unit vector r in spec.safetensors.
  • After each attention and MLP block of layer L the output becomes y - w_L * r * (r . y).
  • EXL3_ABLIT_TORCH=1 forces the PyTorch path instead of the Triton kernel.
  • A direction fitted on refusals makes the model refuse less. Whoever uses a spec is responsible for it. This repository contains no spec file. Steering presets are published separately in the Yamz presets repository (https://github.com/yamz-labs/yamz-presets).
  • Code: exllamav3/modules/ablit_runtime.py. Test: tests/test_ablit_runtime_cpu.py, tests/test_uncensor_bundled_cpu.py.

Tests

pip install pytest
for t in tests/test_ablit_runtime_cpu.py tests/test_uncensor_bundled_cpu.py tools/glm/test_serve.py tools/mimo/test_serve.py tools/mimo/test_toolcalls.py; do PYTHONPATH=. pytest -q $t; done

Run each file separately because two files share a name (test_serve.py). These tests need no GPU and no built extension. Last run (CPU only, clean clone and venv): 3, 16, 7 of 8, 13 and 15 passed (tools/glm/test_serve.py takes a few minutes; its one failure, test_dense_tune_path_follows_the_cpp_tuner, needs the built extension). GPU tests need a built extension and a free GPU.

Build on CUDA

The CUDA build is the upstream one and is unchanged. Install a CUDA 12.4 or newer build of PyTorch, then pip install -r requirements.txt && pip install .. README.upstream.md has the full upstream guide (wheels, PyPI, uv, Windows, architecture list, conversion tool, examples). README.strix-halo.md has the kernel notes and benchmark harnesses.

Credits and licence

ExLlamaV3 by turboderp (MIT, LICENSE unchanged). ROCm decode path for gfx12 from sdougbrown/exllamav3; first gfx1151 port from vcruz305/exllamav3-amd. exllamav3/vendor/fla is flash-linear-attention (MIT). GLM-5.3-Flash is by Z.ai, MiMo-V2.6-Flash by Xiaomi; check each base licence before redistributing weights. Additions: MIT. This project is not affiliated with Z.ai, Xiaomi or turboderp.

amd
exl3
inference
llm
quantization
rocm
strix-halo

Languages

Python

71.5%

Cuda

20.7%

C++

6.9%