Dyluhn/R9V

Shape-specialized ROCm inference for Qwen3.8 Flash Next on RDNA4

Rust

60

205 commits

updated Sep 25, 2026

See the code

See what people are saying

SourceMessageScoreDate

Strata vs R9V running Qwen3.8 Flash Next on dual Radeon R9700 with 96GB DDR4 (r/LocalLLM)

We're hearing a lot about Strata these days, claiming it's the best solution to run Qwen3.8 Flash Next on consumer hardware. I have a dual R9700 config with 96GB DDR4 3200 and AMD 3900X, and so far I've been running Flash Next UD-IQ4\_XS on the custom R9V engine (…

0

Oct 5, 2026

README

R9V

CI

R9V runs Qwen3.8 Flash Next on two AMD Radeon AI PRO R9700 GPUs. It combines a pinned vLLM fork, GGUF loading, specialized gfx1201 kernels, expert offloading and four-token MTP speculative decoding. The server exposes an OpenAI-compatible API for text, tool calls and single-image requests.

Each profile binds a model package, runtime, hardware layout and expert placement. Downloads are checked against pinned revisions and file hashes. Setup records the selected configuration, and first start qualifies its workload and memory headroom before reporting ready.

Current status: both IQ4_XS and Q4_K_XL MTP4 profiles passed ordinary public setup, first-start workload qualification and unchanged-receipt restart on the dual-R9700 reference host. Both profiles remain experimental. Each first start passed all seven checks at 131,072 context, including a 130,941-token prompt, with at least 3 GiB free VRAM per GPU. Setup selects the profile's runtime image bundle, verifies every part and loads the exact image ID. The Q4 profile uses the v0.2.0-rc2 bundle; the IQ4 profile's WMMA-prefill image (release/image-bundle-wmma-prefill-20260915.json) is published under the v0.3.0-rc1-images release tag. See release status and evidence.

New in v0.4.4: qwen38-mtp4-uncensored with the default --ced on now shares one VRAM region per GPU between the vision encoder's weights and the CED projector, as --ced quality does since v0.4.3. GPU 0 keeps 2.07 GiB free at its minimum in first-start qualification instead of 1.65 (GPU 1: 2.52 instead of 2.12), with the same prefill speedup, decode, outputs and image support. Start now needs 60.8 GiB of available RAM with CED on (56.3 GiB with --ced off). See the changelog.

New in v0.4.3: qwen38-mtp4-uncensored --ced quality now passes first-start qualification on the reference host, with image support kept: the vision encoder's weights and the quality projector take turns in one VRAM region per GPU, and the projector's work buffers are smaller. GPU 0 kept 2.04 GiB free (v0.4.2: 0.90, target 1.5), more than on (1.65). See the changelog.

New in v0.4.2: qwen38-mtp4-uncensored has an opt-in CED quality mode, --ced quality, that gives up less long-context quality for less prefill speedup (it did not fit on the reference host until v0.4.3). Right after a first start, ./r9v doctor --runtime no longer fails its KV-pressure check because of qualification's own long prompt.

New in v0.4.1: qwen38-mtp4-uncensored keeps rank 1's 400 most-used experts per layer in VRAM for good, so its host copy of the experts drops from 55.4 to 40.3 GiB and it needs 56.3 GiB of free RAM at start instead of 71.4. In the GPU test its outputs matched 1.3.0 bit for bit on the test prompts and decode speed was unchanged. ./r9v doctor also checks disk space, VRAM held by other programs, API exposure, the runtime's pinned files and the CED projector. See the changelog.

New in v0.4.0: qwen38-mtp4-uncensored runs an uncensored (abliterated) IQ4_XS model on the consolidated 1.3.0 runtime with CED long-prompt prefill on by default. Its public fetch, setup, first-start qualification and restart passed from a clean checkout on the reference host, with CED on and off. See its qualification note.

Profiles and features

AliasModel packageRuntimeStatus
qwen38-mtp4IQ4_XS model bundleMTP4, dual R9700, 128K contextReference setup/start/restart passed
qwen38-q4-xlQ4_K_XL weightsMTP4, dual R9700, 128K contextReference setup/start/restart passed
qwen38-mtp4-uncensoredUncensored IQ4_XS bundle (abliterated, no refusals)Consolidated 1.3.0 runtime with host expert dedupe: MTP4, dual R9700, 128K context, CED on; needs 60.8 GiB free RAM at start (56.3 GiB with --ced off)v0.4.1 clean install (fetch, setup, first start CED on/off, restart) passed on the reference host; v0.4.3 (--ced quality) and v0.4.4 (--ced on) clean installs passed there too (below)

About qwen38-mtp4-uncensored. Its model had its refusal behavior removed and will comply with harmful requests the original model refuses. Use it for research, and add your own moderation before exposing it to anyone. R9V serves it on 127.0.0.1 only; setting R9V_HOST_BIND to expose it gives anyone who can reach the port an unauthenticated model with no refusals.

It has CED (approximate long-prompt prefill) on by default. On prompts of 8,192 tokens or more, layers 0–15 run exactly and a split-16 projector stands in for the later layers on all but the last ~2K prompt tokens. Decode stays exact. The tradeoff, measured on the reference host:

  • About 1.5× faster prefill at ~13K tokens, rising to about 1.8× at 32K tokens and above.
  • About ×1.051 perplexity on prompts that depend on their long context.
  • About 10% fewer MTP tokens per step on the first answer after a CED prefill; later turns are normal.
  • The projector takes 1.76 GiB of VRAM per GPU, so the profile's free-VRAM target is 1.5 GiB per card instead of 3 GiB. Since v0.4.4 it shares that VRAM with the vision encoder's 0.42 GiB (below).

Turn CED off for the server with --ced off in setup or start; keep a single request exact with "vllm_xargs": {"r9v_ced": false}.

CED quality (opt-in, --ced quality). A multi-source projector reads the layer-16 state plus the inputs of full-attention layers 3, 7, 11 and 15. In one GPU grade of both projectors (both loaded as int8; the default ships bf16), on the 16 prompts that depend on their long context:

--cedPerplexityLong-context gain lostPrefill speedup (median)Projector VRAM per GPU
on (default)×1.04917%1.68×1.76 GiB (bf16), shared with the vision encoder's 0.42 GiB
quality×1.02910%1.55×1.79 GiB (stored int8), shared with the vision encoder's 0.42 GiB

The projector math costs 56 ms per 1K approximated tokens instead of 24 ms. Setup downloads the quality projector (1.8 GB) only when you pick it: run setup --ced quality once, then start keeps the choice. on stays the default.

Tested on the compiled server (v0.4.3, v0.4.4). The projector shares one VRAM region per GPU with the vision encoder's weights (quality since v0.4.3, on since v0.4.4): image prompts never use CED, so the region holds whichever the next step needs and is refilled from pinned host RAM (36 ms / 9 ms on GPU 0, 266 ms / 63 ms on GPU 1; only when a CED prompt follows an image or the other way round). In the clean-install GPU test, first start with --ced quality passed all seven checks with 2.04 / 2.49 GiB free (target 1.5; on: 1.65 / 2.12) while desktop apps held 1.23 GiB of GPU 0. Prefill was 1.47× at 12.8K tokens and 1.62× at 32K (on: 1.58× and 1.76×), decode unchanged, exact requests bitwise identical after a CED request, no recompiles; 24 alternating CED and image requests kept VRAM flat and read every image. Quality needs 4.50 GiB more free RAM at start (60.8 GiB) for the pinned copies. v0.4.2 did not fit: GPU 0 fell to 0.90 GiB free. In v0.4.4's clean-install test on kept 2.07 / 2.52 GiB free (v0.4.3: 1.65 / 2.12) with prefill 1.59× at 12.8K tokens and 1.77× at 32K, the vision encoder's output bitwise unchanged, and it now needs the same 60.8 GiB of free RAM at start. See the changelog. The profile uses a fixed expert placement, so it does not accept --headroom.

Use the explicit MTP4 aliases for the current workflow.

  • Per-card headroom: choose the free VRAM to retain on each card with --headroom 3,3 or an asymmetric budget such as --headroom 5,3.
  • Measured expert maps: separate complete catalogs for IQ4 and Q4 rank all 512 experts across 48 layers per rank, using training and separate held-out routing captures. Existing hot prefixes are preserved, while cold-to-hot arrays use measured counts, with deterministic ties for zero-count experts. The planner retains frequently used experts in VRAM within the memory budget.
  • Q4_K_XL support: a dedicated package, packed expert costs, ranked placement and streaming loader. Q4 uses its upstream model bytes and original Q8_0 target head; IQ4 retains its original Q6_K target head.
  • Resumable setup and restart receipts: reuse matching model assets, save configuration and reuse its verified receipt on an unchanged restart after the workflow has qualified.
  • Doctor: inspect model/runtime/placement compatibility, GPU ordering, memory pressure, PCIe information, cache identity and available worker evidence.
  • Support bundles: collect configuration summaries, source identities, worker records, memory and GPU diagnostics, logs and capture tails into a bounded local archive with file hashes.

Unobserved experts are explicit ties in the maps. Routing frequency depends on workload; a different prompt mix can change the best placement. Requested headroom is checked against the qualification workload, and cannot prevent another application from allocating VRAM later.

Model downloads

The ./r9v setup commands below download and verify the required files automatically. For manual downloads, keep every shard and the package directory layout; use the IQ4 package manifest or Q4 package manifest for exact paths, revisions and hashes. The PLE table is extracted locally from the target GGUF; it is not a separate model download.

Hardware and storage

The reference system uses:

  • Two 32 GiB Radeon AI PRO R9700 cards (gfx1201).
  • Linux with working AMD GPU drivers, amd-smi, /dev/kfd and /dev/dri access.
  • 128 GiB host RAM. Smaller hosts are untested; cold expert allocations use host memory. qwen38-mtp4-uncensored checks for 60.8 GiB available before start with CED on or quality, 56.3 GiB with --ced off (was 71.4 GiB in v0.4.0): its host expert copy is 40.3 GiB plus a 16 GiB PLE reserve, plus 4.50 GiB of pinned copies of the CED projector and the vision encoder.
  • An asymmetric PCIe layout: rank 0 on Gen5 x16 and rank 1 across Gen4 x4. GPU ordering matters to placement and performance.
  • Git, Python 3.10+, Docker and the Hugging Face CLI described in the installation guide.

The public image bundle plus its containerd image-store footprint measured roughly 50 GiB; reserve at least 70 GiB for image and cache import space. The IQ4 package occupies approximately 90.36 GiB and the uncensored package 92.39 GiB. The four Q4 target shards alone occupy 103.69 GiB, with auxiliary assets additional. The derived PLE file occupies 26.82 GiB. Leave further room for compilation caches and diagnostics. Reuse verified assets instead of duplicating model files.

Setup and start

Clone the source and its exact dependencies:

git clone --recursive https://github.com/Dyluhn/R9V.git
cd R9V
./r9v show qwen38-mtp4
./r9v doctor qwen38-mtp4 -- --host-only

The profile distribution entries select the matching GitHub Release image bundle and exact image IDs. Docker 29 must use the containerd image store so docker load preserves those IDs. See Docker's containerd image store guide and daemon configuration reference. Verify the store before setup:

docker info -f '{{ .DriverStatus }}'

The output should identify io.containerd.snapshotter.v1. If it does not, follow the configuration steps in the installation guide. For a non-root daemon, see Docker's rootless mode guide and confirm the selected context/socket with docker info; the same containerd image-store check applies. These are the reference identities:

ProfileTested local image ID
qwen38-mtp4sha256:2dac17a215fb5b0e3461e4c3e36a2981eec8ac3d6021e73183d247e819740c03
qwen38-q4-xlsha256:2e50016cfcc9cd22f15d3f69ccf001e4877236e12ebb4ab458cc9c16caaef9e3
qwen38-mtp4-uncensoredsha256:2dac17a215fb5b0e3461e4c3e36a2981eec8ac3d6021e73183d247e819740c03 plus pinned overlays

Install the download CLI in an isolated environment if it is not already available:

python3 -m venv ~/.local/share/r9v/download-tools
~/.local/share/r9v/download-tools/bin/pip install huggingface_hub
export PATH="$HOME/.local/share/r9v/download-tools/bin:$PATH"

For IQ4, choose persistent SSD directories with enough free space:

MODEL_DIR=/path/to/qwen-iq4
STATE_DIR=/path/to/r9v-state/iq4
./r9v setup qwen38-mtp4 \
  --model-dir "$MODEL_DIR" --state-dir "$STATE_DIR" \
  --headroom 3,3 --accept-model-license
./r9v start qwen38-mtp4 --state-dir "$STATE_DIR"

For Q4, use its own model and state directories:

./r9v setup qwen38-q4-xl --model-dir /path/to/qwen-q4 \
  --state-dir /path/to/r9v-state/q4 --headroom 3,3 --accept-model-license
./r9v start qwen38-q4-xl --state-dir /path/to/r9v-state/q4

For the uncensored profile, use its own directories. Its first start compiles the model, so start waits up to 2,400 seconds for it by default (--timeout changes that):

./r9v setup qwen38-mtp4-uncensored --model-dir /path/to/qwen-uncensored \
  --state-dir /path/to/r9v-state/uncensored --accept-model-license
./r9v start qwen38-mtp4-uncensored --state-dir /path/to/r9v-state/uncensored

Run one profile at a time. Before switching, inspect docker ps, save any needed support evidence, and stop the selected R9V container by its exact name using docker stop NAME. Setup supports --reuse-from /path/to/existing/assets for matching assets and --ple-path /path/to/existing/ple.bin for an existing verified PLE file. The profiles select their corresponding expert catalog and reference memory seed automatically.

First start runs the local workload qualification, including context and headroom checks. An unchanged restart can reuse its verified receipt. To change the budget, stop the selected container, then run:

./r9v start qwen38-mtp4 --state-dir "$STATE_DIR" --headroom 5,3

The new placement must qualify. Card order follows the saved GPU selection; inspect it with doctor. The planner reports per-card shortfalls when a request cannot fit, while retaining the configured context.

The default API endpoint is http://127.0.0.1:8004/v1. The API has no authentication, so R9V publishes it on 127.0.0.1 only, in every profile. To reach it from other machines, run setup with R9V_HOST_BIND=0.0.0.0 (or one interface's IPv4/IPv6 address); setup saves it with the port. See API address. Use the address recorded by your selected setup if you override the port. See the release guide, installation guide, and troubleshooting guide. Rebuilding an image does not reproduce a qualified image identity automatically.

Measured results

Prompt processing

Measured with the original release’s corpus, script, warmup and 8K/32K/64K procedure. Rates below are mean prompt tokens/s; every trial and exact reproduction command are linked below.

Profile8K (10 runs)32K (3 runs)64K (2 runs)Raw trials
IQ4_XS, WMMA prefill (current qwen38-mtp4)1,684.61,640.61,605.1JSON
IQ4_XS, v0.2.0 image989.3982.0968.2JSON
Q4_K_XL552.9542.5536.4JSON

The current IQ4 profile runs prompt chunks of 4096 tokens on a gfx12 int8 WMMA grouped MoE prefill kernel; decode kernels, placement policy and the model are unchanged, and the placement assets were re-derived and qualified for the new image. Why v0.2.0 regressed, kernel parity, method and commands; v0.2.0 method. Prefix-cache hits were zero. The older generation references below use different placements and are separate results.

Earlier fixed-prompt generation references

The following fixed-prompt reference samples used MTP4 on the dual-R9700 system:

Profile / placementStatic experts, ranks 0/1Generation tokens/s
IQ4 image7 streaming reference71 / 45089.45196
Q4 measured ranked placement97 / 34953.431
Q4 initial bootstrap placement64 / 32025.285

These are fixed-prompt reference samples with MTP4; they do not measure mixed traffic or generation at full context. Actual user placements depend on the requested headroom and must qualify locally.

The IQ4 image7 streaming reference retained 131,072 context tokens and passed seven bounded workload checks, including an actual 130,941-token prompt, text, tools, three image shapes and idle resume. Its median was 89.45196 TG tok/s, with measured free VRAM of 4,253,020,160 and 4,090,036,224 bytes and a minimum Normal-zone free value of 450,269,184 bytes. The run had a clean 90-second aftermath and GPU reclaim. The supervisor incorrectly reported failure because its cleanup check required exact VRAM equality: rank 0 had 185.203 MiB more free and rank 1 was unchanged. Independent review confirmed no per-card shortfall throughout the aftermath. These reference measurements are separate from the completed ordinary user-flow checks below and do not establish answer quality or a 100 tok/s qualification.

The current IQ4 reference evidence records the measured result and independently verified archive commitments. Historical prefill and comparator results remain in the earlier Qwen qualification; they should not be substituted for measurements of the new placements.

Verified public setup and restart

Both profiles passed ordinary setup into new state directories, first-start text/tool/vision/context/idle-resume checks and restart with the same verified receipt. The first starts used an actual 130,941-token prompt at a 131,072 context limit. Model, head and PLE bytes stayed unchanged.

ProfileStatic experts, ranks 0/1Cache slots, ranks 0/1Minimum free VRAM, ranks 0/1Evidence
IQ4 image776 / 451160 / 03,828,301,824 / 4,120,670,208 bytesUser-flow report
Q4 image699 / 34880 / 04,072,394,752 / 4,013,797,376 bytesUser-flow report

These checks used a fresh anonymous source checkout and new setup state, reusing previously public-downloaded, hash-verified assets and images. The separate anonymous image import recovered from an interrupted load using unchanged verified parts. This was not a second fresh download or a full new-machine download in one uninterrupted run. Both user flows stopped cleanly, reclaimed GPU allocations and completed at least 90 seconds of post-stop observation. Q4 recorded one low Normal-zone memory sample without reaching the unchanged consecutive-sample stop threshold; this is not evidence of a wide host-memory margin.

Diagnostics and reporting a problem

Run doctor for the selected profile:

./r9v doctor qwen38-mtp4 --state-dir "$STATE_DIR"

Collect support evidence before removing a failed container, using the same state directory as setup:

SUPPORT_DIR=/path/to/private/r9v-support
./r9v support qwen38-mtp4 --state-dir "$STATE_DIR" \
  --output "$SUPPORT_DIR/run-001" \
  --archive "$SUPPORT_DIR/run-001.tar.gz"

Collection stays local and never uploads automatically. Configuration summaries hide credentials and personal paths, but raw application and kernel logs can contain identifying information or request content. Review the archive before attaching it to a GitHub issue. Include the profile, failed command, approximate failure time and whether the server, container or whole host stopped responding. Do not attach model weights or private prompts.

Doctor distinguishes configured settings from observed execution. Missing live metrics after a container stops are reported as unavailable; a missing kernel marker alone does not prove the wrong kernel ran. See the configuration reference for available controls and corrective actions.

Source and development

profiles/             model/runtime/hardware compositions and launch settings
packages/models/      pinned model sources, artifact hashes and licenses
packages/placements/  expert maps, memory seeds and placement manifests
runtimes/             runtime descriptors and retained kernel/source overlays
hardware/             GPU, RAM, PCIe and rank contracts
kernels/              pinned R9V kernel submodule
vendor/               pinned vLLM and GGUF-plugin forks
tools/                setup, planning, qualification, doctor and support
tests/                CPU checks and explicit GPU qualification tests

The kernels are specialized for supported shapes, quantizations and gfx1201. A different model, GPU, topology or runtime image requires its own validation. Current reference qualification covers one active sequence at a time.

Run the CPU and static checks with:

python -m pip install --requirement requirements-ci.txt
python -m pytest -q tests
./scripts/ci-static.sh

CPU CI checks tooling and source contracts. GPU parity, graph replay, full-model qualification and throughput measurements require the matching hardware. Both ordinary public setup/start/restart flows passed on the reference machine; a new machine or changed placement still requires local qualification.

Read CONVENTIONS.md before changing code. Dependency gitlinks are release inputs: use the committed revisions rather than replacing them with moving branch heads.

License and provenance

R9V-owned code and kernels use Apache-2.0. The vLLM and GGUF-plugin forks retain their licenses; llama.cpp/ggml-derived quantization primitives retain their MIT notices. Qwen model weights are separately governed by the Qwen Community License 1.0.

See THIRD_PARTY_NOTICES.md, licensing and the provenance audit. Runtime source publication, public image distribution and complete installation qualification are tracked separately; this README reports only the gates that have actually passed.

gguf
qwen
rdna4
rocm
speculative-decoding
vllm

Significant stargazers

silvanshade

37 followers · starred Sep 2026

Dyluhn/R9V

Shape-specialized ROCm inference for Qwen3.8 Flash Next on RDNA4

Rust

60

205 commits

updated Sep 25, 2026

See the code

See what people are saying

SourceMessageScoreDate

Strata vs R9V running Qwen3.8 Flash Next on dual Radeon R9700 with 96GB DDR4 (r/LocalLLM)

We're hearing a lot about Strata these days, claiming it's the best solution to run Qwen3.8 Flash Next on consumer hardware. I have a dual R9700 config with 96GB DDR4 3200 and AMD 3900X, and so far I've been running Flash Next UD-IQ4\_XS on the custom R9V engine (…

0

Oct 5, 2026

README

R9V

CI

R9V runs Qwen3.8 Flash Next on two AMD Radeon AI PRO R9700 GPUs. It combines a pinned vLLM fork, GGUF loading, specialized gfx1201 kernels, expert offloading and four-token MTP speculative decoding. The server exposes an OpenAI-compatible API for text, tool calls and single-image requests.

Each profile binds a model package, runtime, hardware layout and expert placement. Downloads are checked against pinned revisions and file hashes. Setup records the selected configuration, and first start qualifies its workload and memory headroom before reporting ready.

Current status: both IQ4_XS and Q4_K_XL MTP4 profiles passed ordinary public setup, first-start workload qualification and unchanged-receipt restart on the dual-R9700 reference host. Both profiles remain experimental. Each first start passed all seven checks at 131,072 context, including a 130,941-token prompt, with at least 3 GiB free VRAM per GPU. Setup selects the profile's runtime image bundle, verifies every part and loads the exact image ID. The Q4 profile uses the v0.2.0-rc2 bundle; the IQ4 profile's WMMA-prefill image (release/image-bundle-wmma-prefill-20260915.json) is published under the v0.3.0-rc1-images release tag. See release status and evidence.

New in v0.4.4: qwen38-mtp4-uncensored with the default --ced on now shares one VRAM region per GPU between the vision encoder's weights and the CED projector, as --ced quality does since v0.4.3. GPU 0 keeps 2.07 GiB free at its minimum in first-start qualification instead of 1.65 (GPU 1: 2.52 instead of 2.12), with the same prefill speedup, decode, outputs and image support. Start now needs 60.8 GiB of available RAM with CED on (56.3 GiB with --ced off). See the changelog.

New in v0.4.3: qwen38-mtp4-uncensored --ced quality now passes first-start qualification on the reference host, with image support kept: the vision encoder's weights and the quality projector take turns in one VRAM region per GPU, and the projector's work buffers are smaller. GPU 0 kept 2.04 GiB free (v0.4.2: 0.90, target 1.5), more than on (1.65). See the changelog.

New in v0.4.2: qwen38-mtp4-uncensored has an opt-in CED quality mode, --ced quality, that gives up less long-context quality for less prefill speedup (it did not fit on the reference host until v0.4.3). Right after a first start, ./r9v doctor --runtime no longer fails its KV-pressure check because of qualification's own long prompt.

New in v0.4.1: qwen38-mtp4-uncensored keeps rank 1's 400 most-used experts per layer in VRAM for good, so its host copy of the experts drops from 55.4 to 40.3 GiB and it needs 56.3 GiB of free RAM at start instead of 71.4. In the GPU test its outputs matched 1.3.0 bit for bit on the test prompts and decode speed was unchanged. ./r9v doctor also checks disk space, VRAM held by other programs, API exposure, the runtime's pinned files and the CED projector. See the changelog.

New in v0.4.0: qwen38-mtp4-uncensored runs an uncensored (abliterated) IQ4_XS model on the consolidated 1.3.0 runtime with CED long-prompt prefill on by default. Its public fetch, setup, first-start qualification and restart passed from a clean checkout on the reference host, with CED on and off. See its qualification note.

Profiles and features

AliasModel packageRuntimeStatus
qwen38-mtp4IQ4_XS model bundleMTP4, dual R9700, 128K contextReference setup/start/restart passed
qwen38-q4-xlQ4_K_XL weightsMTP4, dual R9700, 128K contextReference setup/start/restart passed
qwen38-mtp4-uncensoredUncensored IQ4_XS bundle (abliterated, no refusals)Consolidated 1.3.0 runtime with host expert dedupe: MTP4, dual R9700, 128K context, CED on; needs 60.8 GiB free RAM at start (56.3 GiB with --ced off)v0.4.1 clean install (fetch, setup, first start CED on/off, restart) passed on the reference host; v0.4.3 (--ced quality) and v0.4.4 (--ced on) clean installs passed there too (below)

About qwen38-mtp4-uncensored. Its model had its refusal behavior removed and will comply with harmful requests the original model refuses. Use it for research, and add your own moderation before exposing it to anyone. R9V serves it on 127.0.0.1 only; setting R9V_HOST_BIND to expose it gives anyone who can reach the port an unauthenticated model with no refusals.

It has CED (approximate long-prompt prefill) on by default. On prompts of 8,192 tokens or more, layers 0–15 run exactly and a split-16 projector stands in for the later layers on all but the last ~2K prompt tokens. Decode stays exact. The tradeoff, measured on the reference host:

  • About 1.5× faster prefill at ~13K tokens, rising to about 1.8× at 32K tokens and above.
  • About ×1.051 perplexity on prompts that depend on their long context.
  • About 10% fewer MTP tokens per step on the first answer after a CED prefill; later turns are normal.
  • The projector takes 1.76 GiB of VRAM per GPU, so the profile's free-VRAM target is 1.5 GiB per card instead of 3 GiB. Since v0.4.4 it shares that VRAM with the vision encoder's 0.42 GiB (below).

Turn CED off for the server with --ced off in setup or start; keep a single request exact with "vllm_xargs": {"r9v_ced": false}.

CED quality (opt-in, --ced quality). A multi-source projector reads the layer-16 state plus the inputs of full-attention layers 3, 7, 11 and 15. In one GPU grade of both projectors (both loaded as int8; the default ships bf16), on the 16 prompts that depend on their long context:

--cedPerplexityLong-context gain lostPrefill speedup (median)Projector VRAM per GPU
on (default)×1.04917%1.68×1.76 GiB (bf16), shared with the vision encoder's 0.42 GiB
quality×1.02910%1.55×1.79 GiB (stored int8), shared with the vision encoder's 0.42 GiB

The projector math costs 56 ms per 1K approximated tokens instead of 24 ms. Setup downloads the quality projector (1.8 GB) only when you pick it: run setup --ced quality once, then start keeps the choice. on stays the default.

Tested on the compiled server (v0.4.3, v0.4.4). The projector shares one VRAM region per GPU with the vision encoder's weights (quality since v0.4.3, on since v0.4.4): image prompts never use CED, so the region holds whichever the next step needs and is refilled from pinned host RAM (36 ms / 9 ms on GPU 0, 266 ms / 63 ms on GPU 1; only when a CED prompt follows an image or the other way round). In the clean-install GPU test, first start with --ced quality passed all seven checks with 2.04 / 2.49 GiB free (target 1.5; on: 1.65 / 2.12) while desktop apps held 1.23 GiB of GPU 0. Prefill was 1.47× at 12.8K tokens and 1.62× at 32K (on: 1.58× and 1.76×), decode unchanged, exact requests bitwise identical after a CED request, no recompiles; 24 alternating CED and image requests kept VRAM flat and read every image. Quality needs 4.50 GiB more free RAM at start (60.8 GiB) for the pinned copies. v0.4.2 did not fit: GPU 0 fell to 0.90 GiB free. In v0.4.4's clean-install test on kept 2.07 / 2.52 GiB free (v0.4.3: 1.65 / 2.12) with prefill 1.59× at 12.8K tokens and 1.77× at 32K, the vision encoder's output bitwise unchanged, and it now needs the same 60.8 GiB of free RAM at start. See the changelog. The profile uses a fixed expert placement, so it does not accept --headroom.

Use the explicit MTP4 aliases for the current workflow.

  • Per-card headroom: choose the free VRAM to retain on each card with --headroom 3,3 or an asymmetric budget such as --headroom 5,3.
  • Measured expert maps: separate complete catalogs for IQ4 and Q4 rank all 512 experts across 48 layers per rank, using training and separate held-out routing captures. Existing hot prefixes are preserved, while cold-to-hot arrays use measured counts, with deterministic ties for zero-count experts. The planner retains frequently used experts in VRAM within the memory budget.
  • Q4_K_XL support: a dedicated package, packed expert costs, ranked placement and streaming loader. Q4 uses its upstream model bytes and original Q8_0 target head; IQ4 retains its original Q6_K target head.
  • Resumable setup and restart receipts: reuse matching model assets, save configuration and reuse its verified receipt on an unchanged restart after the workflow has qualified.
  • Doctor: inspect model/runtime/placement compatibility, GPU ordering, memory pressure, PCIe information, cache identity and available worker evidence.
  • Support bundles: collect configuration summaries, source identities, worker records, memory and GPU diagnostics, logs and capture tails into a bounded local archive with file hashes.

Unobserved experts are explicit ties in the maps. Routing frequency depends on workload; a different prompt mix can change the best placement. Requested headroom is checked against the qualification workload, and cannot prevent another application from allocating VRAM later.

Model downloads

The ./r9v setup commands below download and verify the required files automatically. For manual downloads, keep every shard and the package directory layout; use the IQ4 package manifest or Q4 package manifest for exact paths, revisions and hashes. The PLE table is extracted locally from the target GGUF; it is not a separate model download.

Hardware and storage

The reference system uses:

  • Two 32 GiB Radeon AI PRO R9700 cards (gfx1201).
  • Linux with working AMD GPU drivers, amd-smi, /dev/kfd and /dev/dri access.
  • 128 GiB host RAM. Smaller hosts are untested; cold expert allocations use host memory. qwen38-mtp4-uncensored checks for 60.8 GiB available before start with CED on or quality, 56.3 GiB with --ced off (was 71.4 GiB in v0.4.0): its host expert copy is 40.3 GiB plus a 16 GiB PLE reserve, plus 4.50 GiB of pinned copies of the CED projector and the vision encoder.
  • An asymmetric PCIe layout: rank 0 on Gen5 x16 and rank 1 across Gen4 x4. GPU ordering matters to placement and performance.
  • Git, Python 3.10+, Docker and the Hugging Face CLI described in the installation guide.

The public image bundle plus its containerd image-store footprint measured roughly 50 GiB; reserve at least 70 GiB for image and cache import space. The IQ4 package occupies approximately 90.36 GiB and the uncensored package 92.39 GiB. The four Q4 target shards alone occupy 103.69 GiB, with auxiliary assets additional. The derived PLE file occupies 26.82 GiB. Leave further room for compilation caches and diagnostics. Reuse verified assets instead of duplicating model files.

Setup and start

Clone the source and its exact dependencies:

git clone --recursive https://github.com/Dyluhn/R9V.git
cd R9V
./r9v show qwen38-mtp4
./r9v doctor qwen38-mtp4 -- --host-only

The profile distribution entries select the matching GitHub Release image bundle and exact image IDs. Docker 29 must use the containerd image store so docker load preserves those IDs. See Docker's containerd image store guide and daemon configuration reference. Verify the store before setup:

docker info -f '{{ .DriverStatus }}'

The output should identify io.containerd.snapshotter.v1. If it does not, follow the configuration steps in the installation guide. For a non-root daemon, see Docker's rootless mode guide and confirm the selected context/socket with docker info; the same containerd image-store check applies. These are the reference identities:

ProfileTested local image ID
qwen38-mtp4sha256:2dac17a215fb5b0e3461e4c3e36a2981eec8ac3d6021e73183d247e819740c03
qwen38-q4-xlsha256:2e50016cfcc9cd22f15d3f69ccf001e4877236e12ebb4ab458cc9c16caaef9e3
qwen38-mtp4-uncensoredsha256:2dac17a215fb5b0e3461e4c3e36a2981eec8ac3d6021e73183d247e819740c03 plus pinned overlays

Install the download CLI in an isolated environment if it is not already available:

python3 -m venv ~/.local/share/r9v/download-tools
~/.local/share/r9v/download-tools/bin/pip install huggingface_hub
export PATH="$HOME/.local/share/r9v/download-tools/bin:$PATH"

For IQ4, choose persistent SSD directories with enough free space:

MODEL_DIR=/path/to/qwen-iq4
STATE_DIR=/path/to/r9v-state/iq4
./r9v setup qwen38-mtp4 \
  --model-dir "$MODEL_DIR" --state-dir "$STATE_DIR" \
  --headroom 3,3 --accept-model-license
./r9v start qwen38-mtp4 --state-dir "$STATE_DIR"

For Q4, use its own model and state directories:

./r9v setup qwen38-q4-xl --model-dir /path/to/qwen-q4 \
  --state-dir /path/to/r9v-state/q4 --headroom 3,3 --accept-model-license
./r9v start qwen38-q4-xl --state-dir /path/to/r9v-state/q4

For the uncensored profile, use its own directories. Its first start compiles the model, so start waits up to 2,400 seconds for it by default (--timeout changes that):

./r9v setup qwen38-mtp4-uncensored --model-dir /path/to/qwen-uncensored \
  --state-dir /path/to/r9v-state/uncensored --accept-model-license
./r9v start qwen38-mtp4-uncensored --state-dir /path/to/r9v-state/uncensored

Run one profile at a time. Before switching, inspect docker ps, save any needed support evidence, and stop the selected R9V container by its exact name using docker stop NAME. Setup supports --reuse-from /path/to/existing/assets for matching assets and --ple-path /path/to/existing/ple.bin for an existing verified PLE file. The profiles select their corresponding expert catalog and reference memory seed automatically.

First start runs the local workload qualification, including context and headroom checks. An unchanged restart can reuse its verified receipt. To change the budget, stop the selected container, then run:

./r9v start qwen38-mtp4 --state-dir "$STATE_DIR" --headroom 5,3

The new placement must qualify. Card order follows the saved GPU selection; inspect it with doctor. The planner reports per-card shortfalls when a request cannot fit, while retaining the configured context.

The default API endpoint is http://127.0.0.1:8004/v1. The API has no authentication, so R9V publishes it on 127.0.0.1 only, in every profile. To reach it from other machines, run setup with R9V_HOST_BIND=0.0.0.0 (or one interface's IPv4/IPv6 address); setup saves it with the port. See API address. Use the address recorded by your selected setup if you override the port. See the release guide, installation guide, and troubleshooting guide. Rebuilding an image does not reproduce a qualified image identity automatically.

Measured results

Prompt processing

Measured with the original release’s corpus, script, warmup and 8K/32K/64K procedure. Rates below are mean prompt tokens/s; every trial and exact reproduction command are linked below.

Profile8K (10 runs)32K (3 runs)64K (2 runs)Raw trials
IQ4_XS, WMMA prefill (current qwen38-mtp4)1,684.61,640.61,605.1JSON
IQ4_XS, v0.2.0 image989.3982.0968.2JSON
Q4_K_XL552.9542.5536.4JSON

The current IQ4 profile runs prompt chunks of 4096 tokens on a gfx12 int8 WMMA grouped MoE prefill kernel; decode kernels, placement policy and the model are unchanged, and the placement assets were re-derived and qualified for the new image. Why v0.2.0 regressed, kernel parity, method and commands; v0.2.0 method. Prefix-cache hits were zero. The older generation references below use different placements and are separate results.

Earlier fixed-prompt generation references

The following fixed-prompt reference samples used MTP4 on the dual-R9700 system:

Profile / placementStatic experts, ranks 0/1Generation tokens/s
IQ4 image7 streaming reference71 / 45089.45196
Q4 measured ranked placement97 / 34953.431
Q4 initial bootstrap placement64 / 32025.285

These are fixed-prompt reference samples with MTP4; they do not measure mixed traffic or generation at full context. Actual user placements depend on the requested headroom and must qualify locally.

The IQ4 image7 streaming reference retained 131,072 context tokens and passed seven bounded workload checks, including an actual 130,941-token prompt, text, tools, three image shapes and idle resume. Its median was 89.45196 TG tok/s, with measured free VRAM of 4,253,020,160 and 4,090,036,224 bytes and a minimum Normal-zone free value of 450,269,184 bytes. The run had a clean 90-second aftermath and GPU reclaim. The supervisor incorrectly reported failure because its cleanup check required exact VRAM equality: rank 0 had 185.203 MiB more free and rank 1 was unchanged. Independent review confirmed no per-card shortfall throughout the aftermath. These reference measurements are separate from the completed ordinary user-flow checks below and do not establish answer quality or a 100 tok/s qualification.

The current IQ4 reference evidence records the measured result and independently verified archive commitments. Historical prefill and comparator results remain in the earlier Qwen qualification; they should not be substituted for measurements of the new placements.

Verified public setup and restart

Both profiles passed ordinary setup into new state directories, first-start text/tool/vision/context/idle-resume checks and restart with the same verified receipt. The first starts used an actual 130,941-token prompt at a 131,072 context limit. Model, head and PLE bytes stayed unchanged.

ProfileStatic experts, ranks 0/1Cache slots, ranks 0/1Minimum free VRAM, ranks 0/1Evidence
IQ4 image776 / 451160 / 03,828,301,824 / 4,120,670,208 bytesUser-flow report
Q4 image699 / 34880 / 04,072,394,752 / 4,013,797,376 bytesUser-flow report

These checks used a fresh anonymous source checkout and new setup state, reusing previously public-downloaded, hash-verified assets and images. The separate anonymous image import recovered from an interrupted load using unchanged verified parts. This was not a second fresh download or a full new-machine download in one uninterrupted run. Both user flows stopped cleanly, reclaimed GPU allocations and completed at least 90 seconds of post-stop observation. Q4 recorded one low Normal-zone memory sample without reaching the unchanged consecutive-sample stop threshold; this is not evidence of a wide host-memory margin.

Diagnostics and reporting a problem

Run doctor for the selected profile:

./r9v doctor qwen38-mtp4 --state-dir "$STATE_DIR"

Collect support evidence before removing a failed container, using the same state directory as setup:

SUPPORT_DIR=/path/to/private/r9v-support
./r9v support qwen38-mtp4 --state-dir "$STATE_DIR" \
  --output "$SUPPORT_DIR/run-001" \
  --archive "$SUPPORT_DIR/run-001.tar.gz"

Collection stays local and never uploads automatically. Configuration summaries hide credentials and personal paths, but raw application and kernel logs can contain identifying information or request content. Review the archive before attaching it to a GitHub issue. Include the profile, failed command, approximate failure time and whether the server, container or whole host stopped responding. Do not attach model weights or private prompts.

Doctor distinguishes configured settings from observed execution. Missing live metrics after a container stops are reported as unavailable; a missing kernel marker alone does not prove the wrong kernel ran. See the configuration reference for available controls and corrective actions.

Source and development

profiles/             model/runtime/hardware compositions and launch settings
packages/models/      pinned model sources, artifact hashes and licenses
packages/placements/  expert maps, memory seeds and placement manifests
runtimes/             runtime descriptors and retained kernel/source overlays
hardware/             GPU, RAM, PCIe and rank contracts
kernels/              pinned R9V kernel submodule
vendor/               pinned vLLM and GGUF-plugin forks
tools/                setup, planning, qualification, doctor and support
tests/                CPU checks and explicit GPU qualification tests

The kernels are specialized for supported shapes, quantizations and gfx1201. A different model, GPU, topology or runtime image requires its own validation. Current reference qualification covers one active sequence at a time.

Run the CPU and static checks with:

python -m pip install --requirement requirements-ci.txt
python -m pytest -q tests
./scripts/ci-static.sh

CPU CI checks tooling and source contracts. GPU parity, graph replay, full-model qualification and throughput measurements require the matching hardware. Both ordinary public setup/start/restart flows passed on the reference machine; a new machine or changed placement still requires local qualification.

Read CONVENTIONS.md before changing code. Dependency gitlinks are release inputs: use the committed revisions rather than replacing them with moving branch heads.

License and provenance

R9V-owned code and kernels use Apache-2.0. The vLLM and GGUF-plugin forks retain their licenses; llama.cpp/ggml-derived quantization primitives retain their MIT notices. Qwen model weights are separately governed by the Qwen Community License 1.0.

See THIRD_PARTY_NOTICES.md, licensing and the provenance audit. Runtime source publication, public image distribution and complete installation qualification are tracked separately; this README reports only the gates that have actually passed.

gguf
qwen
rdna4
rocm
speculative-decoding
vllm

Significant stargazers

silvanshade

37 followers · starred Sep 2026