17
stars
23
commits
2
linked in READMEs
Jul 18, 2026
updated

Chadrock-35B Ace Saber is a local AMD Strix Halo release with two additive serving lanes: the original ROCmFP4/MTP speed lane, and a new ROCmFPX MoEQuality-7.07 BPW lane for quality-sensitive coding and agent work. The original lane remains here unchanged; the FPX file is added as a second downloadable GGUF.
This release also includes a tested Qwen3.6 vision projector, so the same full Chadrock language GGUF can run image-text prompts when launched with --mmproj.
The model behavior comes from the Ace Saber build by @DJLougen. The published measurements below use the Chadrock v2 ROCmFPX llama.cpp runner family from ciru-ai/ROCmFPX, with the exact ROCmFP4 and ROCmFPX runner references listed in the settings section.
These GGUFs will not run correctly with stock llama.cpp. Use the Chadrock v2 ROCmFPX runner because these files use ROCmFP4/ROCmFPX tensor types and MTP serving controls that upstream llama.cpp does not currently understand.
The model files are already provided here. You do not need to rebuild or quantize the model.
Ace Saber gives the model its coding, agentic, and tool-use behavior. Chadrock/ROCmFP4 gives it the speed profile needed to feel good locally on AMD unified-memory hardware.
The goal is not just another Qwen3.6 quant. The goal is:
Hugging Face may round the parsed GGUF tensor count to 36B in its automatic badge. This release is the Qwen3.6 35B-A3B MoE family: about 35B-class total parameters with roughly 3B active parameters per token.
| Field | Value |
|---|---|
| model size | 35B-A3B MoE |
| total parameters | 35B class |
| active parameters | ~3B class |
| architecture | qwen35moe |
| direct upstream GGUF | GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER-GGUF-MTP |
| base family | Qwen/Qwen3.6-35B-A3B |
| local runtime format | ROCmFP4 speed GGUF plus ROCmFPX MoEQuality GGUF, with separate GGUF-format vision projector |
Vision is provided by mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj, a GGUF-format Qwen3VL projector converted from the restored Qwen3.6 visual tower sidecar in the upstream Ace Saber release.
This does not replace the language model and does not disable MTP. The validated command used the full Chadrock ROCmFP4 GGUF with native --spec-type draft-mtp enabled and added the projector with --mmproj.
Local validation used two generated images whose answers were not present in the prompt:
| Image gate | Expected | Result |
|---|---|---|
gate_a.png | CIRU-742, red square, blue circle | passed |
gate_b.png | HALO-319, orange triangle, purple star | passed |
The same gate fails without a projector, so this is a real image-read check rather than a metadata-only claim.
| Lane | File | Best use | Notes |
|---|---|---|---|
| ROCmFP4 speed / vision lane | Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf | Fast local ACE/SABER serving with projector support | Original published lane; kept in place. |
| ROCmFPX MoEQuality lane | CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf | Higher-quality coding and agentic work | New additive quality lane, built from the H33 ACE/SABER MoEQuality recipe. |
| Vision projector | mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj | Image-text prompts | Same projector can be used with either language GGUF when memory allows. |
The new FPX lane was built as a quality-focused MoE recipe for the same ACE/SABER 35B-A3B MTP model. The full quantization log reports the 753-tensor qwen35moe topology at 29945.07 MiB (7.07 BPW).
Best local task-eval profile for this lane is Vulkan, one slot, prompt cache disabled, deterministic sampling, MTP on, f16 draft KV, and the current draft cap profile described below.
| Benchmark / row | Setting | Result | Runtime / speed |
|---|---|---|---|
| HermesAgent-20 official | ROCmFPX MoEQuality, Vulkan, 64k, q8_0 target KV, f16 draft KV, draft-MTP n3 | 0.82 avg score | 506.60 s |
| HermesAgent-20 category scores | memory / workspace / skills / scheduling / delegation | 1.00 / 0.80 / 0.80 / 0.75 / 0.73 | same run |
| HumanEval / EvalPlus | ROCmFPX MoEQuality cap2 256k profile, deterministic | 155/164 = 94.51% base, 148/164 = 90.24% plus | 672.05 s, 71.31 tok/s generation |
Useful local comparisons:
| Comparison | Result |
|---|---|
| Stock Unsloth Qwen3.6 35B Q6 K XL, vanilla Vulkan, HermesAgent-20 draftKV16 | 0.71 avg score, so the FPX MoEQuality lane is +0.11 absolute on this agent bench. |
| Original ROCmFP4 ACE/SABER 32k HumanEval guard | 155/164 base and 148/164 plus, matching the new FPX HumanEval row. |
| Stock Unsloth Qwen3.6 35B Q6 K XL HumanEval reference | 157/164 base and 150/164 plus; FPX is within two tasks while generating slightly faster in the recorded EvalPlus row (71.31 vs 68.84 tok/s). |
The tuned Vulkan cap2 profile uses f16 target KV, f16 draft KV, speculative.n_max=2, p_min=0.25, p_split=0.10, gen512, and prompt cache disabled:
| Context | PP tok/s | TG tok/s | TTFP | Total time | Draft accepted |
|---|---|---|---|---|---|
2k | 819.99 | 71.37 | 2.84 s | 10.02 s | 290/412 = 70.4% |
4k | 770.17 | 69.44 | 5.33 s | 12.71 s | 292/416 = 70.2% |
8k | 771.23 | 69.55 | 10.64 s | 18.00 s | 294/415 = 70.8% |
16k | 737.70 | 69.21 | 22.22 s | 29.62 s | 302/405 = 74.6% |
Matched 8k/gen512 comparison against stock Unsloth Qwen3.6 35B Q6 K XL, same prompt, Vulkan, f16 target/draft KV, MTP on, n_max=2, and deterministic sampling:
| Row | PP tok/s | TG tok/s | TTFP | Total time | Draft accepted |
|---|---|---|---|---|---|
| ROCmFPX MoEQuality 7.07 BPW | 771.23 | 69.55 | 10.64 s | 18.00 s | 294/415 = 70.8% |
| Stock Unsloth Q6 K XL | 825.35 | 64.39 | 9.93 s | 17.89 s | 287/434 = 66.1% |
On that matched row, the FPX quality lane is 8.0% faster on decode (69.55 vs 64.39 tok/s), while Q6 keeps a prefill and TTFP edge and the short 512-token wall time is nearly tied.
The first backend sweep for the FPX lane used cap4, n_max=4, and f16 target/draft KV. Vulkan was the better serving path for decode and draft acceptance:
| Context | Vulkan TG | ROCm TG | Vulkan accept | ROCm accept |
|---|---|---|---|---|
2k | 95.82 | 33.25 | 402/428 = 93.9% | 214/987 = 21.7% |
4k | 70.90 | 34.58 | 358/567 = 63.1% | 246/872 = 28.2% |
8k | 56.87 | 41.45 | 325/670 = 48.5% | 287/765 = 37.5% |
16k | 59.66 | 37.86 | 342/644 = 53.1% | 278/820 = 33.9% |
Best current text-only profile on AMD Ryzen AI Max+ 395 / Strix Halo: pinned Chadrock v2 ROCmFPX llama.cpp, Vulkan0 target plus Vulkan0 draft, f16/f16 target and draft KV, one slot, prompt cache disabled, no multimodal projector, deterministic decoding, and request policy speculative.n_max=4, speculative.n_min=0, speculative.p_min=0.25.
| Measurement | Prompt tokens | Generated tokens | Decode tok/s | Prefill tok/s | Total time | Draft accepted |
|---|---|---|---|---|---|---|
| Chadrock v2 best, gen512 | 3,946 | 512 | 143.08 | 1072.34 | 7.26 s | 408 / 408 |
| Chadrock v2 repeat, gen2048 | 3,946 | 2048 | 141.77 | 1064.16 | 18.16 s | 1637 / 1637 |
| Same-run no-draft control, gen512 | 3,946 | 512 | 72.57 | 1064.49 | 10.77 s | 0 / 0 |
| Same-run no-draft control, gen2048 | 3,946 | 2048 | 72.04 | 1067.18 | 32.13 s | 0 / 0 |
Against the same-run no-draft control, the new Chadrock v2 MTP config is 1.97x faster in decode at both gen512 and gen2048: 143.08 vs 72.57 tok/s, and 141.77 vs 72.04 tok/s.
Compared with the older Chadrock v1 card-speed profile (~101.31 tok/s aggregate HumanEval eval speed), the new best served text row is about 1.41x faster (+41.2%). Compared with the older uncached Chadrock TG64 served row (78.28 tok/s server eval), the new gen2048 row is about 1.81x faster (+81.1%). Those older rows are not identical prompts, so treat them as release-to-release runner evidence rather than a strict apples-to-apples benchmark pair.
The older 2026-06-07 HumanEval rerun is still useful as a quality guard: it generated 48,824 completion tokens across 164 tasks at ~101.31 tok/s aggregate llama-server eval speed and produced the 155/164 base pass@1 and 148/164 HumanEval+ pass@1 result below.
This model also posts an exceptional HumanEval result for a local GGUF run:
| Model / row | HumanEval base pass@1 | HumanEval+ pass@1 |
|---|---|---|
| Chadrock-35B Ace Saber ROCmFP4, 32k Vulkan d2 rerun | 155/164 = 94.51% | 148/164 = 90.24% |
| earlier Chadrock-35B Ace Saber ROCmFP4 run | 157/164 = 95.73% | 149/164 = 90.85% |
| recorded stock Qwen3.6-27B UD-Q8_K_XL | 154/164 = 93.90% | 149/164 = 90.85% |
The fresh 32k rerun still beats the recorded stock 27B row on base HumanEval, while the older row remains one task higher on HumanEval+.
The same tuned Chadrock Vulkan d2 family was also run on BigCodeBench-Hard-Instruct:
| Benchmark | Result |
|---|---|
| BigCodeBench-Hard-Instruct pass@1 | 47/148 = 31.76% |
| generation wall time | 799 s |
| aggregate prompt speed | ~624.06 tok/s |
| aggregate generation speed | ~100.12 tok/s |
This is a harder instruction-coding benchmark than HumanEval and is included as a sanity check that the speed-tuned runtime still produces usable code under a broader task mix.
For runner build commands, request-level speculative controls, and the 35B/27B reproduction notes, use the advanced Ciru setup page:
https://llm.ciru.ai/chadrock-rocmfpx/
Runner references used for the published rows:
ROCmFP4 speed lane: ciru-ai/ROCmFPX commit 7aa484a2f0a504dc612a3d74a068024f3e6d6353
ROCmFPX MoEQuality lane: ciru-ai/ROCmFPX branch ciru-ai/fp6-vulkan-dequant-speed, commit fd6555ffb36dcef6419757189012c8c4fd31f91f
Recommended build for current serving: ciru-ai/ROCmFPX main commit 647fd52965a401d7fa8e035fc76a16c94216a794.
For the fastest measured text-only ACE/SABER path on Strix Halo, use:
backend: Vulkan0 target + Vulkan0 draft
context: 32768
batch / ubatch: 2048 / 512
target KV: f16 / f16
draft KV: f16 / f16
MTP: draft-mtp, n_max=4, n_min=0, p_min=0.25, p_split=0.10
serving: one slot, prompt cache disabled for benchmarks, --no-mmproj for text speed
sampler: temperature=0, top_p=0.95, top_k=20
That is the profile that produced 143.08 tok/s at gen512 and repeated at
141.77 tok/s at gen2048 on the 3946-token text prompt. For image-text use,
remove --no-mmproj and add --mmproj mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj;
the headline speed row above is text-only.
For the new FPX quality lane, the current general Strix profile is:
backend: Vulkan0 target + Vulkan0 draft
context: 262144
batch / ubatch: 2048 / 512
target KV: f16 / f16
draft KV: f16 / f16
MTP: draft-mtp, n_max=4 for vision/general use; n_max=2 for the tuned speed row
serving: one slot, prompt cache disabled for benchmarks
sampler: temperature=0, top_p=0.95, top_k=20
The HermesAgent-20 quality row used the text-only 64k task-eval variant: q8_0/q8_0 target KV, f16/f16 draft KV, no projector, and n_max=3.
Build the Chadrock v2 ROCmFPX llama.cpp runner once, download the GGUF lane you want, then run from the ROCmFPX checkout.
For the new ROCmFPX MoEQuality lane:
./build-strix-rocmfp4/bin/llama-server \
-m /path/to/CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf \
--alias chadrock-35b-ace-saber-rocmfpx-moequality-707bpw-cap4 \
--host 127.0.0.1 \
--port 18180 \
--jinja \
-c 262144 \
--reasoning off \
--reasoning-format none \
--reasoning-budget -1 \
--no-context-shift \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-tb 32 \
-ctk f16 \
-ctv f16 \
--temp 0 \
--top-p 0.95 \
--top-k 20 \
--seed 123 \
--parallel 1 \
--mmproj /path/to/mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj \
--metrics \
--no-webui \
--cache-ram 8192 \
--ctx-checkpoints 32 \
--checkpoint-every-n-tokens 8192 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-threads 16 \
--spec-draft-threads-batch 32 \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 4 \
--spec-draft-n-min 0 \
--spec-draft-p-min 0.25 \
--spec-draft-p-split 0.10 \
--no-spec-draft-backend-sampling \
--spec-draft-poll 1 \
--spec-draft-poll-batch 1
For the tuned text-only speed row, use the same FPX command with --no-mmproj, --spec-draft-n-max 2, and the same f16 target/draft KV shown above. For the HermesAgent-20 quality row, use -c 65536, --no-mmproj, -ctk q8_0, -ctv q8_0, and --spec-draft-n-max 3.
For the original ROCmFP4 speed lane:
./build-strix-rocmfp4/bin/llama-server \
-m /path/to/Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf \
--alias chadrock-35b-ace-saber-rocmfp4-cap4 \
--host 127.0.0.1 \
--port 18180 \
--jinja \
-c 32768 \
--reasoning off \
--reasoning-format none \
--reasoning-budget -1 \
--no-context-shift \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-tb 32 \
-ctk f16 \
-ctv f16 \
--temp 0 \
--top-p 0.95 \
--top-k 20 \
--seed 123 \
--parallel 1 \
--no-mmproj \
--metrics \
--no-webui \
--cache-ram 8192 \
--ctx-checkpoints 0 \
--checkpoint-every-n-tokens -1 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-threads 16 \
--spec-draft-threads-batch 32 \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 4 \
--spec-draft-n-min 0 \
--spec-draft-p-min 0.25 \
--spec-draft-p-split 0.10 \
--no-spec-draft-backend-sampling \
--spec-draft-poll 1 \
--spec-draft-poll-batch 1
Use --parallel 1 for MTP. Multi-slot serving changes the MTP behavior and is not the intended profile.
The benchmark table above used -c 32768 and --no-mmproj for the fastest text row. For image-text use, remove --no-mmproj and add the projector:
--mmproj /path/to/mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj
For general local coding and agent work, you can raise context after validating memory headroom on your machine. The advanced page below keeps the copy-paste build and launch blocks current:
https://llm.ciru.ai/chadrock-rocmfpx/
The source checkpoint is GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER, based on Qwen/Qwen3.6-35B-A3B. The direct GGUF-MTP source is GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER-GGUF-MTP.
Training order:
Qwen3.6-35B-A3B -> NSC-ACE -> SABER
NSC-ACE uses multiple steered rollouts from the same model and rewards convergence across latent behavior modes, especially for tool-call structure, reasoning wrappers, self-consistency, and avoiding repeated loops.
SABER is the final calibration pass. The source model card reports 98.33% HarmBench-300 compliance and final KLD 0.025383937664711.
Chadrock v2 uses the Ciru ROCmFPX llama.cpp branches carrying the ROCmFP4 and ROCmFPX tensor/runtime work plus request-level MTP serving controls.
ROCmFP4 is not stock Q4, MXFP4, or NVFP4. It uses custom Codebook10 4-bit weights, finite unsigned E4M3 scale semantics, tensor-aware presets, ROCm/HIP kernels, Vulkan shader support, and MTP regression guards.
ROCmFPX is the higher-bit experimental family used for the new MoEQuality lane. In this release it is applied as a tensor-aware 7.07 BPW quality recipe for the 35B-A3B MoE topology.
Why it matters: Ryzen AI Max+ 395 / Strix Halo has a large unified-memory pool, but decode speed still depends heavily on bandwidth, tensor layout, and draft-token acceptance. ROCmFP4 is designed to make this class of AMD machine fast enough for serious local long-context use, while the FPX lane gives a quality-first option on the same card.
The GGUF files are already provided. You only need to build the custom llama.cpp server once:
git clone https://github.com/ciru-ai/ROCmFPX.git
cd ROCmFPX
git checkout 647fd52965a401d7fa8e035fc76a16c94216a794
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh llama-server llama-bench
The 262K serving profile keeps 32 context checkpoints at 8192-token intervals so long prompts can reuse cached work across the configured context. The pinned runner includes the MTP checkpoint-storage and restore-safety fixes required for this profile.
The server binary will be here:
build-strix-rocmfp4/bin/llama-server
| File | SHA256 |
|---|---|
Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf | 6a635d1d8ac4af8f2c4ca6ff528bc6bad9b3a6d45e8630ef6e5728f04898eeed |
CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf | 1d6a5199dec9c38c103951acf477235f90b6fd26fd9dacd27f24ed6a4d89c3ab |
mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj | 1365c90e3f35cad9c33e09e67ff377af083631c19718ec4c22d251a54c24c6a7 |
ciru-ai/ROCmFPX: Chadrock v2 ROCmFPX llama.cpp runner, ROCmFP4 and ROCmFPX tensor/runtime work, build path, request-level MTP controls, and public reproduction page.Qwen3.6-35B-A3B model.This is an experimental AMD ROCmFP4/ROCmFPX/MTP release. Performance depends on driver version, clocks, prompt shape, MTP acceptance, and serving flags. The numbers above are local reproducible measurements, not universal llama.cpp claims.
23 commits
17
stars
23
commits
2
linked in READMEs
Jul 18, 2026
updated

Chadrock-35B Ace Saber is a local AMD Strix Halo release with two additive serving lanes: the original ROCmFP4/MTP speed lane, and a new ROCmFPX MoEQuality-7.07 BPW lane for quality-sensitive coding and agent work. The original lane remains here unchanged; the FPX file is added as a second downloadable GGUF.
This release also includes a tested Qwen3.6 vision projector, so the same full Chadrock language GGUF can run image-text prompts when launched with --mmproj.
The model behavior comes from the Ace Saber build by @DJLougen. The published measurements below use the Chadrock v2 ROCmFPX llama.cpp runner family from ciru-ai/ROCmFPX, with the exact ROCmFP4 and ROCmFPX runner references listed in the settings section.
These GGUFs will not run correctly with stock llama.cpp. Use the Chadrock v2 ROCmFPX runner because these files use ROCmFP4/ROCmFPX tensor types and MTP serving controls that upstream llama.cpp does not currently understand.
The model files are already provided here. You do not need to rebuild or quantize the model.
Ace Saber gives the model its coding, agentic, and tool-use behavior. Chadrock/ROCmFP4 gives it the speed profile needed to feel good locally on AMD unified-memory hardware.
The goal is not just another Qwen3.6 quant. The goal is:
Hugging Face may round the parsed GGUF tensor count to 36B in its automatic badge. This release is the Qwen3.6 35B-A3B MoE family: about 35B-class total parameters with roughly 3B active parameters per token.
| Field | Value |
|---|---|
| model size | 35B-A3B MoE |
| total parameters | 35B class |
| active parameters | ~3B class |
| architecture | qwen35moe |
| direct upstream GGUF | GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER-GGUF-MTP |
| base family | Qwen/Qwen3.6-35B-A3B |
| local runtime format | ROCmFP4 speed GGUF plus ROCmFPX MoEQuality GGUF, with separate GGUF-format vision projector |
Vision is provided by mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj, a GGUF-format Qwen3VL projector converted from the restored Qwen3.6 visual tower sidecar in the upstream Ace Saber release.
This does not replace the language model and does not disable MTP. The validated command used the full Chadrock ROCmFP4 GGUF with native --spec-type draft-mtp enabled and added the projector with --mmproj.
Local validation used two generated images whose answers were not present in the prompt:
| Image gate | Expected | Result |
|---|---|---|
gate_a.png | CIRU-742, red square, blue circle | passed |
gate_b.png | HALO-319, orange triangle, purple star | passed |
The same gate fails without a projector, so this is a real image-read check rather than a metadata-only claim.
| Lane | File | Best use | Notes |
|---|---|---|---|
| ROCmFP4 speed / vision lane | Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf | Fast local ACE/SABER serving with projector support | Original published lane; kept in place. |
| ROCmFPX MoEQuality lane | CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf | Higher-quality coding and agentic work | New additive quality lane, built from the H33 ACE/SABER MoEQuality recipe. |
| Vision projector | mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj | Image-text prompts | Same projector can be used with either language GGUF when memory allows. |
The new FPX lane was built as a quality-focused MoE recipe for the same ACE/SABER 35B-A3B MTP model. The full quantization log reports the 753-tensor qwen35moe topology at 29945.07 MiB (7.07 BPW).
Best local task-eval profile for this lane is Vulkan, one slot, prompt cache disabled, deterministic sampling, MTP on, f16 draft KV, and the current draft cap profile described below.
| Benchmark / row | Setting | Result | Runtime / speed |
|---|---|---|---|
| HermesAgent-20 official | ROCmFPX MoEQuality, Vulkan, 64k, q8_0 target KV, f16 draft KV, draft-MTP n3 | 0.82 avg score | 506.60 s |
| HermesAgent-20 category scores | memory / workspace / skills / scheduling / delegation | 1.00 / 0.80 / 0.80 / 0.75 / 0.73 | same run |
| HumanEval / EvalPlus | ROCmFPX MoEQuality cap2 256k profile, deterministic | 155/164 = 94.51% base, 148/164 = 90.24% plus | 672.05 s, 71.31 tok/s generation |
Useful local comparisons:
| Comparison | Result |
|---|---|
| Stock Unsloth Qwen3.6 35B Q6 K XL, vanilla Vulkan, HermesAgent-20 draftKV16 | 0.71 avg score, so the FPX MoEQuality lane is +0.11 absolute on this agent bench. |
| Original ROCmFP4 ACE/SABER 32k HumanEval guard | 155/164 base and 148/164 plus, matching the new FPX HumanEval row. |
| Stock Unsloth Qwen3.6 35B Q6 K XL HumanEval reference | 157/164 base and 150/164 plus; FPX is within two tasks while generating slightly faster in the recorded EvalPlus row (71.31 vs 68.84 tok/s). |
The tuned Vulkan cap2 profile uses f16 target KV, f16 draft KV, speculative.n_max=2, p_min=0.25, p_split=0.10, gen512, and prompt cache disabled:
| Context | PP tok/s | TG tok/s | TTFP | Total time | Draft accepted |
|---|---|---|---|---|---|
2k | 819.99 | 71.37 | 2.84 s | 10.02 s | 290/412 = 70.4% |
4k | 770.17 | 69.44 | 5.33 s | 12.71 s | 292/416 = 70.2% |
8k | 771.23 | 69.55 | 10.64 s | 18.00 s | 294/415 = 70.8% |
16k | 737.70 | 69.21 | 22.22 s | 29.62 s | 302/405 = 74.6% |
Matched 8k/gen512 comparison against stock Unsloth Qwen3.6 35B Q6 K XL, same prompt, Vulkan, f16 target/draft KV, MTP on, n_max=2, and deterministic sampling:
| Row | PP tok/s | TG tok/s | TTFP | Total time | Draft accepted |
|---|---|---|---|---|---|
| ROCmFPX MoEQuality 7.07 BPW | 771.23 | 69.55 | 10.64 s | 18.00 s | 294/415 = 70.8% |
| Stock Unsloth Q6 K XL | 825.35 | 64.39 | 9.93 s | 17.89 s | 287/434 = 66.1% |
On that matched row, the FPX quality lane is 8.0% faster on decode (69.55 vs 64.39 tok/s), while Q6 keeps a prefill and TTFP edge and the short 512-token wall time is nearly tied.
The first backend sweep for the FPX lane used cap4, n_max=4, and f16 target/draft KV. Vulkan was the better serving path for decode and draft acceptance:
| Context | Vulkan TG | ROCm TG | Vulkan accept | ROCm accept |
|---|---|---|---|---|
2k | 95.82 | 33.25 | 402/428 = 93.9% | 214/987 = 21.7% |
4k | 70.90 | 34.58 | 358/567 = 63.1% | 246/872 = 28.2% |
8k | 56.87 | 41.45 | 325/670 = 48.5% | 287/765 = 37.5% |
16k | 59.66 | 37.86 | 342/644 = 53.1% | 278/820 = 33.9% |
Best current text-only profile on AMD Ryzen AI Max+ 395 / Strix Halo: pinned Chadrock v2 ROCmFPX llama.cpp, Vulkan0 target plus Vulkan0 draft, f16/f16 target and draft KV, one slot, prompt cache disabled, no multimodal projector, deterministic decoding, and request policy speculative.n_max=4, speculative.n_min=0, speculative.p_min=0.25.
| Measurement | Prompt tokens | Generated tokens | Decode tok/s | Prefill tok/s | Total time | Draft accepted |
|---|---|---|---|---|---|---|
| Chadrock v2 best, gen512 | 3,946 | 512 | 143.08 | 1072.34 | 7.26 s | 408 / 408 |
| Chadrock v2 repeat, gen2048 | 3,946 | 2048 | 141.77 | 1064.16 | 18.16 s | 1637 / 1637 |
| Same-run no-draft control, gen512 | 3,946 | 512 | 72.57 | 1064.49 | 10.77 s | 0 / 0 |
| Same-run no-draft control, gen2048 | 3,946 | 2048 | 72.04 | 1067.18 | 32.13 s | 0 / 0 |
Against the same-run no-draft control, the new Chadrock v2 MTP config is 1.97x faster in decode at both gen512 and gen2048: 143.08 vs 72.57 tok/s, and 141.77 vs 72.04 tok/s.
Compared with the older Chadrock v1 card-speed profile (~101.31 tok/s aggregate HumanEval eval speed), the new best served text row is about 1.41x faster (+41.2%). Compared with the older uncached Chadrock TG64 served row (78.28 tok/s server eval), the new gen2048 row is about 1.81x faster (+81.1%). Those older rows are not identical prompts, so treat them as release-to-release runner evidence rather than a strict apples-to-apples benchmark pair.
The older 2026-06-07 HumanEval rerun is still useful as a quality guard: it generated 48,824 completion tokens across 164 tasks at ~101.31 tok/s aggregate llama-server eval speed and produced the 155/164 base pass@1 and 148/164 HumanEval+ pass@1 result below.
This model also posts an exceptional HumanEval result for a local GGUF run:
| Model / row | HumanEval base pass@1 | HumanEval+ pass@1 |
|---|---|---|
| Chadrock-35B Ace Saber ROCmFP4, 32k Vulkan d2 rerun | 155/164 = 94.51% | 148/164 = 90.24% |
| earlier Chadrock-35B Ace Saber ROCmFP4 run | 157/164 = 95.73% | 149/164 = 90.85% |
| recorded stock Qwen3.6-27B UD-Q8_K_XL | 154/164 = 93.90% | 149/164 = 90.85% |
The fresh 32k rerun still beats the recorded stock 27B row on base HumanEval, while the older row remains one task higher on HumanEval+.
The same tuned Chadrock Vulkan d2 family was also run on BigCodeBench-Hard-Instruct:
| Benchmark | Result |
|---|---|
| BigCodeBench-Hard-Instruct pass@1 | 47/148 = 31.76% |
| generation wall time | 799 s |
| aggregate prompt speed | ~624.06 tok/s |
| aggregate generation speed | ~100.12 tok/s |
This is a harder instruction-coding benchmark than HumanEval and is included as a sanity check that the speed-tuned runtime still produces usable code under a broader task mix.
For runner build commands, request-level speculative controls, and the 35B/27B reproduction notes, use the advanced Ciru setup page:
https://llm.ciru.ai/chadrock-rocmfpx/
Runner references used for the published rows:
ROCmFP4 speed lane: ciru-ai/ROCmFPX commit 7aa484a2f0a504dc612a3d74a068024f3e6d6353
ROCmFPX MoEQuality lane: ciru-ai/ROCmFPX branch ciru-ai/fp6-vulkan-dequant-speed, commit fd6555ffb36dcef6419757189012c8c4fd31f91f
Recommended build for current serving: ciru-ai/ROCmFPX main commit 647fd52965a401d7fa8e035fc76a16c94216a794.
For the fastest measured text-only ACE/SABER path on Strix Halo, use:
backend: Vulkan0 target + Vulkan0 draft
context: 32768
batch / ubatch: 2048 / 512
target KV: f16 / f16
draft KV: f16 / f16
MTP: draft-mtp, n_max=4, n_min=0, p_min=0.25, p_split=0.10
serving: one slot, prompt cache disabled for benchmarks, --no-mmproj for text speed
sampler: temperature=0, top_p=0.95, top_k=20
That is the profile that produced 143.08 tok/s at gen512 and repeated at
141.77 tok/s at gen2048 on the 3946-token text prompt. For image-text use,
remove --no-mmproj and add --mmproj mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj;
the headline speed row above is text-only.
For the new FPX quality lane, the current general Strix profile is:
backend: Vulkan0 target + Vulkan0 draft
context: 262144
batch / ubatch: 2048 / 512
target KV: f16 / f16
draft KV: f16 / f16
MTP: draft-mtp, n_max=4 for vision/general use; n_max=2 for the tuned speed row
serving: one slot, prompt cache disabled for benchmarks
sampler: temperature=0, top_p=0.95, top_k=20
The HermesAgent-20 quality row used the text-only 64k task-eval variant: q8_0/q8_0 target KV, f16/f16 draft KV, no projector, and n_max=3.
Build the Chadrock v2 ROCmFPX llama.cpp runner once, download the GGUF lane you want, then run from the ROCmFPX checkout.
For the new ROCmFPX MoEQuality lane:
./build-strix-rocmfp4/bin/llama-server \
-m /path/to/CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf \
--alias chadrock-35b-ace-saber-rocmfpx-moequality-707bpw-cap4 \
--host 127.0.0.1 \
--port 18180 \
--jinja \
-c 262144 \
--reasoning off \
--reasoning-format none \
--reasoning-budget -1 \
--no-context-shift \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-tb 32 \
-ctk f16 \
-ctv f16 \
--temp 0 \
--top-p 0.95 \
--top-k 20 \
--seed 123 \
--parallel 1 \
--mmproj /path/to/mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj \
--metrics \
--no-webui \
--cache-ram 8192 \
--ctx-checkpoints 32 \
--checkpoint-every-n-tokens 8192 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-threads 16 \
--spec-draft-threads-batch 32 \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 4 \
--spec-draft-n-min 0 \
--spec-draft-p-min 0.25 \
--spec-draft-p-split 0.10 \
--no-spec-draft-backend-sampling \
--spec-draft-poll 1 \
--spec-draft-poll-batch 1
For the tuned text-only speed row, use the same FPX command with --no-mmproj, --spec-draft-n-max 2, and the same f16 target/draft KV shown above. For the HermesAgent-20 quality row, use -c 65536, --no-mmproj, -ctk q8_0, -ctv q8_0, and --spec-draft-n-max 3.
For the original ROCmFP4 speed lane:
./build-strix-rocmfp4/bin/llama-server \
-m /path/to/Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf \
--alias chadrock-35b-ace-saber-rocmfp4-cap4 \
--host 127.0.0.1 \
--port 18180 \
--jinja \
-c 32768 \
--reasoning off \
--reasoning-format none \
--reasoning-budget -1 \
--no-context-shift \
-dev Vulkan0 \
-ngl 999 \
-fa on \
-b 2048 \
-ub 512 \
-t 16 \
-tb 32 \
-ctk f16 \
-ctv f16 \
--temp 0 \
--top-p 0.95 \
--top-k 20 \
--seed 123 \
--parallel 1 \
--no-mmproj \
--metrics \
--no-webui \
--cache-ram 8192 \
--ctx-checkpoints 0 \
--checkpoint-every-n-tokens -1 \
--spec-type draft-mtp \
--spec-draft-device Vulkan0 \
--spec-draft-ngl all \
--spec-draft-threads 16 \
--spec-draft-threads-batch 32 \
--spec-draft-type-k f16 \
--spec-draft-type-v f16 \
--spec-draft-n-max 4 \
--spec-draft-n-min 0 \
--spec-draft-p-min 0.25 \
--spec-draft-p-split 0.10 \
--no-spec-draft-backend-sampling \
--spec-draft-poll 1 \
--spec-draft-poll-batch 1
Use --parallel 1 for MTP. Multi-slot serving changes the MTP behavior and is not the intended profile.
The benchmark table above used -c 32768 and --no-mmproj for the fastest text row. For image-text use, remove --no-mmproj and add the projector:
--mmproj /path/to/mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj
For general local coding and agent work, you can raise context after validating memory headroom on your machine. The advanced page below keeps the copy-paste build and launch blocks current:
https://llm.ciru.ai/chadrock-rocmfpx/
The source checkpoint is GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER, based on Qwen/Qwen3.6-35B-A3B. The direct GGUF-MTP source is GestaltLabs/Qwen3.6-35B-A3B-NSC-ACE-SABER-GGUF-MTP.
Training order:
Qwen3.6-35B-A3B -> NSC-ACE -> SABER
NSC-ACE uses multiple steered rollouts from the same model and rewards convergence across latent behavior modes, especially for tool-call structure, reasoning wrappers, self-consistency, and avoiding repeated loops.
SABER is the final calibration pass. The source model card reports 98.33% HarmBench-300 compliance and final KLD 0.025383937664711.
Chadrock v2 uses the Ciru ROCmFPX llama.cpp branches carrying the ROCmFP4 and ROCmFPX tensor/runtime work plus request-level MTP serving controls.
ROCmFP4 is not stock Q4, MXFP4, or NVFP4. It uses custom Codebook10 4-bit weights, finite unsigned E4M3 scale semantics, tensor-aware presets, ROCm/HIP kernels, Vulkan shader support, and MTP regression guards.
ROCmFPX is the higher-bit experimental family used for the new MoEQuality lane. In this release it is applied as a tensor-aware 7.07 BPW quality recipe for the 35B-A3B MoE topology.
Why it matters: Ryzen AI Max+ 395 / Strix Halo has a large unified-memory pool, but decode speed still depends heavily on bandwidth, tensor layout, and draft-token acceptance. ROCmFP4 is designed to make this class of AMD machine fast enough for serious local long-context use, while the FPX lane gives a quality-first option on the same card.
The GGUF files are already provided. You only need to build the custom llama.cpp server once:
git clone https://github.com/ciru-ai/ROCmFPX.git
cd ROCmFPX
git checkout 647fd52965a401d7fa8e035fc76a16c94216a794
env JOBS=16 scripts/build-strix-rocmfp4-mtp.sh llama-server llama-bench
The 262K serving profile keeps 32 context checkpoints at 8192-token intervals so long prompts can reuse cached work across the configured context. The pinned runner includes the MTP checkpoint-storage and restore-safety fixes required for this profile.
The server binary will be here:
build-strix-rocmfp4/bin/llama-server
| File | SHA256 |
|---|---|
Qwen3.6-35B-A3B-NSC-ACE-SABER-MTP-F16-to-ROCmFP4-STRIX_LEAN.gguf | 6a635d1d8ac4af8f2c4ca6ff528bc6bad9b3a6d45e8630ef6e5728f04898eeed |
CHADROCK-35B-Ace-Saber-MTP-ROCmFPX-MoEQuality-7.07BPW.gguf | 1d6a5199dec9c38c103951acf477235f90b6fd26fd9dacd27f24ed6a4d89c3ab |
mmproj-CHADROCK-35B-Ace-Saber-F32.mmproj | 1365c90e3f35cad9c33e09e67ff377af083631c19718ec4c22d251a54c24c6a7 |
ciru-ai/ROCmFPX: Chadrock v2 ROCmFPX llama.cpp runner, ROCmFP4 and ROCmFPX tensor/runtime work, build path, request-level MTP controls, and public reproduction page.Qwen3.6-35B-A3B model.This is an experimental AMD ROCmFP4/ROCmFPX/MTP release. Performance depends on driver version, clocks, prompt shape, MTP acceptance, and serving flags. The numbers above are local reproducible measurements, not universal llama.cpp claims.
23 commits