ROCmFP4 quantization of Ornith 1.5 35B-A3B for AMD Strix Halo (gfx1151) and RDNA 3.5 GPUs, produced with the Q4_0_ROCMFP4_STRIX_LEAN preset (97.0% of parameters in ROCmFP4, attention K/V in the dual-scale ROCmFP4 variant, token_embd at Q5_K, all norms at F32, and the nextn MTP eh_proj at q8_0).
β οΈ Known issue: output depends on the previous request. With a server reused across requests, this model's output is not reproducible β the same prompt can return different text depending on what ran before it, and can collapse into degenerate repetition (
εε²ηΊΏ,mesh mesh mesh,ζ ζ ζ).The root cause is most likely #29092 β Gated Delta Net recurrent state not being fully cleared between requests on a reused slot. The weights are fine; this is a runtime state bug, and it is not specific to this quant.
Recommended setting:
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.6β ~101 tok/s, and the only setting tested that converges. Counter-intuitively MTP reduces the output-degeneracy problem rather than causing it.--spec-mtp-strict-qwenis optional: it made no measurable difference atn-max 2, though it is required if you use a deeper draft. See Known Issues for the measurements. An earlier revision of this card told you to disable MTP; that advice was wrong and has been retracted.
Revision 2026-08-28: rebuilt from the ornith-ai aug-24 MTP refresh. The retrained MTP head gives 100.96 tok/s at 87.0% acceptance on Strix Halo at the settings above. Note: the new MTP head requires this base revision; do not graft it onto older quants.
| Property | Value |
|---|---|
| Quant format | Q4_0_ROCMFP4_STRIX_LEAN (ROCmFP4) |
| Parameters | 35,505,251,456 (35.51 B) |
| Bits per weight | 4.29 (19,052,438,944 bytes / 35,505,251,456 params) |
| File size | 17.74 GiB (19.05 GB) |
| SHA256 | 0f907917a1bfe4e0ca0d281e5709dcf34b6277063e94fab29491bb5c80fda696 |
| Vision projector | mmproj-Ornith-1.5-35B-BF16.gguf (0.84 GiB) |
| Projector SHA256 | d9ce31026d1cb1f3f8d5152e2e2a014d9d2b302b6c93a7dc07bb0a0487f52837 |
| Max Context | 262,144 tokens (256K) |
| Architecture | qwen35moe, 41 blocks, 1 nextn (MTP) block, d_model 2048, 256 experts |
general.file_type | 106 (ROCmFP4 fork-specific) |
| Source | Ornith-1.5-35B-A3B-GGUF Q8_0 aug-24 MTP refresh (allow-requantize) |
Both checksums above are the content hashes recorded in this repo's git-LFS pointers, which is what huggingface_hub verifies on download. If SHA256SUMS ever disagrees with the table above, trust the table above and the LFS OID β see the note below.
Note 2026-09-27: the
SHA256SUMSfile in this repo was stale, carrying the checksum from an earlier build (b42fb74cβ¦). It has been corrected to match the shipped weights. If you verified againstSHA256SUMSand got a mismatch, your download was fine. The checksum in the table above was always correct.
The factual claims on this card are machine-checked against the shipped files by verify_card.py, which runs in CI on every push. It reads the GGUF header over HTTP range requests and the repo's LFS OIDs, so it needs neither the weights nor a GPU.
# check the card as published on the Hub
python3 verify_card.py --repo julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF
# check a local card and/or a local copy of the weights
python3 verify_card.py --repo julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF \
--card ./README.md --gguf ./Ornith-1.5-35B-A3B-ROCmFP4.gguf
It verifies: both SHA256s (card vs LFS OID vs SHA256SUMS), both file sizes, the parameter count, the bits-per-weight figure, block_count, and that SHA256SUMS lists only files that exist. Exit code is non-zero on any mismatch, so it fails a build rather than shipping a wrong number.
This exists because the card drifted from the artifact several times β a stale SHA256SUMS, a wrong file size, a wrong size-percentage, and a quant description that did not match the actual tensor types. Please run it before publishing a new revision.
| ggml type | Role | Parameters | Share |
|---|---|---|---|
| ROCmFP4 (101) | bulk of the transformer, incl. output.weight | 34,439,168,000 | 97.00% |
| ROCmFP4 (100) | attention K/V, dual-scale variant | 526,385,152 | 1.48% |
| Q5_K (13) | token_embd.weight | 508,559,360 | 1.43% |
| F32 (0) | all norms, incl. blk.40.nextn.{enorm,hnorm,shared_head_norm} | 22,750,336 | 0.06% |
| q8_0 (8) | blk.40.nextn.eh_proj.weight | 8,388,608 | 0.02% |
Note for anyone re-deriving these: earlier revisions of this card described the preset as "FP16 embedding/norm preservation". That was inaccurate β the norms are F32 and the token embedding is Q5_K, not FP16. Only the MTP eh_proj is q8_0; the three MTP norm vectors are F32.
Bare-decode comparison, no speculative decoding on either side:
| Configuration | Size | Decode | Output reproducible? |
|---|---|---|---|
ROCmFP4 + MTP n2 p0.6 (recommended) | 17.74 GiB | 100.96 tok/s | 2 variants, semantically identical |
| ROCmFP4 bare decode | 17.74 GiB | 76.9 tok/s | No β β₯7 variants |
Q4_K_M (base repo) | 20.22 GiB | 71.5β71.7 tok/s | not measured |
ROCmFP4 + MTP n4 p0.6 | 17.74 GiB | 102.77 tok/s | No β β₯5 variants, random |
ROCmFP4 is ~7.3β7.6% faster and 12.3% smaller than the base repo's Q4_K_M (21,713,463,040 bytes vs 19,052,438,944 bytes; equivalently 4.29 vs 4.89 bits per weight).
Note that the fastest row is not the most usable one. n4 is marginally quicker than n2 but diverges randomly, where n2 converges. Treat throughput and stability as separate axes here.
The "reproducible?" column was measured by chandlerma on Strix Halo / Vulkan / RADV, on this quant, by issuing the same request repeatedly and counting distinct outputs. Details and caveats in Known Issues.
Corrected 2026-09-27: this line previously said 16.7% smaller, which is not supported by the actual files. The measured figure is 12.3%, cross-checked two ways β raw bytes and bits-per-weight, which agree. The 16.7% figure could not be reproduced against the
Q4_K_Min the base repo under any assumption about whether that file carries an MTP head, since an MTP head adds only ~9 MB at q8_0. Treat the original number as a transcription error.
Draft acceptance rate is not a health signal for speculative decoding: when the drafter and the target are affected by the same state leak, they are blinded in the same way and acceptance stays high while the output varies. Judge MTP settings by output reproducibility, not acceptance.
halofpx pull downloads and verifies both the ROCmFP4 weights and BF16 vision projector:
halofpx pull ornith-1.5-35b
halofpx serve
halofpx load ornith-1.5-35b
halofpx list reports model-weight and vision-projector readiness separately.
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-mm mmproj-Ornith-1.5-35B-BF16.gguf \
-ngl 999 -c 131072 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-ngl 999 -c 262144 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-ngl 999 -c 131072 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
Note on flags:
- Keep MTP on, at
n-max 2. Disabling MTP makes reproducibility worse, not better. An earlier revision of this card said the opposite; see the retraction below.--spec-mtp-strict-qwenis optional at this depth β it is included in the examples because it is required forn-max > 2, but it measured neutral here.--spec-mtp-strict-qwenis a fork-local flag. It exists in ROCmFPX (PR #54) and not in upstreamggml-org/llama.cpp. You need a ROCmFP4-capable build regardless, since the ROCmFP4 tensor types do not exist upstream β see Requirements.ROCmFP4in ROCmFPX is built as a Vulkan backend plugin, off by default. If you enable it, pass-dev ROCmFPXVulkan0(notVulkan0) and setROCMFPX_PLUGIN_PATH. If your build has ROCmFP4 compiled into the normalVulkan0backend instead, use-dev Vulkan0. Check--list-devicesfor what your build exposes.--mmap/--no-mmapwere removed in favour of--load-modeβ see PR #26934. The default--load-mode automemory-maps unless the device reports it cannot.-np 1is deliberate. The known state leak is per reused slot, and concurrent slots make the cross-talk harder to reason about.
This quant will not load in stock ggml-org/llama.cpp, LM Studio, or Ollama. The q4_0_rocmfp4 tensor types (ggml ids 100 and 101) do not exist upstream. You need a ROCmFP4-capable build:
--spec-mtp-strict-qwenmtp-rocmfp4-strix) β the original ROCmFP4 workBuilding the Vulkan path in ROCmFPX:
cmake -B build -DGGML_VULKAN=ON -DROCMFPX_VULKAN_PLUGIN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target llama-server
Then export ROCMFPX_PLUGIN_PATH=<build>/bin and select the device listed by --list-devices.
Between 2026-09-27 and this revision, this card said "Do not enable --spec-type draft-mtp", pointed users at bare decode, and offered a no-MTP variant as the workaround. That was wrong. It was based on the hypothesis that MTP itself caused the problem. It does not. Disabling MTP makes output reproducibility worse.
I got it wrong because I took a repro from llama.cpp#26425 at face value β in that report, disabling --spec-type made output byte-identical across repeated calls. On this quant on Strix Halo the opposite is observed. The difference is that #26425's clean control was measured on a fresh sequence of requests, whereas the failure here is a per-slot state leak that a short control does not expose.
If you followed the earlier advice, go back to the recommended settings in Serving.
Symptom. The same prompt returns different text depending on which request ran before it. With enough requests it collapses into degenerate repetition β εε²ηΊΏ, mesh mesh mesh, ζ ζ ζ, jijijijiji, op1 op1 op1.
This is not a quant defect. It has been reported on Q4_K_M and other quants of this architecture too.
Upstream: ggml-org/llama.cpp#29092 β "HIP/ROCm β fused Gated Delta Net op carries recurrent state across requests on a reused server slot (qwen35 / qwen35moe); earlier prompts' text is emitted verbatim in later completions". Recurrent state on a Gated Delta Net layer cannot be partially rolled back the way a KV cache can, so a reused slot does not fully return to a clean state.
There is a related open issue, #26425, whose MTP symptom overlaps. Which of the two is the dominant cause here is not settled.
By chandlerma, counting distinct outputs for a repeated identical request:
| Setting | Distinct outputs | Quality |
|---|---|---|
MTP n-max 4 strict-qwen | β₯5 | worst β visible garbage |
MTP off (-ngl 999) | β₯7 | bad β visible garbage |
MTP off, -ngl 998 | β₯6 | bad |
MTP off, GGML_VK_FORCE_MMVQ=1 | β₯5 | bad |
MTP n-max 2 strict-qwen | 2 | good β semantically identical, no garbage |
MTP is a mitigation here, not the cause β but the mechanism is not --spec-mtp-strict-qwen. Tested directly at n-max 2, with and without the flag and nothing else changed:
| Config | Decode | Acceptance | Distinct outputs (8 calls) |
|---|---|---|---|
n-max 2 without --spec-mtp-strict-qwen | 100.76 tok/s | 86.4% | 2 |
n-max 2 with --spec-mtp-strict-qwen | 100.90 tok/s | 87.0% | 2 |
Identical convergence. The flag makes no measurable difference to speed or determinism at this depth, so the mitigation comes from MTP being on at n-max 2, not from the bounded recurrent rollback I originally credited. An earlier revision of this card attributed it to the flag; that attribution is withdrawn.
The 2 remaining variants were semantically equivalent β same answer, slightly different wording. So n-max 2 is stable for practical purposes, not perfectly deterministic.
n-max 2 specificallyDraft depth is the whole story, and the failure has a distinct signature at each depth:
n-max | Decode | Acceptance | Distinct outputs | Pattern |
|---|---|---|---|---|
| 1 | 95.37 tok/s | 92.1% | not measured | β |
| 2 | 100.45β100.96 tok/s | 86.4β87.0% | 2 | monotone convergence |
| 3 | 97.37 tok/s | 80.1% | 3 | strict AβBβCβAβBβC period-3 oscillation |
| 4 | 102.77 tok/s | 86.9% | β₯5 | random divergence |
| off | β | β | β₯7 | random divergence |
n-max 2 is both the fastest well-behaved setting and the boundary: at 3 the output enters a clean period-3 cycle, and by 4 it is fully random. The period-3 oscillation is a useful signature for whoever fixes #29092 β it looks like the leaked state has a small discrete state space that deeper drafts can reach repeatedly.
--spec-draft-p-min was swept at fixed n-max 4: 0.5 β 93.22 tok/s / 70.1% acceptance, 0.6 β 101.15 tok/s / 84.6%, 0.7 β 92.57 tok/s / 88.8%. 0.6 is the throughput optimum; raising it buys acceptance at a net loss.
-ngl 998 (disabling the fused GDN op by moving layer 0 to CPU, the bypass from #29092's HIP-side report) did not help on Vulkan. Do not expect that workaround to transfer.
| Scenario | Distinct outputs | Notes |
|---|---|---|
| Simple factual prompt (TCP/UDP) | 2 | semantically equivalent |
| Code generation (Rust lock-free queue) | 4 | structural divergence β linked list vs ring buffer |
| 5-turn conversation, accumulating context | 0 | fully coherent, each turn correctly referenced prior turns |
Severity scales with how much reasoning the prompt demands. And multi-turn conversation is unaffected β with context accumulating, the anchored history appears to suppress the leak, and 5 consecutive turns produced no degeneration at all. If you are doing agentic or coding work with a conversation history, this configuration is usable as-is; the problem is concentrated in single-shot, stateless requests.
Caveats: all measurements on one machine (Ryzen AI Max+ 395, Mesa 25.3.6, ROCmFPX c49ebdb, Vulkan 1.4.328), -np 1, greedy-ish sampling. Not reproduced on HIP. n-max 1 convergence was not measured. Treat n-max 2 as the best available starting point, not a guarantee.
A large part of the confusion in this thread came from a manually started llama-swap instance holding the port, so the systemd-managed service was crash-looping and every config change was silently served by the old process with the old flags. Symptoms: you edit the config, ps still shows the old arguments, and nothing you do has any effect.
ss -tlnp | grep <port> # is something already listening?
ps aux | grep llama-server # does the running process match your flags?
If you are running under systemd and see bind: address already in use in the journal, kill the stray process before drawing any conclusion from your measurements. It is worth doing this before trusting any A/B result β including the ones on this page.
-np 1 and restart the server between unrelated sessions. Much of the effect is a reused-slot artefact.cache_prompt: false for the requests you care about.Upstream fixed a adjacent bug on 2026-09-26 β
ggml-org/llama.cpp#27530,
"fix K/V and recurrent state cleanup after failed restores". A failed state
restore (corrupt prompt-cache entry, failed checkpoint load) used to leave
partially-written rows behind for the next decode to read. That fix is in
upstream master but not in ROCmFPX (aed0d5f, 2026-09-06) or any other
ROCmFP4-capable build, so it is not protecting this quant yet.
27530-port-rocmfpx.patch
ports it (one test-harness conflict resolved; libllama, llama-server and
the extended test-save-load-state with upstream's new Test 9 all build
clean). It is not runtime-tested. STATE-LEAK-INVESTIGATION.md
documents the port, two theories I checked and withdrew (with reasons), and
the discriminating tests that would settle what remains.
-np 1 and restart the server between unrelated sessions. Much of the effect is a reused-slot artefact.cache_prompt: false for the requests you care about.All four are open as of 2026-09-27. These can produce similar-looking symptoms and are worth ruling out if disabling MTP does not resolve it for you:
| Issue | What it is |
|---|---|
| #29092 | The likely root cause here. Fused Gated Delta Net carries recurrent state across requests on a reused slot. Its own reporter bisects with layer count: -ngl 42 leaks, -ngl 41 (layer 0 on CPU, fused GDN disabled) is clean. Worth trying on Vulkan too. |
| #25618 | On Vulkan (repro hardware: Strix Halo iGPU), speculative decoding against a quantized target is not lossless where the same setup matches on bf16. Ngram speculation stays lossless on the same target, which is what localises it to the draft path. -fa off does not help. |
| #27572 | HIP/gfx1151: a deviceβhost copy race on the MTP hidden state drives draft acceptance to exactly 0.00000 under -np N with long prompts, surfacing as empty completions. Correct at -np 1. The same issue also reports --spec-type draft-mtp aborting during draft-context creation on Vulkan (RADV, b10581) with GGML_ASSERT(tensor->data != NULL), and acceptance 0.00000 from the very first request on newer HIP builds. |
| #26432 | ROCm: context + MTP over budget spills into GTT silently, with no load-time warning, collapsing throughput 60%+. Reported on a 20 GB Radeon 7900 XT rather than a UMA Strix Halo box, and currently labelled stale β treat as a pattern to watch, not a confirmed match. Watch mem_info_gtt_used during prefill. |
turbo3 / turbo4 are not valid -ctk/-ctv values on current llama.cpp. The allowed set is f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1, and the default is f16.
If you want to trade memory for speed, -ctk q8_0 -ctv q8_0 is the safe step down β community testing on Qwen3.5 found q8_0 indistinguishable from f16 in both directions, with real degradation only below that. Going to q4_0/q4_1 is a known cause of degraded long-horizon coherence, so avoid it if you are seeing repetition.
The default for both is f16, so passing nothing is fine.
This is a hybrid reasoning model. Its chat template honours enable_thinking, and llama.cpp exposes it as --reasoning on|off|auto (default auto, which detects from the template). With --reasoning off the template emits a pre-closed empty <think> block, so no thinking tokens are generated.
Two separate flags are easy to confuse:
--reasoning off β controls whether the template starts a thinking block (generation-side).--reasoning-format β controls how thought tags in the response are parsed and returned (none, deepseek, deepseek-legacy; default auto). Use --reasoning-format none if you want any thought text left inline in message.content, or deepseek to get it in message.reasoning_content.--reasoning is a server startup flag and is not settable per request. Note that enable_thinking passed via --chat-template-kwargs is deprecated and ignored on current builds.
If you see raw reasoning-style text in the output despite --reasoning off, that is a different problem from the degeneration above β check that --jinja is on (it is by default on current builds) and that your client is not injecting its own <think> tags.
ROCmFP4 quantization of Ornith 1.5 35B-A3B for AMD Strix Halo (gfx1151) and RDNA 3.5 GPUs, produced with the Q4_0_ROCMFP4_STRIX_LEAN preset (97.0% of parameters in ROCmFP4, attention K/V in the dual-scale ROCmFP4 variant, token_embd at Q5_K, all norms at F32, and the nextn MTP eh_proj at q8_0).
β οΈ Known issue: output depends on the previous request. With a server reused across requests, this model's output is not reproducible β the same prompt can return different text depending on what ran before it, and can collapse into degenerate repetition (
εε²ηΊΏ,mesh mesh mesh,ζ ζ ζ).The root cause is most likely #29092 β Gated Delta Net recurrent state not being fully cleared between requests on a reused slot. The weights are fine; this is a runtime state bug, and it is not specific to this quant.
Recommended setting:
--spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-p-min 0.6β ~101 tok/s, and the only setting tested that converges. Counter-intuitively MTP reduces the output-degeneracy problem rather than causing it.--spec-mtp-strict-qwenis optional: it made no measurable difference atn-max 2, though it is required if you use a deeper draft. See Known Issues for the measurements. An earlier revision of this card told you to disable MTP; that advice was wrong and has been retracted.
Revision 2026-08-28: rebuilt from the ornith-ai aug-24 MTP refresh. The retrained MTP head gives 100.96 tok/s at 87.0% acceptance on Strix Halo at the settings above. Note: the new MTP head requires this base revision; do not graft it onto older quants.
| Property | Value |
|---|---|
| Quant format | Q4_0_ROCMFP4_STRIX_LEAN (ROCmFP4) |
| Parameters | 35,505,251,456 (35.51 B) |
| Bits per weight | 4.29 (19,052,438,944 bytes / 35,505,251,456 params) |
| File size | 17.74 GiB (19.05 GB) |
| SHA256 | 0f907917a1bfe4e0ca0d281e5709dcf34b6277063e94fab29491bb5c80fda696 |
| Vision projector | mmproj-Ornith-1.5-35B-BF16.gguf (0.84 GiB) |
| Projector SHA256 | d9ce31026d1cb1f3f8d5152e2e2a014d9d2b302b6c93a7dc07bb0a0487f52837 |
| Max Context | 262,144 tokens (256K) |
| Architecture | qwen35moe, 41 blocks, 1 nextn (MTP) block, d_model 2048, 256 experts |
general.file_type | 106 (ROCmFP4 fork-specific) |
| Source | Ornith-1.5-35B-A3B-GGUF Q8_0 aug-24 MTP refresh (allow-requantize) |
Both checksums above are the content hashes recorded in this repo's git-LFS pointers, which is what huggingface_hub verifies on download. If SHA256SUMS ever disagrees with the table above, trust the table above and the LFS OID β see the note below.
Note 2026-09-27: the
SHA256SUMSfile in this repo was stale, carrying the checksum from an earlier build (b42fb74cβ¦). It has been corrected to match the shipped weights. If you verified againstSHA256SUMSand got a mismatch, your download was fine. The checksum in the table above was always correct.
The factual claims on this card are machine-checked against the shipped files by verify_card.py, which runs in CI on every push. It reads the GGUF header over HTTP range requests and the repo's LFS OIDs, so it needs neither the weights nor a GPU.
# check the card as published on the Hub
python3 verify_card.py --repo julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF
# check a local card and/or a local copy of the weights
python3 verify_card.py --repo julianmb/Ornith-1.5-35B-A3B-ROCmFP4-GGUF \
--card ./README.md --gguf ./Ornith-1.5-35B-A3B-ROCmFP4.gguf
It verifies: both SHA256s (card vs LFS OID vs SHA256SUMS), both file sizes, the parameter count, the bits-per-weight figure, block_count, and that SHA256SUMS lists only files that exist. Exit code is non-zero on any mismatch, so it fails a build rather than shipping a wrong number.
This exists because the card drifted from the artifact several times β a stale SHA256SUMS, a wrong file size, a wrong size-percentage, and a quant description that did not match the actual tensor types. Please run it before publishing a new revision.
| ggml type | Role | Parameters | Share |
|---|---|---|---|
| ROCmFP4 (101) | bulk of the transformer, incl. output.weight | 34,439,168,000 | 97.00% |
| ROCmFP4 (100) | attention K/V, dual-scale variant | 526,385,152 | 1.48% |
| Q5_K (13) | token_embd.weight | 508,559,360 | 1.43% |
| F32 (0) | all norms, incl. blk.40.nextn.{enorm,hnorm,shared_head_norm} | 22,750,336 | 0.06% |
| q8_0 (8) | blk.40.nextn.eh_proj.weight | 8,388,608 | 0.02% |
Note for anyone re-deriving these: earlier revisions of this card described the preset as "FP16 embedding/norm preservation". That was inaccurate β the norms are F32 and the token embedding is Q5_K, not FP16. Only the MTP eh_proj is q8_0; the three MTP norm vectors are F32.
Bare-decode comparison, no speculative decoding on either side:
| Configuration | Size | Decode | Output reproducible? |
|---|---|---|---|
ROCmFP4 + MTP n2 p0.6 (recommended) | 17.74 GiB | 100.96 tok/s | 2 variants, semantically identical |
| ROCmFP4 bare decode | 17.74 GiB | 76.9 tok/s | No β β₯7 variants |
Q4_K_M (base repo) | 20.22 GiB | 71.5β71.7 tok/s | not measured |
ROCmFP4 + MTP n4 p0.6 | 17.74 GiB | 102.77 tok/s | No β β₯5 variants, random |
ROCmFP4 is ~7.3β7.6% faster and 12.3% smaller than the base repo's Q4_K_M (21,713,463,040 bytes vs 19,052,438,944 bytes; equivalently 4.29 vs 4.89 bits per weight).
Note that the fastest row is not the most usable one. n4 is marginally quicker than n2 but diverges randomly, where n2 converges. Treat throughput and stability as separate axes here.
The "reproducible?" column was measured by chandlerma on Strix Halo / Vulkan / RADV, on this quant, by issuing the same request repeatedly and counting distinct outputs. Details and caveats in Known Issues.
Corrected 2026-09-27: this line previously said 16.7% smaller, which is not supported by the actual files. The measured figure is 12.3%, cross-checked two ways β raw bytes and bits-per-weight, which agree. The 16.7% figure could not be reproduced against the
Q4_K_Min the base repo under any assumption about whether that file carries an MTP head, since an MTP head adds only ~9 MB at q8_0. Treat the original number as a transcription error.
Draft acceptance rate is not a health signal for speculative decoding: when the drafter and the target are affected by the same state leak, they are blinded in the same way and acceptance stays high while the output varies. Judge MTP settings by output reproducibility, not acceptance.
halofpx pull downloads and verifies both the ROCmFP4 weights and BF16 vision projector:
halofpx pull ornith-1.5-35b
halofpx serve
halofpx load ornith-1.5-35b
halofpx list reports model-weight and vision-projector readiness separately.
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-mm mmproj-Ornith-1.5-35B-BF16.gguf \
-ngl 999 -c 131072 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-ngl 999 -c 262144 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
llama-server \
-m Ornith-1.5-35B-A3B-ROCmFP4.gguf \
-ngl 999 -c 131072 -fa on -np 1 \
--spec-type draft-mtp --spec-mtp-strict-qwen \
--spec-draft-n-max 2 --spec-draft-p-min 0.6
Note on flags:
- Keep MTP on, at
n-max 2. Disabling MTP makes reproducibility worse, not better. An earlier revision of this card said the opposite; see the retraction below.--spec-mtp-strict-qwenis optional at this depth β it is included in the examples because it is required forn-max > 2, but it measured neutral here.--spec-mtp-strict-qwenis a fork-local flag. It exists in ROCmFPX (PR #54) and not in upstreamggml-org/llama.cpp. You need a ROCmFP4-capable build regardless, since the ROCmFP4 tensor types do not exist upstream β see Requirements.ROCmFP4in ROCmFPX is built as a Vulkan backend plugin, off by default. If you enable it, pass-dev ROCmFPXVulkan0(notVulkan0) and setROCMFPX_PLUGIN_PATH. If your build has ROCmFP4 compiled into the normalVulkan0backend instead, use-dev Vulkan0. Check--list-devicesfor what your build exposes.--mmap/--no-mmapwere removed in favour of--load-modeβ see PR #26934. The default--load-mode automemory-maps unless the device reports it cannot.-np 1is deliberate. The known state leak is per reused slot, and concurrent slots make the cross-talk harder to reason about.
This quant will not load in stock ggml-org/llama.cpp, LM Studio, or Ollama. The q4_0_rocmfp4 tensor types (ggml ids 100 and 101) do not exist upstream. You need a ROCmFP4-capable build:
--spec-mtp-strict-qwenmtp-rocmfp4-strix) β the original ROCmFP4 workBuilding the Vulkan path in ROCmFPX:
cmake -B build -DGGML_VULKAN=ON -DROCMFPX_VULKAN_PLUGIN=ON -DCMAKE_BUILD_TYPE=Release
cmake --build build -j --target llama-server
Then export ROCMFPX_PLUGIN_PATH=<build>/bin and select the device listed by --list-devices.
Between 2026-09-27 and this revision, this card said "Do not enable --spec-type draft-mtp", pointed users at bare decode, and offered a no-MTP variant as the workaround. That was wrong. It was based on the hypothesis that MTP itself caused the problem. It does not. Disabling MTP makes output reproducibility worse.
I got it wrong because I took a repro from llama.cpp#26425 at face value β in that report, disabling --spec-type made output byte-identical across repeated calls. On this quant on Strix Halo the opposite is observed. The difference is that #26425's clean control was measured on a fresh sequence of requests, whereas the failure here is a per-slot state leak that a short control does not expose.
If you followed the earlier advice, go back to the recommended settings in Serving.
Symptom. The same prompt returns different text depending on which request ran before it. With enough requests it collapses into degenerate repetition β εε²ηΊΏ, mesh mesh mesh, ζ ζ ζ, jijijijiji, op1 op1 op1.
This is not a quant defect. It has been reported on Q4_K_M and other quants of this architecture too.
Upstream: ggml-org/llama.cpp#29092 β "HIP/ROCm β fused Gated Delta Net op carries recurrent state across requests on a reused server slot (qwen35 / qwen35moe); earlier prompts' text is emitted verbatim in later completions". Recurrent state on a Gated Delta Net layer cannot be partially rolled back the way a KV cache can, so a reused slot does not fully return to a clean state.
There is a related open issue, #26425, whose MTP symptom overlaps. Which of the two is the dominant cause here is not settled.
By chandlerma, counting distinct outputs for a repeated identical request:
| Setting | Distinct outputs | Quality |
|---|---|---|
MTP n-max 4 strict-qwen | β₯5 | worst β visible garbage |
MTP off (-ngl 999) | β₯7 | bad β visible garbage |
MTP off, -ngl 998 | β₯6 | bad |
MTP off, GGML_VK_FORCE_MMVQ=1 | β₯5 | bad |
MTP n-max 2 strict-qwen | 2 | good β semantically identical, no garbage |
MTP is a mitigation here, not the cause β but the mechanism is not --spec-mtp-strict-qwen. Tested directly at n-max 2, with and without the flag and nothing else changed:
| Config | Decode | Acceptance | Distinct outputs (8 calls) |
|---|---|---|---|
n-max 2 without --spec-mtp-strict-qwen | 100.76 tok/s | 86.4% | 2 |
n-max 2 with --spec-mtp-strict-qwen | 100.90 tok/s | 87.0% | 2 |
Identical convergence. The flag makes no measurable difference to speed or determinism at this depth, so the mitigation comes from MTP being on at n-max 2, not from the bounded recurrent rollback I originally credited. An earlier revision of this card attributed it to the flag; that attribution is withdrawn.
The 2 remaining variants were semantically equivalent β same answer, slightly different wording. So n-max 2 is stable for practical purposes, not perfectly deterministic.
n-max 2 specificallyDraft depth is the whole story, and the failure has a distinct signature at each depth:
n-max | Decode | Acceptance | Distinct outputs | Pattern |
|---|---|---|---|---|
| 1 | 95.37 tok/s | 92.1% | not measured | β |
| 2 | 100.45β100.96 tok/s | 86.4β87.0% | 2 | monotone convergence |
| 3 | 97.37 tok/s | 80.1% | 3 | strict AβBβCβAβBβC period-3 oscillation |
| 4 | 102.77 tok/s | 86.9% | β₯5 | random divergence |
| off | β | β | β₯7 | random divergence |
n-max 2 is both the fastest well-behaved setting and the boundary: at 3 the output enters a clean period-3 cycle, and by 4 it is fully random. The period-3 oscillation is a useful signature for whoever fixes #29092 β it looks like the leaked state has a small discrete state space that deeper drafts can reach repeatedly.
--spec-draft-p-min was swept at fixed n-max 4: 0.5 β 93.22 tok/s / 70.1% acceptance, 0.6 β 101.15 tok/s / 84.6%, 0.7 β 92.57 tok/s / 88.8%. 0.6 is the throughput optimum; raising it buys acceptance at a net loss.
-ngl 998 (disabling the fused GDN op by moving layer 0 to CPU, the bypass from #29092's HIP-side report) did not help on Vulkan. Do not expect that workaround to transfer.
| Scenario | Distinct outputs | Notes |
|---|---|---|
| Simple factual prompt (TCP/UDP) | 2 | semantically equivalent |
| Code generation (Rust lock-free queue) | 4 | structural divergence β linked list vs ring buffer |
| 5-turn conversation, accumulating context | 0 | fully coherent, each turn correctly referenced prior turns |
Severity scales with how much reasoning the prompt demands. And multi-turn conversation is unaffected β with context accumulating, the anchored history appears to suppress the leak, and 5 consecutive turns produced no degeneration at all. If you are doing agentic or coding work with a conversation history, this configuration is usable as-is; the problem is concentrated in single-shot, stateless requests.
Caveats: all measurements on one machine (Ryzen AI Max+ 395, Mesa 25.3.6, ROCmFPX c49ebdb, Vulkan 1.4.328), -np 1, greedy-ish sampling. Not reproduced on HIP. n-max 1 convergence was not measured. Treat n-max 2 as the best available starting point, not a guarantee.
A large part of the confusion in this thread came from a manually started llama-swap instance holding the port, so the systemd-managed service was crash-looping and every config change was silently served by the old process with the old flags. Symptoms: you edit the config, ps still shows the old arguments, and nothing you do has any effect.
ss -tlnp | grep <port> # is something already listening?
ps aux | grep llama-server # does the running process match your flags?
If you are running under systemd and see bind: address already in use in the journal, kill the stray process before drawing any conclusion from your measurements. It is worth doing this before trusting any A/B result β including the ones on this page.
-np 1 and restart the server between unrelated sessions. Much of the effect is a reused-slot artefact.cache_prompt: false for the requests you care about.Upstream fixed a adjacent bug on 2026-09-26 β
ggml-org/llama.cpp#27530,
"fix K/V and recurrent state cleanup after failed restores". A failed state
restore (corrupt prompt-cache entry, failed checkpoint load) used to leave
partially-written rows behind for the next decode to read. That fix is in
upstream master but not in ROCmFPX (aed0d5f, 2026-09-06) or any other
ROCmFP4-capable build, so it is not protecting this quant yet.
27530-port-rocmfpx.patch
ports it (one test-harness conflict resolved; libllama, llama-server and
the extended test-save-load-state with upstream's new Test 9 all build
clean). It is not runtime-tested. STATE-LEAK-INVESTIGATION.md
documents the port, two theories I checked and withdrew (with reasons), and
the discriminating tests that would settle what remains.
-np 1 and restart the server between unrelated sessions. Much of the effect is a reused-slot artefact.cache_prompt: false for the requests you care about.All four are open as of 2026-09-27. These can produce similar-looking symptoms and are worth ruling out if disabling MTP does not resolve it for you:
| Issue | What it is |
|---|---|
| #29092 | The likely root cause here. Fused Gated Delta Net carries recurrent state across requests on a reused slot. Its own reporter bisects with layer count: -ngl 42 leaks, -ngl 41 (layer 0 on CPU, fused GDN disabled) is clean. Worth trying on Vulkan too. |
| #25618 | On Vulkan (repro hardware: Strix Halo iGPU), speculative decoding against a quantized target is not lossless where the same setup matches on bf16. Ngram speculation stays lossless on the same target, which is what localises it to the draft path. -fa off does not help. |
| #27572 | HIP/gfx1151: a deviceβhost copy race on the MTP hidden state drives draft acceptance to exactly 0.00000 under -np N with long prompts, surfacing as empty completions. Correct at -np 1. The same issue also reports --spec-type draft-mtp aborting during draft-context creation on Vulkan (RADV, b10581) with GGML_ASSERT(tensor->data != NULL), and acceptance 0.00000 from the very first request on newer HIP builds. |
| #26432 | ROCm: context + MTP over budget spills into GTT silently, with no load-time warning, collapsing throughput 60%+. Reported on a 20 GB Radeon 7900 XT rather than a UMA Strix Halo box, and currently labelled stale β treat as a pattern to watch, not a confirmed match. Watch mem_info_gtt_used during prefill. |
turbo3 / turbo4 are not valid -ctk/-ctv values on current llama.cpp. The allowed set is f32, f16, bf16, q8_0, q4_0, q4_1, iq4_nl, q5_0, q5_1, and the default is f16.
If you want to trade memory for speed, -ctk q8_0 -ctv q8_0 is the safe step down β community testing on Qwen3.5 found q8_0 indistinguishable from f16 in both directions, with real degradation only below that. Going to q4_0/q4_1 is a known cause of degraded long-horizon coherence, so avoid it if you are seeing repetition.
The default for both is f16, so passing nothing is fine.
This is a hybrid reasoning model. Its chat template honours enable_thinking, and llama.cpp exposes it as --reasoning on|off|auto (default auto, which detects from the template). With --reasoning off the template emits a pre-closed empty <think> block, so no thinking tokens are generated.
Two separate flags are easy to confuse:
--reasoning off β controls whether the template starts a thinking block (generation-side).--reasoning-format β controls how thought tags in the response are parsed and returned (none, deepseek, deepseek-legacy; default auto). Use --reasoning-format none if you want any thought text left inline in message.content, or deepseek to get it in message.reasoning_content.--reasoning is a server startup flag and is not settable per request. Note that enable_thinking passed via --chat-template-kwargs is deprecated and ignored on current builds.
If you see raw reasoning-style text in the output despite --reasoning off, that is a different problem from the degeneration above β check that --jinja is on (it is by default on current builds) and that your client is not injecting its own <think> tags.