Version 1.5 — 2026-08-20
Ornith-1.5-35B (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to Q4_0_ROCMFP4_STRIX_LEAN (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). Runs the full vision + text multimodal model in ~17.7 GiB.
kyuz0/amd-strix-halo-toolboxes container). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below.Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128, Vulkan (RADV) build of the fork, GPU-exclusive window (production service stopped). Note: the Ornith-1.0 release numbers below were measured on the ROCm (HIP) backend — different backend, not directly comparable.
| Model | Quant | Size | tg128 (tok/s) | pp512 (tok/s) |
|---|---|---|---|---|
| Ornith-1.5-35B (this) | ROCmFP4-STRIX_LEAN | 17.73 GiB | 82.45 ± 1.37 | 1161.00 ± 22.76 |
| Ornith-1.0-35B (ROCm) | ROCmFP4-STRIX_LEAN | 17.32 GiB | 66.68 | 1486 |
| grug-35b-v2 (ROCm) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 70.92 | 1418 |
| Qwen3.6-35B-A3B base (Vulkan ref) | ROCmFP4-STRIX_LEAN | 17.73 GiB | 81.57 | 1164.67 |
Sanity check: Ornith-1.5 tracks its base-architecture sibling Qwen3.6-35B-A3B (same quant, same backend) within 1% — 82.45 vs 81.57 tg128. On this Vulkan build Ornith-1.5 decodes 23.7% faster than Ornith-1.0 on ROCm at the same footprint class.
The source model ships a 1-layer MTP head (nextn_predict_layers=1, tensors blk.40.nextn.*), included in this GGUF and activated at runtime only if you opt in with the fork's --spec-type draft-mtp flags.
Measured on the same GPU-exclusive window, same server/prompt methodology (2 prompts × 2 runs, ctx 16k, Vulkan build of the fork):
| Config | prose (tok/s) | deterministic (tok/s) |
|---|---|---|
| plain (no spec) | 78.0 | 77.2 |
| MTP n-max 2 | 59.1 | 75.4 |
| MTP n-max 3 | 44.6 | 63.5 |
| MTP n-max 5 | 36.9 | 47.3 |
Verdict: speculative decoding does not pay off on Ornith-1.5 — plain inference wins at every n-max (−24% prose at the best MTP setting). The measured draft acceptance explains why: position-1 acceptance is high (0.99 on deterministic tasks) but position-2 collapses to ~0.07, so the mean accepted length (1.4–1.7) never covers the draft+verify cost — the nextn layer is a full MoE layer. For comparison, on the same stack the Qwen3.6-35B-A3B base model accepts (0.87, 0.77, 0.64) and gains +37% with MTP n-max 3: the 1.5 fine-tune degraded the MTP head beyond the first drafted token. If you still want to experiment, use --spec-draft-n-max 2; above that it is pure overhead. MTP stays opt-in: with no spec flags the model runs plain inference at the headline speeds above.
Declared for reproducibility:
balanced (powerprofilesctl get) — default, NOT forced to performance. Representative of an out-of-the-box setup.amd-pstate-epp, scaling_governor performance (amd-pstate-epp default), EPP performanceNote: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.
Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):
token_embd.weight) → Q5_K (preserve vocab fidelity)blk.*.attn_qkv.weight, blk.*.attn_v.weight) → q4_0_rocmfp4 (high-precision path for attention state)blk.*.ffn_*_exps.weight) → q4_0_rocmfp4_fast (max speed path; the bulk of MoE weights)blk.40.nextn.*) kept in the file (BF16/F32 as in source)Reference fork: charlie12345/ROCmFPX (MIT).
Serving runtime: see Runtime.
Precomputed by bartowski on 573 chunks (calibration-v6 dataset), redistributed here as imatrix-Ornith-1.5-35B-bartowski.gguf with explicit attribution. The original is at bartowski/Ornith-1.5-35B-A3B-GGUF (MIT). Unlike the Ornith-1.0 release, which used the unsloth imatrix computed on 1.0 weights.
| File | Size | Description |
|---|---|---|
Ornith-1.5-35B-ROCmFP4-STRIX_LEAN.gguf | ~17.73 GiB | Main model (type 106), MTP head included |
mmproj-Ornith-1.5-35B-BF16.gguf | ~860 MB | Vision projector (BF16, from the official repo) |
imatrix-Ornith-1.5-35B-bartowski.gguf | ~183 MB | Importance matrix (precomputed by bartowski; for re-quantization) |
# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
-v /path/to/models:/models rocmfpx-llm-service \
llama-server \
-m /models/Ornith-1.5-35B-ROCmFP4-STRIX_LEAN.gguf \
--mmproj /models/mmproj-Ornith-1.5-35B-BF16.gguf \
-ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234
Notes:
--spec-type draft-mtp --spec-draft-ngl all --spec-draft-p-min 0.0 --spec-draft-p-split 0.10 --spec-draft-n-max <N> to activate speculative decoding (see the MTP section above for measured n-max guidance). Plain inference (no spec flags) is the default and what the headline benchmark table reports.--mmproj flag is required for the vision tower (multimodal). Without it, text-only still works.docker-llm-service-convert image from kyuz0/amd-strix-halo-toolboxes + charlie12345/ROCmFPX (must contain MODEL_ARCH.QWEN35MOE).ornith-ai/Ornith-1.5-35B-A3B-GGUF.llama-quantize --imatrix imatrix-Ornith-1.5-35B-bartowski.gguf <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.Qwen3.5-VL-MoE (base architecture)
└── ornith-ai/Ornith-1.5-35B (A3B) (MIT)
└── this GGUF (ROCmFP4-STRIX_LEAN)
ornith-ai/Ornith-1.5-35B (MIT)ornith-ai/Ornith-1.5-35B-A3B-GGUF (MIT)bartowski/Ornith-1.5-35B-A3B-GGUF (MIT)charlie12345/ROCmFPX (MIT)kyuz0/amd-strix-halo-toolboxesMIT (inherited from ornith-ai/Ornith-1.5-35B and its GGUF release). Derivative work: original model and its license are preserved. See LICENSE and NOTICE.
Built on the shoulders of giants:
We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.
@misc{ornith152026,
title = {Ornith-1.5-35B},
author = {DeepReinforce Team},
year = {2026},
url = {https://huggingface.co/ornith-ai/Ornith-1.5-35B}
}
No affiliation with AMD, Qwen, DeepReinforce, bartowski, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (MIT).
Benchmark environment: one bare-metal AMD Strix Halo (Ryzen AI MAX+ 395, 128 GB) — full dated configuration and measurement policy: BARE-METAL.md
Requires a ROCmFPX fork build (custom tensor types — stock llama.cpp refuses the file).
Recommended: our lab build (pugant/strix-nebulosa, main) —
reasoning budget and persistent prompt cache on every model; drafter features
where the model ships one: see the engine section of its README.
Honest note: the fine-tune degraded the MTP head (pos-2 acceptance ~0.07) — run it plain, skip the drafter.
Everything here is experimental and provided as-is, at your own risk.
13 commits
Version 1.5 — 2026-08-20
Ornith-1.5-35B (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to Q4_0_ROCMFP4_STRIX_LEAN (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). Runs the full vision + text multimodal model in ~17.7 GiB.
kyuz0/amd-strix-halo-toolboxes container). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below.Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128, Vulkan (RADV) build of the fork, GPU-exclusive window (production service stopped). Note: the Ornith-1.0 release numbers below were measured on the ROCm (HIP) backend — different backend, not directly comparable.
| Model | Quant | Size | tg128 (tok/s) | pp512 (tok/s) |
|---|---|---|---|---|
| Ornith-1.5-35B (this) | ROCmFP4-STRIX_LEAN | 17.73 GiB | 82.45 ± 1.37 | 1161.00 ± 22.76 |
| Ornith-1.0-35B (ROCm) | ROCmFP4-STRIX_LEAN | 17.32 GiB | 66.68 | 1486 |
| grug-35b-v2 (ROCm) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 70.92 | 1418 |
| Qwen3.6-35B-A3B base (Vulkan ref) | ROCmFP4-STRIX_LEAN | 17.73 GiB | 81.57 | 1164.67 |
Sanity check: Ornith-1.5 tracks its base-architecture sibling Qwen3.6-35B-A3B (same quant, same backend) within 1% — 82.45 vs 81.57 tg128. On this Vulkan build Ornith-1.5 decodes 23.7% faster than Ornith-1.0 on ROCm at the same footprint class.
The source model ships a 1-layer MTP head (nextn_predict_layers=1, tensors blk.40.nextn.*), included in this GGUF and activated at runtime only if you opt in with the fork's --spec-type draft-mtp flags.
Measured on the same GPU-exclusive window, same server/prompt methodology (2 prompts × 2 runs, ctx 16k, Vulkan build of the fork):
| Config | prose (tok/s) | deterministic (tok/s) |
|---|---|---|
| plain (no spec) | 78.0 | 77.2 |
| MTP n-max 2 | 59.1 | 75.4 |
| MTP n-max 3 | 44.6 | 63.5 |
| MTP n-max 5 | 36.9 | 47.3 |
Verdict: speculative decoding does not pay off on Ornith-1.5 — plain inference wins at every n-max (−24% prose at the best MTP setting). The measured draft acceptance explains why: position-1 acceptance is high (0.99 on deterministic tasks) but position-2 collapses to ~0.07, so the mean accepted length (1.4–1.7) never covers the draft+verify cost — the nextn layer is a full MoE layer. For comparison, on the same stack the Qwen3.6-35B-A3B base model accepts (0.87, 0.77, 0.64) and gains +37% with MTP n-max 3: the 1.5 fine-tune degraded the MTP head beyond the first drafted token. If you still want to experiment, use --spec-draft-n-max 2; above that it is pure overhead. MTP stays opt-in: with no spec flags the model runs plain inference at the headline speeds above.
Declared for reproducibility:
balanced (powerprofilesctl get) — default, NOT forced to performance. Representative of an out-of-the-box setup.amd-pstate-epp, scaling_governor performance (amd-pstate-epp default), EPP performanceNote: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.
Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):
token_embd.weight) → Q5_K (preserve vocab fidelity)blk.*.attn_qkv.weight, blk.*.attn_v.weight) → q4_0_rocmfp4 (high-precision path for attention state)blk.*.ffn_*_exps.weight) → q4_0_rocmfp4_fast (max speed path; the bulk of MoE weights)blk.40.nextn.*) kept in the file (BF16/F32 as in source)Reference fork: charlie12345/ROCmFPX (MIT).
Serving runtime: see Runtime.
Precomputed by bartowski on 573 chunks (calibration-v6 dataset), redistributed here as imatrix-Ornith-1.5-35B-bartowski.gguf with explicit attribution. The original is at bartowski/Ornith-1.5-35B-A3B-GGUF (MIT). Unlike the Ornith-1.0 release, which used the unsloth imatrix computed on 1.0 weights.
| File | Size | Description |
|---|---|---|
Ornith-1.5-35B-ROCmFP4-STRIX_LEAN.gguf | ~17.73 GiB | Main model (type 106), MTP head included |
mmproj-Ornith-1.5-35B-BF16.gguf | ~860 MB | Vision projector (BF16, from the official repo) |
imatrix-Ornith-1.5-35B-bartowski.gguf | ~183 MB | Importance matrix (precomputed by bartowski; for re-quantization) |
# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
-v /path/to/models:/models rocmfpx-llm-service \
llama-server \
-m /models/Ornith-1.5-35B-ROCmFP4-STRIX_LEAN.gguf \
--mmproj /models/mmproj-Ornith-1.5-35B-BF16.gguf \
-ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234
Notes:
--spec-type draft-mtp --spec-draft-ngl all --spec-draft-p-min 0.0 --spec-draft-p-split 0.10 --spec-draft-n-max <N> to activate speculative decoding (see the MTP section above for measured n-max guidance). Plain inference (no spec flags) is the default and what the headline benchmark table reports.--mmproj flag is required for the vision tower (multimodal). Without it, text-only still works.docker-llm-service-convert image from kyuz0/amd-strix-halo-toolboxes + charlie12345/ROCmFPX (must contain MODEL_ARCH.QWEN35MOE).ornith-ai/Ornith-1.5-35B-A3B-GGUF.llama-quantize --imatrix imatrix-Ornith-1.5-35B-bartowski.gguf <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.Qwen3.5-VL-MoE (base architecture)
└── ornith-ai/Ornith-1.5-35B (A3B) (MIT)
└── this GGUF (ROCmFP4-STRIX_LEAN)
ornith-ai/Ornith-1.5-35B (MIT)ornith-ai/Ornith-1.5-35B-A3B-GGUF (MIT)bartowski/Ornith-1.5-35B-A3B-GGUF (MIT)charlie12345/ROCmFPX (MIT)kyuz0/amd-strix-halo-toolboxesMIT (inherited from ornith-ai/Ornith-1.5-35B and its GGUF release). Derivative work: original model and its license are preserved. See LICENSE and NOTICE.
Built on the shoulders of giants:
We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.
@misc{ornith152026,
title = {Ornith-1.5-35B},
author = {DeepReinforce Team},
year = {2026},
url = {https://huggingface.co/ornith-ai/Ornith-1.5-35B}
}
No affiliation with AMD, Qwen, DeepReinforce, bartowski, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (MIT).
Benchmark environment: one bare-metal AMD Strix Halo (Ryzen AI MAX+ 395, 128 GB) — full dated configuration and measurement policy: BARE-METAL.md
Requires a ROCmFPX fork build (custom tensor types — stock llama.cpp refuses the file).
Recommended: our lab build (pugant/strix-nebulosa, main) —
reasoning budget and persistent prompt cache on every model; drafter features
where the model ships one: see the engine section of its README.
Honest note: the fine-tune degraded the MTP head (pos-2 acceptance ~0.07) — run it plain, skip the drafter.
Everything here is experimental and provided as-is, at your own risk.
13 commits