Version 1.0 — 2026-08-11
Ornith-1.0-35B (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to Q4_0_ROCMFP4_STRIX_LEAN (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). Runs the full vision + text multimodal model in ~17.3 GiB.
kyuz0/amd-strix-halo-toolboxes container). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below.Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128 -mmap 0. All benchmarks for this model were run in ROCm containers (HIP backend). Since then we also benchmarked ROCmFPX quants on the Vulkan (RADV) build of the fork — see the comparison table in Qwen3.6-35B-A3B-MTP-Q6_0_ROCMFPX; those numbers are on a different backend and not directly comparable.
| Model | Quant | Size | tg128 (tok/s) | pp512 (tok/s) |
|---|---|---|---|---|
| Ornith-1.0-35B | ROCmFP4-STRIX_LEAN | 17.32 GiB | 66.68 | 1486 |
| grug-35b-v2 (sibling) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 70.92 | 1418 |
| Qwen3.6-35B-A3B (production ref) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 63 | — |
vs production reference: +5.9% tok/s vs Qwen3.6-35B-A3B (66.68 vs 63), at the same 17.3 GiB footprint. See the sibling grug quant for a +12% variant (same arch family, grug fine-tune).
Declared for reproducibility:
balanced (powerprofilesctl get) — default, NOT forced to performance. Representative of an out-of-the-box setup.amd-pstate-epp, scaling_governor performance (amd-pstate-epp default), EPP performanceNote: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.
Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):
blk.*.attn_qkv.weight, blk.*.attn_v.weight) → q4_0_rocmfp4 (high-precision path for attention state)token_embd.weight) → Q5_K (preserve vocab fidelity)blk.*.ffn_*_exps.weight) → q4_0_rocmfp4_fast (max speed path; the bulk of MoE weights)Reference fork: charlie12345/ROCmFPX commit 00d5452.
Serving runtime: see Runtime.
Precomputed by unsloth (46 chunks), redistributed here as imatrix.dat with explicit attribution. The original is at unsloth/Ornith-1.0-35B-GGUF (MIT).
| File | Size | Description |
|---|---|---|
Ornith-1.0-35B-ROCmFP4-STRIX_LEAN.gguf | ~17.32 GiB | Main model (type 106) |
mmproj-F16.gguf | ~857 MB | Vision projector (F16) |
imatrix.dat | ~183 MB | Importance matrix (precomputed by unsloth; for re-quantization) |
# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
-v /path/to/models:/models rocmfpx-llm-service \
llama-server \
-m /models/Ornith-1.0-35B-ROCmFP4-STRIX_LEAN.gguf \
--mmproj /models/mmproj-F16.gguf \
-ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234
Notes:
mtp_num_hidden_layers=1 (MTP weights are present as blk.40.*), but this quant is aligned with the plain-inference MoE pipeline (no --spec-type draft-mtp). MTP weights remain in the file (~1–2 GiB extra) should a future runtime activate them.--mmproj flag is required for the vision tower (multimodal). Without it, text-only still works.Pipeline described in text only (no published scripts):
docker-llm-service-convert image from kyuz0/amd-strix-halo-toolboxes + charlie12345/ROCmFPX (commit 00d5452 or later main HEAD — must contain MODEL_ARCH.QWEN35MOE).unsloth/Ornith-1.0-35B-GGUF.imatrix.dat: llama-quantize --imatrix imatrix.dat <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.Qwen3.5-VL-MoE (base architecture)
└── ornith-ai/Ornith-1.0-35B (MIT)
└── this GGUF (ROCmFP4-STRIX_LEAN)
ornith-ai/Ornith-1.0-35B (MIT) — alias of deepreinforce-ai/Ornith-1.0-35Bunsloth/Ornith-1.0-35B-GGUF (MIT)charlie12345/ROCmFPX (MIT)kyuz0/amd-strix-halo-toolboxesMIT (inherited from ornith-ai/Ornith-1.0-35B and unsloth/Ornith-1.0-35B-GGUF). Derivative work: original model and its license are preserved. See LICENSE and NOTICE.
Built on the shoulders of giants:
We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.
@misc{ornith102026,
title = {Ornith-1.0-35B},
author = {DeepReinforce Team},
year = {2026},
url = {https://deep-reinforce.com/ornith_1_0.html}
}
No affiliation with AMD, Qwen, DeepReinforce, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (MIT).
Benchmark environment: one bare-metal AMD Strix Halo (Ryzen AI MAX+ 395, 128 GB) — full dated configuration and measurement policy: BARE-METAL.md
Requires a ROCmFPX fork build (custom tensor types — stock llama.cpp refuses the file).
Recommended: our lab build (pugant/strix-nebulosa, main) —
reasoning budget and persistent prompt cache on every model; drafter features
where the model ships one: see the engine section of its README.
Vision quant served plain (MTP not enabled at runtime); prompt cache + reasoning budget still apply.
Everything here is experimental and provided as-is, at your own risk.
16 commits
Version 1.0 — 2026-08-11
Ornith-1.0-35B (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to Q4_0_ROCMFP4_STRIX_LEAN (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) on the ROCmFPX fork family — we serve and benchmark these files on our lab runtime (full source: pugant/strix-nebulosa; upstream: charlie12345/ROCmFPX). Runs the full vision + text multimodal model in ~17.3 GiB.
kyuz0/amd-strix-halo-toolboxes container). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below.Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128 -mmap 0. All benchmarks for this model were run in ROCm containers (HIP backend). Since then we also benchmarked ROCmFPX quants on the Vulkan (RADV) build of the fork — see the comparison table in Qwen3.6-35B-A3B-MTP-Q6_0_ROCMFPX; those numbers are on a different backend and not directly comparable.
| Model | Quant | Size | tg128 (tok/s) | pp512 (tok/s) |
|---|---|---|---|---|
| Ornith-1.0-35B | ROCmFP4-STRIX_LEAN | 17.32 GiB | 66.68 | 1486 |
| grug-35b-v2 (sibling) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 70.92 | 1418 |
| Qwen3.6-35B-A3B (production ref) | ROCmFP4-STRIX_LEAN | 17.31 GiB | 63 | — |
vs production reference: +5.9% tok/s vs Qwen3.6-35B-A3B (66.68 vs 63), at the same 17.3 GiB footprint. See the sibling grug quant for a +12% variant (same arch family, grug fine-tune).
Declared for reproducibility:
balanced (powerprofilesctl get) — default, NOT forced to performance. Representative of an out-of-the-box setup.amd-pstate-epp, scaling_governor performance (amd-pstate-epp default), EPP performanceNote: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.
Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):
blk.*.attn_qkv.weight, blk.*.attn_v.weight) → q4_0_rocmfp4 (high-precision path for attention state)token_embd.weight) → Q5_K (preserve vocab fidelity)blk.*.ffn_*_exps.weight) → q4_0_rocmfp4_fast (max speed path; the bulk of MoE weights)Reference fork: charlie12345/ROCmFPX commit 00d5452.
Serving runtime: see Runtime.
Precomputed by unsloth (46 chunks), redistributed here as imatrix.dat with explicit attribution. The original is at unsloth/Ornith-1.0-35B-GGUF (MIT).
| File | Size | Description |
|---|---|---|
Ornith-1.0-35B-ROCmFP4-STRIX_LEAN.gguf | ~17.32 GiB | Main model (type 106) |
mmproj-F16.gguf | ~857 MB | Vision projector (F16) |
imatrix.dat | ~183 MB | Importance matrix (precomputed by unsloth; for re-quantization) |
# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
-v /path/to/models:/models rocmfpx-llm-service \
llama-server \
-m /models/Ornith-1.0-35B-ROCmFP4-STRIX_LEAN.gguf \
--mmproj /models/mmproj-F16.gguf \
-ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234
Notes:
mtp_num_hidden_layers=1 (MTP weights are present as blk.40.*), but this quant is aligned with the plain-inference MoE pipeline (no --spec-type draft-mtp). MTP weights remain in the file (~1–2 GiB extra) should a future runtime activate them.--mmproj flag is required for the vision tower (multimodal). Without it, text-only still works.Pipeline described in text only (no published scripts):
docker-llm-service-convert image from kyuz0/amd-strix-halo-toolboxes + charlie12345/ROCmFPX (commit 00d5452 or later main HEAD — must contain MODEL_ARCH.QWEN35MOE).unsloth/Ornith-1.0-35B-GGUF.imatrix.dat: llama-quantize --imatrix imatrix.dat <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.Qwen3.5-VL-MoE (base architecture)
└── ornith-ai/Ornith-1.0-35B (MIT)
└── this GGUF (ROCmFP4-STRIX_LEAN)
ornith-ai/Ornith-1.0-35B (MIT) — alias of deepreinforce-ai/Ornith-1.0-35Bunsloth/Ornith-1.0-35B-GGUF (MIT)charlie12345/ROCmFPX (MIT)kyuz0/amd-strix-halo-toolboxesMIT (inherited from ornith-ai/Ornith-1.0-35B and unsloth/Ornith-1.0-35B-GGUF). Derivative work: original model and its license are preserved. See LICENSE and NOTICE.
Built on the shoulders of giants:
We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.
@misc{ornith102026,
title = {Ornith-1.0-35B},
author = {DeepReinforce Team},
year = {2026},
url = {https://deep-reinforce.com/ornith_1_0.html}
}
No affiliation with AMD, Qwen, DeepReinforce, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (MIT).
Benchmark environment: one bare-metal AMD Strix Halo (Ryzen AI MAX+ 395, 128 GB) — full dated configuration and measurement policy: BARE-METAL.md
Requires a ROCmFPX fork build (custom tensor types — stock llama.cpp refuses the file).
Recommended: our lab build (pugant/strix-nebulosa, main) —
reasoning budget and persistent prompt cache on every model; drafter features
where the model ships one: see the engine section of its README.
Vision quant served plain (MTP not enabled at runtime); prompt cache + reasoning budget still apply.
Everything here is experimental and provided as-is, at your own risk.
16 commits