tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF

Model

Ornith-1.5-35B-A3B MXFP4 GGUF

2

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Ornith-1.5-35B-A3B MXFP4 GGUF

MXFP4 GGUF quantization of ornith-ai/Ornith-1.5-35B-A3B, a mixture-of-experts model with ~35B total and ~3B active parameters per token, with the MTP (multi-token prediction) layer kept for self-speculative decoding, plus an F16 vision projector.

Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by the Ornith team (ornith-ai). For the official model, official GGUF builds and evaluations, see ornith-ai/Ornith-1.5-35B-A3B and ornith-ai/Ornith-1.5-35B-A3B-GGUF.

Ornith-1.5-35B-A3B is the Ornith team's coding- and agent-oriented MoE model (architecture qwen35moe: Gated DeltaNet + gated attention hybrid, 40 layers, 256 experts with top-8 routing plus a shared expert). Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates the model's benchmark results (the only quality measurement here is a perplexity comparison between quantizations, see Quality).

Best performance with WHIRL

These files were made for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. WHIRL has dedicated MXFP4 expert kernels (fp8-activation expert prefill, MTP + n-gram speculative decoding whose output is identical to plain greedy decoding).

The files are standard GGUF and also run in llama.cpp: tested with llama.cpp b11214 (ROCm build) — llama-bench, llama-server (plain and --spec-type draft-mtp) and llama-mtmd-cli with the projector. Other llama.cpp builds and other GGUF runtimes were not tested.

Files

FileContentsSize (bytes)SHA256
Ornith-1.5-35B-A3B-MXFP4.gguflanguage model + MTP layer (MXFP4 experts, see recipe)19,819,767,1366735282db893ec925eabe864bca72521563d9f85f7152f0cab7bec5bfaab9960
mmproj-Ornith-1.5-35B-A3B-F16.ggufvision encoder / projector, F16899,283,2967950af74588cfb1acc3dd37974baaa3cceda4acf4d792cc643b4a0adb394e041

The language-model file is 4.46 bits per weight on average and takes 18.45 GiB of VRAM when loaded (plus KV cache). The projector is needed only for image input.

Usage with WHIRL

Requirements (from the WHIRL quick start): AMD Radeon AI PRO R9700, Windows 11 64-bit, AMD Software Adrenalin 26.8.1 or newer; nothing else to install. Unpack whirl-0.1.0-windows-x64.zip from the release page, then in PowerShell:

# one prompt (MTP + n-gram speculative decoding is on by default)
.\whirl.exe chat C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf "用 Python 寫一個 LRU cache,附上 pytest 測試。" --max-tokens 800

# image input
.\whirl.exe chat C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf "Describe this image." `
  --mmproj C:\models\mmproj-Ornith-1.5-35B-A3B-F16.gguf --image photo.png --max-tokens 400

# OpenAI-compatible server (4 slots, continuous batching, prefix cache with RAM / SSD tiers)
.\whirl-server.exe C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf --port 8080 `
  --mmproj C:\models\mmproj-Ornith-1.5-35B-A3B-F16.gguf

The first run with a new model file tunes the prefill kernels once (about 3 seconds for this file; cached afterwards). Every option is in the usage reference.

Usage with llama.cpp

Plain decoding (the configuration that was fastest for this model in llama.cpp b11214, see the performance table):

llama-server -m Ornith-1.5-35B-A3B-MXFP4.gguf -ngl 99 -c 131072 -np 1 -fa on -ub 2048 -b 2048 \
  --port 8080 --jinja --mmproj mmproj-Ornith-1.5-35B-A3B-F16.gguf

The MTP layer can be used with --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --spec-draft-p-min 0.3; on this MoE model it was slower than plain decoding in b11214 for generation prompts, but faster on file-editing prompts with --spec-type draft-mtp,ngram-mod --spec-draft-n-max 1. Without --spec-* flags llama.cpp logs the MTP-layer tensors as unused, which is expected. The loader prints unknown type mxfp4 as a warning; loading succeeds.

Image input:

llama-mtmd-cli -m Ornith-1.5-35B-A3B-MXFP4.gguf --mmproj mmproj-Ornith-1.5-35B-A3B-F16.gguf \
  --image photo.png -ngl 99 -c 8192 -fa on -p "Describe this image."

Sampling recommended by the upstream card: temperature 0.6, top_p 0.95, top_k 20 for general tasks (temperature 1.0 to reproduce its benchmarks). Thinking is on by default in the chat template.

Quantization recipe

  1. Convert the official BF16 checkpoint with llama.cpp convert_hf_to_gguf.py --outtype bf16. The MTP layer is kept (nextn_predict_layers = 1, layer blk.40); tensor names match the official Q4_K_M GGUF. The tokenizer and chat template are embedded by the conversion script.
  2. Quantize with llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build, without an importance matrix (mainline llama.cpp ignores the imatrix for MXFP4 and Q8_0 anyway), using this tensor-type file (first matching rule wins):
blk\.40\.ffn_gate_inp=f32
blk\.40\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate_exps=mxfp4
ffn_up_exps=mxfp4
ffn_down_exps=mxfp4
ffn_gate_shexp=mxfp4
ffn_up_shexp=mxfp4
ffn_down_shexp=mxfp4
ssm_out=mxfp4
llama-quantize --tensor-type-file tt_orn.txt --output-tensor-type q6_k \
  ornith-1.5-35b-bf16.gguf Ornith-1.5-35B-A3B-MXFP4.gguf MXFP4_MOE 16
Tensor groupType
Routed experts (ffn_gate/up/down_exps), shared expert (ffn_*_shexp), and every dense matrix of the 40 main layers (attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out)MXFP4
token_embd; ssm_alpha, ssm_beta; the whole MTP layer blk.40 (except its router)Q8_0
Routers (ffn_gate_inp, including blk.40), ffn_gate_inp_shexp, norms, small SSM tensorsF32
output (LM head)Q6_K
  1. Vision projector: convert_hf_to_gguf.py --mmproj --outtype f16 from the same BF16 weights, no further quantization (the official GGUF repo ships a BF16 projector; this one is F16).

Why these choices: ssm_alpha / ssm_beta in Q8_0 rather than F32 keep WHIRL's single-token and multi-token (MTP verify) paths on the same fused kernel, so speculative output stays bit-identical to plain greedy decoding; a Q6_K head is the fastest head for WHIRL's draft step; Q8_0 for embeddings and the MTP layer costs a little speed for slightly higher draft acceptance. More detail: WHIRL quantization guide §6.

Measured performance (AMD Radeon AI PRO R9700, 32 GB)

WHIRL v0.1.0 vs llama.cpp b11214 (ROCm), same GGUF file, same prompts, greedy decoding in both, one model process on the GPU at a time, measured 2026-10-03. Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram, and of the flag sets tried; named in the table). "no MTP" = plain decoding without any speculation. Full methodology, raw-data locations and the rows where WHIRL does not lead: WHIRL benchmarks. The R9700 in this machine is connected as a USB4 eGPU; this does not affect prefill or decode (they stay in VRAM) but slows model loading and the host-RAM / SSD KV restores.

Decode (tok/s), WHIRL default = MTP + n-gram

ScenarioWHIRL MTP + n-gramllama.cpp fastest (config)Ratio
7 Chinese coding prompts (Chinese question, Chinese explanation, English code), 800 tokens, server244.5118.8 (plain)2.06×
File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output, server618.1239.3 (draft-mtp,ngram-mod, n-max 1)2.58×
128 tokens after a 16k-token code context, server301.5107.5 (plain)2.80×

Decode (tok/s), no MTP (plain decoding in both engines)

ScenarioWHIRL no MTPllama.cpp no MTPRatio
7 Chinese coding prompts, server166.7118.81.40×
File editing, server163.3117.11.40×
After a 16k-token context, server150.5107.51.40×
CLI bench (whirl bench 256 tokens vs llama-bench tg256)193.6119.71.62×

Prefill (tok/s, each engine's own bench tool, KV f16)

Prompt tokensWHIRLllama.cpp (flags)Ratio
882,5462,115 (-ub 512)1.20×
2,04811,7054,876 (-ub 4096 -b 4096)2.40×
8,19210,8584,637 (-ub 2048)2.34×
32,7687,9783,778 (-ub 2048)2.11×
131,0723,7822,239 (-ub 2048)1.69×

Server

TestWHIRLllama-serverRatio
4 concurrent requests, aggregate tok/s (256 tokens each, prefill included) — WHIRL MTP + n-gram vs llama.cpp fastest (plain)396.7191.92.07×
Same, WHIRL no MTP vs llama.cpp no MTP324.9191.91.69×
TTFT, ~12.2k-token system prompt, cold1.314 s3.396 s2.6× faster
TTFT, new conversation reusing that ~12.2k system prompt0.064 s0.259 s4.1× faster
TTFT, ~26.4k-token system prompt, cold3.294 s8.151 s2.5× faster
TTFT, evicted ~27.4k-token session restored from host RAM0.265 s0.950 s3.6× faster
TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier)0.298 sN/A (no equivalent feature)—

VRAM at equal context (CLI bench, ~33k tokens, KV f16): 22.5 GiB in both engines. The WHIRL server deliberately fills free VRAM with its KV pool (414,464 tokens of f16 KV with this file).

These are speed measurements only. In the benchmark runs WHIRL's speculative output was identical to its plain greedy output, so the MTP + n-gram speed-up does not change the generated text.

Quality

  • Perplexity (llama.cpp b11214 ROCm llama-perplexity, -c 2048 -b 512 -ub 512 -fa on -ngl 99, 51 chunks of a 219k-character mixed Chinese/English source-code corpus — the same corpus and settings as our Swift MXFP4 card — R9700, measured 2026-10-03). The usual wikitext-2 test set was not available on the measurement machine.

    filesize (bytes)PPLvs Q4_K_M
    this repo: Ornith-1.5-35B-A3B-MXFP4.gguf19,819,767,1365.5192 ± 0.0653+3.7%
    reference: official Ornith-1.5-35B-Q4_K_M.gguf (ornith-ai GGUF repo)21,713,462,8485.3206 ± 0.0619—

    Interpretation: on this corpus our MXFP4 file has 3.7% higher perplexity than the official Q4_K_M, which is about 10% larger. The ± values are each run's own statistical error; since both runs score the same tokens, a paired per-chunk comparison is the fairer test: MXFP4 is worse on 42 of 51 chunks, by 0.037 ± 0.006 nats/token on average (about 6 standard errors) — a small but measurable difference. The file was built for WHIRL's speed (see the recipe), not to beat Q4_K_M on size-for-quality. No BF16 baseline and no KL-divergence were measured (the BF16 GGUF used for quantization had been deleted from the machine before this test), so the absolute loss against the unquantized model is unknown. Perplexity on one corpus is only a proxy; it was not checked on downstream tasks. These numbers are not comparable with perplexities of other models (e.g. our Swift card), since tokenizers and training differ.

Other checks:

  • In WHIRL, MTP + n-gram, MTP-only and plain greedy decoding give identical text on 7 test prompts (Chinese / English, thinking on and off, 2k / 8k / 32k-token contexts). This shows the speculative paths are exact; it is not a measure of quantization loss.
  • Image input: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, with the exact text, by both WHIRL 0.1.0 (whirl chat --mmproj --image) and llama.cpp b11214 llama-mtmd-cli with this projector (single smoke test, 2026-10-03).

For the model's own evaluations, see the upstream card. A 4.46-bpw MXFP4 file is expected to lose some accuracy against BF16; if that matters for your use, compare against the official Q8_0 / BF16 GGUFs. If you run only llama.cpp and care most about quality, the official Q4_K_M has slightly lower perplexity (about 3.7%) at about 10% larger size.

License

The upstream model is published under the MIT License (declared in the upstream model card metadata). These quantized files are a derivative of it and are distributed under the same license:

  • LICENSE: the MIT License text. At the time of this release the upstream repository declares license: mit but contains no LICENSE file (its license_link returns "Entry not found"), so this is the standard MIT License text with the Ornith team named as copyright holder.
  • NOTICE: attribution and a description of the changes made in this redistribution.

The upstream card states that the Ornith family was developed on top of Qwen3.5 and Gemma 4; see the upstream card for the provenance of the base weights.

Credits

Upstream citation:

@misc{ornith_1_5,
    title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
    url = {https://ornith.ai/ornith_1_5.html},
    author = {{Ornith Team}},
    year = {2026}
}

繁體中文摘要

  • 這是 ornith-ai/Ornith-1.5-35B-A3B 的非官方社群 MXFP4 GGUF 量化,與 Ornith 團隊無關。MoE 模型,總參數約 35B、每個 token 約 3B 啟用;保留 MTP 層(Q8_0),可做自我推測解碼;另附 F16 視覺投影檔。
  • 專為 WHIRL(Apache-2.0,AMD Radeon AI PRO R9700 的原生 Windows C++/HIP 推論引擎,v0.1.0 下載)製作;也能在 llama.cpp 執行(已用 b11214 測試:llama-bench、llama-server、llama-mtmd-cli)。
  • 量化:專家、共享專家與所有稠密矩陣 MXFP4;token_embd、ssm_alpha / beta、整個 MTP 層 Q8_0;router 與 norm 為 F32;輸出頭 Q6_K;不使用 imatrix。平均 4.46 bpw,載入 18.45 GiB。
  • R9700(USB4 外接)WHIRL v0.1.0 對 llama.cpp b11214 最快設定:中文寫程式解碼 244.5 對 118.8 tok/s(2.06×)、檔案編輯 618.1 對 239.3(2.58×)、16k 上下文後 301.5 對 107.5(2.80×);不開 MTP 時 166.7 對 118.8(1.40×);prefill 8k 10,858 對 4,637 tok/s(2.34×);4 個並行請求總吞吐 396.7 對 191.9 tok/s(2.07×)。
  • 品質:困惑度(llama.cpp b11214、中英混合程式碼語料 21.9 萬字元、2048 ctx,與 Swift 卡相同方法)本檔 5.5192 ± 0.0653,官方 Q4_K_M 5.3206 ± 0.0619——本檔高 3.7%(檔案小約 9%),逐段配對比較 51 段中 42 段較差(平均差 0.037 ± 0.006 nats/token,約 6 個標準誤),屬可量到的小幅損失;若只用 llama.cpp 且最在意品質,官方 Q4_K_M 的困惑度略低(約 3.7%),但檔案大約 10%。未量 BF16 基準與 KL(BF16 已刪),對未量化模型的絕對損失未知;困惑度也不能和其他模型(如 Swift)比較。另確認 WHIRL 推測解碼與 plain greedy 輸出逐字相同(7 題)、看圖冒煙測試(WHIRL 與 llama-mtmd-cli)正確。
  • 授權:MIT(上游 metadata 宣告;上游 repo 沒有 LICENSE 檔,本 repo 附標準 MIT 條文);詳見 LICENSE / NOTICE。
conversational
endpoints_compatible
gguf
llama.cpp
mxfp4
rdna4
text-generation
whirl

tsaipifong/Ornith-1.5-35B-A3B-MXFP4-GGUF

Model

Ornith-1.5-35B-A3B MXFP4 GGUF

2

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Ornith-1.5-35B-A3B MXFP4 GGUF

MXFP4 GGUF quantization of ornith-ai/Ornith-1.5-35B-A3B, a mixture-of-experts model with ~35B total and ~3B active parameters per token, with the MTP (multi-token prediction) layer kept for self-speculative decoding, plus an F16 vision projector.

Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by the Ornith team (ornith-ai). For the official model, official GGUF builds and evaluations, see ornith-ai/Ornith-1.5-35B-A3B and ornith-ai/Ornith-1.5-35B-A3B-GGUF.

Ornith-1.5-35B-A3B is the Ornith team's coding- and agent-oriented MoE model (architecture qwen35moe: Gated DeltaNet + gated attention hybrid, 40 layers, 256 experts with top-8 routing plus a shared expert). Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates the model's benchmark results (the only quality measurement here is a perplexity comparison between quantizations, see Quality).

Best performance with WHIRL

These files were made for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. WHIRL has dedicated MXFP4 expert kernels (fp8-activation expert prefill, MTP + n-gram speculative decoding whose output is identical to plain greedy decoding).

The files are standard GGUF and also run in llama.cpp: tested with llama.cpp b11214 (ROCm build) — llama-bench, llama-server (plain and --spec-type draft-mtp) and llama-mtmd-cli with the projector. Other llama.cpp builds and other GGUF runtimes were not tested.

Files

FileContentsSize (bytes)SHA256
Ornith-1.5-35B-A3B-MXFP4.gguflanguage model + MTP layer (MXFP4 experts, see recipe)19,819,767,1366735282db893ec925eabe864bca72521563d9f85f7152f0cab7bec5bfaab9960
mmproj-Ornith-1.5-35B-A3B-F16.ggufvision encoder / projector, F16899,283,2967950af74588cfb1acc3dd37974baaa3cceda4acf4d792cc643b4a0adb394e041

The language-model file is 4.46 bits per weight on average and takes 18.45 GiB of VRAM when loaded (plus KV cache). The projector is needed only for image input.

Usage with WHIRL

Requirements (from the WHIRL quick start): AMD Radeon AI PRO R9700, Windows 11 64-bit, AMD Software Adrenalin 26.8.1 or newer; nothing else to install. Unpack whirl-0.1.0-windows-x64.zip from the release page, then in PowerShell:

# one prompt (MTP + n-gram speculative decoding is on by default)
.\whirl.exe chat C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf "用 Python 寫一個 LRU cache,附上 pytest 測試。" --max-tokens 800

# image input
.\whirl.exe chat C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf "Describe this image." `
  --mmproj C:\models\mmproj-Ornith-1.5-35B-A3B-F16.gguf --image photo.png --max-tokens 400

# OpenAI-compatible server (4 slots, continuous batching, prefix cache with RAM / SSD tiers)
.\whirl-server.exe C:\models\Ornith-1.5-35B-A3B-MXFP4.gguf --port 8080 `
  --mmproj C:\models\mmproj-Ornith-1.5-35B-A3B-F16.gguf

The first run with a new model file tunes the prefill kernels once (about 3 seconds for this file; cached afterwards). Every option is in the usage reference.

Usage with llama.cpp

Plain decoding (the configuration that was fastest for this model in llama.cpp b11214, see the performance table):

llama-server -m Ornith-1.5-35B-A3B-MXFP4.gguf -ngl 99 -c 131072 -np 1 -fa on -ub 2048 -b 2048 \
  --port 8080 --jinja --mmproj mmproj-Ornith-1.5-35B-A3B-F16.gguf

The MTP layer can be used with --spec-type draft-mtp --spec-draft-n-max 2 --spec-draft-n-min 0 --spec-draft-p-min 0.3; on this MoE model it was slower than plain decoding in b11214 for generation prompts, but faster on file-editing prompts with --spec-type draft-mtp,ngram-mod --spec-draft-n-max 1. Without --spec-* flags llama.cpp logs the MTP-layer tensors as unused, which is expected. The loader prints unknown type mxfp4 as a warning; loading succeeds.

Image input:

llama-mtmd-cli -m Ornith-1.5-35B-A3B-MXFP4.gguf --mmproj mmproj-Ornith-1.5-35B-A3B-F16.gguf \
  --image photo.png -ngl 99 -c 8192 -fa on -p "Describe this image."

Sampling recommended by the upstream card: temperature 0.6, top_p 0.95, top_k 20 for general tasks (temperature 1.0 to reproduce its benchmarks). Thinking is on by default in the chat template.

Quantization recipe

  1. Convert the official BF16 checkpoint with llama.cpp convert_hf_to_gguf.py --outtype bf16. The MTP layer is kept (nextn_predict_layers = 1, layer blk.40); tensor names match the official Q4_K_M GGUF. The tokenizer and chat template are embedded by the conversion script.
  2. Quantize with llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build, without an importance matrix (mainline llama.cpp ignores the imatrix for MXFP4 and Q8_0 anyway), using this tensor-type file (first matching rule wins):
blk\.40\.ffn_gate_inp=f32
blk\.40\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate_exps=mxfp4
ffn_up_exps=mxfp4
ffn_down_exps=mxfp4
ffn_gate_shexp=mxfp4
ffn_up_shexp=mxfp4
ffn_down_shexp=mxfp4
ssm_out=mxfp4
llama-quantize --tensor-type-file tt_orn.txt --output-tensor-type q6_k \
  ornith-1.5-35b-bf16.gguf Ornith-1.5-35B-A3B-MXFP4.gguf MXFP4_MOE 16
Tensor groupType
Routed experts (ffn_gate/up/down_exps), shared expert (ffn_*_shexp), and every dense matrix of the 40 main layers (attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out)MXFP4
token_embd; ssm_alpha, ssm_beta; the whole MTP layer blk.40 (except its router)Q8_0
Routers (ffn_gate_inp, including blk.40), ffn_gate_inp_shexp, norms, small SSM tensorsF32
output (LM head)Q6_K
  1. Vision projector: convert_hf_to_gguf.py --mmproj --outtype f16 from the same BF16 weights, no further quantization (the official GGUF repo ships a BF16 projector; this one is F16).

Why these choices: ssm_alpha / ssm_beta in Q8_0 rather than F32 keep WHIRL's single-token and multi-token (MTP verify) paths on the same fused kernel, so speculative output stays bit-identical to plain greedy decoding; a Q6_K head is the fastest head for WHIRL's draft step; Q8_0 for embeddings and the MTP layer costs a little speed for slightly higher draft acceptance. More detail: WHIRL quantization guide §6.

Measured performance (AMD Radeon AI PRO R9700, 32 GB)

WHIRL v0.1.0 vs llama.cpp b11214 (ROCm), same GGUF file, same prompts, greedy decoding in both, one model process on the GPU at a time, measured 2026-10-03. Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram, and of the flag sets tried; named in the table). "no MTP" = plain decoding without any speculation. Full methodology, raw-data locations and the rows where WHIRL does not lead: WHIRL benchmarks. The R9700 in this machine is connected as a USB4 eGPU; this does not affect prefill or decode (they stay in VRAM) but slows model loading and the host-RAM / SSD KV restores.

Decode (tok/s), WHIRL default = MTP + n-gram

ScenarioWHIRL MTP + n-gramllama.cpp fastest (config)Ratio
7 Chinese coding prompts (Chinese question, Chinese explanation, English code), 800 tokens, server244.5118.8 (plain)2.06×
File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output, server618.1239.3 (draft-mtp,ngram-mod, n-max 1)2.58×
128 tokens after a 16k-token code context, server301.5107.5 (plain)2.80×

Decode (tok/s), no MTP (plain decoding in both engines)

ScenarioWHIRL no MTPllama.cpp no MTPRatio
7 Chinese coding prompts, server166.7118.81.40×
File editing, server163.3117.11.40×
After a 16k-token context, server150.5107.51.40×
CLI bench (whirl bench 256 tokens vs llama-bench tg256)193.6119.71.62×

Prefill (tok/s, each engine's own bench tool, KV f16)

Prompt tokensWHIRLllama.cpp (flags)Ratio
882,5462,115 (-ub 512)1.20×
2,04811,7054,876 (-ub 4096 -b 4096)2.40×
8,19210,8584,637 (-ub 2048)2.34×
32,7687,9783,778 (-ub 2048)2.11×
131,0723,7822,239 (-ub 2048)1.69×

Server

TestWHIRLllama-serverRatio
4 concurrent requests, aggregate tok/s (256 tokens each, prefill included) — WHIRL MTP + n-gram vs llama.cpp fastest (plain)396.7191.92.07×
Same, WHIRL no MTP vs llama.cpp no MTP324.9191.91.69×
TTFT, ~12.2k-token system prompt, cold1.314 s3.396 s2.6× faster
TTFT, new conversation reusing that ~12.2k system prompt0.064 s0.259 s4.1× faster
TTFT, ~26.4k-token system prompt, cold3.294 s8.151 s2.5× faster
TTFT, evicted ~27.4k-token session restored from host RAM0.265 s0.950 s3.6× faster
TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier)0.298 sN/A (no equivalent feature)—

VRAM at equal context (CLI bench, ~33k tokens, KV f16): 22.5 GiB in both engines. The WHIRL server deliberately fills free VRAM with its KV pool (414,464 tokens of f16 KV with this file).

These are speed measurements only. In the benchmark runs WHIRL's speculative output was identical to its plain greedy output, so the MTP + n-gram speed-up does not change the generated text.

Quality

  • Perplexity (llama.cpp b11214 ROCm llama-perplexity, -c 2048 -b 512 -ub 512 -fa on -ngl 99, 51 chunks of a 219k-character mixed Chinese/English source-code corpus — the same corpus and settings as our Swift MXFP4 card — R9700, measured 2026-10-03). The usual wikitext-2 test set was not available on the measurement machine.

    filesize (bytes)PPLvs Q4_K_M
    this repo: Ornith-1.5-35B-A3B-MXFP4.gguf19,819,767,1365.5192 ± 0.0653+3.7%
    reference: official Ornith-1.5-35B-Q4_K_M.gguf (ornith-ai GGUF repo)21,713,462,8485.3206 ± 0.0619—

    Interpretation: on this corpus our MXFP4 file has 3.7% higher perplexity than the official Q4_K_M, which is about 10% larger. The ± values are each run's own statistical error; since both runs score the same tokens, a paired per-chunk comparison is the fairer test: MXFP4 is worse on 42 of 51 chunks, by 0.037 ± 0.006 nats/token on average (about 6 standard errors) — a small but measurable difference. The file was built for WHIRL's speed (see the recipe), not to beat Q4_K_M on size-for-quality. No BF16 baseline and no KL-divergence were measured (the BF16 GGUF used for quantization had been deleted from the machine before this test), so the absolute loss against the unquantized model is unknown. Perplexity on one corpus is only a proxy; it was not checked on downstream tasks. These numbers are not comparable with perplexities of other models (e.g. our Swift card), since tokenizers and training differ.

Other checks:

  • In WHIRL, MTP + n-gram, MTP-only and plain greedy decoding give identical text on 7 test prompts (Chinese / English, thinking on and off, 2k / 8k / 32k-token contexts). This shows the speculative paths are exact; it is not a measure of quantization loss.
  • Image input: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, with the exact text, by both WHIRL 0.1.0 (whirl chat --mmproj --image) and llama.cpp b11214 llama-mtmd-cli with this projector (single smoke test, 2026-10-03).

For the model's own evaluations, see the upstream card. A 4.46-bpw MXFP4 file is expected to lose some accuracy against BF16; if that matters for your use, compare against the official Q8_0 / BF16 GGUFs. If you run only llama.cpp and care most about quality, the official Q4_K_M has slightly lower perplexity (about 3.7%) at about 10% larger size.

License

The upstream model is published under the MIT License (declared in the upstream model card metadata). These quantized files are a derivative of it and are distributed under the same license:

  • LICENSE: the MIT License text. At the time of this release the upstream repository declares license: mit but contains no LICENSE file (its license_link returns "Entry not found"), so this is the standard MIT License text with the Ornith team named as copyright holder.
  • NOTICE: attribution and a description of the changes made in this redistribution.

The upstream card states that the Ornith family was developed on top of Qwen3.5 and Gemma 4; see the upstream card for the provenance of the base weights.

Credits

Upstream citation:

@misc{ornith_1_5,
    title = {{Ornith-1.5}: From Self-Scaffolding to Self-Improvement},
    url = {https://ornith.ai/ornith_1_5.html},
    author = {{Ornith Team}},
    year = {2026}
}

繁體中文摘要

  • 這是 ornith-ai/Ornith-1.5-35B-A3B 的非官方社群 MXFP4 GGUF 量化,與 Ornith 團隊無關。MoE 模型,總參數約 35B、每個 token 約 3B 啟用;保留 MTP 層(Q8_0),可做自我推測解碼;另附 F16 視覺投影檔。
  • 專為 WHIRL(Apache-2.0,AMD Radeon AI PRO R9700 的原生 Windows C++/HIP 推論引擎,v0.1.0 下載)製作;也能在 llama.cpp 執行(已用 b11214 測試:llama-bench、llama-server、llama-mtmd-cli)。
  • 量化:專家、共享專家與所有稠密矩陣 MXFP4;token_embd、ssm_alpha / beta、整個 MTP 層 Q8_0;router 與 norm 為 F32;輸出頭 Q6_K;不使用 imatrix。平均 4.46 bpw,載入 18.45 GiB。
  • R9700(USB4 外接)WHIRL v0.1.0 對 llama.cpp b11214 最快設定:中文寫程式解碼 244.5 對 118.8 tok/s(2.06×)、檔案編輯 618.1 對 239.3(2.58×)、16k 上下文後 301.5 對 107.5(2.80×);不開 MTP 時 166.7 對 118.8(1.40×);prefill 8k 10,858 對 4,637 tok/s(2.34×);4 個並行請求總吞吐 396.7 對 191.9 tok/s(2.07×)。
  • 品質:困惑度(llama.cpp b11214、中英混合程式碼語料 21.9 萬字元、2048 ctx,與 Swift 卡相同方法)本檔 5.5192 ± 0.0653,官方 Q4_K_M 5.3206 ± 0.0619——本檔高 3.7%(檔案小約 9%),逐段配對比較 51 段中 42 段較差(平均差 0.037 ± 0.006 nats/token,約 6 個標準誤),屬可量到的小幅損失;若只用 llama.cpp 且最在意品質,官方 Q4_K_M 的困惑度略低(約 3.7%),但檔案大約 10%。未量 BF16 基準與 KL(BF16 已刪),對未量化模型的絕對損失未知;困惑度也不能和其他模型(如 Swift)比較。另確認 WHIRL 推測解碼與 plain greedy 輸出逐字相同(7 題)、看圖冒煙測試(WHIRL 與 llama-mtmd-cli)正確。
  • 授權:MIT(上游 metadata 宣告;上游 repo 沒有 LICENSE 檔,本 repo 附標準 MIT 條文);詳見 LICENSE / NOTICE。
conversational
endpoints_compatible
gguf
llama.cpp
mxfp4
rdna4
text-generation
whirl