tsaipifong/Swift-1.5-Qwen3.8-27b-MXFP4-GGUF

Model

Swift-1.5-Qwen3.8-27b MXFP4 GGUF

2

9 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Swift-1.5-Qwen3.8-27b MXFP4 GGUF

MXFP4 GGUF quantizations of ukisai/Swift-1.5-Qwen3.8-27b, with the MTP (multi-token prediction) layer kept for self-speculative decoding.

Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by UkisAI or Alibaba Cloud. The names "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the origin of the weights. For the official model, official quantizations, evaluations and enterprise licensing, see the upstream repository.

Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient post-trained derivative of Qwen/Qwen3.8-27B. Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates model quality.

Vision is supported through the separate projector file mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf (F16, converted from the Swift-1.5 BF16 weights with convert_hf_to_gguf.py --mmproj --outtype f16). It works with any of the three variants.

Best performance with WHIRL

These files were tuned for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. They are standard GGUF and also run in llama.cpp (tested with b11214, see below).

Also from us: Ornith-1.5-35B-A3B MXFP4 GGUF — a mixture-of-experts model (~3B active of 35B) quantized the same way, with MTP layer and vision projector.

Files

The three language-model files contain the same weights except the output head (output.weight).

VariantFileOutput headSize (bytes)SHA256Recommendation
ASwift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.ggufQ6_K15,815,469,280f12d8dee7d677a0583b243d65c5e75e853b1839dc61e2c101da090bb6163065fRecommended default
CSwift-1.5-Qwen3.8-27B-MXFP4-C-outQ4_K.ggufQ4_K15,487,686,880221aa39e7f7c003ae94903ea686d33d6bb187ed8b943955fb21d0790c7cf6b4aSlightly faster on llama.cpp and in plain (non-speculative) decoding; lower-precision head
BSwift-1.5-Qwen3.8-27B-MXFP4-B-outQ8_0.ggufQ8_016,123,386,080bb075d1cf83352d38751a799feffc0e0c91bb8ccb44a14551c0f485ffd65c72aNot recommended for speed (slowest with MTP); highest-precision head
visionmmproj-Swift-1.5-Qwen3.8-27B-F16.gguf— (vision projector, F16)927,607,58471988a54d379191e1abc26c2b79363b87ef8b3a584ebee882c28a1682ddbfe62Needed only for image input

Quantization recipe

Tensor types (verified by dumping the quantized files):

Tensor groupType
All large matrices of the 64 main layers: attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out, ffn_gate/up/downMXFP4
token_embdQ8_0
ssm_alpha, ssm_beta (48 linear-attention layers × 2)Q8_0
MTP layer blk.64 (all of attn_q/k/v/output, ffn_gate/up/down, nextn.eh_proj)Q8_0
Norms, ssm_a, ssm_conv1d, ssm_dt.bias, nextn.enorm/hnorm/shared_head_normF32
output (LM head)Q6_K (A) / Q8_0 (B) / Q4_K (C)

Steps

  1. BF16 safetensors → BF16 GGUF with llama.cpp convert_hf_to_gguf.py <Swift-1.5 BF16 dir> --outtype bf16 (866 tensors; qwen35.nextn_predict_layers = 1, i.e. the MTP layer is kept). (llama.cpp b10645-47-gef6876693, commit ef6876693.)
  2. Quantize with llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build.

tt_common.txt (passed with --tensor-type-file; the first matching rule wins):

blk\.64\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate=mxfp4
ffn_up=mxfp4
ffn_down=mxfp4
ssm_out=mxfp4

Exact command for variant A (B used --output-tensor-type q8_0, C used --output-tensor-type q4_k):

llama-quantize --imatrix ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf \
  --tensor-type-file tt_common.txt --output-tensor-type q6_k \
  swift-1.5-27b-bf16.gguf Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf MXFP4_MOE 16

Notes:

  • ftype MXFP4_MOE: this llama.cpp build has no dense-model MXFP4 ftype, so MXFP4_MOE is used as the base type and the per-tensor rules above decide every tensor. The GGUF metadata therefore reports general.file_type = 38 (MXFP4_MOE) even though the model is dense.

  • Importance matrix: the imatrix was made and published by bartowski — bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF (ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, 496 entries, 582 chunks). Thank you! Note that mainline llama.cpp ignores the imatrix for MXFP4 and for Q8_0, so in these files it only affects the K-quant output head (Q6_K in A, Q4_K in C). B is effectively imatrix-free.

  • Why ssm_alpha / ssm_beta are Q8_0 and not F32: a first build kept them in F32. With F32 weights for these small matrices, WHIRL computes single-token decode and multi-token MTP verification through different kernels, so MTP output was no longer bit-identical to plain greedy decoding, and the automatic draft length collapsed. Q8_0 has a fused exact path there, restoring bit-identical MTP output. This choice was made for WHIRL; llama.cpp loads and runs the files normally (see below).

  • Differences from the reference Qwen3.8-27B MXFP4 GGUF we compared against (FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF): that file is MXFP4 everywhere, including token_embd, ssm_alpha/beta and the MTP layer, with a Q6_K head. Here those tensors are Q8_0, which costs about 1–2% prefill speed in our measurements.

  • Perplexity (llama.cpp b11214 llama-perplexity, -c 2048 -b 512 -ub 512 -fa on, 51 chunks of a 219 KB mixed Chinese/English source-code corpus, R9700):

    filePPL
    A (Q6_K head)3.8271 ± 0.0394
    B (Q8_0 head)3.8289 ± 0.0395
    C (Q4_K head)3.8386 ± 0.0396
    reference: Qwen3.8-27B MXFP4 (FreedomAISVR)3.7952 ± 0.0384

    The three heads are within noise of each other (C is +0.3%). Swift is a different fine-tune, so its PPL is not directly comparable with the Qwen base quant (a post-trained model typically shifts PPL on generic text). No BF16 baseline / KL-divergence was measured. Beyond PPL, quality was only checked by coherent zh/en greedy output.

Usage with llama.cpp (MTP)

The MTP layer is kept, so llama.cpp can use it for self-speculative decoding with --spec-type draft-mtp. This is the exact llama-server command used for our llama.cpp check (b11214, ROCm):

llama-server -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --device ROCm0 -ngl 99 -c 131072 -np 1 -fa on \
  --port 1234 --no-webui --load-mode none --jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-n-min 0 --spec-draft-p-min 0.3

(Our machine has two GPUs, so the run also set HIP_VISIBLE_DEVICES=1 and GGML_CUDA_NO_PINNED=1; adjust --device / -c to your hardware.) Without the --spec-* flags the model runs as a plain model; llama.cpp then logs the MTP-layer tensors as "unused", which is expected. The loader also prints unknown type mxfp4 as a warning; this appears with other MXFP4 GGUFs as well and loading succeeds.

Image input (tested with llama-mtmd-cli b11214: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, including the exact text):

llama-mtmd-cli -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf   --image photo.png -ngl 99 -c 8192 -p "Describe this image."

llama-server accepts the same --mmproj flag for image input over the OpenAI-compatible API.

Sampling recommended by the upstream card: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0. Thinking is on by default in the chat template; our llama.cpp check disabled it per request with "chat_template_kwargs": {"enable_thinking": false}.

Measured performance (AMD Radeon AI PRO R9700, 32 GB)

Single-machine measurements; speed only, not quality evaluations. The R9700 in this machine is connected as a USB4 eGPU (this slows model loading and host-RAM / SSD KV restores, not prefill or decode).

llama.cpp b11214 (ROCm, llama-server, flags as above; 2026-10-01, single request, short prompts)

Greedy-equivalent sampling (top_k 1), thinking off, 4 short prompts (2 Traditional Chinese, 2 English), max 400 output tokens; prefill measured on a 2,022-token prompt.

llama.cpp b11214A (Q6_K head)B (Q8_0 head)C (Q4_K head)
Plain decode, tok/s (mean of 4 prompts)33.332.833.8
MTP decode, tok/s (--spec-draft-n-max 3)58.155.960.7
MTP draft acceptance73.6% (846/1150)72.2% (840/1164)75.0% (856/1141)
Prefill, 2k prompt, tok/s121512121206
  • Peak dedicated VRAM of llama-server at -c 131072 -np 1: 23.65 GiB.
  • On the English prompts MTP output matched plain output; on the two Chinese prompts the outputs diverged part-way (llama.cpp's batched numerics are not bit-identical between plain and speculative decoding; this is normal).

WHIRL v0.1.0 vs llama.cpp b11214 (variant A)

From the WHIRL v0.1.0 release benchmark (2026-10-03): same GGUF, same prompts, greedy decoding in both engines, one model process on the GPU at a time, median of 3 rounds (file editing / concurrency: 2). Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram and of the flag sets tried; named). "no MTP" = plain decoding without any speculation. llama.cpp here uses -ub 1024 and its best MTP settings (--spec-draft-n-max 4), so its MTP number is a little higher than in the 2026-10-01 check above (60.8 vs 58.1 tok/s).

Decode (tok/s), WHIRL default = MTP + n-gram

Scenario (server)WHIRL MTP + n-gramWHIRL MTPllama.cpp fastest (config)Ratio
7 Chinese coding prompts (Chinese question, Chinese explanation, English code), thinking off, 800 tokens107.5107.360.8 (draft-mtp, n-max 4)1.77×
File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output328.3181.1144.8 (draft-mtp,ngram-mod, n-max 3)2.27×
128 tokens after a 16k-token code context69.572.463.5 (draft-mtp, n-max 4, V cache q8_0)1.09×

After a 16k context WHIRL's lead is small, and its MTP-only mode is faster than its default MTP + n-gram on this model.

Decode (tok/s), no MTP (plain decoding in both engines)

ScenarioWHIRL no MTPllama.cpp no MTPRatio
7 Chinese coding prompts, server37.533.41.13×
File editing, server37.033.01.12×
After a 16k-token context, server35.231.41.12×
CLI bench (whirl bench 256 tokens vs llama-bench tg256)39.433.41.18×

Prefill (tok/s, each engine's own bench tool, KV f16)

Prompt tokensWHIRLllama.cpp (-ub 1024 unless noted)Ratio
881,515928.6 (-ub 512)1.63×
2,0483,5091,3852.53×
8,1923,2781,3382.45×
32,7682,5951,1742.21×
131,0721,405791.31.78×

Server

TestWHIRLllama-serverRatio
4 concurrent requests, aggregate tok/s — WHIRL MTP + n-gram vs llama.cpp fastest (plain)198.665.83.02×
Same, WHIRL no MTP vs llama.cpp no MTP108.665.81.65×
TTFT, ~12.2k-token system prompt, cold4.091 s10.650 s2.6× faster
TTFT, new conversation reusing that system prompt0.102 s0.446 s4.4× faster
TTFT, evicted ~27.4k-token session restored from host RAM0.668 s2.195 s3.3× faster
TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier)0.729 sN/A (no equivalent feature)—
Image encode with the F16 projector, warm, 7 images (220–2,040 image tokens)16.8–253.7 ms58–1,330 ms3.5–5.2× faster

VRAM: at equal context WHIRL's CLI uses more VRAM than llama-bench on this model (23.0 vs 17.8 GiB at ~33k tokens; 26.1 vs 22.7 GiB at ~131k), for larger prefill buffers, the MTP block and the draft head. The WHIRL server deliberately fills free VRAM with its KV pool.

Reading the tables: plain decode (no speculation) is bounded by memory bandwidth — every token reads all ~15.8 GB of weights, and the R9700's ~640 GB/s caps it near 38–40 tok/s in both engines, so WHIRL's no-MTP lead is small (1.12–1.18×). Most of WHIRL's decode lead comes from MTP + n-gram speculative decoding, whose output is identical to plain greedy decoding. n-gram drafts matter most when the output repeats the input (file editing).

Variant comparison in WHIRL (pre-release build, 2026-10-01, mixed 19-prompt protocol)

Used to choose the recommended variant. Server defaults, 19 prompts mixing zh thinking off, zh thinking on and code-edit prompts, 2 interleaved rounds, greedy-equivalent sampling. Because the code-edit prompts are included, the MTP + n-gram row is an average over generation and editing prompts and is not comparable with the scenario numbers above (the zh thinking-off group alone was 106.0 tok/s for A in this run; 107.5 in v0.1.0). "Qwen MXFP4" is the reference Qwen3.8-27B MXFP4 GGUF mentioned above, measured the same way.

WHIRL (pre-release), decode tok/sA (Q6_K head)B (Q8_0 head)C (Q4_K head)Qwen3.8-27B MXFP4
Plain (no MTP, no speculation)36.836.137.536.9
MTP117.098.0112.0118.1
MTP + n-gram — 19-prompt mixed average incl. code-edit prompts157.2143.6157.5156.9
MTP draft acceptance (MTP / MTP + n-gram)65.4% / 66.3%73.2% / 72.9%68.3% / 69.1%60.6% / 62.2%
Prefill tok/s, 2k / 8k / 32k3164 / 3302 / 25703181 / 3291 / 25633138 / 3307 / 25723223 / 3338 / 2602
Plain decode (no MTP) after 16k context34.634.135.234.7

Why A is recommended: B's 8-bit head is the most expensive part of every draft step, so it is the slowest with MTP despite the highest acceptance rate. C is fastest without speculation and on llama.cpp, but in WHIRL A gets a dedicated Q6_K-head path that made it 4.5% faster than C with MTP (and equal with MTP + n-gram) in this comparison, while keeping a more precise head.

Thinking length vs Qwen3.8-27B MXFP4 (small sample)

Swift A vs the reference Qwen3.8-27B MXFP4 GGUF, both served by WHIRL (MTP + n-gram) on the R9700. Settings: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0 (official sampling for both), thinking on (template default), max_tokens 16,384, 9 prompts × seeds 1 and 2 (n = 18 runs per model). Thinking tokens = reasoning_content counted with the model's tokenizer; wall = full non-streaming request time. * = hit the 16,384-token cap (value is a lower bound).

PromptSwift think tokens (s1 / s2)Swift wall sQwen think tokens (s1 / s2)Qwen wall s
你好24 / 240.5 / 0.621 / 210.5 / 0.5
Hi16 / 160.4 / 0.446 / 431.1 / 0.8
zh: Python LRU12639* / 10619161.5 / 135.216382* / 16383*189.8 / 199.2
zh: TypeScript hook2148 / 769547.4 / 109.511795 / 10961179.8 / 166.6
zh: SQL report8961 / 7658174.5 / 138.612893* / 13540*210.9 / 210.4
zh: bug fix2842 / 238651.0 / 40.49598 / 5944140.7 / 97.7
en: ISO-8601 duration parser16384* / 16310*206.5 / 203.616382* / 16384*212.7 / 132.3
en: Go bounded queue16384* / 16384*236.1 / 230.116383* / 16384*245.6 / 240.8
en: JS debounce bug4020 / 478957.6 / 65.51855 / 449530.5 / 68.2
  • Totals (18 runs): Swift 129,299 thinking tokens / 1,859 s / 5 capped; Qwen 169,510 / 2,328 s / 8 capped. Because Qwen hit the cap more often, its totals are lower bounds.
  • 4 Chinese coding prompts: Swift 54,948 vs Qwen 97,496 thinking tokens (−44%), wall 858 s vs 1,395 s (−38%), capped 1 vs 4.
  • 3 English prompts: similar (Swift 74,271 vs Qwen 71,883 tokens); on two of them both models were still thinking at 16k.
  • A capped Swift run (en ISO-8601, seed 1) was re-run and inspected: repeated 32-gram ratio 0.5% (first half) / 1.3% (second half) — genuine long deliberation, not a degenerate repetition loop.
  • This is a small, single-machine sample with a 16k cap; it does not measure answer correctness. See the upstream card for UkisAI's controlled evaluation.

License

These files are a quantized (Object-form) Derivative Work of Swift 1.5 Qwen3.8-27B and are distributed under the same terms as upstream:

  • Swift Open License v1.0 for UkisAI's Swift Contribution — see LICENSE (unmodified copy of the upstream LICENSE).
  • Apache License 2.0 for the Qwen3.8-27B base model — see LICENSE-APACHE-2.0.
  • NOTICE: the upstream NOTICE verbatim, plus a "Changes in this redistribution" section describing the conversion and quantization.

Notice for redistributors, reproduced verbatim from the appendix of the Swift Open License v1.0:

   Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open
   License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE).
   Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.

Commercial-use threshold (summary, not legal advice — the LICENSE text governs): personal, research, educational, evaluation and commercial use are permitted for individuals and organizations whose gross revenue, together with all affiliates (entities controlling, controlled by, or under common control with them), is below US$1,000,000 over the most recently completed fiscal year. Commercial use by an entity at or above that threshold is not licensed under the Swift Open License and requires a separate Swift Enterprise License from UkisAI (contact). The threshold does not apply to qualified non-profit organizations' non-commercial or research use. The license terminates automatically on non-compliance.

Base model: Qwen3.8-27B is Copyright 2026 Alibaba Cloud and licensed under the Apache License 2.0. Per Section 6 of the Swift Open License, nothing in it limits your rights in the Qwen3.8-27B base model under Apache 2.0; the commercial limitation applies only to the Swift Contribution.

Credits

  • Alibaba Cloud / Qwen team — Qwen3.8-27B, the base model.
  • UkisAI — Swift 1.5 Qwen3.8-27B, the post-trained model these files are made from.
  • bartowski — the importance matrix used for the K-quant output heads.
  • llama.cpp / ggml contributors — conversion, quantization and inference tooling.

Upstream citation:

@misc{swift-1.5-qwen3.8-27b,
  title  = {Swift 1.5 Qwen3.8-27B},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}

繁體中文摘要

  • 這是 ukisai/Swift-1.5-Qwen3.8-27b 的非官方社群 MXFP4 GGUF 量化,與 UkisAI 無關;「Swift」「UkisAI」商標僅用來說明權重來源。另附視覺投影檔 mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf(F16),搭配任一版本即可看圖(llama.cpp --mmproj)。
  • 保留 MTP 層(Q8_0),llama.cpp 可用 --spec-type draft-mtp 做自我推測解碼。
  • 三個變體只差輸出頭:A = Q6_K(推薦)、C = Q4_K(在 llama.cpp 與不開推測解碼時稍快)、B = Q8_0(開 MTP 時最慢,不建議)。
  • 大矩陣 MXFP4;token_embd、ssm_alpha / beta、MTP 層 Q8_0;imatrix 來自 bartowski,但 mainline 量 MXFP4 / Q8_0 時不使用 imatrix,只影響 K-quant 輸出頭。
  • AMD Radeon AI PRO R9700、llama.cpp b11214:A 的 plain decode 33.3 tok/s、MTP 58.1 tok/s、2k prefill 1215 tok/s。
  • 最佳效能請用 WHIRL(Apache-2.0 開源,v0.1.0):A 在中文寫程式解碼 107.5 對 llama.cpp 最快設定 60.8 tok/s(1.77×)、檔案編輯 328.3 對 144.8(2.27×)、16k 上下文後 69.5 對 63.5(1.09×);不開 MTP 37.5 對 33.4(1.13×);prefill 8k 3,278 對 1,338 tok/s(2.45×)。舊的 157.2 tok/s 是 19 題混合(含程式碼編輯題)的平均,不能和單一情境數字直接比。
  • 另有 MoE 模型 Ornith-1.5-35B-A3B MXFP4 GGUF(每 token 約 3B 啟用)。
  • 困惑度(中英混合程式碼語料,2048 ctx):A 3.8271、B 3.8289、C 3.8386,三者在誤差內;Swift 是另一個微調模型,不宜直接和 Qwen 量化版(3.7952)比較。
  • 小樣本思考長度比較(16k 上限、9 題 × 2 seed):中文寫程式 4 題 Swift 思考 token 比 Qwen3.8-27B MXFP4 少 44%、時間少 38%。
  • 授權:Swift Open License v1.0(年營收含關係企業達 100 萬美元以上的商業使用需另取得授權)+ 基底模型 Apache 2.0;詳見 LICENSE / NOTICE。
conversational
endpoints_compatible
gguf
imatrix
llama.cpp
mxfp4
qwen3.8
rdna4
text-generation
whirl

tsaipifong/Swift-1.5-Qwen3.8-27b-MXFP4-GGUF

Model

Swift-1.5-Qwen3.8-27b MXFP4 GGUF

2

9 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Swift-1.5-Qwen3.8-27b MXFP4 GGUF

MXFP4 GGUF quantizations of ukisai/Swift-1.5-Qwen3.8-27b, with the MTP (multi-token prediction) layer kept for self-speculative decoding.

Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by UkisAI or Alibaba Cloud. The names "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the origin of the weights. For the official model, official quantizations, evaluations and enterprise licensing, see the upstream repository.

Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient post-trained derivative of Qwen/Qwen3.8-27B. Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates model quality.

Vision is supported through the separate projector file mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf (F16, converted from the Swift-1.5 BF16 weights with convert_hf_to_gguf.py --mmproj --outtype f16). It works with any of the three variants.

Best performance with WHIRL

These files were tuned for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. They are standard GGUF and also run in llama.cpp (tested with b11214, see below).

Also from us: Ornith-1.5-35B-A3B MXFP4 GGUF — a mixture-of-experts model (~3B active of 35B) quantized the same way, with MTP layer and vision projector.

Files

The three language-model files contain the same weights except the output head (output.weight).

VariantFileOutput headSize (bytes)SHA256Recommendation
ASwift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.ggufQ6_K15,815,469,280f12d8dee7d677a0583b243d65c5e75e853b1839dc61e2c101da090bb6163065fRecommended default
CSwift-1.5-Qwen3.8-27B-MXFP4-C-outQ4_K.ggufQ4_K15,487,686,880221aa39e7f7c003ae94903ea686d33d6bb187ed8b943955fb21d0790c7cf6b4aSlightly faster on llama.cpp and in plain (non-speculative) decoding; lower-precision head
BSwift-1.5-Qwen3.8-27B-MXFP4-B-outQ8_0.ggufQ8_016,123,386,080bb075d1cf83352d38751a799feffc0e0c91bb8ccb44a14551c0f485ffd65c72aNot recommended for speed (slowest with MTP); highest-precision head
visionmmproj-Swift-1.5-Qwen3.8-27B-F16.gguf— (vision projector, F16)927,607,58471988a54d379191e1abc26c2b79363b87ef8b3a584ebee882c28a1682ddbfe62Needed only for image input

Quantization recipe

Tensor types (verified by dumping the quantized files):

Tensor groupType
All large matrices of the 64 main layers: attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out, ffn_gate/up/downMXFP4
token_embdQ8_0
ssm_alpha, ssm_beta (48 linear-attention layers × 2)Q8_0
MTP layer blk.64 (all of attn_q/k/v/output, ffn_gate/up/down, nextn.eh_proj)Q8_0
Norms, ssm_a, ssm_conv1d, ssm_dt.bias, nextn.enorm/hnorm/shared_head_normF32
output (LM head)Q6_K (A) / Q8_0 (B) / Q4_K (C)

Steps

  1. BF16 safetensors → BF16 GGUF with llama.cpp convert_hf_to_gguf.py <Swift-1.5 BF16 dir> --outtype bf16 (866 tensors; qwen35.nextn_predict_layers = 1, i.e. the MTP layer is kept). (llama.cpp b10645-47-gef6876693, commit ef6876693.)
  2. Quantize with llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build.

tt_common.txt (passed with --tensor-type-file; the first matching rule wins):

blk\.64\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate=mxfp4
ffn_up=mxfp4
ffn_down=mxfp4
ssm_out=mxfp4

Exact command for variant A (B used --output-tensor-type q8_0, C used --output-tensor-type q4_k):

llama-quantize --imatrix ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf \
  --tensor-type-file tt_common.txt --output-tensor-type q6_k \
  swift-1.5-27b-bf16.gguf Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf MXFP4_MOE 16

Notes:

  • ftype MXFP4_MOE: this llama.cpp build has no dense-model MXFP4 ftype, so MXFP4_MOE is used as the base type and the per-tensor rules above decide every tensor. The GGUF metadata therefore reports general.file_type = 38 (MXFP4_MOE) even though the model is dense.

  • Importance matrix: the imatrix was made and published by bartowski — bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF (ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, 496 entries, 582 chunks). Thank you! Note that mainline llama.cpp ignores the imatrix for MXFP4 and for Q8_0, so in these files it only affects the K-quant output head (Q6_K in A, Q4_K in C). B is effectively imatrix-free.

  • Why ssm_alpha / ssm_beta are Q8_0 and not F32: a first build kept them in F32. With F32 weights for these small matrices, WHIRL computes single-token decode and multi-token MTP verification through different kernels, so MTP output was no longer bit-identical to plain greedy decoding, and the automatic draft length collapsed. Q8_0 has a fused exact path there, restoring bit-identical MTP output. This choice was made for WHIRL; llama.cpp loads and runs the files normally (see below).

  • Differences from the reference Qwen3.8-27B MXFP4 GGUF we compared against (FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF): that file is MXFP4 everywhere, including token_embd, ssm_alpha/beta and the MTP layer, with a Q6_K head. Here those tensors are Q8_0, which costs about 1–2% prefill speed in our measurements.

  • Perplexity (llama.cpp b11214 llama-perplexity, -c 2048 -b 512 -ub 512 -fa on, 51 chunks of a 219 KB mixed Chinese/English source-code corpus, R9700):

    filePPL
    A (Q6_K head)3.8271 ± 0.0394
    B (Q8_0 head)3.8289 ± 0.0395
    C (Q4_K head)3.8386 ± 0.0396
    reference: Qwen3.8-27B MXFP4 (FreedomAISVR)3.7952 ± 0.0384

    The three heads are within noise of each other (C is +0.3%). Swift is a different fine-tune, so its PPL is not directly comparable with the Qwen base quant (a post-trained model typically shifts PPL on generic text). No BF16 baseline / KL-divergence was measured. Beyond PPL, quality was only checked by coherent zh/en greedy output.

Usage with llama.cpp (MTP)

The MTP layer is kept, so llama.cpp can use it for self-speculative decoding with --spec-type draft-mtp. This is the exact llama-server command used for our llama.cpp check (b11214, ROCm):

llama-server -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --device ROCm0 -ngl 99 -c 131072 -np 1 -fa on \
  --port 1234 --no-webui --load-mode none --jinja \
  --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-n-min 0 --spec-draft-p-min 0.3

(Our machine has two GPUs, so the run also set HIP_VISIBLE_DEVICES=1 and GGML_CUDA_NO_PINNED=1; adjust --device / -c to your hardware.) Without the --spec-* flags the model runs as a plain model; llama.cpp then logs the MTP-layer tensors as "unused", which is expected. The loader also prints unknown type mxfp4 as a warning; this appears with other MXFP4 GGUFs as well and loading succeeds.

Image input (tested with llama-mtmd-cli b11214: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, including the exact text):

llama-mtmd-cli -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf   --image photo.png -ngl 99 -c 8192 -p "Describe this image."

llama-server accepts the same --mmproj flag for image input over the OpenAI-compatible API.

Sampling recommended by the upstream card: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0. Thinking is on by default in the chat template; our llama.cpp check disabled it per request with "chat_template_kwargs": {"enable_thinking": false}.

Measured performance (AMD Radeon AI PRO R9700, 32 GB)

Single-machine measurements; speed only, not quality evaluations. The R9700 in this machine is connected as a USB4 eGPU (this slows model loading and host-RAM / SSD KV restores, not prefill or decode).

llama.cpp b11214 (ROCm, llama-server, flags as above; 2026-10-01, single request, short prompts)

Greedy-equivalent sampling (top_k 1), thinking off, 4 short prompts (2 Traditional Chinese, 2 English), max 400 output tokens; prefill measured on a 2,022-token prompt.

llama.cpp b11214A (Q6_K head)B (Q8_0 head)C (Q4_K head)
Plain decode, tok/s (mean of 4 prompts)33.332.833.8
MTP decode, tok/s (--spec-draft-n-max 3)58.155.960.7
MTP draft acceptance73.6% (846/1150)72.2% (840/1164)75.0% (856/1141)
Prefill, 2k prompt, tok/s121512121206
  • Peak dedicated VRAM of llama-server at -c 131072 -np 1: 23.65 GiB.
  • On the English prompts MTP output matched plain output; on the two Chinese prompts the outputs diverged part-way (llama.cpp's batched numerics are not bit-identical between plain and speculative decoding; this is normal).

WHIRL v0.1.0 vs llama.cpp b11214 (variant A)

From the WHIRL v0.1.0 release benchmark (2026-10-03): same GGUF, same prompts, greedy decoding in both engines, one model process on the GPU at a time, median of 3 rounds (file editing / concurrency: 2). Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram and of the flag sets tried; named). "no MTP" = plain decoding without any speculation. llama.cpp here uses -ub 1024 and its best MTP settings (--spec-draft-n-max 4), so its MTP number is a little higher than in the 2026-10-01 check above (60.8 vs 58.1 tok/s).

Decode (tok/s), WHIRL default = MTP + n-gram

Scenario (server)WHIRL MTP + n-gramWHIRL MTPllama.cpp fastest (config)Ratio
7 Chinese coding prompts (Chinese question, Chinese explanation, English code), thinking off, 800 tokens107.5107.360.8 (draft-mtp, n-max 4)1.77×
File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output328.3181.1144.8 (draft-mtp,ngram-mod, n-max 3)2.27×
128 tokens after a 16k-token code context69.572.463.5 (draft-mtp, n-max 4, V cache q8_0)1.09×

After a 16k context WHIRL's lead is small, and its MTP-only mode is faster than its default MTP + n-gram on this model.

Decode (tok/s), no MTP (plain decoding in both engines)

ScenarioWHIRL no MTPllama.cpp no MTPRatio
7 Chinese coding prompts, server37.533.41.13×
File editing, server37.033.01.12×
After a 16k-token context, server35.231.41.12×
CLI bench (whirl bench 256 tokens vs llama-bench tg256)39.433.41.18×

Prefill (tok/s, each engine's own bench tool, KV f16)

Prompt tokensWHIRLllama.cpp (-ub 1024 unless noted)Ratio
881,515928.6 (-ub 512)1.63×
2,0483,5091,3852.53×
8,1923,2781,3382.45×
32,7682,5951,1742.21×
131,0721,405791.31.78×

Server

TestWHIRLllama-serverRatio
4 concurrent requests, aggregate tok/s — WHIRL MTP + n-gram vs llama.cpp fastest (plain)198.665.83.02×
Same, WHIRL no MTP vs llama.cpp no MTP108.665.81.65×
TTFT, ~12.2k-token system prompt, cold4.091 s10.650 s2.6× faster
TTFT, new conversation reusing that system prompt0.102 s0.446 s4.4× faster
TTFT, evicted ~27.4k-token session restored from host RAM0.668 s2.195 s3.3× faster
TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier)0.729 sN/A (no equivalent feature)—
Image encode with the F16 projector, warm, 7 images (220–2,040 image tokens)16.8–253.7 ms58–1,330 ms3.5–5.2× faster

VRAM: at equal context WHIRL's CLI uses more VRAM than llama-bench on this model (23.0 vs 17.8 GiB at ~33k tokens; 26.1 vs 22.7 GiB at ~131k), for larger prefill buffers, the MTP block and the draft head. The WHIRL server deliberately fills free VRAM with its KV pool.

Reading the tables: plain decode (no speculation) is bounded by memory bandwidth — every token reads all ~15.8 GB of weights, and the R9700's ~640 GB/s caps it near 38–40 tok/s in both engines, so WHIRL's no-MTP lead is small (1.12–1.18×). Most of WHIRL's decode lead comes from MTP + n-gram speculative decoding, whose output is identical to plain greedy decoding. n-gram drafts matter most when the output repeats the input (file editing).

Variant comparison in WHIRL (pre-release build, 2026-10-01, mixed 19-prompt protocol)

Used to choose the recommended variant. Server defaults, 19 prompts mixing zh thinking off, zh thinking on and code-edit prompts, 2 interleaved rounds, greedy-equivalent sampling. Because the code-edit prompts are included, the MTP + n-gram row is an average over generation and editing prompts and is not comparable with the scenario numbers above (the zh thinking-off group alone was 106.0 tok/s for A in this run; 107.5 in v0.1.0). "Qwen MXFP4" is the reference Qwen3.8-27B MXFP4 GGUF mentioned above, measured the same way.

WHIRL (pre-release), decode tok/sA (Q6_K head)B (Q8_0 head)C (Q4_K head)Qwen3.8-27B MXFP4
Plain (no MTP, no speculation)36.836.137.536.9
MTP117.098.0112.0118.1
MTP + n-gram — 19-prompt mixed average incl. code-edit prompts157.2143.6157.5156.9
MTP draft acceptance (MTP / MTP + n-gram)65.4% / 66.3%73.2% / 72.9%68.3% / 69.1%60.6% / 62.2%
Prefill tok/s, 2k / 8k / 32k3164 / 3302 / 25703181 / 3291 / 25633138 / 3307 / 25723223 / 3338 / 2602
Plain decode (no MTP) after 16k context34.634.135.234.7

Why A is recommended: B's 8-bit head is the most expensive part of every draft step, so it is the slowest with MTP despite the highest acceptance rate. C is fastest without speculation and on llama.cpp, but in WHIRL A gets a dedicated Q6_K-head path that made it 4.5% faster than C with MTP (and equal with MTP + n-gram) in this comparison, while keeping a more precise head.

Thinking length vs Qwen3.8-27B MXFP4 (small sample)

Swift A vs the reference Qwen3.8-27B MXFP4 GGUF, both served by WHIRL (MTP + n-gram) on the R9700. Settings: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0 (official sampling for both), thinking on (template default), max_tokens 16,384, 9 prompts × seeds 1 and 2 (n = 18 runs per model). Thinking tokens = reasoning_content counted with the model's tokenizer; wall = full non-streaming request time. * = hit the 16,384-token cap (value is a lower bound).

PromptSwift think tokens (s1 / s2)Swift wall sQwen think tokens (s1 / s2)Qwen wall s
你好24 / 240.5 / 0.621 / 210.5 / 0.5
Hi16 / 160.4 / 0.446 / 431.1 / 0.8
zh: Python LRU12639* / 10619161.5 / 135.216382* / 16383*189.8 / 199.2
zh: TypeScript hook2148 / 769547.4 / 109.511795 / 10961179.8 / 166.6
zh: SQL report8961 / 7658174.5 / 138.612893* / 13540*210.9 / 210.4
zh: bug fix2842 / 238651.0 / 40.49598 / 5944140.7 / 97.7
en: ISO-8601 duration parser16384* / 16310*206.5 / 203.616382* / 16384*212.7 / 132.3
en: Go bounded queue16384* / 16384*236.1 / 230.116383* / 16384*245.6 / 240.8
en: JS debounce bug4020 / 478957.6 / 65.51855 / 449530.5 / 68.2
  • Totals (18 runs): Swift 129,299 thinking tokens / 1,859 s / 5 capped; Qwen 169,510 / 2,328 s / 8 capped. Because Qwen hit the cap more often, its totals are lower bounds.
  • 4 Chinese coding prompts: Swift 54,948 vs Qwen 97,496 thinking tokens (−44%), wall 858 s vs 1,395 s (−38%), capped 1 vs 4.
  • 3 English prompts: similar (Swift 74,271 vs Qwen 71,883 tokens); on two of them both models were still thinking at 16k.
  • A capped Swift run (en ISO-8601, seed 1) was re-run and inspected: repeated 32-gram ratio 0.5% (first half) / 1.3% (second half) — genuine long deliberation, not a degenerate repetition loop.
  • This is a small, single-machine sample with a 16k cap; it does not measure answer correctness. See the upstream card for UkisAI's controlled evaluation.

License

These files are a quantized (Object-form) Derivative Work of Swift 1.5 Qwen3.8-27B and are distributed under the same terms as upstream:

  • Swift Open License v1.0 for UkisAI's Swift Contribution — see LICENSE (unmodified copy of the upstream LICENSE).
  • Apache License 2.0 for the Qwen3.8-27B base model — see LICENSE-APACHE-2.0.
  • NOTICE: the upstream NOTICE verbatim, plus a "Changes in this redistribution" section describing the conversion and quantization.

Notice for redistributors, reproduced verbatim from the appendix of the Swift Open License v1.0:

   Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open
   License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE).
   Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.

Commercial-use threshold (summary, not legal advice — the LICENSE text governs): personal, research, educational, evaluation and commercial use are permitted for individuals and organizations whose gross revenue, together with all affiliates (entities controlling, controlled by, or under common control with them), is below US$1,000,000 over the most recently completed fiscal year. Commercial use by an entity at or above that threshold is not licensed under the Swift Open License and requires a separate Swift Enterprise License from UkisAI (contact). The threshold does not apply to qualified non-profit organizations' non-commercial or research use. The license terminates automatically on non-compliance.

Base model: Qwen3.8-27B is Copyright 2026 Alibaba Cloud and licensed under the Apache License 2.0. Per Section 6 of the Swift Open License, nothing in it limits your rights in the Qwen3.8-27B base model under Apache 2.0; the commercial limitation applies only to the Swift Contribution.

Credits

  • Alibaba Cloud / Qwen team — Qwen3.8-27B, the base model.
  • UkisAI — Swift 1.5 Qwen3.8-27B, the post-trained model these files are made from.
  • bartowski — the importance matrix used for the K-quant output heads.
  • llama.cpp / ggml contributors — conversion, quantization and inference tooling.

Upstream citation:

@misc{swift-1.5-qwen3.8-27b,
  title  = {Swift 1.5 Qwen3.8-27B},
  author = {UkisAI},
  year   = {2026},
  url    = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}

繁體中文摘要

  • 這是 ukisai/Swift-1.5-Qwen3.8-27b 的非官方社群 MXFP4 GGUF 量化,與 UkisAI 無關;「Swift」「UkisAI」商標僅用來說明權重來源。另附視覺投影檔 mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf(F16),搭配任一版本即可看圖(llama.cpp --mmproj)。
  • 保留 MTP 層(Q8_0),llama.cpp 可用 --spec-type draft-mtp 做自我推測解碼。
  • 三個變體只差輸出頭:A = Q6_K(推薦)、C = Q4_K(在 llama.cpp 與不開推測解碼時稍快)、B = Q8_0(開 MTP 時最慢,不建議)。
  • 大矩陣 MXFP4;token_embd、ssm_alpha / beta、MTP 層 Q8_0;imatrix 來自 bartowski,但 mainline 量 MXFP4 / Q8_0 時不使用 imatrix,只影響 K-quant 輸出頭。
  • AMD Radeon AI PRO R9700、llama.cpp b11214:A 的 plain decode 33.3 tok/s、MTP 58.1 tok/s、2k prefill 1215 tok/s。
  • 最佳效能請用 WHIRL(Apache-2.0 開源,v0.1.0):A 在中文寫程式解碼 107.5 對 llama.cpp 最快設定 60.8 tok/s(1.77×)、檔案編輯 328.3 對 144.8(2.27×)、16k 上下文後 69.5 對 63.5(1.09×);不開 MTP 37.5 對 33.4(1.13×);prefill 8k 3,278 對 1,338 tok/s(2.45×)。舊的 157.2 tok/s 是 19 題混合(含程式碼編輯題)的平均,不能和單一情境數字直接比。
  • 另有 MoE 模型 Ornith-1.5-35B-A3B MXFP4 GGUF(每 token 約 3B 啟用)。
  • 困惑度(中英混合程式碼語料,2048 ctx):A 3.8271、B 3.8289、C 3.8386,三者在誤差內;Swift 是另一個微調模型,不宜直接和 Qwen 量化版(3.7952)比較。
  • 小樣本思考長度比較(16k 上限、9 題 × 2 seed):中文寫程式 4 題 Swift 思考 token 比 Qwen3.8-27B MXFP4 少 44%、時間少 38%。
  • 授權:Swift Open License v1.0(年營收含關係企業達 100 萬美元以上的商業使用需另取得授權)+ 基底模型 Apache 2.0;詳見 LICENSE / NOTICE。
conversational
endpoints_compatible
gguf
imatrix
llama.cpp
mxfp4
qwen3.8
rdna4
text-generation
whirl