MXFP4 GGUF quantizations of ukisai/Swift-1.5-Qwen3.8-27b, with the MTP (multi-token prediction) layer kept for self-speculative decoding.
Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by UkisAI or Alibaba Cloud. The names "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the origin of the weights. For the official model, official quantizations, evaluations and enterprise licensing, see the upstream repository.
Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient post-trained derivative of Qwen/Qwen3.8-27B. Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates model quality.
Vision is supported through the separate projector file mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf (F16, converted from the Swift-1.5 BF16 weights with convert_hf_to_gguf.py --mmproj --outtype f16). It works with any of the three variants.
Best performance with WHIRL
These files were tuned for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. They are standard GGUF and also run in llama.cpp (tested with b11214, see below).
Also from us: Ornith-1.5-35B-A3B MXFP4 GGUF — a mixture-of-experts model (~3B active of 35B) quantized the same way, with MTP layer and vision projector.
The three language-model files contain the same weights except the output head (output.weight).
| Variant | File | Output head | Size (bytes) | SHA256 | Recommendation |
|---|---|---|---|---|---|
| A | Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf | Q6_K | 15,815,469,280 | f12d8dee7d677a0583b243d65c5e75e853b1839dc61e2c101da090bb6163065f | Recommended default |
| C | Swift-1.5-Qwen3.8-27B-MXFP4-C-outQ4_K.gguf | Q4_K | 15,487,686,880 | 221aa39e7f7c003ae94903ea686d33d6bb187ed8b943955fb21d0790c7cf6b4a | Slightly faster on llama.cpp and in plain (non-speculative) decoding; lower-precision head |
| B | Swift-1.5-Qwen3.8-27B-MXFP4-B-outQ8_0.gguf | Q8_0 | 16,123,386,080 | bb075d1cf83352d38751a799feffc0e0c91bb8ccb44a14551c0f485ffd65c72a | Not recommended for speed (slowest with MTP); highest-precision head |
| vision | mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf | — (vision projector, F16) | 927,607,584 | 71988a54d379191e1abc26c2b79363b87ef8b3a584ebee882c28a1682ddbfe62 | Needed only for image input |
Tensor types (verified by dumping the quantized files):
| Tensor group | Type |
|---|---|
All large matrices of the 64 main layers: attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out, ffn_gate/up/down | MXFP4 |
token_embd | Q8_0 |
ssm_alpha, ssm_beta (48 linear-attention layers × 2) | Q8_0 |
MTP layer blk.64 (all of attn_q/k/v/output, ffn_gate/up/down, nextn.eh_proj) | Q8_0 |
Norms, ssm_a, ssm_conv1d, ssm_dt.bias, nextn.enorm/hnorm/shared_head_norm | F32 |
output (LM head) | Q6_K (A) / Q8_0 (B) / Q4_K (C) |
Steps
convert_hf_to_gguf.py <Swift-1.5 BF16 dir> --outtype bf16 (866 tensors; qwen35.nextn_predict_layers = 1, i.e. the MTP layer is kept). (llama.cpp b10645-47-gef6876693, commit ef6876693.)llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build.tt_common.txt (passed with --tensor-type-file; the first matching rule wins):
blk\.64\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate=mxfp4
ffn_up=mxfp4
ffn_down=mxfp4
ssm_out=mxfp4
Exact command for variant A (B used --output-tensor-type q8_0, C used --output-tensor-type q4_k):
llama-quantize --imatrix ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf \
--tensor-type-file tt_common.txt --output-tensor-type q6_k \
swift-1.5-27b-bf16.gguf Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf MXFP4_MOE 16
Notes:
ftype MXFP4_MOE: this llama.cpp build has no dense-model MXFP4 ftype, so MXFP4_MOE is used as the base type and the per-tensor rules above decide every tensor. The GGUF metadata therefore reports general.file_type = 38 (MXFP4_MOE) even though the model is dense.
Importance matrix: the imatrix was made and published by bartowski — bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF (ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, 496 entries, 582 chunks). Thank you! Note that mainline llama.cpp ignores the imatrix for MXFP4 and for Q8_0, so in these files it only affects the K-quant output head (Q6_K in A, Q4_K in C). B is effectively imatrix-free.
Why ssm_alpha / ssm_beta are Q8_0 and not F32: a first build kept them in F32. With F32 weights for these small matrices, WHIRL computes single-token decode and multi-token MTP verification through different kernels, so MTP output was no longer bit-identical to plain greedy decoding, and the automatic draft length collapsed. Q8_0 has a fused exact path there, restoring bit-identical MTP output. This choice was made for WHIRL; llama.cpp loads and runs the files normally (see below).
Differences from the reference Qwen3.8-27B MXFP4 GGUF we compared against (FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF): that file is MXFP4 everywhere, including token_embd, ssm_alpha/beta and the MTP layer, with a Q6_K head. Here those tensors are Q8_0, which costs about 1–2% prefill speed in our measurements.
Perplexity (llama.cpp b11214 llama-perplexity, -c 2048 -b 512 -ub 512 -fa on, 51 chunks of a 219 KB mixed Chinese/English source-code corpus, R9700):
| file | PPL |
|---|---|
| A (Q6_K head) | 3.8271 ± 0.0394 |
| B (Q8_0 head) | 3.8289 ± 0.0395 |
| C (Q4_K head) | 3.8386 ± 0.0396 |
| reference: Qwen3.8-27B MXFP4 (FreedomAISVR) | 3.7952 ± 0.0384 |
The three heads are within noise of each other (C is +0.3%). Swift is a different fine-tune, so its PPL is not directly comparable with the Qwen base quant (a post-trained model typically shifts PPL on generic text). No BF16 baseline / KL-divergence was measured. Beyond PPL, quality was only checked by coherent zh/en greedy output.
The MTP layer is kept, so llama.cpp can use it for self-speculative decoding with --spec-type draft-mtp. This is the exact llama-server command used for our llama.cpp check (b11214, ROCm):
llama-server -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --device ROCm0 -ngl 99 -c 131072 -np 1 -fa on \
--port 1234 --no-webui --load-mode none --jinja \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-n-min 0 --spec-draft-p-min 0.3
(Our machine has two GPUs, so the run also set HIP_VISIBLE_DEVICES=1 and GGML_CUDA_NO_PINNED=1; adjust --device / -c to your hardware.) Without the --spec-* flags the model runs as a plain model; llama.cpp then logs the MTP-layer tensors as "unused", which is expected. The loader also prints unknown type mxfp4 as a warning; this appears with other MXFP4 GGUFs as well and loading succeeds.
Image input (tested with llama-mtmd-cli b11214: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, including the exact text):
llama-mtmd-cli -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf --image photo.png -ngl 99 -c 8192 -p "Describe this image."
llama-server accepts the same --mmproj flag for image input over the OpenAI-compatible API.
Sampling recommended by the upstream card: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0. Thinking is on by default in the chat template; our llama.cpp check disabled it per request with "chat_template_kwargs": {"enable_thinking": false}.
Single-machine measurements; speed only, not quality evaluations. The R9700 in this machine is connected as a USB4 eGPU (this slows model loading and host-RAM / SSD KV restores, not prefill or decode).
llama-server, flags as above; 2026-10-01, single request, short prompts)Greedy-equivalent sampling (top_k 1), thinking off, 4 short prompts (2 Traditional Chinese, 2 English), max 400 output tokens; prefill measured on a 2,022-token prompt.
| llama.cpp b11214 | A (Q6_K head) | B (Q8_0 head) | C (Q4_K head) |
|---|---|---|---|
| Plain decode, tok/s (mean of 4 prompts) | 33.3 | 32.8 | 33.8 |
MTP decode, tok/s (--spec-draft-n-max 3) | 58.1 | 55.9 | 60.7 |
| MTP draft acceptance | 73.6% (846/1150) | 72.2% (840/1164) | 75.0% (856/1141) |
| Prefill, 2k prompt, tok/s | 1215 | 1212 | 1206 |
llama-server at -c 131072 -np 1: 23.65 GiB.From the WHIRL v0.1.0 release benchmark (2026-10-03): same GGUF, same prompts, greedy decoding in both engines, one model process on the GPU at a time, median of 3 rounds (file editing / concurrency: 2). Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram and of the flag sets tried; named). "no MTP" = plain decoding without any speculation. llama.cpp here uses -ub 1024 and its best MTP settings (--spec-draft-n-max 4), so its MTP number is a little higher than in the 2026-10-01 check above (60.8 vs 58.1 tok/s).
Decode (tok/s), WHIRL default = MTP + n-gram
| Scenario (server) | WHIRL MTP + n-gram | WHIRL MTP | llama.cpp fastest (config) | Ratio |
|---|---|---|---|---|
| 7 Chinese coding prompts (Chinese question, Chinese explanation, English code), thinking off, 800 tokens | 107.5 | 107.3 | 60.8 (draft-mtp, n-max 4) | 1.77× |
| File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output | 328.3 | 181.1 | 144.8 (draft-mtp,ngram-mod, n-max 3) | 2.27× |
| 128 tokens after a 16k-token code context | 69.5 | 72.4 | 63.5 (draft-mtp, n-max 4, V cache q8_0) | 1.09× |
After a 16k context WHIRL's lead is small, and its MTP-only mode is faster than its default MTP + n-gram on this model.
Decode (tok/s), no MTP (plain decoding in both engines)
| Scenario | WHIRL no MTP | llama.cpp no MTP | Ratio |
|---|---|---|---|
| 7 Chinese coding prompts, server | 37.5 | 33.4 | 1.13× |
| File editing, server | 37.0 | 33.0 | 1.12× |
| After a 16k-token context, server | 35.2 | 31.4 | 1.12× |
CLI bench (whirl bench 256 tokens vs llama-bench tg256) | 39.4 | 33.4 | 1.18× |
Prefill (tok/s, each engine's own bench tool, KV f16)
| Prompt tokens | WHIRL | llama.cpp (-ub 1024 unless noted) | Ratio |
|---|---|---|---|
| 88 | 1,515 | 928.6 (-ub 512) | 1.63× |
| 2,048 | 3,509 | 1,385 | 2.53× |
| 8,192 | 3,278 | 1,338 | 2.45× |
| 32,768 | 2,595 | 1,174 | 2.21× |
| 131,072 | 1,405 | 791.3 | 1.78× |
Server
| Test | WHIRL | llama-server | Ratio |
|---|---|---|---|
| 4 concurrent requests, aggregate tok/s — WHIRL MTP + n-gram vs llama.cpp fastest (plain) | 198.6 | 65.8 | 3.02× |
| Same, WHIRL no MTP vs llama.cpp no MTP | 108.6 | 65.8 | 1.65× |
| TTFT, ~12.2k-token system prompt, cold | 4.091 s | 10.650 s | 2.6× faster |
| TTFT, new conversation reusing that system prompt | 0.102 s | 0.446 s | 4.4× faster |
| TTFT, evicted ~27.4k-token session restored from host RAM | 0.668 s | 2.195 s | 3.3× faster |
| TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier) | 0.729 s | N/A (no equivalent feature) | — |
| Image encode with the F16 projector, warm, 7 images (220–2,040 image tokens) | 16.8–253.7 ms | 58–1,330 ms | 3.5–5.2× faster |
VRAM: at equal context WHIRL's CLI uses more VRAM than llama-bench on this model (23.0 vs 17.8 GiB at ~33k tokens; 26.1 vs 22.7 GiB at ~131k), for larger prefill buffers, the MTP block and the draft head. The WHIRL server deliberately fills free VRAM with its KV pool.
Reading the tables: plain decode (no speculation) is bounded by memory bandwidth — every token reads all ~15.8 GB of weights, and the R9700's ~640 GB/s caps it near 38–40 tok/s in both engines, so WHIRL's no-MTP lead is small (1.12–1.18×). Most of WHIRL's decode lead comes from MTP + n-gram speculative decoding, whose output is identical to plain greedy decoding. n-gram drafts matter most when the output repeats the input (file editing).
Used to choose the recommended variant. Server defaults, 19 prompts mixing zh thinking off, zh thinking on and code-edit prompts, 2 interleaved rounds, greedy-equivalent sampling. Because the code-edit prompts are included, the MTP + n-gram row is an average over generation and editing prompts and is not comparable with the scenario numbers above (the zh thinking-off group alone was 106.0 tok/s for A in this run; 107.5 in v0.1.0). "Qwen MXFP4" is the reference Qwen3.8-27B MXFP4 GGUF mentioned above, measured the same way.
| WHIRL (pre-release), decode tok/s | A (Q6_K head) | B (Q8_0 head) | C (Q4_K head) | Qwen3.8-27B MXFP4 |
|---|---|---|---|---|
| Plain (no MTP, no speculation) | 36.8 | 36.1 | 37.5 | 36.9 |
| MTP | 117.0 | 98.0 | 112.0 | 118.1 |
| MTP + n-gram — 19-prompt mixed average incl. code-edit prompts | 157.2 | 143.6 | 157.5 | 156.9 |
| MTP draft acceptance (MTP / MTP + n-gram) | 65.4% / 66.3% | 73.2% / 72.9% | 68.3% / 69.1% | 60.6% / 62.2% |
| Prefill tok/s, 2k / 8k / 32k | 3164 / 3302 / 2570 | 3181 / 3291 / 2563 | 3138 / 3307 / 2572 | 3223 / 3338 / 2602 |
| Plain decode (no MTP) after 16k context | 34.6 | 34.1 | 35.2 | 34.7 |
Why A is recommended: B's 8-bit head is the most expensive part of every draft step, so it is the slowest with MTP despite the highest acceptance rate. C is fastest without speculation and on llama.cpp, but in WHIRL A gets a dedicated Q6_K-head path that made it 4.5% faster than C with MTP (and equal with MTP + n-gram) in this comparison, while keeping a more precise head.
Swift A vs the reference Qwen3.8-27B MXFP4 GGUF, both served by WHIRL (MTP + n-gram) on the R9700. Settings: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0 (official sampling for both), thinking on (template default), max_tokens 16,384, 9 prompts × seeds 1 and 2 (n = 18 runs per model). Thinking tokens = reasoning_content counted with the model's tokenizer; wall = full non-streaming request time. * = hit the 16,384-token cap (value is a lower bound).
| Prompt | Swift think tokens (s1 / s2) | Swift wall s | Qwen think tokens (s1 / s2) | Qwen wall s |
|---|---|---|---|---|
| 你好 | 24 / 24 | 0.5 / 0.6 | 21 / 21 | 0.5 / 0.5 |
| Hi | 16 / 16 | 0.4 / 0.4 | 46 / 43 | 1.1 / 0.8 |
| zh: Python LRU | 12639* / 10619 | 161.5 / 135.2 | 16382* / 16383* | 189.8 / 199.2 |
| zh: TypeScript hook | 2148 / 7695 | 47.4 / 109.5 | 11795 / 10961 | 179.8 / 166.6 |
| zh: SQL report | 8961 / 7658 | 174.5 / 138.6 | 12893* / 13540* | 210.9 / 210.4 |
| zh: bug fix | 2842 / 2386 | 51.0 / 40.4 | 9598 / 5944 | 140.7 / 97.7 |
| en: ISO-8601 duration parser | 16384* / 16310* | 206.5 / 203.6 | 16382* / 16384* | 212.7 / 132.3 |
| en: Go bounded queue | 16384* / 16384* | 236.1 / 230.1 | 16383* / 16384* | 245.6 / 240.8 |
| en: JS debounce bug | 4020 / 4789 | 57.6 / 65.5 | 1855 / 4495 | 30.5 / 68.2 |
These files are a quantized (Object-form) Derivative Work of Swift 1.5 Qwen3.8-27B and are distributed under the same terms as upstream:
LICENSE (unmodified copy of the upstream LICENSE).LICENSE-APACHE-2.0.NOTICE: the upstream NOTICE verbatim, plus a "Changes in this redistribution" section describing the conversion and quantization.Notice for redistributors, reproduced verbatim from the appendix of the Swift Open License v1.0:
Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open
License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE).
Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
Commercial-use threshold (summary, not legal advice — the LICENSE text governs): personal, research, educational, evaluation and commercial use are permitted for individuals and organizations whose gross revenue, together with all affiliates (entities controlling, controlled by, or under common control with them), is below US$1,000,000 over the most recently completed fiscal year. Commercial use by an entity at or above that threshold is not licensed under the Swift Open License and requires a separate Swift Enterprise License from UkisAI (contact). The threshold does not apply to qualified non-profit organizations' non-commercial or research use. The license terminates automatically on non-compliance.
Base model: Qwen3.8-27B is Copyright 2026 Alibaba Cloud and licensed under the Apache License 2.0. Per Section 6 of the Swift Open License, nothing in it limits your rights in the Qwen3.8-27B base model under Apache 2.0; the commercial limitation applies only to the Swift Contribution.
Upstream citation:
@misc{swift-1.5-qwen3.8-27b,
title = {Swift 1.5 Qwen3.8-27B},
author = {UkisAI},
year = {2026},
url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}
mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf(F16),搭配任一版本即可看圖(llama.cpp --mmproj)。--spec-type draft-mtp 做自我推測解碼。MXFP4 GGUF quantizations of ukisai/Swift-1.5-Qwen3.8-27b, with the MTP (multi-token prediction) layer kept for self-speculative decoding.
Unofficial community quantization. This repository is not affiliated with, endorsed by, or supported by UkisAI or Alibaba Cloud. The names "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the origin of the weights. For the official model, official quantizations, evaluations and enterprise licensing, see the upstream repository.
Swift 1.5 Qwen3.8-27B is UkisAI's reasoning-efficient post-trained derivative of Qwen/Qwen3.8-27B. Please read the upstream model card for its training approach and benchmark results; nothing in this card re-evaluates model quality.
Vision is supported through the separate projector file mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf (F16, converted from the Swift-1.5 BF16 weights with convert_hf_to_gguf.py --mmproj --outtype f16). It works with any of the three variants.
Best performance with WHIRL
These files were tuned for WHIRL (Windows HIP Inference for RDNA LLMs), our open-source (Apache-2.0) native-Windows C++/HIP inference engine for the AMD Radeon AI PRO R9700 — download release v0.1.0. They are standard GGUF and also run in llama.cpp (tested with b11214, see below).
Also from us: Ornith-1.5-35B-A3B MXFP4 GGUF — a mixture-of-experts model (~3B active of 35B) quantized the same way, with MTP layer and vision projector.
The three language-model files contain the same weights except the output head (output.weight).
| Variant | File | Output head | Size (bytes) | SHA256 | Recommendation |
|---|---|---|---|---|---|
| A | Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf | Q6_K | 15,815,469,280 | f12d8dee7d677a0583b243d65c5e75e853b1839dc61e2c101da090bb6163065f | Recommended default |
| C | Swift-1.5-Qwen3.8-27B-MXFP4-C-outQ4_K.gguf | Q4_K | 15,487,686,880 | 221aa39e7f7c003ae94903ea686d33d6bb187ed8b943955fb21d0790c7cf6b4a | Slightly faster on llama.cpp and in plain (non-speculative) decoding; lower-precision head |
| B | Swift-1.5-Qwen3.8-27B-MXFP4-B-outQ8_0.gguf | Q8_0 | 16,123,386,080 | bb075d1cf83352d38751a799feffc0e0c91bb8ccb44a14551c0f485ffd65c72a | Not recommended for speed (slowest with MTP); highest-precision head |
| vision | mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf | — (vision projector, F16) | 927,607,584 | 71988a54d379191e1abc26c2b79363b87ef8b3a584ebee882c28a1682ddbfe62 | Needed only for image input |
Tensor types (verified by dumping the quantized files):
| Tensor group | Type |
|---|---|
All large matrices of the 64 main layers: attn_qkv, attn_gate, attn_q/k/v, attn_output, ssm_out, ffn_gate/up/down | MXFP4 |
token_embd | Q8_0 |
ssm_alpha, ssm_beta (48 linear-attention layers × 2) | Q8_0 |
MTP layer blk.64 (all of attn_q/k/v/output, ffn_gate/up/down, nextn.eh_proj) | Q8_0 |
Norms, ssm_a, ssm_conv1d, ssm_dt.bias, nextn.enorm/hnorm/shared_head_norm | F32 |
output (LM head) | Q6_K (A) / Q8_0 (B) / Q4_K (C) |
Steps
convert_hf_to_gguf.py <Swift-1.5 BF16 dir> --outtype bf16 (866 tensors; qwen35.nextn_predict_layers = 1, i.e. the MTP layer is kept). (llama.cpp b10645-47-gef6876693, commit ef6876693.)llama-quantize from llama.cpp build 11214 (commit 2ebd9ae62), ROCm build.tt_common.txt (passed with --tensor-type-file; the first matching rule wins):
blk\.64\.=q8_0
ssm_alpha=q8_0
ssm_beta=q8_0
token_embd=q8_0
attn_gate=mxfp4
attn_qkv=mxfp4
attn_q\.=mxfp4
attn_k\.=mxfp4
attn_v\.=mxfp4
attn_output=mxfp4
ffn_gate=mxfp4
ffn_up=mxfp4
ffn_down=mxfp4
ssm_out=mxfp4
Exact command for variant A (B used --output-tensor-type q8_0, C used --output-tensor-type q4_k):
llama-quantize --imatrix ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf \
--tensor-type-file tt_common.txt --output-tensor-type q6_k \
swift-1.5-27b-bf16.gguf Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf MXFP4_MOE 16
Notes:
ftype MXFP4_MOE: this llama.cpp build has no dense-model MXFP4 ftype, so MXFP4_MOE is used as the base type and the per-tensor rules above decide every tensor. The GGUF metadata therefore reports general.file_type = 38 (MXFP4_MOE) even though the model is dense.
Importance matrix: the imatrix was made and published by bartowski — bartowski/ukisai_Swift-1.5-Qwen3.8-27b-GGUF (ukisai_Swift-1.5-Qwen3.8-27b-imatrix.gguf, 496 entries, 582 chunks). Thank you! Note that mainline llama.cpp ignores the imatrix for MXFP4 and for Q8_0, so in these files it only affects the K-quant output head (Q6_K in A, Q4_K in C). B is effectively imatrix-free.
Why ssm_alpha / ssm_beta are Q8_0 and not F32: a first build kept them in F32. With F32 weights for these small matrices, WHIRL computes single-token decode and multi-token MTP verification through different kernels, so MTP output was no longer bit-identical to plain greedy decoding, and the automatic draft length collapsed. Q8_0 has a fused exact path there, restoring bit-identical MTP output. This choice was made for WHIRL; llama.cpp loads and runs the files normally (see below).
Differences from the reference Qwen3.8-27B MXFP4 GGUF we compared against (FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF): that file is MXFP4 everywhere, including token_embd, ssm_alpha/beta and the MTP layer, with a Q6_K head. Here those tensors are Q8_0, which costs about 1–2% prefill speed in our measurements.
Perplexity (llama.cpp b11214 llama-perplexity, -c 2048 -b 512 -ub 512 -fa on, 51 chunks of a 219 KB mixed Chinese/English source-code corpus, R9700):
| file | PPL |
|---|---|
| A (Q6_K head) | 3.8271 ± 0.0394 |
| B (Q8_0 head) | 3.8289 ± 0.0395 |
| C (Q4_K head) | 3.8386 ± 0.0396 |
| reference: Qwen3.8-27B MXFP4 (FreedomAISVR) | 3.7952 ± 0.0384 |
The three heads are within noise of each other (C is +0.3%). Swift is a different fine-tune, so its PPL is not directly comparable with the Qwen base quant (a post-trained model typically shifts PPL on generic text). No BF16 baseline / KL-divergence was measured. Beyond PPL, quality was only checked by coherent zh/en greedy output.
The MTP layer is kept, so llama.cpp can use it for self-speculative decoding with --spec-type draft-mtp. This is the exact llama-server command used for our llama.cpp check (b11214, ROCm):
llama-server -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --device ROCm0 -ngl 99 -c 131072 -np 1 -fa on \
--port 1234 --no-webui --load-mode none --jinja \
--spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-n-min 0 --spec-draft-p-min 0.3
(Our machine has two GPUs, so the run also set HIP_VISIBLE_DEVICES=1 and GGML_CUDA_NO_PINNED=1; adjust --device / -c to your hardware.) Without the --spec-* flags the model runs as a plain model; llama.cpp then logs the MTP-layer tensors as "unused", which is expected. The loader also prints unknown type mxfp4 as a warning; this appears with other MXFP4 GGUFs as well and loading succeeds.
Image input (tested with llama-mtmd-cli b11214: a synthetic image with a red square, a blue circle and the text "WHIRL 42" was described correctly, including the exact text):
llama-mtmd-cli -m Swift-1.5-Qwen3.8-27B-MXFP4-A-outQ6_K.gguf --mmproj mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf --image photo.png -ngl 99 -c 8192 -p "Describe this image."
llama-server accepts the same --mmproj flag for image input over the OpenAI-compatible API.
Sampling recommended by the upstream card: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0. Thinking is on by default in the chat template; our llama.cpp check disabled it per request with "chat_template_kwargs": {"enable_thinking": false}.
Single-machine measurements; speed only, not quality evaluations. The R9700 in this machine is connected as a USB4 eGPU (this slows model loading and host-RAM / SSD KV restores, not prefill or decode).
llama-server, flags as above; 2026-10-01, single request, short prompts)Greedy-equivalent sampling (top_k 1), thinking off, 4 short prompts (2 Traditional Chinese, 2 English), max 400 output tokens; prefill measured on a 2,022-token prompt.
| llama.cpp b11214 | A (Q6_K head) | B (Q8_0 head) | C (Q4_K head) |
|---|---|---|---|
| Plain decode, tok/s (mean of 4 prompts) | 33.3 | 32.8 | 33.8 |
MTP decode, tok/s (--spec-draft-n-max 3) | 58.1 | 55.9 | 60.7 |
| MTP draft acceptance | 73.6% (846/1150) | 72.2% (840/1164) | 75.0% (856/1141) |
| Prefill, 2k prompt, tok/s | 1215 | 1212 | 1206 |
llama-server at -c 131072 -np 1: 23.65 GiB.From the WHIRL v0.1.0 release benchmark (2026-10-03): same GGUF, same prompts, greedy decoding in both engines, one model process on the GPU at a time, median of 3 rounds (file editing / concurrency: 2). Ratio = WHIRL / llama.cpp's fastest configuration for that row (best of plain, MTP and MTP + n-gram and of the flag sets tried; named). "no MTP" = plain decoding without any speculation. llama.cpp here uses -ub 1024 and its best MTP settings (--spec-draft-n-max 4), so its MTP number is a little higher than in the 2026-10-01 check above (60.8 vs 58.1 tok/s).
Decode (tok/s), WHIRL default = MTP + n-gram
| Scenario (server) | WHIRL MTP + n-gram | WHIRL MTP | llama.cpp fastest (config) | Ratio |
|---|---|---|---|---|
| 7 Chinese coding prompts (Chinese question, Chinese explanation, English code), thinking off, 800 tokens | 107.5 | 107.3 | 60.8 (draft-mtp, n-max 4) | 1.77× |
| File editing: 5 prompts with 1.6–2.3k-token source files, whole file as output | 328.3 | 181.1 | 144.8 (draft-mtp,ngram-mod, n-max 3) | 2.27× |
| 128 tokens after a 16k-token code context | 69.5 | 72.4 | 63.5 (draft-mtp, n-max 4, V cache q8_0) | 1.09× |
After a 16k context WHIRL's lead is small, and its MTP-only mode is faster than its default MTP + n-gram on this model.
Decode (tok/s), no MTP (plain decoding in both engines)
| Scenario | WHIRL no MTP | llama.cpp no MTP | Ratio |
|---|---|---|---|
| 7 Chinese coding prompts, server | 37.5 | 33.4 | 1.13× |
| File editing, server | 37.0 | 33.0 | 1.12× |
| After a 16k-token context, server | 35.2 | 31.4 | 1.12× |
CLI bench (whirl bench 256 tokens vs llama-bench tg256) | 39.4 | 33.4 | 1.18× |
Prefill (tok/s, each engine's own bench tool, KV f16)
| Prompt tokens | WHIRL | llama.cpp (-ub 1024 unless noted) | Ratio |
|---|---|---|---|
| 88 | 1,515 | 928.6 (-ub 512) | 1.63× |
| 2,048 | 3,509 | 1,385 | 2.53× |
| 8,192 | 3,278 | 1,338 | 2.45× |
| 32,768 | 2,595 | 1,174 | 2.21× |
| 131,072 | 1,405 | 791.3 | 1.78× |
Server
| Test | WHIRL | llama-server | Ratio |
|---|---|---|---|
| 4 concurrent requests, aggregate tok/s — WHIRL MTP + n-gram vs llama.cpp fastest (plain) | 198.6 | 65.8 | 3.02× |
| Same, WHIRL no MTP vs llama.cpp no MTP | 108.6 | 65.8 | 1.65× |
| TTFT, ~12.2k-token system prompt, cold | 4.091 s | 10.650 s | 2.6× faster |
| TTFT, new conversation reusing that system prompt | 0.102 s | 0.446 s | 4.4× faster |
| TTFT, evicted ~27.4k-token session restored from host RAM | 0.668 s | 2.195 s | 3.3× faster |
| TTFT, ~25.7k-token session after a server restart (WHIRL SSD tier) | 0.729 s | N/A (no equivalent feature) | — |
| Image encode with the F16 projector, warm, 7 images (220–2,040 image tokens) | 16.8–253.7 ms | 58–1,330 ms | 3.5–5.2× faster |
VRAM: at equal context WHIRL's CLI uses more VRAM than llama-bench on this model (23.0 vs 17.8 GiB at ~33k tokens; 26.1 vs 22.7 GiB at ~131k), for larger prefill buffers, the MTP block and the draft head. The WHIRL server deliberately fills free VRAM with its KV pool.
Reading the tables: plain decode (no speculation) is bounded by memory bandwidth — every token reads all ~15.8 GB of weights, and the R9700's ~640 GB/s caps it near 38–40 tok/s in both engines, so WHIRL's no-MTP lead is small (1.12–1.18×). Most of WHIRL's decode lead comes from MTP + n-gram speculative decoding, whose output is identical to plain greedy decoding. n-gram drafts matter most when the output repeats the input (file editing).
Used to choose the recommended variant. Server defaults, 19 prompts mixing zh thinking off, zh thinking on and code-edit prompts, 2 interleaved rounds, greedy-equivalent sampling. Because the code-edit prompts are included, the MTP + n-gram row is an average over generation and editing prompts and is not comparable with the scenario numbers above (the zh thinking-off group alone was 106.0 tok/s for A in this run; 107.5 in v0.1.0). "Qwen MXFP4" is the reference Qwen3.8-27B MXFP4 GGUF mentioned above, measured the same way.
| WHIRL (pre-release), decode tok/s | A (Q6_K head) | B (Q8_0 head) | C (Q4_K head) | Qwen3.8-27B MXFP4 |
|---|---|---|---|---|
| Plain (no MTP, no speculation) | 36.8 | 36.1 | 37.5 | 36.9 |
| MTP | 117.0 | 98.0 | 112.0 | 118.1 |
| MTP + n-gram — 19-prompt mixed average incl. code-edit prompts | 157.2 | 143.6 | 157.5 | 156.9 |
| MTP draft acceptance (MTP / MTP + n-gram) | 65.4% / 66.3% | 73.2% / 72.9% | 68.3% / 69.1% | 60.6% / 62.2% |
| Prefill tok/s, 2k / 8k / 32k | 3164 / 3302 / 2570 | 3181 / 3291 / 2563 | 3138 / 3307 / 2572 | 3223 / 3338 / 2602 |
| Plain decode (no MTP) after 16k context | 34.6 | 34.1 | 35.2 | 34.7 |
Why A is recommended: B's 8-bit head is the most expensive part of every draft step, so it is the slowest with MTP despite the highest acceptance rate. C is fastest without speculation and on llama.cpp, but in WHIRL A gets a dedicated Q6_K-head path that made it 4.5% faster than C with MTP (and equal with MTP + n-gram) in this comparison, while keeping a more precise head.
Swift A vs the reference Qwen3.8-27B MXFP4 GGUF, both served by WHIRL (MTP + n-gram) on the R9700. Settings: temperature 1.0, top_p 0.95, top_k 20, min_p 0, presence_penalty 0 (official sampling for both), thinking on (template default), max_tokens 16,384, 9 prompts × seeds 1 and 2 (n = 18 runs per model). Thinking tokens = reasoning_content counted with the model's tokenizer; wall = full non-streaming request time. * = hit the 16,384-token cap (value is a lower bound).
| Prompt | Swift think tokens (s1 / s2) | Swift wall s | Qwen think tokens (s1 / s2) | Qwen wall s |
|---|---|---|---|---|
| 你好 | 24 / 24 | 0.5 / 0.6 | 21 / 21 | 0.5 / 0.5 |
| Hi | 16 / 16 | 0.4 / 0.4 | 46 / 43 | 1.1 / 0.8 |
| zh: Python LRU | 12639* / 10619 | 161.5 / 135.2 | 16382* / 16383* | 189.8 / 199.2 |
| zh: TypeScript hook | 2148 / 7695 | 47.4 / 109.5 | 11795 / 10961 | 179.8 / 166.6 |
| zh: SQL report | 8961 / 7658 | 174.5 / 138.6 | 12893* / 13540* | 210.9 / 210.4 |
| zh: bug fix | 2842 / 2386 | 51.0 / 40.4 | 9598 / 5944 | 140.7 / 97.7 |
| en: ISO-8601 duration parser | 16384* / 16310* | 206.5 / 203.6 | 16382* / 16384* | 212.7 / 132.3 |
| en: Go bounded queue | 16384* / 16384* | 236.1 / 230.1 | 16383* / 16384* | 245.6 / 240.8 |
| en: JS debounce bug | 4020 / 4789 | 57.6 / 65.5 | 1855 / 4495 | 30.5 / 68.2 |
These files are a quantized (Object-form) Derivative Work of Swift 1.5 Qwen3.8-27B and are distributed under the same terms as upstream:
LICENSE (unmodified copy of the upstream LICENSE).LICENSE-APACHE-2.0.NOTICE: the upstream NOTICE verbatim, plus a "Changes in this redistribution" section describing the conversion and quantization.Notice for redistributors, reproduced verbatim from the appendix of the Swift Open License v1.0:
Copyright 2026 UkisAI. Swift Contribution licensed under the Swift Open
License v1.0 (https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b/blob/main/LICENSE).
Derivative of Qwen3.8-27B, Copyright 2026 Alibaba Cloud, Apache License 2.0.
Commercial-use threshold (summary, not legal advice — the LICENSE text governs): personal, research, educational, evaluation and commercial use are permitted for individuals and organizations whose gross revenue, together with all affiliates (entities controlling, controlled by, or under common control with them), is below US$1,000,000 over the most recently completed fiscal year. Commercial use by an entity at or above that threshold is not licensed under the Swift Open License and requires a separate Swift Enterprise License from UkisAI (contact). The threshold does not apply to qualified non-profit organizations' non-commercial or research use. The license terminates automatically on non-compliance.
Base model: Qwen3.8-27B is Copyright 2026 Alibaba Cloud and licensed under the Apache License 2.0. Per Section 6 of the Swift Open License, nothing in it limits your rights in the Qwen3.8-27B base model under Apache 2.0; the commercial limitation applies only to the Swift Contribution.
Upstream citation:
@misc{swift-1.5-qwen3.8-27b,
title = {Swift 1.5 Qwen3.8-27B},
author = {UkisAI},
year = {2026},
url = {https://huggingface.co/ukisai/Swift-1.5-Qwen3.8-27b}
}
mmproj-Swift-1.5-Qwen3.8-27B-F16.gguf(F16),搭配任一版本即可看圖(llama.cpp --mmproj)。--spec-type draft-mtp 做自我推測解碼。