Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)
0
5 commits
1 linked in READMEs
updated Sep 28, 2026
[!IMPORTANT]
⚠️ Runtime Compatibility Notice
This is a
qwen4exparchitecture GGUF (Gated-DeltaNet hybrid + 512-expert routed MoE + MTP). Stockllama.cppdoes not support this architecture. Use aqwen4exp-enabled llama.cpp build (e.g. theunslothai/llama.cppfork with qwen4exp support) with--mtp-draftfor speculative decoding.
Swift 1.5 Flash Miroslav2.0 — the Swift-optimized (UkisAI "Swift" reasoning-token reduction, ~58% fewer reasoning tokens, hesitation-loop elimination) variant of Qwen3.8 Flash Next, quantized for local serving on Apple Silicon.
-Q8out-v3 pipeline: 575 tensors upgraded to Q8_0 from the full-precision donor shards),
imatrix-calibrated (imatrix-bartv6, calibration dataset Qwen3.8-Flash-Next-calibration-v6)qwen4exp — 48 layers (GDN recurrent + full attention every 4th), 512 routed experts
(10 active), shared expert, hyper-connections, ngram-PLE embedding, built-in MTP heads-00001-of-00003.gguf … -00003-of-00003.gguf), 95 GiB total| File | Size |
|---|---|
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf | 42.4 GiB |
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00002-of-00003.gguf | 42.8 GiB |
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00003-of-00003.gguf | 10.3 GiB |
Download all three shards into the same directory, then load shard 1.
llama-server -m Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf \
--mmproj mmproj-f16.gguf -ngl 99 -c 262144 --port 8080
Built and benchmarked locally on Apple M5 Pro (64 GB unified memory) with the Slipstream streaming-MoE llama.cpp fork.
Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)
0
5 commits
1 linked in READMEs
updated Sep 28, 2026
[!IMPORTANT]
⚠️ Runtime Compatibility Notice
This is a
qwen4exparchitecture GGUF (Gated-DeltaNet hybrid + 512-expert routed MoE + MTP). Stockllama.cppdoes not support this architecture. Use aqwen4exp-enabled llama.cpp build (e.g. theunslothai/llama.cppfork with qwen4exp support) with--mtp-draftfor speculative decoding.
Swift 1.5 Flash Miroslav2.0 — the Swift-optimized (UkisAI "Swift" reasoning-token reduction, ~58% fewer reasoning tokens, hesitation-loop elimination) variant of Qwen3.8 Flash Next, quantized for local serving on Apple Silicon.
-Q8out-v3 pipeline: 575 tensors upgraded to Q8_0 from the full-precision donor shards),
imatrix-calibrated (imatrix-bartv6, calibration dataset Qwen3.8-Flash-Next-calibration-v6)qwen4exp — 48 layers (GDN recurrent + full attention every 4th), 512 routed experts
(10 active), shared expert, hyper-connections, ngram-PLE embedding, built-in MTP heads-00001-of-00003.gguf … -00003-of-00003.gguf), 95 GiB total| File | Size |
|---|---|
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf | 42.4 GiB |
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00002-of-00003.gguf | 42.8 GiB |
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00003-of-00003.gguf | 10.3 GiB |
Download all three shards into the same directory, then load shard 1.
llama-server -m Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf \
--mmproj mmproj-f16.gguf -ngl 99 -c 262144 --port 8080
Built and benchmarked locally on Apple M5 Pro (64 GB unified memory) with the Slipstream streaming-MoE llama.cpp fork.