nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF

Model

Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)

0

5 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)

[!IMPORTANT]

⚠️ Runtime Compatibility Notice

This is a qwen4exp architecture GGUF (Gated-DeltaNet hybrid + 512-expert routed MoE + MTP). Stock llama.cpp does not support this architecture. Use a qwen4exp-enabled llama.cpp build (e.g. the unslothai/llama.cpp fork with qwen4exp support) with --mtp-draft for speculative decoding.

Overview

Swift 1.5 Flash Miroslav2.0 — the Swift-optimized (UkisAI "Swift" reasoning-token reduction, ~58% fewer reasoning tokens, hesitation-loop elimination) variant of Qwen3.8 Flash Next, quantized for local serving on Apple Silicon.

  • Build: Q4_0 base re-quantized with Q8_0 donor tensors spliced into the output/embedding/residual tensors (-Q8out-v3 pipeline: 575 tensors upgraded to Q8_0 from the full-precision donor shards), imatrix-calibrated (imatrix-bartv6, calibration dataset Qwen3.8-Flash-Next-calibration-v6)
  • Architecture: qwen4exp — 48 layers (GDN recurrent + full attention every 4th), 512 routed experts (10 active), shared expert, hyper-connections, ngram-PLE embedding, built-in MTP heads
  • Context: 262,144 tokens native
  • Shards: 3 (-00001-of-00003.gguf … -00003-of-00003.gguf), 95 GiB total
  • Vision: vision projector tensors included (image-text-to-text capable)

Files

FileSize
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf42.4 GiB
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00002-of-00003.gguf42.8 GiB
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00003-of-00003.gguf10.3 GiB

Download all three shards into the same directory, then load shard 1.

Usage (qwen4exp-enabled llama.cpp)

llama-server -m Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf \
  --mmproj mmproj-f16.gguf -ngl 99 -c 262144 --port 8080

Built and benchmarked locally on Apple M5 Pro (64 GB unified memory) with the Slipstream streaming-MoE llama.cpp fork.

apple-silicon
conversational
endpoints_compatible
gguf
image-text-to-text
imatrix
metal
q4_0
qwen4exp
speculative-decoding
text-generation
vision

nitinpanj/Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-GGUF

Model

Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)

0

5 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

Swift-1.5-Qwen3.8-Flash-Next Q4_0-Q8_0-out v3 (GGUF)

[!IMPORTANT]

⚠️ Runtime Compatibility Notice

This is a qwen4exp architecture GGUF (Gated-DeltaNet hybrid + 512-expert routed MoE + MTP). Stock llama.cpp does not support this architecture. Use a qwen4exp-enabled llama.cpp build (e.g. the unslothai/llama.cpp fork with qwen4exp support) with --mtp-draft for speculative decoding.

Overview

Swift 1.5 Flash Miroslav2.0 — the Swift-optimized (UkisAI "Swift" reasoning-token reduction, ~58% fewer reasoning tokens, hesitation-loop elimination) variant of Qwen3.8 Flash Next, quantized for local serving on Apple Silicon.

  • Build: Q4_0 base re-quantized with Q8_0 donor tensors spliced into the output/embedding/residual tensors (-Q8out-v3 pipeline: 575 tensors upgraded to Q8_0 from the full-precision donor shards), imatrix-calibrated (imatrix-bartv6, calibration dataset Qwen3.8-Flash-Next-calibration-v6)
  • Architecture: qwen4exp — 48 layers (GDN recurrent + full attention every 4th), 512 routed experts (10 active), shared expert, hyper-connections, ngram-PLE embedding, built-in MTP heads
  • Context: 262,144 tokens native
  • Shards: 3 (-00001-of-00003.gguf … -00003-of-00003.gguf), 95 GiB total
  • Vision: vision projector tensors included (image-text-to-text capable)

Files

FileSize
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf42.4 GiB
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00002-of-00003.gguf42.8 GiB
Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00003-of-00003.gguf10.3 GiB

Download all three shards into the same directory, then load shard 1.

Usage (qwen4exp-enabled llama.cpp)

llama-server -m Swift-Qwen3.8-Flash-Next-Q4_0-Q8out-v3-00001-of-00003.gguf \
  --mmproj mmproj-f16.gguf -ngl 99 -c 262144 --port 8080

Built and benchmarked locally on Apple M5 Pro (64 GB unified memory) with the Slipstream streaming-MoE llama.cpp fork.

apple-silicon
conversational
endpoints_compatible
gguf
image-text-to-text
imatrix
metal
q4_0
qwen4exp
speculative-decoding
text-generation
vision