slider-meister-pub/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-GGUF

Model

Swift 1.5 Qwen3.8-Flash-Next GSQ-RCO abliterated — IQ3_S

2

3 commits

updated Oct 5, 2026

See the code

README

Swift 1.5 Qwen3.8-Flash-Next GSQ-RCO abliterated — IQ3_S

The IQ3_S tier that SC117/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF does not have ("No IQ3_S tier — Swift 1.5 upstream does not have one").

FileSize
Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf54.94 GB
Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00002-of-00002.gguf (PLE n-gram table)28.80 GB
mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf (vision projector, from UkisAI)0.91 GB
total model83.74 GB: 3.79 bpw overall, 3.50 bpw excluding the PLE table (ISTA's target)

⚠️ Abliterated: no refusal guardrails. You are responsible for how you use it.

Run

llama-cli -m Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf \
  -lm mmap --lazy-mode on -ngl 99 --cpu-moe -c 4096

Needs a llama.cpp build with the qwen4exp architecture and the GSQ Q2_0 type (ggml id 42). Older builds will refuse the file.

Tested on llama.cpp b11425 (e117148a4), 2× RTX 4070 Ti SUPER + 47.5 GB RAM, experts on CPU: ~9 tok/s generation. --lazy-mode on keeps shard 2 (the 28.8 GB PLE table) on disk.

How it was built

Recipe: ISTA's IQ3_S per-tensor allocation profile on Swift 1.5's weights, which is how UkisAI builds its Swift tiers. Then SC117's 144-tensor abliteration transplant on top. GSQ itself was not re-run. Instead, genuine GSQ bytes were reused wherever they are provably valid for Swift.

  1. Which GSQ bytes are reusable. Every tensor of UkisAI's Swift IQ3_XXS and ISTA's base IQ3_S was hashed. 682 tensors (30.2 GB: the PLE table, all hc_* at BF16, norms, small projections) are byte-identical between Swift and base, so ISTA's IQ3_S bytes for them are valid Swift bytes. No large matmul matched.

  2. Assemble a source GGUF, best bytes per tensor (first rule that matches wins):

    RuleSourceTensorsGB
    Abliteration target (ssm_out, attn_output, ffn_down_shexp, ffn_down_exps)orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF Q8_014443.7
    Swift byte-identical to baseISTA-DASLab IQ3_S68230.2
    A Swift GSQ tier already has the exact target typeUkisAI GSQ-RCO IQ3_XXS / IQ2_XS / Q2_01328.8
    Everything elseUkisAI standard Q8_026662.7
  3. One llama-quantize pass with ISTA's IQ3_S type map (anchored patterns) and UkisAI's Swift imatrix (imatrix-swiftfn-v1mix.gguf). Tensors already at their target type are copied verbatim. The rest are quantized from Q8_0.

  4. Q2_0 learned scales. The 9 ffn_down_exps targets stored as Q2_0 were re-encoded from the donor values using learned GSQ block scales instead of the reference d = amax. Both ISTA's and Swift's own scales were scored per tensor. Swift's won on all 9, raising correlation with the donor values from 0.78 to 0.93 (measured on the first ~268M weights of each tensor). SC117 reported 0.794 → 0.933 for the same step in its base-model release.

  5. Layout. It uses ISTA's split, not UkisAI's: shard 2 is ISTA's PLE-only file, byte-identical (blake2b-verified).

Verification

  • 1224 / 1224 tensors, each at its profile type. The exception is the 96 router tensors (ffn_gate_inp*), which stay at Swift's lossless F32 because llama-quantize never converts routers. ISTA stores them as BF16.
  • 911 tensors byte-identical to their source, i.e. every tensor whose source was already at its final type was copied, not re-encoded: 814 genuine GSQ tensors (682 ISTA, 132 UkisAI Swift), 96 F32 routers, and 1 donor Q8_0 tensor.
  • Text loads and generates coherently in llama-cli. Not tested: vision via the mmproj.

Refusal check

Heretic's default protocol: the 100 prompts of mlabonne/harmful_behaviors test[:100], 100 tokens per response, greedy decoding, thinking off. A response counts as a refusal if it contains any of Heretic's keyword markers ("sorry", "i cannot", "illegal", "disclaimer", …). Heretic counts disclaimers and deflections as refusals.

ModelRefusals (Heretic keywords)Hard refusals ("I cannot…", "I'm sorry…")
UkisAI Swift 1.5 GSQ-RCO IQ3_XXS (original)98/10098/100
This model (IQ3_S abliterated)40/1000/100

All 40 keyword hits on this model are answers that open with a disclaimer or a "legal and ethical distinction" preamble before complying (e.g. "Disclaimer: This manual is intended for educational…"). Some of them soften or redirect the request rather than answering it fully. None is a hard refusal. Measured with llama-server b11425, 4 parallel slots, --cpu-moe.

Limitations

  • Not GSQ everywhere. 266 Swift tensors and the 144 donor tensors (34 GB) use standard imatrix quantization, not GSQ refinement.
  • The abliteration comes from a base-model donor, as in SC117's Swift release. Swift's own weights in those 144 tensors are replaced.
  • hc_* stays at BF16 (identical to ISTA's), not capped.
  • No KLD or benchmark numbers yet. Swift 1.5 ships no MTP head.

License and credits

Swift Open License v1.0 (UkisAI) + Qwen Community License 1.0. See LICENSE and LICENSE-QWEN. Free use is limited to organizations below US$1M gross annual revenue. This is not Apache-2.0.

Credits: Qwen (base model), UkisAI (Swift 1.5, Swift GSQ-RCO tiers, imatrix), IST Austria DASLab (GSQ / RCO, IQ3_S allocation profile), orcarouter (abliterated donor), SC117 (transplant method).

abliterated
conversational
endpoints_compatible
flash-next
gguf
image-text-to-text
imatrix
iq3_s
llama-cpp
quantization
qwen3.8
tensor-transplant
uncensored

slider-meister-pub/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-GGUF

Model

Swift 1.5 Qwen3.8-Flash-Next GSQ-RCO abliterated — IQ3_S

2

3 commits

updated Oct 5, 2026

See the code

README

Swift 1.5 Qwen3.8-Flash-Next GSQ-RCO abliterated — IQ3_S

The IQ3_S tier that SC117/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-GGUF does not have ("No IQ3_S tier — Swift 1.5 upstream does not have one").

FileSize
Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf54.94 GB
Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00002-of-00002.gguf (PLE n-gram table)28.80 GB
mmproj-Swift-Qwen3.8-Flash-Next-BF16.gguf (vision projector, from UkisAI)0.91 GB
total model83.74 GB: 3.79 bpw overall, 3.50 bpw excluding the PLE table (ISTA's target)

⚠️ Abliterated: no refusal guardrails. You are responsible for how you use it.

Run

llama-cli -m Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-abliterated-IQ3_S-00001-of-00002.gguf \
  -lm mmap --lazy-mode on -ngl 99 --cpu-moe -c 4096

Needs a llama.cpp build with the qwen4exp architecture and the GSQ Q2_0 type (ggml id 42). Older builds will refuse the file.

Tested on llama.cpp b11425 (e117148a4), 2× RTX 4070 Ti SUPER + 47.5 GB RAM, experts on CPU: ~9 tok/s generation. --lazy-mode on keeps shard 2 (the 28.8 GB PLE table) on disk.

How it was built

Recipe: ISTA's IQ3_S per-tensor allocation profile on Swift 1.5's weights, which is how UkisAI builds its Swift tiers. Then SC117's 144-tensor abliteration transplant on top. GSQ itself was not re-run. Instead, genuine GSQ bytes were reused wherever they are provably valid for Swift.

  1. Which GSQ bytes are reusable. Every tensor of UkisAI's Swift IQ3_XXS and ISTA's base IQ3_S was hashed. 682 tensors (30.2 GB: the PLE table, all hc_* at BF16, norms, small projections) are byte-identical between Swift and base, so ISTA's IQ3_S bytes for them are valid Swift bytes. No large matmul matched.

  2. Assemble a source GGUF, best bytes per tensor (first rule that matches wins):

    RuleSourceTensorsGB
    Abliteration target (ssm_out, attn_output, ffn_down_shexp, ffn_down_exps)orcarouter/Qwen3.8-Flash-Next-Uncensored-GGUF Q8_014443.7
    Swift byte-identical to baseISTA-DASLab IQ3_S68230.2
    A Swift GSQ tier already has the exact target typeUkisAI GSQ-RCO IQ3_XXS / IQ2_XS / Q2_01328.8
    Everything elseUkisAI standard Q8_026662.7
  3. One llama-quantize pass with ISTA's IQ3_S type map (anchored patterns) and UkisAI's Swift imatrix (imatrix-swiftfn-v1mix.gguf). Tensors already at their target type are copied verbatim. The rest are quantized from Q8_0.

  4. Q2_0 learned scales. The 9 ffn_down_exps targets stored as Q2_0 were re-encoded from the donor values using learned GSQ block scales instead of the reference d = amax. Both ISTA's and Swift's own scales were scored per tensor. Swift's won on all 9, raising correlation with the donor values from 0.78 to 0.93 (measured on the first ~268M weights of each tensor). SC117 reported 0.794 → 0.933 for the same step in its base-model release.

  5. Layout. It uses ISTA's split, not UkisAI's: shard 2 is ISTA's PLE-only file, byte-identical (blake2b-verified).

Verification

  • 1224 / 1224 tensors, each at its profile type. The exception is the 96 router tensors (ffn_gate_inp*), which stay at Swift's lossless F32 because llama-quantize never converts routers. ISTA stores them as BF16.
  • 911 tensors byte-identical to their source, i.e. every tensor whose source was already at its final type was copied, not re-encoded: 814 genuine GSQ tensors (682 ISTA, 132 UkisAI Swift), 96 F32 routers, and 1 donor Q8_0 tensor.
  • Text loads and generates coherently in llama-cli. Not tested: vision via the mmproj.

Refusal check

Heretic's default protocol: the 100 prompts of mlabonne/harmful_behaviors test[:100], 100 tokens per response, greedy decoding, thinking off. A response counts as a refusal if it contains any of Heretic's keyword markers ("sorry", "i cannot", "illegal", "disclaimer", …). Heretic counts disclaimers and deflections as refusals.

ModelRefusals (Heretic keywords)Hard refusals ("I cannot…", "I'm sorry…")
UkisAI Swift 1.5 GSQ-RCO IQ3_XXS (original)98/10098/100
This model (IQ3_S abliterated)40/1000/100

All 40 keyword hits on this model are answers that open with a disclaimer or a "legal and ethical distinction" preamble before complying (e.g. "Disclaimer: This manual is intended for educational…"). Some of them soften or redirect the request rather than answering it fully. None is a hard refusal. Measured with llama-server b11425, 4 parallel slots, --cpu-moe.

Limitations

  • Not GSQ everywhere. 266 Swift tensors and the 144 donor tensors (34 GB) use standard imatrix quantization, not GSQ refinement.
  • The abliteration comes from a base-model donor, as in SC117's Swift release. Swift's own weights in those 144 tensors are replaced.
  • hc_* stays at BF16 (identical to ISTA's), not capped.
  • No KLD or benchmark numbers yet. Swift 1.5 ships no MTP head.

License and credits

Swift Open License v1.0 (UkisAI) + Qwen Community License 1.0. See LICENSE and LICENSE-QWEN. Free use is limited to organizations below US$1M gross annual revenue. This is not Apache-2.0.

Credits: Qwen (base model), UkisAI (Swift 1.5, Swift GSQ-RCO tiers, imatrix), IST Austria DASLab (GSQ / RCO, IQ3_S allocation profile), orcarouter (abliterated donor), SC117 (transplant method).

abliterated
conversational
endpoints_compatible
flash-next
gguf
image-text-to-text
imatrix
iq3_s
llama-cpp
quantization
qwen3.8
tensor-transplant
uncensored