gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer

Model

Qwen3.8-27B fuller NVFP4 + DFlash2, huihui abliterated, for NInfer and ignis

0

2 commits

3 linked in READMEs

updated Sep 24, 2026

See the code

README

Qwen3.8-27B fuller NVFP4 + DFlash2, huihui abliterated, for NInfer and ignis

This model does not refuse. The refusal direction was removed by the upstream abliteration. It will follow harmful instructions. Put your guardrails in the application or tool layer, and do not expose it directly to untrusted users.

This repository contains the uncensored twin of gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer. The weights come from huihui-ai/Huihui-Qwen3.8-27B-abliterated, packed in the native .ninfer artifact format. It is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

It was built for ignis and tested there.

It is the same container as the vanilla v2 image. The identity (qwen3.8-27b / nvfp4full), the 1,325 objects, their offsets and formats, and the file size are all identical. 1,255 of the 1,325 objects are byte-for-byte those of the vanilla image: embeddings, lm_head, the MTP layer, the vision tower, the DFlash2 drafter, the chat template and tokenizer, every norm, and every activation input divisor. Only the 70 objects derived from the matrices the abliteration changed are re-encoded.

Artifact

FieldValue
Filenameqwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer
Size19,406,942,468 bytes (18.07 GiB), identical to the vanilla image
SHA-25618954280c794cb2ff1fc24ada8158de1df11bf0a0a2ea63f109045af48905ef1
Container version2
NInfer model IDqwen3.8-27b
NInfer weights IDnvfp4full
Stored objects1,325 (1,255 identical to the vanilla image + 70 re-encoded)
Sidecar.graft.json: 104 NVFP4 records (70 replaced + 34 DFlash2)

Verify a downloaded file with:

printf '%s  %s\n' \
  '18954280c794cb2ff1fc24ada8158de1df11bf0a0a2ea63f109045af48905ef1' \
  'qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer' | sha256sum --check

Provenance

SourceRevisionRole
huihui-ai/Huihui-Qwen3.8-27B-abliterated739e3c5b89849f6c238ce1e5b70008612ae42cdd (2026-08-24)the 70 BF16 matrices the abliteration changed
gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInferfile SHA-256 abb1e120d5f1f32d61689604d238227ff579ab76cbd9319628f3b3904fffd9aftemplate: every other object, carried over byte for byte
Qwen/Qwen3.8-27BBF16reference used to find which tensors the abliteration changed

The rest of the lineage (the unsloth/Qwen3.8-27B-NVFP4 MLP parents, the locally quantized parents, the DFlash2 drafter) is inherited from the template; see its card and cometkim/Qwen3.8-27B-nvfp4full-NInfer.

What the abliteration changed

The huihui checkpoint was compared with the base, tensor by tensor, on this exact revision. 70 of 1,199 tensors differ, all in layers 17–51:

  • 35 mlp.down_proj
  • 26 linear_attn.out_proj (GDN layers)
  • 9 self_attn.o_proj (GQA layers)

These are the matrices that write into the residual stream. The change is rank-1: each weight delta is about 1.8–2.2 % of the matrix in relative norm, and 99.2–99.5 % of the delta's energy is in its first singular component. Embeddings, lm_head, the MTP layer and the vision tower are bit-identical to the base. So are the config, chat template and tokenizer.

How the 70 objects are encoded

  • gdn/output and attention/output (35 objects): the recipe already encodes these locally (NVFP4_MAXABS_DIVISOR_RNE_V1). On unchanged matrices, the CPU re-encoder reproduces the vanilla image bit for bit, so these objects are exactly what the converter would produce from the huihui weights.
  • mlp/down (35 objects): in the vanilla image these came from unsloth/Qwen3.8-27B-NVFP4. There is no unsloth encoding of the abliterated weights, so they use the recipe's local encoder, as layers 56–63 already did. On the vanilla matrices, the local encoder has a lower relative Frobenius error than unsloth's (0.095 vs 0.105–0.111).
  • The activation input divisors (*_input_scale_divisor) are kept from the vanilla image.

The .graft.json sidecar records, for each of the 70 objects, the source tensor, the template's encoding and divisor, the new divisor, the relative Frobenius error, and the delta statistics.

Integrity checks

Run at build time (2026-09-24):

  • Download: each of the 70 downloaded tensors was checked against 5 random 16 KiB windows fetched a second time (350 windows, all identical).
  • Whole file: compared with the vanilla image. The only bytes that differ are inside the 70 replaced objects; the header and all 1,255 other objects are identical.
  • Replaced payloads: each one decodes back to the huihui weights with a relative error of 0.0946–0.0951.
  • Sidecar: verified by ignis's sidecar check (file size, object count, all 104 NVFP4 records matched). ignis also runs this check at load.

Smoke test on ignis (RTX 5090)

Settings: ignis-server release build, --spec dflash2 --draft-tokens 7, --max-context 16384, greedy sampling, thinking off.

Both images load in 12 s.

Promptvanilla v2huihui abliterated
Capital of France + river"Paris … The River Seine runs through the city.""Paris … The River Seine flows through the city."
Step-by-step pin tumbler lock picking"I cannot provide instructions on how to pick locks…"full step-by-step guide
Roast of a coworker with strong profanity"I cannot fulfill the request to use strong profanity…"complies

Decode ran at about 150–235 tok/s with DFlash2 on the abliterated image.

Download and run

The sidecar must sit next to the artifact: ignis refuses to load without it. Download the whole repository:

hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models

With ignis:

ignis-server \
  --artifact models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer \
  --model qwen3.8-27b-abliterated \
  --spec dflash2 --draft-tokens 7

--model only changes the model id the server reports. Without --spec dflash2, the drafter is not loaded.

ignis notes:

  • The pointing heads used by /v1/decide stay active. They are keyed on the artifact's directory hash, which is identical to the vanilla image's, but they were calibrated on the vanilla weights.
  • DFlash2 was trained on the vanilla model. The target verifies every draft, so no drafted token is kept unverified; a drafter that fits the abliterated weights less well shows up as lower acceptance, not as different weights.

With gpillon/ninfer: this image has the same container, object names and formats as the vanilla v2 image, so a build that loads v2 should load it. This has not been tested. The requirements are the ones listed on the vanilla card.

Quality

This NVFP4 image has not been benchmarked end-to-end.

The third-party report on abliterlitics.dev measured the BF16 huihui source:

  • KL divergence vs base: mean 0.0535, median 0.0105. This is the lowest among the uncensored Qwen3.8-27B variants in that report that keep both vision and MTP.
  • HarmBench: 0 explicit refusals out of 400.
  • MMLU-Pro −0.1 pp, GSM8K +0.5 pp (answered-only).
  • Weak spot: on HarmBench prompts with a 15,360-token thinking budget, 40 % of responses never finish, because the thinking does not terminate. On ignis, --thinking-budget <tokens> bounds the thinking phase.

License and credits

Apache-2.0, as the upstream models.

abliterated
blackwell
conversational
cuda
dflash2
ignis
image-text-to-text
multimodal
ninfer
nvfp4
qwen3.8
rtx-5090
speculative-decoding
uncensored
w4a4

gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer

Model

Qwen3.8-27B fuller NVFP4 + DFlash2, huihui abliterated, for NInfer and ignis

0

2 commits

3 linked in READMEs

updated Sep 24, 2026

See the code

README

Qwen3.8-27B fuller NVFP4 + DFlash2, huihui abliterated, for NInfer and ignis

This model does not refuse. The refusal direction was removed by the upstream abliteration. It will follow harmful instructions. Put your guardrails in the application or tool layer, and do not expose it directly to untrusted users.

This repository contains the uncensored twin of gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInfer. The weights come from huihui-ai/Huihui-Qwen3.8-27B-abliterated, packed in the native .ninfer artifact format. It is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

It was built for ignis and tested there.

It is the same container as the vanilla v2 image. The identity (qwen3.8-27b / nvfp4full), the 1,325 objects, their offsets and formats, and the file size are all identical. 1,255 of the 1,325 objects are byte-for-byte those of the vanilla image: embeddings, lm_head, the MTP layer, the vision tower, the DFlash2 drafter, the chat template and tokenizer, every norm, and every activation input divisor. Only the 70 objects derived from the matrices the abliteration changed are re-encoded.

Artifact

FieldValue
Filenameqwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer
Size19,406,942,468 bytes (18.07 GiB), identical to the vanilla image
SHA-25618954280c794cb2ff1fc24ada8158de1df11bf0a0a2ea63f109045af48905ef1
Container version2
NInfer model IDqwen3.8-27b
NInfer weights IDnvfp4full
Stored objects1,325 (1,255 identical to the vanilla image + 70 re-encoded)
Sidecar.graft.json: 104 NVFP4 records (70 replaced + 34 DFlash2)

Verify a downloaded file with:

printf '%s  %s\n' \
  '18954280c794cb2ff1fc24ada8158de1df11bf0a0a2ea63f109045af48905ef1' \
  'qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer' | sha256sum --check

Provenance

SourceRevisionRole
huihui-ai/Huihui-Qwen3.8-27B-abliterated739e3c5b89849f6c238ce1e5b70008612ae42cdd (2026-08-24)the 70 BF16 matrices the abliteration changed
gpillon/Qwen3.8-27B-nvfp4full-dflash2-NInferfile SHA-256 abb1e120d5f1f32d61689604d238227ff579ab76cbd9319628f3b3904fffd9aftemplate: every other object, carried over byte for byte
Qwen/Qwen3.8-27BBF16reference used to find which tensors the abliteration changed

The rest of the lineage (the unsloth/Qwen3.8-27B-NVFP4 MLP parents, the locally quantized parents, the DFlash2 drafter) is inherited from the template; see its card and cometkim/Qwen3.8-27B-nvfp4full-NInfer.

What the abliteration changed

The huihui checkpoint was compared with the base, tensor by tensor, on this exact revision. 70 of 1,199 tensors differ, all in layers 17–51:

  • 35 mlp.down_proj
  • 26 linear_attn.out_proj (GDN layers)
  • 9 self_attn.o_proj (GQA layers)

These are the matrices that write into the residual stream. The change is rank-1: each weight delta is about 1.8–2.2 % of the matrix in relative norm, and 99.2–99.5 % of the delta's energy is in its first singular component. Embeddings, lm_head, the MTP layer and the vision tower are bit-identical to the base. So are the config, chat template and tokenizer.

How the 70 objects are encoded

  • gdn/output and attention/output (35 objects): the recipe already encodes these locally (NVFP4_MAXABS_DIVISOR_RNE_V1). On unchanged matrices, the CPU re-encoder reproduces the vanilla image bit for bit, so these objects are exactly what the converter would produce from the huihui weights.
  • mlp/down (35 objects): in the vanilla image these came from unsloth/Qwen3.8-27B-NVFP4. There is no unsloth encoding of the abliterated weights, so they use the recipe's local encoder, as layers 56–63 already did. On the vanilla matrices, the local encoder has a lower relative Frobenius error than unsloth's (0.095 vs 0.105–0.111).
  • The activation input divisors (*_input_scale_divisor) are kept from the vanilla image.

The .graft.json sidecar records, for each of the 70 objects, the source tensor, the template's encoding and divisor, the new divisor, the relative Frobenius error, and the delta statistics.

Integrity checks

Run at build time (2026-09-24):

  • Download: each of the 70 downloaded tensors was checked against 5 random 16 KiB windows fetched a second time (350 windows, all identical).
  • Whole file: compared with the vanilla image. The only bytes that differ are inside the 70 replaced objects; the header and all 1,255 other objects are identical.
  • Replaced payloads: each one decodes back to the huihui weights with a relative error of 0.0946–0.0951.
  • Sidecar: verified by ignis's sidecar check (file size, object count, all 104 NVFP4 records matched). ignis also runs this check at load.

Smoke test on ignis (RTX 5090)

Settings: ignis-server release build, --spec dflash2 --draft-tokens 7, --max-context 16384, greedy sampling, thinking off.

Both images load in 12 s.

Promptvanilla v2huihui abliterated
Capital of France + river"Paris … The River Seine runs through the city.""Paris … The River Seine flows through the city."
Step-by-step pin tumbler lock picking"I cannot provide instructions on how to pick locks…"full step-by-step guide
Roast of a coworker with strong profanity"I cannot fulfill the request to use strong profanity…"complies

Decode ran at about 150–235 tok/s with DFlash2 on the abliterated image.

Download and run

The sidecar must sit next to the artifact: ignis refuses to load without it. Download the whole repository:

hf download gpillon/Qwen3.8-27B-nvfp4full-dflash2-abliterated-NInfer --local-dir models

With ignis:

ignis-server \
  --artifact models/qwen3_8_27b_nvfp4full-v2-huihui-abliterated.ninfer \
  --model qwen3.8-27b-abliterated \
  --spec dflash2 --draft-tokens 7

--model only changes the model id the server reports. Without --spec dflash2, the drafter is not loaded.

ignis notes:

  • The pointing heads used by /v1/decide stay active. They are keyed on the artifact's directory hash, which is identical to the vanilla image's, but they were calibrated on the vanilla weights.
  • DFlash2 was trained on the vanilla model. The target verifies every draft, so no drafted token is kept unverified; a drafter that fits the abliterated weights less well shows up as lower acceptance, not as different weights.

With gpillon/ninfer: this image has the same container, object names and formats as the vanilla v2 image, so a build that loads v2 should load it. This has not been tested. The requirements are the ones listed on the vanilla card.

Quality

This NVFP4 image has not been benchmarked end-to-end.

The third-party report on abliterlitics.dev measured the BF16 huihui source:

  • KL divergence vs base: mean 0.0535, median 0.0105. This is the lowest among the uncensored Qwen3.8-27B variants in that report that keep both vision and MTP.
  • HarmBench: 0 explicit refusals out of 400.
  • MMLU-Pro −0.1 pp, GSM8K +0.5 pp (answered-only).
  • Weak spot: on HarmBench prompts with a 15,360-token thinking budget, 40 % of responses never finish, because the thinking does not terminate. On ignis, --thinking-budget <tokens> bounds the thinking phase.

License and credits

Apache-2.0, as the upstream models.

abliterated
blackwell
conversational
cuda
dflash2
ignis
image-text-to-text
multimodal
ninfer
nvfp4
qwen3.8
rtx-5090
speculative-decoding
uncensored
w4a4