perkel/Qwen3.8-27B-Uncensored-MC

Model

Qwen3.8 27B-Uncensored-MC

0

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.8 27B-Uncensored-MC

Qwen3.8-27B-Uncensored, OrcaRouter's abliterated BF16 build of Qwen's Qwen3.8-27B, converted to MegaCapybara's MX weight formats. It comes in five sizes for one NVIDIA RTX 5090, made the same way as perkel/Qwen3.8-27B-MC.

Safety alignment largely removed. The source model had its refusal behaviour removed by abliteration. It follows requests that the original Qwen3.8-27B would refuse, including harmful ones, and has no built-in guardrails.

It is meant for research, red-teaming, refusal and interpretability studies, and controlled experiments. Do not serve it to others without your own safety and moderation layers. You are responsible for how you use it and for what it generates, under the Apache 2.0 license and the laws that apply to you.

Run it with MegaCapybara

These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).

  1. Download MegaCapybara from its releases and unpack it.
  2. Run MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository's list (Qwen3.8-27B-Uncensored-MC); what it needs beside it comes with it.
  3. Pick a preset, press Load server, and point your client at http://127.0.0.1:8080/v1.

Without the launcher, put the files in the package's model folder and start a preset: the usage guide has the details.

The source

  • What OrcaRouter changed (from their model card, after Arditi et al. 2024, Refusal in Language Models Is Mediated by a Single Direction):
    • One refusal direction was estimated at layer 38.
    • It was orthogonalized out of every matrix that writes to the residual stream: the attention and DeltaNet output projections, the MLP down projections, the embedding and the MTP head's. That is 131 matrices.
    • The vision tower is untouched.
  • What stays the same: the architecture, tensor names, tokenizer and chat template are Qwen's.
  • On our test text, its full-precision outputs match the original's within 0.5% in perplexity on reasoning, math, code and wiki. They differ on chat (where refusals live) and on tool-call text.

Files

FileLauncher nameWeight formatsOn diskIn VRAMKLTop-1
qwen3.8-27b-uncensored-tiny.mcapyQwen3.8 27B-Uncensored-MCTinyall MXFP416.6 GB12.70 GiB0.066292.7%
qwen3.8-27b-uncensored-small.mcapyQwen3.8 27B-Uncensored-MCSmallMXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs17.6 GB13.67 GiB0.057493.7%
qwen3.8-27b-uncensored-medium.mcapyQwen3.8 27B-Uncensored-MCMediumMXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs19.5 GB15.39 GiB0.026195.6%
qwen3.8-27b-uncensored-large.mcapyQwen3.8 27B-Uncensored-MCLargeall MXFP622.9 GB18.66 GiB0.007497.9%
qwen3.8-27b-uncensored-xxl.mcapyQwen3.8 27B-Uncensored-MCXXLMXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms)28.2 GB23.58 GiB0.005798.2%

You need one model file. These support files are shared by all five:

FileWhat it isOn disk
qwen3.8-27b-dflash2.mcapyDFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster)1.3 GB
qwen3.8-27b-uncensored-mtp.mcapythe uncensored model's own MTP layer (abliterated with it) as a smaller, slower drafter0.28 GB
qwen3.8-27b-vision-fp8.mcapythe vision tower in FP8 (images; nearly exact; abliteration left it unchanged)0.57 GB
qwen3.8-27b-vision-bf16.mcapythe original BF16 vision tower1.0 GB
  • Using both repositories: they can share one model\ folder. The launcher lists each model by its name, and gives each the MTP layer of its own source.
  • The drafter: DFlash2 was trained on the original model and is used here as is. Drafting never changes the answers, since the model checks every drafted token. It may change speed; see below.

Accuracy against its own original

Top-1 by size: our uncensored models beside Unsloth's GGUFs of the original

  • What KL and top-1 measure here: KL divergence and top-1 agreement against the uncensored model's own BF16 weights, computed with FP32 math. This is how much the conversion loses; the uncensoring itself is not counted.
  • The test text: 81,880 held-out tokens of reasoning traces, math, code, chat and wiki.
  • Settings: the engine at 16-bit activations and an unquantized KV cache.
  • The chart:
    • Unsloth's points are their published top-1 for GGUFs of Qwen's original model, on their own text. They are there to show what each size costs, not to compare the two models.
    • Sizes are the bytes on the GPU (no embedding table or MTP layer).
  • Against our conversion of the original: top-1 is the same at every size, within 0.4 points.
KL per kind of text
SizeReasoningMathCodeChatWiki
Tiny0.01380.01560.05080.20940.0414
Small0.01070.01390.04500.18990.0277
Medium0.00590.00800.02820.06890.0196
Large0.00090.00150.00920.02270.0028
XXL0.00070.00140.00810.01620.0022

Speed on an RTX 5090

Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.

SizeCoding, 1 conversationProse, 1 conversationCoding, 8 at once (total)Reading an 8K-token prompt
Tiny4932501,8567,854
Small4882361,8287,697
Medium4452191,6697,365
Large3751781,4446,982
XXL3121491,3146,212
  • Against our conversion of the original: the same within run-to-run differences, except Tiny's coding: 493 against 540. These are single runs.
  • Conditions: measured with nothing else on the card. Its memory was overclocked to 16.8 GHz (+20%); a stock card is slower.

How the files were made

  • Source: orcarouter/Qwen3.8-27B-Uncensored, revision 404ea47, BF16. It is the reference for every score above.

  • Recipes: the same as for perkel/Qwen3.8-27B-MC.

    • The weights are rounded layer by layer with GPTQ and MX block scales (32 weights share a power-of-two scale).
    • The calibration text is the same 514K tokens.
    • Each size uses the same per-layer format choices.
  • Storage changes for the engine:

    • norm gains are folded into the following matrices;
    • the attention value rows are Hadamard-rotated;
    • the output head's rows are ordered by token frequency.

    None of these change the model's function. XXL's 16-bit parts are stored as two FP8 E4M3 terms.

License

Apache 2.0, as the original Qwen3.8-27B, OrcaRouter's Qwen3.8-27B-Uncensored and Qwen3.8-27B-DFlash2. The changes are listed in NOTICE. Converted by Perkel's Software Corner.

abliterated
blackwell
dflash2
image-text-to-text
megacapybara
mxfp4
mxfp6
quantized
qwen3.8
rtx-5090
speculative-decoding
uncensored

perkel/Qwen3.8-27B-Uncensored-MC

Model

Qwen3.8 27B-Uncensored-MC

0

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.8 27B-Uncensored-MC

Qwen3.8-27B-Uncensored, OrcaRouter's abliterated BF16 build of Qwen's Qwen3.8-27B, converted to MegaCapybara's MX weight formats. It comes in five sizes for one NVIDIA RTX 5090, made the same way as perkel/Qwen3.8-27B-MC.

Safety alignment largely removed. The source model had its refusal behaviour removed by abliteration. It follows requests that the original Qwen3.8-27B would refuse, including harmful ones, and has no built-in guardrails.

It is meant for research, red-teaming, refusal and interpretability studies, and controlled experiments. Do not serve it to others without your own safety and moderation layers. You are responsible for how you use it and for what it generates, under the Apache 2.0 license and the laws that apply to you.

Run it with MegaCapybara

These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).

  1. Download MegaCapybara from its releases and unpack it.
  2. Run MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository's list (Qwen3.8-27B-Uncensored-MC); what it needs beside it comes with it.
  3. Pick a preset, press Load server, and point your client at http://127.0.0.1:8080/v1.

Without the launcher, put the files in the package's model folder and start a preset: the usage guide has the details.

The source

  • What OrcaRouter changed (from their model card, after Arditi et al. 2024, Refusal in Language Models Is Mediated by a Single Direction):
    • One refusal direction was estimated at layer 38.
    • It was orthogonalized out of every matrix that writes to the residual stream: the attention and DeltaNet output projections, the MLP down projections, the embedding and the MTP head's. That is 131 matrices.
    • The vision tower is untouched.
  • What stays the same: the architecture, tensor names, tokenizer and chat template are Qwen's.
  • On our test text, its full-precision outputs match the original's within 0.5% in perplexity on reasoning, math, code and wiki. They differ on chat (where refusals live) and on tool-call text.

Files

FileLauncher nameWeight formatsOn diskIn VRAMKLTop-1
qwen3.8-27b-uncensored-tiny.mcapyQwen3.8 27B-Uncensored-MCTinyall MXFP416.6 GB12.70 GiB0.066292.7%
qwen3.8-27b-uncensored-small.mcapyQwen3.8 27B-Uncensored-MCSmallMXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs17.6 GB13.67 GiB0.057493.7%
qwen3.8-27b-uncensored-medium.mcapyQwen3.8 27B-Uncensored-MCMediumMXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs19.5 GB15.39 GiB0.026195.6%
qwen3.8-27b-uncensored-large.mcapyQwen3.8 27B-Uncensored-MCLargeall MXFP622.9 GB18.66 GiB0.007497.9%
qwen3.8-27b-uncensored-xxl.mcapyQwen3.8 27B-Uncensored-MCXXLMXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms)28.2 GB23.58 GiB0.005798.2%

You need one model file. These support files are shared by all five:

FileWhat it isOn disk
qwen3.8-27b-dflash2.mcapyDFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster)1.3 GB
qwen3.8-27b-uncensored-mtp.mcapythe uncensored model's own MTP layer (abliterated with it) as a smaller, slower drafter0.28 GB
qwen3.8-27b-vision-fp8.mcapythe vision tower in FP8 (images; nearly exact; abliteration left it unchanged)0.57 GB
qwen3.8-27b-vision-bf16.mcapythe original BF16 vision tower1.0 GB
  • Using both repositories: they can share one model\ folder. The launcher lists each model by its name, and gives each the MTP layer of its own source.
  • The drafter: DFlash2 was trained on the original model and is used here as is. Drafting never changes the answers, since the model checks every drafted token. It may change speed; see below.

Accuracy against its own original

Top-1 by size: our uncensored models beside Unsloth's GGUFs of the original

  • What KL and top-1 measure here: KL divergence and top-1 agreement against the uncensored model's own BF16 weights, computed with FP32 math. This is how much the conversion loses; the uncensoring itself is not counted.
  • The test text: 81,880 held-out tokens of reasoning traces, math, code, chat and wiki.
  • Settings: the engine at 16-bit activations and an unquantized KV cache.
  • The chart:
    • Unsloth's points are their published top-1 for GGUFs of Qwen's original model, on their own text. They are there to show what each size costs, not to compare the two models.
    • Sizes are the bytes on the GPU (no embedding table or MTP layer).
  • Against our conversion of the original: top-1 is the same at every size, within 0.4 points.
KL per kind of text
SizeReasoningMathCodeChatWiki
Tiny0.01380.01560.05080.20940.0414
Small0.01070.01390.04500.18990.0277
Medium0.00590.00800.02820.06890.0196
Large0.00090.00150.00920.02270.0028
XXL0.00070.00140.00810.01620.0022

Speed on an RTX 5090

Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.

SizeCoding, 1 conversationProse, 1 conversationCoding, 8 at once (total)Reading an 8K-token prompt
Tiny4932501,8567,854
Small4882361,8287,697
Medium4452191,6697,365
Large3751781,4446,982
XXL3121491,3146,212
  • Against our conversion of the original: the same within run-to-run differences, except Tiny's coding: 493 against 540. These are single runs.
  • Conditions: measured with nothing else on the card. Its memory was overclocked to 16.8 GHz (+20%); a stock card is slower.

How the files were made

  • Source: orcarouter/Qwen3.8-27B-Uncensored, revision 404ea47, BF16. It is the reference for every score above.

  • Recipes: the same as for perkel/Qwen3.8-27B-MC.

    • The weights are rounded layer by layer with GPTQ and MX block scales (32 weights share a power-of-two scale).
    • The calibration text is the same 514K tokens.
    • Each size uses the same per-layer format choices.
  • Storage changes for the engine:

    • norm gains are folded into the following matrices;
    • the attention value rows are Hadamard-rotated;
    • the output head's rows are ordered by token frequency.

    None of these change the model's function. XXL's 16-bit parts are stored as two FP8 E4M3 terms.

License

Apache 2.0, as the original Qwen3.8-27B, OrcaRouter's Qwen3.8-27B-Uncensored and Qwen3.8-27B-DFlash2. The changes are listed in NOTICE. Converted by Perkel's Software Corner.

abliterated
blackwell
dflash2
image-text-to-text
megacapybara
mxfp4
mxfp6
quantized
qwen3.8
rtx-5090
speculative-decoding
uncensored