Qwen3.8-27B-Uncensored, OrcaRouter's abliterated BF16 build of Qwen's Qwen3.8-27B, converted to MegaCapybara's MX weight formats. It comes in five sizes for one NVIDIA RTX 5090, made the same way as perkel/Qwen3.8-27B-MC.
Safety alignment largely removed. The source model had its refusal behaviour removed by abliteration. It follows requests that the original Qwen3.8-27B would refuse, including harmful ones, and has no built-in guardrails.
It is meant for research, red-teaming, refusal and interpretability studies, and controlled experiments. Do not serve it to others without your own safety and moderation layers. You are responsible for how you use it and for what it generates, under the Apache 2.0 license and the laws that apply to you.
These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).
MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository's
list (Qwen3.8-27B-Uncensored-MC); what it needs beside it comes with it.http://127.0.0.1:8080/v1.Without the launcher, put the files in the package's model folder and start a preset: the
usage guide has the details.
| File | Launcher name | Weight formats | On disk | In VRAM | KL | Top-1 |
|---|---|---|---|---|---|---|
qwen3.8-27b-uncensored-tiny.mcapy | Qwen3.8 27B-Uncensored-MCTiny | all MXFP4 | 16.6 GB | 12.70 GiB | 0.0662 | 92.7% |
qwen3.8-27b-uncensored-small.mcapy | Qwen3.8 27B-Uncensored-MCSmall | MXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs | 17.6 GB | 13.67 GiB | 0.0574 | 93.7% |
qwen3.8-27b-uncensored-medium.mcapy | Qwen3.8 27B-Uncensored-MCMedium | MXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs | 19.5 GB | 15.39 GiB | 0.0261 | 95.6% |
qwen3.8-27b-uncensored-large.mcapy | Qwen3.8 27B-Uncensored-MCLarge | all MXFP6 | 22.9 GB | 18.66 GiB | 0.0074 | 97.9% |
qwen3.8-27b-uncensored-xxl.mcapy | Qwen3.8 27B-Uncensored-MCXXL | MXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms) | 28.2 GB | 23.58 GiB | 0.0057 | 98.2% |
You need one model file. These support files are shared by all five:
| File | What it is | On disk |
|---|---|---|
qwen3.8-27b-dflash2.mcapy | DFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster) | 1.3 GB |
qwen3.8-27b-uncensored-mtp.mcapy | the uncensored model's own MTP layer (abliterated with it) as a smaller, slower drafter | 0.28 GB |
qwen3.8-27b-vision-fp8.mcapy | the vision tower in FP8 (images; nearly exact; abliteration left it unchanged) | 0.57 GB |
qwen3.8-27b-vision-bf16.mcapy | the original BF16 vision tower | 1.0 GB |
model\ folder. The launcher lists each model by its name, and gives
each the MTP layer of its own source.
| Size | Reasoning | Math | Code | Chat | Wiki |
|---|---|---|---|---|---|
| Tiny | 0.0138 | 0.0156 | 0.0508 | 0.2094 | 0.0414 |
| Small | 0.0107 | 0.0139 | 0.0450 | 0.1899 | 0.0277 |
| Medium | 0.0059 | 0.0080 | 0.0282 | 0.0689 | 0.0196 |
| Large | 0.0009 | 0.0015 | 0.0092 | 0.0227 | 0.0028 |
| XXL | 0.0007 | 0.0014 | 0.0081 | 0.0162 | 0.0022 |
Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.
| Size | Coding, 1 conversation | Prose, 1 conversation | Coding, 8 at once (total) | Reading an 8K-token prompt |
|---|---|---|---|---|
| Tiny | 493 | 250 | 1,856 | 7,854 |
| Small | 488 | 236 | 1,828 | 7,697 |
| Medium | 445 | 219 | 1,669 | 7,365 |
| Large | 375 | 178 | 1,444 | 6,982 |
| XXL | 312 | 149 | 1,314 | 6,212 |
Source: orcarouter/Qwen3.8-27B-Uncensored, revision
404ea47, BF16. It is the reference for every score above.
Recipes: the same as for perkel/Qwen3.8-27B-MC.
Storage changes for the engine:
None of these change the model's function. XXL's 16-bit parts are stored as two FP8 E4M3 terms.
Apache 2.0, as the original Qwen3.8-27B, OrcaRouter's
Qwen3.8-27B-Uncensored and
Qwen3.8-27B-DFlash2. The changes are listed in NOTICE.
Converted by Perkel's Software Corner.
Qwen3.8-27B-Uncensored, OrcaRouter's abliterated BF16 build of Qwen's Qwen3.8-27B, converted to MegaCapybara's MX weight formats. It comes in five sizes for one NVIDIA RTX 5090, made the same way as perkel/Qwen3.8-27B-MC.
Safety alignment largely removed. The source model had its refusal behaviour removed by abliteration. It follows requests that the original Qwen3.8-27B would refuse, including harmful ones, and has no built-in guardrails.
It is meant for research, red-teaming, refusal and interpretability studies, and controlled experiments. Do not serve it to others without your own safety and moderation layers. You are responsible for how you use it and for what it generates, under the Apache 2.0 license and the laws that apply to you.
These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).
MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository's
list (Qwen3.8-27B-Uncensored-MC); what it needs beside it comes with it.http://127.0.0.1:8080/v1.Without the launcher, put the files in the package's model folder and start a preset: the
usage guide has the details.
| File | Launcher name | Weight formats | On disk | In VRAM | KL | Top-1 |
|---|---|---|---|---|---|---|
qwen3.8-27b-uncensored-tiny.mcapy | Qwen3.8 27B-Uncensored-MCTiny | all MXFP4 | 16.6 GB | 12.70 GiB | 0.0662 | 92.7% |
qwen3.8-27b-uncensored-small.mcapy | Qwen3.8 27B-Uncensored-MCSmall | MXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs | 17.6 GB | 13.67 GiB | 0.0574 | 93.7% |
qwen3.8-27b-uncensored-medium.mcapy | Qwen3.8 27B-Uncensored-MCMedium | MXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs | 19.5 GB | 15.39 GiB | 0.0261 | 95.6% |
qwen3.8-27b-uncensored-large.mcapy | Qwen3.8 27B-Uncensored-MCLarge | all MXFP6 | 22.9 GB | 18.66 GiB | 0.0074 | 97.9% |
qwen3.8-27b-uncensored-xxl.mcapy | Qwen3.8 27B-Uncensored-MCXXL | MXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms) | 28.2 GB | 23.58 GiB | 0.0057 | 98.2% |
You need one model file. These support files are shared by all five:
| File | What it is | On disk |
|---|---|---|
qwen3.8-27b-dflash2.mcapy | DFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster) | 1.3 GB |
qwen3.8-27b-uncensored-mtp.mcapy | the uncensored model's own MTP layer (abliterated with it) as a smaller, slower drafter | 0.28 GB |
qwen3.8-27b-vision-fp8.mcapy | the vision tower in FP8 (images; nearly exact; abliteration left it unchanged) | 0.57 GB |
qwen3.8-27b-vision-bf16.mcapy | the original BF16 vision tower | 1.0 GB |
model\ folder. The launcher lists each model by its name, and gives
each the MTP layer of its own source.
| Size | Reasoning | Math | Code | Chat | Wiki |
|---|---|---|---|---|---|
| Tiny | 0.0138 | 0.0156 | 0.0508 | 0.2094 | 0.0414 |
| Small | 0.0107 | 0.0139 | 0.0450 | 0.1899 | 0.0277 |
| Medium | 0.0059 | 0.0080 | 0.0282 | 0.0689 | 0.0196 |
| Large | 0.0009 | 0.0015 | 0.0092 | 0.0227 | 0.0028 |
| XXL | 0.0007 | 0.0014 | 0.0081 | 0.0162 | 0.0022 |
Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.
| Size | Coding, 1 conversation | Prose, 1 conversation | Coding, 8 at once (total) | Reading an 8K-token prompt |
|---|---|---|---|---|
| Tiny | 493 | 250 | 1,856 | 7,854 |
| Small | 488 | 236 | 1,828 | 7,697 |
| Medium | 445 | 219 | 1,669 | 7,365 |
| Large | 375 | 178 | 1,444 | 6,982 |
| XXL | 312 | 149 | 1,314 | 6,212 |
Source: orcarouter/Qwen3.8-27B-Uncensored, revision
404ea47, BF16. It is the reference for every score above.
Recipes: the same as for perkel/Qwen3.8-27B-MC.
Storage changes for the engine:
None of these change the model's function. XXL's 16-bit parts are stored as two FP8 E4M3 terms.
Apache 2.0, as the original Qwen3.8-27B, OrcaRouter's
Qwen3.8-27B-Uncensored and
Qwen3.8-27B-DFlash2. The changes are listed in NOTICE.
Converted by Perkel's Software Corner.