perkel/Qwen3.8-27B-MC

Model

Qwen3.8 27B-MC

0

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.8 27B-MC

Qwen3.8-27B (Qwen's BF16 release) converted to MegaCapybara's MX weight formats, in five sizes for one NVIDIA RTX 5090. The formats are MXFP4 and MXFP6, which Blackwell tensor cores read at full rate.

Run it with MegaCapybara

These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).

  1. Download MegaCapybara from its releases and unpack it.
  2. Run MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository; the drafter and the vision tower come with it.
  3. Pick a preset, press Load server, and point your client at http://127.0.0.1:8080/v1.

Without the launcher, put the files in the package's model folder and start a preset: the usage guide has the details.

An uncensored version, converted the same way from an abliterated release, is perkel/Qwen3.8-27B-Uncensored-MC.

Files

FileLauncher nameWeight formatsOn diskIn VRAMKLTop-1
qwen3.8-27b-tiny.mcapyQwen3.8 27B-MCTinyall MXFP416.6 GB12.70 GiB0.058292.7%
qwen3.8-27b-small.mcapyQwen3.8 27B-MCSmallMXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs17.6 GB13.67 GiB0.047893.8%
qwen3.8-27b-medium.mcapyQwen3.8 27B-MCMediumMXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs19.5 GB15.39 GiB0.029895.2%
qwen3.8-27b-large.mcapyQwen3.8 27B-MCLargeall MXFP622.9 GB18.66 GiB0.008197.7%
qwen3.8-27b-xxl.mcapyQwen3.8 27B-MCXXLMXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms)28.2 GB23.58 GiB0.005198.3%

You need one model file. These support files are shared by all five:

FileWhat it isOn disk
qwen3.8-27b-dflash2.mcapyDFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster)1.3 GB
qwen3.8-27b-mtp.mcapyQwen's own MTP layer as a smaller, slower drafter0.28 GB
qwen3.8-27b-vision-fp8.mcapythe vision tower in FP8 (images; nearly exact)0.57 GB
qwen3.8-27b-vision-bf16.mcapythe original BF16 vision tower1.0 GB

The launcher lists every model file in its model\ folder by its name, smallest first, and projects each one's accuracy, speed and memory for the chosen settings from scores stored in the file.

Accuracy against the original

Top-1 by size: our models and Unsloth's GGUFs

  • What KL and top-1 measure: KL divergence and top-1 agreement (how often the most likely next token is the same) against Qwen's BF16 weights computed with FP32 math.
  • The test text: 81,880 held-out tokens: reasoning traces, math, code, chat and wiki. It is disjoint from the calibration text. Tool calls are not scored: they differ even between BF16 and FP32 math.
  • Settings: the engine at 16-bit activations and an unquantized KV cache. FP8 activations, the engine's default, add a little KL; the launcher shows each setting's cost.
  • The chart:
    • Unsloth's points are their published top-1 for their Qwen3.8-27B GGUFs, measured on their own text.
    • Sizes are the bytes each puts on the GPU, read from the GGUF headers. The token embedding table and the MTP layer are not counted, because llama.cpp keeps the table in RAM as MegaCapybara does.
    • As a cross-check, their UD-Q4_K_XL scores 96.2% on our test, against the 96.4% they publish.
KL per kind of text
SizeReasoningMathCodeChatWiki
Tiny0.01380.01560.04710.17500.0395
Small0.01070.01380.04240.14560.0262
Medium0.00580.00780.02900.08740.0190
Large0.00090.00100.00750.02830.0026
XXL0.00070.00080.00690.01510.0021

Speed on an RTX 5090

Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.

SizeCoding, 1 conversationProse, 1 conversationCoding, 8 at once (total)Reading an 8K-token prompt
Tiny5402481,8907,789
Small4952321,7547,619
Medium4402081,6387,377
Large3771761,4276,955
XXL3131481,2546,178

Measured with nothing else on the card. Its memory was overclocked to 16.8 GHz (+20%); decoding is bound by memory bandwidth, so a stock card is slower.

Which size

  • Tiny and Small: the fastest, and the most room for context.
  • Medium: the balance; the launcher's default.
  • Large and XXL: the closest to the original, with less room for context.

How the files were made

  • Source: Qwen/Qwen3.8-27B, revision 1d4bf0f. Its BF16 weights are the reference for every score above.

  • Rounding: layer by layer, GPTQ with MX block scales (32 weights share a power-of-two scale), on 514K calibration tokens. The calibration text is 2K and 16K-token sequences of reasoning, math, code, chat and wiki.

  • Storage changes for the engine:

    • norm gains are folded into the following matrices;
    • the attention value rows are Hadamard-rotated;
    • the output head's rows are ordered by token frequency.

    None of these change the model's function.

  • XXL's 16-bit parts are stored as two FP8 E4M3 terms (value and residual), about 8 significant bits.

License

Apache 2.0, as the original Qwen3.8-27B and Qwen3.8-27B-DFlash2. The changes are listed in NOTICE. Converted by Perkel's Software Corner.

blackwell
dflash2
image-text-to-text
megacapybara
mxfp4
mxfp6
quantized
qwen3.8
rtx-5090
speculative-decoding

perkel/Qwen3.8-27B-MC

Model

Qwen3.8 27B-MC

0

5 commits

1 linked in READMEs

updated Oct 3, 2026

See the code

README

Qwen3.8 27B-MC

Qwen3.8-27B (Qwen's BF16 release) converted to MegaCapybara's MX weight formats, in five sizes for one NVIDIA RTX 5090. The formats are MXFP4 and MXFP6, which Blackwell tensor cores read at full rate.

Run it with MegaCapybara

These files run in MegaCapybara, the fastest inference engine for Qwen3.8-27B on the NVIDIA RTX 5090, for Windows 11 and Linux: up to 500 tokens/s for one coding agent and up to 2,000 tokens/s for a team of them, with OpenAI's and Anthropic's APIs (Claude Code runs on it).

  1. Download MegaCapybara from its releases and unpack it.
  2. Run MegaCapybaraLauncher and press the arrow next to the model to download a size from this repository; the drafter and the vision tower come with it.
  3. Pick a preset, press Load server, and point your client at http://127.0.0.1:8080/v1.

Without the launcher, put the files in the package's model folder and start a preset: the usage guide has the details.

An uncensored version, converted the same way from an abliterated release, is perkel/Qwen3.8-27B-Uncensored-MC.

Files

FileLauncher nameWeight formatsOn diskIn VRAMKLTop-1
qwen3.8-27b-tiny.mcapyQwen3.8 27B-MCTinyall MXFP416.6 GB12.70 GiB0.058292.7%
qwen3.8-27b-small.mcapyQwen3.8 27B-MCSmallMXFP4; MXFP6 for the head, the mixer outputs of 24 layers and the last 8 layers' MLPs17.6 GB13.67 GiB0.047893.8%
qwen3.8-27b-medium.mcapyQwen3.8 27B-MCMediumMXFP4; MXFP6 for the head, every mixer output, half the mixer inputs and MLP outputs19.5 GB15.39 GiB0.029895.2%
qwen3.8-27b-large.mcapyQwen3.8 27B-MCLargeall MXFP622.9 GB18.66 GiB0.008197.7%
qwen3.8-27b-xxl.mcapyQwen3.8 27B-MCXXLMXFP6; the head and layers 24-39's MLP inputs at ~16 bits (two FP8 terms)28.2 GB23.58 GiB0.005198.3%

You need one model file. These support files are shared by all five:

FileWhat it isOn disk
qwen3.8-27b-dflash2.mcapyDFlash2 speculative drafter, converted from incoai/Qwen3.8-27B-DFlash2 (the default: the same answers, faster)1.3 GB
qwen3.8-27b-mtp.mcapyQwen's own MTP layer as a smaller, slower drafter0.28 GB
qwen3.8-27b-vision-fp8.mcapythe vision tower in FP8 (images; nearly exact)0.57 GB
qwen3.8-27b-vision-bf16.mcapythe original BF16 vision tower1.0 GB

The launcher lists every model file in its model\ folder by its name, smallest first, and projects each one's accuracy, speed and memory for the chosen settings from scores stored in the file.

Accuracy against the original

Top-1 by size: our models and Unsloth's GGUFs

  • What KL and top-1 measure: KL divergence and top-1 agreement (how often the most likely next token is the same) against Qwen's BF16 weights computed with FP32 math.
  • The test text: 81,880 held-out tokens: reasoning traces, math, code, chat and wiki. It is disjoint from the calibration text. Tool calls are not scored: they differ even between BF16 and FP32 math.
  • Settings: the engine at 16-bit activations and an unquantized KV cache. FP8 activations, the engine's default, add a little KL; the launcher shows each setting's cost.
  • The chart:
    • Unsloth's points are their published top-1 for their Qwen3.8-27B GGUFs, measured on their own text.
    • Sizes are the bytes each puts on the GPU, read from the GGUF headers. The token embedding table and the MTP layer are not counted, because llama.cpp keeps the table in RAM as MegaCapybara does.
    • As a cross-check, their UD-Q4_K_XL scores 96.2% on our test, against the 96.4% they publish.
KL per kind of text
SizeReasoningMathCodeChatWiki
Tiny0.01380.01560.04710.17500.0395
Small0.01070.01380.04240.14560.0262
Medium0.00580.00780.02900.08740.0190
Large0.00090.00100.00750.02830.0026
XXL0.00070.00080.00690.01510.0021

Speed on an RTX 5090

Tokens per second; greedy, thinking off, answers up to 2,000 tokens, FP8 activations, DFlash2.

SizeCoding, 1 conversationProse, 1 conversationCoding, 8 at once (total)Reading an 8K-token prompt
Tiny5402481,8907,789
Small4952321,7547,619
Medium4402081,6387,377
Large3771761,4276,955
XXL3131481,2546,178

Measured with nothing else on the card. Its memory was overclocked to 16.8 GHz (+20%); decoding is bound by memory bandwidth, so a stock card is slower.

Which size

  • Tiny and Small: the fastest, and the most room for context.
  • Medium: the balance; the launcher's default.
  • Large and XXL: the closest to the original, with less room for context.

How the files were made

  • Source: Qwen/Qwen3.8-27B, revision 1d4bf0f. Its BF16 weights are the reference for every score above.

  • Rounding: layer by layer, GPTQ with MX block scales (32 weights share a power-of-two scale), on 514K calibration tokens. The calibration text is 2K and 16K-token sequences of reasoning, math, code, chat and wiki.

  • Storage changes for the engine:

    • norm gains are folded into the following matrices;
    • the attention value rows are Hadamard-rotated;
    • the output head's rows are ordered by token frequency.

    None of these change the model's function.

  • XXL's 16-bit parts are stored as two FP8 E4M3 terms (value and residual), about 8 significant bits.

License

Apache 2.0, as the original Qwen3.8-27B and Qwen3.8-27B-DFlash2. The changes are listed in NOTICE. Converted by Perkel's Software Corner.

blackwell
dflash2
image-text-to-text
megacapybara
mxfp4
mxfp6
quantized
qwen3.8
rtx-5090
speculative-decoding