FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF

Model

Qwen3.8-27B MXFP4 GGUF

2

5 commits

1 linked in READMEs

updated Sep 16, 2026

See the code

README

Qwen3.8-27B MXFP4 GGUF

GGUF quantization of Qwen/Qwen3.8-27B for llama.cpp inference on NVIDIA Blackwell GPUs.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, Qwen3.8 is the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.

Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

Qwen3.8 Highlights

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
  • Flexible Thinking Control: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking.
  • Vision-Language Understanding: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.

Model Overview

PropertyValue
Base ModelQwen/Qwen3.8-27B
ArchitectureQwen3_5ForConditionalGeneration (vision + language)
Language Model Parameters27B
Hidden Dimension5120
Token Embedding248,320 (Padded)
Number of Layers64
Hidden Layout16 Γ— (3 Γ— (Gated DeltaNet β†’ FFN) β†’ 1 Γ— (Gated Attention β†’ FFN))
Gated DeltaNet Heads48 for V, 16 for QK (head dim 128)
Gated Attention Heads24 for Q, 4 for KV (head dim 256, RoPE dim 64)
Feed Forward Intermediate Dim17,408
Context Length262,144 natively (extensible to 1M)
MTPIncluded (next-token prediction head)
LicenseApache-2.0

Quantization Details

PropertyValue
Quantization FormatMXFP4 (OCP Microscaled FP4)
Bits Per Weight4.36 BPW
Original Model Size~52 GB (F16)
Quantized Size~14.2 GB
KV Cache (recommended)Q8_0
Quantized Withllama.cpp llama-quantize (MXFP4 ftype)

What is MXFP4?

MXFP4 uses the OCP (Open Compute Project) microscaled FP4 format with E8M0 power-of-two scaling per 32 values. The power-of-two scaling means dequantization is just bit-shifts β€” no floating-point multiply needed β€” making decode significantly faster than NVFP4 on consumer Blackwell GPUs.

MXFP4 has 50% less scale metadata per weight (0.25 bits overhead) compared to NVFP4 (0.5 bits overhead), further reducing memory bandwidth pressure during decode.

MXFP4 is the recommended format for consumer RTX 50-series GPUs due to its superior decode speed and smaller file size.

Files

FileSizeDescription
qwen3.8-27b-mxfp4.gguf~14.2 GBText model (MXFP4 quantized)
mmproj-qwen3.8-27b-f16.gguf~928 MBVision projector (F16)

Hardware Requirements

  • GPU: NVIDIA RTX 50-series (Blackwell) with 16+ GB VRAM
  • RAM: 16+ GB system RAM recommended
  • Storage: ~20 GB free disk space

Usage with llama.cpp

Text Only

llama-cli -m qwen3.8-27b-mxfp4.gguf -ngl 99 -p "Hello"

With Vision

llama-cli -m qwen3.8-27b-mxfp4.gguf   --mmproj mmproj-qwen3.8-27b-f16.gguf   -ngl 99   --chat-template chatml   --conversation
llama-server   -m qwen3.8-27b-mxfp4.gguf   --mmproj mmproj-qwen3.8-27b-f16.gguf   -ngl 99   --ctx-size 32768   -ctk q8_0 -ctv q8_0   -fa on   -b 512   --ubatch-size 128
SettingValueWhy
-ngl 99All layers on GPUFull GPU offload for speed
--ctx-size 3276832K contextSweet spot for speed/quality
-ctk q8_0 -ctv q8_0Q8 KV cacheNear-lossless, 2x VRAM savings vs F16
-fa onFlash attentionFaster attention kernels
-b 512Batch sizeOptimal for single-user
--ubatch-size 128Micro-batchBalances throughput and latency

Performance (RTX 5060 Ti 16GB)

MetricValue
Prompt Processing~106 t/s
Generation~25 t/s
VRAM Usage~15.9 GB

Context Length vs Speed

ContextGeneration Speed
8K~26 t/s
32K~25 t/s
64K~10 t/s

Why MXFP4 Over NVFP4?

PropertyMXFP4NVFP4
Generation Speed~25 t/s~10 t/s
Prompt Processing~106 t/s~30 t/s
File Size14.2 GB15 GB
Scale TypeE8M0 power-of-twoE4M3 FP8
Scale Overhead0.25 bits/weight0.5 bits/weight
Best ForConsumer GPUsDatacenter Blackwell

Citation

@misc{qwen38,
    title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
    url = {https://qwen.ai/blog?id=qwen3.8},
    author = {{Qwen Team}},
    month = {August},
    year = {2026}
}

License

Apache-2.0 (same as base model)

conversational
endpoints_compatible
gguf
image-text-to-text
mxfp4
qwen3
qwen3.8
vision

FreedomAISVR/Qwen3.8-27B-MXFP4-GGUF

Model

Qwen3.8-27B MXFP4 GGUF

2

5 commits

1 linked in READMEs

updated Sep 16, 2026

See the code

README

Qwen3.8-27B MXFP4 GGUF

GGUF quantization of Qwen/Qwen3.8-27B for llama.cpp inference on NVIDIA Blackwell GPUs.

Following the widespread community adoption of the Qwen3.5 and Qwen3.6 series, Qwen3.8 is the most capable generation in the Qwen open-model family to date. Built on the architectural foundation of Qwen3.5, Qwen3.8 delivers substantial gains across coding, professional work, research, and long-horizon agentic tasks.

Qwen3.8-27B brings these advances to a compact, deployment-friendly dense model: a native vision-language model that understands images and videos, with flexible thinking control, designed to carry complex, multi-step tasks through to completion with greater reliability.

Qwen3.8 Highlights

  • Core Capabilities: Comprehensive improvements across coding, professional work, research, and long-horizon agentic tasks.
  • Agent Execution: Stronger autonomous planning and better handling of environment feedback, leading to more reliable end-to-end task completion.
  • Downstream Compatibility: Broader support for popular harnesses and development tools, making it easier to integrate into your existing stack.
  • Flexible Thinking Control: Thinking mode is on by default and can be disabled per request; reasoning depth can be tuned with reasoning_effort, and reasoning context from historical messages is retained via preserve_thinking.
  • Vision-Language Understanding: Native support for image and video understanding, from STEM diagrams and documents to hour-scale videos.

Model Overview

PropertyValue
Base ModelQwen/Qwen3.8-27B
ArchitectureQwen3_5ForConditionalGeneration (vision + language)
Language Model Parameters27B
Hidden Dimension5120
Token Embedding248,320 (Padded)
Number of Layers64
Hidden Layout16 Γ— (3 Γ— (Gated DeltaNet β†’ FFN) β†’ 1 Γ— (Gated Attention β†’ FFN))
Gated DeltaNet Heads48 for V, 16 for QK (head dim 128)
Gated Attention Heads24 for Q, 4 for KV (head dim 256, RoPE dim 64)
Feed Forward Intermediate Dim17,408
Context Length262,144 natively (extensible to 1M)
MTPIncluded (next-token prediction head)
LicenseApache-2.0

Quantization Details

PropertyValue
Quantization FormatMXFP4 (OCP Microscaled FP4)
Bits Per Weight4.36 BPW
Original Model Size~52 GB (F16)
Quantized Size~14.2 GB
KV Cache (recommended)Q8_0
Quantized Withllama.cpp llama-quantize (MXFP4 ftype)

What is MXFP4?

MXFP4 uses the OCP (Open Compute Project) microscaled FP4 format with E8M0 power-of-two scaling per 32 values. The power-of-two scaling means dequantization is just bit-shifts β€” no floating-point multiply needed β€” making decode significantly faster than NVFP4 on consumer Blackwell GPUs.

MXFP4 has 50% less scale metadata per weight (0.25 bits overhead) compared to NVFP4 (0.5 bits overhead), further reducing memory bandwidth pressure during decode.

MXFP4 is the recommended format for consumer RTX 50-series GPUs due to its superior decode speed and smaller file size.

Files

FileSizeDescription
qwen3.8-27b-mxfp4.gguf~14.2 GBText model (MXFP4 quantized)
mmproj-qwen3.8-27b-f16.gguf~928 MBVision projector (F16)

Hardware Requirements

  • GPU: NVIDIA RTX 50-series (Blackwell) with 16+ GB VRAM
  • RAM: 16+ GB system RAM recommended
  • Storage: ~20 GB free disk space

Usage with llama.cpp

Text Only

llama-cli -m qwen3.8-27b-mxfp4.gguf -ngl 99 -p "Hello"

With Vision

llama-cli -m qwen3.8-27b-mxfp4.gguf   --mmproj mmproj-qwen3.8-27b-f16.gguf   -ngl 99   --chat-template chatml   --conversation
llama-server   -m qwen3.8-27b-mxfp4.gguf   --mmproj mmproj-qwen3.8-27b-f16.gguf   -ngl 99   --ctx-size 32768   -ctk q8_0 -ctv q8_0   -fa on   -b 512   --ubatch-size 128
SettingValueWhy
-ngl 99All layers on GPUFull GPU offload for speed
--ctx-size 3276832K contextSweet spot for speed/quality
-ctk q8_0 -ctv q8_0Q8 KV cacheNear-lossless, 2x VRAM savings vs F16
-fa onFlash attentionFaster attention kernels
-b 512Batch sizeOptimal for single-user
--ubatch-size 128Micro-batchBalances throughput and latency

Performance (RTX 5060 Ti 16GB)

MetricValue
Prompt Processing~106 t/s
Generation~25 t/s
VRAM Usage~15.9 GB

Context Length vs Speed

ContextGeneration Speed
8K~26 t/s
32K~25 t/s
64K~10 t/s

Why MXFP4 Over NVFP4?

PropertyMXFP4NVFP4
Generation Speed~25 t/s~10 t/s
Prompt Processing~106 t/s~30 t/s
File Size14.2 GB15 GB
Scale TypeE8M0 power-of-twoE4M3 FP8
Scale Overhead0.25 bits/weight0.5 bits/weight
Best ForConsumer GPUsDatacenter Blackwell

Citation

@misc{qwen38,
    title = {{Qwen3.8-Max}: A New Bar for Coding and Cowork},
    url = {https://qwen.ai/blog?id=qwen3.8},
    author = {{Qwen Team}},
    month = {August},
    year = {2026}
}

License

Apache-2.0 (same as base model)

conversational
endpoints_compatible
gguf
image-text-to-text
mxfp4
qwen3
qwen3.8
vision