Low-bit quantization of Qwen/Qwen3.8-27B
produced with GSQ (Gumbel-Softmax Quantization).
This checkpoint applies GSQ-based post-training quantization to the model weights, reducing precision while preserving the original model's reasoning, coding, multilingual, long-context, and agentic capabilities.
The transformer weights are quantized to 3-bit GSQ with group size 128. The embedding layer and LM head are quantized separately to 4-bit RTN with group size 64 to preserve output quality and embedding fidelity.
We evaluate the quantized checkpoint against the original
Qwen/Qwen3.8-27B. Both models were evaluated with xhigh thinking enabled.
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 100.00 | 100.00 |
| GPQA Diamond | 89.90 | 91.41 |
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 0.603M | 0.615M |
| GPQA Diamond | 3.721M | 3.705M |
Note: Due to the stochastic nature of these tasks, benchmark results can exhibit significant variance across runs. The reported scores correspond to a single evaluation run and should not be interpreted as definitive estimates of model performance.
The GSQ quantization calibration dataset was constructed to represent a broad range of LLM workloads, including reasoning, coding, scientific tasks, multilingual understanding, long-context processing, and agentic behaviour.
The calibration mixture consists of:
| Category | Percentage |
|---|---|
| Math | 13.5% |
| Code | 17.5% |
| Science | 20.0% |
| General | 12.5% |
| Multilingual | 12.5% |
| Long context | 14.0% |
| Agentic trajectories | 10.0% |
Serving this checkpoint requires a patched vLLM installation.
Requirements:
patch_vllm_qwen35_embedding.py
This patch enables vLLM support for quantized embedding weights.
Install vLLM:
pip install vllm==0.27.1
Apply the patch in the same Python environment:
python patch_vllm_qwen35_embedding.py
Then serve the model:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
Important: If vLLM is reinstalled or the environment is recreated, run the patch again before serving the checkpoint.
The full checkpoint size is approximately 11.83 GB when deployed with vision capabilities enabled.
The quantization calibration dataset used for this release did not include vision samples. Therefore, while the vision components are preserved in the checkpoint and can be loaded, they were not calibrated using multimodal calibration data.
For text-only deployment, the vision components are not required. The model can
be loaded with the --language-model-only option:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--language-model-only
This removes the vision-related components from the loaded model and reduces the checkpoint size to approximately 10.90 GB.
This release does not currently support speculative decoding. The MTP (Multi-Token Prediction) components have been removed from the published checkpoint and are not available for MTP-based inference.
Actual VRAM usage during serving will be higher than the raw checkpoint size and depends on:
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurti{\'c}, Eldar and Kleinegger, Max and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026},
url = {https://arxiv.org/abs/2604.18556}
}
Low-bit quantization of Qwen/Qwen3.8-27B
produced with GSQ (Gumbel-Softmax Quantization).
This checkpoint applies GSQ-based post-training quantization to the model weights, reducing precision while preserving the original model's reasoning, coding, multilingual, long-context, and agentic capabilities.
The transformer weights are quantized to 3-bit GSQ with group size 128. The embedding layer and LM head are quantized separately to 4-bit RTN with group size 64 to preserve output quality and embedding fidelity.
We evaluate the quantized checkpoint against the original
Qwen/Qwen3.8-27B. Both models were evaluated with xhigh thinking enabled.
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 100.00 | 100.00 |
| GPQA Diamond | 89.90 | 91.41 |
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 0.603M | 0.615M |
| GPQA Diamond | 3.721M | 3.705M |
Note: Due to the stochastic nature of these tasks, benchmark results can exhibit significant variance across runs. The reported scores correspond to a single evaluation run and should not be interpreted as definitive estimates of model performance.
The GSQ quantization calibration dataset was constructed to represent a broad range of LLM workloads, including reasoning, coding, scientific tasks, multilingual understanding, long-context processing, and agentic behaviour.
The calibration mixture consists of:
| Category | Percentage |
|---|---|
| Math | 13.5% |
| Code | 17.5% |
| Science | 20.0% |
| General | 12.5% |
| Multilingual | 12.5% |
| Long context | 14.0% |
| Agentic trajectories | 10.0% |
Serving this checkpoint requires a patched vLLM installation.
Requirements:
patch_vllm_qwen35_embedding.py
This patch enables vLLM support for quantized embedding weights.
Install vLLM:
pip install vllm==0.27.1
Apply the patch in the same Python environment:
python patch_vllm_qwen35_embedding.py
Then serve the model:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
Important: If vLLM is reinstalled or the environment is recreated, run the patch again before serving the checkpoint.
The full checkpoint size is approximately 11.83 GB when deployed with vision capabilities enabled.
The quantization calibration dataset used for this release did not include vision samples. Therefore, while the vision components are preserved in the checkpoint and can be loaded, they were not calibrated using multimodal calibration data.
For text-only deployment, the vision components are not required. The model can
be loaded with the --language-model-only option:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--language-model-only
This removes the vision-related components from the loaded model and reduces the checkpoint size to approximately 10.90 GB.
This release does not currently support speculative decoding. The MTP (Multi-Token Prediction) components have been removed from the published checkpoint and are not available for MTP-based inference.
Actual VRAM usage during serving will be higher than the raw checkpoint size and depends on:
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurti{\'c}, Eldar and Kleinegger, Max and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026},
url = {https://arxiv.org/abs/2604.18556}
}