Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels
0
2 commits
1 linked in READMEs
updated Oct 5, 2026
A 4-bit copy of Qwen/Qwen2.5-7B-Instruct in the AutoAWQ GEMM layout
(4 bits, group size 64, asymmetric with integer zero points), so that vLLM's awq / awq_marlin kernels load it directly.
Measured on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05, same box and protocol for every row:
| model | WikiText-2 perplexity | HumanEval pass@1 |
|---|---|---|
| fp16 | 7.145 | 70.1% (115/164) |
| this model | 7.288 | 67.1% (110/164) |
| Qwen's official AWQ 4-bit | 7.583 | 64.6% (106/164) |
Read the numbers with these limits in mind: one run and one seed per row; the calibration text (128 x 512 tokens of WikiText-2 train) is in-domain for the perplexity test, which favours this model by an amount measured at about 0.3 on the 1.5B model; the HumanEval standard error at 164 problems is about 3.6 points, so the code-generation difference between the two 4-bit models is not significant. Perplexity protocol: WikiText-2 test, first 20 non-overlapping windows of 2048 tokens, scored through the serving kernel. HumanEval: greedy, completion-style prompts.
vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16
The recipe from github.com/dfed25/mlx-gptq: error-feedback rounding (GPTQ), an
alternating grid fit, coordinate-descent refinement and a weighted refit in the layer metric, with every group's offset
constrained to a whole number of steps from the start (integer zero points). Because of that constraint the conversion
to this format changes the weights by 0.05% (mean relative). This checkpoint was produced with a PyTorch/CUDA port of
that code written by M. Federico (--mse --lloyd --refine 3 --refit 2 --int-zero), about 7 minutes per layer on one
A10G. Embedding and output head are kept in fp16.
The smaller sibling is dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq.
Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), with M. Federico (PyTorch port, NVIDIA evaluation), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).
Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels
0
2 commits
1 linked in READMEs
updated Oct 5, 2026
A 4-bit copy of Qwen/Qwen2.5-7B-Instruct in the AutoAWQ GEMM layout
(4 bits, group size 64, asymmetric with integer zero points), so that vLLM's awq / awq_marlin kernels load it directly.
Measured on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05, same box and protocol for every row:
| model | WikiText-2 perplexity | HumanEval pass@1 |
|---|---|---|
| fp16 | 7.145 | 70.1% (115/164) |
| this model | 7.288 | 67.1% (110/164) |
| Qwen's official AWQ 4-bit | 7.583 | 64.6% (106/164) |
Read the numbers with these limits in mind: one run and one seed per row; the calibration text (128 x 512 tokens of WikiText-2 train) is in-domain for the perplexity test, which favours this model by an amount measured at about 0.3 on the 1.5B model; the HumanEval standard error at 164 problems is about 3.6 points, so the code-generation difference between the two 4-bit models is not significant. Perplexity protocol: WikiText-2 test, first 20 non-overlapping windows of 2048 tokens, scored through the serving kernel. HumanEval: greedy, completion-style prompts.
vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16
The recipe from github.com/dfed25/mlx-gptq: error-feedback rounding (GPTQ), an
alternating grid fit, coordinate-descent refinement and a weighted refit in the layer metric, with every group's offset
constrained to a whole number of steps from the start (integer zero points). Because of that constraint the conversion
to this format changes the weights by 0.05% (mean relative). This checkpoint was produced with a PyTorch/CUDA port of
that code written by M. Federico (--mse --lloyd --refine 3 --refit 2 --int-zero), about 7 minutes per layer on one
A10G. Embedding and output head are kept in fp16.
The smaller sibling is dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq.
Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), with M. Federico (PyTorch port, NVIDIA evaluation), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).