dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq

Model

Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

0

2 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

A 4-bit copy of Qwen/Qwen2.5-7B-Instruct in the AutoAWQ GEMM layout (4 bits, group size 64, asymmetric with integer zero points), so that vLLM's awq / awq_marlin kernels load it directly.

Measured on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05, same box and protocol for every row:

modelWikiText-2 perplexityHumanEval pass@1
fp167.14570.1% (115/164)
this model7.28867.1% (110/164)
Qwen's official AWQ 4-bit7.58364.6% (106/164)

Read the numbers with these limits in mind: one run and one seed per row; the calibration text (128 x 512 tokens of WikiText-2 train) is in-domain for the perplexity test, which favours this model by an amount measured at about 0.3 on the 1.5B model; the HumanEval standard error at 164 problems is about 3.6 points, so the code-generation difference between the two 4-bit models is not significant. Perplexity protocol: WikiText-2 test, first 20 non-overlapping windows of 2048 tokens, scored through the serving kernel. HumanEval: greedy, completion-style prompts.

vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16

How it was made

The recipe from github.com/dfed25/mlx-gptq: error-feedback rounding (GPTQ), an alternating grid fit, coordinate-descent refinement and a weighted refit in the layer metric, with every group's offset constrained to a whole number of steps from the start (integer zero points). Because of that constraint the conversion to this format changes the weights by 0.05% (mean relative). This checkpoint was produced with a PyTorch/CUDA port of that code written by M. Federico (--mse --lloyd --refine 3 --refit 2 --int-zero), about 7 minutes per layer on one A10G. Embedding and output head are kept in fp16.

The smaller sibling is dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq.

Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), with M. Federico (PyTorch port, NVIDIA evaluation), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).

4-bit
conversational
quantized
qwen2
safetensors
text-generation
vllm

dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq

Model

Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

0

2 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen2.5-7B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

A 4-bit copy of Qwen/Qwen2.5-7B-Instruct in the AutoAWQ GEMM layout (4 bits, group size 64, asymmetric with integer zero points), so that vLLM's awq / awq_marlin kernels load it directly.

Measured on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05, same box and protocol for every row:

modelWikiText-2 perplexityHumanEval pass@1
fp167.14570.1% (115/164)
this model7.28867.1% (110/164)
Qwen's official AWQ 4-bit7.58364.6% (106/164)

Read the numbers with these limits in mind: one run and one seed per row; the calibration text (128 x 512 tokens of WikiText-2 train) is in-domain for the perplexity test, which favours this model by an amount measured at about 0.3 on the 1.5B model; the HumanEval standard error at 164 problems is about 3.6 points, so the code-generation difference between the two 4-bit models is not significant. Perplexity protocol: WikiText-2 test, first 20 non-overlapping windows of 2048 tokens, scored through the serving kernel. HumanEval: greedy, completion-style prompts.

vllm serve dfed24/Qwen2.5-7B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16

How it was made

The recipe from github.com/dfed25/mlx-gptq: error-feedback rounding (GPTQ), an alternating grid fit, coordinate-descent refinement and a weighted refit in the layer metric, with every group's offset constrained to a whole number of steps from the start (integer zero points). Because of that constraint the conversion to this format changes the weights by 0.05% (mean relative). This checkpoint was produced with a PyTorch/CUDA port of that code written by M. Federico (--mse --lloyd --refine 3 --refit 2 --int-zero), about 7 minutes per layer on one A10G. Embedding and output head are kept in fp16.

The smaller sibling is dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq.

Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), with M. Federico (PyTorch port, NVIDIA evaluation), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).

4-bit
conversational
quantized
qwen2
safetensors
text-generation
vllm