dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq

Model

Qwen2.5-1.5B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

0

3 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

Status: verified on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05: WikiText-2 perplexity 9.6625, HumanEval pass@1 33.5%, identical to the MLX measurements below. On the same box and protocol Qwen's official AWQ 4-bit scores 10.161 / 34.1% and its GPTQ-Int4 10.398 / 27.4% (fp16: 9.372 / 37.2%). This checkpoint is in the AutoAWQ GEMM layout (4 bits, group size 64, asymmetric with integer zero points) so that vLLM's awq / awq_marlin kernels can load it. The weights come from the recipe in github.com/dfed25/mlx-gptq run with --zero-point int: error-feedback rounding (GPTQ), an alternating grid fit, coordinate-descent refinement and a weighted refit, with every group's offset constrained to a whole number of steps from the start, so that conversion to this format changes the weights by 0.05% (mean relative), against about 10% when a free-offset model is converted afterwards.

Measured on the MLX version of the same weights (MacBook Pro M4 Pro): WikiText-2 test perplexity 9.663 (fp16 9.379; the free-offset MLX model 9.586; Qwen's official AWQ 4-bit 10.16 and GPTQ-Int4 10.40 as measured independently on an A10G), HumanEval pass@1 33.5% (fp16 38.4%; standard error about 3.5 points).

vllm serve dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16

Calibration: 128 x 512 tokens of WikiText-2 train (in-domain for the perplexity above). Converter: mlx2hf.py --awq by M. Federico. Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).

4-bit
conversational
quantized
qwen2
safetensors
text-generation
vllm

dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq

Model

Qwen2.5-1.5B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

0

3 commits

1 linked in READMEs

updated Oct 5, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 4-bit, AWQ file format (integer zero points), for vLLM / int4 GPU kernels

Status: verified on an NVIDIA A10G with vLLM 0.29 (awq_marlin), 2026-10-05: WikiText-2 perplexity 9.6625, HumanEval pass@1 33.5%, identical to the MLX measurements below. On the same box and protocol Qwen's official AWQ 4-bit scores 10.161 / 34.1% and its GPTQ-Int4 10.398 / 27.4% (fp16: 9.372 / 37.2%). This checkpoint is in the AutoAWQ GEMM layout (4 bits, group size 64, asymmetric with integer zero points) so that vLLM's awq / awq_marlin kernels can load it. The weights come from the recipe in github.com/dfed25/mlx-gptq run with --zero-point int: error-feedback rounding (GPTQ), an alternating grid fit, coordinate-descent refinement and a weighted refit, with every group's offset constrained to a whole number of steps from the start, so that conversion to this format changes the weights by 0.05% (mean relative), against about 10% when a free-offset model is converted afterwards.

Measured on the MLX version of the same weights (MacBook Pro M4 Pro): WikiText-2 test perplexity 9.663 (fp16 9.379; the free-offset MLX model 9.586; Qwen's official AWQ 4-bit 10.16 and GPTQ-Int4 10.40 as measured independently on an A10G), HumanEval pass@1 33.5% (fp16 38.4%; standard error about 3.5 points).

vllm serve dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-int4-awq --quantization awq_marlin --dtype float16

Calibration: 128 x 512 tokens of WikiText-2 train (in-domain for the perplexity above). Converter: mlx2hf.py --awq by M. Federico. Author: Domenic Federico (Cal Poly San Luis Obispo; domfederico21@gmail.com), built with Claude (Anthropic) as a coding and research assistant. Licence follows the base model (Apache 2.0).

4-bit
conversational
quantized
qwen2
safetensors
text-generation
vllm