dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx

Model

Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX

0

4 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX

A 4-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct made with error-feedback rounding (GPTQ, Frantar et al. 2022) and a per-group grid chosen by alternating nearest-level assignment with a least-squares fit of the grid (see the repository), instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64, 8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelperplexitysize
original, fp169.379
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding)10.669~1.0 GB
previous version of this model (grid fit only)9.596951 MB
this model (full recipe --mse --lloyd --refine 3 --refit 2, 8-bit embedding)9.586951 MB
same method with the range-search grid only (--mse)9.658951 MB
mlx_lm.gptq, current main, 6-bit embedding, its default calibration text10.408 (seed 7: 10.152)903 MB
mlx_lm.gptq, current main, 4-bit embedding10.878
mlx_lm.awq (4-bit embedding)10.498853 MB

The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.

Code generation (HumanEval pass@1, 164 problems, greedy, 1 run; standard error about 3.5 points)

fp16 38.4% | this model 39.0% | previous version 32.3% | Qwen's official AWQ 4-bit 34.1% and GPTQ-Int4 27.4% (measured independently on an NVIDIA A10G through vLLM by M. Federico, who also reproduced the perplexities above to three decimals). The 3-bit models from this pipeline score about 14% and are text models only.

Three calibration seeds of the full recipe at 4 bits give 9.575–9.586 in memory; the refit's perplexity gain at 4 bits is at the pipeline's noise floor (0.02), its code gain is about 2 standard errors.

Calibration caveat

This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).

Method

128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

4-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation

dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx

Model

Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX

0

4 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX

A 4-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct made with error-feedback rounding (GPTQ, Frantar et al. 2022) and a per-group grid chosen by alternating nearest-level assignment with a least-squares fit of the grid (see the repository), instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64, 8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelperplexitysize
original, fp169.379
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding)10.669~1.0 GB
previous version of this model (grid fit only)9.596951 MB
this model (full recipe --mse --lloyd --refine 3 --refit 2, 8-bit embedding)9.586951 MB
same method with the range-search grid only (--mse)9.658951 MB
mlx_lm.gptq, current main, 6-bit embedding, its default calibration text10.408 (seed 7: 10.152)903 MB
mlx_lm.gptq, current main, 4-bit embedding10.878
mlx_lm.awq (4-bit embedding)10.498853 MB

The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.

Code generation (HumanEval pass@1, 164 problems, greedy, 1 run; standard error about 3.5 points)

fp16 38.4% | this model 39.0% | previous version 32.3% | Qwen's official AWQ 4-bit 34.1% and GPTQ-Int4 27.4% (measured independently on an NVIDIA A10G through vLLM by M. Federico, who also reproduced the perplexities above to three decimals). The 3-bit models from this pipeline score about 14% and are text models only.

Three calibration seeds of the full recipe at 4 bits give 9.575–9.586 in memory; the refit's perplexity gain at 4 bits is at the pipeline's noise floor (0.02), its code gain is about 2 standard errors.

Calibration caveat

This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).

Method

128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

4-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation