Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX
0
4 commits
1 linked in READMEs
updated Oct 4, 2026
A 4-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct made with error-feedback rounding (GPTQ, Frantar et al.
2022) and a per-group grid chosen by alternating nearest-level assignment with a least-squares fit of the grid (see the
repository), instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64,
8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx --prompt "Hello"
| model | perplexity | size |
|---|---|---|
| original, fp16 | 9.379 | |
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding) | 10.669 | ~1.0 GB |
| previous version of this model (grid fit only) | 9.596 | 951 MB |
this model (full recipe --mse --lloyd --refine 3 --refit 2, 8-bit embedding) | 9.586 | 951 MB |
same method with the range-search grid only (--mse) | 9.658 | 951 MB |
mlx_lm.gptq, current main, 6-bit embedding, its default calibration text | 10.408 (seed 7: 10.152) | 903 MB |
mlx_lm.gptq, current main, 4-bit embedding | 10.878 | |
mlx_lm.awq (4-bit embedding) | 10.498 | 853 MB |
The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.
fp16 38.4% | this model 39.0% | previous version 32.3% | Qwen's official AWQ 4-bit 34.1% and GPTQ-Int4 27.4% (measured independently on an NVIDIA A10G through vLLM by M. Federico, who also reproduced the perplexities above to three decimals). The 3-bit models from this pipeline score about 14% and are text models only.
Three calibration seeds of the full recipe at 4 bits give 9.575–9.586 in memory; the refit's perplexity gain at 4 bits is at the pipeline's noise floor (0.02), its code gain is about 2 standard errors.
This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on
mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest
and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).
128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).
Qwen2.5-1.5B-Instruct, 4-bit GPTQ for MLX
0
4 commits
1 linked in READMEs
updated Oct 4, 2026
A 4-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct made with error-feedback rounding (GPTQ, Frantar et al.
2022) and a per-group grid chosen by alternating nearest-level assignment with a least-squares fit of the grid (see the
repository), instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64,
8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-4bit-mlx --prompt "Hello"
| model | perplexity | size |
|---|---|---|
| original, fp16 | 9.379 | |
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding) | 10.669 | ~1.0 GB |
| previous version of this model (grid fit only) | 9.596 | 951 MB |
this model (full recipe --mse --lloyd --refine 3 --refit 2, 8-bit embedding) | 9.586 | 951 MB |
same method with the range-search grid only (--mse) | 9.658 | 951 MB |
mlx_lm.gptq, current main, 6-bit embedding, its default calibration text | 10.408 (seed 7: 10.152) | 903 MB |
mlx_lm.gptq, current main, 4-bit embedding | 10.878 | |
mlx_lm.awq (4-bit embedding) | 10.498 | 853 MB |
The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.
fp16 38.4% | this model 39.0% | previous version 32.3% | Qwen's official AWQ 4-bit 34.1% and GPTQ-Int4 27.4% (measured independently on an NVIDIA A10G through vLLM by M. Federico, who also reproduced the perplexities above to three decimals). The 3-bit models from this pipeline score about 14% and are text models only.
Three calibration seeds of the full recipe at 4 bits give 9.575–9.586 in memory; the refit's perplexity gain at 4 bits is at the pipeline's noise floor (0.02), its code gain is about 2 standard errors.
This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on
mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest
and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).
128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).