Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune
0
3 commits
1 linked in READMEs
updated Oct 4, 2026
A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx --prompt "Hello"
| model | bits/weight | perplexity | size |
|---|---|---|---|
| original, fp16 | 16 | 9.38 | 3.1 GB |
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding) | 4.5 | 10.67 | ~1.0 GB |
mlx_lm.awq 3-bit, 8-bit embedding | 3.5 | 12.43 | |
| GPTQ 3-bit with a range-search grid (this pipeline, earlier) | 3.5 | 10.90 | 794 MB |
| the same rounding without fine-tuning (gptq-3bit-mlx) | 3.5 | 10.38 | 794 MB |
mlx_lm.dwq 3-bit fine-tuned from the round-to-nearest start, same budget as below | 3.5 | 13.55 | |
this model: our rounding, then mlx_lm.dwq (256 sequences x 512 tokens of WikiText-2 train, batch 1) | 3.5 | 10.18 | 794 MB |
Same recipe on SmolLM2-1.7B-Instruct: 11.49 β 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.
Starting Apple's distillation fine-tune (mlx_lm.dwq) from this rounding instead of from round-to-nearest is what makes
the difference: 10.18 against 13.55 at the same (small) training budget. Their default budget is about eight times larger
and was not run here for memory reasons.
Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32β35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.
Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).
Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune
0
3 commits
1 linked in READMEs
updated Oct 4, 2026
A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx --prompt "Hello"
| model | bits/weight | perplexity | size |
|---|---|---|---|
| original, fp16 | 16 | 9.38 | 3.1 GB |
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding) | 4.5 | 10.67 | ~1.0 GB |
mlx_lm.awq 3-bit, 8-bit embedding | 3.5 | 12.43 | |
| GPTQ 3-bit with a range-search grid (this pipeline, earlier) | 3.5 | 10.90 | 794 MB |
| the same rounding without fine-tuning (gptq-3bit-mlx) | 3.5 | 10.38 | 794 MB |
mlx_lm.dwq 3-bit fine-tuned from the round-to-nearest start, same budget as below | 3.5 | 13.55 | |
this model: our rounding, then mlx_lm.dwq (256 sequences x 512 tokens of WikiText-2 train, batch 1) | 3.5 | 10.18 | 794 MB |
Same recipe on SmolLM2-1.7B-Instruct: 11.49 β 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.
Starting Apple's distillation fine-tune (mlx_lm.dwq) from this rounding instead of from round-to-nearest is what makes
the difference: 10.18 against 13.55 at the same (small) training budget. Their default budget is about eight times larger
and was not run here for memory reasons.
Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32β35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.
Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).