dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx

Model

Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune

A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16169.383.1 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.67~1.0 GB
mlx_lm.awq 3-bit, 8-bit embedding3.512.43
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.510.90794 MB
the same rounding without fine-tuning (gptq-3bit-mlx)3.510.38794 MB
mlx_lm.dwq 3-bit fine-tuned from the round-to-nearest start, same budget as below3.513.55
this model: our rounding, then mlx_lm.dwq (256 sequences x 512 tokens of WikiText-2 train, batch 1)3.510.18794 MB

Same recipe on SmolLM2-1.7B-Instruct: 11.49 β†’ 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.

Fine-tuning note

Starting Apple's distillation fine-tune (mlx_lm.dwq) from this rounding instead of from round-to-nearest is what makes the difference: 10.18 against 13.55 at the same (small) training budget. Their default budget is about eight times larger and was not run here for memory reasons.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32–35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation

dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx

Model

Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 3-bit for MLX: grid fit + refinement + weighted refit, then a short DWQ fine-tune

A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-dwq-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16169.383.1 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.67~1.0 GB
mlx_lm.awq 3-bit, 8-bit embedding3.512.43
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.510.90794 MB
the same rounding without fine-tuning (gptq-3bit-mlx)3.510.38794 MB
mlx_lm.dwq 3-bit fine-tuned from the round-to-nearest start, same budget as below3.513.55
this model: our rounding, then mlx_lm.dwq (256 sequences x 512 tokens of WikiText-2 train, batch 1)3.510.18794 MB

Same recipe on SmolLM2-1.7B-Instruct: 11.49 β†’ 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.

Fine-tuning note

Starting Apple's distillation fine-tune (mlx_lm.dwq) from this rounding instead of from round-to-nearest is what makes the difference: 10.18 against 13.55 at the same (small) training budget. Their default budget is about eight times larger and was not run here for memory reasons.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32–35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation