dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-mlx

Model

Qwen2.5-1.5B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16169.383.1 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.67~1.0 GB
mlx_lm.awq 3-bit, 8-bit embedding3.512.43
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.510.90794 MB
this model3.510.38794 MB

Same recipe on SmolLM2-1.7B-Instruct: 11.49 → 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32–35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation

dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-mlx

Model

Qwen2.5-1.5B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

Qwen2.5-1.5B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

A 3-bit MLX quantization of Qwen/Qwen2.5-1.5B-Instruct: group size 64, 8-bit tied embedding, float16, 794 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/Qwen2.5-1.5B-Instruct-gptq-3bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16169.383.1 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.67~1.0 GB
mlx_lm.awq 3-bit, 8-bit embedding3.512.43
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.510.90794 MB
this model3.510.38794 MB

Same recipe on SmolLM2-1.7B-Instruct: 11.49 → 10.20 (fp16 8.94). Three-seed spread of the pipeline: 0.02. Decode speed in the same session: about 10% below the community 4-bit model (81 vs 90 tok/s under load), for 20% less memory.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: an independent evaluation on an NVIDIA A10G through vLLM 0.29 (M. Federico, 2026-10-04), which reproduced the perplexities above to three decimals, measured HumanEval pass@1 (164 problems, greedy): fp16 37.2%, our 4-bit models 32–35%, Qwen's official AWQ 4-bit 34.1%, Qwen's GPTQ-Int4 27.4%, this 3-bit model about 14%. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
quantized
qwen2
safetensors
text-generation