SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)
0
3 commits
1 linked in READMEs
updated Oct 4, 2026
A 3-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct: group size 64, 8-bit tied embedding, float16, 778 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx --prompt "Hello"
| model | bits/weight | perplexity | size |
|---|---|---|---|
| original, fp16 | 16 | 8.94 | 3.4 GB |
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding) | 4.5 | 10.54 | 922 MB |
| GPTQ 3-bit with a range-search grid (this pipeline, earlier) | 3.5 | 11.49 | |
| this model | 3.5 | 10.20 | 778 MB |
Same recipe on Qwen2.5-1.5B-Instruct: 10.90 → 10.38 (fp16 9.38; Apple's AWQ 3-bit 12.43). Three-seed spread of the pipeline: 0.02.
Perplexity measures text prediction and hides a large loss on code: on Qwen2.5-1.5B an independent evaluation (NVIDIA A10G, vLLM 0.29, M. Federico, 2026-10-04) measured HumanEval pass@1: fp16 37.2%, 4-bit models 32–35%, 3-bit models about 14%; the same collapse should be expected here. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.
Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).
SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)
0
3 commits
1 linked in READMEs
updated Oct 4, 2026
A 3-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct: group size 64, 8-bit tied embedding, float16, 778 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx --prompt "Hello"
| model | bits/weight | perplexity | size |
|---|---|---|---|
| original, fp16 | 16 | 8.94 | 3.4 GB |
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding) | 4.5 | 10.54 | 922 MB |
| GPTQ 3-bit with a range-search grid (this pipeline, earlier) | 3.5 | 11.49 | |
| this model | 3.5 | 10.20 | 778 MB |
Same recipe on Qwen2.5-1.5B-Instruct: 10.90 → 10.38 (fp16 9.38; Apple's AWQ 3-bit 12.43). Three-seed spread of the pipeline: 0.02.
Perplexity measures text prediction and hides a large loss on code: on Qwen2.5-1.5B an independent evaluation (NVIDIA A10G, vLLM 0.29, M. Federico, 2026-10-04) measured HumanEval pass@1: fp16 37.2%, 4-bit models 32–35%, 3-bit models about 14%; the same collapse should be expected here. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.
Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).