SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX
0
3 commits
1 linked in READMEs
updated Sep 29, 2026
A 4-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct made with error-feedback rounding (GPTQ, Frantar et al.
2022) and a least-error grid per group, instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64,
8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx --prompt "Hello"
| model | perplexity | size |
|---|---|---|
| original, fp16 | 8.939 | |
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding) | 10.537 | 922 MB |
| this model | 9.412 | 970 MB |
The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.
This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on
mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest
and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).
128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).
SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX
0
3 commits
1 linked in READMEs
updated Sep 29, 2026
A 4-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct made with error-feedback rounding (GPTQ, Frantar et al.
2022) and a least-error grid per group, instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64,
8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.
pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx --prompt "Hello"
| model | perplexity | size |
|---|---|---|
| original, fp16 | 8.939 | |
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding) | 10.537 | 922 MB |
| this model | 9.412 | 970 MB |
The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.
This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on
mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest
and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).
128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.
Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).