dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx

Model

SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX

0

3 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX

A 4-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct made with error-feedback rounding (GPTQ, Frantar et al. 2022) and a least-error grid per group, instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64, 8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelperplexitysize
original, fp168.939
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding)10.537922 MB
this model9.412970 MB

The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.

Calibration caveat

This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).

Method

128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

4-bit
conversational
gptq
llama
quantized
safetensors
text-generation

dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx

Model

SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX

0

3 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

SmolLM2-1.7B-Instruct, 4-bit GPTQ for MLX

A 4-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct made with error-feedback rounding (GPTQ, Frantar et al. 2022) and a least-error grid per group, instead of the round-to-nearest that mlx_lm convert -q uses. Group size 64, 8-bit tied embedding, float16. It loads and runs with mlx_lm like any other MLX model.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-4bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelperplexitysize
original, fp168.939
4-bit round to nearest (mlx_lm convert -q, group 64, 4-bit embedding)10.537922 MB
this model9.412970 MB

The 8-bit embedding is why this model is about 5% larger and 5% slower to decode than the round-to-nearest one; with a 4-bit embedding the perplexity is about 0.5 higher and size and speed match.

Calibration caveat

This model was calibrated on WikiText-2 train and is evaluated on WikiText-2 test, which is in-domain. Calibrated on mlx-lm's generic calibration text instead, the same method scores 9.945 on Qwen (still 0.7 better than round to nearest and 0.45 better than mlx_lm.gptq on the same text). Seed-to-seed spread of this pipeline: 0.02 perplexity (three seeds).

Method

128 calibration sequences of 512 tokens from WikiText-2 train; per linear layer H = XᵀX with 1% damping; columns rounded in order with the error fed back into the remaining columns through the Cholesky factor of H⁻¹ (block size 128); each layer calibrated on the outputs of the already-quantized layers before it; each group's grid range searched for least squared error. Then packed into MLX's affine format. Code, scripts and all measurements: github.com/dfed25/mlx-gptq.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

4-bit
conversational
gptq
llama
quantized
safetensors
text-generation