dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx

Model

SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

A 3-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct: group size 64, 8-bit tied embedding, float16, 778 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16168.943.4 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.54922 MB
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.511.49
this model3.510.20778 MB

Same recipe on Qwen2.5-1.5B-Instruct: 10.90 → 10.38 (fp16 9.38; Apple's AWQ 3-bit 12.43). Three-seed spread of the pipeline: 0.02.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: on Qwen2.5-1.5B an independent evaluation (NVIDIA A10G, vLLM 0.29, M. Federico, 2026-10-04) measured HumanEval pass@1: fp16 37.2%, 4-bit models 32–35%, 3-bit models about 14%; the same collapse should be expected here. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
llama
quantized
safetensors
text-generation

dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx

Model

SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

0

3 commits

1 linked in READMEs

updated Oct 4, 2026

See the code

README

SmolLM2-1.7B-Instruct, 3-bit for MLX (grid fit + refinement + weighted refit)

A 3-bit MLX quantization of HuggingFaceTB/SmolLM2-1.7B-Instruct: group size 64, 8-bit tied embedding, float16, 778 MB. Made with error-feedback rounding (GPTQ) plus three additions described in github.com/dfed25/mlx-gptq: an alternating grid fit per group, coordinate-descent refinement of the codes on the layer objective, and a least-squares refit of every group's grid in the same objective.

pip install mlx-lm
python -m mlx_lm generate --model dfed24/SmolLM2-1.7B-Instruct-gptq-3bit-mlx --prompt "Hello"

Quality (WikiText-2 test perplexity, 20 windows of 2048 tokens, lower is better; MacBook Pro M4 Pro, MLX 0.32)

modelbits/weightperplexitysize
original, fp16168.943.4 GB
4-bit round to nearest (mlx_lm convert -q, 4-bit embedding)4.510.54922 MB
GPTQ 3-bit with a range-search grid (this pipeline, earlier)3.511.49
this model3.510.20778 MB

Same recipe on Qwen2.5-1.5B-Instruct: 10.90 → 10.38 (fp16 9.38; Apple's AWQ 3-bit 12.43). Three-seed spread of the pipeline: 0.02.

Code generation caveat

Perplexity measures text prediction and hides a large loss on code: on Qwen2.5-1.5B an independent evaluation (NVIDIA A10G, vLLM 0.29, M. Federico, 2026-10-04) measured HumanEval pass@1: fp16 37.2%, 4-bit models 32–35%, 3-bit models about 14%; the same collapse should be expected here. Treat this model as a text model; for code generation use a 4-bit model. A code metric will be added to this pipeline's own evaluation.

Calibration caveat

Calibrated on WikiText-2 train and evaluated on WikiText-2 test (in-domain). The AWQ comparison above used mlx-lm's generic calibration text. Numbers on other text will differ.

Author

Domenic Federico, undergraduate at Cal Poly San Luis Obispo (on exchange at University College Cork, 2026), domfederico21@gmail.com. Built with Claude (Anthropic) as a coding and research assistant; the grid fit and the weighted refit were worked out by the author. Every number above is reproducible with the linked scripts. Licence and usage terms follow the base model (Apache 2.0).

3-bit
conversational
gptq
llama
quantized
safetensors
text-generation