Model Quantization and MLX Deployment

10 repos

Techniques and implementations for quantizing large language models to run efficiently on Apple Silicon using the MLX framework. The cluster focuses on reducing model size through various quantization schemes (4-bit, NVFP4, DWQ) applied to models like Lance-3B, enabling inference on resource-constrained devices. Most repositories appear to be model variants or configuration implementations rather than foundational libraries, but collectively demonstrate practical approaches to deploying conversational AI on Apple hardware with quantized weights distributed via safetensors format.

Python · 1
apple-silicon ·246
mlx ·246
omlx-alternative ·225
gguf ·225
omlx ·225
jang-quantization ·225
llamacpp ·225
mlxllm ·225
llm ·225
quantization ·225