4 repos
Quantized language model variants optimized for Apple Silicon and resource-constrained devices. This cluster contains multiple compressed versions of open-source language models (GPT, DeepSeek, MiniCPM, Devstral) post-training quantized to 2-bit, 4-bit, and 6-bit precision using the MLX framework, enabling efficient inference on edge hardware. Repos focus on reducing model size and computational requirements while maintaining usable performance for on-device deployment.