Model Compression for Edge AI

4 repos

Quantized language model variants optimized for Apple Silicon and resource-constrained devices. This cluster contains multiple compressed versions of open-source language models (GPT, DeepSeek, MiniCPM, Devstral) post-training quantized to 2-bit, 4-bit, and 6-bit precision using the MLX framework, enabling efficient inference on edge hardware. Repos focus on reducing model size and computational requirements while maintaining usable performance for on-device deployment.

2-bit ·0
apple-silicon ·0
axq ·0
axquant ·0
conversational ·0
deepseek ·0
deepseek-v4 ·0
deepseek_v4 ·0
development ·0
experimental ·0