Apple Silicon LLM Quantization

9 repos

Model quantization and optimization for large language models running on Apple Silicon hardware (MLX framework). These repositories contain pre-quantized versions of popular models like GPT-OSS and Gemma at 4-bit and 6-bit precision levels, enabling efficient inference on Apple devices. The cluster focuses on making state-of-the-art language models accessible and performant on consumer Mac hardware through aggressive quantization techniques.

4-bit ·0
4bit ·0
6-bit ·0
6bit ·0
apple-silicon ·0
axquant ·0
conversational ·0
development ·0
en ·0
gemma-4 ·0