9 repos
Model quantization and optimization for large language models running on Apple Silicon hardware (MLX framework). These repositories contain pre-quantized versions of popular models like GPT-OSS and Gemma at 4-bit and 6-bit precision levels, enabling efficient inference on Apple devices. The cluster focuses on making state-of-the-art language models accessible and performant on consumer Mac hardware through aggressive quantization techniques.