Quantized LLM Inference on Apple Silicon

32 repos

Optimized implementations of large language models running on Apple's MLX framework with aggressive quantization techniques (AXQ/mixed-precision formats) to fit models like Qwen on resource-constrained devices. These repositories focus on model compression, inference optimization, and making state-of-the-art LLMs practically deployable on Mac and iOS hardware through low-bit quantization schemes.

Python · 1
mlx ·18
apple-silicon ·18
mixed-precision ·18
safetensors ·18
development ·17
quantized ·17
axq ·17
axquant ·17
4-bit ·17
conversational ·17