32 repos
Optimized implementations of large language models running on Apple's MLX framework with aggressive quantization techniques (AXQ/mixed-precision formats) to fit models like Qwen on resource-constrained devices. These repositories focus on model compression, inference optimization, and making state-of-the-art LLMs practically deployable on Mac and iOS hardware through low-bit quantization schemes.