25 repos
Optimized inference implementations of large language models using MLX (Apple's machine learning framework) with AXQuant quantization techniques to reduce model size and memory requirements. The cluster centers on quantized versions of popular models like Qwen, Mistral, and Ministral, enabling efficient text generation and conversational AI on resource-constrained devices. Most repos are model artifact collections and deployment configurations rather than algorithmic frameworks, making this primarily a catalog of pre-optimized model checkpoints and inference recipes for Apple Silicon and edge inference scenarios.