39 repos
Quantized machine learning models and inference frameworks optimized for Apple Silicon using the MLX library and Metal acceleration. The cluster centers on efficient implementations of large language models and video generation systems in 4-bit, 8-bit, and bfloat16 precision, enabling on-device inference on Mac hardware. Repositories here demonstrate practical deployment strategies for models like MiniMax-H3 and LongCat video avatars using MLX's native Apple Silicon support.