6 repos
Optimized inference implementations of Qwen vision and speech models using Apple's MLX framework, with aggressive quantization (4-bit and 6-bit) for efficient deployment on Apple Silicon. The cluster focuses on making large multimodal and language models practical for edge inference through post-training quantization and hardware-specific optimization, with repositories representing different model variants (vision language models like Qwen3-VL and automatic speech recognition models like Qwen3-ASR) all compiled to run efficiently on Mac hardware.