2 repos
Optimized implementations and variants of large language models and multimodal models designed for efficient deployment, particularly on resource-constrained or specialized hardware. The cluster centers on quantized and distilled versions of popular LLM architectures (MiniCPM, Qwen, Gemma, Nanbeige) alongside platform-specific integrations—notably Apple's Core AI framework for on-device inference on iOS and macOS. These repositories reflect the practical engineering of making state-of-the-art models deployable in production environments with strict latency and memory constraints.