Quantized LLM Model Optimization

12 repos

Low-bit quantization (2-bit, 4-bit, 6-bit) techniques for compressing and optimizing large language and vision models, particularly for deployment on resource-constrained hardware like Apple Silicon via MLX. The cluster focuses on making state-of-the-art models (Qwen, DeepSeek, MiniMax, Muse) inference-efficient through aggressive quantization while maintaining usable accuracy, with AXQ emerging as a common quantization standard across implementations.

4-bit ·0
6-bit ·0
agent ·0
axq ·0
axquant ·0
conversational ·0
custom_code ·0
deepseek ·0
deepseekocr_2 ·0
2-bit ·0