10 repos
Optimized implementations and quantized variants of Alibaba's Qwen large language models, focusing on inference efficiency through techniques like INT4/INT8 weight and activation quantization. These repositories provide pre-quantized model checkpoints and integration with vLLM for deployment, alongside standardized chat templates for consistent model behavior across different serving frameworks.