9 repos
Quantized and optimized variants of the Qwen language model family, focusing on reducing model size and memory footprint through techniques like INT4/INT8 weight quantization, activation quantization, and AWQ optimization. These repositories contain specialized model checkpoints and configurations designed for efficient inference with vLLM and similar serving frameworks, enabling deployment of large language models on resource-constrained hardware while maintaining reasonable performance.