Qwen Model Quantization & Optimization

9 repos

Quantized and optimized variants of the Qwen language model family, focusing on reducing model size and memory footprint through techniques like INT4/INT8 weight quantization, activation quantization, and AWQ optimization. These repositories contain specialized model checkpoints and configurations designed for efficient inference with vLLM and similar serving frameworks, enabling deployment of large language models on resource-constrained hardware while maintaining reasonable performance.

vllm ·1,776
qwen3.6 ·1,657
qwen ·1,646
qwen3.5 ·1,646
chat-template ·1,646
jinja ·1,646
mlx ·1,646
llama.cpp ·1,646
lm-studio ·1,646
qwen3.8 ·1,646