Qwen LLM Quantization & Deployment

10 repos

Optimized implementations and quantized variants of Alibaba's Qwen large language models, focusing on inference efficiency through techniques like INT4/INT8 weight and activation quantization. These repositories provide pre-quantized model checkpoints and integration with vLLM for deployment, alongside standardized chat templates for consistent model behavior across different serving frameworks.

vllm ·1,848
qwen3.6 ·1,716
qwen3.5 ·1,705
lm-studio ·1,705
chat-template ·1,705
llama.cpp ·1,705
jinja ·1,705
qwen ·1,705
qwen3.8 ·1,705
thinking ·1,705