Quantized LLM Model Weights

20 repos

Optimized and quantized versions of large language models, primarily focusing on 4-bit and 6-bit quantization formats for reduced memory footprint and faster inference. The cluster contains numerous variants of Qwen models (base, coding, and vision versions) packaged in AWQ and other quantization schemes, compatible with inference frameworks like SGLang and optimized for hardware like RDNA4 GPUs. These repositories represent the practical engineering of model compression and deployment artifacts rather than training or architecture research.

quantized ·497
4-bit ·497
conversational ·487
moe ·483
text-generation ·481
endpoints_compatible ·477
deepseek ·469
deepseek-v4 ·469
2-bit ·469
deepseek-v4-flash ·469