20 repos
Optimized and quantized versions of large language models, primarily focusing on 4-bit and 6-bit quantization formats for reduced memory footprint and faster inference. The cluster contains numerous variants of Qwen models (base, coding, and vision versions) packaged in AWQ and other quantization schemes, compatible with inference frameworks like SGLang and optimized for hardware like RDNA4 GPUs. These repositories represent the practical engineering of model compression and deployment artifacts rather than training or architecture research.