18 repos
Model quantization and optimization techniques for large language models, particularly focused on AMD ROCm hardware acceleration and GGUF format compatibility. The cluster centers on optimized model variants (Qwen, Qwopus, KAT-Coder, AEON) with multiple quantization schemes (FP4, FPX, MTP) targeting specific AMD architectures like Strix Halo. While most repos appear to be model artifacts or configuration repositories rather than frameworks, they reflect active work in making LLMs efficient for inference on AMD GPUs through aggressive quantization strategies.