LLM Model Quantization & ROCm Optimization

18 repos

Model quantization and optimization techniques for large language models, particularly focused on AMD ROCm hardware acceleration and GGUF format compatibility. The cluster centers on optimized model variants (Qwen, Qwopus, KAT-Coder, AEON) with multiple quantization schemes (FP4, FPX, MTP) targeting specific AMD architectures like Strix Halo. While most repos appear to be model artifacts or configuration repositories rather than frameworks, they reflect active work in making LLMs efficient for inference on AMD GPUs through aggressive quantization strategies.

Python · 1
mtp ·379
gguf ·379
strix-halo ·365
rocm ·346
speculative-decoding ·313
vulkan ·298
qwen ·223
radv ·205
llama-cpp ·205
amd-ryzen-ai ·205