38 repos
Model quantization formats and GPU-accelerated inference for large language models, with a focus on ROCm (AMD GPU) support and GGUF serialization. The cluster centers on optimized model weights, speculative decoding techniques, and Vulkan-based acceleration for efficient LLM deployment on AMD hardware. Central repositories feature quantized variants of models like Qwen across different precision levels (FP4, FP8) paired with ROCm compilation targets.