LLM Quantization and ROCm Optimization

38 repos

Model quantization formats and GPU-accelerated inference for large language models, with a focus on ROCm (AMD GPU) support and GGUF serialization. The cluster centers on optimized model weights, speculative decoding techniques, and Vulkan-based acceleration for efficient LLM deployment on AMD hardware. Central repositories feature quantized variants of models like Qwen across different precision levels (FP4, FP8) paired with ROCm compilation targets.

Python · 2
C++ · 1
Shell · 1
gguf ·505
strix-halo ·480
rocm ·437
vulkan ·379
speculative-decoding ·360
qwen ·268
llama-cpp ·255
mtp ·250
fp4 ·225
radv ·225