GGUF Model Quantization & ROCm

31 repos

Quantized large language model weights distributed in GGUF format, optimized for AMD ROCm accelerators. This cluster contains model variants (Qwen, MiMo, Tess) quantized to different precision levels (FP4, Q4_0) and compiled for ROCm-compatible hardware, enabling efficient local inference on AMD GPUs. The repositories represent pre-built model artifacts rather than framework code, with minimal language diversity reflecting their nature as model weight distributions.

C++ · 1
Python · 1
Shell · 1
gguf ·183
strix-halo ·157
endpoints_compatible ·148
rocmfpx ·139
rocm ·135
conversational ·123
rocmfp4 ·111
text-generation ·110
amd ·109
mtp ·81