GGUF Model Quantization & Deployment

93 repos across 5 sub-areas

Quantized large language model variants packaged in GGUF format, optimized for inference on consumer hardware with different precision levels (FP4, Q4_0) and hardware backends (ROCm). The cluster primarily consists of model repositories using GGUF serialization with conversational AI capabilities, along with supporting tools for quantization (imatrix) and API compatibility layers. These repos represent practical implementations of model compression and inference optimization, serving as reference deployments for running large language models efficiently.

LLM Quantization and ROCm Optimization

38 repos

Model quantization formats and GPU-accelerated inference for large language models, with a focus on ROCm (AMD GPU) support and GGUF serialization. The cluster centers on optimized model weights, speculative decoding techniques, and Vulkan-based acceleration for efficient LLM deployment on AMD hardware. Central repositories feature quantized variants of models like Qwen across different precision levels (FP4, FP8) paired with ROCm compilation targets.

GGUF quantized LLM models

23 repos

Quantized versions of large language models in GGUF format, optimized for efficient inference and local deployment. This collection focuses on conversational AI models from providers like Qwen, Google Gemma, and others, with imatrix quantization techniques and endpoint compatibility for various inference frameworks. Repos here serve as model repositories and deployment targets for running capable LLMs with reduced memory footprint.

Text-to-Speech & GGUF Model Variants

14 repos

Quantized language models and text-to-speech systems optimized for efficient inference, primarily centered on GGUF-format model variants. The cluster features multiple quantization levels (Q4, Q8) and language-specific implementations, particularly for French and Spanish, alongside infrastructure for model endpoints and optimization. Repositories here focus on making neural speech synthesis and language models lightweight and deployable across different hardware constraints.

Cluster 639558

13 repos

Cluster 639556

5 repos

GGUF Model Quantization & Deployment — Shadowgraph