93 repos across 5 sub-areas
Quantized large language model variants packaged in GGUF format, optimized for inference on consumer hardware with different precision levels (FP4, Q4_0) and hardware backends (ROCm). The cluster primarily consists of model repositories using GGUF serialization with conversational AI capabilities, along with supporting tools for quantization (imatrix) and API compatibility layers. These repos represent practical implementations of model compression and inference optimization, serving as reference deployments for running large language models efficiently.
LLM Quantization and ROCm Optimization
38 repos
Model quantization formats and GPU-accelerated inference for large language models, with a focus on ROCm (AMD GPU) support and GGUF serialization. The cluster centers on optimized model weights, speculative decoding techniques, and Vulkan-based acceleration for efficient LLM deployment on AMD hardware. Central repositories feature quantized variants of models like Qwen across different precision levels (FP4, FP8) paired with ROCm compilation targets.
GGUF quantized LLM models
23 repos
Quantized versions of large language models in GGUF format, optimized for efficient inference and local deployment. This collection focuses on conversational AI models from providers like Qwen, Google Gemma, and others, with imatrix quantization techniques and endpoint compatibility for various inference frameworks. Repos here serve as model repositories and deployment targets for running capable LLMs with reduced memory footprint.
Text-to-Speech & GGUF Model Variants
14 repos
Quantized language models and text-to-speech systems optimized for efficient inference, primarily centered on GGUF-format model variants. The cluster features multiple quantization levels (Q4, Q8) and language-specific implementations, particularly for French and Spanish, alongside infrastructure for model endpoints and optimization. Repositories here focus on making neural speech synthesis and language models lightweight and deployable across different hardware constraints.
Cluster 639558
13 repos
Cluster 639556
5 repos