GGUF Model Quantization & Serving

29 repos

Quantized language model distributions and inference optimization, primarily focused on the GGUF format for efficient local LLM deployment. The cluster centers on pre-quantized Gemma model variants (ranging from 2B to 31B parameters) in different bit-widths and formats, alongside tools and frameworks like Unsloth and LocalAI for running these models efficiently on consumer hardware. Developers here are working with model compression techniques, GGUF serialization, and practical inference infrastructure for bringing large language models to edge and local environments.

Rust · 1
Swift · 1
llm ·41,801
gguf ·41,445
unsloth ·40,725
skill ·35,128
localai ·35,128
mlx ·35,128
gemma ·11,106
gemma4 ·9,665
endpoints_compatible ·8,478
conversational ·7,384