29 repos
Quantized language model distributions and inference optimization, primarily focused on the GGUF format for efficient local LLM deployment. The cluster centers on pre-quantized Gemma model variants (ranging from 2B to 31B parameters) in different bit-widths and formats, alongside tools and frameworks like Unsloth and LocalAI for running these models efficiently on consumer hardware. Developers here are working with model compression techniques, GGUF serialization, and practical inference infrastructure for bringing large language models to edge and local environments.