23 repos
Pre-quantized and converted large language models in GGUF format, optimized for efficient inference on consumer hardware. These repositories primarily collect model weights and conversion artifacts for various open-source LLMs (Llama, Qwen, Gemma, MiniCPM) rather than containing source code, making them distribution points for practitioners looking to run LLMs locally via llama.cpp or Open WebUI. The cluster reflects the practical infrastructure around making state-of-the-art models accessible without GPU resources.