Quantized LLM Model Collections

23 repos

Pre-quantized and converted large language models in GGUF format, optimized for efficient inference on consumer hardware. These repositories primarily collect model weights and conversion artifacts for various open-source LLMs (Llama, Qwen, Gemma, MiniCPM) rather than containing source code, making them distribution points for practitioners looking to run LLMs locally via llama.cpp or Open WebUI. The cluster reflects the practical infrastructure around making state-of-the-art models accessible without GPU resources.

llama.cpp ·2,394
gguf ·2,394
endpoints_compatible ·1,935
conversational ·1,855
text-generation ·1,820
en ·1,128
gemma4 ·1,064
tool-use ·990
coding ·989
ko ·984