GGUF quantized language models

36 repos

Pre-quantized and optimized large language models in GGUF format, designed for efficient inference on consumer hardware via llama.cpp and compatible endpoints. The cluster contains various open-weight model variants—including MiniCPM, Gemma, Granite, LLaVA, and others—spanning different sizes and capabilities, all optimized for conversational and text-generation tasks with minimal computational requirements. Repositories here serve as distribution points for ready-to-run model weights rather than training or architecture frameworks.

gguf ·3,527
llama.cpp ·3,508
endpoints_compatible ·3,037
conversational ·2,960
text-generation ·2,186
en ·1,467
tool-use ·1,240
gemma4 ·1,073
coding ·1,036
ko ·1,015