36 repos
Pre-quantized and optimized large language models in GGUF format, designed for efficient inference on consumer hardware via llama.cpp and compatible endpoints. The cluster contains various open-weight model variants—including MiniCPM, Gemma, Granite, LLaVA, and others—spanning different sizes and capabilities, all optimized for conversational and text-generation tasks with minimal computational requirements. Repositories here serve as distribution points for ready-to-run model weights rather than training or architecture frameworks.