GGUF quantized LLM models

23 repos

Quantized versions of large language models in GGUF format, optimized for efficient inference and local deployment. This collection focuses on conversational AI models from providers like Qwen, Google Gemma, and others, with imatrix quantization techniques and endpoint compatibility for various inference frameworks. Repos here serve as model repositories and deployment targets for running capable LLMs with reduced memory footprint.

conversational ·21,475
gguf ·21,475
endpoints_compatible ·21,164
imatrix ·18,738
image-text-to-text ·15,618
unsloth ·11,822
zh ·8,914
en ·8,914
uncensored ·8,771
vision ·8,206