23 repos
Quantized versions of large language models in GGUF format, optimized for efficient inference and local deployment. This collection focuses on conversational AI models from providers like Qwen, Google Gemma, and others, with imatrix quantization techniques and endpoint compatibility for various inference frameworks. Repos here serve as model repositories and deployment targets for running capable LLMs with reduced memory footprint.