28 repos
Model quantization techniques and GGUF file format implementations for efficiently running large language models on consumer hardware. This cluster covers tools and pre-quantized model variants (primarily Gemma models in various bit-widths like E2B, E4B, and 4-bit formats) that enable deployment of LLMs with reduced memory and computational requirements. Repositories focus on model compression, inference optimization, and the GGUF format ecosystem that underpins local AI inference.