LLM Model Quantization & GGUF Formats

28 repos

Model quantization techniques and GGUF file format implementations for efficiently running large language models on consumer hardware. This cluster covers tools and pre-quantized model variants (primarily Gemma models in various bit-widths like E2B, E4B, and 4-bit formats) that enable deployment of LLMs with reduced memory and computational requirements. Repositories focus on model compression, inference optimization, and the GGUF format ecosystem that underpins local AI inference.

Rust · 1
gguf ·49,448
unsloth ·42,563
skill ·36,719
localai ·36,719
llm ·36,719
mlx ·36,719
endpoints_compatible ·14,908
transformers ·12,573
gemma ·11,413
text-generation ·9,478