39 repos
Optimized implementations and quantized variants of large language models, primarily focused on the GGUF format for efficient inference and deployment. The cluster centers on model compression techniques (particularly 4-bit and 8-bit quantization), multilingual model variants, and infrastructure for serving conversational AI systems with reduced computational footprint. Repositories here span model distributions, quantization tools, and frameworks that make LLM inference practical on resource-constrained hardware.