Quantized Language Models and Inference

39 repos

Optimized implementations and quantized variants of large language models, primarily focused on the GGUF format for efficient inference and deployment. The cluster centers on model compression techniques (particularly 4-bit and 8-bit quantization), multilingual model variants, and infrastructure for serving conversational AI systems with reduced computational footprint. Repositories here span model distributions, quantization tools, and frameworks that make LLM inference practical on resource-constrained hardware.

imatrix ·14,243
gguf ·14,243
conversational ·14,155
endpoints_compatible ·13,935
unsloth ·12,544
image-text-to-text ·9,677
qwen3_5 ·6,573
transformers ·5,177
qwen ·4,768
qwen3_5_moe ·3,367