Quantized LLM inference and deployment

9 repos

Optimized implementations and model variants for running large language models with 4-bit quantization and other compression techniques. The cluster centers on quantized versions of models like Qwen and Lance, alongside frameworks and tools for efficient text generation and conversational AI at reduced precision. Resources here focus on making LLMs practical for inference on resource-constrained hardware while maintaining model quality through careful quantization strategies.

4-bit ·3,586
safetensors ·3,586
qwen3_5_moe ·3,586
conversational ·3,585
text-generation ·3,398
edge-inference ·3,397
lora ·3,397
prerouter ·3,397
ssd-offload ·3,397
image-text-to-text ·187