9 repos
Optimized implementations and model variants for running large language models with 4-bit quantization and other compression techniques. The cluster centers on quantized versions of models like Qwen and Lance, alongside frameworks and tools for efficient text generation and conversational AI at reduced precision. Resources here focus on making LLMs practical for inference on resource-constrained hardware while maintaining model quality through careful quantization strategies.