19 repos
Quantized versions of large language models optimized for memory efficiency and inference speed, primarily using 4-bit quantization techniques with bitsandbytes and safetensors for model storage and loading. The cluster centers on pre-trained and instruction-tuned models across multiple languages and architectures—including Llama, Mistral, Qwen, and vision-capable variants—that enable deployment on resource-constrained hardware. Repositories here serve as model hubs and checkpoints ready for inference or fine-tuning via the Hugging Face transformers ecosystem.