Quantized LLM Model Checkpoints

19 repos

Quantized versions of large language models optimized for memory efficiency and inference speed, primarily using 4-bit quantization techniques with bitsandbytes and safetensors for model storage and loading. The cluster centers on pre-trained and instruction-tuned models across multiple languages and architectures—including Llama, Mistral, Qwen, and vision-capable variants—that enable deployment on resource-constrained hardware. Repositories here serve as model hubs and checkpoints ready for inference or fine-tuning via the Hugging Face transformers ecosystem.

4-bit ·263
bitsandbytes ·263
safetensors ·214
transformers ·192
image-text-to-text ·178
conversational ·169
custom_code ·144
minicpmv ·139
minicpm-v ·100
endpoints_compatible ·99