Vision-Language Model Checkpoints

17 repos

Pre-trained model weights and implementations for vision-language transformers that perform zero-shot image classification and multimodal understanding. The cluster centers on CLIP variants (Chinese CLIP and SigLIP2 architectures) optimized for different patch sizes and resolutions, with Hugging Face transformers integration and safetensors format support for easy deployment and inference.

Python · 1
transformers ·1,515
endpoints_compatible ·1,515
zero-shot-image-classification ·1,494
vision ·1,494
safetensors ·1,246
siglip ·1,113
pytorch ·371
chinese_clip ·227
siglip2 ·123
align ·31