17 repos
Pre-trained model weights and implementations for vision-language transformers that perform zero-shot image classification and multimodal understanding. The cluster centers on CLIP variants (Chinese CLIP and SigLIP2 architectures) optimized for different patch sizes and resolutions, with Hugging Face transformers integration and safetensors format support for easy deployment and inference.