31 repos
Pre-trained and fine-tuned CLIP (Contrastive Language-Image Pre-training) models for zero-shot image classification and multimodal learning. The cluster centers on OpenCLIP implementations and model variants trained on different scales of data (LAION, CommonPool), spanning different Vision Transformer architectures (ViT-B, ViT-L, ViT-H, ViT-bigG) and checkpoint sizes. Repositories here provide model weights, training infrastructure, and implementations for leveraging vision-language representations across computer vision and multimodal tasks.