Vision Transformers & Image Feature Extraction

11 repos

Self-supervised learning models for extracting visual features from images using Vision Transformer (ViT) architectures, particularly the DINO family of pre-trained models at various scales. These repositories provide pretrained weights and implementations optimized for image-feature-extraction tasks, leveraging the transformers library and safetensors format for efficient model distribution and inference across compatible endpoints.

dinov3 ·1,698
transformers ·1,698
endpoints_compatible ·1,698
safetensors ·1,698
dino ·1,675
image-feature-extraction ·1,675
en ·1,673
dinov3_vit ·1,650
dinov3_convnext ·25
chmv2 ·23