Vision Transformers and Scalable Vision Encoders

15 repos

Libraries and models for building efficient vision transformers and scalable vision encoders, primarily using PyTorch. The cluster focuses on feature extraction and visual representation learning, with particular emphasis on Vision Transformer (ViT) variants at different scales and resolutions. Resources here support both research and production deployment of vision-based machine learning models.

Python · 1
scalable-vision-encoder ·211
vlms ·211
feature-extraction ·71
transformers ·61
custom_code ·61
pytorch ·61
safetensors ·45
endpoints_compatible ·35
image-feature-extraction ·35
internvl ·35