timm/vit_base_patch16_siglip_512.v2_webli

Model

4

stars

2

commits

1

linked in READMEs

Feb 21, 2025

updated

image-feature-extraction
pytorch
safetensors
siglip
siglip2
timm
transformers

README

Model card for vit_base_patch16_siglip_512.v2_webli

A SigLIP 2 ViT (image encoder only) for timm. Equivalent to image tower from https://huggingface.co/timm/ViT-B-16-SigLIP2-512.

Model Details

Citation

@article{tschannen2025siglip,
          title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
          author={Tschannen, Michael and Gritsenko, Alexey and Wang, Xiao and Naeem, Muhammad Ferjad and Alabdulmohsin, Ibrahim and Parthasarathy, Nikhil and Evans, Talfan and Beyer, Lucas and Xia, Ye and Mustafa, Basil and H'enaff, Olivier and Harmsen, Jeremiah and Steiner, Andreas and Zhai, Xiaohua},
          year={2025},
          journal={arXiv preprint arXiv:2502.14786}
        }
        
@inproceedings{zhai2023sigmoid,
          title={Sigmoid loss for language image pre-training},
          author={Zhai, Xiaohua and Mustafa, Basil and Kolesnikov, Alexander and Beyer, Lucas},
          booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
          pages={11975--11986},
          year={2023}
        }
        

Contributors

rwightman

2 commits

timm/vit_base_patch16_siglip_512.v2_webli

Model

4

stars

2

commits

1

linked in READMEs

Feb 21, 2025

updated

image-feature-extraction
pytorch
safetensors
siglip
siglip2
timm
transformers

README

Model card for vit_base_patch16_siglip_512.v2_webli

A SigLIP 2 ViT (image encoder only) for timm. Equivalent to image tower from https://huggingface.co/timm/ViT-B-16-SigLIP2-512.

Model Details

Citation

@article{tschannen2025siglip,
          title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features},
          author={Tschannen, Michael and Gritsenko, Alexey and Wang, Xiao and Naeem, Muhammad Ferjad and Alabdulmohsin, Ibrahim and Parthasarathy, Nikhil and Evans, Talfan and Beyer, Lucas and Xia, Ye and Mustafa, Basil and H'enaff, Olivier and Harmsen, Jeremiah and Steiner, Andreas and Zhai, Xiaohua},
          year={2025},
          journal={arXiv preprint arXiv:2502.14786}
        }
        
@inproceedings{zhai2023sigmoid,
          title={Sigmoid loss for language image pre-training},
          author={Zhai, Xiaohua and Mustafa, Basil and Kolesnikov, Alexander and Beyer, Lucas},
          booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision},
          pages={11975--11986},
          year={2023}
        }
        

Contributors

rwightman

2 commits