Florence-2 Vision-Language Models

16 repos

Model checkpoints and fine-tuned variants of Florence-2, a vision-language foundation model designed for image-text understanding tasks. The cluster contains base and large model versions along with fine-tuned adaptations, distributed via safetensors format and compatible with Hugging Face endpoints. Repositories here provide different model scales and specializations for tasks requiring joint reasoning over images and text.

Python · 2
safetensors ·3,367
custom_code ·3,367
florence2 ·3,367
transformers ·3,000
endpoints_compatible ·3,000
image-text-to-text ·3,000
vision ·2,794
pytorch ·2,794
en ·206
art ·206