16 repos
Model checkpoints and fine-tuned variants of Florence-2, a vision-language foundation model designed for image-text understanding tasks. The cluster contains base and large model versions along with fine-tuned adaptations, distributed via safetensors format and compatible with Hugging Face endpoints. Repositories here provide different model scales and specializations for tasks requiring joint reasoning over images and text.