kanashi6/UFO

Model

This repository contains the model presented in the paper UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface.

3

52 commits

1 linked in READMEs

updated Sep 27, 2025

See the code

README

This repository contains the model presented in the paper UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface.

UFO unifies object-level detection, pixel-level segmentation, and image-level vision-language tasks into a single model by transforming all perception targets into the language space. It introduces a novel embedding retrieval approach that relies solely on the language interface to support segmentation tasks.

For more details, please refer to the original paper and the GitHub repository:

endpoints_compatible
image-text-to-text
transformers

Contributors

kanashi6

51 commits

nielsr

1 commits

kanashi6/UFO

Model

This repository contains the model presented in the paper UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface.

3

52 commits

1 linked in READMEs

updated Sep 27, 2025

See the code

README

This repository contains the model presented in the paper UFO: A Unified Approach to Fine-grained Visual Perception via Open-ended Language Interface.

UFO unifies object-level detection, pixel-level segmentation, and image-level vision-language tasks into a single model by transforming all perception targets into the language space. It introduces a novel embedding retrieval approach that relies solely on the language interface to support segmentation tasks.

For more details, please refer to the original paper and the GitHub repository:

endpoints_compatible
image-text-to-text
transformers

Contributors

kanashi6

51 commits

nielsr

1 commits