Visual Referential Understanding & COCO Datasets

8 repos

Datasets and validation resources for visual reference understanding tasks, centered around the COCO (Common Objects in Context) family of datasets. The cluster includes RefCOCO variants and related captioning benchmarks used to train and evaluate models that understand referring expressions — where humans use natural language to identify specific objects in images. These resources serve as standard evaluation sets for computer vision and vision-language research.