The SeeTRUE-Feedback dataset is a diverse benchmark for the meta-evaluation of image-text matching/alignment feedback. It aims to overcome limitations in current benchmarks, which primarily focus on predicting a matching score between 0-1. SeeTRUE provides, for each row, the original caption, feedback related to text-image misalignment, and the caption+visual source of misalignments (including a bounding box for the visual misalignment).
The dataset supports English language.
feedback field.SeeTRUE-Feedback contains a single split: TEST, and should not be used for training.
The dataset has been created by sourcing and matching images and text from multiple datasets. More information in the paper:
The dataset is under the CC-By 4.0 license.
TODO
11 commits
The SeeTRUE-Feedback dataset is a diverse benchmark for the meta-evaluation of image-text matching/alignment feedback. It aims to overcome limitations in current benchmarks, which primarily focus on predicting a matching score between 0-1. SeeTRUE provides, for each row, the original caption, feedback related to text-image misalignment, and the caption+visual source of misalignments (including a bounding box for the visual misalignment).
The dataset supports English language.
feedback field.SeeTRUE-Feedback contains a single split: TEST, and should not be used for training.
The dataset has been created by sourcing and matching images and text from multiple datasets. More information in the paper:
The dataset is under the CC-By 4.0 license.
TODO
11 commits