SugarCrepe++ allows us to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. The instances from SugarCrepe++ dataset represent images from MS-COCO and their associated text captions, negative captions from SugarCrepe and newly introduced positive captions.
Despite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understand precise semantics. For example, semantically equivalent sentences expressed using different lexical compositions elicit diverging representations. The degree of this divergence and its impact on encoded semantics is not very well understood. In this paper, we introduce the SugarCrepe++ dataset to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. Each sample in SugarCrepe++ dataset consists of an image and a corresponding triplet of captions: a pair of semantically equivalent but lexically different positive captions and one hard negative caption. This poses a 3-way semantic (in)equivalence problem to the language models. Given the importance of the property that the SugarCrepe++ dataset targets, it serves as a new challenge to the vision-and-language community
This work is licensed under a Creative Commons Attribution 4.0 International License.
The primary use case of our benchmark is to evaluate the sensitivity of VLMs and ULMs to semantic and lexical alterations.
This is only an evaluation dataset and should only be used for evaluating LLMs. This includes both vision language models (VLMs) and unimodel language models (ULMs). Refer to the paper for information (link to be added).
There are five categories separated by the type of negative caption. Each file contains the following fields:
Images of the COCO-2017 validation set can be downloaded from here.
SugerCrepe++ was created using the generative model and validated using expert humans. The complete process of generation is described in the paper (to be added). Refer to the Github repository for more info.
The SugarCrepe++ dataset was created to evaluate the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations. The SugarCrepe dataset consists of (only) one positive and one hard negative caption for each image. Relative to the negative caption, a single positive caption can either have low or high lexical overlap. The original SugarCrepe only captures the high overlap case. To evaluate the sensitivity of encoded semantics to lexical alteration, we require an additional positive caption with a different lexical composition. SugarCrepe++ fills this gap by adding an additional positive caption enabling a more thorough assessment of models’ abilities to handle semantic content and lexical variation
we source part of our dataset, such as image-caption pairs from MS-COCO and negative captions from SugarCrepe, both of which are open-source datasets.
Refer to the Github repository for more info.
Please refer to the original source datasets for more info MS-COCO and SugarCrepe.
The creators of this paper generated the textual content using generative AI and manually validated it. More info in the paper (to be added).
Expert humans and generative AI. We used the mistral model to generate the positive caption. Refer to the Github repository for more info.
Some images might contain identifiable individual faces.
Due to the reliance on the MS-COCO and SugarCrepe datasets, SUGAR- CREPE++ may contain offensive material or biases present in these source datasets. Users of SugarCrepe++ should carefully consider how these limitations may impact their potential use case and exercise discretion in their dataset application.
The dataset should be avoided for a task if the limitations discussed above are unacceptable or potentially problematic for the intended use case.
BibTeX:
@misc{dumpala2024sugarcrepe,
title={SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations},
author={Sri Harsha Dumpala and Aman Jaiswal and Chandramouli Sastry and Evangelos Milios and Sageev Oore and Hassan Sajjad},
year={2024},
eprint={2406.11171},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
email: aman.jaiswal@dal.ca
5 commits
SugarCrepe++ allows us to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. The instances from SugarCrepe++ dataset represent images from MS-COCO and their associated text captions, negative captions from SugarCrepe and newly introduced positive captions.
Despite their remarkable successes, state-of-the-art large language models (LLMs), including vision-and-language models (VLMs) and unimodal language models (ULMs), fail to understand precise semantics. For example, semantically equivalent sentences expressed using different lexical compositions elicit diverging representations. The degree of this divergence and its impact on encoded semantics is not very well understood. In this paper, we introduce the SugarCrepe++ dataset to analyze the sensitivity of VLMs and ULMs to lexical and semantic alterations. Each sample in SugarCrepe++ dataset consists of an image and a corresponding triplet of captions: a pair of semantically equivalent but lexically different positive captions and one hard negative caption. This poses a 3-way semantic (in)equivalence problem to the language models. Given the importance of the property that the SugarCrepe++ dataset targets, it serves as a new challenge to the vision-and-language community
This work is licensed under a Creative Commons Attribution 4.0 International License.
The primary use case of our benchmark is to evaluate the sensitivity of VLMs and ULMs to semantic and lexical alterations.
This is only an evaluation dataset and should only be used for evaluating LLMs. This includes both vision language models (VLMs) and unimodel language models (ULMs). Refer to the paper for information (link to be added).
There are five categories separated by the type of negative caption. Each file contains the following fields:
Images of the COCO-2017 validation set can be downloaded from here.
SugerCrepe++ was created using the generative model and validated using expert humans. The complete process of generation is described in the paper (to be added). Refer to the Github repository for more info.
The SugarCrepe++ dataset was created to evaluate the sensitivity of vision language models (VLMs) and unimodal language models (ULMs) to semantic and lexical alterations. The SugarCrepe dataset consists of (only) one positive and one hard negative caption for each image. Relative to the negative caption, a single positive caption can either have low or high lexical overlap. The original SugarCrepe only captures the high overlap case. To evaluate the sensitivity of encoded semantics to lexical alteration, we require an additional positive caption with a different lexical composition. SugarCrepe++ fills this gap by adding an additional positive caption enabling a more thorough assessment of models’ abilities to handle semantic content and lexical variation
we source part of our dataset, such as image-caption pairs from MS-COCO and negative captions from SugarCrepe, both of which are open-source datasets.
Refer to the Github repository for more info.
Please refer to the original source datasets for more info MS-COCO and SugarCrepe.
The creators of this paper generated the textual content using generative AI and manually validated it. More info in the paper (to be added).
Expert humans and generative AI. We used the mistral model to generate the positive caption. Refer to the Github repository for more info.
Some images might contain identifiable individual faces.
Due to the reliance on the MS-COCO and SugarCrepe datasets, SUGAR- CREPE++ may contain offensive material or biases present in these source datasets. Users of SugarCrepe++ should carefully consider how these limitations may impact their potential use case and exercise discretion in their dataset application.
The dataset should be avoided for a task if the limitations discussed above are unacceptable or potentially problematic for the intended use case.
BibTeX:
@misc{dumpala2024sugarcrepe,
title={SUGARCREPE++ Dataset: Vision-Language Model Sensitivity to Semantic and Lexical Alterations},
author={Sri Harsha Dumpala and Aman Jaiswal and Chandramouli Sastry and Evangelos Milios and Sageev Oore and Hassan Sajjad},
year={2024},
eprint={2406.11171},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
email: aman.jaiswal@dal.ca
5 commits