TROHN-Text is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the caption. Its purpose has been to train contrastive models by adding only hard negatives in the form of text to improve compositional understanding. You can find the fine-tuned CLIP model in CLIP_TROHN-Text.
Each instance of the dataset consists of three fields:
To load data with datasets:
>>> data = load_dataset("imirandam/TROHN-Text")
Each instance has the following structure:
{
'image_id': '000000391979.jpg' ,
'caption': 'A bird is flying over the water of a beach.',
'negative_caption': 'A bird is flying over the snow of a mountain.',
}
TROHN-Text has 3,652,846 instances consisting of 1 image and 2 captions. It is divided into two splits, 80% train and 20% validation.
This dataset has been created semi-automatically using the LLM OpenCHAT-3.5 and templates. Instances are not checked and may contain incorrect, duplicate, etc. information.
If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.
This work is licensed under a MIT License.
If you find this dataset useful, please consider citing our paper:
@misc{miranda2024bivlc,
title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval},
author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
year={2024},
eprint={2406.09952},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
8 commits
TROHN-Text is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the caption. Its purpose has been to train contrastive models by adding only hard negatives in the form of text to improve compositional understanding. You can find the fine-tuned CLIP model in CLIP_TROHN-Text.
Each instance of the dataset consists of three fields:
To load data with datasets:
>>> data = load_dataset("imirandam/TROHN-Text")
Each instance has the following structure:
{
'image_id': '000000391979.jpg' ,
'caption': 'A bird is flying over the water of a beach.',
'negative_caption': 'A bird is flying over the snow of a mountain.',
}
TROHN-Text has 3,652,846 instances consisting of 1 image and 2 captions. It is divided into two splits, 80% train and 20% validation.
This dataset has been created semi-automatically using the LLM OpenCHAT-3.5 and templates. Instances are not checked and may contain incorrect, duplicate, etc. information.
If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.
This work is licensed under a MIT License.
If you find this dataset useful, please consider citing our paper:
@misc{miranda2024bivlc,
title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval},
author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
year={2024},
eprint={2406.09952},
archivePrefix={arXiv},
primaryClass={cs.CV}
}
8 commits