imirandam/TROHN-Img

Dataset

3

stars

7

commits

1

linked in READMEs

Jun 17, 2024

updated

README

Dataset Card for TROHN-Img

Dataset Description

Dataset Summary

TROHN-Img is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the COCO caption and subsequently a negative image is created from the generated negative caption using the SD-XL model. Its objective has been to train contrastive models by adding negative pairs, i.e., caption and negative images, to improve compositional understanding. The fine-tuned CLIP model can be found in CLIP_TROHN-Img.

Dataset instances

Each instance of the dataset consists of three fields:

  • image_id: COCO 2017 train image id.
  • caption: COCO 2017 train text describing the COCO image.
  • negative_caption: Negative caption generated from the COCO 2017 train text description by BiVLC.
  • negative_image: Negative image generated from the negative_caption by BiVLC.

How to use

To load data with datasets:

>>> data = load_dataset("imirandam/TROHN-Img")

Instance example

Each instance has the following structure:

{
    'image_id': '000000103673.jpg' ,
    'caption': 'Three monkeys sit on a fence eating bananas.',
    'negative_caption': 'Three monkeys sit on a fence drinking water.',
    'negative_image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=512x512 at 0x7F9BE45571C0>
}

Dataset statistics

TROHN-Img has 296,070 instances consisting of 2 images and 2 captions. It is divided into two splits, 80% train and 20% validation.

Source Data

  • image and caption are from COCO 2017 train split.

Dataset curation

This dataset was created by filtering the TROHN-Text dataset based on plausibility and linguistic acceptability scores; images are then generated from the negative captions. Instances are not checked and may contain incorrect, duplicate, etc. information.

Evaluation Data

If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.

Licensing Information

This work is licensed under a MIT License.

Citation Information

If you find this dataset useful, please consider citing our paper:

@misc{miranda2024bivlc,
      title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval}, 
      author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
      year={2024},
      eprint={2406.09952},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

imirandam

7 commits

imirandam/TROHN-Img

Dataset

3

stars

7

commits

1

linked in READMEs

Jun 17, 2024

updated

README

Dataset Card for TROHN-Img

Dataset Description

Dataset Summary

TROHN-Img is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the COCO caption and subsequently a negative image is created from the generated negative caption using the SD-XL model. Its objective has been to train contrastive models by adding negative pairs, i.e., caption and negative images, to improve compositional understanding. The fine-tuned CLIP model can be found in CLIP_TROHN-Img.

Dataset instances

Each instance of the dataset consists of three fields:

  • image_id: COCO 2017 train image id.
  • caption: COCO 2017 train text describing the COCO image.
  • negative_caption: Negative caption generated from the COCO 2017 train text description by BiVLC.
  • negative_image: Negative image generated from the negative_caption by BiVLC.

How to use

To load data with datasets:

>>> data = load_dataset("imirandam/TROHN-Img")

Instance example

Each instance has the following structure:

{
    'image_id': '000000103673.jpg' ,
    'caption': 'Three monkeys sit on a fence eating bananas.',
    'negative_caption': 'Three monkeys sit on a fence drinking water.',
    'negative_image': <PIL.JpegImagePlugin.JpegImageFile image mode=RGB size=512x512 at 0x7F9BE45571C0>
}

Dataset statistics

TROHN-Img has 296,070 instances consisting of 2 images and 2 captions. It is divided into two splits, 80% train and 20% validation.

Source Data

  • image and caption are from COCO 2017 train split.

Dataset curation

This dataset was created by filtering the TROHN-Text dataset based on plausibility and linguistic acceptability scores; images are then generated from the negative captions. Instances are not checked and may contain incorrect, duplicate, etc. information.

Evaluation Data

If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.

Licensing Information

This work is licensed under a MIT License.

Citation Information

If you find this dataset useful, please consider citing our paper:

@misc{miranda2024bivlc,
      title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval}, 
      author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
      year={2024},
      eprint={2406.09952},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

imirandam

7 commits