imirandam/TROHN-Text

Dataset

2

stars

8

commits

1

linked in READMEs

Jun 17, 2024

updated

README

Dataset Card for TROHN-Text

Dataset Description

Dataset Summary

TROHN-Text is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the caption. Its purpose has been to train contrastive models by adding only hard negatives in the form of text to improve compositional understanding. You can find the fine-tuned CLIP model in CLIP_TROHN-Text.

Dataset instances

Each instance of the dataset consists of three fields:

  • image_id: COCO 2017 train image id.
  • caption: COCO 2017 train text describing the COCO image.
  • negative_caption: Negative caption generated from the COCO 2017 train text description by BiVLC.

How to use

To load data with datasets:

>>> data = load_dataset("imirandam/TROHN-Text")

Instance example

Each instance has the following structure:

{
    'image_id': '000000391979.jpg' ,
    'caption': 'A bird is flying over the water of a beach.',
    'negative_caption': 'A bird is flying over the snow of a mountain.',
}

Dataset statistics

TROHN-Text has 3,652,846 instances consisting of 1 image and 2 captions. It is divided into two splits, 80% train and 20% validation.

Source Data

  • image and caption are from COCO 2017 train split.

Dataset curation

This dataset has been created semi-automatically using the LLM OpenCHAT-3.5 and templates. Instances are not checked and may contain incorrect, duplicate, etc. information.

Evaluation Data

If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.

Licensing Information

This work is licensed under a MIT License.

Citation Information

If you find this dataset useful, please consider citing our paper:

@misc{miranda2024bivlc,
      title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval}, 
      author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
      year={2024},
      eprint={2406.09952},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

imirandam

8 commits

imirandam/TROHN-Text

Dataset

2

stars

8

commits

1

linked in READMEs

Jun 17, 2024

updated

README

Dataset Card for TROHN-Text

Dataset Description

Dataset Summary

TROHN-Text is a dataset presented in the BiVLC paper for experimentation. It is based on the COCO 2017 train split, a negative caption with an LLM is created from the caption. Its purpose has been to train contrastive models by adding only hard negatives in the form of text to improve compositional understanding. You can find the fine-tuned CLIP model in CLIP_TROHN-Text.

Dataset instances

Each instance of the dataset consists of three fields:

  • image_id: COCO 2017 train image id.
  • caption: COCO 2017 train text describing the COCO image.
  • negative_caption: Negative caption generated from the COCO 2017 train text description by BiVLC.

How to use

To load data with datasets:

>>> data = load_dataset("imirandam/TROHN-Text")

Instance example

Each instance has the following structure:

{
    'image_id': '000000391979.jpg' ,
    'caption': 'A bird is flying over the water of a beach.',
    'negative_caption': 'A bird is flying over the snow of a mountain.',
}

Dataset statistics

TROHN-Text has 3,652,846 instances consisting of 1 image and 2 captions. It is divided into two splits, 80% train and 20% validation.

Source Data

  • image and caption are from COCO 2017 train split.

Dataset curation

This dataset has been created semi-automatically using the LLM OpenCHAT-3.5 and templates. Instances are not checked and may contain incorrect, duplicate, etc. information.

Evaluation Data

If you need evaluation data, you can use the dataset proposed in the paper in the following link, BiVLC.

Licensing Information

This work is licensed under a MIT License.

Citation Information

If you find this dataset useful, please consider citing our paper:

@misc{miranda2024bivlc,
      title={BiVLC: Extending Vision-Language Compositionality Evaluation with Text-to-Image Retrieval}, 
      author={Imanol Miranda and Ander Salaberria and Eneko Agirre and Gorka Azkune},
      year={2024},
      eprint={2406.09952},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

imirandam

8 commits