mair-lab/vismin-bench

Dataset

VisMin Dataset

0

8 commits

1 linked in READMEs

updated Jan 22, 2025

See the code

README

VisMin Dataset

Overview

The VisMin dataset is designed for evaluating models on minimal-change tasks involving image-caption pairs. It consists of four types of minimal changes: object, attribute, count, and spatial relation. The dataset is used to benchmark models' ability to predict the correct image-caption match given two images and one caption, or two captions and one image.

Dataset Structure

  • Total Samples: 2,084
    • Object Changes: 579 samples
    • Attribute Changes: 294 samples
    • Counting Changes: 589 samples
    • Spatial Relation Changes: 622 samples

Data Format

Each sample in the dataset includes:

  • image_0: Image 1.
  • text_0: Caption 1.
  • image_1: Image 2.
  • text_1: Caption 2.
  • question_1: Question 1.
  • question_2: Question 2.
  • question_3: Question 3.
  • question_4: Question 4.
  • id: Unique identifier for the sample.

Usage

The dataset can be used with different model types, such as CLIP or MLLM, to evaluate their performance on the VisMin benchmark tasks.

from datasets import load_dataset

dataset = load_dataset("https://huggingface.co/datasets/rabiulawal/vismin-bench")

Evaluation

To evaluate the performance of a model on the VisMin benchmark, you can utilize the vismin repository.

The following CSV file serve as the ground truth for evaluation and is available in the repository:

  • vqa_solution.csv: Provides ground truth for Multi-Modal Language Model (MLLM) tasks.

We recommend using the example script in the repository to facilitate the evaluation process.

Citation Information

If you are using this dataset, please cite

 @article{vismin2024,
    title={VisMin: Visual Minimal-Change Understanding},
    author={Awal, Rabiul and Ahmadi, Saba and Zhang, Le and Agrawal, Aishwarya},
    year={2024}
}

Contributors

rabiulawal

8 commits

mair-lab/vismin-bench

Dataset

VisMin Dataset

0

8 commits

1 linked in READMEs

updated Jan 22, 2025

See the code

README

VisMin Dataset

Overview

The VisMin dataset is designed for evaluating models on minimal-change tasks involving image-caption pairs. It consists of four types of minimal changes: object, attribute, count, and spatial relation. The dataset is used to benchmark models' ability to predict the correct image-caption match given two images and one caption, or two captions and one image.

Dataset Structure

  • Total Samples: 2,084
    • Object Changes: 579 samples
    • Attribute Changes: 294 samples
    • Counting Changes: 589 samples
    • Spatial Relation Changes: 622 samples

Data Format

Each sample in the dataset includes:

  • image_0: Image 1.
  • text_0: Caption 1.
  • image_1: Image 2.
  • text_1: Caption 2.
  • question_1: Question 1.
  • question_2: Question 2.
  • question_3: Question 3.
  • question_4: Question 4.
  • id: Unique identifier for the sample.

Usage

The dataset can be used with different model types, such as CLIP or MLLM, to evaluate their performance on the VisMin benchmark tasks.

from datasets import load_dataset

dataset = load_dataset("https://huggingface.co/datasets/rabiulawal/vismin-bench")

Evaluation

To evaluate the performance of a model on the VisMin benchmark, you can utilize the vismin repository.

The following CSV file serve as the ground truth for evaluation and is available in the repository:

  • vqa_solution.csv: Provides ground truth for Multi-Modal Language Model (MLLM) tasks.

We recommend using the example script in the repository to facilitate the evaluation process.

Citation Information

If you are using this dataset, please cite

 @article{vismin2024,
    title={VisMin: Visual Minimal-Change Understanding},
    author={Awal, Rabiul and Ahmadi, Saba and Zhang, Le and Agrawal, Aishwarya},
    year={2024}
}

Contributors

rabiulawal

8 commits