turing-motors/Japanese-Heron-Bench

Dataset

11

stars

4

commits

2

linked in READMEs

Apr 12, 2024

updated

Browse cluster: Japanese Vision-Language Models

README

Japanese-Heron-Bench

Dataset Description

Japanese-Heron-Bench is a benchmark for evaluating Japanese VLMs (Vision-Language Models). We collected 21 images related to Japan. We then set up three categories for each image: Conversation, Detail, and Complex, and prepared one or two questions for each category. The final evaluation dataset consists of 102 questions. Furthermore, each image is assigned one of seven subcategories: anime, art, culture, food, landscape, landmark, and transportation.

For more details and the run script, please visit to our GitHub repository.

Uses

We have collected images that are either in the public domain or licensed under Creative Commons Attribution 1.0 (CC BY 1.0) or Creative Commons Attribution 2.0 (CC BY 2.0). Please refer to the LICENSE.md file for details on the licenses.

Citation

@misc{inoue2024heronbench,
      title={Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese}, 
      author={Yuichi Inoue and Kento Sasaki and Yuma Ochi and Kazuki Fujii and Kotaro Tanahashi and Yu Yamaguchi},
      year={2024},
      eprint={2404.07824},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

Inoichan

2 commits

KentoSasaki

2 commits

turing-motors/Japanese-Heron-Bench

Dataset

11

stars

4

commits

2

linked in READMEs

Apr 12, 2024

updated

Browse cluster: Japanese Vision-Language Models

README

Japanese-Heron-Bench

Dataset Description

Japanese-Heron-Bench is a benchmark for evaluating Japanese VLMs (Vision-Language Models). We collected 21 images related to Japan. We then set up three categories for each image: Conversation, Detail, and Complex, and prepared one or two questions for each category. The final evaluation dataset consists of 102 questions. Furthermore, each image is assigned one of seven subcategories: anime, art, culture, food, landscape, landmark, and transportation.

For more details and the run script, please visit to our GitHub repository.

Uses

We have collected images that are either in the public domain or licensed under Creative Commons Attribution 1.0 (CC BY 1.0) or Creative Commons Attribution 2.0 (CC BY 2.0). Please refer to the LICENSE.md file for details on the licenses.

Citation

@misc{inoue2024heronbench,
      title={Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese}, 
      author={Yuichi Inoue and Kento Sasaki and Yuma Ochi and Kazuki Fujii and Kotaro Tanahashi and Yu Yamaguchi},
      year={2024},
      eprint={2404.07824},
      archivePrefix={arXiv},
      primaryClass={cs.CV}
}

Contributors

Inoichan

2 commits

KentoSasaki

2 commits