A unified, multi-modal evaluation benchmark for controllable captioning across images, videos, and audio. AnyCapEval is designed to test both content adherence (how well captions follow explicit user instructions) and style consistency (fluency, tone, and expressiveness) under a diversity of control directives.
AnyCapEval/
├── anycapeval_image/ # Test examples for image captioning (instruction, reference, candidate)
├── anycapeval_video/ # Test examples for video captioning
├── anycapeval_audio/ # Test examples for audio captioning
└── LICENSE # Apache-2.0 license for data
(instruction, high_quality_caption, low_quality_caption)pip install datasets
from datasets import load_dataset
ds = load_dataset("qishisuren/AnyCapEval", split="test")
print(ds[0])
This dataset is released under the Apache-2.0 license. See LICENSE for details.
A unified, multi-modal evaluation benchmark for controllable captioning across images, videos, and audio. AnyCapEval is designed to test both content adherence (how well captions follow explicit user instructions) and style consistency (fluency, tone, and expressiveness) under a diversity of control directives.
AnyCapEval/
├── anycapeval_image/ # Test examples for image captioning (instruction, reference, candidate)
├── anycapeval_video/ # Test examples for video captioning
├── anycapeval_audio/ # Test examples for audio captioning
└── LICENSE # Apache-2.0 license for data
(instruction, high_quality_caption, low_quality_caption)pip install datasets
from datasets import load_dataset
ds = load_dataset("qishisuren/AnyCapEval", split="test")
print(ds[0])
This dataset is released under the Apache-2.0 license. See LICENSE for details.