AVGen-Bench Generated Videos Data Card
6
12 commits
1 linked in READMEs
updated May 27, 2026
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a model-generated video produced from a prompt defined in prompts/*.json.
For Hugging Face Hub compatibility, the repository includes a root-level metadata.parquet file so the Dataset Viewer can expose each video as a structured row with prompt metadata instead of treating the repo as an unindexed file dump.
The relative video path is stored as a plain string column (video_path) rather than a media-typed file_name column, which avoids current Dataset Viewer post-processing failures on video rows.
As described in the GitHub repository, you can generate videos from the benchmark prompts using the following command:
python batch_generate.py \
--provider sora2 \
--task_type video_generation \
--prompts_dir ./prompts \
--out_dir ./generated_videos/sora2 \
--concurrency 2 \
--seconds 12 \
--size 1280x720
The dataset is organized by:
.mp4 filesA typical top-level structure is:
AVGen-Bench/
βββ Kling_2.6/
βββ LTX-2/
βββ LTX-2.3/
βββ MOVA_360p_Emu3.5/
βββ MOVA_360p_NanoBanana_2/
βββ Ovi_11/
βββ Seedance_1.5_pro/
βββ Seedance_2/
βββ Sora_2/
βββ Veo_3.1_fast/
βββ Veo_3.1_quality/
βββ Wan_2.2_HunyuanVideo-Foley/
βββ Wan_2.6/
βββ metadata.parquet
βββ prompts/
βββ reference_image/ # optional, depending on generation pipeline
Within each model directory, videos are grouped by category, for example:
Veo_3.1_fast/
βββ ads/
βββ animals/
βββ asmr/
βββ chemical_reaction/
βββ cooking/
βββ gameplays/
βββ movie_trailer/
βββ musical_instrument_tutorial/
βββ news/
βββ physical_experiment/
βββ sports/
Prompt definitions are stored in prompts/*.json.
The current prompt set contains 235 prompts across 11 categories:
| Category | Prompt count |
|---|---|
ads | 20 |
animals | 20 |
asmr | 20 |
chemical_reaction | 20 |
cooking | 20 |
gameplays | 20 |
movie_trailer | 20 |
musical_instrument_tutorial | 35 |
news | 20 |
physical_experiment | 20 |
sports | 20 |
Prompt JSON entries typically contain:
content: a short content descriptor used for naming or indexingprompt: the full generation promptEach generated item is typically:
.mp4 file<model>/<category>/The filename is usually derived from prompt content after sanitization. Exact naming may vary by generation script or provider wrapper.
In the standard export pipeline, the filename is derived from the prompt's content field using the following logic:
def safe_filename(name: str, max_len: int = 180) -> str:
name = str(name).strip()
name = re.sub(r"[/\\:*?\"<>|\
\\r\\t]", "_", name)
name = re.sub(r"\\s+", " ", name).strip()
if not name:
name = "untitled"
if len(name) > max_len:
name = name[:max_len].rstrip()
return name
So the expected output path pattern is:
<model>/<category>/<safe_filename(content)>.mp4
For Dataset Viewer indexing, metadata.parquet stores one row per exported video with:
video_path: relative path to the .mp4 stored as a plain stringmodel: model directory namecategory: benchmark categorycontent: prompt short nameprompt: full generation promptprompt_id: index inside prompts/<category>.jsonThe videos were generated by running different T2AV systems on a shared benchmark prompt set.
Important properties:
This dataset is intended for:
This dataset is not intended for:
Because these are generated videos:
Anyone redistributing results should clearly label them as synthetic model outputs.
If you find AVGen-Bench useful, please cite:
@misc{zhou2026avgenbenchtaskdrivenbenchmarkmultigranular,
title={AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation},
author={Ziwei Zhou and Zeyuan Lai and Rui Wang and Yifan Yang and Zhen Xing and Yuqing Yang and Qi Dai and Lili Qiu and Chong Luo},
year={2026},
eprint={2604.08540},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.08540},
}
AVGen-Bench Generated Videos Data Card
6
12 commits
1 linked in READMEs
updated May 27, 2026
This data card describes the generated audio-video outputs stored directly in the repository root by model directory.
The collection is intended for benchmarking and qualitative/quantitative evaluation of text-to-audio-video (T2AV) systems. It was presented in the paper AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation. It is not a training dataset. Each item is a model-generated video produced from a prompt defined in prompts/*.json.
For Hugging Face Hub compatibility, the repository includes a root-level metadata.parquet file so the Dataset Viewer can expose each video as a structured row with prompt metadata instead of treating the repo as an unindexed file dump.
The relative video path is stored as a plain string column (video_path) rather than a media-typed file_name column, which avoids current Dataset Viewer post-processing failures on video rows.
As described in the GitHub repository, you can generate videos from the benchmark prompts using the following command:
python batch_generate.py \
--provider sora2 \
--task_type video_generation \
--prompts_dir ./prompts \
--out_dir ./generated_videos/sora2 \
--concurrency 2 \
--seconds 12 \
--size 1280x720
The dataset is organized by:
.mp4 filesA typical top-level structure is:
AVGen-Bench/
βββ Kling_2.6/
βββ LTX-2/
βββ LTX-2.3/
βββ MOVA_360p_Emu3.5/
βββ MOVA_360p_NanoBanana_2/
βββ Ovi_11/
βββ Seedance_1.5_pro/
βββ Seedance_2/
βββ Sora_2/
βββ Veo_3.1_fast/
βββ Veo_3.1_quality/
βββ Wan_2.2_HunyuanVideo-Foley/
βββ Wan_2.6/
βββ metadata.parquet
βββ prompts/
βββ reference_image/ # optional, depending on generation pipeline
Within each model directory, videos are grouped by category, for example:
Veo_3.1_fast/
βββ ads/
βββ animals/
βββ asmr/
βββ chemical_reaction/
βββ cooking/
βββ gameplays/
βββ movie_trailer/
βββ musical_instrument_tutorial/
βββ news/
βββ physical_experiment/
βββ sports/
Prompt definitions are stored in prompts/*.json.
The current prompt set contains 235 prompts across 11 categories:
| Category | Prompt count |
|---|---|
ads | 20 |
animals | 20 |
asmr | 20 |
chemical_reaction | 20 |
cooking | 20 |
gameplays | 20 |
movie_trailer | 20 |
musical_instrument_tutorial | 35 |
news | 20 |
physical_experiment | 20 |
sports | 20 |
Prompt JSON entries typically contain:
content: a short content descriptor used for naming or indexingprompt: the full generation promptEach generated item is typically:
.mp4 file<model>/<category>/The filename is usually derived from prompt content after sanitization. Exact naming may vary by generation script or provider wrapper.
In the standard export pipeline, the filename is derived from the prompt's content field using the following logic:
def safe_filename(name: str, max_len: int = 180) -> str:
name = str(name).strip()
name = re.sub(r"[/\\:*?\"<>|\
\\r\\t]", "_", name)
name = re.sub(r"\\s+", " ", name).strip()
if not name:
name = "untitled"
if len(name) > max_len:
name = name[:max_len].rstrip()
return name
So the expected output path pattern is:
<model>/<category>/<safe_filename(content)>.mp4
For Dataset Viewer indexing, metadata.parquet stores one row per exported video with:
video_path: relative path to the .mp4 stored as a plain stringmodel: model directory namecategory: benchmark categorycontent: prompt short nameprompt: full generation promptprompt_id: index inside prompts/<category>.jsonThe videos were generated by running different T2AV systems on a shared benchmark prompt set.
Important properties:
This dataset is intended for:
This dataset is not intended for:
Because these are generated videos:
Anyone redistributing results should clearly label them as synthetic model outputs.
If you find AVGen-Bench useful, please cite:
@misc{zhou2026avgenbenchtaskdrivenbenchmarkmultigranular,
title={AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation},
author={Ziwei Zhou and Zeyuan Lai and Rui Wang and Yifan Yang and Zhen Xing and Yuqing Yang and Qi Dai and Lili Qiu and Chong Luo},
year={2026},
eprint={2604.08540},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.08540},
}