T2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video (T2AV) generation, targeting not only unimodal quality (video/audio) but also cross-modal alignment & synchronization, complex instruction following, and perceptual realism grounded in physical/common-sense constraints.
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts.
This dataset includes 500 taxonomy-driven prompts and fine-grained checklists for an MLLM-as-a-Judge protocol.
T2AV-Compass is designed for:
data/prompts_with_checklist.json: Core benchmark data (500 prompts + checklists)Each sample is a JSON object with the following fields:
| Field | Type | Description |
|---|---|---|
index | int | Sample ID (1–500) |
source | str | Source tag (e.g., LMArena, RealVideo, VidProM, Kling, Shot2Story) |
subject_matter | str | Theme/genre |
core_subject | list[str] | Subject taxonomy (People/Objects/Animals/…) |
event_scenario | list[str] | Scenario taxonomy (Urban/Living/Natural/Virtual/…) |
sound_type | list[str] | Sound taxonomy (Ambient/Musical/Speech/…) |
camera_movement | list[str] | Camera motion taxonomy (Static/Translation/Zoom/…) |
prompt | str | Integrated prompt (visual + audio + speech) |
video_prompt | str | Video-only prompt |
audio_prompt | str | Non-speech audio prompt (can be empty) |
speech_prompt | list[object] | Structured speech with speaker/description/text |
video | str | Reference video path (optional) |
checklist_info | object | Checklist for MLLM-as-a-Judge |
promptvideo_promptaudio_promptspeech_promptExisting T2AV evaluation benchmarks often necessitate trade-offs between:
T2AV-Compass addresses these limitations with a taxonomy-driven curation pipeline and dual-level evaluation framework.
The dataset is constructed through a three-stage pipeline:
Data Collection: Aggregated raw prompts from high-quality sources including VidProM, Kling AI community, LMArena, and Shot2Story. Semantic clustering with cosine similarity threshold of 0.8 was applied for deduplication.
Prompt Refinement and Alignment: Used Gemini-2.5-Pro to restructure and enrich prompts with visual subjects, motion dynamics, acoustic events, and cinematographic constraints. Manual audit filtered out static scenes or illogical compositions, resulting in 400 complex prompts.
Real-world Video Inversion: Selected 100 diverse, high-fidelity video clips (4–10s) from YouTube with dense captioning via Gemini-2.5-Pro and human-in-the-loop verification.
Each prompt is annotated with:
Research team members with expertise in multimodal generation and evaluation.
| Category | Metric | Description |
|---|---|---|
| Video Quality | VT (Video Technological) | Low-level visual integrity via DOVER++ |
| VA (Video Aesthetic) | High-level perceptual attributes via LAION-Aesthetic V2.5 | |
| Audio Quality | PQ (Perceptual Quality) | Signal fidelity and acoustic realism |
| CU (Content Usefulness) | Semantic validity and information density | |
| Cross-modal Alignment | T-A | Text–Audio alignment via CLAP |
| T-V | Text–Video alignment via VideoCLIP-XL-V2 | |
| A-V | Audio–Video alignment via ImageBind | |
| DS (DeSync) | Temporal synchronization error (lower is better) | |
| LS (LatentSync) | Lip-sync quality for talking-face scenarios |
Instruction Following (IF) - 7 dimensions, 17 sub-dimensions:
Realism - 5 metrics:
Users should:
If you find this work useful, please cite:
@misc{cao2025t2avcompass,
title = {T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation},
author = {Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Chen, Jialu and Deng, Miao and Wang, Jiahao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng},
year = {2025},
eprint = {2512.21094},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2512.21094},
}
NJU-LINK Team, Nanjing University
zhecao@smail.nju.edu.cnliujiaheng@nju.edu.cnT2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video (T2AV) generation, targeting not only unimodal quality (video/audio) but also cross-modal alignment & synchronization, complex instruction following, and perceptual realism grounded in physical/common-sense constraints.
Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped benchmarks that fail to capture cross-modal alignment, instruction following, and perceptual realism under complex prompts.
This dataset includes 500 taxonomy-driven prompts and fine-grained checklists for an MLLM-as-a-Judge protocol.
T2AV-Compass is designed for:
data/prompts_with_checklist.json: Core benchmark data (500 prompts + checklists)Each sample is a JSON object with the following fields:
| Field | Type | Description |
|---|---|---|
index | int | Sample ID (1–500) |
source | str | Source tag (e.g., LMArena, RealVideo, VidProM, Kling, Shot2Story) |
subject_matter | str | Theme/genre |
core_subject | list[str] | Subject taxonomy (People/Objects/Animals/…) |
event_scenario | list[str] | Scenario taxonomy (Urban/Living/Natural/Virtual/…) |
sound_type | list[str] | Sound taxonomy (Ambient/Musical/Speech/…) |
camera_movement | list[str] | Camera motion taxonomy (Static/Translation/Zoom/…) |
prompt | str | Integrated prompt (visual + audio + speech) |
video_prompt | str | Video-only prompt |
audio_prompt | str | Non-speech audio prompt (can be empty) |
speech_prompt | list[object] | Structured speech with speaker/description/text |
video | str | Reference video path (optional) |
checklist_info | object | Checklist for MLLM-as-a-Judge |
promptvideo_promptaudio_promptspeech_promptExisting T2AV evaluation benchmarks often necessitate trade-offs between:
T2AV-Compass addresses these limitations with a taxonomy-driven curation pipeline and dual-level evaluation framework.
The dataset is constructed through a three-stage pipeline:
Data Collection: Aggregated raw prompts from high-quality sources including VidProM, Kling AI community, LMArena, and Shot2Story. Semantic clustering with cosine similarity threshold of 0.8 was applied for deduplication.
Prompt Refinement and Alignment: Used Gemini-2.5-Pro to restructure and enrich prompts with visual subjects, motion dynamics, acoustic events, and cinematographic constraints. Manual audit filtered out static scenes or illogical compositions, resulting in 400 complex prompts.
Real-world Video Inversion: Selected 100 diverse, high-fidelity video clips (4–10s) from YouTube with dense captioning via Gemini-2.5-Pro and human-in-the-loop verification.
Each prompt is annotated with:
Research team members with expertise in multimodal generation and evaluation.
| Category | Metric | Description |
|---|---|---|
| Video Quality | VT (Video Technological) | Low-level visual integrity via DOVER++ |
| VA (Video Aesthetic) | High-level perceptual attributes via LAION-Aesthetic V2.5 | |
| Audio Quality | PQ (Perceptual Quality) | Signal fidelity and acoustic realism |
| CU (Content Usefulness) | Semantic validity and information density | |
| Cross-modal Alignment | T-A | Text–Audio alignment via CLAP |
| T-V | Text–Video alignment via VideoCLIP-XL-V2 | |
| A-V | Audio–Video alignment via ImageBind | |
| DS (DeSync) | Temporal synchronization error (lower is better) | |
| LS (LatentSync) | Lip-sync quality for talking-face scenarios |
Instruction Following (IF) - 7 dimensions, 17 sub-dimensions:
Realism - 5 metrics:
Users should:
If you find this work useful, please cite:
@misc{cao2025t2avcompass,
title = {T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation},
author = {Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Chen, Jialu and Deng, Miao and Wang, Jiahao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng},
year = {2025},
eprint = {2512.21094},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2512.21094},
}
NJU-LINK Team, Nanjing University
zhecao@smail.nju.edu.cnliujiaheng@nju.edu.cn