Project Page | Paper | Evaluation Code
AVE-Compass is a benchmark for evaluating instruction-based audio-video editing abilities. It targets realistic editing requests where visual and audio streams are tightly coupled, such as changing a visible event together with its sound, preserving background ambience during visual edits, or editing speech while maintaining audio-visual consistency.
This dataset release contains the source videos, edit instruction JSON files, and checklist JSON files used by AVE-Compass.

AVE-Compass contains:
The benchmark covers four major editing branches:
AVE-Compass evaluates edited audio-video outputs along complementary dimensions:
The paper also reports automated metrics for cross-modal, visual, and audio quality, including AV Sync, Lip Sync, Video Aesthetic, Subject Consistency, Motion Smoothness, Audio Aesthetic, and Speech Quality.

Models are ranked by Overall Editing Intent, the primary metric of AVE-Compass. Scores are reported on a 0-100 scale, and higher is better.
| Rank | Model | Overall | Video | Audio |
|---|---|---|---|---|
| 1 | AVE-Agent (Wan) | 59.8 | 66.7 | 50.2 |
| 2 | Wan2.7 | 42.4 | 60.1 | 24.8 |
| 3 | HappyHorse | 41.3 | 56.7 | 18.8 |
| 4 | Gemini-Omni* | 38.0 | 56.1 | 10.0 |
| 5 | Seedance | 26.6 | 36.1 | 13.5 |
| 6 | LTX2 | 15.2 | 10.7 | 26.4 |
*Gemini-Omni misses 16 speech edits due to content moderation.
AVE-Compass/
assets/
bench.png # benchmark overview image
evaluation_matrix.png # evaluation matrix rendered from the paper
videos/
*.mp4 # 145 source videos
metadata.jsonl # 196 Dataset Viewer rows
edit_instructions/ # 196 edit instruction JSON files
checklists/ # 196 checklist JSON files
The Dataset Viewer uses videos/metadata.jsonl to display each source video together with its edit instruction and instruction_index.
Each file in edit_instructions/ corresponds to one edit instruction:
{
"video": "example.mp4",
"task": "example",
"instruction_index": 1,
"total_instructions_for_video": 1,
"instruction": {
"category_label": "joint",
"category": "J1",
"operation": "J1.1 New Source Insertion",
"prompt_en": "Add ...",
"audio_label": {
"audio_op": "add",
"sound_type": "event_sfx",
"edit_aspect": "content"
},
"difficulty": {
"d1_object_localization": "hard",
"d1_reason": "...",
"d2_audio_complexity": "complex",
"d3_cross_modal_linkage": "explicit",
"d3_reason": "..."
}
}
}
Important fields:
video: source video filename.instruction_index: 1-based index of the edit instruction for the source video.prompt_en: English edit instruction.category_label: one of joint, speech, video_only, or audio_only.audio_label: structured annotation of the audio-side edit target.difficulty: difficulty annotations for object localization, audio complexity, and cross-modal linkage.Each file in checklists/ corresponds to the edit instruction with the same video and instruction_index. Checklist files contain atomic Yes/No questions for evaluating instruction following and fidelity preservation.
Typical fields include:
video: source video filename.instruction_index: 1-based edit instruction index for the source video.edit_prompt: the edit instruction.edit_category: editing branch.questions: modality-tagged diagnostic questions.Each checklist question includes:
question_iddimensionsubdimensionmodality_tagquestionAVE-Compass was constructed through a human-in-the-loop pipeline. Candidate edit instructions were generated from structured source-video descriptions using modality-aware LLM generators. A critic model filtered out unnatural, infeasible, or ambiguous instructions, and human annotators verified naturalness, executability, and target specificity. The verified instructions were then converted into fine-grained checklist items, followed by human deduplication and refinement.
video field.video and instruction_index.If you use AVE-Compass, please cite:
@article{wen2026avecompass,
title = {AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities},
author = {Wen, Yuqing and Huang, Yukai and Xie, Qianqian and Wu, Jiangtao and Lin, Yibin and Gu, Yikai and Chen, Jialu and Zhang, Yuanxing and Liu, Jiaheng},
year = {2026},
eprint = {2607.24821},
archivePrefix = {arXiv},
primaryClass = {cs.MM},
url = {https://arxiv.org/abs/2607.24821}
}
Project Page | Paper | Evaluation Code
AVE-Compass is a benchmark for evaluating instruction-based audio-video editing abilities. It targets realistic editing requests where visual and audio streams are tightly coupled, such as changing a visible event together with its sound, preserving background ambience during visual edits, or editing speech while maintaining audio-visual consistency.
This dataset release contains the source videos, edit instruction JSON files, and checklist JSON files used by AVE-Compass.

AVE-Compass contains:
The benchmark covers four major editing branches:
AVE-Compass evaluates edited audio-video outputs along complementary dimensions:
The paper also reports automated metrics for cross-modal, visual, and audio quality, including AV Sync, Lip Sync, Video Aesthetic, Subject Consistency, Motion Smoothness, Audio Aesthetic, and Speech Quality.

Models are ranked by Overall Editing Intent, the primary metric of AVE-Compass. Scores are reported on a 0-100 scale, and higher is better.
| Rank | Model | Overall | Video | Audio |
|---|---|---|---|---|
| 1 | AVE-Agent (Wan) | 59.8 | 66.7 | 50.2 |
| 2 | Wan2.7 | 42.4 | 60.1 | 24.8 |
| 3 | HappyHorse | 41.3 | 56.7 | 18.8 |
| 4 | Gemini-Omni* | 38.0 | 56.1 | 10.0 |
| 5 | Seedance | 26.6 | 36.1 | 13.5 |
| 6 | LTX2 | 15.2 | 10.7 | 26.4 |
*Gemini-Omni misses 16 speech edits due to content moderation.
AVE-Compass/
assets/
bench.png # benchmark overview image
evaluation_matrix.png # evaluation matrix rendered from the paper
videos/
*.mp4 # 145 source videos
metadata.jsonl # 196 Dataset Viewer rows
edit_instructions/ # 196 edit instruction JSON files
checklists/ # 196 checklist JSON files
The Dataset Viewer uses videos/metadata.jsonl to display each source video together with its edit instruction and instruction_index.
Each file in edit_instructions/ corresponds to one edit instruction:
{
"video": "example.mp4",
"task": "example",
"instruction_index": 1,
"total_instructions_for_video": 1,
"instruction": {
"category_label": "joint",
"category": "J1",
"operation": "J1.1 New Source Insertion",
"prompt_en": "Add ...",
"audio_label": {
"audio_op": "add",
"sound_type": "event_sfx",
"edit_aspect": "content"
},
"difficulty": {
"d1_object_localization": "hard",
"d1_reason": "...",
"d2_audio_complexity": "complex",
"d3_cross_modal_linkage": "explicit",
"d3_reason": "..."
}
}
}
Important fields:
video: source video filename.instruction_index: 1-based index of the edit instruction for the source video.prompt_en: English edit instruction.category_label: one of joint, speech, video_only, or audio_only.audio_label: structured annotation of the audio-side edit target.difficulty: difficulty annotations for object localization, audio complexity, and cross-modal linkage.Each file in checklists/ corresponds to the edit instruction with the same video and instruction_index. Checklist files contain atomic Yes/No questions for evaluating instruction following and fidelity preservation.
Typical fields include:
video: source video filename.instruction_index: 1-based edit instruction index for the source video.edit_prompt: the edit instruction.edit_category: editing branch.questions: modality-tagged diagnostic questions.Each checklist question includes:
question_iddimensionsubdimensionmodality_tagquestionAVE-Compass was constructed through a human-in-the-loop pipeline. Candidate edit instructions were generated from structured source-video descriptions using modality-aware LLM generators. A critic model filtered out unnatural, infeasible, or ambiguous instructions, and human annotators verified naturalness, executability, and target specificity. The verified instructions were then converted into fine-grained checklist items, followed by human deduplication and refinement.
video field.video and instruction_index.If you use AVE-Compass, please cite:
@article{wen2026avecompass,
title = {AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities},
author = {Wen, Yuqing and Huang, Yukai and Xie, Qianqian and Wu, Jiangtao and Lin, Yibin and Gu, Yikai and Chen, Jialu and Zhang, Yuanxing and Liu, Jiaheng},
year = {2026},
eprint = {2607.24821},
archivePrefix = {arXiv},
primaryClass = {cs.MM},
url = {https://arxiv.org/abs/2607.24821}
}