The Source Code for T2AV-Compass @ ICML 2026
See the codeObjective evaluation on Linux is now wrapped by two top-level scripts:
setup_objective.shandrun_objective_batch.sh. T2AV-Compass is accepted to ICML 2026.
This flow was validated on a Linux server with a single 4090 GPU. The public interface is repository-relative by default and can be overridden with environment variables when you need a different cache or Conda location.
Default relative layout:
T2AV-Compass/input/t2av-compass/Data/prompts.jsonOutput/.cache/t2av-cache.cache/conda/envsThe objective pipeline now uses a compact default environment layout:
t2av-core: VA, AA, SQ, T-V, T-A, and A-Vt2av-dover: VTt2av-synchformer: DeSynct2av-latentsync: LSUse submodules.
git clone --recurse-submodules https://github.com/NJU-LINK/T2AV-Compass.git
cd T2AV-Compass
If GitHub is slow in your region, you can optionally clone through a mirror instead. Keep the checked-out repository layout unchanged.
Optional environment overrides before setup:
export T2AV_CACHE_ROOT=/path/to/cache-root
export T2AV_CONDA_ROOT=/path/to/conda-root
export T2AV_CORE_ENV=t2av-core
# Default is https://hf-mirror.com for server-side reproducibility in mainland China.
# Override with the official endpoint when it is reachable in your region:
# export HF_ENDPOINT=https://huggingface.co
# optional when GitHub downloads need a mirror
export T2AV_GITHUB_MIRROR_PREFIX=https://your-mirror.example
bash setup_objective.sh
What this script does:
ffmpegt2av-core environment for compatible objective metrics.cache/ by defaultThe script is safe to re-run.
Put videos into input/.
Supported video naming conventions for prompt-linked metrics (T-V, T-A) include:
sample_0001.mp4sample_0002.mp41.mp40001.mp4video_0001.mp4The index field in prompts.json must match the video file index.
Example layout:
T2AV-Compass/
├── input/
│ ├── sample_0001.mp4
│ └── sample_0002.mp4
├── Output/
├── setup_objective.sh
├── run_objective_batch.sh
└── t2av-compass/
└── Data/
└── prompts.json
Minimal t2av-compass/Data/prompts.json example:
[
{
"index": 1,
"prompt": "A person speaking directly to the camera.",
"video_prompt": "A person speaking directly to the camera.",
"audio_prompt": "clean speech from a person speaking indoors",
"speech_prompt": []
},
{
"index": 2,
"prompt": "A person speaking directly to the camera.",
"video_prompt": "A person speaking directly to the camera.",
"audio_prompt": "clean speech from a person speaking indoors",
"speech_prompt": []
}
]
Default paths:
bash run_objective_batch.sh
Custom paths:
bash run_objective_batch.sh /abs/path/to/input /abs/path/to/prompts.json /abs/path/to/output
The batch runs all objective metrics:
VT: video technical qualityVA: video aesthetic qualityAA: audio aesthetic qualitySQ: speech qualityT-V: text-video alignmentT-A: text-audio alignmentA-V: audio-video alignmentDeSync: audio-video synchronization errorLS: lip-sync qualityAfter a successful run, Output/ contains:
video_technical.jsonvideo_aesthetic.jsonaudio_aesthetic.jsonspeech_quality.jsontext_video_alignment.jsontext_audio_alignment.jsonaudio_video_alignment.jsonav_sync.jsonlipsync.jsonevaluation_summary.jsonRun these from the repository root.
bash t2av-compass/scripts/eval_video_technical.sh input Output
bash t2av-compass/scripts/eval_video_aesthetic.sh input Output
bash t2av-compass/scripts/eval_audio_aesthetic.sh input Output
bash t2av-compass/scripts/eval_speech_quality.sh input Output
bash t2av-compass/scripts/eval_text_video_alignment.sh input t2av-compass/Data/prompts.json Output
bash t2av-compass/scripts/eval_text_audio_alignment.sh input t2av-compass/Data/prompts.json Output
bash t2av-compass/scripts/eval_audio_video_alignment.sh input Output
bash t2av-compass/scripts/eval_av_sync.sh input Output
bash t2av-compass/scripts/eval_lipsync.sh input Output
https://hf-mirror.com to keep server-side reproduction stable in mainland China. Set HF_ENDPOINT=https://huggingface.co when the official endpoint is reachable and preferred.T2AV_CACHE_ROOT, T2AV_CONDA_ROOT, T2AV_CONDA_EXE, or rename the shared objective environment with T2AV_CORE_ENV.DeSync run downloads an additional large MotionFormer checkpoint.LS is intended for talking-face videos. For non-talking-face content, the score is not meaningful even if the script finishes.setup_objective.sh or run_objective_batch.sh is supported.视频文件未找到 (index: N): rename the file to match one of the supported index patterns, or fix the index field in prompts.json.ffmpeg not found: run bash setup_objective.sh again on a Debian/Ubuntu-like system with package manager access.HF_ENDPOINT or T2AV_GITHUB_MIRROR_PREFIX before running setup.LS fails on a batch with no visible speaking face: use talking-face videos for this metric.@inproceedings{cao2026t2avcompass,
title = {T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation},
author = {Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Chen, Jialu and Deng, Miao and Wang, Jiahao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng},
booktitle = {International Conference on Machine Learning (ICML)},
year = {2026},
eprint = {2512.21094},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2512.21094},
}
T2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video generation across:
The ICML 2026 version evaluates 15 representative T2AV systems: 7 closed-source end-to-end models, 3 open-source end-to-end models, and 5 composed generation pipelines. The benchmark includes 500 prompts and associated checklist annotations. For subjective evaluation and repository internals, see t2av-compass/README.md.
Jupyter Notebook
64.0%
Python
34.6%
Shell
1.3%
The Source Code for T2AV-Compass @ ICML 2026
See the codeObjective evaluation on Linux is now wrapped by two top-level scripts:
setup_objective.shandrun_objective_batch.sh. T2AV-Compass is accepted to ICML 2026.
This flow was validated on a Linux server with a single 4090 GPU. The public interface is repository-relative by default and can be overridden with environment variables when you need a different cache or Conda location.
Default relative layout:
T2AV-Compass/input/t2av-compass/Data/prompts.jsonOutput/.cache/t2av-cache.cache/conda/envsThe objective pipeline now uses a compact default environment layout:
t2av-core: VA, AA, SQ, T-V, T-A, and A-Vt2av-dover: VTt2av-synchformer: DeSynct2av-latentsync: LSUse submodules.
git clone --recurse-submodules https://github.com/NJU-LINK/T2AV-Compass.git
cd T2AV-Compass
If GitHub is slow in your region, you can optionally clone through a mirror instead. Keep the checked-out repository layout unchanged.
Optional environment overrides before setup:
export T2AV_CACHE_ROOT=/path/to/cache-root
export T2AV_CONDA_ROOT=/path/to/conda-root
export T2AV_CORE_ENV=t2av-core
# Default is https://hf-mirror.com for server-side reproducibility in mainland China.
# Override with the official endpoint when it is reachable in your region:
# export HF_ENDPOINT=https://huggingface.co
# optional when GitHub downloads need a mirror
export T2AV_GITHUB_MIRROR_PREFIX=https://your-mirror.example
bash setup_objective.sh
What this script does:
ffmpegt2av-core environment for compatible objective metrics.cache/ by defaultThe script is safe to re-run.
Put videos into input/.
Supported video naming conventions for prompt-linked metrics (T-V, T-A) include:
sample_0001.mp4sample_0002.mp41.mp40001.mp4video_0001.mp4The index field in prompts.json must match the video file index.
Example layout:
T2AV-Compass/
├── input/
│ ├── sample_0001.mp4
│ └── sample_0002.mp4
├── Output/
├── setup_objective.sh
├── run_objective_batch.sh
└── t2av-compass/
└── Data/
└── prompts.json
Minimal t2av-compass/Data/prompts.json example:
[
{
"index": 1,
"prompt": "A person speaking directly to the camera.",
"video_prompt": "A person speaking directly to the camera.",
"audio_prompt": "clean speech from a person speaking indoors",
"speech_prompt": []
},
{
"index": 2,
"prompt": "A person speaking directly to the camera.",
"video_prompt": "A person speaking directly to the camera.",
"audio_prompt": "clean speech from a person speaking indoors",
"speech_prompt": []
}
]
Default paths:
bash run_objective_batch.sh
Custom paths:
bash run_objective_batch.sh /abs/path/to/input /abs/path/to/prompts.json /abs/path/to/output
The batch runs all objective metrics:
VT: video technical qualityVA: video aesthetic qualityAA: audio aesthetic qualitySQ: speech qualityT-V: text-video alignmentT-A: text-audio alignmentA-V: audio-video alignmentDeSync: audio-video synchronization errorLS: lip-sync qualityAfter a successful run, Output/ contains:
video_technical.jsonvideo_aesthetic.jsonaudio_aesthetic.jsonspeech_quality.jsontext_video_alignment.jsontext_audio_alignment.jsonaudio_video_alignment.jsonav_sync.jsonlipsync.jsonevaluation_summary.jsonRun these from the repository root.
bash t2av-compass/scripts/eval_video_technical.sh input Output
bash t2av-compass/scripts/eval_video_aesthetic.sh input Output
bash t2av-compass/scripts/eval_audio_aesthetic.sh input Output
bash t2av-compass/scripts/eval_speech_quality.sh input Output
bash t2av-compass/scripts/eval_text_video_alignment.sh input t2av-compass/Data/prompts.json Output
bash t2av-compass/scripts/eval_text_audio_alignment.sh input t2av-compass/Data/prompts.json Output
bash t2av-compass/scripts/eval_audio_video_alignment.sh input Output
bash t2av-compass/scripts/eval_av_sync.sh input Output
bash t2av-compass/scripts/eval_lipsync.sh input Output
https://hf-mirror.com to keep server-side reproduction stable in mainland China. Set HF_ENDPOINT=https://huggingface.co when the official endpoint is reachable and preferred.T2AV_CACHE_ROOT, T2AV_CONDA_ROOT, T2AV_CONDA_EXE, or rename the shared objective environment with T2AV_CORE_ENV.DeSync run downloads an additional large MotionFormer checkpoint.LS is intended for talking-face videos. For non-talking-face content, the score is not meaningful even if the script finishes.setup_objective.sh or run_objective_batch.sh is supported.视频文件未找到 (index: N): rename the file to match one of the supported index patterns, or fix the index field in prompts.json.ffmpeg not found: run bash setup_objective.sh again on a Debian/Ubuntu-like system with package manager access.HF_ENDPOINT or T2AV_GITHUB_MIRROR_PREFIX before running setup.LS fails on a batch with no visible speaking face: use talking-face videos for this metric.@inproceedings{cao2026t2avcompass,
title = {T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation},
author = {Cao, Zhe and Wang, Tao and Wang, Jiaming and Wang, Yanghai and Zhang, Yuanxing and Chen, Jialu and Deng, Miao and Wang, Jiahao and Guo, Yubin and Liao, Chenxi and Zhang, Yize and Zhang, Zhaoxiang and Liu, Jiaheng},
booktitle = {International Conference on Machine Learning (ICML)},
year = {2026},
eprint = {2512.21094},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
url = {https://arxiv.org/abs/2512.21094},
}
T2AV-Compass is a unified benchmark for evaluating Text-to-Audio-Video generation across:
The ICML 2026 version evaluates 15 representative T2AV systems: 7 closed-source end-to-end models, 3 open-source end-to-end models, and 5 composed generation pipelines. The benchmark includes 500 prompts and associated checklist annotations. For subjective evaluation and repository internals, see t2av-compass/README.md.
Jupyter Notebook
64.0%
Python
34.6%
Shell
1.3%