The TalkingHeadBench(THB) is a curated dataset designed to support the training and evaluation of deepfake detection models, especially in audio-visual and cross-method generalization scenarios. It includes synthetic videos generated using six modern face animation techniques:
Each video is named using the format:
[image]--[driving_signals]--[generation_method].mp4
image: identity image from FFHQdriving_signals: facial motion and optionally audio from CelebV-HQgeneration_method: the name of the generator usedTalkingHeadBench/
├── fake/
│ ├── [generator_name]/[split]/*.mp4
│ ├── additional_dataset/[generator_name]/*.mp4 # Additional evaluation-only dataset generated using MAGI-1 and Hallo3.
├── audio/
│ ├── fake/*.wav # Extracted from generated fake videos
│ ├── fake_celebvhq/*.wav # Original driving audio from CelebV-HQ
│ ├── ff++/*.wav # Original audio from FaceForensics++ YouTube videos
├── real/
│ ├── real_dataset_split.json # Default split used in this work
│ ├── real_dataset_split_official_ff++.json # Official FF++ split version (for compatibility)
train, val, and testfaceforensics++/original_sequences/youtube/raw/videos) for our source of real videosWe provide two versions of real-data splits:
real/real_dataset_split.jsonreal/real_dataset_split_official_ff++.jsonFake Audio (extracted) (audio/fake/):
Fake Audio (source) (audio/fake_celebvhq/):
Real Audio (audio/ff++/):
faceforensics++/original_sequences/youtube/raw/videosNew audio paths:
Extracted audio from generated videos:
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fakeOriginal source audio (from CelebV-HQ):
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake_celebvhqAdded an alternative real-data split using the official FF++ protocol:
real/real_dataset_split_official_ff++.jsonNotes:
real_dataset_split.json).Updated path:
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/ff++To support reproducibility, we will release the checkpoints for all models used in this work, along with the evaluation code, in a future update.
Please ensure compliance with the original licenses:
If you use this dataset in your research, please cite the relevant original sources (FFHQ, CelebV-HQ, FaceForensics++) and the associated paper.
The TalkingHeadBench(THB) is a curated dataset designed to support the training and evaluation of deepfake detection models, especially in audio-visual and cross-method generalization scenarios. It includes synthetic videos generated using six modern face animation techniques:
Each video is named using the format:
[image]--[driving_signals]--[generation_method].mp4
image: identity image from FFHQdriving_signals: facial motion and optionally audio from CelebV-HQgeneration_method: the name of the generator usedTalkingHeadBench/
├── fake/
│ ├── [generator_name]/[split]/*.mp4
│ ├── additional_dataset/[generator_name]/*.mp4 # Additional evaluation-only dataset generated using MAGI-1 and Hallo3.
├── audio/
│ ├── fake/*.wav # Extracted from generated fake videos
│ ├── fake_celebvhq/*.wav # Original driving audio from CelebV-HQ
│ ├── ff++/*.wav # Original audio from FaceForensics++ YouTube videos
├── real/
│ ├── real_dataset_split.json # Default split used in this work
│ ├── real_dataset_split_official_ff++.json # Official FF++ split version (for compatibility)
train, val, and testfaceforensics++/original_sequences/youtube/raw/videos) for our source of real videosWe provide two versions of real-data splits:
real/real_dataset_split.jsonreal/real_dataset_split_official_ff++.jsonFake Audio (extracted) (audio/fake/):
Fake Audio (source) (audio/fake_celebvhq/):
Real Audio (audio/ff++/):
faceforensics++/original_sequences/youtube/raw/videosNew audio paths:
Extracted audio from generated videos:
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fakeOriginal source audio (from CelebV-HQ):
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake_celebvhqAdded an alternative real-data split using the official FF++ protocol:
real/real_dataset_split_official_ff++.jsonNotes:
real_dataset_split.json).Updated path:
/playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/ff++To support reproducibility, we will release the checkpoints for all models used in this work, along with the evaluation code, in a future update.
Please ensure compliance with the original licenses:
If you use this dataset in your research, please cite the relevant original sources (FFHQ, CelebV-HQ, FaceForensics++) and the associated paper.