luchaoqi/TalkingHeadBench

Dataset

TalkingHeadBench

11

20 commits

4 linked in READMEs

updated Apr 6, 2026

See the code

README

TalkingHeadBench

Project page | Paper

Overview

The TalkingHeadBench(THB) is a curated dataset designed to support the training and evaluation of deepfake detection models, especially in audio-visual and cross-method generalization scenarios. It includes synthetic videos generated using six modern face animation techniques:

  • LivePortrait
  • AniPortraitAudio
  • AniPortraitVideo
  • Hallo
  • Hallo2
  • EmoPortrait

Each video is named using the format:

[image]--[driving_signals]--[generation_method].mp4
  • image: identity image from FFHQ
  • driving_signals: facial motion and optionally audio from CelebV-HQ
  • generation_method: the name of the generator used

Dataset Structure

TalkingHeadBench/
├── fake/
│   ├── [generator_name]/[split]/*.mp4
│   ├── additional_dataset/[generator_name]/*.mp4  # Additional evaluation-only dataset generated using MAGI-1 and Hallo3.
├── audio/
│   ├── fake/*.wav               # Extracted from generated fake videos
│   ├── fake_celebvhq/*.wav      # Original driving audio from CelebV-HQ
│   ├── ff++/*.wav               # Original audio from FaceForensics++ YouTube videos
├── real/
│   ├── real_dataset_split.json                     # Default split used in this work
│   ├── real_dataset_split_official_ff++.json       # Official FF++ split version (for compatibility)

  • Each generator has three splits: train, val, and test
  • Training and testing sets come from disjoint identity pools
  • ~300 fake videos per generator are used for training
  • 50 videos per generator are held out as validation
  • Testing uses entirely unseen identities

Real Dataset

  • For training and evaluating purposes, we added real (non-deepfake/true) videos to the process at an approximately 1:1 ratio
  • We used CelebV-HQ and FaceForensics++ (faceforensics++/original_sequences/youtube/raw/videos) for our source of real videos
  • All the real videos are checked against both driving signals and images to ensure no id leakage.

Real Data Splits

We provide two versions of real-data splits:

Default split (used in this work)

  • File: real/real_dataset_split.json
  • Custom split constructed for this benchmark.
  • Ensures identity disjointness across training and testing.

Official FF++ split version

  • File: real/real_dataset_split_official_ff++.json
  • Uses the official FaceForensics++ train/val/test split for FF++ videos.
  • Provided for compatibility with prior work.

Notes

  • Results reported in this paper are based on the default split.
  • Different split definitions are not directly comparable.

❗️Disclaimer

  • The default split used in this work does not follow the official FF++ split. See Real Data Splits for details.

Audio Details

  • Fake Audio (extracted) (audio/fake/):

    • Extracted from generated fake videos for aligned audio–visual evaluation.
  • Fake Audio (source) (audio/fake_celebvhq/):

    • Original driving audio from CelebV-HQ used for generation.
    • May not fully match the final generated videos.
  • Real Audio (audio/ff++/):

    • Original audio from FaceForensics++ YouTube videos.
    • Corresponds to: faceforensics++/original_sequences/youtube/raw/videos
    • 704 audio clips are provided due to public availability.

Notes

  • Earlier versions may contain audio–video misalignment due to preprocessing differences (e.g., FPS mismatch, trimming).
  • Audio has been re-extracted to improve alignment.
  • Results across versions may not be directly comparable.

Applications

  • Audio-visual deepfake detection
  • Modality-specific detection (audio-only or video-only)
  • Cross-generator generalization testing
  • Audio-video consistency evaluation

Update Log

2026-04-06

  • Fixed video–audio synchronization issues in:
    • EmoPortrait
    • AniPortraitVideo
  • Resolved missing-audio problems for some generated videos.
  • Re-extracted audio tracks from all fake videos.

New audio paths:

  • Extracted audio from generated videos:

    • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake
  • Original source audio (from CelebV-HQ):

    • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake_celebvhq
  • Added an alternative real-data split using the official FF++ protocol:

    • real/real_dataset_split_official_ff++.json

Notes:

  • Some generated videos do not use the full original audio clip.
  • This update may affect audio-visual alignment-sensitive models.
  • The official FF++ split is provided for compatibility only; results in this work are based on the default split (real_dataset_split.json).
  • Results across different split definitions are not directly comparable.

2026-03-15

  • Replaced audio for FF++ subset to fix video–audio misalignment issues.

Updated path:

  • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/ff++

Upcoming Release

To support reproducibility, we will release the checkpoints for all models used in this work, along with the evaluation code, in a future update.

Licensing and Attribution

Please ensure compliance with the original licenses:

Citation

If you use this dataset in your research, please cite the relevant original sources (FFHQ, CelebV-HQ, FaceForensics++) and the associated paper.

luchaoqi/TalkingHeadBench

Dataset

TalkingHeadBench

11

20 commits

4 linked in READMEs

updated Apr 6, 2026

See the code

README

TalkingHeadBench

Project page | Paper

Overview

The TalkingHeadBench(THB) is a curated dataset designed to support the training and evaluation of deepfake detection models, especially in audio-visual and cross-method generalization scenarios. It includes synthetic videos generated using six modern face animation techniques:

  • LivePortrait
  • AniPortraitAudio
  • AniPortraitVideo
  • Hallo
  • Hallo2
  • EmoPortrait

Each video is named using the format:

[image]--[driving_signals]--[generation_method].mp4
  • image: identity image from FFHQ
  • driving_signals: facial motion and optionally audio from CelebV-HQ
  • generation_method: the name of the generator used

Dataset Structure

TalkingHeadBench/
├── fake/
│   ├── [generator_name]/[split]/*.mp4
│   ├── additional_dataset/[generator_name]/*.mp4  # Additional evaluation-only dataset generated using MAGI-1 and Hallo3.
├── audio/
│   ├── fake/*.wav               # Extracted from generated fake videos
│   ├── fake_celebvhq/*.wav      # Original driving audio from CelebV-HQ
│   ├── ff++/*.wav               # Original audio from FaceForensics++ YouTube videos
├── real/
│   ├── real_dataset_split.json                     # Default split used in this work
│   ├── real_dataset_split_official_ff++.json       # Official FF++ split version (for compatibility)

  • Each generator has three splits: train, val, and test
  • Training and testing sets come from disjoint identity pools
  • ~300 fake videos per generator are used for training
  • 50 videos per generator are held out as validation
  • Testing uses entirely unseen identities

Real Dataset

  • For training and evaluating purposes, we added real (non-deepfake/true) videos to the process at an approximately 1:1 ratio
  • We used CelebV-HQ and FaceForensics++ (faceforensics++/original_sequences/youtube/raw/videos) for our source of real videos
  • All the real videos are checked against both driving signals and images to ensure no id leakage.

Real Data Splits

We provide two versions of real-data splits:

Default split (used in this work)

  • File: real/real_dataset_split.json
  • Custom split constructed for this benchmark.
  • Ensures identity disjointness across training and testing.

Official FF++ split version

  • File: real/real_dataset_split_official_ff++.json
  • Uses the official FaceForensics++ train/val/test split for FF++ videos.
  • Provided for compatibility with prior work.

Notes

  • Results reported in this paper are based on the default split.
  • Different split definitions are not directly comparable.

❗️Disclaimer

  • The default split used in this work does not follow the official FF++ split. See Real Data Splits for details.

Audio Details

  • Fake Audio (extracted) (audio/fake/):

    • Extracted from generated fake videos for aligned audio–visual evaluation.
  • Fake Audio (source) (audio/fake_celebvhq/):

    • Original driving audio from CelebV-HQ used for generation.
    • May not fully match the final generated videos.
  • Real Audio (audio/ff++/):

    • Original audio from FaceForensics++ YouTube videos.
    • Corresponds to: faceforensics++/original_sequences/youtube/raw/videos
    • 704 audio clips are provided due to public availability.

Notes

  • Earlier versions may contain audio–video misalignment due to preprocessing differences (e.g., FPS mismatch, trimming).
  • Audio has been re-extracted to improve alignment.
  • Results across versions may not be directly comparable.

Applications

  • Audio-visual deepfake detection
  • Modality-specific detection (audio-only or video-only)
  • Cross-generator generalization testing
  • Audio-video consistency evaluation

Update Log

2026-04-06

  • Fixed video–audio synchronization issues in:
    • EmoPortrait
    • AniPortraitVideo
  • Resolved missing-audio problems for some generated videos.
  • Re-extracted audio tracks from all fake videos.

New audio paths:

  • Extracted audio from generated videos:

    • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake
  • Original source audio (from CelebV-HQ):

    • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/fake_celebvhq
  • Added an alternative real-data split using the official FF++ protocol:

    • real/real_dataset_split_official_ff++.json

Notes:

  • Some generated videos do not use the full original audio clip.
  • This update may affect audio-visual alignment-sensitive models.
  • The official FF++ split is provided for compatibility only; results in this work are based on the default split (real_dataset_split.json).
  • Results across different split definitions are not directly comparable.

2026-03-15

  • Replaced audio for FF++ subset to fix video–audio misalignment issues.

Updated path:

  • /playpen-nas-ssd3/anaxxq/TalkingHeadBench/audio/ff++

Upcoming Release

To support reproducibility, we will release the checkpoints for all models used in this work, along with the evaluation code, in a future update.

Licensing and Attribution

Please ensure compliance with the original licenses:

Citation

If you use this dataset in your research, please cite the relevant original sources (FFHQ, CelebV-HQ, FaceForensics++) and the associated paper.