Soul-AILab/VividHead

Dataset

Tan Yu*, Qian Qiao*✉, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu ✉

66

27 commits

1 linked in READMEs

updated Feb 12, 2026

See the code

README

SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads

Tan Yu*, Qian Qiao*βœ‰, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu βœ‰

*Equal Contribution βœ‰Corresponding Author

VividHead Dataset

Highlights

  • πŸ”₯ Large-scale, high-quality talking-head dataset with 330K clips and 782 hours of head-cropped videos
  • πŸ”₯ Broad diversity across 15+ languages and a wide age range (0–60+)
  • πŸ”₯ Rich annotations including age, gender, ethnicity, and language
  • πŸ”₯ Unified and standardized processing, with a consistent FPS = 25 and resolution = 512 Γ— 512

ShowCase

🌰 Examples

Dataset Statistics

This dataset exhibits strong diversity across multiple dimensions:

  • Duration: 3s–60s+, bimodal (peaks ~5s, ~10s), mean 8.37s; most clips in 3–15s.
  • Age: 31–45 (432.5h), 19–30 (277.2h), 46–60 (61.3h), 60+ (10.4h), 0–19 (0.2h).
  • Language (Top 10): English (651.4h), Chinese (67.5h), Russian (8.7h), Spanish (7.1h), Portuguese (6.4h), Welsh (5.4h), Hindi (5.3h), German (3.6h), French (3.0h), Korean (2.7h); 15+ languages in total.
  • Gender & ethnicity: Male (552.8h), Female (229.0h); White (506.7h), Asian (113.1h), Latino/Hispanic (56.5h), Middle Eastern (42.9h), Black (36.4h).

Duration

Age group

Language (Top 10)

Gender & ethnicity

Comparison with Other Datasets

DatasetSpeakersFace CropClipsHoursResolutionLanguageAgeEthnicitySource
MEAD60βœ…281.4K39384pEnglish20–35–Lab
HDTF362βœ…10K15.8512p–––Wild
AVSpeech150K❌2.5M4700720p, 1080p–––Wild
Hallo3β€“βœ…101.5K70720p–––Wild
OpenHumanVidβ€“βŒ13.4M16.7K720p–––Wild
TalkVid7,729❌281.4K12441080p, 2160p15 lang.0–60+3Wild
SpeakerVid83K❌5.2M8.7K1080p–––Wild
Ours60Kβœ…330K782512p15 lang.0–60+3Wild

Data Processing Pipeline

Our data processing pipeline is designed to construct a large-scale, high-quality talking-head dataset through systematic preprocessing, filtering, and annotation, ensuring sample uniqueness, temporal consistency, and reliable multi-modal supervision.

Data Preprocessing Stage

  1. Data collection: Aggregates initial content from Web videos and various Open-source videos to build a diverse raw data pool.
  2. Deduplication & Slicing: Employs MD5 hash verification to eliminate redundant content and uses PySceneDetect to divide long videos into coherent clips ranging from 3 to 60+ seconds.
  3. Standardize to 25 FPS: Normalizes all video clips to a uniform frame rate of 25 FPS using FFMPEG to ensure temporal consistency for model training.

Data Filter & Annotation Stage

  1. Face detection & crop: Detects face visibility and crops valid sequences into a centered $512 \times 512$ resolution.
  2. Jump cut detection: Uses optical flow analysis to identify and exclude sequences containing scene discontinuities or abrupt transitions.
  3. Faceless filter: Screens and excludes frames where a detectable face is missing or the head region is improperly framed.
  4. DWpose extraction & hand-filter: Extracts body keypoints and strictly removes clips featuring hand-over-face occlusion to prevent generation artifacts.
  5. Lip-sync: Utilizes the SyncNet model to calculate confidence scores (LSE-C and LSE-D), discarding any samples with poor audio-visual alignment.
  6. Audio feature & attribute labeling: Extracts robust streaming features via Wav2Vec and annotates metadata including language, ethnicity, age, and gender.

πŸ“š Citation

If you find our work useful in your research, please consider citing:

@misc{yu2026soulxflashheadoracleguidedgenerationinfinite,
      title={SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads}, 
      author={Tan Yu and Qian Qiao and Le Shen and Ke Zhou and Jincheng Hu and Dian Sheng and Bo Hu and Haoming Qin and Jun Gao and Changhai Zhou and Shunshun Yin and Siyuan Liu},
      year={2026},
      eprint={2602.07449},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2602.07449}, 
}

License

Our VividHead dataset is released under the CC-BY-4.0 license and is intended for research and non-commercial purposes. The video samples are collected from publicly available datasets.

Soul-AILab/VividHead

Dataset

Tan Yu*, Qian Qiao*✉, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu ✉

66

27 commits

1 linked in READMEs

updated Feb 12, 2026

See the code

README

SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads

Tan Yu*, Qian Qiao*βœ‰, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu βœ‰

*Equal Contribution βœ‰Corresponding Author

VividHead Dataset

Highlights

  • πŸ”₯ Large-scale, high-quality talking-head dataset with 330K clips and 782 hours of head-cropped videos
  • πŸ”₯ Broad diversity across 15+ languages and a wide age range (0–60+)
  • πŸ”₯ Rich annotations including age, gender, ethnicity, and language
  • πŸ”₯ Unified and standardized processing, with a consistent FPS = 25 and resolution = 512 Γ— 512

ShowCase

🌰 Examples

Dataset Statistics

This dataset exhibits strong diversity across multiple dimensions:

  • Duration: 3s–60s+, bimodal (peaks ~5s, ~10s), mean 8.37s; most clips in 3–15s.
  • Age: 31–45 (432.5h), 19–30 (277.2h), 46–60 (61.3h), 60+ (10.4h), 0–19 (0.2h).
  • Language (Top 10): English (651.4h), Chinese (67.5h), Russian (8.7h), Spanish (7.1h), Portuguese (6.4h), Welsh (5.4h), Hindi (5.3h), German (3.6h), French (3.0h), Korean (2.7h); 15+ languages in total.
  • Gender & ethnicity: Male (552.8h), Female (229.0h); White (506.7h), Asian (113.1h), Latino/Hispanic (56.5h), Middle Eastern (42.9h), Black (36.4h).

Duration

Age group

Language (Top 10)

Gender & ethnicity

Comparison with Other Datasets

DatasetSpeakersFace CropClipsHoursResolutionLanguageAgeEthnicitySource
MEAD60βœ…281.4K39384pEnglish20–35–Lab
HDTF362βœ…10K15.8512p–––Wild
AVSpeech150K❌2.5M4700720p, 1080p–––Wild
Hallo3β€“βœ…101.5K70720p–––Wild
OpenHumanVidβ€“βŒ13.4M16.7K720p–––Wild
TalkVid7,729❌281.4K12441080p, 2160p15 lang.0–60+3Wild
SpeakerVid83K❌5.2M8.7K1080p–––Wild
Ours60Kβœ…330K782512p15 lang.0–60+3Wild

Data Processing Pipeline

Our data processing pipeline is designed to construct a large-scale, high-quality talking-head dataset through systematic preprocessing, filtering, and annotation, ensuring sample uniqueness, temporal consistency, and reliable multi-modal supervision.

Data Preprocessing Stage

  1. Data collection: Aggregates initial content from Web videos and various Open-source videos to build a diverse raw data pool.
  2. Deduplication & Slicing: Employs MD5 hash verification to eliminate redundant content and uses PySceneDetect to divide long videos into coherent clips ranging from 3 to 60+ seconds.
  3. Standardize to 25 FPS: Normalizes all video clips to a uniform frame rate of 25 FPS using FFMPEG to ensure temporal consistency for model training.

Data Filter & Annotation Stage

  1. Face detection & crop: Detects face visibility and crops valid sequences into a centered $512 \times 512$ resolution.
  2. Jump cut detection: Uses optical flow analysis to identify and exclude sequences containing scene discontinuities or abrupt transitions.
  3. Faceless filter: Screens and excludes frames where a detectable face is missing or the head region is improperly framed.
  4. DWpose extraction & hand-filter: Extracts body keypoints and strictly removes clips featuring hand-over-face occlusion to prevent generation artifacts.
  5. Lip-sync: Utilizes the SyncNet model to calculate confidence scores (LSE-C and LSE-D), discarding any samples with poor audio-visual alignment.
  6. Audio feature & attribute labeling: Extracts robust streaming features via Wav2Vec and annotates metadata including language, ethnicity, age, and gender.

πŸ“š Citation

If you find our work useful in your research, please consider citing:

@misc{yu2026soulxflashheadoracleguidedgenerationinfinite,
      title={SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads}, 
      author={Tan Yu and Qian Qiao and Le Shen and Ke Zhou and Jincheng Hu and Dian Sheng and Bo Hu and Haoming Qin and Jun Gao and Changhai Zhou and Shunshun Yin and Siyuan Liu},
      year={2026},
      eprint={2602.07449},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2602.07449}, 
}

License

Our VividHead dataset is released under the CC-BY-4.0 license and is intended for research and non-commercial purposes. The video samples are collected from publicly available datasets.