Tan Yu*, Qian Qiao*✉, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu ✉
66
27 commits
1 linked in READMEs
updated Feb 12, 2026
Tan Yu*, Qian Qiao*β, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu β
*Equal Contribution βCorresponding Author
This dataset exhibits strong diversity across multiple dimensions:
![]() Duration | ![]() Age group |
![]() Language (Top 10) | ![]() Gender & ethnicity |
| Dataset | Speakers | Face Crop | Clips | Hours | Resolution | Language | Age | Ethnicity | Source |
|---|---|---|---|---|---|---|---|---|---|
| MEAD | 60 | β | 281.4K | 39 | 384p | English | 20β35 | β | Lab |
| HDTF | 362 | β | 10K | 15.8 | 512p | β | β | β | Wild |
| AVSpeech | 150K | β | 2.5M | 4700 | 720p, 1080p | β | β | β | Wild |
| Hallo3 | β | β | 101.5K | 70 | 720p | β | β | β | Wild |
| OpenHumanVid | β | β | 13.4M | 16.7K | 720p | β | β | β | Wild |
| TalkVid | 7,729 | β | 281.4K | 1244 | 1080p, 2160p | 15 lang. | 0β60+ | 3 | Wild |
| SpeakerVid | 83K | β | 5.2M | 8.7K | 1080p | β | β | β | Wild |
| Ours | 60K | β | 330K | 782 | 512p | 15 lang. | 0β60+ | 3 | Wild |
Our data processing pipeline is designed to construct a large-scale, high-quality talking-head dataset through systematic preprocessing, filtering, and annotation, ensuring sample uniqueness, temporal consistency, and reliable multi-modal supervision.
If you find our work useful in your research, please consider citing:
@misc{yu2026soulxflashheadoracleguidedgenerationinfinite,
title={SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads},
author={Tan Yu and Qian Qiao and Le Shen and Ke Zhou and Jincheng Hu and Dian Sheng and Bo Hu and Haoming Qin and Jun Gao and Changhai Zhou and Shunshun Yin and Siyuan Liu},
year={2026},
eprint={2602.07449},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.07449},
}
Our VividHead dataset is released under the CC-BY-4.0 license and is intended for research and non-commercial purposes. The video samples are collected from publicly available datasets.
Tan Yu*, Qian Qiao*✉, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu ✉
66
27 commits
1 linked in READMEs
updated Feb 12, 2026
Tan Yu*, Qian Qiao*β, Le Shen*, Ke Zhou, Jincheng Hu, Dian Sheng, Bo Hu, Haoming Qin, Jun Gao, Changhai Zhou, Shunshun Yin, Siyuan Liu β
*Equal Contribution βCorresponding Author
This dataset exhibits strong diversity across multiple dimensions:
![]() Duration | ![]() Age group |
![]() Language (Top 10) | ![]() Gender & ethnicity |
| Dataset | Speakers | Face Crop | Clips | Hours | Resolution | Language | Age | Ethnicity | Source |
|---|---|---|---|---|---|---|---|---|---|
| MEAD | 60 | β | 281.4K | 39 | 384p | English | 20β35 | β | Lab |
| HDTF | 362 | β | 10K | 15.8 | 512p | β | β | β | Wild |
| AVSpeech | 150K | β | 2.5M | 4700 | 720p, 1080p | β | β | β | Wild |
| Hallo3 | β | β | 101.5K | 70 | 720p | β | β | β | Wild |
| OpenHumanVid | β | β | 13.4M | 16.7K | 720p | β | β | β | Wild |
| TalkVid | 7,729 | β | 281.4K | 1244 | 1080p, 2160p | 15 lang. | 0β60+ | 3 | Wild |
| SpeakerVid | 83K | β | 5.2M | 8.7K | 1080p | β | β | β | Wild |
| Ours | 60K | β | 330K | 782 | 512p | 15 lang. | 0β60+ | 3 | Wild |
Our data processing pipeline is designed to construct a large-scale, high-quality talking-head dataset through systematic preprocessing, filtering, and annotation, ensuring sample uniqueness, temporal consistency, and reliable multi-modal supervision.
If you find our work useful in your research, please consider citing:
@misc{yu2026soulxflashheadoracleguidedgenerationinfinite,
title={SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads},
author={Tan Yu and Qian Qiao and Le Shen and Ke Zhou and Jincheng Hu and Dian Sheng and Bo Hu and Haoming Qin and Jun Gao and Changhai Zhou and Shunshun Yin and Siyuan Liu},
year={2026},
eprint={2602.07449},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2602.07449},
}
Our VividHead dataset is released under the CC-BY-4.0 license and is intended for research and non-commercial purposes. The video samples are collected from publicly available datasets.