Jason-cs18/awesome-avatar

📖 A curated list of resources dedicated to avatar.

Jupyter Notebook

61

171 commits

updated Nov 8, 2024

See the code

README

awesome-avatar

This is a repository for organizing papers, codes and other resources related to the topic of Avatar (talking-face and talking-body).

🔆 This project is still on-going, pull requests are welcomed!!

If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and pull a request.

News

  • 2024.09.07: add ASR and TTS tool
  • 2024.08.24: add backgrounds for image/video generations
  • 2024.08.24: re-organize paper list with table formating
  • 2024.08.24: add works about full-body avatar synthesis

TO DO LIST

  • Main paper list
  • Researchers list
  • Toolbox for avatar
  • Add paper link
  • Add paper notes
  • Add codes if have
  • Add project page if have
  • Datasets and metrics
  • Related links

Researchers and labs

  1. NVIDIA Research
  2. Aliaksandr Siarohin @ Snap Research
  3. Ziwei Liu @ Nanyang Technological University
  4. Xiaodong Cun @ Tencent AI Lab:
  1. Max Planck Institute for Informatics:

Papers

Image and video generation

3D Avatar (face+body)

2D talking-face synthesis

ConferencePaperAffiliationCodebaseTraining CodeNotes
MM 2020Wav2Lip: Accurately Lip-sync Videos to Any SpeechThe International Institute of Islamic Thought (IIIT), IndiaCode Github stars Github forks✅most accurate lip-sync model, bad video quality 96*96, pre-trained on ~180 hours video data from LRS2
MM 2021Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisTsinghua UniversityCode, Github stars Github forks
CVPR 2021Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationThe Chinese University of Hong KongCode Github stars Github forkscontrastive learning on audio-lip
ICCV 2021PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingPeking UniversityCode Github stars Github forks
ECCV 2022StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANTsinghua UniversityCode Github stars Github forksHigh-fidenity synthesis via StyleGAN
SIGGRAPH Asia 2022VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildXidian UniversityCode Github stars Github forks
AAAI 2023DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoVirtual Human Group, Netease Fuxi AI LabCodeGithub stars Github forks✅accurate lip-sync and high-quality synthesis (256*256)
CVPR 2023SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationXi'an Jiaotong UniversityCode Github stars Github forks, Note
arXiv 2023DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic ModelsTsinghua UniversityCode, Github stars Github forksdiffusion
Tencent TMElyralabMuseTalk: Real-Time High Quality Lip Synchorization with Latent Space Inpainting Github stars Github forks
arXiv 2024LivePortrait: Efficient Portrait Animation with Stitching and Retargeting ControlKuaishou TechnologyCode Github stars Github forksface reenactment with micro-expression
arXiv 2024EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsAnt GroupCode Github stars Github forksaccurate lip-sync on Chinese speakers, diffusion, pre-trained on 540 hours cleaned video data (collected from internet)
arXiv 2024Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image AnimationFudan UniversityCode, Github stars Github forks✅accurate lip-sync, diffusion, pre-trained on 264 hours of cleaned video data (155 hours from internet and 9 hours from HDTF)
[arXiv 2024]Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion DependencyZhejiang University and ByteDanceexpressive animation driven by audio only, pre-trained on 160 hours of cleaned video data (collected from internet)

3D talking-face synthesis

Talking-body synthesis

Pose2video

Datasets

Talking-face

Audio-Visual Datasets for Enlish Speakers
Dataset nameEnvironmentYearResolutionSubjectDurationSentence
VoxCeleb1Wild2017360p~720p1251352 hours100k
VoxCeleb2Wild2018360p~720p61122442 hours1128k
HDTFWild2020720p~1080p300+15.8 hours
LSPWild2021720p~1080p418 minutes100k
Audio-Visual Datasets for Chinese Speakers
Dataset nameEnvironmentYearResolutionSubjectDurationSentence
CMLRLab201911102k
MAVDLab20231920x10806424 hours12k
CN-CelebWild202030001200 hours
CN-Celeb-AVWild20231136660 hours
CN-CVSWild20232500+300+ hours

Metrics

Talking-face

Lip-Sync
Metric nameDescriptionCode/Paper
LMD↓Mouth landmark distance
LMD↓Mouth landmark distance
MA↑The Insertion-over-Union (IoU) for the overlap between the predicted mouth area and the ground truth area
Sync↑The confidence score from SyncNet (Sync)wav2lip
LSE-C↑Lip Sync Error - Confidencewav2lip
LSE-D↓Lip Sync Error - Distancewav2lip
Image Quality (identity preserving)
Metric nameDescriptionCode/Paper
MAE↓Mean Absolute Error metric for imagemmagic
MSE↓Mean Squared Error metric for imagemmagic
PSNR↑Peak Signal-to-Noise Ratiommagic
SSIM↑Structural similarity for imagemmagic
FID↓Frchet Inception Distancemmagic
IS↑Inception score mmagic
NIQE↓Natural Image Quality Evaluator metricmmagic
CSIM↑The cosine similarity of identity embeddingInsightFace
CPBD↑The cumulative probability blur detectionpython-cpbd
Diversity
Metric nameDescriptionCode/Paper
Diversity of head motions↑A standard deviation of the head motion feature embeddings extracted from the generated frames using Hopenet (Ruiz et al., 2018) is calculatedSadTalker
Beat Align Score↑The alignment of the audio and generated head motions is calculated in Bailando (Siyao et al., 2022)SadTalker

Toolbox

  1. A general toolbox for AIGC, including common metrics and models https://github.com/open-mmlab/mmagic
  2. face3d: Python tools for processing 3D face https://github.com/yfeng95/face3d
  3. 3DMM model fitting using Pytorch https://github.com/ascust/3DMM-Fitting-Pytorch
  4. OpenFace: a facial behavior analysis toolkit https://github.com/TadasBaltrusaitis/OpenFace
  5. autocrop: Automatically detects and crops faces from batches of pictures https://github.com/leblancfg/autocrop
  6. OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation https://github.com/CMU-Perceptual-Computing-Lab/openpose
  7. GFPGAN: Practical Algorithm for Real-world Face Restoration https://github.com/TencentARC/GFPGAN
  8. CodeFormer: Robust Blind Face Restoration https://github.com/sczhou/CodeFormer
  9. metahuman-stream: Real time interactive streaming digital human https://github.com/lipku/metahuman-stream
  10. EasyVolcap: a PyTorch library for accelerating neural volumetric video research https://github.com/zju3dv/EasyVolcap
  11. 3D Model in gradio https://www.gradio.app/guides/how-to-use-3D-model-component

Automatic Speech Recognition (ASR)

  1. BELLE-2/Belle-whisper-large-v3-zh https://huggingface.co/BELLE-2/Belle-whisper-large-v3-zh
  2. SenseVoice (multilingual) https://github.com/FunAudioLLM/SenseVoice 👍👍

Text to Speech (TTS)

  1. CosyVoice, Alibaba Tongyi SpeechTeam https://github.com/FunAudioLLM/CosyVoice 👍👍
  2. FireRedTTS, FireReadTeam https://github.com/FireRedTeam/FireRedTTS
  3. GPT-SoVITS https://github.com/RVC-Boss/GPT-SoVITS?tab=readme-ov-file

Speech to Speech (GPT4-o)

  1. Mini-Omni, Tsinghua University https://github.com/gpt-omni/mini-omni
  2. Speech To Speech, HuggingFace https://github.com/huggingface/speech-to-speech

If you are interested in avatar and digital human, we would also like to recommend you to check out other related collections:

avatar
awesome-list
co-speech-gesture
deep-generative-models
digital-human
pose2img
talking-head

Contributors

Jason-cs18

167 commits

SuperGoodGame

4 commits

Jason-cs18/awesome-avatar

📖 A curated list of resources dedicated to avatar.

Jupyter Notebook

61

171 commits

updated Nov 8, 2024

See the code

README

awesome-avatar

This is a repository for organizing papers, codes and other resources related to the topic of Avatar (talking-face and talking-body).

🔆 This project is still on-going, pull requests are welcomed!!

If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and pull a request.

News

  • 2024.09.07: add ASR and TTS tool
  • 2024.08.24: add backgrounds for image/video generations
  • 2024.08.24: re-organize paper list with table formating
  • 2024.08.24: add works about full-body avatar synthesis

TO DO LIST

  • Main paper list
  • Researchers list
  • Toolbox for avatar
  • Add paper link
  • Add paper notes
  • Add codes if have
  • Add project page if have
  • Datasets and metrics
  • Related links

Researchers and labs

  1. NVIDIA Research
  2. Aliaksandr Siarohin @ Snap Research
  3. Ziwei Liu @ Nanyang Technological University
  4. Xiaodong Cun @ Tencent AI Lab:
  1. Max Planck Institute for Informatics:

Papers

Image and video generation

3D Avatar (face+body)

2D talking-face synthesis

ConferencePaperAffiliationCodebaseTraining CodeNotes
MM 2020Wav2Lip: Accurately Lip-sync Videos to Any SpeechThe International Institute of Islamic Thought (IIIT), IndiaCode Github stars Github forks✅most accurate lip-sync model, bad video quality 96*96, pre-trained on ~180 hours video data from LRS2
MM 2021Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisTsinghua UniversityCode, Github stars Github forks
CVPR 2021Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationThe Chinese University of Hong KongCode Github stars Github forkscontrastive learning on audio-lip
ICCV 2021PIRenderer: Controllable Portrait Image Generation via Semantic Neural RenderingPeking UniversityCode Github stars Github forks
ECCV 2022StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANTsinghua UniversityCode Github stars Github forksHigh-fidenity synthesis via StyleGAN
SIGGRAPH Asia 2022VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildXidian UniversityCode Github stars Github forks
AAAI 2023DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoVirtual Human Group, Netease Fuxi AI LabCodeGithub stars Github forks✅accurate lip-sync and high-quality synthesis (256*256)
CVPR 2023SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationXi'an Jiaotong UniversityCode Github stars Github forks, Note
arXiv 2023DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic ModelsTsinghua UniversityCode, Github stars Github forksdiffusion
Tencent TMElyralabMuseTalk: Real-Time High Quality Lip Synchorization with Latent Space Inpainting Github stars Github forks
arXiv 2024LivePortrait: Efficient Portrait Animation with Stitching and Retargeting ControlKuaishou TechnologyCode Github stars Github forksface reenactment with micro-expression
arXiv 2024EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditionsAnt GroupCode Github stars Github forksaccurate lip-sync on Chinese speakers, diffusion, pre-trained on 540 hours cleaned video data (collected from internet)
arXiv 2024Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image AnimationFudan UniversityCode, Github stars Github forks✅accurate lip-sync, diffusion, pre-trained on 264 hours of cleaned video data (155 hours from internet and 9 hours from HDTF)
[arXiv 2024]Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion DependencyZhejiang University and ByteDanceexpressive animation driven by audio only, pre-trained on 160 hours of cleaned video data (collected from internet)

3D talking-face synthesis

Talking-body synthesis

Pose2video

Datasets

Talking-face

Audio-Visual Datasets for Enlish Speakers
Dataset nameEnvironmentYearResolutionSubjectDurationSentence
VoxCeleb1Wild2017360p~720p1251352 hours100k
VoxCeleb2Wild2018360p~720p61122442 hours1128k
HDTFWild2020720p~1080p300+15.8 hours
LSPWild2021720p~1080p418 minutes100k
Audio-Visual Datasets for Chinese Speakers
Dataset nameEnvironmentYearResolutionSubjectDurationSentence
CMLRLab201911102k
MAVDLab20231920x10806424 hours12k
CN-CelebWild202030001200 hours
CN-Celeb-AVWild20231136660 hours
CN-CVSWild20232500+300+ hours

Metrics

Talking-face

Lip-Sync
Metric nameDescriptionCode/Paper
LMD↓Mouth landmark distance
LMD↓Mouth landmark distance
MA↑The Insertion-over-Union (IoU) for the overlap between the predicted mouth area and the ground truth area
Sync↑The confidence score from SyncNet (Sync)wav2lip
LSE-C↑Lip Sync Error - Confidencewav2lip
LSE-D↓Lip Sync Error - Distancewav2lip
Image Quality (identity preserving)
Metric nameDescriptionCode/Paper
MAE↓Mean Absolute Error metric for imagemmagic
MSE↓Mean Squared Error metric for imagemmagic
PSNR↑Peak Signal-to-Noise Ratiommagic
SSIM↑Structural similarity for imagemmagic
FID↓Frchet Inception Distancemmagic
IS↑Inception score mmagic
NIQE↓Natural Image Quality Evaluator metricmmagic
CSIM↑The cosine similarity of identity embeddingInsightFace
CPBD↑The cumulative probability blur detectionpython-cpbd
Diversity
Metric nameDescriptionCode/Paper
Diversity of head motions↑A standard deviation of the head motion feature embeddings extracted from the generated frames using Hopenet (Ruiz et al., 2018) is calculatedSadTalker
Beat Align Score↑The alignment of the audio and generated head motions is calculated in Bailando (Siyao et al., 2022)SadTalker

Toolbox

  1. A general toolbox for AIGC, including common metrics and models https://github.com/open-mmlab/mmagic
  2. face3d: Python tools for processing 3D face https://github.com/yfeng95/face3d
  3. 3DMM model fitting using Pytorch https://github.com/ascust/3DMM-Fitting-Pytorch
  4. OpenFace: a facial behavior analysis toolkit https://github.com/TadasBaltrusaitis/OpenFace
  5. autocrop: Automatically detects and crops faces from batches of pictures https://github.com/leblancfg/autocrop
  6. OpenPose: Real-time multi-person keypoint detection library for body, face, hands, and foot estimation https://github.com/CMU-Perceptual-Computing-Lab/openpose
  7. GFPGAN: Practical Algorithm for Real-world Face Restoration https://github.com/TencentARC/GFPGAN
  8. CodeFormer: Robust Blind Face Restoration https://github.com/sczhou/CodeFormer
  9. metahuman-stream: Real time interactive streaming digital human https://github.com/lipku/metahuman-stream
  10. EasyVolcap: a PyTorch library for accelerating neural volumetric video research https://github.com/zju3dv/EasyVolcap
  11. 3D Model in gradio https://www.gradio.app/guides/how-to-use-3D-model-component

Automatic Speech Recognition (ASR)

  1. BELLE-2/Belle-whisper-large-v3-zh https://huggingface.co/BELLE-2/Belle-whisper-large-v3-zh
  2. SenseVoice (multilingual) https://github.com/FunAudioLLM/SenseVoice 👍👍

Text to Speech (TTS)

  1. CosyVoice, Alibaba Tongyi SpeechTeam https://github.com/FunAudioLLM/CosyVoice 👍👍
  2. FireRedTTS, FireReadTeam https://github.com/FireRedTeam/FireRedTTS
  3. GPT-SoVITS https://github.com/RVC-Boss/GPT-SoVITS?tab=readme-ov-file

Speech to Speech (GPT4-o)

  1. Mini-Omni, Tsinghua University https://github.com/gpt-omni/mini-omni
  2. Speech To Speech, HuggingFace https://github.com/huggingface/speech-to-speech

If you are interested in avatar and digital human, we would also like to recommend you to check out other related collections:

avatar
awesome-list
co-speech-gesture
deep-generative-models
digital-human
pose2img
talking-head

Contributors

Jason-cs18

167 commits

SuperGoodGame

4 commits

Languages

Jupyter Notebook

73.2%

Python

26.8%