📖 A curated list of resources dedicated to avatar.
Jupyter Notebook
61
171 commits
updated Nov 8, 2024
This is a repository for organizing papers, codes and other resources related to the topic of Avatar (talking-face and talking-body).
If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and pull a request.
| Conference | Paper | Affiliation | Codebase | Notes |
|---|---|---|---|---|
| CVPR 2021 | Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors | Tsinghua University | Dataset | |
| ECCV 2022 | HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling | Shanghai Artificial Intelligence Laboratory | Dataset | |
| SIGGRAPH 2023 | AvatarReX: Real-time Expressive Full-body Avatars | Tsinghua University | Dataset | |
| arXiv 2024 | A Survey on 3D Human Avatar Modeling - From Reconstruction to Generation | The University of Hong Kong | ||
| arXiv 2024 | From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations | Meta Reality Labs Research | Code | conversational avatar |
| CVPR 2024 | Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling | Tsinghua Univserity | Code | |
| CVPR 2024 | 4K4D: Real-Time 4D View Synthesis at 4K Resolution | Zhejiang University | Code | real-time synthesis with 3DGS |
| Conference | Paper | Affiliation | Codebase | Notes |
|---|---|---|---|---|
| ICCV 2021 | AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis | University of Science and Technology of China | Code | |
| ECCV 2022 | Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis | Tsinghua University | Code | |
| ICLR 2023 | GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis | Zhejiang University | Code | |
| ICCV 2023 | Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis | Beihang University | Code | |
| arXiv 2023 | GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation | Zhejiang University | Code | |
| CVPR 2024 | SyncTalk: The Devil is in the Synchronization for Talking Head Synthesi | Renmin University of China | Code | |
| ECCV 2024 | TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting | Beihang University | Code |
| Audio-Visual Datasets for Enlish Speakers | ||||||
|---|---|---|---|---|---|---|
| Dataset name | Environment | Year | Resolution | Subject | Duration | Sentence |
| VoxCeleb1 | Wild | 2017 | 360p~720p | 1251 | 352 hours | 100k |
| VoxCeleb2 | Wild | 2018 | 360p~720p | 6112 | 2442 hours | 1128k |
| HDTF | Wild | 2020 | 720p~1080p | 300+ | 15.8 hours | |
| LSP | Wild | 2021 | 720p~1080p | 4 | 18 minutes | 100k |
| Audio-Visual Datasets for Chinese Speakers | ||||||
| Dataset name | Environment | Year | Resolution | Subject | Duration | Sentence |
| CMLR | Lab | 2019 | 11 | 102k | ||
| MAVD | Lab | 2023 | 1920x1080 | 64 | 24 hours | 12k |
| CN-Celeb | Wild | 2020 | 3000 | 1200 hours | ||
| CN-Celeb-AV | Wild | 2023 | 1136 | 660 hours | ||
| CN-CVS | Wild | 2023 | 2500+ | 300+ hours | ||
| Lip-Sync | ||
|---|---|---|
| Metric name | Description | Code/Paper |
| LMD↓ | Mouth landmark distance | |
| LMD↓ | Mouth landmark distance | |
| MA↑ | The Insertion-over-Union (IoU) for the overlap between the predicted mouth area and the ground truth area | |
| Sync↑ | The confidence score from SyncNet (Sync) | wav2lip |
| LSE-C↑ | Lip Sync Error - Confidence | wav2lip |
| LSE-D↓ | Lip Sync Error - Distance | wav2lip |
| Image Quality (identity preserving) | ||
| Metric name | Description | Code/Paper |
| MAE↓ | Mean Absolute Error metric for image | mmagic |
| MSE↓ | Mean Squared Error metric for image | mmagic |
| PSNR↑ | Peak Signal-to-Noise Ratio | mmagic |
| SSIM↑ | Structural similarity for image | mmagic |
| FID↓ | Frchet Inception Distance | mmagic |
| IS↑ | Inception score | mmagic |
| NIQE↓ | Natural Image Quality Evaluator metric | mmagic |
| CSIM↑ | The cosine similarity of identity embedding | InsightFace |
| CPBD↑ | The cumulative probability blur detection | python-cpbd |
| Diversity | ||
| Metric name | Description | Code/Paper |
| Diversity of head motions↑ | A standard deviation of the head motion feature embeddings extracted from the generated frames using Hopenet (Ruiz et al., 2018) is calculated | SadTalker |
| Beat Align Score↑ | The alignment of the audio and generated head motions is calculated in Bailando (Siyao et al., 2022) | SadTalker |
If you are interested in avatar and digital human, we would also like to recommend you to check out other related collections:
167 commits
4 commits
Jupyter Notebook
73.2%
Python
26.8%
📖 A curated list of resources dedicated to avatar.
Jupyter Notebook
61
171 commits
updated Nov 8, 2024
This is a repository for organizing papers, codes and other resources related to the topic of Avatar (talking-face and talking-body).
If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and pull a request.
| Conference | Paper | Affiliation | Codebase | Notes |
|---|---|---|---|---|
| CVPR 2021 | Function4D: Real-time Human Volumetric Capture from Very Sparse Consumer RGBD Sensors | Tsinghua University | Dataset | |
| ECCV 2022 | HuMMan: Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling | Shanghai Artificial Intelligence Laboratory | Dataset | |
| SIGGRAPH 2023 | AvatarReX: Real-time Expressive Full-body Avatars | Tsinghua University | Dataset | |
| arXiv 2024 | A Survey on 3D Human Avatar Modeling - From Reconstruction to Generation | The University of Hong Kong | ||
| arXiv 2024 | From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations | Meta Reality Labs Research | Code | conversational avatar |
| CVPR 2024 | Animatable Gaussians: Learning Pose-dependent Gaussian Maps for High-fidelity Human Avatar Modeling | Tsinghua Univserity | Code | |
| CVPR 2024 | 4K4D: Real-Time 4D View Synthesis at 4K Resolution | Zhejiang University | Code | real-time synthesis with 3DGS |
| Conference | Paper | Affiliation | Codebase | Notes |
|---|---|---|---|---|
| ICCV 2021 | AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis | University of Science and Technology of China | Code | |
| ECCV 2022 | Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head Synthesis | Tsinghua University | Code | |
| ICLR 2023 | GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis | Zhejiang University | Code | |
| ICCV 2023 | Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait Synthesis | Beihang University | Code | |
| arXiv 2023 | GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation | Zhejiang University | Code | |
| CVPR 2024 | SyncTalk: The Devil is in the Synchronization for Talking Head Synthesi | Renmin University of China | Code | |
| ECCV 2024 | TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting | Beihang University | Code |
| Audio-Visual Datasets for Enlish Speakers | ||||||
|---|---|---|---|---|---|---|
| Dataset name | Environment | Year | Resolution | Subject | Duration | Sentence |
| VoxCeleb1 | Wild | 2017 | 360p~720p | 1251 | 352 hours | 100k |
| VoxCeleb2 | Wild | 2018 | 360p~720p | 6112 | 2442 hours | 1128k |
| HDTF | Wild | 2020 | 720p~1080p | 300+ | 15.8 hours | |
| LSP | Wild | 2021 | 720p~1080p | 4 | 18 minutes | 100k |
| Audio-Visual Datasets for Chinese Speakers | ||||||
| Dataset name | Environment | Year | Resolution | Subject | Duration | Sentence |
| CMLR | Lab | 2019 | 11 | 102k | ||
| MAVD | Lab | 2023 | 1920x1080 | 64 | 24 hours | 12k |
| CN-Celeb | Wild | 2020 | 3000 | 1200 hours | ||
| CN-Celeb-AV | Wild | 2023 | 1136 | 660 hours | ||
| CN-CVS | Wild | 2023 | 2500+ | 300+ hours | ||
| Lip-Sync | ||
|---|---|---|
| Metric name | Description | Code/Paper |
| LMD↓ | Mouth landmark distance | |
| LMD↓ | Mouth landmark distance | |
| MA↑ | The Insertion-over-Union (IoU) for the overlap between the predicted mouth area and the ground truth area | |
| Sync↑ | The confidence score from SyncNet (Sync) | wav2lip |
| LSE-C↑ | Lip Sync Error - Confidence | wav2lip |
| LSE-D↓ | Lip Sync Error - Distance | wav2lip |
| Image Quality (identity preserving) | ||
| Metric name | Description | Code/Paper |
| MAE↓ | Mean Absolute Error metric for image | mmagic |
| MSE↓ | Mean Squared Error metric for image | mmagic |
| PSNR↑ | Peak Signal-to-Noise Ratio | mmagic |
| SSIM↑ | Structural similarity for image | mmagic |
| FID↓ | Frchet Inception Distance | mmagic |
| IS↑ | Inception score | mmagic |
| NIQE↓ | Natural Image Quality Evaluator metric | mmagic |
| CSIM↑ | The cosine similarity of identity embedding | InsightFace |
| CPBD↑ | The cumulative probability blur detection | python-cpbd |
| Diversity | ||
| Metric name | Description | Code/Paper |
| Diversity of head motions↑ | A standard deviation of the head motion feature embeddings extracted from the generated frames using Hopenet (Ruiz et al., 2018) is calculated | SadTalker |
| Beat Align Score↑ | The alignment of the audio and generated head motions is calculated in Bailando (Siyao et al., 2022) | SadTalker |
If you are interested in avatar and digital human, we would also like to recommend you to check out other related collections:
167 commits
4 commits
Jupyter Notebook
73.2%
Python
26.8%