YunjinPark/awesome_talking_face_generation

836

41 commits

updated Nov 19, 2025

See the code

README

Awesome talking face generation

papers & codes

🔍 Survey / Review Papers

YearTitleFocus/TagPaper
2025Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functionsmethods, datasets, metricspaper
2024A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editinggeneration/driving/editing, datasets/metrics, applicationspaper

2025

TitlePaperCodeDatasetsKeywords
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionCVPR(25)paperprojectHDTF, MEAD[emotion
InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoCVPR(25)papercode-Few-shot Learning
Perceptually Accurate 3D Talking Head Generation: NewDefinitions, Speech-Mesh Representation, and Evaluation MetricsCVPR(25)papercodeMEAD-3D, VOCASET3D
Monocular and Generalizable Gaussian Talking Head AnimationCVPR(25)paperprojectHDTF, NeRSemble3D
DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsCVPR(25)papercode-3D, dual

2023

TitlePaperCodeDatasetsKeywords
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorCVPR(23)papercodeBIWI, VOCA
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationCVPR(23)paperHDTFDiffusion
AVFace: Towards Detailed Audio-Visual 4D Face ReconstructionCVPR(23)paperMultiface3D
Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertCVPR(23)papercodeLRS2
LipFormer: High-fidelity and Generalizable Talking Face Generation with A Pre-learned Facial CodebookCVPR(23)paperLRS2, FFHQ
Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentCVPR(23)paperHDTF
Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsCVPR(23)papercodeLRS2, LRS3
High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space LearningCVPR(23)paperMEADemotion
Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksInterSpeech(23)paperMEADemotion
EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationICCV(23)papercode(not yet)emotion
Emotionally Enhanced Talking Face GenerationpapercodeCREMA-Demotion
DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoAAAI(23)papercode
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Priorpapercode3D
GENEFACE: GENERALIZED AND HIGH-FIDELITY AUDIO-DRIVEN 3D TALKING FACE SYNTHESISICLR (23)papercodeNeRF
OPT: ONE-SHOT POSE-CONTROLLABLE TALKING HEAD GENERATIONpaper
LipNeRF: What is the right feature space to lip-sync a NeRF?paperNeRF
Audio-Visual Face ReenactmentWACV (23)papercode
Towards Generating Ultra-High Resolution Talking-Face Videos With Lip SynchronizationWACV (23)paper
StyleTalk: One-shot Talking Head Generation with Controllable Speaking StylesAAAI(23)papercode
DiffTalk: Crafting Diffusion Models for Generalized Talking Head SynthesispaperprojDiffusion
Diffused Heads: Diffusion Models Beat GANs on Talking-Face GenerationpaperprojDiffusion
Speech Driven Video Editing via an Audio-Conditioned Diffusion ModelpapercodeDiffusion
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking StylespaperText-Annotated MEADText

2022

titlepapercodedatasetkeywords
Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priorspaperproj
SPACE: Speech-driven Portrait Animation with Controllable ExpressionICCV(23)paperPose, Emotion
SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationCVPR(23)papercode
Compressing Video Calls using Synthetic Talking HeadsBMVC (22)paperapplication
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelSIGGRAPH (22)paperemotion
Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head SynthesisECCV(22)papercode
Expressive Talking Head Generation with Granular Audio-Visual ControlCVPR(22)paper
Talking Face Generation With Multilingual TTSCVPR(22)papercode-
Deep Learning for Visual Speech Analysis: A Surveypapersurvey
StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANpapercodestylegan
Semantic-Aware Implicit Neural Audio-Driven Video Portrait GenerationECCV(22)papercode(coming soon)NeRF
Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and Manipulationpaper
SyncTalkFace: Talking Face Generation with Precise Lip-syncing via Audio-Lip MemoryAAAI(22)paper(temp)LRW, LRS2, BBC News
DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural RenderingpaperNeRF
Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videospaper
Dynamic Neural Textures: Generating Talking-Face Videos with Continuously Controllable Expressionspaper
DialogueNeRF: Towards Realistic Avatar Face-to-face Conversation Video Generationpaper
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusionpaper
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generationpaper-
AUTOLV: AUTOMATIC LECTURE VIDEO GENERATORpaper
Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglementpaper

2021

titlepapercodedataset
Depth-Aware Generative Adversarial Network for Talking Head Video Generationpapercode
papercode
Parallel and High-Fidelity Text-to-Lip Generationpaper
[Survey]Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis-paper
FaceFormer: Speech-Driven 3D Facial Animation with TransformersCVPR(22)papercode
Voice2Mesh: Cross-Modal 3D Face Model Generation from Voicespapercode
FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningICCVpapercode
Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesispapercode
Audio-Driven Emotional Video PortraitsCVPRpapercodeMEAD, LRW
LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting NormalizationCVPRpaper
Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationCVPRpapercodeVoxCeleb2, LRW
Flow-guided One-shot Talking Face Generation with a High-resolution Audio-visual DatasetCVPRpapercodeHDTF
MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementICCVpapercode(coming soon)
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisICCVpapercode
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationAAAIpapercode(coming soon)Mocap dataset
Visual Speech Enhancement Without A Real Visual Streampaper
Text2Video: Text-driven Talking-head Video Synthesis with Phonetic Dictionarypapercode
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head MotionIJCAIpapercodeVoxCeleb, GRID, LRW
3D-TalkEmo: Learning to Synthesize 3D Emotional Talking Headpaper
AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary PersonpaperVoxCeleb2, Obama

2020

titlepapercodedataset
[Survey]What comprises a good talking-head video generation?: A survey and benchmarkpapercode
One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingCVPR(21)papercode
Speech Driven Talking Face Generation from a Single Image and an Emotion ConditionpapercodeCREMA-D
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The WildACMMMpapercodeLRS2
Talking-head Generation with Rhythmic Head MotionECCVpapercodeCrema, Grid, Voxceleb, Lrs3
MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face GenerationECCVpapercodeVoxCeleb2, AffectNet
Neural voice puppetry:Audio-driven facial reenactmentECCVpaper
Fast Bi-layer Neural Synthesis of One-Shot Realistic Head AvatarsECCVpapercode
HeadGAN:Video-and-Audio-Driven Talking Head SynthesispaperVoxCeleb2
MakeItTalk: Speaker-Aware Talking Head Animationpapercode, codeVoxCeleb2, VCTK
Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose-papercodeImageNet, FaceWarehouse, LRW
Photorealistic Lip Sync with Adversarial Temporal Convolutional Networkspaper
SPEECH-DRIVEN FACIAL ANIMATION USING POLYNOMIAL FUSION OF FEATURESpaperLRW
Animating Face using Disentangled Audio RepresentationsWACVpaper
Everybody’s Talkin’: Let Me Talk as You Wantpaper
Multimodal Inputs Driven Talking Face Generation With Spatial-Temporal Dependencypaper
Speech Driven Talking Face Generation from a Single Image and an Emotion Conditionpaper

2019

titlepapercodedataset
Hierarchical Cross-Modal Talking Face Generation with Dynamic Pixel-Wise LossCVPRpapercodeVGG Face, LRW

datasets

metrics

  • PSNR (peak signal-to-noise ratio)
  • SSIM (structural similarity index measure)
  • LMD (landmark distance error)
  • LRA (lip-reading accuracy) -
  • FID (Fréchet inception distance)
  • LSE-D (Lip Sync Error - Distance)
  • LSE-C (Lip Sync Error - Confidence)
  • LPIPS (Learned Perceptual Image Patch Similarity) -
  • NIQE (Natural Image Quality Evaluator) -

YunjinPark/awesome_talking_face_generation

836

41 commits

updated Nov 19, 2025

See the code

README

Awesome talking face generation

papers & codes

🔍 Survey / Review Papers

YearTitleFocus/TagPaper
2025Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss Functionsmethods, datasets, metricspaper
2024A Comprehensive Taxonomy and Analysis of Talking Head Synthesis: Techniques for Portrait Generation, Driving Mechanisms, and Editinggeneration/driving/editing, datasets/metrics, applicationspaper

2025

TitlePaperCodeDatasetsKeywords
EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionCVPR(25)paperprojectHDTF, MEAD[emotion
InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoCVPR(25)papercode-Few-shot Learning
Perceptually Accurate 3D Talking Head Generation: NewDefinitions, Speech-Mesh Representation, and Evaluation MetricsCVPR(25)papercodeMEAD-3D, VOCASET3D
Monocular and Generalizable Gaussian Talking Head AnimationCVPR(25)paperprojectHDTF, NeRSemble3D
DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsCVPR(25)papercode-3D, dual

2023

TitlePaperCodeDatasetsKeywords
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorCVPR(23)papercodeBIWI, VOCA
DiffTalk: Crafting Diffusion Models for Generalized Audio-Driven Portraits AnimationCVPR(23)paperHDTFDiffusion
AVFace: Towards Detailed Audio-Visual 4D Face ReconstructionCVPR(23)paperMultiface3D
Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertCVPR(23)papercodeLRS2
LipFormer: High-fidelity and Generalizable Talking Face Generation with A Pre-learned Facial CodebookCVPR(23)paperLRS2, FFHQ
Parametric Implicit Face Representation for Audio-Driven Facial ReenactmentCVPR(23)paperHDTF
Identity-Preserving Talking Face Generation with Landmark and Appearance PriorsCVPR(23)papercodeLRS2, LRS3
High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space LearningCVPR(23)paperMEADemotion
Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksInterSpeech(23)paperMEADemotion
EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationICCV(23)papercode(not yet)emotion
Emotionally Enhanced Talking Face GenerationpapercodeCREMA-Demotion
DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoAAAI(23)papercode
CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Priorpapercode3D
GENEFACE: GENERALIZED AND HIGH-FIDELITY AUDIO-DRIVEN 3D TALKING FACE SYNTHESISICLR (23)papercodeNeRF
OPT: ONE-SHOT POSE-CONTROLLABLE TALKING HEAD GENERATIONpaper
LipNeRF: What is the right feature space to lip-sync a NeRF?paperNeRF
Audio-Visual Face ReenactmentWACV (23)papercode
Towards Generating Ultra-High Resolution Talking-Face Videos With Lip SynchronizationWACV (23)paper
StyleTalk: One-shot Talking Head Generation with Controllable Speaking StylesAAAI(23)papercode
DiffTalk: Crafting Diffusion Models for Generalized Talking Head SynthesispaperprojDiffusion
Diffused Heads: Diffusion Models Beat GANs on Talking-Face GenerationpaperprojDiffusion
Speech Driven Video Editing via an Audio-Conditioned Diffusion ModelpapercodeDiffusion
TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking StylespaperText-Annotated MEADText

2022

titlepapercodedatasetkeywords
Talking Head Generation with Probabilistic Audio-to-Visual Diffusion Priorspaperproj
SPACE: Speech-driven Portrait Animation with Controllable ExpressionICCV(23)paperPose, Emotion
SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationCVPR(23)papercode
Compressing Video Calls using Synthetic Talking HeadsBMVC (22)paperapplication
EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelSIGGRAPH (22)paperemotion
Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head SynthesisECCV(22)papercode
Expressive Talking Head Generation with Granular Audio-Visual ControlCVPR(22)paper
Talking Face Generation With Multilingual TTSCVPR(22)papercode-
Deep Learning for Visual Speech Analysis: A Surveypapersurvey
StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANpapercodestylegan
Semantic-Aware Implicit Neural Audio-Driven Video Portrait GenerationECCV(22)papercode(coming soon)NeRF
Cross-Modal Mutual Learning for Audio-Visual Speech Recognition and Manipulationpaper
SyncTalkFace: Talking Face Generation with Precise Lip-syncing via Audio-Lip MemoryAAAI(22)paper(temp)LRW, LRS2, BBC News
DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural RenderingpaperNeRF
Face-Dubbing++: Lip-Synchronous, Voice Preserving Translation of Videospaper
Dynamic Neural Textures: Generating Talking-Face Videos with Continuously Controllable Expressionspaper
DialogueNeRF: Towards Realistic Avatar Face-to-face Conversation Video Generationpaper
Talking Head Generation Driven by Speech-Related Facial Action Units and Audio- Based on Multimodal Representation Fusionpaper
StyleTalker: One-shot Style-based Audio-driven Talking Head Video Generationpaper-
AUTOLV: AUTOMATIC LECTURE VIDEO GENERATORpaper
Synthesizing Photorealistic Virtual Humans Through Cross-modal Disentanglementpaper

2021

titlepapercodedataset
Depth-Aware Generative Adversarial Network for Talking Head Video Generationpapercode
papercode
Parallel and High-Fidelity Text-to-Lip Generationpaper
[Survey]Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth Synthesis-paper
FaceFormer: Speech-Driven 3D Facial Animation with TransformersCVPR(22)papercode
Voice2Mesh: Cross-Modal 3D Face Model Generation from Voicespapercode
FACIAL: Synthesizing Dynamic Talking Face with Implicit Attribute LearningICCVpapercode
Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face Synthesispapercode
Audio-Driven Emotional Video PortraitsCVPRpapercodeMEAD, LRW
LipSync3D: Data-Efficient Learning of Personalized 3D Talking Faces from Video using Pose and Lighting NormalizationCVPRpaper
Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationCVPRpapercodeVoxCeleb2, LRW
Flow-guided One-shot Talking Face Generation with a High-resolution Audio-visual DatasetCVPRpapercodeHDTF
MeshTalk: 3D Face Animation from Speech using Cross-Modality DisentanglementICCVpapercode(coming soon)
AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisICCVpapercode
Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationAAAIpapercode(coming soon)Mocap dataset
Visual Speech Enhancement Without A Real Visual Streampaper
Text2Video: Text-driven Talking-head Video Synthesis with Phonetic Dictionarypapercode
Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head MotionIJCAIpapercodeVoxCeleb, GRID, LRW
3D-TalkEmo: Learning to Synthesize 3D Emotional Talking Headpaper
AnyoneNet: Synchronized Speech and Talking Head Generation for Arbitrary PersonpaperVoxCeleb2, Obama

2020

titlepapercodedataset
[Survey]What comprises a good talking-head video generation?: A survey and benchmarkpapercode
One-Shot Free-View Neural Talking-Head Synthesis for Video ConferencingCVPR(21)papercode
Speech Driven Talking Face Generation from a Single Image and an Emotion ConditionpapercodeCREMA-D
A Lip Sync Expert Is All You Need for Speech to Lip Generation In The WildACMMMpapercodeLRS2
Talking-head Generation with Rhythmic Head MotionECCVpapercodeCrema, Grid, Voxceleb, Lrs3
MEAD: A Large-scale Audio-visual Dataset for Emotional Talking-face GenerationECCVpapercodeVoxCeleb2, AffectNet
Neural voice puppetry:Audio-driven facial reenactmentECCVpaper
Fast Bi-layer Neural Synthesis of One-Shot Realistic Head AvatarsECCVpapercode
HeadGAN:Video-and-Audio-Driven Talking Head SynthesispaperVoxCeleb2
MakeItTalk: Speaker-Aware Talking Head Animationpapercode, codeVoxCeleb2, VCTK
Audio-driven Talking Face Video Generation with Learning-based Personalized Head Pose-papercodeImageNet, FaceWarehouse, LRW
Photorealistic Lip Sync with Adversarial Temporal Convolutional Networkspaper
SPEECH-DRIVEN FACIAL ANIMATION USING POLYNOMIAL FUSION OF FEATURESpaperLRW
Animating Face using Disentangled Audio RepresentationsWACVpaper
Everybody’s Talkin’: Let Me Talk as You Wantpaper
Multimodal Inputs Driven Talking Face Generation With Spatial-Temporal Dependencypaper
Speech Driven Talking Face Generation from a Single Image and an Emotion Conditionpaper

2019

titlepapercodedataset
Hierarchical Cross-Modal Talking Face Generation with Dynamic Pixel-Wise LossCVPRpapercodeVGG Face, LRW

datasets

metrics

  • PSNR (peak signal-to-noise ratio)
  • SSIM (structural similarity index measure)
  • LMD (landmark distance error)
  • LRA (lip-reading accuracy) -
  • FID (Fréchet inception distance)
  • LSE-D (Lip Sync Error - Distance)
  • LSE-C (Lip Sync Error - Confidence)
  • LPIPS (Learned Perceptual Image Patch Similarity) -
  • NIQE (Natural Image Quality Evaluator) -