Kedreamix/Awesome-Talking-Head-Synthesis

💬 An extensive collection of exceptional resources dedicated to the captivating world of talking face synthesis! ⭐ If you find this repo useful, please give it a star! 🤩

Python

1,555

131 commits

updated Sep 24, 2026

See the code

README

Awesome-Talking-Head-Synthesis

🌐 Website: https://kedreamix.github.io/Awesome-Talking-Head-Synthesis/

The README below starts with runnable open-source systems that do not have a paper, then the complete paper list. For search, filters, bilingual browsing, and category overview, please use the website.

This repository organizes papers, codes, datasets and project pages for Talking Head Synthesis, covering audio-driven avatars, portrait animation, NeRF / 3D heads, 3D Gaussian Splatting, conversational agents, talking body, and more. 👤

Papers for Talking Head Synthesis, released codes collections. ✍️

Most papers are linked to PDFs on "arXiv" or journal/conference websites 📚. However, some papers require an academic license to view 🔐.

🔆 This project Awesome-Talking-Head-Synthesis is ongoing - pull requests are welcome! If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and submit a PR. You can also open an issue or contact me directly via email. 📩

⭐ If you find this repo useful, please give it a star! 🤩

2026.09 Update 📆

The paper list has grown large enough that browsing only in the README is no longer convenient, so I launched an interactive website:

👉 https://kedreamix.github.io/Awesome-Talking-Head-Synthesis/

You can search papers, browse runnable open-source projects, filter by category and year, switch between English / 中文, and use dark mode. If you find a missing paper or project, newly released code, an accepted venue, or want to suggest a new category, please submit it through the website or GitHub Issues. 🙌

I also added Open-source Projects for runnable talking-head systems that are useful to try, but do not have a paper (for example Linly-Talker and NanoAvatar). SadTalker, MuseTalk, Hallo, and similar works stay in the paper tables.

⭐ If the site helps you, please star the repo — it really motivates continued updates!

2023.12 Update 📆

Thank you to https://github.com/Curated-Awesome-Lists/awesome-ai-talking-heads, I have added some of its contents, such as Tools & Software and Slides & Presentations. 🙏 I hope this will be helpful.😊

If you have any feedback or ideas on extending this aggregated resource, please open an issue or PR - community contributions are vital to advancing this shared knowledge. 🤝

Let's keep pushing forward to recreate ever more realistic digital human faces! 💪 We've come so far but still have a long way to go. With continued research 🔬 and collaboration, I'm sure we'll get there! 🤗

Please feel free to star ⭐ and share this repo if you find it a valuable resource. Your support helps motivate me to keep maintaining and improving it. 🥰 Let me know if you have any other questions!


Open-source Projects

Runnable talking-head systems, apps, and integration frameworks that do not have a paper listing. Research papers with official code stay in the sections below (Audio-driven, Conversational, and so on).

YearProjectCodeResourcesDescription
2026NanoAvatarCodeWeights · APKsOn-device audio-driven talking avatars for Android and local NVIDIA GPUs, with offline APKs.
2026Linly-Talker-StreamCodeFull-duplex, low-latency conversational digital human built on a real-time WebRTC streaming pipeline.
2026CyberVerseCodeSiteSelf-hosted real-time digital-human agent platform with WebRTC, memory, tools, RAG, and optional avatar video.
2025OpenAvatarChatCodeDocs · DemoModular interactive avatar chat with replaceable ASR, LLM, TTS, and avatar backends.
2025LiteAvatarCodeGalleryReal-time CPU audio-to-face 2D chat avatar and an avatar backend for OpenAvatarChat.
2024Duix-AvatarCodeSiteOffline avatar toolkit for appearance and voice cloning plus text- or audio-driven video generation.
2024Ultralight-Digital-HumanCodeFeatherTalkMobile-friendly real-time 2D digital human, with FeatherTalk as its lighter successor.
2024DH_liveCodeMatesXLightweight real-time 2D digital human for web and mobile, followed by the multi-platform MatesX engine.
2023Linly-TalkerCodeWeights · PageConversational digital human WebUI combining LLM, ASR, TTS, voice cloning, and multiple talking-head backends.
2023LiveTalkingCodeSiteReal-time interactive streaming digital human engine supporting Wav2Lip, MuseTalk, ER-NeRF, and other backends.
2023TalkingHeadCodeBrowser JavaScript class for real-time lip-sync using full-body 3D avatars.
2022VU-VRMCodeDemoBrowser-based real-time lip-sync VRM avatar driven by a microphone without a webcam.

Datasets

dataset描述

YearDatasetConference/JournalDownload LinkDescription
2026The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction DatasetArXiv 2026DownloadLarge-scale disentangled human-evaluation challenge for speech-driven co-speech gesture generation (GENEA 2026).
2026REACT 2026: The Fourth Multiple Appropriate Facial Reaction Generation Challenge: Personalised MAFRG and Appropriate EEG Reaction PredictionArXiv 2026DownloadFourth Multiple Appropriate Facial Reaction Generation Challenge for personalized listener reactions in dyadic interaction.
2026MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video GenerationArXiv 2026DownloadBenchmark diagnosing cinematic-expressiveness failure modes in multi-talker audio-visual generation.
2026AVAPrintDB: Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and BenchmarkArXiv 2026N/AMulti-generator benchmark for talking-head avatar fingerprinting and photorealistic avatar generation evaluation.
2026A Near-Raw Talking-Head Video Dataset for Various Computer Vision TasksArXiv 2026N/ANear-raw talking-head video dataset for computer vision tasks including detection, recognition, and generation.
2026Face-to-Face: A Video Dataset for Multi-Person Interaction ModelingArXiv 2026Download70-hour, 14k-clip dataset of two-person talk-show exchanges for multi-person interaction modeling.
2026SFQAArXivN/AA dataset for singing face generation quality assessment with 5,184 videos generated from 100 photographs and 36 music clips using 12 generation methods.
2025TalkCutsArXiv 2025DownloadA large-scale dataset with 164k clips totaling over 500 hours of human speech videos featuring diverse camera shots and detailed annotations including textual descriptions, 2D keypoints, and 3D SMPL-X motions for multi-shot speech video generation.
2025EmojiBench++IJCV 2025DownloadA comprehensive benchmark for portrait animation comprising diverse portraits, driving videos, and landmark sequences.
2025Multi-human InteractiveArXiv 2025Download12 hours of high-res footage with 2-4 speakers, fine-grained body pose and speech interaction annotations.
2025THQA-10KArXiv 2025DownloadLargest AGTH quality assessment dataset with 10,457 samples from 12 T2I models and 14 talkers.
2025SpeakerVid-5MArXivDownloadLarge-scale dataset with 5.2M video clips (8,743 hours) for audio-visual dyadic interactive virtual human generation, covering monadic talking, listening, and dyadic conversations, with pre-training and SFT subsets.
2025TalkingHeadBenchWACV 2026DownloadComprehensive benchmark for talking-head deepfake detection with multi-model generators.
2025Motion-X++ArXiv 2025N/A19.5M 3D whole-body pose annotations covering 120.5K motion sequences with 80.8K RGB videos.
2024GLCF (MSTF)ArXiv 2024N/AFirst large-scale multi-scenario talking face dataset with 22 audio/video forgery techniques.
2024SAVEEArXiv 2024Download480 British English utterances from 4 male actors expressing 7 emotions.
2024Allo-AVAArXiv 2024N/A~1,250 hours of conversational content for allocentric avatar gesture animation.
2024MMHeadACMMM 2024DownloadLarge-scale multi-modal 3D facial animation dataset with 49 hours of 3D facial motion sequences, speech audios, and hierarchical text annotations for text-induced 3D talking head animation and text-to-3D facial motion generation.
2024DH-FaceVid-1KICCV 2025Download1,200 hours, 270K+ clips from 20K+ individuals with speech audio, keypoints, and text annotations.
2024MultiTalkCVPR 2024Download420+ hours across 20 languages, 293K clips (512x512, 25fps, avg 5.19s duration).
2024THQAArXiv 2024Download800 talking head videos from 8 speech-driven methods with subjective quality assessments.
2023GRIDArXiv 2023Download34 volunteers each speaking 1000 phrases (34K utterances) with 6-word sentence structures.
2023ViCoArXiv 2023DownloadViCo and ViCo-X are datasets for conversational head generation, with ViCo for sentence-level independent talking and listening tasks, and ViCo-X for multi-turn conversational scenarios.
2023TalkingHead-1KHArXiv 2023Download500K video clips with ~80K greater than 512x512 resolution. Only permissive license videos included.
2023CelebVCVPR 2023DownloadIncludes CelebV-Text with 70,000 in-the-wild face video clips for text-to-video generation.
2023MMFace4DArXiv 2023DownloadLarge-scale multi-modal 4D dataset with 35,000+ sequences from 431 subjects (age 15-68).
2022CelebV-HQECCV 2022Download35,666 clips with 15,653 identities, each labeled with 83 facial attributes.
2022MultifaceNeurIPS 2022DownloadHigh-quality multi-view recordings of 13 people with 12K-23K frames per subject at 30fps. 65TB dataset.
2022VFHQCVPRW 2022Download16,000+ high-fidelity clips for video face super-resolution research.
2021HDTFCVPR 2021DownloadHigh-definition Talking-Face Dataset with ~362 videos (15.8 hours) in 720P/1080P resolution.
2020MEADECCV 2020DownloadLarge-scale audio-visual dataset with 60 actors expressing 8 emotions at 3 intensity levels.
2019BIWIArXiv 2019Download3D Audiovisual Corpus of Affective Communication with 40 sentences spoken by 14 subjects.
2019VOCASIGGRAPH 2019Download4D-face dataset with ~29 minutes of 4D face scans and synchronized audio from 12 speakers.
2019CN-CVSArXiv 2019DownloadLarge-scale continuous visual-speech dataset in Mandarin Chinese from TV news and speech shows.
2019FaceForensics++ICCV 2019DownloadLarge-scale dataset for detecting manipulated facial images with over 1.8M images.
2018VoxCeleb2Interspeech 2018DownloadLargest public audio-visual dataset with video URLs and timestamps. Requires 300GB+ storage.
2018LRS2ArXiv 2018DownloadLip reading dataset with videos recorded in diverse settings from BBC television.
2018LRWACCV 2018DownloadDiverse English-speaking dataset from BBC with 1000+ speakers. Each video is 1.16s (29 frames).
2017VoxCeleb1Interspeech 2017DownloadContains over 100,000 utterances for 1,251 celebrities, extracted from YouTube videos.
2017ObamaSetSIGGRAPH 2017DownloadSpecialized audio-visual dataset focused on analyzing visual speech of Barack Obama from weekly address footage.
2014CREMA-DACM TOCC 2014DownloadDiverse dataset with 7,442 clips featuring 91 actors (48 male, 43 female) aged 20-74, expressing six emotions at four intensity levels.

Survey

YearTitleConference/Journal
2026How to Build Digital Humans? From Priors to Photorealistic AvatarsEurographics 2026
2025A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative TechniquesArXiv 2025
2025Human Motion Video Generation: A SurveyTPAMI
2025A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and GenerationArXiv 2025
2025Controllable Video Generation: A SurveyArXiv 2025
2025Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss FunctionsArXiv 2025
2025Survey of Video Diffusion Models: Foundations, Implementations, and ApplicationsTMLR
2025A Survey on Human Interaction Motion GenerationArXiv 2025
2024Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyEMNLP 2025
2024Passive Deepfake Detection Across Multi-modalities: A Comprehensive SurveyArXiv 2024
20243D Gaussian Splatting: Survey, Technologies, Challenges, and OpportunitiesArXiv 2024
2024A Comprehensive Survey on Human Video Generation: Challenges, Methods, and InsightsArXiv 2024
2024A Survey on 3D Human Avatar Modeling — From Reconstruction to GenerationArXiv 2024
2024Video Diffusion Models: A SurveyArXiv 2024
2024Deepfake Generation and Detection: A Benchmark and SurveyACM Computing Surveys
2024ADTH-QA: A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head VideosArXiv 2024
2024How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a SurveyArXiv 2024
20243D Gaussian as a New Vision Era: A SurveyIEEE TVCG
2024Advances in 3D Generation: A SurveyArXiv 2024
2024A Survey on 3D Gaussian SplattingArXiv 2024
2024Neural Radiance Fields: Past, Present, and FutureArXiv 2024
2023From Pixels to Portraits: A Comprehensive Survey of Talking Head Generation Techniques and ApplicationsArXiv 2023
2023Human Motion Generation: A SurveyTPAMI
2023Human 3D Avatar Modeling with Implicit Neural Representation: A Brief SurveyArXiv 2023
2023Human-Computer Interaction System: A Survey of Talking-Head GenerationIEEE
2023Talking human face generation: A surveyACM
2022Face Generation and Editing with StyleGAN: A SurveyArXiv 2022
2022Deep Learning for Visual Speech Analysis: A SurveyArXiv 2022
2022A Survey on Applications of Digital Human Avatars toward Virtual Co-presenceArXiv 2022
2021A Review of 3D Face Reconstruction From a Single ImageArXiv 2021
2021Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth SynthesisArXiv 2021
2021AudioVisual Speech Synthesis: A brief literature reviewArXiv 2021
2020What comprises a good talking-head video generation?: A Survey and BenchmarkArXiv 2020

Funny Work


Audio-driven

YearTitleConference/JournalCodeProjectKeywords
2026EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture MixingArXiv 2026Projectlip synchronization, audio-driven, texture mixing, real-time
2026DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar GenerationACM MM 2026audio-driven, streaming avatar, lip-sync, self-forcing
2026Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait SynthesisArXiv 2026Codeaudio-driven, emotion control, talking portrait, lip-sync
2026Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar GenerationArXiv 2026CodeProjectjoint audio-video, streaming avatar, real-time, prompt planning
2026Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite AvatarsArXiv 2026CodeProjectaudio-driven, streaming avatar, real-time, long-horizon, distillation
2026Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait AnimationArXiv 2026audio-driven, portrait animation, emotion control, one-shot, real-time
2026GemTalk: Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face GenerationACM MM 2026audio-driven, emotional talking face, blendshape prior, diffusion
2026LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head GenerationArXiv 2026CodeProjectaudio-driven, talking head, real-time, diffusion distillation
2026TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human GenerationArXiv 2026CodeProjectdigital human, audio-video, real-time, talking head, TaoMate
2026AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready AvatarsArXiv 2026CodeProjectaudio-driven, avatar generation, distillation, long-form, real-time
2026SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait AnimationECCV 2026audio-driven, portrait animation, caching, lip-sync, DiT
2026ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech GuidanceArXiv 2026speech-driven, portrait animation, lip-sync, co-speech
2026Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip SynchronizationArXiv 2026CodeProjectlip synchronization, audio-driven, autoregressive diffusion, real-time
2026Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait AnimationICME 2026audio-driven, portrait animation, implicit motion, diffusion
2026Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEsCVPR 2026audio-driven, talking portrait, streaming, causal VAE, real-time
2026FreeTalkDiff: IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face GenerationArXiv 2026Codetalking face, diffusion, IP-Adapter, lip-sync
2026CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent PlanningArXiv 2026audio-driven, portrait animation, eye control, lip-sync, DiT
2026Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head GenerationArXiv 2026audio-driven, talking head, test-time adaptation, identity stability
2026HighSync: High-Quality Lip Synchronization via Latent Diffusion ModelsArXiv 2026Codelip synchronization, diffusion model, talking face, high-resolution, audio-driven
2026MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head GenerationArXiv 2026talking head, multi-conditional, diffusion, 3DMM, controllable
2026AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language ModelsArXiv 2026speech-driven, facial animation, blendshape, multimodal LLM, language-assisted
2026TAVR: Generate Your Talking Avatar from Video ReferenceArXiv 2026Projecttalking avatar, video reference, cross-scene, reinforcement learning, identity preservation
2026Talking Slide Avatars: Open-Source Multimodal Communication Approach for TeachingArXiv 2026talking slide avatars, text-to-speech, audio-driven synthesis, educational
2026Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal RetrievalArXiv 2026causal audio-driven, facial motion, multi-modal retrieval, personalization
2026Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference DistillationArXiv 2026Codereal-time avatar, audio-video generation, diffusion, streaming
2026Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion ModelingArXiv 2026CodeProjectjoint audio-video generation, autoregressive diffusion, talking head synthesis
2026EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal CoherenceICMR 2026emotion-aware, talking head generation, lip-sync, temporal coherence
2026Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression ManipulationArXiv 2026speech-preserving, facial expression manipulation, spatial-temporal correlation, emotion editing
2026Polyglot: Multilingual Style Preserving Speech-Driven Facial AnimationArXiv 2026Projectmultilingual, speaker style, speech-driven, facial animation
2026TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar GenerationArXiv 2026distillation, audio-driven, avatar, head avatar
2026SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion DiarizationArXiv 2026CodeProjectaudio-driven, 3D, emotion, facial animation
2026C-MET: Cross-Modal Emotion Transfer for Emotion Editing in Talking Face VideoCVPR 2026CodeProjectemotion transfer, emotion editing, talking face, cross-modal
2026MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature FusionArXiv 20263D, Audio-Driven, Multimodal Fusion, Mesh Parameterization
2026EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationCVPR 2026CodeProjectGaussian Splatting, 3DGS, Audio-Driven, Emotion, Few-Shot
2026EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise ControlArXiv 2026Autoregressive, GPT-style, Talking Head
2026FreeTalk: Emotional Topology-Free 3D Talking HeadsArXiv 20263D Talking Heads, Emotional, Topology-Free
2026AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window DenoisingArXiv 2026CodeProjectStreaming Avatar, Real-Time, Diffusion
2026DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and SynchronizationCVPR 2026 FindingsCodeProjectVideo Dubbing, Flow Matching, Cross-Modal
2026OmniEdit: A Training-free Framework for Lip Synchronization and Audio-Visual EditingArXiv 2026CodeLip Sync, Audio-Visual Editing, Training-Free
2026EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian DeformationPreprintGaussian Splatting, 3DGS, Audio-Driven, Talking Head
2026TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head GenerationArXiv 2026CodeProjectDiffusion, Audio-Driven, Talking Head, VAE, Latent
2026UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging ScenariosArXiv 2026Lip Sync, Pose-Anchored, Generalizable
2026UniTalking: A Unified Audio-Video Framework for Talking Portrait GenerationCVPR 2026Audio-Driven, Portrait Animation, Talking Head, CVPR, Transformer, Attention
2026FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video GenerationArXiv 2026Audio-Driven, Portrait Animation, Reinforcement Learning, GRPO
2026Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent SpaceWACV 2026CodeAudio-Driven, WACV, Latent
2026VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head ReconstructionArXiv 2026audio-driven, video conferencing, talking head
2026DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationArXiv 2026CodeProjectAudio-Driven, Transformer, Attention
20263DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking AvatarsArXiv 2026CodeProject3D, Emotional, Lip Sync, Avatar, Talking Head, Transformer
2026AUHead: Realistic Emotional Talking Head Generation via Action Units ControlArXiv 2026CodeAction Units, Audio-Driven Generation, Emotion Control, Diffusion Model
2026MOVA: Towards Scalable and Synchronized Video-Audio GenerationArXiv 2026CodeProjectAudio-Driven
2026VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational AvatarsArXiv 2026CodeProjectAvatar, Talking Head
2026SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking HeadsArXiv 2026CodeProjectReal-time, Streaming, Talking Head
2026Asymmetric Hierarchical Anchoring for Audio-Visual Joint RepresentationArXiv 2026Audio-Driven
2026JoyAvatar: Unlocking Highly Expressive Avatars via Harmonized Text-Audio ConditioningArXiv 2026ProjectAudio-Driven, Avatar
2026LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the WildArXiv 2026CodeProjectLip Sync, Audio-Driven, Talking Head, Latent
2026MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion ControlICASSP 2026Personalized Avatars, Lip Sync, Style Disentanglement, Diffusion Model
2026JUST-DUB-IT: Video Dubbing via Joint Audio-Visual DiffusionArXiv 2026CodeProjectAudio-Visual Diffusion, LoRA, Lip Sync
2026EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion TransformersArXiv 2026ProjectDiffusion, Audio-Driven, Talking Head
2026UA-3DTalk: Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior DistillationICASSP 2026CodeProject3D, Emotional, Talking Head, ICASSP, Attention
2026Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks EncodingArXiv 2026Audio-Driven, Talking Head, Transformer
2026SkyReels-V3 Technique ReportArXiv 2026CodeVideo Generation, Audio-Guided, Talking Avatar, Diffusion Transformers
2026THFEM: Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression ManipulationACM Trans. MultimediaCodeProjectSpeech-Driven, Talking Head
2026Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head VideosArXiv 2026Talking Head
2026EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression EditingArXiv 20263D, Speech-Driven
2026ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian SplattingArXiv 20263D, Gaussian Splatting, 3DGS, Emotional, Audio-Driven
2026DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive ModelArXiv 2026CodeProjectStreaming, Talking Head, Flow Matching
2026SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional DistillationArXiv 2026CodeProjectReal-time, Streaming, Audio-Driven, Avatar, Attention, VAE
2026SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wildArXiv 2026ProjectTransformer
2026JoyAvatar-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive DiffusionArXiv 2026ProjectDiffusion, Real-time, Audio-Driven, Avatar
2026REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming DistillationArXiv 2026Diffusion, Real-time, Streaming, Talking Head, Latent
2026Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head GenerationArXiv 2026Talking Head, Attention, Latent
2025X-Dub: From Inpainting to Editing: A Self-Bootstrapping Framework for Context-Rich Visual DubbingArXiv 2025CodeProjectVisual dubbing, Diffusion Transformer, Self-bootstrapping, Lip sync
2025PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentArXiv 20253D, Speech-Driven, Talking Head, Attention
2025FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANsArXiv 2025Diffusion, Transformer, GAN, Latent
2025In-Context Audio Control of Video Diffusion TransformersArXiv 2025Diffusion, Audio-Driven, Transformer, Attention
2025Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication SystemsIEEE Big Data 2025lip sync, real-time, multilingual
2025TalkVerse: Democratizing Minute-Long Audio-Driven Video GenerationArXiv 2025CodeProjectAudio-Driven, VAE, Latent
2025VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageNeurIPS 20253D, Gaussian Splatting, Audio-Driven, Avatar
2025FacEDiT: Unified Talking Face Editing and Generation via Facial Motion InfillingArXiv 2025ProjectTalking Head, Transformer, Attention, Flow Matching
2025JoVA: Unified Multimodal Learning for Joint Video-Audio GenerationArXiv 2025CodeProjectAudio-Driven, Transformer, Attention, GAN
2025Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation ModelTech ReportAudio-Driven, Transformer, Reinforcement Learning
2025STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking PortraitsArXiv 2025CodeProjectDiffusion, Portrait Animation, Talking Head
2025GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian SplattingWACV 20263D, Gaussian Splatting, 3DGS, Audio-Driven, Talking Head
2025UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingCVPR 2026CodeProjectAudio-Driven, Avatar
2025Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite LengthArXiv 2025CodeProjectReal-time, Streaming, Audio-Driven, Avatar
2025EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking HumansArXiv 2025Portrait Animation, Talking Head
2025AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity RefinementArXiv 2025CodeProjectTalking Head, Transformer, Attention
2025AI killed the video star. Audio-driven diffusion model for expressive talking head generationArXiv 2025audio-driven, diffusion, talking head
2025IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion TransferArXiv 2025Audio-Driven, Talking Head, Attention, Latent
2025Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyArXiv 2025Audio-Driven, Attention, Latent
2025StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion ModelArXiv 2025Project3D, Diffusion, Streaming, Audio-Driven
2025ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise SearchAAAI 2026Diffusion, Talking Head, AAAI, Knowledge Distillation
2025Shared Latent Representation for Joint Text-to-Audio-Visual SynthesisArXiv 2025Audio-Driven, Latent
2025UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsArXiv 2025Audio-Driven, Transformer, Attention, Latent
2025See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region RefinementTASLP 2025High-Resolution, Talking Faces, Speech-to-Face, Diffusion
2025Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face AnimationICXR 2025Blendshapes, FLAME, Disentanglement, 3D Animation
2025MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity ControlArXiv 2025
2025LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature RepresentationSIGGRAPH Asia 2025CodeLabel-Free, Speech-Driven, Facial Animation, FLAME
2025Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward FeedbackArXiv 2025Diffusion, Audio-Driven, AAAI, Transformer
2025DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait SynthesisArXiv 2025Disentangled Motion, Flow Matching, Talking Portrait, Controllable
2025SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face RepresentationArXiv 2025Contrastive Masked Pretraining, Audio-Visual, Talking-Face
2025EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian DeformationIEEE SMC 2025Real-Time, Audio-Driven, Gaussian Deformation, Talking Head
2025A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple LanguagesArXiv 2025Phoneme-Viseme Alignment, Multilingual TFS, Mixture-of-Experts
2025IASA: Input-Aware Sparse Attention for Real-Time Co-Speech Video GenerationArXiv 2025CodeProjectDiffusion models, co-speech video, real-time, sparse attention
2025Audio Driven Real-Time Facial Animation for Social TelepresenceSIGGRAPH Asia 2025ProjectReal-time, Audio-Driven, SIGGRAPH, Transformer, Latent
2025StableDub: Taming Diffusion Prior for Generalized and Efficient Visual DubbingArXiv 2025ProjectVisual Dubbing, Diffusion, Mamba-Transformer
2025KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial AnimationArXiv 2025Keyframe, Diffusion, Dual-Path, Facial Animation
2025SynchroRaMa: Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion EmbeddingWACV 2026CodeProjectMulti-Modal, Emotion-Aware, LLM
2025Talking Head Generation via AU-Guided Landmark PredictionArXiv 2025Action Units, Landmark Prediction, Diffusion
2025PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density ControlICONIP 20253DGS, Real-Time, Pixel-Aware, Audio-Driven
2025A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync SynthesisArXiv 2025lip sync, voice cloning
2025Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation SynthesisArXiv 2025ProjectMultimodal Instructions, Avatar Synthesis, Lip Synchronization
2025Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head AnimationArXiv 2025Singing-Driven, 3D Head, Diffusion
2025EmoCAST: Emotional Talking Portrait via Emotive Text DescriptionArXiv 2025CodeProjectEmotional, Portrait Animation, Talking Head, Attention
2025Wan-S2V: Audio-Driven Cinematic Video GenerationArXiv 2025Cinematic, Audio-Driven, Video Generation
2025Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured GuidanceArXiv 2025 (Withdrawn)Emotional, Avatar, Talking Head, Transformer, Latent
2025Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital AvatarsArXiv 2025Audio-driven Realistic Facial Animation, Digital Avatars
2025D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head SynthesisECAI 2025Few-Shot, 3DGS, Deformation Fields
2025InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video DubbingArXiv 2025Sparse-Frame Dubbing, Full-Body
2025CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face GenerationArXiv 2025Cross-emotion memory, audio emotion enhancement, expression displacement, lip sync
2025RealTalk: Realistic Emotion-Aware Lifelike Talking-Head SynthesisICCV 2025 WorkshopEmotion, NeRF, VAE
2025FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait AnimationArXiv 2025CodeProjectAudio-Driven, Portrait Animation, Preference Optimization
2025HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head SynthesisArXiv 2025Hybrid Motion, High-Fidelity, Talking Head
2025StableAvatar: Infinite-Length Audio-Driven Avatar Video GenerationArXiv 2025CodeProjectStable Diffusion
2025DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait AnimationArXiv 2025ProjectDiT, Portrait Animation, Speaking Styles
2025READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationArXiv 2025ProjectDiffusion, Real-time, Audio-Driven, Talking Head, Transformer, VAE
2025X-Actor: Emotional and Expressive Long-Range Portrait Acting from AudioArXiv 2025ProjectEmotional Portrait, Long-range, Audio-driven
2025SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationACM MM 2025Spatial Audio, Video Generation, MLLM
2025Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity PreservationArXiv 2025Mask-Free, Identity Preservation, Audio-driven
2025Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial AnimationInterspeech 2025CodeProjectPhonetic Context, Viseme, 3D Facial Animation
2025MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationICCV 2025CodeProjectPersonalized, 3D Facial Animation, Memory
2025JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-SyncArXiv 20253DMM, Joint Learning, Talking Head
2025Livatar-1: Real-Time Talking Heads Generation with Tailored Flow MatchingTechnical ReportProjectreal-time, flow matching, lip-sync
2025ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise DiffusionArXiv 2025CodeDiffusion, Landmarks-Guide, Real-time, Identity Preservation
2025MOSPA: Human Motion Generation Driven by Spatial AudioNeurIPS 2025CodeSpatial Audio, Human Motion Generation, Virtual Human
2025M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head GenerationArXiv 2025ProjectMulti-granular Motion, Decoupling, Optimization
2025MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingArXiv 2025CodeMultimodal, 3D Facial Animation, Dynamic Emotions
2025MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head GenerationArXiv 20253DMM, Diffusion Transformer, Temporal Consistency, Blinking Dynamics
2025MoDA: Multi-modal Diffusion Architecture for Talking Head GenerationArXiv 2025CodeProjectMulti-modal, Diffusion, Talking Head Generation
2025FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesArXiv 2025Identity Leakage, Extreme Cases
2025JAM-Flow: Joint Audio-Motion Synthesis with Flow MatchingArXiv 2025Flow Matching, Audio-Motion
2025FIAG: Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian FieldArXiv 2025CodeFew-Shot, Global Gaussian Field, 3DGS
2025GGTalker: Talking Head Synthesis with Generalizable Gaussian Priors and Identity-Specific AdaptationICCV 2025CodeProject3D Talking Head, Gaussian Priors, Identity Adaptation
2025SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian SplattingArXiv 2025CodeProject3DGS, Synchronization
2025LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion ModelsArXiv 2025Low-Latency, Real-Time, Interactive
2025TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion ModelsArXiv 2025CodeProjectReal-Time, Autoregressive Diffusion
2025SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion TransformersArXiv 2025video diffusion transformer, multimodal, talking portrait, audio-conditioned
2025Video Editing for Audio-Visual DubbingArXiv 2025Video Editing, Dubbing
2025Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationCVPR 2025Code3D, Semantic Decoupling
2025MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video GenerationArXiv 2025CodeCo-Speech Gesture, Two-Stage
2025FaceEditTalker: Interactive Talking Head Generation with Facial Attribute EditingArXiv 2025ProjectAttribute Editing, Interactive
2025OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersArXiv 2025Lip Sync, Universal, Visual Prosody
2025OT-Talk: Animating 3D Talking Head with Optimal TransportationArXiv 2025FLAME, 3D
2025GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian SplattingCVPRW 20253DGS
2025Model See Model Do: Speech-Driven Facial Animation with Style ControlSIGGRAPH 2025
2025KeySync: A Robust Approach for Leakage-free Lip Synchronization in High ResolutionArXiv 2025CodeProject
2025IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCVPR 2025Project3D-aware, Video Diffusion
2025Audio-Driven Talking Face Video Generation with Joint Uncertainty LearningArXiv 2025Joint uncertainty learning, Audio-driven talking face, Lip sync, Visual uncertainty
2025DICE-Talk: Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait GenerationACM MM 2025Emotional Portrait, Identity Preservation, Emotion Cooperation
2025PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head SynthesisArXiv 2025Lip Sync, Phoneme-Aware, Speech Encoder
2025Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head AnimationTMM 2025Talking Head Animation, Temporal Correlation, One-Shot
2025FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisArXiv 2025CodeProjecttalking portrait, motion synthesis, video diffusion, audio-visual alignment
2025ACTalker: Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationArXiv 2025CodeProject
2025MGGTalk:Monocular and Generalizable Gaussian Talking Head AnimationCVPR 2025ProjectOne Shot, 3DGS
2025STSA: Spatial-Temporal Semantic Alignment for Visual DubbingICME 2025CodeSpatial-Temporal Alignment, Semantic Features, Visual Dubbing, Stability
2025Dual Audio-Centric Modality Coupling for Talking Head GenerationArXiv 2025NeRF
2025Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head SynthesisArXiv 2025Project3DGS
2025Audio-driven Gesture Generation via Deviation Feature in the Latent SpaceArXiv 2025Gesture
2025Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation MetricsCVPR 2025CodeProject
2025AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion TransformersCVPR 2025ProjectDiT
2025EmoHead: Emotional Talking Head via Manipulating Semantic Expression ParametersArXiv 2025Neural Radiance Fields, Expression Parameters, Emotion Control, Audio-Driven
2025DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion ModelICME 2025Project
2025Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion GenerationCVPR 2025ProjectAutoregressive
2025DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided DistillationICME 2025CodeDiffusion, 3D
2025KDTalker: Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking PortraitArXiv 2025Codeimplicit keypoint, spatiotemporal diffusion, audio-driven, talking portrait
2025StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial AnimationArXiv 20253D
2025MagicInfinite: Generating Infinite Talking Videos with Your Words and VoiceArXiv 2025Project
2025FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait SynthesisICMR 2025
2025KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame InterpolationCVPR 2025Diffusion, Long Sequences
2025TexTalk: Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureCVPR 2025CodeProjectTexture
2025InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoCVPR 2025CodeProjectFew Shot, 3DGS
2025ARTalk: Speech-Driven 3D Head Animation via Autoregressive ModelSIGGRAPH AsiaCodeProjectAutoregressive, FLAME, 3D
2025FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion modelArXiv 2025Diffusion
2025Dimitra: Audio-driven Diffusion model for Expressive Talking Head GenerationArXiv 2025Audio-driven, Diffusion model, Motion Diffusion Transformer, Lip sync
2025NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head SynthesisICASSP 2025CodeProject
2025Emotional Face-to-SpeechArXiv 2025Projectemotion, face2speech
2025EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head SynthesisArXiv 2025emotion, 3DGS
2025Identity-Preserving Video Dubbing Using Motion WarpingArXiv 2025Video Dubbing
2025MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationArXiv 2025Audio-driven, Emotion Synthesis, Mixture of Experts, Portrait Animation
2025DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face SynthesisICASSP 2025CodeHair-Preserving
2025UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting ControlArXiv 2025SD, Lighting control
2025FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG DistillationCVPR 2025ProjectFast Diffusion 12.5X speedup
2025Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationCVPR 2025CodeProject
2025V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified FlowICASSP 2025CodeVideo-to-Speech, Speech Decomposition
2025Sonic: Shifting Focus to Global Audio Perception in Portrait AnimationCVPR 2025CodeProjectGlobal Audio Perception, Portrait Animation
2025MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal SamplingArXiv 2025Code
2025Lipschitz-Driven Noise Robustness in VQ-AE for High-Frequency Texture Repair in ID-Specific Talking HeadsArXiv 2025Noise Robustness, VQ-AE, High-Frequency
2025Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion DependencyICLR 2025Project
2025EmoFace: Emotion-Content Disentangled Speech-Driven 3D Talking Face AnimationArXiv 2025emotion,3D
2025DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsICCV 2025ProjectGaussian, Latent Space
2025Talking Head Generation via Viewpoint and Lighting Simulation Based on Global RepresentationACM MM 2025Depth-based
2025PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesACM MM 2025FLAME
2025DisenEmo: Learning disentangled emotional representation from facial motion for 3D talking head generationICIP 2025Disentangled Emotional Representation, 3D Talking Head Generation
2025ExpTalk: Diverse Emotional Expression via Adaptive Disentanglement and Refined Alignment for Speech-Driven 3D Facial AnimationIJCAI 2025Adaptive Disentanglement, Refined Alignment, 3D Facial Animation
2025SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeatureIJCAI 2025Stable 3D Gaussian-Based Talking Head Generation, Enhanced Lip Sync, Discriminative Speech Feature
2025SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head AnimationArXiv 2025Huaman Pose
2025JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video EditingArXiv 2025Depth, JD work
2024VQTalker: Towards Multilingual Talking Avatars through Facial Motion TokenizationArXiv 2024CodeProjectvisemes, code book
2024LatentSync: Audio Conditioned Latent Diffusion Models for Lip SyncArXiv 2024Diffusion, SyncNet
2024PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head SynthesisAAAI 2025Point Cloud, Gaussian Splatting
2024PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face GenerationArXiv 2024Diffusion, Attention, One-Shot
2024MEMO: Memory-Guided Diffusion for Expressive Talking Video GenerationArXiv 2024CodeProjectMemory
2024IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head GenerationArXiv 2024Motion Diffusion Model
2024FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking PortraitICCV 2025CodeProjectFlow Matching
2024LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head SynthesisArXiv 2024ProjectNeRF
2024Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head SynthesisArXiv 2024CodeProjectDiffusion
2024GaussianSpeech: Audio-Driven Gaussian AvatarsArXiv 2024CodeProject3DGS, 3D
2024LetsTalk: Latent Diffusion Transformer for Talking Video SynthesisArXiv 2024
2024EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionCVPR 2025ProjectEmotion, Expressive, Diffusion
2024Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait SynthesisArXiv 2024Audio Feature Extraction, Whisper, Real-time processing, Talking portrait synthesis
2024LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion SpaceArXiv 2024Fine-Grained Emotion
2024JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion GenerationArXiv 2024CodeDiffusion, VASA
2024Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-ExpertsArXiv 2024
2024Audio-Driven Emotional 3D Talking-Head GenerationArXiv 2024Emotion
2024Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss OptimizationArXiv 2024
2024DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video GenerationArXiv 2024CodeNon-autoregressive Diffusion
2024Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationICLR 2025CodeProjectDiffusion, Hallo
2024Diverse Code Query Learning for Speech-Driven Facial AnimationArXiv 2024
2024TalkinNeRF: Animatable Neural Fields for Full-Body Talking HumansECCVW 2024CodeProjectNeRF
2024JoyHallo: Digital human model for MandarinArXiv 2024CodeProjectDiffusion, Hallo
2024JEAN: Joint Expression and Audio-guided NeRF-based Talking Face GenerationBMVC 2024ProjectNeRF
20243DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion PolicyArXiv 2024
2024LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping DeformationArXiv 2024
2024StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking HeadsTPAMI 2024
2024ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAESIGGRAPH MIG 2024Code3D
2024DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech GesturesArXiv 2024diffusion
2024EMOdiffhead: Continuously Emotional Control in Talking Head Generation via DiffusionArXiv 2024Diffusion
2024PersonaTalk: Bring Attention to Your Persona in Visual DubbingSIGGRAPH Asia 2024Project
2024KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks GenerationArXiv 2024KAN
2024SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion ModelArXiv 2024Diffusion, Style
2024PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head GenerationArXiv 2024ProjectPose Latent Diffusion, Lip Synchronization, Text-Audio Control
2024Mini-Omni: Language Models Can Hear, Talk While Thinking in StreamingTech ReportCodeOmni!!!
2024TalkLoRA: Low-Rank Adaptation for Speech-Driven AnimationArXiv 2024LoRA
2024Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face AnimationArXiv 2024
2024S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head SynthesisECCV 2024
2024DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationAAAI 2025CodeProject3D face, FLAME, Emotion
2024High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion ModelIEEE TIP
2024Style-Preserving Lip Sync via Audio-Aware Style ReferenceIEEE TIP
2024MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture GenerationArXiv 2024Co-Speech Gesture
2024GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion TransformerArXiv 2024
2024Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head GenerationArXiv 2024
2024JambaTalk: Speech-Driven 3D Talking Head Generation Based on Hybrid Transformer-Mamba ModelArXiv 20243D
2024ICCA: Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMsCOLM 2024CodeLLM
2024UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified ModelArXiv 2024Code
2024DiM-Gesture: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2 frameworkArXiv 2024
2024EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking HeadECCV 2024Project
2024LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control EnhancementArXiv 2024
2024Text-based Talking Video Editing with Cascaded Conditional DiffusionArXiv 2024
2024EmoFace: Audio-driven Emotional 3D Face AnimationIEEE VR 2024Code
2024Learning Online Scale Transformation for Talking Head Video GenerationArXiv 2024
2024Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGANArXiv 2024StyleGAN
2024Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading ExpertInterspeech 2024CodeProject3D
2024RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment NetworkArXiv 2024
2024MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video DatasetInterspeech 2024CodeProject3D, Dataset
2024RITA: A Real-time Interactive Talking Avatars FrameworkArXiv 2024Real-time, Interactive, Talking Avatar
2024NLDF: Neural Light Dynamic Fields for Efficient 3D Talking Head GenerationArXiv 2024NeRF
2024Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image AnimationArXiv 2024CodeProject🔥EMO, Diffusion, Open-source
2024MyTalk: Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance DisentanglementArXiv 2024Project
2024Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose GenerationArXiv 2024Emotion
2024ControlTalk: Controllable Talking Face Generation by Implicit Facial Keypoints EditingArXiv 2024CodeFace Edit
2024SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head GenerationArXiv 2024
2024NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative PriorCVPRW 2024SadTalker+NeRF
2024SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent SpaceICASSP 2025
2024AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingArXiv 2024Code
2024GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian SplattingArXiv 2024🔥Gaussian Splatting
2024CSTalk: Correlation Supervised Speech-driven 3D Emotional Facial Animation GenerationArXiv 2024Emotion
2024GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian SplattingACMM 2024CodeProject🔥Gaussian Splatting
2024TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian SplattingECCV 2024CodeProject🔥Gaussian Splatting
2024GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingACMM 2024Project🔥Gaussian Splatting
2024Learn2Talk: 3D Talking Face Learns from 2D Talking FaceArXiv 2024🔥Gaussian Splatting
2024VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeNeurIPS 2024🔥🔥🔥Awesome,Microsoft
2024EDTalk: Efficient Disentanglement for Emotional Talking Head SynthesisECCV 2024CodeProjectEmotion
2024Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative PriorArXiv 2024CodeProject
2024AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait AnimationArXiv 2024Code🔥🔥🔥Similar to EMO
2024Adaptive Super Resolution For One-Shot Talking-Head GenerationICASSP 2024Code
2024EmoVOCA: Speech-Driven Emotional 3D Talking HeadsArXiv 20243D, VOCA
2024ScanTalk: 3D Talking Heads from Unregistered ScansECCV 2024Code3D
2024FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationArXiv 2024Normalizing Flow, Vector-Quantization, Lip Sync, Emotional Talking Faces
2024Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleArXiv 2024
2024FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioArXiv 2024Code
2024G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal AlignmentArXiv 2024A Generic Framework
2024EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak ConditionsArXiv 2024🔥🔥🔥Amazing, Diffusion
2024Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisCVPR 2024High-Quality
2024DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerArXiv 2024Code3D
2024EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face GenerationArXiv 2024CodeProjectEmotion
2024NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face SynthesisICASSP 2024CodeProjectAU
2024Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisICLR 2024CodeProject3D, One-Shot, Realistic
2024Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering PriorsArXiv 2024Projectvisual dubbing, lip sync, neural rendering, data-efficient
2024DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face GenerationArXiv 2024ProjectEmotion
2024VectorTalker: SVG Talking Face Generation with Progressive VectorisationArXiv 2024SVG
2024AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisAAAI 2024
2024Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial AnimationAAAI 2024CodeProject3D
2024DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic ModelsArXiv 2024CodeProjectDiffusion
2024FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head ModelsArXiv 2024CodeProject
2024GMTalker: Gaussian Mixture based Emotional talking video PortraitsArXiv 2024ProjectEmotion
2024GSmoothFace: Generalized Smooth Talking Face Generation via Fine Grained 3D Face GuidanceArXiv 2024CodeProject3D
2024R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer ConditioningArXiv 2024based-RAD-NeRF
2024VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorArXiv 2024Mesh
2024SyncTalk: The Devil😈 is in the Synchronization for Talking Head SynthesisCVPR 2024CodeProject😈Talking Head
2024GAIA: Zero-shot Talking Avatar GenerationArXiv 2024Project😲😲😲
2024AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial AnimationIEEE Transactions on MultimediaCodeProject3D, Mesh
2024DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisICASSP 2024ER-NeRF
2024EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditioningAAAI 2025🔥阿里
2024Pose-Aware 3D Talking Face Synthesis using Geometry-guided Audio-Vertices AttentionIEEE 2024
2023DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoderICASSP 2024CodeProjectvisual dubbing, diffusion, inpainting, person-generic
2023Towards Streaming Speech-to-Avatar SynthesisArXiv 2023Streaming Synthesis, Articulatory Inversion, Real-time, Speech-driven
2023OSM-Net: One-to-Many One-shot Talking Head Generation with Spontaneous Head MotionsArXiv 2023One-shot Talking Head, Head Motions, One-to-Many Mapping, Audio-driven
2023EAT: Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationICCV 2023CodeProject-
2023Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric SynthesisTCSVT 2023
2023Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationICCV 2023-
2023Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksInterSpeech 2023Emotion
2023StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based GeneratorCVPR 2023CodeProject-
2023High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space LearningCVPR 2023Emotion
2023FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsICME 2023Natural Head Motions, Flow-guided, Audio-driven Pose Prediction, One-shot Talking Head
2023DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderACMM 2023🔥Diffusion
2023TalkLip: Seeing What You Said - Talking Face Generation Guided by a Lip Reading ExpertCVPR 2023
2023OTAvatar: One-shot Talking Face Avatar with Controllable Tri-plane RenderingCVPR 2023CodeTri-plane Rendering, One-shot Avatar, Controllable, 3D Consistency
2023Emotionally Enhanced Talking Face GenerationArXiv 2023Emotion
2023EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationICCV 2023CodeProject3D, Emotion
2023READ Avatars: Realistic Emotion-controllable Audio Driven AvatarsArXiv 2023-
2023OPT: One-shot Pose-Controllable Talking Head GenerationICASSP 2023pose control, identity preservation, audio feature disentanglement
2023DiffTalk: Crafting Diffusion Models for Generalized Talking Head SynthesisCVPR 2023CodeProject🔥Diffusion
2023CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorCVPR 2023CodeProject3D, codebook
2023StyleTalk: One-shot Talking Head Generation with Controllable Speaking StylesAAAI 2023CodeStyle
2023SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationCVPR 2023CodeProject3D, Single Image
2023Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisICCV 2023Tri-plane
2023LipNeRF: What is the right feature space to lip-sync a NeRF?FG 2023Wav2lip
2023ToonTalker: Cross-Domain Face ReenactmentICCV 2023-
2023EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationICCV 2023Emotion
2023Facediffuser: Speech-driven 3d facial animation synthesis using diffusionACM SIGGRAPH MIG 2023🔥Diffusion, 3D
2023DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoAAAI 2023
2023Diffused Heads: Diffusion Models Beat GANs on Talking-Face GenerationArXiv 2023🔥Diffusion
2022Memories are One-to-Many Mapping Alleviators in Talking Face GenerationArXiv 2022Project-
2022Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in TransformersSIGGRAPH Asia 2022-
2022VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildSIGGRAPH 2022
2022Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisArXiv 2022disentangled representation, contrastive learning, multi-motion control
2022SPACEx 🚀: Speech-driven Portrait Animation with Controllable ExpressionArXiv 2022Project-
2022Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking AvatarICTAI2022talking avatar, presentation
2022EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelSIGGRAPH 2022Emotion
2022Emotion-Controllable Generalized Talking Face GenerationIJCAI 2022emotion control, graph convolutional network, geometry-aware
2022StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANArXiv 2022StyleGAN, high-resolution, one-shot, lip sync
2022Expressive Talking Head Generation with Granular Audio-Visual ControlCVPR 2022-
2022Talking Face Generation with Multilingual TTSCVPR 2022-
2021One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation LearningAAAI 2022one-shot, audio-visual correlation, keypoint-based motion, lip sync
2021Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisACM MM 2021-
2021Talking Head Generation with Audio and Speech Related Facial Action UnitsBMVC 2021AU
2021Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head MotionArXiv 2021Audio-driven, Talking-head, Head Motion, Keypoint-based Motion
20213D-TalkEmo: Learning to Synthesize 3D Emotional Talking HeadArXiv 20213D Talking Head, Emotion, Geometry Map, Audio-driven
2021PC-AVS: Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationCVPR 2021-
2021MakeItTalk: Speaker-Aware Talking-Head AnimationSIGGRAPH Asia 2020Speaker-Aware, Audio-Driven, Facial Landmarks, Photorealistic
2021Speech2Talking-Face: Inferring and Driving a Face with Synchronized Audio-Visual RepresentationIJCAI 2021-
2021Audio-Driven Emotional Video PortraitsCVPR 2021Emotion
2021IATS: Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisACM Multimedia 2021-
2020Multi Modal Adaptive Normalization for Audio to Video GenerationArXiv 2020Audio-to-Video, Multi-Modal Adaptive Normalization, Facial Video Generation, Keypoint Heatmap
2020A Lip Sync Expert Is All You Need for Speech to Lip Generation In The WildACM Multimedia 2020-
2020Talking-head Generation with Rhythmic Head MotionECCV 2020-
2020Speaker-Aware Talking-Head AnimationSIGGRAPH Asia 2020-
2020Neural Voice Puppetry: Audio-driven Facial ReenactmentECCV 2020-
2020Realistic Speech-Driven Facial Animation with GANsIJCV 2020-
2020A Large-scale Audio-visual Dataset for Emotional Talking-face GenerationECCV 2020-
2019Talking Face Generation by Adversarially Disentangled Audio-Visual RepresentationAAAI 2019-
2019Hierarchical Cross-modal Talking Face Generation with Dynamic Pixel-wise LossCVPR 2019-
2018Audio-Driven Animator-Centric Speech AnimationSIGGRAPH 2018-
2018Lip Movements Generation at a GlanceECCV 2018-
2017You Said That? Synthesising Talking Faces From AudioBMVC 2019-
2017Synthesizing Obama: Learning Lip Sync From AudioSIGGRAPH 2017-
2017Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and EmotionSIGGRAPH 2017-
2017A Deep Learning Approach for Generalized Speech AnimationSIGGRAPH 2017-

Portrait Animation

YearTitleConference/JournalCodeProjectKeywords
2026TongueReenact: Geometry-Anchored Tongue Synthesis for Face ReenactmentArXiv 2026face reenactment, tongue, video-driven, diffusion, cross-identity
2026ViDS: Video Diffusion Shader using 3D Face TrackingArXiv 2026Projectportrait animation, video-driven, 3DMM, video diffusion, face tracking
2026MagPlus: Bridging Micro-to-Regular Facial Expressions through Learnable MagnificationArXiv 2026micro-expression, facial animation, motion magnification, portrait animation
2026Loki: Representation over Architecture for Diffusion-Based Portrait AnimationArXiv 2026face reenactment, video-driven, portrait animation, expression, head pose
2026PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial ReenactmentCVPR 2026face reenactment, disentanglement, real-time, CVPR 2026
2026MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face GenerationCVPR 2026CodeProjectDiffusion Transformer, Multimodal, Face Generation
2026FG-Portrait: 3D Flow Guided Editable Portrait AnimationCVPR 2026portrait animation, 3D flow, CVPR
2026MoCha:End-to-End Video Character Replacement without Structural GuidanceArXiv 2026Talking Head
2025Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait AnimationArXiv 2025CodeProjectDiffusion, Real-time, Portrait Animation, Attention
2025SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait AnimationArXiv 2025Portrait Animation, ICASSP, Attention
2025FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionArXiv 2025Portrait Animation, Transformer, Latent
2025DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsArXiv 2025ProjectPortrait Animation, Disentangled, Expressive
2025FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and ViewpointArXiv 2025ProjectPortrait Animation, Transformer, Latent
2025PersonaLive! Expressive Portrait Image Animation for Live StreamingArXiv 2025Streaming, Portrait Animation
2025Beat on Gaze: Learning Stylized Generation of Gaze and Head DynamicsArXiv 2025Gaze Control, Head Motion, Style-Aware, 3D
2025Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion ModelArXiv 2025Face Reenactment, Large-Pose, Video Diffusion
2025Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory GuidanceArXiv 2025
2025HunyuanPortrait: Implicit Condition Control for Enhanced Portrait AnimationCVPR 2025Hunyuan
2025Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided DiffusionArXiv 2025CodeDiffusion, 3D
2025MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile DevicesCVPR 2025100+fps
2025GOES: 3D Gaussian-based One-shot Head Animation with Any Emotion and Any StyleACM MM 2025One-Shot, 3DGS
2024GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionAAAI 2025Gaze-oriented
2024Avatar Concept Slider: Manipulate Concepts In Your Human Avatar With Fine-grained ControlArXiv 2024
2024G3FA: Geometry-guided GAN for Face AnimationBMVC 2024
2024Anchored Diffusion for Video Face ReenactmentArXiv 2024Face Reenactment, Anchored Diffusion
2024One-Shot Pose-Driving Face Animation PlatformArXiv 2024One-Shot, Pose-Driving, Face Animation, Talking Head
2024V-Express: Conditional Dropout for Progressive Training of Portrait Video GenerationTech Report🔥EMO, Diffusion, Open-source
2024EMOPortraits: Emotion-enhanced Multimodal One-shot Head AvatarsArXiv 2024CodeEmotional, One Shot, Cross-Driving
2024EMOPortraits: Emotion-enhanced Multimodal One-shot Head AvatarsArXiv 2024EMO
2024FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression FeaturesCVPR 2024Face Reenactment, Transformer
2024DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face ReenactmentArXiv 2024CodeProjectface reenactment, diffusion autoencoder, one-shot, controllable
2024One-shot Neural Face Reenactment via Finding Directions in GAN's Latent SpaceIJCVface reenactment, GAN, one-shot, latent space
2024CVTHead: One-shot Controllable Head Avatar with Vertex-feature TransformerWACV 2024
2024VLOGGER: Multimodal Diffusion for EmbodiedArXiv 2024Embodied
2023MaskRenderer: 3D-Infused Multi-Mask Realistic Face ReenactmentArXiv 2023face reenactment, 3D-infused, multi-mask, real-time
2023Controllable One-Shot Face Video Synthesis With Semantic Aware PriorArXiv 2023One-shot Talking Head, Semantic Aware Prior, Controllable Generation, Pose Alignment

Text-driven

YearTitleConference/JournalCode/Proj
2026RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and EditingArXiv 2026
2026IaD: Customizing Video Portraits via Identity-Action DecouplingArXiv 2026
2026High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained ExpressionsArXiv 2026
2026Text-Driven Emotionally Continuous Talking Face GenerationArXiv 2026
2026Dual Diffusion Models for Multi-modal Guided 3D Avatar GenerationArXiv 2026
2026InteractAvatar: Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking AvatarsArXiv 2026Code Project
2026ActAvatar: Temporally-Aware Precise Action Control for Talking AvatarsCVPR 2026Project
2025KeyframeFace: From Text to Expressive Facial KeyframesArXiv 2025Code Project
2025Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided RenderingArXiv 2025Project
2025Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head GenerationArXiv 2025
2025OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style ReplicationArXiv 2025Project
2025EmoAva: When Words Smile: Generating Diverse Emotional Facial Expressions from TextEMNLP 2025Code Project
2024GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec AvatarsArXiv 2024
2024InstructAvatar: Text-Guided Emotion and Motion Control for Avatar GenerationArXiv 2024Code Project
2024FT2TF: First-Person Statement Text-To-Talking Face GenerationWACV 2025
2024Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven ModelsICASSP 2024
2024HeadStudio: Text to Animatable Head Avatars with 3D Gaussian SplattingECCV 2024Code Project
2023Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realismArXiv 2023
2023AgentAvatar: Disentangling Planning, Driving and Rendering for Photorealistic Avatar AgentsArXiv 2023Code Project
2023Text-to-Video: A Two-stage Framework for Zero-shot Identity-agnostic Talking-head GenerationArXivCode
2023Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar SynthesisICML 2023 Workshop
2023TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking StylesArXiv
2022Text2Video: Text-driven Talking-head Video Synthesis with Phonetic DictionaryICASSP 2022Code Project
2021Txt2vid: Ultra-low bitrate compression of talking-head videos via textArXivCode
2021Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationAAAICode

NeRF & 3D Head Avatar

YearTitleConference/JournalCodeProjectKeywords
2026AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation ModelIEEE TVCGProjectblendshape, 3D speech animation, video diffusion, lip-sync
2026CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head GenerationArXiv 2026FLAME, audio-driven, valence-arousal, 3D talking head
2026SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow MatchingArXiv 2026CodeProject3D facial animation, audio-driven, flow matching, facial dynamics
2026ETHead: Generating Expressive 3D Facial Animation and Head Movement from SpeechArXiv 2026CodeProject3D facial animation, speech-driven, expressive motion, head movement
2026KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue LocalizationArXiv 20263D facial animation, speech-driven, style control, visual dubbing
2026From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial AnimationInterspeech 2026CodeProject3D facial animation, speech representation, audio-driven, AVTTS, mesh
2026AutoFaceARKit: Deploying Speech-Driven 3D Facial Animation in Unreal Engine for Production-Ready Digital HumansSIGGRAPH 2026 PostersProjectspeech-driven, 3D facial animation, ARKit blendshape, Unreal Engine, digital human
2026TokTalk: Expressive Real-time Facial Animation from Audio-LLM TokensArXiv 2026FLAME, 3D facial animation, Audio-LLM, real-time, flow matching
2026CapTalk: Text-Guided Stylization and Speech-Driven 3D Head AnimationArXiv 20263D head animation, speech-driven, text-guided style, emotion
2026MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar ReconstructionCVPR 2026Project3D mesh avatar, one-shot reconstruction, animatable head, feed-forward
2026CompHairHead: One-shot Compositional 3D Head Avatars with Deformable HairArXiv 2026CodeProject3D, avatar, head avatar, hair, one-shot, compositional
2026Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head AvatarsArXiv 20263D, avatar, emotion, head avatar
2026PartNerFace: Part-based Neural Radiance Fields for Animatable Facial Avatar ReconstructionArXiv 2026NeRF, part-based, animatable avatar, deformation
20263DRealHead: Few-Shot Detailed Head AvatarArXiv 2026Project3D, avatar, head avatar, few-shot
2026PerformRecast: Expression and Head Pose Disentanglement for Portrait Video EditingCVPR 2026ProjectPortrait Editing, Expression Control, Head Pose Disentanglement
2026TDMM-LM: Bridging Facial Understanding and Animation via Language ModelsArXiv 2026ProjectFacial Animation, Language Models, Text-guided
2026NBAvatar: Neural Billboards Avatars with Realistic Hand-Face InteractionArXiv 2026Hand-Face Interaction, Neural Billboards, Avatar
2026Motion Manipulation via Unsupervised Keypoint Positioning in Face AnimationArXiv 2026Face Animation, Unsupervised Keypoint, Motion Manipulation
2026Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal PredictionArXiv 2026Project3D Head Reconstruction, Multi-View Normal, Feed-forward
2026Toward Fine-Grained Facial Control in 3D Talking Head GenerationArXiv 20263D, Talking Head
2026Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language ModelsArXiv 20263D Facial Animation, Speech-Driven, Omni-modal LLMs, Token-as-query Fusion
2026SuperHead: From Blurry to Believable: Enhancing Low-quality Talking Heads with 3D Generative Priors3DV 2026CodeProject3D, Talking Head, 3DV, Latent
2026Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video ConferenceArXiv 20263D, Talking Head
2026GAT-NeRF: Geometry-Aware-Transformer Enhanced Neural Radiance Fields for High-Fidelity 4D Facial AvatarsArXiv 2026NeRF, Geometry-Aware Transformer, 4D Facial Avatar
2026REFA: Real-time Egocentric Facial Animations for Virtual RealityArXiv 2026Egocentric, Facial Animation, VR, Real-time
2026MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted EnhancementArXiv 20263D, Talking Head, Transformer
2025FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionArXiv 2025Project3D head avatar, partial supervision, transformer, monocular training
2025PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesArXiv 20253D facial animation, speech-driven, emotion
2025Is It Truly Necessary to Process and Fit Minutes-Long Reference Videos for Personalized Talking Face Generation?ArXiv 2025Talking Head, Attention
2025Capturing Head Avatar with Hand Contacts from a Monocular VideoICCV 2025Head Avatar, Hand Contacts, Monocular Video, 3D Reconstruction
2025HRM²Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone ScansSIGGRAPH Asia 2025CodeProjectMobile, Real-Time, Monocular, Avatar
2025MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D AvatarsArXiv 2025Multi-View, Portrait Video, Diffusion, 4D Avatar
20253DiFACE: Synthesizing and Editing Holistic 3D Facial AnimationArXiv 2025CodeProject3D Facial Animation, Diffusion, Editing, Speech-Driven
2025SIE3D: Single-image Expressive 3D Avatar generation via Semantic Embedding and Perceptual Expression LossArXiv 2025ProjectExpressive, Text-Driven, Single Image
2025Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space DiffusionArXiv 2025Dynamic Avatar, Weight-Space Diffusion
2025TeRA: Rethinking Text-guided Realistic 3D Avatar GenerationICCV 2025Text-to-Avatar, Latent Diffusion
2025DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual PerspectiveArXiv 2025Avatar Reconstruction, Video Generation
2025EPSilon: Efficient Point Sampling for Lightening of Hybrid-based 3D Avatar GenerationArXiv 2025CodeEfficient Point Sampling, Hybrid 3D Avatar
2025VisualSpeaker: Visually-Guided 3D Avatar Lip SynthesisICCV 2025 WorkshopVisually-Guided, 3D Avatar, Lip Synthesis
2025GenHMC: Generative Head-Mounted Camera Captures for Photorealistic AvatarsSIGGRAPH Asia 2025ProjectHead-Mounted Camera, Photorealistic, Avatar
2025AvatarMakeup: Realistic Makeup Transfer for 3D Animatable Head AvatarsArXiv 20253D Makeup Transfer, Avatar
2025Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding RouterArXiv 2025Multi-Character, 3D-mask
2025Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal SpaceICME 2025Project3D, Diffusion, Multimodal
2025SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI AgentsArXiv 2025Text-Guided, VLM Agents
2025UMA: Ultra-detailed Human Avatars via Multi-level Surface AlignmentArXiv 2025Ultra-detailed, Surface Alignment
2025Total-Editing: Head Avatar with Editable Appearance, Motion, and LightingArXiv 2025Neural Radiance Fields, Intrinsic Decomposition, Portrait Editing, Motion Control
2025AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion ModelsArXiv 2025CodeAvatar, Human-Centric Animation
2025Eye-See-You: Reverse Pass-Through VR and Head AvatarsIJCAI 2025VR, Head Avatars, Pass-Through
2025MAGE:A Multi-stage Avatar Generator with Sparse ObservationsArXiv 2025Avatar, AR/VR
2025MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric GuidanceArXiv 2025CodeLatent Diffusion, FLAME, 3D Geometric Guidance, Face Reenactment
2025Supervising 3D Talking Head Avatars with Analysis-by-Audio-SynthesisArXiv 2025Project3D, Avatar, Audio-Synthesis
2025EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion ModelsArXiv 2025Project3D facial animation, latent diffusion, emotional expression, speech-driven
2025Better Together: Unified Motion Capture and 3D Avatar ReconstructionArXiv 2025
2025Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorCVPR 2025CodeProject
2025LUCAS: Layered Universal Codec AvatarsArXiv 2025
2025GAS: Generative Avatar Synthesis from a Single ImageICCV 2025CodeProjectSingle Image, 3D Avatar, NeRF, Diffusion
2025MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual LearningAAAI 2025NeRF
2025Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular VideosArXiv 2025Motion Blur, Animatable Avatars
2025TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze AnimationArXiv 2025CodeProject3D Eye Gaze, Speech-Driven, Pluralistic
2025Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID GuidanceArXiv 2025CodeProjectAvatars, Single Image
2025LayerAvatar: Disentangled Clothed Avatar Generation with Layered RepresentationICCV 2025 (Highlight)CodeProject
2025L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the WildICASSP 2025Project
2025Barbie: Text to Barbie-Style 3D AvatarsArXiv 2025CodeProjectText to Avatar, Barbie-Style
2025Hybrid Explicit Representation for Ultra-Realistic Head AvatarsArXiv 2025
2024Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with AdaptersArXiv 2024Co-speech
2024CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsArXiv 2024CodeProjectMulti-View Diffusion
2024StrandHead: Text to Strand-Disentangled 3D Head Avatars Using Hair Geometric PriorsICCVCodeProject
2024SimAvatar: Simulation-Ready Avatars with Layered Hair and ClothingArXiv 2024ProjectNVIDIA, Hair and Clothing
2024DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion ModelsArXiv 2024
2024ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceArXiv 2024CodeProject
2024DanceFusion: A Spatio-Temporal Skeleton Diffusion Transformer for Audio-Driven Dance Motion ReconstructionArXiv 2024Project
2024InstantGeoAvatar: Effective Geometry and Appearance Modeling of Animatable Avatars from Monocular VideoACCV 2024Code
2024EgoAvatar: Egocentric View-Driven and Photorealistic Full-body AvatarsArXiv 2024
2024Towards Native Generative Model for 3D Head AvatarArXiv 2024
2024Stable Video PortraitsECCV 2024ProjectDiffusion
2024LightAvatar: Efficient Head Avatar as Dynamic Neural Light FieldECCV'24 CADLCode
2024FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation ModelArXiv 20243D Facial Animation, Expression Transfer, Foundation Model, Video-driven
2024KMTalk: Speech-Driven 3D Facial Animation with Key Motion EmbeddingECCV 2024Code3D Facial Animation, Key Motion, Speech-Driven
2024Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual RealityIEEE 2024
2024PAV: Personalized Head Avatar from Unstructured Video CollectionECCV 2024Project
2024XHand: Real-time Expressive Hand AvatarArXiv 2024CodeHand
2024Bridging the Gap: Studio-like Avatar Creation from a Monocular Phone CaptureECCV 2024Project
2024Universal Facial Encoding of Codec Avatars from VR HeadsetsSIGGRAPH 2024Facial Encoding, VR Headset, Real-time Animation, 3D Avatar
2024CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple ImagesECCV 2024Code
2024WildAvatar: Web-scale In-the-wild Video Dataset for 3D Avatar CreationArXiv 2024CodeProjectDataset
2024AniFaceDiff: Animating Stylized Avatars via Parametric Conditioned Diffusion ModelsArXiv 2024
2024Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion ModelsNIPS 2024CodeProjectDiffusion
2024Instant 3D Human Avatar Generation using Image Diffusion ModelsArXiv 2024Project
2024Representing Animatable Avatar via Factorized Neural FieldsArXiv 2024
2024Stratified Avatar Generation from Sparse ObservationsCVPR 2024 (Oral)
2024E3Gen: Efficient, Expressive and Editable Avatars GenerationArXiv 2024CodeProject
2024X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar GenerationICML 2024CodeProject
2024GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single ImageCVPR 2024CodeProjectEditing
2024MonoAvatar++: Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table BlendshapesCVPR 2024ProjectBlendshapes
2024MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search SpaceArXiv 2024Project
2024MI-NeRF: Learning a Single Face NeRF from Multiple IdentitiesArXiv 2024ProjectNeRF, Multi-Identity, Face Modeling
2024NECA: Neural Customizable Human AvatarCVPR 2024Code
2024Magic-Me: Identity-Specific Video Customized DiffusionArXiv 2024CodeProject
2024ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance FieldsNIPS 2023CodeProject3D Edit
2024Sketch2NeRF: Multi-view Sketch-guided Text-to-3D GenerationArXiv 2024Text to 3D
2024UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided TexturesArXiv 2024ProjectDiffusion, Avatar
2024High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse RenderingArXiv 2024Codemesh blendshape, neural inverse rendering, face reconstruction
2024FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRFArXiv 2024Code4D face video editor
2024Learning Dense Correspondence for NeRF-Based Face ReenactmentAAAI 2024one-shot multi-view face reenactmen
2024VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style TransferICCV2023
2024AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view VideosECCV 2024CodeProject
2024What You See Is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANsArXiv 2024Project
2023INFAMOUS-NeRF: ImproviNg FAce MOdeling Using Semantically-Aligned Hypernetworks with Neural Radiance FieldsArXiv 2023NeRF, Face Modeling, Hypernetworks, Semantically-Aligned
2023AvatarStudio: High-fidelity and Animatable 3D Avatar Creation from TextArXiv 2023ProjectText-to-3D, NeRF, SMPL, Diffusion Model
20233D Face Style Transfer with a Hybrid Solution of NeRF and Mesh RasterizationArXiv 20233D Face, Style Transfer, NeRF, Mesh
2023HAvatar: High-fidelity Head Avatar via Facial Model Conditioned Neural Radiance FieldArXiv 2023Neural Radiance Field, Facial Model Conditioning, 3D Head Avatar, Expression Control
2023NOFA: NeRF-based One-shot Facial Avatar ReconstructionArXiv 2023NeRF, One-shot, Facial Avatar
2023Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with InstructionsArXiv 2023
2023MA-NeRF: Motion-Assisted Neural Radiance Fields for Face Synthesis from Sparse ImagesArXiv 2023NeRF, Motion-Assisted, Face Synthesis
2023GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face GenerationArXiv 2023CodeProject-
2023GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisICLR 2023CodeProject-
2023SD-NeRF: Towards Lifelike Talking Head Animation via Spatially-adaptive Dual-driven NeRFsIEEE 2023
2022NeRFInvertor: High Fidelity NeRF-GAN Inversion for Single-shot Real Image AnimationArXiv 2022CodeProject-
2022RAD-NeRF: Real-time Neural Talking Portrait SynthesisArXiv 2022CodeProjectInstantNGP
2022Next3D: Generative Neural Texture Rasterization for 3D-Aware Head AvatarsArXiv 2022CodeProject-
2022FNeVR: Neural Volume Rendering for Face AnimationArXiv 2022Code-
20223DFaceShop: Explicitly Controllable 3D-Aware Portrait GenerationArXiv 2022CodeProject-
2022DFRF:Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head SynthesisECCV 2022CodeProject
2022ROME: Realistic One-shot Mesh-based Head AvatarsECCV 2022CodeProject-
2022SSP-NeRF: Semantic-Aware Implicit Neural Audio-Driven Video Portrait GenerationArXiv 2022CodeProject-
2022IMavatar: Implicit Morphable Head Avatars from VideosCVPR 2022CodeProject-
2022HeadNeRF: A Real-time NeRF-based Parametric Head ModelCVPR 2022CodeProject-
2021DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural RenderingArXiv 2021Code-
2021AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisICCV 2021CodeProject-
2021NerFACE: Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar ReconstructionCVPR 2021 OralCodeProject-

3D Gaussian Splatting

YearTitleConference/JournalCodeProjectKeywords
2026PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking HeadsACM MM 20263DGS, phoneme-driven, audio-driven, lip articulation
2026S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single ImageArXiv 20263DGS, single-image, FLAME, animatable head, diffusion
2026SpiD: Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head AvatarsArXiv 20263DGS, single-image, animatable head, real-time
2026DynHair: Head Avatars with Dynamic Explicit HairArXiv 2026CodeProject3DGS, head avatar, dynamic hair, animatable
2026URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head AvatarsECCV 2026Project3DGS, mesh, animatable head, UV-space
2026FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian HeadArXiv 20263DGS, one-shot, 4D head, animatable avatar
2026GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian SplattingArXiv 2026Project3DGS, audio-driven, emotional talking head, blendshapes
2026FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait ImagesECCV 2026CodeProject3DGS, 4D head, FLAME, feed-forward, animatable
2026FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait ImageArXiv 2026Project3DGS, codec avatar, single-image, drivable, feed-forward
2026Edit3DGS: Unified Framework for Dynamic Head Editing via 2D Instruction-Guided Diffusion and 3D Gaussian SplattingSOICT 2025text editing, 3DGS, dynamic head, instruction-guided
2026EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action UnitsArXiv 20263DGS, audio-driven, action units, expression control
2026SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage ReconstructionArXiv 2026Project3DGS, 4D head, FLAME, feed-forward
2026LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic CommunicationArXiv 20263DGS, head avatar, telepresence, controllable
2026SAGE: Self-Learning Expression Deformations for Data-Efficient Gaussian AvatarsArXiv 20263DGS, Gaussian avatar, expression, animatable, few-shot
2026SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian SplittingArXiv 20263DGS, one-shot, animatable head, Gaussian splitting, 3DMM
2026PiG-Avatar: Hierarchical Neural-Field-Guided Gaussian AvatarsArXiv 2026gaussian splatting, neural field, full-body avatar, clothing, hierarchical
2026FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar ReconstructionArXiv 2026Projectgaussian splatting, head avatar, few-shot, FLAME, feed-forward
2026SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head SynthesisArXiv 2026gaussian splatting, talking head, one-shot, facial priors, lip sync
2026HeadsUp: Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View CapturesArXiv 2026Projectgaussian splatting, head reconstruction, multi-view, large-scale, animatable
2026HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting AvatarArXiv 2026gaussian splatting, human mesh recovery, full-body avatar, novel view synthesis, pose refinement
2026High-Fidelity Mobile Avatars with Pruned Local BlendshapesArXiv 2026Projectgaussian splatting, mobile rendering, full-body avatar, blendshapes, real-time
2026SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian SplattingCVPR 2026 (Highlight)Code3DGS, sketch-driven, face editing, real-time, CVPR 2026
2026Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait ImageArXiv 2026single-image, 3DGS head, feed-forward, full-head
2026F3G-Avatar : Face Focused Full-body Gaussian AvatarCVPRW 2026Codefull-body avatar, face-focused, gaussian splatting, multi-view
2026SFGS: Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar ReconstructionArXiv 2026Codefine-grained, structure-aware, gaussian splatting, expressive avatar
2026PhysHead: Simulation-Ready Gaussian Head AvatarsCVPR 2026Project3D Gaussian, Head Avatar, Physics Simulation
2026AvatarPointillist: AutoRegressive 4D Gaussian AvatarizationCVPR 2026CodeProject4D Gaussian, Autoregressive, Avatar
2026Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian AvatarsArXiv 2026CodeBody Model, Gaussian Avatars, Ablation
2026AAP-3DGA: Autoregressive Appearance Prediction for 3D Gaussian AvatarsArXiv 2026Project3D Gaussian, Autoregressive, Appearance Prediction
2026DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular VideoArXiv 20263D Head Avatar, Gaussian Splatting, Monocular Video, Personalized
2026FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head AvatarArXiv 2026Face-and-Hair Avatar, 3D Gaussian, Composable Reconstruction
2026ProgressiveAvatars: Progressive Animatable 3D Gaussian AvatarsCVPR 2026Project3D Gaussian, Progressive, Animatable Avatars
2026Feed-forward Gaussian Registration for Head Avatar Creation and EditingArXiv 2026ProjectGaussian Registration, Head Avatar, Feed-forward
2026Retrieval-Augmented Gaussian Avatars: Improving Expression GeneralizationArXiv 2026Gaussian Splatting, Expression Generalization, Retrieval Augmentation, 3D Avatars
2026Gaussian Wardrobe: Compositional 3D Gaussian Avatars for Free-Form Virtual Try-On3DV 2026CodeProject3D Gaussian, Compositional Avatar, Virtual Try-On
2026LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar AnimationArXiv 2026Kinematic-Space Completion, Expression Control, 3D Gaussian Splatting, Video Diffusion
2026OMG-Avatar: One-shot Multi-LOD Gaussian Head AvatarArXiv 2026Project3D Gaussian, One-Shot, Head Avatar
2026GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar ReconstructionArXiv 2026CodeProjectGeometry-Aware Diffusion, 4D Avatar Reconstruction, 3D Gaussian Splatting, Surface Normals
2026OMEGA-Avatar: One-shot Modeling of 360° Gaussian AvatarsArXiv 2026ProjectOne-Shot Avatar, 360° Full-Head, 3D Gaussian Splatting, Multi-View Feature Splatting
2026OFERA: Blendshape-driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VRArXiv 2026CodeProjectBlendshape Control, Gaussian Avatars, VR Telepresence, Real-time Expression
2026VRGaussianAvatar: Integrating 3D Gaussian Avatars into VRArXiv 2026CodeProject3D Gaussian, VR Avatar, Integration
2026The Gaussian-Head OFL Family: One-Shot Federated Learning from Client Global Statisticsthe International Conference on Learning Representations (ICLR) 2026Gaussian-Head, Federated Learning, One-Shot
2026Splat-Portrait: Generalizing Talking Heads with Gaussian SplattingArXiv 2026CodeProjectGaussian Splatting, 3DGS, Portrait Animation, Talking Head
2026GlassesGB: Controllable 2D GAN-Based Eyewear Personalization for 3D Gaussian Blendshapes Head AvatarsIEEE VR 2026Projectgaussian blendshapes, virtual try-on, eyewear, head avatar
2026CAG-Avatar: Cross-Attention Guided Gaussian Avatars for High-Fidelity Head ReconstructionArXiv 20263D Gaussian Splatting, cross-attention, head reconstruction, drivable avatars
2026FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationICLR 20263D Gaussian Splatting, Head Avatars, Real-Time Animation, Few-Shot Learning
2026FHAvatar: Generalizable and Animatable 3D Full-Head Gaussian Avatar from a Single ImageArXiv 2026CodeProject3D full-head avatar, Gaussian primitives, UV space, single-image reconstruction
2026ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative AdaptationArXiv 2026CodeProjectGaussian Avatar, Test-time Adaptation, Diffusion, Monocular Video
2026UIKA: Fast Universal Head Avatar from Pose-Free ImagesArXiv 2026CodeProjectGaussian Splatting, UV Mapping, Head Avatar, Feed-forward
2026LayerGS: Decomposition and Inpainting of Layered 3D Human Avatars via 2D Gaussian SplattingArXiv 2026Code2D Gaussian Splatting, Layered Avatar, Decomposition
2026GaussianSwap: Animatable Video Face Swapping with 3D Gaussian SplattingArXiv 20263D Gaussian Splatting, Face Swapping, Animatable
2026RelightAnyone: A Generalized Relightable 3D Gaussian Head ModelArXiv 20263D Gaussian Splatting, relightable avatars, single-image fitting, cross-subject generalization
2026CaricatureGS: Exaggerating 3D Gaussian Splatting Faces With Gaussian CurvatureArXiv 2026CodeProject3D Gaussian Splatting, Caricature, Gaussian Curvature
2026GTAvatar: Bridging Gaussian Splatting and Texture Mapping for Relightable and Editable Gaussian AvatarsEurographics 2026CodeProjectGaussian Splatting, Texture Mapping, Relightable Avatars, 3D Reconstruction
2026STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars ReconstructionCVPR 2026CodeProjectGaussian Splatting, 3D Head Avatars, Soft Binding, Temporal Density Control
2025TexAvatars: Hybrid Texel-3D Representations for Stable Rigging of Photorealistic Gaussian Head AvatarsArXiv 2025CodeProject3D Gaussian Splatting, hybrid representation, analytic rigging, UV space
2025FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed DeformationArXiv 2025Project3D avatar, Gaussian Splatting, deformation, reconstruction
2025Instant Expressive Gaussian Head Avatar via 3D-Aware Expression DistillationArXiv 20253D, Gaussian Splatting, Avatar, Attention
2025Gaussian Pixel Codec Avatars: A Hybrid Representation for Efficient RenderingTech Report 2025Gaussian Splatting, head avatar, hybrid representation, efficient rendering
2025AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head AvatarsArXiv 2025CodeProject3D Gaussian Splatting, Animatable Avatars, FLAME, Real-time Rendering
2025EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadCVPR 2026Project3D, Diffusion, Emotional, Talking Head
2025AHA! Animating Human Avatars in Diverse Scenes with Gaussian SplattingArXiv 2025CodeProjectGaussian Splatting, Human Avatar, Scene Animation
2025Densemarks: Learning Canonical Embeddings for Human Heads Images via Point TracksArXiv 2025CodeProjectHead Correspondence, Canonical Embedding, Tracking, Avatar
2025STG-Avatar: Animatable Human Avatars via Spacetime GaussianIROS 2025CodeProjectSpacetime Gaussian, Animatable Avatar, 3DGS
2025Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone ImagesICCV 2025Zero-Shot, 3D Gaussian Avatars, Phone Images
2025Instant Skinned Gaussian Avatars for Web, Mobile and VR ApplicationsSUI 2025CodeProjectReal-Time, Cross-Platform, 3D Avatar, Gaussian Splatting
2025Towards Efficient 3D Gaussian Human Avatar Compression: A Prior-Guided FrameworkArXiv 20253D Gaussian, Human Avatar, Compression
2025ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head AvatarsArXiv 20253D Gaussian Head Avatars, Level of Detail Control, Continuous LOD
2025MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based DynamicsNeurIPS 2025CodeProjectPhysics-Based, 3DGS, Garments
2025FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar ReconstructionArXiv 20252D Gaussian, Mesh-Guided, Avatar
2025Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single ImageArXiv 2025Text-Driven, Single Image, 3DGS
2025PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed ImageArXiv 2025ProjectGaussian, Full-Head Synthesis, One-shot
2025GaussianGAN: Real-Time Photorealistic controllable Human AvatarsFG 20253DGS, Real-Time, Photorealistic
2025Im2Haircut: Single-view Strand-based Hair Reconstruction for Human AvatarsArXiv 2025CodeProjectHair Reconstruction, Gaussian Splatting
2025AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View ImagesArXiv 20253D Gaussian Splatting, Back-Head Generation, Avatar Reconstruction, Spatial Alignment
2025FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained PosesArXiv 2025CodeProject3DGS, Instant, Single Image
2025EAvatar: Expression-Aware Head Avatar Reconstruction with Generative Geometry PriorsArXiv 20253D Gaussian Splatting, expression-aware, deformation-aware, generative priors
2025SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time EditingArXiv 2025CodeProjectGaussian Splatting, 3D Avatar, Texture Editing, FLAME
2025MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial GeometryArXiv 2025Gaussian Avatars, FLAME Meshes, Geometry Refinement
2025MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar ReconstructionICCV 2025CodeProject3D Generative Avatar, Monocular Reconstruction
2025HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsArXiv 2025Project3D Gaussian Avatars, Hair Compositionality, Disentangled Prior, Few-shot Fine-tuning
2025GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarArXiv 2025CodeProjectAdaptive Gaussian Splatting, 3D Head Avatar, Mouth Structure, Deformation Strategy
2025StreamME: Simplify 3D Gaussian Avatar within Live StreamArXiv 2025CodeProject3D Gaussian Splatting, avatar reconstruction, on-the-fly training
2025ScaffoldAvatar: High-Fidelity Gaussian Avatars with Patch ExpressionsSIGGRAPH 2025ProjectHigh-Fidelity, Gaussian Avatars, Patch Expressions
2025HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face AvatarsArXiv 2025CodeProjectHigh-Dimensional, Gaussian Splatting
2025BecomingLit: Relightable Gaussian Avatars with Hybrid Neural ShadingArXiv 2025CodeProject3DGS, Relightable, Neural Shading
2025EVA: Expressive Virtual Avatars from Multi-view VideosSIGGRAPH 2025ProjectAvatar, 3D Gaussian
2025ToonifyGB: StyleGAN-based Gaussian Blendshapes for 3D Stylized Head AvatarsIEEE VR 2026CodeProjectgaussian blendshapes, stylization, toonify, head avatar
2025TeGA: Texture Space Gaussian Avatars for High-Resolution Dynamic Head ModelingSIGGRAPH 2025Project3DGS, Avatar, High-Resolution
2025SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data AugmentationCVPRW 2025CodeProjectSingle Image, 3D Avatar, 3DGS, Video Diffusion
2025GUAVA: Generalizable Upper Body 3D Gaussian AvatarICCV 2025CodeProject3D Gaussian Avatar, Upper Body, SMPLX
2025DNF-Avatar: Distilling Neural Fields for Real-time Animatable Avatar RelightingICCV 2025CodeProjectRelightable Avatar, 2DGS Distillation
2025TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian SplattingCVPR 2025(Highlight🚀)ProjectAR
2025RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head AvatarsArXiv 2025Codegaussian blendshapes, real-time, compact, animatable
20252DGS-Avatar: Animatable High-fidelity Clothed Avatar via 2D Gaussian SplattingICVRV 20242DGS
2025Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head AvatarsArXiv 2025Project
2025LAM: Large Avatar Model for One-shot Animatable Gaussian HeadArXiv 2025CodeProject
2025RFGCA: Relightable Full-Body Gaussian Codec AvatarsArXiv 2025ProjectFull-Body, Avatars
2025PERSE: Personalized 3D Generative Avatars from A Single PortraitCVPR 2025CodeProjectPersonalized, 3DGS, Single Image
2025EGG3D: Generating Editable Head Avatars with 3D Gaussian GANsArXiv 2025CodeProject3DGS
2025FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionArXiv 2025ProjectPose-free, Sparse-view, 3DGS
2025InteractRAGA: Interactive Rendering of Relightable and Animatable Gaussian AvatarsArXiv 2025CodeProjectGaussian Splatting, Relightable Avatars, Interactive Rendering, Pose-driven Animation
2024GraphAvatar: Compact Head Avatars with GNN-Generated 3D GaussiansAAAI 2025CodeGNN-Generated, 3DGS
20243D$^2$-Actor: Learning Pose-Conditioned 3D-Aware Denoiser for Realistic Gaussian Avatar ModelingAAAI 2025CodeProject3DGS
2024GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view DiffusionArXiv 2024CodeProjectDiffusion
2024GASP: Gaussian Avatars with Synthetic PriorsArXiv 2024Project
2024MixedGaussianAvatar: Realistically and Geometrically Accurate Head Avatar via Mixed 2D-3D Gaussian SplattingArXiv 2024Project3DGS, 2D-3D
2024PBDyG: Position Based Dynamic Gaussians for Motion-Aware Clothed Human AvatarsArXiv 2024Clothed Avatar
2024GAST: Sequential Gaussian Avatars with Hierarchical Spatio-temporal ContextArXiv 2024CodeProject
2024Bundle Adjusted Gaussian Avatars DeblurringCVPRCode
2024FATE: Full-head Gaussian Avatar with Textural Editing from Monocular VideoCVPRCodeProject
2024DAGSM: Disentangled Avatar Generation with GS-enhanced MeshCVPR
2024DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D DiffusionArXiv 2024CodeProject
2024Gaussian Déjà-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization AbilitiesWACV 2025CodeProject
2024GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine RepresentationsSIGGRAPH Asia 2024CodeProject🔥Gaussian Splatting
2024DEGAS: Detailed Expressions on Full-Body Gaussian AvatarsArXiv 2024🔥Gaussian Splatting
2024Topology-aware Human Avatars with Semantically-guided Gaussian SplattingArXiv 2024
2024CHASE: 3D-Consistent Human Avatars with Sparse Inputs via Gaussian Splatting and Contrastive LearningArXiv 2024
2024ExAvatar: Expressive Whole-Body 3D Gaussian AvatarECCV 2024CodeProject
2024GEM: Gaussian Eigen Models for Human HeadsCVPR 2025CodeProject
2024EVA: Expressive Gaussian Human Avatars from Monocular RGB VideoArXiv 2024CodeProject
2024NPGA: Neural Parametric Gaussian AvatarsArXiv 2024Project
2024GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image PromptingOn going workTry-ON
20243D Gaussian Blendshapes for Head Avatar AnimationACM SIGGRAPH 2024Gaussian splatting, blendshapes, head avatar, real-time rendering
2024MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head EditingArXiv 2024CodeProject🔥Gaussian Splatting
2024DG-Mesh: Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular VideosArXiv 2024CodeProject🔥Gaussian Splatting
2024HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh PriorArXiv 2024🔥Gaussian Splatting
2024UV Gaussians: Joint Learning of Mesh Deformation and Gaussian Textures for Human Avatar ModelingArXiv 2024Project🔥Gaussian Splatting
2024DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth NormalizationCVPR 2024CodeProject🔥Gaussian Splatting, Sparse-View
2024V3D: Video Diffusion Models are Effective 3D GeneratorsArXiv 2024CodeProject🔥Gaussian Splatting, Video
2024SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian SplattingCVPR 2024CodeProject🔥Gaussian Splatting
2024GEA: Reconstructing Expressive 3D Gaussian Avatar from Monocular VideoArXiv 2024Project🔥Gaussian Splatting, Avatar
2024Consolidating Attention Features for Multi-view Image EditingArXiv 2024Project🔥Gaussian Splatting, Edit
2024GaussianHair: Hair Modeling and Rendering with Light-aware GaussiansArXiv 2024🔥Gaussian Splatting
2024ImplicitDeepfake: Plausible Face-Swapping through Implicit Deepfake Generation using NeRF and Gaussian SplattingArXiv 2024🔥Gaussian Splatting, Deepfake
2024HeadStudio: Text to Animatable Head Avatars with 3D Gaussian SplattingECCVCode🔥Gaussian Splatting, Avatar
2024Rig3DGS: Creating Controllable Portraits from Casual Monocular VideosArXiv 2024ProjectPortraits
20244D Gaussian Splatting: Towards Efficient Novel View Synthesis for Dynamic ScenesArXiv 2024Dynamic Scenes
2024PSAvatar: A Point-based Shape Model for Real-Time Head Avatar Animation with 3D Gaussian SplattingArXiv 20243D Gaussian Splatting, Head Avatar Animation, Point-based Shape Model, Real-time Rendering
2024GaussianBody: Clothed Human Reconstruction via 3d Gaussian SplattingArXiv 2024🔥Gaussian Splatting
2024Gaussian Shadow Casting for Neural CharactersArXiv 2024🔥Gaussian Splatting
2024CoSSegGaussians: Compact and Swift Scene Segmenting 3D Gaussians with Dual Feature FusionArXiv 2024CodeProjectSegmentic
2024AGG: Amortized Generative 3D Gaussians for Single Image to 3DArXiv 2024Project🔥Gaussian Splatting
20244DGen: Grounded 4D Content Generation with Spatial-temporal ConsistencyArXiv 2024CodeProject🔥Gaussian Splatting
2024Human101: Training 100+FPS Human Gaussians in 100s from 1 ViewArXiv 2024CodeProject🔥Gaussian Splatting
2024Deformable 3D Gaussian Splatting for Animatable Human AvatarsArXiv 2024🔥Gaussian Splatting
20243DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian SplattingArXiv 2024CodeProject🔥Gaussian Splatting
2024HHAvatar: Gaussian Head Avatar with Dynamic HairsArXiv 2024CodeProjectHair
2024HeadGaS: Real-Time Animatable Head Avatars via 3D Gaussian SplattingECCV🔥Gaussian Splatting
2024GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansCVPR 2024CodeProject🔥Gaussian Splatting
2024GaussianHead: Impressive 3D Gaussian-based Head Avatars with Dynamic Hybrid Neural FieldArXiv 2024Code🔥Gaussian Splatting
2024MonoGaussianAvatar: Monocular Gaussian Point-based Head AvatarArXiv 2024🔥Gaussian Splatting

Conversational & Dialogue

YearTitleConference/JournalCodeProjectKeywords
2026EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live ChatbotArXiv 2026CodeProjectconversational avatar, lip-sync, empathetic chatbot, live interaction
2026STEER: Steerable Dyadic Head AvatarsArXiv 2026CodeProject3DGS, dyadic, conversational head, Gaussian avatar, listener
2026OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive AvatarsArXiv 2026interactive avatar, multi-turn, listening, streaming, audio-visual
2026MaAI: Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue SystemsICMI 2026CodeProjectlistener nodding, dyadic, avatar dialogue, VAP, real-time
2026Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic PriorsArXiv 2026Projectdyadic, conversational motion, talking head, interaction
2026CHAT: Conversational Human Audio-visual Talking Dialogue GenerationECCV 2026dyadic dialogue, talking face, interactive avatar, audio-visual
2026InterTalk: Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face GenerationArXiv 2026Projectconversational talking face, multi-party, listener feedback, real-time
2026FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational AvatarsArXiv 2026Projectfull-duplex, conversational avatar, facial motion, speech generation
2026MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic ConversationsECCV 2026Projectdyadic, listener, facial animation, talking-and-listening, flow matching
2026InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware AvatarsArXiv 2026interactive avatar, intent-aware, streaming, listening, audio-driven
2026Resonant Minds: Closed-Loop Social Avatars with Theory of MindArXiv 2026CodeProjectsocial avatar, theory of mind, listener, talking face
2026DyaPlex: Full-Duplex Speech-Motion Model for Dyadic InteractionArXiv 2026Projectdyadic interaction, full-duplex, speech-motion, DyaPlex, streaming
2026EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational AgentsArXiv 2026listening-speaking, DiT, rectified flow, real-time avatar
2026Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware KernelsArXiv 2026Projecttalking-listening, interactive, full-duplex, conversational
2026PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic InteractionArXiv 2026Codepolyadic interaction, speaking-listening, multimodal reaction
2026GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy OptimizationArXiv 2026listener, interactive, flow matching, RLHF
2026InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual GuidanceArXiv 2026Projectdyadic, speech-to-video, interactive
2026ECHO: Towards Emotionally Appropriate and Contextually Aware Interactive Head GenerationArXiv 2026ProjectInteractive Head, Emotion, Context-Aware, Dialogue
2026ReactMotion: Generating Reactive Listener Motions from Speaker UtteranceArXiv 2026CodeProjectlistener motion, reactive, body motion, speaker utterance
2026Talking Together: Synthesizing Co-Located 3D Conversations from AudioCVPR 20263D, Conversations, Talking Head, CVPR
2026A²-LLM: An End-to-end Conversational Audio Avatar Large Language ModelArXiv 2026CodeConversational, Avatar, LLM
2026HoverAI: An Embodied Aerial Agent for Natural Human-Drone InteractionArXiv 2026lip-synced avatars, real-time conversational AI, multimodal pipeline
2026RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationArXiv 2026Multi-Turn Conversation, Talking Head
2026Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationArXiv 2026CodeProjectInteractive avatar, Diffusion forcing, Real-time, Preference optimization
2025ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time InteractionArXiv 2025neural talking-head synthesis, content-aware retrieval, real-time interaction, LLM
2025TAVID: Text-Driven Audio-Visual Interactive Dialogue GenerationArXiv 2025text-driven, audio-visual, interactive dialogue, cross-modal mappers
2025ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyArXiv 2025CodeProjectconversational agent, 3D avatar, multimodal interaction, joint language-motion
2025Towards Interactive Intelligence for Digital HumansArXiv 2025interactive intelligence, digital human, multimodal embodiment, real-time interaction
2025Hi-Reco: High-Fidelity Real-Time Conversational Digital HumansCGI 2025real-time, conversational, high-fidelity, 3D avatar
2025Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text GuidanceArXiv 2025CodeProjectDialogue Generation, Speech Language Models
2025UniTalker: Conversational Speech-Visual SynthesisACM MM 2025Conversational, Multimodal, Emotion
2025MaAI: Real-time Generation of Various Types of Nodding for Avatar Attentive Listening SystemICMI 2025CodeReal-time, Nodding Generation, Avatar Interaction
2025MultiTalk: Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationArXiv 2025CodeProjectMulti-Person, Conversational
2025DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsCVPR 2025CodeProject3D, Interaction, Dual-Speaker, Conversations, FLAME
2025VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionArXiv 2025Listener dynamics, 3D dyadic conversation, Expressive control, Multi-modal conditions
2024INFP: Audio-Driven Interactive Head Generation in Dyadic ConversationsArXiv 2024ProjectDyadic Conversations
2024PerceptiveAgent: Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionACL 2024CodeEmpathetic Dialogue
2024MultiDialog: Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationACL 2024CodeProjectDialogue, Face-to-Face Conversation
2023Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationICCV 2023Emotion, LHG
2022DialogueNeRF: Towards Realistic Avatar Face-to-face Conversation Video GenerationArXiv 2022Dialogue, Face-to-face Conversation

Talking Body & Avatar

YearTitleConference/JournalCodeProjectKeywords
2026Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture GenerationArXiv 2026co-speech gesture, object-grounded, diffusion, posture-aware
2026InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture ControlECCV 2026 WorkshopCodeProjectco-speech gesture, streaming, spatial control, diffusion, InteractGesture
2026Super Star: Towards Streaming Real-time Interactive Agents for Digital HumansACM MM 2026CodeProjectco-speech gesture, streaming, digital human, real-time, Super Star
2026Multi-View Face and Gesture Animation with Dynamic GaussiansSCA 2026Project3DGS, upper-body avatar, face and hands, animatable, multi-view
2026StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose AnchoringECCV 2026CodeProjectco-speech gesture, streaming, key-pose, StreamTalk, DiT
2026SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L DatasetECCV 2026co-speech gesture, culture-aware, SICAGE, TED4C-L, diffusion
2026EMOSH: Expressive Motion and Shape Disentanglement for Human AnimationECCV 2026CodeProjectfull-body avatar, video-driven, expression, shape disentanglement, human animation
2026SiGnature: Explicit Motion Diffusion for Stylized Semantic GestureArXiv 2026co-speech gesture, semantic gesture, style, diffusion, SiGnature
2026Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech GesturesArXiv 2026co-speech gesture, text-to-gesture, retrieval, semantic anchors, BEAT2
2026EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsSIGGRAPH 2026CodeProjectfull-body motion, co-speech, streaming, speech and music, 3D character
2026LongCat-Video-Avatar 1.5 Technical ReportArXiv 2026CodeProjectaudio-driven, full-body avatar, lip-sync, long video, human animation
2026DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture GenerationArXiv 2026Projectco-speech gesture, DuoGesture, semantic-beat, dual-stream
2026UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech AvatarsArXiv 2026co-speech avatar, real-time, mixture-of-experts, gesture generation, unified motion
2026PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion RepresentationArXiv 2026Projectco-speech gesture, personalization, VQ-VAE, semantic-aware, motion generation
2026PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen SpeakersArXiv 2026Projectco-speech gesture, personalization, diffusion, style transfer, single-reference
2026Reality Check: How Avatar and Face Representation Affect the Perceptual Evaluation of Synthesized GesturesArXiv 2026co-speech gesture, perceptual evaluation, avatar representation, user study, benchmarking
2026D-Rex : Diffusion Rendering for Relightable Expressive AvatarsArXiv 2026relighting, full-body avatar, diffusion, expressive animation, light stage
2026Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture GenerationFG 2026Projectgesture generation
2026LiveGesture Streamable Co-Speech Gesture Generation ModelArXiv 2026Projectco-speech gesture, streaming, full-body, real-time
2026GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild VideosArXiv 2026full-body avatar, 3D diffusion, photorealistic
2026SentiAvatar: Towards Expressive and Interactive Digital HumansArXiv 2026CodeProjectDigital Human, Expressive, Interactive, Sentiment
2026HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-MatchingArXiv 2026CodeProjectCo-Speech Gesture, Semantic Grounding, Flow-Matching
2026SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation with Neighbor Forcing and ConvKV MemoryArXiv 2026AR diffusion, real-time, hour-scale, human animation
2026MIBURI: Towards Expressive Interactive Gesture SynthesisCVPR 2026CodeProjectGesture Synthesis, Real-Time, LLM-Conditioned, Whole-Body
2026DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture GenerationArXiv 2026ProjectDyadic Gesture, Diffusion Transformer, Multi-Modal, Social Interaction
20263DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action ControlArXiv 2026co-speech gesture, diffusion policy, phoneme-aware, holistic motion
2026[SmoothSync: Dual-Stream Diffusion Transformers for Jitter

Truncated — view the full README on GitHub.

arxiv
audio-driven
paper
synthesis
talking-face-generation
talking-head
talking-head-video-generation

Contributors

Kedreamix

107 commits

satoooh

1 commits

xg-chu

1 commits

Kedreamix/Awesome-Talking-Head-Synthesis

💬 An extensive collection of exceptional resources dedicated to the captivating world of talking face synthesis! ⭐ If you find this repo useful, please give it a star! 🤩

Python

1,555

131 commits

updated Sep 24, 2026

See the code

README

Awesome-Talking-Head-Synthesis

🌐 Website: https://kedreamix.github.io/Awesome-Talking-Head-Synthesis/

The README below starts with runnable open-source systems that do not have a paper, then the complete paper list. For search, filters, bilingual browsing, and category overview, please use the website.

This repository organizes papers, codes, datasets and project pages for Talking Head Synthesis, covering audio-driven avatars, portrait animation, NeRF / 3D heads, 3D Gaussian Splatting, conversational agents, talking body, and more. 👤

Papers for Talking Head Synthesis, released codes collections. ✍️

Most papers are linked to PDFs on "arXiv" or journal/conference websites 📚. However, some papers require an academic license to view 🔐.

🔆 This project Awesome-Talking-Head-Synthesis is ongoing - pull requests are welcome! If you have any suggestions (missing papers, new papers, key researchers or typos), please feel free to edit and submit a PR. You can also open an issue or contact me directly via email. 📩

⭐ If you find this repo useful, please give it a star! 🤩

2026.09 Update 📆

The paper list has grown large enough that browsing only in the README is no longer convenient, so I launched an interactive website:

👉 https://kedreamix.github.io/Awesome-Talking-Head-Synthesis/

You can search papers, browse runnable open-source projects, filter by category and year, switch between English / 中文, and use dark mode. If you find a missing paper or project, newly released code, an accepted venue, or want to suggest a new category, please submit it through the website or GitHub Issues. 🙌

I also added Open-source Projects for runnable talking-head systems that are useful to try, but do not have a paper (for example Linly-Talker and NanoAvatar). SadTalker, MuseTalk, Hallo, and similar works stay in the paper tables.

⭐ If the site helps you, please star the repo — it really motivates continued updates!

2023.12 Update 📆

Thank you to https://github.com/Curated-Awesome-Lists/awesome-ai-talking-heads, I have added some of its contents, such as Tools & Software and Slides & Presentations. 🙏 I hope this will be helpful.😊

If you have any feedback or ideas on extending this aggregated resource, please open an issue or PR - community contributions are vital to advancing this shared knowledge. 🤝

Let's keep pushing forward to recreate ever more realistic digital human faces! 💪 We've come so far but still have a long way to go. With continued research 🔬 and collaboration, I'm sure we'll get there! 🤗

Please feel free to star ⭐ and share this repo if you find it a valuable resource. Your support helps motivate me to keep maintaining and improving it. 🥰 Let me know if you have any other questions!


Open-source Projects

Runnable talking-head systems, apps, and integration frameworks that do not have a paper listing. Research papers with official code stay in the sections below (Audio-driven, Conversational, and so on).

YearProjectCodeResourcesDescription
2026NanoAvatarCodeWeights · APKsOn-device audio-driven talking avatars for Android and local NVIDIA GPUs, with offline APKs.
2026Linly-Talker-StreamCodeFull-duplex, low-latency conversational digital human built on a real-time WebRTC streaming pipeline.
2026CyberVerseCodeSiteSelf-hosted real-time digital-human agent platform with WebRTC, memory, tools, RAG, and optional avatar video.
2025OpenAvatarChatCodeDocs · DemoModular interactive avatar chat with replaceable ASR, LLM, TTS, and avatar backends.
2025LiteAvatarCodeGalleryReal-time CPU audio-to-face 2D chat avatar and an avatar backend for OpenAvatarChat.
2024Duix-AvatarCodeSiteOffline avatar toolkit for appearance and voice cloning plus text- or audio-driven video generation.
2024Ultralight-Digital-HumanCodeFeatherTalkMobile-friendly real-time 2D digital human, with FeatherTalk as its lighter successor.
2024DH_liveCodeMatesXLightweight real-time 2D digital human for web and mobile, followed by the multi-platform MatesX engine.
2023Linly-TalkerCodeWeights · PageConversational digital human WebUI combining LLM, ASR, TTS, voice cloning, and multiple talking-head backends.
2023LiveTalkingCodeSiteReal-time interactive streaming digital human engine supporting Wav2Lip, MuseTalk, ER-NeRF, and other backends.
2023TalkingHeadCodeBrowser JavaScript class for real-time lip-sync using full-body 3D avatars.
2022VU-VRMCodeDemoBrowser-based real-time lip-sync VRM avatar driven by a microphone without a webcam.

Datasets

dataset描述

YearDatasetConference/JournalDownload LinkDescription
2026The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction DatasetArXiv 2026DownloadLarge-scale disentangled human-evaluation challenge for speech-driven co-speech gesture generation (GENEA 2026).
2026REACT 2026: The Fourth Multiple Appropriate Facial Reaction Generation Challenge: Personalised MAFRG and Appropriate EEG Reaction PredictionArXiv 2026DownloadFourth Multiple Appropriate Facial Reaction Generation Challenge for personalized listener reactions in dyadic interaction.
2026MTAVG-Bench 2.0: Diagnosing Failure Modes of Cinematic Expressiveness in Multi-Talker Audio-Video GenerationArXiv 2026DownloadBenchmark diagnosing cinematic-expressiveness failure modes in multi-talker audio-visual generation.
2026AVAPrintDB: Leveraging Avatar Fingerprinting: A Multi-Generator Photorealistic Talking-Head Public Database and BenchmarkArXiv 2026N/AMulti-generator benchmark for talking-head avatar fingerprinting and photorealistic avatar generation evaluation.
2026A Near-Raw Talking-Head Video Dataset for Various Computer Vision TasksArXiv 2026N/ANear-raw talking-head video dataset for computer vision tasks including detection, recognition, and generation.
2026Face-to-Face: A Video Dataset for Multi-Person Interaction ModelingArXiv 2026Download70-hour, 14k-clip dataset of two-person talk-show exchanges for multi-person interaction modeling.
2026SFQAArXivN/AA dataset for singing face generation quality assessment with 5,184 videos generated from 100 photographs and 36 music clips using 12 generation methods.
2025TalkCutsArXiv 2025DownloadA large-scale dataset with 164k clips totaling over 500 hours of human speech videos featuring diverse camera shots and detailed annotations including textual descriptions, 2D keypoints, and 3D SMPL-X motions for multi-shot speech video generation.
2025EmojiBench++IJCV 2025DownloadA comprehensive benchmark for portrait animation comprising diverse portraits, driving videos, and landmark sequences.
2025Multi-human InteractiveArXiv 2025Download12 hours of high-res footage with 2-4 speakers, fine-grained body pose and speech interaction annotations.
2025THQA-10KArXiv 2025DownloadLargest AGTH quality assessment dataset with 10,457 samples from 12 T2I models and 14 talkers.
2025SpeakerVid-5MArXivDownloadLarge-scale dataset with 5.2M video clips (8,743 hours) for audio-visual dyadic interactive virtual human generation, covering monadic talking, listening, and dyadic conversations, with pre-training and SFT subsets.
2025TalkingHeadBenchWACV 2026DownloadComprehensive benchmark for talking-head deepfake detection with multi-model generators.
2025Motion-X++ArXiv 2025N/A19.5M 3D whole-body pose annotations covering 120.5K motion sequences with 80.8K RGB videos.
2024GLCF (MSTF)ArXiv 2024N/AFirst large-scale multi-scenario talking face dataset with 22 audio/video forgery techniques.
2024SAVEEArXiv 2024Download480 British English utterances from 4 male actors expressing 7 emotions.
2024Allo-AVAArXiv 2024N/A~1,250 hours of conversational content for allocentric avatar gesture animation.
2024MMHeadACMMM 2024DownloadLarge-scale multi-modal 3D facial animation dataset with 49 hours of 3D facial motion sequences, speech audios, and hierarchical text annotations for text-induced 3D talking head animation and text-to-3D facial motion generation.
2024DH-FaceVid-1KICCV 2025Download1,200 hours, 270K+ clips from 20K+ individuals with speech audio, keypoints, and text annotations.
2024MultiTalkCVPR 2024Download420+ hours across 20 languages, 293K clips (512x512, 25fps, avg 5.19s duration).
2024THQAArXiv 2024Download800 talking head videos from 8 speech-driven methods with subjective quality assessments.
2023GRIDArXiv 2023Download34 volunteers each speaking 1000 phrases (34K utterances) with 6-word sentence structures.
2023ViCoArXiv 2023DownloadViCo and ViCo-X are datasets for conversational head generation, with ViCo for sentence-level independent talking and listening tasks, and ViCo-X for multi-turn conversational scenarios.
2023TalkingHead-1KHArXiv 2023Download500K video clips with ~80K greater than 512x512 resolution. Only permissive license videos included.
2023CelebVCVPR 2023DownloadIncludes CelebV-Text with 70,000 in-the-wild face video clips for text-to-video generation.
2023MMFace4DArXiv 2023DownloadLarge-scale multi-modal 4D dataset with 35,000+ sequences from 431 subjects (age 15-68).
2022CelebV-HQECCV 2022Download35,666 clips with 15,653 identities, each labeled with 83 facial attributes.
2022MultifaceNeurIPS 2022DownloadHigh-quality multi-view recordings of 13 people with 12K-23K frames per subject at 30fps. 65TB dataset.
2022VFHQCVPRW 2022Download16,000+ high-fidelity clips for video face super-resolution research.
2021HDTFCVPR 2021DownloadHigh-definition Talking-Face Dataset with ~362 videos (15.8 hours) in 720P/1080P resolution.
2020MEADECCV 2020DownloadLarge-scale audio-visual dataset with 60 actors expressing 8 emotions at 3 intensity levels.
2019BIWIArXiv 2019Download3D Audiovisual Corpus of Affective Communication with 40 sentences spoken by 14 subjects.
2019VOCASIGGRAPH 2019Download4D-face dataset with ~29 minutes of 4D face scans and synchronized audio from 12 speakers.
2019CN-CVSArXiv 2019DownloadLarge-scale continuous visual-speech dataset in Mandarin Chinese from TV news and speech shows.
2019FaceForensics++ICCV 2019DownloadLarge-scale dataset for detecting manipulated facial images with over 1.8M images.
2018VoxCeleb2Interspeech 2018DownloadLargest public audio-visual dataset with video URLs and timestamps. Requires 300GB+ storage.
2018LRS2ArXiv 2018DownloadLip reading dataset with videos recorded in diverse settings from BBC television.
2018LRWACCV 2018DownloadDiverse English-speaking dataset from BBC with 1000+ speakers. Each video is 1.16s (29 frames).
2017VoxCeleb1Interspeech 2017DownloadContains over 100,000 utterances for 1,251 celebrities, extracted from YouTube videos.
2017ObamaSetSIGGRAPH 2017DownloadSpecialized audio-visual dataset focused on analyzing visual speech of Barack Obama from weekly address footage.
2014CREMA-DACM TOCC 2014DownloadDiverse dataset with 7,442 clips featuring 91 actors (48 male, 43 female) aged 20-74, expressing six emotions at four intensity levels.

Survey

YearTitleConference/Journal
2026How to Build Digital Humans? From Priors to Photorealistic AvatarsEurographics 2026
2025A Survey of Body and Face Motion: Datasets, Performance Evaluation Metrics and Generative TechniquesArXiv 2025
2025Human Motion Video Generation: A SurveyTPAMI
2025A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and GenerationArXiv 2025
2025Controllable Video Generation: A SurveyArXiv 2025
2025Advancing Talking Head Generation: A Comprehensive Survey of Multi-Modal Methodologies, Datasets, Evaluation Metrics, and Loss FunctionsArXiv 2025
2025Survey of Video Diffusion Models: Foundations, Implementations, and ApplicationsTMLR
2025A Survey on Human Interaction Motion GenerationArXiv 2025
2024Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic SurveyEMNLP 2025
2024Passive Deepfake Detection Across Multi-modalities: A Comprehensive SurveyArXiv 2024
20243D Gaussian Splatting: Survey, Technologies, Challenges, and OpportunitiesArXiv 2024
2024A Comprehensive Survey on Human Video Generation: Challenges, Methods, and InsightsArXiv 2024
2024A Survey on 3D Human Avatar Modeling — From Reconstruction to GenerationArXiv 2024
2024Video Diffusion Models: A SurveyArXiv 2024
2024Deepfake Generation and Detection: A Benchmark and SurveyACM Computing Surveys
2024ADTH-QA: A Comparative Study of Perceptual Quality Metrics for Audio-driven Talking Head VideosArXiv 2024
2024How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a SurveyArXiv 2024
20243D Gaussian as a New Vision Era: A SurveyIEEE TVCG
2024Advances in 3D Generation: A SurveyArXiv 2024
2024A Survey on 3D Gaussian SplattingArXiv 2024
2024Neural Radiance Fields: Past, Present, and FutureArXiv 2024
2023From Pixels to Portraits: A Comprehensive Survey of Talking Head Generation Techniques and ApplicationsArXiv 2023
2023Human Motion Generation: A SurveyTPAMI
2023Human 3D Avatar Modeling with Implicit Neural Representation: A Brief SurveyArXiv 2023
2023Human-Computer Interaction System: A Survey of Talking-Head GenerationIEEE
2023Talking human face generation: A surveyACM
2022Face Generation and Editing with StyleGAN: A SurveyArXiv 2022
2022Deep Learning for Visual Speech Analysis: A SurveyArXiv 2022
2022A Survey on Applications of Digital Human Avatars toward Virtual Co-presenceArXiv 2022
2021A Review of 3D Face Reconstruction From a Single ImageArXiv 2021
2021Deep Person Generation: A Survey from the Perspective of Face, Pose and Cloth SynthesisArXiv 2021
2021AudioVisual Speech Synthesis: A brief literature reviewArXiv 2021
2020What comprises a good talking-head video generation?: A Survey and BenchmarkArXiv 2020

Funny Work


Audio-driven

YearTitleConference/JournalCodeProjectKeywords
2026EfficientSync: Real-Time Lip Synchronization via Deformation-Based Reference Texture MixingArXiv 2026Projectlip synchronization, audio-driven, texture mixing, real-time
2026DynaForcing: Overcoming Dynamic Collapse in Self-Forcing Distillation for Streaming Avatar GenerationACM MM 2026audio-driven, streaming avatar, lip-sync, self-forcing
2026Xemo-Talker: Unlock Emotions Explicitly for Audio-Driven Talking Portrait SynthesisArXiv 2026Codeaudio-driven, emotion control, talking portrait, lip-sync
2026Omni-LiveAvatar: Minute-Level Real-Time Streaming Joint Audio-Video Avatar GenerationArXiv 2026CodeProjectjoint audio-video, streaming avatar, real-time, prompt planning
2026Avatar-Forever: Decoupled Parallel Training for High-Quality Real-Time Infinite AvatarsArXiv 2026CodeProjectaudio-driven, streaming avatar, real-time, long-horizon, distillation
2026Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait AnimationArXiv 2026audio-driven, portrait animation, emotion control, one-shot, real-time
2026GemTalk: Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face GenerationACM MM 2026audio-driven, emotional talking face, blendshape prior, diffusion
2026LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head GenerationArXiv 2026CodeProjectaudio-driven, talking head, real-time, diffusion distillation
2026TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human GenerationArXiv 2026CodeProjectdigital human, audio-video, real-time, talking head, TaoMate
2026AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready AvatarsArXiv 2026CodeProjectaudio-driven, avatar generation, distillation, long-form, real-time
2026SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait AnimationECCV 2026audio-driven, portrait animation, caching, lip-sync, DiT
2026ReFree: Towards Realistic Co-Speech Video Generation via Reward-Free RL and Multilevel Speech GuidanceArXiv 2026speech-driven, portrait animation, lip-sync, co-speech
2026Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip SynchronizationArXiv 2026CodeProjectlip synchronization, audio-driven, autoregressive diffusion, real-time
2026Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait AnimationICME 2026audio-driven, portrait animation, implicit motion, diffusion
2026Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEsCVPR 2026audio-driven, talking portrait, streaming, causal VAE, real-time
2026FreeTalkDiff: IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face GenerationArXiv 2026Codetalking face, diffusion, IP-Adapter, lip-sync
2026CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent PlanningArXiv 2026audio-driven, portrait animation, eye control, lip-sync, DiT
2026Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head GenerationArXiv 2026audio-driven, talking head, test-time adaptation, identity stability
2026HighSync: High-Quality Lip Synchronization via Latent Diffusion ModelsArXiv 2026Codelip synchronization, diffusion model, talking face, high-resolution, audio-driven
2026MoCoTalk: Multi-Conditional Diffusion with Adaptive Router for Controllable Talking Head GenerationArXiv 2026talking head, multi-conditional, diffusion, 3DMM, controllable
2026AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language ModelsArXiv 2026speech-driven, facial animation, blendshape, multimodal LLM, language-assisted
2026TAVR: Generate Your Talking Avatar from Video ReferenceArXiv 2026Projecttalking avatar, video reference, cross-scene, reinforcement learning, identity preservation
2026Talking Slide Avatars: Open-Source Multimodal Communication Approach for TeachingArXiv 2026talking slide avatars, text-to-speech, audio-driven synthesis, educational
2026Personalizing Causal Audio-Driven Facial Motion via Dynamic Multi-modal RetrievalArXiv 2026causal audio-driven, facial motion, multi-modal retrieval, personalization
2026Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference DistillationArXiv 2026Codereal-time avatar, audio-video generation, diffusion, streaming
2026Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion ModelingArXiv 2026CodeProjectjoint audio-video generation, autoregressive diffusion, talking head synthesis
2026EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal CoherenceICMR 2026emotion-aware, talking head generation, lip-sync, temporal coherence
2026Learning Spatial-Temporal Coherent Correlations for Speech-Preserving Facial Expression ManipulationArXiv 2026speech-preserving, facial expression manipulation, spatial-temporal correlation, emotion editing
2026Polyglot: Multilingual Style Preserving Speech-Driven Facial AnimationArXiv 2026Projectmultilingual, speaker style, speech-driven, facial animation
2026TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar GenerationArXiv 2026distillation, audio-driven, avatar, head avatar
2026SEDTalker: Emotion-Aware 3D Facial Animation Using Frame-Level Speech Emotion DiarizationArXiv 2026CodeProjectaudio-driven, 3D, emotion, facial animation
2026C-MET: Cross-Modal Emotion Transfer for Emotion Editing in Talking Face VideoCVPR 2026CodeProjectemotion transfer, emotion editing, talking face, cross-modal
2026MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature FusionArXiv 20263D, Audio-Driven, Multimodal Fusion, Mesh Parameterization
2026EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot PersonalizationCVPR 2026CodeProjectGaussian Splatting, 3DGS, Audio-Driven, Emotion, Few-Shot
2026EARTalking: End-to-end GPT-style Autoregressive Talking Head Synthesis with Frame-wise ControlArXiv 2026Autoregressive, GPT-style, Talking Head
2026FreeTalk: Emotional Topology-Free 3D Talking HeadsArXiv 20263D Talking Heads, Emotional, Topology-Free
2026AvatarForcing: One-Step Streaming Talking Avatars via Local-Future Sliding-Window DenoisingArXiv 2026CodeProjectStreaming Avatar, Real-Time, Diffusion
2026DiFlowDubber: Discrete Flow Matching for Automated Video Dubbing via Cross-Modal Alignment and SynchronizationCVPR 2026 FindingsCodeProjectVideo Dubbing, Flow Matching, Cross-Modal
2026OmniEdit: A Training-free Framework for Lip Synchronization and Audio-Visual EditingArXiv 2026CodeLip Sync, Audio-Visual Editing, Training-Free
2026EmbedTalk: Triplane-Free Talking Head Synthesis using Embedding-Driven Gaussian DeformationPreprintGaussian Splatting, 3DGS, Audio-Driven, Talking Head
2026TempoSyncDiff: Distilled Temporally-Consistent Diffusion for Low-Latency Audio-Driven Talking Head GenerationArXiv 2026CodeProjectDiffusion, Audio-Driven, Talking Head, VAE, Latent
2026UniSync: Towards Generalizable and High-Fidelity Lip Synchronization for Challenging ScenariosArXiv 2026Lip Sync, Pose-Anchored, Generalizable
2026UniTalking: A Unified Audio-Video Framework for Talking Portrait GenerationCVPR 2026Audio-Driven, Portrait Animation, Talking Head, CVPR, Transformer, Attention
2026FlowPortrait: Reinforcement Learning for Audio-Driven Portrait Video GenerationArXiv 2026Audio-Driven, Portrait Animation, Reinforcement Learning, GRPO
2026Narrating For You: Prompt-guided Audio-visual Narrating Face Generation Employing Multi-entangled Latent SpaceWACV 2026CodeAudio-Driven, WACV, Latent
2026VineetVC: Adaptive Video Conferencing Under Severe Bandwidth Constraints Using Audio-Driven Talking-Head ReconstructionArXiv 2026audio-driven, video conferencing, talking head
2026DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video GenerationArXiv 2026CodeProjectAudio-Driven, Transformer, Attention
20263DXTalker: Unifying Identity, Lip Sync, Emotion, and Spatial Dynamics in Expressive 3D Talking AvatarsArXiv 2026CodeProject3D, Emotional, Lip Sync, Avatar, Talking Head, Transformer
2026AUHead: Realistic Emotional Talking Head Generation via Action Units ControlArXiv 2026CodeAction Units, Audio-Driven Generation, Emotion Control, Diffusion Model
2026MOVA: Towards Scalable and Synchronized Video-Audio GenerationArXiv 2026CodeProjectAudio-Driven
2026VedicTHG: Symbolic Vedic Computation for Low-Resource Talking-Head Generation in Educational AvatarsArXiv 2026CodeProjectAvatar, Talking Head
2026SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking HeadsArXiv 2026CodeProjectReal-time, Streaming, Talking Head
2026Asymmetric Hierarchical Anchoring for Audio-Visual Joint RepresentationArXiv 2026Audio-Driven
2026JoyAvatar: Unlocking Highly Expressive Avatars via Harmonized Text-Audio ConditioningArXiv 2026ProjectAudio-Driven, Avatar
2026LPIPS-AttnWav2Lip: Generic Audio-Driven lip synchronization for Talking Head Generation in the WildArXiv 2026CodeProjectLip Sync, Audio-Driven, Talking Head, Latent
2026MIRRORTALK: Forging Personalized Avatars Via Disentangled Style and Hierarchical Motion ControlICASSP 2026Personalized Avatars, Lip Sync, Style Disentanglement, Diffusion Model
2026JUST-DUB-IT: Video Dubbing via Joint Audio-Visual DiffusionArXiv 2026CodeProjectAudio-Visual Diffusion, LoRA, Lip Sync
2026EditYourself: Audio-Driven Generation and Manipulation of Talking Head Videos with Diffusion TransformersArXiv 2026ProjectDiffusion, Audio-Driven, Talking Head
2026UA-3DTalk: Uncertainty-Aware 3D Emotional Talking Face Synthesis with Emotion Prior DistillationICASSP 2026CodeProject3D, Emotional, Talking Head, ICASSP, Attention
2026Audio-Driven Talking Face Generation with Blink Embedding and Hash Grid Landmarks EncodingArXiv 2026Audio-Driven, Talking Head, Transformer
2026SkyReels-V3 Technique ReportArXiv 2026CodeVideo Generation, Audio-Guided, Talking Avatar, Diffusion Transformers
2026THFEM: Exploring Talking Head Models With Adjacent Frame Prior for Speech-Preserving Facial Expression ManipulationACM Trans. MultimediaCodeProjectSpeech-Driven, Talking Head
2026Now You See Me, Now You Don't: A Unified Framework for Expression Consistent Anonymization in Talking Head VideosArXiv 2026Talking Head
2026EditEmoTalk: Controllable Speech-Driven 3D Facial Animation with Continuous Expression EditingArXiv 20263D, Speech-Driven
2026ESGaussianFace: Emotional and Stylized Audio-Driven Facial Animation via 3D Gaussian SplattingArXiv 20263D, Gaussian Splatting, 3DGS, Emotional, Audio-Driven
2026DyStream: Streaming Dyadic Talking Heads Generation via Flow Matching-based Autoregressive ModelArXiv 2026CodeProjectStreaming, Talking Head, Flow Matching
2026SoulX-FlashTalk: Real-Time Infinite Streaming of Audio-Driven Avatars via Self-Correcting Bidirectional DistillationArXiv 2026CodeProjectReal-time, Streaming, Audio-Driven, Avatar, Attention, VAE
2026SyncAnyone: Implicit Disentanglement via Progressive Self-Correction for Lip-Syncing in the wildArXiv 2026ProjectTransformer
2026JoyAvatar-Flash: Real-time and Infinite Audio-Driven Avatar Generation with Autoregressive DiffusionArXiv 2026ProjectDiffusion, Real-time, Audio-Driven, Avatar
2026REST: Diffusion-based Real-time End-to-end Streaming Talking Head Generation via ID-Context Caching and Asynchronous Streaming DistillationArXiv 2026Diffusion, Real-time, Streaming, Talking Head, Latent
2026Lightning Fast Caching-based Parallel Denoising Prediction for Accelerating Talking Head GenerationArXiv 2026Talking Head, Attention, Latent
2025X-Dub: From Inpainting to Editing: A Self-Bootstrapping Framework for Context-Rich Visual DubbingArXiv 2025CodeProjectVisual dubbing, Diffusion Transformer, Self-bootstrapping, Lip sync
2025PTalker: Personalized Speech-Driven 3D Talking Head Animation via Style Disentanglement and Modality AlignmentArXiv 20253D, Speech-Driven, Talking Head, Attention
2025FlashLips: 100-FPS Mask-Free Latent Lip-Sync using Reconstruction Instead of Diffusion or GANsArXiv 2025Diffusion, Transformer, GAN, Latent
2025In-Context Audio Control of Video Diffusion TransformersArXiv 2025Diffusion, Audio-Driven, Transformer, Attention
2025Asynchronous Pipeline Parallelism for Real-Time Multilingual Lip Synchronization in Video Communication SystemsIEEE Big Data 2025lip sync, real-time, multilingual
2025TalkVerse: Democratizing Minute-Long Audio-Driven Video GenerationArXiv 2025CodeProjectAudio-Driven, VAE, Latent
2025VASA-3D: Lifelike Audio-Driven Gaussian Head Avatars from a Single ImageNeurIPS 20253D, Gaussian Splatting, Audio-Driven, Avatar
2025FacEDiT: Unified Talking Face Editing and Generation via Facial Motion InfillingArXiv 2025ProjectTalking Head, Transformer, Attention, Flow Matching
2025JoVA: Unified Multimodal Learning for Joint Video-Audio GenerationArXiv 2025CodeProjectAudio-Driven, Transformer, Attention, GAN
2025Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation ModelTech ReportAudio-Driven, Transformer, Reinforcement Learning
2025STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking PortraitsArXiv 2025CodeProjectDiffusion, Portrait Animation, Talking Head
2025GaussianHeadTalk: Wobble-Free 3D Talking Heads with Audio Driven Gaussian SplattingWACV 20263D, Gaussian Splatting, 3DGS, Audio-Driven, Talking Head
2025UniLS: End-to-End Audio-Driven Avatars for Unified Listening and SpeakingCVPR 2026CodeProjectAudio-Driven, Avatar
2025Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite LengthArXiv 2025CodeProjectReal-time, Streaming, Audio-Driven, Avatar
2025EvalTalker: Learning to Evaluate Real-Portrait-Driven Multi-Subject Talking HumansArXiv 2025Portrait Animation, Talking Head
2025AnyTalker: Scaling Multi-Person Talking Video Generation with Interactivity RefinementArXiv 2025CodeProjectTalking Head, Transformer, Attention
2025AI killed the video star. Audio-driven diffusion model for expressive talking head generationArXiv 2025audio-driven, diffusion, talking head
2025IMTalker: Efficient Audio-driven Talking Face Generation with Implicit Motion TransferArXiv 2025Audio-Driven, Talking Head, Attention, Latent
2025Harmony: Harmonizing Audio and Video Generation through Cross-Task SynergyArXiv 2025Audio-Driven, Attention, Latent
2025StreamingTalker: Audio-driven 3D Facial Animation with Autoregressive Diffusion ModelArXiv 2025Project3D, Diffusion, Streaming, Audio-Driven
2025ConsistTalk: Intensity Controllable Temporally Consistent Talking Head Generation with Diffusion Noise SearchAAAI 2026Diffusion, Talking Head, AAAI, Knowledge Distillation
2025Shared Latent Representation for Joint Text-to-Audio-Visual SynthesisArXiv 2025Audio-Driven, Latent
2025UniAVGen: Unified Audio and Video Generation with Asymmetric Cross-Modal InteractionsArXiv 2025Audio-Driven, Transformer, Attention, Latent
2025See the Speaker: Crafting High-Resolution Talking Faces from Speech with Prior Guidance and Region RefinementTASLP 2025High-Resolution, Talking Faces, Speech-to-Face, Diffusion
2025Learning Disentangled Speech- and Expression-Driven Blendshapes for 3D Talking Face AnimationICXR 2025Blendshapes, FLAME, Disentanglement, 3D Animation
2025MAGIC-Talk: Motion-aware Audio-Driven Talking Face Generation with Customizable Identity ControlArXiv 2025
2025LSF-Animation: Label-Free Speech-Driven Facial Animation via Implicit Feature RepresentationSIGGRAPH Asia 2025CodeLabel-Free, Speech-Driven, Facial Animation, FLAME
2025Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward FeedbackArXiv 2025Diffusion, Audio-Driven, AAAI, Transformer
2025DEMO: Disentangled Motion Latent Flow Matching for Fine-Grained Controllable Talking Portrait SynthesisArXiv 2025Disentangled Motion, Flow Matching, Talking Portrait, Controllable
2025SyncLipMAE: Contrastive Masked Pretraining for Audio-Visual Talking-Face RepresentationArXiv 2025Contrastive Masked Pretraining, Audio-Visual, Talking-Face
2025EGSTalker: Real-Time Audio-Driven Talking Head Generation with Efficient Gaussian DeformationIEEE SMC 2025Real-Time, Audio-Driven, Gaussian Deformation, Talking Head
2025A Bridge from Audio to Video: Phoneme-Viseme Alignment Allows Every Face to Speak Multiple LanguagesArXiv 2025Phoneme-Viseme Alignment, Multilingual TFS, Mixture-of-Experts
2025IASA: Input-Aware Sparse Attention for Real-Time Co-Speech Video GenerationArXiv 2025CodeProjectDiffusion models, co-speech video, real-time, sparse attention
2025Audio Driven Real-Time Facial Animation for Social TelepresenceSIGGRAPH Asia 2025ProjectReal-time, Audio-Driven, SIGGRAPH, Transformer, Latent
2025StableDub: Taming Diffusion Prior for Generalized and Efficient Visual DubbingArXiv 2025ProjectVisual Dubbing, Diffusion, Mamba-Transformer
2025KSDiff: Keyframe-Augmented Speech-Aware Dual-Path Diffusion for Facial AnimationArXiv 2025Keyframe, Diffusion, Dual-Path, Facial Animation
2025SynchroRaMa: Lip-Synchronized and Emotion-Aware Talking Face Generation via Multi-Modal Emotion EmbeddingWACV 2026CodeProjectMulti-Modal, Emotion-Aware, LLM
2025Talking Head Generation via AU-Guided Landmark PredictionArXiv 2025Action Units, Landmark Prediction, Diffusion
2025PGSTalker: Real-Time Audio-Driven Talking Head Generation via 3D Gaussian Splatting with Pixel-Aware Density ControlICONIP 20253DGS, Real-Time, Pixel-Aware, Audio-Driven
2025A Lightweight Pipeline for Noisy Speech Voice Cloning and Accurate Lip Sync SynthesisArXiv 2025lip sync, voice cloning
2025Kling-Avatar: Grounding Multimodal Instructions for Cascaded Long-Duration Avatar Animation SynthesisArXiv 2025ProjectMultimodal Instructions, Avatar Synthesis, Lip Synchronization
2025Think2Sing: Orchestrating Structured Motion Subtitles for Singing-Driven 3D Head AnimationArXiv 2025Singing-Driven, 3D Head, Diffusion
2025EmoCAST: Emotional Talking Portrait via Emotive Text DescriptionArXiv 2025CodeProjectEmotional, Portrait Animation, Talking Head, Attention
2025Wan-S2V: Audio-Driven Cinematic Video GenerationArXiv 2025Cinematic, Audio-Driven, Video Generation
2025Warm Chat: Diffuse Emotion-aware Interactive Talking Head Avatar with Tree-Structured GuidanceArXiv 2025 (Withdrawn)Emotional, Avatar, Talking Head, Transformer, Latent
2025Audio2Face-3D: Audio-driven Realistic Facial Animation For Digital AvatarsArXiv 2025Audio-driven Realistic Facial Animation, Digital Avatars
2025D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head SynthesisECAI 2025Few-Shot, 3DGS, Deformation Fields
2025InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video DubbingArXiv 2025Sparse-Frame Dubbing, Full-Body
2025CEM-Net: Cross-Emotion Memory Network for Emotional Talking Face GenerationArXiv 2025Cross-emotion memory, audio emotion enhancement, expression displacement, lip sync
2025RealTalk: Realistic Emotion-Aware Lifelike Talking-Head SynthesisICCV 2025 WorkshopEmotion, NeRF, VAE
2025FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait AnimationArXiv 2025CodeProjectAudio-Driven, Portrait Animation, Preference Optimization
2025HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head SynthesisArXiv 2025Hybrid Motion, High-Fidelity, Talking Head
2025StableAvatar: Infinite-Length Audio-Driven Avatar Video GenerationArXiv 2025CodeProjectStable Diffusion
2025DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait AnimationArXiv 2025ProjectDiT, Portrait Animation, Speaking Styles
2025READ: Real-time and Efficient Asynchronous Diffusion for Audio-driven Talking Head GenerationArXiv 2025ProjectDiffusion, Real-time, Audio-Driven, Talking Head, Transformer, VAE
2025X-Actor: Emotional and Expressive Long-Range Portrait Acting from AudioArXiv 2025ProjectEmotional Portrait, Long-range, Audio-driven
2025SpA2V: Harnessing Spatial Auditory Cues for Audio-driven Spatially-aware Video GenerationACM MM 2025Spatial Audio, Video Generation, MLLM
2025Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity PreservationArXiv 2025Mask-Free, Identity Preservation, Audio-driven
2025Learning Phonetic Context-Dependent Viseme for Enhancing Speech-Driven 3D Facial AnimationInterspeech 2025CodeProjectPhonetic Context, Viseme, 3D Facial Animation
2025MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided StylizationICCV 2025CodeProjectPersonalized, 3D Facial Animation, Memory
2025JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-SyncArXiv 20253DMM, Joint Learning, Talking Head
2025Livatar-1: Real-Time Talking Heads Generation with Tailored Flow MatchingTechnical ReportProjectreal-time, flow matching, lip-sync
2025ATL-Diff: Audio-Driven Talking Head Generation with Early Landmarks-Guide Noise DiffusionArXiv 2025CodeDiffusion, Landmarks-Guide, Real-time, Identity Preservation
2025MOSPA: Human Motion Generation Driven by Spatial AudioNeurIPS 2025CodeSpatial Audio, Human Motion Generation, Virtual Human
2025M2DAO-Talker: Harmonizing Multi-granular Motion Decoupling and Alternating Optimization for Talking-head GenerationArXiv 2025ProjectMulti-granular Motion, Decoupling, Optimization
2025MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled EmbeddingArXiv 2025CodeMultimodal, 3D Facial Animation, Dynamic Emotions
2025MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head GenerationArXiv 20253DMM, Diffusion Transformer, Temporal Consistency, Blinking Dynamics
2025MoDA: Multi-modal Diffusion Architecture for Talking Head GenerationArXiv 2025CodeProjectMulti-modal, Diffusion, Talking Head Generation
2025FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme CasesArXiv 2025Identity Leakage, Extreme Cases
2025JAM-Flow: Joint Audio-Motion Synthesis with Flow MatchingArXiv 2025Flow Matching, Audio-Motion
2025FIAG: Few-Shot Identity Adaptation for 3D Talking Heads via Global Gaussian FieldArXiv 2025CodeFew-Shot, Global Gaussian Field, 3DGS
2025GGTalker: Talking Head Synthesis with Generalizable Gaussian Priors and Identity-Specific AdaptationICCV 2025CodeProject3D Talking Head, Gaussian Priors, Identity Adaptation
2025SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian SplattingArXiv 2025CodeProject3DGS, Synchronization
2025LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion ModelsArXiv 2025Low-Latency, Real-Time, Interactive
2025TalkingMachines: Real-Time Audio-Driven FaceTime-Style Video via Autoregressive Diffusion ModelsArXiv 2025CodeProjectReal-Time, Autoregressive Diffusion
2025SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion TransformersArXiv 2025video diffusion transformer, multimodal, talking portrait, audio-conditioned
2025Video Editing for Audio-Visual DubbingArXiv 2025Video Editing, Dubbing
2025Wav2Sem: Plug-and-Play Audio Semantic Decoupling for 3D Speech-Driven Facial AnimationCVPR 2025Code3D, Semantic Decoupling
2025MMGT: Motion Mask Guided Two-Stage Network for Co-Speech Gesture Video GenerationArXiv 2025CodeCo-Speech Gesture, Two-Stage
2025FaceEditTalker: Interactive Talking Head Generation with Facial Attribute EditingArXiv 2025ProjectAttribute Editing, Interactive
2025OmniSync: Towards Universal Lip Synchronization via Diffusion TransformersArXiv 2025Lip Sync, Universal, Visual Prosody
2025OT-Talk: Animating 3D Talking Head with Optimal TransportationArXiv 2025FLAME, 3D
2025GenSync: A Generalized Talking Head Framework for Audio-driven Multi-Subject Lip-Sync using 3D Gaussian SplattingCVPRW 20253DGS
2025Model See Model Do: Speech-Driven Facial Animation with Style ControlSIGGRAPH 2025
2025KeySync: A Robust Approach for Leakage-free Lip Synchronization in High ResolutionArXiv 2025CodeProject
2025IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular VideosCVPR 2025Project3D-aware, Video Diffusion
2025Audio-Driven Talking Face Video Generation with Joint Uncertainty LearningArXiv 2025Joint uncertainty learning, Audio-driven talking face, Lip sync, Visual uncertainty
2025DICE-Talk: Disentangle Identity, Cooperate Emotion: Correlation-Aware Emotional Talking Portrait GenerationACM MM 2025Emotional Portrait, Identity Preservation, Emotion Cooperation
2025PASE: Phoneme-Aware Speech Encoder to Improve Lip Sync Accuracy for Talking Head SynthesisArXiv 2025Lip Sync, Phoneme-Aware, Speech Encoder
2025Exploiting Temporal Audio-Visual Correlation Embedding for Audio-Driven One-Shot Talking Head AnimationTMM 2025Talking Head Animation, Temporal Correlation, One-Shot
2025FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion SynthesisArXiv 2025CodeProjecttalking portrait, motion synthesis, video diffusion, audio-visual alignment
2025ACTalker: Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head GenerationArXiv 2025CodeProject
2025MGGTalk:Monocular and Generalizable Gaussian Talking Head AnimationCVPR 2025ProjectOne Shot, 3DGS
2025STSA: Spatial-Temporal Semantic Alignment for Visual DubbingICME 2025CodeSpatial-Temporal Alignment, Semantic Features, Visual Dubbing, Stability
2025Dual Audio-Centric Modality Coupling for Talking Head GenerationArXiv 2025NeRF
2025Audio-Plane: Audio Factorization Plane Gaussian Splatting for Real-Time Talking Head SynthesisArXiv 2025Project3DGS
2025Audio-driven Gesture Generation via Deviation Feature in the Latent SpaceArXiv 2025Gesture
2025Perceptually Accurate 3D Talking Head Generation: New Definitions, Speech-Mesh Representation, and Evaluation MetricsCVPR 2025CodeProject
2025AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion TransformersCVPR 2025ProjectDiT
2025EmoHead: Emotional Talking Head via Manipulating Semantic Expression ParametersArXiv 2025Neural Radiance Fields, Expression Parameters, Emotion Control, Audio-Driven
2025DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion ModelICME 2025Project
2025Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion GenerationCVPR 2025ProjectAutoregressive
2025DiffusionTalker: Efficient and Compact Speech-Driven 3D Talking Head via Personalizer-Guided DistillationICME 2025CodeDiffusion, 3D
2025KDTalker: Unlock Pose Diversity: Accurate and Efficient Implicit Keypoint-based Spatiotemporal Diffusion for Audio-driven Talking PortraitArXiv 2025Codeimplicit keypoint, spatiotemporal diffusion, audio-driven, talking portrait
2025StyleSpeaker: Audio-Enhanced Fine-Grained Style Modeling for Speech-Driven 3D Facial AnimationArXiv 20253D
2025MagicInfinite: Generating Infinite Talking Videos with Your Words and VoiceArXiv 2025Project
2025FREAK: Frequency-modulated High-fidelity and Real-time Audio-driven Talking Portrait SynthesisICMR 2025
2025KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame InterpolationCVPR 2025Diffusion, Long Sequences
2025TexTalk: Towards High-fidelity 3D Talking Avatar with Personalized Dynamic TextureCVPR 2025CodeProjectTexture
2025InsTaG: Learning Personalized 3D Talking Head from Few-Second VideoCVPR 2025CodeProjectFew Shot, 3DGS
2025ARTalk: Speech-Driven 3D Head Animation via Autoregressive ModelSIGGRAPH AsiaCodeProjectAutoregressive, FLAME, 3D
2025FLAP: Fully-controllable Audio-driven Portrait Video Generation through 3D head conditioned diffusion modelArXiv 2025Diffusion
2025Dimitra: Audio-driven Diffusion model for Expressive Talking Head GenerationArXiv 2025Audio-driven, Diffusion model, Motion Diffusion Transformer, Lip sync
2025NeRF-3DTalker: Neural Radiance Field with 3D Prior Aided Audio Disentanglement for Talking Head SynthesisICASSP 2025CodeProject
2025Emotional Face-to-SpeechArXiv 2025Projectemotion, face2speech
2025EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head SynthesisArXiv 2025emotion, 3DGS
2025Identity-Preserving Video Dubbing Using Motion WarpingArXiv 2025Video Dubbing
2025MoEE: Mixture of Emotion Experts for Audio-Driven Portrait AnimationArXiv 2025Audio-driven, Emotion Synthesis, Mixture of Experts, Portrait Animation
2025DEGSTalk: Decomposed Per-Embedding Gaussian Fields for Hair-Preserving Talking Face SynthesisICASSP 2025CodeHair-Preserving
2025UniAvatar: Taming Lifelike Audio-Driven Talking Head Generation with Comprehensive Motion and Lighting ControlArXiv 2025SD, Lighting control
2025FADA: Fast Diffusion Avatar Synthesis with Mixed-Supervised Multi-CFG DistillationCVPR 2025ProjectFast Diffusion 12.5X speedup
2025Synergizing Motion and Appearance: Multi-Scale Compensatory Codebooks for Talking Head Video GenerationCVPR 2025CodeProject
2025V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified FlowICASSP 2025CodeVideo-to-Speech, Speech Decomposition
2025Sonic: Shifting Focus to Global Audio Perception in Portrait AnimationCVPR 2025CodeProjectGlobal Audio Perception, Portrait Animation
2025MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal SamplingArXiv 2025Code
2025Lipschitz-Driven Noise Robustness in VQ-AE for High-Frequency Texture Repair in ID-Specific Talking HeadsArXiv 2025Noise Robustness, VQ-AE, High-Frequency
2025Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion DependencyICLR 2025Project
2025EmoFace: Emotion-Content Disentangled Speech-Driven 3D Talking Face AnimationArXiv 2025emotion,3D
2025DGTalker: Disentangled Generative Latent Space Learning for Audio-Driven Gaussian Talking HeadsICCV 2025ProjectGaussian, Latent Space
2025Talking Head Generation via Viewpoint and Lighting Simulation Based on Global RepresentationACM MM 2025Depth-based
2025PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesACM MM 2025FLAME
2025DisenEmo: Learning disentangled emotional representation from facial motion for 3D talking head generationICIP 2025Disentangled Emotional Representation, 3D Talking Head Generation
2025ExpTalk: Diverse Emotional Expression via Adaptive Disentanglement and Refined Alignment for Speech-Driven 3D Facial AnimationIJCAI 2025Adaptive Disentanglement, Refined Alignment, 3D Facial Animation
2025SyncGaussian: Stable 3D Gaussian-Based Talking Head Generation with Enhanced Lip Sync via Discriminative Speech FeatureIJCAI 2025Stable 3D Gaussian-Based Talking Head Generation, Enhanced Lip Sync, Discriminative Speech Feature
2025SyncAnimation: A Real-Time End-to-End Framework for Audio-Driven Human Pose and Talking Head AnimationArXiv 2025Huaman Pose
2025JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video EditingArXiv 2025Depth, JD work
2024VQTalker: Towards Multilingual Talking Avatars through Facial Motion TokenizationArXiv 2024CodeProjectvisemes, code book
2024LatentSync: Audio Conditioned Latent Diffusion Models for Lip SyncArXiv 2024Diffusion, SyncNet
2024PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head SynthesisAAAI 2025Point Cloud, Gaussian Splatting
2024PortraitTalk: Towards Customizable One-Shot Audio-to-Talking Face GenerationArXiv 2024Diffusion, Attention, One-Shot
2024MEMO: Memory-Guided Diffusion for Expressive Talking Video GenerationArXiv 2024CodeProjectMemory
2024IF-MDM: Implicit Face Motion Diffusion Model for High-Fidelity Realtime Talking Head GenerationArXiv 2024Motion Diffusion Model
2024FLOAT: Generative Motion Latent Flow Matching for Audio-driven Talking PortraitICCV 2025CodeProjectFlow Matching
2024LokiTalk: Learning Fine-Grained and Generalizable Correspondences to Enhance NeRF-based Talking Head SynthesisArXiv 2024ProjectNeRF
2024Ditto: Motion-Space Diffusion for Controllable Realtime Talking Head SynthesisArXiv 2024CodeProjectDiffusion
2024GaussianSpeech: Audio-Driven Gaussian AvatarsArXiv 2024CodeProject3DGS, 3D
2024LetsTalk: Latent Diffusion Transformer for Talking Video SynthesisArXiv 2024
2024EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video DiffusionCVPR 2025ProjectEmotion, Expressive, Diffusion
2024Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait SynthesisArXiv 2024Audio Feature Extraction, Whisper, Real-time processing, Talking portrait synthesis
2024LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion SpaceArXiv 2024Fine-Grained Emotion
2024JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion GenerationArXiv 2024CodeDiffusion, VASA
2024Stereo-Talker: Audio-driven 3D Human Synthesis with Prior-Guided Mixture-of-ExpertsArXiv 2024
2024Audio-Driven Emotional 3D Talking-Head GenerationArXiv 2024Emotion
2024Takin-ADA: Emotion Controllable Audio-Driven Animation with Canonical and Landmark Loss OptimizationArXiv 2024
2024DAWN: Dynamic Frame Avatar with Non-autoregressive Diffusion Framework for Talking Head Video GenerationArXiv 2024CodeNon-autoregressive Diffusion
2024Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image AnimationICLR 2025CodeProjectDiffusion, Hallo
2024Diverse Code Query Learning for Speech-Driven Facial AnimationArXiv 2024
2024TalkinNeRF: Animatable Neural Fields for Full-Body Talking HumansECCVW 2024CodeProjectNeRF
2024JoyHallo: Digital human model for MandarinArXiv 2024CodeProjectDiffusion, Hallo
2024JEAN: Joint Expression and Audio-guided NeRF-based Talking Face GenerationBMVC 2024ProjectNeRF
20243DFacePolicy: Speech-Driven 3D Facial Animation with Diffusion PolicyArXiv 2024
2024LawDNet: Enhanced Audio-Driven Lip Synthesis via Local Affine Warping DeformationArXiv 2024
2024StyleTalk++: A Unified Framework for Controlling the Speaking Styles of Talking HeadsTPAMI 2024
2024ProbTalk3D: Non-Deterministic Emotion Controllable Speech-Driven 3D Facial Animation Synthesis Using VQ-VAESIGGRAPH MIG 2024Code3D
2024DiffTED: One-shot Audio-driven TED Talk Video Generation with Diffusion-based Co-speech GesturesArXiv 2024diffusion
2024EMOdiffhead: Continuously Emotional Control in Talking Head Generation via DiffusionArXiv 2024Diffusion
2024PersonaTalk: Bring Attention to Your Persona in Visual DubbingSIGGRAPH Asia 2024Project
2024KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks GenerationArXiv 2024KAN
2024SVP: Style-Enhanced Vivid Portrait Talking Head Diffusion ModelArXiv 2024Diffusion, Style
2024PoseTalk: Text-and-Audio-based Pose Control and Motion Refinement for One-Shot Talking Head GenerationArXiv 2024ProjectPose Latent Diffusion, Lip Synchronization, Text-Audio Control
2024Mini-Omni: Language Models Can Hear, Talk While Thinking in StreamingTech ReportCodeOmni!!!
2024TalkLoRA: Low-Rank Adaptation for Speech-Driven AnimationArXiv 2024LoRA
2024Meta-Learning Empowered Meta-Face: Personalized Speaking Style Adaptation for Audio-Driven 3D Talking Face AnimationArXiv 2024
2024S^3D-NeRF: Single-Shot Speech-Driven Neural Radiance Field for High Fidelity Talking Head SynthesisECCV 2024
2024DEEPTalk: Dynamic Emotion Embedding for Probabilistic Speech-Driven 3D Face AnimationAAAI 2025CodeProject3D face, FLAME, Emotion
2024High-fidelity and Lip-synced Talking Face Synthesis via Landmark-based Diffusion ModelIEEE TIP
2024Style-Preserving Lip Sync via Audio-Aware Style ReferenceIEEE TIP
2024MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture GenerationArXiv 2024Co-Speech Gesture
2024GLDiTalker: Speech-Driven 3D Facial Animation with Graph Latent Diffusion TransformerArXiv 2024
2024Landmark-guided Diffusion Model for High-fidelity and Temporally Coherent Talking Head GenerationArXiv 2024
2024JambaTalk: Speech-Driven 3D Talking Head Generation Based on Hybrid Transformer-Mamba ModelArXiv 20243D
2024ICCA: Talk Less, Interact Better: Evaluating In-context Conversational Adaptation in Multimodal LLMsCOLM 2024CodeLLM
2024UniTalker: Scaling up Audio-Driven 3D Facial Animation through A Unified ModelArXiv 2024Code
2024DiM-Gesture: Co-Speech Gesture Generation with Adaptive Layer Normalization Mamba-2 frameworkArXiv 2024
2024EmoTalk3D: High-Fidelity Free-View Synthesis of Emotional 3D Talking HeadECCV 2024Project
2024LinguaLinker: Audio-Driven Portraits Animation with Implicit Facial Control EnhancementArXiv 2024
2024Text-based Talking Video Editing with Cascaded Conditional DiffusionArXiv 2024
2024EmoFace: Audio-driven Emotional 3D Face AnimationIEEE VR 2024Code
2024Learning Online Scale Transformation for Talking Head Video GenerationArXiv 2024
2024Audio-driven High-resolution Seamless Talking Head Video Editing via StyleGANArXiv 2024StyleGAN
2024Enhancing Speech-Driven 3D Facial Animation with Audio-Visual Guidance from Lip Reading ExpertInterspeech 2024CodeProject3D
2024RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment NetworkArXiv 2024
2024MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video DatasetInterspeech 2024CodeProject3D, Dataset
2024RITA: A Real-time Interactive Talking Avatars FrameworkArXiv 2024Real-time, Interactive, Talking Avatar
2024NLDF: Neural Light Dynamic Fields for Efficient 3D Talking Head GenerationArXiv 2024NeRF
2024Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image AnimationArXiv 2024CodeProject🔥EMO, Diffusion, Open-source
2024MyTalk: Make Your Actor Talk: Generalizable and High-Fidelity Lip Sync with Motion and Appearance DisentanglementArXiv 2024Project
2024Emotional Conversation: Empowering Talking Faces with Cohesive Expression, Gaze and Pose GenerationArXiv 2024Emotion
2024ControlTalk: Controllable Talking Face Generation by Implicit Facial Keypoints EditingArXiv 2024CodeFace Edit
2024SPEAK: Speech-Driven Pose and Emotion-Adjustable Talking Head GenerationArXiv 2024
2024NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative PriorCVPRW 2024SadTalker+NeRF
2024SwapTalk: Audio-Driven Talking Face Generation with One-Shot Customization in Latent SpaceICASSP 2025
2024AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion EncodingArXiv 2024Code
2024GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian SplattingArXiv 2024🔥Gaussian Splatting
2024CSTalk: Correlation Supervised Speech-driven 3D Emotional Facial Animation GenerationArXiv 2024Emotion
2024GaussianTalker: Real-Time High-Fidelity Talking Head Synthesis with Audio-Driven 3D Gaussian SplattingACMM 2024CodeProject🔥Gaussian Splatting
2024TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian SplattingECCV 2024CodeProject🔥Gaussian Splatting
2024GaussianTalker: Speaker-specific Talking Head Synthesis via 3D Gaussian SplattingACMM 2024Project🔥Gaussian Splatting
2024Learn2Talk: 3D Talking Face Learns from 2D Talking FaceArXiv 2024🔥Gaussian Splatting
2024VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real TimeNeurIPS 2024🔥🔥🔥Awesome,Microsoft
2024EDTalk: Efficient Disentanglement for Emotional Talking Head SynthesisECCV 2024CodeProjectEmotion
2024Talk3D: High-Fidelity Talking Portrait Synthesis via Personalized 3D Generative PriorArXiv 2024CodeProject
2024AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait AnimationArXiv 2024Code🔥🔥🔥Similar to EMO
2024Adaptive Super Resolution For One-Shot Talking-Head GenerationICASSP 2024Code
2024EmoVOCA: Speech-Driven Emotional 3D Talking HeadsArXiv 20243D, VOCA
2024ScanTalk: 3D Talking Heads from Unregistered ScansECCV 2024Code3D
2024FlowVQTalker: High-Quality Emotional Talking Face Generation through Normalizing Flow and QuantizationArXiv 2024Normalizing Flow, Vector-Quantization, Lip Sync, Emotional Talking Faces
2024Style2Talker: High-Resolution Talking Head Generation with Emotion Style and Art StyleArXiv 2024
2024FaceChain-ImagineID: Freely Crafting High-Fidelity Diverse Talking Faces from Disentangled AudioArXiv 2024Code
2024G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal AlignmentArXiv 2024A Generic Framework
2024EMO: Emote Portrait Alive - Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak ConditionsArXiv 2024🔥🔥🔥Amazing, Diffusion
2024Learning Dynamic Tetrahedra for High-Quality Talking Head SynthesisCVPR 2024High-Quality
2024DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion TransformerArXiv 2024Code3D
2024EmoSpeaker: One-shot Fine-grained Emotion-Controlled Talking Face GenerationArXiv 2024CodeProjectEmotion
2024NeRF-AD: Neural Radiance Field with Attention-based Disentanglement for Talking Face SynthesisICASSP 2024CodeProjectAU
2024Real3D-Portrait: One-shot Realistic 3D Talking Portrait SynthesisICLR 2024CodeProject3D, One-Shot, Realistic
2024Dubbing for Everyone: Data-Efficient Visual Dubbing using Neural Rendering PriorsArXiv 2024Projectvisual dubbing, lip sync, neural rendering, data-efficient
2024DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face GenerationArXiv 2024ProjectEmotion
2024VectorTalker: SVG Talking Face Generation with Progressive VectorisationArXiv 2024SVG
2024AE-NeRF: Audio Enhanced Neural Radiance Field for Few Shot Talking Head SynthesisAAAI 2024
2024Mimic: Speaking Style Disentanglement for Speech-Driven 3D Facial AnimationAAAI 2024CodeProject3D
2024DreamTalk: When Expressive Talking Head Generation Meets Diffusion Probabilistic ModelsArXiv 2024CodeProjectDiffusion
2024FaceTalk: Audio-Driven Motion Diffusion for Neural Parametric Head ModelsArXiv 2024CodeProject
2024GMTalker: Gaussian Mixture based Emotional talking video PortraitsArXiv 2024ProjectEmotion
2024GSmoothFace: Generalized Smooth Talking Face Generation via Fine Grained 3D Face GuidanceArXiv 2024CodeProject3D
2024R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer ConditioningArXiv 2024based-RAD-NeRF
2024VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid PriorArXiv 2024Mesh
2024SyncTalk: The Devil😈 is in the Synchronization for Talking Head SynthesisCVPR 2024CodeProject😈Talking Head
2024GAIA: Zero-shot Talking Avatar GenerationArXiv 2024Project😲😲😲
2024AdaMesh: Personalized Facial Expressions and Head Poses for Adaptive Speech-Driven 3D Facial AnimationIEEE Transactions on MultimediaCodeProject3D, Mesh
2024DT-NeRF: Decomposed Triplane-Hash Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisICASSP 2024ER-NeRF
2024EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark ConditioningAAAI 2025🔥阿里
2024Pose-Aware 3D Talking Face Synthesis using Geometry-guided Audio-Vertices AttentionIEEE 2024
2023DiffDub: Person-generic Visual Dubbing Using Inpainting Renderer with Diffusion Auto-encoderICASSP 2024CodeProjectvisual dubbing, diffusion, inpainting, person-generic
2023Towards Streaming Speech-to-Avatar SynthesisArXiv 2023Streaming Synthesis, Articulatory Inversion, Real-time, Speech-driven
2023OSM-Net: One-to-Many One-shot Talking Head Generation with Spontaneous Head MotionsArXiv 2023One-shot Talking Head, Head Motions, One-to-Many Mapping, Audio-driven
2023EAT: Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationICCV 2023CodeProject-
2023Audio-Driven Dubbing for User Generated Contents via Style-Aware Semi-Parametric SynthesisTCSVT 2023
2023Implicit Identity Representation Conditioned Memory Compensation Network for Talking Head Video GenerationICCV 2023-
2023Emotional Talking Head Generation based on Memory-Sharing and Attention-Augmented NetworksInterSpeech 2023Emotion
2023StyleSync: High-Fidelity Generalized and Personalized Lip Sync in Style-based GeneratorCVPR 2023CodeProject-
2023High-fidelity Generalized Emotional Talking Face Generation with Multi-modal Emotion Space LearningCVPR 2023Emotion
2023FONT: Flow-guided One-shot Talking Head Generation with Natural Head MotionsICME 2023Natural Head Motions, Flow-guided, Audio-driven Pose Prediction, One-shot Talking Head
2023DAE-Talker: High Fidelity Speech-Driven Talking Face Generation with Diffusion AutoencoderACMM 2023🔥Diffusion
2023TalkLip: Seeing What You Said - Talking Face Generation Guided by a Lip Reading ExpertCVPR 2023
2023OTAvatar: One-shot Talking Face Avatar with Controllable Tri-plane RenderingCVPR 2023CodeTri-plane Rendering, One-shot Avatar, Controllable, 3D Consistency
2023Emotionally Enhanced Talking Face GenerationArXiv 2023Emotion
2023EmoTalk: Speech-Driven Emotional Disentanglement for 3D Face AnimationICCV 2023CodeProject3D, Emotion
2023READ Avatars: Realistic Emotion-controllable Audio Driven AvatarsArXiv 2023-
2023OPT: One-shot Pose-Controllable Talking Head GenerationICASSP 2023pose control, identity preservation, audio feature disentanglement
2023DiffTalk: Crafting Diffusion Models for Generalized Talking Head SynthesisCVPR 2023CodeProject🔥Diffusion
2023CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion PriorCVPR 2023CodeProject3D, codebook
2023StyleTalk: One-shot Talking Head Generation with Controllable Speaking StylesAAAI 2023CodeStyle
2023SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face AnimationCVPR 2023CodeProject3D, Single Image
2023Efficient Region-Aware Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisICCV 2023Tri-plane
2023LipNeRF: What is the right feature space to lip-sync a NeRF?FG 2023Wav2lip
2023ToonTalker: Cross-Domain Face ReenactmentICCV 2023-
2023EMMN: Emotional Motion Memory Network for Audio-driven Emotional Talking Face GenerationICCV 2023Emotion
2023Facediffuser: Speech-driven 3d facial animation synthesis using diffusionACM SIGGRAPH MIG 2023🔥Diffusion, 3D
2023DINet: Deformation Inpainting Network for Realistic Face Visually Dubbing on High Resolution VideoAAAI 2023
2023Diffused Heads: Diffusion Models Beat GANs on Talking-Face GenerationArXiv 2023🔥Diffusion
2022Memories are One-to-Many Mapping Alleviators in Talking Face GenerationArXiv 2022Project-
2022Masked Lip-Sync Prediction by Audio-Visual Contextual Exploitation in TransformersSIGGRAPH Asia 2022-
2022VideoReTalking: Audio-based Lip Synchronization for Talking Head Video Editing In the WildSIGGRAPH 2022
2022Progressive Disentangled Representation Learning for Fine-Grained Controllable Talking Head SynthesisArXiv 2022disentangled representation, contrastive learning, multi-motion control
2022SPACEx 🚀: Speech-driven Portrait Animation with Controllable ExpressionArXiv 2022Project-
2022Pre-Avatar: An Automatic Presentation Generation Framework Leveraging Talking AvatarICTAI2022talking avatar, presentation
2022EAMM: One-Shot Emotional Talking Face via Audio-Based Emotion-Aware Motion ModelSIGGRAPH 2022Emotion
2022Emotion-Controllable Generalized Talking Face GenerationIJCAI 2022emotion control, graph convolutional network, geometry-aware
2022StyleHEAT: One-Shot High-Resolution Editable Talking Face Generation via Pre-trained StyleGANArXiv 2022StyleGAN, high-resolution, one-shot, lip sync
2022Expressive Talking Head Generation with Granular Audio-Visual ControlCVPR 2022-
2022Talking Face Generation with Multilingual TTSCVPR 2022-
2021One-shot Talking Face Generation from Single-speaker Audio-Visual Correlation LearningAAAI 2022one-shot, audio-visual correlation, keypoint-based motion, lip sync
2021Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisACM MM 2021-
2021Talking Head Generation with Audio and Speech Related Facial Action UnitsBMVC 2021AU
2021Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head MotionArXiv 2021Audio-driven, Talking-head, Head Motion, Keypoint-based Motion
20213D-TalkEmo: Learning to Synthesize 3D Emotional Talking HeadArXiv 20213D Talking Head, Emotion, Geometry Map, Audio-driven
2021PC-AVS: Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual RepresentationCVPR 2021-
2021MakeItTalk: Speaker-Aware Talking-Head AnimationSIGGRAPH Asia 2020Speaker-Aware, Audio-Driven, Facial Landmarks, Photorealistic
2021Speech2Talking-Face: Inferring and Driving a Face with Synchronized Audio-Visual RepresentationIJCAI 2021-
2021Audio-Driven Emotional Video PortraitsCVPR 2021Emotion
2021IATS: Imitating Arbitrary Talking Style for Realistic Audio-Driven Talking Face SynthesisACM Multimedia 2021-
2020Multi Modal Adaptive Normalization for Audio to Video GenerationArXiv 2020Audio-to-Video, Multi-Modal Adaptive Normalization, Facial Video Generation, Keypoint Heatmap
2020A Lip Sync Expert Is All You Need for Speech to Lip Generation In The WildACM Multimedia 2020-
2020Talking-head Generation with Rhythmic Head MotionECCV 2020-
2020Speaker-Aware Talking-Head AnimationSIGGRAPH Asia 2020-
2020Neural Voice Puppetry: Audio-driven Facial ReenactmentECCV 2020-
2020Realistic Speech-Driven Facial Animation with GANsIJCV 2020-
2020A Large-scale Audio-visual Dataset for Emotional Talking-face GenerationECCV 2020-
2019Talking Face Generation by Adversarially Disentangled Audio-Visual RepresentationAAAI 2019-
2019Hierarchical Cross-modal Talking Face Generation with Dynamic Pixel-wise LossCVPR 2019-
2018Audio-Driven Animator-Centric Speech AnimationSIGGRAPH 2018-
2018Lip Movements Generation at a GlanceECCV 2018-
2017You Said That? Synthesising Talking Faces From AudioBMVC 2019-
2017Synthesizing Obama: Learning Lip Sync From AudioSIGGRAPH 2017-
2017Audio-Driven Facial Animation by Joint End-to-End Learning of Pose and EmotionSIGGRAPH 2017-
2017A Deep Learning Approach for Generalized Speech AnimationSIGGRAPH 2017-

Portrait Animation

YearTitleConference/JournalCodeProjectKeywords
2026TongueReenact: Geometry-Anchored Tongue Synthesis for Face ReenactmentArXiv 2026face reenactment, tongue, video-driven, diffusion, cross-identity
2026ViDS: Video Diffusion Shader using 3D Face TrackingArXiv 2026Projectportrait animation, video-driven, 3DMM, video diffusion, face tracking
2026MagPlus: Bridging Micro-to-Regular Facial Expressions through Learnable MagnificationArXiv 2026micro-expression, facial animation, motion magnification, portrait animation
2026Loki: Representation over Architecture for Diffusion-Based Portrait AnimationArXiv 2026face reenactment, video-driven, portrait animation, expression, head pose
2026PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial ReenactmentCVPR 2026face reenactment, disentanglement, real-time, CVPR 2026
2026MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face GenerationCVPR 2026CodeProjectDiffusion Transformer, Multimodal, Face Generation
2026FG-Portrait: 3D Flow Guided Editable Portrait AnimationCVPR 2026portrait animation, 3D flow, CVPR
2026MoCha:End-to-End Video Character Replacement without Structural GuidanceArXiv 2026Talking Head
2025Knot Forcing: Taming Autoregressive Video Diffusion Models for Real-time Infinite Interactive Portrait AnimationArXiv 2025CodeProjectDiffusion, Real-time, Portrait Animation, Attention
2025SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait AnimationArXiv 2025Portrait Animation, ICASSP, Attention
2025FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent PredictionArXiv 2025Portrait Animation, Transformer, Latent
2025DeX-Portrait: Disentangled and Expressive Portrait Animation via Explicit and Latent Motion RepresentationsArXiv 2025ProjectPortrait Animation, Disentangled, Expressive
2025FactorPortrait: Controllable Portrait Animation via Disentangled Expression, Pose, and ViewpointArXiv 2025ProjectPortrait Animation, Transformer, Latent
2025PersonaLive! Expressive Portrait Image Animation for Live StreamingArXiv 2025Streaming, Portrait Animation
2025Beat on Gaze: Learning Stylized Generation of Gaze and Head DynamicsArXiv 2025Gaze Control, Head Motion, Style-Aware, 3D
2025Navigating Large-Pose Challenge for High-Fidelity Face Reenactment with Video Diffusion ModelArXiv 2025Face Reenactment, Large-Pose, Video Diffusion
2025Follow Your Motion: A Generic Temporal Consistency Portrait Editing Framework with Trajectory GuidanceArXiv 2025
2025HunyuanPortrait: Implicit Condition Control for Enhanced Portrait AnimationCVPR 2025Hunyuan
2025Playmate: Flexible Control of Portrait Animation via 3D-Implicit Space Guided DiffusionArXiv 2025CodeDiffusion, 3D
2025MobilePortrait: Real-Time One-Shot Neural Head Avatars on Mobile DevicesCVPR 2025100+fps
2025GOES: 3D Gaussian-based One-shot Head Animation with Any Emotion and Any StyleACM MM 2025One-Shot, 3DGS
2024GoHD: Gaze-oriented and Highly Disentangled Portrait Animation with Rhythmic Poses and Realistic ExpressionAAAI 2025Gaze-oriented
2024Avatar Concept Slider: Manipulate Concepts In Your Human Avatar With Fine-grained ControlArXiv 2024
2024G3FA: Geometry-guided GAN for Face AnimationBMVC 2024
2024Anchored Diffusion for Video Face ReenactmentArXiv 2024Face Reenactment, Anchored Diffusion
2024One-Shot Pose-Driving Face Animation PlatformArXiv 2024One-Shot, Pose-Driving, Face Animation, Talking Head
2024V-Express: Conditional Dropout for Progressive Training of Portrait Video GenerationTech Report🔥EMO, Diffusion, Open-source
2024EMOPortraits: Emotion-enhanced Multimodal One-shot Head AvatarsArXiv 2024CodeEmotional, One Shot, Cross-Driving
2024EMOPortraits: Emotion-enhanced Multimodal One-shot Head AvatarsArXiv 2024EMO
2024FSRT: Facial Scene Representation Transformer for Face Reenactment from Factorized Appearance, Head-pose, and Facial Expression FeaturesCVPR 2024Face Reenactment, Transformer
2024DiffusionAct: Controllable Diffusion Autoencoder for One-shot Face ReenactmentArXiv 2024CodeProjectface reenactment, diffusion autoencoder, one-shot, controllable
2024One-shot Neural Face Reenactment via Finding Directions in GAN's Latent SpaceIJCVface reenactment, GAN, one-shot, latent space
2024CVTHead: One-shot Controllable Head Avatar with Vertex-feature TransformerWACV 2024
2024VLOGGER: Multimodal Diffusion for EmbodiedArXiv 2024Embodied
2023MaskRenderer: 3D-Infused Multi-Mask Realistic Face ReenactmentArXiv 2023face reenactment, 3D-infused, multi-mask, real-time
2023Controllable One-Shot Face Video Synthesis With Semantic Aware PriorArXiv 2023One-shot Talking Head, Semantic Aware Prior, Controllable Generation, Pose Alignment

Text-driven

YearTitleConference/JournalCode/Proj
2026RAGMesh with FaME-G2E: Long-Form Text-Driven 3D Face Generation and EditingArXiv 2026
2026IaD: Customizing Video Portraits via Identity-Action DecouplingArXiv 2026
2026High-Fidelity 3D Facial Avatar Synthesis with Controllable Fine-Grained ExpressionsArXiv 2026
2026Text-Driven Emotionally Continuous Talking Face GenerationArXiv 2026
2026Dual Diffusion Models for Multi-modal Guided 3D Avatar GenerationArXiv 2026
2026InteractAvatar: Making Avatars Interact: Towards Text-Driven Human-Object Interaction for Controllable Talking AvatarsArXiv 2026Code Project
2026ActAvatar: Temporally-Aware Precise Action Control for Talking AvatarsCVPR 2026Project
2025KeyframeFace: From Text to Expressive Facial KeyframesArXiv 2025Code Project
2025Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided RenderingArXiv 2025Project
2025Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head GenerationArXiv 2025
2025OmniTalker: Real-Time Text-Driven Talking Head Generation with In-Context Audio-Visual Style ReplicationArXiv 2025Project
2025EmoAva: When Words Smile: Generating Diverse Emotional Facial Expressions from TextEMNLP 2025Code Project
2024GenCA: A Text-conditioned Generative Model for Realistic and Drivable Codec AvatarsArXiv 2024
2024InstructAvatar: Text-Guided Emotion and Motion Control for Avatar GenerationArXiv 2024Code Project
2024FT2TF: First-Person Statement Text-To-Talking Face GenerationWACV 2025
2024Text-Driven Talking Face Synthesis by Reprogramming Audio-Driven ModelsICASSP 2024
2024HeadStudio: Text to Animatable Head Avatars with 3D Gaussian SplattingECCV 2024Code Project
2023Neural Text to Articulate Talk: Deep Text to Audiovisual Speech Synthesis achieving both Auditory and Photo-realismArXiv 2023
2023AgentAvatar: Disentangling Planning, Driving and Rendering for Photorealistic Avatar AgentsArXiv 2023Code Project
2023Text-to-Video: A Two-stage Framework for Zero-shot Identity-agnostic Talking-head GenerationArXivCode
2023Ada-TTA: Towards Adaptive High-Quality Text-to-Talking Avatar SynthesisICML 2023 Workshop
2023TalkCLIP: Talking Head Generation with Text-Guided Expressive Speaking StylesArXiv
2022Text2Video: Text-driven Talking-head Video Synthesis with Phonetic DictionaryICASSP 2022Code Project
2021Txt2vid: Ultra-low bitrate compression of talking-head videos via textArXivCode
2021Write-a-speaker: Text-based Emotional and Rhythmic Talking-head GenerationAAAICode

NeRF & 3D Head Avatar

YearTitleConference/JournalCodeProjectKeywords
2026AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation ModelIEEE TVCGProjectblendshape, 3D speech animation, video diffusion, lip-sync
2026CETalk: Continuous Valence-Arousal Control for Audio-Driven 3D Talking Head GenerationArXiv 2026FLAME, audio-driven, valence-arousal, 3D talking head
2026SubtleTalk: Generating Controllable Weakly-correlated Facial Dynamics for 3D Talking Heads via Residual Flow MatchingArXiv 2026CodeProject3D facial animation, audio-driven, flow matching, facial dynamics
2026ETHead: Generating Expressive 3D Facial Animation and Head Movement from SpeechArXiv 2026CodeProject3D facial animation, speech-driven, expressive motion, head movement
2026KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue LocalizationArXiv 20263D facial animation, speech-driven, style control, visual dubbing
2026From Tokens to Faces: Investigating Discrete Speech Representations for 3D Facial AnimationInterspeech 2026CodeProject3D facial animation, speech representation, audio-driven, AVTTS, mesh
2026AutoFaceARKit: Deploying Speech-Driven 3D Facial Animation in Unreal Engine for Production-Ready Digital HumansSIGGRAPH 2026 PostersProjectspeech-driven, 3D facial animation, ARKit blendshape, Unreal Engine, digital human
2026TokTalk: Expressive Real-time Facial Animation from Audio-LLM TokensArXiv 2026FLAME, 3D facial animation, Audio-LLM, real-time, flow matching
2026CapTalk: Text-Guided Stylization and Speech-Driven 3D Head AnimationArXiv 20263D head animation, speech-driven, text-guided style, emotion
2026MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar ReconstructionCVPR 2026Project3D mesh avatar, one-shot reconstruction, animatable head, feed-forward
2026CompHairHead: One-shot Compositional 3D Head Avatars with Deformable HairArXiv 2026CodeProject3D, avatar, head avatar, hair, one-shot, compositional
2026Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head AvatarsArXiv 20263D, avatar, emotion, head avatar
2026PartNerFace: Part-based Neural Radiance Fields for Animatable Facial Avatar ReconstructionArXiv 2026NeRF, part-based, animatable avatar, deformation
20263DRealHead: Few-Shot Detailed Head AvatarArXiv 2026Project3D, avatar, head avatar, few-shot
2026PerformRecast: Expression and Head Pose Disentanglement for Portrait Video EditingCVPR 2026ProjectPortrait Editing, Expression Control, Head Pose Disentanglement
2026TDMM-LM: Bridging Facial Understanding and Animation via Language ModelsArXiv 2026ProjectFacial Animation, Language Models, Text-guided
2026NBAvatar: Neural Billboards Avatars with Realistic Hand-Face InteractionArXiv 2026Hand-Face Interaction, Neural Billboards, Avatar
2026Motion Manipulation via Unsupervised Keypoint Positioning in Face AnimationArXiv 2026Face Animation, Unsupervised Keypoint, Motion Manipulation
2026Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal PredictionArXiv 2026Project3D Head Reconstruction, Multi-View Normal, Feed-forward
2026Toward Fine-Grained Facial Control in 3D Talking Head GenerationArXiv 20263D, Talking Head
2026Ex-Omni: Enabling 3D Facial Animation Generation for Omni-modal Large Language ModelsArXiv 20263D Facial Animation, Speech-Driven, Omni-modal LLMs, Token-as-query Fusion
2026SuperHead: From Blurry to Believable: Enhancing Low-quality Talking Heads with 3D Generative Priors3DV 2026CodeProject3D, Talking Head, 3DV, Latent
2026Lightweight High-Fidelity Low-Bitrate Talking Face Compression for 3D Video ConferenceArXiv 20263D, Talking Head
2026GAT-NeRF: Geometry-Aware-Transformer Enhanced Neural Radiance Fields for High-Fidelity 4D Facial AvatarsArXiv 2026NeRF, Geometry-Aware Transformer, 4D Facial Avatar
2026REFA: Real-time Egocentric Facial Animations for Virtual RealityArXiv 2026Egocentric, Facial Animation, VR, Real-time
2026MANGO:Natural Multi-speaker 3D Talking Head Generation via 2D-Lifted EnhancementArXiv 20263D, Talking Head, Transformer
2025FlexAvatar: Learning Complete 3D Head Avatars with Partial SupervisionArXiv 2025Project3D head avatar, partial supervision, transformer, monocular training
2025PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional StylesArXiv 20253D facial animation, speech-driven, emotion
2025Is It Truly Necessary to Process and Fit Minutes-Long Reference Videos for Personalized Talking Face Generation?ArXiv 2025Talking Head, Attention
2025Capturing Head Avatar with Hand Contacts from a Monocular VideoICCV 2025Head Avatar, Hand Contacts, Monocular Video, 3D Reconstruction
2025HRM²Avatar: High-Fidelity Real-Time Mobile Avatars from Monocular Phone ScansSIGGRAPH Asia 2025CodeProjectMobile, Real-Time, Monocular, Avatar
2025MVP4D: Multi-View Portrait Video Diffusion for Animatable 4D AvatarsArXiv 2025Multi-View, Portrait Video, Diffusion, 4D Avatar
20253DiFACE: Synthesizing and Editing Holistic 3D Facial AnimationArXiv 2025CodeProject3D Facial Animation, Diffusion, Editing, Speech-Driven
2025SIE3D: Single-image Expressive 3D Avatar generation via Semantic Embedding and Perceptual Expression LossArXiv 2025ProjectExpressive, Text-Driven, Single Image
2025Hyper Diffusion Avatars: Dynamic Human Avatar Generation using Network Weight Space DiffusionArXiv 2025Dynamic Avatar, Weight-Space Diffusion
2025TeRA: Rethinking Text-guided Realistic 3D Avatar GenerationICCV 2025Text-to-Avatar, Latent Diffusion
2025DevilSight: Augmenting Monocular Human Avatar Reconstruction through a Virtual PerspectiveArXiv 2025Avatar Reconstruction, Video Generation
2025EPSilon: Efficient Point Sampling for Lightening of Hybrid-based 3D Avatar GenerationArXiv 2025CodeEfficient Point Sampling, Hybrid 3D Avatar
2025VisualSpeaker: Visually-Guided 3D Avatar Lip SynthesisICCV 2025 WorkshopVisually-Guided, 3D Avatar, Lip Synthesis
2025GenHMC: Generative Head-Mounted Camera Captures for Photorealistic AvatarsSIGGRAPH Asia 2025ProjectHead-Mounted Camera, Photorealistic, Avatar
2025AvatarMakeup: Realistic Makeup Transfer for 3D Animatable Head AvatarsArXiv 20253D Makeup Transfer, Avatar
2025Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding RouterArXiv 2025Multi-Character, 3D-mask
2025Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal SpaceICME 2025Project3D, Diffusion, Multimodal
2025SmartAvatar: Text- and Image-Guided Human Avatar Generation with VLM AI AgentsArXiv 2025Text-Guided, VLM Agents
2025UMA: Ultra-detailed Human Avatars via Multi-level Surface AlignmentArXiv 2025Ultra-detailed, Surface Alignment
2025Total-Editing: Head Avatar with Editable Appearance, Motion, and LightingArXiv 2025Neural Radiance Fields, Intrinsic Decomposition, Portrait Editing, Motion Control
2025AniCrafter: Customizing Realistic Human-Centric Animation via Avatar-Background Conditioning in Video Diffusion ModelsArXiv 2025CodeAvatar, Human-Centric Animation
2025Eye-See-You: Reverse Pass-Through VR and Head AvatarsIJCAI 2025VR, Head Avatars, Pass-Through
2025MAGE:A Multi-stage Avatar Generator with Sparse ObservationsArXiv 2025Avatar, AR/VR
2025MagicPortrait: Temporally Consistent Face Reenactment with 3D Geometric GuidanceArXiv 2025CodeLatent Diffusion, FLAME, 3D Geometric Guidance, Face Reenactment
2025Supervising 3D Talking Head Avatars with Analysis-by-Audio-SynthesisArXiv 2025Project3D, Avatar, Audio-Synthesis
2025EmoDiffusion: Enhancing Emotional 3D Facial Animation with Latent Diffusion ModelsArXiv 2025Project3D facial animation, latent diffusion, emotional expression, speech-driven
2025Better Together: Unified Motion Capture and 3D Avatar ReconstructionArXiv 2025
2025Vid2Avatar-Pro: Authentic Avatar from Videos in the Wild via Universal PriorCVPR 2025CodeProject
2025LUCAS: Layered Universal Codec AvatarsArXiv 2025
2025GAS: Generative Avatar Synthesis from a Single ImageICCV 2025CodeProjectSingle Image, 3D Avatar, NeRF, Diffusion
2025MaintaAvatar: A Maintainable Avatar Based on Neural Radiance Fields by Continual LearningAAAI 2025NeRF
2025Deblur-Avatar: Animatable Avatars from Motion-Blurred Monocular VideosArXiv 2025Motion Blur, Animatable Avatars
2025TalkingEyes: Pluralistic Speech-Driven 3D Eye Gaze AnimationArXiv 2025CodeProject3D Eye Gaze, Speech-Driven, Pluralistic
2025Arc2Avatar: Generating Expressive 3D Avatars from a Single Image via ID GuidanceArXiv 2025CodeProjectAvatars, Single Image
2025LayerAvatar: Disentangled Clothed Avatar Generation with Layered RepresentationICCV 2025 (Highlight)CodeProject
2025L3D-Pose: Lifting Pose for 3D Avatars from a Single Camera in the WildICASSP 2025Project
2025Barbie: Text to Barbie-Style 3D AvatarsArXiv 2025CodeProjectText to Avatar, Barbie-Style
2025Hybrid Explicit Representation for Ultra-Realistic Head AvatarsArXiv 2025
2024Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with AdaptersArXiv 2024Co-speech
2024CAP4D: Creating Animatable 4D Portrait Avatars with Morphable Multi-View Diffusion ModelsArXiv 2024CodeProjectMulti-View Diffusion
2024StrandHead: Text to Strand-Disentangled 3D Head Avatars Using Hair Geometric PriorsICCVCodeProject
2024SimAvatar: Simulation-Ready Avatars with Layered Hair and ClothingArXiv 2024ProjectNVIDIA, Hair and Clothing
2024DynamicAvatars: Accurate Dynamic Facial Avatars Reconstruction and Precise Editing with Diffusion ModelsArXiv 2024
2024ConsistentAvatar: Learning to Diffuse Fully Consistent Talking Head Avatar with Temporal GuidanceArXiv 2024CodeProject
2024DanceFusion: A Spatio-Temporal Skeleton Diffusion Transformer for Audio-Driven Dance Motion ReconstructionArXiv 2024Project
2024InstantGeoAvatar: Effective Geometry and Appearance Modeling of Animatable Avatars from Monocular VideoACCV 2024Code
2024EgoAvatar: Egocentric View-Driven and Photorealistic Full-body AvatarsArXiv 2024
2024Towards Native Generative Model for 3D Head AvatarArXiv 2024
2024Stable Video PortraitsECCV 2024ProjectDiffusion
2024LightAvatar: Efficient Head Avatar as Dynamic Neural Light FieldECCV'24 CADLCode
2024FreeAvatar: Robust 3D Facial Animation Transfer by Learning an Expression Foundation ModelArXiv 20243D Facial Animation, Expression Transfer, Foundation Model, Video-driven
2024KMTalk: Speech-Driven 3D Facial Animation with Key Motion EmbeddingECCV 2024Code3D Facial Animation, Key Motion, Speech-Driven
2024Subjective and Objective Quality Assessment of Rendered Human Avatar Videos in Virtual RealityIEEE 2024
2024PAV: Personalized Head Avatar from Unstructured Video CollectionECCV 2024Project
2024XHand: Real-time Expressive Hand AvatarArXiv 2024CodeHand
2024Bridging the Gap: Studio-like Avatar Creation from a Monocular Phone CaptureECCV 2024Project
2024Universal Facial Encoding of Codec Avatars from VR HeadsetsSIGGRAPH 2024Facial Encoding, VR Headset, Real-time Animation, 3D Avatar
2024CanonicalFusion: Generating Drivable 3D Human Avatars from Multiple ImagesECCV 2024Code
2024WildAvatar: Web-scale In-the-wild Video Dataset for 3D Avatar CreationArXiv 2024CodeProjectDataset
2024AniFaceDiff: Animating Stylized Avatars via Parametric Conditioned Diffusion ModelsArXiv 2024
2024Human-3Diffusion: Realistic Avatar Creation via Explicit 3D Consistent Diffusion ModelsNIPS 2024CodeProjectDiffusion
2024Instant 3D Human Avatar Generation using Image Diffusion ModelsArXiv 2024Project
2024Representing Animatable Avatar via Factorized Neural FieldsArXiv 2024
2024Stratified Avatar Generation from Sparse ObservationsCVPR 2024 (Oral)
2024E3Gen: Efficient, Expressive and Editable Avatars GenerationArXiv 2024CodeProject
2024X-Oscar: A Progressive Framework for High-quality Text-guided 3D Animatable Avatar GenerationICML 2024CodeProject
2024GeneAvatar: Generic Expression-Aware Volumetric Head Avatar Editing from a Single ImageCVPR 2024CodeProjectEditing
2024MonoAvatar++: Efficient 3D Implicit Head Avatar with Mesh-anchored Hash Table BlendshapesCVPR 2024ProjectBlendshapes
2024MagicMirror: Fast and High-Quality Avatar Generation with a Constrained Search SpaceArXiv 2024Project
2024MI-NeRF: Learning a Single Face NeRF from Multiple IdentitiesArXiv 2024ProjectNeRF, Multi-Identity, Face Modeling
2024NECA: Neural Customizable Human AvatarCVPR 2024Code
2024Magic-Me: Identity-Specific Video Customized DiffusionArXiv 2024CodeProject
2024ViCA-NeRF: View-Consistency-Aware 3D Editing of Neural Radiance FieldsNIPS 2023CodeProject3D Edit
2024Sketch2NeRF: Multi-view Sketch-guided Text-to-3D GenerationArXiv 2024Text to 3D
2024UltrAvatar: A Realistic Animatable 3D Avatar Diffusion Model with Authenticity Guided TexturesArXiv 2024ProjectDiffusion, Avatar
2024High-Quality Mesh Blendshape Generation from Face Videos via Neural Inverse RenderingArXiv 2024Codemesh blendshape, neural inverse rendering, face reconstruction
2024FED-NeRF: Achieve High 3D Consistency and Temporal Coherence for Face Video Editing on Dynamic NeRFArXiv 2024Code4D face video editor
2024Learning Dense Correspondence for NeRF-Based Face ReenactmentAAAI 2024one-shot multi-view face reenactmen
2024VAST: Vivify Your Talking Avatar via Zero-Shot Expressive Facial Style TransferICCV2023
2024AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view VideosECCV 2024CodeProject
2024What You See Is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANsArXiv 2024Project
2023INFAMOUS-NeRF: ImproviNg FAce MOdeling Using Semantically-Aligned Hypernetworks with Neural Radiance FieldsArXiv 2023NeRF, Face Modeling, Hypernetworks, Semantically-Aligned
2023AvatarStudio: High-fidelity and Animatable 3D Avatar Creation from TextArXiv 2023ProjectText-to-3D, NeRF, SMPL, Diffusion Model
20233D Face Style Transfer with a Hybrid Solution of NeRF and Mesh RasterizationArXiv 20233D Face, Style Transfer, NeRF, Mesh
2023HAvatar: High-fidelity Head Avatar via Facial Model Conditioned Neural Radiance FieldArXiv 2023Neural Radiance Field, Facial Model Conditioning, 3D Head Avatar, Expression Control
2023NOFA: NeRF-based One-shot Facial Avatar ReconstructionArXiv 2023NeRF, One-shot, Facial Avatar
2023Instruct-NeuralTalker: Editing Audio-Driven Talking Radiance Fields with InstructionsArXiv 2023
2023MA-NeRF: Motion-Assisted Neural Radiance Fields for Face Synthesis from Sparse ImagesArXiv 2023NeRF, Motion-Assisted, Face Synthesis
2023GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face GenerationArXiv 2023CodeProject-
2023GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face SynthesisICLR 2023CodeProject-
2023SD-NeRF: Towards Lifelike Talking Head Animation via Spatially-adaptive Dual-driven NeRFsIEEE 2023
2022NeRFInvertor: High Fidelity NeRF-GAN Inversion for Single-shot Real Image AnimationArXiv 2022CodeProject-
2022RAD-NeRF: Real-time Neural Talking Portrait SynthesisArXiv 2022CodeProjectInstantNGP
2022Next3D: Generative Neural Texture Rasterization for 3D-Aware Head AvatarsArXiv 2022CodeProject-
2022FNeVR: Neural Volume Rendering for Face AnimationArXiv 2022Code-
20223DFaceShop: Explicitly Controllable 3D-Aware Portrait GenerationArXiv 2022CodeProject-
2022DFRF:Learning Dynamic Facial Radiance Fields for Few-Shot Talking Head SynthesisECCV 2022CodeProject
2022ROME: Realistic One-shot Mesh-based Head AvatarsECCV 2022CodeProject-
2022SSP-NeRF: Semantic-Aware Implicit Neural Audio-Driven Video Portrait GenerationArXiv 2022CodeProject-
2022IMavatar: Implicit Morphable Head Avatars from VideosCVPR 2022CodeProject-
2022HeadNeRF: A Real-time NeRF-based Parametric Head ModelCVPR 2022CodeProject-
2021DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural RenderingArXiv 2021Code-
2021AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head SynthesisICCV 2021CodeProject-
2021NerFACE: Dynamic Neural Radiance Fields for Monocular 4D Facial Avatar ReconstructionCVPR 2021 OralCodeProject-

3D Gaussian Splatting

YearTitleConference/JournalCodeProjectKeywords
2026PD-GS: Phoneme-Driven 3DGS for Audio-Driven Talking HeadsACM MM 20263DGS, phoneme-driven, audio-driven, lip articulation
2026S-Avatar: Diffusion-Guided Gaussian Head Avatars from a Single ImageArXiv 20263DGS, single-image, FLAME, animatable head, diffusion
2026SpiD: Split and Drive: Dual-Axis Disentanglement for Real-Time Gaussian Head AvatarsArXiv 20263DGS, single-image, animatable head, real-time
2026DynHair: Head Avatars with Dynamic Explicit HairArXiv 2026CodeProject3DGS, head avatar, dynamic hair, animatable
2026URHead: A Unified UV-Space Representation for Joint Mesh-3DGS Optimization in Head AvatarsECCV 2026Project3DGS, mesh, animatable head, UV-space
2026FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian HeadArXiv 20263DGS, one-shot, 4D head, animatable avatar
2026GaussianEmoTalker: Real-Time Emotional Talking Head Synthesis with Audio-Driven and Blendshape-Based 3D Gaussian SplattingArXiv 2026Project3DGS, audio-driven, emotional talking head, blendshapes
2026FFAvatar: Feed-Forward 4D Head Avatar Reconstruction from Sparse Portrait ImagesECCV 2026CodeProject3DGS, 4D head, FLAME, feed-forward, animatable
2026FiCA: Feed-forward instant Gaussian Codec Avatars from a Single Portrait ImageArXiv 2026Project3DGS, codec avatar, single-image, drivable, feed-forward
2026Edit3DGS: Unified Framework for Dynamic Head Editing via 2D Instruction-Guided Diffusion and 3D Gaussian SplattingSOICT 2025text editing, 3DGS, dynamic head, instruction-guided
2026EmoZone-Talker: Regional Semantic Control of Audio-Driven 3DGS Talking Heads via Facial Action UnitsArXiv 20263DGS, audio-driven, action units, expression control
2026SpatialAvatar-0: High-Quality 4D Head Avatar with Multi-Stage ReconstructionArXiv 2026Project3DGS, 4D head, FLAME, feed-forward
2026LentiAvatar: Pseudo-Multiview Reconstruction and Subpixel Prism Rendering for Real-Time Stereoscopic CommunicationArXiv 20263DGS, head avatar, telepresence, controllable
2026SAGE: Self-Learning Expression Deformations for Data-Efficient Gaussian AvatarsArXiv 20263DGS, Gaussian avatar, expression, animatable, few-shot
2026SplitAvatar: One-shot Head Avatar with Autoregressive Gaussian SplittingArXiv 20263DGS, one-shot, animatable head, Gaussian splitting, 3DMM
2026PiG-Avatar: Hierarchical Neural-Field-Guided Gaussian AvatarsArXiv 2026gaussian splatting, neural field, full-body avatar, clothing, hierarchical
2026FFAvatar: Few-Shot, Feed-Forward, and Generalizable Avatar ReconstructionArXiv 2026Projectgaussian splatting, head avatar, few-shot, FLAME, feed-forward
2026SDTalk: Structured Facial Priors and Dual-Branch Motion Fields for Generalizable Gaussian Talking Head SynthesisArXiv 2026gaussian splatting, talking head, one-shot, facial priors, lip sync
2026HeadsUp: Large-Scale High-Quality 3D Gaussian Head Reconstruction from Multi-View CapturesArXiv 2026Projectgaussian splatting, head reconstruction, multi-view, large-scale, animatable
2026HumanSplatHMR: Closing the Loop Between Human Mesh Recovery and Gaussian Splatting AvatarArXiv 2026gaussian splatting, human mesh recovery, full-body avatar, novel view synthesis, pose refinement
2026High-Fidelity Mobile Avatars with Pruned Local BlendshapesArXiv 2026Projectgaussian splatting, mobile rendering, full-body avatar, blendshapes, real-time
2026SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian SplattingCVPR 2026 (Highlight)Code3DGS, sketch-driven, face editing, real-time, CVPR 2026
2026Any3DAvatar: Fast and High-Quality Full-Head 3D Avatar Reconstruction from Single Portrait ImageArXiv 2026single-image, 3DGS head, feed-forward, full-head
2026F3G-Avatar : Face Focused Full-body Gaussian AvatarCVPRW 2026Codefull-body avatar, face-focused, gaussian splatting, multi-view
2026SFGS: Structure-Aware Fine-Grained Gaussian Splatting for Expressive Avatar ReconstructionArXiv 2026Codefine-grained, structure-aware, gaussian splatting, expressive avatar
2026PhysHead: Simulation-Ready Gaussian Head AvatarsCVPR 2026Project3D Gaussian, Head Avatar, Physics Simulation
2026AvatarPointillist: AutoRegressive 4D Gaussian AvatarizationCVPR 2026CodeProject4D Gaussian, Autoregressive, Avatar
2026Better Rigs, Not Bigger Networks: A Body Model Ablation for Gaussian AvatarsArXiv 2026CodeBody Model, Gaussian Avatars, Ablation
2026AAP-3DGA: Autoregressive Appearance Prediction for 3D Gaussian AvatarsArXiv 2026Project3D Gaussian, Autoregressive, Appearance Prediction
2026DipGuava: Disentangling Personalized Gaussian Features for 3D Head Avatars from Monocular VideoArXiv 20263D Head Avatar, Gaussian Splatting, Monocular Video, Personalized
2026FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head AvatarArXiv 2026Face-and-Hair Avatar, 3D Gaussian, Composable Reconstruction
2026ProgressiveAvatars: Progressive Animatable 3D Gaussian AvatarsCVPR 2026Project3D Gaussian, Progressive, Animatable Avatars
2026Feed-forward Gaussian Registration for Head Avatar Creation and EditingArXiv 2026ProjectGaussian Registration, Head Avatar, Feed-forward
2026Retrieval-Augmented Gaussian Avatars: Improving Expression GeneralizationArXiv 2026Gaussian Splatting, Expression Generalization, Retrieval Augmentation, 3D Avatars
2026Gaussian Wardrobe: Compositional 3D Gaussian Avatars for Free-Form Virtual Try-On3DV 2026CodeProject3D Gaussian, Compositional Avatar, Virtual Try-On
2026LiftAvatar: Kinematic-Space Completion for Expression-Controlled 3D Gaussian Avatar AnimationArXiv 2026Kinematic-Space Completion, Expression Control, 3D Gaussian Splatting, Video Diffusion
2026OMG-Avatar: One-shot Multi-LOD Gaussian Head AvatarArXiv 2026Project3D Gaussian, One-Shot, Head Avatar
2026GeoDiff4D: Geometry-Aware Diffusion for 4D Head Avatar ReconstructionArXiv 2026CodeProjectGeometry-Aware Diffusion, 4D Avatar Reconstruction, 3D Gaussian Splatting, Surface Normals
2026OMEGA-Avatar: One-shot Modeling of 360° Gaussian AvatarsArXiv 2026ProjectOne-Shot Avatar, 360° Full-Head, 3D Gaussian Splatting, Multi-View Feature Splatting
2026OFERA: Blendshape-driven 3D Gaussian Control for Occluded Facial Expression to Realistic Avatars in VRArXiv 2026CodeProjectBlendshape Control, Gaussian Avatars, VR Telepresence, Real-time Expression
2026VRGaussianAvatar: Integrating 3D Gaussian Avatars into VRArXiv 2026CodeProject3D Gaussian, VR Avatar, Integration
2026The Gaussian-Head OFL Family: One-Shot Federated Learning from Client Global Statisticsthe International Conference on Learning Representations (ICLR) 2026Gaussian-Head, Federated Learning, One-Shot
2026Splat-Portrait: Generalizing Talking Heads with Gaussian SplattingArXiv 2026CodeProjectGaussian Splatting, 3DGS, Portrait Animation, Talking Head
2026GlassesGB: Controllable 2D GAN-Based Eyewear Personalization for 3D Gaussian Blendshapes Head AvatarsIEEE VR 2026Projectgaussian blendshapes, virtual try-on, eyewear, head avatar
2026CAG-Avatar: Cross-Attention Guided Gaussian Avatars for High-Fidelity Head ReconstructionArXiv 20263D Gaussian Splatting, cross-attention, head reconstruction, drivable avatars
2026FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time AnimationICLR 20263D Gaussian Splatting, Head Avatars, Real-Time Animation, Few-Shot Learning
2026FHAvatar: Generalizable and Animatable 3D Full-Head Gaussian Avatar from a Single ImageArXiv 2026CodeProject3D full-head avatar, Gaussian primitives, UV space, single-image reconstruction
2026ELITE: Efficient Gaussian Head Avatar from a Monocular Video via Learned Initialization and TEst-time Generative AdaptationArXiv 2026CodeProjectGaussian Avatar, Test-time Adaptation, Diffusion, Monocular Video
2026UIKA: Fast Universal Head Avatar from Pose-Free ImagesArXiv 2026CodeProjectGaussian Splatting, UV Mapping, Head Avatar, Feed-forward
2026LayerGS: Decomposition and Inpainting of Layered 3D Human Avatars via 2D Gaussian SplattingArXiv 2026Code2D Gaussian Splatting, Layered Avatar, Decomposition
2026GaussianSwap: Animatable Video Face Swapping with 3D Gaussian SplattingArXiv 20263D Gaussian Splatting, Face Swapping, Animatable
2026RelightAnyone: A Generalized Relightable 3D Gaussian Head ModelArXiv 20263D Gaussian Splatting, relightable avatars, single-image fitting, cross-subject generalization
2026CaricatureGS: Exaggerating 3D Gaussian Splatting Faces With Gaussian CurvatureArXiv 2026CodeProject3D Gaussian Splatting, Caricature, Gaussian Curvature
2026GTAvatar: Bridging Gaussian Splatting and Texture Mapping for Relightable and Editable Gaussian AvatarsEurographics 2026CodeProjectGaussian Splatting, Texture Mapping, Relightable Avatars, 3D Reconstruction
2026STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars ReconstructionCVPR 2026CodeProjectGaussian Splatting, 3D Head Avatars, Soft Binding, Temporal Density Control
2025TexAvatars: Hybrid Texel-3D Representations for Stable Rigging of Photorealistic Gaussian Head AvatarsArXiv 2025CodeProject3D Gaussian Splatting, hybrid representation, analytic rigging, UV space
2025FlexAvatar: Flexible Large Reconstruction Model for Animatable Gaussian Head Avatars with Detailed DeformationArXiv 2025Project3D avatar, Gaussian Splatting, deformation, reconstruction
2025Instant Expressive Gaussian Head Avatar via 3D-Aware Expression DistillationArXiv 20253D, Gaussian Splatting, Avatar, Attention
2025Gaussian Pixel Codec Avatars: A Hybrid Representation for Efficient RenderingTech Report 2025Gaussian Splatting, head avatar, hybrid representation, efficient rendering
2025AGORA: Adversarial Generation Of Real-time Animatable 3D Gaussian Head AvatarsArXiv 2025CodeProject3D Gaussian Splatting, Animatable Avatars, FLAME, Real-time Rendering
2025EmoDiffTalk:Emotion-aware Diffusion for Editable 3D Gaussian Talking HeadCVPR 2026Project3D, Diffusion, Emotional, Talking Head
2025AHA! Animating Human Avatars in Diverse Scenes with Gaussian SplattingArXiv 2025CodeProjectGaussian Splatting, Human Avatar, Scene Animation
2025Densemarks: Learning Canonical Embeddings for Human Heads Images via Point TracksArXiv 2025CodeProjectHead Correspondence, Canonical Embedding, Tracking, Avatar
2025STG-Avatar: Animatable Human Avatars via Spacetime GaussianIROS 2025CodeProjectSpacetime Gaussian, Animatable Avatar, 3DGS
2025Capture, Canonicalize, Splat: Zero-Shot 3D Gaussian Avatars from Unstructured Phone ImagesICCV 2025Zero-Shot, 3D Gaussian Avatars, Phone Images
2025Instant Skinned Gaussian Avatars for Web, Mobile and VR ApplicationsSUI 2025CodeProjectReal-Time, Cross-Platform, 3D Avatar, Gaussian Splatting
2025Towards Efficient 3D Gaussian Human Avatar Compression: A Prior-Guided FrameworkArXiv 20253D Gaussian, Human Avatar, Compression
2025ArchitectHead: Continuous Level of Detail Control for 3D Gaussian Head AvatarsArXiv 20253D Gaussian Head Avatars, Level of Detail Control, Continuous LOD
2025MPMAvatar: Learning 3D Gaussian Avatars with Accurate and Robust Physics-Based DynamicsNeurIPS 2025CodeProjectPhysics-Based, 3DGS, Garments
2025FMGS-Avatar: Mesh-Guided 2D Gaussian Splatting with Foundation Model Priors for 3D Monocular Avatar ReconstructionArXiv 20252D Gaussian, Mesh-Guided, Avatar
2025Dream3DAvatar: Text-Controlled 3D Avatar Reconstruction from a Single ImageArXiv 2025Text-Driven, Single Image, 3DGS
2025PanoLAM: Large Avatar Model for Gaussian Full-Head Synthesis from One-shot Unposed ImageArXiv 2025ProjectGaussian, Full-Head Synthesis, One-shot
2025GaussianGAN: Real-Time Photorealistic controllable Human AvatarsFG 20253DGS, Real-Time, Photorealistic
2025Im2Haircut: Single-view Strand-based Hair Reconstruction for Human AvatarsArXiv 2025CodeProjectHair Reconstruction, Gaussian Splatting
2025AvatarBack: Back-Head Generation for Complete 3D Avatars from Front-View ImagesArXiv 20253D Gaussian Splatting, Back-Head Generation, Avatar Reconstruction, Spatial Alignment
2025FastAvatar: Instant 3D Gaussian Splatting for Faces from Single Unconstrained PosesArXiv 2025CodeProject3DGS, Instant, Single Image
2025EAvatar: Expression-Aware Head Avatar Reconstruction with Generative Geometry PriorsArXiv 20253D Gaussian Splatting, expression-aware, deformation-aware, generative priors
2025SVG-Head: Hybrid Surface-Volumetric Gaussians for High-Fidelity Head Reconstruction and Real-Time EditingArXiv 2025CodeProjectGaussian Splatting, 3D Avatar, Texture Editing, FLAME
2025MoGaFace: Momentum-Guided and Texture-Aware Gaussian Avatars for Consistent Facial GeometryArXiv 2025Gaussian Avatars, FLAME Meshes, Geometry Refinement
2025MoGA: 3D Generative Avatar Prior for Monocular Gaussian Avatar ReconstructionICCV 2025CodeProject3D Generative Avatar, Monocular Reconstruction
2025HairCUP: Hair Compositional Universal Prior for 3D Gaussian AvatarsArXiv 2025Project3D Gaussian Avatars, Hair Compositionality, Disentangled Prior, Few-shot Fine-tuning
2025GeoAvatar: Adaptive Geometrical Gaussian Splatting for 3D Head AvatarArXiv 2025CodeProjectAdaptive Gaussian Splatting, 3D Head Avatar, Mouth Structure, Deformation Strategy
2025StreamME: Simplify 3D Gaussian Avatar within Live StreamArXiv 2025CodeProject3D Gaussian Splatting, avatar reconstruction, on-the-fly training
2025ScaffoldAvatar: High-Fidelity Gaussian Avatars with Patch ExpressionsSIGGRAPH 2025ProjectHigh-Fidelity, Gaussian Avatars, Patch Expressions
2025HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face AvatarsArXiv 2025CodeProjectHigh-Dimensional, Gaussian Splatting
2025BecomingLit: Relightable Gaussian Avatars with Hybrid Neural ShadingArXiv 2025CodeProject3DGS, Relightable, Neural Shading
2025EVA: Expressive Virtual Avatars from Multi-view VideosSIGGRAPH 2025ProjectAvatar, 3D Gaussian
2025ToonifyGB: StyleGAN-based Gaussian Blendshapes for 3D Stylized Head AvatarsIEEE VR 2026CodeProjectgaussian blendshapes, stylization, toonify, head avatar
2025TeGA: Texture Space Gaussian Avatars for High-Resolution Dynamic Head ModelingSIGGRAPH 2025Project3DGS, Avatar, High-Resolution
2025SVAD: From Single Image to 3D Avatar via Synthetic Data Generation with Video Diffusion and Data AugmentationCVPRW 2025CodeProjectSingle Image, 3D Avatar, 3DGS, Video Diffusion
2025GUAVA: Generalizable Upper Body 3D Gaussian AvatarICCV 2025CodeProject3D Gaussian Avatar, Upper Body, SMPLX
2025DNF-Avatar: Distilling Neural Fields for Real-time Animatable Avatar RelightingICCV 2025CodeProjectRelightable Avatar, 2DGS Distillation
2025TaoAvatar: Real-Time Lifelike Full-Body Talking Avatars for Augmented Reality via 3D Gaussian SplattingCVPR 2025(Highlight🚀)ProjectAR
2025RGBAvatar: Reduced Gaussian Blendshapes for Online Modeling of Head AvatarsArXiv 2025Codegaussian blendshapes, real-time, compact, animatable
20252DGS-Avatar: Animatable High-fidelity Clothed Avatar via 2D Gaussian SplattingICVRV 20242DGS
2025Avat3r: Large Animatable Gaussian Reconstruction Model for High-fidelity 3D Head AvatarsArXiv 2025Project
2025LAM: Large Avatar Model for One-shot Animatable Gaussian HeadArXiv 2025CodeProject
2025RFGCA: Relightable Full-Body Gaussian Codec AvatarsArXiv 2025ProjectFull-Body, Avatars
2025PERSE: Personalized 3D Generative Avatars from A Single PortraitCVPR 2025CodeProjectPersonalized, 3DGS, Single Image
2025EGG3D: Generating Editable Head Avatars with 3D Gaussian GANsArXiv 2025CodeProject3DGS
2025FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionArXiv 2025ProjectPose-free, Sparse-view, 3DGS
2025InteractRAGA: Interactive Rendering of Relightable and Animatable Gaussian AvatarsArXiv 2025CodeProjectGaussian Splatting, Relightable Avatars, Interactive Rendering, Pose-driven Animation
2024GraphAvatar: Compact Head Avatars with GNN-Generated 3D GaussiansAAAI 2025CodeGNN-Generated, 3DGS
20243D$^2$-Actor: Learning Pose-Conditioned 3D-Aware Denoiser for Realistic Gaussian Avatar ModelingAAAI 2025CodeProject3DGS
2024GAF: Gaussian Avatar Reconstruction from Monocular Videos via Multi-view DiffusionArXiv 2024CodeProjectDiffusion
2024GASP: Gaussian Avatars with Synthetic PriorsArXiv 2024Project
2024MixedGaussianAvatar: Realistically and Geometrically Accurate Head Avatar via Mixed 2D-3D Gaussian SplattingArXiv 2024Project3DGS, 2D-3D
2024PBDyG: Position Based Dynamic Gaussians for Motion-Aware Clothed Human AvatarsArXiv 2024Clothed Avatar
2024GAST: Sequential Gaussian Avatars with Hierarchical Spatio-temporal ContextArXiv 2024CodeProject
2024Bundle Adjusted Gaussian Avatars DeblurringCVPRCode
2024FATE: Full-head Gaussian Avatar with Textural Editing from Monocular VideoCVPRCodeProject
2024DAGSM: Disentangled Avatar Generation with GS-enhanced MeshCVPR
2024DreamWaltz-G: Expressive 3D Gaussian Avatars from Skeleton-Guided 2D DiffusionArXiv 2024CodeProject
2024Gaussian Déjà-vu: Creating Controllable 3D Gaussian Head-Avatars with Enhanced Generalization and Personalization AbilitiesWACV 2025CodeProject
2024GaussianHeads: End-to-End Learning of Drivable Gaussian Head Avatars from Coarse-to-fine RepresentationsSIGGRAPH Asia 2024CodeProject🔥Gaussian Splatting
2024DEGAS: Detailed Expressions on Full-Body Gaussian AvatarsArXiv 2024🔥Gaussian Splatting
2024Topology-aware Human Avatars with Semantically-guided Gaussian SplattingArXiv 2024
2024CHASE: 3D-Consistent Human Avatars with Sparse Inputs via Gaussian Splatting and Contrastive LearningArXiv 2024
2024ExAvatar: Expressive Whole-Body 3D Gaussian AvatarECCV 2024CodeProject
2024GEM: Gaussian Eigen Models for Human HeadsCVPR 2025CodeProject
2024EVA: Expressive Gaussian Human Avatars from Monocular RGB VideoArXiv 2024CodeProject
2024NPGA: Neural Parametric Gaussian AvatarsArXiv 2024Project
2024GaussianVTON: 3D Human Virtual Try-ON via Multi-Stage Gaussian Splatting Editing with Image PromptingOn going workTry-ON
20243D Gaussian Blendshapes for Head Avatar AnimationACM SIGGRAPH 2024Gaussian splatting, blendshapes, head avatar, real-time rendering
2024MeGA: Hybrid Mesh-Gaussian Head Avatar for High-Fidelity Rendering and Head EditingArXiv 2024CodeProject🔥Gaussian Splatting
2024DG-Mesh: Dynamic Gaussians Mesh: Consistent Mesh Reconstruction from Monocular VideosArXiv 2024CodeProject🔥Gaussian Splatting
2024HAHA: Highly Articulated Gaussian Human Avatars with Textured Mesh PriorArXiv 2024🔥Gaussian Splatting
2024UV Gaussians: Joint Learning of Mesh Deformation and Gaussian Textures for Human Avatar ModelingArXiv 2024Project🔥Gaussian Splatting
2024DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth NormalizationCVPR 2024CodeProject🔥Gaussian Splatting, Sparse-View
2024V3D: Video Diffusion Models are Effective 3D GeneratorsArXiv 2024CodeProject🔥Gaussian Splatting, Video
2024SplattingAvatar: Realistic Real-Time Human Avatars with Mesh-Embedded Gaussian SplattingCVPR 2024CodeProject🔥Gaussian Splatting
2024GEA: Reconstructing Expressive 3D Gaussian Avatar from Monocular VideoArXiv 2024Project🔥Gaussian Splatting, Avatar
2024Consolidating Attention Features for Multi-view Image EditingArXiv 2024Project🔥Gaussian Splatting, Edit
2024GaussianHair: Hair Modeling and Rendering with Light-aware GaussiansArXiv 2024🔥Gaussian Splatting
2024ImplicitDeepfake: Plausible Face-Swapping through Implicit Deepfake Generation using NeRF and Gaussian SplattingArXiv 2024🔥Gaussian Splatting, Deepfake
2024HeadStudio: Text to Animatable Head Avatars with 3D Gaussian SplattingECCVCode🔥Gaussian Splatting, Avatar
2024Rig3DGS: Creating Controllable Portraits from Casual Monocular VideosArXiv 2024ProjectPortraits
20244D Gaussian Splatting: Towards Efficient Novel View Synthesis for Dynamic ScenesArXiv 2024Dynamic Scenes
2024PSAvatar: A Point-based Shape Model for Real-Time Head Avatar Animation with 3D Gaussian SplattingArXiv 20243D Gaussian Splatting, Head Avatar Animation, Point-based Shape Model, Real-time Rendering
2024GaussianBody: Clothed Human Reconstruction via 3d Gaussian SplattingArXiv 2024🔥Gaussian Splatting
2024Gaussian Shadow Casting for Neural CharactersArXiv 2024🔥Gaussian Splatting
2024CoSSegGaussians: Compact and Swift Scene Segmenting 3D Gaussians with Dual Feature FusionArXiv 2024CodeProjectSegmentic
2024AGG: Amortized Generative 3D Gaussians for Single Image to 3DArXiv 2024Project🔥Gaussian Splatting
20244DGen: Grounded 4D Content Generation with Spatial-temporal ConsistencyArXiv 2024CodeProject🔥Gaussian Splatting
2024Human101: Training 100+FPS Human Gaussians in 100s from 1 ViewArXiv 2024CodeProject🔥Gaussian Splatting
2024Deformable 3D Gaussian Splatting for Animatable Human AvatarsArXiv 2024🔥Gaussian Splatting
20243DGS-Avatar: Animatable Avatars via Deformable 3D Gaussian SplattingArXiv 2024CodeProject🔥Gaussian Splatting
2024HHAvatar: Gaussian Head Avatar with Dynamic HairsArXiv 2024CodeProjectHair
2024HeadGaS: Real-Time Animatable Head Avatars via 3D Gaussian SplattingECCV🔥Gaussian Splatting
2024GaussianAvatars: Photorealistic Head Avatars with Rigged 3D GaussiansCVPR 2024CodeProject🔥Gaussian Splatting
2024GaussianHead: Impressive 3D Gaussian-based Head Avatars with Dynamic Hybrid Neural FieldArXiv 2024Code🔥Gaussian Splatting
2024MonoGaussianAvatar: Monocular Gaussian Point-based Head AvatarArXiv 2024🔥Gaussian Splatting

Conversational & Dialogue

YearTitleConference/JournalCodeProjectKeywords
2026EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live ChatbotArXiv 2026CodeProjectconversational avatar, lip-sync, empathetic chatbot, live interaction
2026STEER: Steerable Dyadic Head AvatarsArXiv 2026CodeProject3DGS, dyadic, conversational head, Gaussian avatar, listener
2026OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive AvatarsArXiv 2026interactive avatar, multi-turn, listening, streaming, audio-visual
2026MaAI: Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue SystemsICMI 2026CodeProjectlistener nodding, dyadic, avatar dialogue, VAP, real-time
2026Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic PriorsArXiv 2026Projectdyadic, conversational motion, talking head, interaction
2026CHAT: Conversational Human Audio-visual Talking Dialogue GenerationECCV 2026dyadic dialogue, talking face, interactive avatar, audio-visual
2026InterTalk: Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face GenerationArXiv 2026Projectconversational talking face, multi-party, listener feedback, real-time
2026FacePlex: Full-Duplex Joint Speech-Facial Motion Generation for Conversational AvatarsArXiv 2026Projectfull-duplex, conversational avatar, facial motion, speech generation
2026MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic ConversationsECCV 2026Projectdyadic, listener, facial animation, talking-and-listening, flow matching
2026InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware AvatarsArXiv 2026interactive avatar, intent-aware, streaming, listening, audio-driven
2026Resonant Minds: Closed-Loop Social Avatars with Theory of MindArXiv 2026CodeProjectsocial avatar, theory of mind, listener, talking face
2026DyaPlex: Full-Duplex Speech-Motion Model for Dyadic InteractionArXiv 2026Projectdyadic interaction, full-duplex, speech-motion, DyaPlex, streaming
2026EmbodiedHead: Real-Time Listening and Speaking Avatar for Conversational AgentsArXiv 2026listening-speaking, DiT, rectified flow, real-time avatar
2026Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware KernelsArXiv 2026Projecttalking-listening, interactive, full-duplex, conversational
2026PolySLGen: Online Multimodal Speaking-Listening Reaction Generation in Polyadic InteractionArXiv 2026Codepolyadic interaction, speaking-listening, multimodal reaction
2026GDPO-Listener: Expressive Interactive Head Generation via Auto-Regressive Flow Matching and Group reward-Decoupled Policy OptimizationArXiv 2026listener, interactive, flow matching, RLHF
2026InterDyad: Interactive Dyadic Speech-to-Video Generation by Querying Intermediate Visual GuidanceArXiv 2026Projectdyadic, speech-to-video, interactive
2026ECHO: Towards Emotionally Appropriate and Contextually Aware Interactive Head GenerationArXiv 2026ProjectInteractive Head, Emotion, Context-Aware, Dialogue
2026ReactMotion: Generating Reactive Listener Motions from Speaker UtteranceArXiv 2026CodeProjectlistener motion, reactive, body motion, speaker utterance
2026Talking Together: Synthesizing Co-Located 3D Conversations from AudioCVPR 20263D, Conversations, Talking Head, CVPR
2026A²-LLM: An End-to-end Conversational Audio Avatar Large Language ModelArXiv 2026CodeConversational, Avatar, LLM
2026HoverAI: An Embodied Aerial Agent for Natural Human-Drone InteractionArXiv 2026lip-synced avatars, real-time conversational AI, multimodal pipeline
2026RSATalker: Realistic Socially-Aware Talking Head Generation for Multi-Turn ConversationArXiv 2026Multi-Turn Conversation, Talking Head
2026Avatar Forcing: Real-Time Interactive Head Avatar Generation for Natural ConversationArXiv 2026CodeProjectInteractive avatar, Diffusion forcing, Real-time, Preference optimization
2025ALIVE: An Avatar-Lecture Interactive Video Engine with Content-Aware Retrieval for Real-Time InteractionArXiv 2025neural talking-head synthesis, content-aware retrieval, real-time interaction, LLM
2025TAVID: Text-Driven Audio-Visual Interactive Dialogue GenerationArXiv 2025text-driven, audio-visual, interactive dialogue, cross-modal mappers
2025ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual BodyArXiv 2025CodeProjectconversational agent, 3D avatar, multimodal interaction, joint language-motion
2025Towards Interactive Intelligence for Digital HumansArXiv 2025interactive intelligence, digital human, multimodal embodiment, real-time interaction
2025Hi-Reco: High-Fidelity Real-Time Conversational Digital HumansCGI 2025real-time, conversational, high-fidelity, 3D avatar
2025Think Before You Talk: Enhancing Meaningful Dialogue Generation in Full-Duplex Speech Language Models with Planning-Inspired Text GuidanceArXiv 2025CodeProjectDialogue Generation, Speech Language Models
2025UniTalker: Conversational Speech-Visual SynthesisACM MM 2025Conversational, Multimodal, Emotion
2025MaAI: Real-time Generation of Various Types of Nodding for Avatar Attentive Listening SystemICMI 2025CodeReal-time, Nodding Generation, Avatar Interaction
2025MultiTalk: Let Them Talk: Audio-Driven Multi-Person Conversational Video GenerationArXiv 2025CodeProjectMulti-Person, Conversational
2025DualTalk: Dual-Speaker Interaction for 3D Talking Head ConversationsCVPR 2025CodeProject3D, Interaction, Dual-Speaker, Conversations, FLAME
2025VividListener: Expressive and Controllable Listener Dynamics Modeling for Multi-Modal Responsive InteractionArXiv 2025Listener dynamics, 3D dyadic conversation, Expressive control, Multi-modal conditions
2024INFP: Audio-Driven Interactive Head Generation in Dyadic ConversationsArXiv 2024ProjectDyadic Conversations
2024PerceptiveAgent: Talk With Human-like Agents: Empathetic Dialogue Through Perceptible Acoustic Reception and ReactionACL 2024CodeEmpathetic Dialogue
2024MultiDialog: Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationACL 2024CodeProjectDialogue, Face-to-Face Conversation
2023Emotional Listener Portrait: Realistic Listener Motion Simulation in ConversationICCV 2023Emotion, LHG
2022DialogueNeRF: Towards Realistic Avatar Face-to-face Conversation Video GenerationArXiv 2022Dialogue, Face-to-face Conversation

Talking Body & Avatar

YearTitleConference/JournalCodeProjectKeywords
2026Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture GenerationArXiv 2026co-speech gesture, object-grounded, diffusion, posture-aware
2026InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture ControlECCV 2026 WorkshopCodeProjectco-speech gesture, streaming, spatial control, diffusion, InteractGesture
2026Super Star: Towards Streaming Real-time Interactive Agents for Digital HumansACM MM 2026CodeProjectco-speech gesture, streaming, digital human, real-time, Super Star
2026Multi-View Face and Gesture Animation with Dynamic GaussiansSCA 2026Project3DGS, upper-body avatar, face and hands, animatable, multi-view
2026StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose AnchoringECCV 2026CodeProjectco-speech gesture, streaming, key-pose, StreamTalk, DiT
2026SICAGE: Speaker-Independent Culture-Aware Gesture Generation using TED4C-L DatasetECCV 2026co-speech gesture, culture-aware, SICAGE, TED4C-L, diffusion
2026EMOSH: Expressive Motion and Shape Disentanglement for Human AnimationECCV 2026CodeProjectfull-body avatar, video-driven, expression, shape disentanglement, human animation
2026SiGnature: Explicit Motion Diffusion for Stylized Semantic GestureArXiv 2026co-speech gesture, semantic gesture, style, diffusion, SiGnature
2026Semantic Motion Anchors: Bridging Motion and Meaning in Co-Speech GesturesArXiv 2026co-speech gesture, text-to-gesture, retrieval, semantic anchors, BEAT2
2026EchoAvatar: Real-time Generative Avatar Animation from Audio StreamsSIGGRAPH 2026CodeProjectfull-body motion, co-speech, streaming, speech and music, 3D character
2026LongCat-Video-Avatar 1.5 Technical ReportArXiv 2026CodeProjectaudio-driven, full-body avatar, lip-sync, long video, human animation
2026DuoGesture: Neuro-Inspired and Biomechanically Informed Dual-Stream Co-Speech Gesture GenerationArXiv 2026Projectco-speech gesture, DuoGesture, semantic-beat, dual-stream
2026UMo: Unified Sparse Motion Modeling for Real-Time Co-Speech AvatarsArXiv 2026co-speech avatar, real-time, mixture-of-experts, gesture generation, unified motion
2026PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion RepresentationArXiv 2026Projectco-speech gesture, personalization, VQ-VAE, semantic-aware, motion generation
2026PersonaGesture: Single-Reference Co-Speech Gesture Personalization for Unseen SpeakersArXiv 2026Projectco-speech gesture, personalization, diffusion, style transfer, single-reference
2026Reality Check: How Avatar and Face Representation Affect the Perceptual Evaluation of Synthesized GesturesArXiv 2026co-speech gesture, perceptual evaluation, avatar representation, user study, benchmarking
2026D-Rex : Diffusion Rendering for Relightable Expressive AvatarsArXiv 2026relighting, full-body avatar, diffusion, expressive animation, light stage
2026Prompt-to-Gesture: Measuring the Capabilities of Image-to-Video Deictic Gesture GenerationFG 2026Projectgesture generation
2026LiveGesture Streamable Co-Speech Gesture Generation ModelArXiv 2026Projectco-speech gesture, streaming, full-body, real-time
2026GenLCA: 3D Diffusion for Full-Body Avatars from In-the-Wild VideosArXiv 2026full-body avatar, 3D diffusion, photorealistic
2026SentiAvatar: Towards Expressive and Interactive Digital HumansArXiv 2026CodeProjectDigital Human, Expressive, Interactive, Sentiment
2026HolisticSemGes: Semantic Grounding of Holistic Co-Speech Gesture Generation with Contrastive Flow-MatchingArXiv 2026CodeProjectCo-Speech Gesture, Semantic Grounding, Flow-Matching
2026SoulX-LiveAct: Towards Hour-Scale Real-Time Human Animation with Neighbor Forcing and ConvKV MemoryArXiv 2026AR diffusion, real-time, hour-scale, human animation
2026MIBURI: Towards Expressive Interactive Gesture SynthesisCVPR 2026CodeProjectGesture Synthesis, Real-Time, LLM-Conditioned, Whole-Body
2026DyaDiT: A Multi-Modal Diffusion Transformer for Socially Favorable Dyadic Gesture GenerationArXiv 2026ProjectDyadic Gesture, Diffusion Transformer, Multi-Modal, Social Interaction
20263DGesPolicy: Phoneme-Aware Holistic Co-Speech Gesture Generation Based on Action ControlArXiv 2026co-speech gesture, diffusion policy, phoneme-aware, holistic motion
2026[SmoothSync: Dual-Stream Diffusion Transformers for Jitter

Truncated — view the full README on GitHub.

arxiv
audio-driven
paper
synthesis
talking-face-generation
talking-head
talking-head-video-generation

Contributors

Kedreamix

107 commits

satoooh

1 commits

xg-chu

1 commits

Languages

Python

100.0%