Awesome Human Motion
An aggregation of human motion understanding research; feel free to contribute.
Reviews & Surveys
(JEB 2025) McAllister et al : Behavioural energetics in human locomotion: how energy use influences how we move, McAllister et al.
(ICER 2025) Zhao et al : Motion Generation Review: Exploring Deep Learning for Lifelike Animation with Manifold, Zhao et al.
(ArXiv 2025) Advances in 4D Representation : Geometry, Motion, and Interaction, Zhao et al.
(ArXiv 2025) Motion Generation : A Survey of Generative Approaches and Benchmarks, Khani et al.
(ArXiv 2025) Segado et al : Grounding Intelligence in Movement, Segado et al.
(ArXiv 2025) Multimodal Generative AI with Autoregressive LLMs for Human Motion Understanding and Generation : A Way Forward, Islam et al.
(ArXiv 2025) Generative AI for Character Animation : A Comprehensive Survey of Techniques, Applications, and Future Directions, Abootorabi et al.
(ArXiv 2025) Sui et al : A Survey on Human Interaction Motion Generation, Sui et al.
(ArXiv 2025) 3D Human Interaction Generation : A Survey, Fan et al.
(ArXiv 2025) Human-Centric Foundation Models : Perception, Generation and Agentic Modeling, Tang et al.
(ArXiv 2025) Humanoid Locomotion and Manipulation : Current Progress and Challenges in Control, Planning, and Learning, Gu et al.
(ArXiv 2024) A Comprehensive Survey on Human Video Generation : Challenges, Methods, and Insights, Lei et al.
(ArXiv 2024) Human Motion Video Generation : A survey, Xue et al.
(Neurocomputing) Deep Learning for 3D Human Pose Estimation and Mesh Recovery : A survey, Liu et al.
(TVCG 2024) Loi et al : Machine Learning Approaches for 3D Motion Synthesis and Musculoskeletal Dynamics Estimation: A Survey, Loi et al.
(T-PAMI 2023) Zhu et al : Human Motion Generation: A Survey, Zhu et al.
Motion Generation, Text/Speech/Music-Driven
2026
(ArXiv 2026) MoRAE : Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation, Zhu et al.
(SIGGRAPH 2026) ARDY : Autoregressive Diffusion with Hybrid Representation for Interactive Human Motion Generation, Zhao et al.
(ECCV 2026) IRG-MotionLLM : Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation, Li et al.
(ECCV 2026) Odoriko : A Shape-Aware Multimodal Diffusion Framework for Human Motion, Shim et al.
(ArXiv 2026) MotionVLA : Vision-Language-Action Model for Humanoid Motion, Zhang et al.
(ArXiv 2026) DC-Motion : Decoupling Semantics and Details via Discrete-Continuous Tokens for Human Motion Generation, Wang et al.
(ArXiv 2026) VideoMDM : Towards 3D Human Motion Generation From 2D Supervision, Mann et al.
(ArXiv 2026) Sketch2Motion : Text-driven 2D Sketch to 3D Animation via Diffusion-guided Skeleton Optimization, Rai et al.
(ArXiv 2026) EchoAvatar : Real-time Generative Avatar Animation from Audio Streams, Chen et al.
(CVPR 2026) RoMo : A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation, Zhang et al.
(ArXiv 2026) Latent Dynamics for Full Body Avatar Animation , Peng et al.
(ArXiv 2026) DrawMotion : Generating 3D Human Motions by Freehand Drawing, Wang et al.
(ArXiv 2026) ScaleMoGen : Autoregressive Next-Scale Prediction for Human Motion Generation, Hwang et al.
(ArXiv 2026) IAM : Identity-Aware Human Motion and Shape Joint Generation, Jia et al.
(ArXiv 2026) Re2 MoGen : Open-Vocabulary Motion Generation via LLM Reasoning and Physics-Aware Refinement, Zheng et al.
(ArXiv 2026) Marrying Text-to-Motion Generation with Skeleton-Based Action Recognition , Kuang et al.
(ArXiv 2026) Motion-Adapter : A Diffusion Model Adapter for Text-to-Motion Generation of Compound Actions, Jiang et al.
(ArXiv 2026) A Unified Conditional Flow for Motion Generation, Editing, and Intra-Structural Retargeting : unifies text-driven motion editing and retargeting via conditional transport and rectified flow, Li et al.
(CVPR 2026) Next-Scale Autoregressive Models : for Text-to-Motion Generation, Zheng et al.
(ArXiv 2026) BiTDiff : Fine-Grained 3D Conducting Motion Generation via BiMamba-Transformer Diffusion, Jia et al.
(ArXiv 2026) FlowCoMotion : Text-to-Motion Generation via Token-Latent Flow Modeling, Guan et al.
(ArXiv 2026) Exploring Motion-Language Alignment : for Text-driven Motion Generation, Gu et al.
(ArXiv 2026) MotionRFT : Unified Reinforcement Fine-Tuning for Text-to-Motion Generation, Tan et al.
(ArXiv 2026) Unified Number-Free Text-to-Motion Generation Via Flow Matching , Huang et al.
(ArXiv 2026) From Diffusion To Flow : Efficient Motion Generation In MotionGPT3, Ban et al.
(ArXiv 2026) Bilingual Text-to-Motion Generation : A New Benchmark and Baselines, Weng et al.
(ArXiv 2026) UniMotion : A Unified Framework for Motion-Text-Vision Understanding and Generation, Wang et al.
(ArXiv 2026) Controllable Text-to-Motion Generation : via Modular Body-Part Phase Control, Dai et al.
(ArXiv 2026) MoTok : Bridging Semantic and Kinematic Conditions with Diffusion-based Discrete Motion Tokenizer, Gu et al.
(ArXiv 2026) OpenT2M : No-frill Motion Generation with Open-source, Large-scale, High-quality Data, Cao et al.
(ArXiv 2026) UMO : Unified In-Context Learning Unlocks Motion Foundation Model Priors, Cong et al.
(ArXiv 2026) Kimodo : Scaling Controllable Human Motion Generation, Rempe et al.
(ArXiv 2026) Riemannian Motion Generation : A Unified Framework for Human Motion Representation and Generation via Riemannian Flow Matching, Miao et al.
(ArXiv 2026) ActionPlan : Future-Aware Streaming Motion Synthesis via Frame-Level Action Planning, Nazarenus et al.
(CVPR 2026) LaMoGen : Language to Motion Generation Through LLM-Guided Symbolic Inference, Jiang et al.
(ArXiv 2026) ParTY : Part-Guidance for Expressive Text-to-Motion Synthesis, Heo et al.
(ArXiv 2026) PRISM : Streaming Human Motion Generation with Per-Joint Latent Decomposition, Ling et al.
(CVPR 2026) CMDM : Causal Motion Diffusion Models for Autoregressive Motion Generation, Yu et al.
(ArXiv 2026) TCA-T2M : Temporal Consistency-Aware Text-to-Motion Generation, Wang et al.
(ArXiv 2026) DMC : A Self-Supervised Approach on Motion Calibration for Enhancing Physical Plausibility in Text-to-Motion, Shim et al.
(ArXiv 2026) DiMo : Discrete Diffusion Modeling for Motion Generation and Understanding, Zhang et al.
(ArXiv 2026) LG-Tok : Language-Guided Transformer Tokenizer for Human Motion Generation, Yan et al.
(ArXiv 2026) TriC-Motion : Tri-Domain Causal Modeling Grounded Text-to-Motion Generation, Cao et al.
(ArXiv 2026) FrankenMotion : Part-level Human Motion Generation and Composition, Li et al.
(ArXiv 2026) CoMoVi : Co-Generation of 3D Human Motions and Realistic Videos, Zhao et al.
(ICLR 2026) EasyTune : Efficient Step-Aware Fine-Tuning for Diffusion-Based Motion Generation, Tan et al.
(WACV 2026) SegMo : Segment-aligned Text to 3D Human Motion Generation, Dang et al.
(AAAI 2026) ReAlign : Bilingual Text-to-Motion Generation via Step-Aware Reward-Guided Alignment, Weng et al.
(AAAI 2026) FineXtrol : Controllable Motion Generation via Fine-Grained Text, Shen et al.
2025
(NeurIPS 2025) HMVLM : Human Motion-Vision-Lanuage Model via MoE LoRA, Hu et al.
(NeurIPS 2025) TransPhase : Deep Compositional Phase Diffusion for Long Motion Sequence Generation, Au et al.
(NeurIPS 2025) MEGADance : Mixture-of-experts architecture for genre-aware 3d dance generation, Yang et al.
(SIGGRAPH Asia 2025) TCM : Learning Human Motion with Temporally Conditional Mamba, Nguyen et al.
(TMLR 2025) MoReact : Generating Reactive Motion from Textual Descriptions, Xu et al.
(ICCV 2025) Align Your Rhythm : Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation, Fan et al.
(ICCV 2025) UniEgoMotion : A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation, Patel et al.
(ICCV 2025) FineMotion : A Dataset and Benchmark with both Spatial and Temporal Annotation for Fine-grained Motion Generation and Editing, Wu et al.
(ICCV 2025) PUMPS : Skeleton-Agnostic Point-based Universal Motion Pre-Training for Synthesis in Human Motion Tasks, Mo et al.
(ICCV 2025) GENMO : A GENeralist Model for Human MOtion, Li et al.
(ICCV 2025) InfiniDreamer : Arbitrarily Long Human Motion Generation via Segment Score Distillation, Zhuo et al.
(ICCV 2025) Go to Zero : Towards Zero-shot Motion Generation with Million-scale Data, Fan et al.
(ICCV 2025) Morph : A Motion-free Physics Optimization Framework for Human Motion Generation, Li et al.
(ICCV 2025) DisCoRD : Discrete Tokens to Continuous Motion via Rectified Flow Decoding, Cho et al.
(ICCV 2025) SemTalk : Holistic Co-speech Motion Generation with Frame-level Semantic Emphasis, Zhang et al.
(ICCV 2025) KinMo : Kinematic-aware Human Motion Understanding and Generation, Zhang et al.
(ICCV 2025) GestureLSM : Latent Shortcut-based Co-Speech Gesture Generation with Spatial-Temporal Modeling, Liu et al.
(ICCV 2025) Motion-2-to-3 : Leveraging 2D Motion Data to Boost 3D Motion Generation, Pi et al.
(ICCV 2025) MotionLab : Unified Human Motion Generation and Editing via the Motion-Condition-Motion Paradigm, Guo et al.
(ICCV 2025) SFControl : Motion Synthesis with Sparse and Flexible Keyjoint Control, Hwang et al.
(ICCV 2025) Less Is More : Improving Motion Diffusion Models with Sparse Keyframes, Bae et al.
(ICCV 2025) ControlMM : Controllable Masked Motion Generation, Pinyoanuntapong et al.
(ICCV 2025) PRIMAL : Physically Reactive and Interactive Motor Model for Avatar Learning, Zhang et al.
(ICCV 2025) HERO : Human Reaction Generation from Videos, Yu et al.
(ICCV 2025) MotionStreamer : Streaming Motion Generation via Diffusion-based Autoregressive Model in Causal Latent Space, Xiao et al.
(ICCV 2025) GenM3 : Generative Pretrained Multi-path Motion Model for Text Conditional Human Motion Generation, Shi et al.
(ACM MM 2025) ChoreoMuse : Robust Music-to-Dance Video Generation with Style Transfer and Beat-Adherent Motion, Wang et al.
(ICML 2025) Being-M0 : Scaling Motion Generation Models with Million-Level Human Motions, Wang et al.
(TOG 2025) Sketch2Anim : Towards Transferring Sketch Storyboards into 3D Animation, Zhong et al.
(SIGGRAPH 2025) MECo : Motion-example-controlled Co-speech Gesture Generation Leveraging Large Language Models, Chen et al.
(SIGGRAPH 2025) Chang et al. : Large-Scale Multi-Character Interaction Synthesis, Chang et al.
(SIGGRAPH 2025) AnyTop : Character Animation Diffusion with Any Topology, Gat et al.
(CVPR 2025) DSDFM : Deterministic-to-Stochastic Diverse Latent Feature Mapping for Human Motion Synthesis, Hua et al.
(CVPR 2025) EchoMimicV2 : Towards Striking, Simplified, and Semi-Body Human Animation, Hua et al.
(CVPR 2025) UniPose : A Unified Multimodal Framework for Human Pose Comprehension, Generation and Editing, Li et al.
(CVPR 2025) From Sparse Signal to Smooth Motion : Real-Time Motion Generation with Rolling Prediction Models, Barquero et al.
(CVPR 2025) Shape My Moves : Text-Driven Shape-Aware Synthesis of Human Motions, Liao et al.
(CVPR 2025) MG-MotionLLM : A Unified Framework for Motion Comprehension and Generation across Multiple Granularities, Wu et al.
(CVPR 2025) SALAD : Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing, Hong et al.
(CVPR 2025) PersonalBooth : Personalized Text-to-Motion Generation, Kim et al.
(CVPR 2025) MARDM : Rethinking Diffusion for Text-Driven Human Motion Generation, Meng et al.
(CVPR 2025) StickMotion : Generating 3D Human Motions by Drawing a Stickman, Wang et al.
(CVPR 2025) LLaMo : Human Motion Instruction Tuning, Li et al.
(CVPR 2025) HOP : Heterogeneous Topology-based Multimodal Entanglement for Co-Speech Gesture Generation, Cheng et al.
(CVPR 2025) AtoM : Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward, Han et al.
(CVPR 2025) EnergyMoGen : Compositional Human Motion Generation with Energy-Based Diffusion Model in Latent Space, Zhang et al.
(CVPR 2025) The Languate of Motion : Unifying Verbal and Non-verbal Language of 3D Human Motion, Chen et al.
(CVPR 2025) ScaMo : Exploring the Scaling Law in Autoregressive Motion Generation Model, Lu et al.
(CVPR 2025) Move in 2D : 2D-Conditioned Human Motion Generation, Huang et al.
(CVPR 2025) SOLAMI : Social Vision-Language-Action Modeling for Immersive Interaction with 3D Autonomous Characters, Jiang et al.
(CVPR 2025) MVLift : Lifting Motion to the 3D World via 2D Diffusion, Li et al.
(CVPR 2025 Workshop) MoCLIP : Motion-Aware Fine-Tuning and Distillation of CLIP for Human Motion Generation, Maldonado et al.
(CVPR 2025 Workshop) Dyadic Mamba : Long-term Dyadic Human Motion Synthesis, Tanke et al.
(ACM Sensys 2025) SHADE-AD : An LLM-Based Framework for Synthesizing Activity Data of Alzheimer’s Patients, Fu et al.
(ICRA 2025) MotionGlot : A Multi-Embodied Motion Generation Model, Harithas et al.
(ICLR 2025) CLoSD : Closing the Loop between Simulation and Diffusion for Multi-Task Character Control, Tevet et al.
(ICLR 2025) PedGen : Learning to Generate Diverse Pedestrian Movements from Web Videos with Noisy Labels, Liu et al.
(ICLR 2025) HGM³ : Hierarchical Generative Masked Motion Modeling with Hard Token Mining, Jeong et al.
(ICLR 2025) LaMP : Language-Motion Pretraining for Motion Generation, Retrieval, and Captioning, Li et al.
(ICLR 2025) MotionDreamer : One-to-Many Motion Synthesis with Localized Generative Masked Transformer, Wang et al.
(ICLR 2025) Lyu et al : Towards Unified Human Motion-Language Understanding via Sparse Interpretable Characterization, Lyu et al.
(ICLR 2025) DART : A Diffusion-Based Autoregressive Motion Model for Real-Time Text-Driven Motion Control, Zhao et al.
(ICLR 2025) Motion-Agent : A Conversational Framework for Human Motion Generation with LLMs, Wu et al.
(TMM 2025) MCG-IMM : A Plug-and-Play Multi-Criteria Guidance for Diverse In-Betweening Human Motion Generation, Yu et al.
(IJCV 2025) Fg-T2M++ : LLMs-Augmented Fine-Grained Text Driven Human Motion Generation, Wang et al.
(TCSVT 2025) Zeng et al : Progressive Human Motion Generation Based on Text and Few Motion Frames, Zeng et al.
(Arxiv 2025) HY-Motion 1.0 : Scaling Flow Matching Models for Text-To-Motion Generation, Wen et al.
(Arxiv 2025) DeMoGen : Towards Decompositional Human Motion Generation with Energy-Based Diffusion Models, Zhang et al.
(Arxiv 2025) Jeong et al : Pose-Guided Residual Refinement for Interpretable Text-to-Motion Generation and Editing, Jeong et al.
(Arxiv 2025) TempoMOE : Tempo as the Stable Cue: Hierarchical Mixture of Tempo and Beat Experts for Music to 3D Dance Generation, Lyu et al.
(Arxiv 2025) FlowerDance : MeanFlow for Efficient and Refined 3D Dance Generation, Yang et al.
(ArXiv 2025) OmniMoGen : Unifying Human Motion Generation via Learning from Interleaved Text-Motion Instructions, Bu et al.
(ArXiv 2025) MoLingo : Motion–Language Alignment for Text-to-Human Motion Generation, He et al.
(ArXiv 2025) FunPhase : A Periodic Functional Autoencoder for Motion Generation via Phase Manifolds, Pegoraro et al.
(ArXiv 2025) Kinetic Mining in Context : Few-Shot Action Synthesis via Text-to-Motion Distillations, Cazzola et al.
(ArXiv 2025) COMET : Controllable Long-term Motion Generation with Extended Joint Targets, Li et al.
(ArXiv 2025) Back to Basics : Motion Representation Matters for Human Motion Generation Using Diffusion Model, Jin et al.
(ArXiv 2025) UniMo : Unifying 2D Video and 3D Human Motion with an Autoregressive Framework, Pang et al.
(ArXiv 2025) Free3D : 3D Human Motion Emerges from Single-View 2D Supervision, Liu et al.
(ArXiv 2025) Pressure2Motion : Hierarchical Motion Synthesis from Ground Pressure with Text Guidance, Li et al.
(ArXiv 2025) Mem-MLP : Real-Time 3D Human Motion Generation from Sparse Inputs, Mutlu et al.
(ArXiv 2025) The Quest for Generalizable Motion Generation : Data, Model, and Evaluation, Lin et al.
(ArXiv 2025) MoSa : Motion Generation with Scalable Autoregressive Modeling, Liu et al.
(ArXiv 2025) OmniMotion-X : Versatile Multimodal Whole-Body Motion Generation, Xu et al.
(ArXiv 2025) OmniMotion : Multimodal Motion Generation with Continuous Masked Autoregression, Li et al.
(ArXiv 2025) No MoCap Needed : Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts, Girolamo et al.
(ArXiv 2025) Pulp Motion : Framing-aware multimodal camera and human motion generation, Courant et al.
(ArXiv 2025) MonSTeR : a Unified Model for Motion, Scene, Text Retrieval, Collorone et al.
(ArXiv 2025) MoGIC : Boosting Motion Generation via Intention Understanding and Visual Context, Shi et al.
(ArXiv 2025) Gupta et al : Unified Multi-Modal Interactive & Reactive 3D Motion Generation via Rectified Flow, Gupta et al.
(ArXiv 2025) LaMoGen : Laban Movement-Guided Diffusion for Text-to-Motion Generation, Kim et al.
(ArXiv 2025) LUMA : Low-Dimension Unified Motion Alignment with Dual-Path Anchoring for Text-to-Motion Diffusion Model, Jia et al.
(ArXiv 2025) SimDiff : Simulator-constrained Diffusion Model for Physically Plausible Motion Generation, Watanabe et al.
(ArXiv 2025) SmooGPT : Stylized Motion Generation using Large Language Models, Zhong et al.
(ArXiv 2025) Embracing Aleatoric Uncertainty : Generating Diverse 3D Human Motion, Qin et al.
(ArXiv 2025) MotionFLUX : Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment, Gao et al.
(ArXiv 2025) VimoRAG : Video-based Retrieval-augmented 3D Motion Generation for Motion Language Models, Xu et al.
(ArXiv 2025) MSQ : Spatial-Temporal Multi-Scale Quantizationfor Flexible Motion Generation, Wang et al.
(ArXiv 2025) X-MoGen : Unified Motion Generation across Humans and Animals, Wang et al.
(ArXiv 2025) ReMoMask : Retrieval-Augmented Masked Motion Generation, Li et al.
(ArXiv 2025) OmniAvatar : Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation, Gan et al.
(ArXiv 2025) SpeakerVid-5M : A Large-Scale High-Quality Dataset for audio-visual Dyadic Interactive Human Generation, Zhang et al.
(ArXiv 2025) EchoMimicV3 : 1.3B Parameters are All You Need for Unified Multi-Modal and Multi-Task Human Animation, Meng et al.
(ArXiv 2025) MOSPA : Human Motion Generation Driven by Spatial Audio, Xu et al.
(ArXiv 2025) SnapMoGen : Human Motion Generation from Expressive Texts, Wang et al.
(ArXiv 2025) MOST : Motion Diffusion Model for Rare Text via Temporal Clip Banzhaf Interaction, Wang et al.
(ArXiv 2025) Grounded Gestures : Language, Motion and Space, Deichler et al.
(ArXiv 2025) MotionGPT3 : Human Motion as a Second Modality, Zhu et al.
(ArXiv 2025) HumanAttr : Generating Attribute-Aware Human Motions from Textual Prompt, Wang et al.
(ArXiv 2025) PlanMoGPT : Flow-Enhanced Progressive Planning for Text to Motion Synthesis, Jin et al.
(ArXiv 2025) Motion-R1 : Chain-of-Thought Reasoning and Reinforcement Learning for Human Motion Generation, Ouyang et al.
(ArXiv 2025) ANT : Adaptive Neural Temporal-Aware Text-to-Motion Model, Chen et al.
(ArXiv 2025) MotionRAG-Diff : A Retrieval-Augmented Diffusion Framework for Long-Term Music-to-Dance Generation, Huang et al.
(ArXiv 2025) IKMo : Image-Keyframed Motion Generation with Trajectory-Pose Conditioned Motion Diffusion Model, Zhao et al.
(ArXiv 2025) Li et al : How Much Do Large Language Models Know about Human Motion? A Case Study in 3D Avatar Control, Li et al.
(ArXiv 2025) UniMoGen : Universal Motion Generation, Khani et al.
(ArXiv 2025) Wang et al : Semantics-Aware Human Motion Generation from Audio Instructions, Wang et al.
(ArXiv 2025) ACMDM : Absolute Coordinates Make Motion Generation Easy, Meng et al.
(ArXiv 2025) PAMD : Plausibility-Aware Motion Diffusion Model for Long Dance Generation, Zhu et al.
(ArXiv 2025) Intentional Gesture : Deliver Your Intentions with Gestures for Speech, Liu et al.
(ArXiv 2025) MatchDance : Collaborative Mamba-Transformer Architecture Matching for High-Quality 3D Dance Synthesis, Yang et al.
(ArXiv 2025) M3G : Multi-Granular Gesture Generator for Audio-Driven Full-Body Human Motion Synthesis, Yin et al.
(ArXiv 2025) ReactDance : Progressive-Granular Representation for Long-Term Coherent Reactive Dance Generation, Lin et al.
(ArXiv 2025) PMG : Progressive Motion Generation via Sparse Anchor Postures Curriculum Learning, Xi et al.
(ArXiv 2025) DanceMosaic : High-Fidelity Dance Generation with Multimodal Editability, Shah et al.
(ArXiv 2025) ReCoM : Realistic Co-Speech Motion Generation with Recurrent Embedded Transformer, Xie et al.
(ArXiv 2025) HMU : Human Motion Unlearning, Matteis et al.
(ArXiv 2025) ACMo : Attribute Controllable Motion Generation, Wei et al.
(ArXiv 2025) BioMoDiffuse : Physics-Guided Biomechanical Diffusion for Controllable and Authentic Human Motion Synthesis, Kang et al.
(ArXiv 2025) ExGes : Expressive Human Motion Retrieval and Modulation for Audio-Driven Gesture Synthesis, Zhou et al.
(ArXiv 2025) Motion Anything : Any to Motion Generation, Zhang et al.
(ArXiv 2025) GCDance : Genre-Controlled 3D Full Body Dance Generation Driven By Music, Liu et al.
(ArXiv 2025) CASIM : Composite Aware Semantic Injection for Text to Motion Generation, Chang et al.
(ArXiv 2025) MotionPCM : Real-Time Motion Synthesis with Phased Consistency Model, Jiang et al.
(ArXiv 2025) Free-T2M : Frequency Enhanced Text-to-Motion Diffusion Model With Consistency Loss, Chen et al.
(ArXiv 2025) FlexMotion : Lightweight, Physics-Aware, and Controllable Human Motion Generation, Tashakori et al.
(ArXiv 2025) HiSTF Mamba : Hierarchical Spatiotemporal Fusion with Multi-Granular Body-Spatial Modeling for High-Fidelity Text-to-Motion Generation, Zhan et al.
(ArXiv 2025) PackDiT : Joint Human Motion and Text Generation via Mutual Prompting, Jiang et al.
(3DV 2025) Unimotion : Unifying 3D Human Motion Synthesis and Understanding, Li et al.
(3DV 2025) HoloGest : Decoupled Diffusion and Motion Priors for Generating Holisticly Expressive Co-speech Gestures, Cheng et al.
(AAAI 2025) RemoGPT : Part-Level Retrieval-Augmented Motion-Language Models, Yu et al.
(AAAI 2025) UniMuMo : Unified Text, Music and Motion Generation, Yang et al.
(AAAI 2025) EchoMimic : Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditioning, Chen et al.
(AAAI 2025) ALERT-Motion : Autonomous LLM-Enhanced Adversarial Attack for Text-to-Motion, Miao et al.
(AAAI 2025) MotionCraft : Crafting Whole-Body Motion with Plug-and-Play Multimodal Controls, Bian et al.
(AAAI 2025) Light-T2M : A Lightweight and Fast Model for Text-to-Motion Generation, Zeng et al.
(WACV 2025 Worhshop) LS-GAN : Human Motion Synthesis with Latent-space GANs, Amballa et al.
(WACV 2025) ReinDiffuse : Crafting Physically Plausible Motions with Reinforced Diffusion Model, Han et al.
(WACV 2025) MoRAG : Multi-Fusion Retrieval Augmented Generation for Human Motion, Shashank et al.
(WACV 2025) Mandelli et al : Generation of Complex 3D Human Motion by Temporal and Spatial Composition of Diffusion Models, Mandelli et al.
2024
(ArXiv 2024) MMoFusion : Multi-modal Co-Speech Motion Generation with Diffusion Model, Wang et al.
(ArXiv 2024) InterDance : Reactive 3D Dance Generation with Realistic Duet Interactions, Li et al.
(ArXiv 2024) Mogo : RQ Hierarchical Causal Transformer for High-Quality 3D Human Motion Generation, Fu et al.
(ArXiv 2024) CoMA : Compositional Human Motion Generation with Multi-modal Agents, Sun et al.
(ArXiv 2024) SoPo : Text-to-Motion Generation Using Semi-Online Preference Optimization, Tan et al.
(ArXiv 2024) RMD : A Simple Baseline for More General Human Motion Generation via Training-free Retrieval-Augmented Motion Diffuse, Liao et al.
(ArXiv 2024) BiPO : Bidirectional Partial Occlusion Network for Text-to-Motion Synthesis, Hong et al.
(ArXiv 2024) MoTe : Learning Motion-Text Diffusion Model for Multiple Generation Tasks, Wue et al.
(ArXiv 2024) FTMoMamba : Motion Generation with Frequency and Text State Space Models, Li et al.
(ArXiv 2024) KMM : Key Frame Mask Mamba for Extended Motion Generation, Zhang et al.
(ArXiv 2024) MotionGPT-2 : A General-Purpose Motion-Language Model for Motion Generation and Understanding, Wang et al.
(ArXiv 2024) Lodge++ : High-quality and Long Dance Generation with Vivid Choreography Patterns, Li et al.
(ArXiv 2024) MotionCLR : Motion Generation and Training-Free Editing via Understanding Attention Mechanisms, Chen et al.
(ArXiv 2024) LEAD : Latent Realignment for Human Motion Diffusion, Andreou et al.
(ArXiv 2024) Leite et al. Enhancing Motion Variation in Text-to-Motion Models via Pose and Video Conditioned Editing, Leite et al.
(ArXiv 2024) MotionRL : Align Text-to-Motion Generation to Human Preferences with Multi-Reward Reinforcement Learning, Liu et al.
(ArXiv 2024) MotionLLM : Understanding Human Behaviors from Human Motions and Videos, Chen et al.
(ArXiv 2024) T2M-X : Learning Expressive Text-to-Motion Generation from Partially Annotated Data, Liu et al.
(ArXiv 2024) BAD : Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation, Hosseyni et al.
(ArXiv 2024) synNsync : Synergy and Synchrony in Couple Dances, Manukele et al.
(EMNLP 2024) Dong et al : Word-Conditioned 3D American Sign Language Motion Generation, Dong et al.
(NeurIPS D&B 2024) Kim et al : Text to Blind Motion, Kim et al.
(NeurIPS 2024) UniMTS : Unified Pre-training for Motion Time Series, Zhang et al.
(NeurIPS 2024) Christopher et al. : Constrained Synthesis with Projected Diffusion Models, Christopher et al.
(NeurIPS 2024) MoMu-Diffusion : On Learning Long-Term Motion-Music Synchronization and Correspondence, You et al.
(NeurIPS 2024) MoGenTS : Motion Generation based on Spatial-Temporal Joint Modeling, Yuan et al.
(NeurIPS 2024) M3GPT : An Advanced Multimodal, Multitask Framework for Motion Comprehension and Generation, Luo et al.
(NeurIPS Workshop 2024) Bikov et al : Fitness Aware Human Motion Generation with Fine-Tuning, Bikov et al.
(NeurIPS Workshop 2024) DGFM : Full Body Dance Generation Driven by Music Foundation Models, Liu et al.
(ICPR 2024) FG-MDM : Towards Zero-Shot Human Motion Generation via ChatGPT-Refined Descriptions, Shi et al.
(ACM MM 2024) SynTalker : Enabling Synergistic Full-Body Control in Prompt-Based Co-Speech Motion Generation, Chen et al.
(ACM MM 2024) L3EM : Towards Emotion-enriched Text-to-Motion Generation via LLM-guided Limb-level Emotion Manipulating. Yu et al.
(ACM MM 2024) StableMoFusion : Towards Robust and Efficient Diffusion-based Motion Generation Framework, Huang et al.
(ACM MM 2024) SATO : Stable Text-to-Motion Framework, Chen et al.
(ICANN 2024) PIDM : Personality-Aware Interaction Diffusion Model for Gesture Generation, Shibasaki et al.
(HFES 2024) Macwan et al : High-Fidelity Worker Motion Simulation With Generative AI, Macwan et al.
(ECCV 2024) Jin et al : Local Action-Guided Motion Diffusion Model for Text-to-Motion Generation, Jin et al.
(ECCV 2024) Motion Mamba : Efficient and Long Sequence Motion Generation, Zhong et al.
(ECCV 2024) EMDM : Efficient Motion Diffusion Model for Fast, High-Quality Human Motion Generation, Zhou et al.
(ECCV 2024) CoMo : Controllable Motion Generation through Language Guided Pose Code Editing, Huang et al.
(ECCV 2024) CoMusion : Towards Consistent Stochastic Human Motion Prediction via Motion Diffusion, Sun et al.
(ECCV 2024) Shan et al : Towards Open Domain Text-Driven Synthesis of Multi-Person Motions, Shan et al.
(ECCV 2024) ParCo : Part-Coordinating Text-to-Motion Synthesis, Zou et al.
(ECCV 2024) Sampieri et al : Length-Aware Motion Synthesis via Latent Diffusion, Sampieri et al.
(ECCV 2024) ChroAccRet : Chronologically Accurate Retrieval for Temporal Grounding of Motion-Language Models, Fujiwara et al.
(ECCV 2024) MHC : Generating Physically Realistic and Directable Human Motions from Multi-Modal Inputs, Liu et al.
(ECCV 2024) ProMotion : Plan, Posture and Go: Towards Open-vocabulary Text-to-Motion Generation, Liu et al.
(ECCV 2024) FreeMotion : MoCap-Free Human Motion Synthesis with Multimodal Large Language Models, Zhang et al.
(ECCV 2024) Text Motion Translator : A Bi-Directional Model for Enhanced 3D Human Motion Generation from Open-Vocabulary Descriptions, Qian et al.
(ECCV 2024) FreeMotion : A Unified Framework for Number-free Text-to-Motion Synthesis, Fan et al.
(ECCV 2024) Kinematic Phrases : Bridging the Gap between Human Motion and Action Semantics via Kinematic Phrases, Liu et al.
(ECCV 2024) MotionChain : Conversational Motion Controllers via Multimodal Prompts, Jiang et al.
(ECCV 2024) SMooDi : Stylized Motion Diffusion Model, Zhong et al.
(ECCV 2024) BAMM : Bidirectional Autoregressive Motion Model, Pinyoanuntapong et al.
(ECCV 2024) MotionLCM : Real-time Controllable Motion Generation via Latent Consistency Model, Dai et al.
(ECCV 2024) Ren et al : Realistic Human Motion Generation with Cross-Diffusion Models, Ren et al.
(ECCV 2024) M2D2M : Multi-Motion Generation from Text with Discrete Diffusion Models, Chi et al.
(ECCV 2024) LMM : Large Motion Model for Unified Multi-Modal Motion Generation, Zhang et al.
(ECCV 2024) TesMo : Generating Human Interaction Motions in Scenes with Text Control, Yi et al.
(ECCV 2024) TLcontrol : Trajectory and Language Control for Human Motion Synthesis, Wan et al.
(ICME 2024) ExpGest : Expressive Speaker Generation Using Diffusion Model and Hybrid Audio-Text Guidance, Cheng et al.
(ICME Workshop 2024) Chen et al : Anatomically-Informed Vector Quantization Variational Auto-Encoder for Text-to-Motion Generation, Chen et al.
(ICML 2024) HumanTOMATO : Text-aligned Whole-body Motion Generation, Lu et al.
(ICML 2024) GPHLVM : Bringing Motion Taxonomies to Continuous Domains via GPLVM on Hyperbolic Manifolds, Jaquier et al.
(SIGGRAPH 2024) DiffPoseTalk : Speech-Driven Stylistic 3D Facial Animation and Head Pose Generation via Diffusion Models, Sun et al.
(SIGGRAPH 2024) CondMDI : Flexible Motion In-betweening with Diffusion Models, Cohan et al.
(SIGGRAPH 2024) CAMDM : Taming Diffusion Probabilistic Models for Character Control, Chen et al.
(SIGGRAPH 2024) LGTM : Local-to-Global Text-Driven Human Motion Diffusion Models, Sun et al.
(SIGGRAPH 2024) TEDi : Temporally-Entangled Diffusion for Long-Term Motion Synthesis, Zhang et al.
(SIGGRAPH 2024) A-MDM : Interactive Character Control with Auto-Regressive Motion Diffusion Models, Shi et al.
(SIGGRAPH 2024) Starke et al : Categorical Codebook Matching for Embodied Character Controllers, Starke et al.
(SIGGRAPH 2024) SuperPADL : Scaling Language-Directed Physics-Based Control with Progressive Supervised Distillation, Juravsky et al.
(CVPR 2024) ProgMoGen : Programmable Motion Generation for Open-set Motion Control Tasks, Liu et al.
(CVPR 2024) PACER+ : On-Demand Pedestrian Animation Controller in Driving Scenarios, Wang et al.
(CVPR 2024) AMUSE : Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion, Chhatre et al.
(CVPR 2024) Liu et al : Towards Variable and Coordinated Holistic Co-Speech Motion Generation, Liu et al.
(CVPR 2024) MAS : Multi-view Ancestral Sampling for 3D motion generation using 2D diffusion, Kapon et al.
(CVPR 2024) WANDR : Intention-guided Human Motion Generation, Diomataris et al.
(CVPR 2024) MoMask : Generative Masked Modeling of 3D Human Motions, Guo et al.
(CVPR 2024) ChatPose : Chatting about 3D Human Pose, Feng et al.
(CVPR 2024) AvatarGPT : All-in-One Framework for Motion Understanding, Planning, Generation and Beyond, Zhou et al.
(CVPR 2024) MMM : Generative Masked Motion Model, Pinyoanuntapong et al.
(CVPR 2024) AAMDM : Accelerated Auto-regressive Motion Diffusion Model, Li et al.
(CVPR 2024) OMG : Towards Open-vocabulary Motion Generation via Mixture of Controllers, Liang et al.
(CVPR 2024) FlowMDM : Seamless Human Motion Composition with Blended Positional Encodings, Barquero et al.
(CVPR 2024) Digital Life Project : Autonomous 3D Characters with Social Intelligence, Cai et al.
(CVPR 2024) EMAGE : Towards Unified Holistic Co-Speech Gesture Generation via Expressive Masked Audio Gesture Modeling, Liu et al.
(CVPR Workshop 2024) STMC : Multi-Track Timeline Control for Text-Driven 3D Human Motion Generation, Petrovich et al.
(CVPR Workshop 2024) InstructMotion : Exploring Text-to-Motion Generation with Human Preference, Sheng et al.
(ICLR 2024) Single Motion Diffusion : Raab et al.
(ICLR 2024) NeRM : Learning Neural Representations for High-Framerate Human Motion Synthesis, Wei et al.
(ICLR 2024) PriorMDM : Human Motion Diffusion as a Generative Prior, Shafir et al.
(ICLR 2024) OmniControl : Control Any Joint at Any Time for Human Motion Generation, Xie et al.
(ICLR 2024) Adiya et al. : Bidirectional Temporal Diffusion Model for Temporally Consistent Human Animation, Adiya et al.
(ICLR 2024) Duolando : Follower GPT with Off-Policy Reinforcement Learning for Dance Accompaniment, Li et al.
(AAAI 2024) HuTuDiffusion : Human-Tuned Navigation of Latent Motion Diffusion Models with Minimal Feedback, Han et al.
(AAAI 2024) AMD : Anatomical Motion Diffusion with Interpretable Motion Decomposition and Fusion, Jing et al.
(AAAI 2024) MotionMix : Weakly-Supervised Diffusion for Controllable Motion Generation, Hoang et al.
(AAAI 2024) B2A-HDM : Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion Model, Xie et al.
(AAAI 2024) Everything2Motion : Everything2Motion: Synchronizing Diverse Inputs via a Unified Framework for Human Motion Synthesis, Fan et al.
(AAAI 2024) MotionGPT : Finetuned LLMs are General-Purpose Motion Generators, Zhang et al.
(AAAI 2024) Dong et al : Enhanced Fine-grained Motion Diffusion for Text-driven Human Motion Synthesis, Dong et al.
(AAAI 2024) UNIMASKM : A Unified Masked Autoencoder with Patchified Skeletons for Motion Synthesis, Mascaro et al.
(AAAI 2024) B2A-HDM : Towards Detailed Text-to-Motion Synthesis via Basic-to-Advanced Hierarchical Diffusion Model, Xie et al.
(TPAMI 2024) GUESS : GradUally Enriching SyntheSis for Text-Driven Human Motion Generation, Gao et al.
(WACV 2024) Xie et al. : Sign Language Production with Latent Motion Transformer, Xie et al.
2023
(NeurIPS 2023) GraphMotion : Act As You Wish: Fine-grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs, Jin et al.
(NeurIPS 2023) MotionGPT : Human Motion as Foreign Language, Jiang et al.
(NeurIPS 2023) FineMoGen : Fine-Grained Spatio-Temporal Motion Generation and Editing, Zhang et al.
(NeurIPS 2023) InsActor : Instruction-driven Physics-based Characters, Ren et al.
(ICCV 2023) AttT2M : Text-Driven Human Motion Generation with Multi-Perspective Attention Mechanism, Zhong et al.
(ICCV 2023) TMR : Text-to-Motion Retrieval Using Contrastive 3D Human Motion Synthesis, Petrovich et al.
(ICCV 2023) MAA : Make-An-Animation: Large-Scale Text-conditional 3D Human Motion Generation, Azadi et al.
(ICCV 2023) PhysDiff : Physics-Guided Human Motion Diffusion Model, Yuan et al.
(ICCV 2023) ReMoDiffuse : Retrieval-Augmented Motion Diffusion Model, Zhang et al.
(ICCV 2023) BelFusion : Latent Diffusion for Behavior-Driven Human Motion Prediction, Barquero et al.
(ICCV 2023) GMD : Guided Motion Diffusion for Controllable Human Motion Synthesis, Karunratanakul et al.
(ICCV 2023) HMD-NeMo : Online 3D Avatar Motion Generation From Sparse Observations, Aliakbarian et al.
(ICCV 2023) SINC : Spatial Composition of 3D Human Motions for Simultaneous Action Generation, Athanasiou et al.
(ICCV 2023) Kong et al. : Priority-Centric Human Motion Generation in Discrete Latent Space, Kong et al.
(ICCV 2023) Fg-T2M : Fine-Grained Text-Driven Human Motion Generation via Diffusion Model, Wang et al.
(ICCV 2023) EMS : Breaking The Limits of Text-conditioned 3D Motion Synthesis with Elaborative Descriptions, Qian et al.
(SIGGRAPH 2023) GenMM : Example-based Motion Synthesis via Generative Motion Matching, Li et al.
(SIGGRAPH 2023) GestureDiffuCLIP : Gesture Diffusion Model with CLIP Latents, Ao et al.
(SIGGRAPH 2023) BodyFormer : Semantics-guided 3D Body Gesture Synthesis with Transformer, Pang et al.
(SIGGRAPH 2023) Alexanderson et al. : Listen, denoise, action! Audio-driven motion synthesis with diffusion models, Alexanderson et al.
(CVPR 2023) AGroL : Avatars Grow Legs: Generating Smooth Human Motion from Sparse Tracking Inputs with Diffusion Model, Du et al.
(CVPR 2023) TALKSHOW : Generating Holistic 3D Human Motion from Speech, Yi et al.
(CVPR 2023) T2M-GPT : Generating Human Motion from Textual Descriptions with Discrete Representations, Zhang et al.
(CVPR 2023) UDE : A Unified Driving Engine for Human Motion Generation, Zhou et al.
(CVPR 2023) OOHMG : Being Comes from Not-being: Open-vocabulary Text-to-Motion Generation with Wordless Training, Lin et al.
(CVPR 2023) EDGE : Editable Dance Generation From Music, Tseng et al.
(CVPR 2023) MLD : Executing your Commands via Motion Diffusion in Latent Space, Chen et al.
(CVPR 2023) MoDi : Unconditional Motion Synthesis from Diverse Data, Raab et al.
(CVPR 2023) MoFusion : A Framework for Denoising-Diffusion-based Motion Synthesis, Dabral et al.
(CVPR 2023) Mo et al. : Continuous Intermediate Token Learning with Implicit Motion Manifold for Keyframe Based Motion Interpolation, Mo et al.
(ICLR 2023) HMDM : Human Motion Diffusion Model, Tevet et al.
(TPAMI 2023) MotionDiffuse : Text-Driven Human Motion Generation with Diffusion Model, Zhang et al.
(TPAMI 2023) Bailando++ : 3D Dance GPT with Choreographic Memory, Li et al.
(ArXiv 2023) UDE-2 : A Unified Framework for Multimodal, Multi-Part Human Motion Synthesis, Zhou et al.
(ArXiv 2023) Motion Script : Natural Language Descriptions for Expressive 3D Human Motions, Yazdian et al.
2022 and earlier
(NeurIPS 2022) NeMF : Neural Motion Fields for Kinematic Animation, He et al.
(SIGGRAPH Asia 2022) PADL : Language-Directed Physics-Based Character, Juravsky et al.
(SIGGRAPH Asia 2022) Rhythmic Gesticulator : Rhythm-Aware Co-Speech Gesture Synthesis with Hierarchical Neural Embeddings, Ao et al.
(3DV 2022) TEACH : Temporal Action Composition for 3D Human, Athanasiou et al.
(ECCV 2022) Implicit Motion : Implicit Neural Representations for Variable Length Human Motion Generation, Cervantes et al.
(ECCV 2022) Zhong et al. : Learning Uncoupled-Modulation CVAE for 3D Action-Conditioned Human Motion Synthesis, Zhong et al.
(ECCV 2022) MotionCLIP : Exposing Human Motion Generation to CLIP Space, Tevet et al.
(ECCV 2022) PoseGPT : Quantizing human motion for large scale generative modeling, Lucas et al.
(ECCV 2022) TEMOS : Generating diverse human motions from textual descriptions, Petrovich et al.
(ECCV 2022) TM2T : Stochastic and Tokenized Modeling for the Reciprocal Generation of 3D Human Motions and Texts, Guo et al.
(SIGGRAPH 2022) AvatarCLIP : Zero-Shot Text-Driven Generation and Animation of 3D Avatars, Hong et al.
(SIGGRAPH 2022) DeepPhase : Periodic autoencoders for learning motion phase manifolds, Starke et al.
(CVPR 2022) Guo et al. : Generating Diverse and Natural 3D Human Motions from Text, Guo et al.
(CVPR 2022) Bailando : 3D Dance Generation by Actor-Critic GPT with Choreographic Memory, Li et al.
(ICCV 2021) ACTOR : Action-Conditioned 3D Human Motion Synthesis with Transformer VAE, Petrovich et al.
(ICCV 2021) AIST++ : AI Choreographer: Music Conditioned 3D Dance Generation with AIST++, Li et al.
(SIGGRAPH 2021) Starke et al. : Neural animation layering for synthesizing martial arts movements, Starke et al.
(CVPR 2021) MOJO : We are More than Our Joints: Predicting how 3D Bodies Move, Zhang et al.
(ECCV 2020) DLow : Diversifying Latent Flows for Diverse Human Motion Prediction, Yuan et al.
(SIGGRAPH 2020) Starke et al. : Local motion phases for learning multi-contact character movements, Starke et al.
Motion Editing
(ArXiv 2026) Skinned Motion Retargeting with Spatially Adaptive Interaction Guidance , Choi et al.
(ArXiv 2026) MotionMERGE : A Multi-granular Framework for Human Motion Editing, Reasoning, Generation, and Explanation, Wu et al.
(ArXiv 2026) ExpertEdit : Learning Skill-Aware Motion Editing from Expert Videos, Somayazulu et al.
(ArXiv 2026) InterEdit : Navigating Text-Guided Multi-Human 3D Motion Editing, Yang et al.
(IVA 2025) TF-JAX-IK : Real-Time Inverse Kinematics for Generating Multi-Constrained Movements of Virtual Human Characters, Voss et al.
(ICCV 2025) PRIMAL : Physically Reactive and Interactive Motor Model for Avatar Learning, Zhang et al.
(CVPR 2025) SALAD : Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing, Hong et al.
(CVPR 2025) MixerMDM : Learnable Composition of Human Motion Diffusion Models, Ruiz-Ponce et al.
(CVPR 2025) AnyMoLe : Any Character Motion In-Betweening Leveraging Video Diffusion Models, Yun et al.
(CVPR 2025) SimMotionEdit : Text-Based Human Motion Editing with Motion Similarity Prediction, Li et al.
(CVPR 2025) MotionReFit : Dynamic Motion Blending for Versatile Motion Editing, Jiang et al.
(ArXiv 2025) StableMotion : Training Motion Cleanup Models with Unpaired Corrupted Data, Mu et al.
(ArXiv 2025) Dai et al : Towards Synthesized and Editable Motion In-Betweening Through Part-Wise Phase Representation, Dai et al.
(SIGGRAPH Asia 2024) MotionFix : Text-Driven 3D Human Motion Editing, Athanasiou et al.
(NeurIPS 2024) CigTime : Corrective Instruction Generation Through Inverse Motion Editing, Fang et al.
(SIGGRAPH 2024) Iterative Motion Editing : Iterative Motion Editing with Natural Language, Goel et al.
(CVPR 2024) DNO : Optimizing Diffusion Noise Can Serve As Universal Motion Priors, Karunratanakul et al.
Motion Stylization
(ICCV 2025) StyleMotif : Multi-Modal Motion Stylization using Style-Content Cross Fusion, Guo et al.
(CVPR 2025) Visual Persona : Foundation Model for Full-Body Human Customization, Nam et al.
(ArXiv 2025) ClusterStyle : Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation, Chen et al.
(ArXiv 2025) MotionPersona : Characteristics-aware Locomotion Control, Shi et al.
(ArXiv 2025) AStF : Motion Style Transfer via Adaptive Statistics Fusor, Chen et al.
(ArXiv 2025) Dance Like a Chicken : Low-Rank Stylization for Human Motion Diffusion, Sawdayee et al.
(ArXiv 2024) MulSMo : Multimodal Stylized Motion Generation by Bidirectional Control Flow, Li et al.
(TSMC 2024) D-LORD : D-LORD for Motion Stylization, Gupta et al.
(ECCV 2024) HUMOS : Human Motion Model Conditioned on Body Shape, Tripathi et al.
(SIGGRAPH 2024) SMEAR : Stylized Motion Exaggeration with ARt-direction, Basset et al.
(SIGGRAPH 2024) Portrait3D : Text-Guided High-Quality 3D Portrait Generation Using Pyramid Representation and GANs Prior, Wu et al.
(CVPR 2024) MCM-LDM : Arbitrary Motion Style Transfer with Multi-condition Motion Latent Diffusion Model, Song et al.
(CVPR 2024) MoST : Motion Style Transformer between Diverse Action Contents, Kim et al.
(ICLR 2024) GenMoStyle : Generative Human Motion Stylization in Latent Space, Guo et al.
Human-Object Interaction
2026
(CVPR 2026) ViHOI : Human-Object Interaction Synthesis with Visual Priors, Cai et al.
(CVPR 2026) InterPrior : Scaling Generative Control for Physics-Based Human-Object Interactions, Xu et al.
(CVPR 2026) TeamHOI : Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size, Lionar et al.
(ArXiv 2026) MaMi-HOI : Harmonizing Global Kinematics and Local Geometry for Human-Object Interaction Generation, Wang et al.
(ArXiv 2026) Hoi3DGen : Generating High-Quality Human-Object-Interactions in 3D, Sharma et al.
(ArXiv 2026) InterReal : A Unified Physics-Based Imitation Framework for Learning Human-Object Interaction Skills, Liang et al.
2025
(NeurIPS 2025) HHOI : Learning to Generate Human-Human-Object Interactions from Textual Descriptions, Na et al.
(ACM MM 2025) PA-HOI : A Physics-Aware Human and Object Interaction Dataset, Wang et al.
(ACM MM 2025) OnlineHOI : Towards Online Human-Object Interaction Generation and Perception, Ji et al.
(ICCV 2025) Perceiving and Acting in First-Person : A Dataset and Benchmark for Egocentric Human-Object-Human Interactions, Xu et al.
(ICCV 2025) TriDi : Trilateral Diffusion of 3D Humans, Objects and Interactions, Petrov et al.
(ICCV 2025) SMGDiff : Soccer Motion Generation using diffusion probabilistic models, Yang et al.
(ICCV 2025) SyncDiff : Synchronized Motion Diffusion for Multi-Body Human-Object Interaction Synthesis, He et al.
(ICCV 2025) Wu et al : Human-Object Interaction from Human-Level Instructions, Wu et al.
(ICCV 2025) HUMOTO : A 4D Dataset of Mocap Human Object Interactions, Lu et al.
(SIGGRAPH 2025) PhysicsFC : Learning User-Controlled Skills for a Physics-Based Football Player Controller, Kim et al.
(SIGGRAPH 2025) SkillMimic-v2 : Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations, Yu et al.
(Bioengineering 2025) MeLLO : The Utah Manipulation and Locomotion of Large Objects (MeLLO) Data Library, Luttmer et al.
(CVPR 2025) ChainHOI : Joint-based Kinematic Chain Modeling for Human-Object Interaction Generation, Zeng et al.
(CVPR 2025) HOIGPT : Learning Long Sequence Hand-Object Interaction with Language Models, Huang et al.
(CVPR 2025) Hui et al : An Image-like Diffusion Method for Human-Object Interaction Detection, Hui et al.
(CVPR 2025) PersonaHOI : Effortlessly Improving Personalized Face with Human-Object Interaction Generation, Hu et al.
(CVPR 2025) InteractVLM : 3D Interaction Reasoning from 2D Foundational Models, Dwivedi et al.
(CVPR 2025) PICO : Reconstructing 3D People In Contact with Objects, Cseke et al.
(CVPR 2025) EasyHOI : Unleashing the Power of Large Models for Reconstructing Hand-Object Interactions in the Wild, Liu et al.
(CVPR 2025) FIction : 4D Future Interaction Prediction from Video, Ashutosh et al.
(CVPR 2025) ROG : Guiding Human-Object Interactions with Rich Geometry and Relations, Xue et al.
(CVPR 2025) SemGeoMo : Dynamic Contextual Human Motion Generation with Semantic and Geometric Guidance, Cong et al.
(CVPR 2025) Phys-Reach-Grasp : Learning Physics-Based Full-Body Human Reaching and Grasping from Brief Walking References, Li et al.
(CVPR 2025) ParaHome : Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions, Kim et al.
(CVPR 2025) InterMimic : Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions, Xu et al.
(CVPR 2025) CORE4D : A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement, Zhang et al.
(CVPR 2025) InteractAnything : Zero-shot Human Object-Interaction Synthesis via LLM Feedback and Object Affordance Parsing, Zhang et al.
(CVPR 2025) SkillMimic : Learning Reusable Basketball Skills from Demonstrations, Wang et al.
(CVPR 2025) MobileH2R : Learning Generalizable Human to Mobile Robot Handover Exclusively from Scalable and Diverse Synthetic Data, Wang et al.
(AAAI 2025) ARDHOI : Auto-Regressive Diffusion for Generating 3D Human-Object Interactions, Geng et al.
(AAAI 2025) DiffGrasp : Whole-Body Grasping Synthesis Guided by Object Motion Using a Diffusion Model, Zhang et al.
(3DV 2025) Paschalidis et al : 3D Whole-body Grasp Synthesis with Directional Controllability, Paschalidis et al.
(3DV 2025) InterTrack : Tracking Human Object Interaction without Object Templates, Xie et al.
(3DV 2025) FORCE : Dataset and Method for Intuitive Physics Guided Human-object Interaction, Zhang et al.
(PAMI 2025) MotionVerse : A Unified Multimodal Framework for Motion Comprehension, Generation and Editing, Hou et al.
(PAMI 2025) EigenActor : Variant Body-Object Interaction Generation Evolved from Invariant Action Basis Reasoning, Guo et al.
(ArXiv 2025) InteractMove : Text-Controlled Human-Object Interaction Generation in 3D Scenes with Movable Objects, Cai et al.
(ArXiv 2025) InterPose : Learning to Generate Human-Object Interactions from Large-Scale Web Videos, Zhang et al.
(ArXiv 2025) ECHO : Ego-Centric modeling of Human-Object interactions, Petrov et al.
(ArXiv 2025) CoopDiff : Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion, Lin et al.
(ArXiv 2025) HOI-Dyn : Learning Interaction Dynamics for Human-Object Motion Diffusion, Wu et al.
(ArXiv 2025) HOIDiNi : Human-Object Interaction through Diffusion Noise Optimization, Ron et al.
(ArXiv 2025) GenHOI : Generalizing Text-driven 4D Human-Object Interaction Synthesis for Unseen Objects, Li et al.
(ArXiv 2025) HOI-PAGE : Zero-Shot Human-Object Interaction Generation with Part Affordance Guidance, Li et al.
(ArXiv 2025) HOSIG : Full-Body Human-Object-Scene Interaction Generation, Yao et al.
(ArXiv 2025) CoDA : Coordinated Diffusion Noise Optimization for Whole-Body Manipulation of Articulated Objects, Pi et al.
(ArXiv 2025) MaskedManipulator : Versatile Whole-Body Control for Loco-Manipulation, Tessler et al.
(ArXiv 2025) UniHM : Universal Human Motion Generation with Object Interactions in Indoor Scenes, Geng et al.
(ArXiv 2025) EJIM : Efficient Explicit Joint-level Interaction Modeling with Mamba for Text-guided HOI Generation, Huang et al.
(ArXiv 2025) ZeroHOI : Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors, Lou et al.
(ArXiv 2025) RMD-HOI : Human-Object Interaction with Vision-Language Model Guided Relative Movement Dynamics, Deng et al.
(ArXiv 2025) Kaiwu : A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction, Jiang et al.
2024
(ArXiv 2024) CHOICE : Coordinated Human-Object Interaction in Cluttered Environments for Pick-and-Place Actions, Lu et al.
(ArXiv 2024) OOD-HOI : Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains, Zhang et al.
(ArXiv 2024) COLLAGE : Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models, Daiya et al.
(NeurIPS 2024) HumanVLA : Towards Vision-Language Directed Object Rearrangement by Physical Humanoid, Xu et al.
(NeurIPS 2024) OmniGrasp : Grasping Diverse Objects with Simulated Humanoids, Luo et al.
(NeurIPS 2024) EgoChoir : Capturing 3D Human-Object Interaction Regions from Egocentric Views, Yang et al.
(NeurIPS 2024) CooHOI : Learning Cooperative Human-Object Interaction with Manipulated Object Dynamics, Gao et al.
(NeurIPS 2024) InterDreamer : Zero-Shot Text to 3D Dynamic Human-Object Interaction, Xu et al.
(NeurIPS 2024) PiMForce : Posture-Informed Muscular Force Learning for Robust Hand Pressure Estimation, Seo et al.
(ECCV 2024) InterFusion : Text-Driven Generation of 3D Human-Object Interaction, Dai et al.
(ECCV 2024) CHOIS : Controllable Human-Object Interaction Synthesis, Li et al.
(ECCV 2024) F-HOI : Toward Fine-grained Semantic-Aligned 3D Human-Object Interactions, Yang et al.
(ECCV 2024) HIMO : A New Benchmark for Full-Body Human Interacting with Multiple Objects, Lv et al.
(SIGGRAPH 2024) PhysicsPingPong : Strategy and Skill Learning for Physics-based Table Tennis Animation, Wang et al.
(CVPR 2024) NIFTY : Neural Object Interaction Fields for Guided Human Motion Synthesis, Kulkarni et al.
(CVPR 2024) HOI Animator : Generating Text-Prompt Human-Object Animations using Novel Perceptive Diffusion Models, Son et al.
(CVPR 2024) CG-HOI : Contact-Guided 3D Human-Object Interaction Generation, Diller et al.
(IJCV 2024) InterCap : Joint Markerless 3D Tracking of Humans and Objects in Interaction, Huang et al.
(3DV 2024) Phys-Fullbody-Grasp : Physically Plausible Full-Body Hand-Object Interaction Synthesis, Braun et al.
(3DV 2024) GRIP : Generating Interaction Poses Using Spatial Cues and Latent Consistency, Taheri et al.
(AAAI 2024) FAVOR : Full-Body AR-driven Virtual Object Rearrangement Guided by Instruction Text, Li et al.
2023 and earlier
(SIGGRAPH Asia 2023) OMOMO : Object Motion Guided Human Motion Synthesis, Li et al.
(ICCV 2023) CHAIRS : Full-Body Articulated Human-Object Interaction, Jiang et al.
(ICCV 2023) HGHOI : Hierarchical Generation of Human-Object Interactions with Diffusion Probabilistic Models, Pi et al.
(ICCV 2023) InterDiff : Generating 3D Human-Object Interactions with Physics-Informed Diffusion, Xu et al.
(CVPR 2023) Object Pop Up : Can we infer 3D objects and their poses from human interactions alone? Petrov et al.
(CVPR 2023) ARCTIC : A Dataset for Dexterous Bimanual Hand-Object Manipulation, Fan et al.
(ECCV 2022) TOCH : Spatio-Temporal Object-to-Hand Correspondence for Motion Refinement, Zhou et al.
(ECCV 2022) COUCH : Towards Controllable Human-Chair Interactions, Zhang et al.
(ECCV 2022) SAGA : Stochastic Whole-Body Grasping with Contact, Wu et al.
(CVPR 2022) GOAL : Generating 4D Whole-Body Motion for Hand-Object Grasping, Taheri et al.
(CVPR 2022) BEHAVE : Dataset and Method for Tracking Human Object Interactions, Bhatnagar et al.
(ECCV 2020) GRAB : A Dataset of Whole-Body Human Grasping of Objects, Taheri et al.
Human-Scene Interaction
2026
(ICLR 2026) InfBaGel : Human-Object-Scene Interaction Generation with Dynamic Perception and Iterative Refinement, Zou et al.
(ArXiv 2026) ArtHOI : Articulated Human-Object Interaction Synthesis by 4D Reconstruction from Video Priors, Huang et al.
(ArXiv 2026) SceMoS : Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens, Ghosh et al.
(ArXiv 2026) Dynamic Worlds, Dynamic Humans : Generating Virtual Human-Scene Interaction Motion in Dynamic Scenes, Wang et al.
2025
(ICCV 2025) Being-M0.5 : A Real-Time Controllable Vision-Language-Motion Model, Cao et al.
(ICCV 2025) SceneMI : Motion In-Betweening for Modeling Human-Scene Interactions, Hwang et al.
(ICCV 2025) SIMS : Simulating Human-Scene Interactions with Real World Script Planning, Wang et al.
(ICCV 2025) Lim et al : Event-Driven Storytelling with Multiple Lifelike Humans in a 3D scene, Lim et al.
(ICME 2025) TSTMotion : Training-free Scene-aware Text-to-motion Generation, Guo et al.
(CVPR 2025) HSI-GPT : A General-Purpose Large Scene-Motion-Language Model for Human Scene Interaction. Wang et al.
(CVPR 2025) Vision-Guided Action : Enhancing 3D Human Motion Prediction with Gaze-informed Affordance in 3D Scenes. Yu et al.
(CVPR 2025) Yi et al : Estimating Body and Hand Motion in an Ego‑sensed World, Yi et al.
(CVPR 2025) EnvPoser : Environment-aware Realistic Human Motion Estimation from Sparse Observations with Uncertainty Modeling. Xia et al.
(CVPR 2025) TokenHSI : Unified Synthesis of Physical Human-Scene Interactions through Task Tokenization, Pan et al.
(ICLR 2025) Sitcom-Crafter : A Plot-Driven Human Motion Generation System in 3D Scenes, Chen et al.
(3DV 2025) Paschalidis et al : 3D Whole-body Grasp Synthesis with Directional Controllability, Paschalidis et al.
(WACV 2025) GHOST : Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts, Milacski et al.
(ArXiv 2025) Prime and Reach : Synthesising Body Motion for Gaze-Primed Object Reach, Hatano et al.
(ArXiv 2025) Uni-Inter : Unifying 3D Human Motion Synthesis Across Diverse Interaction Contexts, Liu et al.
(ArXiv 2025) SSOMotion : HumanMotion Synthesis in 3D Scenes via Unified Scene Semantic Occupancy, Cho et al.
(ArXiv 2025) SceneAdapt : Scene-aware Adaptation of Human Motion Diffusion, Cho et al.
(ArXiv 2025) FantasyHSI : Video-Generation-Centric 4D Human Synthesis In Any Scene through A Graph-based Multi-Agent Framework, Mu et al.
(ArXiv 2025) Half-Physics : Enabling Kinematic 3D Human Model with Physical Interactions, Li et al.
(ArXiv 2025) GenHSI : Controllable Generation of Human-Scene Interaction Videos, Li et al.
(ArXiv 2025) RMD-HOI : Human-Object Interaction with Vision-Language Model Guided Relative Movement Dynamics, Deng et al.
(ArXiv 2025) HIS-GPT : Towards 3D Human-In-Scene Multimodal Understanding, Zhao et al.
(ArXiv 2025) Jointly Understand Your Command and Intention : Reciprocal Co-Evolution between Scene-Aware 3D Human Motion Synthesis and Analysis, Gao et al.
2024
(ArXiv 2024) ZeroHSI : Zero-Shot 4D Human-Scene Interaction by Video Generation, Li et al.
(ArXiv 2024) Mimicking-Bench : A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking, Liu et al.
(ArXiv 2024) SCENIC : Scene-aware Semantic Navigation with Instruction-guided Control, Zhang et al.
(ArXiv 2024) Diffusion Implicit Policy : Diffusion Implicit Policy for Unpaired Scene-aware Motion synthesis, Gong et al.
(ArXiv 2024) LaserHuman : Language-guided Scene-aware Human Motion Generation in Free Environment, Cong et al.
(SIGGRAPH Asia 2024) LINGO : Autonomous Character-Scene Interaction Synthesis from Text Instruction, Jiang et al.
(NeurIPS 2024) DiMoP3D : Harmonizing Stochasticity and Determinism: Scene-responsive Diverse Human Motion Prediction, Lou et al.
(ECCV 2024) MOB : Revisit Human-Scene Interaction via Space Occupancy, Liu et al.
(ECCV 2024) TesMo : Generating Human Interaction Motions in Scenes with Text Control, Yi et al.
(ECCV 2024 Workshop) SAST : Massively Multi-Person 3D Human Motion Forecasting with Scene Context, Mueller et al.
(Eurographics 2024) Kang et al : Learning Climbing Controllers for Physics-Based Characters, Kang et al.
(CVPR 2024) Afford-Motion : Move as You Say, Interact as You Can: Language-guided Human Motion Generation with Scene Affordance, Wang et al.
(CVPR 2024) GenZI : Zero-Shot 3D Human-Scene Interaction Generation, Li et al.
(CVPR 2024) Cen et al. : Generating Human Motion in 3D Scenes from Text Descriptions, Cen et al.
(CVPR 2024) TRUMANS : Scaling Up Dynamic Human-Scene Interaction Modeling, Jiang et al.
(ICLR 2024) UniHSI : Unified Human-Scene Interaction via Prompted Chain-of-Contacts, Xiao et al.
(3DV 2024) Purposer : Putting Human Motion Generation in Context, Ugrinovic et al.
(3DV 2024) InterScene : Synthesizing Physically Plausible Human Motions in 3D Scenes, Pan et al.
(3DV 2024) Mir et al : Generating Continual Human Motion in Diverse 3D Scenes, Mir et al.
2023 and earlier
(ICCV 2023) DIMOS : Synthesizing Diverse Human Motions in 3D Indoor Scenes, Zhao et al.
(ICCV 2023) LAMA : Locomotion-Action-Manipulation: Synthesizing Human-Scene Interactions in Complex 3D Environments, Lee et al.
(ICCV 2023) Narrator : Towards Natural Control of Human-Scene Interaction Generation via Relationship Reasoning, Xuan et al.
(CVPR 2023) CIMI4D : A Large Multimodal Climbing Motion Dataset under Human-Scene Interactions, Yan et al.
(CVPR 2023) Scene-Ego : Scene-aware Egocentric 3D Human Pose Estimation, Wang et al.
(CVPR 2023) SLOPER4D : A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments, Dai et al.
(CVPR 2023) CIRCLE : Capture in Rich Contextual Environments, Araujo et al.
(CVPR 2023) SceneDiffuser : Diffusion-based Generation, Optimization, and Planning in 3D Scenes, Huang et al.
(CVPR 2023) MIME : Human-Aware 3D Scene Generation, Yi et al.
(SIGGRAPH 2023) PMP : Learning to Physically Interact with Environments using Part-wise Motion Priors, Bae et al.
(SIGGRAPH 2023) QuestEnvSim : Environment-Aware Simulated Motion Tracking from Sparse Sensors, Lee et al.
(SIGGRAPH 2023) Hassan et al. : Synthesizing Physical Character-Scene Interactions, Hassan et al.
(NeurIPS 2022) Mao et al. : Contact-Aware Human Motion Forecasting, Mao et al.
(NeurIPS 2022) HUMANISE : Language-conditioned Human Motion Generation in 3D Scenes, Wang et al.
(NeurIPS 2022) EmbodiedPose : Embodied Scene-aware Human Pose Estimation, Luo et al.
(ECCV 2022) GIMO : Gaze-Informed Human Motion Prediction in Context, Zheng et al.
(ECCV 2022) COINS : Compositional Human-Scene Interaction Synthesis with Semantic Control, Zhao et al.
(CVPR 2022) Wang et al. : Towards Diverse and Natural Scene-aware 3D Human Motion Synthesis, Wang et al.
(CVPR 2022) GAMMA : The Wanderings of Odysseus in 3D Scenes, Zhang et al.
(ICCV 2021) SAMP : Stochastic Scene-Aware Motion Prediction, Hassan et al.
(ICCV 2021) LEMO : Learning Motion Priors for 4D Human Body Capture in 3D Scenes, Zhang et al.
(3DV 2020) PLACE : Proximity Learning of Articulation and Contact in 3D Environments, Zhang et al.
(SIGGRAPH 2020) Starke et al. : Local motion phases for learning multi-contact character movements, Starke et al.
(CVPR 2020) PSI : Generating 3D People in Scenes without People, Zhang et al.
(SIGGRAPH Asia 2019) NSM : Neural State Machine for Character-Scene Interactions, Starke et al.
(ICCV 2019) PROX : Resolving 3D Human Pose Ambiguities with 3D Scene Constraints, Hassan et al.
Human-Human Interaction
(ArXiv 2026) Contact Matrix : Enhancing Dance Motion Synthesis with Precise Interaction Modeling, Chen et al.
(ArXiv 2026) Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation , Xu et al.
(CVPR 2026) ReMoGen : Real-time Human Interaction-to-Reaction Generation via Modular Learning from Diverse Data, Ye et al.
(CVPR 2026) Learning to Assist : Physics-Grounded Human-Human Control via Multi-Agent Reinforcement Learning, Shibata et al.
(ArXiv 2026) Rhythm : Learning Interactive Whole-Body Control for Dual Humanoids, Chen et al.
(ArXiv 2026) HINT : Hierarchical Interaction Modeling for Autoregressive Multi-Human Motion Generation, Liu et al.
(AAAI 2026) InterMoE : Individual-Specific 3D Human Interaction Generation via Dynamic Temporal-Selective MoE, Wang et al.
(ICCV 2025) Ponimator : Unfolding Interactive Pose for Versatile Human-human Interaction Animation, Liu et al.
(ICCV 2025) Perceiving and Acting in First-Person : A Dataset and Benchmark for Egocentric Human-Object-Human Interactions, Xu et al.
(ICCV 2025) Towards Immersive Human-X Interaction : A Real-Time Framework for Physically Plausible Motion Synthesis, Ji et al.
(ICCV 2025) PINO : Person-Interaction Noise Optimization for Long-Duration and Customizable Motion Generation of Arbitrary-Sized Groups, Ota et al.
(SIGGRAPH 2025) Xu et al : Multi-Person Interaction Generation from Two-Person Motion Priors, Xu et al.
(SIGGRAPH 2025) DuetGen : Music Driven Two-Person Dance Generation via Hierarchical Masked Modeling, Ghosh et al.
(CVPR 2025) TIMotion : Temporal and Interactive Framework for Efficient Human-Human Motion Generation, Wang et al.
(ICLR 2025) Think Then React : Towards Unconstrained Action-to-Reaction Motion Generation, Tan et al.
(ICLR 2025) Ready-to-React : Online Reaction Policy for Two-Character Interaction Generation, Cen et al.
(ICLR 2025) InterMask : 3D Human Interaction Generation via Collaborative Masked Modelling, Javed et al.
(3DV 2025) Interactive Humanoid : Online Full-Body Motion Reaction Synthesis with Social Affordance Canonicalization and Forecasting, Liu et al.
(ArXiv 2025) Interact2Ar : Full-Body Human-Human Interaction Generation via Autoregressive Diffusion Models, Ruiz-Ponce et al.
(ArXiv 2025) Text2Interact : High-Fidelity and Diverse Text-to-Two-Person Interaction Generation, Wu et al.
(ArXiv 2025) InterAct : A Large-Scale Dataset of Dynamic, Expressive and Interactive Activities between Two People in Daily Scenarios, Ho et al.
(ArXiv 2025) E-React : Towards Emotionally Controlled Synthesis of Human Reactions, Zhu et al.
(ArXiv 2025) Seamless Interaction : Dyadic Audiovisual Motion Modeling and Large-Scale Dataset, Agrawal et al.
(ArXiv 2025) MAMMA : Markerless & Automatic Multi-Person Motion Action Capture, Cuevas-Velasquez et al.
(ArXiv 2025) PhysInter : Integrating Physical Mapping for High-Fidelity Human Interaction Generation, Yao et al.
(ArXiv 2025) InterMamba : Efficient Human-Human Interaction Generation with Adaptive Spatio-Temporal Mamba, Wu et al.
(ArXiv 2025) MARRS : MaskedAutoregressive Unit-based Reaction Synthesis, Wang et al.
(ArXiv 2025) SocialGen : Modeling Multi-Human Social Interaction with Language Models, Yu et al.
(ArXiv 2025) ARFlow : Human Action-Reaction Flow Matching with Physical Guidance, Jiang et al.
(ArXiv 2025) Fan et al : 3D Human Interaction Generation: A Survey, Fan et al.
(ArXiv 2025) Invisible Strings : Revealing Latent Dancer-to-Dancer Interactions with Graph Neural Networks, Zerkowski et al.
(ArXiv 2025) Leader and Follower : Interactive Motion Generation under Trajectory Constraints, Wang et al.
(ArXiv 2024) Two in One : Unified Multi-Person Interactive Motion Generation by Latent Diffusion Transformer, Li et al.
(ArXiv 2024) It Takes Two : Real-time Co-Speech Two-person’s Interaction Generation via Reactive Auto-regressive Diffusion Model, Shi et al.
(ArXiv 2024) COLLAGE : Collaborative Human-Agent Interaction Generation using Hierarchical Latent Diffusion and Language Models, Daiya et al.
(NeurIPS 2024) InterControl : Generate Human Motion Interactions by Controlling Every Joint, Wang et al.
(ACM MM 2024) PhysReaction : Physically Plausible Real-Time Humanoid Reaction Synthesis via Forward Dynamics Guided 4D Imitation, Liu et al.
(ECCV 2024) Shan et al : Towards Open Domain Text-Driven Synthesis of Multi-Person Motions, Shan et al.
(ECCV 2024) ReMoS : 3D Motion-Conditioned Reaction Synthesis for Two-Person Interactions, Ghosh et al.
(CVPR 2024) Inter-X : Towards Versatile Human-Human Interaction Analysis, Xu et al.
(CVPR 2024) ReGenNet : Towards Human Action-Reaction Synthesis, Xu et al.
(CVPR Workshop 2024) in2IN : in2IN: Leveraging Individual Information to Generate Human INteractions, Ruiz-Ponce et al.
(IJCV 2024) InterGen : Diffusion-based Multi-human Motion Generation under Complex Interactions, Liang et al.
(ICCV 2023) ActFormer : A GAN-based Transformer towards General Action-Conditioned 3D Human Motion Generation, Xu et al.
(ICCV 2023) Tanaka et al. : Role-aware Interaction Generation from Textual Description, Tanaka et al.
(CVPR 2023) Hi4D : 4D Instance Segmentation of Close Human Interaction, Yin et al.
(CVPR 2022) ExPI : Multi-Person Extreme Motion Prediction, Guo et al.
(CVPR 2020) CHI3D : Three-Dimensional Reconstruction of Human Interactions, Fieraru et al.
Datasets & Benchmarks
2026
(ArXiv 2026) HumanCLAW : Can Vision-Language Models Act Through a Body?, Li et al.
(ArXiv 2026) EgoHTR : Egocentric 4D Demonstrations of Human Terrain Traversal, Brandes et al.
(ArXiv 2026) Humanoid-OmniOcc : Stereo-Based Full-View Occupancy Dataset for Embodied AI, Guo et al.
(ArXiv 2026) Data Standards for Humanoid Robotics : The Missing Infrastructure for Physical AI, Liu et al.
(ArXiv 2026) SIMPLE : Simulation-Based Policy Learning and Evaluation for Humanoid Loco-manipulation, Wei et al.
(CVPR 2026) RoMo : A Large-Scale, Richly Organized Dataset and Semantic Taxonomy for Human Motion Generation, Zhang et al.
(ArXiv 2026) HumanScore : Benchmarking Human Motions in Generated Videos, Fang et al.
(ArXiv 2026) A Dataset and Evaluation for Complex 4D Markerless Human Motion Capture , Park et al.
(ArXiv 2026) Towards Motion Turing Test : Evaluating Human-Likeness in Humanoid Robots, Li et al.
(ArXiv 2026) Moving Through Clutter : Scaling Data Collection and Benchmarking for 3D Scene-Aware Humanoid Locomotion via Virtual Reality, Wang et al.
2025
(ICCV 2025) PP-Motion : Physical-Perceptual Fidelity Evaluation for Human Motion Generation, Zhao et al.
(ICCV 2025) MDD : A Dataset for Text-and-Music Conditioned Duet Dance Generation, Gupta et al.
(ACM MM 2025) Perceiving and Acting in First-Person : A Dataset and Benchmark for Egocentric Human-Object-Human Interactions, Xu et al.
(Bioengineering 2025) MeLLO : The Utah Manipulation and Locomotion of Large Objects (MeLLO) Data Library, Luttmer et al.
(CVPR 2025) OpenHumanVid : A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation, Xu et al.
(CVPR 2025) InterAct : Advancing Large-Scale Versatile 3D Human-Object Interaction Generation, Xu et al.
(CVPR 2025) MotionPro : Exploring the Role of Pressure in Human MoCap and Beyond, Ren et al.
(CVPR 2025) GORP : Real-Time Motion Generation with Rolling Prediction Models, Barquero et al.
(CVPR 2025) ClimbingCap : ClimbingCap: Multi-Modal Dataset and Method for Rock Climbing in World Coordinate, Yan et al.
(CVPR 2025) AtoM : AToM: Aligning Text-to-Motion Model at Event-Level with GPT-4Vision Reward, Han et al.
(CVPR 2025) CORE4D : CORE4D: A 4D Human-Object-Human Interaction Dataset for Collaborative Object REarrangement, Zhang et al.
(ICLR 2025) MotionCritic : Aligning Human Motion Generation with Human Perceptions, Wang et al.
(ICLR 2025) LocoVR : LocoVR: Multiuser Indoor Locomotion Dataset in Virtual Reality, Takeyama et al.
(ICLR 2025) PMR : Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities, Wang et al.
(AAAI 2025) EMHI : A Multimodal Egocentric Human Motion Dataset with HMD and Body-Worn IMUs, Fan et al.
(ArXiv 2025) H2IAD : 3D Human-Human Interaction Anomaly Detection, Maeda et al.
(ArXiv 2025) RoleMotion : A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions, Peng et al.
(ArXiv 2025) Embody 3D : A Large-scale Multimodal Motion and Behavior Dataset, McLean et al.
(ArXiv 2025) CaddieSet : A Golf Swing Dataset with Human Joint Features and Ball Information, Jung et al.
(ArXiv 2025) Waymo-3DSkelMo : A Multi-Agent 3D Skeletal Motion Dataset for Pedestrian Interaction Modeling in Autonomous Driving, Zhu et al.
(ArXiv 2025) SpeakerVid-5M : A Large-Scale High-Quality Dataset for audio-visual Dyadic Interactive Human Generation, Zhang et al.
(ArXiv 2025) AthleticsPose : Authentic Sports Motion Dataset on Athletic Field and Evaluation of Monocular 3D Pose Estimation Ability, Suzuki et al.
(ArXiv 2025) MMHU : A Massive-Scale Multimodal Benchmark for Human Behavior Understanding, Li et al.
(ArXiv 2025) FLEX : A Large-Scale Multi-Modal Multi-Action Dataset for Fitness Action Quality Assessment, Yin et al.
(ArXiv 2025) From Motion to Behavior : Hierarchical Modeling of Humanoid Generative Behavior Control, Zhang et al.
(ArXiv 2025) Rekik et al : Quality assessment of 3D human animation: Subjective and objective evaluation, Rekik et al.
(ArXiv 2025) K2MUSE : A Large-scale Human Lower limb Dataset of Kinematics, Kinetics, amplitude Mode Ultrasound and Surface Electromyography, Li et al.
(ArXiv 2025) RMD-HOI : Human-Object Interaction with Vision-Language Model Guided Relative Movement Dynamics, Deng et al.
(ArXiv 2025) SGA-INTERACT : SGA-INTERACT: A3DSkeleton-based Benchmark for Group Activity Understanding in Modern Basketball Tactic, Yang et al.
(ArXiv 2025) Kaiwu : Kaiwu: A Multimodal Manipulation Dataset and Framework for Robot Learning and Human-Robot Interaction, Jiang et al.
(ArXiv 2025) Motion-X++ : Motion-X++: A Large-Scale Multimodal 3D Whole-body Human Motion Dataset, Zhang et al.
2024
(ArXiv 2024) Mimicking-Bench : A Benchmark for Generalizable Humanoid-Scene Interaction Learning via Human Mimicking, Liu et al.
(ArXiv 2024) LaserHuman : Language-guided Scene-aware Human Motion Generation in Free Environment, Cong et al.
(ArXiv 2024) SCENIC : Scene-aware Semantic Navigation with Instruction-guided Control, Zhang et al.
(ArXiv 2024) synNsync : Synergy and Synchrony in Couple Dances, Manukele et al.
(ArXiv 2024) MotionBank : A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations, Xu et al.
(Github 2024) CMP & CMR : AnimationGPT: An AIGC tool for generating game combat motion assets, Liao et al.
(Scientific Data 2024) Evans et al : Synchronized Video, Motion Capture and Force Plate Dataset for Validating Markerless Human Movement Analysis, Evans et al.
(Scientific Data 2024) MultiSenseBadminton : MultiSenseBadminton: Wearable Sensor–Based Biomechanical Dataset for Evaluation of Badminton Performance, Seong et al.
(SIGGRAPH Asia 2024) LINGO : Autonomous Character-Scene Interaction Synthesis from Text Instruction, Jiang et al.
(NeurIPS 2024) Harmony4D : Harmony4D: A Video Dataset for In-The-Wild Close Human Interactions, Khirodkar et al.
(NeurIPS D&B 2024) EgoSim : EgoSim: An Egocentric Multi-view Simulator for Body-worn Cameras during Human Motion, Hollidt et al.
(NeurIPS D&B 2024) Muscles in Time : Muscles in Time: Learning to Understand Human Motion by Simulating Muscle Activations, Schneider et al.
(NeurIPS D&B 2024) Text to blind motion : Text to blind motion, Kim et al.
(ACM MM 2024) CLaM : CLaM: An Open-Source Library for Performance Evaluation of Text-driven Human Motion Generation, Chen et al.
(ECCV 2024) AddBiomechanics : AddBiomechanics Dataset: Capturing the Physics of Human Motion at Scale, Werling et al.
(ECCV 2024) LiveHPS++ : Robust and Coherent Motion Capture in Dynamic Free Environment, Ren et al.
(ECCV 2024) SignAvatars : A Large-scale 3D Sign Language Holistic Motion Dataset and Benchmark, Yu et al.
(ECCV 2024) Nymeria : A massive collection of multimodal egocentric daily motion in the wild, Ma et al.
(Multibody System Dynamics 2024) Human3.6M+ : Using musculoskeletal models to generate physically-consistent data for 3D human pose, kinematic, dynamic, and muscle estimation, Nasr et al.
(CVPR 2024) Inter-X : Towards Versatile Human-Human Interaction Analysis, Xu et al.
(CVPR 2024) HardMo : A Large-Scale Hardcase Dataset for Motion Capture, Liao et al.
(CVPR 2024) Xie et al : Template Free Reconstruction of Human-object Interaction with Procedural Interaction Generation, Xie et al.
(CVPR 2024) MMVP : MMVP: A Multimodal MoCap Dataset with Vision and Pressure Sensors, Zhang et al.
(CVPR 2024) RELI11D : RELI11D: A Comprehensive Multimodal Human Motion Dataset and Method, Yan et al.
2023 and earlier
(SIGGRAPH Asia 2023) GroundLink : A Dataset Unifying Human Body Movement and Ground Reaction Dynamics, Han et al.
(NeurIPS D&B 2023) HOH : Markerless Multimodal Human-Object-Human Handover Dataset with Large Object Count, Wiederhold et al.
(NeurIPS D&B 2023) Motion-X : A Large-scale 3D Expressive Whole-body Human Motion Dataset, Lin et al.
(NeurIPS D&B 2023) Humans in Kitchens : A Dataset for Multi-Person Human Motion Forecasting with Scene Context, Tanke et al.
(ICCV 2023) CHAIRS : Full-Body Articulated Human-Object Interaction, Jiang et al.
(ICCV 2023) EMDB : The Electromagnetic Database of Global 3D Human Pose and Shape in the Wild, Kaufmann et al.
(CVPR 2023) MOYO : 3D Human Pose Estimation via Intuitive Physics, Tripathi et al.
(CVPR 2023) CIMI4D : A Large Multimodal Climbing Motion Dataset under Human-Scene Interactions, Yan et al.
(CVPR 2023) FLAG3D : A 3D Fitness Activity Dataset with Language Instruction, Tang et al.
(CVPR 2023) Hi4D : 4D Instance Segmentation of Close Human Interaction, Yin et al.
(CVPR 2023) CIRCLE : Capture in Rich Contextual Environments, Araujo et al.
(CVPR 2023) BEDLAM : A Synthetic Dataset of Bodies Exhibiting Detailed Lifelike Animated Motion, Black et al.
(CVPR 2023) SLOPER4D : A Scene-Aware Dataset for Global 4D Human Pose Estimation in Urban Environments, Dai et al.
(CVPR 2023) MIME : Human-Aware 3D Scene Generation, Yi et al.
(NeurIPS 2022) MoCapAct : A Multi-Task Dataset for Simulated Humanoid Control, Wagener et al.
(ACM MM 2022) ForcePose : Learning to Estimate External Forces of Human Motion in Video, Louis et al.
(ECCV 2022) BEAT : A Large-Scale Semantic and Emotional Multi-Modal Dataset for Conversational Gestures Synthesis, Liu et al.
(ECCV 2022) BRACE : The Breakdancing Competition Dataset for Dance Motion Synthesis, Moltisanti et al.
(ECCV 2022) EgoBody : Human body shape and motion of interacting people from head-mounted devices, Zhang et al.
(ECCV 2022) GIMO : Gaze-Informed Human Motion Prediction in Context, Zheng et al.
(ECCV 2022) HuMMan : Multi-Modal 4D Human Dataset for Versatile Sensing and Modeling, Cai et al.
(CVPR 2022) ExPI : Multi-Person Extreme Motion Prediction, Guo et al.
(CVPR 2022) HumanML3D : Generating Diverse and Natural 3D Human Motions from Text, Guo et al.
(CVPR 2022) Putting People in their Place : Monocular Regression of 3D People in Depth, Sun et al.
(CVPR 2022) BEHAVE : Dataset and Method for Tracking Human Object Interactions, Bhatnagar et al.
(ICCV 2021) AIST++ : AI Choreographer: Music Conditioned 3D Dance Generation with AIST++, Li et al.
(CVPR 2021) Fit3D : AIFit: Automatic 3D Human-Interpretable Feedback Models for Fitness Training, Fieraru et al.
(CVPR 2021) BABEL : Bodies, Action, and Behavior with English Labels, Punnakkal et al.
(AAAI 2021) HumanSC3D : Learning complex 3d human self-contact, Fieraru et al.
(CVPR 2020) CHI3D : Three-Dimensional Reconstruction of Human Interactions, Fieraru et al.
(ICCV 2019) PROX : Resolving 3D Human Pose Ambiguities with 3D Scene Constraints, Hassan et al.
(ICCV 2019) AMASS : Archive of Motion Capture As Surface Shapes, Mahmood et al.
Humanoid, Simulated or Real
2026
(ArXiv 2026) Sample, Simulate, Select : Physics-in-the-Loop Text-to-Motion for Humanoids Without Training, Memmesheimer et al.
(ArXiv 2026) Brace Yourself : Task-Conditioned Environmental Bracing for Forceful Humanoid Manipulation, Zhang et al.
(ArXiv 2026) HOTICE : Whole-Body Humanoid Object Transportation in Cluttered Environments, Nguyen et al.
(ArXiv 2026) PredActor : Predictive Action Diffusion for Steerable Onboard Humanoid Control, Ding et al.
(ArXiv 2026) MoSAT : Human Motion Generation from Spatial Audio and Textual Description, Komura et al.
(ArXiv 2026) PRIMO : Prior-Informed Odometry from Human-Motion Tracking for Humanoid Robots, Lan et al.
(ArXiv 2026) STRIDER : Stepping-Enabled Multi-Gait Hierarchical 3D Loco-Manipulation Framework for Humanoid Robots, Guo et al.
(ArXiv 2026) EmoPose : Vision-Language Model Guided Emotion-Aware Gesture Generation for Humanoid Robots, Ma et al.
(ArXiv 2026) Whole-Body UMI : Transferring UMI Manipulation Skills to Humanoid Whole-Body Manipulation via Real-Time Motion Generation, Li et al.
(ArXiv 2026) CHOREO : Every Humanoid Skill as a Trajectory, Dong et al.
(ArXiv 2026) LIMBO : Learning and Internalizing Model-Free Barrier Objectives for Agile and Safe Whole-Body Control, Gonzales et al.
(ArXiv 2026) Beyond Kinematics : Benchmarking Simulation Fidelity for Muscle-Driven Imitation Learning, Ahmad et al.
(ArXiv 2026) Learning Scene-Aware Humanoid Locomotion through 3D Clutter from Immersive Human Demonstrations , Wang et al.
(ArXiv 2026) KINO : A Keyframe Interface for VLM Planning and Whole-Body Control in Humanoid Loco-Manipulation, Chen et al.
(ArXiv 2026) Gated Residual Body-Hand Coordination for Whole-Body Humanoid Teleoperation , Wu et al.
(ArXiv 2026) PASSAGE : Scaling Scene-Aligned Motion Learning for Perceptive Humanoid Traversal in Cluttered Environments, Ma et al.
(ArXiv 2026) WholeBodyWAM : Generalizing Pre-trained World-Action Priors to Humanoid Loco-Manipulation via WBC-Grounded Coordination, Li et al.
(ArXiv 2026) X-WBC : A Cross-Embodiment Foundation Model for Humanoid Whole-Body Control, Zhang et al.
(ArXiv 2026) EMoG : Emotion-Modulated Gait Generation for Expressive Humanoid Locomotion, Lu et al.
(ArXiv 2026) Learning Terrain-Adaptive Humanoid Locomotion on Granular Terrain , Kamohara et al.
(CoRL 2026) SwingBot : Learning Whole-Body Brachiation for Humanoid Robots, Xiong et al.
(ArXiv 2026) ViBe : Visual Behavior Adaptation for Perceptive Humanoid Whole-Body Control, Krishna et al.
(ArXiv 2026) TANGO : Humanoid Navigation in Cluttered Environments with a Whole-Body Vision-Language-Action Model, Li et al.
(ArXiv 2026) PGMT : Perceptive General Motion Tracking for Humanoid Robots, Li et al.
(ArXiv 2026) RoboDreamer : Anticipatory Humanoid Locomotion with Predictive State-Space Models, Li et al.
(ArXiv 2026) SkillX : Unified Multi-Skill Policy Learning for Humanoid Soccer, Ye et al.
(ArXiv 2026) Unifying Physics-Based Humanoid Interaction with a Context-Conditioned Interaction Prior , Li et al.
(ArXiv 2026) GLoRI : Closed-Loop Whole-Body Tracking with Global-Local Reference Interaction for Humanoid Loco-Manipulation, Xu et al.
(ArXiv 2026) World-Model-Augmented Visual Locomotion for Humanoids on Foothold-Constrained Terrain , Liu et al.
(ArXiv 2026) Contact-Constrained Lower-Limb Joint-Offset Calibration for Humanoid Robots , Lu et al.
(ArXiv 2026) FOCUS : Foot Observation Confidence for Robust Humanoid Proprioceptive Odometry, Feng et al.
(ArXiv 2026) Unified Motion Retargeting for Humanoids with Learned Point Cloud Correspondence , Cao et al.
(IROS 2026) Reinforcement Learning-Based Control for an Inline Skating Humanoid Robot , Marot et al.
(CLAWAR 2026) Cheng et al. : Learning Roller-Skating Motions of Humanoid Robots Based on Adversarial Motion Priors, Cheng et al.
(RSS 2026) PRIME : Physically-consistent Robotic Inertial and Motion Estimation for Legged and Humanoid Robots, Kang et al.
(RSS 2026) Mind Your Steps : A General Learning Framework for Accurate Humanoid Foothold Tracking, Montenegro et al.
(SIGGRAPH 2026) ReActor : Reinforcement Learning for Physics-Aware Motion Retargeting, Muumlller et al.
(SIGGRAPH 2026) MotionBricks : Scalable Real-Time Motions with Modular Latent Generative Model and Smart Primitives, Wang et al.
(Nature 2026) In vivo feasibility study of humanoid robots in surgery , Liang et al.
(CVPR 2026) MaskAdapt : Learning Flexible Motion Adaptation via Mask-Invariant Prior for Physics-Based Characters, Park et al.
(CVPR 2026) InterPrior : Scaling Generative Control for Physics-Based Human-Object Interactions, Xu et al.
(ICRA 2026) TrajBooster : Boosting Humanoid Whole-Body Manipulation via Trajectory-Centric Learning, Liu et al.
(ICLR 2026) WholeBodyVLA : Towards Unified Latent VLA for Whole-Body Loco-Manipulation Control, Jiang et al.
(L4DC 2026) FALCON : Learning Force-Adaptive Humanoid Loco-Manipulation, Zhang et al.
(Github 2026) UFO : A General Unsupervised Reinforcement Learning Framework for Humanoid Control.
(ArXiv 2026) AdaPT : Towards Professional Tennis Styles for Humanoid Robots with Adaptive Motion Planning and Tracking, Huang et al.
(ArXiv 2026) HumanTracker : Towards Comprehensive and Human-Aligned Motion Tracking Benchmark, Liu et al.
(ArXiv 2026) HumanoidVLN : A Physics-Grounded Simulator and Benchmark for Vision-Language Navigation Across Diverse Humanoid Embodiments, Pham et al.
(ArXiv 2026) LUCID : Latent-Skill Unified Control via Imagined Dynamics for Long-Horizon Humanoid Loco-Manipulation, Guo et al.
(ArXiv 2026) ω-0 : A Latent Predictive World Action Model for Concurrent Humanoid Loco-Manipulation, Li et al.
(ArXiv 2026) Learning Context-Aware Motion Priors for Humanoid Control , Mo et al.
(ArXiv 2026) PFM-HR : Pose Flow Matching for Humanoid Robots, Gao et al.
(ArXiv 2026) StableMimic : Smooth Human-Like Recovery for Humanoid Motion Tracking - Learning Beyond the Tracking Distribution for Structured Post-Fall Behavior, Wu et al.
(ArXiv 2026) Teleopit : A Full-Embodiment Humanoid Teleoperation System, Wu et al.
(ArXiv 2026) LooperMuscle : Fast and Stable Learning of Humanoid Whole-Body Tracking via Structured Mixture-of-Experts, Liu et al.
(ArXiv 2026) PAC-MAN : Perception-Aware CBF-RL for Whole-Body Safety in Humanoid Dodgeball, Yang et al.
(ArXiv 2026) Closing the Lab-to-Store Gap : A Data-Efficient Post-Training and Experience-Driven Learning VLA Framework for Retail Humanoids, Sisó et al.
(ArXiv 2026) Extreme-RGMT : Continual Learning of Highly Dynamic Skills for Robust Generalist Humanoid Control, Ma et al.
(ArXiv 2026) What Matters in Humanoid General Motion Tracking? : An Empirical Study, Amadio et al.
(ArXiv 2026) POT-VLA : Closing the Loop in Humanoid VLA: Persistent 3D Object Tokens for Verifiable Loco-Manipulation, Ren et al.
(ArXiv 2026) Handroid : Bridging Dexterous Hand and Humanoid, Li et al.
(ArXiv 2026) Scaling Behavior Foundation Model for Humanoid Robots , Zeng et al.
(ArXiv 2026) GaitSpan : Growing Humanoid Locomotion from Walking to Running, Lin et al.
(ArXiv 2026) ContactMimic : Humanoid Object Interaction via Contact Control, Li et al.
(ArXiv 2026) LingBot-VLA 2.0 : From Foundation to Application: Improving VLA Models in Practice, Wu et al.
(ArXiv 2026) Athena-WBC : Capability-Aligned Policy Experts for Long-Tail Humanoid Whole-Body Control, Jiang et al.
(ArXiv 2026) ADP : Adversarial Dynamics Priors for Physically Grounded Humanoid Locomotion, Lee et al.
(ArXiv 2026) HEFT : Heavy-Payload Full-size Humanoid Teleoperation with Privileged Motion Guidance and Windowed Payload Curriculum, Liu et al.
(ArXiv 2026) FastDSAC : Enhancing Policy Plasticity via Constrained Exploration for Scalable Humanoid Locomotion, Lu et al.
(ArXiv 2026) VLK : Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes, Wang et al.
(ArXiv 2026) ReactiveBFM : Reactive Closed-Loop Motion Planning Towards Universal Humanoid Whole-Body Control, Chen et al.
(ArXiv 2026) X-Morph : Human Motion Priors for Scalable Robot Learning Across Morphologies, Sharma et al.
(ArXiv 2026) AnyBody : Free-Form Whole-Body Humanoid Control from Arbitrary Keypoint Guidance, Li et al.
(ArXiv 2026) FADA : Few-Shot Domain Adaptation via Dynamics Alignment for Humanoid Control, Xie et al.
(ArXiv 2026) Booster Lab : A Data-Centric Pipeline for Learning Deployable Humanoid Locomotion Policies, Chen et al.
(ArXiv 2026) CWI : Composite Humanoid Whole-Body Imitation System for Loco-manipulation, Ge et al.
(ArXiv 2026) SceneBot : Contact-Prompted General Humanoid Whole Body Tracking with Scene-Interaction, Chen et al.
(ArXiv 2026) HumanoidUMI : Bridging Robot-Free Demonstrations and Humanoid Whole-Body Manipulation, Wang et al.
(ArXiv 2026) Humanoid-DART : Humanoid Loco-Manipulation using Diffusion-guided Augmentation through Relabeling and Tracking, Debbad et al.
(ArXiv 2026) PressMimic : Pressure-Guided Motion Capture and Control for Humanoid Robot Imitation, Lu et al.
(ArXiv 2026) TaskNPoint : How to Teach Your Humanoid to Hit a Backhand in Minutes, Werner et al.
(ArXiv 2026) OmniContact : Chaining Meta-Skills via Contact Flow for Generalizable Humanoid Loco-Manipulation, Yu et al.
(ArXiv 2026) Learning Asynchronous Upper-body Task-space Trajectory Tracking Policy for Humanoid Robots , Liu et al.
(ArXiv 2026) WOLF-VLA : Whole-Body Humanoid Optimal Locomotion Framework for Vision-Language-Action Learning, Boukheddimi et al.
(ArXiv 2026) RGB : RL Guided Whole-Body MPPI for Humanoid Control, Seo et al.
(ArXiv 2026) CoorDex : Coordinating Body and Hand Priors for Continuous Dexterous Humanoid Loco-Manipulation, Li et al.
(ArXiv 2026) TEXEDO : Test Time Scaling for Controller-aware Language-conditioned Humanoid Motion Generation, Cao et al.
(ArXiv 2026) OpenHLM : An Empirical Recipe for Whole-Body Humanoid Loco-Manipulation, Hu et al.
(ArXiv 2026) TACT-ful : Multi-Channel Terrain Affordance and Compliance Training for Payload-Robust Perceptive Humanoid Locomotion, Ly et al.
(ArXiv 2026) Proprioceptive Invariant State Estimation for Humanoid Robots on Non-Inertial Ground , Mandali et al.
(ArXiv 2026) HALOMI : Learning Humanoid Loco-Manipulation with Active Perception from Human Demonstrations, Zhao et al.
(ArXiv 2026) VENOM : Versatile Embodied Network for Omni-bodied Motion Tracking, Padmanabhan et al.
(ArXiv 2026) WaveSync : Constrained Wavefront Optimization for Synchronized Co-Speech Gestures in Humanoid Robots, Tran et al.
(ArXiv 2026) ADAPT : Analytical Disturbance-Aware Policy Training for Humanoid Locomotion, Lyu et al.
(ArXiv 2026) λ-Reachability : Geometric-Horizon Safety Bellman Equations for Humanoid Safety, Chen et al.
(ArXiv 2026) WT-UMI : Tactile-based Whole-Body Manipulation via Force-Supervised Contact-Aware Planning, Jang et al.
(ArXiv 2026) Proprioceptive-visual correspondence enables self-other distinction in humanoid robots , Chen et al.
(ArXiv 2026) GenHOI : Contact-Aware Humanoid-Object Interaction by Imitating Generated Videos without Task-Specific Training, Bi et al.
(ArXiv 2026) Stubborn : A Streamlined and Unified Reinforcement Learning Framework for Robust Motion Tracking and Fall Recovery for Humanoids, Ren et al.
(ArXiv 2026) Critic Architecture Matters : Dual vs. Unified Critics for Humanoid Loco-Manipulation, Yardimci.
(ArXiv 2026) RoboNaldo : Accurate, Stable and Powerful Humanoid Soccer Shooting via Motion-Guided Curriculum Reinforcement Learning, Zhong et al.
(ArXiv 2026) MARCH : Model-Assisted Reinforcement Learning for the Perceptive Control of Humanoids over Sparse Footholds, Crismariu et al.
(ArXiv 2026) VAIC : Vision-Guided Humanoid Agile Object Interaction Control via Decoupled Commands, Li et al.
(ArXiv 2026) MotionWAM : Towards Foundation World Action Models for Real-Time Humanoid Loco-Manipulation, Zheng et al.
(ArXiv 2026) OASIS : From Simulation Data Collection to Real-World Humanoid Loco-Manipulation, Yu et al.
(ArXiv 2026) EgoPriMo : Egocentric Motion Generation for Interactive Humanoid Control, Ge et al.
(ArXiv 2026) Perceptive Behavior Foundation Model : Adapting Human Motion Priors to Robot-Centric Terrain, Wang et al.
(ArXiv 2026) Predictive Style Matching : Natural and Robust Humanoid Locomotion, Nedelchev et al.
(ArXiv 2026) T-GMP : Terrain-conditioned Generative Motion Priors for Versatile and Natural Humanoid Locomotion, Guo et al.
(ArXiv 2026) HANDOFF : Humanoid Agentic Task-Space Whole-Body Control via Distilled Complementary Teachers, Yang et al.
(ArXiv 2026) MotionDisco : Motion Discovery for Extreme Humanoid Loco-Manipulation, Taouil et al.
(ArXiv 2026) TAGA : Terrain-aware Active Gaze Learning for Generalizable Agile Humanoid Locomotion, Li et al.
(ArXiv 2026) LadderMan : Learning Humanoid Perceptive Ladder Climbing, Zhao et al.
(ArXiv 2026) Accelerating and Scaling MPC-Guided Reinforcement Learning for Humanoid Locomotion and Manipulation , Li et al.
(ArXiv 2026) GRAIL : Generating Humanoid Loco-Manipulation from 3D Assets and Video Priors, Xie et al.
(ArXiv 2026) CoRe-MoE : Contrastive Reweighted Mixture of Experts for Multi-Terrain Humanoid Locomotion with Gait Adaptation, Huang et al.
(ArXiv 2026) Humanoid-GPT : Scaling Data and Structure for Zero-Shot Motion Tracking, Qi et al.
(ArXiv 2026) Bionic Human-Motion Style Transfer : for Physically Executable Whole-Body Control of Humanoid Robots, Huang et al.
(ArXiv 2026) Human2Humanoid : Physics-Aware Cross-Morphology Motion Retargeting for Humanoid Robots, Huang et al.
(ArXiv 2026) SplitAdapter : Load-Aware Humanoid Loco-Manipulation via Factorized Adaptation, Kang et al.
(ArXiv 2026) PHASOR : Phase-Anchored Universal Action Representations for Humanoid Embodiments, Kim et al.
(ArXiv 2026) LEGS : Fine-Tuning Teleop-Free VLAs for Humanoid Loco-manipulation in an Embodied Gaussian Splatting World, Kim et al.
(ArXiv 2026) GLAD : Global-Local Attention Decomposition for Terrain Encoding in Humanoid Perceptive Locomotion, Fu et al.
(ArXiv 2026) ConstrainedMimic : Constrained Whole-Body Tracking for Humanoid Robots, Morton et al.
(ArXiv 2026) HOIST : Humanoid Optimization with Imitation and Sample-efficient Tuning for Manipulating Suspended Loads, Liu et al.
(ArXiv 2026) SSR : Scaling Surefooted and Symmetric Humanoid Traversal to the Open World, Yu et al.
(ArXiv 2026) SPRINT : Efficient Spectral Priors for Humanoid Athletic Sprints, Wei et al.
(ArXiv 2026) HumanoidMimicGen : Data Generation for Loco-Manipulation via Whole-Body Planning, Lin et al.
(ArXiv 2026) MIND : Multi-Scale Intent Diffusion for Text-Driven Physics-Based Humanoid Control, Li et al.
(ArXiv 2026) Lee et al : Safety-Critical Whole-Body Control for Humanoid Robots via Input-to-State Safe Control Barrier Functions, Lee et al.
(ArXiv 2026) MuGen : Multi-Skill Generative Locomotion Controller for Humanoid Robots, Feng et al.
(ArXiv 2026) Direct Dynamic Retargeting for Humanoid Imitation Learning from Videos , Roux et al.
(ArXiv 2026) Any2Any : Efficient Cross-Embodiment Transfer for Humanoid Whole-Body Tracking, Yang et al.
(ArXiv 2026) SCRIPT : Scalable Diffusion Policy with Multi-stage Training for Language-driven Physics-Based Humanoid Control, Zhang et al.
(ArXiv 2026) Humanoid Whole-Body Manipulation via Active Spatial Brain and Generalizable Action Cerebellum , Liang et al.
(ArXiv 2026) SUGAR : A Scalable Human-Video-Driven Generalizable Humanoid Loco-Manipulation Learning Framework, Wu et al.
(ArXiv 2026) CEER : Compliant End-Effector and Root Control as a Unified Interface for Hierarchical Humanoid Loco-Manipulation, Luo et al.
(ArXiv 2026) Xu et al : Domain-Adaptive Communication-Rate Optimization for Sim-to-Real Humanoid-Robot Wireless XR Teleoperation, Xu et al.
(ArXiv 2026) Ghosh et al : Adversarial Stress Testing of SPARK Humanoid Safety Filters, Ghosh et al.
(ArXiv 2026) Unified Walking, Running, and Recovery for Humanoids : via State-Dependent Adversarial Motion Priors, Lu et al.
(ArXiv 2026) Terrain Consistent Reference-Guided RL : for Humanoid Navigation Autonomy, Compton et al.
(ArXiv 2026) HoloMotion-1 : Technical Report, Chen et al.
(ArXiv 2026) DAJI : Before the Body Moves: Learning Anticipatory Joint Intent for Language-Conditioned Humanoid Control, Jia et al.
(ArXiv 2026) Durrani et al : Real-Time Whole-Body Teleoperation of a Humanoid Robot Using IMU-Based Motion Capture with Sim2Sim and Sim2Real Validation, Durrani et al.
(ArXiv 2026) Zhang et al : Explicit Stair Geometry Conditioning for Robust Humanoid Locomotion, Zhang et al.
(ArXiv 2026) SixthSense : Task-Agnostic Proprioception-Only Whole-Body Wrench Estimation for Humanoids, Chen et al.
(ArXiv 2026) Egocentric Tactile and Proximity Sensors as Observation Priors for Humanoid Collision Avoidance , Kohlbrenner et al.
(ArXiv 2026) SynAgent : Generalizable Cooperative Humanoid Manipulation via Solo-to-Cooperative Agent Synergy, Yao et al.
(ArXiv 2026) Learning Whole-Body Humanoid Locomotion via Motion Generation and Motion Tracking , Zhang et al.
(ArXiv 2026) Switch : Learning Agile Skills Switching for Humanoid Robots, Lau et al.
(ArXiv 2026) Learning Versatile Humanoid Manipulation with Touch Dreaming , Niu et al.
(ArXiv 2026) Tree Learning : A Multi-Skill Continual Learning Framework for Humanoid Robots, Yan et al.
(ArXiv 2026) Safe Human-to-Humanoid Motion Imitation Using Control Barrier Functions , Cai et al.
(ArXiv 2026) HEX : Humanoid-Aligned Experts for Cross-Embodiment Whole-Body Manipulation, Bai et al.
(ArXiv 2026) RoSHI : A Versatile Robot-oriented Suit for Human Data In-the-Wild, Mao et al.
(ArXiv 2026) Dynamic Whole-Body Dancing with Humanoid Robots : A Model-Based Control Approach, Zhang et al.
(ArXiv 2026) SMASH : Mastering Scalable Whole-Body Skills for Humanoid Ping-Pong with Egocentric Vision, Ren et al.
(ArXiv 2026) BAT : Balancing Agility and Stability via Online Policy Switching for Long-Horizon Whole-Body Humanoid Control, Baek et al.
(ArXiv 2026) Learning Humanoid Navigation from Human Data , Wang et al.
(ArXiv 2026) Heracles : Bridging Precise Tracking and Generative Synthesis for General Humanoid Control, Tao et al.
(ArXiv 2026) Chasing Autonomy : Dynamic Retargeting and Control Guided RL for Performant and Controllable Humanoid Running, Olkin et al.
(ArXiv 2026) PCHC : Enabling Preference Conditioned Humanoid Control via Multi-Objective Reinforcement Learning, Li et al.
(ArXiv 2026) SafeFlow : Real-Time Text-Driven Humanoid Whole-Body Control via Physics-Guided Rectified Flow and Selective Safety Gating, Cho et al.
(ArXiv 2026) Sun et al : Learning Safe-Stoppability Monitors for Humanoid Robots, Sun et al.
(ArXiv 2026) Make Tracking Easy : Neural Motion Retargeting for Humanoid Whole-body Control, Zhao et al.
(ArXiv 2026) Cha et al : Sim-to-Real of Humanoid Locomotion Policies via Joint Torque Space Perturbation Injection, Cha et al.
(ArXiv 2026) AGILE : A Comprehensive Workflow for Humanoid Loco-Manipulation Learning, Zhao et al.
(ArXiv 2026) Morphology-Consistent Humanoid Interaction : through Robot-Centric Video Synthesis, Xu et al.
(ArXiv 2026) PhyGile : Physics-Prefix Guided Motion Generation for Agile General Humanoid Motion Tracking, Bao et al.
(ArXiv 2026) PRIOR : Perceptive Learning for Humanoid Locomotion with Reference Gait Priors, Han et al.
(ArXiv 2026) RoboForge : Physically Optimized Text-guided Whole-Body Locomotion for Humanoids, Yuan et al.
(ArXiv 2026) ECHO : Edge-Cloud Humanoid Orchestration for Language-to-Motion Control, Jia et al.
(ArXiv 2026) He et al : Enforcing Task-Specified Compliance Bounds for Humanoids via Anisotropic Lipschitz-Constrained Policies, He et al.
(ArXiv 2026) HALO : Closing Sim-to-Real Gap for Heavy-loaded Humanoid Agile Motion Skills via Differentiable Simulation, Wang et al.
(ArXiv 2026) CyboRacket : A Perception-to-Action Framework for Humanoid Racket Sports, Ren et al.
(ArXiv 2026) OmniClone : Engineering a Robust, All-Rounder Whole-Body Humanoid Teleoperation System, Li et al.
(ArXiv 2026) PhysMoDPO : Physically-Plausible Humanoid Motion with Preference Optimization, Zhang et al.
(ArXiv 2026) LATENT : Learning Athletic Humanoid Tennis Skills from Imperfect Human Motion Data, Zhang et al.
(ArXiv 2026) Psi_0 : An Open Foundation Model Towards Universal Humanoid Loco-Manipulation, Wei et al.
(ArXiv 2026) HumDex : Humanoid Dexterous Manipulation Made Easy, Heng et al.
(ArXiv 2026) SPARK : Skeleton-Parameter Aligned Retargeting on Humanoid Robots with Kinodynamic Trajectory Optimization, Wang et al.
(ArXiv 2026) Cybo-Waiter : A Physical Agentic Framework for Humanoid Whole-Body Locomotion-Manipulation, Ren et al.
(ArXiv 2026) SteadyTray : Learning Object Balancing Tasks in Humanoid Tray Transport via Residual Reinforcement Learning, Huang et al.
(ArXiv 2026) KDMR : Kinodynamic Motion Retargeting for Humanoid Locomotion via Multi-Contact Whole-Body Trajectory Optimization, Zhang et al.
(ArXiv 2026) SCDP : Learning Humanoid Locomotion from Partial Observations via Mixed-Observation Distillation, Carroll et al.
(ArXiv 2026) ZeroWBC : Learning Natural Visuomotor Humanoid Control Directly from Human Egocentric Video, Yang et al.
(ArXiv 2026) FAME : Force-Adaptive RL for Expanding the Manipulation Envelope of a Full-Scale Humanoid, Pudasaini et al.
(ArXiv 2026) Poddar et al : Embedding Classical Balance Control Principles in Reinforcement Learning for Humanoid Recovery, Poddar et al.
(ArXiv 2026) MetaWorld-X : Hierarchical World Modeling via VLM-Orchestrated Experts for Humanoid Loco-Manipulation, Shen et al.
(ArXiv 2026) Jiang et al : Omnidirectional Humanoid Locomotion on Stairs via Unsafe Stepping Penalty and Sparse LiDAR Elevation Mapping, Jiang et al.
(ArXiv 2026) GeoLoco : Leveraging 3D Geometric Priors from Visual Foundation Model for Robust RGB-Only Humanoid Locomotion, Liu et al.
(ArXiv 2026) Xiang et al : Perceptive Variable-Timing Footstep Planning for Humanoid Locomotion on Disconnected Footholds, Xiang et al.
(ArXiv 2026) HybridMimic : Hybrid RL-Centroidal Control for Humanoid Motion Mimicking, Tay et al.
(ICRA 2026) CMoE : Contrastive Mixture of Experts for Motion Control and Terrain Adaptation of Humanoid Robots, Ma et al.
(ArXiv 2026) IO-WBC : Interaction-Aware Whole-Body Control for Compliant Object Transport, Zhang et al.
(ArXiv 2026) Cognition to Control : Multi-Agent Learning for Human-Humanoid Collaborative Transport, Zhang et al.
(ArXiv 2026) Omni-Manip : Beyond-FOV Large-Workspace Humanoid Manipulation with Omnidirectional 3D Perception, Qu et al.
(ArXiv 2026) PhysiFlow : Physics-Aware Humanoid Whole-Body VLA via Multi-Brain Latent Flow Matching and Robust Tracking, Qin et al.
(ArXiv 2026) X-Loco : Towards Generalist Humanoid Locomotion Control via Synergetic Policy Distillation, Wang et al.
(ArXiv 2026) ULTRA : Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation, He et al.
(ArXiv 2026) Rhythm : Learning Interactive Whole-Body Control for Dual Humanoids, Chen et al.
(ArXiv 2026) SLMP : Spherical Latent Motion Prior for Physics-Based Simulated Humanoid Control, Tan et al.
(ArXiv 2026) Scaling Tasks, Not Samples : Mastering Humanoid Control through Multi-Task Model-Based Reinforcement Learning, Liu et al.
(ArXiv 2026) Shi et al : Minimalist Compliance Control, Shi et al.
(ArXiv 2026) Pro-HOI : Perceptive Root-guided Humanoid-Object Interaction, Lin et al.
(ArXiv 2026) OmniXtreme : Breaking the Generality Barrier in High-Dynamic Humanoid Control, Wang et al.
(ArXiv 2026) OmniTrack : General Motion Tracking via Physics-Consistent Reference, Li et al.
(ArXiv 2026) LessMimic : Long-Horizon Humanoid Interaction with Unified Distance Field Representations, Lin et al.
(ArXiv 2026) Feng et al : Biomechanical Comparisons Reveal Divergence of Human and Humanoid Gaits, Feng et al.
(ArXiv 2026) Xu et al : Iterative Closed-Loop Motion Synthesis for Scaling the Capabilities of Humanoid Control, Xu et al.
(ArXiv 2026) MeshMimic : Geometry-Aware Humanoid Motion Learning through 3D Scene Reconstruction, Zhang et al.
(ArXiv 2026) Perceptive Humanoid Parkour : Chaining Dynamic Human Skills via Motion Matching, Wu et al.
(ArXiv 2026) VIGOR : Visual Goal-In-Context Inference for Unified Humanoid Fall Safety, Azulay et al.
(ArXiv 2026) HERO : Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation, Dong et al.
(ArXiv 2026) EgoActor : Grounding Task Planning into Spatial-aware Egocentric Actions for Humanoid Robots via Visual-Language Models, Bai et al.
(ArXiv 2026) HoRD : Robust Humanoid Control via History-Conditioned Reinforcement Learning and Online Distillation, Wang et al.
(ArXiv 2026) CMR : Contractive Mapping Embeddings for Robust Humanoid Locomotion on Unstructured Terrains, Zeng et al.
(ArXiv 2026) HUSKY : Humanoid Skateboarding System via Physics-Aware Whole-Body Control, Han et al.
(ArXiv 2026) RPL : Learning Robust Humanoid Perceptive Locomotion on Challenging Terrains, Zhang et al.
(ArXiv 2026) EAGLE : Embodiment-Aware Generalist Specialist Distillation for Unified Humanoid Whole-Body Control, Peng et al.
(ArXiv 2026) PDF-HR : Pose Distance Fields for Humanoid Robots, Gu et al.
(ArXiv 2026) Kong et al : Learning Soccer Skills for Humanoid Robots: A Progressive Perception-Action Framework, Kong et al.
(ArXiv 2026) XHugWBC : Scalable and General Whole-Body Control for Cross-Humanoid Locomotion, Xue et al.
(ArXiv 2026) TextOp : Real-time Interactive Text-Driven Humanoid Robot Motion Generation and Control, Xie et al.
(ArXiv 2026) Now You See That : Learning End-to-End Humanoid Locomotion from Raw Pixels, Sun et al.
(ArXiv 2026) Bridging Speech, Emotion, and Motion : a VLM-based Multimodal Edge-deployable Framework for Humanoid Robots, Yang et al.
(ArXiv 2026) MOSAIC : Bridging the Sim-to-Real Gap in Generalist Humanoid Motion Tracking and Teleoperation with Rapid Residual Adaptation, Sun et al.
(ArXiv 2026) Chen et al : Learning Human-Like Badminton Skills for Humanoid Robots, Chen et al.
(ArXiv 2026) RoboStriker : Hierarchical Decision-Making for Autonomous Humanoid Boxing, Yin et al.
(ArXiv 2026) RGMT : Robust and Generalized Humanoid Motion Tracking, Ma et al.
(ArXiv 2026) RAPT : Model-Predictive Out-of-Distribution Detection and Failure Diagnosis for Sim-to-Real Humanoid Robots, Munn et al.
(ArXiv 2026) HumanX : Toward Agile and Generalizable Humanoid Interaction Skills from Human Videos, Wang et al.
(ArXiv 2026) Yi et al : Flow Policy Gradients for Robot Control, Yi et al.
(ArXiv 2026) CAT : Collision-Free Humanoid Traversal in Cluttered Indoor Scenes, Xue et al.
(ArXiv 2026) Fauna Sprout : A lightweight, approachable, developer-ready humanoid robot, Fauna Robotics Team.
(ArXiv 2026) Li et al : Generalizable Geometric Prior and Recurrent Spiking Feature Learning for Humanoid Robot Manipulation, Li et al.
(ArXiv 2026) Huang et al : Learning Whole-Body Human-Humanoid Interaction from Human-Human Demonstrations, Huang et al.
(ArXiv 2026) FastStair : Learning to Run Up Stairs with Humanoid Robots, Liu et al.
(ArXiv 2026) Zhuang et al : Deep Whole-body Parkour, Zhuang et al.
(ArXiv 2026) WaveMan : mmWave-Based Room-Scale Human Interaction Perception for Humanoid Robots, Hu et al.
(ArXiv 2026) Hiking in the Wild : A Scalable Perceptive Parkour Framework for Humanoids, Zhu et al.
(ArXiv 2026) Walk the PLANC : Physics-Guided RL for Agile Humanoid LocomotioN on Constrained Footholds, Dai et al.
(ArXiv 2026) Yang et al : Locomotion Beyond Feet, Yang et al.
(ArXiv 2026) SKATER : Synthesized Kinematics for Advanced Traversing Efficiency on a Humanoid Robot via Roller Skate Swizzles, Gu et al.
2025
(SIGGRAPH Asia 2025) MaskedManipulator : Versatile Whole-Body Control for Loco-Manipulation, Tessler et al.
(CoRL 2025) Hold My Beer🍻 : Learning Gentle Humanoid Locomotion and End-Effector Stabilization Control, Li et al.
(CoRL 2025) Robot Trains Robot : Automatic Real-World Policy Adaptation and Learning for Humanoids, Hu et al.
(CoRL 2025) HuB : Learning Extreme Humanoid Balance, Zhang et al.
(CoRL 2025) CLONE : Closed-Loop Whole-Body Humanoid Teleoperation for Long-Horizon Tasks, Li et al.
(CoRL 2025) Hand-Eye Autonomous Delivery : Learning Humanoid Navigation, Locomotion and Reaching, Ye et al.
(ICCV 2025) PDC : Emergent Active Perception and Dexterity of Simulated Humanoids from Visual Reinforcement Learning, Luo et al.
(ICCV 2025) SIMS : Simulating Human-Scene Interactions with Real World Script Planning, Wang et al.
(ICCV 2025) ModSkill : Physical Character Skill Modularization, Huang et al.
(ICCV 2025) UniPhys : Unified Planner and Controller with Diffusion for Flexible Physics-Based Character Control, Wu et al.
(RSS 2025) HOMIE : Humanoid Loco-Manipulation with Isomorphic Exoskeleton Cockpit, Ben et al.
(RSS 2025) BeamDojo : Learning Agile Humanoid Locomotion on Sparse Footholds, Wang et al.
(RSS 2025) ASAP : Aligning Simulation and Real-World Physics for Learning Agile Humanoid Whole-Body Skills, He et al.
(RSS 2025) HumanUP : Learning Getting-Up Policies for Real-World Humanoid Robots, He et al.
(RSS 2025) Demonstrating Berkeley Humanoid Lite : An Open-source, Accessible, and Customizable 3D-printed Humanoid Robot, Chi et al.
(RSS 2025) AMO : Adaptive Motion Optimization for Hyper-Dexterous Humanoid Whole-Body Control, Li et al.
(RSS 2025) HoST : Learning Humanoid Standing-up Control across Diverse Postures, Huang et al.
(RSS 2025 Workshop) Exbody2 : Advanced Expressive Humanoid Whole-Body Control, Ji et al.
(SIGGRPAH 2025) Diffuse-CLoC : Guided Diffusion for Physics-based Character Look-ahead Control, Huang et al.
(SIGGRAPH 2025) AMOR : Adaptive Character Control through Multi-Objective Reinforcement Learning, Alegre et al.
(SIGGRAPH 2025) PARC : Physics-based Augmentation with Reinforcement Learning for Character Controllers, Xu et al.
(SIGGRAPH 2025) SkillMimic-v2 : Learning Robust and Generalizable Interaction Skills from Sparse and Noisy Demonstrations, Yu et al.
(CVPR 2025) POMP : Physics-constrainable Motion Generative Model through Phase Manifolds, Ji et al.
(CVPR 2025) Let Humanoids Hike! Integrative Skill Development on Complex Trails, Lin et al.
(CVPR 2025) GROVE : A Generalized Reward for Learning Open-Vocabulary Physical Skill, Cui et al.
(CVPR 2025) InterMimic : Towards Universal Whole-Body Control for Physics-Based Human-Object Interactions, Xu et al.
(CVPR 2025) SkillMimic : Learning Reusable Basketball Skills from Demonstrations, Wang et al.
(CVPR 2025) Neural Motion Simulator : Pushing the Limit of World Models in Reinforcement Learning, Hao et al.
(Eurographics 2025) Bae et al : Versatile Physics-based Character Control with Hybrid Latent Representation, Bae et al.
(ICRA 2025) Boguslavskii et al : Human-Robot Collaboration for the Remote Control of Mobile Humanoid Robots with Torso-Arm Coordination, Boguslavskii et al.
(ICRA 2025) HOVER : Versatile Neural Whole-Body Controller for Humanoid Robots, He et al.
(ICRA 2025) PIM : Learning Humanoid Locomotion with Perceptive Internal Model, Long et al.
(ICRA 2025) Think on your feet : Seamless Transition between Human-like Locomotion in Response to Changing Commands, Huang et al.
(ICLR 2025) MimicLabs : What Matters in Learning from Large-Scale Datasets for Robot Manipulation, Saxena et al.
(ICLR 2025) Puppeteer : Hierarchical World Models as Visual Whole-Body Humanoid Controllers, Hansen et al.
(ICLR 2025) FB-CPR : Zero-Shot Whole-Body Humanoid Control via Behavioral Foundation Models, Tirinzoni et al.
(ICLR 2025) MPC2 : Motion Control of High-Dimensional Musculoskeletal System with Hierarchical Model-Based Planning, Wei et al.
(ICLR 2025) CLoSD : Closing the Loop between Simulation and Diffusion for multi-task character control, Tevet et al.
(ICLR 2025) HiLo : Learning Whole-Body Human-like Locomotion with Motion Tracking Controller, Zhang et al.
(Github 2025) MobilityGen : MobilityGen.
(ArXiv 2025) UniAct : Unified Motion Generation and Action Streaming for Humanoid Robots, Jiang et al.
(ArXiv 2025) Do You Have Freestyle? : Expressive Humanoid Locomotion via Audio Control, Li et al.
(ArXiv 2025) RoboMirror : Understand Before You Imitate for Video to Humanoid Locomotion, Li et al.
(ArXiv 2025) EGM : Efficiently Learning General Motion Tracking Policy for High Dynamic Humanoid Whole-Body Control, Yang et al.
(ArXiv 2025) CHIP : Adaptive Compliance for Humanoid Control through Hindsight Perturbation, Chen et al.
(ArXiv 2025) Spraggett et al : Learning to Get Up Across Morphologies: Zero-Shot Recovery with a Unified Humanoid Policy, Spraggett et al.
(ArXiv 2025) PvP : Data-Efficient Humanoid Robot Learning with Proprioceptive-Privileged Contrastive Representations, Yuan et al.
(ArXiv 2025) Mimic2DM : Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions, Li et al.
(ArXiv 2025) Xu et al. : Learning Agile Striker Skills for Humanoid Soccer Robots from Noisy Sensory Input, Xu et al.
(ArXiv 2025) Song et al. : Gait-Adaptive Perceptive Humanoid Locomotion with Real-Time Under-Base Terrain Reconstruction, Song et al.
(ArXiv 2025) Toward Seamless Physical Human-Humanoid Interaction : Insights from Control, Intent, and Modeling with a Vision for What Comes Next, Cardona et al.
(ArXiv 2025) Kumbhar et al : Efficient and Compliant Control Framework for Versatile Human-Humanoid Collaborative Transportation, Kumbhar et al.
(ArXiv 2025) SMP : Reusable Score-Matching Motion Priors for Physics-Based Character Control, Mu et al.
(ArXiv 2025) GenMimic : From Generated Human Videos to Physically Plausible Robot Trajectories, Ni et al.
(ArXiv 2025) H-Zero : Cross-Humanoid Locomotion Pretraining Enables Few-shot Novel Embodiment Transfer, Lin et al.
(ArXiv 2025) Xue et al : Opening the Sim-to-Real Door for Humanoid Pixel-to-Action Policy Transfer, Xue et al.
(ArXiv 2025) Seo et al : Learning Sim-to-Real Humanoid Locomotion in 15 Minutes, Seo et al.
(ArXiv 2025) Seo et al : Discovering Self-Protective Falling Policy for Humanoid Robot via Deep Reinforcement Learning, Shi et al.
(ArXiv 2025) SafeHumanoid : VLM-RAG-driven Control of Upper Body Impedance for Humanoid Robot, Mahmoud et al.
(ArXiv 2025) Commanding Humanoid by Free-form Language : A Large Language Action Model with Unified Motion Vocabulary, Liu et al.
(ArXiv 2025) Agility Meets Stability : Versatile Humanoid Control with Heterogeneous Data, Pan et al.
(ArXiv 2025) SafeFall : Learning Protective Control for Humanoid Robots, Meng et al.
(ArXiv 2025) SENTINEL : A Fully End-to-End Language-Action Model for Humanoid Whole Body Control, Wang et al.
(ArXiv 2025) HAFO : Humanoid Force-Adaptive Control for Intense External Force Interaction Environments, Dong et al.
(ArXiv 2025) Xiao et al. : Kinematics-Aware Multi-Policy Reinforcement Learning for Force-Capable Humanoid Loco-Manipulation, Xiao et al.
(ArXiv 2025) Switch-JustDance : Benchmarking Whole-Body Motion Tracking Policies Using a Commercial Console Game, Kim et al.
(ArXiv 2025) VIRAL : Visual Sim-to-Real at Scale for Humanoid Loco-Manipulation, He et al.
(ArXiv 2025) HMC : Learning Heterogeneous Meta-Control for Contact-Rich Loco-Manipulation, Wei et al.
(ArXiv 2025) Gallant : Voxel Grid-based Humanoid Locomotion and Local-navigation across 3D Constrained Terrains, Ben et al.
(ArXiv 2025) Liu et al. : Humanoid Whole-Body Badminton via Multi-Stage Reinforcement Learning, Liu et al.
(ArXiv 2025) SPIDER : Scalable Physics-Informed DExterous Retargeting, Pan et al.
(ArXiv 2025) SCHUR : Unveiling the Impact of Data and Model Scaling on High-Level Control for Humanoid Robots, Wei et al.
(ArXiv 2025) RGMP : Recurrent Geometric-prior Multimodal Policy for Generalizable Humanoid Robot Manipulation, Li et al.
(ArXiv 2025) SONIC : Supersizing Motion Tracking for Natural Humanoid Whole-Body Control, Luo et al.
(ArXiv 2025) FIRM : Unified Humanoid Fall-Safety Policy from a Few Demonstrations, Xu et al.
(ArXiv 2025) AHC : Towards Adaptive Humanoid Control via Multi-Behavior Distillation and Reinforced Fine-Tuning, Zhao et al.
(ArXiv 2025) BFM-Zero : A Promptable Behavioral Foundation Model for Humanoid Control Using Unsupervised Reinforcement Learning, Ze et al.
(ArXiv 2025) GentleHumanoid : Learning Upper-body Compliance for Contact-rich Human and Object Interaction, Lu et al.
(ArXiv 2025) Wang et al : Learning Vision-Driven Reactive Soccer Skills for Humanoid Robots, Wang et al.
(ArXiv 2025) TWIST2 : Scalable, Portable, and Holistic Humanoid Data Collection System, Ze et al.
(ArXiv 2025) Huang et al : One-shot Humanoid Whole-body Motion Learning, Huang et al.
(ArXiv 2025) Thor : Towards Human-Level Whole-Body Reactions for Intense Contact-Rich Environments, Li et al.
(ArXiv 2025) PHUMA : Building the Bridge Between Off-the-Shelf VLMs and the Physical World, Lee et al.
(ArXiv 2025) Kwon et al : A Humanoid Visual-Tactile-Action Dataset for Contact-Rich Manipulation, Kwon et al.
(ArXiv 2025) Endowing GPT-4 with a Humanoid Body : Building the Bridge Between Off-the-Shelf VLMs and the Physical World, Jian et al.
(ArXiv 2025) Humanoid Goalkeeper : Learning from Position Conditioned Task-Motion Constraints, Ren et al.
(ArXiv 2025) SoftMimic : Learning Compliant Whole-body Control from Examples, Margolis et al.
(ArXiv 2025) AdaMimic : Towards Adaptable Humanoid Control via Adaptive Motion Tracking, Huang et al.
(ArXiv 2025) COLA : Learning Human-Humanoid Coordination for Collaborative Object Carrying, Du et al.
(ArXiv 2025) Architecture Is All You Need : Diversity-Enabled Sweet Spots for Robust Humanoid Locomotion, Werner et al.
(ArXiv 2025) From Language to Locomotion : Retargeting-free Humanoid Control via Motion Latent Guidance, Li et al.
(ArXiv 2025) Wu et al : Path and Motion Optimization for Efficient Multi-Location Inspection with Humanoid Robots, Wu et al.
(ArXiv 2025) DemoHLM : From One Demonstration to Generalizable Humanoid Loco-Manipulation, Fu et al.
(ArXiv 2025) PhysHSI : Towards a Real-World Generalizable and Natural Humanoid-Scene Interaction System, Wang et al.
(ArXiv 2025) Ego-VCP : Ego-Vision World Model for Humanoid Contact Planning, Liu et al.
(ArXiv 2025) Humanoid Everyday : A Comprehensive Robotic Dataset for Open-World Humanoid Manipulation, Zhao et al.
(ArXiv 2025) DPL : Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction, Sun et al.
(ArXiv 2025) ResMimic : From General Motion Tracking to Humanoid Whole-body Loco-Manipulation via Residual Learning, Zhao et al.
(ArXiv 2025) Retargeting Matters : General Motion Retargeting for Humanoid Motion Tracking, Ara´ujo et al.
(ArXiv 2025) PolySim : Bridging the Sim-to-Real Gap for Humanoid Control via Multi-Simulator Dynamics Randomization, Lei et al.
(ArXiv 2025) D'Elia et al : Stabilizing Humanoid Robot Trajectory Generation via Physics-Informed Learning and Control-Informed Steering, D'Elia et al.
(ArXiv 2025) OmniRetarget : Interaction-Preserving Data Generation for Humanoid Whole-Body Loco-Manipulation and Scene Interaction, Yang et al.
(ArXiv 2025) MoReFlow : Motion Retargeting Learning through Unsupervised Flow Matching, Kim et al.
(ArXiv 2025) Towards Versatile Humanoid Table Tennis : Unified Reinforcement Learning with Prediction Augmentation, Hu et al.
(ArXiv 2025) SEEC : Stable End-Effector Control with Model-Enhanced Residual Learning for Humanoid Loco-Manipulation, Jang et al.
(ArXiv 2025) RuN : Residual Policy for Natural Humanoid Locomotion, Li et al.
(ArXiv 2025) RobotDancing : Residual-Action Reinforcement Learning Enables Robust Long-Horizon Humanoid Motion Tracking, Sun et al.
(ArXiv 2025) VisualMimic : Visual Humanoid Loco-Manipulation via Motion Tracking and Generation, Yin et al.
(ArXiv 2025) Chasing Stability : Humanoid Running via Control Lyapunov Function Guided Reinforcement Learning, Olkin et al.
(ArXiv 2025) RoMoCo : Robotic Motion Control Toolbox for Reduced-Order Model-Based Locomotion on Bipedal and Humanoid Robots, Dai et al.
(ArXiv 2025) HDMI : Learning Interactive Humanoid Whole-Body Control from Human Videos, Weng et al.
(ArXiv 2025) HuMam : Humanoid Motion Control via End-to-End Deep Reinforcement Learning with Mamba, Wang et al.
(ArXiv 2025) KungfuBot2 : Learning Versatile Motion Skills for Humanoid Whole-Body Contro, Han et al.
(ArXiv 2025) IKMR : Implicit Kinodynamic Motion Retargeting for Human-to-humanoid Imitation Learning, Chen et al.
(ArXiv 2025) DreamControl : Human-Inspired Whole-Body Humanoid Control for Scene Interaction via Guided Diffusion, Kalaria et al.
(ArXiv 2025) BFM : Behavior Foundation Model for Humanoid Robots, Zeng et al.
(ArXiv 2025) Zheng et al : Embracing Bulky Objects with Humanoid Robots: Whole-Body Manipulation with Reinforcement Learning, Zheng et al.
(ArXiv 2025) StageAC
Show more Truncated — view the full README on GitHub .