RoboScholar: A Comprehensive Paper List of Embodied AI and Robotics Research
See the code
It's RoboScholar Project Here, started by Tianxing Chen.
Related Information:
Lumina Embodied AI Community
Embodied-AI-Guide
Open-Vocabulary 3D Articulated Objects Modeling https://arxiv.org/pdf/2507.02747
[] [arXiv 25] LEMON: Learning 3D Human-Object Interaction Relation from 2D Images, arXiv
[] [arXiv 25] Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation, arXiv
[] [RSS 25] Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation, arXiv
[] [arXiv 24] GRAPE: Generalizing Robot Policy via Preference Alignment, arXiv
[] [arXiv 25] GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill, arXiv
[] [arXiv 24] Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers, arXiv
[] [arXiv 24] Generative Image as Action Models, website
[] [arXiv 24] Genie: Generative Interactive Environments, website
[] [CVPR 24 (Highlight)] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects, website
[] [CVPR 23 (Highlight)] GAPartNet: Cross-Category Domain-Generalizable Object Perception and Manipulation via Generalizable and Actionable Parts, website
[] [arXiv 23] GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects, website
[] [arXiv 24] ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics, website
[] [ICCV 23] AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand Pose, website
[] [CVPR 23] BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects, website
[] [arXiv 24] WiLoR: End-to-end 3D hand localization and reconstruction in-the-wild, website
Where2Act: From Pixels to Actions for Articulated 3D Objects
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
Decision Transformer: Reinforcement Learning via Sequence Modeling
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
AO-Grasp: Articulated Object Grasp Generation
Human-to-Robot Imitation in the Wild
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation https://sam-embodied.github.io/, ICML24
PerAct, Act3D
Probing the 3D Awareness of Visual Foundation Model: https://arxiv.org/pdf/2404.08636
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
CLIP: Zero-shot Jack of All Trades, website, CLIP GradCAM CLIP_GradCAM_Visualization
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise: https://arxiv.org/pdf/2402.18699
3D-VLA: A 3D Vision-Language-Action Generative World Model
PDDLGym: Gym Environments from PDDL Problems: https://arxiv.org/abs/2002.06432
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
VisionLLM: https://arxiv.org/abs/2305.11175
Ferret: Refer and Ground Anything Anywhere at Any Granularity: https://github.com/apple/ml-ferret
LangSplat
Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity
SparseDFF
ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics
Stabilizing Transformers for Reinforcement Learning
CoBERL: Contrastive BERT for Reinforcement Learning
Adaptive Transformers in RL
Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation
Deep Transformer Q-Networks for Partially Observable Reinforcement Learning
CtrlFormer: Learning Transferable State Representation for Visual Control via Transformer
Sapiens: Foundation for Human Vision Models: https://about.meta.com/realitylabs/codecavatars/sapiens General Flow as Foundation Affordance for Scalable Robot Learning https://general-flow.github.io/
RoboScholar: A Comprehensive Paper List of Embodied AI and Robotics Research
See the code
It's RoboScholar Project Here, started by Tianxing Chen.
Related Information:
Lumina Embodied AI Community
Embodied-AI-Guide
Open-Vocabulary 3D Articulated Objects Modeling https://arxiv.org/pdf/2507.02747
[] [arXiv 25] LEMON: Learning 3D Human-Object Interaction Relation from 2D Images, arXiv
[] [arXiv 25] Merging and Disentangling Views in Visual Reinforcement Learning for Robotic Manipulation, arXiv
[] [RSS 25] Reactive Diffusion Policy: Slow-Fast Visual-Tactile Policy Learning for Contact-Rich Manipulation, arXiv
[] [arXiv 24] GRAPE: Generalizing Robot Policy via Preference Alignment, arXiv
[] [arXiv 25] GROVE: A Generalized Reward for Learning Open-Vocabulary Physical Skill, arXiv
[] [arXiv 24] Scaling Proprioceptive-Visual Learning with Heterogeneous Pre-trained Transformers, arXiv
[] [arXiv 24] Generative Image as Action Models, website
[] [arXiv 24] Genie: Generative Interactive Environments, website
[] [CVPR 24 (Highlight)] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects, website
[] [CVPR 23 (Highlight)] GAPartNet: Cross-Category Domain-Generalizable Object Perception and Manipulation via Generalizable and Actionable Parts, website
[] [arXiv 23] GAMMA: Generalizable Articulation Modeling and Manipulation for Articulated Objects, website
[] [arXiv 24] ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics, website
[] [ICCV 23] AffordPose: A Large-scale Dataset of Hand-Object Interactions with Affordance-driven Hand Pose, website
[] [CVPR 23] BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects, website
[] [arXiv 24] WiLoR: End-to-end 3D hand localization and reconstruction in-the-wild, website
Where2Act: From Pixels to Actions for Articulated 3D Objects
PreAfford: Universal Affordance-Based Pre-Grasping for Diverse Objects and Environments
Decision Transformer: Reinforcement Learning via Sequence Modeling
Toward General-Purpose Robots via Foundation Models: A Survey and Meta-Analysis
AO-Grasp: Articulated Object Grasp Generation
Human-to-Robot Imitation in the Wild
RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots
SAM-E: Leveraging Visual Foundation Model with Sequence Imitation for Embodied Manipulation https://sam-embodied.github.io/, ICML24
PerAct, Act3D
Probing the 3D Awareness of Visual Foundation Model: https://arxiv.org/pdf/2404.08636
ManipVQA: Injecting Robotic Affordance and Physically Grounded Information into Multi-Modal Large Language Models
CLIP: Zero-shot Jack of All Trades, website, CLIP GradCAM CLIP_GradCAM_Visualization
Articulated Object Manipulation with Coarse-to-fine Affordance for Mitigating the Effect of Point Cloud Noise: https://arxiv.org/pdf/2402.18699
3D-VLA: A 3D Vision-Language-Action Generative World Model
PDDLGym: Gym Environments from PDDL Problems: https://arxiv.org/abs/2002.06432
TravelPlanner: A Benchmark for Real-World Planning with Language Agents
VisionLLM: https://arxiv.org/abs/2305.11175
Ferret: Refer and Ground Anything Anywhere at Any Granularity: https://github.com/apple/ml-ferret
LangSplat
Embodied AI with Two Arms: Zero-shot Learning, Safety and Modularity
SparseDFF
ManiPose: A Comprehensive Benchmark for Pose-aware Object Manipulation in Robotics
Stabilizing Transformers for Reinforcement Learning
CoBERL: Contrastive BERT for Reinforcement Learning
Adaptive Transformers in RL
Efficient Transformers in Reinforcement Learning using Actor-Learner Distillation
Deep Transformer Q-Networks for Partially Observable Reinforcement Learning
CtrlFormer: Learning Transferable State Representation for Visual Control via Transformer
Sapiens: Foundation for Human Vision Models: https://about.meta.com/realitylabs/codecavatars/sapiens General Flow as Foundation Affordance for Scalable Robot Learning https://general-flow.github.io/