Awesome 3D Generation for Embodied AI

A curated list of papers, datasets, benchmarks, simulators, and evaluation resources for 3D generation in embodied AI and robotic simulation.
This repository accompanies our survey on 3D generation for embodied AI and robotic simulation. It organizes the literature around how generative 3D methods support embodied systems, from simulation-ready asset creation, to interactive scene generation, to sim-to-real transfer. In addition to research papers, we collect resources that are useful for building and evaluating embodied 3D generation systems, including datasets, benchmarks, simulator platforms, and asset formats.
Contributions are welcome. Please open an issue or submit a pull request if you would like to add a paper, dataset, benchmark, simulator, or evaluation resource.
Highlights
- Survey-driven taxonomy for 3D generation in embodied AI
- Curated coverage of papers, datasets, benchmarks, simulators, and evaluation
- Emphasis on simulation-ready assets and environments rather than visual realism alone
Table of Contents
Scope and Taxonomy
We focus on 3D generation methods that are useful for embodied AI, robot learning, and robotic simulation. Compared with general-purpose 3D generation, embodied settings require more than visual realism: generated assets and environments should be executable, physically meaningful, interactive, and compatible with simulation platforms.
We organize the field into three major roles:
- Data Generator: generating simulation-ready objects and assets, including articulated, physically grounded, and deformable content.
- Simulation Environments: generating interactive, controllable, and task-oriented 3D scenes for embodied training and evaluation.
- Sim2Real Bridge: using 3D generation for digital twins, data augmentation, and synthetic demonstrations to support transfer to real robots.
Awesome List
In chronological order, from the earliest to the latest.
Data Generator
Articulated Object Generation
| Model | Paper | Venue | Website | GitHub |
|---|
NAP | NAP: Neural 3D Articulated Object Prior | NeurIPS '23 |  | - |
CAGE | CAGE: Controllable Articulation Generation | CVPR '24 |  |  |
URDFormer | URDFormer: A Pipeline for Constructing Articulated Simulation Environments from Real-World Images | RSS '24 |  |  |
SINGAPO | SINGAPO: Single Image Controlled Generation of Articulated Parts in Objects | ICLR '25 |  |  |
Real2Code | Real2code: Reconstruct Articulated Objects via Code Generation | ICLR '25 | - | - |
Articulate-Anything | Articulate-Anything: Automatic Modeling of Articulated Objects via a Vision-Language Foundation Model | ICLR '25 |  |  |
MagicArticulate | MagicArticulate: Make Your 3D Models Articulation-Ready | CVPR '25 |  |  |
MeshArt | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | CVPR '25 |  |  |
PartRM | PartRM: Modeling Part-Level Dynamics with Large Cross-State Reconstruction Model | CVPR '25 |  | - |
ArtFormer | ArtFormer: Controllable Generation of Diverse 3D Articulated Objects | CVPR '25 | - | - |
ArtiWorld | ArtiWorld: LLM-Driven Articulation of 3D Objects in Scenes | arXiv '25 | - | - |
Articulate Anymesh | Articulate Anymesh: Open-Vocabulary 3D Articulated Objects Modeling | CoRL '25 |  | - |
ATOP | Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization | arXiv '25 | - | - |
URDF-Anything | URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model | NeurIPS '25 |  |  |
ArtiLatent | ArtiLatent: Realistic Articulated 3D Object Generation via Structured Latents | SIG. Asia '25 |  |  |
ArtGen | ArtGen: Conditional Generative Modeling of Articulated Objects in Arbitrary Part-Level States | arXiv '25 | - | - |
DreamArt | DreamArt: Generating Interactable Articulated Objects from a Single Image | arXiv '25 |  | - |
GAOT | GAOT: Generating Articulated Objects Through Text-Guided Diffusion Models | ACM MM Asia '25 | - | - |
UniArt | UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation | arXiv '25 | - | - |
Infinite Mobility | Infinite Mobility: Scalable High-Fidelity Synthesis of Articulated Objects via Procedural Generation | arXiv '25 |  |  |
Kinematify | Kinematify: Open-Vocabulary Synthesis of High-DoF Articulated Objects | ICRA '26 |  | - |
Particulate | Particulate: Feed-Forward 3D Object Articulation | CVPR '26 |  | - |
URDF-Anything+ | URDF-Anything+: Autoregressive Articulated 3D Models Generation for Physical Simulation | arXiv '26 |  |  |
ArtLLM | ArtLLM: Generating Articulated Assets via 3D LLM | CVPR '26 |  | - |
SPARK | SPARK: Sim-Ready Part-Level Articulated Reconstruction with VLM Knowledge | CVPR '26 |  | - |
PAct | PAct: Part-Decomposed Single-View Articulated Object Generation | arXiv '26 |  | - |
Physically-Grounded Object Generation
| Model | Paper | Venue | Website | GitHub |
|---|
NeRF2Physics | Physical Property Understanding from Language-Embedded Feature Fields | CVPR '24 |  |  |
Atlas3D | Atlas3D: Physically Constrained Self-Supporting Text-To-3D for Simulation and Fabrication | NeurIPS '24 |  |  |
Physically Compatible 3D | Physically Compatible 3D Object Modeling from a Single Image | NeurIPS '24 | - | - |
PhyCAGE | PhyCAGE: Physically Plausible Compositional 3D Asset Generation from a Single Image | NeurIPS '24 |  |  |
PhysGaussian | PhysGaussian: Physics-Integrated 3D Gaussians for Generative Dynamics | CVPR '24 |  | - |
PhysPart | PhysPart: Physically Plausible Part Completion for Interactable Objects | ICRA '25 |  | - |
DSO | DSO: Aligning 3D Generators with Simulation Feedback for Physical Soundness | ICCV '25 |  | - |
GaussianProperty | GaussianProperty: Integrating Physical Properties to 3D Gaussians with LMMs | ICCV '25 |  | - |
PhysX-3D | PhysX-3D: Physical-Grounded 3D Asset Generation | NeurIPS '25 |  |  |
SOPHY | SOPHY: Learning to Generate Simulation-Ready Objects with Physical Materials | arXiv '25 |  |  |
Pixie | Pixie: Fast and Generalizable Supervised Learning of 3D Physics from Pixels | arXiv '25 |  |  |
DensiCrafter | DensiCrafter: Physically-Constrained Generation and Fabrication of Self-Supporting Hollow Structures | arXiv '25 | - | - |
PhysX-Anything | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | CVPR '26 |  |  |
| Model | Paper | Venue | Website | GitHub |
|---|
DressCode | DressCode: Autoregressively Sewing and Generating Garments from Text Guidance | ToG (SIG. '24) |  |  |
PhysDreamer | PhysDreamer: Physics-Based Interaction with 3D Objects via Video Generation | ECCV '24 |  |  |
GarmentDreamer | GarmentDreamer: 3DGS Guided Garment Synthesis with Diverse Geometry and Texture Details | arXiv '24 |  | - |
Dress-1-to-3 | Dress-1-to-3: Single Image to Simulation-Ready 3D Outfit with Diffusion Prior and Differentiable Physics | ToG (SIG. '25) |  | - |
Image2Garment | Image2Garment: Simulation-Ready Garment Generation from a Single Image | arXiv '26 |  | - |
End-to-End Sim-Ready Asset Pipelines
| Model | Paper | Venue | Website | GitHub |
|---|
TRELLIS | Structured 3D Latents for Scalable and Versatile 3D Generation | CVPR '25 |  |  |
TRELLIS.2 | Native and Compact Structured Latents for 3D Generation | arXiv '25 |  |  |
EmbodiedGen | EmbodiedGen: Towards a Generative 3D World Engine for Embodied Intelligence | NeurIPS '25 |  | - |
Seed3D | Seed3D 1.0: From Images to High-Fidelity Simulation-Ready 3D Assets | arXiv '25 |  | - |
Simulation Environments
Structure-Driven Scene Generation
| Model | Paper | Venue | Website | GitHub |
|---|
Graph-to-3D | Graph-to-3D: End-To-End Generation and Manipulation of 3D Scenes Using Scene Graphs | ICCV '21 |  |  |
ATISS | ATISS: Autoregressive Transformers for Indoor Scene Synthesis | NeurIPS '21 |  |  |
ProcTHOR | ProcTHOR: Large-Scale Embodied AI Using Procedural Generation | NeurIPS '22 |  |  |
CC3D | CC3D: Layout-Conditioned Generation of Compositional 3D Scenes | ICCV '23 |  |  |
LayoutGPT | LayoutGPT: Compositional Visual Planning and Generation with Large Language Models | NeurIPS '23 |  |  |
Infinigen Indoors | Infinigen Indoors: Photorealistic Indoor Scenes Using Procedural Generation | CVPR '24 |  | - |
Holodeck | Holodeck: Language Guided Generation of 3D Embodied AI Environments | CVPR '24 |  |  |
DiffuScene | DiffuScene: Denoising Diffusion Models for Generative Indoor Scene Synthesis | CVPR '24 |  |  |
Controllable Scene Generation
| Model | Paper | Venue | Website | GitHub |
|---|
PhyScene | PhyScene: Physically Interactable 3D Scene Synthesis for Embodied AI | CVPR '24 |  |  |
InstructScene | InstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph Prior | ICLR '24 |  |  |
DepR | DepR: Depth Guided Single-View Scene Reconstruction with Instance-Level Diffusion | ICCV '25 |  |  |
MIDI | MIDI: Multi-Instance Diffusion for Single Image to 3D Scene Generation | CVPR '25 |  |  |
DynScene | DynScene: Scalable Generation of Dynamic Robotic Manipulation Scenes for Embodied AI | CVPR '25 | - | - |
FactoredScenes | From Programs to Poses: Factored Real-World Scene Generation via Learned Program Libraries | NeurIPS '25 |  |  |
Steerable Scene Generation | Steerable Scene Generation with Post Training and Inference-Time Search | CoRL '25 |  |  |
SceneFoundry | SceneFoundry: Generating Interactive Infinite 3D Worlds | arXiv '26 |  | - |
SceneGen | SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass | 3DV '26 |  |  |
Agentic and Task-Oriented Scene Generation
| Model | Paper | Venue | Website | GitHub |
|---|
SceneCraft | Scenecraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code | ICML '24 | - | - |
Architect | Architect: Generating Vivid and Interactive 3D Scenes with Hierarchical 2D Inpainting | NeurIPS '24 |  |  |
OptiScene | OptiScene: LLM-Driven Indoor Scene Layout Generation via Scaled Human-Aligned Data Synthesis and Multi-Stage Preference Optimization | NeurIPS '25 |  |  |
LayoutVLM | LayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language Models | CVPR '25 |  |  |
SceneWeaver | SceneWeaver: All-In-One 3D Scene Synthesis with an Extensible and Self-Reflective Agent | NeurIPS '25 |  |  |
CAST | CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image | TOG '25 |  | - |
3D-RE-GEN | 3D-RE-GEN: 3D Reconstruction of Indoor Scenes with a Generative Framework | arXiv '25 |  | - |
MetaScenes | MetaScenes: Towards Automated Replica Creation for Real-World 3D Scans | CVPR '25 |  |  |
Disco-Layout | Disco-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesis | arXiv '25 |  | - |
MarketGen | MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments | arXiv '25 |  | - |
MesaTask | MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning | NeurIPS '25 |  |  |
TabletopGen | TabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single Image | arXiv '25 |  |  |
PAT3D | PAT3D: Physics-Augmented Text-To-3D Scene Generation | ICLR '26 | - | - |
PhyScensis | PhyScensis: Physics-Augmented LLM Agents for Complex Physical Scene Arrangement | ICLR '26 |  | - |
SAGE | SAGE: Scalable Agentic 3D Scene Generation for Embodied AI | arXiv '26 |  |  |
SceneSmith | SceneSmith: Agentic Generation of Simulation-Ready Indoor Scenes | arXiv '26 |  |  |
Scenethesis | Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation | ICLR '26 | - | - |
3D-Generalist | 3D-Generalist: Vision-Language-Action Models for Crafting 3D Worlds | 3DV '26 |  | - |
Sim2Real Bridge
Digital Twin Construction
| Model | Paper | Venue | Website | GitHub |
|---|
Ditto | Ditto: Building Digital Twins of Articulated Objects from Interaction | CVPR '22 |  |  |
PARIS | PARIS: Part-Level Reconstruction and Motion Analysis for Articulated Objects | ICCV '23 |  |  |
DRAWER | DRAWER: Digital Reconstruction and Articulation with Environment Realism | CVPR '25 |  |  |
LiteReality | LiteReality: Graphics-Ready 3D Scene Reconstruction from RGB-D Scans | NeurIPS '25 |  |  |
LatticeWorld | LatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generation | arXiv '25 | - | - |
RoboSimGS | High-Fidelity Simulated Data Generation for Real-World Zero-Shot Robotic Manipulation Learning with Gaussian Splatting | arXiv '25 |  |  |
Real-is-Sim | Real-is-Sim: Bridging the Sim-To-Real Gap with a Dynamic Digital Twin | arXiv '25 |  | - |
ArticFlow | ArticFlow: Generative Simulation of Articulated Mechanisms | arXiv '25 | - | - |
ArticulateGS | ArticulateGS: Self-Supervised Digital Twin Modeling of Articulated Objects Using 3D Gaussian Splatting | CVPR '25 | - | - |
ArtGS | ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting | ICLR '25 |  |  |
PhysTwin | PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos | ICCV '25 |  |  |
RoboScape | RoboScape: Physics-informed embodied world model | arXiv '25 | - |  |
TwinAligner | TwinAligner: Visual-Dynamic Alignment Empowers Physics-aware Real2Sim2Real for Robotic Manipulation | arXiv '25 |  |  |
Scalable Real2Sim | Scalable Real2Sim: Physics-Aware Asset Generation via Robotic Pick-And-Place Setups | IROS '25 |  |  |
ReartGS++ | ReartGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting | CVPR '26 |  | - |
Data Augmentation
| Model | Paper | Venue | Website | GitHub |
|---|
RoboGSim | RoboGSim: A Real2Sim2Real Robotic Gaussian Splatting Simulator | arXiv '24 |  | - |
Splat-MOVER | Splat-MOVER: Multi-Stage, Open-Vocabulary Robotic Manipulation via Editable Gaussian Splatting | arXiv '24 | - | - |
Manipulate Anywhere | Learning to Manipulate Anywhere: A Visual Generalizable Framework for Reinforcement Learning | arXiv '24 | - | - |
RoboSplat | Novel Demonstration Generation with Gaussian Splatting Enables Robust One-Shot Manipulation | RSS '25 |  |  |
SplatSim | SplatSim: Zero-Shot Sim2Real Transfer of RGB Manipulation Policies Using Gaussian Splatting | ICRA '25 |  |  |
SIGHT | SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories | arXiv '25 | - | - |
RoboTransfer | RoboTransfer: Controllable Geometry-Consistent Video Diffusion for Manipulation Policy Transfer | arXiv '25 | - | - |
GAF | GAF: Gaussian Action Field as a 4D Representation for Dynamic World Modeling in Robotic Manipulation | arXiv '25 |  | - |
EgoDemoGen | EgoDemoGen: Novel Egocentric Demonstration Generation Enables Viewpoint-Robust Manipulation | arXiv '25 | - | - |
ExoGS | ExoGS: A 4D Real-To-Sim-To-Real Framework for Scalable Manipulation Data Collection | arXiv '26 | - |  |
Synthetic Demonstration and Task Generation
| Model | Paper | Venue | Website | GitHub |
|---|
MimicGen | MimicGen: A Data Generation System for Scalable Robot Learning Using Human Demonstrations | CoRL '23 |  |  |
Gen2Sim | Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models | ICRA '24 |  |  |
GenSim2 | GenSim2: Scaling Robot Data Generation with Multi-Modal and Reasoning LLMs | CoRL '24 |  |  |
DemoGen | DemoGen: Synthetic Demonstration Generation for Data-Efficient Visuomotor Policy Learning | arXiv '25 |  |  |
Real2Render2Real | Real2Render2Real: Scaling Robot Data Without Dynamics Simulation or Robot Hardware | CoRL '25 |  | - |
ManipDreamer3D | ManipDreamer3D: Synthesizing Plausible Robotic Manipulation Video with Occupancy-Aware 3D Trajectory | arXiv '25 | - | - |
DreamGen | DreamGen: Unlocking Generalization in Robot Learning Through Video World Models | arXiv '25 |  |  |
AnchorDream | AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis | arXiv '25 |  | - |
PhysWorld | Robot Learning from a Physical World Model | arXiv '25 |  |  |
DRAW2ACT | DRAW2ACT: Turning Depth-Encoded Trajectories into Robotic Demonstration Videos | arXiv '25 | - | - |
Video2Act | Video2Act: A Dual-System Video Diffusion Policy with Robotic Spatio-Motional Modeling | arXiv '25 | - | - |
GWM | GWM: Towards Scalable Gaussian World Models for Robotic Manipulation | ICCV '25 |  |  |
GraspVLA | GraspVLA: A Grasping Foundation Model Pre-Trained on Billion-Scale Synthetic Action Data | arXiv '25 |  |  |
AOMGen | AOMGen: Photoreal, Physics-Consistent Demonstration Generation for Articulated Object Manipulation | CVPR Find. '26 | - | - |
Datasets and Benchmarks
Object-Level Datasets
| Model | Paper | Venue | Details | Website |
|---|
ShapeNet | ShapeNet: An Information-Rich 3D Model Repository | arXiv '15 | 51K objects / 55 cat. |  |
PartNet | PartNet: A Large-scale Benchmark for Fine-grained and Hierarchical Part-level 3D Object Understanding | CVPR '19 | 26.7K objects / 24 cat. |  |
PartNet-Mobility | SAPIEN: A SimulAted Part-based Interactive ENvironment | CVPR '20 | 2,346 objects / 46 cat. |  |
3D-FUTURE | 3D-FUTURE: 3D Furniture shape with TextURE | IJCV '21 | 13.1K objects / 43 cat. |  |
GSO | Google Scanned Objects: A High-Quality Dataset of 3D Scanned Household Items | ICRA '22 | 1,030 objects |  |
ABO | ABO: Dataset and Benchmarks for Real-World 3D Object Understanding | CVPR '22 | 7,953 objects / 63 cat. |  |
AKB-48 | AKB-48: A Real-World Articulated Object Knowledge Base | CVPR '22 | 2,037 objects / 48 cat. |  |
Objaverse | Objaverse: A Universe of Annotated 3D Objects | CVPR '23 | 800K+ objects |  |
GAPartNet | GAPartNet: Cross-Category Domain-Generalizable Object Perception and Manipulation via Generalizable and Actionable Parts | CVPR '23 | 1,166 objects / 27 cat. |  |
ClothesNet | ClothesNet: An Information-Rich 3D Garment Model Repository with Simulated Clothes | ICCV '23 | 4,400+ garments / 11 cat. |  |
Objaverse-XL | Objaverse-XL: A Universe of 10M+ 3D Objects | NeurIPS '23 | 10.2M+ objects |  |
GarmentLab | GarmentLab: A Unified Simulation and Benchmark for Garment Manipulation | NeurIPS '24 | 9K+ objects / 11 garment cat. |  |
Articulation-XL | MagicArticulate: Make Your 3D Models Articulation-Ready | CVPR '25 | 48K+ objects |  |
DTC | Digital Twin Catalog: A Large-Scale Photorealistic 3D Object Digital Twin Dataset | CVPR '25 | 2K objects / 40 cat. |  |
PhysXNet | PhysX-3D: Physical-Grounded 3D Asset Generation | NeurIPS '25 | 26K objects |  |
PhysX-Mobility | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | CVPR '26 | 2K+ objects / 47 cat. |  |
ManiTwin | ManiTwin: Scaling Data-Generation-Ready Digital Object Dataset to 100K | arXiv '26 | 100K+ objects |  |
Scene-Level Datasets
| Model | Paper | Venue | Details | Website |
|---|
Matterport3D | Matterport3D: Learning from RGB-D Data in Indoor Environments | CVPR '17 | 90 scenes / 10.8K pans / 194K RGB-D |  |
ScanNet | ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes | 3DV '17 | 1.5K scans / 2.5M views |  |
Replica | The Replica Dataset: A Digital Replica of Indoor Spaces | arXiv '19 | 18 scenes |  |
3DSSG | Learning 3D Semantic Scene Graphs from 3D Indoor Reconstructions | CVPR '20 | 1.5K SGs / 48K objs / 544K rels |  |
3D-FRONT | 3D-FRONT: 3D Furnished Rooms with layOuts and semaNTics | ICCV '21 | 18.9K rooms / 6.8K houses |  |
HM3D | Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI | NeurIPS '21 | 1K scans |  |
ProcTHOR-10K | ProcTHOR: Large-Scale Embodied AI Using Procedural Generation | NeurIPS '22 | 10K houses |  |
MetaScenes | MetaScenes: Towards Automated Replica Creation for Real-world 3D Scans | CVPR '25 | 706 scenes / 15.4K objs |  |
MesaTask-10K | MesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoning | arXiv '25 | 10.7K tabletop scenes |  |
MarketGen | MarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environments | arXiv '25 | Supermarket envs / 1.1K+ assets |  |
SAGE-10K | SAGE: Scalable Agentic 3D Scene Generation for Embodied AI | arXiv '26 | 10K scenes / 565K objs |  |
Robot Demonstration Datasets
| Model | Paper | Venue | Details | Website |
|---|
RLBench | RLBench: The Robot Learning Benchmark & Learning Environment | RA-L '20 | Franka Panda (CoppeliaSim) · 100 tasks / unlimited demos |  |
CALVIN | CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks | RA-L '22 | Franka Panda (PyBullet) · 34 tasks / 24h play data |  |
ManiSkill2/3 | ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills | ICLR '23 | Multi-robot (SAPIEN) · 20+ task families / 2K+ objects |  |
BridgeData V2 | BridgeData V2: A Dataset for Robot Learning at Scale | CoRL '23 | WidowX 250 · 60K trajectories / 24 envs / 13 skills |  |
LIBERO | LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning | NeurIPS '23 | Franka Panda (robosuite) · 130 tasks / 6.5K demos |  |
MimicGen | MimicGen: A Data Generation System for Scalable Robot Learning using Human Demonstrations | CoRL '23 | Franka Panda (robosuite) · 50K+ demos from ~200 seeds |  |
Open X-Embodiment | Open X-Embodiment: Robotic Learning Datasets and RT-X Models | ICRA '24 | 22 embodiments (60 datasets) · 1M+ trajectories / 527 skills |  |
DROID | DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset | RSS '24 | Franka Panda (18 robots) · 76K trajectories / 86 tasks |  |
RH20T | RH20T: A Comprehensive Robotic Dataset for Learning Diverse Skills in One-Shot | ICRA '24 | 4 robots / 7 configurations · 110K+ sequences / 147 tasks |  |
RoboTwin | RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins | ECCV '24 | ALOHA dual-arm · 50 bimanual tasks / 731 objects |  |
RoboCasa | RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots | RSS '24 | Franka (robosuite/MuJoCo) · 100K+ trajectories / 100 tasks |  |
RoboTwin 2.0 | RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation | arXiv '25 | 5 bimanual embodiments · 100K+ trajectories / 50 tasks |  |
Embodied Benchmarks
| Model | Paper | Venue | Details | Website |
|---|
RLBench | RLBench: The Robot Learning Benchmark & Learning Environment | RA-L '20 | Franka Panda (CoppeliaSim) · 100 tasks / unlimited demos |  |
CALVIN | CALVIN: A Benchmark for Language-Conditioned Policy Learning for Long-Horizon Robot Manipulation Tasks | RA-L '22 | Franka Panda (PyBullet) · 34 tasks / 24h play data |  |
ManiSkill2/3 | ManiSkill2: A Unified Benchmark for Generalizable Manipulation Skills | ICLR '23 | Multi-robot (SAPIEN) · 20+ task families / 2K+ objects |  |
LIBERO | LIBERO: Benchmarking Knowledge Transfer for Lifelong Robot Learning | NeurIPS '23 | Franka Panda (robosuite) · 130 tasks / 6.5K demos |  |
RoboTwin | RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins | ECCV '24 | ALOHA dual-arm · 50 bimanual tasks / 731 objects |  |
RoboCasa | RoboCasa: Large-Scale Simulation of Everyday Tasks for Generalist Robots | RSS '24 | Franka (robosuite/MuJoCo) · 100K+ trajectories / 100 tasks |  |
RoboTwin 2.0 | RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation | arXiv '25 | 5 bimanual embodiments · 100K+ trajectories / 50 tasks |  |
| Platform | Type | Strengths | Link |
|---|
| MuJoCo | Physics simulator | Stable dynamics and articulated control benchmarking |  |
| Isaac Sim | Robotics simulator | Photorealistic rendering and PhysX-based simulation |  |
| Habitat | Embodied AI platform | Large-scale navigation and interactive embodied environments |  |
| ManiSkill | Robotics benchmark and simulator | High-throughput embodied manipulation evaluation |  |
| OmniGibson | Embodied simulation platform | Rich object semantics and interactive scene modeling |  |
| AI2-THOR | Embodied AI simulator | Procedural indoor environments for navigation and interaction |  |
| Format | Use Case | Notes |
|---|
| URDF | Robot and articulated asset description | Widely used for kinematic structures in robotics |
| MJCF | Physics-centric simulation description | Common in MuJoCo-based embodied benchmarks |
| USD | Scene and asset interchange | Increasingly used in modern simulation pipelines |
| glTF / GLB | 3D asset exchange | Lightweight format for geometry and materials |
| OBJ | Mesh storage | Common geometry format but limited semantics |
Evaluation
We recommend evaluating embodied 3D generation at three levels:
- Geometric Fidelity: shape and appearance quality
- Physical Plausibility and Simulation Compatibility: stability, articulation correctness, collision validity, material realism, and simulator readiness
- Downstream Embodied Task Performance: grasp success, articulated manipulation success, navigation success, task completion, and sim-to-real transfer performance
Citation
If you find this repository useful, please consider citing our survey.
@article{to_be_updated,
title = {3D Generation for Embodied AI and Robotic Simulation: A Survey},
author = {To be updated},
journal = {To be updated},
year = {2025}
}