hzxie/Awesome-3D-Scene-Generation

A curated list of awesome 3D scene generation papers. (IJCV 2026)

1,096

79 commits

updated Sep 18, 2026

See the code

README

Awesome Logo Counter arXiv YouTube PR's Welcome

Overview

This repository collects summaries of over 300 recent studies on 3D scene generation, along with the downstream applications, and will be continuously updated.

If you have suggestions for new resources, improvements to methodologies, or corrections for broken links, please don't hesitate to open an issue or submit a pull request. Contributions of all kinds are welcome and greatly appreciated.

3D-Scene-Generation-Teaser

Table of Contents

Methods: A Hierarchical Taxonomy

Procedural Generation

Rule-based Generation

YearVenueAcronymPaperProjectRepo@GitHub
1988SIGGRAPHTerrain simulation using a model of stream erosion
1989SIGGRAPHThe synthesis and rendering of eroded fractal terrains
1993Graphics InterfaceA fractal model of mountains and riverslink
1998SIGGRAPHRealistic modeling and rendering of plant ecosystemslink
2001SIGGRAPHCityEngineProcedural modeling of cities
2005VRSTModeling Landscapes with Ridges and Rivers
2006TOGProcedural modeling of buildings
2007GDTWCitygenCitygen: An Interactive System for Procedural City Generationlink
2007I3DExample-based model synthesislinkGitHub
2007TVCGTerrain Synthesis from Digital Elevation Modelslink
2008CGFReal-Time Rendering and Editing of Vector-based Terrainslink
2008TOGContinuous model synthesislink
2008TOGInteractive Procedural Street Modelinglink
2009CGFArches: a Framework for Modeling Complex Terrains
2009CGFInteractive Geometric Simulation of 4D Cities
2009TOGInteractive design of urban spaces using geometrical and behavioral modelingyoutube
2010CGFProcedural Generation of Roads
2011CGFInteractive Modeling of City Layouts using Layers of Procedural Content
2011SI3DUrban Ecosystem Design
2011TOGMetropolis procedural modelinglink
2012CGFProcedural Generation of Parcels in Urban Modeling
2012TOGInverse design of urban procedural modelslink
2013TOGTerrain Generation Using Procedural Models Based on Hydrologyyoutube
2013TOGUrban PatternUrban Pattern: Layout Design by Hierarchical Domain Splittinglink
2015TOGWorldBrushWorldBrush: Interactive Example-Based Synthesis of Procedural Virtual Worldslink
2016CGFExample-Driven Procedural Urban Roads
20163DVProceduralization for Editing 3D Architectural Models
2016TOGInteractive Sketching of Urban Procedural Modelslink
2017TOGAuthoring landscapes by combining ecosystem and terrain erosion simulation
2017TOGFast Weather Simulation for Inverse Procedural Design of 3D Urban ModelslinkGitHub
2017TOGInteractive Example-Based Terrain Authoring with Conditional Generative Adversarial Networkslink
2019TOGSynthetic Silviculture: Multi-scale Modeling of Plant Ecosystemslink
2021TOGAuthoring Consistent Landscapes with Flora and Faunalink
2022TOGEcoclimatesEcoclimates: Climate-Response Modeling of Vegetationlink
2022TOGProcedural Urban Forestrylink
2023CVPRInfinigenInfinite Photorealistic Worlds using Procedural GenerationlinkGitHub
2023TOGForming Terrains by Glacial Erosionlinklink
2023TOGLarge-scale terrain authoring through interactive erosion simulationyoutubeGitHub
2023TOGAuthoring and Simulating Meandering RiverslinkGitHub
2025CVPRWProc-GSProc-GS: Procedural Building Generation for City Assembly with 3D GaussianslinkGitHub
2025CEUSVoxCityVoxCity: A seamless framework for open geospatial data integration, grid-based semantic 3D city model generation, and urban environment simulationlinkGitHub

Optimization-based Generation

LLM-based Generation

YearVenueAcronymPaperProjectRepo@GitHub
2023NeurIPSLayoutGPTLayoutGPT: Compositional Visual Planning and Generation with Large Language ModelslinkGitHub
2024CVPRGraphDreamerGraphDreamer: Compositional 3D Scene Synthesis from Scene GraphslinkGitHub
2024ECCVAnyHomeAnyHome: Open-Vocabulary Generation of Structured and Textured 3D HomeslinkGitHub
2024ECCVSceneTellerSceneTeller: Language-to-3D Scene GenerationlinkGitHub
2024ECCVI-DesignI-Design: Personalized LLM Interior DesignerlinkGitHub
2024ICMLSceneCraftSceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code
2024MMControllable Procedural Generation of LandscapesGitHub
2024SIGGRAPH AsiaDISceneDIScene: Object Decoupling and Interaction Modeling for Complex Scene Generationlink
20253DV3D-GPT3D-GPT: Procedural 3D Modeling with Large Language ModelslinkGitHub
2025AAAISceneXSceneX: Procedural Controllable Large-scale Scene GenerationlinkGitHub
2025AAAIHierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language ModelGitHub
2025CVPRGlobal-Local Tree Search in VLMs for 3D Indoor Scene GenerationGitHub
2025CVPRLayoutVLMLayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language ModelslinkGitHub
2025CVPRThe Scene LanguageThe Scene Language: Representing Scenes with Programs, Words, and EmbeddingslinkGitHub
2025ACL FindingsUnrealLLMUnrealLLM: Towards Highly Controllable and Interactable 3D Scene Generation by LLM-powered Procedural Content Generation
2025UISTEchoLadderEchoLadder: Progressive AI-Assisted Design of Immersive VR Scenes
2025NeurIPSSceneWeaverSceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective AgentlinkGitHub
2025NeurIPSMesaTaskMesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoninglinkGitHub
2025NeurIPSOptiSceneOptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference OptimizationlinkGitHub
2025NeurIPSDirectLayoutDirect Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial ReasoninglinkGitHub
2025Nature Computational ScienceUrban planning in the era of large language models
2025SIGGRAPHBuildingBlockBuildingBlockGitHub
2026ICLRScenethesisScenethesis: A Language and Vision Agentic Framework for 3D Scene Generationlink
2026ICLRReSpaceReSpace: Text-Driven 3D Scene Synthesis and Editing with Preference AlignmentlinkGitHub
2026ICMLSceneSmithSceneSmith: Agentic Generation of Simulation-Ready Indoor SceneslinkGitHub
20263DV3D-Generalist3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D WorldslinkGitHub
2026AAAILandCraftLandCraft: Designing the Structured 3D Landscapes via Text GuidanceGitHub
2026AAAIReason-3DText-to-Scene with Large Reasoning ModelslinkGitHub
2026CVPRRAISECityRAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-ScaleGitHub
2026CVPRHOG-LayoutHOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
2026CVPRLaviGenRepurposing 3D Generative Model for Autoregressive Layout GenerationlinkGitHub
2026CVPRYo'CityYo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
2026CVPRMajutsuCityMajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and LayoutslinkGitHub
2026CVPRGardenDesignerGardenDesigner: Encoding Aesthetic Principles into Jiangnan Garden Construction via a Chain of Agentslink
2026CVPRSAGESAGE: Scalable Agentic 3D Scene Generation for Embodied AIlinkGitHub
2026ECCVSceneOrchestraSceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
2026ECCVNaLANaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene GenerationlinkGitHub
2026ACL FindingsSceneLMSceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis
2026ICMLR3LR$^3$L: Reasoning 3D Layouts from Relative Spatial RelationsGitHub
2026ICMLCode2WorldsCode2Worlds: Empowering Coding LLMs for 4D World GenerationlinkGitHub
2024arXivCityCraftCityCraft: A Real Crafter for 3D City GenerationGitHub
2024arXivCityXCityX: Controllable Procedural Content Generation for Unbounded 3D CitieslinkGitHub
2024arXivGraphCanvas3DGraph Canvas for Controllable 3D Scene GenerationGitHub
2024arXivUrbanWorldUrbanWorld: An Urban World Model for 3D City GenerationGitHub
2025arXivCubeCube: A Roblox View of 3D IntelligencelinkGitHub
2025arXivAgentic 3D Scene Generation with Spatially Contextualized VLMslink
2025arXivHLGHLG: Comprehensive 3D Room Construction via Hierarchical Layout GenerationGitHub
2025arXivLatticeWorldLatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generationyoutube
2025arXivCausalStructCausal Reasoning Elicits Controllable 3D Scene GenerationlinkGitHub
2025arXivDisCo-LayoutDisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesislink
2025arXivWorldGenWorldGen: From Text to Traversable and Interactive 3D Worldslink
2025arXivMarketGenMarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environmentslink
2026arXivSceneFoundrySceneFoundry: Generating Interactive Infinite 3D WorldslinkGitHub
2026arXivMANSIONMANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon taskslinkGitHub
2026arXivSceneAssistantSceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene GenerationGitHub
2026arXivSceneCodeSceneCode: Executable World Programs for Editable Indoor Scenes with Articulated ObjectslinkGitHub
2026arXivCode-as-RoomCode-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code SynthesislinkGitHub
2026arXivSpatialGrammarSpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generationlink
2026arXivClosing the Loop: Unified 3D Scene Generation and Immersive Interaction via LLM-RL Couplinglink
2026arXivGlobal-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene GenerationGitHub
2026arXivSceneConductorSceneConductor: 3D Scene Generation from a Single Image with Multi-Agent OrchestrationlinkGitHub
2026arXivSceneReVisSceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RLlinkGitHub
2026arXivWorldClawWorldClaw: Agentic 3D Open-World Generation at ScalelinkGitHub
2026arXivScenePilotScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generationlink
2026arXivSceneMosaicSceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout EvolutionlinkGitHub
2026arXiv4DSynth4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied SimulationlinkGitHub

Neural-3D Generation

Scene Parameters

YearVenueAcronymPaperProjectRepo@GitHub
2018SIGGRAPHDeepSynthDeep Convolutional Priors for Indoor Scene SynthesislinkGitHub
2019CVPRFastSynthFast and Flexible Indoor Scene Synthesis via Deep Convolutional Generative ModelsGitHub
2020SIGGRAPHDeep Generative Modeling for Scene Synthesis via Hybrid Representations
20213DVSceneFormerSceneFormer: Indoor Scene Generation with TransformerslinkGitHub
2021ICCVSync2GenScene Synthesis via Uncertainty-Driven Attribute SynchronizationGitHub
2021NeurIPSATISSATISS: Autoregressive Transformers for Indoor Scene SynthesislinkGitHub
2022ECCVPose2RoomPose2Room: Understanding 3D Scenes from Human ActivitieslinkGitHub
2022SIGGRAPH AsiaSUMMONScene Synthesis from Human MotionlinkGitHub
2023CVPRLearning 3D Scene Priors with 2D SupervisionlinkGitHub
2023CVPRMIMEMIME: Human-Aware 3D Scene GenerationlinkGitHub
2023SIGGRAPHCOFSCOFS: COntrollable Furniture layout Synthesis
2023NeurIPSLanguage-driven Scene Synthesis using Multi-conditional Diffusion ModellinkGitHub
20243DVRoomDesignerRoomDesigner: Encoding Anchor-latents for Style-consistent and Shape-compatible Indoor Scene GenerationGitHub
2024CVPRDiffuSceneDiffuScene: Denoising Diffusion Models for Generative Indoor Scene SynthesislinkGitHub
2024CVPRSceneWiz3DSceneWiz3D: Towards Text-guided 3D Scene CompositionlinkGitHub
2024CVPRPhyScenePhyScene: Physically Interactable 3D Scene Synthesis for Embodied AIlinkGitHub
2024ECCVDreamSceneDreamScene: 3D Gaussian-Based Text-to-3D Scene Generation via Formation Pattern SamplinglinkGitHub
2024ICMLGALA3DGALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian SplattinglinkGitHub
2024ICMLDisentangled 3D Scene Generation with Layout Learninglink
2024MMRelSceneRelScene: A Benchmark and baseline for Spatial Relations in text-driven 3D Scene Generation
2024NeurIPSDeBaRADeBaRA: Denoising-Based 3D Room Arrangement Generation
2024SIGGRAPHINFERACTPhysics-based Scene Layout Generation From Human Motionlink
20253DVCtrl-RoomCtrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout ConstraintslinkGitHub
2025CVPRSceneFactorSceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene GenerationlinkGitHub
2025CVPRCASAGPTCASAGPT: Cuboid Arrangement and Scene Assembly for Interior DesignGitHub
2025CoRLSteerable Scene Generation with Post Training and Inference-Time SearchlinkGitHub
2025NeurIPSFactoredScenesFrom Programs to Poses: Factored Real-World Scene Generation via Learned Program LibrarieslinkGitHub

Scene Graph

YearVenueAcronymPaperProjectRepo@GitHub
2014EMNLPLearning Spatial Knowledge for Text to 3D Scene Generation
2016CGFLearning 3D Scene Synthesis from Annotated RGB-D Images
2017TOGAdaptive synthesis of indoor scenes via activity-associated object relation graphsyoutube
2018TOGLanguage-Driven Synthesis of 3D Scenes from Scene Databaseslink
2019ICCVMeta-SimMeta-Sim: Learning to Generate Synthetic DatasetslinkGitHub
2019SIGGRAPHGRAINSGRAINS: Generative Recursive Autoencoders for INdoor SceneslinkGitHub
2019SIGGRAPHPlanITPlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior NetworksGitHub
2020CVPR3D-SLNEnd-to-End Optimization of Scene LayoutlinkGitHub
2020ECCVMeta-Sim 2Meta-Sim 2 Unsupervised Learning of Scene Structure for Synthetic Data GenerationlinkGitHub
2021ICCVGraph-to-3DGraph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphslinkGitHub
2023NeurIPSCommonScenesCommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph DiffusionlinkGitHub
2023TPAMISceneHGNSceneHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation With Fine-Grained GeometrylinkGitHub
2024ECCVSEKExternal Knowledge Enhanced 3D Scene Generation from Sketch
2024ECCVForest2SeqForest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
2024ECCVEchoSceneEchoScene: Indoor Scene Generation via Information Echo over Scene Graph DiffusionlinkGitHub
2024ICLRInstructSceneInstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph PriorlinkGitHub
2025AAAIMMGDreamerMMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationlinkGitHub
2025MMHiSceneHiScene: Creating Hierarchical 3D Scenes with Isometric View Generationlink
2025CVPRFreeSceneFreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free PromptslinkGitHub
2025ICCVControllable 3D Outdoor Scene Generation via Scene GraphsGitHub
2025TOGImaginariumImaginarium: Vision-guided High-Quality 3D Scene Layout GenerationlinkGitHub
2026TOGCasLayoutCasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling
2026TVCGSceneLinkerSceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequenceslink

Semantic Layout

YearVenueAcronymPaperProjectRepo@GitHub
2021ICCVSGSDIIndoor Scene Generation from a Collection of Semantic-Segmented Depth ImageslinkGitHub
2021ICCVGANcraftGANcraft: Unsupervised 3D Neural Rendering of Minecraft WorldslinkGitHub
2023CVPRDisCoSceneDisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene SynthesislinkGitHub
2023ICCVInfiniCityInfiniCity: Infinite-Scale City Synthesislink
2023ICCVCC3DCC3D: Layout-Conditioned Generation of Compositional 3D SceneslinkGitHub
2023ICCVSet-the-SceneSet-the-Scene: Global-Local Training for Generating Controllable NeRF SceneslinkGitHub
2023ICCVUrbanGIRAFFEUrbanGIRAFFE: Representing Urban Scenes as Compositional Generative Neural Feature FieldslinkGitHub
2023TPAMISceneDreamerSceneDreamer: Unbounded 3D Scene Generation From 2D Image CollectionslinkGitHub
20243DVComp3DCompositional 3D Scene Generation using Locally Conditioned Diffusionlink
2024CVPRCityDreamerCityDreamer: Compositional Generative Model of Unbounded 3D CitieslinkGitHub
2024CVPRBerfSceneBerfScene: Bev-conditioned Equivariant Radiance Fields for Infinite 3D Scene GenerationlinkGitHub
2024NeurIPSSceneCraftSceneCraft: Layout-Guided 3D Scene GenerationlinkGitHub
2024SIGGRAPHBlockFusionBlockFusion: Expandable 3D Scene Generation Using Latent Tri-plane ExtrapolationlinkGitHub
2024SIGGRAPH AsiaFrankensteinFrankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-PlanelinkGitHub
2025CVPRGaussianCityGenerative Gaussian Splatting for Unbounded 3D City GenerationlinkGitHub
2025ICLRLayout-your-3DLayout-your-3D: Controllable and Precise 3D Generation with 2D BlueprintlinkGitHub
2025TPAMICityDreamer4DCityDreamer4D: Compositional Generative Model of Unbounded 4D CitieslinkGitHub
2025TPAMIUrbanGenUrbanGen: Urban Generation with Compositional and Controllable Neural Fieldslink
2025ICCVSat2CitySat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent DiffusionlinkGitHub
2025NeurIPSX-SceneX-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilitylinkGitHub
2026AAAIEarthCrafterEarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent DiffusionlinkGitHub
20263DVSPATIALGENSPATIALGEN: Layout-guided 3D Indoor Scene GenerationlinkGitHub
2026CVPRPrITTIPrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic SceneslinkGitHub
20263DVSemLayoutDiffSemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesislink
2023arXivCompoNeRFCompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene LayoutlinkGitHub
2024arXivUrban ArchitectUrban Architect: Steerable 3D Urban Scene Generation with Layout PriorlinkGitHub
2025arXivLayout2SceneLayout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion PriorslinkGitHub

Implicit Layout

YearVenueAcronymPaperProjectRepo@GitHub
2021CVPRGIRAFFEGIRAFFE: Representing Scenes as Compositional Generative Neural Feature FieldsGitHub
2021ICCVGSNUnconstrained Scene Generation With Locally Conditioned Radiance FieldslinkGitHub
2021ICMLNeRF-VAENeRF-VAE: A geometry aware 3d scene generative model
2022NeurIPSGAUDIGAUDI: A Neural Architect for Immersive 3D Scene GenerationGitHub
2023CVPRPersistent NaturePersistent Nature: A generative model of unbounded 3D worldslinkGitHub
2023CVPRNeuralField-LDMNeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Modelslink
2024CVPRDiffInDSceneDiffInDScene: Diffusion-based High-Quality 3D Indoor Scene GenerationlinkGitHub
2024CVPRXCubeXCube: Large-Scale 3D Generative Modeling using Sparse Voxel HierarchieslinkGitHub
2024CVPRSemCitySemCity: Semantic Scene Generation with Triplane DiffusionlinkGitHub
2024ECCVPDDPyramid Diffusion for Fine 3D Large Scene GenerationlinkGitHub
2024NeurIPSDirector3DDirector3D: Real-world Camera Trajectory and 3D Scene Generation from TextlinkGitHub
2025CVPRLT3SDLT3SD: Latent Trees for 3D Scene DiffusionlinkGitHub
2025CVPRSplatFlowSplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting SynthesislinkGitHub
2025CVPRPrometheusPrometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene GenerationlinkGitHub
2025ICLRDynamicCityDynamicCity: Large-Scale Occupancy Generation from Dynamic SceneslinkGitHub
2025ICCVNuiSceneNuiScene: Exploring Efficient Generation of Unbounded Outdoor SceneslinkGitHub
2025ICCVVideoRFSplatVideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelinglinkGitHub
2026ICLRFlashWorldFlashWorld: High-quality 3D Scene Generation within SecondslinkGitHub
2026ICLRUniUGGUniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodinglinkGitHub
2026ICMLPERSISTBeyond Pixel Histories: World Models with Persistent 3D StatelinkGitHub
2026AAAIWorldGrowWorldGrow: Generating Infinite 3D WorldlinkGitHub
2026AAAILSD-3DLSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Groundinglink
20263DVSceneGenSceneGen: Single-Image 3D Scene Generation in One Feedforward Passlink
2026CVMExCellGenExCellGen: Fast, Controllable, Photorealistic 3D Scene Generation from a Single Real-World ExemplarlinkGitHub
2026CVPRDiff4SplatDiff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction ModelslinkGitHub
2026CVPRScenDiScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generationlink
2026ECCVGaussianGPTGaussianGPT: Towards Autoregressive 3D Gaussian Scene GenerationlinkGitHub
2026ECCVSPAR3SSparse auto-regressive modeling for scene generation from multi-view images
2026SIGGRAPH AsiaSpatialCrafterSpatialCrafter: Single Image World Modeling with Generative 3D ProxieslinkGitHub
2026SIGGRAPHInfiniteDiffusionInfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain GenerationlinkGitHub
2023arXivDiffusion Probabilistic Models for Scene-Scale 3D Categorical DatalinkGitHub
2025arXivTerraTerra: Explorable Native 3D World Model with Point LatentslinkGitHub
2026arXivABot-Earth-0.5ABot-Earth-0.5: Generative 3D Earth ModellinkGitHub

Image-based Generation

Holistic Generation

YearVenueAcronymPaperProjectRepo@GitHub
2019ICIP360-Degree Image Completion by Two-Stage Conditional Gans
2020CVPRSat2GroundGeometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasGitHub
2020WACV360 Panorama Synthesis from a Sparse Set of Images with Unknown Field of View
2021AAAISIG-SSSpherical Image Generation from a Single Image by Considering Scene SymmetryGitHub
2021CVPREnvMapNetHDR Environment Map Estimation for Real-Time Augmented RealitylinkGitHub
2021ICCVSat2vidSat2vid: Street-view panoramic video synthesis from a single satellite image
20223DVImmerseGANGuided Co-Modulated GAN for 360° Field of View Extrapolationlink
2022CVPROmniDreamerDiverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background CreationlinkGitHub
2022ECCVBIPSBIPS: Bi-modal Indoor Panorama Synthesis via Residual Depth-aided Adversarial LearningGitHub
2022SIGGRAPH AsiaText2LightText2Light: Zero-Shot Text-Driven HDR Panorama GenerationlinkGitHub
2022TMMPanoGANCross-View Panorama Image SynthesisGitHub
2022TPAMISat2StrGeometry-Guided Street-View Panorama Synthesis from Satellite ImageryGitHub
2023CVPRDiffCollageDiffCollage: Parallel Generation of Large Content with Diffusion Modelslink
2023ICCVSat2DensitySat2Density: Faithful Density Learning from Satellite-Ground Image PairslinkGitHub
2023MMPanoDiff360-Degree Panorama Generation from Few Unregistered NFoV ImagesGitHub
2023NeurIPSMVDiffusionMVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionlinkGitHub
2023TPAMISpherical Image Generation From a Few Normal-Field-of-View Images by Considering Scene SymmetryGitHub
2024ICLRPanoDiffusionPanoDiffusion: 360-degree Panorama Outpainting via DiffusionlinkGitHub
2024CVPRControlRoom3DControlRoom3D 🤖Room Generation using Semantic Proxy Roomslink
2024CVPRSat2SceneSat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionGitHub
2024CVPRPanFusionTaming stable diffusion for text to 360◦ panorama image generationlinkGitHub
2024ECCVDreamScene360DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian SplattinglinkGitHub
2024ECCVGeospecific View Generation - Geometry-Context Aware High-resolution Ground View Inference from Satellite Viewslink
2024IJCAIFastSceneFastScene: Text-Driven Fast Indoor 3D Scene Generation via Panoramic Gaussian SplattingGitHub
2024NeurIPSDiffPanoDiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionlinkGitHub
2024TPAMIPERFPERF: Panoramic Neural Radiance Field from a Single PanoramalinkGitHub
2024TVCGDream360Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpainting
2024WACVStitchDiffusionCustomizing 360-Degree Panoramas through Text-to-Image Diffusion ModelslinkGitHub
2025ICLRCubeDiffCubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generationlink
2025SIGGRAPHLayerPano3DLayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene GenerationlinkGitHub
2025ICCVA Recipe for Generating 3D Worlds From a Single Imagelink
2025ICCVDreamCubeDreamCube: 3D Panorama Generation via Multi-plane SynchronizationlinkGitHub
2026TPAMISat2Density++Seeing through Satellite Images at Street ViewslinkGitHub
2026ICLROne2SceneOne2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single ImagelinkGitHub
2026ICLRSat3DGenSat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImagelinkGitHub
2026TVCGImmerseGenImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxieslink
2026ECCVGuidedSceneGenScene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene GenerationlinkGitHub
2026ECCVOmniXOmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D SceneslinkGitHub
2026TVCGHoloDreamerHoloDreamer: Holistic 3D Panoramic World Generation from Text DescriptionslinkGitHub
2026TIPCGGSCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene GenerationlinkGitHub
2023arXivDiffusion360Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion ModelsGitHub
2024arXivSceneDreamer360SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian SplattinglinkGitHub
2025arXivEmbodiedGenEmbodiedGen: Towards a Generative 3D World Engine for Embodied IntelligencelinkGitHub
2025arXivHunyuanWorld 1.0HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelslinkGitHub
2025arXivMatrix-3DMatrix-3D: Omnidirectional Explorable 3D World GenerationlinkGitHub
2026arXivRoamScene3DRoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware RoamingGitHub
2026arXivWorldComposerFrom Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXivPixWorldPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel SpacelinkGitHub

Iterative Generation

YearVenueAcronymPaperProjectRepo@GitHub
2019TOG3D Ken Burns Effect from a Single ImagelinkGitHub
2020CVPRSynSinSynSin: End-to-end view synthesis from a single imagelinkGitHub
2020CVPR3D Photo3D Photography Using Context-Aware Layered Depth InpaintinglinkGitHub
2020CVPRSingle-View View Synthesis with Multiplane ImageslinkGitHub
2020NeurIPSGVSGenerative View Synthesis: From Single-view Semantics to Novel-view ImageslinkGitHub
2021ICCVWorldsheetWorldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImagelinkGitHub
2021ICCVInfiniteNatureInfinite Nature: Perpetual View Generation of Natural Scenes from a Single ImagelinkGitHub
2021ICCVGFVSGeometry-free view synthesis: Transformers and no 3d priorslinkGitHub
2021ICCVPathdreamerPathdreamer: A World Model for Indoor NavigationlinkGitHub
2021ICCVPixelSynthPixelSynth: Generating a 3D-Consistent Experience from a Single ImagelinkGitHub
2022CVPRLOTRLook outside the room: Synthesizing a consistent long-term 3d scene video from a single imagelinkGitHub
2022ECCVInfiniteNature-ZeroInfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single ImageslinkGitHub
2022NeurIPSSGAMSGAM: Building a Virtual 3D World through Simultaneous Generation and MappinglinkGitHub
2023AAAISE3DSSimple and Effective Synthesis of Indoor 3D ScenesGitHub
2023CVPR3D Cinemagraphy3D Cinemagraphy from a Single ImagelinkGitHub
2023CVPRConsistent View Synthesis with Pose-Guided Diffusion Modelslink
2023ICCVDiffDreamerDiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion ModelslinkGitHub
2023ICCVText2RoomText2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelslinkGitHub
2023ICCVLong-Term Photometric Consistent Novel View Synthesis with Diffusion ModelslinkGitHub
2023MMMake-It-4DMake-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single ImageGitHub
2023NeurIPSSceneScapeSceneScape: Text-Driven Consistent Scene GenerationlinkGitHub
2023NeurIPSPanoGenPanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationlinkGitHub
2024AAAIAOG-NetAutoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image GenerationGitHub
2024CVPRWonderJourneyWonderJourney: Going from Anywhere to EverywherelinkGitHub
2024CVPR3D-SceneDreamer3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generationlink
2024ECCVPanoFreePanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-view Self-GuidancelinkGitHub
2024MMiControl3DiControl3D: An Interactive System for Controllable 3D Scene GenerationGitHub
2024NeurIPSODINFrom an Image to a Scene: Learning to Imagine the World from a Million 360° VideoslinkGitHub
2024NeurIPSCAT3DCAT3D: Create Anything in 3D with Multi-View Diffusion Modelslink
2024TVCGText2NeRFText2NeRF: Text-Driven 3D Scene Generation with Neural Radiance FieldslinkGitHub
2025TVCGLucidDreamerLucidDreamer: Domain-free Generation of 3D Gaussian Splatting SceneslinkGitHub
20253DVRealmDreamerRealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth DiffusionlinkGitHub
20253DVInvisible StitchInvisible Stitch: Generating Smooth 3D Scenes with Depth InpaintinglinkGitHub
2025AAAIBloomSceneBloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene GenerationGitHub
2025CVPRWonderWorldWonderWorld: Interactive 3D Scene Generation from a Single ImagelinkGitHub
2025CVPRArtiSceneArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediarylinkGitHub
2025ICLR3D-MOMOptimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImageslinkGitHub
2025MMScene123Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAElinkGitHub
2025ICCVWonderTurboWonderTurbo: Generating Interactive 3D World in 0.72 SecondslinkGitHub
2025ICCVSynCitySynCity: Training-Free Generation of 3D WorldslinkGitHub
2025ICCVBolt3DBolt3D: Generating 3D Scenes in Secondslink
2025ICCVScenePainterScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentlinkGitHub
2026TIPCGGSCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene GenerationlinkGitHub
2026CVPREvoSceneSelf-Evolving 3D Scene Generation from a Single ImageGitHub
2025ECCVSynCity 3000SynCity 3000: Bootstrapping Scene-Scale 3D DiffusionlinkGitHub
2026SIGGRAPH AsiaLyra 2.0Lyra 2.0: Explorable Generative 3D WorldslinkGitHub
2023arXivText2ImmersionText2Immersion: Generative Immersive Scene with 3D Gaussianslink
2024arXivOPa-MaOPa-Ma: Text Guided Mamba for 360-degree Image Out-paintingGitHub
2025arXivMeSSMeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusionlink
2025arXivCausNVSCausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesislink
2025arXivWonderZoomWonderZoom: Multi-Scale 3D World GenerationlinkGitHub
2026arXivSceneFrom3DSceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level ControllinkGitHub

Video-based Generation

Two-stage Generation

One-stage Generation

YearVenueAcronymPaperProjectRepo@GitHub
2024ICLRMagicDriveMagicDrive: Street View Generation with Diverse 3D Geometry ControllinkGitHub
2024CVPRPanaceaPanacea: Panoramic and Controllable Video Generation for Autonomous DrivinglinkGitHub
2024CVPRDrive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivinglinkGitHub
2024CVPR360DVD360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModellinkGitHub
2024ECCVDriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivinglinkGitHub
2024ECCVDrivingDiffusionDrivingDiffusion: Layout-Guided Multi-View Driving Scenarios Video Generation with Latent Diffusion ModellinkGitHub
2024ECCVWoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene GenerationGitHub
2024NeurIPSVistaVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilitylinkGitHub
2024NeurIPSDIAMONDDiffusion for World Modeling: Visual Details Matter in AtarilinkGitHub
2025AAAIDriveDreamer-2DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationlinkGitHub
2025ICLR4K4DGen4K4DGen: Panoramic 4D Generation at 4K ResolutionlinkGitHub
2025ICLRGameGen-XGameGen-X: Interactive Open-world Game Video GenerationlinkGitHub
2025ICLRGameNGenDiffusion Models Are Real-Time Game Engineslink
2025ICLRGenexGenerative World ExplorerlinkGitHub
2025ICLRGLADGlad: A Streaming Scene Generator for Autonomous Driving
2025CVPRDrivingSphereDrivingSphere: Building a High-fidelity 4D World for Closed-loop SimulationlinkGitHub
2025CVPRStreetCrafterStreetCrafter: Street View Synthesiswith Controllable Video Diffusion ModelslinkGitHub
2025CVPRDriveScapeDriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generationlink
2025CVPRUniSceneUniScene: Unified Occupancy-centric Driving Scene GenerationlinkGitHub
2025CVPRGEMGEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControllinkGitHub
2025CVPRUMGenGenerating Multimodal Driving Scenes via Next-Scene PredictionlinkGitHub
2025CVPRCAT4DCAT4D: Create Anything in 4D with Multi-View Video Diffusion Modelslink
2025CVPRWonderlandWonderland: Navigating 3D Scenes from a Single ImagelinkGitHub
2025CVPRVideoSceneVideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One SteplinkGitHub
2025CVPRScene SplatterScene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion ModellinkGitHub
2025CVPRDynamicScalerDynamicScaler: Seamless and Scalable Video Generation for Panoramic SceneslinkGitHub
2025ICMLAdaWorldAdaWorld: Learning Adaptable World Models with Latent ActionslinkGitHub
2025NatureWHAMWorld and Human Action Models towards gameplay ideationlink
2025ICCVGameFactoryGameFactory: Creating New Games with Generative Interactive VideoslinkGitHub
2025ICCVWonderPlayWonderPlay: Dynamic 3D Scene Generation from a Single Image and ActionslinkGitHub
2025ICCVMagicDrive-V2MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControllinkGitHub
2025ICCVDynamicVoyagerVoyaging into Unbounded Dynamic Scenes from a Single ViewlinkGitHub
2025ICCVInfiniCubeInfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelslinkGitHub
2025ICCVVMemVMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View MemorylinkGitHub
2025SIGGRAPH AsiaVideoFrom3DVideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion ModelslinkGitHub
2025SIGGRAPH AsiaWorldExplorerWorldExplorer: Towards Generating Fully Navigable 3D SceneslinkGitHub
2025TOGVoyagerVoyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
2026ICLRFantasyWorldFantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionlinkGitHub
2026CVPRCaptain SafariCaptain Safari: A World EnginelinkGitHub
2026CVPRPerpetualWonderPerpetualWonder: Long-Horizon Action-Conditioned 4D Scene GenerationlinkGitHub
2026CVPRWorldForgeWorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free GuidancelinkGitHub
2026CVPRWorldReelWorldReel: 4D Video Generation with Consistent Geometry and Motion Modelinglink
2026ICRAUniFutureSeeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and PerceptionlinkGitHub
2026ECCVOne4DOne4D: Unified 4D Generation and Reconstruction via Decoupled LoRA ControllinkGitHub
2026ECCVLivingWorldLivingWorld: Interactive 4D World Generation with Environmental Dynamicslink
2023arXivGAIA-1GAIA-1: A Generative World Model for Autonomous Drivinglink
2024arXivMagicDrive3DMagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street SceneslinkGitHub
2024arXivDelphiUnleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video GenerationlinkGitHub
2024arXivBEVWorldBEVWorld: A Multimodal World Model for Autonomous Driving via Unified BEV Latent SpaceGitHub
2024arXivDriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivinglinkGitHub
2024arXivDiVEDiVE: DiT-based Video Generation with Enhanced ControllinkGitHub
2024arXivDreamForgeDreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Sceneslink
2024arXivSyntheOccSyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIslinkGitHub
2024arXivCogDrivingSeeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attentionlink
2024arXivImagine360Imagine360: Immersive 360 Video Generation from Perspective AnchorlinkGitHub
2024arXivDrivingWorldDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTlinkGitHub
2024arXivViewCrafterViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View SynthesislinkGitHub
2024arXivViewExtrapolatorNovel View Extrapolation with Video Diffusion PriorslinkGitHub
2025arXivDreamDriveDreamDrive: Generative 4D Scene Modeling from Street View Imageslink
2025arXivMaskGWMMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionlinkGitHub
2025arXivSimWorldSimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World ModelGitHub
2025arXivDiST-4DDiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationlinkGitHub
2025arXivGAIA-2GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Drivinglink
2025arXivSteerXSteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric SteeringlinkGitHub
2025arXivFlexWorldFlexWorld: Progressively Expanding 3D Scenes for Flexiable-View SynthesislinkGitHub
2025arXivWORLDMEMWORLDMEM: Long-term Consistent World Simulation with MemorylinkGitHub
2025arXivHoloTimeHoloTime: Taming Video Diffusion Models for Panoramic 4D Scene GenerationlinkGitHub
2025arXivMineWorldMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftlinkGitHub
2025arXivCoGenCoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Drivinglink
2025arXivDreamlandDreamland: Controllable World Creation with Simulator and Generative Modelslink
2025arXivMatrix-GameMatrix-Game: Interactive World Foundation ModellinkGitHub
2025arXivMatrix-Game 2.0Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModellinkGitHub
2025arXivHunyuan-GameCraftHunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Conditionlink
2025arXivCoCo4DCoCo4D: Comprehensive and Complex 4D Scene GenerationlinkGitHub
2025arXivWonderFreeWonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene ExplorationlinkGitHub
2025arXiv4DVD4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generationlink
2025arXivIDCNetIDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Controllink
2025arXiv4DNeX4DNeX: Feed-Forward 4D Generative Modeling Made EasylinkGitHub
2025arXivFrom Virtual Games to Real-World PlaylinkGitHub
2025arXivEvoWorldEvoWorld: Evolving Panoramic World Generation with Explicit 3D MemoryGitHub
2025arXivMagicWorldMagicWorld: Interactive Geometry-driven Video World ExplorationlinkGitHub
2025arXivHY-World 1.5HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric ConsistencylinkGitHub

Datasets

Indoor Datasets

YearTypeSourceAcronymPaperProject
2012Indoor, NatureRealSUN360Recognizing scene viewpoint using panoramic place representationlink
2012IndoorRealNYUv2Indoor Segmentation and Support Inference From RGBD Imageslink
2015IndoorRealSunRGBDSun RGB-D: A RGB-D scene understanding benchmark suitelink
2016IndoorRealSceneNNSceneNN: A Scene Meshes Dataset with aNNotationslink
2017IndoorReal2D-3D-SJoint 2D-3D-Semantic Data for Indoor Scene Understandinglink
2017IndoorRealMatterport3DMatterport3D: Learning from RGB-D Data in Indoor Environmentslink
2017IndoorRealScanNetScanNet: Richly-annotated 3D Reconstructions of Indoor Sceneslink
2017IndoorRealLaval IndoorLearning to Predict Indoor Illumination from a Single Imagelink
2018Indoor, UrbanRealRealEstate10KStereo Magnification: Learning View Synthesis using Multiplane Imageslink
2019IndoorRealReplicaThe Replica Dataset: A Digital Replica of Indoor Spaceslink
2020IndoorReal3DSSGLearning 3D Semantic Scene Graphs from 3D Indoor Reconstructionslink
2021IndoorRealHM3DHabitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AIlink
2023IndoorRealScanNet++ScanNet++: A high-fidelity dataset of 3D indoor sceneslink
2023Indoor, Nature, UrbanRealDL3DV-10KDL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Visionlink
2012IndoorSyntheticSceneSynthExample-based synthesis of 3D object arrangementslink
2017IndoorSyntheticSUNCGSemantic Scene Completion from a Single Depth Imagelink
2020IndoorSyntheticStructured3DStructured3D: A Large Photo-realistic Dataset for Structured 3D Modelinglink
2020IndoorSyntheticHyperSimHyperSim: A photorealistic synthetic dataset for holistic indoor scene understandinglink
2021IndoorSynthetic3D-FRONT3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicslink
2021IndoorSynthetic3D-Future3D-FUTURE: 3D Furniture shape with TextURElink
2023IndoorSyntheticSG-FRONTCommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusionlink
2025IndoorSyntheticSE(3) SceneSteerable Scene Generation with Post Training and Inference-Time Searchlink
2025IndoorSyntheticSYNBUILD-3DSYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4link
2025IndoorSyntheticInternScenesInternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layoutslink
2025IndoorSyntheticSPATIALGENSPATIALGEN: Layout-guided 3D Indoor Scene Generationlink
2025IndoorSyntheticMesaTaskMesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoninglink

Natural Datasets

Urban Datasets

Tasks and Applications

Downstream Tasks

3D Scene Editing

YearVenueAcronymPaperProjectRepo@GitHub
2022CVPRStyleMeshStyleMesh: Style Transfer for Indoor 3D Scene ReconstructionslinkGitHub
2023CVPRDisCoSceneDisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene SynthesislinkGitHub
2023CVPRLEGO-NetLEGO-Net: Learning Regular Rearrangements of Objects in RoomslinkGitHub
2023CVPRLift3DLift3D: Synthesize 3D Training Data by Lifting 2D GAN to 3D Generative Radiance FieldlinkGitHub
2023CVPRText2SceneText2Scene: Text-driven Indoor Scene Stylization with Part-aware Details
2023ICRACabiNetCabiNet: Scaling Neural Collision Detection for Object Rearrangement with Procedural Scene GenerationlinkGitHub
2023MMRoomDreamerRoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture
2024CVPRSceneTexSceneTex: High-Quality Texture Synthesis for Indoor Scenes via Diffusion PriorslinkGitHub
2024CVPRControlRoom3DControlRoom3D 🤖Room Generation using Semantic Proxy Roomslink
2024ECCVStyleCityStyleCity: Large-Scale 3D Urban Scenes StylizationlinkGitHub
2024ECCVRoomTexRoomTex: Texturing Compositional Indoor Scenes via Iterative InpaintinglinkGitHub
2024ECCV3D-GOI3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editinglink
2024MMSceneExpanderSceneExpander: Real-Time Scene Synthesis for Interactive Floor Plan EditingGitHub
2024NeurIPSNeural AssetsNeural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Modelslink
2024NeurIPSDeBaRADeBaRA: Denoising-Based 3D Room Arrangement Generation
2024SIGGRAPH AsiaInstanceTexInstanceTex: Instance-level Controllable Texture Synthesis for 3D Scenes via Diffusion Priorslink
2024TVCGSceneDirectorSceneDirector: Interactive Scene Synthesis by Simultaneously Editing Multiple Objects in Real-TimeGitHub
2024VRDreamSpaceDreamSpace: Dreaming Your Room Space with Text-Driven Panoramic Texture PropagationlinkGitHub
20253DVCtrl-RoomCtrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout ConstraintslinkGitHub
2025CVPRRoomPainterRoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing
2025SIGGRAPHReStyle3DReStyle3D: Scene-Level Appearance Transfer with Semantic CorrespondenceslinkGitHub
2025NeurIPSStyl3RStyl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and StyleslinkGitHub
2026CVPRCatalyst4DCatalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagationlink
2026CVPREdit-As-ActEdit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene EditinglinkGitHub

Human-Scene Interaction

Embodied Navigation

Application Domains

Robotics

YearVenueAcronymPaperProjectRepo@GitHub
2023NeurIPSUniPiLearning Universal Policies via Text-Guided Video Generationlink
2023NeurIPSHiPCompositional Foundation Models for Hierarchical PlanninglinkGitHub
2024CoRLImagination PolicyImagination Policy: Using Generative Point Cloud Models for Learning Manipulation PolicieslinkGitHub
2024CoRLEurekaverseEurekaverse: Environment Curriculum Generation via Large Language ModelslinkGitHub
2024ICLRGR-1Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationlinkGitHub
2024ICMLRoboGenRoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationlinkGitHub
2024ICMLVLPUsing Left and Right Brains Together: Towards Vision and Language Planning
2024IROSActNeRFUncertainty-aware Active Learning of NeRF-based Object Models for Robot Manipulators using Visual and Re-orientation ActionslinkGitHub
2024NeurIPSCLOVERClosed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationGitHub
2025ICLRSlowFast-VGenSlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video GenerationlinkGitHub
2025ICLRReGenReGen: Generative Robot Simulation via Inverse Designlink
2025ICMLVideo Prediction PolicyVideo Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationslinkGitHub
2025RSSRoboVerseRoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot LearninglinkGitHub
2024arXivGR-2GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulationlink
2025arXivVideoWorldVideoWorld: Exploring Knowledge Learning from Unlabeled VideoslinkGitHub
2025arXivCosmos-Transfer1Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControllinkGitHub
2025arXivTesserActTesserAct: Learning 4D Embodied World ModelslinkGitHub
2025arXivMarketGenMarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environmentslink
2025arXivTabletopGenTabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single ImagelinkGitHub
2026arXivWorldComposerFrom Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXivSimFoundrySimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluationlink
2026arXivRoboSnapRoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXiv4DSynth4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied SimulationlinkGitHub

Autonomous Driving

YearVenueAcronymPaperProjectRepo@GitHub
2024CVPRCam4DOccCam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsGitHub
2024CVPRDrive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivinglinkGitHub
2024ECCVDriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivinglinkGitHub
2024ECCVOccWorldOccWorld: Learning a 3D Occupancy World Model for Autonomous DrivinglinkGitHub
2024ECCVWoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene GenerationGitHub
2024ICLRMagicDriveMagicDrive: Street View Generation with Diverse 3D Geometry ControllinkGitHub
2024NeurIPSVistaVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilitylinkGitHub
2025AAAIDrive-OccWorldDriving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivinglinkGitHub
2025CVPRDrivingSphereDrivingSphere: Building a High-fidelity 4D World for Closed-loop SimulationlinkGitHub
2025ICLRGLADGlad: A Streaming Scene Generator for Autonomous Driving
2025NeurIPSOrbisOrbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelslinkGitHub
2026CVPRDriveLaWDriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldGitHub
2023arXivGAIA-1GAIA-1: A Generative World Model for Autonomous Drivinglink
2024arXivOccSoraOccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivinglinkGitHub
2024arXivDelphiUnleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video GenerationlinkGitHub
2024arXivDriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivinglinkGitHub
2024arXivDiVEDiVE: DiT-based Video Generation with Enhanced ControllinkGitHub
2024arXivDreamForgeDreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Sceneslink
2024arXivDrivingWorldDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTlinkGitHub
2025arXivDreamDriveDreamDrive: Generative 4D Scene Modeling from Street View Imageslink
2025arXivCosmos-Transfer1Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControllinkGitHub
2025arXivGenieDriveGenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationlinkGitHub
2026arXivVectorWorldVectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector GraphsGitHub
2026arXivOccSimOccSim: Multi-kilometer Simulation with Long-horizon Occupancy World ModelslinkGitHub
3d-aigc
3d-scene-generation
awesome-list
paper-review

Contributors

wenbc21

46 commits

hzxie

25 commits

Zhu-Liyuan

2 commits

Cylrx

1 commits

hzxie/Awesome-3D-Scene-Generation

A curated list of awesome 3D scene generation papers. (IJCV 2026)

1,096

79 commits

updated Sep 18, 2026

See the code

README

Awesome Logo Counter arXiv YouTube PR's Welcome

Overview

This repository collects summaries of over 300 recent studies on 3D scene generation, along with the downstream applications, and will be continuously updated.

If you have suggestions for new resources, improvements to methodologies, or corrections for broken links, please don't hesitate to open an issue or submit a pull request. Contributions of all kinds are welcome and greatly appreciated.

3D-Scene-Generation-Teaser

Table of Contents

Methods: A Hierarchical Taxonomy

Procedural Generation

Rule-based Generation

YearVenueAcronymPaperProjectRepo@GitHub
1988SIGGRAPHTerrain simulation using a model of stream erosion
1989SIGGRAPHThe synthesis and rendering of eroded fractal terrains
1993Graphics InterfaceA fractal model of mountains and riverslink
1998SIGGRAPHRealistic modeling and rendering of plant ecosystemslink
2001SIGGRAPHCityEngineProcedural modeling of cities
2005VRSTModeling Landscapes with Ridges and Rivers
2006TOGProcedural modeling of buildings
2007GDTWCitygenCitygen: An Interactive System for Procedural City Generationlink
2007I3DExample-based model synthesislinkGitHub
2007TVCGTerrain Synthesis from Digital Elevation Modelslink
2008CGFReal-Time Rendering and Editing of Vector-based Terrainslink
2008TOGContinuous model synthesislink
2008TOGInteractive Procedural Street Modelinglink
2009CGFArches: a Framework for Modeling Complex Terrains
2009CGFInteractive Geometric Simulation of 4D Cities
2009TOGInteractive design of urban spaces using geometrical and behavioral modelingyoutube
2010CGFProcedural Generation of Roads
2011CGFInteractive Modeling of City Layouts using Layers of Procedural Content
2011SI3DUrban Ecosystem Design
2011TOGMetropolis procedural modelinglink
2012CGFProcedural Generation of Parcels in Urban Modeling
2012TOGInverse design of urban procedural modelslink
2013TOGTerrain Generation Using Procedural Models Based on Hydrologyyoutube
2013TOGUrban PatternUrban Pattern: Layout Design by Hierarchical Domain Splittinglink
2015TOGWorldBrushWorldBrush: Interactive Example-Based Synthesis of Procedural Virtual Worldslink
2016CGFExample-Driven Procedural Urban Roads
20163DVProceduralization for Editing 3D Architectural Models
2016TOGInteractive Sketching of Urban Procedural Modelslink
2017TOGAuthoring landscapes by combining ecosystem and terrain erosion simulation
2017TOGFast Weather Simulation for Inverse Procedural Design of 3D Urban ModelslinkGitHub
2017TOGInteractive Example-Based Terrain Authoring with Conditional Generative Adversarial Networkslink
2019TOGSynthetic Silviculture: Multi-scale Modeling of Plant Ecosystemslink
2021TOGAuthoring Consistent Landscapes with Flora and Faunalink
2022TOGEcoclimatesEcoclimates: Climate-Response Modeling of Vegetationlink
2022TOGProcedural Urban Forestrylink
2023CVPRInfinigenInfinite Photorealistic Worlds using Procedural GenerationlinkGitHub
2023TOGForming Terrains by Glacial Erosionlinklink
2023TOGLarge-scale terrain authoring through interactive erosion simulationyoutubeGitHub
2023TOGAuthoring and Simulating Meandering RiverslinkGitHub
2025CVPRWProc-GSProc-GS: Procedural Building Generation for City Assembly with 3D GaussianslinkGitHub
2025CEUSVoxCityVoxCity: A seamless framework for open geospatial data integration, grid-based semantic 3D city model generation, and urban environment simulationlinkGitHub

Optimization-based Generation

LLM-based Generation

YearVenueAcronymPaperProjectRepo@GitHub
2023NeurIPSLayoutGPTLayoutGPT: Compositional Visual Planning and Generation with Large Language ModelslinkGitHub
2024CVPRGraphDreamerGraphDreamer: Compositional 3D Scene Synthesis from Scene GraphslinkGitHub
2024ECCVAnyHomeAnyHome: Open-Vocabulary Generation of Structured and Textured 3D HomeslinkGitHub
2024ECCVSceneTellerSceneTeller: Language-to-3D Scene GenerationlinkGitHub
2024ECCVI-DesignI-Design: Personalized LLM Interior DesignerlinkGitHub
2024ICMLSceneCraftSceneCraft: An LLM Agent for Synthesizing 3D Scenes as Blender Code
2024MMControllable Procedural Generation of LandscapesGitHub
2024SIGGRAPH AsiaDISceneDIScene: Object Decoupling and Interaction Modeling for Complex Scene Generationlink
20253DV3D-GPT3D-GPT: Procedural 3D Modeling with Large Language ModelslinkGitHub
2025AAAISceneXSceneX: Procedural Controllable Large-scale Scene GenerationlinkGitHub
2025AAAIHierarchically-Structured Open-Vocabulary Indoor Scene Synthesis with Pre-trained Large Language ModelGitHub
2025CVPRGlobal-Local Tree Search in VLMs for 3D Indoor Scene GenerationGitHub
2025CVPRLayoutVLMLayoutVLM: Differentiable Optimization of 3D Layout via Vision-Language ModelslinkGitHub
2025CVPRThe Scene LanguageThe Scene Language: Representing Scenes with Programs, Words, and EmbeddingslinkGitHub
2025ACL FindingsUnrealLLMUnrealLLM: Towards Highly Controllable and Interactable 3D Scene Generation by LLM-powered Procedural Content Generation
2025UISTEchoLadderEchoLadder: Progressive AI-Assisted Design of Immersive VR Scenes
2025NeurIPSSceneWeaverSceneWeaver: All-in-One 3D Scene Synthesis with an Extensible and Self-Reflective AgentlinkGitHub
2025NeurIPSMesaTaskMesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial ReasoninglinkGitHub
2025NeurIPSOptiSceneOptiScene: LLM-driven Indoor Scene Layout Generation via Scaled Human-aligned Data Synthesis and Multi-Stage Preference OptimizationlinkGitHub
2025NeurIPSDirectLayoutDirect Numerical Layout Generation for 3D Indoor Scene Synthesis via Spatial ReasoninglinkGitHub
2025Nature Computational ScienceUrban planning in the era of large language models
2025SIGGRAPHBuildingBlockBuildingBlockGitHub
2026ICLRScenethesisScenethesis: A Language and Vision Agentic Framework for 3D Scene Generationlink
2026ICLRReSpaceReSpace: Text-Driven 3D Scene Synthesis and Editing with Preference AlignmentlinkGitHub
2026ICMLSceneSmithSceneSmith: Agentic Generation of Simulation-Ready Indoor SceneslinkGitHub
20263DV3D-Generalist3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D WorldslinkGitHub
2026AAAILandCraftLandCraft: Designing the Structured 3D Landscapes via Text GuidanceGitHub
2026AAAIReason-3DText-to-Scene with Large Reasoning ModelslinkGitHub
2026CVPRRAISECityRAISECity: A Multimodal Agent Framework for Reality-Aligned 3D World Generation at City-ScaleGitHub
2026CVPRHOG-LayoutHOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
2026CVPRLaviGenRepurposing 3D Generative Model for Autoregressive Layout GenerationlinkGitHub
2026CVPRYo'CityYo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
2026CVPRMajutsuCityMajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and LayoutslinkGitHub
2026CVPRGardenDesignerGardenDesigner: Encoding Aesthetic Principles into Jiangnan Garden Construction via a Chain of Agentslink
2026CVPRSAGESAGE: Scalable Agentic 3D Scene Generation for Embodied AIlinkGitHub
2026ECCVSceneOrchestraSceneOrchestra: Efficient Agentic 3D Scene Synthesis via Full Tool-Call Trajectory Generation
2026ECCVNaLANaLA: A 3D Native LLM Layout Agent for High-quality 3D Scene GenerationlinkGitHub
2026ACL FindingsSceneLMSceneLM: 3D-Aware Language Models for Editable 3D Scene Synthesis
2026ICMLR3LR$^3$L: Reasoning 3D Layouts from Relative Spatial RelationsGitHub
2026ICMLCode2WorldsCode2Worlds: Empowering Coding LLMs for 4D World GenerationlinkGitHub
2024arXivCityCraftCityCraft: A Real Crafter for 3D City GenerationGitHub
2024arXivCityXCityX: Controllable Procedural Content Generation for Unbounded 3D CitieslinkGitHub
2024arXivGraphCanvas3DGraph Canvas for Controllable 3D Scene GenerationGitHub
2024arXivUrbanWorldUrbanWorld: An Urban World Model for 3D City GenerationGitHub
2025arXivCubeCube: A Roblox View of 3D IntelligencelinkGitHub
2025arXivAgentic 3D Scene Generation with Spatially Contextualized VLMslink
2025arXivHLGHLG: Comprehensive 3D Room Construction via Hierarchical Layout GenerationGitHub
2025arXivLatticeWorldLatticeWorld: A Multimodal Large Language Model-Empowered Framework for Interactive Complex World Generationyoutube
2025arXivCausalStructCausal Reasoning Elicits Controllable 3D Scene GenerationlinkGitHub
2025arXivDisCo-LayoutDisCo-Layout: Disentangling and Coordinating Semantic and Physical Refinement in a Multi-Agent Framework for 3D Indoor Layout Synthesislink
2025arXivWorldGenWorldGen: From Text to Traversable and Interactive 3D Worldslink
2025arXivMarketGenMarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environmentslink
2026arXivSceneFoundrySceneFoundry: Generating Interactive Infinite 3D WorldslinkGitHub
2026arXivMANSIONMANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon taskslinkGitHub
2026arXivSceneAssistantSceneAssistant: A Visual Feedback Agent for Open-Vocabulary 3D Scene GenerationGitHub
2026arXivSceneCodeSceneCode: Executable World Programs for Editable Indoor Scenes with Articulated ObjectslinkGitHub
2026arXivCode-as-RoomCode-as-Room: Generating 3D Rooms from Top-Down View Images via Agentic Code SynthesislinkGitHub
2026arXivSpatialGrammarSpatialGrammar: A Domain-Specific Language for LLM-Based 3D Indoor Scene Generationlink
2026arXivClosing the Loop: Unified 3D Scene Generation and Immersive Interaction via LLM-RL Couplinglink
2026arXivGlobal-Local Monte Carlo Tree Search in Vision-Language Models for Text-to-3D Indoor Scene GenerationGitHub
2026arXivSceneConductorSceneConductor: 3D Scene Generation from a Single Image with Multi-Agent OrchestrationlinkGitHub
2026arXivSceneReVisSceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RLlinkGitHub
2026arXivWorldClawWorldClaw: Agentic 3D Open-World Generation at ScalelinkGitHub
2026arXivScenePilotScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generationlink
2026arXivSceneMosaicSceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout EvolutionlinkGitHub
2026arXiv4DSynth4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied SimulationlinkGitHub

Neural-3D Generation

Scene Parameters

YearVenueAcronymPaperProjectRepo@GitHub
2018SIGGRAPHDeepSynthDeep Convolutional Priors for Indoor Scene SynthesislinkGitHub
2019CVPRFastSynthFast and Flexible Indoor Scene Synthesis via Deep Convolutional Generative ModelsGitHub
2020SIGGRAPHDeep Generative Modeling for Scene Synthesis via Hybrid Representations
20213DVSceneFormerSceneFormer: Indoor Scene Generation with TransformerslinkGitHub
2021ICCVSync2GenScene Synthesis via Uncertainty-Driven Attribute SynchronizationGitHub
2021NeurIPSATISSATISS: Autoregressive Transformers for Indoor Scene SynthesislinkGitHub
2022ECCVPose2RoomPose2Room: Understanding 3D Scenes from Human ActivitieslinkGitHub
2022SIGGRAPH AsiaSUMMONScene Synthesis from Human MotionlinkGitHub
2023CVPRLearning 3D Scene Priors with 2D SupervisionlinkGitHub
2023CVPRMIMEMIME: Human-Aware 3D Scene GenerationlinkGitHub
2023SIGGRAPHCOFSCOFS: COntrollable Furniture layout Synthesis
2023NeurIPSLanguage-driven Scene Synthesis using Multi-conditional Diffusion ModellinkGitHub
20243DVRoomDesignerRoomDesigner: Encoding Anchor-latents for Style-consistent and Shape-compatible Indoor Scene GenerationGitHub
2024CVPRDiffuSceneDiffuScene: Denoising Diffusion Models for Generative Indoor Scene SynthesislinkGitHub
2024CVPRSceneWiz3DSceneWiz3D: Towards Text-guided 3D Scene CompositionlinkGitHub
2024CVPRPhyScenePhyScene: Physically Interactable 3D Scene Synthesis for Embodied AIlinkGitHub
2024ECCVDreamSceneDreamScene: 3D Gaussian-Based Text-to-3D Scene Generation via Formation Pattern SamplinglinkGitHub
2024ICMLGALA3DGALA3D: Towards Text-to-3D Complex Scene Generation via Layout-guided Generative Gaussian SplattinglinkGitHub
2024ICMLDisentangled 3D Scene Generation with Layout Learninglink
2024MMRelSceneRelScene: A Benchmark and baseline for Spatial Relations in text-driven 3D Scene Generation
2024NeurIPSDeBaRADeBaRA: Denoising-Based 3D Room Arrangement Generation
2024SIGGRAPHINFERACTPhysics-based Scene Layout Generation From Human Motionlink
20253DVCtrl-RoomCtrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout ConstraintslinkGitHub
2025CVPRSceneFactorSceneFactor: Factored Latent 3D Diffusion for Controllable 3D Scene GenerationlinkGitHub
2025CVPRCASAGPTCASAGPT: Cuboid Arrangement and Scene Assembly for Interior DesignGitHub
2025CoRLSteerable Scene Generation with Post Training and Inference-Time SearchlinkGitHub
2025NeurIPSFactoredScenesFrom Programs to Poses: Factored Real-World Scene Generation via Learned Program LibrarieslinkGitHub

Scene Graph

YearVenueAcronymPaperProjectRepo@GitHub
2014EMNLPLearning Spatial Knowledge for Text to 3D Scene Generation
2016CGFLearning 3D Scene Synthesis from Annotated RGB-D Images
2017TOGAdaptive synthesis of indoor scenes via activity-associated object relation graphsyoutube
2018TOGLanguage-Driven Synthesis of 3D Scenes from Scene Databaseslink
2019ICCVMeta-SimMeta-Sim: Learning to Generate Synthetic DatasetslinkGitHub
2019SIGGRAPHGRAINSGRAINS: Generative Recursive Autoencoders for INdoor SceneslinkGitHub
2019SIGGRAPHPlanITPlanIT: Planning and Instantiating Indoor Scenes with Relation Graph and Spatial Prior NetworksGitHub
2020CVPR3D-SLNEnd-to-End Optimization of Scene LayoutlinkGitHub
2020ECCVMeta-Sim 2Meta-Sim 2 Unsupervised Learning of Scene Structure for Synthetic Data GenerationlinkGitHub
2021ICCVGraph-to-3DGraph-to-3D: End-to-End Generation and Manipulation of 3D Scenes Using Scene GraphslinkGitHub
2023NeurIPSCommonScenesCommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph DiffusionlinkGitHub
2023TPAMISceneHGNSceneHGN: Hierarchical Graph Networks for 3D Indoor Scene Generation With Fine-Grained GeometrylinkGitHub
2024ECCVSEKExternal Knowledge Enhanced 3D Scene Generation from Sketch
2024ECCVForest2SeqForest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
2024ECCVEchoSceneEchoScene: Indoor Scene Generation via Information Echo over Scene Graph DiffusionlinkGitHub
2024ICLRInstructSceneInstructScene: Instruction-Driven 3D Indoor Scene Synthesis with Semantic Graph PriorlinkGitHub
2025AAAIMMGDreamerMMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene GenerationlinkGitHub
2025MMHiSceneHiScene: Creating Hierarchical 3D Scenes with Isometric View Generationlink
2025CVPRFreeSceneFreeScene: Mixed Graph Diffusion for 3D Scene Synthesis from Free PromptslinkGitHub
2025ICCVControllable 3D Outdoor Scene Generation via Scene GraphsGitHub
2025TOGImaginariumImaginarium: Vision-guided High-Quality 3D Scene Layout GenerationlinkGitHub
2026TOGCasLayoutCasLayout: Cascaded 3D Layout Diffusion for Indoor Scene Synthesis with Implicit Relation Modeling
2026TVCGSceneLinkerSceneLinker: Compositional 3D Scene Generation via Semantic Scene Graph from RGB Sequenceslink

Semantic Layout

YearVenueAcronymPaperProjectRepo@GitHub
2021ICCVSGSDIIndoor Scene Generation from a Collection of Semantic-Segmented Depth ImageslinkGitHub
2021ICCVGANcraftGANcraft: Unsupervised 3D Neural Rendering of Minecraft WorldslinkGitHub
2023CVPRDisCoSceneDisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene SynthesislinkGitHub
2023ICCVInfiniCityInfiniCity: Infinite-Scale City Synthesislink
2023ICCVCC3DCC3D: Layout-Conditioned Generation of Compositional 3D SceneslinkGitHub
2023ICCVSet-the-SceneSet-the-Scene: Global-Local Training for Generating Controllable NeRF SceneslinkGitHub
2023ICCVUrbanGIRAFFEUrbanGIRAFFE: Representing Urban Scenes as Compositional Generative Neural Feature FieldslinkGitHub
2023TPAMISceneDreamerSceneDreamer: Unbounded 3D Scene Generation From 2D Image CollectionslinkGitHub
20243DVComp3DCompositional 3D Scene Generation using Locally Conditioned Diffusionlink
2024CVPRCityDreamerCityDreamer: Compositional Generative Model of Unbounded 3D CitieslinkGitHub
2024CVPRBerfSceneBerfScene: Bev-conditioned Equivariant Radiance Fields for Infinite 3D Scene GenerationlinkGitHub
2024NeurIPSSceneCraftSceneCraft: Layout-Guided 3D Scene GenerationlinkGitHub
2024SIGGRAPHBlockFusionBlockFusion: Expandable 3D Scene Generation Using Latent Tri-plane ExtrapolationlinkGitHub
2024SIGGRAPH AsiaFrankensteinFrankenstein: Generating Semantic-Compositional 3D Scenes in One Tri-PlanelinkGitHub
2025CVPRGaussianCityGenerative Gaussian Splatting for Unbounded 3D City GenerationlinkGitHub
2025ICLRLayout-your-3DLayout-your-3D: Controllable and Precise 3D Generation with 2D BlueprintlinkGitHub
2025TPAMICityDreamer4DCityDreamer4D: Compositional Generative Model of Unbounded 4D CitieslinkGitHub
2025TPAMIUrbanGenUrbanGen: Urban Generation with Compositional and Controllable Neural Fieldslink
2025ICCVSat2CitySat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent DiffusionlinkGitHub
2025NeurIPSX-SceneX-Scene: Large-Scale Driving Scene Generation with High Fidelity and Flexible ControllabilitylinkGitHub
2026AAAIEarthCrafterEarthCrafter: Scalable 3D Earth Generation via Dual-Sparse Latent DiffusionlinkGitHub
20263DVSPATIALGENSPATIALGEN: Layout-guided 3D Indoor Scene GenerationlinkGitHub
2026CVPRPrITTIPrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic SceneslinkGitHub
20263DVSemLayoutDiffSemLayoutDiff: Semantic Layout Generation with Diffusion Model for Indoor Scene Synthesislink
2023arXivCompoNeRFCompoNeRF: Text-guided Multi-object Compositional NeRF with Editable 3D Scene LayoutlinkGitHub
2024arXivUrban ArchitectUrban Architect: Steerable 3D Urban Scene Generation with Layout PriorlinkGitHub
2025arXivLayout2SceneLayout2Scene: 3D Semantic Layout Guided Scene Generation via Geometry and Appearance Diffusion PriorslinkGitHub

Implicit Layout

YearVenueAcronymPaperProjectRepo@GitHub
2021CVPRGIRAFFEGIRAFFE: Representing Scenes as Compositional Generative Neural Feature FieldsGitHub
2021ICCVGSNUnconstrained Scene Generation With Locally Conditioned Radiance FieldslinkGitHub
2021ICMLNeRF-VAENeRF-VAE: A geometry aware 3d scene generative model
2022NeurIPSGAUDIGAUDI: A Neural Architect for Immersive 3D Scene GenerationGitHub
2023CVPRPersistent NaturePersistent Nature: A generative model of unbounded 3D worldslinkGitHub
2023CVPRNeuralField-LDMNeuralField-LDM: Scene Generation with Hierarchical Latent Diffusion Modelslink
2024CVPRDiffInDSceneDiffInDScene: Diffusion-based High-Quality 3D Indoor Scene GenerationlinkGitHub
2024CVPRXCubeXCube: Large-Scale 3D Generative Modeling using Sparse Voxel HierarchieslinkGitHub
2024CVPRSemCitySemCity: Semantic Scene Generation with Triplane DiffusionlinkGitHub
2024ECCVPDDPyramid Diffusion for Fine 3D Large Scene GenerationlinkGitHub
2024NeurIPSDirector3DDirector3D: Real-world Camera Trajectory and 3D Scene Generation from TextlinkGitHub
2025CVPRLT3SDLT3SD: Latent Trees for 3D Scene DiffusionlinkGitHub
2025CVPRSplatFlowSplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting SynthesislinkGitHub
2025CVPRPrometheusPrometheus: 3D-Aware Latent Diffusion Models for Feed-Forward Text-to-3D Scene GenerationlinkGitHub
2025ICLRDynamicCityDynamicCity: Large-Scale Occupancy Generation from Dynamic SceneslinkGitHub
2025ICCVNuiSceneNuiScene: Exploring Efficient Generation of Unbounded Outdoor SceneslinkGitHub
2025ICCVVideoRFSplatVideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint ModelinglinkGitHub
2026ICLRFlashWorldFlashWorld: High-quality 3D Scene Generation within SecondslinkGitHub
2026ICLRUniUGGUniUGG: Unified 3D Understanding and Generation via Geometric-Semantic EncodinglinkGitHub
2026ICMLPERSISTBeyond Pixel Histories: World Models with Persistent 3D StatelinkGitHub
2026AAAIWorldGrowWorldGrow: Generating Infinite 3D WorldlinkGitHub
2026AAAILSD-3DLSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Groundinglink
20263DVSceneGenSceneGen: Single-Image 3D Scene Generation in One Feedforward Passlink
2026CVMExCellGenExCellGen: Fast, Controllable, Photorealistic 3D Scene Generation from a Single Real-World ExemplarlinkGitHub
2026CVPRDiff4SplatDiff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction ModelslinkGitHub
2026CVPRScenDiScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generationlink
2026ECCVGaussianGPTGaussianGPT: Towards Autoregressive 3D Gaussian Scene GenerationlinkGitHub
2026ECCVSPAR3SSparse auto-regressive modeling for scene generation from multi-view images
2026SIGGRAPH AsiaSpatialCrafterSpatialCrafter: Single Image World Modeling with Generative 3D ProxieslinkGitHub
2026SIGGRAPHInfiniteDiffusionInfiniteDiffusion: Bridging Learned Fidelity and Procedural Utility for Open-World Terrain GenerationlinkGitHub
2023arXivDiffusion Probabilistic Models for Scene-Scale 3D Categorical DatalinkGitHub
2025arXivTerraTerra: Explorable Native 3D World Model with Point LatentslinkGitHub
2026arXivABot-Earth-0.5ABot-Earth-0.5: Generative 3D Earth ModellinkGitHub

Image-based Generation

Holistic Generation

YearVenueAcronymPaperProjectRepo@GitHub
2019ICIP360-Degree Image Completion by Two-Stage Conditional Gans
2020CVPRSat2GroundGeometry-Aware Satellite-to-Ground Image Synthesis for Urban AreasGitHub
2020WACV360 Panorama Synthesis from a Sparse Set of Images with Unknown Field of View
2021AAAISIG-SSSpherical Image Generation from a Single Image by Considering Scene SymmetryGitHub
2021CVPREnvMapNetHDR Environment Map Estimation for Real-Time Augmented RealitylinkGitHub
2021ICCVSat2vidSat2vid: Street-view panoramic video synthesis from a single satellite image
20223DVImmerseGANGuided Co-Modulated GAN for 360° Field of View Extrapolationlink
2022CVPROmniDreamerDiverse Plausible 360-Degree Image Outpainting for Efficient 3DCG Background CreationlinkGitHub
2022ECCVBIPSBIPS: Bi-modal Indoor Panorama Synthesis via Residual Depth-aided Adversarial LearningGitHub
2022SIGGRAPH AsiaText2LightText2Light: Zero-Shot Text-Driven HDR Panorama GenerationlinkGitHub
2022TMMPanoGANCross-View Panorama Image SynthesisGitHub
2022TPAMISat2StrGeometry-Guided Street-View Panorama Synthesis from Satellite ImageryGitHub
2023CVPRDiffCollageDiffCollage: Parallel Generation of Large Content with Diffusion Modelslink
2023ICCVSat2DensitySat2Density: Faithful Density Learning from Satellite-Ground Image PairslinkGitHub
2023MMPanoDiff360-Degree Panorama Generation from Few Unregistered NFoV ImagesGitHub
2023NeurIPSMVDiffusionMVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionlinkGitHub
2023TPAMISpherical Image Generation From a Few Normal-Field-of-View Images by Considering Scene SymmetryGitHub
2024ICLRPanoDiffusionPanoDiffusion: 360-degree Panorama Outpainting via DiffusionlinkGitHub
2024CVPRControlRoom3DControlRoom3D 🤖Room Generation using Semantic Proxy Roomslink
2024CVPRSat2SceneSat2Scene: 3D Urban Scene Generation from Satellite Images with DiffusionGitHub
2024CVPRPanFusionTaming stable diffusion for text to 360◦ panorama image generationlinkGitHub
2024ECCVDreamScene360DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian SplattinglinkGitHub
2024ECCVGeospecific View Generation - Geometry-Context Aware High-resolution Ground View Inference from Satellite Viewslink
2024IJCAIFastSceneFastScene: Text-Driven Fast Indoor 3D Scene Generation via Panoramic Gaussian SplattingGitHub
2024NeurIPSDiffPanoDiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionlinkGitHub
2024TPAMIPERFPERF: Panoramic Neural Radiance Field from a Single PanoramalinkGitHub
2024TVCGDream360Dream360: Diverse and Immersive Outdoor Virtual Scene Creation via Transformer-Based 360° Image Outpainting
2024WACVStitchDiffusionCustomizing 360-Degree Panoramas through Text-to-Image Diffusion ModelslinkGitHub
2025ICLRCubeDiffCubeDiff: Repurposing Diffusion-Based Image Models for Panorama Generationlink
2025SIGGRAPHLayerPano3DLayerPano3D: Layered 3D Panorama for Hyper-Immersive Scene GenerationlinkGitHub
2025ICCVA Recipe for Generating 3D Worlds From a Single Imagelink
2025ICCVDreamCubeDreamCube: 3D Panorama Generation via Multi-plane SynchronizationlinkGitHub
2026TPAMISat2Density++Seeing through Satellite Images at Street ViewslinkGitHub
2026ICLROne2SceneOne2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single ImagelinkGitHub
2026ICLRSat3DGenSat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite ImagelinkGitHub
2026TVCGImmerseGenImmerseGen: Agent-Guided Immersive World Generation with Alpha-Textured Proxieslink
2026ECCVGuidedSceneGenScene Generation at Absolute Scale: Utilizing Semantic and Geometric Guidance From Text for Accurate and Interpretable 3D Indoor Scene GenerationlinkGitHub
2026ECCVOmniXOmniX: From Unified Panoramic Generation and Perception to Graphics-Ready 3D SceneslinkGitHub
2026TVCGHoloDreamerHoloDreamer: Holistic 3D Panoramic World Generation from Text DescriptionslinkGitHub
2026TIPCGGSCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene GenerationlinkGitHub
2023arXivDiffusion360Diffusion360: Seamless 360 Degree Panoramic Image Generation based on Diffusion ModelsGitHub
2024arXivSceneDreamer360SceneDreamer360: Text-Driven 3D-Consistent Scene Generation with Panoramic Gaussian SplattinglinkGitHub
2025arXivEmbodiedGenEmbodiedGen: Towards a Generative 3D World Engine for Embodied IntelligencelinkGitHub
2025arXivHunyuanWorld 1.0HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelslinkGitHub
2025arXivMatrix-3DMatrix-3D: Omnidirectional Explorable 3D World GenerationlinkGitHub
2026arXivRoamScene3DRoamScene3D: Immersive Text-to-3D Scene Generation via Adaptive Object-aware RoamingGitHub
2026arXivWorldComposerFrom Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXivPixWorldPixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel SpacelinkGitHub

Iterative Generation

YearVenueAcronymPaperProjectRepo@GitHub
2019TOG3D Ken Burns Effect from a Single ImagelinkGitHub
2020CVPRSynSinSynSin: End-to-end view synthesis from a single imagelinkGitHub
2020CVPR3D Photo3D Photography Using Context-Aware Layered Depth InpaintinglinkGitHub
2020CVPRSingle-View View Synthesis with Multiplane ImageslinkGitHub
2020NeurIPSGVSGenerative View Synthesis: From Single-view Semantics to Novel-view ImageslinkGitHub
2021ICCVWorldsheetWorldsheet: Wrapping the World in a 3D Sheet for View Synthesis from a Single ImagelinkGitHub
2021ICCVInfiniteNatureInfinite Nature: Perpetual View Generation of Natural Scenes from a Single ImagelinkGitHub
2021ICCVGFVSGeometry-free view synthesis: Transformers and no 3d priorslinkGitHub
2021ICCVPathdreamerPathdreamer: A World Model for Indoor NavigationlinkGitHub
2021ICCVPixelSynthPixelSynth: Generating a 3D-Consistent Experience from a Single ImagelinkGitHub
2022CVPRLOTRLook outside the room: Synthesizing a consistent long-term 3d scene video from a single imagelinkGitHub
2022ECCVInfiniteNature-ZeroInfiniteNature-Zero: Learning Perpetual View Generation of Natural Scenes from Single ImageslinkGitHub
2022NeurIPSSGAMSGAM: Building a Virtual 3D World through Simultaneous Generation and MappinglinkGitHub
2023AAAISE3DSSimple and Effective Synthesis of Indoor 3D ScenesGitHub
2023CVPR3D Cinemagraphy3D Cinemagraphy from a Single ImagelinkGitHub
2023CVPRConsistent View Synthesis with Pose-Guided Diffusion Modelslink
2023ICCVDiffDreamerDiffDreamer: Towards Consistent Unsupervised Single-view Scene Extrapolation with Conditional Diffusion ModelslinkGitHub
2023ICCVText2RoomText2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelslinkGitHub
2023ICCVLong-Term Photometric Consistent Novel View Synthesis with Diffusion ModelslinkGitHub
2023MMMake-It-4DMake-It-4D: Synthesizing a Consistent Long-Term Dynamic Scene Video from a Single ImageGitHub
2023NeurIPSSceneScapeSceneScape: Text-Driven Consistent Scene GenerationlinkGitHub
2023NeurIPSPanoGenPanoGen: Text-Conditioned Panoramic Environment Generation for Vision-and-Language NavigationlinkGitHub
2024AAAIAOG-NetAutoregressive Omni-Aware Outpainting for Open-Vocabulary 360-Degree Image GenerationGitHub
2024CVPRWonderJourneyWonderJourney: Going from Anywhere to EverywherelinkGitHub
2024CVPR3D-SceneDreamer3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generationlink
2024ECCVPanoFreePanoFree: Tuning-Free Holistic Multi-view Image Generation with Cross-view Self-GuidancelinkGitHub
2024MMiControl3DiControl3D: An Interactive System for Controllable 3D Scene GenerationGitHub
2024NeurIPSODINFrom an Image to a Scene: Learning to Imagine the World from a Million 360° VideoslinkGitHub
2024NeurIPSCAT3DCAT3D: Create Anything in 3D with Multi-View Diffusion Modelslink
2024TVCGText2NeRFText2NeRF: Text-Driven 3D Scene Generation with Neural Radiance FieldslinkGitHub
2025TVCGLucidDreamerLucidDreamer: Domain-free Generation of 3D Gaussian Splatting SceneslinkGitHub
20253DVRealmDreamerRealmDreamer: Text-Driven 3D Scene Generation with Inpainting and Depth DiffusionlinkGitHub
20253DVInvisible StitchInvisible Stitch: Generating Smooth 3D Scenes with Depth InpaintinglinkGitHub
2025AAAIBloomSceneBloomScene: Lightweight Structured 3D Gaussian Splatting for Crossmodal Scene GenerationGitHub
2025CVPRWonderWorldWonderWorld: Interactive 3D Scene Generation from a Single ImagelinkGitHub
2025CVPRArtiSceneArtiScene: Language-Driven Artistic 3D Scene Generation Through Image IntermediarylinkGitHub
2025ICLR3D-MOMOptimizing 4D Gaussians for Dynamic Scene Video from Single Landscape ImageslinkGitHub
2025MMScene123Scene123: One Prompt to 3D Scene Generation via Video-Assisted and Consistency-Enhanced MAElinkGitHub
2025ICCVWonderTurboWonderTurbo: Generating Interactive 3D World in 0.72 SecondslinkGitHub
2025ICCVSynCitySynCity: Training-Free Generation of 3D WorldslinkGitHub
2025ICCVBolt3DBolt3D: Generating 3D Scenes in Secondslink
2025ICCVScenePainterScenePainter: Semantically Consistent Perpetual 3D Scene Generation with Concept Relation AlignmentlinkGitHub
2026TIPCGGSCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-Centric 3D Scene GenerationlinkGitHub
2026CVPREvoSceneSelf-Evolving 3D Scene Generation from a Single ImageGitHub
2025ECCVSynCity 3000SynCity 3000: Bootstrapping Scene-Scale 3D DiffusionlinkGitHub
2026SIGGRAPH AsiaLyra 2.0Lyra 2.0: Explorable Generative 3D WorldslinkGitHub
2023arXivText2ImmersionText2Immersion: Generative Immersive Scene with 3D Gaussianslink
2024arXivOPa-MaOPa-Ma: Text Guided Mamba for 360-degree Image Out-paintingGitHub
2025arXivMeSSMeSS: City Mesh-Guided Outdoor Scene Generation with Cross-View Consistent Diffusionlink
2025arXivCausNVSCausNVS: Autoregressive Multi-view Diffusion for Flexible 3D Novel View Synthesislink
2025arXivWonderZoomWonderZoom: Multi-Scale 3D World GenerationlinkGitHub
2026arXivSceneFrom3DSceneFrom3D: Geometry-Conditioned Outdoor 3D Scene Generation via View Scheduling with Object-Level ControllinkGitHub

Video-based Generation

Two-stage Generation

One-stage Generation

YearVenueAcronymPaperProjectRepo@GitHub
2024ICLRMagicDriveMagicDrive: Street View Generation with Diverse 3D Geometry ControllinkGitHub
2024CVPRPanaceaPanacea: Panoramic and Controllable Video Generation for Autonomous DrivinglinkGitHub
2024CVPRDrive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivinglinkGitHub
2024CVPR360DVD360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModellinkGitHub
2024ECCVDriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivinglinkGitHub
2024ECCVDrivingDiffusionDrivingDiffusion: Layout-Guided Multi-View Driving Scenarios Video Generation with Latent Diffusion ModellinkGitHub
2024ECCVWoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene GenerationGitHub
2024NeurIPSVistaVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilitylinkGitHub
2024NeurIPSDIAMONDDiffusion for World Modeling: Visual Details Matter in AtarilinkGitHub
2025AAAIDriveDreamer-2DriveDreamer-2: LLM-Enhanced World Models for Diverse Driving Video GenerationlinkGitHub
2025ICLR4K4DGen4K4DGen: Panoramic 4D Generation at 4K ResolutionlinkGitHub
2025ICLRGameGen-XGameGen-X: Interactive Open-world Game Video GenerationlinkGitHub
2025ICLRGameNGenDiffusion Models Are Real-Time Game Engineslink
2025ICLRGenexGenerative World ExplorerlinkGitHub
2025ICLRGLADGlad: A Streaming Scene Generator for Autonomous Driving
2025CVPRDrivingSphereDrivingSphere: Building a High-fidelity 4D World for Closed-loop SimulationlinkGitHub
2025CVPRStreetCrafterStreetCrafter: Street View Synthesiswith Controllable Video Diffusion ModelslinkGitHub
2025CVPRDriveScapeDriveScape: Towards High-Resolution Controllable Multi-View Driving Video Generationlink
2025CVPRUniSceneUniScene: Unified Occupancy-centric Driving Scene GenerationlinkGitHub
2025CVPRGEMGEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition ControllinkGitHub
2025CVPRUMGenGenerating Multimodal Driving Scenes via Next-Scene PredictionlinkGitHub
2025CVPRCAT4DCAT4D: Create Anything in 4D with Multi-View Video Diffusion Modelslink
2025CVPRWonderlandWonderland: Navigating 3D Scenes from a Single ImagelinkGitHub
2025CVPRVideoSceneVideoScene: Distilling Video Diffusion Model to Generate 3D Scenes in One SteplinkGitHub
2025CVPRScene SplatterScene Splatter: Momentum 3D Scene Generation from Single Image with Video Diffusion ModellinkGitHub
2025CVPRDynamicScalerDynamicScaler: Seamless and Scalable Video Generation for Panoramic SceneslinkGitHub
2025ICMLAdaWorldAdaWorld: Learning Adaptable World Models with Latent ActionslinkGitHub
2025NatureWHAMWorld and Human Action Models towards gameplay ideationlink
2025ICCVGameFactoryGameFactory: Creating New Games with Generative Interactive VideoslinkGitHub
2025ICCVWonderPlayWonderPlay: Dynamic 3D Scene Generation from a Single Image and ActionslinkGitHub
2025ICCVMagicDrive-V2MagicDrive-V2: High-Resolution Long Video Generation for Autonomous Driving with Adaptive ControllinkGitHub
2025ICCVDynamicVoyagerVoyaging into Unbounded Dynamic Scenes from a Single ViewlinkGitHub
2025ICCVInfiniCubeInfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video ModelslinkGitHub
2025ICCVVMemVMem: Consistent Interactive Video Scene Generation with Surfel-Indexed View MemorylinkGitHub
2025SIGGRAPH AsiaVideoFrom3DVideoFrom3D: 3D Scene Video Generation via Complementary Image and Video Diffusion ModelslinkGitHub
2025SIGGRAPH AsiaWorldExplorerWorldExplorer: Towards Generating Fully Navigable 3D SceneslinkGitHub
2025TOGVoyagerVoyager: Long-Range and World-Consistent Video Diffusion for Explorable 3D Scene Generation
2026ICLRFantasyWorldFantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D PredictionlinkGitHub
2026CVPRCaptain SafariCaptain Safari: A World EnginelinkGitHub
2026CVPRPerpetualWonderPerpetualWonder: Long-Horizon Action-Conditioned 4D Scene GenerationlinkGitHub
2026CVPRWorldForgeWorldForge: Unlocking Emergent 3D/4D Generation in Video Diffusion Model via Training-Free GuidancelinkGitHub
2026CVPRWorldReelWorldReel: 4D Video Generation with Consistent Geometry and Motion Modelinglink
2026ICRAUniFutureSeeing the Future, Perceiving the Future: A Unified Driving World Model for Future Generation and PerceptionlinkGitHub
2026ECCVOne4DOne4D: Unified 4D Generation and Reconstruction via Decoupled LoRA ControllinkGitHub
2026ECCVLivingWorldLivingWorld: Interactive 4D World Generation with Environmental Dynamicslink
2023arXivGAIA-1GAIA-1: A Generative World Model for Autonomous Drivinglink
2024arXivMagicDrive3DMagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street SceneslinkGitHub
2024arXivDelphiUnleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video GenerationlinkGitHub
2024arXivBEVWorldBEVWorld: A Multimodal World Model for Autonomous Driving via Unified BEV Latent SpaceGitHub
2024arXivDriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivinglinkGitHub
2024arXivDiVEDiVE: DiT-based Video Generation with Enhanced ControllinkGitHub
2024arXivDreamForgeDreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Sceneslink
2024arXivSyntheOccSyntheOcc: Synthesize Geometric-Controlled Street View Images through 3D Semantic MPIslinkGitHub
2024arXivCogDrivingSeeing Beyond Views: Multi-View Driving Scene Video Generation with Holistic Attentionlink
2024arXivImagine360Imagine360: Immersive 360 Video Generation from Perspective AnchorlinkGitHub
2024arXivDrivingWorldDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTlinkGitHub
2024arXivViewCrafterViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View SynthesislinkGitHub
2024arXivViewExtrapolatorNovel View Extrapolation with Video Diffusion PriorslinkGitHub
2025arXivDreamDriveDreamDrive: Generative 4D Scene Modeling from Street View Imageslink
2025arXivMaskGWMMaskGWM: A Generalizable Driving World Model with Video Mask ReconstructionlinkGitHub
2025arXivSimWorldSimWorld: A Unified Benchmark for Simulator-Conditioned Scene Generation via World ModelGitHub
2025arXivDiST-4DDiST-4D: Disentangled Spatiotemporal Diffusion with Metric Depth for 4D Driving Scene GenerationlinkGitHub
2025arXivGAIA-2GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Drivinglink
2025arXivSteerXSteerX: Creating Any Camera-Free 3D and 4D Scenes with Geometric SteeringlinkGitHub
2025arXivFlexWorldFlexWorld: Progressively Expanding 3D Scenes for Flexiable-View SynthesislinkGitHub
2025arXivWORLDMEMWORLDMEM: Long-term Consistent World Simulation with MemorylinkGitHub
2025arXivHoloTimeHoloTime: Taming Video Diffusion Models for Panoramic 4D Scene GenerationlinkGitHub
2025arXivMineWorldMineWorld: a Real-Time and Open-Source Interactive World Model on MinecraftlinkGitHub
2025arXivCoGenCoGen: 3D Consistent Video Generation via Adaptive Conditioning for Autonomous Drivinglink
2025arXivDreamlandDreamland: Controllable World Creation with Simulator and Generative Modelslink
2025arXivMatrix-GameMatrix-Game: Interactive World Foundation ModellinkGitHub
2025arXivMatrix-Game 2.0Matrix-Game 2.0: An Open-Source, Real-Time, and Streaming Interactive World ModellinkGitHub
2025arXivHunyuan-GameCraftHunyuan-GameCraft: High-dynamic Interactive Game Video Generation with Hybrid History Conditionlink
2025arXivCoCo4DCoCo4D: Comprehensive and Complex 4D Scene GenerationlinkGitHub
2025arXivWonderFreeWonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene ExplorationlinkGitHub
2025arXiv4DVD4DVD: Cascaded Dense-view Video Diffusion Model for High-quality 4D Content Generationlink
2025arXivIDCNetIDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Controllink
2025arXiv4DNeX4DNeX: Feed-Forward 4D Generative Modeling Made EasylinkGitHub
2025arXivFrom Virtual Games to Real-World PlaylinkGitHub
2025arXivEvoWorldEvoWorld: Evolving Panoramic World Generation with Explicit 3D MemoryGitHub
2025arXivMagicWorldMagicWorld: Interactive Geometry-driven Video World ExplorationlinkGitHub
2025arXivHY-World 1.5HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric ConsistencylinkGitHub

Datasets

Indoor Datasets

YearTypeSourceAcronymPaperProject
2012Indoor, NatureRealSUN360Recognizing scene viewpoint using panoramic place representationlink
2012IndoorRealNYUv2Indoor Segmentation and Support Inference From RGBD Imageslink
2015IndoorRealSunRGBDSun RGB-D: A RGB-D scene understanding benchmark suitelink
2016IndoorRealSceneNNSceneNN: A Scene Meshes Dataset with aNNotationslink
2017IndoorReal2D-3D-SJoint 2D-3D-Semantic Data for Indoor Scene Understandinglink
2017IndoorRealMatterport3DMatterport3D: Learning from RGB-D Data in Indoor Environmentslink
2017IndoorRealScanNetScanNet: Richly-annotated 3D Reconstructions of Indoor Sceneslink
2017IndoorRealLaval IndoorLearning to Predict Indoor Illumination from a Single Imagelink
2018Indoor, UrbanRealRealEstate10KStereo Magnification: Learning View Synthesis using Multiplane Imageslink
2019IndoorRealReplicaThe Replica Dataset: A Digital Replica of Indoor Spaceslink
2020IndoorReal3DSSGLearning 3D Semantic Scene Graphs from 3D Indoor Reconstructionslink
2021IndoorRealHM3DHabitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AIlink
2023IndoorRealScanNet++ScanNet++: A high-fidelity dataset of 3D indoor sceneslink
2023Indoor, Nature, UrbanRealDL3DV-10KDL3DV-10K: A Large-Scale Scene Dataset for Deep Learning-based 3D Visionlink
2012IndoorSyntheticSceneSynthExample-based synthesis of 3D object arrangementslink
2017IndoorSyntheticSUNCGSemantic Scene Completion from a Single Depth Imagelink
2020IndoorSyntheticStructured3DStructured3D: A Large Photo-realistic Dataset for Structured 3D Modelinglink
2020IndoorSyntheticHyperSimHyperSim: A photorealistic synthetic dataset for holistic indoor scene understandinglink
2021IndoorSynthetic3D-FRONT3D-FRONT: 3D Furnished Rooms with layOuts and semaNTicslink
2021IndoorSynthetic3D-Future3D-FUTURE: 3D Furniture shape with TextURElink
2023IndoorSyntheticSG-FRONTCommonScenes: Generating Commonsense 3D Indoor Scenes with Scene Graph Diffusionlink
2025IndoorSyntheticSE(3) SceneSteerable Scene Generation with Post Training and Inference-Time Searchlink
2025IndoorSyntheticSYNBUILD-3DSYNBUILD-3D: A large, multi-modal, and semantically rich synthetic dataset of 3D building models at Level of Detail 4link
2025IndoorSyntheticInternScenesInternScenes: A Large-scale Simulatable Indoor Scene Dataset with Realistic Layoutslink
2025IndoorSyntheticSPATIALGENSPATIALGEN: Layout-guided 3D Indoor Scene Generationlink
2025IndoorSyntheticMesaTaskMesaTask: Towards Task-Driven Tabletop Scene Generation via 3D Spatial Reasoninglink

Natural Datasets

Urban Datasets

Tasks and Applications

Downstream Tasks

3D Scene Editing

YearVenueAcronymPaperProjectRepo@GitHub
2022CVPRStyleMeshStyleMesh: Style Transfer for Indoor 3D Scene ReconstructionslinkGitHub
2023CVPRDisCoSceneDisCoScene: Spatially Disentangled Generative Radiance Fields for Controllable 3D-aware Scene SynthesislinkGitHub
2023CVPRLEGO-NetLEGO-Net: Learning Regular Rearrangements of Objects in RoomslinkGitHub
2023CVPRLift3DLift3D: Synthesize 3D Training Data by Lifting 2D GAN to 3D Generative Radiance FieldlinkGitHub
2023CVPRText2SceneText2Scene: Text-driven Indoor Scene Stylization with Part-aware Details
2023ICRACabiNetCabiNet: Scaling Neural Collision Detection for Object Rearrangement with Procedural Scene GenerationlinkGitHub
2023MMRoomDreamerRoomDreamer: Text-Driven 3D Indoor Scene Synthesis with Coherent Geometry and Texture
2024CVPRSceneTexSceneTex: High-Quality Texture Synthesis for Indoor Scenes via Diffusion PriorslinkGitHub
2024CVPRControlRoom3DControlRoom3D 🤖Room Generation using Semantic Proxy Roomslink
2024ECCVStyleCityStyleCity: Large-Scale 3D Urban Scenes StylizationlinkGitHub
2024ECCVRoomTexRoomTex: Texturing Compositional Indoor Scenes via Iterative InpaintinglinkGitHub
2024ECCV3D-GOI3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editinglink
2024MMSceneExpanderSceneExpander: Real-Time Scene Synthesis for Interactive Floor Plan EditingGitHub
2024NeurIPSNeural AssetsNeural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion Modelslink
2024NeurIPSDeBaRADeBaRA: Denoising-Based 3D Room Arrangement Generation
2024SIGGRAPH AsiaInstanceTexInstanceTex: Instance-level Controllable Texture Synthesis for 3D Scenes via Diffusion Priorslink
2024TVCGSceneDirectorSceneDirector: Interactive Scene Synthesis by Simultaneously Editing Multiple Objects in Real-TimeGitHub
2024VRDreamSpaceDreamSpace: Dreaming Your Room Space with Text-Driven Panoramic Texture PropagationlinkGitHub
20253DVCtrl-RoomCtrl-Room: Controllable Text-to-3D Room Meshes Generation with Layout ConstraintslinkGitHub
2025CVPRRoomPainterRoomPainter: View-Integrated Diffusion for Consistent Indoor Scene Texturing
2025SIGGRAPHReStyle3DReStyle3D: Scene-Level Appearance Transfer with Semantic CorrespondenceslinkGitHub
2025NeurIPSStyl3RStyl3R: Instant 3D Stylized Reconstruction for Arbitrary Scenes and StyleslinkGitHub
2026CVPRCatalyst4DCatalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagationlink
2026CVPREdit-As-ActEdit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene EditinglinkGitHub

Human-Scene Interaction

Embodied Navigation

Application Domains

Robotics

YearVenueAcronymPaperProjectRepo@GitHub
2023NeurIPSUniPiLearning Universal Policies via Text-Guided Video Generationlink
2023NeurIPSHiPCompositional Foundation Models for Hierarchical PlanninglinkGitHub
2024CoRLImagination PolicyImagination Policy: Using Generative Point Cloud Models for Learning Manipulation PolicieslinkGitHub
2024CoRLEurekaverseEurekaverse: Environment Curriculum Generation via Large Language ModelslinkGitHub
2024ICLRGR-1Unleashing Large-Scale Video Generative Pre-training for Visual Robot ManipulationlinkGitHub
2024ICMLRoboGenRoboGen: Towards Unleashing Infinite Data for Automated Robot Learning via Generative SimulationlinkGitHub
2024ICMLVLPUsing Left and Right Brains Together: Towards Vision and Language Planning
2024IROSActNeRFUncertainty-aware Active Learning of NeRF-based Object Models for Robot Manipulators using Visual and Re-orientation ActionslinkGitHub
2024NeurIPSCLOVERClosed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationGitHub
2025ICLRSlowFast-VGenSlowFast-VGen: Slow-Fast Learning for Action-Driven Long Video GenerationlinkGitHub
2025ICLRReGenReGen: Generative Robot Simulation via Inverse Designlink
2025ICMLVideo Prediction PolicyVideo Prediction Policy: A Generalist Robot Policy with Predictive Visual RepresentationslinkGitHub
2025RSSRoboVerseRoboVerse: Towards a Unified Platform, Dataset and Benchmark for Scalable and Generalizable Robot LearninglinkGitHub
2024arXivGR-2GR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot Manipulationlink
2025arXivVideoWorldVideoWorld: Exploring Knowledge Learning from Unlabeled VideoslinkGitHub
2025arXivCosmos-Transfer1Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControllinkGitHub
2025arXivTesserActTesserAct: Learning 4D Embodied World ModelslinkGitHub
2025arXivMarketGenMarketGen: A Scalable Simulation Platform with Auto-Generated Embodied Supermarket Environmentslink
2025arXivTabletopGenTabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single ImagelinkGitHub
2026arXivWorldComposerFrom Seeing to Simulating: Generative High-Fidelity Simulation with Digital Cousins for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXivSimFoundrySimFoundry: Modular and Automated Scene Generation for Policy Learning and Evaluationlink
2026arXivRoboSnapRoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and EvaluationlinkGitHub
2026arXiv4DSynth4DSynth: Controllable Procedural World Synthesis for Dynamic Embodied SimulationlinkGitHub

Autonomous Driving

YearVenueAcronymPaperProjectRepo@GitHub
2024CVPRCam4DOccCam4DOcc: Benchmark for Camera-Only 4D Occupancy Forecasting in Autonomous Driving ApplicationsGitHub
2024CVPRDrive-WMDriving into the Future: Multiview Visual Forecasting and Planning with World Model for Autonomous DrivinglinkGitHub
2024ECCVDriveDreamerDriveDreamer: Towards Real-world-driven World Models for Autonomous DrivinglinkGitHub
2024ECCVOccWorldOccWorld: Learning a 3D Occupancy World Model for Autonomous DrivinglinkGitHub
2024ECCVWoVoGenWoVoGen: World Volume-Aware Diffusion for Controllable Multi-camera Driving Scene GenerationGitHub
2024ICLRMagicDriveMagicDrive: Street View Generation with Diverse 3D Geometry ControllinkGitHub
2024NeurIPSVistaVista: A Generalizable Driving World Model with High Fidelity and Versatile ControllabilitylinkGitHub
2025AAAIDrive-OccWorldDriving in the Occupancy World: Vision-Centric 4D Occupancy Forecasting and Planning via World Models for Autonomous DrivinglinkGitHub
2025CVPRDrivingSphereDrivingSphere: Building a High-fidelity 4D World for Closed-loop SimulationlinkGitHub
2025ICLRGLADGlad: A Streaming Scene Generator for Autonomous Driving
2025NeurIPSOrbisOrbis: Overcoming Challenges of Long-Horizon Prediction in Driving World ModelslinkGitHub
2026CVPRDriveLaWDriveLaW: Unifying Planning and Video Generation in a Latent Driving WorldGitHub
2023arXivGAIA-1GAIA-1: A Generative World Model for Autonomous Drivinglink
2024arXivOccSoraOccSora: 4D Occupancy Generation Models as World Simulators for Autonomous DrivinglinkGitHub
2024arXivDelphiUnleashing Generalization of End-to-End Autonomous Driving with Controllable Long Video GenerationlinkGitHub
2024arXivDriveArenaDriveArena: A Closed-loop Generative Simulation Platform for Autonomous DrivinglinkGitHub
2024arXivDiVEDiVE: DiT-based Video Generation with Enhanced ControllinkGitHub
2024arXivDreamForgeDreamForge: Motion-Aware Autoregressive Video Generation for Multi-View Driving Sceneslink
2024arXivDrivingWorldDrivingWorld: Constructing World Model for Autonomous Driving via Video GPTlinkGitHub
2025arXivDreamDriveDreamDrive: Generative 4D Scene Modeling from Street View Imageslink
2025arXivCosmos-Transfer1Cosmos-Transfer1: Conditional World Generation with Adaptive Multimodal ControllinkGitHub
2025arXivGenieDriveGenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video GenerationlinkGitHub
2026arXivVectorWorldVectorWorld: Efficient Streaming World Model via Diffusion Flow on Vector GraphsGitHub
2026arXivOccSimOccSim: Multi-kilometer Simulation with Long-horizon Occupancy World ModelslinkGitHub
3d-aigc
3d-scene-generation
awesome-list
paper-review

Contributors

wenbc21

46 commits

hzxie

25 commits

Zhu-Liyuan

2 commits

Cylrx

1 commits