chang-xinhai/Awesome-Embodied-3DV

A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.

Python

13

151 commits

updated Sep 23, 2026

See the code

README

Awesome-Embodied-3DV

A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.

Perception Representation Reconstruction Generation Embodiment Infrastructure

About

Awesome-Embodied-3DV is a curated list for the research space where 3D vision, 3D/4D reconstruction, 3D generation, simulation-ready assets, and embodied world models meet.

This repository focuses on:

  • Data Perception: depth, normals, active imaging, panoramic/fisheye, dense mapping, and 3D semantic understanding (detection, segmentation, grounding) that feed downstream systems
  • 3D/4D Representation: 3DGS, 4DGS, NeRF/SDF, mesh, voxel, point, and hybrid structures
  • 3D Reconstruction: offline, feed-forward, streaming, online, semantic, instance-level, and dynamic reconstruction
  • 3D Generation: object-level, part-level, articulated, scene-level, editable, and simulation-ready asset generation
  • Embodiment & World Models: reconstruction-based world models, dynamic scene graphs, physical interaction & affordance, human-centric 3D, robotics integration, and agent-facing 3D grounding
  • Datasets, Benchmarks & Infrastructure: datasets, metrics, simulators, toolchains, and surveys for fast research orientation

This list is intentionally embodied-3DV-first. It includes 3D generation and 3DGS work only when it helps understand, build, evaluate, or deploy 3D assets and world models for embodied agents. It is not a generic catalog of all 3D generation, editing, rendering, compression, or graphics papers.

Daily candidate feed. The automatically updated arXiv Daily is a high-recall, topic-tagged candidate archive across the six areas above. It is deliberately broader than this curated README: papers are promoted here only after manual primary-source verification.

Must Read

Start here if you want the shortest path through the field.

GoalStart with
Estimate temporally consistent video depthVideo Depth Anything, FlashDepth, ICDepth, ViGeo, DVD, RollingDepth
Deploy zero-shot monocular depth at the edgeZipDepth, DepthART
Recover metric-scale geometry from RGBMetricAnything, Metric DAv2, Depth Pro, UniDepthV2
Understand feed-forward 3D reconstructionAdvances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey, DUSt3R, MASt3R, VGGT
Track online / streaming 3D reconstructionDynamic 3D Gaussians, CUT3R, Spann3R, SLAM3R, StreamSplat, UniSim-SLAM
Study dynamic 4D feed-forward geometryPAGE-4D, 4D-VGGT, DynamicVGGT
Learn representation foundations3D Gaussian Splatting, NeRF, Instant-NGP, TensoRF
Generate object and simulation-ready assetsTRELLIS, Hunyuan3D 2.0, TripoSR, PartCrafter
Generate scenes and worldsInfinigen, SceneDreamer, Text2Room, WonderWorld, EmbodiedGen V2
Build articulated objectsArtGS, Articulate Anything, URDF-Anything, URDF-Anything+, PartNet-Mobility
Build an open-vocabulary 3D robot memoryCLIP-Fields, ConceptFusion, Open-Fusion, ConceptGraphs
Build persistent, action-conditioned 3D world modelsLearning 3D Persistent Embodied World Models, GEN3C, NeoVerse, Genie 4D, BWM
Pick datasets and benchmarksScanNet, Tanks and Temples, DTU, Objaverse, GSO, HM3D, TransBiolab

News

  • [2026-08-04] Added a six-topic, automatically refreshed arXiv Daily candidate archive, with manual verification required before promotion to this curated list.
  • [2026-06-15] Added a dedicated 3D Editing taxonomy under 3D Generation, covering object-level, scene-level, and dynamic / 4D editing methods.
  • [2026-04-30] Initialized Awesome-Embodied-3DV with a six-part taxonomy for data perception, representations, reconstruction, generation, embodied world models, and infrastructure.

Contents

📡 1. Data Perception

Data perception covers the sensor-facing and semantic layers: extracting geometric priors (depth, normals), understanding 3D semantics (detection, segmentation, grounding), active-imaging signals, and dense maps from 2D images, video, or physical sensors.

1.1 Geometric Priors

1.1.1 Monocular High-Fidelity Depth

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Metric Depth, 3DGS Relocalization, Sparse PnP Anchors, Temporal MemoryThe Chinese University of Hong Kong, ShenzhenRIDE: Relocalization-Informed Depth Estimation with 3D Gaussian SplattingarXivpaper
2026-09-08Diffusion Transformer, Single-Step Depth, Sharp Details, Dense PredictionEPFL / HUAWEI Bayer LabMarigold V2: Revisiting Diffusion Transformers for Monocular Depth EstimationSIGGRAPH Asia 2026project / github
2026-09-08Any Camera, Metric Point Cloud, Optional Intrinsics/Sparse DepthGoogle DeepMindOmniPoint: Universal Monocular Metric Pointcloud from Any CameraECCV 2026project
2026-08-30Transparent/Reflective Scenes, Bias-Aware Training, 30M ParametersThe University of Hong KongOptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging ScenesarXivproject
2026-08-17Pixel-Space Prediction, Fine Structures, Sharp Boundaries, Efficient DepthThe Chinese University of Hong Kong, ShenzhenPXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth EstimationarXivproject
2026-08-03Geometry-Invariant Adaptation, Non-Lambertian Surfaces, Mirror/Glass DepthChangchun Institute of Optics, Fine Mechanics and Physics, CASGIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth EstimationarXivpaper
2026-07-23UAV Depth, Arbitrary Camera Pose, Metric GeometryAerospace Information Research Institute, CASDAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOVarXivgithub
2026-07-20Fine-Detail Geometry, Sparse Volumetric Refinement, Metric ScaleTsinghua UniversityMoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivproject
2026-07-19Lightweight Foundation Depth, Camera-Conditioned Metric Depth, Edge DeploymentUniversity of TrentoDepthART: Scaling Foundation Monocular Depth to Tiny ModelsACM Multimedia 2026project / github
2026-07-19Metric Depth, Odometry Anchor, Recurrent SLAMUC BerkeleyDROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth EstimationarXivpaper
2026-07-17Stereo Distillation, Epipolar Cues, Metric DepthMichigan State UniversityGeometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular DeptharXivpaper
2026-07-14Auto-Regressive Depth, Coarse-to-Fine, Semantic GuidanceSun Yat-sen UniversityARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual ConditioningarXivpaper
2026-07-13Metric Point Map, Pixel-Wise Calibration, Camera DiversityThe University of Hong KongFoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric GeometryECCV 2026project / github / model
2026-07-09Lightweight Zero-Shot, 6.1M Parameters, On-Device DepthUniversity of BolognaZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any DeviceECCV 2026project / github
2026-05-27Multi-Layer Depth, Transparent Surfaces, Point ProcessPrinceton UniversitySeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined GroupingCVPR 2026github
2026-05-15VLM, Dense Metric Depth, Spatial ReasoningZhejiang UnivUnlocking Dense Metric Depth Estimation in VLMsarXivproject / github
2026-05-12Sparse 3D Anchors, Relative-to-Metric, Graph OptimizationTongji UniversityThe Midas Touch for Metric DepthCVPR 2026 Highlightproject
2026-03-28Universal Camera, Metric Depth, Zero-ShotMichigan State UniversityUniDAC: Universal Metric Depth Estimation for Any CameraCVPR 2026paper
2026-03-20Transparent Objects, Generative Opacification, Monocular Depth, SeeClear-396kUniversity of California, Los AngelesSeeClear: Reliable Transparent Object Depth Estimation via Generative OpacificationECCV 2026project / dataset
2026-03-17Diffusion Prior, Real-World Data, Monocular DepthNanjing University of Science and TechnologyIris: Bringing Real-World Priors into Diffusion Model for Monocular Depth EstimationCVPR 2026paper
2026-03-04Fine-Grained Geometry, Dual-Stream, EfficientUMass AmherstDAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationCVPR 2026paper
2026-01-29Sparse Metric Prompt, 20M Image-Depth Pairs, Metric Foundation ModelLi Auto IncMetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous SourcesECCV 2026project / github
2026-01-06Arbitrary-Resolution, Neural Implicit, Fine DetailsZhejiang UnivInfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsCVPR 2026project / github
2025-12-13Defocus Cue, Bokeh Stack, Metric DepthNanyang Tech UnivBoosting Monocular Metric Depth Estimation via Bokeh RenderingICML 2026project / github
2025-11-30Deterministic Diffusion, Dense Geometry, Fine DetailsHKUST(GZ)Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative ModelarXivproject / github
2025-11-13Any-View Depth, Metric Geometry, Multi-ViewByteDanceDepth Anything 3: Recovering the Visual Space from Any ViewsICLR 2026project / github
2025-10-27Unified Generation+Depth, Diffusion Prior, Zero-ShotHUSTMore Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion ModelsNeurIPS 2025github
2025-10-08Pixel-Space Diffusion, Flying-Pixel-Free, Point CloudsHUSTPixel-Perfect Depth with Semantics-Prompted Diffusion TransformersNeurIPS 2025project / github
2025-09-29VLM, Metric Depth, Sparse SupervisionMeta AIDepthLM: Metric Depth From Vision Language ModelsICLR 2026 Oralgithub
2025-07-03Monocular Geometry, Metric Scale, Sharp DetailsUSTC / MicrosoftMoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPS 2025project / github
2025-04-16Sliding Anchor, Unknown Intrinsics, Metric ScaleShanghai UnivMetric-Solver: Sliding Anchored Metric Depth Estimation from a Single ImagearXivproject / github
2025-02-27Metric 3D Points, Self-Prompt Camera, UncertaintyETH ZurichUniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerTPAMI 2026github
2025-02-26Cross-Context Distillation, Multi-Teacher, Fine-Detail DepthZhejiang Univ of TechnologyDistill Any Depth: Distillation Creates a Stronger Monocular Depth EstimatorarXivproject / github
2024-11-27Diffusion Distillation, Metric + Sharp, Boundary DetailQualcomm AI ResearchSharpDepth: Sharpening Metric Depth Predictions Using Diffusion DistillationCVPR 2025project / github
2024-10-02Metric Depth, Zero-Shot, Image PriorsAppleDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLR 2025github
2024-09-26Diffusion, Single-Step, Dense GeometryHKUST(GZ)Lotus: Diffusion-based Visual Foundation Model for High-quality Dense PredictionICLR 2025project / github
2024-06-13Monocular Depth, Foundation Model, Metric DAv2TikTok / HKUDepth Anything V2NeurIPS 2024project / github / metric models
2024-03-27Metric Depth, Universal, Zero-ShotETH ZurichUniDepth: Universal Monocular Metric Depth EstimationCVPR 2024github
2024-01-19Monocular Depth, Relative Depth, Foundation ModelTikTok / HKUDepth Anything: Unleashing the Power of Large-Scale Unlabeled DataCVPR 2024project
2023-12-04Monocular Depth, Zero-Shot, Affine-InvariantIntel LabsMarigold: Repurposing Diffusion-Based Image Generators for Monocular Depth EstimationCVPR 2024project
2023-07-20Metric 3D, Zero-Shot, Canonical SpaceAlibabaMetric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageICCV 2023github
2021-03-24ViT, Dense Prediction, Foundation, DepthIntel LabsDPT: Vision Transformers for Dense PredictionICCV 2021github
2019-07-02Robust Depth, Zero-Shot, Cross-Dataset, MiDaSIntel LabsMiDaS: Towards Robust Monocular Depth EstimationTPAMI 2022github

1.1.2 Temporally Consistent Video Depth

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-23Unified Video Model, Depth+Normals, Temporal ConsistencyAdobe ResearchUnified Video Dense Prediction from Disjoint DataarXivproject
2026-07-02Video Diffusion, In-Context Conditioning, Zero-ShotHKUSTICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context ConditioningECCV 2026project
2026-05-28Streaming Geometry, Dynamic Chunking, Depth+NormalsZhejiang UniversityTowards Consistent Video Geometry EstimationarXivproject
2026-05-11Camera Motion, 3D Consistency, Geometry EmbeddingHUSTGemDepth: Geometry-Embedded Features for 3D-Consistent Video DeptharXivgithub
2026-04-08Post-Processing, Scalable, Single-Image BackboneSeoul National UniversityVDPP: Video Depth Post-Processing for Speed and ScalabilityCVPR 2026 ECV Workshoppaper
2026-04-02Pose Refinement, Temporal Consistency, Monocular VideoAjou UniversityPTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal ConsistencyCVPR 2026paper
2026-03-12Deterministic Diffusion, Generative Prior, Long VideoHKUST(GZ)DVD: Deterministic Video Depth Estimation with Generative PriorsarXivproject
2026-01-06Temporal Stability, Monocular Video, Long SequenceETH ZurichStableDPT: Temporal Stable Monocular Video Depth EstimationarXivpaper
2025-12-20Endoscopic Geometry, Streaming Mamba, Metric DepthVanderbilt UniversityEndoStreamDepth: Temporally Consistent Monocular Depth Estimation for Endoscopic Video StreamsarXivgithub / paper
2025-12-11Sparse Keyframes, Propagation, Long-Video ConsistencyETH ZurichVideo Depth Propagation3DV 2026paper
2025-10-10Online Inference, Low Memory, Temporal ConsistencyHeidelberg UniversityOnline Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory ConsumptionarXivpaper
2025-07-02Diffusion Guidance, Scale Synchronization, Geometry ConsistencyTsinghua UniversityDepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationICCV 2025project
2025-04-092K Streaming, Mamba, 24 FPSNetflix Eyeline StudiosFlashDepth: Real-time Streaming Video Depth Estimation at 2K ResolutionICCV 2025 Highlightproject / github
2025-01-21Super-Long Video, Temporal Gradient, 30 FPSByteDanceVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPR 2025project / github
2024-11-28Long Video, Diffusion, Multi-Resolution AlignmentETH ZurichVideo Depth without Video ModelsCVPR 2025project

1.1.3 Geometric Consistency Prior

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-06Surface Normals, Transparent Objects, Rectified Flow, Edge RefinementZhejiang UniversityTransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal EstimationarXivproject
2026-07-28Stereo Depth, Walsh-Hadamard Mixing, Efficient InferenceAuthorsWHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token MixingarXivpaper
2026-07-22Stereo Diffusion Transformer, Flow Matching, Progressive RefinementBeihang UniversitySTEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow MatchingarXivpaper
2026-07-15Vision Features, SE(3) Latent Geometry, Visual NavigationGoogle DeepMindSeeSE3: Emergence of 3D Space in Vision FeaturesarXivpaper
2026-07-14Heterogeneous Cameras, Metric Depth, Real-TimeD-RoboticsX-Lens: Real-Time Metric Depth Estimation with Heterogeneous CamerasarXivproject / github
2026-07-10Video Generative Pretraining, Depth/Normals/Pose, Grounded 4DGoogle DeepMindVideo Generation Models are General-Purpose Vision LearnersECCV 2026project
2026-07-06Boundary-Centric Pretraining, Dense Spatial PerceptionRobbyantLingBot-Vision: Vision Pretraining for Dense Spatial PerceptionarXivproject / github
2026-03-02Zero-Shot Stereo, Structure Prompt, Motion PromptHUSTPromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsCVPR 2026paper
2025-12-11Real-Time Stereo, Zero-Shot, Foundation ModelNVIDIAFast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingCVPR 2026paper
2025-07-22Foundation Model, Depth/Normal/Pointmap, Multi-ViewSJTUDens3R: A Foundation Model for 3D Geometry PredictionICLR 2026paper
2025-04-15Surface Normals, Foundation Model, Video, TemporalAuthorsNormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsICCV 2025project / paper
2025-03-21Scalable Depth, Autoregressive, 2B ParametersBaiduDAR: Scalable Autoregressive Monocular Depth EstimationCVPR 2025project
2025-03-20Universal Camera, Spherical 3D, Any CameraETH ZurichUniK3D: Universal Camera Monocular 3D EstimationCVPR 2025github / paper
2025-03-11LiDAR Surface Normal, Dataset, Point CloudTU GrazLiSu: A Dataset and Method for LiDAR Surface Normal EstimationCVPR 2025github
2025-01-17Stereo, Foundation Model, RAFT-Style, Zero-ShotNVIDIAFoundationStereo: Zero-Shot Stereo MatchingCVPR 2025 Oral / Best Paper Nominationgithub
2025-01-17Stereo, Robust, Zero-Shot, Non-LambertianUniv of BolognaStereo Anywhere: Robust Zero-Shot Deep Stereo MatchingCVPR 2025project
2024-12-11Panoramic / Fisheye Depth, Zero-Shot MetricIntelDepth Any Camera: Zero-Shot Metric Depth from Any CameraCVPR 2025project
2024-10-24Monocular Geometry, Pointmap, Affine-InvariantUSTC / MicrosoftMoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain ImagesCVPR 2025 Oralgithub
2024-06-24Normal Estimation, 3D Priors, Surface GeometryNvidiaStableNormal: Reducing Diffusion Variance for Stable and Sharp NormalSIGGRAPH Asia 2024project
2024-03-27Diffusion, Effective Conditioning, ViT PriorsIIT DelhiECoDepth: Effective Conditioning of Diffusion Models for Monocular DepthCVPR 2024github
2024-03-22Metric 3D, Multi-Task, Geometry FoundationShanghai AI LabMetric3D v2: A Versatile Monocular Geometric Foundation ModelTPAMI 2025project
2024-03-18Geometry, Normals, Depth, Multi-TaskAppleGeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single ImageECCV 2024project
2024-03-01Inductive Biases, Surface Normal, OralImperial College LondonDSINE: Rethinking Inductive Biases for Surface Normal EstimationCVPR 2024 Oralgithub
2023-12-04High-Res, Patch-Wise, Model-AgnosticKAUSTPatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric DepthCVPR 2024project
2023-09-25Iterative Bins, Elastic, GRU, Classification-RegressionBeihang UnivIEBins: Iterative Elastic Bins for Monocular Depth EstimationNeurIPS 2023github
2023-04-13Internal Discretization, Continuous-Discrete, DepthETH ZurichiDisc: Internal Discretization for Monocular Depth EstimationCVPR 2023github

1.2 3D Semantic Understanding

1.2.1 Open-Vocabulary 3D Segmentation & Grounding

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Annotation-Free, Open-Vocabulary 3D, Language-Space LiftingNational Technical University of AthensGoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space LiftingarXivpaper
2026-07-21Referring Segmentation, 3DGS, Generalized GroundingPeking UniversityZeroSplat: Generalized Referring Segmentation in 3D Gaussian SplattingarXivproject
2026-06-23Open-Vocabulary BEV, 3DGS, Geometric ConstraintsKAIST AIOpen-Vocabulary BEV Segmentation with 3D-Aware Geometric ConstraintsECCV 2026paper
2026-06-04Open-Vocabulary, Functionality Segmentation, RoboticsAuthorsT-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality SegmentationarXivpaper
2026-05-07Open-Vocabulary, Gaussian Feature Field, CodebookTU Munich / GoogleOpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook AttentionarXivpaper
2025-11-20Open-Vocabulary, SAM3, Promptable, DETRMetaSAM 3: Segment Anything with ConceptsICLR 2026github / paper
2025-04-03Open-Vocabulary, Dual-Level Contrastive, Instance-AwareBIGAI / TsinghuaMPEC: Masked Point-Entity Contrast for Open-Vocabulary 3D Scene UnderstandingCVPR 2025project
2025-03-22Training-Free, MLLM Caption, Voxel GroupingNVIDIA Research TaiwanOpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene UnderstandingCVPR 2026project
2025-03-19SAM-2, 3D Tracking, Dynamic ProgrammingVinAIAny3DIS: Open-Vocabulary 3D Instance Segmentation with SAM-2CVPR 2025paper
2025-03-13Open-Vocabulary, LLM Canonical, Part SegmentationShandong Univ / TencentCoSMo3D: Open-World Promptable 3D Semantic Part Segmentation via LLM-Guided Canonical Spatial ModelingCVPR 2026 Oralgithub
2025-03-12Functional 3D, CoT, VLM, Training-Free, HighlightAuthorsFun3DU: Functional 3D Scene Understanding via Chain-of-ThoughtCVPR 2025 Highlightpaper
2025-01-02Panoptic, Open-Vocabulary, 3D Gaussian SplattingNUSPanoGS: Gaussian-based Panoptic Open-Vocabulary 3D Scene UnderstandingCVPR 2025project
2024-12-13Open-Vocabulary 3D, Structured Super-Gaussians, SegmentationGoogleSuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians3DV 2026paper
2024-12-12Open-Vocabulary, Foundation Dataset, Mask-Text PairsNVIDIAMosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationCVPR 2025github
2024-07-02Open-Vocabulary, Mask-Snap-Lookup, Indoor/OutdoorHKUSTOpenIns3D: Open-Vocabulary 3D Segmentation with Mask-Snap-LookupECCV 2024github
2024-04-01Region-Level, Point-Language Contrastive, Multi-VLMHKURegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingCVPR 2024project / github
2024-03-19Open-Vocabulary, MLLM, Point-Entity-Text, nuScenesMPIOV3D: Open-Vocabulary 3D Understanding with Multi-Modal AlignmentCVPR 2024paper
2024-03-192D-Guided, 3D Proposals, SAM, InstanceVinAI / IBMOpen3DIS: Open-Vocabulary 3D Instance Segmentation with 2D-Guided Mask GenerationCVPR 2024github
2024-03-15Training-Free, View-Consensus, InstancePKUMaskClustering: View-Consensus for 3D Instance SegmentationCVPR 2024paper
2024-03-11Zero-Shot, Superpoint, SAM, Scene GraphPKUSAI3D: Zero-Shot 3D Instance Segmentation by Scene-Aware Incremental MergingCVPR 2024github
2024-01-31Segment Anything, 3D Gaussians, InteractiveZhejiang UniversitySAGD: Boundary-Enhanced Segment Anything in 3D GaussiansarXivgithub
2024-01-042D+3D Unified, Single Model, HighlightAuthorsODIN: A Single Model for 2D and 3D SegmentationCVPR 2024 Highlightgithub
2023-12-26Open-Vocabulary 3DGS, Segmentation, LanguageETH ZurichLangSplat: 3D Language Gaussian SplattingCVPR 2024 Highlightproject / github
2023-11-17Unified, Semantic+Instance+Panoptic, Single TransformerSamsungOneFormer3D: One Transformer for Unified 3D SegmentationCVPR 2024github
2023-06-23Open-Vocabulary 3D Instance, Mask Proposals, CLIPETH ZurichOpenMask3D: Open-Vocabulary 3D Instance SegmentationNeurIPS 2023project
2023-06-06Segment Anything 3D, Point Cloud, InteractiveVAST AISAM3D: Segment Anything in 3D ScenesCVPR 2024github
2023-03-16Open-Vocabulary, 3D Scene, Language FieldUC BerkeleyLERF: Language Embedded Radiance FieldsICCV 2023project / github
2022-11-28Open-Vocabulary 3D, Point Cloud, SegmentationETH ZurichOpenScene: 3D Scene Understanding with Open VocabulariesCVPR 2023project / github

1.2.2 3D Instance & Panoptic Segmentation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-08Feed-Forward, Open-Vocabulary, PanopticAuthorsEPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationICML 2026github
2025-07Feed-Forward, Panoramic, DUSt3R-BasedAuthorsPanSt3R: Single-Feed 3D Geometry and Panoptic SegmentationICCV 2025paper
2025-05-14MoE, Multi-Dataset, PTv3, CLIP AlignmentUVAPoint-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic SegmentationICLR 2026project
2025-03-01Bayesian 3DGS, Training-Free, EIGSonyB3-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIGCVPR 2026project
2025-01-14End-to-End, 2D-to-3D Lifting, 3DGSCUHKUnified-Lift: Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian SplattingCVPR 2025github
2025-01-10Unsupervised Panoptic, Scene-CentricTU DarmstadtCUPS: Scene-Centric Unsupervised Panoptic SegmentationCVPR 2025 Highlightgithub
2025-01-06Zero-Shot Instance, SAM Prompts in 3DCUHK-SZ / MSRASAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation3DV 2025github
2024-07-033D Panoptic, LiDAR, Multi-SceneTsinghuaUniSeg3D: Unified 3D Panoptic SegmentationCVPR 2023github
2024-03-25Unsupervised 3D Instance, IndoorAuthorsUnScene3D: Unsupervised 3D Instance SegmentationCVPR 2024paper
2024-03-24Unified 6 Tasks, Single TransformerHUSTUniSeg3D: Unified 3D Segmentation with TransformerNeurIPS 2024paper
2023-12-15Point Cloud, Foundation Model, 3D UnderstandingShanghai AI LabPoint Transformer V3: Simpler, Faster, StrongerCVPR 2024github
2022-10-063D Instance Segmentation, Transformer, Point CloudETH ZurichMask3D: Mask Transformer for 3D Semantic Instance SegmentationICRA 2023github

1.2.3 3D Visual Grounding & Spatial Reasoning

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-233D-Aware VLM, Implicit+Explicit Geometry, RGB VideoNanyang Technological University3D-Aware VLMs with Implicit and Explicit GeometriesECCV 2026github
2026-07-143D Language Fields, Ambiguity Awareness, Object RetrievalAuthorsSaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object RetrievalarXivpaper
2026-06-23Agentic, Cognitive Map, Zero-Shot 3DSichuan UniversityAgentic Collaborative Cognition for Zero-Shot 3D UnderstandingECCV 2026project
2026-06-22Map-Grounded, MV3D-VQA, Dense RewardKAISTDense Reward for Multi-View 3D Reasoning with Global Maps and Local ViewsECCV 2026paper
2026-06-17Panoramic Reprojection, 3D VLM, Spatial ReasoningTechnical University of MunichOneCanvas: 3D Scene Understanding via Panoramic ReprojectionarXivproject
2026-06-04Part-Aware 3D-MLLM, Scene Understanding, GroundingXiamen UniversityPAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene UnderstandingarXivproject
2025-10-19Grounded CoT, 3D Reasoning, SceneCOT-185K DatasetBIGAI / PKU / TsinghuaSceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D ScenesICLR 2026paper
2025-07Scene Graph, LLM, Relation EncodingAIRI3DGraphLLM: 3D Scene Graph Learning with LLMsICCV 2025paper
2025-03-19Gesture+Language, Embodied Reference, +30%AuthorsGes3ViG: 3D Embodied Reference Understanding with Pointing GesturesCVPR 2025paper
2025-03-19Geometry VLM, Unified 3D Recon + Spatial ReasoningShanghai AI Lab / UCLA / SJTUG2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningCVPR 2026github
2025-03-12Dense Grounding, 6.2M Pairs, Hallucination BenchmarkUMich3D-GRAND: A Million-Scale Dataset for 3D GroundingCVPR 2025project
2025-03-113D-Informed, Spatial Reasoning, HighlightJHUSpatialLLM: Spatial Reasoning with 3D-Informed LLMsCVPR 2025 Highlightpaper
2025-03-10LLM Attention, Scene Magnifier, Cross-RoomSCUTLSceneLLM: LLM-Attention Adaptive 3D Scene UnderstandingCVPR 2025paper
2025-01-12LVLM-Guided, Hierarchical Feature, 3DGSFudan / NTUReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual GroundingCVPR 2025project
2025-01-04Generalist 3D LMM, Omni Superpoint TransformerAdelaide / Microsoft3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint TransformerCVPR 2025github
2024-07-18Object Identifiers, 3D VL, Unified TasksZJUChat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersNeurIPS 2024github
2024-07-15Million-Scale 3D VL, Multi-Level ContrastiveBIGAISceneVerse: Scaling 3D Vision-Language Learning for GroundingECCV 2024project
2024-06-06Zero-Shot 3D Grounding, 2D VLM TransferNUSSeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingCVPR 2025project
2024-04-03Open-Vocabulary 3D Scene Graph, Open RelationsBoschOpen3DSG: Open-Vocabulary 3D Scene Graph GenerationCVPR 2024paper
2024-03-08Interactive 3D, Language Assistant, Point CloudCUHKLL3DA: Visual Interactive Instruction Tuning for 3D Language AssistantCVPR 2024github
2024-03CLIP Cross-Modal, Contrastive, Scene GraphAuthorsCCL-3DSGG: CLIP-Driven Contrastive Learning for 3D Scene Graph GenerationCVPR 2024paper
2023-11-21Generalist Embodied 3D Agent, VLA, ICMLBIGAILEO: An Embodied Generalist Agent in 3D WorldICML 2024project
2023-08-083D Visual Grounding, Referring Expression, Point CloudPeking University3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentICCV 2023github
2023-07-243D VQA, 3D Captioning, Scene UnderstandingShanghai AI Lab3D-LLM: Injecting the 3D World into Large Language ModelsNeurIPS 2023project / github
2023-033D Language Pre-training, Captioning, QAAuthors3D-VLP: 3D Vision-Language Pre-training with Contextual SceneCVPR 2023paper

1.3 Active Imaging & Sensors

1.3.1 Structured Light & Active Stereo

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-29iToF, Sensor-Intrinsic Uncertainty, Heteroscedastic RestorationTsinghua UniversityReliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space RestorationarXivpaper
2026-07-27Neural Structured Light, Metric Depth, Online SLAMPeking UniversityNSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and ReconstructionACM MM 2026paper
2026-07-20Projector-Camera, Feed-Forward 3DGS, Active IlluminationNingbo UniversityFF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera SystemarXivgithub
2026-06Active Stereo, 2DGS Supervision, RealSense DatasetHangzhou Dianzi UniversityGS-ASM: 2DGS-Supervised Active Stereo MatchingCVPR 2026paper
2026-05-07Adaptive 4D Illumination, Shape+Reflectance, Differentiable CaptureZhejiang UniversityDifferentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and ReflectanceCVPR 2026paper
2026-03Multi-Projector Structured Light, One-Shot Scan, Neural SDFKyushu UniversityMulti-view Stereo with Multiple Projectors for Oneshot Entire Shape Scan based on Neural SDF and DSSS DemultiplexingWACV 2026paper
2026-02-05Latent Diffusion, Fringe Projection, Reflective ObjectsYonsei UniversityLD-SLRO: Latent Diffusion Structured Light for 3-D Reconstruction of Highly Reflective ObjectsarXivpaper
2025-12-16Single-Shot, Neural Feature Decoding, Robust CorrespondencePeking UniversityRobust Single-shot Structured Light 3D Imaging via Neural Feature DecodingSIGGRAPH Asia 2025project
2025-12Event Camera, HDR Measurement, Confidence StereoUSTCEvent-based HDR Structured LightNeurIPS 2025github
2025-03-08Active Stereo, Phase Speckle, Cross-Scene GeneralizationSouthwest Jiaotong UniversityRGB-Phase Speckle: Cross-Scene Stereo 3D Reconstruction via Wrapped Pre-NormalizationarXivpaper
2025-02Unsupervised Structured Light, Neural SDF, Shadow-AwareKyushu UniversityNeural SDF for Shadow-Aware Unsupervised Structured LightWACV 2025paper
2025-01-13Matching-Free, Volume Rendering, Monocular Structured LightUSTCMatching-Free Depth Recovery from Structured LightarXivpaper
2024-10-20Neural SDF, One-Shot Scan, Low-Light/UnderwaterKyushu UniversityActiveNeuS: Neural Signed Distance Fields for Active Stereo3DV 2024paper
2024-06-06Virtual Pattern Projection, Depth Fusion, In-the-Wild StereoUniversity of BolognaActive Stereo in the Wild through Virtual Pattern ProjectionarXivgithub
2024-06Neural Inverse, Dense Depth, 3-4 PatternsUniversity of TorontoTurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured LightCVPR 2024project
2023-06-17Structured Light, Phase Unwrapping, LearningNanjing UniversityDeep Learning-Based Structured Light 3D Imaging: A SurveyarXivsurvey
2018-11-27Event Camera, Active Stereo, DepthTsinghua UniversityEvent-Based Structured Light for Depth ReconstructionIJCAS 2024paper

1.3.2 Panoramic & Fisheye Perception

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-16Wide-FOV, Egocentric, 4D Hand-ObjectRice UniversityEgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot LearningarXivproject / dataset
2026-03-19Panoramic Depth, VGGT, Geometry ConsistencySingapore University of Technology and DesignVGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth EstimationCVPR 2026paper
2026-03-18Panoramic Reconstruction, Permutation-Equivariant, 360°CornellPanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryCVPR 2026paper
2026-01-25RGB-D, Depth Completion, Reflective/TransparentRobbyantLingBot-Depth: Masked Depth Modeling for Spatial PerceptionarXivgithub / paper
2025-12-18Panoramic Foundation Model, Multi-Camera, Metric DepthInsta360 ResearchDepth Any Panoramas: A Foundation Model for Panoramic Depth EstimationCVPR 2026paper
2025-01Fisheye, Real-Time, Cassini, Multi-ViewSun Yat-Sen UnivOmniStereo: Real-time Omnidirectional Depth with Multiview Fisheye CamerasCVPR 2025github
2024-06-19Panoramic Depth, Semi-Supervised, MobiusHKUST(GZ)PanDA: Panoramic Depth Anything with Mobius Spatial AugmentationCVPR 2025project
2024-03-25360 Depth, Bi-Projection, ERP+ICOSAPHKUST(GZ)Elite360D: Efficient 360 Depth Estimation via Bi-Projection FusionCVPR 2024github
2021-09-06360 Depth, Indoor, Panoramic ImagesCERTHPano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth EstimationCVPRW 2021project / github

1.3.3 Event-Based & Time-of-Flight Imaging

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-26Active Event Stereo, High-Speed Depth, 150 FPSAuthorsTowards Ultrafast Depth Sensing Via Active Event-based Stereo VisionTPAMI 2026paper
2026-07-17Event Camera, Feed-Forward 3D, Temporal AggregationZhejiang UniversityEvent3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature AggregationarXivpaper
2026-06RGB+ToF Histogram, High-Resolution Metric Depth, LightweightTongji UniversityLiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth EstimationCVPR 2026 Highlightpaper
2026-06Sparse dToF, Zero-Shot Completion, Sensor GeneralizationKAISTDense Metric Depth Completion from Sparse Direct Time-of-Flight SensorsCVPR 2026paper
2026-06Image-Event Fusion, Monocular Depth, Linear ComplexityPeking UniversityAIMDepth: Asymmetric Image-Event Mamba for Monocular Depth EstimationCVPR 2026paper
2026-06Event-Image Depth, Hypothesis Volume, Iterative RefinementSoutheast UniversityDepth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth EstimationCVPR 2026paper
2026-04-16Event-Frame Stereo, Cross-Modal Prompting, High Dynamic RangeSoutheast UniversityBidirectional Cross-Modal Prompting for Event-Frame Asymmetric StereoCVPR 2026paper
2026-04-02Event Stereo, Data Factory, Cross-Modal DistillationUniversity of BolognaEventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active SensorsCVPR 2026project
2026-02-03Event Camera, Neural SDF, Single-Camera MeshSaarland UniversityEventNeuS: 3D Mesh Reconstruction from a Single Event Camera3DV 2026project / github / dataset
2025-12-20Event Camera, Structured Light, Real-Time RGB-DPolytechnique MontrealE-RGB-D: Real-Time Event-Based Perception with Structured LightarXivgithub
2025-09-18Event Camera, Depth Estimation, Any-to-AnyShanghai AI LabDepth AnyEvent: Event Camera Based Monocular Depth Estimation via Dense Correspondence DistillationarXivpaper
2025-09-08Event Camera, Multispectral, Structured LightETH ZurichEvent Spectroscopy: Event-based Multispectral and Depth Sensing using Structured LightarXivpaper
2025-05-28Burst-Encodable ToF, Long-Range Depth, Hardware-Aware CodingNanjing UniversityLearnable Burst-Encodable Time-of-Flight Imaging for High-Fidelity Long-Distance Depth SensingNeurIPS 2025github
2025-05Event, Distillation, Confidence-Guided, Pseudo-LabelsNUSDistil-E2D: Distilling Image-to-Depth Priors for Event-Based DepthNeurIPS 2025paper
2025-04-23ToF, Sparse Depth, 3DGS, SLAMAuthorsToF-Splatting: Dense SLAM using Sparse Time-of-Flight DepthICCV 2025paper / paper
2025-04-22Event Camera, Ray Density, 3D Conv, SpotlightTU BerlinDERD-Net: Learning Depth from Event-based Ray DensitiesNeurIPS 2025 Spotlightgithub
2025-03-03dToF, Video Depth Completion, Frequency SelectiveAuthorsSVDC: Consistent Direct Time-of-Flight Video Depth CompletionCVPR 2025paper / paper
2024-10-10Event Camera, Pose-Free, Gaussian Splatting, HighlightZhejiang UnivIncEventGS: Pose-Free Gaussian Splatting from a Single Event CameraCVPR 2025 Highlightgithub

1.3.4 Radar & RF Imaging

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10mmWave Radar, Complex-Valued Point Splatting, Material Model, Novel ViewsCornell Tech3D Point Splatting for mmWave Radar Novel View SynthesisarXivpaper
2026-09-09mmWave Radar, Metric Depth, Smoke/Fog/Darkness, 95K FramesRice UniversityGRADE: Single-Frame Generative Radar Depth Estimation Under Visual DegradationMobiCom 2026project+data

1.4 Dense Mapping Systems

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-09RGB-D 3DGS, Online Reconstruction, Reactive ControlMitsubishi Electric Research LaboratoriesSplatCtrl: Perception-Action Coupling via Gaussian Scene Representations and Reactive Robot ControlICRA 2026paper
2026-07-08Geometry-Only 3DGS, Dense Monocular SLAMBeihang UniversityGeoGS-SLAM: Geometry-Only Gaussian Splatting for Dense Monocular SLAMarXivpaper
2026-07-05LiDAR 3DGS SLAM, Geometry-Aware Covariance, Real-TimeAuthorsReal-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance CouplingarXivgithub
2026-07-02Dynamic Gaussian SLAM, Dual-Level Probability, Semantic MapAuthorsDL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level ProbabilityarXivpaper
2026-06-29Task-Conditioned 3DGS, Real-Time Mapping, Multi-Agent FusionMITGaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic MappingarXivpaper
2026-06-29RGB-Only Gaussian SLAM, Closed-Loop Geometry, Scale FeedbackCAS / USTCMyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAMIROS 2026paper
2026-06-27Object-Centric 3DGS, Lifelong Mapping, Dynamic MaintenanceBeijing Institute of TechnologyCubifyGS: Object-Centric 3D Gaussian Splatting for Lifelong Dynamic Scene MaintenanceIROS 2026paper
2026-03-10Uncertainty-Aware 3DGS, RGB-D SLAM, Loop DetectionGeorge Mason UniversityVarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAMCVPR 2026project / github
2025-11-20Language-Embedded, Open-Vocabulary, 3DGS SLAMKAISTLEGO-SLAM: Language-Embedded Gaussian Optimization SLAMarXivpaper
2025-07-25Neural SLAM, Dense Mapping, Self-SupervisedShanghai AI LabDINO-SLAM: Dense Tracking and Mapping with Self-Supervised Feature LearningarXivpaper
2025-07Dynamic Surface, Non-Rigid, 4D TrackingImperial College4DTAM: Dynamic Surface Gaussian SLAMCVPR 2025paper
2025-03-204DGS SLAM, Dynamic/Static, TrackingAuthors4D Gaussian Splatting SLAMICCV 2025paper / paper
2025-03-11Gaussian SLAM, Dense Reconstruction, RGB-DZhejiang UniversityGigaSLAM: Gaussian Splatting-based Large-Scale Dense SLAMarXivgithub
2025-03Multi-Agent, 3DGS SLAM, Loop ClosureAuthorsMAGiC-SLAM: Multi-Agent 3DGS SLAMCVPR 2025paper
2025-01-25Gaussian SLAM, Dynamic Environments, MonocularStanford / ETHWildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic EnvironmentsCVPR 2025project
2025-01-24Multi-Robot, Semantic, Heterogeneous, 3DGSStanfordHAMMER: Heterogeneous, Multi-Robot Semantic Gaussian SplattingRAL 2025project
2024-11-03Gaussian SLAM, Global BA, Monocular RGBETH / MetaSplat-SLAM: Globally Optimized RGB-only SLAM with 3D GaussiansCVPR 2025project
2024-09-10Single-Image Calibration, Geometric OptimizationETH / MetaGeoCalib: Learning Single-image CalibrationECCV 2024paper
2024-04-11Detector-Free SfM, Texture-PoorZhejiang UnivDetector-Free Structure from MotionCVPR 2024github
2024-04Pose Regression, Map-Relative, Multi-Scene, HighlightNiantic / OxfordMarepo: Map-Relative Pose Regression for Visual Re-LocalizationCVPR 2024 Highlightgithub
2024-02-20Neural SLAM, Survey, Radiance FieldsTUMHow NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a SurveyarXivproject
2023-12-11Dense SLAM, Gaussian Splatting, RGB-DTUMGaussian Splatting SLAMCVPR 2024github
2023-12-07End-to-End SfM, Differentiable BAMeta AI / OxfordVGGSfM: Visual Geometry Grounded Deep SfMCVPR 2024github
2023-12-04Gaussian SLAM, RGB-D, VolumetricCMU / MITSplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAMCVPR 2024project / github
2021-12-22Neural Mapping, Dense RGB-D SLAM, SDFHKUSTNICE-SLAM: Neural Implicit Scalable Encoding for SLAMCVPR 2022project

🧱 2. 3D/4D Representation

This section focuses on the mathematical and data-structure layer used to represent geometry, appearance, and motion.

2.1 Explicit & Hybrid Representations

2.1.1 Structure-Aware 3D Gaussian Splatting

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-21Gaussian Surfels, Topology Recovery, Mesh Extraction, Surface ReconstructionUniversity of Science and Technology of ChinaTopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface ReconstructionarXivgithub
2026-08-20Sparse Light Field, Casual Capture, 3D/4D Reconstruction, 3DGSCornell UniversitySparse Light Field Sampling Improves Casual 3D and 4D ReconstructionarXivproject
2026-08-17Pose-Free NVS, 3DGS Geometry, Visibility-Aware GuidanceAalto UniversitySplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View SynthesisarXivpaper
2026-08-13Feed-Forward 3DGS, 3D Anchors, Spatially Grounded TokensNational University of Defense TechnologyLocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian SplattingarXivproject
2026-08-03Streaming Feed-Forward, Persistent Geometry, Memory-Bounded 3DGSUniversity of Science and Technology of ChinaStreamSplat: Streaming Feed-Forward 3D Gaussian SplattingarXivpaper
2026-08-03View-Conditioned, Feed-Forward 3DGS, Generalizable ReconstructionTsinghua UniversityUniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D ReconstructionarXivpaper
2026-07-28Dynamic 3DGS, Adaptive Streaming, Volumetric VideoAuthorsSplatStream: Fine Granular Scalable Gaussian Splatting for Adaptive 3D Scene StreamingarXivpaper
2026-07-22Feed-Forward 3DGS, Adaptive Tokens, Compact RepresentationYonsei UniversityATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token ExpansionarXivproject
2026-04-16Feed-Forward 3DGS, Global Scene Tokens, CompactTel Aviv UniversityGlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene TokensarXivpaper
2026-04-16TokenGS, Learnable Gaussian Tokens, Pose-RobustNVIDIATokenGS: Decoupling 3D Gaussian Prediction from PixelsCVPR 2026 Highlightproject
2026-04-12UniSplat, Unposed Multi-View, Feed-ForwardUC BerkeleyUniSplat: Learning 3D Representations from Unposed Multi-View ImagesCVPR 2026paper
2026-02-02Feed-Forward 2DGS, Surface Continuity, Sparse ViewsShanghai Jiao TongSurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity PriorsICLR 2026paper
2025-07-31Sparse-View 4D, Monocular Fusion, Cross-VideoMeta AIMonoFusion: Sparse-View 4D Reconstruction via Monocular FusionICCV 2025paper
2025-06-17Anti-Aliasing, Gaussian, 3D, AdaptiveCUHKAAA-Gaussians: Anti-Aliasing 3D Gaussian SplattingICCV 2025paper
2025-06-05Feed-Forward 3DGS, Depth Parameterization, GeometryZhejiang UniversityRevisiting Depth Representations for Feed-Forward 3D Gaussian Splatting3DV 2026paper
2025-05-21Sparse-View Surface, Geometry-Prioritized, 2DGSNankai UnivSparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsCVPR 2025paper
2025-04-03Feed-Forward, No Camera, FreeSplatterZJUFreeSplatter: Pose-Free 3DGS from Sparse ViewsCVPR 2025paper
2025-03-13Few-Shot, Diffusion Prior, Repair + InpaintingZJU / AlibabaRI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion PriorsICCV 2025paper
2025-01-11Scene-Level, 3DGS, Open-Vocabulary, 4D LangSplatUPenn4D LangSplat: 4D Language Gaussian SplattingCVPR 2025paper
2024-12-09Large-Scale Dynamic, City, 4DGS, SpotlightZJUDynamicCity: Large-Scale 4D Gaussian City ModelingICLR 2025 Spotlightpaper
2024-09-26Progressive, Pruning, 3DGS, CVPRS-LabPUP 3D-GS: Progressive Pruning for 3DGSCVPR 2025paper
2024-03-262DGS, Surface Reconstruction, GeometryTUM2D Gaussian Splatting for Geometrically Accurate Radiance FieldsSIGGRAPH 2024project
2024-03-21Feed-Forward 3DGS, Sparse Views, GeneralizableMonash / TübingenMVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View ImagesECCV 2024 Oralproject / github
2024-03-14Multi-Scale, 3DGS, Level-of-DetailShanghai Jiao TongMulti-Scale 3DGSCVPR 2024paper
2023-12-19Feed-Forward 3DGS, Image Pairs, EpipolarMITpixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D ReconstructionCVPR 2024 Oralgithub
2023-12-01Distillation, Lightweight, 3DGS, SpotlightVirginia TechLightGaussian: Distilled 3D Gaussian SplattingNeurIPS 2024 Spotlightpaper
2023-11-303DGS, Structure, Sparse ViewsUSTCSparseGS: Real-Time 360 Sparse View Synthesis using Gaussian Splatting3DV 2025project
2023-11-27Anti-Aliasing, Mip, 3DGS, Best Student PaperInriaMip-Splatting: Alias-free 3DGSCVPR 2024 Oral / Best Student Papergithub
2023-11-27Compression, Compact, 3DGS, HighlightTsinghuaCompact 3D Gaussian SplattingCVPR 2024 Highlightpaper
2023-11-213DGS, Mesh Extraction, SurfaceETH ZurichSuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh ReconstructionCVPR 2024project
2023-08-083DGS, Real-Time Rendering, Explicit RadianceInria3D Gaussian Splatting for Real-Time Radiance Field RenderingSIGGRAPH 2023github / paper

2.1.2 Voxel, Mesh & Point-Cloud Innovations

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-18Sparse Voxel Latent, Fitting-Free 3DGS, Large-Area GenerationAmap, AlibabaGS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS GenerationarXivpaper
2026-07-22Shape Completion, Unified 3D Representation, Faithful GeometryAuthorsAxolotl3D: A Unified Framework for Faithful 3D Shape CompletionarXivpaper
2024-12-02Structured Latent, 3D Generation, MeshMicrosoftTRELLIS: Structured 3D Latents for Scalable and Versatile 3D GenerationCVPR 2025project / github
2024-03-04Sparse Voxel, Text/Image-to-3D, MeshStability AITripoSR: Fast 3D Object Reconstruction from a Single ImagearXivgithub
2022-12-16Point Cloud, Diffusion, Shape GenerationOpenAIPoint-E: A System for Generating 3D Point Clouds from Complex PromptsarXivgithub

2.2 Neural Implicit Representations

2.2.1 Advanced NeRF & SDF Architectures

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2025-07-15Unbiased SDF, Neural Implicit Surface, SDF-to-DensityCUHKUNIS: Unified Framework for Unbiased Neural Implicit SurfacesICCV 2025paper
2024-09-06α-NeuS, Alpha, SDF, Volume RenderingZJUα-NeuS: Alpha-Governed Neural Implicit SurfacesNeurIPS 2024paper
2023-12-29Objects as Volumes, NeRF, 3D Reconstruction, OralUPennObjects as Volumes: Feed-Forward 3D from a Single ImageCVPR 2024 Oralpaper
2023-04-13Zip-NeRF, Anti-Aliasing, Mip, GridGoogle ResearchZip-NeRF: Anti-Aliased Grid-Based NeRFICCV 2023project
2022-03-17Tensor Factorization, Radiance Fields, CompressionTsinghuaTensoRF: Tensorial Radiance FieldsECCV 2022project
2022-01-16Hash Grid, Real-Time NeRF, Neural GraphicsNVIDIAInstant Neural Graphics Primitives with a Multiresolution Hash EncodingSIGGRAPH 2022github / paper
2021-12-15EG3D, Triplane, 3D GAN, GenerativeNVIDIAEG3D: Efficient Geometry-aware 3D Generative Adversarial NetworksCVPR 2022 Oralgithub
2021-12-09Plenoxels, No Neural Network, Fast, OralUC BerkeleyPlenoxels: Radiance Fields without Neural NetworksCVPR 2022 Oralgithub
2021-12-07Ref-NeRF, Reflection, Specular, Best Student Paper HMGoogle ResearchRef-NeRF: Structured View-Dependent Appearance for NeRFCVPR 2022 Best Student Paper HMproject
2021-11-23Unbounded Scenes, Anti-Aliasing, NeRFGoogle ResearchMip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsCVPR 2022project
2021-06-20SDF, Surface Reconstruction, Neural RenderingMPI-ISNeuS: Learning Neural Implicit Surfaces by Volume RenderingNeurIPS 2021project
2021-04-13BARF, Bundle-Adjusting NeRF, OralUC BerkeleyBARF: Bundle-Adjusting Neural Radiance FieldsICCV 2021 Oralgithub
2020-08-05NeRF-W, Unbounded, In-the-WildGoogle ResearchNeRF in the Wild: Neural Radiance Fields for Unconstrained Photo CollectionsCVPR 2021github
2020-03-19NeRF, View Synthesis, Neural RenderingUC BerkeleyNeRF: Representing Scenes as Neural Radiance Fields for View SynthesisECCV 2020github / paper

2.3 Dynamic & Spatiotemporal 4D

2.3.1 4D Gaussian Splatting (4DGS)

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-23Dynamic 3DGS, Gradient Decoupling, Novel View SynthesisAuthorsGrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View SynthesisarXivpaper
2026-06-22Dynamic 3DGS, Visibility-Aware Densification, Temporal LifespanIndian Institute of ScienceTemporally Aware Densification for Dynamic 3D Gaussian SplattingarXivpaper
2025-077DGS, Spatial-Temporal-Angular, UnifiedAuthors7D Gaussian Splatting: Unified Spatial-Temporal-Angular GSICCV 2025paper
2024-10Event Camera, High-Speed, 4D, STD-GSAuthorsSTD-GS: SpatioTemporal-Disentangled Gaussian Splatting with Event CamerasICCV 2025paper
2023-10-124DGS, Dynamic Scenes, Real-Time RenderingZhejiang University4D Gaussian Splatting for Real-Time Dynamic Scene RenderingCVPR 2024project
2023-08-18Dynamic 3DGS, Scene Motion, Multi-View VideoCornellDynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis3DV 2025github / paper
2023-01-24K-Planes, Explicit, 4D, Space-TimeAuthorsK-Planes: Explicit Radiance Fields in Space, Time, and AppearanceCVPR 2023github
2023-01-23HexPlane, 4D Representation, Space-TimeCMUHexPlane: A Fast Representation for Dynamic ScenesCVPR 2023project
2021-06-24Dynamic NeRF, Deformation, Canonical SpaceGoogle ResearchHyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance FieldsSIGGRAPH Asia 2021project

2.3.2 Deformation Graphs & Canonical Spaces

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-24Structured Motion, SE(3), 4D ReconstructionTsinghua UniversitySM4RT: Learning Structured Motion Geometry for 4D ReconstructionarXivproject / paper
2023-12-04Deformation, Dynamic Radiance Fields, CanonicalETH ZurichSC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic ScenesCVPR 2024project
2023-06-05Non-Rigid Tracking, Neural Deformation, 4DTsinghuaNeuralangelo: High-Fidelity Neural Surface ReconstructionCVPR 2023project

🏗️ 3. 3D Reconstruction

Reconstruction systems recover objects or scenes from images, video, RGB-D, or multi-sensor streams under offline, online, and dynamic conditions.

3.1 Static Reconstruction

3.1.1 Object-Centric Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Active In-Hand Reconstruction, Uncertainty, Next-Best ViewShanghaiTech UniversityAURORA: Active Uncertainty-Driven Re-Orientation for In-Hand ReconstructionCoRL 2026project
2026-07-21Single-View 3D, Object Perception, Generative ReconstructionDeakin UniversitySeeing Before Generating: Object Perception Enhances Single-View 3D ReconstructionarXivproject
2026-06-17Sparse-View Object, Flow Steering, 3DGS RefinementGraz University of TechnologyFlowObject: Flow Steering for Bridging Generative Priors and Reconstruction FidelityarXivproject
2026-05-05Generative Reconstruction, Multi-View Alignment, PoseTsinghuaMix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose EstimationarXivproject
2025-11-19Single Image, Object Mesh, SAM 3DMeta AISAM 3D ObjectsGitHubgithub
2025-10-23Pose-Free Online, Free-Moving Objects, Constant MemorySUTDOnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving ObjectsNeurIPS 2025 Spotlightproject
2025-06-05RGB-D Object Completion, Novel Depth, Feed-ForwardCarnegie Mellon UniversityRaySt3R: Predicting Novel Depth Maps for Zero-Shot Object CompletionNeurIPS 2025project
2025-06Feed-Forward, Densification, Gaussian, DetailAuthorsGenerative Densification: Feed-Forward 3DGS DensificationCVPR 2025paper
2025-06Photogrammetry Foundation, Multi-Task, HighlightAuthorsMatrix3D: A Foundation Model for PhotogrammetryCVPR 2025 Highlightpaper
2025-04-04Sparse-View, Feed-Forward, Camera, GeometryAnt Research / StanfordFLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsCVPR 2025project
2024-07Ego-Centric, Autonomous Driving, Sparse-ViewAuthorsOmni-Scene: Omni-Gaussian for Ego-Centric Sparse-View ReconstructionCVPR 2025paper
2024-06-14Multi-View, Stereo, Feed-ForwardNAVER LabsMASt3R: Grounding Image Matching in 3D with MASt3RECCV 2024project / github
2023-12-21Multi-View, Pointmap, Pose-FreeNAVER LabsDUSt3R: Geometric 3D Vision Made EasyCVPR 2024project / github
2023-12-13Object Pose, Reconstruction, Model-BasedNVIDIAFoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsCVPR 2024project
2023-03-24Unknown Object, RGB-D, 6-DoF Tracking, Neural SDFNVIDIABundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown ObjectsCVPR 2023project / github

3.1.2 Large-Scale Scene Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Feed-Forward, Compositional Scene, Complete Meshes, Simulation-ReadyUniversity of Illinois Urbana-ChampaignFIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A MinutearXivproject / github
2026-09-04Bundle Adjustment, Multi-View Matching, Monocular Priors, Online+OfflineNAVER Labs EuropeBLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular PriorsECCV 2026paper
2026-08-31Real-to-Sim, Parse-Generate-Place, Composable Object AssetsByteDance SeedLucida: Parse, Generate, and Place for Composable Real-to-Sim Scene ModelingarXivproject
2026-08-25Single Image, Generative Reconstruction, Complete Object Assets, Scene AssemblyHuaweiSceneReGen: Generative Reconstruction of 3D Scenes from a Single ImagearXivpaper
2026-08-18Instance-Grouped 3DGS, Semantic Reconstruction, Referential Scene GraphShanghai Jiao Tong UniversityGroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian SplattingarXivpaper
2026-08-18Generative NVS, Reconstruction/Generation Split, Scene CoordinatesKAISTGenRec: Knowing Where to Reconstruct and Where to GeneratearXivproject
2026-08-18Long Sequence, Chunk Priors, Sim(3) Assembly, Test-Time AdaptationKosmo ResearchGeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric AssemblyarXivproject
2026-08-15Long Sequence, Scale-Consistent Alignment, Test-Time AdaptationNorthwestern Polytechnical UniversityVGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D ReconstructionACM Multimedia 2026github
2026-08-12Streaming Multi-View, Metric 3D, Feed-Forward Prior, Object DetectionETH ZurichMap-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming InputsECCV 2026project
2026-08-07Sparse View, Editable Indoor Scenes, Executable Scene ProgramsCity University of Hong KongScenix: Sparse-View 3D Scene Reconstruction via Executable Scene ProgramsarXivpaper
2026-07-31Active Reconstruction, Next-Best-View, Predictive EntropyFudan UniversityGO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D ReconstructionarXivpaper
2026-07-23Underwater 3D, Feed-Forward Reconstruction, Degradation AdaptationHKUSTWAT3R: Feedforward Underwater 3D ReconstructionarXivproject
2026-07-15Feed-Forward Driving Reconstruction, Layered 3DGS, Dynamic ActorsNVIDIAInstant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene SimulationarXivproject / github / docs
2026-07-103D Foundation Model, Global SfM, Bundle AdjustmentHKUSTGlob3R: Global Structure-from-Motion with 3D Foundation ModelsarXivproject
2026-07-08Feed-Forward 3D, Unposed Images, Drift-RobustAuthorsNoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D ReconstructionECCV 2026paper
2026-06-09RGB-T, Thermal Geometry, Low-LightUniversity of MinnesotaDarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight TaxarXivproject
2026-06-02Single Image, Physics-in-the-Loop, Simulation-ReadySeoul National UniversitySimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single ImagearXivproject
2026-05-14VGGT, Scaling, Static+DynamicUniversity of OxfordVGGT-Omega: Scaling VGGT to Large-Scale 3D ReconstructionCVPR 2026 Oralproject / github / demo / model
2026-05-07Feed-Forward 3D, Token Reduction, Long SequencePeking UniversitySpark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D ReconstructionarXivpaper
2026-04-30Generalizable, Sparse-View, Unposed Images, OutdoorUIUC / NVIDIAGenWildSplat: Generalizable Sparse-View 3D Reconstruction from Unconstrained ImagesarXivpaper
2026-03-24Panoramic Video, Pose-Free 3DGS, Consistent Depth PriorUniversity of Chinese Academy of SciencesPose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth PriorsCVPR 2026github
2026-03-16Event-to-Edge, Pose-Free, Gaussian ReconstructionKAISTE2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D ReconstructionCVPR 2026paper
2026-02-26VGGT, TTT, Large-ScaleNVIDIAVGG-T^3: Offline Feed-Forward 3D Reconstruction at ScaleCVPR 2026project
2026-02-03Single Image, Object Decomposition, Occlusion-Aware Scene ReconstructionUniversity of California, San DiegoSeeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal3DV 2026project
2025-09-24Mirror Stereo, Single-View 3D, Symmetry ConstraintUniversity of OxfordReflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections3DV 2026project / github / dataset
2025-09-16Universal 3D, Metric Reconstruction, Optional PriorsMeta AIMapAnything: Universal Feed-Forward Metric 3D Reconstruction3DV 2026project
2025-08-05Unposed Multi-View, 3DGS, Semantic ReconstructionSungkyunkwan UniversityUni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesCVPR 2026paper
2025-07-17Permutation-Equivariant, Visual Geometry, Point MapsOxford / Metaπ³: Permutation-Equivariant Visual Geometry LearningICLR 2026github
2025-07Monocular Prior, MVS, DTU/Tanks SOTAAuthorsMonoMVSNet: Monocular Prior Guided MVSICCV 2025paper
2025-07Latent Align, Stereo+Monocular, HighlightAuthorsBridgeDepth: Unified Monocular and Stereo DepthICCV 2025 Highlightpaper
2025-06-30Video-Depth Augmentation, Scalable Training, Feed-Forward 3DAustralian National UniversityPuzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D ReconstructionNeurIPS 2025project
2025-06-03Monocular Depth, Dynamic Video, AlignmentHKUST / CUHK / HKUAlign3R: Aligned Monocular Depth Estimation for Dynamic VideosCVPR 2025github
2025-06Aerial-Ground, Large-Scale, 3DGSAuthorsHorizon-GS: Unified Aerial-Ground 3DGSCVPR 2025paper
2025-06Autonomous Driving, Feed-Forward, 3DGSAuthorsEVolSplat: Feed-Forward 3DGS for Urban DrivingCVPR 2025paper
2025-06Sparse-View, Super-Resolution, 3DGSAuthorsS2Gaussian: Sparse-View Super-Resolution 3DGSCVPR 2025paper
2025-05-05Relative Camera Pose, Regression, LocalizationAalto / HKUReloc3r: Large-Scale Training of Relative Camera Pose RegressionCVPR 2025github
2025-03-28Feed-Forward, Surface, MVS, Multi-ViewAuthorsMVSAnywhere: Zero-Shot Multi-View StereoCVPR 2025paper / paper
2025-03-17Feed-Forward, Multi-View, Auxiliary PriorsETH / MicrosoftPow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene PriorsCVPR 2025project
2025-03-14Feed-Forward 3D, Pose-Free, Point MapMeta AIVGGT: Visual Geometry Grounded TransformerCVPR 2025 Best Paperproject / github
2025-03-03Multi-View, Symmetric, 1000+ Images, O(N)NAVER LabsMUSt3R: Multi-view Network for Stereo 3D ReconstructionCVPR 2025project
2025-01-23Feed-Forward, 1500+ Images, 251 FPSMeta AIFast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassCVPR 2025project
2024-12-12Feed-Forward, Online, Dense, MonocularShanghai AI Lab / PKUSLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosCVPR 2025 Highlightproject
2024-12-09Sparse View, Single-Stage, 2 SecondsMeta AIMV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsCVPR 2025 Oralproject
2024-09-27Multi-View Reconstruction, Matching, MVSNAVER LabsMASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion3DV 2025github
2024-08-28Spatial Memory, Feed-Forward, Multi-ViewHKUSpann3R: 3D Reconstruction with Spatial Memory3DV 2025 Oral (Best Paper Candidate)project

3.2 Streaming & Online Reconstruction

3.2.1 Dense Semantic/Instance Mapping

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-07Online SLAM, Functional Scene Graph, Interaction Elements, Map MemoryTsinghua UniversityFunctional-SLAM: Interaction-Aware Mapping with Online Functional Scene GraphsarXivgithub
2026-09-01Training-Free, Open-Vocabulary Instance Map, RGB-D/Monocular SLAMUniversity of Technology SydneyVOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAMarXivpaper
2026-08-24Spatio-Temporal SLAM, Open-Vocabulary, 4D Scene Graph, VLNCarnegie Mellon UniversitySuperMap: A Spatio-Temporal SLAM System for Visual-Language NavigationarXivproject
2026-08-18Open-Vocabulary Map, Instance Preservation, Fine-Grained Retrieval, Target AbsenceXi'an Jiaotong UniversityOVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained ObjectsarXivgithub
2026-06-23Object-Level Map, Open-Vocabulary, RelocalizationZhejiang UniversityCompact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual RelocalizationarXivpaper
2026-05-05Online, Voxel+Instance, Open-Vocabulary MappingÖrebro UniversityFUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic MappingarXivproject
2026-05-03Gaussian-Language Map, Zero-Shot Navigation, Multi-ScaleCASIAMulti-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and ReasoningarXivgithub
2026-03-04Semantic 3DGS, Online, CLIP, Open-VocabularyNational University of SingaporeEmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingCVPR 2026project
2025-08-02Open-Vocabulary, Hybrid 3DGS+TSDF, Dense MappingTsinghuaOpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian SplattingarXivproject
2025-07Feed-Forward, Panoramic Segmentation, DUSt3RAuthorsPanSt3R: Single-Feed 3D and Panoptic SegmentationICCV 2025paper
2023-10-05Open-Vocabulary, RGB-D, TSDF, Real-Time MappingUniversity of ArkansasOpen-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationIROS 2024project / github
2023-02-14Open-Set, Multimodal, 3D Map, Language QueryMITConceptFusion: Open-set Multimodal 3D MappingICRA 2023project / github
2022-10-11Implicit Field, CLIP, Semantic Search, Robot MemoryNew York UniversityCLIP-Fields: Weakly Supervised Semantic Fields for Robotic MemoryICRA 2023project

3.2.2 Continuous Neural Tracking & Mapping

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-03Online 3R, Multi-Relative Pose Query, Pose-Graph OptimizationNational Yang Ming Chiao Tung UniversityScal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D ReconstructionECCV 2026project
2026-09-01Online Feed-Forward 3R, Unordered UAV Images, Retrieval+RetryWuhan UniversityOn-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV ScenariosarXivgithub
2026-08-03Active Reconstruction, Ergodic Coverage, Trajectory OptimizationJohns Hopkins UniversityTRACE: Ergodic Trajectory Optimization for Active Scene ReconstructionarXivgithub
2026-08-03Feed-Forward SLAM, Sim(3) Factor Graph, Persistent MappingUlsan National Institute of Science and TechnologyUniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) OptimizationECCV 2026project
2026-07-25Semantic SLAM, Data Association, Object LandmarksMITSemantic Semi-Incremental Data-Association-Free Object SLAMarXivpaper
2026-07-23Gaussian SLAM, Large-Scale Mapping, Real-TimeAthena Research CenterGLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial DecompositionIROS 2026github
2026-07-16Multi-Agent 3R, RGB Video, Point-Map FusionUniversity of BolognaMAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB VideosarXivproject
2026-07-16Incremental 3DGS, Unordered Capture, Global ConsistencyInriaImmediate 3D Gaussian Splat Reconstruction of Unordered Input with Global ConsistencySIGGRAPH 2026paper
2026-07-01Long-Sequence, Instance Anchors, Persistent Spatial MemoryBeijing Jiaotong UniversityLIST3R: Long-sequence Instance-aware 3D ReconstructionarXivproject
2026-06-233DGS-SLAM, Memory-Efficient, Outdoor MappingUniversity of MinnesotaPocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAMICRA 2026github
2026-06-20RGB+Pose, 3DGS Scene Regression, Robot CapturePeking UniversityACEsplat: Accelerated 3D Gaussian Scene Regression via RGB and Poses OnlyarXivpaper
2026-06-193DGS-SLAM, Degeneracy-Robust, Real-Time TrackingNanyang Technological UniversitySpectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAMIROS 2026paper
2026-06-18LiDAR-Inertial-Thermal, 3DGS Mapping, Illumination-RobustShenzhen UniversityLIT-GS: LiDAR-Inertial-Thermal Gaussian Splatting for Illumination-Robust MappingIROS 2026paper
2026-06-03Streaming, Transient Anchors, Long-Horizon MappingAuthorsAnchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual MappingarXivpaper
2026-06Spatial-Difference Sensor, Edge-Guided Tracking, 3DGSTsinghua UniversitySDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D ReconstructionCVPR 2026paper
2026-06Stereo 3DGS-SLAM, Auto-Exposure Robustness, Photometric MappingSouth China University of TechnologyAERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAMCVPR 2026github
2026-05-10VGGT, Retrieval, Constant MemoryFudan UniversityRetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity RetrievalarXivgithub
2026-04-244DGS-SLAM, Optical Flow, Dynamic MappingNational University of SingaporeFlow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAMCVPR 2026github
2026-04-15Streaming, Feed-Forward, Long VideoRobbyantLingBot-Map: Geometric Context Transformer for Streaming 3D ReconstructionarXivgithub
2026-02-13Streaming, Autoregressive, Long Sequence3DAgentWorldLongStream: Long-Sequence Streaming Autoregressive Visual GeometryCVPR 2026project
2026-01-03StreamVGGT, KV Cache, Memory CompressionSun Yat-sen UniversityXStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded TransformerarXivgithub
2025-09-30TTT, Online, Long ContextShanghai AI LabTTT3R: 3D Reconstruction as Test-Time TrainingarXivproject
2025-08-14Streaming, Causal Transformer, SequentialNTU / Shanghai AI LabSTream3R: Scalable Sequential 3D Reconstruction with Causal TransformerarXivproject
2025-01-21Online 3D, Recurrent Pointmap, StreamingMeta AICUT3R: Continuous 3D Perception Model with Persistent StateCVPR 2025 Oralproject / github
2024-12-16MASt3R, Dense SLAM, Real-TimeImperial College LondonMASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsCVPR 2025github
2023-09-05Global BA, Neural Implicit, Dense RGB-D SLAMUniversity of BolognaGO-SLAM: Global Optimization for Consistent 3D Instant ReconstructionICCV 2023project / github
2022-11-21Hybrid SDF, Dense RGB-D SLAM, KeyframesIdiap Research InstituteESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance FieldsCVPR 2023project / github
2021-08-24Deep SLAM, Dense BA, Monocular/Stereo/RGB-DPrinceton UniversityDROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasNeurIPS 2021github
2021-03-23Neural Implicit, Online RGB-D, Dense SLAMImperial College LondoniMAP: Implicit Mapping and Positioning in Real-TimeICCV 2021project / github

3.3 Dynamic 3D Reconstruction

3.3.1 Non-Rigid Tracking & Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Long-Range 4D Motion, 3D Queries, Occlusion-Robust Trajectory ChainingCarnegie Mellon UniversityPoint4D: Long-range 4D Motion ReconstructionarXivproject
2026-09-08Event Stream, Extreme-Low-Frame-Rate RGB, Dynamic 3DGS, Real-TimeMacau University of Science and TechnologyEdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene ReconstructionNeurocomputinggithub+dataset
2026-09-05Sparse-View 4D, Spatio-Temporal Depth Alignment, Dynamic 3DGSBIGAIUniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth AlignmentECCV 2026project
2026-08-18Query-Conditioned 4D, Scene Flow, Dynamic Points, Sparse-to-DenseKosmo ResearchUniQuery4R: Unified 4D Scene Reconstruction from a Single QueryarXivproject
2026-07-29Articulated Objects, Structure-aware 3DGS, Part ConnectivityPOSTECHStructureGS: Structure-aware Gaussian Splatting for Articulated Object ReconstructionarXivpaper
2026-07-21Streaming 4D, Instance Grounding, Geometry TransformerHorizon RoboticsIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerarXivproject
2026-07-16Online Dynamic NVS, Space-Time Memory, Real-TimeUniversity of WashingtonOnline Neural Space Time Memory for Dynamic Novel View SynthesisarXivproject
2026-07-01Dynamic Gaussian Reconstruction, Monocular Video, GenerativeStanford UniversityWorld from Motion: Generative Dynamic Gaussian Reconstruction from Monocular VideoarXivproject
2026-06-23Articulated Digital Twin, RGB-D, URDF ExportETH ZurichArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D VideosICRA 2026 Workshoppaper
2026-06-22Monocular Video, 4DGS, In-the-Wild Non-RigidCarnegie Mellon UniversityLift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-WildarXivproject
2026-06-22Dynamic Driving, Sparse Voxels, LiDAR-GuidedHuawei Paris Research CenterDrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene ReconstructionarXivpaper
2026-06-09Future Extrapolation, 4DGS, Autonomous DrivingTsinghua UniversityEnvision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous DrivingarXivproject
2026-06-09Manipulation Video, Decoupled 3DGS, Scene GraphAuthorsManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian SplattingarXivpaper
2026-06-02Object Permanence, Differentiable Physics, 4DGSAuthorsPersistGS: Differentiable Physics for Object Permanence in 4D Gaussian SplattingCVPR 2026 Workshoppaper
2026-06Single Event Camera, Deformable 3DGS, High-Speed 4DShanghaiTech UniversityFastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event CameraCVPR 2026paper
2026-04-10Feed-Forward, Unconstrained Views, Semantic-GeometryNanyang TechFF3R: Feedforward Feature 3D Reconstruction from Unconstrained ViewsCVPR 2026 Findingspaper
2026-04-10Dynamic 4D, Semantic Prior, Gaussian SLAM, Action-ControlUniversity of ZurichGenie 4D: Semantic-Prior-Guided 4D Dynamic Scene ReconstructionarXivpaper
2026-04-10Dynamic/Static Disentanglement, Uncertainty-Aware, Feed-ForwardZhejiang UniversityRobust 4D VGT: Robust 4D Visual Geometry Transformer with Uncertainty-Aware PriorsarXivpaper
2026-04-07Functional Scenes, Egocentric Interaction, URDF/USDStanfordFunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction VideosCVPR 2026paper
2026-04-05Sparse Camera, 4DGS, Neural Decay, CVPR 2026Authors4C4D: 4 Camera 4D Gaussian SplattingCVPR 2026paper / paper
2026-03-30Dynamic Surface, Explicit Geometry, High FidelityAustralian National University4DSurf: High-Fidelity Dynamic Scene Surface ReconstructionCVPR 2026paper
2026-03-21RayMap, Dynamic, StreamingUniversity of Illinois ChicagoRayMap3R: Inference-Time RayMap for Dynamic 3D ReconstructionarXivproject / github
2026-03-09Dynamic VGGT, Autonomous Driving, 4DFudan UniversityDynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingarXivpaper
2025-11-23Dynamic Geometry, Spatiotemporal, VGGTHuazhong University of Science and Technology4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry EstimationarXivpaper
2025-11-07Motion-Aware, Monocular Video, Bundle AdjustmentKAIST4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic ScenesNeurIPS 2025paper
2025-10-20VGGT-4D, Pose/Geometry, Dynamic MaskHarvardPAGE-4D: Disentangled Pose and Geometry Estimation for 4D PerceptionICLR 2026project
2025-08-133D Reconstruction, Human Motion, VideoShanghai AI LabHuman3R: Reconstructing 3D Human Avatars from Monocular VideoCVPR 2025project
2025-07Human, Multi-View, Sparse, Robust, RoGSplatAuthorsRoGSplat: Robust Generalizable Human Gaussian SplattingCVPR 2025paper
2025-06-11Dynamic Human, Temporal Consistency, 4DTsinghuaCARI4D: Cross-Modal Alignment and Reconstruction for Interactive 4D HumanCVPR 2025project
2025-06-10Online, Dynamic 3DGS, Uncalibrated VideoUniversity of British ColumbiaStreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video StreamsICLR 2026project / github
2025-06-094DGS, Transformer, Monocular VideoMeta Reality Labs4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosNeurIPS 2025 Spotlightproject
2025-06-02Video Generators, 4D GeometryOxford VGG / NAVER LABSGeo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionICCV 2025 Highlightpaper
2025-06Self-Supervised, Dynamic, Driving, FlowAuthorsSplatFlow: Self-Supervised Dynamic 3DGS with Neural Motion FlowCVPR 2025paper
2025-06Few-Shot, Personal Avatar, HighlightAuthorsFRESA: Personalized 3D Human Avatar from Few ImagesCVPR 2025 Highlightpaper
2025-06Real-Time Avatar, 166fps, Gaussian, HighlightAuthorsMMLP-Human: Real-Time High-Fidelity Gaussian Human AvatarCVPR 2025 Highlightpaper
2025-05-274D, Dual Correspondences, Dynamic VideoNUS / Shanghai AI LabC4D: 4D Made from 3D through Dual CorrespondencesICCV 2025project
2025-05-144D Pointmaps, Dynamic-Static DisentanglementKAIST / ETH / SonyD2USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic ScenesNeurIPS 2025project
2025-05-02Training-Free, Motion Disentangle, DUSt3RWestlake / MPIEasi3R: Estimating Disentangled Motion from DUSt3R Without TrainingICCV 2025project
2025-04-174D Tracking, Feed-Forward, Pointmap, TrackingMPI / UC BerkeleySt4RTrack: Simultaneous 4D Reconstruction and TrackingICCV 2025paper / paper
2025-04-07Neural Rendering, Human, No EyesUniversity of CambridgeSeeing Without Eyes: Neural Human Rendering from Monocular VideoCVPR 2025project
2025-03-24Multi-Object, 4D, In-the-Wild VideosCMUGenMOJO: Robust Multi-Object 4D Generation for In-the-wild VideosCVPR 2025project
2025-02-27Layered Avatar, Hair, Face, MetaAuthorsLUCAS: Layered Universal Codec AvatarsCVPR 2025paper / paper
2025-01-223D Reconstruction, Canonical, Multi-ViewStanfordUniCon3R: Unified 3D Reconstruction and RecognitionCVPR 2025project
2024-12-03Single Image, Animatable, Avatar, 4DGSAuthorsAniGS: Animatable Gaussian Avatar from a Single ImageCVPR 2025paper / paper
2024-11-274D Generation, Multi-View Video, DiffusionGoogleCAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsCVPR 2025project
2024-10-28Dynamic Geometry, DUSt3R, MotionUniversity of OxfordMonST3R: Estimating Geometry in the Presence of MotionICLR 2025project / github

3.3.2 Deformation Graphs & Canonical Spaces

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-02-124D Dynamic, Monocular Video, Tree-ChainsCornellWorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-ChainsICLR 2026paper
2025-11-014D Scene, Feed-Forward, Controllable, Video DiffusionShanghai AI LabDiff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction ModelsCVPR 2026paper
2025-10-15Dynamic 4D, Gaussian, CanonicalZhejiang UniversityDirector: Directed Generative Models for 4D Scene EvolutionCVPR 2025project
2025-08-06Multi-Baseline, Generalizable, GaussianAuthorsMuGS: Multi-Baseline Generalizable Gaussian SplattingICCV 2025paper / paper
2025-07Surface, Gaussian Surfels, 2DGS, Sparse-View, SpotlightAuthorsMAtCha Gaussians: Atlas Charting with 2D Gaussian SurfelsCVPR 2025 Spotlightpaper
2025-07SDF+3DGS, Hybrid, Surface, ICCVAuthorsSurfaceSplat: SDF+3DGS Hybrid Surface ReconstructionICCV 2025paper
2025-07Sparse-View, Implicit, Voxel, ConsistencyAuthorsSparseRecon: Sparse-View Implicit Surface ReconstructionICCV 2025paper
2025-07Low-Texture, Reflection, Unified, +21%AuthorsHiNeuS: Unified Neural Implicit Surface ReconstructionICCV 2025paper
2025-06Joint Human+Scene, MASt3R ExtensionAuthorsHAMSt3R: Joint Human and Scene 3D ReconstructionICCV 2025paper
2024-11-204D Reconstruction, Gaussian Splatting, ForwardShanghai AI LabForge4D: Gaussian Splatting for Forward Facing 4D ReconstructionarXivpaper

🎛️ 4. 3D Generation

This section tracks methods that create new 3D assets, parts, articulated objects, scenes, and editable 3D worlds, with emphasis on physical and simulation use.

4.1 Object-Level Generation

4.1.1 Image/Text to 3D Mesh

Single Image / Text-Conditioned Object Mesh
DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Multi-View, Reconstruction Prior, Noise Inversion, Faithful Asset CompletionHKUSTReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and ModulationarXivpaper
2026-09-09Image-to-3D, Partial Geometry, Test-Time Guidance, SAM 3DUniversity of OxfordGuiding Image-to-3D Generation with Test-Time Partial ObservationsarXivpaper
2026-07-22Code-Native, Programmable, 3D AssetsAuthorsNova3D: Code-Native Generation of Programmable 3D AssetsarXivpaper
2026-06-23Image-to-3DGS, Sparse Voxel, High-Fidelity AssetsThe Australian National UniversityFLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse RepresentationarXivpaper
2026-06-23Vehicle Assets, 3D-Consistent Views, SimulationShanghai Jiao Tong University3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view SynthesisarXivpaper
2026-06-22Constraint Meshes, Controllable Assets, TRELLISUniversity of TübingenArbor: Explicit Geometric Conditioning for Controllable 3D Asset GenerationarXivproject
2026-05-22Image-to-3D, Deployable Mesh, UV, Real-TimeMeta Reality LabsAssetGen: Deployable 3D Asset Generation at Interactive SpeedarXivpaper
2026-05-11Pixel-Aligned, Image-to-3D, PBR, Multi-ViewTencent ARCPixal3D: Pixel-Aligned 3D Generation from ImagesSIGGRAPH 2026project / github / demo
2026-05-01Pose-Aware, Diffusion, 3D Geometry, DirectRenmin UniversityPAD: Pose-Aware Diffusion for 3D GenerationarXivpaper
2026-04-22Simulation-Ready, PBR, Part-Aware, ArticulationByteDance SeedSeed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content GenerationarXivdemo
2025-12-16Image-to-3D, O-Voxel, PBR MaterialsMicrosoftNative and Compact Structured Latents for 3D GenerationarXivgithub / model
2025-11-20Single Image, Multi-Object, Layout, SAM 3DMeta AISAM 3D: 3Dfy Anything in ImagesarXivgithub
2025-10-23Single Image, Pose-Grounded, Flow MatchingMeta AICUPID: Generative 3D Reconstruction via Joint Object and Pose ModelingarXivproject
2025-10-09PBR, Material Diffusion, Relighting, Single ImageStability AISViM3D: Stable Video Material Diffusion for Single Image 3D GenerationarXivpaper
2025-06-18Image-to-3D, PBR Materials, Production AssetsTencentHunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR MaterialarXivgithub / model
2025-05-23Gigascale 3D, Sparse Volume, Image ConditioningNanjing UniversityDirect3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionNeurIPS 2025project / github
2025-05-12Textured Assets, Controllable 3D, Open FrameworkStepFunStep1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D AssetsarXivgithub
2025-05-07Image-to-3D, PBR Textured Mesh, Render-Enhanced Auto-EncoderTsinghuaMeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationCVPR 2025project
2025-03-28Geometry Detail, Normal Bridging, Image-to-3DCornellHi3DGen: High-fidelity 3D Geometry Generation from Images via Normal BridgingarXivproject
2025-03-03Text/Image-to-3D, Bundle Image, Data-Efficient (147K)HKUST(GZ)Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset GenerationarXivgithub
2025-02-10Shape Synthesis, Rectified Flow, Image-to-3DVAST AI / TripoTripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow ModelsarXivgithub / model
2025-01-21Image/Text-to-3D, Mesh, TextureTencentHunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets GenerationarXivgithub
2025-01-08Point-Aware, Interactive Editing, Single ImageStability AI / UIUCSPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single ImagesarXivproject / model
2024-11-11Text/Image-to-3D, PBR, Multi-View DiffusionNVIDIAEdify 3D: Scalable High-Quality 3D Asset GenerationarXivproject
2024-09-193DTopia-XL, Primitive Diffusion, Large Scale, HighlightNTU / PKU3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionCVPR 2025 Highlightproject
2024-08-01Fast Mesh, UV Unwrap, Material DisentanglementStability AIStable Fast 3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination DisentanglementarXivproject / github
2024-07-02Text-to-Mesh, PBR, Geometry+TextureMetaMeta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR MaterialsNeurIPS 2024project / paper
2024-06-26GaussianDreamer, Text-to-3DGS, 2D+3D DiffusionHUSTGaussianDreamer: Fast Generation from Text to 3D GaussiansCVPR 2024github
2024-05-30High-Quality Mesh, Multi-View Normals, ISOMERTsinghuaUnique3D: High-Quality and Efficient 3D Mesh Generation from a Single ImagearXivproject
2024-05-23Text-to-3D, 3D-DiT, Interactive RefinementHKUSTCraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry RefinerarXivgithub
2024-05-19Multi-View Diffusion, Row-Wise Attention, MeshUniversity of Hong KongEra3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionarXivproject
2024-04-10Single Image, Sparse-View LRM, MeshTencent ARCInstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction ModelsarXivgithub
2024-03-18Single Image, Orbital Video, Multi-View PriorStability AISV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video DiffusionarXivproject
2024-03-08Textured Mesh, Convolutional Reconstruction, Single ImageShanghai AI LabCRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction ModelarXivproject
2024-03-043DTopia, Hybrid Diffusion, Text-to-3DNTU / PKU3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion PriorsCVPR 2024project
2024-02-20MVDiffusion++, Dense High-Res, Sparse-ViewSimon Fraser / MetaMVDiffusion++: A Dense High-Resolution Multi-view Diffusion ModelECCV 2024paper
2024-02-07Multi-View Gaussian, Fast 3D Content, Image/TextPeking UniversityLGM: Large Multi-View Gaussian Model for High-Resolution 3D Content CreationECCV 2024 Oralgithub
2023-12-06XCube, Sparse Voxel, Large-Scale, HighlightNVIDIAXCube: Large-Scale 3D Generative Modeling using Sparse Voxel HierarchiesCVPR 2024 Highlightproject
2023-12-05ReconFusion, Diffusion Prior, 3D ReconstructionColumbia / GoogleReconFusion: 3D Reconstruction with Diffusion PriorsCVPR 2024paper
2023-11-27MeshGPT, Triangle Mesh, Decoder-Only, TransformerTUMMeshGPT: Generating Triangle Meshes with Decoder-Only TransformersCVPR 2024project
2023-11-19LucidDreamer, Interval Score, Text-to-3D, HighlightHKUST(GZ)LucidDreamer: Towards High-Fidelity Text-to-3D via Interval Score MatchingCVPR 2024 Highlightproject
2023-11-15DMV3D, Multi-View Diffusion, 3D LRM, NeRFAdobe / StanfordDMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction ModelNeurIPS 2023paper
2023-10-26MVDiffusion, Multi-View, Consistent, Single ImageSimon FraserMVDiffusion: Multi-view Consistent Image Generation from a Single ImageCVPR 2024github
2023-10-25DreamCraft3D, Bootstrapped, Hierarchical, 3DDeepSeek / TsinghuaDreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorICLR 2024project
2023-10-23Cross-Domain Diffusion, Multi-View Normals, MeshHKUSTWonder3D: Single Image to 3D using Cross-Domain DiffusionCVPR 2024 Highlightgithub
2023-10-23Single Image, Consistent Multi-View DiffusionUC San DiegoZero123++: a Single Image to Consistent Multi-view Diffusion Base ModelarXivgithub
2023-09-28GS, SDS, Efficient 3D Content CreationPKU / NTUDreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationICLR 2024 Oralgithub
2023-09-04Multi-View Consistent Images, Single ImageHKUSyncDreamer: Generating Multiview-consistent Images from a Single-view ImageICLR 2024 Spotlightgithub
2023-08-31Multi-View Diffusion, 3D GenerationUCSD / ByteDanceMVDream: Multi-view Diffusion for 3D GenerationICLR 2024project
2023-08-22IT3D, Explicit View, Text-to-3DNTUIT3D: Improved Text-to-3D with Explicit View SynthesisNeurIPS 2023paper
2023-06-30Magic123, One Image to 3D, 2D+3D PriorsKAUST / SnapMagic123: One Image to High-Quality 3D Object Using Both 2D and 3D Diffusion PriorsCVPR 2024project
2023-06-29Single Image, Optimization-Free, MeshShanghai AI LabOne-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationNeurIPS 2023github
2023-06-21DreamTime, Improved SDS, OptimizationIDEA / HKUDreamTime: An Improved Optimization for Diffusion-Guided 3DICLR 2024paper
2023-06-06ATT3D, Amortized, Text-to-3DNVIDIAATT3D: Amortized Text-to-3D Object SynthesisICCV 2023paper
2023-05-30HiFA, High-Fidelity, Text-to-3D, Advanced GuidanceUIUCHiFA: High-fidelity Text-to-3D with Advanced Diffusion GuidanceICLR 2024paper
2023-05-25Text-to-3D, Variational SDS, High-Fidelity, SpotlightTsinghuaProlificDreamer: High-Fidelity and Diverse Text-to-3D GenerationNeurIPS 2023 Spotlightproject
2023-03-24Make-It-3D, Single Image, Diffusion PriorSJTU / MicrosoftMake-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorICCV 2023project
2023-03-24Text-to-3D, Geometry+Appearance DisentangleSCUTFantasia3D: Disentangling Geometry and Appearance for Text-to-3DICCV 2023project
2023-03-20Zero-1-to-3, Zero-Shot, Single Image to 3DColumbia / TRIZero-1-to-3: Zero-shot One Image to 3D ObjectICCV 2023github
2023-02-21RealFusion, Single Image, 360°, 3DOxford VGGRealFusion: 360° Reconstruction of Any Object from a Single ImageCVPR 2023project
2023-01-263DShape2VecSet, Neural Fields, Diffusion, MeshKAUST / TUM3DShape2VecSet: A 3D Shape Representation for Neural Fields and DiffusionSIGGRAPH 2023 (ACM TOG)github
2022-12-02DiffRF, Rendering-Guided, Radiance Field Diffusion, HighlightTUM / MetaDiffRF: Rendering-Guided 3D Radiance Field DiffusionCVPR 2023 Highlightpaper
2022-12-01Text-to-3D, Score Jacobian Chaining, 2D DiffusionTTI-ChicagoScore Jacobian Chaining: Lifting Pretrained 2D Diffusion for 3DCVPR 2023paper
2022-11-18Text-to-3D, High-Res, Coarse-to-Fine, HighlightNVIDIAMagic3D: High-Resolution Text-to-3D Content CreationCVPR 2023 Highlightpaper
2022-10-12LION, Latent Point Diffusion, ShapeNVIDIALION: Latent Point Diffusion Models for 3D Shape GenerationNeurIPS 2023github
2022-09-29Text-to-3D, Score Distillation, SDS, 2D Diffusion, Outstanding PaperGoogle ResearchDreamFusion: Text-to-3D using 2D DiffusionICLR 2023 Outstanding Paperproject
2022-09-22GET3D, Generative, Textured Shapes, Images OnlyNVIDIAGET3D: A Generative Model of High Quality 3D Textured ShapesNeurIPS 2022github
2022-03-17AutoSDF, Shape Prior, 3D CompletionCMUAutoSDF: Shape Priors for 3D Completion, Reconstruction and GenerationCVPR 2022paper
2020-02-23PolyGen, Autoregressive, Mesh, DeepMindDeepMindPolyGen: An Autoregressive Generative Model of 3D MeshesICML 2020paper
Multi-Image / Multi-View Object Mesh
DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-23Multi-View + LiDAR, Vehicle Assets, TRELLISShanghai Jiao Tong UniversityMM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous DrivingarXivgithub
2026-03-12Multi-View, SAM3D, Layout-Aware, Physical PlausibilityPeking UniversityMV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D GenerationarXivgithub
2026-01-16Casual Capture, Posed Multi-View, Metric ShapeMeta AIShapeR: Robust Conditional 3D Shape Generation from Casual CapturesarXivpaper
2025-11-12Multi-Image Fusion, Region Control, TRELLISZhejiang UniversityFuse3D: Generating 3D Assets Controlled by Multi-Image FusionSIGGRAPH Asia 2025project / github
2025-03-18Multi-View Image-to-Shape, Hunyuan3D-DiTTencentHunyuan3D 2.0 MVModelgithub / model
2024-02-06Scalable View Synthesis, Single/Multi-Image 3DKAUSTEscherNet: A Generative Model for Scalable View SynthesisCVPR 2024project
2019-08-05Multi-View Images, Mesh Deformation, Shape RefinementNational Tsing Hua UniversityPixel2Mesh++: Multi-View 3D Mesh Generation via DeformationICCV 2019github

4.1.2 Texture & Material Generation (PBR)

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-25Relightable 3D Assets, PBR Maps, Gaussian RepresentationAppleLuce: Relightable Gaussians for 3D Asset GenerationarXivpaper
2026-08-24Material Decomposition, Physical Properties, Watertight Sub-Meshes, Sim-ReadyUniversity of BristolGen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material DecompositionarXivpaper
2026-07-01Complex Textures, Video Generative Prior, 3D AssetsAuthorsInk3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative ModelsarXivpaper
2024-11-10VLM-Guided, PBR Texture3D AIGCTexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian SplattingCVPR 2025project
2024-01-17TextureDreamer, Geometry-Aware, DiffusionUCSD / MetaTextureDreamer: Image-Guided Texture SynthesisCVPR 2024paper
2023-12-21Texture Generation, Mesh, Multi-View ConsistencyTencentPaint3D: Paint Anything 3D with Lighting-Less Texture Diffusion ModelsCVPR 2024project
2023-11-28SceneTex, Indoor, Texture, Diffusion, HighlightTUM / SnapSceneTex: High-Quality Texture Synthesis for Indoor ScenesCVPR 2024 Highlightpaper
2023-11-21SyncMVD, Multi-View, Text-to-TextureCUHKSyncMVD: Text-Guided Texturing by Synchronized Multi-View DiffusionCVPR 2024paper
2023-08-22PBR Material, SVBRDF, Text-to-MaterialAdobeMatFuse: Controllable Material Generation with Diffusion ModelsSIGGRAPH Asia 2024project
2023-03-20Texture, Material, Text-to-TextureKAISTText2Tex: Text-driven Texture Synthesis via Diffusion ModelsICCV 2023project
2023-02-03TEXTure, Text-Guided, 3D Texture, DiffusionTel Aviv UniversityTEXTure: Text-Guided Texturing of 3D ShapesSIGGRAPH 2023project
2022-07-06nvdiffrec, 3D Mesh, Material, Lighting, OralNVIDIAnvdiffrec: Extracting Triangular 3D Models, Materials, and LightingCVPR 2022 Oralgithub

4.2 Part-Level & Articulated Generation

4.2.1 Kinematic Structure Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-08Single Image, Physical CoT, URDF, Simulation-ReadyAerospace Information Research Institute, CASPhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D AssetsarXivpaper
2026-07-15Articulation + Physics, 40K Assets, Simulation-ReadyZhejiang UniversityUniPhysGen: Unified Physical Grounding for Simulation-Ready 3D AssetsarXivgithub
2026-05-20Rigid/Deformable/Articulated, Physical Attributes, Sim-ReadyNanyang Technological UniversityPhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated ObjectsarXivproject / dataset
2026-05-14Agentic Generation, Articraft-10K, URDF AssetsAuthorsArticraft: An Agentic System for Scalable Articulated 3D Asset GenerationarXivproject / github
2026-05-06Physics-Grounded, Kinematic, Simulation-Ready AssetsHKUPhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual WorldICML 2026project / github
2026-03-14URDF, Autoregressive, Simulation-Ready AssetsAuthorsURDF-Anything+: End-to-End Generation for Simulation-Ready Articulated AssetsarXivpaper
2026-03-01Articulated Assets, 3D LLM, Kinematic StructureTsinghuaArtLLM: Generating Articulated Assets via 3D LLMCVPR 2026paper
2025-12-12Articulation, Kinematic Tree, Feed-Forward, URDF-ReadyUniversity of OxfordParticulate: Feed-Forward 3D Object ArticulationarXivproject
2025-11-26Single Image, Open-Set Articulation, Unified LatentShanghaiTechUniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set ArticulationarXivpaper
2025-11-17Sim-Ready Assets, Physical Properties, Single ImageNTUPhysX-Anything: Simulation-Ready Physical 3D Assets from Single ImageCVPR 2026paper
2025-11-02URDF, 3D MLLM, Articulated ObjectsTsinghuaURDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelNeurIPS 2025paper
2025-08-20Articulated Geometry, Motion Modeling, Gaussian RepresentationTsinghua UniversityGaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects3DV 2026paper
2025-07-16Physical Properties, Scale, Material, AffordanceNanyang Technological UniversityPhysX-3D: Physical-Grounded 3D Asset GenerationNeurIPS 2025 Spotlightproject / github
2025-06-10Interactable Digital Twin, Articulated Object, RGB-D VideoShanghai Jiao Tong UniversityiTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos3DV 2026paper
2025-04-17Simulation-Ready, Physical Materials, DynamicsUniversity of Massachusetts AmherstSOPHY: Generating Simulation-Ready Objects with Physical MaterialsWACV 2026project / github
2025-03-11Part-Level Digital Twin, Joint Estimation, Self-Supervised 3DGSUSTCArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian SplattingCVPR 2025paper
2025-02-26Articulated Objects, 3DGS, Joint EstimationTsinghuaArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian SplattingICLR 2025project
2025-02-17Articulation-Ready, Skeleton, Skinning, BenchmarkNanyang Technological UniversityMagicArticulate: Make Your 3D Models Articulation-ReadyCVPR 2025project / github
2024-09-26Open-Vocabulary, URDF, ArticulationStanfordArticulate Anything: Open-vocabulary 3D Articulated Object GenerationICLR 2025project

4.2.2 Part-Aware Assembly & Editing

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-20Compositional 3D, Part Semantics, Spatial Control, Reassemblable AssetsRobloxMultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial ControlarXivproject
2026-08-14Up to 300 Parts, Token-Efficient VQ, Autoregressive 3D, Structured AssetsThe University of Hong KongMegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive ModelingarXivproject
2026-08-13Part-Aware Generation, Recursive Decomposition, Editable AssetsShanghai Jiao Tong UniversitySCULPT: Subtractive Composition for 3D Part GenerationarXivproject
2026-07-18Category-Agnostic, Neural Shape Editing, Coupled RepresentationAuthorsCNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape RepresentationarXivpaper
2026-06-23Garment Patterns, Simulation-Ready, EditingUniversity of Hong KongPatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D GarmentsarXivgithub
2026-05-27Part-Controllable, Open-Vocabulary, Game-ReadyRobloxCubePart: An Open-Vocabulary Part-Controllable 3D GeneratorSIGGRAPH 2026project / model
2025-09-10Part Decomposition, Editable, Production-Ready AssetsTencentX-Part: High-Fidelity and Structure-Coherent Shape DecompositionTech Reportproject
2025-08-14Rigging, Animation, Skeleton, SkinningNanyang Technological UniversityPuppeteer: Rig and Animate Your 3D ModelsNeurIPS 2025 Spotlightproject / github
2025-06-05Part-Level Mesh, Compositional DiT, Single ImageUniversity of WaterlooPartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersarXivproject
2024-12-16Articulated Mesh, Part-by-Part, Hierarchical TransformerCornellMeshArt: Generating Articulated Meshes with Structure-Guided TransformersarXivpaper
2023-12-13Shape Program, Structure, Editable AssetsMITShape2Program: Learning to Infer Shape Programs from 3D ShapesarXivproject
2023-06-29Part-Aware, Shape Assembly, 3D GenerationShanghai AI LabMichelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationNeurIPS 2023github

4.3 Scene-Level Generation

4.3.1 Layout & Procedural Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Single Image, Executable Scene Programs, Recursive Construction, Editable 3DGeorgia Institute of TechnologyRecursive Code World Models: Building Complex Worlds through Recursive Scene ProgramsarXivpaper
2026-09-05Agentic 3D Composition, Functional Objects, Executable Robot ScenesPeking University / GalbotGIF: Agentic Generation of Interactive and Functional Object Compositions for Robot LearningarXivpaper
2026-09-04Image-to-Scene, Agentic Layout Evolution, Simulation-Ready DiversityThe University of Hong KongSceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout EvolutionarXivproject / github
2026-08-31Text-to-Scene, Grow-and-Repair, Functional Groups, SceneReverse-17KSoutheast UniversityScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene GenerationarXivproject
2026-08-27Single Image, Generative 3D Proxy, RGB-D World Expansion, Explorable SceneHong Kong University of Science and TechnologySpatialCrafter: Single Image World Modeling with Generative 3D ProxiesarXivproject
2026-08-25Monocular Image, Interactive Scene Programming, Articulation+Physics, Embodied SimulationShanghai Jiao Tong UniversityNeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied SimulationarXivproject
2026-08-19Usage-Driven Code Scenes, Multi-Part Interaction, Executable SimulationShanghai Jiao Tong UniversityBeyond Placement and Articulation: Usage-Driven Code Scenes for Embodied InteractionarXivpaper
2026-07-29Panoramic Video, 3DGS, Simulation-ready WorldAgiBotGenie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and SimulationarXivgithub
2026-07-15Indoor Layout, Progressive VLM Reasoning, Interactive EditingCity University of Hong KongThinkBLOX: 3D Indoor Scene Generation with Progressive ReasoningarXivpaper
2026-07-08Simulation-Ready Assets, Affordances, Cross-Simulator WorldsHorizon RoboticsEmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AIarXivproject / github
2026-07-07Real-to-Sim, One-Shot Scene Generation, Robot EvaluationShanghai AI LabRoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and EvaluationarXivproject
2026-07-04Egocentric Scene Generation, Geometric 3DGS, ConsistencySouth China Univ. of TechnologyCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene GenerationarXivpaper
2026-06-23Triangle Splatting, Single-Image Scene, Game-ReadyGoogle ResearchFLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene GenerationarXivproject
2026-06-23Text-to-Scene, Video Priors, 3DGS OrbitUniversity of BernOrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video SynthesisarXivpaper
2026-06-23Compositional 3D, Physical Interaction, Multi-View ConsistencyChina University of Petroleum (East China)Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D GenerationarXivpaper
2026-06-23Satellite-to-City, Textured Mesh, Urban SimulationHKUST(GZ)Sat2City v2: Native 3D City Asset Generation from a Single Satellite ImagearXivpaper
2026-06-08Satellite-to-3D, 3DGS, UAV SimulationAmap-cvlab / AlibabaABot-Earth 0.5: Generative 3D Earth ModelarXivproject
2026-06-04Whole-Home Scenes, Floorplans, InteractiveAuthorsHomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home ScenesarXivpaper
2026-05-28Physical Stability, Single Image, Scene Tree, SimulationCarnegie Mellon UniversityREST3D: Reconstructing Physically Stable 3D Scenes from a Single ImagearXivproject
2026-05-01Segment Map, Text-to-World, Controllable 3D WorldsSeoul National UniversityMap2World: Segment Map Conditioned Text to 3D World GenerationarXivpaper
2026-04-14Explorable 3D Worlds, Long Trajectory, 3DGSNVIDIALyra 2.0: Explorable Generative 3D WorldsarXivproject
2026-04-06Single-Image Scene, In-Place Completion, ARSG-110KNankai University3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageCVPR 2026project / github / dataset
2026-03-31Town-Scale, Single Image, Latent ExtensionSeoul National UniversityExtend3D: Town-Scale 3D GenerationCVPR 2026project
2026-03-31Unbounded World, Flow Matching, LayoutsPrinceton UniversityWorldFlow3D: Flowing Through 3D Distributions for Unbounded World GenerationarXivproject
2026-03-27Autoregressive 3DGS, Token Generation, Completion/OutpaintingTechnical University of MunichGaussianGPT: Towards Autoregressive 3D Gaussian Scene GenerationarXivproject
2026-03-12Multi-Floor, Language-to-3D, Long-Horizon TasksTsinghuaMANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasksCVPR 2026paper
2026-03-06Compositional Scene, Panoramic Image, Feed-ForwardNTUPano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic ImageCVPR 2026paper
2026-02-10Agentic Scene Generation, Sim-Ready, SAGE-10kNVIDIASAGE: Scalable Agentic 3D Scene Generation for Embodied AICVPR 2026project / github
2026-01-09Language-Guided, Infinite Worlds, Articulated FurnitureNational Taiwan UnivSceneFoundry: Generating Interactive Infinite 3D WorldsarXivproject
2025-12-01Tabletop, Instance-Level, Interactive Scene, Text/ImageD-RoboticsTabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single ImagearXivproject / paper
2025-11-18Single-Image Scene, Gaussian World, Scene GenerationTsinghua UniversityGEN3D: Generating Domain-Free 3D Scenes from a Single ImagearXivpaper
2025-09-18Layout-Guided, Indoor Scenes, Decoupled Geometry/AppearanceHong Kong University of Science and TechnologySPATIALGEN: Layout-guided 3D Indoor Scene Generation3DV 2026paper
2025-08-21Single Image, Multi-Asset Scene, Feed-ForwardShanghai Jiao Tong UniversitySceneGen: Single-Image 3D Scene Generation in One Feedforward Pass3DV 2026project / github
2025-08-11Panoramic, Explorable World, Matrix-PanoKunlun WanweiMatrix-3D: Omnidirectional Explorable 3D World GenerationarXivproject / github
2025-07-29Panoramic, Text/Image-to-World, Mesh ExportTencent HunyuanHunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelsarXivgithub
2025-07-09VLA Scene Authoring, Simulation-Ready Worlds, Synthetic DataNVIDIA / Stanford University3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds3DV 2026project / github
2025-06-25Explorable Scene, Novel-View Restoration, ConsistencyBeijing Academy of AIWonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene ExplorationarXivpaper
2025-03-13Scene Layout, Optimization, GenerationTsinghuaHOG-Layout: Layout-Enhanced Scene Generation via Hierarchical OptimizationarXivproject
2025-02-20Octree, 3D Diffusion, Scene GenerationZhejiang UniversityOctree Diffusion: Hierarchical Scene Generation via Octree StructuresarXivproject
2025-01-15Gaussian, GPT, Scene GenerationShanghai AI LabGaussianGPT: Language-Driven Scene Generation with Gaussian RepresentationarXivpaper
2024-12-19Scene Generation, Growing, IncrementalTsinghuaWorldGrow: Incremental 3D Scene GenerationCVPR 2025project
2024-11-05Splatting, Fluents, Scene UnderstandingUniversity of CambridgeFluSplat: Fluent Scene Generation via Gaussian SplattingCVPR 2025project
2024-11-04GenXD, Any 3D and 4D, Scene GenerationNUSGenXD: Generating Any 3D and 4D ScenesICLR 2025paper
2024-06-17Procedural Scenes, Synthetic Data, Embodied AIPrincetonInfinigen Indoors: Photorealistic Indoor Scenes using Procedural GenerationNeurIPS 2024project / github
2024-06-13Interactive 3D Scene, FLAGS, Single ImageStanford UniversityWonderWorld: Interactive 3D Scene Generation from a Single ImageCVPR 2025project
2024-05-02EchoScene, Scene Graph, Diffusion, IndoorTUM / JHUEchoScene: Indoor Scene Generation via Information EchoECCV 2024paper
2024-02-12SceneScape, Text-Driven, Consistent, SceneWeizmannSceneScape: Text-Driven Consistent Scene GenerationNeurIPS 2024paper
2024-01-30BlockFusion, Expandable, Tri-plane, SIGGRAPHTencent / UTokyoBlockFusion: Expandable 3D Scene GenerationSIGGRAPH 2024 (ACM TOG)paper
2023-12-01ControlRoom3D, Semantic Proxy, Room GenerationTUM / MetaControlRoom3D: Room Generation using Semantic Proxy RoomsCVPR 2024paper
2023-10-05Ctrl-Room, Text-to-3D, Layout ConstraintsSimon FraserCtrl-Room: Controllable Text-to-3D Room Meshes GenerationECCV 2024paper
2023-10-04MagicDrive, Street View, 3D Geometry ControlCUHK / HKUSTMagicDrive: Street View Generation with Diverse 3D Geometry ControlICLR 2024project
2023-06-15Procedural World, Synthetic Data, SimulationPrincetonInfinite Photorealistic Worlds using Procedural GenerationCVPR 2023project
2023-03-24DiffuScene, Diffusion, Indoor Scene SynthesisTUMDiffuScene: Denoising Diffusion for Generative Indoor Scene SynthesisCVPR 2024paper
2023-03-21Text-to-3D Room, Indoor Scenes, MeshLMU MunichText2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsICCV 2023project
2023-02-02Unbounded 3D Scene, Generative Model, DrivingNVIDIASceneDreamer: Unbounded 3D Scene Generation from 2D Image CollectionsCVPR 2023project

4.3.2 Semantic Scene Generation & Spatial Intelligence

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-18Multi-View Generation, Scene Assets, Training-FreeThe University of QueenslandScene-SAM3D: Multi-View Scene Asset Generation Without Fine-TuningarXivgithub
2025-01-03Rover, Semantic, 3D SceneCarnegie Mellon UniversitySEM-ROVER: Semantic Scene Exploration with Hierarchical Spatial ReasoningICLR 2025project
2024-09-30Spatial, Generation, LanguageTsinghuaSpatialGen: Language-Driven Spatial Scene GenerationNeurIPS 2024project

4.3.3 4D / Dynamic Scene Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-11Layout-Conditioned, 3DGS, Mixed RealityAuthorsSyncSpace: Layout-Conditioned 3D Gaussian Splatting for Space Reskinning in Mixed RealityarXivpaper
2025-04-17Image-Pair-to-4D, Diffusion, Explicit 3D MotionTechnical University of MunichTwoSquared: 4D Generation from 2D Image Pairs3DV 2026 Oralpaper
2024-12-054Real-Video, Photo-Realistic, Video Diffusion, CVPRSnap / KAUST4Real-Video: Generalizable Photo-Realistic 4D Video DiffusionCVPR 2025paper
2024-07-16Animate3D, Multi-View, Video Diffusion, NeurIPSCASIA / AlibabaAnimate3D: Animating Any 3D Model with Multi-view Video DiffusionNeurIPS 2024paper
2024-05-314Diffusion, Multi-View Video, 4D, NeurIPSCASIA / Shanghai AI Lab4Diffusion: Multi-view Video Diffusion Model for 4D GenerationNeurIPS 2024paper
2024-05-26Diffusion4D, Video Diffusion, 4D, NeurIPSToronto / BJTUDiffusion4D: Fast Spatial-temporal Consistent 4D GenerationNeurIPS 2024paper
2024-05-03DreamScene4D, Multi-Object, Dynamic, NeurIPSCMUDreamScene4D: Dynamic Multi-Object Scene GenerationNeurIPS 2024paper
2024-03-22STAG4D, Spatial-Temporal, 4D Gaussians, ECCVNanjing / CASIASTAG4D: Spatial-Temporal Anchored Generative 4D GaussiansECCV 2024paper
2023-11-294D-fy, Text-to-4D, Score Distillation, CVPRKAUST / Snap4D-fy: Text-to-4D Generation Using Hybrid Score Distillation SamplingCVPR 2024project
2023-11-24Animate124, Image-to-4D, Animation, ICLRNUS / HuaweiAnimate124: Animating One Image to 4D Dynamic SceneICLR 2024paper
2023-11-17Consistent4D, 360° Dynamic, Monocular Video, ICLRCASIA / NanjingConsistent4D: Consistent 360° Dynamic Object GenerationICLR 2024project

4.4 3D Editing

3D editing covers methods that modify existing 3D assets, Gaussian / NeRF fields, meshes, voxel or latent states, and dynamic scenes. Entries are grouped by editable state: object-level, scene-level, and dynamic / 4D.

4.4.1 Object-Level Editing

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-27NeRF Editing, Object Removal, Robot ManipulationAuthorsNEO: NeRF It Once, Edit It Many Times for Continuous Object ManipulationarXivpaper
2026-06-05Mesh Editing, Image-Guided, Local MorphingLeiden University3DMorph: Single-Image-Guided Local 3D Shape Editing and MorphingIJCNN 2026github
2026-05-26PartFlow, Semantic-Part Transformation, Mask-FreeNanyang Technological UniversityFeedforward 3D Editing Learns from Semantic-Part TransformationarXivproject / github / benchmark (Steer3D)
2026-05-08VS3D, Velocity-Space, Mask-FreeTsinghua UniversityVelocity-Space 3D Asset EditingarXivpaper
2026-05-01Latent Editing, Object-Level, Structured 3D LatentsSeoul National UniversityInpaintSLat: Inpainting Structured 3D Latents via Initial Noise OptimizationarXivproject
2026-04-30MeshReGen, VecSet Regeneration, Image-Guided EditingKAISTMeshReGen: A Unified 3D Geometry Regeneration FrameworkarXivproject / benchmark (VoxHammer)
2026-04-26Primitive Proxy, Shape Editing, Fine-Grained ControlTel Aviv UniversityProx-E: Fine-Grained 3D Shape Editing via Primitive-Based AbstractionsSIGGRAPH 2026project / github / benchmark (VoxHammer)
2026-03-303DGS Editing, Object-Level, Single-ViewZhejiang Gongshang UniversitySVGS: Single-View to 3D Object Editing via Gaussian SplattingACM TOMM 2026project
2026-02-25Voxel Editing, Object-Level, Rectified Voxel FlowUSTCEasy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowCVPR 2026project
2026-02-05Native 3D Editing, Image-Conditioned, Latent-to-LatentAigency.ai / Tel Aviv UniversityShapeUP: Scalable Image-Conditioned 3D EditingSIGGRAPH 2026project / github
2026-02-04Mesh Editing, Object-Level, Single-Image LRMNational Yang Ming Chiao Tung UniversityVecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single ImagearXivgithub / benchmark (VoxHammer)
2025-12-15Steer3D, Text-Steerable, Edit3D-Bench (Ma)CaltechFeedforward 3D Editing via Text-Steerable Image-to-3DarXivproject / github / benchmark (Steer3D)
2025-11-27Latent Anchor, Object-Level, Mask-FreeZhejiang UniversityAnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned FlowsCVPR 2026 Oralproject / github / benchmark
2025-11-21Native Editing, Object-Level, Full AttentionFudan University / StepFunNative 3D Editing with Full AttentionarXivpaper
2025-10-16FlowEdit, Object-Level, Mask-FreeTsinghua UniversityNANO3D: A Training-Free Approach for Efficient 3D Editing Without MasksICLR 2026project / github / dataset
2025-10-033DEditFormer, Paired Dataset, Mask-FreeEast China Univ of Science & TechnologyTowards Scalable and Consistent 3D EditingarXivproject / github
2025-08-293D-LATTE, Text Instructions, 3D Diffusion LatentUniversity of Tübingen3D-LATTE: Latent Space 3D Editing from Textual InstructionsCVPR 2026 Oralproject / CVF
2025-08-26VoxHammer, Edit3D-Bench (Li), Training-FreeRenmin University / Beihang UniversityVoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space3DV 2026 Oralproject / github / benchmark (VoxHammer)
2025-07-153DGS Editing, Part-Level, Regularized SDSSeoul National UniversityRobust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation SamplingICCV 2025CVF
2025-06-25Image Prompt, Multi-View Propagation, Mask-FreeTel Aviv UniversityEditP23: 3D Editing via Propagation of Image Prompts to Multi-ViewACM TOG 2025project / github
2025-05-11CMD, Local Editing, Progressive GenerationHKUSTCMD: Controllable Multiview Diffusion for 3D Editing and Progressive GenerationSIGGRAPH 2025project
2024-12-11Mesh Editing, Object-Level, Masked LR

Truncated — view the full README on GitHub.

3d-generation
3d-reconstruction
3d-vision
awesome-list
computer-vision
embodied-ai
gaussian-splatting
nerf
robotics
world-models

Contributors

chang-xinhai/Awesome-Embodied-3DV

A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.

Python

13

151 commits

updated Sep 23, 2026

See the code

README

Awesome-Embodied-3DV

A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.

Perception Representation Reconstruction Generation Embodiment Infrastructure

About

Awesome-Embodied-3DV is a curated list for the research space where 3D vision, 3D/4D reconstruction, 3D generation, simulation-ready assets, and embodied world models meet.

This repository focuses on:

  • Data Perception: depth, normals, active imaging, panoramic/fisheye, dense mapping, and 3D semantic understanding (detection, segmentation, grounding) that feed downstream systems
  • 3D/4D Representation: 3DGS, 4DGS, NeRF/SDF, mesh, voxel, point, and hybrid structures
  • 3D Reconstruction: offline, feed-forward, streaming, online, semantic, instance-level, and dynamic reconstruction
  • 3D Generation: object-level, part-level, articulated, scene-level, editable, and simulation-ready asset generation
  • Embodiment & World Models: reconstruction-based world models, dynamic scene graphs, physical interaction & affordance, human-centric 3D, robotics integration, and agent-facing 3D grounding
  • Datasets, Benchmarks & Infrastructure: datasets, metrics, simulators, toolchains, and surveys for fast research orientation

This list is intentionally embodied-3DV-first. It includes 3D generation and 3DGS work only when it helps understand, build, evaluate, or deploy 3D assets and world models for embodied agents. It is not a generic catalog of all 3D generation, editing, rendering, compression, or graphics papers.

Daily candidate feed. The automatically updated arXiv Daily is a high-recall, topic-tagged candidate archive across the six areas above. It is deliberately broader than this curated README: papers are promoted here only after manual primary-source verification.

Must Read

Start here if you want the shortest path through the field.

GoalStart with
Estimate temporally consistent video depthVideo Depth Anything, FlashDepth, ICDepth, ViGeo, DVD, RollingDepth
Deploy zero-shot monocular depth at the edgeZipDepth, DepthART
Recover metric-scale geometry from RGBMetricAnything, Metric DAv2, Depth Pro, UniDepthV2
Understand feed-forward 3D reconstructionAdvances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey, DUSt3R, MASt3R, VGGT
Track online / streaming 3D reconstructionDynamic 3D Gaussians, CUT3R, Spann3R, SLAM3R, StreamSplat, UniSim-SLAM
Study dynamic 4D feed-forward geometryPAGE-4D, 4D-VGGT, DynamicVGGT
Learn representation foundations3D Gaussian Splatting, NeRF, Instant-NGP, TensoRF
Generate object and simulation-ready assetsTRELLIS, Hunyuan3D 2.0, TripoSR, PartCrafter
Generate scenes and worldsInfinigen, SceneDreamer, Text2Room, WonderWorld, EmbodiedGen V2
Build articulated objectsArtGS, Articulate Anything, URDF-Anything, URDF-Anything+, PartNet-Mobility
Build an open-vocabulary 3D robot memoryCLIP-Fields, ConceptFusion, Open-Fusion, ConceptGraphs
Build persistent, action-conditioned 3D world modelsLearning 3D Persistent Embodied World Models, GEN3C, NeoVerse, Genie 4D, BWM
Pick datasets and benchmarksScanNet, Tanks and Temples, DTU, Objaverse, GSO, HM3D, TransBiolab

News

  • [2026-08-04] Added a six-topic, automatically refreshed arXiv Daily candidate archive, with manual verification required before promotion to this curated list.
  • [2026-06-15] Added a dedicated 3D Editing taxonomy under 3D Generation, covering object-level, scene-level, and dynamic / 4D editing methods.
  • [2026-04-30] Initialized Awesome-Embodied-3DV with a six-part taxonomy for data perception, representations, reconstruction, generation, embodied world models, and infrastructure.

Contents

📡 1. Data Perception

Data perception covers the sensor-facing and semantic layers: extracting geometric priors (depth, normals), understanding 3D semantics (detection, segmentation, grounding), active-imaging signals, and dense maps from 2D images, video, or physical sensors.

1.1 Geometric Priors

1.1.1 Monocular High-Fidelity Depth

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Metric Depth, 3DGS Relocalization, Sparse PnP Anchors, Temporal MemoryThe Chinese University of Hong Kong, ShenzhenRIDE: Relocalization-Informed Depth Estimation with 3D Gaussian SplattingarXivpaper
2026-09-08Diffusion Transformer, Single-Step Depth, Sharp Details, Dense PredictionEPFL / HUAWEI Bayer LabMarigold V2: Revisiting Diffusion Transformers for Monocular Depth EstimationSIGGRAPH Asia 2026project / github
2026-09-08Any Camera, Metric Point Cloud, Optional Intrinsics/Sparse DepthGoogle DeepMindOmniPoint: Universal Monocular Metric Pointcloud from Any CameraECCV 2026project
2026-08-30Transparent/Reflective Scenes, Bias-Aware Training, 30M ParametersThe University of Hong KongOptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging ScenesarXivproject
2026-08-17Pixel-Space Prediction, Fine Structures, Sharp Boundaries, Efficient DepthThe Chinese University of Hong Kong, ShenzhenPXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth EstimationarXivproject
2026-08-03Geometry-Invariant Adaptation, Non-Lambertian Surfaces, Mirror/Glass DepthChangchun Institute of Optics, Fine Mechanics and Physics, CASGIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth EstimationarXivpaper
2026-07-23UAV Depth, Arbitrary Camera Pose, Metric GeometryAerospace Information Research Institute, CASDAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOVarXivgithub
2026-07-20Fine-Detail Geometry, Sparse Volumetric Refinement, Metric ScaleTsinghua UniversityMoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivproject
2026-07-19Lightweight Foundation Depth, Camera-Conditioned Metric Depth, Edge DeploymentUniversity of TrentoDepthART: Scaling Foundation Monocular Depth to Tiny ModelsACM Multimedia 2026project / github
2026-07-19Metric Depth, Odometry Anchor, Recurrent SLAMUC BerkeleyDROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth EstimationarXivpaper
2026-07-17Stereo Distillation, Epipolar Cues, Metric DepthMichigan State UniversityGeometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular DeptharXivpaper
2026-07-14Auto-Regressive Depth, Coarse-to-Fine, Semantic GuidanceSun Yat-sen UniversityARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual ConditioningarXivpaper
2026-07-13Metric Point Map, Pixel-Wise Calibration, Camera DiversityThe University of Hong KongFoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric GeometryECCV 2026project / github / model
2026-07-09Lightweight Zero-Shot, 6.1M Parameters, On-Device DepthUniversity of BolognaZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any DeviceECCV 2026project / github
2026-05-27Multi-Layer Depth, Transparent Surfaces, Point ProcessPrinceton UniversitySeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined GroupingCVPR 2026github
2026-05-15VLM, Dense Metric Depth, Spatial ReasoningZhejiang UnivUnlocking Dense Metric Depth Estimation in VLMsarXivproject / github
2026-05-12Sparse 3D Anchors, Relative-to-Metric, Graph OptimizationTongji UniversityThe Midas Touch for Metric DepthCVPR 2026 Highlightproject
2026-03-28Universal Camera, Metric Depth, Zero-ShotMichigan State UniversityUniDAC: Universal Metric Depth Estimation for Any CameraCVPR 2026paper
2026-03-20Transparent Objects, Generative Opacification, Monocular Depth, SeeClear-396kUniversity of California, Los AngelesSeeClear: Reliable Transparent Object Depth Estimation via Generative OpacificationECCV 2026project / dataset
2026-03-17Diffusion Prior, Real-World Data, Monocular DepthNanjing University of Science and TechnologyIris: Bringing Real-World Priors into Diffusion Model for Monocular Depth EstimationCVPR 2026paper
2026-03-04Fine-Grained Geometry, Dual-Stream, EfficientUMass AmherstDAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationCVPR 2026paper
2026-01-29Sparse Metric Prompt, 20M Image-Depth Pairs, Metric Foundation ModelLi Auto IncMetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous SourcesECCV 2026project / github
2026-01-06Arbitrary-Resolution, Neural Implicit, Fine DetailsZhejiang UnivInfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsCVPR 2026project / github
2025-12-13Defocus Cue, Bokeh Stack, Metric DepthNanyang Tech UnivBoosting Monocular Metric Depth Estimation via Bokeh RenderingICML 2026project / github
2025-11-30Deterministic Diffusion, Dense Geometry, Fine DetailsHKUST(GZ)Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative ModelarXivproject / github
2025-11-13Any-View Depth, Metric Geometry, Multi-ViewByteDanceDepth Anything 3: Recovering the Visual Space from Any ViewsICLR 2026project / github
2025-10-27Unified Generation+Depth, Diffusion Prior, Zero-ShotHUSTMore Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion ModelsNeurIPS 2025github
2025-10-08Pixel-Space Diffusion, Flying-Pixel-Free, Point CloudsHUSTPixel-Perfect Depth with Semantics-Prompted Diffusion TransformersNeurIPS 2025project / github
2025-09-29VLM, Metric Depth, Sparse SupervisionMeta AIDepthLM: Metric Depth From Vision Language ModelsICLR 2026 Oralgithub
2025-07-03Monocular Geometry, Metric Scale, Sharp DetailsUSTC / MicrosoftMoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPS 2025project / github
2025-04-16Sliding Anchor, Unknown Intrinsics, Metric ScaleShanghai UnivMetric-Solver: Sliding Anchored Metric Depth Estimation from a Single ImagearXivproject / github
2025-02-27Metric 3D Points, Self-Prompt Camera, UncertaintyETH ZurichUniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerTPAMI 2026github
2025-02-26Cross-Context Distillation, Multi-Teacher, Fine-Detail DepthZhejiang Univ of TechnologyDistill Any Depth: Distillation Creates a Stronger Monocular Depth EstimatorarXivproject / github
2024-11-27Diffusion Distillation, Metric + Sharp, Boundary DetailQualcomm AI ResearchSharpDepth: Sharpening Metric Depth Predictions Using Diffusion DistillationCVPR 2025project / github
2024-10-02Metric Depth, Zero-Shot, Image PriorsAppleDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLR 2025github
2024-09-26Diffusion, Single-Step, Dense GeometryHKUST(GZ)Lotus: Diffusion-based Visual Foundation Model for High-quality Dense PredictionICLR 2025project / github
2024-06-13Monocular Depth, Foundation Model, Metric DAv2TikTok / HKUDepth Anything V2NeurIPS 2024project / github / metric models
2024-03-27Metric Depth, Universal, Zero-ShotETH ZurichUniDepth: Universal Monocular Metric Depth EstimationCVPR 2024github
2024-01-19Monocular Depth, Relative Depth, Foundation ModelTikTok / HKUDepth Anything: Unleashing the Power of Large-Scale Unlabeled DataCVPR 2024project
2023-12-04Monocular Depth, Zero-Shot, Affine-InvariantIntel LabsMarigold: Repurposing Diffusion-Based Image Generators for Monocular Depth EstimationCVPR 2024project
2023-07-20Metric 3D, Zero-Shot, Canonical SpaceAlibabaMetric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageICCV 2023github
2021-03-24ViT, Dense Prediction, Foundation, DepthIntel LabsDPT: Vision Transformers for Dense PredictionICCV 2021github
2019-07-02Robust Depth, Zero-Shot, Cross-Dataset, MiDaSIntel LabsMiDaS: Towards Robust Monocular Depth EstimationTPAMI 2022github

1.1.2 Temporally Consistent Video Depth

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-23Unified Video Model, Depth+Normals, Temporal ConsistencyAdobe ResearchUnified Video Dense Prediction from Disjoint DataarXivproject
2026-07-02Video Diffusion, In-Context Conditioning, Zero-ShotHKUSTICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context ConditioningECCV 2026project
2026-05-28Streaming Geometry, Dynamic Chunking, Depth+NormalsZhejiang UniversityTowards Consistent Video Geometry EstimationarXivproject
2026-05-11Camera Motion, 3D Consistency, Geometry EmbeddingHUSTGemDepth: Geometry-Embedded Features for 3D-Consistent Video DeptharXivgithub
2026-04-08Post-Processing, Scalable, Single-Image BackboneSeoul National UniversityVDPP: Video Depth Post-Processing for Speed and ScalabilityCVPR 2026 ECV Workshoppaper
2026-04-02Pose Refinement, Temporal Consistency, Monocular VideoAjou UniversityPTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal ConsistencyCVPR 2026paper
2026-03-12Deterministic Diffusion, Generative Prior, Long VideoHKUST(GZ)DVD: Deterministic Video Depth Estimation with Generative PriorsarXivproject
2026-01-06Temporal Stability, Monocular Video, Long SequenceETH ZurichStableDPT: Temporal Stable Monocular Video Depth EstimationarXivpaper
2025-12-20Endoscopic Geometry, Streaming Mamba, Metric DepthVanderbilt UniversityEndoStreamDepth: Temporally Consistent Monocular Depth Estimation for Endoscopic Video StreamsarXivgithub / paper
2025-12-11Sparse Keyframes, Propagation, Long-Video ConsistencyETH ZurichVideo Depth Propagation3DV 2026paper
2025-10-10Online Inference, Low Memory, Temporal ConsistencyHeidelberg UniversityOnline Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory ConsumptionarXivpaper
2025-07-02Diffusion Guidance, Scale Synchronization, Geometry ConsistencyTsinghua UniversityDepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth EstimationICCV 2025project
2025-04-092K Streaming, Mamba, 24 FPSNetflix Eyeline StudiosFlashDepth: Real-time Streaming Video Depth Estimation at 2K ResolutionICCV 2025 Highlightproject / github
2025-01-21Super-Long Video, Temporal Gradient, 30 FPSByteDanceVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPR 2025project / github
2024-11-28Long Video, Diffusion, Multi-Resolution AlignmentETH ZurichVideo Depth without Video ModelsCVPR 2025project

1.1.3 Geometric Consistency Prior

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-06Surface Normals, Transparent Objects, Rectified Flow, Edge RefinementZhejiang UniversityTransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal EstimationarXivproject
2026-07-28Stereo Depth, Walsh-Hadamard Mixing, Efficient InferenceAuthorsWHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token MixingarXivpaper
2026-07-22Stereo Diffusion Transformer, Flow Matching, Progressive RefinementBeihang UniversitySTEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow MatchingarXivpaper
2026-07-15Vision Features, SE(3) Latent Geometry, Visual NavigationGoogle DeepMindSeeSE3: Emergence of 3D Space in Vision FeaturesarXivpaper
2026-07-14Heterogeneous Cameras, Metric Depth, Real-TimeD-RoboticsX-Lens: Real-Time Metric Depth Estimation with Heterogeneous CamerasarXivproject / github
2026-07-10Video Generative Pretraining, Depth/Normals/Pose, Grounded 4DGoogle DeepMindVideo Generation Models are General-Purpose Vision LearnersECCV 2026project
2026-07-06Boundary-Centric Pretraining, Dense Spatial PerceptionRobbyantLingBot-Vision: Vision Pretraining for Dense Spatial PerceptionarXivproject / github
2026-03-02Zero-Shot Stereo, Structure Prompt, Motion PromptHUSTPromptStereo: Zero-Shot Stereo Matching via Structure and Motion PromptsCVPR 2026paper
2025-12-11Real-Time Stereo, Zero-Shot, Foundation ModelNVIDIAFast-FoundationStereo: Real-Time Zero-Shot Stereo MatchingCVPR 2026paper
2025-07-22Foundation Model, Depth/Normal/Pointmap, Multi-ViewSJTUDens3R: A Foundation Model for 3D Geometry PredictionICLR 2026paper
2025-04-15Surface Normals, Foundation Model, Video, TemporalAuthorsNormalCrafter: Learning Temporally Consistent Normals from Video Diffusion PriorsICCV 2025project / paper
2025-03-21Scalable Depth, Autoregressive, 2B ParametersBaiduDAR: Scalable Autoregressive Monocular Depth EstimationCVPR 2025project
2025-03-20Universal Camera, Spherical 3D, Any CameraETH ZurichUniK3D: Universal Camera Monocular 3D EstimationCVPR 2025github / paper
2025-03-11LiDAR Surface Normal, Dataset, Point CloudTU GrazLiSu: A Dataset and Method for LiDAR Surface Normal EstimationCVPR 2025github
2025-01-17Stereo, Foundation Model, RAFT-Style, Zero-ShotNVIDIAFoundationStereo: Zero-Shot Stereo MatchingCVPR 2025 Oral / Best Paper Nominationgithub
2025-01-17Stereo, Robust, Zero-Shot, Non-LambertianUniv of BolognaStereo Anywhere: Robust Zero-Shot Deep Stereo MatchingCVPR 2025project
2024-12-11Panoramic / Fisheye Depth, Zero-Shot MetricIntelDepth Any Camera: Zero-Shot Metric Depth from Any CameraCVPR 2025project
2024-10-24Monocular Geometry, Pointmap, Affine-InvariantUSTC / MicrosoftMoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain ImagesCVPR 2025 Oralgithub
2024-06-24Normal Estimation, 3D Priors, Surface GeometryNvidiaStableNormal: Reducing Diffusion Variance for Stable and Sharp NormalSIGGRAPH Asia 2024project
2024-03-27Diffusion, Effective Conditioning, ViT PriorsIIT DelhiECoDepth: Effective Conditioning of Diffusion Models for Monocular DepthCVPR 2024github
2024-03-22Metric 3D, Multi-Task, Geometry FoundationShanghai AI LabMetric3D v2: A Versatile Monocular Geometric Foundation ModelTPAMI 2025project
2024-03-18Geometry, Normals, Depth, Multi-TaskAppleGeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single ImageECCV 2024project
2024-03-01Inductive Biases, Surface Normal, OralImperial College LondonDSINE: Rethinking Inductive Biases for Surface Normal EstimationCVPR 2024 Oralgithub
2023-12-04High-Res, Patch-Wise, Model-AgnosticKAUSTPatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric DepthCVPR 2024project
2023-09-25Iterative Bins, Elastic, GRU, Classification-RegressionBeihang UnivIEBins: Iterative Elastic Bins for Monocular Depth EstimationNeurIPS 2023github
2023-04-13Internal Discretization, Continuous-Discrete, DepthETH ZurichiDisc: Internal Discretization for Monocular Depth EstimationCVPR 2023github

1.2 3D Semantic Understanding

1.2.1 Open-Vocabulary 3D Segmentation & Grounding

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Annotation-Free, Open-Vocabulary 3D, Language-Space LiftingNational Technical University of AthensGoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space LiftingarXivpaper
2026-07-21Referring Segmentation, 3DGS, Generalized GroundingPeking UniversityZeroSplat: Generalized Referring Segmentation in 3D Gaussian SplattingarXivproject
2026-06-23Open-Vocabulary BEV, 3DGS, Geometric ConstraintsKAIST AIOpen-Vocabulary BEV Segmentation with 3D-Aware Geometric ConstraintsECCV 2026paper
2026-06-04Open-Vocabulary, Functionality Segmentation, RoboticsAuthorsT-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality SegmentationarXivpaper
2026-05-07Open-Vocabulary, Gaussian Feature Field, CodebookTU Munich / GoogleOpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook AttentionarXivpaper
2025-11-20Open-Vocabulary, SAM3, Promptable, DETRMetaSAM 3: Segment Anything with ConceptsICLR 2026github / paper
2025-04-03Open-Vocabulary, Dual-Level Contrastive, Instance-AwareBIGAI / TsinghuaMPEC: Masked Point-Entity Contrast for Open-Vocabulary 3D Scene UnderstandingCVPR 2025project
2025-03-22Training-Free, MLLM Caption, Voxel GroupingNVIDIA Research TaiwanOpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene UnderstandingCVPR 2026project
2025-03-19SAM-2, 3D Tracking, Dynamic ProgrammingVinAIAny3DIS: Open-Vocabulary 3D Instance Segmentation with SAM-2CVPR 2025paper
2025-03-13Open-Vocabulary, LLM Canonical, Part SegmentationShandong Univ / TencentCoSMo3D: Open-World Promptable 3D Semantic Part Segmentation via LLM-Guided Canonical Spatial ModelingCVPR 2026 Oralgithub
2025-03-12Functional 3D, CoT, VLM, Training-Free, HighlightAuthorsFun3DU: Functional 3D Scene Understanding via Chain-of-ThoughtCVPR 2025 Highlightpaper
2025-01-02Panoptic, Open-Vocabulary, 3D Gaussian SplattingNUSPanoGS: Gaussian-based Panoptic Open-Vocabulary 3D Scene UnderstandingCVPR 2025project
2024-12-13Open-Vocabulary 3D, Structured Super-Gaussians, SegmentationGoogleSuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians3DV 2026paper
2024-12-12Open-Vocabulary, Foundation Dataset, Mask-Text PairsNVIDIAMosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D SegmentationCVPR 2025github
2024-07-02Open-Vocabulary, Mask-Snap-Lookup, Indoor/OutdoorHKUSTOpenIns3D: Open-Vocabulary 3D Segmentation with Mask-Snap-LookupECCV 2024github
2024-04-01Region-Level, Point-Language Contrastive, Multi-VLMHKURegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene UnderstandingCVPR 2024project / github
2024-03-19Open-Vocabulary, MLLM, Point-Entity-Text, nuScenesMPIOV3D: Open-Vocabulary 3D Understanding with Multi-Modal AlignmentCVPR 2024paper
2024-03-192D-Guided, 3D Proposals, SAM, InstanceVinAI / IBMOpen3DIS: Open-Vocabulary 3D Instance Segmentation with 2D-Guided Mask GenerationCVPR 2024github
2024-03-15Training-Free, View-Consensus, InstancePKUMaskClustering: View-Consensus for 3D Instance SegmentationCVPR 2024paper
2024-03-11Zero-Shot, Superpoint, SAM, Scene GraphPKUSAI3D: Zero-Shot 3D Instance Segmentation by Scene-Aware Incremental MergingCVPR 2024github
2024-01-31Segment Anything, 3D Gaussians, InteractiveZhejiang UniversitySAGD: Boundary-Enhanced Segment Anything in 3D GaussiansarXivgithub
2024-01-042D+3D Unified, Single Model, HighlightAuthorsODIN: A Single Model for 2D and 3D SegmentationCVPR 2024 Highlightgithub
2023-12-26Open-Vocabulary 3DGS, Segmentation, LanguageETH ZurichLangSplat: 3D Language Gaussian SplattingCVPR 2024 Highlightproject / github
2023-11-17Unified, Semantic+Instance+Panoptic, Single TransformerSamsungOneFormer3D: One Transformer for Unified 3D SegmentationCVPR 2024github
2023-06-23Open-Vocabulary 3D Instance, Mask Proposals, CLIPETH ZurichOpenMask3D: Open-Vocabulary 3D Instance SegmentationNeurIPS 2023project
2023-06-06Segment Anything 3D, Point Cloud, InteractiveVAST AISAM3D: Segment Anything in 3D ScenesCVPR 2024github
2023-03-16Open-Vocabulary, 3D Scene, Language FieldUC BerkeleyLERF: Language Embedded Radiance FieldsICCV 2023project / github
2022-11-28Open-Vocabulary 3D, Point Cloud, SegmentationETH ZurichOpenScene: 3D Scene Understanding with Open VocabulariesCVPR 2023project / github

1.2.2 3D Instance & Panoptic Segmentation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-08Feed-Forward, Open-Vocabulary, PanopticAuthorsEPS3D: End-to-End Feed-Forward 3D Panoptic SegmentationICML 2026github
2025-07Feed-Forward, Panoramic, DUSt3R-BasedAuthorsPanSt3R: Single-Feed 3D Geometry and Panoptic SegmentationICCV 2025paper
2025-05-14MoE, Multi-Dataset, PTv3, CLIP AlignmentUVAPoint-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic SegmentationICLR 2026project
2025-03-01Bayesian 3DGS, Training-Free, EIGSonyB3-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIGCVPR 2026project
2025-01-14End-to-End, 2D-to-3D Lifting, 3DGSCUHKUnified-Lift: Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian SplattingCVPR 2025github
2025-01-10Unsupervised Panoptic, Scene-CentricTU DarmstadtCUPS: Scene-Centric Unsupervised Panoptic SegmentationCVPR 2025 Highlightgithub
2025-01-06Zero-Shot Instance, SAM Prompts in 3DCUHK-SZ / MSRASAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation3DV 2025github
2024-07-033D Panoptic, LiDAR, Multi-SceneTsinghuaUniSeg3D: Unified 3D Panoptic SegmentationCVPR 2023github
2024-03-25Unsupervised 3D Instance, IndoorAuthorsUnScene3D: Unsupervised 3D Instance SegmentationCVPR 2024paper
2024-03-24Unified 6 Tasks, Single TransformerHUSTUniSeg3D: Unified 3D Segmentation with TransformerNeurIPS 2024paper
2023-12-15Point Cloud, Foundation Model, 3D UnderstandingShanghai AI LabPoint Transformer V3: Simpler, Faster, StrongerCVPR 2024github
2022-10-063D Instance Segmentation, Transformer, Point CloudETH ZurichMask3D: Mask Transformer for 3D Semantic Instance SegmentationICRA 2023github

1.2.3 3D Visual Grounding & Spatial Reasoning

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-233D-Aware VLM, Implicit+Explicit Geometry, RGB VideoNanyang Technological University3D-Aware VLMs with Implicit and Explicit GeometriesECCV 2026github
2026-07-143D Language Fields, Ambiguity Awareness, Object RetrievalAuthorsSaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object RetrievalarXivpaper
2026-06-23Agentic, Cognitive Map, Zero-Shot 3DSichuan UniversityAgentic Collaborative Cognition for Zero-Shot 3D UnderstandingECCV 2026project
2026-06-22Map-Grounded, MV3D-VQA, Dense RewardKAISTDense Reward for Multi-View 3D Reasoning with Global Maps and Local ViewsECCV 2026paper
2026-06-17Panoramic Reprojection, 3D VLM, Spatial ReasoningTechnical University of MunichOneCanvas: 3D Scene Understanding via Panoramic ReprojectionarXivproject
2026-06-04Part-Aware 3D-MLLM, Scene Understanding, GroundingXiamen UniversityPAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene UnderstandingarXivproject
2025-10-19Grounded CoT, 3D Reasoning, SceneCOT-185K DatasetBIGAI / PKU / TsinghuaSceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D ScenesICLR 2026paper
2025-07Scene Graph, LLM, Relation EncodingAIRI3DGraphLLM: 3D Scene Graph Learning with LLMsICCV 2025paper
2025-03-19Gesture+Language, Embodied Reference, +30%AuthorsGes3ViG: 3D Embodied Reference Understanding with Pointing GesturesCVPR 2025paper
2025-03-19Geometry VLM, Unified 3D Recon + Spatial ReasoningShanghai AI Lab / UCLA / SJTUG2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningCVPR 2026github
2025-03-12Dense Grounding, 6.2M Pairs, Hallucination BenchmarkUMich3D-GRAND: A Million-Scale Dataset for 3D GroundingCVPR 2025project
2025-03-113D-Informed, Spatial Reasoning, HighlightJHUSpatialLLM: Spatial Reasoning with 3D-Informed LLMsCVPR 2025 Highlightpaper
2025-03-10LLM Attention, Scene Magnifier, Cross-RoomSCUTLSceneLLM: LLM-Attention Adaptive 3D Scene UnderstandingCVPR 2025paper
2025-01-12LVLM-Guided, Hierarchical Feature, 3DGSFudan / NTUReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual GroundingCVPR 2025project
2025-01-04Generalist 3D LMM, Omni Superpoint TransformerAdelaide / Microsoft3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint TransformerCVPR 2025github
2024-07-18Object Identifiers, 3D VL, Unified TasksZJUChat-Scene: Bridging 3D Scene and Large Language Models with Object IdentifiersNeurIPS 2024github
2024-07-15Million-Scale 3D VL, Multi-Level ContrastiveBIGAISceneVerse: Scaling 3D Vision-Language Learning for GroundingECCV 2024project
2024-06-06Zero-Shot 3D Grounding, 2D VLM TransferNUSSeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingCVPR 2025project
2024-04-03Open-Vocabulary 3D Scene Graph, Open RelationsBoschOpen3DSG: Open-Vocabulary 3D Scene Graph GenerationCVPR 2024paper
2024-03-08Interactive 3D, Language Assistant, Point CloudCUHKLL3DA: Visual Interactive Instruction Tuning for 3D Language AssistantCVPR 2024github
2024-03CLIP Cross-Modal, Contrastive, Scene GraphAuthorsCCL-3DSGG: CLIP-Driven Contrastive Learning for 3D Scene Graph GenerationCVPR 2024paper
2023-11-21Generalist Embodied 3D Agent, VLA, ICMLBIGAILEO: An Embodied Generalist Agent in 3D WorldICML 2024project
2023-08-083D Visual Grounding, Referring Expression, Point CloudPeking University3D-VisTA: Pre-trained Transformer for 3D Vision and Text AlignmentICCV 2023github
2023-07-243D VQA, 3D Captioning, Scene UnderstandingShanghai AI Lab3D-LLM: Injecting the 3D World into Large Language ModelsNeurIPS 2023project / github
2023-033D Language Pre-training, Captioning, QAAuthors3D-VLP: 3D Vision-Language Pre-training with Contextual SceneCVPR 2023paper

1.3 Active Imaging & Sensors

1.3.1 Structured Light & Active Stereo

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-29iToF, Sensor-Intrinsic Uncertainty, Heteroscedastic RestorationTsinghua UniversityReliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space RestorationarXivpaper
2026-07-27Neural Structured Light, Metric Depth, Online SLAMPeking UniversityNSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and ReconstructionACM MM 2026paper
2026-07-20Projector-Camera, Feed-Forward 3DGS, Active IlluminationNingbo UniversityFF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera SystemarXivgithub
2026-06Active Stereo, 2DGS Supervision, RealSense DatasetHangzhou Dianzi UniversityGS-ASM: 2DGS-Supervised Active Stereo MatchingCVPR 2026paper
2026-05-07Adaptive 4D Illumination, Shape+Reflectance, Differentiable CaptureZhejiang UniversityDifferentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and ReflectanceCVPR 2026paper
2026-03Multi-Projector Structured Light, One-Shot Scan, Neural SDFKyushu UniversityMulti-view Stereo with Multiple Projectors for Oneshot Entire Shape Scan based on Neural SDF and DSSS DemultiplexingWACV 2026paper
2026-02-05Latent Diffusion, Fringe Projection, Reflective ObjectsYonsei UniversityLD-SLRO: Latent Diffusion Structured Light for 3-D Reconstruction of Highly Reflective ObjectsarXivpaper
2025-12-16Single-Shot, Neural Feature Decoding, Robust CorrespondencePeking UniversityRobust Single-shot Structured Light 3D Imaging via Neural Feature DecodingSIGGRAPH Asia 2025project
2025-12Event Camera, HDR Measurement, Confidence StereoUSTCEvent-based HDR Structured LightNeurIPS 2025github
2025-03-08Active Stereo, Phase Speckle, Cross-Scene GeneralizationSouthwest Jiaotong UniversityRGB-Phase Speckle: Cross-Scene Stereo 3D Reconstruction via Wrapped Pre-NormalizationarXivpaper
2025-02Unsupervised Structured Light, Neural SDF, Shadow-AwareKyushu UniversityNeural SDF for Shadow-Aware Unsupervised Structured LightWACV 2025paper
2025-01-13Matching-Free, Volume Rendering, Monocular Structured LightUSTCMatching-Free Depth Recovery from Structured LightarXivpaper
2024-10-20Neural SDF, One-Shot Scan, Low-Light/UnderwaterKyushu UniversityActiveNeuS: Neural Signed Distance Fields for Active Stereo3DV 2024paper
2024-06-06Virtual Pattern Projection, Depth Fusion, In-the-Wild StereoUniversity of BolognaActive Stereo in the Wild through Virtual Pattern ProjectionarXivgithub
2024-06Neural Inverse, Dense Depth, 3-4 PatternsUniversity of TorontoTurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured LightCVPR 2024project
2023-06-17Structured Light, Phase Unwrapping, LearningNanjing UniversityDeep Learning-Based Structured Light 3D Imaging: A SurveyarXivsurvey
2018-11-27Event Camera, Active Stereo, DepthTsinghua UniversityEvent-Based Structured Light for Depth ReconstructionIJCAS 2024paper

1.3.2 Panoramic & Fisheye Perception

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-16Wide-FOV, Egocentric, 4D Hand-ObjectRice UniversityEgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot LearningarXivproject / dataset
2026-03-19Panoramic Depth, VGGT, Geometry ConsistencySingapore University of Technology and DesignVGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth EstimationCVPR 2026paper
2026-03-18Panoramic Reconstruction, Permutation-Equivariant, 360°CornellPanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryCVPR 2026paper
2026-01-25RGB-D, Depth Completion, Reflective/TransparentRobbyantLingBot-Depth: Masked Depth Modeling for Spatial PerceptionarXivgithub / paper
2025-12-18Panoramic Foundation Model, Multi-Camera, Metric DepthInsta360 ResearchDepth Any Panoramas: A Foundation Model for Panoramic Depth EstimationCVPR 2026paper
2025-01Fisheye, Real-Time, Cassini, Multi-ViewSun Yat-Sen UnivOmniStereo: Real-time Omnidirectional Depth with Multiview Fisheye CamerasCVPR 2025github
2024-06-19Panoramic Depth, Semi-Supervised, MobiusHKUST(GZ)PanDA: Panoramic Depth Anything with Mobius Spatial AugmentationCVPR 2025project
2024-03-25360 Depth, Bi-Projection, ERP+ICOSAPHKUST(GZ)Elite360D: Efficient 360 Depth Estimation via Bi-Projection FusionCVPR 2024github
2021-09-06360 Depth, Indoor, Panoramic ImagesCERTHPano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth EstimationCVPRW 2021project / github

1.3.3 Event-Based & Time-of-Flight Imaging

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-26Active Event Stereo, High-Speed Depth, 150 FPSAuthorsTowards Ultrafast Depth Sensing Via Active Event-based Stereo VisionTPAMI 2026paper
2026-07-17Event Camera, Feed-Forward 3D, Temporal AggregationZhejiang UniversityEvent3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature AggregationarXivpaper
2026-06RGB+ToF Histogram, High-Resolution Metric Depth, LightweightTongji UniversityLiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth EstimationCVPR 2026 Highlightpaper
2026-06Sparse dToF, Zero-Shot Completion, Sensor GeneralizationKAISTDense Metric Depth Completion from Sparse Direct Time-of-Flight SensorsCVPR 2026paper
2026-06Image-Event Fusion, Monocular Depth, Linear ComplexityPeking UniversityAIMDepth: Asymmetric Image-Event Mamba for Monocular Depth EstimationCVPR 2026paper
2026-06Event-Image Depth, Hypothesis Volume, Iterative RefinementSoutheast UniversityDepth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth EstimationCVPR 2026paper
2026-04-16Event-Frame Stereo, Cross-Modal Prompting, High Dynamic RangeSoutheast UniversityBidirectional Cross-Modal Prompting for Event-Frame Asymmetric StereoCVPR 2026paper
2026-04-02Event Stereo, Data Factory, Cross-Modal DistillationUniversity of BolognaEventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active SensorsCVPR 2026project
2026-02-03Event Camera, Neural SDF, Single-Camera MeshSaarland UniversityEventNeuS: 3D Mesh Reconstruction from a Single Event Camera3DV 2026project / github / dataset
2025-12-20Event Camera, Structured Light, Real-Time RGB-DPolytechnique MontrealE-RGB-D: Real-Time Event-Based Perception with Structured LightarXivgithub
2025-09-18Event Camera, Depth Estimation, Any-to-AnyShanghai AI LabDepth AnyEvent: Event Camera Based Monocular Depth Estimation via Dense Correspondence DistillationarXivpaper
2025-09-08Event Camera, Multispectral, Structured LightETH ZurichEvent Spectroscopy: Event-based Multispectral and Depth Sensing using Structured LightarXivpaper
2025-05-28Burst-Encodable ToF, Long-Range Depth, Hardware-Aware CodingNanjing UniversityLearnable Burst-Encodable Time-of-Flight Imaging for High-Fidelity Long-Distance Depth SensingNeurIPS 2025github
2025-05Event, Distillation, Confidence-Guided, Pseudo-LabelsNUSDistil-E2D: Distilling Image-to-Depth Priors for Event-Based DepthNeurIPS 2025paper
2025-04-23ToF, Sparse Depth, 3DGS, SLAMAuthorsToF-Splatting: Dense SLAM using Sparse Time-of-Flight DepthICCV 2025paper / paper
2025-04-22Event Camera, Ray Density, 3D Conv, SpotlightTU BerlinDERD-Net: Learning Depth from Event-based Ray DensitiesNeurIPS 2025 Spotlightgithub
2025-03-03dToF, Video Depth Completion, Frequency SelectiveAuthorsSVDC: Consistent Direct Time-of-Flight Video Depth CompletionCVPR 2025paper / paper
2024-10-10Event Camera, Pose-Free, Gaussian Splatting, HighlightZhejiang UnivIncEventGS: Pose-Free Gaussian Splatting from a Single Event CameraCVPR 2025 Highlightgithub

1.3.4 Radar & RF Imaging

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10mmWave Radar, Complex-Valued Point Splatting, Material Model, Novel ViewsCornell Tech3D Point Splatting for mmWave Radar Novel View SynthesisarXivpaper
2026-09-09mmWave Radar, Metric Depth, Smoke/Fog/Darkness, 95K FramesRice UniversityGRADE: Single-Frame Generative Radar Depth Estimation Under Visual DegradationMobiCom 2026project+data

1.4 Dense Mapping Systems

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-09RGB-D 3DGS, Online Reconstruction, Reactive ControlMitsubishi Electric Research LaboratoriesSplatCtrl: Perception-Action Coupling via Gaussian Scene Representations and Reactive Robot ControlICRA 2026paper
2026-07-08Geometry-Only 3DGS, Dense Monocular SLAMBeihang UniversityGeoGS-SLAM: Geometry-Only Gaussian Splatting for Dense Monocular SLAMarXivpaper
2026-07-05LiDAR 3DGS SLAM, Geometry-Aware Covariance, Real-TimeAuthorsReal-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance CouplingarXivgithub
2026-07-02Dynamic Gaussian SLAM, Dual-Level Probability, Semantic MapAuthorsDL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level ProbabilityarXivpaper
2026-06-29Task-Conditioned 3DGS, Real-Time Mapping, Multi-Agent FusionMITGaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic MappingarXivpaper
2026-06-29RGB-Only Gaussian SLAM, Closed-Loop Geometry, Scale FeedbackCAS / USTCMyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAMIROS 2026paper
2026-06-27Object-Centric 3DGS, Lifelong Mapping, Dynamic MaintenanceBeijing Institute of TechnologyCubifyGS: Object-Centric 3D Gaussian Splatting for Lifelong Dynamic Scene MaintenanceIROS 2026paper
2026-03-10Uncertainty-Aware 3DGS, RGB-D SLAM, Loop DetectionGeorge Mason UniversityVarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAMCVPR 2026project / github
2025-11-20Language-Embedded, Open-Vocabulary, 3DGS SLAMKAISTLEGO-SLAM: Language-Embedded Gaussian Optimization SLAMarXivpaper
2025-07-25Neural SLAM, Dense Mapping, Self-SupervisedShanghai AI LabDINO-SLAM: Dense Tracking and Mapping with Self-Supervised Feature LearningarXivpaper
2025-07Dynamic Surface, Non-Rigid, 4D TrackingImperial College4DTAM: Dynamic Surface Gaussian SLAMCVPR 2025paper
2025-03-204DGS SLAM, Dynamic/Static, TrackingAuthors4D Gaussian Splatting SLAMICCV 2025paper / paper
2025-03-11Gaussian SLAM, Dense Reconstruction, RGB-DZhejiang UniversityGigaSLAM: Gaussian Splatting-based Large-Scale Dense SLAMarXivgithub
2025-03Multi-Agent, 3DGS SLAM, Loop ClosureAuthorsMAGiC-SLAM: Multi-Agent 3DGS SLAMCVPR 2025paper
2025-01-25Gaussian SLAM, Dynamic Environments, MonocularStanford / ETHWildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic EnvironmentsCVPR 2025project
2025-01-24Multi-Robot, Semantic, Heterogeneous, 3DGSStanfordHAMMER: Heterogeneous, Multi-Robot Semantic Gaussian SplattingRAL 2025project
2024-11-03Gaussian SLAM, Global BA, Monocular RGBETH / MetaSplat-SLAM: Globally Optimized RGB-only SLAM with 3D GaussiansCVPR 2025project
2024-09-10Single-Image Calibration, Geometric OptimizationETH / MetaGeoCalib: Learning Single-image CalibrationECCV 2024paper
2024-04-11Detector-Free SfM, Texture-PoorZhejiang UnivDetector-Free Structure from MotionCVPR 2024github
2024-04Pose Regression, Map-Relative, Multi-Scene, HighlightNiantic / OxfordMarepo: Map-Relative Pose Regression for Visual Re-LocalizationCVPR 2024 Highlightgithub
2024-02-20Neural SLAM, Survey, Radiance FieldsTUMHow NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a SurveyarXivproject
2023-12-11Dense SLAM, Gaussian Splatting, RGB-DTUMGaussian Splatting SLAMCVPR 2024github
2023-12-07End-to-End SfM, Differentiable BAMeta AI / OxfordVGGSfM: Visual Geometry Grounded Deep SfMCVPR 2024github
2023-12-04Gaussian SLAM, RGB-D, VolumetricCMU / MITSplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAMCVPR 2024project / github
2021-12-22Neural Mapping, Dense RGB-D SLAM, SDFHKUSTNICE-SLAM: Neural Implicit Scalable Encoding for SLAMCVPR 2022project

🧱 2. 3D/4D Representation

This section focuses on the mathematical and data-structure layer used to represent geometry, appearance, and motion.

2.1 Explicit & Hybrid Representations

2.1.1 Structure-Aware 3D Gaussian Splatting

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-21Gaussian Surfels, Topology Recovery, Mesh Extraction, Surface ReconstructionUniversity of Science and Technology of ChinaTopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface ReconstructionarXivgithub
2026-08-20Sparse Light Field, Casual Capture, 3D/4D Reconstruction, 3DGSCornell UniversitySparse Light Field Sampling Improves Casual 3D and 4D ReconstructionarXivproject
2026-08-17Pose-Free NVS, 3DGS Geometry, Visibility-Aware GuidanceAalto UniversitySplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View SynthesisarXivpaper
2026-08-13Feed-Forward 3DGS, 3D Anchors, Spatially Grounded TokensNational University of Defense TechnologyLocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian SplattingarXivproject
2026-08-03Streaming Feed-Forward, Persistent Geometry, Memory-Bounded 3DGSUniversity of Science and Technology of ChinaStreamSplat: Streaming Feed-Forward 3D Gaussian SplattingarXivpaper
2026-08-03View-Conditioned, Feed-Forward 3DGS, Generalizable ReconstructionTsinghua UniversityUniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D ReconstructionarXivpaper
2026-07-28Dynamic 3DGS, Adaptive Streaming, Volumetric VideoAuthorsSplatStream: Fine Granular Scalable Gaussian Splatting for Adaptive 3D Scene StreamingarXivpaper
2026-07-22Feed-Forward 3DGS, Adaptive Tokens, Compact RepresentationYonsei UniversityATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token ExpansionarXivproject
2026-04-16Feed-Forward 3DGS, Global Scene Tokens, CompactTel Aviv UniversityGlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene TokensarXivpaper
2026-04-16TokenGS, Learnable Gaussian Tokens, Pose-RobustNVIDIATokenGS: Decoupling 3D Gaussian Prediction from PixelsCVPR 2026 Highlightproject
2026-04-12UniSplat, Unposed Multi-View, Feed-ForwardUC BerkeleyUniSplat: Learning 3D Representations from Unposed Multi-View ImagesCVPR 2026paper
2026-02-02Feed-Forward 2DGS, Surface Continuity, Sparse ViewsShanghai Jiao TongSurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity PriorsICLR 2026paper
2025-07-31Sparse-View 4D, Monocular Fusion, Cross-VideoMeta AIMonoFusion: Sparse-View 4D Reconstruction via Monocular FusionICCV 2025paper
2025-06-17Anti-Aliasing, Gaussian, 3D, AdaptiveCUHKAAA-Gaussians: Anti-Aliasing 3D Gaussian SplattingICCV 2025paper
2025-06-05Feed-Forward 3DGS, Depth Parameterization, GeometryZhejiang UniversityRevisiting Depth Representations for Feed-Forward 3D Gaussian Splatting3DV 2026paper
2025-05-21Sparse-View Surface, Geometry-Prioritized, 2DGSNankai UnivSparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse ViewsCVPR 2025paper
2025-04-03Feed-Forward, No Camera, FreeSplatterZJUFreeSplatter: Pose-Free 3DGS from Sparse ViewsCVPR 2025paper
2025-03-13Few-Shot, Diffusion Prior, Repair + InpaintingZJU / AlibabaRI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion PriorsICCV 2025paper
2025-01-11Scene-Level, 3DGS, Open-Vocabulary, 4D LangSplatUPenn4D LangSplat: 4D Language Gaussian SplattingCVPR 2025paper
2024-12-09Large-Scale Dynamic, City, 4DGS, SpotlightZJUDynamicCity: Large-Scale 4D Gaussian City ModelingICLR 2025 Spotlightpaper
2024-09-26Progressive, Pruning, 3DGS, CVPRS-LabPUP 3D-GS: Progressive Pruning for 3DGSCVPR 2025paper
2024-03-262DGS, Surface Reconstruction, GeometryTUM2D Gaussian Splatting for Geometrically Accurate Radiance FieldsSIGGRAPH 2024project
2024-03-21Feed-Forward 3DGS, Sparse Views, GeneralizableMonash / TübingenMVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View ImagesECCV 2024 Oralproject / github
2024-03-14Multi-Scale, 3DGS, Level-of-DetailShanghai Jiao TongMulti-Scale 3DGSCVPR 2024paper
2023-12-19Feed-Forward 3DGS, Image Pairs, EpipolarMITpixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D ReconstructionCVPR 2024 Oralgithub
2023-12-01Distillation, Lightweight, 3DGS, SpotlightVirginia TechLightGaussian: Distilled 3D Gaussian SplattingNeurIPS 2024 Spotlightpaper
2023-11-303DGS, Structure, Sparse ViewsUSTCSparseGS: Real-Time 360 Sparse View Synthesis using Gaussian Splatting3DV 2025project
2023-11-27Anti-Aliasing, Mip, 3DGS, Best Student PaperInriaMip-Splatting: Alias-free 3DGSCVPR 2024 Oral / Best Student Papergithub
2023-11-27Compression, Compact, 3DGS, HighlightTsinghuaCompact 3D Gaussian SplattingCVPR 2024 Highlightpaper
2023-11-213DGS, Mesh Extraction, SurfaceETH ZurichSuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh ReconstructionCVPR 2024project
2023-08-083DGS, Real-Time Rendering, Explicit RadianceInria3D Gaussian Splatting for Real-Time Radiance Field RenderingSIGGRAPH 2023github / paper

2.1.2 Voxel, Mesh & Point-Cloud Innovations

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-18Sparse Voxel Latent, Fitting-Free 3DGS, Large-Area GenerationAmap, AlibabaGS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS GenerationarXivpaper
2026-07-22Shape Completion, Unified 3D Representation, Faithful GeometryAuthorsAxolotl3D: A Unified Framework for Faithful 3D Shape CompletionarXivpaper
2024-12-02Structured Latent, 3D Generation, MeshMicrosoftTRELLIS: Structured 3D Latents for Scalable and Versatile 3D GenerationCVPR 2025project / github
2024-03-04Sparse Voxel, Text/Image-to-3D, MeshStability AITripoSR: Fast 3D Object Reconstruction from a Single ImagearXivgithub
2022-12-16Point Cloud, Diffusion, Shape GenerationOpenAIPoint-E: A System for Generating 3D Point Clouds from Complex PromptsarXivgithub

2.2 Neural Implicit Representations

2.2.1 Advanced NeRF & SDF Architectures

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2025-07-15Unbiased SDF, Neural Implicit Surface, SDF-to-DensityCUHKUNIS: Unified Framework for Unbiased Neural Implicit SurfacesICCV 2025paper
2024-09-06α-NeuS, Alpha, SDF, Volume RenderingZJUα-NeuS: Alpha-Governed Neural Implicit SurfacesNeurIPS 2024paper
2023-12-29Objects as Volumes, NeRF, 3D Reconstruction, OralUPennObjects as Volumes: Feed-Forward 3D from a Single ImageCVPR 2024 Oralpaper
2023-04-13Zip-NeRF, Anti-Aliasing, Mip, GridGoogle ResearchZip-NeRF: Anti-Aliased Grid-Based NeRFICCV 2023project
2022-03-17Tensor Factorization, Radiance Fields, CompressionTsinghuaTensoRF: Tensorial Radiance FieldsECCV 2022project
2022-01-16Hash Grid, Real-Time NeRF, Neural GraphicsNVIDIAInstant Neural Graphics Primitives with a Multiresolution Hash EncodingSIGGRAPH 2022github / paper
2021-12-15EG3D, Triplane, 3D GAN, GenerativeNVIDIAEG3D: Efficient Geometry-aware 3D Generative Adversarial NetworksCVPR 2022 Oralgithub
2021-12-09Plenoxels, No Neural Network, Fast, OralUC BerkeleyPlenoxels: Radiance Fields without Neural NetworksCVPR 2022 Oralgithub
2021-12-07Ref-NeRF, Reflection, Specular, Best Student Paper HMGoogle ResearchRef-NeRF: Structured View-Dependent Appearance for NeRFCVPR 2022 Best Student Paper HMproject
2021-11-23Unbounded Scenes, Anti-Aliasing, NeRFGoogle ResearchMip-NeRF 360: Unbounded Anti-Aliased Neural Radiance FieldsCVPR 2022project
2021-06-20SDF, Surface Reconstruction, Neural RenderingMPI-ISNeuS: Learning Neural Implicit Surfaces by Volume RenderingNeurIPS 2021project
2021-04-13BARF, Bundle-Adjusting NeRF, OralUC BerkeleyBARF: Bundle-Adjusting Neural Radiance FieldsICCV 2021 Oralgithub
2020-08-05NeRF-W, Unbounded, In-the-WildGoogle ResearchNeRF in the Wild: Neural Radiance Fields for Unconstrained Photo CollectionsCVPR 2021github
2020-03-19NeRF, View Synthesis, Neural RenderingUC BerkeleyNeRF: Representing Scenes as Neural Radiance Fields for View SynthesisECCV 2020github / paper

2.3 Dynamic & Spatiotemporal 4D

2.3.1 4D Gaussian Splatting (4DGS)

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-23Dynamic 3DGS, Gradient Decoupling, Novel View SynthesisAuthorsGrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View SynthesisarXivpaper
2026-06-22Dynamic 3DGS, Visibility-Aware Densification, Temporal LifespanIndian Institute of ScienceTemporally Aware Densification for Dynamic 3D Gaussian SplattingarXivpaper
2025-077DGS, Spatial-Temporal-Angular, UnifiedAuthors7D Gaussian Splatting: Unified Spatial-Temporal-Angular GSICCV 2025paper
2024-10Event Camera, High-Speed, 4D, STD-GSAuthorsSTD-GS: SpatioTemporal-Disentangled Gaussian Splatting with Event CamerasICCV 2025paper
2023-10-124DGS, Dynamic Scenes, Real-Time RenderingZhejiang University4D Gaussian Splatting for Real-Time Dynamic Scene RenderingCVPR 2024project
2023-08-18Dynamic 3DGS, Scene Motion, Multi-View VideoCornellDynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis3DV 2025github / paper
2023-01-24K-Planes, Explicit, 4D, Space-TimeAuthorsK-Planes: Explicit Radiance Fields in Space, Time, and AppearanceCVPR 2023github
2023-01-23HexPlane, 4D Representation, Space-TimeCMUHexPlane: A Fast Representation for Dynamic ScenesCVPR 2023project
2021-06-24Dynamic NeRF, Deformation, Canonical SpaceGoogle ResearchHyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance FieldsSIGGRAPH Asia 2021project

2.3.2 Deformation Graphs & Canonical Spaces

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-24Structured Motion, SE(3), 4D ReconstructionTsinghua UniversitySM4RT: Learning Structured Motion Geometry for 4D ReconstructionarXivproject / paper
2023-12-04Deformation, Dynamic Radiance Fields, CanonicalETH ZurichSC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic ScenesCVPR 2024project
2023-06-05Non-Rigid Tracking, Neural Deformation, 4DTsinghuaNeuralangelo: High-Fidelity Neural Surface ReconstructionCVPR 2023project

🏗️ 3. 3D Reconstruction

Reconstruction systems recover objects or scenes from images, video, RGB-D, or multi-sensor streams under offline, online, and dynamic conditions.

3.1 Static Reconstruction

3.1.1 Object-Centric Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Active In-Hand Reconstruction, Uncertainty, Next-Best ViewShanghaiTech UniversityAURORA: Active Uncertainty-Driven Re-Orientation for In-Hand ReconstructionCoRL 2026project
2026-07-21Single-View 3D, Object Perception, Generative ReconstructionDeakin UniversitySeeing Before Generating: Object Perception Enhances Single-View 3D ReconstructionarXivproject
2026-06-17Sparse-View Object, Flow Steering, 3DGS RefinementGraz University of TechnologyFlowObject: Flow Steering for Bridging Generative Priors and Reconstruction FidelityarXivproject
2026-05-05Generative Reconstruction, Multi-View Alignment, PoseTsinghuaMix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose EstimationarXivproject
2025-11-19Single Image, Object Mesh, SAM 3DMeta AISAM 3D ObjectsGitHubgithub
2025-10-23Pose-Free Online, Free-Moving Objects, Constant MemorySUTDOnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving ObjectsNeurIPS 2025 Spotlightproject
2025-06-05RGB-D Object Completion, Novel Depth, Feed-ForwardCarnegie Mellon UniversityRaySt3R: Predicting Novel Depth Maps for Zero-Shot Object CompletionNeurIPS 2025project
2025-06Feed-Forward, Densification, Gaussian, DetailAuthorsGenerative Densification: Feed-Forward 3DGS DensificationCVPR 2025paper
2025-06Photogrammetry Foundation, Multi-Task, HighlightAuthorsMatrix3D: A Foundation Model for PhotogrammetryCVPR 2025 Highlightpaper
2025-04-04Sparse-View, Feed-Forward, Camera, GeometryAnt Research / StanfordFLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsCVPR 2025project
2024-07Ego-Centric, Autonomous Driving, Sparse-ViewAuthorsOmni-Scene: Omni-Gaussian for Ego-Centric Sparse-View ReconstructionCVPR 2025paper
2024-06-14Multi-View, Stereo, Feed-ForwardNAVER LabsMASt3R: Grounding Image Matching in 3D with MASt3RECCV 2024project / github
2023-12-21Multi-View, Pointmap, Pose-FreeNAVER LabsDUSt3R: Geometric 3D Vision Made EasyCVPR 2024project / github
2023-12-13Object Pose, Reconstruction, Model-BasedNVIDIAFoundationPose: Unified 6D Pose Estimation and Tracking of Novel ObjectsCVPR 2024project
2023-03-24Unknown Object, RGB-D, 6-DoF Tracking, Neural SDFNVIDIABundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown ObjectsCVPR 2023project / github

3.1.2 Large-Scale Scene Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Feed-Forward, Compositional Scene, Complete Meshes, Simulation-ReadyUniversity of Illinois Urbana-ChampaignFIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A MinutearXivproject / github
2026-09-04Bundle Adjustment, Multi-View Matching, Monocular Priors, Online+OfflineNAVER Labs EuropeBLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular PriorsECCV 2026paper
2026-08-31Real-to-Sim, Parse-Generate-Place, Composable Object AssetsByteDance SeedLucida: Parse, Generate, and Place for Composable Real-to-Sim Scene ModelingarXivproject
2026-08-25Single Image, Generative Reconstruction, Complete Object Assets, Scene AssemblyHuaweiSceneReGen: Generative Reconstruction of 3D Scenes from a Single ImagearXivpaper
2026-08-18Instance-Grouped 3DGS, Semantic Reconstruction, Referential Scene GraphShanghai Jiao Tong UniversityGroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian SplattingarXivpaper
2026-08-18Generative NVS, Reconstruction/Generation Split, Scene CoordinatesKAISTGenRec: Knowing Where to Reconstruct and Where to GeneratearXivproject
2026-08-18Long Sequence, Chunk Priors, Sim(3) Assembly, Test-Time AdaptationKosmo ResearchGeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric AssemblyarXivproject
2026-08-15Long Sequence, Scale-Consistent Alignment, Test-Time AdaptationNorthwestern Polytechnical UniversityVGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D ReconstructionACM Multimedia 2026github
2026-08-12Streaming Multi-View, Metric 3D, Feed-Forward Prior, Object DetectionETH ZurichMap-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming InputsECCV 2026project
2026-08-07Sparse View, Editable Indoor Scenes, Executable Scene ProgramsCity University of Hong KongScenix: Sparse-View 3D Scene Reconstruction via Executable Scene ProgramsarXivpaper
2026-07-31Active Reconstruction, Next-Best-View, Predictive EntropyFudan UniversityGO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D ReconstructionarXivpaper
2026-07-23Underwater 3D, Feed-Forward Reconstruction, Degradation AdaptationHKUSTWAT3R: Feedforward Underwater 3D ReconstructionarXivproject
2026-07-15Feed-Forward Driving Reconstruction, Layered 3DGS, Dynamic ActorsNVIDIAInstant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene SimulationarXivproject / github / docs
2026-07-103D Foundation Model, Global SfM, Bundle AdjustmentHKUSTGlob3R: Global Structure-from-Motion with 3D Foundation ModelsarXivproject
2026-07-08Feed-Forward 3D, Unposed Images, Drift-RobustAuthorsNoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D ReconstructionECCV 2026paper
2026-06-09RGB-T, Thermal Geometry, Low-LightUniversity of MinnesotaDarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight TaxarXivproject
2026-06-02Single Image, Physics-in-the-Loop, Simulation-ReadySeoul National UniversitySimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single ImagearXivproject
2026-05-14VGGT, Scaling, Static+DynamicUniversity of OxfordVGGT-Omega: Scaling VGGT to Large-Scale 3D ReconstructionCVPR 2026 Oralproject / github / demo / model
2026-05-07Feed-Forward 3D, Token Reduction, Long SequencePeking UniversitySpark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D ReconstructionarXivpaper
2026-04-30Generalizable, Sparse-View, Unposed Images, OutdoorUIUC / NVIDIAGenWildSplat: Generalizable Sparse-View 3D Reconstruction from Unconstrained ImagesarXivpaper
2026-03-24Panoramic Video, Pose-Free 3DGS, Consistent Depth PriorUniversity of Chinese Academy of SciencesPose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth PriorsCVPR 2026github
2026-03-16Event-to-Edge, Pose-Free, Gaussian ReconstructionKAISTE2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D ReconstructionCVPR 2026paper
2026-02-26VGGT, TTT, Large-ScaleNVIDIAVGG-T^3: Offline Feed-Forward 3D Reconstruction at ScaleCVPR 2026project
2026-02-03Single Image, Object Decomposition, Occlusion-Aware Scene ReconstructionUniversity of California, San DiegoSeeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal3DV 2026project
2025-09-24Mirror Stereo, Single-View 3D, Symmetry ConstraintUniversity of OxfordReflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections3DV 2026project / github / dataset
2025-09-16Universal 3D, Metric Reconstruction, Optional PriorsMeta AIMapAnything: Universal Feed-Forward Metric 3D Reconstruction3DV 2026project
2025-08-05Unposed Multi-View, 3DGS, Semantic ReconstructionSungkyunkwan UniversityUni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View ImagesCVPR 2026paper
2025-07-17Permutation-Equivariant, Visual Geometry, Point MapsOxford / Metaπ³: Permutation-Equivariant Visual Geometry LearningICLR 2026github
2025-07Monocular Prior, MVS, DTU/Tanks SOTAAuthorsMonoMVSNet: Monocular Prior Guided MVSICCV 2025paper
2025-07Latent Align, Stereo+Monocular, HighlightAuthorsBridgeDepth: Unified Monocular and Stereo DepthICCV 2025 Highlightpaper
2025-06-30Video-Depth Augmentation, Scalable Training, Feed-Forward 3DAustralian National UniversityPuzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D ReconstructionNeurIPS 2025project
2025-06-03Monocular Depth, Dynamic Video, AlignmentHKUST / CUHK / HKUAlign3R: Aligned Monocular Depth Estimation for Dynamic VideosCVPR 2025github
2025-06Aerial-Ground, Large-Scale, 3DGSAuthorsHorizon-GS: Unified Aerial-Ground 3DGSCVPR 2025paper
2025-06Autonomous Driving, Feed-Forward, 3DGSAuthorsEVolSplat: Feed-Forward 3DGS for Urban DrivingCVPR 2025paper
2025-06Sparse-View, Super-Resolution, 3DGSAuthorsS2Gaussian: Sparse-View Super-Resolution 3DGSCVPR 2025paper
2025-05-05Relative Camera Pose, Regression, LocalizationAalto / HKUReloc3r: Large-Scale Training of Relative Camera Pose RegressionCVPR 2025github
2025-03-28Feed-Forward, Surface, MVS, Multi-ViewAuthorsMVSAnywhere: Zero-Shot Multi-View StereoCVPR 2025paper / paper
2025-03-17Feed-Forward, Multi-View, Auxiliary PriorsETH / MicrosoftPow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene PriorsCVPR 2025project
2025-03-14Feed-Forward 3D, Pose-Free, Point MapMeta AIVGGT: Visual Geometry Grounded TransformerCVPR 2025 Best Paperproject / github
2025-03-03Multi-View, Symmetric, 1000+ Images, O(N)NAVER LabsMUSt3R: Multi-view Network for Stereo 3D ReconstructionCVPR 2025project
2025-01-23Feed-Forward, 1500+ Images, 251 FPSMeta AIFast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassCVPR 2025project
2024-12-12Feed-Forward, Online, Dense, MonocularShanghai AI Lab / PKUSLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosCVPR 2025 Highlightproject
2024-12-09Sparse View, Single-Stage, 2 SecondsMeta AIMV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsCVPR 2025 Oralproject
2024-09-27Multi-View Reconstruction, Matching, MVSNAVER LabsMASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion3DV 2025github
2024-08-28Spatial Memory, Feed-Forward, Multi-ViewHKUSpann3R: 3D Reconstruction with Spatial Memory3DV 2025 Oral (Best Paper Candidate)project

3.2 Streaming & Online Reconstruction

3.2.1 Dense Semantic/Instance Mapping

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-07Online SLAM, Functional Scene Graph, Interaction Elements, Map MemoryTsinghua UniversityFunctional-SLAM: Interaction-Aware Mapping with Online Functional Scene GraphsarXivgithub
2026-09-01Training-Free, Open-Vocabulary Instance Map, RGB-D/Monocular SLAMUniversity of Technology SydneyVOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAMarXivpaper
2026-08-24Spatio-Temporal SLAM, Open-Vocabulary, 4D Scene Graph, VLNCarnegie Mellon UniversitySuperMap: A Spatio-Temporal SLAM System for Visual-Language NavigationarXivproject
2026-08-18Open-Vocabulary Map, Instance Preservation, Fine-Grained Retrieval, Target AbsenceXi'an Jiaotong UniversityOVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained ObjectsarXivgithub
2026-06-23Object-Level Map, Open-Vocabulary, RelocalizationZhejiang UniversityCompact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual RelocalizationarXivpaper
2026-05-05Online, Voxel+Instance, Open-Vocabulary MappingÖrebro UniversityFUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic MappingarXivproject
2026-05-03Gaussian-Language Map, Zero-Shot Navigation, Multi-ScaleCASIAMulti-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and ReasoningarXivgithub
2026-03-04Semantic 3DGS, Online, CLIP, Open-VocabularyNational University of SingaporeEmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene UnderstandingCVPR 2026project
2025-08-02Open-Vocabulary, Hybrid 3DGS+TSDF, Dense MappingTsinghuaOpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian SplattingarXivproject
2025-07Feed-Forward, Panoramic Segmentation, DUSt3RAuthorsPanSt3R: Single-Feed 3D and Panoptic SegmentationICCV 2025paper
2023-10-05Open-Vocabulary, RGB-D, TSDF, Real-Time MappingUniversity of ArkansasOpen-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene RepresentationIROS 2024project / github
2023-02-14Open-Set, Multimodal, 3D Map, Language QueryMITConceptFusion: Open-set Multimodal 3D MappingICRA 2023project / github
2022-10-11Implicit Field, CLIP, Semantic Search, Robot MemoryNew York UniversityCLIP-Fields: Weakly Supervised Semantic Fields for Robotic MemoryICRA 2023project

3.2.2 Continuous Neural Tracking & Mapping

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-03Online 3R, Multi-Relative Pose Query, Pose-Graph OptimizationNational Yang Ming Chiao Tung UniversityScal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D ReconstructionECCV 2026project
2026-09-01Online Feed-Forward 3R, Unordered UAV Images, Retrieval+RetryWuhan UniversityOn-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV ScenariosarXivgithub
2026-08-03Active Reconstruction, Ergodic Coverage, Trajectory OptimizationJohns Hopkins UniversityTRACE: Ergodic Trajectory Optimization for Active Scene ReconstructionarXivgithub
2026-08-03Feed-Forward SLAM, Sim(3) Factor Graph, Persistent MappingUlsan National Institute of Science and TechnologyUniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) OptimizationECCV 2026project
2026-07-25Semantic SLAM, Data Association, Object LandmarksMITSemantic Semi-Incremental Data-Association-Free Object SLAMarXivpaper
2026-07-23Gaussian SLAM, Large-Scale Mapping, Real-TimeAthena Research CenterGLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial DecompositionIROS 2026github
2026-07-16Multi-Agent 3R, RGB Video, Point-Map FusionUniversity of BolognaMAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB VideosarXivproject
2026-07-16Incremental 3DGS, Unordered Capture, Global ConsistencyInriaImmediate 3D Gaussian Splat Reconstruction of Unordered Input with Global ConsistencySIGGRAPH 2026paper
2026-07-01Long-Sequence, Instance Anchors, Persistent Spatial MemoryBeijing Jiaotong UniversityLIST3R: Long-sequence Instance-aware 3D ReconstructionarXivproject
2026-06-233DGS-SLAM, Memory-Efficient, Outdoor MappingUniversity of MinnesotaPocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAMICRA 2026github
2026-06-20RGB+Pose, 3DGS Scene Regression, Robot CapturePeking UniversityACEsplat: Accelerated 3D Gaussian Scene Regression via RGB and Poses OnlyarXivpaper
2026-06-193DGS-SLAM, Degeneracy-Robust, Real-Time TrackingNanyang Technological UniversitySpectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAMIROS 2026paper
2026-06-18LiDAR-Inertial-Thermal, 3DGS Mapping, Illumination-RobustShenzhen UniversityLIT-GS: LiDAR-Inertial-Thermal Gaussian Splatting for Illumination-Robust MappingIROS 2026paper
2026-06-03Streaming, Transient Anchors, Long-Horizon MappingAuthorsAnchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual MappingarXivpaper
2026-06Spatial-Difference Sensor, Edge-Guided Tracking, 3DGSTsinghua UniversitySDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D ReconstructionCVPR 2026paper
2026-06Stereo 3DGS-SLAM, Auto-Exposure Robustness, Photometric MappingSouth China University of TechnologyAERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAMCVPR 2026github
2026-05-10VGGT, Retrieval, Constant MemoryFudan UniversityRetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity RetrievalarXivgithub
2026-04-244DGS-SLAM, Optical Flow, Dynamic MappingNational University of SingaporeFlow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAMCVPR 2026github
2026-04-15Streaming, Feed-Forward, Long VideoRobbyantLingBot-Map: Geometric Context Transformer for Streaming 3D ReconstructionarXivgithub
2026-02-13Streaming, Autoregressive, Long Sequence3DAgentWorldLongStream: Long-Sequence Streaming Autoregressive Visual GeometryCVPR 2026project
2026-01-03StreamVGGT, KV Cache, Memory CompressionSun Yat-sen UniversityXStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded TransformerarXivgithub
2025-09-30TTT, Online, Long ContextShanghai AI LabTTT3R: 3D Reconstruction as Test-Time TrainingarXivproject
2025-08-14Streaming, Causal Transformer, SequentialNTU / Shanghai AI LabSTream3R: Scalable Sequential 3D Reconstruction with Causal TransformerarXivproject
2025-01-21Online 3D, Recurrent Pointmap, StreamingMeta AICUT3R: Continuous 3D Perception Model with Persistent StateCVPR 2025 Oralproject / github
2024-12-16MASt3R, Dense SLAM, Real-TimeImperial College LondonMASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsCVPR 2025github
2023-09-05Global BA, Neural Implicit, Dense RGB-D SLAMUniversity of BolognaGO-SLAM: Global Optimization for Consistent 3D Instant ReconstructionICCV 2023project / github
2022-11-21Hybrid SDF, Dense RGB-D SLAM, KeyframesIdiap Research InstituteESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance FieldsCVPR 2023project / github
2021-08-24Deep SLAM, Dense BA, Monocular/Stereo/RGB-DPrinceton UniversityDROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasNeurIPS 2021github
2021-03-23Neural Implicit, Online RGB-D, Dense SLAMImperial College LondoniMAP: Implicit Mapping and Positioning in Real-TimeICCV 2021project / github

3.3 Dynamic 3D Reconstruction

3.3.1 Non-Rigid Tracking & Reconstruction

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-08Long-Range 4D Motion, 3D Queries, Occlusion-Robust Trajectory ChainingCarnegie Mellon UniversityPoint4D: Long-range 4D Motion ReconstructionarXivproject
2026-09-08Event Stream, Extreme-Low-Frame-Rate RGB, Dynamic 3DGS, Real-TimeMacau University of Science and TechnologyEdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene ReconstructionNeurocomputinggithub+dataset
2026-09-05Sparse-View 4D, Spatio-Temporal Depth Alignment, Dynamic 3DGSBIGAIUniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth AlignmentECCV 2026project
2026-08-18Query-Conditioned 4D, Scene Flow, Dynamic Points, Sparse-to-DenseKosmo ResearchUniQuery4R: Unified 4D Scene Reconstruction from a Single QueryarXivproject
2026-07-29Articulated Objects, Structure-aware 3DGS, Part ConnectivityPOSTECHStructureGS: Structure-aware Gaussian Splatting for Articulated Object ReconstructionarXivpaper
2026-07-21Streaming 4D, Instance Grounding, Geometry TransformerHorizon RoboticsIGGT4D: Streaming 4D Instance-Grounded Geometry TransformerarXivproject
2026-07-16Online Dynamic NVS, Space-Time Memory, Real-TimeUniversity of WashingtonOnline Neural Space Time Memory for Dynamic Novel View SynthesisarXivproject
2026-07-01Dynamic Gaussian Reconstruction, Monocular Video, GenerativeStanford UniversityWorld from Motion: Generative Dynamic Gaussian Reconstruction from Monocular VideoarXivproject
2026-06-23Articulated Digital Twin, RGB-D, URDF ExportETH ZurichArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D VideosICRA 2026 Workshoppaper
2026-06-22Monocular Video, 4DGS, In-the-Wild Non-RigidCarnegie Mellon UniversityLift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-WildarXivproject
2026-06-22Dynamic Driving, Sparse Voxels, LiDAR-GuidedHuawei Paris Research CenterDrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene ReconstructionarXivpaper
2026-06-09Future Extrapolation, 4DGS, Autonomous DrivingTsinghua UniversityEnvision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous DrivingarXivproject
2026-06-09Manipulation Video, Decoupled 3DGS, Scene GraphAuthorsManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian SplattingarXivpaper
2026-06-02Object Permanence, Differentiable Physics, 4DGSAuthorsPersistGS: Differentiable Physics for Object Permanence in 4D Gaussian SplattingCVPR 2026 Workshoppaper
2026-06Single Event Camera, Deformable 3DGS, High-Speed 4DShanghaiTech UniversityFastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event CameraCVPR 2026paper
2026-04-10Feed-Forward, Unconstrained Views, Semantic-GeometryNanyang TechFF3R: Feedforward Feature 3D Reconstruction from Unconstrained ViewsCVPR 2026 Findingspaper
2026-04-10Dynamic 4D, Semantic Prior, Gaussian SLAM, Action-ControlUniversity of ZurichGenie 4D: Semantic-Prior-Guided 4D Dynamic Scene ReconstructionarXivpaper
2026-04-10Dynamic/Static Disentanglement, Uncertainty-Aware, Feed-ForwardZhejiang UniversityRobust 4D VGT: Robust 4D Visual Geometry Transformer with Uncertainty-Aware PriorsarXivpaper
2026-04-07Functional Scenes, Egocentric Interaction, URDF/USDStanfordFunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction VideosCVPR 2026paper
2026-04-05Sparse Camera, 4DGS, Neural Decay, CVPR 2026Authors4C4D: 4 Camera 4D Gaussian SplattingCVPR 2026paper / paper
2026-03-30Dynamic Surface, Explicit Geometry, High FidelityAustralian National University4DSurf: High-Fidelity Dynamic Scene Surface ReconstructionCVPR 2026paper
2026-03-21RayMap, Dynamic, StreamingUniversity of Illinois ChicagoRayMap3R: Inference-Time RayMap for Dynamic 3D ReconstructionarXivproject / github
2026-03-09Dynamic VGGT, Autonomous Driving, 4DFudan UniversityDynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingarXivpaper
2025-11-23Dynamic Geometry, Spatiotemporal, VGGTHuazhong University of Science and Technology4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry EstimationarXivpaper
2025-11-07Motion-Aware, Monocular Video, Bundle AdjustmentKAIST4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic ScenesNeurIPS 2025paper
2025-10-20VGGT-4D, Pose/Geometry, Dynamic MaskHarvardPAGE-4D: Disentangled Pose and Geometry Estimation for 4D PerceptionICLR 2026project
2025-08-133D Reconstruction, Human Motion, VideoShanghai AI LabHuman3R: Reconstructing 3D Human Avatars from Monocular VideoCVPR 2025project
2025-07Human, Multi-View, Sparse, Robust, RoGSplatAuthorsRoGSplat: Robust Generalizable Human Gaussian SplattingCVPR 2025paper
2025-06-11Dynamic Human, Temporal Consistency, 4DTsinghuaCARI4D: Cross-Modal Alignment and Reconstruction for Interactive 4D HumanCVPR 2025project
2025-06-10Online, Dynamic 3DGS, Uncalibrated VideoUniversity of British ColumbiaStreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video StreamsICLR 2026project / github
2025-06-094DGS, Transformer, Monocular VideoMeta Reality Labs4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosNeurIPS 2025 Spotlightproject
2025-06-02Video Generators, 4D GeometryOxford VGG / NAVER LABSGeo4D: Leveraging Video Generators for Geometric 4D Scene ReconstructionICCV 2025 Highlightpaper
2025-06Self-Supervised, Dynamic, Driving, FlowAuthorsSplatFlow: Self-Supervised Dynamic 3DGS with Neural Motion FlowCVPR 2025paper
2025-06Few-Shot, Personal Avatar, HighlightAuthorsFRESA: Personalized 3D Human Avatar from Few ImagesCVPR 2025 Highlightpaper
2025-06Real-Time Avatar, 166fps, Gaussian, HighlightAuthorsMMLP-Human: Real-Time High-Fidelity Gaussian Human AvatarCVPR 2025 Highlightpaper
2025-05-274D, Dual Correspondences, Dynamic VideoNUS / Shanghai AI LabC4D: 4D Made from 3D through Dual CorrespondencesICCV 2025project
2025-05-144D Pointmaps, Dynamic-Static DisentanglementKAIST / ETH / SonyD2USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic ScenesNeurIPS 2025project
2025-05-02Training-Free, Motion Disentangle, DUSt3RWestlake / MPIEasi3R: Estimating Disentangled Motion from DUSt3R Without TrainingICCV 2025project
2025-04-174D Tracking, Feed-Forward, Pointmap, TrackingMPI / UC BerkeleySt4RTrack: Simultaneous 4D Reconstruction and TrackingICCV 2025paper / paper
2025-04-07Neural Rendering, Human, No EyesUniversity of CambridgeSeeing Without Eyes: Neural Human Rendering from Monocular VideoCVPR 2025project
2025-03-24Multi-Object, 4D, In-the-Wild VideosCMUGenMOJO: Robust Multi-Object 4D Generation for In-the-wild VideosCVPR 2025project
2025-02-27Layered Avatar, Hair, Face, MetaAuthorsLUCAS: Layered Universal Codec AvatarsCVPR 2025paper / paper
2025-01-223D Reconstruction, Canonical, Multi-ViewStanfordUniCon3R: Unified 3D Reconstruction and RecognitionCVPR 2025project
2024-12-03Single Image, Animatable, Avatar, 4DGSAuthorsAniGS: Animatable Gaussian Avatar from a Single ImageCVPR 2025paper / paper
2024-11-274D Generation, Multi-View Video, DiffusionGoogleCAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsCVPR 2025project
2024-10-28Dynamic Geometry, DUSt3R, MotionUniversity of OxfordMonST3R: Estimating Geometry in the Presence of MotionICLR 2025project / github

3.3.2 Deformation Graphs & Canonical Spaces

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-02-124D Dynamic, Monocular Video, Tree-ChainsCornellWorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-ChainsICLR 2026paper
2025-11-014D Scene, Feed-Forward, Controllable, Video DiffusionShanghai AI LabDiff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction ModelsCVPR 2026paper
2025-10-15Dynamic 4D, Gaussian, CanonicalZhejiang UniversityDirector: Directed Generative Models for 4D Scene EvolutionCVPR 2025project
2025-08-06Multi-Baseline, Generalizable, GaussianAuthorsMuGS: Multi-Baseline Generalizable Gaussian SplattingICCV 2025paper / paper
2025-07Surface, Gaussian Surfels, 2DGS, Sparse-View, SpotlightAuthorsMAtCha Gaussians: Atlas Charting with 2D Gaussian SurfelsCVPR 2025 Spotlightpaper
2025-07SDF+3DGS, Hybrid, Surface, ICCVAuthorsSurfaceSplat: SDF+3DGS Hybrid Surface ReconstructionICCV 2025paper
2025-07Sparse-View, Implicit, Voxel, ConsistencyAuthorsSparseRecon: Sparse-View Implicit Surface ReconstructionICCV 2025paper
2025-07Low-Texture, Reflection, Unified, +21%AuthorsHiNeuS: Unified Neural Implicit Surface ReconstructionICCV 2025paper
2025-06Joint Human+Scene, MASt3R ExtensionAuthorsHAMSt3R: Joint Human and Scene 3D ReconstructionICCV 2025paper
2024-11-204D Reconstruction, Gaussian Splatting, ForwardShanghai AI LabForge4D: Gaussian Splatting for Forward Facing 4D ReconstructionarXivpaper

🎛️ 4. 3D Generation

This section tracks methods that create new 3D assets, parts, articulated objects, scenes, and editable 3D worlds, with emphasis on physical and simulation use.

4.1 Object-Level Generation

4.1.1 Image/Text to 3D Mesh

Single Image / Text-Conditioned Object Mesh
DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Multi-View, Reconstruction Prior, Noise Inversion, Faithful Asset CompletionHKUSTReconPlusGen: Injecting Reconstruction Prior into Multi-view 3D Generation through Noise Inversion and ModulationarXivpaper
2026-09-09Image-to-3D, Partial Geometry, Test-Time Guidance, SAM 3DUniversity of OxfordGuiding Image-to-3D Generation with Test-Time Partial ObservationsarXivpaper
2026-07-22Code-Native, Programmable, 3D AssetsAuthorsNova3D: Code-Native Generation of Programmable 3D AssetsarXivpaper
2026-06-23Image-to-3DGS, Sparse Voxel, High-Fidelity AssetsThe Australian National UniversityFLUX3D: High-Fidelity 3D Gaussian Generation with Diffusion-Aligned Sparse RepresentationarXivpaper
2026-06-23Vehicle Assets, 3D-Consistent Views, SimulationShanghai Jiao Tong University3DCarGen: Scalable 3D Car Generation via 3D-consistent Multi-view SynthesisarXivpaper
2026-06-22Constraint Meshes, Controllable Assets, TRELLISUniversity of TübingenArbor: Explicit Geometric Conditioning for Controllable 3D Asset GenerationarXivproject
2026-05-22Image-to-3D, Deployable Mesh, UV, Real-TimeMeta Reality LabsAssetGen: Deployable 3D Asset Generation at Interactive SpeedarXivpaper
2026-05-11Pixel-Aligned, Image-to-3D, PBR, Multi-ViewTencent ARCPixal3D: Pixel-Aligned 3D Generation from ImagesSIGGRAPH 2026project / github / demo
2026-05-01Pose-Aware, Diffusion, 3D Geometry, DirectRenmin UniversityPAD: Pose-Aware Diffusion for 3D GenerationarXivpaper
2026-04-22Simulation-Ready, PBR, Part-Aware, ArticulationByteDance SeedSeed3D 2.0: Advancing High-Fidelity Simulation-Ready 3D Content GenerationarXivdemo
2025-12-16Image-to-3D, O-Voxel, PBR MaterialsMicrosoftNative and Compact Structured Latents for 3D GenerationarXivgithub / model
2025-11-20Single Image, Multi-Object, Layout, SAM 3DMeta AISAM 3D: 3Dfy Anything in ImagesarXivgithub
2025-10-23Single Image, Pose-Grounded, Flow MatchingMeta AICUPID: Generative 3D Reconstruction via Joint Object and Pose ModelingarXivproject
2025-10-09PBR, Material Diffusion, Relighting, Single ImageStability AISViM3D: Stable Video Material Diffusion for Single Image 3D GenerationarXivpaper
2025-06-18Image-to-3D, PBR Materials, Production AssetsTencentHunyuan3D 2.1: From Images to High-Fidelity 3D Assets with Production-Ready PBR MaterialarXivgithub / model
2025-05-23Gigascale 3D, Sparse Volume, Image ConditioningNanjing UniversityDirect3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionNeurIPS 2025project / github
2025-05-12Textured Assets, Controllable 3D, Open FrameworkStepFunStep1X-3D: Towards High-Fidelity and Controllable Generation of Textured 3D AssetsarXivgithub
2025-05-07Image-to-3D, PBR Textured Mesh, Render-Enhanced Auto-EncoderTsinghuaMeshGen: Generating PBR Textured Mesh with Render-Enhanced Auto-Encoder and Generative Data AugmentationCVPR 2025project
2025-03-28Geometry Detail, Normal Bridging, Image-to-3DCornellHi3DGen: High-fidelity 3D Geometry Generation from Images via Normal BridgingarXivproject
2025-03-03Text/Image-to-3D, Bundle Image, Data-Efficient (147K)HKUST(GZ)Kiss3DGen: Repurposing Image Diffusion Models for 3D Asset GenerationarXivgithub
2025-02-10Shape Synthesis, Rectified Flow, Image-to-3DVAST AI / TripoTripoSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow ModelsarXivgithub / model
2025-01-21Image/Text-to-3D, Mesh, TextureTencentHunyuan3D 2.0: Scaling Diffusion Models for High Resolution Textured 3D Assets GenerationarXivgithub
2025-01-08Point-Aware, Interactive Editing, Single ImageStability AI / UIUCSPAR3D: Stable Point-Aware Reconstruction of 3D Objects from Single ImagesarXivproject / model
2024-11-11Text/Image-to-3D, PBR, Multi-View DiffusionNVIDIAEdify 3D: Scalable High-Quality 3D Asset GenerationarXivproject
2024-09-193DTopia-XL, Primitive Diffusion, Large Scale, HighlightNTU / PKU3DTopia-XL: Scaling High-quality 3D Asset Generation via Primitive DiffusionCVPR 2025 Highlightproject
2024-08-01Fast Mesh, UV Unwrap, Material DisentanglementStability AIStable Fast 3D: Stable Fast 3D Mesh Reconstruction with UV-unwrapping and Illumination DisentanglementarXivproject / github
2024-07-02Text-to-Mesh, PBR, Geometry+TextureMetaMeta 3D AssetGen: Text-to-Mesh Generation with High-Quality Geometry, Texture, and PBR MaterialsNeurIPS 2024project / paper
2024-06-26GaussianDreamer, Text-to-3DGS, 2D+3D DiffusionHUSTGaussianDreamer: Fast Generation from Text to 3D GaussiansCVPR 2024github
2024-05-30High-Quality Mesh, Multi-View Normals, ISOMERTsinghuaUnique3D: High-Quality and Efficient 3D Mesh Generation from a Single ImagearXivproject
2024-05-23Text-to-3D, 3D-DiT, Interactive RefinementHKUSTCraftsMan3D: High-fidelity Mesh Generation with 3D Native Generation and Interactive Geometry RefinerarXivgithub
2024-05-19Multi-View Diffusion, Row-Wise Attention, MeshUniversity of Hong KongEra3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionarXivproject
2024-04-10Single Image, Sparse-View LRM, MeshTencent ARCInstantMesh: Efficient 3D Mesh Generation from a Single Image with Sparse-view Large Reconstruction ModelsarXivgithub
2024-03-18Single Image, Orbital Video, Multi-View PriorStability AISV3D: Novel Multi-view Synthesis and 3D Generation from a Single Image using Latent Video DiffusionarXivproject
2024-03-08Textured Mesh, Convolutional Reconstruction, Single ImageShanghai AI LabCRM: Single Image to 3D Textured Mesh with Convolutional Reconstruction ModelarXivproject
2024-03-043DTopia, Hybrid Diffusion, Text-to-3DNTU / PKU3DTopia: Large Text-to-3D Generation Model with Hybrid Diffusion PriorsCVPR 2024project
2024-02-20MVDiffusion++, Dense High-Res, Sparse-ViewSimon Fraser / MetaMVDiffusion++: A Dense High-Resolution Multi-view Diffusion ModelECCV 2024paper
2024-02-07Multi-View Gaussian, Fast 3D Content, Image/TextPeking UniversityLGM: Large Multi-View Gaussian Model for High-Resolution 3D Content CreationECCV 2024 Oralgithub
2023-12-06XCube, Sparse Voxel, Large-Scale, HighlightNVIDIAXCube: Large-Scale 3D Generative Modeling using Sparse Voxel HierarchiesCVPR 2024 Highlightproject
2023-12-05ReconFusion, Diffusion Prior, 3D ReconstructionColumbia / GoogleReconFusion: 3D Reconstruction with Diffusion PriorsCVPR 2024paper
2023-11-27MeshGPT, Triangle Mesh, Decoder-Only, TransformerTUMMeshGPT: Generating Triangle Meshes with Decoder-Only TransformersCVPR 2024project
2023-11-19LucidDreamer, Interval Score, Text-to-3D, HighlightHKUST(GZ)LucidDreamer: Towards High-Fidelity Text-to-3D via Interval Score MatchingCVPR 2024 Highlightproject
2023-11-15DMV3D, Multi-View Diffusion, 3D LRM, NeRFAdobe / StanfordDMV3D: Denoising Multi-View Diffusion using 3D Large Reconstruction ModelNeurIPS 2023paper
2023-10-26MVDiffusion, Multi-View, Consistent, Single ImageSimon FraserMVDiffusion: Multi-view Consistent Image Generation from a Single ImageCVPR 2024github
2023-10-25DreamCraft3D, Bootstrapped, Hierarchical, 3DDeepSeek / TsinghuaDreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorICLR 2024project
2023-10-23Cross-Domain Diffusion, Multi-View Normals, MeshHKUSTWonder3D: Single Image to 3D using Cross-Domain DiffusionCVPR 2024 Highlightgithub
2023-10-23Single Image, Consistent Multi-View DiffusionUC San DiegoZero123++: a Single Image to Consistent Multi-view Diffusion Base ModelarXivgithub
2023-09-28GS, SDS, Efficient 3D Content CreationPKU / NTUDreamGaussian: Generative Gaussian Splatting for Efficient 3D Content CreationICLR 2024 Oralgithub
2023-09-04Multi-View Consistent Images, Single ImageHKUSyncDreamer: Generating Multiview-consistent Images from a Single-view ImageICLR 2024 Spotlightgithub
2023-08-31Multi-View Diffusion, 3D GenerationUCSD / ByteDanceMVDream: Multi-view Diffusion for 3D GenerationICLR 2024project
2023-08-22IT3D, Explicit View, Text-to-3DNTUIT3D: Improved Text-to-3D with Explicit View SynthesisNeurIPS 2023paper
2023-06-30Magic123, One Image to 3D, 2D+3D PriorsKAUST / SnapMagic123: One Image to High-Quality 3D Object Using Both 2D and 3D Diffusion PriorsCVPR 2024project
2023-06-29Single Image, Optimization-Free, MeshShanghai AI LabOne-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationNeurIPS 2023github
2023-06-21DreamTime, Improved SDS, OptimizationIDEA / HKUDreamTime: An Improved Optimization for Diffusion-Guided 3DICLR 2024paper
2023-06-06ATT3D, Amortized, Text-to-3DNVIDIAATT3D: Amortized Text-to-3D Object SynthesisICCV 2023paper
2023-05-30HiFA, High-Fidelity, Text-to-3D, Advanced GuidanceUIUCHiFA: High-fidelity Text-to-3D with Advanced Diffusion GuidanceICLR 2024paper
2023-05-25Text-to-3D, Variational SDS, High-Fidelity, SpotlightTsinghuaProlificDreamer: High-Fidelity and Diverse Text-to-3D GenerationNeurIPS 2023 Spotlightproject
2023-03-24Make-It-3D, Single Image, Diffusion PriorSJTU / MicrosoftMake-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorICCV 2023project
2023-03-24Text-to-3D, Geometry+Appearance DisentangleSCUTFantasia3D: Disentangling Geometry and Appearance for Text-to-3DICCV 2023project
2023-03-20Zero-1-to-3, Zero-Shot, Single Image to 3DColumbia / TRIZero-1-to-3: Zero-shot One Image to 3D ObjectICCV 2023github
2023-02-21RealFusion, Single Image, 360°, 3DOxford VGGRealFusion: 360° Reconstruction of Any Object from a Single ImageCVPR 2023project
2023-01-263DShape2VecSet, Neural Fields, Diffusion, MeshKAUST / TUM3DShape2VecSet: A 3D Shape Representation for Neural Fields and DiffusionSIGGRAPH 2023 (ACM TOG)github
2022-12-02DiffRF, Rendering-Guided, Radiance Field Diffusion, HighlightTUM / MetaDiffRF: Rendering-Guided 3D Radiance Field DiffusionCVPR 2023 Highlightpaper
2022-12-01Text-to-3D, Score Jacobian Chaining, 2D DiffusionTTI-ChicagoScore Jacobian Chaining: Lifting Pretrained 2D Diffusion for 3DCVPR 2023paper
2022-11-18Text-to-3D, High-Res, Coarse-to-Fine, HighlightNVIDIAMagic3D: High-Resolution Text-to-3D Content CreationCVPR 2023 Highlightpaper
2022-10-12LION, Latent Point Diffusion, ShapeNVIDIALION: Latent Point Diffusion Models for 3D Shape GenerationNeurIPS 2023github
2022-09-29Text-to-3D, Score Distillation, SDS, 2D Diffusion, Outstanding PaperGoogle ResearchDreamFusion: Text-to-3D using 2D DiffusionICLR 2023 Outstanding Paperproject
2022-09-22GET3D, Generative, Textured Shapes, Images OnlyNVIDIAGET3D: A Generative Model of High Quality 3D Textured ShapesNeurIPS 2022github
2022-03-17AutoSDF, Shape Prior, 3D CompletionCMUAutoSDF: Shape Priors for 3D Completion, Reconstruction and GenerationCVPR 2022paper
2020-02-23PolyGen, Autoregressive, Mesh, DeepMindDeepMindPolyGen: An Autoregressive Generative Model of 3D MeshesICML 2020paper
Multi-Image / Multi-View Object Mesh
DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-06-23Multi-View + LiDAR, Vehicle Assets, TRELLISShanghai Jiao Tong UniversityMM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous DrivingarXivgithub
2026-03-12Multi-View, SAM3D, Layout-Aware, Physical PlausibilityPeking UniversityMV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D GenerationarXivgithub
2026-01-16Casual Capture, Posed Multi-View, Metric ShapeMeta AIShapeR: Robust Conditional 3D Shape Generation from Casual CapturesarXivpaper
2025-11-12Multi-Image Fusion, Region Control, TRELLISZhejiang UniversityFuse3D: Generating 3D Assets Controlled by Multi-Image FusionSIGGRAPH Asia 2025project / github
2025-03-18Multi-View Image-to-Shape, Hunyuan3D-DiTTencentHunyuan3D 2.0 MVModelgithub / model
2024-02-06Scalable View Synthesis, Single/Multi-Image 3DKAUSTEscherNet: A Generative Model for Scalable View SynthesisCVPR 2024project
2019-08-05Multi-View Images, Mesh Deformation, Shape RefinementNational Tsing Hua UniversityPixel2Mesh++: Multi-View 3D Mesh Generation via DeformationICCV 2019github

4.1.2 Texture & Material Generation (PBR)

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-25Relightable 3D Assets, PBR Maps, Gaussian RepresentationAppleLuce: Relightable Gaussians for 3D Asset GenerationarXivpaper
2026-08-24Material Decomposition, Physical Properties, Watertight Sub-Meshes, Sim-ReadyUniversity of BristolGen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material DecompositionarXivpaper
2026-07-01Complex Textures, Video Generative Prior, 3D AssetsAuthorsInk3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative ModelsarXivpaper
2024-11-10VLM-Guided, PBR Texture3D AIGCTexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian SplattingCVPR 2025project
2024-01-17TextureDreamer, Geometry-Aware, DiffusionUCSD / MetaTextureDreamer: Image-Guided Texture SynthesisCVPR 2024paper
2023-12-21Texture Generation, Mesh, Multi-View ConsistencyTencentPaint3D: Paint Anything 3D with Lighting-Less Texture Diffusion ModelsCVPR 2024project
2023-11-28SceneTex, Indoor, Texture, Diffusion, HighlightTUM / SnapSceneTex: High-Quality Texture Synthesis for Indoor ScenesCVPR 2024 Highlightpaper
2023-11-21SyncMVD, Multi-View, Text-to-TextureCUHKSyncMVD: Text-Guided Texturing by Synchronized Multi-View DiffusionCVPR 2024paper
2023-08-22PBR Material, SVBRDF, Text-to-MaterialAdobeMatFuse: Controllable Material Generation with Diffusion ModelsSIGGRAPH Asia 2024project
2023-03-20Texture, Material, Text-to-TextureKAISTText2Tex: Text-driven Texture Synthesis via Diffusion ModelsICCV 2023project
2023-02-03TEXTure, Text-Guided, 3D Texture, DiffusionTel Aviv UniversityTEXTure: Text-Guided Texturing of 3D ShapesSIGGRAPH 2023project
2022-07-06nvdiffrec, 3D Mesh, Material, Lighting, OralNVIDIAnvdiffrec: Extracting Triangular 3D Models, Materials, and LightingCVPR 2022 Oralgithub

4.2 Part-Level & Articulated Generation

4.2.1 Kinematic Structure Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-08Single Image, Physical CoT, URDF, Simulation-ReadyAerospace Information Research Institute, CASPhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D AssetsarXivpaper
2026-07-15Articulation + Physics, 40K Assets, Simulation-ReadyZhejiang UniversityUniPhysGen: Unified Physical Grounding for Simulation-Ready 3D AssetsarXivgithub
2026-05-20Rigid/Deformable/Articulated, Physical Attributes, Sim-ReadyNanyang Technological UniversityPhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated ObjectsarXivproject / dataset
2026-05-14Agentic Generation, Articraft-10K, URDF AssetsAuthorsArticraft: An Agentic System for Scalable Articulated 3D Asset GenerationarXivproject / github
2026-05-06Physics-Grounded, Kinematic, Simulation-Ready AssetsHKUPhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual WorldICML 2026project / github
2026-03-14URDF, Autoregressive, Simulation-Ready AssetsAuthorsURDF-Anything+: End-to-End Generation for Simulation-Ready Articulated AssetsarXivpaper
2026-03-01Articulated Assets, 3D LLM, Kinematic StructureTsinghuaArtLLM: Generating Articulated Assets via 3D LLMCVPR 2026paper
2025-12-12Articulation, Kinematic Tree, Feed-Forward, URDF-ReadyUniversity of OxfordParticulate: Feed-Forward 3D Object ArticulationarXivproject
2025-11-26Single Image, Open-Set Articulation, Unified LatentShanghaiTechUniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set ArticulationarXivpaper
2025-11-17Sim-Ready Assets, Physical Properties, Single ImageNTUPhysX-Anything: Simulation-Ready Physical 3D Assets from Single ImageCVPR 2026paper
2025-11-02URDF, 3D MLLM, Articulated ObjectsTsinghuaURDF-Anything: Constructing Articulated Objects with 3D Multimodal Language ModelNeurIPS 2025paper
2025-08-20Articulated Geometry, Motion Modeling, Gaussian RepresentationTsinghua UniversityGaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects3DV 2026paper
2025-07-16Physical Properties, Scale, Material, AffordanceNanyang Technological UniversityPhysX-3D: Physical-Grounded 3D Asset GenerationNeurIPS 2025 Spotlightproject / github
2025-06-10Interactable Digital Twin, Articulated Object, RGB-D VideoShanghai Jiao Tong UniversityiTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos3DV 2026paper
2025-04-17Simulation-Ready, Physical Materials, DynamicsUniversity of Massachusetts AmherstSOPHY: Generating Simulation-Ready Objects with Physical MaterialsWACV 2026project / github
2025-03-11Part-Level Digital Twin, Joint Estimation, Self-Supervised 3DGSUSTCArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian SplattingCVPR 2025paper
2025-02-26Articulated Objects, 3DGS, Joint EstimationTsinghuaArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian SplattingICLR 2025project
2025-02-17Articulation-Ready, Skeleton, Skinning, BenchmarkNanyang Technological UniversityMagicArticulate: Make Your 3D Models Articulation-ReadyCVPR 2025project / github
2024-09-26Open-Vocabulary, URDF, ArticulationStanfordArticulate Anything: Open-vocabulary 3D Articulated Object GenerationICLR 2025project

4.2.2 Part-Aware Assembly & Editing

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-08-20Compositional 3D, Part Semantics, Spatial Control, Reassemblable AssetsRobloxMultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial ControlarXivproject
2026-08-14Up to 300 Parts, Token-Efficient VQ, Autoregressive 3D, Structured AssetsThe University of Hong KongMegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive ModelingarXivproject
2026-08-13Part-Aware Generation, Recursive Decomposition, Editable AssetsShanghai Jiao Tong UniversitySCULPT: Subtractive Composition for 3D Part GenerationarXivproject
2026-07-18Category-Agnostic, Neural Shape Editing, Coupled RepresentationAuthorsCNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape RepresentationarXivpaper
2026-06-23Garment Patterns, Simulation-Ready, EditingUniversity of Hong KongPatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D GarmentsarXivgithub
2026-05-27Part-Controllable, Open-Vocabulary, Game-ReadyRobloxCubePart: An Open-Vocabulary Part-Controllable 3D GeneratorSIGGRAPH 2026project / model
2025-09-10Part Decomposition, Editable, Production-Ready AssetsTencentX-Part: High-Fidelity and Structure-Coherent Shape DecompositionTech Reportproject
2025-08-14Rigging, Animation, Skeleton, SkinningNanyang Technological UniversityPuppeteer: Rig and Animate Your 3D ModelsNeurIPS 2025 Spotlightproject / github
2025-06-05Part-Level Mesh, Compositional DiT, Single ImageUniversity of WaterlooPartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersarXivproject
2024-12-16Articulated Mesh, Part-by-Part, Hierarchical TransformerCornellMeshArt: Generating Articulated Meshes with Structure-Guided TransformersarXivpaper
2023-12-13Shape Program, Structure, Editable AssetsMITShape2Program: Learning to Infer Shape Programs from 3D ShapesarXivproject
2023-06-29Part-Aware, Shape Assembly, 3D GenerationShanghai AI LabMichelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationNeurIPS 2023github

4.3 Scene-Level Generation

4.3.1 Layout & Procedural Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-09-10Single Image, Executable Scene Programs, Recursive Construction, Editable 3DGeorgia Institute of TechnologyRecursive Code World Models: Building Complex Worlds through Recursive Scene ProgramsarXivpaper
2026-09-05Agentic 3D Composition, Functional Objects, Executable Robot ScenesPeking University / GalbotGIF: Agentic Generation of Interactive and Functional Object Compositions for Robot LearningarXivpaper
2026-09-04Image-to-Scene, Agentic Layout Evolution, Simulation-Ready DiversityThe University of Hong KongSceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout EvolutionarXivproject / github
2026-08-31Text-to-Scene, Grow-and-Repair, Functional Groups, SceneReverse-17KSoutheast UniversityScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene GenerationarXivproject
2026-08-27Single Image, Generative 3D Proxy, RGB-D World Expansion, Explorable SceneHong Kong University of Science and TechnologySpatialCrafter: Single Image World Modeling with Generative 3D ProxiesarXivproject
2026-08-25Monocular Image, Interactive Scene Programming, Articulation+Physics, Embodied SimulationShanghai Jiao Tong UniversityNeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied SimulationarXivproject
2026-08-19Usage-Driven Code Scenes, Multi-Part Interaction, Executable SimulationShanghai Jiao Tong UniversityBeyond Placement and Articulation: Usage-Driven Code Scenes for Embodied InteractionarXivpaper
2026-07-29Panoramic Video, 3DGS, Simulation-ready WorldAgiBotGenie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and SimulationarXivgithub
2026-07-15Indoor Layout, Progressive VLM Reasoning, Interactive EditingCity University of Hong KongThinkBLOX: 3D Indoor Scene Generation with Progressive ReasoningarXivpaper
2026-07-08Simulation-Ready Assets, Affordances, Cross-Simulator WorldsHorizon RoboticsEmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AIarXivproject / github
2026-07-07Real-to-Sim, One-Shot Scene Generation, Robot EvaluationShanghai AI LabRoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and EvaluationarXivproject
2026-07-04Egocentric Scene Generation, Geometric 3DGS, ConsistencySouth China Univ. of TechnologyCGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene GenerationarXivpaper
2026-06-23Triangle Splatting, Single-Image Scene, Game-ReadyGoogle ResearchFLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene GenerationarXivproject
2026-06-23Text-to-Scene, Video Priors, 3DGS OrbitUniversity of BernOrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video SynthesisarXivpaper
2026-06-23Compositional 3D, Physical Interaction, Multi-View ConsistencyChina University of Petroleum (East China)Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D GenerationarXivpaper
2026-06-23Satellite-to-City, Textured Mesh, Urban SimulationHKUST(GZ)Sat2City v2: Native 3D City Asset Generation from a Single Satellite ImagearXivpaper
2026-06-08Satellite-to-3D, 3DGS, UAV SimulationAmap-cvlab / AlibabaABot-Earth 0.5: Generative 3D Earth ModelarXivproject
2026-06-04Whole-Home Scenes, Floorplans, InteractiveAuthorsHomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home ScenesarXivpaper
2026-05-28Physical Stability, Single Image, Scene Tree, SimulationCarnegie Mellon UniversityREST3D: Reconstructing Physically Stable 3D Scenes from a Single ImagearXivproject
2026-05-01Segment Map, Text-to-World, Controllable 3D WorldsSeoul National UniversityMap2World: Segment Map Conditioned Text to 3D World GenerationarXivpaper
2026-04-14Explorable 3D Worlds, Long Trajectory, 3DGSNVIDIALyra 2.0: Explorable Generative 3D WorldsarXivproject
2026-04-06Single-Image Scene, In-Place Completion, ARSG-110KNankai University3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageCVPR 2026project / github / dataset
2026-03-31Town-Scale, Single Image, Latent ExtensionSeoul National UniversityExtend3D: Town-Scale 3D GenerationCVPR 2026project
2026-03-31Unbounded World, Flow Matching, LayoutsPrinceton UniversityWorldFlow3D: Flowing Through 3D Distributions for Unbounded World GenerationarXivproject
2026-03-27Autoregressive 3DGS, Token Generation, Completion/OutpaintingTechnical University of MunichGaussianGPT: Towards Autoregressive 3D Gaussian Scene GenerationarXivproject
2026-03-12Multi-Floor, Language-to-3D, Long-Horizon TasksTsinghuaMANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasksCVPR 2026paper
2026-03-06Compositional Scene, Panoramic Image, Feed-ForwardNTUPano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic ImageCVPR 2026paper
2026-02-10Agentic Scene Generation, Sim-Ready, SAGE-10kNVIDIASAGE: Scalable Agentic 3D Scene Generation for Embodied AICVPR 2026project / github
2026-01-09Language-Guided, Infinite Worlds, Articulated FurnitureNational Taiwan UnivSceneFoundry: Generating Interactive Infinite 3D WorldsarXivproject
2025-12-01Tabletop, Instance-Level, Interactive Scene, Text/ImageD-RoboticsTabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single ImagearXivproject / paper
2025-11-18Single-Image Scene, Gaussian World, Scene GenerationTsinghua UniversityGEN3D: Generating Domain-Free 3D Scenes from a Single ImagearXivpaper
2025-09-18Layout-Guided, Indoor Scenes, Decoupled Geometry/AppearanceHong Kong University of Science and TechnologySPATIALGEN: Layout-guided 3D Indoor Scene Generation3DV 2026paper
2025-08-21Single Image, Multi-Asset Scene, Feed-ForwardShanghai Jiao Tong UniversitySceneGen: Single-Image 3D Scene Generation in One Feedforward Pass3DV 2026project / github
2025-08-11Panoramic, Explorable World, Matrix-PanoKunlun WanweiMatrix-3D: Omnidirectional Explorable 3D World GenerationarXivproject / github
2025-07-29Panoramic, Text/Image-to-World, Mesh ExportTencent HunyuanHunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or PixelsarXivgithub
2025-07-09VLA Scene Authoring, Simulation-Ready Worlds, Synthetic DataNVIDIA / Stanford University3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds3DV 2026project / github
2025-06-25Explorable Scene, Novel-View Restoration, ConsistencyBeijing Academy of AIWonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene ExplorationarXivpaper
2025-03-13Scene Layout, Optimization, GenerationTsinghuaHOG-Layout: Layout-Enhanced Scene Generation via Hierarchical OptimizationarXivproject
2025-02-20Octree, 3D Diffusion, Scene GenerationZhejiang UniversityOctree Diffusion: Hierarchical Scene Generation via Octree StructuresarXivproject
2025-01-15Gaussian, GPT, Scene GenerationShanghai AI LabGaussianGPT: Language-Driven Scene Generation with Gaussian RepresentationarXivpaper
2024-12-19Scene Generation, Growing, IncrementalTsinghuaWorldGrow: Incremental 3D Scene GenerationCVPR 2025project
2024-11-05Splatting, Fluents, Scene UnderstandingUniversity of CambridgeFluSplat: Fluent Scene Generation via Gaussian SplattingCVPR 2025project
2024-11-04GenXD, Any 3D and 4D, Scene GenerationNUSGenXD: Generating Any 3D and 4D ScenesICLR 2025paper
2024-06-17Procedural Scenes, Synthetic Data, Embodied AIPrincetonInfinigen Indoors: Photorealistic Indoor Scenes using Procedural GenerationNeurIPS 2024project / github
2024-06-13Interactive 3D Scene, FLAGS, Single ImageStanford UniversityWonderWorld: Interactive 3D Scene Generation from a Single ImageCVPR 2025project
2024-05-02EchoScene, Scene Graph, Diffusion, IndoorTUM / JHUEchoScene: Indoor Scene Generation via Information EchoECCV 2024paper
2024-02-12SceneScape, Text-Driven, Consistent, SceneWeizmannSceneScape: Text-Driven Consistent Scene GenerationNeurIPS 2024paper
2024-01-30BlockFusion, Expandable, Tri-plane, SIGGRAPHTencent / UTokyoBlockFusion: Expandable 3D Scene GenerationSIGGRAPH 2024 (ACM TOG)paper
2023-12-01ControlRoom3D, Semantic Proxy, Room GenerationTUM / MetaControlRoom3D: Room Generation using Semantic Proxy RoomsCVPR 2024paper
2023-10-05Ctrl-Room, Text-to-3D, Layout ConstraintsSimon FraserCtrl-Room: Controllable Text-to-3D Room Meshes GenerationECCV 2024paper
2023-10-04MagicDrive, Street View, 3D Geometry ControlCUHK / HKUSTMagicDrive: Street View Generation with Diverse 3D Geometry ControlICLR 2024project
2023-06-15Procedural World, Synthetic Data, SimulationPrincetonInfinite Photorealistic Worlds using Procedural GenerationCVPR 2023project
2023-03-24DiffuScene, Diffusion, Indoor Scene SynthesisTUMDiffuScene: Denoising Diffusion for Generative Indoor Scene SynthesisCVPR 2024paper
2023-03-21Text-to-3D Room, Indoor Scenes, MeshLMU MunichText2Room: Extracting Textured 3D Meshes from 2D Text-to-Image ModelsICCV 2023project
2023-02-02Unbounded 3D Scene, Generative Model, DrivingNVIDIASceneDreamer: Unbounded 3D Scene Generation from 2D Image CollectionsCVPR 2023project

4.3.2 Semantic Scene Generation & Spatial Intelligence

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-18Multi-View Generation, Scene Assets, Training-FreeThe University of QueenslandScene-SAM3D: Multi-View Scene Asset Generation Without Fine-TuningarXivgithub
2025-01-03Rover, Semantic, 3D SceneCarnegie Mellon UniversitySEM-ROVER: Semantic Scene Exploration with Hierarchical Spatial ReasoningICLR 2025project
2024-09-30Spatial, Generation, LanguageTsinghuaSpatialGen: Language-Driven Spatial Scene GenerationNeurIPS 2024project

4.3.3 4D / Dynamic Scene Generation

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-11Layout-Conditioned, 3DGS, Mixed RealityAuthorsSyncSpace: Layout-Conditioned 3D Gaussian Splatting for Space Reskinning in Mixed RealityarXivpaper
2025-04-17Image-Pair-to-4D, Diffusion, Explicit 3D MotionTechnical University of MunichTwoSquared: 4D Generation from 2D Image Pairs3DV 2026 Oralpaper
2024-12-054Real-Video, Photo-Realistic, Video Diffusion, CVPRSnap / KAUST4Real-Video: Generalizable Photo-Realistic 4D Video DiffusionCVPR 2025paper
2024-07-16Animate3D, Multi-View, Video Diffusion, NeurIPSCASIA / AlibabaAnimate3D: Animating Any 3D Model with Multi-view Video DiffusionNeurIPS 2024paper
2024-05-314Diffusion, Multi-View Video, 4D, NeurIPSCASIA / Shanghai AI Lab4Diffusion: Multi-view Video Diffusion Model for 4D GenerationNeurIPS 2024paper
2024-05-26Diffusion4D, Video Diffusion, 4D, NeurIPSToronto / BJTUDiffusion4D: Fast Spatial-temporal Consistent 4D GenerationNeurIPS 2024paper
2024-05-03DreamScene4D, Multi-Object, Dynamic, NeurIPSCMUDreamScene4D: Dynamic Multi-Object Scene GenerationNeurIPS 2024paper
2024-03-22STAG4D, Spatial-Temporal, 4D Gaussians, ECCVNanjing / CASIASTAG4D: Spatial-Temporal Anchored Generative 4D GaussiansECCV 2024paper
2023-11-294D-fy, Text-to-4D, Score Distillation, CVPRKAUST / Snap4D-fy: Text-to-4D Generation Using Hybrid Score Distillation SamplingCVPR 2024project
2023-11-24Animate124, Image-to-4D, Animation, ICLRNUS / HuaweiAnimate124: Animating One Image to 4D Dynamic SceneICLR 2024paper
2023-11-17Consistent4D, 360° Dynamic, Monocular Video, ICLRCASIA / NanjingConsistent4D: Consistent 360° Dynamic Object GenerationICLR 2024project

4.4 3D Editing

3D editing covers methods that modify existing 3D assets, Gaussian / NeRF fields, meshes, voxel or latent states, and dynamic scenes. Entries are grouped by editable state: object-level, scene-level, and dynamic / 4D.

4.4.1 Object-Level Editing

DateKeywordsInstitute (first)Paper / ResourcePublicationOthers
2026-07-27NeRF Editing, Object Removal, Robot ManipulationAuthorsNEO: NeRF It Once, Edit It Many Times for Continuous Object ManipulationarXivpaper
2026-06-05Mesh Editing, Image-Guided, Local MorphingLeiden University3DMorph: Single-Image-Guided Local 3D Shape Editing and MorphingIJCNN 2026github
2026-05-26PartFlow, Semantic-Part Transformation, Mask-FreeNanyang Technological UniversityFeedforward 3D Editing Learns from Semantic-Part TransformationarXivproject / github / benchmark (Steer3D)
2026-05-08VS3D, Velocity-Space, Mask-FreeTsinghua UniversityVelocity-Space 3D Asset EditingarXivpaper
2026-05-01Latent Editing, Object-Level, Structured 3D LatentsSeoul National UniversityInpaintSLat: Inpainting Structured 3D Latents via Initial Noise OptimizationarXivproject
2026-04-30MeshReGen, VecSet Regeneration, Image-Guided EditingKAISTMeshReGen: A Unified 3D Geometry Regeneration FrameworkarXivproject / benchmark (VoxHammer)
2026-04-26Primitive Proxy, Shape Editing, Fine-Grained ControlTel Aviv UniversityProx-E: Fine-Grained 3D Shape Editing via Primitive-Based AbstractionsSIGGRAPH 2026project / github / benchmark (VoxHammer)
2026-03-303DGS Editing, Object-Level, Single-ViewZhejiang Gongshang UniversitySVGS: Single-View to 3D Object Editing via Gaussian SplattingACM TOMM 2026project
2026-02-25Voxel Editing, Object-Level, Rectified Voxel FlowUSTCEasy3E: Feed-Forward 3D Asset Editing via Rectified Voxel FlowCVPR 2026project
2026-02-05Native 3D Editing, Image-Conditioned, Latent-to-LatentAigency.ai / Tel Aviv UniversityShapeUP: Scalable Image-Conditioned 3D EditingSIGGRAPH 2026project / github
2026-02-04Mesh Editing, Object-Level, Single-Image LRMNational Yang Ming Chiao Tung UniversityVecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single ImagearXivgithub / benchmark (VoxHammer)
2025-12-15Steer3D, Text-Steerable, Edit3D-Bench (Ma)CaltechFeedforward 3D Editing via Text-Steerable Image-to-3DarXivproject / github / benchmark (Steer3D)
2025-11-27Latent Anchor, Object-Level, Mask-FreeZhejiang UniversityAnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned FlowsCVPR 2026 Oralproject / github / benchmark
2025-11-21Native Editing, Object-Level, Full AttentionFudan University / StepFunNative 3D Editing with Full AttentionarXivpaper
2025-10-16FlowEdit, Object-Level, Mask-FreeTsinghua UniversityNANO3D: A Training-Free Approach for Efficient 3D Editing Without MasksICLR 2026project / github / dataset
2025-10-033DEditFormer, Paired Dataset, Mask-FreeEast China Univ of Science & TechnologyTowards Scalable and Consistent 3D EditingarXivproject / github
2025-08-293D-LATTE, Text Instructions, 3D Diffusion LatentUniversity of Tübingen3D-LATTE: Latent Space 3D Editing from Textual InstructionsCVPR 2026 Oralproject / CVF
2025-08-26VoxHammer, Edit3D-Bench (Li), Training-FreeRenmin University / Beihang UniversityVoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space3DV 2026 Oralproject / github / benchmark (VoxHammer)
2025-07-153DGS Editing, Part-Level, Regularized SDSSeoul National UniversityRobust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation SamplingICCV 2025CVF
2025-06-25Image Prompt, Multi-View Propagation, Mask-FreeTel Aviv UniversityEditP23: 3D Editing via Propagation of Image Prompts to Multi-ViewACM TOG 2025project / github
2025-05-11CMD, Local Editing, Progressive GenerationHKUSTCMD: Controllable Multiview Diffusion for 3D Editing and Progressive GenerationSIGGRAPH 2025project
2024-12-11Mesh Editing, Object-Level, Masked LR

Truncated — view the full README on GitHub.

3d-generation
3d-reconstruction
3d-vision
awesome-list
computer-vision
embodied-ai
gaussian-splatting
nerf
robotics
world-models

Contributors

Languages

Python

100.0%