A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.
Python
13
151 commits
updated Sep 23, 2026
A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.
Awesome-Embodied-3DV is a curated list for the research space where 3D vision, 3D/4D reconstruction, 3D generation, simulation-ready assets, and embodied world models meet.
This repository focuses on:
This list is intentionally embodied-3DV-first. It includes 3D generation and 3DGS work only when it helps understand, build, evaluate, or deploy 3D assets and world models for embodied agents. It is not a generic catalog of all 3D generation, editing, rendering, compression, or graphics papers.
Daily candidate feed. The automatically updated arXiv Daily is a high-recall, topic-tagged candidate archive across the six areas above. It is deliberately broader than this curated README: papers are promoted here only after manual primary-source verification.
Start here if you want the shortest path through the field.
Data perception covers the sensor-facing and semantic layers: extracting geometric priors (depth, normals), understanding 3D semantics (detection, segmentation, grounding), active-imaging signals, and dense maps from 2D images, video, or physical sensors.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | Metric Depth, 3DGS Relocalization, Sparse PnP Anchors, Temporal Memory | The Chinese University of Hong Kong, Shenzhen | RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting | arXiv | paper |
| 2026-09-08 | Diffusion Transformer, Single-Step Depth, Sharp Details, Dense Prediction | EPFL / HUAWEI Bayer Lab | Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation | SIGGRAPH Asia 2026 | project / github |
| 2026-09-08 | Any Camera, Metric Point Cloud, Optional Intrinsics/Sparse Depth | Google DeepMind | OmniPoint: Universal Monocular Metric Pointcloud from Any Camera | ECCV 2026 | project |
| 2026-08-30 | Transparent/Reflective Scenes, Bias-Aware Training, 30M Parameters | The University of Hong Kong | OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes | arXiv | project |
| 2026-08-17 | Pixel-Space Prediction, Fine Structures, Sharp Boundaries, Efficient Depth | The Chinese University of Hong Kong, Shenzhen | PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation | arXiv | project |
| 2026-08-03 | Geometry-Invariant Adaptation, Non-Lambertian Surfaces, Mirror/Glass Depth | Changchun Institute of Optics, Fine Mechanics and Physics, CAS | GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation | arXiv | paper |
| 2026-07-23 | UAV Depth, Arbitrary Camera Pose, Metric Geometry | Aerospace Information Research Institute, CAS | DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV | arXiv | github |
| 2026-07-20 | Fine-Detail Geometry, Sparse Volumetric Refinement, Metric Scale | Tsinghua University | MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement | arXiv | project |
| 2026-07-19 | Lightweight Foundation Depth, Camera-Conditioned Metric Depth, Edge Deployment | University of Trento | DepthART: Scaling Foundation Monocular Depth to Tiny Models | ACM Multimedia 2026 | project / github |
| 2026-07-19 | Metric Depth, Odometry Anchor, Recurrent SLAM | UC Berkeley | DROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth Estimation | arXiv | paper |
| 2026-07-17 | Stereo Distillation, Epipolar Cues, Metric Depth | Michigan State University | Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth | arXiv | paper |
| 2026-07-14 | Auto-Regressive Depth, Coarse-to-Fine, Semantic Guidance | Sun Yat-sen University | ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning | arXiv | paper |
| 2026-07-13 | Metric Point Map, Pixel-Wise Calibration, Camera Diversity | The University of Hong Kong | FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry | ECCV 2026 | project / github / model |
| 2026-07-09 | Lightweight Zero-Shot, 6.1M Parameters, On-Device Depth | University of Bologna | ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device | ECCV 2026 | project / github |
| 2026-05-27 | Multi-Layer Depth, Transparent Surfaces, Point Process | Princeton University | SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping | CVPR 2026 | github |
| 2026-05-15 | VLM, Dense Metric Depth, Spatial Reasoning | Zhejiang Univ | Unlocking Dense Metric Depth Estimation in VLMs | arXiv | project / github |
| 2026-05-12 | Sparse 3D Anchors, Relative-to-Metric, Graph Optimization | Tongji University | The Midas Touch for Metric Depth | CVPR 2026 Highlight | project |
| 2026-03-28 | Universal Camera, Metric Depth, Zero-Shot | Michigan State University | UniDAC: Universal Metric Depth Estimation for Any Camera | CVPR 2026 | paper |
| 2026-03-20 | Transparent Objects, Generative Opacification, Monocular Depth, SeeClear-396k | University of California, Los Angeles | SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification | ECCV 2026 | project / dataset |
| 2026-03-17 | Diffusion Prior, Real-World Data, Monocular Depth | Nanjing University of Science and Technology | Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-03-04 | Fine-Grained Geometry, Dual-Stream, Efficient | UMass Amherst | DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation | CVPR 2026 | paper |
| 2026-01-29 | Sparse Metric Prompt, 20M Image-Depth Pairs, Metric Foundation Model | Li Auto Inc | MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources | ECCV 2026 | project / github |
| 2026-01-06 | Arbitrary-Resolution, Neural Implicit, Fine Details | Zhejiang Univ | InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields | CVPR 2026 | project / github |
| 2025-12-13 | Defocus Cue, Bokeh Stack, Metric Depth | Nanyang Tech Univ | Boosting Monocular Metric Depth Estimation via Bokeh Rendering | ICML 2026 | project / github |
| 2025-11-30 | Deterministic Diffusion, Dense Geometry, Fine Details | HKUST(GZ) | Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model | arXiv | project / github |
| 2025-11-13 | Any-View Depth, Metric Geometry, Multi-View | ByteDance | Depth Anything 3: Recovering the Visual Space from Any Views | ICLR 2026 | project / github |
| 2025-10-27 | Unified Generation+Depth, Diffusion Prior, Zero-Shot | HUST | More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models | NeurIPS 2025 | github |
| 2025-10-08 | Pixel-Space Diffusion, Flying-Pixel-Free, Point Clouds | HUST | Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers | NeurIPS 2025 | project / github |
| 2025-09-29 | VLM, Metric Depth, Sparse Supervision | Meta AI | DepthLM: Metric Depth From Vision Language Models | ICLR 2026 Oral | github |
| 2025-07-03 | Monocular Geometry, Metric Scale, Sharp Details | USTC / Microsoft | MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details | NeurIPS 2025 | project / github |
| 2025-04-16 | Sliding Anchor, Unknown Intrinsics, Metric Scale | Shanghai Univ | Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image | arXiv | project / github |
| 2025-02-27 | Metric 3D Points, Self-Prompt Camera, Uncertainty | ETH Zurich | UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler | TPAMI 2026 | github |
| 2025-02-26 | Cross-Context Distillation, Multi-Teacher, Fine-Detail Depth | Zhejiang Univ of Technology | Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator | arXiv | project / github |
| 2024-11-27 | Diffusion Distillation, Metric + Sharp, Boundary Detail | Qualcomm AI Research | SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation | CVPR 2025 | project / github |
| 2024-10-02 | Metric Depth, Zero-Shot, Image Priors | Apple | Depth Pro: Sharp Monocular Metric Depth in Less Than a Second | ICLR 2025 | github |
| 2024-09-26 | Diffusion, Single-Step, Dense Geometry | HKUST(GZ) | Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction | ICLR 2025 | project / github |
| 2024-06-13 | Monocular Depth, Foundation Model, Metric DAv2 | TikTok / HKU | Depth Anything V2 | NeurIPS 2024 | project / github / metric models |
| 2024-03-27 | Metric Depth, Universal, Zero-Shot | ETH Zurich | UniDepth: Universal Monocular Metric Depth Estimation | CVPR 2024 | github |
| 2024-01-19 | Monocular Depth, Relative Depth, Foundation Model | TikTok / HKU | Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data | CVPR 2024 | project |
| 2023-12-04 | Monocular Depth, Zero-Shot, Affine-Invariant | Intel Labs | Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation | CVPR 2024 | project |
| 2023-07-20 | Metric 3D, Zero-Shot, Canonical Space | Alibaba | Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image | ICCV 2023 | github |
| 2021-03-24 | ViT, Dense Prediction, Foundation, Depth | Intel Labs | DPT: Vision Transformers for Dense Prediction | ICCV 2021 | github |
| 2019-07-02 | Robust Depth, Zero-Shot, Cross-Dataset, MiDaS | Intel Labs | MiDaS: Towards Robust Monocular Depth Estimation | TPAMI 2022 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | Unified Video Model, Depth+Normals, Temporal Consistency | Adobe Research | Unified Video Dense Prediction from Disjoint Data | arXiv | project |
| 2026-07-02 | Video Diffusion, In-Context Conditioning, Zero-Shot | HKUST | ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning | ECCV 2026 | project |
| 2026-05-28 | Streaming Geometry, Dynamic Chunking, Depth+Normals | Zhejiang University | Towards Consistent Video Geometry Estimation | arXiv | project |
| 2026-05-11 | Camera Motion, 3D Consistency, Geometry Embedding | HUST | GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth | arXiv | github |
| 2026-04-08 | Post-Processing, Scalable, Single-Image Backbone | Seoul National University | VDPP: Video Depth Post-Processing for Speed and Scalability | CVPR 2026 ECV Workshop | paper |
| 2026-04-02 | Pose Refinement, Temporal Consistency, Monocular Video | Ajou University | PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency | CVPR 2026 | paper |
| 2026-03-12 | Deterministic Diffusion, Generative Prior, Long Video | HKUST(GZ) | DVD: Deterministic Video Depth Estimation with Generative Priors | arXiv | project |
| 2026-01-06 | Temporal Stability, Monocular Video, Long Sequence | ETH Zurich | StableDPT: Temporal Stable Monocular Video Depth Estimation | arXiv | paper |
| 2025-12-20 | Endoscopic Geometry, Streaming Mamba, Metric Depth | Vanderbilt University | EndoStreamDepth: Temporally Consistent Monocular Depth Estimation for Endoscopic Video Streams | arXiv | github / paper |
| 2025-12-11 | Sparse Keyframes, Propagation, Long-Video Consistency | ETH Zurich | Video Depth Propagation | 3DV 2026 | paper |
| 2025-10-10 | Online Inference, Low Memory, Temporal Consistency | Heidelberg University | Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption | arXiv | paper |
| 2025-07-02 | Diffusion Guidance, Scale Synchronization, Geometry Consistency | Tsinghua University | DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation | ICCV 2025 | project |
| 2025-04-09 | 2K Streaming, Mamba, 24 FPS | Netflix Eyeline Studios | FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution | ICCV 2025 Highlight | project / github |
| 2025-01-21 | Super-Long Video, Temporal Gradient, 30 FPS | ByteDance | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos | CVPR 2025 | project / github |
| 2024-11-28 | Long Video, Diffusion, Multi-Resolution Alignment | ETH Zurich | Video Depth without Video Models | CVPR 2025 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-06 | Surface Normals, Transparent Objects, Rectified Flow, Edge Refinement | Zhejiang University | TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation | arXiv | project |
| 2026-07-28 | Stereo Depth, Walsh-Hadamard Mixing, Efficient Inference | Authors | WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing | arXiv | paper |
| 2026-07-22 | Stereo Diffusion Transformer, Flow Matching, Progressive Refinement | Beihang University | STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching | arXiv | paper |
| 2026-07-15 | Vision Features, SE(3) Latent Geometry, Visual Navigation | Google DeepMind | SeeSE3: Emergence of 3D Space in Vision Features | arXiv | paper |
| 2026-07-14 | Heterogeneous Cameras, Metric Depth, Real-Time | D-Robotics | X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras | arXiv | project / github |
| 2026-07-10 | Video Generative Pretraining, Depth/Normals/Pose, Grounded 4D | Google DeepMind | Video Generation Models are General-Purpose Vision Learners | ECCV 2026 | project |
| 2026-07-06 | Boundary-Centric Pretraining, Dense Spatial Perception | Robbyant | LingBot-Vision: Vision Pretraining for Dense Spatial Perception | arXiv | project / github |
| 2026-03-02 | Zero-Shot Stereo, Structure Prompt, Motion Prompt | HUST | PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts | CVPR 2026 | paper |
| 2025-12-11 | Real-Time Stereo, Zero-Shot, Foundation Model | NVIDIA | Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching | CVPR 2026 | paper |
| 2025-07-22 | Foundation Model, Depth/Normal/Pointmap, Multi-View | SJTU | Dens3R: A Foundation Model for 3D Geometry Prediction | ICLR 2026 | paper |
| 2025-04-15 | Surface Normals, Foundation Model, Video, Temporal | Authors | NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors | ICCV 2025 | project / paper |
| 2025-03-21 | Scalable Depth, Autoregressive, 2B Parameters | Baidu | DAR: Scalable Autoregressive Monocular Depth Estimation | CVPR 2025 | project |
| 2025-03-20 | Universal Camera, Spherical 3D, Any Camera | ETH Zurich | UniK3D: Universal Camera Monocular 3D Estimation | CVPR 2025 | github / paper |
| 2025-03-11 | LiDAR Surface Normal, Dataset, Point Cloud | TU Graz | LiSu: A Dataset and Method for LiDAR Surface Normal Estimation | CVPR 2025 | github |
| 2025-01-17 | Stereo, Foundation Model, RAFT-Style, Zero-Shot | NVIDIA | FoundationStereo: Zero-Shot Stereo Matching | CVPR 2025 Oral / Best Paper Nomination | github |
| 2025-01-17 | Stereo, Robust, Zero-Shot, Non-Lambertian | Univ of Bologna | Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching | CVPR 2025 | project |
| 2024-12-11 | Panoramic / Fisheye Depth, Zero-Shot Metric | Intel | Depth Any Camera: Zero-Shot Metric Depth from Any Camera | CVPR 2025 | project |
| 2024-10-24 | Monocular Geometry, Pointmap, Affine-Invariant | USTC / Microsoft | MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images | CVPR 2025 Oral | github |
| 2024-06-24 | Normal Estimation, 3D Priors, Surface Geometry | Nvidia | StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal | SIGGRAPH Asia 2024 | project |
| 2024-03-27 | Diffusion, Effective Conditioning, ViT Priors | IIT Delhi | ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth | CVPR 2024 | github |
| 2024-03-22 | Metric 3D, Multi-Task, Geometry Foundation | Shanghai AI Lab | Metric3D v2: A Versatile Monocular Geometric Foundation Model | TPAMI 2025 | project |
| 2024-03-18 | Geometry, Normals, Depth, Multi-Task | Apple | GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image | ECCV 2024 | project |
| 2024-03-01 | Inductive Biases, Surface Normal, Oral | Imperial College London | DSINE: Rethinking Inductive Biases for Surface Normal Estimation | CVPR 2024 Oral | github |
| 2023-12-04 | High-Res, Patch-Wise, Model-Agnostic | KAUST | PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth | CVPR 2024 | project |
| 2023-09-25 | Iterative Bins, Elastic, GRU, Classification-Regression | Beihang Univ | IEBins: Iterative Elastic Bins for Monocular Depth Estimation | NeurIPS 2023 | github |
| 2023-04-13 | Internal Discretization, Continuous-Discrete, Depth | ETH Zurich | iDisc: Internal Discretization for Monocular Depth Estimation | CVPR 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Annotation-Free, Open-Vocabulary 3D, Language-Space Lifting | National Technical University of Athens | GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting | arXiv | paper |
| 2026-07-21 | Referring Segmentation, 3DGS, Generalized Grounding | Peking University | ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting | arXiv | project |
| 2026-06-23 | Open-Vocabulary BEV, 3DGS, Geometric Constraints | KAIST AI | Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints | ECCV 2026 | paper |
| 2026-06-04 | Open-Vocabulary, Functionality Segmentation, Robotics | Authors | T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation | arXiv | paper |
| 2026-05-07 | Open-Vocabulary, Gaussian Feature Field, Codebook | TU Munich / Google | OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention | arXiv | paper |
| 2025-11-20 | Open-Vocabulary, SAM3, Promptable, DETR | Meta | SAM 3: Segment Anything with Concepts | ICLR 2026 | github / paper |
| 2025-04-03 | Open-Vocabulary, Dual-Level Contrastive, Instance-Aware | BIGAI / Tsinghua | MPEC: Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding | CVPR 2025 | project |
| 2025-03-22 | Training-Free, MLLM Caption, Voxel Grouping | NVIDIA Research Taiwan | OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding | CVPR 2026 | project |
| 2025-03-19 | SAM-2, 3D Tracking, Dynamic Programming | VinAI | Any3DIS: Open-Vocabulary 3D Instance Segmentation with SAM-2 | CVPR 2025 | paper |
| 2025-03-13 | Open-Vocabulary, LLM Canonical, Part Segmentation | Shandong Univ / Tencent | CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation via LLM-Guided Canonical Spatial Modeling | CVPR 2026 Oral | github |
| 2025-03-12 | Functional 3D, CoT, VLM, Training-Free, Highlight | Authors | Fun3DU: Functional 3D Scene Understanding via Chain-of-Thought | CVPR 2025 Highlight | paper |
| 2025-01-02 | Panoptic, Open-Vocabulary, 3D Gaussian Splatting | NUS | PanoGS: Gaussian-based Panoptic Open-Vocabulary 3D Scene Understanding | CVPR 2025 | project |
| 2024-12-13 | Open-Vocabulary 3D, Structured Super-Gaussians, Segmentation | SuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians | 3DV 2026 | paper | |
| 2024-12-12 | Open-Vocabulary, Foundation Dataset, Mask-Text Pairs | NVIDIA | Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation | CVPR 2025 | github |
| 2024-07-02 | Open-Vocabulary, Mask-Snap-Lookup, Indoor/Outdoor | HKUST | OpenIns3D: Open-Vocabulary 3D Segmentation with Mask-Snap-Lookup | ECCV 2024 | github |
| 2024-04-01 | Region-Level, Point-Language Contrastive, Multi-VLM | HKU | RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding | CVPR 2024 | project / github |
| 2024-03-19 | Open-Vocabulary, MLLM, Point-Entity-Text, nuScenes | MPI | OV3D: Open-Vocabulary 3D Understanding with Multi-Modal Alignment | CVPR 2024 | paper |
| 2024-03-19 | 2D-Guided, 3D Proposals, SAM, Instance | VinAI / IBM | Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D-Guided Mask Generation | CVPR 2024 | github |
| 2024-03-15 | Training-Free, View-Consensus, Instance | PKU | MaskClustering: View-Consensus for 3D Instance Segmentation | CVPR 2024 | paper |
| 2024-03-11 | Zero-Shot, Superpoint, SAM, Scene Graph | PKU | SAI3D: Zero-Shot 3D Instance Segmentation by Scene-Aware Incremental Merging | CVPR 2024 | github |
| 2024-01-31 | Segment Anything, 3D Gaussians, Interactive | Zhejiang University | SAGD: Boundary-Enhanced Segment Anything in 3D Gaussians | arXiv | github |
| 2024-01-04 | 2D+3D Unified, Single Model, Highlight | Authors | ODIN: A Single Model for 2D and 3D Segmentation | CVPR 2024 Highlight | github |
| 2023-12-26 | Open-Vocabulary 3DGS, Segmentation, Language | ETH Zurich | LangSplat: 3D Language Gaussian Splatting | CVPR 2024 Highlight | project / github |
| 2023-11-17 | Unified, Semantic+Instance+Panoptic, Single Transformer | Samsung | OneFormer3D: One Transformer for Unified 3D Segmentation | CVPR 2024 | github |
| 2023-06-23 | Open-Vocabulary 3D Instance, Mask Proposals, CLIP | ETH Zurich | OpenMask3D: Open-Vocabulary 3D Instance Segmentation | NeurIPS 2023 | project |
| 2023-06-06 | Segment Anything 3D, Point Cloud, Interactive | VAST AI | SAM3D: Segment Anything in 3D Scenes | CVPR 2024 | github |
| 2023-03-16 | Open-Vocabulary, 3D Scene, Language Field | UC Berkeley | LERF: Language Embedded Radiance Fields | ICCV 2023 | project / github |
| 2022-11-28 | Open-Vocabulary 3D, Point Cloud, Segmentation | ETH Zurich | OpenScene: 3D Scene Understanding with Open Vocabularies | CVPR 2023 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-08 | Feed-Forward, Open-Vocabulary, Panoptic | Authors | EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation | ICML 2026 | github |
| 2025-07 | Feed-Forward, Panoramic, DUSt3R-Based | Authors | PanSt3R: Single-Feed 3D Geometry and Panoptic Segmentation | ICCV 2025 | paper |
| 2025-05-14 | MoE, Multi-Dataset, PTv3, CLIP Alignment | UVA | Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation | ICLR 2026 | project |
| 2025-03-01 | Bayesian 3DGS, Training-Free, EIG | Sony | B3-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG | CVPR 2026 | project |
| 2025-01-14 | End-to-End, 2D-to-3D Lifting, 3DGS | CUHK | Unified-Lift: Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting | CVPR 2025 | github |
| 2025-01-10 | Unsupervised Panoptic, Scene-Centric | TU Darmstadt | CUPS: Scene-Centric Unsupervised Panoptic Segmentation | CVPR 2025 Highlight | github |
| 2025-01-06 | Zero-Shot Instance, SAM Prompts in 3D | CUHK-SZ / MSRA | SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation | 3DV 2025 | github |
| 2024-07-03 | 3D Panoptic, LiDAR, Multi-Scene | Tsinghua | UniSeg3D: Unified 3D Panoptic Segmentation | CVPR 2023 | github |
| 2024-03-25 | Unsupervised 3D Instance, Indoor | Authors | UnScene3D: Unsupervised 3D Instance Segmentation | CVPR 2024 | paper |
| 2024-03-24 | Unified 6 Tasks, Single Transformer | HUST | UniSeg3D: Unified 3D Segmentation with Transformer | NeurIPS 2024 | paper |
| 2023-12-15 | Point Cloud, Foundation Model, 3D Understanding | Shanghai AI Lab | Point Transformer V3: Simpler, Faster, Stronger | CVPR 2024 | github |
| 2022-10-06 | 3D Instance Segmentation, Transformer, Point Cloud | ETH Zurich | Mask3D: Mask Transformer for 3D Semantic Instance Segmentation | ICRA 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | 3D-Aware VLM, Implicit+Explicit Geometry, RGB Video | Nanyang Technological University | 3D-Aware VLMs with Implicit and Explicit Geometries | ECCV 2026 | github |
| 2026-07-14 | 3D Language Fields, Ambiguity Awareness, Object Retrieval | Authors | SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval | arXiv | paper |
| 2026-06-23 | Agentic, Cognitive Map, Zero-Shot 3D | Sichuan University | Agentic Collaborative Cognition for Zero-Shot 3D Understanding | ECCV 2026 | project |
| 2026-06-22 | Map-Grounded, MV3D-VQA, Dense Reward | KAIST | Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views | ECCV 2026 | paper |
| 2026-06-17 | Panoramic Reprojection, 3D VLM, Spatial Reasoning | Technical University of Munich | OneCanvas: 3D Scene Understanding via Panoramic Reprojection | arXiv | project |
| 2026-06-04 | Part-Aware 3D-MLLM, Scene Understanding, Grounding | Xiamen University | PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding | arXiv | project |
| 2025-10-19 | Grounded CoT, 3D Reasoning, SceneCOT-185K Dataset | BIGAI / PKU / Tsinghua | SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes | ICLR 2026 | paper |
| 2025-07 | Scene Graph, LLM, Relation Encoding | AIRI | 3DGraphLLM: 3D Scene Graph Learning with LLMs | ICCV 2025 | paper |
| 2025-03-19 | Gesture+Language, Embodied Reference, +30% | Authors | Ges3ViG: 3D Embodied Reference Understanding with Pointing Gestures | CVPR 2025 | paper |
| 2025-03-19 | Geometry VLM, Unified 3D Recon + Spatial Reasoning | Shanghai AI Lab / UCLA / SJTU | G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning | CVPR 2026 | github |
| 2025-03-12 | Dense Grounding, 6.2M Pairs, Hallucination Benchmark | UMich | 3D-GRAND: A Million-Scale Dataset for 3D Grounding | CVPR 2025 | project |
| 2025-03-11 | 3D-Informed, Spatial Reasoning, Highlight | JHU | SpatialLLM: Spatial Reasoning with 3D-Informed LLMs | CVPR 2025 Highlight | paper |
| 2025-03-10 | LLM Attention, Scene Magnifier, Cross-Room | SCUT | LSceneLLM: LLM-Attention Adaptive 3D Scene Understanding | CVPR 2025 | paper |
| 2025-01-12 | LVLM-Guided, Hierarchical Feature, 3DGS | Fudan / NTU | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding | CVPR 2025 | project |
| 2025-01-04 | Generalist 3D LMM, Omni Superpoint Transformer | Adelaide / Microsoft | 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer | CVPR 2025 | github |
| 2024-07-18 | Object Identifiers, 3D VL, Unified Tasks | ZJU | Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers | NeurIPS 2024 | github |
| 2024-07-15 | Million-Scale 3D VL, Multi-Level Contrastive | BIGAI | SceneVerse: Scaling 3D Vision-Language Learning for Grounding | ECCV 2024 | project |
| 2024-06-06 | Zero-Shot 3D Grounding, 2D VLM Transfer | NUS | SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding | CVPR 2025 | project |
| 2024-04-03 | Open-Vocabulary 3D Scene Graph, Open Relations | Bosch | Open3DSG: Open-Vocabulary 3D Scene Graph Generation | CVPR 2024 | paper |
| 2024-03-08 | Interactive 3D, Language Assistant, Point Cloud | CUHK | LL3DA: Visual Interactive Instruction Tuning for 3D Language Assistant | CVPR 2024 | github |
| 2024-03 | CLIP Cross-Modal, Contrastive, Scene Graph | Authors | CCL-3DSGG: CLIP-Driven Contrastive Learning for 3D Scene Graph Generation | CVPR 2024 | paper |
| 2023-11-21 | Generalist Embodied 3D Agent, VLA, ICML | BIGAI | LEO: An Embodied Generalist Agent in 3D World | ICML 2024 | project |
| 2023-08-08 | 3D Visual Grounding, Referring Expression, Point Cloud | Peking University | 3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment | ICCV 2023 | github |
| 2023-07-24 | 3D VQA, 3D Captioning, Scene Understanding | Shanghai AI Lab | 3D-LLM: Injecting the 3D World into Large Language Models | NeurIPS 2023 | project / github |
| 2023-03 | 3D Language Pre-training, Captioning, QA | Authors | 3D-VLP: 3D Vision-Language Pre-training with Contextual Scene | CVPR 2023 | paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-29 | iToF, Sensor-Intrinsic Uncertainty, Heteroscedastic Restoration | Tsinghua University | Reliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space Restoration | arXiv | paper |
| 2026-07-27 | Neural Structured Light, Metric Depth, Online SLAM | Peking University | NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction | ACM MM 2026 | paper |
| 2026-07-20 | Projector-Camera, Feed-Forward 3DGS, Active Illumination | Ningbo University | FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System | arXiv | github |
| 2026-06 | Active Stereo, 2DGS Supervision, RealSense Dataset | Hangzhou Dianzi University | GS-ASM: 2DGS-Supervised Active Stereo Matching | CVPR 2026 | paper |
| 2026-05-07 | Adaptive 4D Illumination, Shape+Reflectance, Differentiable Capture | Zhejiang University | Differentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and Reflectance | CVPR 2026 | paper |
| 2026-03 | Multi-Projector Structured Light, One-Shot Scan, Neural SDF | Kyushu University | Multi-view Stereo with Multiple Projectors for Oneshot Entire Shape Scan based on Neural SDF and DSSS Demultiplexing | WACV 2026 | paper |
| 2026-02-05 | Latent Diffusion, Fringe Projection, Reflective Objects | Yonsei University | LD-SLRO: Latent Diffusion Structured Light for 3-D Reconstruction of Highly Reflective Objects | arXiv | paper |
| 2025-12-16 | Single-Shot, Neural Feature Decoding, Robust Correspondence | Peking University | Robust Single-shot Structured Light 3D Imaging via Neural Feature Decoding | SIGGRAPH Asia 2025 | project |
| 2025-12 | Event Camera, HDR Measurement, Confidence Stereo | USTC | Event-based HDR Structured Light | NeurIPS 2025 | github |
| 2025-03-08 | Active Stereo, Phase Speckle, Cross-Scene Generalization | Southwest Jiaotong University | RGB-Phase Speckle: Cross-Scene Stereo 3D Reconstruction via Wrapped Pre-Normalization | arXiv | paper |
| 2025-02 | Unsupervised Structured Light, Neural SDF, Shadow-Aware | Kyushu University | Neural SDF for Shadow-Aware Unsupervised Structured Light | WACV 2025 | paper |
| 2025-01-13 | Matching-Free, Volume Rendering, Monocular Structured Light | USTC | Matching-Free Depth Recovery from Structured Light | arXiv | paper |
| 2024-10-20 | Neural SDF, One-Shot Scan, Low-Light/Underwater | Kyushu University | ActiveNeuS: Neural Signed Distance Fields for Active Stereo | 3DV 2024 | paper |
| 2024-06-06 | Virtual Pattern Projection, Depth Fusion, In-the-Wild Stereo | University of Bologna | Active Stereo in the Wild through Virtual Pattern Projection | arXiv | github |
| 2024-06 | Neural Inverse, Dense Depth, 3-4 Patterns | University of Toronto | TurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured Light | CVPR 2024 | project |
| 2023-06-17 | Structured Light, Phase Unwrapping, Learning | Nanjing University | Deep Learning-Based Structured Light 3D Imaging: A Survey | arXiv | survey |
| 2018-11-27 | Event Camera, Active Stereo, Depth | Tsinghua University | Event-Based Structured Light for Depth Reconstruction | IJCAS 2024 | paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-16 | Wide-FOV, Egocentric, 4D Hand-Object | Rice University | EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning | arXiv | project / dataset |
| 2026-03-19 | Panoramic Depth, VGGT, Geometry Consistency | Singapore University of Technology and Design | VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation | CVPR 2026 | paper |
| 2026-03-18 | Panoramic Reconstruction, Permutation-Equivariant, 360° | Cornell | PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery | CVPR 2026 | paper |
| 2026-01-25 | RGB-D, Depth Completion, Reflective/Transparent | Robbyant | LingBot-Depth: Masked Depth Modeling for Spatial Perception | arXiv | github / paper |
| 2025-12-18 | Panoramic Foundation Model, Multi-Camera, Metric Depth | Insta360 Research | Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation | CVPR 2026 | paper |
| 2025-01 | Fisheye, Real-Time, Cassini, Multi-View | Sun Yat-Sen Univ | OmniStereo: Real-time Omnidirectional Depth with Multiview Fisheye Cameras | CVPR 2025 | github |
| 2024-06-19 | Panoramic Depth, Semi-Supervised, Mobius | HKUST(GZ) | PanDA: Panoramic Depth Anything with Mobius Spatial Augmentation | CVPR 2025 | project |
| 2024-03-25 | 360 Depth, Bi-Projection, ERP+ICOSAP | HKUST(GZ) | Elite360D: Efficient 360 Depth Estimation via Bi-Projection Fusion | CVPR 2024 | github |
| 2021-09-06 | 360 Depth, Indoor, Panoramic Images | CERTH | Pano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth Estimation | CVPRW 2021 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-26 | Active Event Stereo, High-Speed Depth, 150 FPS | Authors | Towards Ultrafast Depth Sensing Via Active Event-based Stereo Vision | TPAMI 2026 | paper |
| 2026-07-17 | Event Camera, Feed-Forward 3D, Temporal Aggregation | Zhejiang University | Event3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature Aggregation | arXiv | paper |
| 2026-06 | RGB+ToF Histogram, High-Resolution Metric Depth, Lightweight | Tongji University | LiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth Estimation | CVPR 2026 Highlight | paper |
| 2026-06 | Sparse dToF, Zero-Shot Completion, Sensor Generalization | KAIST | Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors | CVPR 2026 | paper |
| 2026-06 | Image-Event Fusion, Monocular Depth, Linear Complexity | Peking University | AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-06 | Event-Image Depth, Hypothesis Volume, Iterative Refinement | Southeast University | Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-04-16 | Event-Frame Stereo, Cross-Modal Prompting, High Dynamic Range | Southeast University | Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo | CVPR 2026 | paper |
| 2026-04-02 | Event Stereo, Data Factory, Cross-Modal Distillation | University of Bologna | EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors | CVPR 2026 | project |
| 2026-02-03 | Event Camera, Neural SDF, Single-Camera Mesh | Saarland University | EventNeuS: 3D Mesh Reconstruction from a Single Event Camera | 3DV 2026 | project / github / dataset |
| 2025-12-20 | Event Camera, Structured Light, Real-Time RGB-D | Polytechnique Montreal | E-RGB-D: Real-Time Event-Based Perception with Structured Light | arXiv | github |
| 2025-09-18 | Event Camera, Depth Estimation, Any-to-Any | Shanghai AI Lab | Depth AnyEvent: Event Camera Based Monocular Depth Estimation via Dense Correspondence Distillation | arXiv | paper |
| 2025-09-08 | Event Camera, Multispectral, Structured Light | ETH Zurich | Event Spectroscopy: Event-based Multispectral and Depth Sensing using Structured Light | arXiv | paper |
| 2025-05-28 | Burst-Encodable ToF, Long-Range Depth, Hardware-Aware Coding | Nanjing University | Learnable Burst-Encodable Time-of-Flight Imaging for High-Fidelity Long-Distance Depth Sensing | NeurIPS 2025 | github |
| 2025-05 | Event, Distillation, Confidence-Guided, Pseudo-Labels | NUS | Distil-E2D: Distilling Image-to-Depth Priors for Event-Based Depth | NeurIPS 2025 | paper |
| 2025-04-23 | ToF, Sparse Depth, 3DGS, SLAM | Authors | ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth | ICCV 2025 | paper / paper |
| 2025-04-22 | Event Camera, Ray Density, 3D Conv, Spotlight | TU Berlin | DERD-Net: Learning Depth from Event-based Ray Densities | NeurIPS 2025 Spotlight | github |
| 2025-03-03 | dToF, Video Depth Completion, Frequency Selective | Authors | SVDC: Consistent Direct Time-of-Flight Video Depth Completion | CVPR 2025 | paper / paper |
| 2024-10-10 | Event Camera, Pose-Free, Gaussian Splatting, Highlight | Zhejiang Univ | IncEventGS: Pose-Free Gaussian Splatting from a Single Event Camera | CVPR 2025 Highlight | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | mmWave Radar, Complex-Valued Point Splatting, Material Model, Novel Views | Cornell Tech | 3D Point Splatting for mmWave Radar Novel View Synthesis | arXiv | paper |
| 2026-09-09 | mmWave Radar, Metric Depth, Smoke/Fog/Darkness, 95K Frames | Rice University | GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation | MobiCom 2026 | project+data |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-09 | RGB-D 3DGS, Online Reconstruction, Reactive Control | Mitsubishi Electric Research Laboratories | SplatCtrl: Perception-Action Coupling via Gaussian Scene Representations and Reactive Robot Control | ICRA 2026 | paper |
| 2026-07-08 | Geometry-Only 3DGS, Dense Monocular SLAM | Beihang University | GeoGS-SLAM: Geometry-Only Gaussian Splatting for Dense Monocular SLAM | arXiv | paper |
| 2026-07-05 | LiDAR 3DGS SLAM, Geometry-Aware Covariance, Real-Time | Authors | Real-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance Coupling | arXiv | github |
| 2026-07-02 | Dynamic Gaussian SLAM, Dual-Level Probability, Semantic Map | Authors | DL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level Probability | arXiv | paper |
| 2026-06-29 | Task-Conditioned 3DGS, Real-Time Mapping, Multi-Agent Fusion | MIT | GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping | arXiv | paper |
| 2026-06-29 | RGB-Only Gaussian SLAM, Closed-Loop Geometry, Scale Feedback | CAS / USTC | MyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAM | IROS 2026 | paper |
| 2026-06-27 | Object-Centric 3DGS, Lifelong Mapping, Dynamic Maintenance | Beijing Institute of Technology | CubifyGS: Object-Centric 3D Gaussian Splatting for Lifelong Dynamic Scene Maintenance | IROS 2026 | paper |
| 2026-03-10 | Uncertainty-Aware 3DGS, RGB-D SLAM, Loop Detection | George Mason University | VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM | CVPR 2026 | project / github |
| 2025-11-20 | Language-Embedded, Open-Vocabulary, 3DGS SLAM | KAIST | LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM | arXiv | paper |
| 2025-07-25 | Neural SLAM, Dense Mapping, Self-Supervised | Shanghai AI Lab | DINO-SLAM: Dense Tracking and Mapping with Self-Supervised Feature Learning | arXiv | paper |
| 2025-07 | Dynamic Surface, Non-Rigid, 4D Tracking | Imperial College | 4DTAM: Dynamic Surface Gaussian SLAM | CVPR 2025 | paper |
| 2025-03-20 | 4DGS SLAM, Dynamic/Static, Tracking | Authors | 4D Gaussian Splatting SLAM | ICCV 2025 | paper / paper |
| 2025-03-11 | Gaussian SLAM, Dense Reconstruction, RGB-D | Zhejiang University | GigaSLAM: Gaussian Splatting-based Large-Scale Dense SLAM | arXiv | github |
| 2025-03 | Multi-Agent, 3DGS SLAM, Loop Closure | Authors | MAGiC-SLAM: Multi-Agent 3DGS SLAM | CVPR 2025 | paper |
| 2025-01-25 | Gaussian SLAM, Dynamic Environments, Monocular | Stanford / ETH | WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments | CVPR 2025 | project |
| 2025-01-24 | Multi-Robot, Semantic, Heterogeneous, 3DGS | Stanford | HAMMER: Heterogeneous, Multi-Robot Semantic Gaussian Splatting | RAL 2025 | project |
| 2024-11-03 | Gaussian SLAM, Global BA, Monocular RGB | ETH / Meta | Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians | CVPR 2025 | project |
| 2024-09-10 | Single-Image Calibration, Geometric Optimization | ETH / Meta | GeoCalib: Learning Single-image Calibration | ECCV 2024 | paper |
| 2024-04-11 | Detector-Free SfM, Texture-Poor | Zhejiang Univ | Detector-Free Structure from Motion | CVPR 2024 | github |
| 2024-04 | Pose Regression, Map-Relative, Multi-Scene, Highlight | Niantic / Oxford | Marepo: Map-Relative Pose Regression for Visual Re-Localization | CVPR 2024 Highlight | github |
| 2024-02-20 | Neural SLAM, Survey, Radiance Fields | TUM | How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey | arXiv | project |
| 2023-12-11 | Dense SLAM, Gaussian Splatting, RGB-D | TUM | Gaussian Splatting SLAM | CVPR 2024 | github |
| 2023-12-07 | End-to-End SfM, Differentiable BA | Meta AI / Oxford | VGGSfM: Visual Geometry Grounded Deep SfM | CVPR 2024 | github |
| 2023-12-04 | Gaussian SLAM, RGB-D, Volumetric | CMU / MIT | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM | CVPR 2024 | project / github |
| 2021-12-22 | Neural Mapping, Dense RGB-D SLAM, SDF | HKUST | NICE-SLAM: Neural Implicit Scalable Encoding for SLAM | CVPR 2022 | project |
This section focuses on the mathematical and data-structure layer used to represent geometry, appearance, and motion.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-21 | Gaussian Surfels, Topology Recovery, Mesh Extraction, Surface Reconstruction | University of Science and Technology of China | TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction | arXiv | github |
| 2026-08-20 | Sparse Light Field, Casual Capture, 3D/4D Reconstruction, 3DGS | Cornell University | Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction | arXiv | project |
| 2026-08-17 | Pose-Free NVS, 3DGS Geometry, Visibility-Aware Guidance | Aalto University | SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis | arXiv | paper |
| 2026-08-13 | Feed-Forward 3DGS, 3D Anchors, Spatially Grounded Tokens | National University of Defense Technology | LocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian Splatting | arXiv | project |
| 2026-08-03 | Streaming Feed-Forward, Persistent Geometry, Memory-Bounded 3DGS | University of Science and Technology of China | StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting | arXiv | paper |
| 2026-08-03 | View-Conditioned, Feed-Forward 3DGS, Generalizable Reconstruction | Tsinghua University | UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction | arXiv | paper |
| 2026-07-28 | Dynamic 3DGS, Adaptive Streaming, Volumetric Video | Authors | SplatStream: Fine Granular Scalable Gaussian Splatting for Adaptive 3D Scene Streaming | arXiv | paper |
| 2026-07-22 | Feed-Forward 3DGS, Adaptive Tokens, Compact Representation | Yonsei University | ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion | arXiv | project |
| 2026-04-16 | Feed-Forward 3DGS, Global Scene Tokens, Compact | Tel Aviv University | GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens | arXiv | paper |
| 2026-04-16 | TokenGS, Learnable Gaussian Tokens, Pose-Robust | NVIDIA | TokenGS: Decoupling 3D Gaussian Prediction from Pixels | CVPR 2026 Highlight | project |
| 2026-04-12 | UniSplat, Unposed Multi-View, Feed-Forward | UC Berkeley | UniSplat: Learning 3D Representations from Unposed Multi-View Images | CVPR 2026 | paper |
| 2026-02-02 | Feed-Forward 2DGS, Surface Continuity, Sparse Views | Shanghai Jiao Tong | SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors | ICLR 2026 | paper |
| 2025-07-31 | Sparse-View 4D, Monocular Fusion, Cross-Video | Meta AI | MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion | ICCV 2025 | paper |
| 2025-06-17 | Anti-Aliasing, Gaussian, 3D, Adaptive | CUHK | AAA-Gaussians: Anti-Aliasing 3D Gaussian Splatting | ICCV 2025 | paper |
| 2025-06-05 | Feed-Forward 3DGS, Depth Parameterization, Geometry | Zhejiang University | Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting | 3DV 2026 | paper |
| 2025-05-21 | Sparse-View Surface, Geometry-Prioritized, 2DGS | Nankai Univ | Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse Views | CVPR 2025 | paper |
| 2025-04-03 | Feed-Forward, No Camera, FreeSplatter | ZJU | FreeSplatter: Pose-Free 3DGS from Sparse Views | CVPR 2025 | paper |
| 2025-03-13 | Few-Shot, Diffusion Prior, Repair + Inpainting | ZJU / Alibaba | RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion Priors | ICCV 2025 | paper |
| 2025-01-11 | Scene-Level, 3DGS, Open-Vocabulary, 4D LangSplat | UPenn | 4D LangSplat: 4D Language Gaussian Splatting | CVPR 2025 | paper |
| 2024-12-09 | Large-Scale Dynamic, City, 4DGS, Spotlight | ZJU | DynamicCity: Large-Scale 4D Gaussian City Modeling | ICLR 2025 Spotlight | paper |
| 2024-09-26 | Progressive, Pruning, 3DGS, CVPR | S-Lab | PUP 3D-GS: Progressive Pruning for 3DGS | CVPR 2025 | paper |
| 2024-03-26 | 2DGS, Surface Reconstruction, Geometry | TUM | 2D Gaussian Splatting for Geometrically Accurate Radiance Fields | SIGGRAPH 2024 | project |
| 2024-03-21 | Feed-Forward 3DGS, Sparse Views, Generalizable | Monash / Tübingen | MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images | ECCV 2024 Oral | project / github |
| 2024-03-14 | Multi-Scale, 3DGS, Level-of-Detail | Shanghai Jiao Tong | Multi-Scale 3DGS | CVPR 2024 | paper |
| 2023-12-19 | Feed-Forward 3DGS, Image Pairs, Epipolar | MIT | pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction | CVPR 2024 Oral | github |
| 2023-12-01 | Distillation, Lightweight, 3DGS, Spotlight | Virginia Tech | LightGaussian: Distilled 3D Gaussian Splatting | NeurIPS 2024 Spotlight | paper |
| 2023-11-30 | 3DGS, Structure, Sparse Views | USTC | SparseGS: Real-Time 360 Sparse View Synthesis using Gaussian Splatting | 3DV 2025 | project |
| 2023-11-27 | Anti-Aliasing, Mip, 3DGS, Best Student Paper | Inria | Mip-Splatting: Alias-free 3DGS | CVPR 2024 Oral / Best Student Paper | github |
| 2023-11-27 | Compression, Compact, 3DGS, Highlight | Tsinghua | Compact 3D Gaussian Splatting | CVPR 2024 Highlight | paper |
| 2023-11-21 | 3DGS, Mesh Extraction, Surface | ETH Zurich | SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction | CVPR 2024 | project |
| 2023-08-08 | 3DGS, Real-Time Rendering, Explicit Radiance | Inria | 3D Gaussian Splatting for Real-Time Radiance Field Rendering | SIGGRAPH 2023 | github / paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-18 | Sparse Voxel Latent, Fitting-Free 3DGS, Large-Area Generation | Amap, Alibaba | GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation | arXiv | paper |
| 2026-07-22 | Shape Completion, Unified 3D Representation, Faithful Geometry | Authors | Axolotl3D: A Unified Framework for Faithful 3D Shape Completion | arXiv | paper |
| 2024-12-02 | Structured Latent, 3D Generation, Mesh | Microsoft | TRELLIS: Structured 3D Latents for Scalable and Versatile 3D Generation | CVPR 2025 | project / github |
| 2024-03-04 | Sparse Voxel, Text/Image-to-3D, Mesh | Stability AI | TripoSR: Fast 3D Object Reconstruction from a Single Image | arXiv | github |
| 2022-12-16 | Point Cloud, Diffusion, Shape Generation | OpenAI | Point-E: A System for Generating 3D Point Clouds from Complex Prompts | arXiv | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2025-07-15 | Unbiased SDF, Neural Implicit Surface, SDF-to-Density | CUHK | UNIS: Unified Framework for Unbiased Neural Implicit Surfaces | ICCV 2025 | paper |
| 2024-09-06 | α-NeuS, Alpha, SDF, Volume Rendering | ZJU | α-NeuS: Alpha-Governed Neural Implicit Surfaces | NeurIPS 2024 | paper |
| 2023-12-29 | Objects as Volumes, NeRF, 3D Reconstruction, Oral | UPenn | Objects as Volumes: Feed-Forward 3D from a Single Image | CVPR 2024 Oral | paper |
| 2023-04-13 | Zip-NeRF, Anti-Aliasing, Mip, Grid | Google Research | Zip-NeRF: Anti-Aliased Grid-Based NeRF | ICCV 2023 | project |
| 2022-03-17 | Tensor Factorization, Radiance Fields, Compression | Tsinghua | TensoRF: Tensorial Radiance Fields | ECCV 2022 | project |
| 2022-01-16 | Hash Grid, Real-Time NeRF, Neural Graphics | NVIDIA | Instant Neural Graphics Primitives with a Multiresolution Hash Encoding | SIGGRAPH 2022 | github / paper |
| 2021-12-15 | EG3D, Triplane, 3D GAN, Generative | NVIDIA | EG3D: Efficient Geometry-aware 3D Generative Adversarial Networks | CVPR 2022 Oral | github |
| 2021-12-09 | Plenoxels, No Neural Network, Fast, Oral | UC Berkeley | Plenoxels: Radiance Fields without Neural Networks | CVPR 2022 Oral | github |
| 2021-12-07 | Ref-NeRF, Reflection, Specular, Best Student Paper HM | Google Research | Ref-NeRF: Structured View-Dependent Appearance for NeRF | CVPR 2022 Best Student Paper HM | project |
| 2021-11-23 | Unbounded Scenes, Anti-Aliasing, NeRF | Google Research | Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields | CVPR 2022 | project |
| 2021-06-20 | SDF, Surface Reconstruction, Neural Rendering | MPI-IS | NeuS: Learning Neural Implicit Surfaces by Volume Rendering | NeurIPS 2021 | project |
| 2021-04-13 | BARF, Bundle-Adjusting NeRF, Oral | UC Berkeley | BARF: Bundle-Adjusting Neural Radiance Fields | ICCV 2021 Oral | github |
| 2020-08-05 | NeRF-W, Unbounded, In-the-Wild | Google Research | NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections | CVPR 2021 | github |
| 2020-03-19 | NeRF, View Synthesis, Neural Rendering | UC Berkeley | NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis | ECCV 2020 | github / paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | Dynamic 3DGS, Gradient Decoupling, Novel View Synthesis | Authors | GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis | arXiv | paper |
| 2026-06-22 | Dynamic 3DGS, Visibility-Aware Densification, Temporal Lifespan | Indian Institute of Science | Temporally Aware Densification for Dynamic 3D Gaussian Splatting | arXiv | paper |
| 2025-07 | 7DGS, Spatial-Temporal-Angular, Unified | Authors | 7D Gaussian Splatting: Unified Spatial-Temporal-Angular GS | ICCV 2025 | paper |
| 2024-10 | Event Camera, High-Speed, 4D, STD-GS | Authors | STD-GS: SpatioTemporal-Disentangled Gaussian Splatting with Event Cameras | ICCV 2025 | paper |
| 2023-10-12 | 4DGS, Dynamic Scenes, Real-Time Rendering | Zhejiang University | 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering | CVPR 2024 | project |
| 2023-08-18 | Dynamic 3DGS, Scene Motion, Multi-View Video | Cornell | Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis | 3DV 2025 | github / paper |
| 2023-01-24 | K-Planes, Explicit, 4D, Space-Time | Authors | K-Planes: Explicit Radiance Fields in Space, Time, and Appearance | CVPR 2023 | github |
| 2023-01-23 | HexPlane, 4D Representation, Space-Time | CMU | HexPlane: A Fast Representation for Dynamic Scenes | CVPR 2023 | project |
| 2021-06-24 | Dynamic NeRF, Deformation, Canonical Space | Google Research | HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields | SIGGRAPH Asia 2021 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-24 | Structured Motion, SE(3), 4D Reconstruction | Tsinghua University | SM4RT: Learning Structured Motion Geometry for 4D Reconstruction | arXiv | project / paper |
| 2023-12-04 | Deformation, Dynamic Radiance Fields, Canonical | ETH Zurich | SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes | CVPR 2024 | project |
| 2023-06-05 | Non-Rigid Tracking, Neural Deformation, 4D | Tsinghua | Neuralangelo: High-Fidelity Neural Surface Reconstruction | CVPR 2023 | project |
Reconstruction systems recover objects or scenes from images, video, RGB-D, or multi-sensor streams under offline, online, and dynamic conditions.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Active In-Hand Reconstruction, Uncertainty, Next-Best View | ShanghaiTech University | AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction | CoRL 2026 | project |
| 2026-07-21 | Single-View 3D, Object Perception, Generative Reconstruction | Deakin University | Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction | arXiv | project |
| 2026-06-17 | Sparse-View Object, Flow Steering, 3DGS Refinement | Graz University of Technology | FlowObject: Flow Steering for Bridging Generative Priors and Reconstruction Fidelity | arXiv | project |
| 2026-05-05 | Generative Reconstruction, Multi-View Alignment, Pose | Tsinghua | Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation | arXiv | project |
| 2025-11-19 | Single Image, Object Mesh, SAM 3D | Meta AI | SAM 3D Objects | GitHub | github |
| 2025-10-23 | Pose-Free Online, Free-Moving Objects, Constant Memory | SUTD | OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects | NeurIPS 2025 Spotlight | project |
| 2025-06-05 | RGB-D Object Completion, Novel Depth, Feed-Forward | Carnegie Mellon University | RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion | NeurIPS 2025 | project |
| 2025-06 | Feed-Forward, Densification, Gaussian, Detail | Authors | Generative Densification: Feed-Forward 3DGS Densification | CVPR 2025 | paper |
| 2025-06 | Photogrammetry Foundation, Multi-Task, Highlight | Authors | Matrix3D: A Foundation Model for Photogrammetry | CVPR 2025 Highlight | paper |
| 2025-04-04 | Sparse-View, Feed-Forward, Camera, Geometry | Ant Research / Stanford | FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | CVPR 2025 | project |
| 2024-07 | Ego-Centric, Autonomous Driving, Sparse-View | Authors | Omni-Scene: Omni-Gaussian for Ego-Centric Sparse-View Reconstruction | CVPR 2025 | paper |
| 2024-06-14 | Multi-View, Stereo, Feed-Forward | NAVER Labs | MASt3R: Grounding Image Matching in 3D with MASt3R | ECCV 2024 | project / github |
| 2023-12-21 | Multi-View, Pointmap, Pose-Free | NAVER Labs | DUSt3R: Geometric 3D Vision Made Easy | CVPR 2024 | project / github |
| 2023-12-13 | Object Pose, Reconstruction, Model-Based | NVIDIA | FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects | CVPR 2024 | project |
| 2023-03-24 | Unknown Object, RGB-D, 6-DoF Tracking, Neural SDF | NVIDIA | BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects | CVPR 2023 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Feed-Forward, Compositional Scene, Complete Meshes, Simulation-Ready | University of Illinois Urbana-Champaign | FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute | arXiv | project / github |
| 2026-09-04 | Bundle Adjustment, Multi-View Matching, Monocular Priors, Online+Offline | NAVER Labs Europe | BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors | ECCV 2026 | paper |
| 2026-08-31 | Real-to-Sim, Parse-Generate-Place, Composable Object Assets | ByteDance Seed | Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling | arXiv | project |
| 2026-08-25 | Single Image, Generative Reconstruction, Complete Object Assets, Scene Assembly | Huawei | SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image | arXiv | paper |
| 2026-08-18 | Instance-Grouped 3DGS, Semantic Reconstruction, Referential Scene Graph | Shanghai Jiao Tong University | GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting | arXiv | paper |
| 2026-08-18 | Generative NVS, Reconstruction/Generation Split, Scene Coordinates | KAIST | GenRec: Knowing Where to Reconstruct and Where to Generate | arXiv | project |
| 2026-08-18 | Long Sequence, Chunk Priors, Sim(3) Assembly, Test-Time Adaptation | Kosmo Research | GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly | arXiv | project |
| 2026-08-15 | Long Sequence, Scale-Consistent Alignment, Test-Time Adaptation | Northwestern Polytechnical University | VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction | ACM Multimedia 2026 | github |
| 2026-08-12 | Streaming Multi-View, Metric 3D, Feed-Forward Prior, Object Detection | ETH Zurich | Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs | ECCV 2026 | project |
| 2026-08-07 | Sparse View, Editable Indoor Scenes, Executable Scene Programs | City University of Hong Kong | Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs | arXiv | paper |
| 2026-07-31 | Active Reconstruction, Next-Best-View, Predictive Entropy | Fudan University | GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction | arXiv | paper |
| 2026-07-23 | Underwater 3D, Feed-Forward Reconstruction, Degradation Adaptation | HKUST | WAT3R: Feedforward Underwater 3D Reconstruction | arXiv | project |
| 2026-07-15 | Feed-Forward Driving Reconstruction, Layered 3DGS, Dynamic Actors | NVIDIA | Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation | arXiv | project / github / docs |
| 2026-07-10 | 3D Foundation Model, Global SfM, Bundle Adjustment | HKUST | Glob3R: Global Structure-from-Motion with 3D Foundation Models | arXiv | project |
| 2026-07-08 | Feed-Forward 3D, Unposed Images, Drift-Robust | Authors | NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction | ECCV 2026 | paper |
| 2026-06-09 | RGB-T, Thermal Geometry, Low-Light | University of Minnesota | DarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax | arXiv | project |
| 2026-06-02 | Single Image, Physics-in-the-Loop, Simulation-Ready | Seoul National University | SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image | arXiv | project |
| 2026-05-14 | VGGT, Scaling, Static+Dynamic | University of Oxford | VGGT-Omega: Scaling VGGT to Large-Scale 3D Reconstruction | CVPR 2026 Oral | project / github / demo / model |
| 2026-05-07 | Feed-Forward 3D, Token Reduction, Long Sequence | Peking University | Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction | arXiv | paper |
| 2026-04-30 | Generalizable, Sparse-View, Unposed Images, Outdoor | UIUC / NVIDIA | GenWildSplat: Generalizable Sparse-View 3D Reconstruction from Unconstrained Images | arXiv | paper |
| 2026-03-24 | Panoramic Video, Pose-Free 3DGS, Consistent Depth Prior | University of Chinese Academy of Sciences | Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors | CVPR 2026 | github |
| 2026-03-16 | Event-to-Edge, Pose-Free, Gaussian Reconstruction | KAIST | E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction | CVPR 2026 | paper |
| 2026-02-26 | VGGT, TTT, Large-Scale | NVIDIA | VGG-T^3: Offline Feed-Forward 3D Reconstruction at Scale | CVPR 2026 | project |
| 2026-02-03 | Single Image, Object Decomposition, Occlusion-Aware Scene Reconstruction | University of California, San Diego | Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal | 3DV 2026 | project |
| 2025-09-24 | Mirror Stereo, Single-View 3D, Symmetry Constraint | University of Oxford | Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections | 3DV 2026 | project / github / dataset |
| 2025-09-16 | Universal 3D, Metric Reconstruction, Optional Priors | Meta AI | MapAnything: Universal Feed-Forward Metric 3D Reconstruction | 3DV 2026 | project |
| 2025-08-05 | Unposed Multi-View, 3DGS, Semantic Reconstruction | Sungkyunkwan University | Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images | CVPR 2026 | paper |
| 2025-07-17 | Permutation-Equivariant, Visual Geometry, Point Maps | Oxford / Meta | π³: Permutation-Equivariant Visual Geometry Learning | ICLR 2026 | github |
| 2025-07 | Monocular Prior, MVS, DTU/Tanks SOTA | Authors | MonoMVSNet: Monocular Prior Guided MVS | ICCV 2025 | paper |
| 2025-07 | Latent Align, Stereo+Monocular, Highlight | Authors | BridgeDepth: Unified Monocular and Stereo Depth | ICCV 2025 Highlight | paper |
| 2025-06-30 | Video-Depth Augmentation, Scalable Training, Feed-Forward 3D | Australian National University | Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction | NeurIPS 2025 | project |
| 2025-06-03 | Monocular Depth, Dynamic Video, Alignment | HKUST / CUHK / HKU | Align3R: Aligned Monocular Depth Estimation for Dynamic Videos | CVPR 2025 | github |
| 2025-06 | Aerial-Ground, Large-Scale, 3DGS | Authors | Horizon-GS: Unified Aerial-Ground 3DGS | CVPR 2025 | paper |
| 2025-06 | Autonomous Driving, Feed-Forward, 3DGS | Authors | EVolSplat: Feed-Forward 3DGS for Urban Driving | CVPR 2025 | paper |
| 2025-06 | Sparse-View, Super-Resolution, 3DGS | Authors | S2Gaussian: Sparse-View Super-Resolution 3DGS | CVPR 2025 | paper |
| 2025-05-05 | Relative Camera Pose, Regression, Localization | Aalto / HKU | Reloc3r: Large-Scale Training of Relative Camera Pose Regression | CVPR 2025 | github |
| 2025-03-28 | Feed-Forward, Surface, MVS, Multi-View | Authors | MVSAnywhere: Zero-Shot Multi-View Stereo | CVPR 2025 | paper / paper |
| 2025-03-17 | Feed-Forward, Multi-View, Auxiliary Priors | ETH / Microsoft | Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | CVPR 2025 | project |
| 2025-03-14 | Feed-Forward 3D, Pose-Free, Point Map | Meta AI | VGGT: Visual Geometry Grounded Transformer | CVPR 2025 Best Paper | project / github |
| 2025-03-03 | Multi-View, Symmetric, 1000+ Images, O(N) | NAVER Labs | MUSt3R: Multi-view Network for Stereo 3D Reconstruction | CVPR 2025 | project |
| 2025-01-23 | Feed-Forward, 1500+ Images, 251 FPS | Meta AI | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | CVPR 2025 | project |
| 2024-12-12 | Feed-Forward, Online, Dense, Monocular | Shanghai AI Lab / PKU | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | CVPR 2025 Highlight | project |
| 2024-12-09 | Sparse View, Single-Stage, 2 Seconds | Meta AI | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | CVPR 2025 Oral | project |
| 2024-09-27 | Multi-View Reconstruction, Matching, MVS | NAVER Labs | MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion | 3DV 2025 | github |
| 2024-08-28 | Spatial Memory, Feed-Forward, Multi-View | HKU | Spann3R: 3D Reconstruction with Spatial Memory | 3DV 2025 Oral (Best Paper Candidate) | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-07 | Online SLAM, Functional Scene Graph, Interaction Elements, Map Memory | Tsinghua University | Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs | arXiv | github |
| 2026-09-01 | Training-Free, Open-Vocabulary Instance Map, RGB-D/Monocular SLAM | University of Technology Sydney | VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM | arXiv | paper |
| 2026-08-24 | Spatio-Temporal SLAM, Open-Vocabulary, 4D Scene Graph, VLN | Carnegie Mellon University | SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation | arXiv | project |
| 2026-08-18 | Open-Vocabulary Map, Instance Preservation, Fine-Grained Retrieval, Target Absence | Xi'an Jiaotong University | OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects | arXiv | github |
| 2026-06-23 | Object-Level Map, Open-Vocabulary, Relocalization | Zhejiang University | Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization | arXiv | paper |
| 2026-05-05 | Online, Voxel+Instance, Open-Vocabulary Mapping | Örebro University | FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping | arXiv | project |
| 2026-05-03 | Gaussian-Language Map, Zero-Shot Navigation, Multi-Scale | CASIA | Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning | arXiv | github |
| 2026-03-04 | Semantic 3DGS, Online, CLIP, Open-Vocabulary | National University of Singapore | EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding | CVPR 2026 | project |
| 2025-08-02 | Open-Vocabulary, Hybrid 3DGS+TSDF, Dense Mapping | Tsinghua | OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting | arXiv | project |
| 2025-07 | Feed-Forward, Panoramic Segmentation, DUSt3R | Authors | PanSt3R: Single-Feed 3D and Panoptic Segmentation | ICCV 2025 | paper |
| 2023-10-05 | Open-Vocabulary, RGB-D, TSDF, Real-Time Mapping | University of Arkansas | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation | IROS 2024 | project / github |
| 2023-02-14 | Open-Set, Multimodal, 3D Map, Language Query | MIT | ConceptFusion: Open-set Multimodal 3D Mapping | ICRA 2023 | project / github |
| 2022-10-11 | Implicit Field, CLIP, Semantic Search, Robot Memory | New York University | CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory | ICRA 2023 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-03 | Online 3R, Multi-Relative Pose Query, Pose-Graph Optimization | National Yang Ming Chiao Tung University | Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction | ECCV 2026 | project |
| 2026-09-01 | Online Feed-Forward 3R, Unordered UAV Images, Retrieval+Retry | Wuhan University | On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios | arXiv | github |
| 2026-08-03 | Active Reconstruction, Ergodic Coverage, Trajectory Optimization | Johns Hopkins University | TRACE: Ergodic Trajectory Optimization for Active Scene Reconstruction | arXiv | github |
| 2026-08-03 | Feed-Forward SLAM, Sim(3) Factor Graph, Persistent Mapping | Ulsan National Institute of Science and Technology | UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization | ECCV 2026 | project |
| 2026-07-25 | Semantic SLAM, Data Association, Object Landmarks | MIT | Semantic Semi-Incremental Data-Association-Free Object SLAM | arXiv | paper |
| 2026-07-23 | Gaussian SLAM, Large-Scale Mapping, Real-Time | Athena Research Center | GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition | IROS 2026 | github |
| 2026-07-16 | Multi-Agent 3R, RGB Video, Point-Map Fusion | University of Bologna | MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos | arXiv | project |
| 2026-07-16 | Incremental 3DGS, Unordered Capture, Global Consistency | Inria | Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency | SIGGRAPH 2026 | paper |
| 2026-07-01 | Long-Sequence, Instance Anchors, Persistent Spatial Memory | Beijing Jiaotong University | LIST3R: Long-sequence Instance-aware 3D Reconstruction | arXiv | project |
| 2026-06-23 | 3DGS-SLAM, Memory-Efficient, Outdoor Mapping | University of Minnesota | Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM | ICRA 2026 | github |
| 2026-06-20 | RGB+Pose, 3DGS Scene Regression, Robot Capture | Peking University | ACEsplat: Accelerated 3D Gaussian Scene Regression via RGB and Poses Only | arXiv | paper |
| 2026-06-19 | 3DGS-SLAM, Degeneracy-Robust, Real-Time Tracking | Nanyang Technological University | Spectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAM | IROS 2026 | paper |
| 2026-06-18 | LiDAR-Inertial-Thermal, 3DGS Mapping, Illumination-Robust | Shenzhen University | LIT-GS: LiDAR-Inertial-Thermal Gaussian Splatting for Illumination-Robust Mapping | IROS 2026 | paper |
| 2026-06-03 | Streaming, Transient Anchors, Long-Horizon Mapping | Authors | Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping | arXiv | paper |
| 2026-06 | Spatial-Difference Sensor, Edge-Guided Tracking, 3DGS | Tsinghua University | SDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D Reconstruction | CVPR 2026 | paper |
| 2026-06 | Stereo 3DGS-SLAM, Auto-Exposure Robustness, Photometric Mapping | South China University of Technology | AERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAM | CVPR 2026 | github |
| 2026-05-10 | VGGT, Retrieval, Constant Memory | Fudan University | RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval | arXiv | github |
| 2026-04-24 | 4DGS-SLAM, Optical Flow, Dynamic Mapping | National University of Singapore | Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM | CVPR 2026 | github |
| 2026-04-15 | Streaming, Feed-Forward, Long Video | Robbyant | LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction | arXiv | github |
| 2026-02-13 | Streaming, Autoregressive, Long Sequence | 3DAgentWorld | LongStream: Long-Sequence Streaming Autoregressive Visual Geometry | CVPR 2026 | project |
| 2026-01-03 | StreamVGGT, KV Cache, Memory Compression | Sun Yat-sen University | XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer | arXiv | github |
| 2025-09-30 | TTT, Online, Long Context | Shanghai AI Lab | TTT3R: 3D Reconstruction as Test-Time Training | arXiv | project |
| 2025-08-14 | Streaming, Causal Transformer, Sequential | NTU / Shanghai AI Lab | STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer | arXiv | project |
| 2025-01-21 | Online 3D, Recurrent Pointmap, Streaming | Meta AI | CUT3R: Continuous 3D Perception Model with Persistent State | CVPR 2025 Oral | project / github |
| 2024-12-16 | MASt3R, Dense SLAM, Real-Time | Imperial College London | MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors | CVPR 2025 | github |
| 2023-09-05 | Global BA, Neural Implicit, Dense RGB-D SLAM | University of Bologna | GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction | ICCV 2023 | project / github |
| 2022-11-21 | Hybrid SDF, Dense RGB-D SLAM, Keyframes | Idiap Research Institute | ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields | CVPR 2023 | project / github |
| 2021-08-24 | Deep SLAM, Dense BA, Monocular/Stereo/RGB-D | Princeton University | DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras | NeurIPS 2021 | github |
| 2021-03-23 | Neural Implicit, Online RGB-D, Dense SLAM | Imperial College London | iMAP: Implicit Mapping and Positioning in Real-Time | ICCV 2021 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Long-Range 4D Motion, 3D Queries, Occlusion-Robust Trajectory Chaining | Carnegie Mellon University | Point4D: Long-range 4D Motion Reconstruction | arXiv | project |
| 2026-09-08 | Event Stream, Extreme-Low-Frame-Rate RGB, Dynamic 3DGS, Real-Time | Macau University of Science and Technology | EdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene Reconstruction | Neurocomputing | github+dataset |
| 2026-09-05 | Sparse-View 4D, Spatio-Temporal Depth Alignment, Dynamic 3DGS | BIGAI | UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment | ECCV 2026 | project |
| 2026-08-18 | Query-Conditioned 4D, Scene Flow, Dynamic Points, Sparse-to-Dense | Kosmo Research | UniQuery4R: Unified 4D Scene Reconstruction from a Single Query | arXiv | project |
| 2026-07-29 | Articulated Objects, Structure-aware 3DGS, Part Connectivity | POSTECH | StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction | arXiv | paper |
| 2026-07-21 | Streaming 4D, Instance Grounding, Geometry Transformer | Horizon Robotics | IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer | arXiv | project |
| 2026-07-16 | Online Dynamic NVS, Space-Time Memory, Real-Time | University of Washington | Online Neural Space Time Memory for Dynamic Novel View Synthesis | arXiv | project |
| 2026-07-01 | Dynamic Gaussian Reconstruction, Monocular Video, Generative | Stanford University | World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video | arXiv | project |
| 2026-06-23 | Articulated Digital Twin, RGB-D, URDF Export | ETH Zurich | ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D Videos | ICRA 2026 Workshop | paper |
| 2026-06-22 | Monocular Video, 4DGS, In-the-Wild Non-Rigid | Carnegie Mellon University | Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild | arXiv | project |
| 2026-06-22 | Dynamic Driving, Sparse Voxels, LiDAR-Guided | Huawei Paris Research Center | DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction | arXiv | paper |
| 2026-06-09 | Future Extrapolation, 4DGS, Autonomous Driving | Tsinghua University | Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving | arXiv | project |
| 2026-06-09 | Manipulation Video, Decoupled 3DGS, Scene Graph | Authors | ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting | arXiv | paper |
| 2026-06-02 | Object Permanence, Differentiable Physics, 4DGS | Authors | PersistGS: Differentiable Physics for Object Permanence in 4D Gaussian Splatting | CVPR 2026 Workshop | paper |
| 2026-06 | Single Event Camera, Deformable 3DGS, High-Speed 4D | ShanghaiTech University | FastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event Camera | CVPR 2026 | paper |
| 2026-04-10 | Feed-Forward, Unconstrained Views, Semantic-Geometry | Nanyang Tech | FF3R: Feedforward Feature 3D Reconstruction from Unconstrained Views | CVPR 2026 Findings | paper |
| 2026-04-10 | Dynamic 4D, Semantic Prior, Gaussian SLAM, Action-Control | University of Zurich | Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction | arXiv | paper |
| 2026-04-10 | Dynamic/Static Disentanglement, Uncertainty-Aware, Feed-Forward | Zhejiang University | Robust 4D VGT: Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors | arXiv | paper |
| 2026-04-07 | Functional Scenes, Egocentric Interaction, URDF/USD | Stanford | FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos | CVPR 2026 | paper |
| 2026-04-05 | Sparse Camera, 4DGS, Neural Decay, CVPR 2026 | Authors | 4C4D: 4 Camera 4D Gaussian Splatting | CVPR 2026 | paper / paper |
| 2026-03-30 | Dynamic Surface, Explicit Geometry, High Fidelity | Australian National University | 4DSurf: High-Fidelity Dynamic Scene Surface Reconstruction | CVPR 2026 | paper |
| 2026-03-21 | RayMap, Dynamic, Streaming | University of Illinois Chicago | RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction | arXiv | project / github |
| 2026-03-09 | Dynamic VGGT, Autonomous Driving, 4D | Fudan University | DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving | arXiv | paper |
| 2025-11-23 | Dynamic Geometry, Spatiotemporal, VGGT | Huazhong University of Science and Technology | 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation | arXiv | paper |
| 2025-11-07 | Motion-Aware, Monocular Video, Bundle Adjustment | KAIST | 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes | NeurIPS 2025 | paper |
| 2025-10-20 | VGGT-4D, Pose/Geometry, Dynamic Mask | Harvard | PAGE-4D: Disentangled Pose and Geometry Estimation for 4D Perception | ICLR 2026 | project |
| 2025-08-13 | 3D Reconstruction, Human Motion, Video | Shanghai AI Lab | Human3R: Reconstructing 3D Human Avatars from Monocular Video | CVPR 2025 | project |
| 2025-07 | Human, Multi-View, Sparse, Robust, RoGSplat | Authors | RoGSplat: Robust Generalizable Human Gaussian Splatting | CVPR 2025 | paper |
| 2025-06-11 | Dynamic Human, Temporal Consistency, 4D | Tsinghua | CARI4D: Cross-Modal Alignment and Reconstruction for Interactive 4D Human | CVPR 2025 | project |
| 2025-06-10 | Online, Dynamic 3DGS, Uncalibrated Video | University of British Columbia | StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams | ICLR 2026 | project / github |
| 2025-06-09 | 4DGS, Transformer, Monocular Video | Meta Reality Labs | 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos | NeurIPS 2025 Spotlight | project |
| 2025-06-02 | Video Generators, 4D Geometry | Oxford VGG / NAVER LABS | Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction | ICCV 2025 Highlight | paper |
| 2025-06 | Self-Supervised, Dynamic, Driving, Flow | Authors | SplatFlow: Self-Supervised Dynamic 3DGS with Neural Motion Flow | CVPR 2025 | paper |
| 2025-06 | Few-Shot, Personal Avatar, Highlight | Authors | FRESA: Personalized 3D Human Avatar from Few Images | CVPR 2025 Highlight | paper |
| 2025-06 | Real-Time Avatar, 166fps, Gaussian, Highlight | Authors | MMLP-Human: Real-Time High-Fidelity Gaussian Human Avatar | CVPR 2025 Highlight | paper |
| 2025-05-27 | 4D, Dual Correspondences, Dynamic Video | NUS / Shanghai AI Lab | C4D: 4D Made from 3D through Dual Correspondences | ICCV 2025 | project |
| 2025-05-14 | 4D Pointmaps, Dynamic-Static Disentanglement | KAIST / ETH / Sony | D2USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic Scenes | NeurIPS 2025 | project |
| 2025-05-02 | Training-Free, Motion Disentangle, DUSt3R | Westlake / MPI | Easi3R: Estimating Disentangled Motion from DUSt3R Without Training | ICCV 2025 | project |
| 2025-04-17 | 4D Tracking, Feed-Forward, Pointmap, Tracking | MPI / UC Berkeley | St4RTrack: Simultaneous 4D Reconstruction and Tracking | ICCV 2025 | paper / paper |
| 2025-04-07 | Neural Rendering, Human, No Eyes | University of Cambridge | Seeing Without Eyes: Neural Human Rendering from Monocular Video | CVPR 2025 | project |
| 2025-03-24 | Multi-Object, 4D, In-the-Wild Videos | CMU | GenMOJO: Robust Multi-Object 4D Generation for In-the-wild Videos | CVPR 2025 | project |
| 2025-02-27 | Layered Avatar, Hair, Face, Meta | Authors | LUCAS: Layered Universal Codec Avatars | CVPR 2025 | paper / paper |
| 2025-01-22 | 3D Reconstruction, Canonical, Multi-View | Stanford | UniCon3R: Unified 3D Reconstruction and Recognition | CVPR 2025 | project |
| 2024-12-03 | Single Image, Animatable, Avatar, 4DGS | Authors | AniGS: Animatable Gaussian Avatar from a Single Image | CVPR 2025 | paper / paper |
| 2024-11-27 | 4D Generation, Multi-View Video, Diffusion | CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models | CVPR 2025 | project | |
| 2024-10-28 | Dynamic Geometry, DUSt3R, Motion | University of Oxford | MonST3R: Estimating Geometry in the Presence of Motion | ICLR 2025 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-02-12 | 4D Dynamic, Monocular Video, Tree-Chains | Cornell | WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains | ICLR 2026 | paper |
| 2025-11-01 | 4D Scene, Feed-Forward, Controllable, Video Diffusion | Shanghai AI Lab | Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models | CVPR 2026 | paper |
| 2025-10-15 | Dynamic 4D, Gaussian, Canonical | Zhejiang University | Director: Directed Generative Models for 4D Scene Evolution | CVPR 2025 | project |
| 2025-08-06 | Multi-Baseline, Generalizable, Gaussian | Authors | MuGS: Multi-Baseline Generalizable Gaussian Splatting | ICCV 2025 | paper / paper |
| 2025-07 | Surface, Gaussian Surfels, 2DGS, Sparse-View, Spotlight | Authors | MAtCha Gaussians: Atlas Charting with 2D Gaussian Surfels | CVPR 2025 Spotlight | paper |
| 2025-07 | SDF+3DGS, Hybrid, Surface, ICCV | Authors | SurfaceSplat: SDF+3DGS Hybrid Surface Reconstruction | ICCV 2025 | paper |
| 2025-07 | Sparse-View, Implicit, Voxel, Consistency | Authors | SparseRecon: Sparse-View Implicit Surface Reconstruction | ICCV 2025 | paper |
| 2025-07 | Low-Texture, Reflection, Unified, +21% | Authors | HiNeuS: Unified Neural Implicit Surface Reconstruction | ICCV 2025 | paper |
| 2025-06 | Joint Human+Scene, MASt3R Extension | Authors | HAMSt3R: Joint Human and Scene 3D Reconstruction | ICCV 2025 | paper |
| 2024-11-20 | 4D Reconstruction, Gaussian Splatting, Forward | Shanghai AI Lab | Forge4D: Gaussian Splatting for Forward Facing 4D Reconstruction | arXiv | paper |
This section tracks methods that create new 3D assets, parts, articulated objects, scenes, and editable 3D worlds, with emphasis on physical and simulation use.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-23 | Multi-View + LiDAR, Vehicle Assets, TRELLIS | Shanghai Jiao Tong University | MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving | arXiv | github |
| 2026-03-12 | Multi-View, SAM3D, Layout-Aware, Physical Plausibility | Peking University | MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation | arXiv | github |
| 2026-01-16 | Casual Capture, Posed Multi-View, Metric Shape | Meta AI | ShapeR: Robust Conditional 3D Shape Generation from Casual Captures | arXiv | paper |
| 2025-11-12 | Multi-Image Fusion, Region Control, TRELLIS | Zhejiang University | Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion | SIGGRAPH Asia 2025 | project / github |
| 2025-03-18 | Multi-View Image-to-Shape, Hunyuan3D-DiT | Tencent | Hunyuan3D 2.0 MV | Model | github / model |
| 2024-02-06 | Scalable View Synthesis, Single/Multi-Image 3D | KAUST | EscherNet: A Generative Model for Scalable View Synthesis | CVPR 2024 | project |
| 2019-08-05 | Multi-View Images, Mesh Deformation, Shape Refinement | National Tsing Hua University | Pixel2Mesh++: Multi-View 3D Mesh Generation via Deformation | ICCV 2019 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-25 | Relightable 3D Assets, PBR Maps, Gaussian Representation | Apple | Luce: Relightable Gaussians for 3D Asset Generation | arXiv | paper |
| 2026-08-24 | Material Decomposition, Physical Properties, Watertight Sub-Meshes, Sim-Ready | University of Bristol | Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition | arXiv | paper |
| 2026-07-01 | Complex Textures, Video Generative Prior, 3D Assets | Authors | Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models | arXiv | paper |
| 2024-11-10 | VLM-Guided, PBR Texture | 3D AIGC | TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting | CVPR 2025 | project |
| 2024-01-17 | TextureDreamer, Geometry-Aware, Diffusion | UCSD / Meta | TextureDreamer: Image-Guided Texture Synthesis | CVPR 2024 | paper |
| 2023-12-21 | Texture Generation, Mesh, Multi-View Consistency | Tencent | Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models | CVPR 2024 | project |
| 2023-11-28 | SceneTex, Indoor, Texture, Diffusion, Highlight | TUM / Snap | SceneTex: High-Quality Texture Synthesis for Indoor Scenes | CVPR 2024 Highlight | paper |
| 2023-11-21 | SyncMVD, Multi-View, Text-to-Texture | CUHK | SyncMVD: Text-Guided Texturing by Synchronized Multi-View Diffusion | CVPR 2024 | paper |
| 2023-08-22 | PBR Material, SVBRDF, Text-to-Material | Adobe | MatFuse: Controllable Material Generation with Diffusion Models | SIGGRAPH Asia 2024 | project |
| 2023-03-20 | Texture, Material, Text-to-Texture | KAIST | Text2Tex: Text-driven Texture Synthesis via Diffusion Models | ICCV 2023 | project |
| 2023-02-03 | TEXTure, Text-Guided, 3D Texture, Diffusion | Tel Aviv University | TEXTure: Text-Guided Texturing of 3D Shapes | SIGGRAPH 2023 | project |
| 2022-07-06 | nvdiffrec, 3D Mesh, Material, Lighting, Oral | NVIDIA | nvdiffrec: Extracting Triangular 3D Models, Materials, and Lighting | CVPR 2022 Oral | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-08 | Single Image, Physical CoT, URDF, Simulation-Ready | Aerospace Information Research Institute, CAS | PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets | arXiv | paper |
| 2026-07-15 | Articulation + Physics, 40K Assets, Simulation-Ready | Zhejiang University | UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets | arXiv | github |
| 2026-05-20 | Rigid/Deformable/Articulated, Physical Attributes, Sim-Ready | Nanyang Technological University | PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects | arXiv | project / dataset |
| 2026-05-14 | Agentic Generation, Articraft-10K, URDF Assets | Authors | Articraft: An Agentic System for Scalable Articulated 3D Asset Generation | arXiv | project / github |
| 2026-05-06 | Physics-Grounded, Kinematic, Simulation-Ready Assets | HKU | PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World | ICML 2026 | project / github |
| 2026-03-14 | URDF, Autoregressive, Simulation-Ready Assets | Authors | URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets | arXiv | paper |
| 2026-03-01 | Articulated Assets, 3D LLM, Kinematic Structure | Tsinghua | ArtLLM: Generating Articulated Assets via 3D LLM | CVPR 2026 | paper |
| 2025-12-12 | Articulation, Kinematic Tree, Feed-Forward, URDF-Ready | University of Oxford | Particulate: Feed-Forward 3D Object Articulation | arXiv | project |
| 2025-11-26 | Single Image, Open-Set Articulation, Unified Latent | ShanghaiTech | UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation | arXiv | paper |
| 2025-11-17 | Sim-Ready Assets, Physical Properties, Single Image | NTU | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | CVPR 2026 | paper |
| 2025-11-02 | URDF, 3D MLLM, Articulated Objects | Tsinghua | URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model | NeurIPS 2025 | paper |
| 2025-08-20 | Articulated Geometry, Motion Modeling, Gaussian Representation | Tsinghua University | GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects | 3DV 2026 | paper |
| 2025-07-16 | Physical Properties, Scale, Material, Affordance | Nanyang Technological University | PhysX-3D: Physical-Grounded 3D Asset Generation | NeurIPS 2025 Spotlight | project / github |
| 2025-06-10 | Interactable Digital Twin, Articulated Object, RGB-D Video | Shanghai Jiao Tong University | iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos | 3DV 2026 | paper |
| 2025-04-17 | Simulation-Ready, Physical Materials, Dynamics | University of Massachusetts Amherst | SOPHY: Generating Simulation-Ready Objects with Physical Materials | WACV 2026 | project / github |
| 2025-03-11 | Part-Level Digital Twin, Joint Estimation, Self-Supervised 3DGS | USTC | ArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splatting | CVPR 2025 | paper |
| 2025-02-26 | Articulated Objects, 3DGS, Joint Estimation | Tsinghua | ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting | ICLR 2025 | project |
| 2025-02-17 | Articulation-Ready, Skeleton, Skinning, Benchmark | Nanyang Technological University | MagicArticulate: Make Your 3D Models Articulation-Ready | CVPR 2025 | project / github |
| 2024-09-26 | Open-Vocabulary, URDF, Articulation | Stanford | Articulate Anything: Open-vocabulary 3D Articulated Object Generation | ICLR 2025 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-20 | Compositional 3D, Part Semantics, Spatial Control, Reassemblable Assets | Roblox | MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control | arXiv | project |
| 2026-08-14 | Up to 300 Parts, Token-Efficient VQ, Autoregressive 3D, Structured Assets | The University of Hong Kong | MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling | arXiv | project |
| 2026-08-13 | Part-Aware Generation, Recursive Decomposition, Editable Assets | Shanghai Jiao Tong University | SCULPT: Subtractive Composition for 3D Part Generation | arXiv | project |
| 2026-07-18 | Category-Agnostic, Neural Shape Editing, Coupled Representation | Authors | CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation | arXiv | paper |
| 2026-06-23 | Garment Patterns, Simulation-Ready, Editing | University of Hong Kong | PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments | arXiv | github |
| 2026-05-27 | Part-Controllable, Open-Vocabulary, Game-Ready | Roblox | CubePart: An Open-Vocabulary Part-Controllable 3D Generator | SIGGRAPH 2026 | project / model |
| 2025-09-10 | Part Decomposition, Editable, Production-Ready Assets | Tencent | X-Part: High-Fidelity and Structure-Coherent Shape Decomposition | Tech Report | project |
| 2025-08-14 | Rigging, Animation, Skeleton, Skinning | Nanyang Technological University | Puppeteer: Rig and Animate Your 3D Models | NeurIPS 2025 Spotlight | project / github |
| 2025-06-05 | Part-Level Mesh, Compositional DiT, Single Image | University of Waterloo | PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers | arXiv | project |
| 2024-12-16 | Articulated Mesh, Part-by-Part, Hierarchical Transformer | Cornell | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | arXiv | paper |
| 2023-12-13 | Shape Program, Structure, Editable Assets | MIT | Shape2Program: Learning to Infer Shape Programs from 3D Shapes | arXiv | project |
| 2023-06-29 | Part-Aware, Shape Assembly, 3D Generation | Shanghai AI Lab | Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation | NeurIPS 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | Single Image, Executable Scene Programs, Recursive Construction, Editable 3D | Georgia Institute of Technology | Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs | arXiv | paper |
| 2026-09-05 | Agentic 3D Composition, Functional Objects, Executable Robot Scenes | Peking University / Galbot | GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning | arXiv | paper |
| 2026-09-04 | Image-to-Scene, Agentic Layout Evolution, Simulation-Ready Diversity | The University of Hong Kong | SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution | arXiv | project / github |
| 2026-08-31 | Text-to-Scene, Grow-and-Repair, Functional Groups, SceneReverse-17K | Southeast University | ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation | arXiv | project |
| 2026-08-27 | Single Image, Generative 3D Proxy, RGB-D World Expansion, Explorable Scene | Hong Kong University of Science and Technology | SpatialCrafter: Single Image World Modeling with Generative 3D Proxies | arXiv | project |
| 2026-08-25 | Monocular Image, Interactive Scene Programming, Articulation+Physics, Embodied Simulation | Shanghai Jiao Tong University | NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation | arXiv | project |
| 2026-08-19 | Usage-Driven Code Scenes, Multi-Part Interaction, Executable Simulation | Shanghai Jiao Tong University | Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction | arXiv | paper |
| 2026-07-29 | Panoramic Video, 3DGS, Simulation-ready World | AgiBot | Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation | arXiv | github |
| 2026-07-15 | Indoor Layout, Progressive VLM Reasoning, Interactive Editing | City University of Hong Kong | ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning | arXiv | paper |
| 2026-07-08 | Simulation-Ready Assets, Affordances, Cross-Simulator Worlds | Horizon Robotics | EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI | arXiv | project / github |
| 2026-07-07 | Real-to-Sim, One-Shot Scene Generation, Robot Evaluation | Shanghai AI Lab | RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation | arXiv | project |
| 2026-07-04 | Egocentric Scene Generation, Geometric 3DGS, Consistency | South China Univ. of Technology | CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation | arXiv | paper |
| 2026-06-23 | Triangle Splatting, Single-Image Scene, Game-Ready | Google Research | FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation | arXiv | project |
| 2026-06-23 | Text-to-Scene, Video Priors, 3DGS Orbit | University of Bern | OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis | arXiv | paper |
| 2026-06-23 | Compositional 3D, Physical Interaction, Multi-View Consistency | China University of Petroleum (East China) | Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation | arXiv | paper |
| 2026-06-23 | Satellite-to-City, Textured Mesh, Urban Simulation | HKUST(GZ) | Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image | arXiv | paper |
| 2026-06-08 | Satellite-to-3D, 3DGS, UAV Simulation | Amap-cvlab / Alibaba | ABot-Earth 0.5: Generative 3D Earth Model | arXiv | project |
| 2026-06-04 | Whole-Home Scenes, Floorplans, Interactive | Authors | HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes | arXiv | paper |
| 2026-05-28 | Physical Stability, Single Image, Scene Tree, Simulation | Carnegie Mellon University | REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image | arXiv | project |
| 2026-05-01 | Segment Map, Text-to-World, Controllable 3D Worlds | Seoul National University | Map2World: Segment Map Conditioned Text to 3D World Generation | arXiv | paper |
| 2026-04-14 | Explorable 3D Worlds, Long Trajectory, 3DGS | NVIDIA | Lyra 2.0: Explorable Generative 3D Worlds | arXiv | project |
| 2026-04-06 | Single-Image Scene, In-Place Completion, ARSG-110K | Nankai University | 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image | CVPR 2026 | project / github / dataset |
| 2026-03-31 | Town-Scale, Single Image, Latent Extension | Seoul National University | Extend3D: Town-Scale 3D Generation | CVPR 2026 | project |
| 2026-03-31 | Unbounded World, Flow Matching, Layouts | Princeton University | WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation | arXiv | project |
| 2026-03-27 | Autoregressive 3DGS, Token Generation, Completion/Outpainting | Technical University of Munich | GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | arXiv | project |
| 2026-03-12 | Multi-Floor, Language-to-3D, Long-Horizon Tasks | Tsinghua | MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks | CVPR 2026 | paper |
| 2026-03-06 | Compositional Scene, Panoramic Image, Feed-Forward | NTU | Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic Image | CVPR 2026 | paper |
| 2026-02-10 | Agentic Scene Generation, Sim-Ready, SAGE-10k | NVIDIA | SAGE: Scalable Agentic 3D Scene Generation for Embodied AI | CVPR 2026 | project / github |
| 2026-01-09 | Language-Guided, Infinite Worlds, Articulated Furniture | National Taiwan Univ | SceneFoundry: Generating Interactive Infinite 3D Worlds | arXiv | project |
| 2025-12-01 | Tabletop, Instance-Level, Interactive Scene, Text/Image | D-Robotics | TabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single Image | arXiv | project / paper |
| 2025-11-18 | Single-Image Scene, Gaussian World, Scene Generation | Tsinghua University | GEN3D: Generating Domain-Free 3D Scenes from a Single Image | arXiv | paper |
| 2025-09-18 | Layout-Guided, Indoor Scenes, Decoupled Geometry/Appearance | Hong Kong University of Science and Technology | SPATIALGEN: Layout-guided 3D Indoor Scene Generation | 3DV 2026 | paper |
| 2025-08-21 | Single Image, Multi-Asset Scene, Feed-Forward | Shanghai Jiao Tong University | SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass | 3DV 2026 | project / github |
| 2025-08-11 | Panoramic, Explorable World, Matrix-Pano | Kunlun Wanwei | Matrix-3D: Omnidirectional Explorable 3D World Generation | arXiv | project / github |
| 2025-07-29 | Panoramic, Text/Image-to-World, Mesh Export | Tencent Hunyuan | HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels | arXiv | github |
| 2025-07-09 | VLA Scene Authoring, Simulation-Ready Worlds, Synthetic Data | NVIDIA / Stanford University | 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds | 3DV 2026 | project / github |
| 2025-06-25 | Explorable Scene, Novel-View Restoration, Consistency | Beijing Academy of AI | WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration | arXiv | paper |
| 2025-03-13 | Scene Layout, Optimization, Generation | Tsinghua | HOG-Layout: Layout-Enhanced Scene Generation via Hierarchical Optimization | arXiv | project |
| 2025-02-20 | Octree, 3D Diffusion, Scene Generation | Zhejiang University | Octree Diffusion: Hierarchical Scene Generation via Octree Structures | arXiv | project |
| 2025-01-15 | Gaussian, GPT, Scene Generation | Shanghai AI Lab | GaussianGPT: Language-Driven Scene Generation with Gaussian Representation | arXiv | paper |
| 2024-12-19 | Scene Generation, Growing, Incremental | Tsinghua | WorldGrow: Incremental 3D Scene Generation | CVPR 2025 | project |
| 2024-11-05 | Splatting, Fluents, Scene Understanding | University of Cambridge | FluSplat: Fluent Scene Generation via Gaussian Splatting | CVPR 2025 | project |
| 2024-11-04 | GenXD, Any 3D and 4D, Scene Generation | NUS | GenXD: Generating Any 3D and 4D Scenes | ICLR 2025 | paper |
| 2024-06-17 | Procedural Scenes, Synthetic Data, Embodied AI | Princeton | Infinigen Indoors: Photorealistic Indoor Scenes using Procedural Generation | NeurIPS 2024 | project / github |
| 2024-06-13 | Interactive 3D Scene, FLAGS, Single Image | Stanford University | WonderWorld: Interactive 3D Scene Generation from a Single Image | CVPR 2025 | project |
| 2024-05-02 | EchoScene, Scene Graph, Diffusion, Indoor | TUM / JHU | EchoScene: Indoor Scene Generation via Information Echo | ECCV 2024 | paper |
| 2024-02-12 | SceneScape, Text-Driven, Consistent, Scene | Weizmann | SceneScape: Text-Driven Consistent Scene Generation | NeurIPS 2024 | paper |
| 2024-01-30 | BlockFusion, Expandable, Tri-plane, SIGGRAPH | Tencent / UTokyo | BlockFusion: Expandable 3D Scene Generation | SIGGRAPH 2024 (ACM TOG) | paper |
| 2023-12-01 | ControlRoom3D, Semantic Proxy, Room Generation | TUM / Meta | ControlRoom3D: Room Generation using Semantic Proxy Rooms | CVPR 2024 | paper |
| 2023-10-05 | Ctrl-Room, Text-to-3D, Layout Constraints | Simon Fraser | Ctrl-Room: Controllable Text-to-3D Room Meshes Generation | ECCV 2024 | paper |
| 2023-10-04 | MagicDrive, Street View, 3D Geometry Control | CUHK / HKUST | MagicDrive: Street View Generation with Diverse 3D Geometry Control | ICLR 2024 | project |
| 2023-06-15 | Procedural World, Synthetic Data, Simulation | Princeton | Infinite Photorealistic Worlds using Procedural Generation | CVPR 2023 | project |
| 2023-03-24 | DiffuScene, Diffusion, Indoor Scene Synthesis | TUM | DiffuScene: Denoising Diffusion for Generative Indoor Scene Synthesis | CVPR 2024 | paper |
| 2023-03-21 | Text-to-3D Room, Indoor Scenes, Mesh | LMU Munich | Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models | ICCV 2023 | project |
| 2023-02-02 | Unbounded 3D Scene, Generative Model, Driving | NVIDIA | SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections | CVPR 2023 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-18 | Multi-View Generation, Scene Assets, Training-Free | The University of Queensland | Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning | arXiv | github |
| 2025-01-03 | Rover, Semantic, 3D Scene | Carnegie Mellon University | SEM-ROVER: Semantic Scene Exploration with Hierarchical Spatial Reasoning | ICLR 2025 | project |
| 2024-09-30 | Spatial, Generation, Language | Tsinghua | SpatialGen: Language-Driven Spatial Scene Generation | NeurIPS 2024 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-11 | Layout-Conditioned, 3DGS, Mixed Reality | Authors | SyncSpace: Layout-Conditioned 3D Gaussian Splatting for Space Reskinning in Mixed Reality | arXiv | paper |
| 2025-04-17 | Image-Pair-to-4D, Diffusion, Explicit 3D Motion | Technical University of Munich | TwoSquared: 4D Generation from 2D Image Pairs | 3DV 2026 Oral | paper |
| 2024-12-05 | 4Real-Video, Photo-Realistic, Video Diffusion, CVPR | Snap / KAUST | 4Real-Video: Generalizable Photo-Realistic 4D Video Diffusion | CVPR 2025 | paper |
| 2024-07-16 | Animate3D, Multi-View, Video Diffusion, NeurIPS | CASIA / Alibaba | Animate3D: Animating Any 3D Model with Multi-view Video Diffusion | NeurIPS 2024 | paper |
| 2024-05-31 | 4Diffusion, Multi-View Video, 4D, NeurIPS | CASIA / Shanghai AI Lab | 4Diffusion: Multi-view Video Diffusion Model for 4D Generation | NeurIPS 2024 | paper |
| 2024-05-26 | Diffusion4D, Video Diffusion, 4D, NeurIPS | Toronto / BJTU | Diffusion4D: Fast Spatial-temporal Consistent 4D Generation | NeurIPS 2024 | paper |
| 2024-05-03 | DreamScene4D, Multi-Object, Dynamic, NeurIPS | CMU | DreamScene4D: Dynamic Multi-Object Scene Generation | NeurIPS 2024 | paper |
| 2024-03-22 | STAG4D, Spatial-Temporal, 4D Gaussians, ECCV | Nanjing / CASIA | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians | ECCV 2024 | paper |
| 2023-11-29 | 4D-fy, Text-to-4D, Score Distillation, CVPR | KAUST / Snap | 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling | CVPR 2024 | project |
| 2023-11-24 | Animate124, Image-to-4D, Animation, ICLR | NUS / Huawei | Animate124: Animating One Image to 4D Dynamic Scene | ICLR 2024 | paper |
| 2023-11-17 | Consistent4D, 360° Dynamic, Monocular Video, ICLR | CASIA / Nanjing | Consistent4D: Consistent 360° Dynamic Object Generation | ICLR 2024 | project |
3D editing covers methods that modify existing 3D assets, Gaussian / NeRF fields, meshes, voxel or latent states, and dynamic scenes. Entries are grouped by editable state: object-level, scene-level, and dynamic / 4D.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-27 | NeRF Editing, Object Removal, Robot Manipulation | Authors | NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation | arXiv | paper |
| 2026-06-05 | Mesh Editing, Image-Guided, Local Morphing | Leiden University | 3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing | IJCNN 2026 | github |
| 2026-05-26 | PartFlow, Semantic-Part Transformation, Mask-Free | Nanyang Technological University | Feedforward 3D Editing Learns from Semantic-Part Transformation | arXiv | project / github / benchmark (Steer3D) |
| 2026-05-08 | VS3D, Velocity-Space, Mask-Free | Tsinghua University | Velocity-Space 3D Asset Editing | arXiv | paper |
| 2026-05-01 | Latent Editing, Object-Level, Structured 3D Latents | Seoul National University | InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization | arXiv | project |
| 2026-04-30 | MeshReGen, VecSet Regeneration, Image-Guided Editing | KAIST | MeshReGen: A Unified 3D Geometry Regeneration Framework | arXiv | project / benchmark (VoxHammer) |
| 2026-04-26 | Primitive Proxy, Shape Editing, Fine-Grained Control | Tel Aviv University | Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions | SIGGRAPH 2026 | project / github / benchmark (VoxHammer) |
| 2026-03-30 | 3DGS Editing, Object-Level, Single-View | Zhejiang Gongshang University | SVGS: Single-View to 3D Object Editing via Gaussian Splatting | ACM TOMM 2026 | project |
| 2026-02-25 | Voxel Editing, Object-Level, Rectified Voxel Flow | USTC | Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow | CVPR 2026 | project |
| 2026-02-05 | Native 3D Editing, Image-Conditioned, Latent-to-Latent | Aigency.ai / Tel Aviv University | ShapeUP: Scalable Image-Conditioned 3D Editing | SIGGRAPH 2026 | project / github |
| 2026-02-04 | Mesh Editing, Object-Level, Single-Image LRM | National Yang Ming Chiao Tung University | VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image | arXiv | github / benchmark (VoxHammer) |
| 2025-12-15 | Steer3D, Text-Steerable, Edit3D-Bench (Ma) | Caltech | Feedforward 3D Editing via Text-Steerable Image-to-3D | arXiv | project / github / benchmark (Steer3D) |
| 2025-11-27 | Latent Anchor, Object-Level, Mask-Free | Zhejiang University | AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows | CVPR 2026 Oral | project / github / benchmark |
| 2025-11-21 | Native Editing, Object-Level, Full Attention | Fudan University / StepFun | Native 3D Editing with Full Attention | arXiv | paper |
| 2025-10-16 | FlowEdit, Object-Level, Mask-Free | Tsinghua University | NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks | ICLR 2026 | project / github / dataset |
| 2025-10-03 | 3DEditFormer, Paired Dataset, Mask-Free | East China Univ of Science & Technology | Towards Scalable and Consistent 3D Editing | arXiv | project / github |
| 2025-08-29 | 3D-LATTE, Text Instructions, 3D Diffusion Latent | University of Tübingen | 3D-LATTE: Latent Space 3D Editing from Textual Instructions | CVPR 2026 Oral | project / CVF |
| 2025-08-26 | VoxHammer, Edit3D-Bench (Li), Training-Free | Renmin University / Beihang University | VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space | 3DV 2026 Oral | project / github / benchmark (VoxHammer) |
| 2025-07-15 | 3DGS Editing, Part-Level, Regularized SDS | Seoul National University | Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling | ICCV 2025 | CVF |
| 2025-06-25 | Image Prompt, Multi-View Propagation, Mask-Free | Tel Aviv University | EditP23: 3D Editing via Propagation of Image Prompts to Multi-View | ACM TOG 2025 | project / github |
| 2025-05-11 | CMD, Local Editing, Progressive Generation | HKUST | CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation | SIGGRAPH 2025 | project |
| 2024-12-11 | Mesh Editing, Object-Level, Masked LR |
Truncated — view the full README on GitHub.
98 commits
53 commits
Python
100.0%
A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.
Python
13
151 commits
updated Sep 23, 2026
A curated map of 3D/4D perception, reconstruction, generation, simulation-ready assets, and world models for embodied AI.
Awesome-Embodied-3DV is a curated list for the research space where 3D vision, 3D/4D reconstruction, 3D generation, simulation-ready assets, and embodied world models meet.
This repository focuses on:
This list is intentionally embodied-3DV-first. It includes 3D generation and 3DGS work only when it helps understand, build, evaluate, or deploy 3D assets and world models for embodied agents. It is not a generic catalog of all 3D generation, editing, rendering, compression, or graphics papers.
Daily candidate feed. The automatically updated arXiv Daily is a high-recall, topic-tagged candidate archive across the six areas above. It is deliberately broader than this curated README: papers are promoted here only after manual primary-source verification.
Start here if you want the shortest path through the field.
Data perception covers the sensor-facing and semantic layers: extracting geometric priors (depth, normals), understanding 3D semantics (detection, segmentation, grounding), active-imaging signals, and dense maps from 2D images, video, or physical sensors.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | Metric Depth, 3DGS Relocalization, Sparse PnP Anchors, Temporal Memory | The Chinese University of Hong Kong, Shenzhen | RIDE: Relocalization-Informed Depth Estimation with 3D Gaussian Splatting | arXiv | paper |
| 2026-09-08 | Diffusion Transformer, Single-Step Depth, Sharp Details, Dense Prediction | EPFL / HUAWEI Bayer Lab | Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation | SIGGRAPH Asia 2026 | project / github |
| 2026-09-08 | Any Camera, Metric Point Cloud, Optional Intrinsics/Sparse Depth | Google DeepMind | OmniPoint: Universal Monocular Metric Pointcloud from Any Camera | ECCV 2026 | project |
| 2026-08-30 | Transparent/Reflective Scenes, Bias-Aware Training, 30M Parameters | The University of Hong Kong | OptiGeo: Efficient Monocular Geometry for Embodied Perception in Optically Challenging Scenes | arXiv | project |
| 2026-08-17 | Pixel-Space Prediction, Fine Structures, Sharp Boundaries, Efficient Depth | The Chinese University of Hong Kong, Shenzhen | PXDepth: Pixel-Space Modeling for Structure Preserving Monocular Depth Estimation | arXiv | project |
| 2026-08-03 | Geometry-Invariant Adaptation, Non-Lambertian Surfaces, Mirror/Glass Depth | Changchun Institute of Optics, Fine Mechanics and Physics, CAS | GIFT: Geometry-Invariant Fine-Tuning for Non-Lambertian Monocular Depth Estimation | arXiv | paper |
| 2026-07-23 | UAV Depth, Arbitrary Camera Pose, Metric Geometry | Aerospace Information Research Institute, CAS | DAPM: UAV Monocular Depth Estimation from Any Height, Pitch, Roll and FOV | arXiv | github |
| 2026-07-20 | Fine-Detail Geometry, Sparse Volumetric Refinement, Metric Scale | Tsinghua University | MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric Refinement | arXiv | project |
| 2026-07-19 | Lightweight Foundation Depth, Camera-Conditioned Metric Depth, Edge Deployment | University of Trento | DepthART: Scaling Foundation Monocular Depth to Tiny Models | ACM Multimedia 2026 | project / github |
| 2026-07-19 | Metric Depth, Odometry Anchor, Recurrent SLAM | UC Berkeley | DROID-ANCHOR: Odometry-Anchored Recurrent Metric Depth Estimation | arXiv | paper |
| 2026-07-17 | Stereo Distillation, Epipolar Cues, Metric Depth | Michigan State University | Geometric Distillation from Rectified Stereo: Leveraging Epipolar Cues for Monocular Depth | arXiv | paper |
| 2026-07-14 | Auto-Regressive Depth, Coarse-to-Fine, Semantic Guidance | Sun Yat-sen University | ARDepth: Auto-regressive Monocular Depth Estimation with Progressive Visual Conditioning | arXiv | paper |
| 2026-07-13 | Metric Point Map, Pixel-Wise Calibration, Camera Diversity | The University of Hong Kong | FoundationGeo: Learning Spatial Pixel-Wise Fields for Monocular Metric Geometry | ECCV 2026 | project / github / model |
| 2026-07-09 | Lightweight Zero-Shot, 6.1M Parameters, On-Device Depth | University of Bologna | ZipDepth: Bringing Lightweight Zero-Shot Monocular Depth Anywhere, on Any Device | ECCV 2026 | project / github |
| 2026-05-27 | Multi-Layer Depth, Transparent Surfaces, Point Process | Princeton University | SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping | CVPR 2026 | github |
| 2026-05-15 | VLM, Dense Metric Depth, Spatial Reasoning | Zhejiang Univ | Unlocking Dense Metric Depth Estimation in VLMs | arXiv | project / github |
| 2026-05-12 | Sparse 3D Anchors, Relative-to-Metric, Graph Optimization | Tongji University | The Midas Touch for Metric Depth | CVPR 2026 Highlight | project |
| 2026-03-28 | Universal Camera, Metric Depth, Zero-Shot | Michigan State University | UniDAC: Universal Metric Depth Estimation for Any Camera | CVPR 2026 | paper |
| 2026-03-20 | Transparent Objects, Generative Opacification, Monocular Depth, SeeClear-396k | University of California, Los Angeles | SeeClear: Reliable Transparent Object Depth Estimation via Generative Opacification | ECCV 2026 | project / dataset |
| 2026-03-17 | Diffusion Prior, Real-World Data, Monocular Depth | Nanjing University of Science and Technology | Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-03-04 | Fine-Grained Geometry, Dual-Stream, Efficient | UMass Amherst | DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation | CVPR 2026 | paper |
| 2026-01-29 | Sparse Metric Prompt, 20M Image-Depth Pairs, Metric Foundation Model | Li Auto Inc | MetricAnything: Scaling Metric Depth Pretraining with Noisy Heterogeneous Sources | ECCV 2026 | project / github |
| 2026-01-06 | Arbitrary-Resolution, Neural Implicit, Fine Details | Zhejiang Univ | InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields | CVPR 2026 | project / github |
| 2025-12-13 | Defocus Cue, Bokeh Stack, Metric Depth | Nanyang Tech Univ | Boosting Monocular Metric Depth Estimation via Bokeh Rendering | ICML 2026 | project / github |
| 2025-11-30 | Deterministic Diffusion, Dense Geometry, Fine Details | HKUST(GZ) | Lotus-2: Advancing Geometric Dense Prediction with Powerful Image Generative Model | arXiv | project / github |
| 2025-11-13 | Any-View Depth, Metric Geometry, Multi-View | ByteDance | Depth Anything 3: Recovering the Visual Space from Any Views | ICLR 2026 | project / github |
| 2025-10-27 | Unified Generation+Depth, Diffusion Prior, Zero-Shot | HUST | More Than Generation: Unifying Generation and Depth Estimation via Text-to-Image Diffusion Models | NeurIPS 2025 | github |
| 2025-10-08 | Pixel-Space Diffusion, Flying-Pixel-Free, Point Clouds | HUST | Pixel-Perfect Depth with Semantics-Prompted Diffusion Transformers | NeurIPS 2025 | project / github |
| 2025-09-29 | VLM, Metric Depth, Sparse Supervision | Meta AI | DepthLM: Metric Depth From Vision Language Models | ICLR 2026 Oral | github |
| 2025-07-03 | Monocular Geometry, Metric Scale, Sharp Details | USTC / Microsoft | MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp Details | NeurIPS 2025 | project / github |
| 2025-04-16 | Sliding Anchor, Unknown Intrinsics, Metric Scale | Shanghai Univ | Metric-Solver: Sliding Anchored Metric Depth Estimation from a Single Image | arXiv | project / github |
| 2025-02-27 | Metric 3D Points, Self-Prompt Camera, Uncertainty | ETH Zurich | UniDepthV2: Universal Monocular Metric Depth Estimation Made Simpler | TPAMI 2026 | github |
| 2025-02-26 | Cross-Context Distillation, Multi-Teacher, Fine-Detail Depth | Zhejiang Univ of Technology | Distill Any Depth: Distillation Creates a Stronger Monocular Depth Estimator | arXiv | project / github |
| 2024-11-27 | Diffusion Distillation, Metric + Sharp, Boundary Detail | Qualcomm AI Research | SharpDepth: Sharpening Metric Depth Predictions Using Diffusion Distillation | CVPR 2025 | project / github |
| 2024-10-02 | Metric Depth, Zero-Shot, Image Priors | Apple | Depth Pro: Sharp Monocular Metric Depth in Less Than a Second | ICLR 2025 | github |
| 2024-09-26 | Diffusion, Single-Step, Dense Geometry | HKUST(GZ) | Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction | ICLR 2025 | project / github |
| 2024-06-13 | Monocular Depth, Foundation Model, Metric DAv2 | TikTok / HKU | Depth Anything V2 | NeurIPS 2024 | project / github / metric models |
| 2024-03-27 | Metric Depth, Universal, Zero-Shot | ETH Zurich | UniDepth: Universal Monocular Metric Depth Estimation | CVPR 2024 | github |
| 2024-01-19 | Monocular Depth, Relative Depth, Foundation Model | TikTok / HKU | Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data | CVPR 2024 | project |
| 2023-12-04 | Monocular Depth, Zero-Shot, Affine-Invariant | Intel Labs | Marigold: Repurposing Diffusion-Based Image Generators for Monocular Depth Estimation | CVPR 2024 | project |
| 2023-07-20 | Metric 3D, Zero-Shot, Canonical Space | Alibaba | Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image | ICCV 2023 | github |
| 2021-03-24 | ViT, Dense Prediction, Foundation, Depth | Intel Labs | DPT: Vision Transformers for Dense Prediction | ICCV 2021 | github |
| 2019-07-02 | Robust Depth, Zero-Shot, Cross-Dataset, MiDaS | Intel Labs | MiDaS: Towards Robust Monocular Depth Estimation | TPAMI 2022 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | Unified Video Model, Depth+Normals, Temporal Consistency | Adobe Research | Unified Video Dense Prediction from Disjoint Data | arXiv | project |
| 2026-07-02 | Video Diffusion, In-Context Conditioning, Zero-Shot | HKUST | ICDepth: Taming Video Diffusion Models for Video Depth Estimation via In-Context Conditioning | ECCV 2026 | project |
| 2026-05-28 | Streaming Geometry, Dynamic Chunking, Depth+Normals | Zhejiang University | Towards Consistent Video Geometry Estimation | arXiv | project |
| 2026-05-11 | Camera Motion, 3D Consistency, Geometry Embedding | HUST | GemDepth: Geometry-Embedded Features for 3D-Consistent Video Depth | arXiv | github |
| 2026-04-08 | Post-Processing, Scalable, Single-Image Backbone | Seoul National University | VDPP: Video Depth Post-Processing for Speed and Scalability | CVPR 2026 ECV Workshop | paper |
| 2026-04-02 | Pose Refinement, Temporal Consistency, Monocular Video | Ajou University | PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency | CVPR 2026 | paper |
| 2026-03-12 | Deterministic Diffusion, Generative Prior, Long Video | HKUST(GZ) | DVD: Deterministic Video Depth Estimation with Generative Priors | arXiv | project |
| 2026-01-06 | Temporal Stability, Monocular Video, Long Sequence | ETH Zurich | StableDPT: Temporal Stable Monocular Video Depth Estimation | arXiv | paper |
| 2025-12-20 | Endoscopic Geometry, Streaming Mamba, Metric Depth | Vanderbilt University | EndoStreamDepth: Temporally Consistent Monocular Depth Estimation for Endoscopic Video Streams | arXiv | github / paper |
| 2025-12-11 | Sparse Keyframes, Propagation, Long-Video Consistency | ETH Zurich | Video Depth Propagation | 3DV 2026 | paper |
| 2025-10-10 | Online Inference, Low Memory, Temporal Consistency | Heidelberg University | Online Video Depth Anything: Temporally-Consistent Depth Prediction with Low Memory Consumption | arXiv | paper |
| 2025-07-02 | Diffusion Guidance, Scale Synchronization, Geometry Consistency | Tsinghua University | DepthSync: Diffusion Guidance-Based Depth Synchronization for Scale- and Geometry-Consistent Video Depth Estimation | ICCV 2025 | project |
| 2025-04-09 | 2K Streaming, Mamba, 24 FPS | Netflix Eyeline Studios | FlashDepth: Real-time Streaming Video Depth Estimation at 2K Resolution | ICCV 2025 Highlight | project / github |
| 2025-01-21 | Super-Long Video, Temporal Gradient, 30 FPS | ByteDance | Video Depth Anything: Consistent Depth Estimation for Super-Long Videos | CVPR 2025 | project / github |
| 2024-11-28 | Long Video, Diffusion, Multi-Resolution Alignment | ETH Zurich | Video Depth without Video Models | CVPR 2025 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-06 | Surface Normals, Transparent Objects, Rectified Flow, Edge Refinement | Zhejiang University | TransNormal-2: Geometry-Grounded Rectified Flow with Edge-Aware Decoding for Precise Normal Estimation | arXiv | project |
| 2026-07-28 | Stereo Depth, Walsh-Hadamard Mixing, Efficient Inference | Authors | WHTMix: Efficient Stereo Depth Estimation via Walsh-Hadamard Token Mixing | arXiv | paper |
| 2026-07-22 | Stereo Diffusion Transformer, Flow Matching, Progressive Refinement | Beihang University | STEREOFLOW: Progressive Stereo Matching with StereoDiT and Transition Flow Matching | arXiv | paper |
| 2026-07-15 | Vision Features, SE(3) Latent Geometry, Visual Navigation | Google DeepMind | SeeSE3: Emergence of 3D Space in Vision Features | arXiv | paper |
| 2026-07-14 | Heterogeneous Cameras, Metric Depth, Real-Time | D-Robotics | X-Lens: Real-Time Metric Depth Estimation with Heterogeneous Cameras | arXiv | project / github |
| 2026-07-10 | Video Generative Pretraining, Depth/Normals/Pose, Grounded 4D | Google DeepMind | Video Generation Models are General-Purpose Vision Learners | ECCV 2026 | project |
| 2026-07-06 | Boundary-Centric Pretraining, Dense Spatial Perception | Robbyant | LingBot-Vision: Vision Pretraining for Dense Spatial Perception | arXiv | project / github |
| 2026-03-02 | Zero-Shot Stereo, Structure Prompt, Motion Prompt | HUST | PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts | CVPR 2026 | paper |
| 2025-12-11 | Real-Time Stereo, Zero-Shot, Foundation Model | NVIDIA | Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching | CVPR 2026 | paper |
| 2025-07-22 | Foundation Model, Depth/Normal/Pointmap, Multi-View | SJTU | Dens3R: A Foundation Model for 3D Geometry Prediction | ICLR 2026 | paper |
| 2025-04-15 | Surface Normals, Foundation Model, Video, Temporal | Authors | NormalCrafter: Learning Temporally Consistent Normals from Video Diffusion Priors | ICCV 2025 | project / paper |
| 2025-03-21 | Scalable Depth, Autoregressive, 2B Parameters | Baidu | DAR: Scalable Autoregressive Monocular Depth Estimation | CVPR 2025 | project |
| 2025-03-20 | Universal Camera, Spherical 3D, Any Camera | ETH Zurich | UniK3D: Universal Camera Monocular 3D Estimation | CVPR 2025 | github / paper |
| 2025-03-11 | LiDAR Surface Normal, Dataset, Point Cloud | TU Graz | LiSu: A Dataset and Method for LiDAR Surface Normal Estimation | CVPR 2025 | github |
| 2025-01-17 | Stereo, Foundation Model, RAFT-Style, Zero-Shot | NVIDIA | FoundationStereo: Zero-Shot Stereo Matching | CVPR 2025 Oral / Best Paper Nomination | github |
| 2025-01-17 | Stereo, Robust, Zero-Shot, Non-Lambertian | Univ of Bologna | Stereo Anywhere: Robust Zero-Shot Deep Stereo Matching | CVPR 2025 | project |
| 2024-12-11 | Panoramic / Fisheye Depth, Zero-Shot Metric | Intel | Depth Any Camera: Zero-Shot Metric Depth from Any Camera | CVPR 2025 | project |
| 2024-10-24 | Monocular Geometry, Pointmap, Affine-Invariant | USTC / Microsoft | MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images | CVPR 2025 Oral | github |
| 2024-06-24 | Normal Estimation, 3D Priors, Surface Geometry | Nvidia | StableNormal: Reducing Diffusion Variance for Stable and Sharp Normal | SIGGRAPH Asia 2024 | project |
| 2024-03-27 | Diffusion, Effective Conditioning, ViT Priors | IIT Delhi | ECoDepth: Effective Conditioning of Diffusion Models for Monocular Depth | CVPR 2024 | github |
| 2024-03-22 | Metric 3D, Multi-Task, Geometry Foundation | Shanghai AI Lab | Metric3D v2: A Versatile Monocular Geometric Foundation Model | TPAMI 2025 | project |
| 2024-03-18 | Geometry, Normals, Depth, Multi-Task | Apple | GeoWizard: Unleashing the Diffusion Priors for 3D Geometry Estimation from a Single Image | ECCV 2024 | project |
| 2024-03-01 | Inductive Biases, Surface Normal, Oral | Imperial College London | DSINE: Rethinking Inductive Biases for Surface Normal Estimation | CVPR 2024 Oral | github |
| 2023-12-04 | High-Res, Patch-Wise, Model-Agnostic | KAUST | PatchFusion: An End-to-End Tile-Based Framework for High-Resolution Monocular Metric Depth | CVPR 2024 | project |
| 2023-09-25 | Iterative Bins, Elastic, GRU, Classification-Regression | Beihang Univ | IEBins: Iterative Elastic Bins for Monocular Depth Estimation | NeurIPS 2023 | github |
| 2023-04-13 | Internal Discretization, Continuous-Discrete, Depth | ETH Zurich | iDisc: Internal Discretization for Monocular Depth Estimation | CVPR 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Annotation-Free, Open-Vocabulary 3D, Language-Space Lifting | National Technical University of Athens | GoDeep: Annotation-Free Open-Vocabulary 3D Scene Understanding via Language-Space Lifting | arXiv | paper |
| 2026-07-21 | Referring Segmentation, 3DGS, Generalized Grounding | Peking University | ZeroSplat: Generalized Referring Segmentation in 3D Gaussian Splatting | arXiv | project |
| 2026-06-23 | Open-Vocabulary BEV, 3DGS, Geometric Constraints | KAIST AI | Open-Vocabulary BEV Segmentation with 3D-Aware Geometric Constraints | ECCV 2026 | paper |
| 2026-06-04 | Open-Vocabulary, Functionality Segmentation, Robotics | Authors | T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation | arXiv | paper |
| 2026-05-07 | Open-Vocabulary, Gaussian Feature Field, Codebook | TU Munich / Google | OpenGaFF: Open-Vocabulary Gaussian Feature Field with Codebook Attention | arXiv | paper |
| 2025-11-20 | Open-Vocabulary, SAM3, Promptable, DETR | Meta | SAM 3: Segment Anything with Concepts | ICLR 2026 | github / paper |
| 2025-04-03 | Open-Vocabulary, Dual-Level Contrastive, Instance-Aware | BIGAI / Tsinghua | MPEC: Masked Point-Entity Contrast for Open-Vocabulary 3D Scene Understanding | CVPR 2025 | project |
| 2025-03-22 | Training-Free, MLLM Caption, Voxel Grouping | NVIDIA Research Taiwan | OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding | CVPR 2026 | project |
| 2025-03-19 | SAM-2, 3D Tracking, Dynamic Programming | VinAI | Any3DIS: Open-Vocabulary 3D Instance Segmentation with SAM-2 | CVPR 2025 | paper |
| 2025-03-13 | Open-Vocabulary, LLM Canonical, Part Segmentation | Shandong Univ / Tencent | CoSMo3D: Open-World Promptable 3D Semantic Part Segmentation via LLM-Guided Canonical Spatial Modeling | CVPR 2026 Oral | github |
| 2025-03-12 | Functional 3D, CoT, VLM, Training-Free, Highlight | Authors | Fun3DU: Functional 3D Scene Understanding via Chain-of-Thought | CVPR 2025 Highlight | paper |
| 2025-01-02 | Panoptic, Open-Vocabulary, 3D Gaussian Splatting | NUS | PanoGS: Gaussian-based Panoptic Open-Vocabulary 3D Scene Understanding | CVPR 2025 | project |
| 2024-12-13 | Open-Vocabulary 3D, Structured Super-Gaussians, Segmentation | SuperGSeg: Open-Vocabulary 3D Segmentation with Structured Super-Gaussians | 3DV 2026 | paper | |
| 2024-12-12 | Open-Vocabulary, Foundation Dataset, Mask-Text Pairs | NVIDIA | Mosaic3D: Foundation Dataset and Model for Open-Vocabulary 3D Segmentation | CVPR 2025 | github |
| 2024-07-02 | Open-Vocabulary, Mask-Snap-Lookup, Indoor/Outdoor | HKUST | OpenIns3D: Open-Vocabulary 3D Segmentation with Mask-Snap-Lookup | ECCV 2024 | github |
| 2024-04-01 | Region-Level, Point-Language Contrastive, Multi-VLM | HKU | RegionPLC: Regional Point-Language Contrastive Learning for Open-World 3D Scene Understanding | CVPR 2024 | project / github |
| 2024-03-19 | Open-Vocabulary, MLLM, Point-Entity-Text, nuScenes | MPI | OV3D: Open-Vocabulary 3D Understanding with Multi-Modal Alignment | CVPR 2024 | paper |
| 2024-03-19 | 2D-Guided, 3D Proposals, SAM, Instance | VinAI / IBM | Open3DIS: Open-Vocabulary 3D Instance Segmentation with 2D-Guided Mask Generation | CVPR 2024 | github |
| 2024-03-15 | Training-Free, View-Consensus, Instance | PKU | MaskClustering: View-Consensus for 3D Instance Segmentation | CVPR 2024 | paper |
| 2024-03-11 | Zero-Shot, Superpoint, SAM, Scene Graph | PKU | SAI3D: Zero-Shot 3D Instance Segmentation by Scene-Aware Incremental Merging | CVPR 2024 | github |
| 2024-01-31 | Segment Anything, 3D Gaussians, Interactive | Zhejiang University | SAGD: Boundary-Enhanced Segment Anything in 3D Gaussians | arXiv | github |
| 2024-01-04 | 2D+3D Unified, Single Model, Highlight | Authors | ODIN: A Single Model for 2D and 3D Segmentation | CVPR 2024 Highlight | github |
| 2023-12-26 | Open-Vocabulary 3DGS, Segmentation, Language | ETH Zurich | LangSplat: 3D Language Gaussian Splatting | CVPR 2024 Highlight | project / github |
| 2023-11-17 | Unified, Semantic+Instance+Panoptic, Single Transformer | Samsung | OneFormer3D: One Transformer for Unified 3D Segmentation | CVPR 2024 | github |
| 2023-06-23 | Open-Vocabulary 3D Instance, Mask Proposals, CLIP | ETH Zurich | OpenMask3D: Open-Vocabulary 3D Instance Segmentation | NeurIPS 2023 | project |
| 2023-06-06 | Segment Anything 3D, Point Cloud, Interactive | VAST AI | SAM3D: Segment Anything in 3D Scenes | CVPR 2024 | github |
| 2023-03-16 | Open-Vocabulary, 3D Scene, Language Field | UC Berkeley | LERF: Language Embedded Radiance Fields | ICCV 2023 | project / github |
| 2022-11-28 | Open-Vocabulary 3D, Point Cloud, Segmentation | ETH Zurich | OpenScene: 3D Scene Understanding with Open Vocabularies | CVPR 2023 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-08 | Feed-Forward, Open-Vocabulary, Panoptic | Authors | EPS3D: End-to-End Feed-Forward 3D Panoptic Segmentation | ICML 2026 | github |
| 2025-07 | Feed-Forward, Panoramic, DUSt3R-Based | Authors | PanSt3R: Single-Feed 3D Geometry and Panoptic Segmentation | ICCV 2025 | paper |
| 2025-05-14 | MoE, Multi-Dataset, PTv3, CLIP Alignment | UVA | Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation | ICLR 2026 | project |
| 2025-03-01 | Bayesian 3DGS, Training-Free, EIG | Sony | B3-Seg: Camera-Free, Training-Free 3DGS Segmentation via Analytic EIG | CVPR 2026 | project |
| 2025-01-14 | End-to-End, 2D-to-3D Lifting, 3DGS | CUHK | Unified-Lift: Rethinking End-to-End 2D to 3D Scene Segmentation in Gaussian Splatting | CVPR 2025 | github |
| 2025-01-10 | Unsupervised Panoptic, Scene-Centric | TU Darmstadt | CUPS: Scene-Centric Unsupervised Panoptic Segmentation | CVPR 2025 Highlight | github |
| 2025-01-06 | Zero-Shot Instance, SAM Prompts in 3D | CUHK-SZ / MSRA | SAMPro3D: Locating SAM Prompts in 3D for Zero-Shot Instance Segmentation | 3DV 2025 | github |
| 2024-07-03 | 3D Panoptic, LiDAR, Multi-Scene | Tsinghua | UniSeg3D: Unified 3D Panoptic Segmentation | CVPR 2023 | github |
| 2024-03-25 | Unsupervised 3D Instance, Indoor | Authors | UnScene3D: Unsupervised 3D Instance Segmentation | CVPR 2024 | paper |
| 2024-03-24 | Unified 6 Tasks, Single Transformer | HUST | UniSeg3D: Unified 3D Segmentation with Transformer | NeurIPS 2024 | paper |
| 2023-12-15 | Point Cloud, Foundation Model, 3D Understanding | Shanghai AI Lab | Point Transformer V3: Simpler, Faster, Stronger | CVPR 2024 | github |
| 2022-10-06 | 3D Instance Segmentation, Transformer, Point Cloud | ETH Zurich | Mask3D: Mask Transformer for 3D Semantic Instance Segmentation | ICRA 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | 3D-Aware VLM, Implicit+Explicit Geometry, RGB Video | Nanyang Technological University | 3D-Aware VLMs with Implicit and Explicit Geometries | ECCV 2026 | github |
| 2026-07-14 | 3D Language Fields, Ambiguity Awareness, Object Retrieval | Authors | SaaF: Scene-Specific Ambiguity-Aware 3D Language Fields towards Interactive Real-World Object Retrieval | arXiv | paper |
| 2026-06-23 | Agentic, Cognitive Map, Zero-Shot 3D | Sichuan University | Agentic Collaborative Cognition for Zero-Shot 3D Understanding | ECCV 2026 | project |
| 2026-06-22 | Map-Grounded, MV3D-VQA, Dense Reward | KAIST | Dense Reward for Multi-View 3D Reasoning with Global Maps and Local Views | ECCV 2026 | paper |
| 2026-06-17 | Panoramic Reprojection, 3D VLM, Spatial Reasoning | Technical University of Munich | OneCanvas: 3D Scene Understanding via Panoramic Reprojection | arXiv | project |
| 2026-06-04 | Part-Aware 3D-MLLM, Scene Understanding, Grounding | Xiamen University | PAR3D: A Unified 3D-MLLM with Part-Aware Representation for Scene Understanding | arXiv | project |
| 2025-10-19 | Grounded CoT, 3D Reasoning, SceneCOT-185K Dataset | BIGAI / PKU / Tsinghua | SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes | ICLR 2026 | paper |
| 2025-07 | Scene Graph, LLM, Relation Encoding | AIRI | 3DGraphLLM: 3D Scene Graph Learning with LLMs | ICCV 2025 | paper |
| 2025-03-19 | Gesture+Language, Embodied Reference, +30% | Authors | Ges3ViG: 3D Embodied Reference Understanding with Pointing Gestures | CVPR 2025 | paper |
| 2025-03-19 | Geometry VLM, Unified 3D Recon + Spatial Reasoning | Shanghai AI Lab / UCLA / SJTU | G2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning | CVPR 2026 | github |
| 2025-03-12 | Dense Grounding, 6.2M Pairs, Hallucination Benchmark | UMich | 3D-GRAND: A Million-Scale Dataset for 3D Grounding | CVPR 2025 | project |
| 2025-03-11 | 3D-Informed, Spatial Reasoning, Highlight | JHU | SpatialLLM: Spatial Reasoning with 3D-Informed LLMs | CVPR 2025 Highlight | paper |
| 2025-03-10 | LLM Attention, Scene Magnifier, Cross-Room | SCUT | LSceneLLM: LLM-Attention Adaptive 3D Scene Understanding | CVPR 2025 | paper |
| 2025-01-12 | LVLM-Guided, Hierarchical Feature, 3DGS | Fudan / NTU | ReasonGrounder: LVLM-Guided Hierarchical Feature Splatting for Open-Vocabulary 3D Visual Grounding | CVPR 2025 | project |
| 2025-01-04 | Generalist 3D LMM, Omni Superpoint Transformer | Adelaide / Microsoft | 3D-LLaVA: Towards Generalist 3D LMMs with Omni Superpoint Transformer | CVPR 2025 | github |
| 2024-07-18 | Object Identifiers, 3D VL, Unified Tasks | ZJU | Chat-Scene: Bridging 3D Scene and Large Language Models with Object Identifiers | NeurIPS 2024 | github |
| 2024-07-15 | Million-Scale 3D VL, Multi-Level Contrastive | BIGAI | SceneVerse: Scaling 3D Vision-Language Learning for Grounding | ECCV 2024 | project |
| 2024-06-06 | Zero-Shot 3D Grounding, 2D VLM Transfer | NUS | SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual Grounding | CVPR 2025 | project |
| 2024-04-03 | Open-Vocabulary 3D Scene Graph, Open Relations | Bosch | Open3DSG: Open-Vocabulary 3D Scene Graph Generation | CVPR 2024 | paper |
| 2024-03-08 | Interactive 3D, Language Assistant, Point Cloud | CUHK | LL3DA: Visual Interactive Instruction Tuning for 3D Language Assistant | CVPR 2024 | github |
| 2024-03 | CLIP Cross-Modal, Contrastive, Scene Graph | Authors | CCL-3DSGG: CLIP-Driven Contrastive Learning for 3D Scene Graph Generation | CVPR 2024 | paper |
| 2023-11-21 | Generalist Embodied 3D Agent, VLA, ICML | BIGAI | LEO: An Embodied Generalist Agent in 3D World | ICML 2024 | project |
| 2023-08-08 | 3D Visual Grounding, Referring Expression, Point Cloud | Peking University | 3D-VisTA: Pre-trained Transformer for 3D Vision and Text Alignment | ICCV 2023 | github |
| 2023-07-24 | 3D VQA, 3D Captioning, Scene Understanding | Shanghai AI Lab | 3D-LLM: Injecting the 3D World into Large Language Models | NeurIPS 2023 | project / github |
| 2023-03 | 3D Language Pre-training, Captioning, QA | Authors | 3D-VLP: 3D Vision-Language Pre-training with Contextual Scene | CVPR 2023 | paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-29 | iToF, Sensor-Intrinsic Uncertainty, Heteroscedastic Restoration | Tsinghua University | Reliable iToF Depth Sensing via Sensor-Intrinsic Uncertainty Modeling and State-Space Restoration | arXiv | paper |
| 2026-07-27 | Neural Structured Light, Metric Depth, Online SLAM | Peking University | NSL-SLAM: High-Fidelity Neural Structured-Light Depth for Practical SLAM and Reconstruction | ACM MM 2026 | paper |
| 2026-07-20 | Projector-Camera, Feed-Forward 3DGS, Active Illumination | Ningbo University | FF-ProCams: Feed-Forward Gaussian Splatting for Projector-Camera System | arXiv | github |
| 2026-06 | Active Stereo, 2DGS Supervision, RealSense Dataset | Hangzhou Dianzi University | GS-ASM: 2DGS-Supervised Active Stereo Matching | CVPR 2026 | paper |
| 2026-05-07 | Adaptive 4D Illumination, Shape+Reflectance, Differentiable Capture | Zhejiang University | Differentiable Adaptive 4D Structured Illumination for Joint Capture of Shape and Reflectance | CVPR 2026 | paper |
| 2026-03 | Multi-Projector Structured Light, One-Shot Scan, Neural SDF | Kyushu University | Multi-view Stereo with Multiple Projectors for Oneshot Entire Shape Scan based on Neural SDF and DSSS Demultiplexing | WACV 2026 | paper |
| 2026-02-05 | Latent Diffusion, Fringe Projection, Reflective Objects | Yonsei University | LD-SLRO: Latent Diffusion Structured Light for 3-D Reconstruction of Highly Reflective Objects | arXiv | paper |
| 2025-12-16 | Single-Shot, Neural Feature Decoding, Robust Correspondence | Peking University | Robust Single-shot Structured Light 3D Imaging via Neural Feature Decoding | SIGGRAPH Asia 2025 | project |
| 2025-12 | Event Camera, HDR Measurement, Confidence Stereo | USTC | Event-based HDR Structured Light | NeurIPS 2025 | github |
| 2025-03-08 | Active Stereo, Phase Speckle, Cross-Scene Generalization | Southwest Jiaotong University | RGB-Phase Speckle: Cross-Scene Stereo 3D Reconstruction via Wrapped Pre-Normalization | arXiv | paper |
| 2025-02 | Unsupervised Structured Light, Neural SDF, Shadow-Aware | Kyushu University | Neural SDF for Shadow-Aware Unsupervised Structured Light | WACV 2025 | paper |
| 2025-01-13 | Matching-Free, Volume Rendering, Monocular Structured Light | USTC | Matching-Free Depth Recovery from Structured Light | arXiv | paper |
| 2024-10-20 | Neural SDF, One-Shot Scan, Low-Light/Underwater | Kyushu University | ActiveNeuS: Neural Signed Distance Fields for Active Stereo | 3DV 2024 | paper |
| 2024-06-06 | Virtual Pattern Projection, Depth Fusion, In-the-Wild Stereo | University of Bologna | Active Stereo in the Wild through Virtual Pattern Projection | arXiv | github |
| 2024-06 | Neural Inverse, Dense Depth, 3-4 Patterns | University of Toronto | TurboSL: Dense, Accurate and Fast 3D by Neural Inverse Structured Light | CVPR 2024 | project |
| 2023-06-17 | Structured Light, Phase Unwrapping, Learning | Nanjing University | Deep Learning-Based Structured Light 3D Imaging: A Survey | arXiv | survey |
| 2018-11-27 | Event Camera, Active Stereo, Depth | Tsinghua University | Event-Based Structured Light for Depth Reconstruction | IJCAS 2024 | paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-16 | Wide-FOV, Egocentric, 4D Hand-Object | Rice University | EgoInfinity: A Web-Scale 4D Hand-Object Interaction Data Engine for Any-View Robot Retargeting and Video-to-Action Robot Learning | arXiv | project / dataset |
| 2026-03-19 | Panoramic Depth, VGGT, Geometry Consistency | Singapore University of Technology and Design | VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation | CVPR 2026 | paper |
| 2026-03-18 | Panoramic Reconstruction, Permutation-Equivariant, 360° | Cornell | PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery | CVPR 2026 | paper |
| 2026-01-25 | RGB-D, Depth Completion, Reflective/Transparent | Robbyant | LingBot-Depth: Masked Depth Modeling for Spatial Perception | arXiv | github / paper |
| 2025-12-18 | Panoramic Foundation Model, Multi-Camera, Metric Depth | Insta360 Research | Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation | CVPR 2026 | paper |
| 2025-01 | Fisheye, Real-Time, Cassini, Multi-View | Sun Yat-Sen Univ | OmniStereo: Real-time Omnidirectional Depth with Multiview Fisheye Cameras | CVPR 2025 | github |
| 2024-06-19 | Panoramic Depth, Semi-Supervised, Mobius | HKUST(GZ) | PanDA: Panoramic Depth Anything with Mobius Spatial Augmentation | CVPR 2025 | project |
| 2024-03-25 | 360 Depth, Bi-Projection, ERP+ICOSAP | HKUST(GZ) | Elite360D: Efficient 360 Depth Estimation via Bi-Projection Fusion | CVPR 2024 | github |
| 2021-09-06 | 360 Depth, Indoor, Panoramic Images | CERTH | Pano3D: A Holistic Benchmark and a Solid Baseline for 360 Depth Estimation | CVPRW 2021 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-26 | Active Event Stereo, High-Speed Depth, 150 FPS | Authors | Towards Ultrafast Depth Sensing Via Active Event-based Stereo Vision | TPAMI 2026 | paper |
| 2026-07-17 | Event Camera, Feed-Forward 3D, Temporal Aggregation | Zhejiang University | Event3R: Asynchronous-to-Global 3D Reconstruction from Event Camera via Spatial-Temporal Feature Aggregation | arXiv | paper |
| 2026-06 | RGB+ToF Histogram, High-Resolution Metric Depth, Lightweight | Tongji University | LiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth Estimation | CVPR 2026 Highlight | paper |
| 2026-06 | Sparse dToF, Zero-Shot Completion, Sensor Generalization | KAIST | Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors | CVPR 2026 | paper |
| 2026-06 | Image-Event Fusion, Monocular Depth, Linear Complexity | Peking University | AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-06 | Event-Image Depth, Hypothesis Volume, Iterative Refinement | Southeast University | Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation | CVPR 2026 | paper |
| 2026-04-16 | Event-Frame Stereo, Cross-Modal Prompting, High Dynamic Range | Southeast University | Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo | CVPR 2026 | paper |
| 2026-04-02 | Event Stereo, Data Factory, Cross-Modal Distillation | University of Bologna | EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors | CVPR 2026 | project |
| 2026-02-03 | Event Camera, Neural SDF, Single-Camera Mesh | Saarland University | EventNeuS: 3D Mesh Reconstruction from a Single Event Camera | 3DV 2026 | project / github / dataset |
| 2025-12-20 | Event Camera, Structured Light, Real-Time RGB-D | Polytechnique Montreal | E-RGB-D: Real-Time Event-Based Perception with Structured Light | arXiv | github |
| 2025-09-18 | Event Camera, Depth Estimation, Any-to-Any | Shanghai AI Lab | Depth AnyEvent: Event Camera Based Monocular Depth Estimation via Dense Correspondence Distillation | arXiv | paper |
| 2025-09-08 | Event Camera, Multispectral, Structured Light | ETH Zurich | Event Spectroscopy: Event-based Multispectral and Depth Sensing using Structured Light | arXiv | paper |
| 2025-05-28 | Burst-Encodable ToF, Long-Range Depth, Hardware-Aware Coding | Nanjing University | Learnable Burst-Encodable Time-of-Flight Imaging for High-Fidelity Long-Distance Depth Sensing | NeurIPS 2025 | github |
| 2025-05 | Event, Distillation, Confidence-Guided, Pseudo-Labels | NUS | Distil-E2D: Distilling Image-to-Depth Priors for Event-Based Depth | NeurIPS 2025 | paper |
| 2025-04-23 | ToF, Sparse Depth, 3DGS, SLAM | Authors | ToF-Splatting: Dense SLAM using Sparse Time-of-Flight Depth | ICCV 2025 | paper / paper |
| 2025-04-22 | Event Camera, Ray Density, 3D Conv, Spotlight | TU Berlin | DERD-Net: Learning Depth from Event-based Ray Densities | NeurIPS 2025 Spotlight | github |
| 2025-03-03 | dToF, Video Depth Completion, Frequency Selective | Authors | SVDC: Consistent Direct Time-of-Flight Video Depth Completion | CVPR 2025 | paper / paper |
| 2024-10-10 | Event Camera, Pose-Free, Gaussian Splatting, Highlight | Zhejiang Univ | IncEventGS: Pose-Free Gaussian Splatting from a Single Event Camera | CVPR 2025 Highlight | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | mmWave Radar, Complex-Valued Point Splatting, Material Model, Novel Views | Cornell Tech | 3D Point Splatting for mmWave Radar Novel View Synthesis | arXiv | paper |
| 2026-09-09 | mmWave Radar, Metric Depth, Smoke/Fog/Darkness, 95K Frames | Rice University | GRADE: Single-Frame Generative Radar Depth Estimation Under Visual Degradation | MobiCom 2026 | project+data |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-09 | RGB-D 3DGS, Online Reconstruction, Reactive Control | Mitsubishi Electric Research Laboratories | SplatCtrl: Perception-Action Coupling via Gaussian Scene Representations and Reactive Robot Control | ICRA 2026 | paper |
| 2026-07-08 | Geometry-Only 3DGS, Dense Monocular SLAM | Beihang University | GeoGS-SLAM: Geometry-Only Gaussian Splatting for Dense Monocular SLAM | arXiv | paper |
| 2026-07-05 | LiDAR 3DGS SLAM, Geometry-Aware Covariance, Real-Time | Authors | Real-Time LiDAR Gaussian Splatting SLAM via Geometry-Aware Covariance Coupling | arXiv | github |
| 2026-07-02 | Dynamic Gaussian SLAM, Dual-Level Probability, Semantic Map | Authors | DL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level Probability | arXiv | paper |
| 2026-06-29 | Task-Conditioned 3DGS, Real-Time Mapping, Multi-Agent Fusion | MIT | GaussLite: Online Task-Conditioned 3D Gaussian Splatting for Real-Time Robotic Mapping | arXiv | paper |
| 2026-06-29 | RGB-Only Gaussian SLAM, Closed-Loop Geometry, Scale Feedback | CAS / USTC | MyGO-Splat: Multi-Objective Closed-Loop Geometric Feedback for RGB-Only Gaussian SLAM | IROS 2026 | paper |
| 2026-06-27 | Object-Centric 3DGS, Lifelong Mapping, Dynamic Maintenance | Beijing Institute of Technology | CubifyGS: Object-Centric 3D Gaussian Splatting for Lifelong Dynamic Scene Maintenance | IROS 2026 | paper |
| 2026-03-10 | Uncertainty-Aware 3DGS, RGB-D SLAM, Loop Detection | George Mason University | VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM | CVPR 2026 | project / github |
| 2025-11-20 | Language-Embedded, Open-Vocabulary, 3DGS SLAM | KAIST | LEGO-SLAM: Language-Embedded Gaussian Optimization SLAM | arXiv | paper |
| 2025-07-25 | Neural SLAM, Dense Mapping, Self-Supervised | Shanghai AI Lab | DINO-SLAM: Dense Tracking and Mapping with Self-Supervised Feature Learning | arXiv | paper |
| 2025-07 | Dynamic Surface, Non-Rigid, 4D Tracking | Imperial College | 4DTAM: Dynamic Surface Gaussian SLAM | CVPR 2025 | paper |
| 2025-03-20 | 4DGS SLAM, Dynamic/Static, Tracking | Authors | 4D Gaussian Splatting SLAM | ICCV 2025 | paper / paper |
| 2025-03-11 | Gaussian SLAM, Dense Reconstruction, RGB-D | Zhejiang University | GigaSLAM: Gaussian Splatting-based Large-Scale Dense SLAM | arXiv | github |
| 2025-03 | Multi-Agent, 3DGS SLAM, Loop Closure | Authors | MAGiC-SLAM: Multi-Agent 3DGS SLAM | CVPR 2025 | paper |
| 2025-01-25 | Gaussian SLAM, Dynamic Environments, Monocular | Stanford / ETH | WildGS-SLAM: Monocular Gaussian Splatting SLAM in Dynamic Environments | CVPR 2025 | project |
| 2025-01-24 | Multi-Robot, Semantic, Heterogeneous, 3DGS | Stanford | HAMMER: Heterogeneous, Multi-Robot Semantic Gaussian Splatting | RAL 2025 | project |
| 2024-11-03 | Gaussian SLAM, Global BA, Monocular RGB | ETH / Meta | Splat-SLAM: Globally Optimized RGB-only SLAM with 3D Gaussians | CVPR 2025 | project |
| 2024-09-10 | Single-Image Calibration, Geometric Optimization | ETH / Meta | GeoCalib: Learning Single-image Calibration | ECCV 2024 | paper |
| 2024-04-11 | Detector-Free SfM, Texture-Poor | Zhejiang Univ | Detector-Free Structure from Motion | CVPR 2024 | github |
| 2024-04 | Pose Regression, Map-Relative, Multi-Scene, Highlight | Niantic / Oxford | Marepo: Map-Relative Pose Regression for Visual Re-Localization | CVPR 2024 Highlight | github |
| 2024-02-20 | Neural SLAM, Survey, Radiance Fields | TUM | How NeRFs and 3D Gaussian Splatting are Reshaping SLAM: a Survey | arXiv | project |
| 2023-12-11 | Dense SLAM, Gaussian Splatting, RGB-D | TUM | Gaussian Splatting SLAM | CVPR 2024 | github |
| 2023-12-07 | End-to-End SfM, Differentiable BA | Meta AI / Oxford | VGGSfM: Visual Geometry Grounded Deep SfM | CVPR 2024 | github |
| 2023-12-04 | Gaussian SLAM, RGB-D, Volumetric | CMU / MIT | SplaTAM: Splat, Track & Map 3D Gaussians for Dense RGB-D SLAM | CVPR 2024 | project / github |
| 2021-12-22 | Neural Mapping, Dense RGB-D SLAM, SDF | HKUST | NICE-SLAM: Neural Implicit Scalable Encoding for SLAM | CVPR 2022 | project |
This section focuses on the mathematical and data-structure layer used to represent geometry, appearance, and motion.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-21 | Gaussian Surfels, Topology Recovery, Mesh Extraction, Surface Reconstruction | University of Science and Technology of China | TopoSurfel: Closing the Loop between Gaussian Surfels and Meshes for Surface Reconstruction | arXiv | github |
| 2026-08-20 | Sparse Light Field, Casual Capture, 3D/4D Reconstruction, 3DGS | Cornell University | Sparse Light Field Sampling Improves Casual 3D and 4D Reconstruction | arXiv | project |
| 2026-08-17 | Pose-Free NVS, 3DGS Geometry, Visibility-Aware Guidance | Aalto University | SplatGuide: Geometric Priors from 3D Gaussians for Pose-Free Novel View Synthesis | arXiv | paper |
| 2026-08-13 | Feed-Forward 3DGS, 3D Anchors, Spatially Grounded Tokens | National University of Defense Technology | LocusGS: Spatially Grounded Tokens for Feed-Forward 3D Gaussian Splatting | arXiv | project |
| 2026-08-03 | Streaming Feed-Forward, Persistent Geometry, Memory-Bounded 3DGS | University of Science and Technology of China | StreamSplat: Streaming Feed-Forward 3D Gaussian Splatting | arXiv | paper |
| 2026-08-03 | View-Conditioned, Feed-Forward 3DGS, Generalizable Reconstruction | Tsinghua University | UniqueSplat: View-conditioned 3D Gaussian Splatting for Generalizable 3D Reconstruction | arXiv | paper |
| 2026-07-28 | Dynamic 3DGS, Adaptive Streaming, Volumetric Video | Authors | SplatStream: Fine Granular Scalable Gaussian Splatting for Adaptive 3D Scene Streaming | arXiv | paper |
| 2026-07-22 | Feed-Forward 3DGS, Adaptive Tokens, Compact Representation | Yonsei University | ATSplat: Compact Feed-forward 3D Gaussian Splatting with Adaptive Token Expansion | arXiv | project |
| 2026-04-16 | Feed-Forward 3DGS, Global Scene Tokens, Compact | Tel Aviv University | GlobalSplat: Efficient Feed-Forward 3D Gaussian Splatting via Global Scene Tokens | arXiv | paper |
| 2026-04-16 | TokenGS, Learnable Gaussian Tokens, Pose-Robust | NVIDIA | TokenGS: Decoupling 3D Gaussian Prediction from Pixels | CVPR 2026 Highlight | project |
| 2026-04-12 | UniSplat, Unposed Multi-View, Feed-Forward | UC Berkeley | UniSplat: Learning 3D Representations from Unposed Multi-View Images | CVPR 2026 | paper |
| 2026-02-02 | Feed-Forward 2DGS, Surface Continuity, Sparse Views | Shanghai Jiao Tong | SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors | ICLR 2026 | paper |
| 2025-07-31 | Sparse-View 4D, Monocular Fusion, Cross-Video | Meta AI | MonoFusion: Sparse-View 4D Reconstruction via Monocular Fusion | ICCV 2025 | paper |
| 2025-06-17 | Anti-Aliasing, Gaussian, 3D, Adaptive | CUHK | AAA-Gaussians: Anti-Aliasing 3D Gaussian Splatting | ICCV 2025 | paper |
| 2025-06-05 | Feed-Forward 3DGS, Depth Parameterization, Geometry | Zhejiang University | Revisiting Depth Representations for Feed-Forward 3D Gaussian Splatting | 3DV 2026 | paper |
| 2025-05-21 | Sparse-View Surface, Geometry-Prioritized, 2DGS | Nankai Univ | Sparse2DGS: Geometry-Prioritized Gaussian Splatting for Surface Reconstruction from Sparse Views | CVPR 2025 | paper |
| 2025-04-03 | Feed-Forward, No Camera, FreeSplatter | ZJU | FreeSplatter: Pose-Free 3DGS from Sparse Views | CVPR 2025 | paper |
| 2025-03-13 | Few-Shot, Diffusion Prior, Repair + Inpainting | ZJU / Alibaba | RI3D: Few-Shot Gaussian Splatting With Repair and Inpainting Diffusion Priors | ICCV 2025 | paper |
| 2025-01-11 | Scene-Level, 3DGS, Open-Vocabulary, 4D LangSplat | UPenn | 4D LangSplat: 4D Language Gaussian Splatting | CVPR 2025 | paper |
| 2024-12-09 | Large-Scale Dynamic, City, 4DGS, Spotlight | ZJU | DynamicCity: Large-Scale 4D Gaussian City Modeling | ICLR 2025 Spotlight | paper |
| 2024-09-26 | Progressive, Pruning, 3DGS, CVPR | S-Lab | PUP 3D-GS: Progressive Pruning for 3DGS | CVPR 2025 | paper |
| 2024-03-26 | 2DGS, Surface Reconstruction, Geometry | TUM | 2D Gaussian Splatting for Geometrically Accurate Radiance Fields | SIGGRAPH 2024 | project |
| 2024-03-21 | Feed-Forward 3DGS, Sparse Views, Generalizable | Monash / Tübingen | MVSplat: Efficient 3D Gaussian Splatting from Sparse Multi-View Images | ECCV 2024 Oral | project / github |
| 2024-03-14 | Multi-Scale, 3DGS, Level-of-Detail | Shanghai Jiao Tong | Multi-Scale 3DGS | CVPR 2024 | paper |
| 2023-12-19 | Feed-Forward 3DGS, Image Pairs, Epipolar | MIT | pixelSplat: 3D Gaussian Splats from Image Pairs for Scalable Generalizable 3D Reconstruction | CVPR 2024 Oral | github |
| 2023-12-01 | Distillation, Lightweight, 3DGS, Spotlight | Virginia Tech | LightGaussian: Distilled 3D Gaussian Splatting | NeurIPS 2024 Spotlight | paper |
| 2023-11-30 | 3DGS, Structure, Sparse Views | USTC | SparseGS: Real-Time 360 Sparse View Synthesis using Gaussian Splatting | 3DV 2025 | project |
| 2023-11-27 | Anti-Aliasing, Mip, 3DGS, Best Student Paper | Inria | Mip-Splatting: Alias-free 3DGS | CVPR 2024 Oral / Best Student Paper | github |
| 2023-11-27 | Compression, Compact, 3DGS, Highlight | Tsinghua | Compact 3D Gaussian Splatting | CVPR 2024 Highlight | paper |
| 2023-11-21 | 3DGS, Mesh Extraction, Surface | ETH Zurich | SuGaR: Surface-Aligned Gaussian Splatting for Efficient 3D Mesh Reconstruction | CVPR 2024 | project |
| 2023-08-08 | 3DGS, Real-Time Rendering, Explicit Radiance | Inria | 3D Gaussian Splatting for Real-Time Radiance Field Rendering | SIGGRAPH 2023 | github / paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-18 | Sparse Voxel Latent, Fitting-Free 3DGS, Large-Area Generation | Amap, Alibaba | GS-Voxel: Fitting-Free Structured Latents for Large-Scale 3DGS Generation | arXiv | paper |
| 2026-07-22 | Shape Completion, Unified 3D Representation, Faithful Geometry | Authors | Axolotl3D: A Unified Framework for Faithful 3D Shape Completion | arXiv | paper |
| 2024-12-02 | Structured Latent, 3D Generation, Mesh | Microsoft | TRELLIS: Structured 3D Latents for Scalable and Versatile 3D Generation | CVPR 2025 | project / github |
| 2024-03-04 | Sparse Voxel, Text/Image-to-3D, Mesh | Stability AI | TripoSR: Fast 3D Object Reconstruction from a Single Image | arXiv | github |
| 2022-12-16 | Point Cloud, Diffusion, Shape Generation | OpenAI | Point-E: A System for Generating 3D Point Clouds from Complex Prompts | arXiv | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2025-07-15 | Unbiased SDF, Neural Implicit Surface, SDF-to-Density | CUHK | UNIS: Unified Framework for Unbiased Neural Implicit Surfaces | ICCV 2025 | paper |
| 2024-09-06 | α-NeuS, Alpha, SDF, Volume Rendering | ZJU | α-NeuS: Alpha-Governed Neural Implicit Surfaces | NeurIPS 2024 | paper |
| 2023-12-29 | Objects as Volumes, NeRF, 3D Reconstruction, Oral | UPenn | Objects as Volumes: Feed-Forward 3D from a Single Image | CVPR 2024 Oral | paper |
| 2023-04-13 | Zip-NeRF, Anti-Aliasing, Mip, Grid | Google Research | Zip-NeRF: Anti-Aliased Grid-Based NeRF | ICCV 2023 | project |
| 2022-03-17 | Tensor Factorization, Radiance Fields, Compression | Tsinghua | TensoRF: Tensorial Radiance Fields | ECCV 2022 | project |
| 2022-01-16 | Hash Grid, Real-Time NeRF, Neural Graphics | NVIDIA | Instant Neural Graphics Primitives with a Multiresolution Hash Encoding | SIGGRAPH 2022 | github / paper |
| 2021-12-15 | EG3D, Triplane, 3D GAN, Generative | NVIDIA | EG3D: Efficient Geometry-aware 3D Generative Adversarial Networks | CVPR 2022 Oral | github |
| 2021-12-09 | Plenoxels, No Neural Network, Fast, Oral | UC Berkeley | Plenoxels: Radiance Fields without Neural Networks | CVPR 2022 Oral | github |
| 2021-12-07 | Ref-NeRF, Reflection, Specular, Best Student Paper HM | Google Research | Ref-NeRF: Structured View-Dependent Appearance for NeRF | CVPR 2022 Best Student Paper HM | project |
| 2021-11-23 | Unbounded Scenes, Anti-Aliasing, NeRF | Google Research | Mip-NeRF 360: Unbounded Anti-Aliased Neural Radiance Fields | CVPR 2022 | project |
| 2021-06-20 | SDF, Surface Reconstruction, Neural Rendering | MPI-IS | NeuS: Learning Neural Implicit Surfaces by Volume Rendering | NeurIPS 2021 | project |
| 2021-04-13 | BARF, Bundle-Adjusting NeRF, Oral | UC Berkeley | BARF: Bundle-Adjusting Neural Radiance Fields | ICCV 2021 Oral | github |
| 2020-08-05 | NeRF-W, Unbounded, In-the-Wild | Google Research | NeRF in the Wild: Neural Radiance Fields for Unconstrained Photo Collections | CVPR 2021 | github |
| 2020-03-19 | NeRF, View Synthesis, Neural Rendering | UC Berkeley | NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis | ECCV 2020 | github / paper |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-23 | Dynamic 3DGS, Gradient Decoupling, Novel View Synthesis | Authors | GrainGS: Gradient-Decoupled Gaussian Splatting for Efficient Dynamic Novel View Synthesis | arXiv | paper |
| 2026-06-22 | Dynamic 3DGS, Visibility-Aware Densification, Temporal Lifespan | Indian Institute of Science | Temporally Aware Densification for Dynamic 3D Gaussian Splatting | arXiv | paper |
| 2025-07 | 7DGS, Spatial-Temporal-Angular, Unified | Authors | 7D Gaussian Splatting: Unified Spatial-Temporal-Angular GS | ICCV 2025 | paper |
| 2024-10 | Event Camera, High-Speed, 4D, STD-GS | Authors | STD-GS: SpatioTemporal-Disentangled Gaussian Splatting with Event Cameras | ICCV 2025 | paper |
| 2023-10-12 | 4DGS, Dynamic Scenes, Real-Time Rendering | Zhejiang University | 4D Gaussian Splatting for Real-Time Dynamic Scene Rendering | CVPR 2024 | project |
| 2023-08-18 | Dynamic 3DGS, Scene Motion, Multi-View Video | Cornell | Dynamic 3D Gaussians: Tracking by Persistent Dynamic View Synthesis | 3DV 2025 | github / paper |
| 2023-01-24 | K-Planes, Explicit, 4D, Space-Time | Authors | K-Planes: Explicit Radiance Fields in Space, Time, and Appearance | CVPR 2023 | github |
| 2023-01-23 | HexPlane, 4D Representation, Space-Time | CMU | HexPlane: A Fast Representation for Dynamic Scenes | CVPR 2023 | project |
| 2021-06-24 | Dynamic NeRF, Deformation, Canonical Space | Google Research | HyperNeRF: A Higher-Dimensional Representation for Topologically Varying Neural Radiance Fields | SIGGRAPH Asia 2021 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-24 | Structured Motion, SE(3), 4D Reconstruction | Tsinghua University | SM4RT: Learning Structured Motion Geometry for 4D Reconstruction | arXiv | project / paper |
| 2023-12-04 | Deformation, Dynamic Radiance Fields, Canonical | ETH Zurich | SC-GS: Sparse-Controlled Gaussian Splatting for Editable Dynamic Scenes | CVPR 2024 | project |
| 2023-06-05 | Non-Rigid Tracking, Neural Deformation, 4D | Tsinghua | Neuralangelo: High-Fidelity Neural Surface Reconstruction | CVPR 2023 | project |
Reconstruction systems recover objects or scenes from images, video, RGB-D, or multi-sensor streams under offline, online, and dynamic conditions.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Active In-Hand Reconstruction, Uncertainty, Next-Best View | ShanghaiTech University | AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction | CoRL 2026 | project |
| 2026-07-21 | Single-View 3D, Object Perception, Generative Reconstruction | Deakin University | Seeing Before Generating: Object Perception Enhances Single-View 3D Reconstruction | arXiv | project |
| 2026-06-17 | Sparse-View Object, Flow Steering, 3DGS Refinement | Graz University of Technology | FlowObject: Flow Steering for Bridging Generative Priors and Reconstruction Fidelity | arXiv | project |
| 2026-05-05 | Generative Reconstruction, Multi-View Alignment, Pose | Tsinghua | Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose Estimation | arXiv | project |
| 2025-11-19 | Single Image, Object Mesh, SAM 3D | Meta AI | SAM 3D Objects | GitHub | github |
| 2025-10-23 | Pose-Free Online, Free-Moving Objects, Constant Memory | SUTD | OnlineSplatter: Pose-Free Online 3D Reconstruction for Free-Moving Objects | NeurIPS 2025 Spotlight | project |
| 2025-06-05 | RGB-D Object Completion, Novel Depth, Feed-Forward | Carnegie Mellon University | RaySt3R: Predicting Novel Depth Maps for Zero-Shot Object Completion | NeurIPS 2025 | project |
| 2025-06 | Feed-Forward, Densification, Gaussian, Detail | Authors | Generative Densification: Feed-Forward 3DGS Densification | CVPR 2025 | paper |
| 2025-06 | Photogrammetry Foundation, Multi-Task, Highlight | Authors | Matrix3D: A Foundation Model for Photogrammetry | CVPR 2025 Highlight | paper |
| 2025-04-04 | Sparse-View, Feed-Forward, Camera, Geometry | Ant Research / Stanford | FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | CVPR 2025 | project |
| 2024-07 | Ego-Centric, Autonomous Driving, Sparse-View | Authors | Omni-Scene: Omni-Gaussian for Ego-Centric Sparse-View Reconstruction | CVPR 2025 | paper |
| 2024-06-14 | Multi-View, Stereo, Feed-Forward | NAVER Labs | MASt3R: Grounding Image Matching in 3D with MASt3R | ECCV 2024 | project / github |
| 2023-12-21 | Multi-View, Pointmap, Pose-Free | NAVER Labs | DUSt3R: Geometric 3D Vision Made Easy | CVPR 2024 | project / github |
| 2023-12-13 | Object Pose, Reconstruction, Model-Based | NVIDIA | FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects | CVPR 2024 | project |
| 2023-03-24 | Unknown Object, RGB-D, 6-DoF Tracking, Neural SDF | NVIDIA | BundleSDF: Neural 6-DoF Tracking and 3D Reconstruction of Unknown Objects | CVPR 2023 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Feed-Forward, Compositional Scene, Complete Meshes, Simulation-Ready | University of Illinois Urbana-Champaign | FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute | arXiv | project / github |
| 2026-09-04 | Bundle Adjustment, Multi-View Matching, Monocular Priors, Online+Offline | NAVER Labs Europe | BLASt3R: Bundle Adjustment of Any Image Set with Multi-View Matching and Monocular Priors | ECCV 2026 | paper |
| 2026-08-31 | Real-to-Sim, Parse-Generate-Place, Composable Object Assets | ByteDance Seed | Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling | arXiv | project |
| 2026-08-25 | Single Image, Generative Reconstruction, Complete Object Assets, Scene Assembly | Huawei | SceneReGen: Generative Reconstruction of 3D Scenes from a Single Image | arXiv | paper |
| 2026-08-18 | Instance-Grouped 3DGS, Semantic Reconstruction, Referential Scene Graph | Shanghai Jiao Tong University | GroupForward: Building Referable 3D Scenes via Instance-Grouped Feed-Forward Gaussian Splatting | arXiv | paper |
| 2026-08-18 | Generative NVS, Reconstruction/Generation Split, Scene Coordinates | KAIST | GenRec: Knowing Where to Reconstruct and Where to Generate | arXiv | project |
| 2026-08-18 | Long Sequence, Chunk Priors, Sim(3) Assembly, Test-Time Adaptation | Kosmo Research | GeoWeaver: Accurate Long-Sequence 3D Reconstruction via Hierarchical Geometric Assembly | arXiv | project |
| 2026-08-15 | Long Sequence, Scale-Consistent Alignment, Test-Time Adaptation | Northwestern Polytechnical University | VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction | ACM Multimedia 2026 | github |
| 2026-08-12 | Streaming Multi-View, Metric 3D, Feed-Forward Prior, Object Detection | ETH Zurich | Map-Det3D: Metric Feed-Forward 3D Reconstruction Prior for Multi-view 3D Object Detection from Streaming Inputs | ECCV 2026 | project |
| 2026-08-07 | Sparse View, Editable Indoor Scenes, Executable Scene Programs | City University of Hong Kong | Scenix: Sparse-View 3D Scene Reconstruction via Executable Scene Programs | arXiv | paper |
| 2026-07-31 | Active Reconstruction, Next-Best-View, Predictive Entropy | Fudan University | GO-PRE: Goal-Oriented Next-Best-View Selection via Predictive Rendering Entropy for Active 3D Reconstruction | arXiv | paper |
| 2026-07-23 | Underwater 3D, Feed-Forward Reconstruction, Degradation Adaptation | HKUST | WAT3R: Feedforward Underwater 3D Reconstruction | arXiv | project |
| 2026-07-15 | Feed-Forward Driving Reconstruction, Layered 3DGS, Dynamic Actors | NVIDIA | Instant NuRec: Feed-Forward 3D Gaussian Reconstruction for Driving Scene Simulation | arXiv | project / github / docs |
| 2026-07-10 | 3D Foundation Model, Global SfM, Bundle Adjustment | HKUST | Glob3R: Global Structure-from-Motion with 3D Foundation Models | arXiv | project |
| 2026-07-08 | Feed-Forward 3D, Unposed Images, Drift-Robust | Authors | NoDrift3R: Raymap-Guided Coupling for Drift-Robust Unposed Feed-Forward 3D Reconstruction | ECCV 2026 | paper |
| 2026-06-09 | RGB-T, Thermal Geometry, Low-Light | University of Minnesota | DarkVGGT: Seeing Through Darkness Using Thermal Geometry without Daylight Tax | arXiv | project |
| 2026-06-02 | Single Image, Physics-in-the-Loop, Simulation-Ready | Seoul National University | SimuScene: Simulation-Ready Compositional 3D Scene Reconstruction from a Single Image | arXiv | project |
| 2026-05-14 | VGGT, Scaling, Static+Dynamic | University of Oxford | VGGT-Omega: Scaling VGGT to Large-Scale 3D Reconstruction | CVPR 2026 Oral | project / github / demo / model |
| 2026-05-07 | Feed-Forward 3D, Token Reduction, Long Sequence | Peking University | Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction | arXiv | paper |
| 2026-04-30 | Generalizable, Sparse-View, Unposed Images, Outdoor | UIUC / NVIDIA | GenWildSplat: Generalizable Sparse-View 3D Reconstruction from Unconstrained Images | arXiv | paper |
| 2026-03-24 | Panoramic Video, Pose-Free 3DGS, Consistent Depth Prior | University of Chinese Academy of Sciences | Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors | CVPR 2026 | github |
| 2026-03-16 | Event-to-Edge, Pose-Free, Gaussian Reconstruction | KAIST | E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction | CVPR 2026 | paper |
| 2026-02-26 | VGGT, TTT, Large-Scale | NVIDIA | VGG-T^3: Offline Feed-Forward 3D Reconstruction at Scale | CVPR 2026 | project |
| 2026-02-03 | Single Image, Object Decomposition, Occlusion-Aware Scene Reconstruction | University of California, San Diego | Seeing Through Clutter: Structured 3D Scene Reconstruction via Iterative Object Removal | 3DV 2026 | project |
| 2025-09-24 | Mirror Stereo, Single-View 3D, Symmetry Constraint | University of Oxford | Reflect3r: Single-View 3D Stereo Reconstruction Aided by Mirror Reflections | 3DV 2026 | project / github / dataset |
| 2025-09-16 | Universal 3D, Metric Reconstruction, Optional Priors | Meta AI | MapAnything: Universal Feed-Forward Metric 3D Reconstruction | 3DV 2026 | project |
| 2025-08-05 | Unposed Multi-View, 3DGS, Semantic Reconstruction | Sungkyunkwan University | Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images | CVPR 2026 | paper |
| 2025-07-17 | Permutation-Equivariant, Visual Geometry, Point Maps | Oxford / Meta | π³: Permutation-Equivariant Visual Geometry Learning | ICLR 2026 | github |
| 2025-07 | Monocular Prior, MVS, DTU/Tanks SOTA | Authors | MonoMVSNet: Monocular Prior Guided MVS | ICCV 2025 | paper |
| 2025-07 | Latent Align, Stereo+Monocular, Highlight | Authors | BridgeDepth: Unified Monocular and Stereo Depth | ICCV 2025 Highlight | paper |
| 2025-06-30 | Video-Depth Augmentation, Scalable Training, Feed-Forward 3D | Australian National University | Puzzles: Unbounded Video-Depth Augmentation for Scalable End-to-End 3D Reconstruction | NeurIPS 2025 | project |
| 2025-06-03 | Monocular Depth, Dynamic Video, Alignment | HKUST / CUHK / HKU | Align3R: Aligned Monocular Depth Estimation for Dynamic Videos | CVPR 2025 | github |
| 2025-06 | Aerial-Ground, Large-Scale, 3DGS | Authors | Horizon-GS: Unified Aerial-Ground 3DGS | CVPR 2025 | paper |
| 2025-06 | Autonomous Driving, Feed-Forward, 3DGS | Authors | EVolSplat: Feed-Forward 3DGS for Urban Driving | CVPR 2025 | paper |
| 2025-06 | Sparse-View, Super-Resolution, 3DGS | Authors | S2Gaussian: Sparse-View Super-Resolution 3DGS | CVPR 2025 | paper |
| 2025-05-05 | Relative Camera Pose, Regression, Localization | Aalto / HKU | Reloc3r: Large-Scale Training of Relative Camera Pose Regression | CVPR 2025 | github |
| 2025-03-28 | Feed-Forward, Surface, MVS, Multi-View | Authors | MVSAnywhere: Zero-Shot Multi-View Stereo | CVPR 2025 | paper / paper |
| 2025-03-17 | Feed-Forward, Multi-View, Auxiliary Priors | ETH / Microsoft | Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | CVPR 2025 | project |
| 2025-03-14 | Feed-Forward 3D, Pose-Free, Point Map | Meta AI | VGGT: Visual Geometry Grounded Transformer | CVPR 2025 Best Paper | project / github |
| 2025-03-03 | Multi-View, Symmetric, 1000+ Images, O(N) | NAVER Labs | MUSt3R: Multi-view Network for Stereo 3D Reconstruction | CVPR 2025 | project |
| 2025-01-23 | Feed-Forward, 1500+ Images, 251 FPS | Meta AI | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | CVPR 2025 | project |
| 2024-12-12 | Feed-Forward, Online, Dense, Monocular | Shanghai AI Lab / PKU | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | CVPR 2025 Highlight | project |
| 2024-12-09 | Sparse View, Single-Stage, 2 Seconds | Meta AI | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | CVPR 2025 Oral | project |
| 2024-09-27 | Multi-View Reconstruction, Matching, MVS | NAVER Labs | MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion | 3DV 2025 | github |
| 2024-08-28 | Spatial Memory, Feed-Forward, Multi-View | HKU | Spann3R: 3D Reconstruction with Spatial Memory | 3DV 2025 Oral (Best Paper Candidate) | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-07 | Online SLAM, Functional Scene Graph, Interaction Elements, Map Memory | Tsinghua University | Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs | arXiv | github |
| 2026-09-01 | Training-Free, Open-Vocabulary Instance Map, RGB-D/Monocular SLAM | University of Technology Sydney | VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM | arXiv | paper |
| 2026-08-24 | Spatio-Temporal SLAM, Open-Vocabulary, 4D Scene Graph, VLN | Carnegie Mellon University | SuperMap: A Spatio-Temporal SLAM System for Visual-Language Navigation | arXiv | project |
| 2026-08-18 | Open-Vocabulary Map, Instance Preservation, Fine-Grained Retrieval, Target Absence | Xi'an Jiaotong University | OVIP-SG: Open-Vocabulary Instance-Preserving Scene Graphs for Mapping and Retrieval of Small, Fine-Grained Objects | arXiv | github |
| 2026-06-23 | Object-Level Map, Open-Vocabulary, Relocalization | Zhejiang University | Compact Object-Level Representations with Open-Vocabulary Understanding for Indoor Visual Relocalization | arXiv | paper |
| 2026-05-05 | Online, Voxel+Instance, Open-Vocabulary Mapping | Örebro University | FUS3DMaps: Scalable and Accurate Open-Vocabulary Semantic Mapping | arXiv | project |
| 2026-05-03 | Gaussian-Language Map, Zero-Shot Navigation, Multi-Scale | CASIA | Multi-Scale Gaussian-Language Map for Zero-shot Embodied Navigation and Reasoning | arXiv | github |
| 2026-03-04 | Semantic 3DGS, Online, CLIP, Open-Vocabulary | National University of Singapore | EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding | CVPR 2026 | project |
| 2025-08-02 | Open-Vocabulary, Hybrid 3DGS+TSDF, Dense Mapping | Tsinghua | OpenGS-Fusion: Open-Vocabulary Dense Mapping with Hybrid 3D Gaussian Splatting | arXiv | project |
| 2025-07 | Feed-Forward, Panoramic Segmentation, DUSt3R | Authors | PanSt3R: Single-Feed 3D and Panoptic Segmentation | ICCV 2025 | paper |
| 2023-10-05 | Open-Vocabulary, RGB-D, TSDF, Real-Time Mapping | University of Arkansas | Open-Fusion: Real-time Open-Vocabulary 3D Mapping and Queryable Scene Representation | IROS 2024 | project / github |
| 2023-02-14 | Open-Set, Multimodal, 3D Map, Language Query | MIT | ConceptFusion: Open-set Multimodal 3D Mapping | ICRA 2023 | project / github |
| 2022-10-11 | Implicit Field, CLIP, Semantic Search, Robot Memory | New York University | CLIP-Fields: Weakly Supervised Semantic Fields for Robotic Memory | ICRA 2023 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-03 | Online 3R, Multi-Relative Pose Query, Pose-Graph Optimization | National Yang Ming Chiao Tung University | Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction | ECCV 2026 | project |
| 2026-09-01 | Online Feed-Forward 3R, Unordered UAV Images, Retrieval+Retry | Wuhan University | On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios | arXiv | github |
| 2026-08-03 | Active Reconstruction, Ergodic Coverage, Trajectory Optimization | Johns Hopkins University | TRACE: Ergodic Trajectory Optimization for Active Scene Reconstruction | arXiv | github |
| 2026-08-03 | Feed-Forward SLAM, Sim(3) Factor Graph, Persistent Mapping | Ulsan National Institute of Science and Technology | UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization | ECCV 2026 | project |
| 2026-07-25 | Semantic SLAM, Data Association, Object Landmarks | MIT | Semantic Semi-Incremental Data-Association-Free Object SLAM | arXiv | paper |
| 2026-07-23 | Gaussian SLAM, Large-Scale Mapping, Real-Time | Athena Research Center | GLAM-SLAM: Real-time Gaussian Large-scale Mapping via Flow Densification and Spatial Decomposition | IROS 2026 | github |
| 2026-07-16 | Multi-Agent 3R, RGB Video, Point-Map Fusion | University of Bologna | MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos | arXiv | project |
| 2026-07-16 | Incremental 3DGS, Unordered Capture, Global Consistency | Inria | Immediate 3D Gaussian Splat Reconstruction of Unordered Input with Global Consistency | SIGGRAPH 2026 | paper |
| 2026-07-01 | Long-Sequence, Instance Anchors, Persistent Spatial Memory | Beijing Jiaotong University | LIST3R: Long-sequence Instance-aware 3D Reconstruction | arXiv | project |
| 2026-06-23 | 3DGS-SLAM, Memory-Efficient, Outdoor Mapping | University of Minnesota | Pocket-SLAM: Rendering-Area-Aware Pruning for Memory-Efficient 3DGS-SLAM | ICRA 2026 | github |
| 2026-06-20 | RGB+Pose, 3DGS Scene Regression, Robot Capture | Peking University | ACEsplat: Accelerated 3D Gaussian Scene Regression via RGB and Poses Only | arXiv | paper |
| 2026-06-19 | 3DGS-SLAM, Degeneracy-Robust, Real-Time Tracking | Nanyang Technological University | Spectral GS-SLAM: Observability-Aware, Degeneracy-Robust Tracking for Real-Time 3D Gaussian Splatting SLAM | IROS 2026 | paper |
| 2026-06-18 | LiDAR-Inertial-Thermal, 3DGS Mapping, Illumination-Robust | Shenzhen University | LIT-GS: LiDAR-Inertial-Thermal Gaussian Splatting for Illumination-Robust Mapping | IROS 2026 | paper |
| 2026-06-03 | Streaming, Transient Anchors, Long-Horizon Mapping | Authors | Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping | arXiv | paper |
| 2026-06 | Spatial-Difference Sensor, Edge-Guided Tracking, 3DGS | Tsinghua University | SDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D Reconstruction | CVPR 2026 | paper |
| 2026-06 | Stereo 3DGS-SLAM, Auto-Exposure Robustness, Photometric Mapping | South China University of Technology | AERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAM | CVPR 2026 | github |
| 2026-05-10 | VGGT, Retrieval, Constant Memory | Fudan University | RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval | arXiv | github |
| 2026-04-24 | 4DGS-SLAM, Optical Flow, Dynamic Mapping | National University of Singapore | Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM | CVPR 2026 | github |
| 2026-04-15 | Streaming, Feed-Forward, Long Video | Robbyant | LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction | arXiv | github |
| 2026-02-13 | Streaming, Autoregressive, Long Sequence | 3DAgentWorld | LongStream: Long-Sequence Streaming Autoregressive Visual Geometry | CVPR 2026 | project |
| 2026-01-03 | StreamVGGT, KV Cache, Memory Compression | Sun Yat-sen University | XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer | arXiv | github |
| 2025-09-30 | TTT, Online, Long Context | Shanghai AI Lab | TTT3R: 3D Reconstruction as Test-Time Training | arXiv | project |
| 2025-08-14 | Streaming, Causal Transformer, Sequential | NTU / Shanghai AI Lab | STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer | arXiv | project |
| 2025-01-21 | Online 3D, Recurrent Pointmap, Streaming | Meta AI | CUT3R: Continuous 3D Perception Model with Persistent State | CVPR 2025 Oral | project / github |
| 2024-12-16 | MASt3R, Dense SLAM, Real-Time | Imperial College London | MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors | CVPR 2025 | github |
| 2023-09-05 | Global BA, Neural Implicit, Dense RGB-D SLAM | University of Bologna | GO-SLAM: Global Optimization for Consistent 3D Instant Reconstruction | ICCV 2023 | project / github |
| 2022-11-21 | Hybrid SDF, Dense RGB-D SLAM, Keyframes | Idiap Research Institute | ESLAM: Efficient Dense SLAM System Based on Hybrid Representation of Signed Distance Fields | CVPR 2023 | project / github |
| 2021-08-24 | Deep SLAM, Dense BA, Monocular/Stereo/RGB-D | Princeton University | DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D Cameras | NeurIPS 2021 | github |
| 2021-03-23 | Neural Implicit, Online RGB-D, Dense SLAM | Imperial College London | iMAP: Implicit Mapping and Positioning in Real-Time | ICCV 2021 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-08 | Long-Range 4D Motion, 3D Queries, Occlusion-Robust Trajectory Chaining | Carnegie Mellon University | Point4D: Long-range 4D Motion Reconstruction | arXiv | project |
| 2026-09-08 | Event Stream, Extreme-Low-Frame-Rate RGB, Dynamic 3DGS, Real-Time | Macau University of Science and Technology | EdMCGS: Event-Driven Markov Chain Gaussian Splatting for Extreme-Low-Frame-Rate Dynamic Scene Reconstruction | Neurocomputing | github+dataset |
| 2026-09-05 | Sparse-View 4D, Spatio-Temporal Depth Alignment, Dynamic 3DGS | BIGAI | UniFusion: Sparse-View 4D Reconstruction via Unified Spatio-temporal Depth Alignment | ECCV 2026 | project |
| 2026-08-18 | Query-Conditioned 4D, Scene Flow, Dynamic Points, Sparse-to-Dense | Kosmo Research | UniQuery4R: Unified 4D Scene Reconstruction from a Single Query | arXiv | project |
| 2026-07-29 | Articulated Objects, Structure-aware 3DGS, Part Connectivity | POSTECH | StructureGS: Structure-aware Gaussian Splatting for Articulated Object Reconstruction | arXiv | paper |
| 2026-07-21 | Streaming 4D, Instance Grounding, Geometry Transformer | Horizon Robotics | IGGT4D: Streaming 4D Instance-Grounded Geometry Transformer | arXiv | project |
| 2026-07-16 | Online Dynamic NVS, Space-Time Memory, Real-Time | University of Washington | Online Neural Space Time Memory for Dynamic Novel View Synthesis | arXiv | project |
| 2026-07-01 | Dynamic Gaussian Reconstruction, Monocular Video, Generative | Stanford University | World from Motion: Generative Dynamic Gaussian Reconstruction from Monocular Video | arXiv | project |
| 2026-06-23 | Articulated Digital Twin, RGB-D, URDF Export | ETH Zurich | ArtiTwinSplat: Interactable Digital Twin Reconstruction via Gaussian Splatting from RGB-D Videos | ICRA 2026 Workshop | paper |
| 2026-06-22 | Monocular Video, 4DGS, In-the-Wild Non-Rigid | Carnegie Mellon University | Lift4D: Harmonizing Single-View 3D Estimation for 4D Reconstruction In-the-Wild | arXiv | project |
| 2026-06-22 | Dynamic Driving, Sparse Voxels, LiDAR-Guided | Huawei Paris Research Center | DrivingVoxels: Compositional Sparse Voxel Rasterization for Dynamic Driving Scene Reconstruction | arXiv | paper |
| 2026-06-09 | Future Extrapolation, 4DGS, Autonomous Driving | Tsinghua University | Envision4D: Envisioning Visual Futures via Feed-forward 4D Gaussian Splatting for Autonomous Driving | arXiv | project |
| 2026-06-09 | Manipulation Video, Decoupled 3DGS, Scene Graph | Authors | ManiSplat: Manipulation Trajectory Synthesis from Monocular Video via Decoupled 3D Gaussian Splatting | arXiv | paper |
| 2026-06-02 | Object Permanence, Differentiable Physics, 4DGS | Authors | PersistGS: Differentiable Physics for Object Permanence in 4D Gaussian Splatting | CVPR 2026 Workshop | paper |
| 2026-06 | Single Event Camera, Deformable 3DGS, High-Speed 4D | ShanghaiTech University | FastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event Camera | CVPR 2026 | paper |
| 2026-04-10 | Feed-Forward, Unconstrained Views, Semantic-Geometry | Nanyang Tech | FF3R: Feedforward Feature 3D Reconstruction from Unconstrained Views | CVPR 2026 Findings | paper |
| 2026-04-10 | Dynamic 4D, Semantic Prior, Gaussian SLAM, Action-Control | University of Zurich | Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction | arXiv | paper |
| 2026-04-10 | Dynamic/Static Disentanglement, Uncertainty-Aware, Feed-Forward | Zhejiang University | Robust 4D VGT: Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors | arXiv | paper |
| 2026-04-07 | Functional Scenes, Egocentric Interaction, URDF/USD | Stanford | FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos | CVPR 2026 | paper |
| 2026-04-05 | Sparse Camera, 4DGS, Neural Decay, CVPR 2026 | Authors | 4C4D: 4 Camera 4D Gaussian Splatting | CVPR 2026 | paper / paper |
| 2026-03-30 | Dynamic Surface, Explicit Geometry, High Fidelity | Australian National University | 4DSurf: High-Fidelity Dynamic Scene Surface Reconstruction | CVPR 2026 | paper |
| 2026-03-21 | RayMap, Dynamic, Streaming | University of Illinois Chicago | RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction | arXiv | project / github |
| 2026-03-09 | Dynamic VGGT, Autonomous Driving, 4D | Fudan University | DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving | arXiv | paper |
| 2025-11-23 | Dynamic Geometry, Spatiotemporal, VGGT | Huazhong University of Science and Technology | 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation | arXiv | paper |
| 2025-11-07 | Motion-Aware, Monocular Video, Bundle Adjustment | KAIST | 4D3R: Motion-Aware Neural Reconstruction and Rendering of Dynamic Scenes | NeurIPS 2025 | paper |
| 2025-10-20 | VGGT-4D, Pose/Geometry, Dynamic Mask | Harvard | PAGE-4D: Disentangled Pose and Geometry Estimation for 4D Perception | ICLR 2026 | project |
| 2025-08-13 | 3D Reconstruction, Human Motion, Video | Shanghai AI Lab | Human3R: Reconstructing 3D Human Avatars from Monocular Video | CVPR 2025 | project |
| 2025-07 | Human, Multi-View, Sparse, Robust, RoGSplat | Authors | RoGSplat: Robust Generalizable Human Gaussian Splatting | CVPR 2025 | paper |
| 2025-06-11 | Dynamic Human, Temporal Consistency, 4D | Tsinghua | CARI4D: Cross-Modal Alignment and Reconstruction for Interactive 4D Human | CVPR 2025 | project |
| 2025-06-10 | Online, Dynamic 3DGS, Uncalibrated Video | University of British Columbia | StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams | ICLR 2026 | project / github |
| 2025-06-09 | 4DGS, Transformer, Monocular Video | Meta Reality Labs | 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular Videos | NeurIPS 2025 Spotlight | project |
| 2025-06-02 | Video Generators, 4D Geometry | Oxford VGG / NAVER LABS | Geo4D: Leveraging Video Generators for Geometric 4D Scene Reconstruction | ICCV 2025 Highlight | paper |
| 2025-06 | Self-Supervised, Dynamic, Driving, Flow | Authors | SplatFlow: Self-Supervised Dynamic 3DGS with Neural Motion Flow | CVPR 2025 | paper |
| 2025-06 | Few-Shot, Personal Avatar, Highlight | Authors | FRESA: Personalized 3D Human Avatar from Few Images | CVPR 2025 Highlight | paper |
| 2025-06 | Real-Time Avatar, 166fps, Gaussian, Highlight | Authors | MMLP-Human: Real-Time High-Fidelity Gaussian Human Avatar | CVPR 2025 Highlight | paper |
| 2025-05-27 | 4D, Dual Correspondences, Dynamic Video | NUS / Shanghai AI Lab | C4D: 4D Made from 3D through Dual Correspondences | ICCV 2025 | project |
| 2025-05-14 | 4D Pointmaps, Dynamic-Static Disentanglement | KAIST / ETH / Sony | D2USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic Scenes | NeurIPS 2025 | project |
| 2025-05-02 | Training-Free, Motion Disentangle, DUSt3R | Westlake / MPI | Easi3R: Estimating Disentangled Motion from DUSt3R Without Training | ICCV 2025 | project |
| 2025-04-17 | 4D Tracking, Feed-Forward, Pointmap, Tracking | MPI / UC Berkeley | St4RTrack: Simultaneous 4D Reconstruction and Tracking | ICCV 2025 | paper / paper |
| 2025-04-07 | Neural Rendering, Human, No Eyes | University of Cambridge | Seeing Without Eyes: Neural Human Rendering from Monocular Video | CVPR 2025 | project |
| 2025-03-24 | Multi-Object, 4D, In-the-Wild Videos | CMU | GenMOJO: Robust Multi-Object 4D Generation for In-the-wild Videos | CVPR 2025 | project |
| 2025-02-27 | Layered Avatar, Hair, Face, Meta | Authors | LUCAS: Layered Universal Codec Avatars | CVPR 2025 | paper / paper |
| 2025-01-22 | 3D Reconstruction, Canonical, Multi-View | Stanford | UniCon3R: Unified 3D Reconstruction and Recognition | CVPR 2025 | project |
| 2024-12-03 | Single Image, Animatable, Avatar, 4DGS | Authors | AniGS: Animatable Gaussian Avatar from a Single Image | CVPR 2025 | paper / paper |
| 2024-11-27 | 4D Generation, Multi-View Video, Diffusion | CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models | CVPR 2025 | project | |
| 2024-10-28 | Dynamic Geometry, DUSt3R, Motion | University of Oxford | MonST3R: Estimating Geometry in the Presence of Motion | ICLR 2025 | project / github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-02-12 | 4D Dynamic, Monocular Video, Tree-Chains | Cornell | WorldTree: Towards 4D Dynamic Worlds from Monocular Video using Tree-Chains | ICLR 2026 | paper |
| 2025-11-01 | 4D Scene, Feed-Forward, Controllable, Video Diffusion | Shanghai AI Lab | Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models | CVPR 2026 | paper |
| 2025-10-15 | Dynamic 4D, Gaussian, Canonical | Zhejiang University | Director: Directed Generative Models for 4D Scene Evolution | CVPR 2025 | project |
| 2025-08-06 | Multi-Baseline, Generalizable, Gaussian | Authors | MuGS: Multi-Baseline Generalizable Gaussian Splatting | ICCV 2025 | paper / paper |
| 2025-07 | Surface, Gaussian Surfels, 2DGS, Sparse-View, Spotlight | Authors | MAtCha Gaussians: Atlas Charting with 2D Gaussian Surfels | CVPR 2025 Spotlight | paper |
| 2025-07 | SDF+3DGS, Hybrid, Surface, ICCV | Authors | SurfaceSplat: SDF+3DGS Hybrid Surface Reconstruction | ICCV 2025 | paper |
| 2025-07 | Sparse-View, Implicit, Voxel, Consistency | Authors | SparseRecon: Sparse-View Implicit Surface Reconstruction | ICCV 2025 | paper |
| 2025-07 | Low-Texture, Reflection, Unified, +21% | Authors | HiNeuS: Unified Neural Implicit Surface Reconstruction | ICCV 2025 | paper |
| 2025-06 | Joint Human+Scene, MASt3R Extension | Authors | HAMSt3R: Joint Human and Scene 3D Reconstruction | ICCV 2025 | paper |
| 2024-11-20 | 4D Reconstruction, Gaussian Splatting, Forward | Shanghai AI Lab | Forge4D: Gaussian Splatting for Forward Facing 4D Reconstruction | arXiv | paper |
This section tracks methods that create new 3D assets, parts, articulated objects, scenes, and editable 3D worlds, with emphasis on physical and simulation use.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-06-23 | Multi-View + LiDAR, Vehicle Assets, TRELLIS | Shanghai Jiao Tong University | MM-TRELLIS: Point-Cloud Guided Multi-Modal 3D Vehicle Generation in Autonomous Driving | arXiv | github |
| 2026-03-12 | Multi-View, SAM3D, Layout-Aware, Physical Plausibility | Peking University | MV-SAM3D: Adaptive Multi-View Fusion for Layout-Aware 3D Generation | arXiv | github |
| 2026-01-16 | Casual Capture, Posed Multi-View, Metric Shape | Meta AI | ShapeR: Robust Conditional 3D Shape Generation from Casual Captures | arXiv | paper |
| 2025-11-12 | Multi-Image Fusion, Region Control, TRELLIS | Zhejiang University | Fuse3D: Generating 3D Assets Controlled by Multi-Image Fusion | SIGGRAPH Asia 2025 | project / github |
| 2025-03-18 | Multi-View Image-to-Shape, Hunyuan3D-DiT | Tencent | Hunyuan3D 2.0 MV | Model | github / model |
| 2024-02-06 | Scalable View Synthesis, Single/Multi-Image 3D | KAUST | EscherNet: A Generative Model for Scalable View Synthesis | CVPR 2024 | project |
| 2019-08-05 | Multi-View Images, Mesh Deformation, Shape Refinement | National Tsing Hua University | Pixel2Mesh++: Multi-View 3D Mesh Generation via Deformation | ICCV 2019 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-25 | Relightable 3D Assets, PBR Maps, Gaussian Representation | Apple | Luce: Relightable Gaussians for 3D Asset Generation | arXiv | paper |
| 2026-08-24 | Material Decomposition, Physical Properties, Watertight Sub-Meshes, Sim-Ready | University of Bristol | Gen2Physics: Grounding Generated 3D Meshes in Physics via Multi-View Material Decomposition | arXiv | paper |
| 2026-07-01 | Complex Textures, Video Generative Prior, 3D Assets | Authors | Ink3D: Sculpting 3D Assets with Extremely Complex Textures via Video Generative Models | arXiv | paper |
| 2024-11-10 | VLM-Guided, PBR Texture | 3D AIGC | TexGaussian: Generating High-quality PBR Material via Octree-based 3D Gaussian Splatting | CVPR 2025 | project |
| 2024-01-17 | TextureDreamer, Geometry-Aware, Diffusion | UCSD / Meta | TextureDreamer: Image-Guided Texture Synthesis | CVPR 2024 | paper |
| 2023-12-21 | Texture Generation, Mesh, Multi-View Consistency | Tencent | Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models | CVPR 2024 | project |
| 2023-11-28 | SceneTex, Indoor, Texture, Diffusion, Highlight | TUM / Snap | SceneTex: High-Quality Texture Synthesis for Indoor Scenes | CVPR 2024 Highlight | paper |
| 2023-11-21 | SyncMVD, Multi-View, Text-to-Texture | CUHK | SyncMVD: Text-Guided Texturing by Synchronized Multi-View Diffusion | CVPR 2024 | paper |
| 2023-08-22 | PBR Material, SVBRDF, Text-to-Material | Adobe | MatFuse: Controllable Material Generation with Diffusion Models | SIGGRAPH Asia 2024 | project |
| 2023-03-20 | Texture, Material, Text-to-Texture | KAIST | Text2Tex: Text-driven Texture Synthesis via Diffusion Models | ICCV 2023 | project |
| 2023-02-03 | TEXTure, Text-Guided, 3D Texture, Diffusion | Tel Aviv University | TEXTure: Text-Guided Texturing of 3D Shapes | SIGGRAPH 2023 | project |
| 2022-07-06 | nvdiffrec, 3D Mesh, Material, Lighting, Oral | NVIDIA | nvdiffrec: Extracting Triangular 3D Models, Materials, and Lighting | CVPR 2022 Oral | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-08 | Single Image, Physical CoT, URDF, Simulation-Ready | Aerospace Information Research Institute, CAS | PhysX-CoT: Structured Physical Reasoning from a Single Image to Simulation-Ready 3D Assets | arXiv | paper |
| 2026-07-15 | Articulation + Physics, 40K Assets, Simulation-Ready | Zhejiang University | UniPhysGen: Unified Physical Grounding for Simulation-Ready 3D Assets | arXiv | github |
| 2026-05-20 | Rigid/Deformable/Articulated, Physical Attributes, Sim-Ready | Nanyang Technological University | PhysX-Omni: Unified Simulation-Ready Physical 3D Generation for Rigid, Deformable, and Articulated Objects | arXiv | project / dataset |
| 2026-05-14 | Agentic Generation, Articraft-10K, URDF Assets | Authors | Articraft: An Agentic System for Scalable Articulated 3D Asset Generation | arXiv | project / github |
| 2026-05-06 | Physics-Grounded, Kinematic, Simulation-Ready Assets | HKU | PhysForge: Generating Physics-Grounded 3D Assets for Interactive Virtual World | ICML 2026 | project / github |
| 2026-03-14 | URDF, Autoregressive, Simulation-Ready Assets | Authors | URDF-Anything+: End-to-End Generation for Simulation-Ready Articulated Assets | arXiv | paper |
| 2026-03-01 | Articulated Assets, 3D LLM, Kinematic Structure | Tsinghua | ArtLLM: Generating Articulated Assets via 3D LLM | CVPR 2026 | paper |
| 2025-12-12 | Articulation, Kinematic Tree, Feed-Forward, URDF-Ready | University of Oxford | Particulate: Feed-Forward 3D Object Articulation | arXiv | project |
| 2025-11-26 | Single Image, Open-Set Articulation, Unified Latent | ShanghaiTech | UniArt: Unified 3D Representation for Generating 3D Articulated Objects with Open-Set Articulation | arXiv | paper |
| 2025-11-17 | Sim-Ready Assets, Physical Properties, Single Image | NTU | PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image | CVPR 2026 | paper |
| 2025-11-02 | URDF, 3D MLLM, Articulated Objects | Tsinghua | URDF-Anything: Constructing Articulated Objects with 3D Multimodal Language Model | NeurIPS 2025 | paper |
| 2025-08-20 | Articulated Geometry, Motion Modeling, Gaussian Representation | Tsinghua University | GaussianArt: Unified Modeling of Geometry and Motion for Articulated Objects | 3DV 2026 | paper |
| 2025-07-16 | Physical Properties, Scale, Material, Affordance | Nanyang Technological University | PhysX-3D: Physical-Grounded 3D Asset Generation | NeurIPS 2025 Spotlight | project / github |
| 2025-06-10 | Interactable Digital Twin, Articulated Object, RGB-D Video | Shanghai Jiao Tong University | iTACO: Interactable Digital Twins of Articulated Objects from Casually Captured RGBD Videos | 3DV 2026 | paper |
| 2025-04-17 | Simulation-Ready, Physical Materials, Dynamics | University of Massachusetts Amherst | SOPHY: Generating Simulation-Ready Objects with Physical Materials | WACV 2026 | project / github |
| 2025-03-11 | Part-Level Digital Twin, Joint Estimation, Self-Supervised 3DGS | USTC | ArticulatedGS: Self-supervised Digital Twin Modeling of Articulated Objects using 3D Gaussian Splatting | CVPR 2025 | paper |
| 2025-02-26 | Articulated Objects, 3DGS, Joint Estimation | Tsinghua | ArtGS: Building Interactable Replicas of Complex Articulated Objects via Gaussian Splatting | ICLR 2025 | project |
| 2025-02-17 | Articulation-Ready, Skeleton, Skinning, Benchmark | Nanyang Technological University | MagicArticulate: Make Your 3D Models Articulation-Ready | CVPR 2025 | project / github |
| 2024-09-26 | Open-Vocabulary, URDF, Articulation | Stanford | Articulate Anything: Open-vocabulary 3D Articulated Object Generation | ICLR 2025 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-08-20 | Compositional 3D, Part Semantics, Spatial Control, Reassemblable Assets | Roblox | MultiCube: Compositional 3D Generation With Part-Level Semantic and Spatial Control | arXiv | project |
| 2026-08-14 | Up to 300 Parts, Token-Efficient VQ, Autoregressive 3D, Structured Assets | The University of Hong Kong | MegaParts: Scaling Part-Aware 3D Object Generation to 300 Parts via Token-Efficient Autoregressive Modeling | arXiv | project |
| 2026-08-13 | Part-Aware Generation, Recursive Decomposition, Editable Assets | Shanghai Jiao Tong University | SCULPT: Subtractive Composition for 3D Part Generation | arXiv | project |
| 2026-07-18 | Category-Agnostic, Neural Shape Editing, Coupled Representation | Authors | CNS-Edit++: Category-Agnostic 3D Editing with Coupled Neural Shape Representation | arXiv | paper |
| 2026-06-23 | Garment Patterns, Simulation-Ready, Editing | University of Hong Kong | PatternGSL: A Structured Specification Language for Template-Free and Simulation-Ready 3D Garments | arXiv | github |
| 2026-05-27 | Part-Controllable, Open-Vocabulary, Game-Ready | Roblox | CubePart: An Open-Vocabulary Part-Controllable 3D Generator | SIGGRAPH 2026 | project / model |
| 2025-09-10 | Part Decomposition, Editable, Production-Ready Assets | Tencent | X-Part: High-Fidelity and Structure-Coherent Shape Decomposition | Tech Report | project |
| 2025-08-14 | Rigging, Animation, Skeleton, Skinning | Nanyang Technological University | Puppeteer: Rig and Animate Your 3D Models | NeurIPS 2025 Spotlight | project / github |
| 2025-06-05 | Part-Level Mesh, Compositional DiT, Single Image | University of Waterloo | PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers | arXiv | project |
| 2024-12-16 | Articulated Mesh, Part-by-Part, Hierarchical Transformer | Cornell | MeshArt: Generating Articulated Meshes with Structure-Guided Transformers | arXiv | paper |
| 2023-12-13 | Shape Program, Structure, Editable Assets | MIT | Shape2Program: Learning to Infer Shape Programs from 3D Shapes | arXiv | project |
| 2023-06-29 | Part-Aware, Shape Assembly, 3D Generation | Shanghai AI Lab | Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent Representation | NeurIPS 2023 | github |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-09-10 | Single Image, Executable Scene Programs, Recursive Construction, Editable 3D | Georgia Institute of Technology | Recursive Code World Models: Building Complex Worlds through Recursive Scene Programs | arXiv | paper |
| 2026-09-05 | Agentic 3D Composition, Functional Objects, Executable Robot Scenes | Peking University / Galbot | GIF: Agentic Generation of Interactive and Functional Object Compositions for Robot Learning | arXiv | paper |
| 2026-09-04 | Image-to-Scene, Agentic Layout Evolution, Simulation-Ready Diversity | The University of Hong Kong | SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution | arXiv | project / github |
| 2026-08-31 | Text-to-Scene, Grow-and-Repair, Functional Groups, SceneReverse-17K | Southeast University | ScenePilot: Grow-and-Repair Policy for Text-Driven 3D Indoor Scene Generation | arXiv | project |
| 2026-08-27 | Single Image, Generative 3D Proxy, RGB-D World Expansion, Explorable Scene | Hong Kong University of Science and Technology | SpatialCrafter: Single Image World Modeling with Generative 3D Proxies | arXiv | project |
| 2026-08-25 | Monocular Image, Interactive Scene Programming, Articulation+Physics, Embodied Simulation | Shanghai Jiao Tong University | NeoWorld-Pro: Programming Interactive Scenes from Monocular Images for Embodied Simulation | arXiv | project |
| 2026-08-19 | Usage-Driven Code Scenes, Multi-Part Interaction, Executable Simulation | Shanghai Jiao Tong University | Beyond Placement and Articulation: Usage-Driven Code Scenes for Embodied Interaction | arXiv | paper |
| 2026-07-29 | Panoramic Video, 3DGS, Simulation-ready World | AgiBot | Genie Sim PanoWorld: An Infinite Indoor 3D World Generation Pipeline via Panoramic Scene Modeling and Simulation | arXiv | github |
| 2026-07-15 | Indoor Layout, Progressive VLM Reasoning, Interactive Editing | City University of Hong Kong | ThinkBLOX: 3D Indoor Scene Generation with Progressive Reasoning | arXiv | paper |
| 2026-07-08 | Simulation-Ready Assets, Affordances, Cross-Simulator Worlds | Horizon Robotics | EmbodiedGen V2: An Agentic, Simulation-Ready 3D World Engine for Embodied AI | arXiv | project / github |
| 2026-07-07 | Real-to-Sim, One-Shot Scene Generation, Robot Evaluation | Shanghai AI Lab | RoboSnap: One-Shot Real-to-Sim Scene Generation for Generalizable Robot Learning and Evaluation | arXiv | project |
| 2026-07-04 | Egocentric Scene Generation, Geometric 3DGS, Consistency | South China Univ. of Technology | CGGS: Consistency-Augmented Geometric Gaussian Splatting for Ego-centric 3D Scene Generation | arXiv | paper |
| 2026-06-23 | Triangle Splatting, Single-Image Scene, Game-Ready | Google Research | FLAT: Feedforward Latent Triangle Splatting for Geometrically Accurate Scene Generation | arXiv | project |
| 2026-06-23 | Text-to-Scene, Video Priors, 3DGS Orbit | University of Bern | OrbitForge: Text-to-3D Scene Generation via Reconstruction-Anchored Video Synthesis | arXiv | paper |
| 2026-06-23 | Compositional 3D, Physical Interaction, Multi-View Consistency | China University of Petroleum (East China) | Inclusive Interactive Collisions for Multi-View Consistent Compositional 3D Generation | arXiv | paper |
| 2026-06-23 | Satellite-to-City, Textured Mesh, Urban Simulation | HKUST(GZ) | Sat2City v2: Native 3D City Asset Generation from a Single Satellite Image | arXiv | paper |
| 2026-06-08 | Satellite-to-3D, 3DGS, UAV Simulation | Amap-cvlab / Alibaba | ABot-Earth 0.5: Generative 3D Earth Model | arXiv | project |
| 2026-06-04 | Whole-Home Scenes, Floorplans, Interactive | Authors | HomeWorld: A Unified Floorplan-to-Furnished Framework for Generating Controllable, Densely Interactive Whole-Home Scenes | arXiv | paper |
| 2026-05-28 | Physical Stability, Single Image, Scene Tree, Simulation | Carnegie Mellon University | REST3D: Reconstructing Physically Stable 3D Scenes from a Single Image | arXiv | project |
| 2026-05-01 | Segment Map, Text-to-World, Controllable 3D Worlds | Seoul National University | Map2World: Segment Map Conditioned Text to 3D World Generation | arXiv | paper |
| 2026-04-14 | Explorable 3D Worlds, Long Trajectory, 3DGS | NVIDIA | Lyra 2.0: Explorable Generative 3D Worlds | arXiv | project |
| 2026-04-06 | Single-Image Scene, In-Place Completion, ARSG-110K | Nankai University | 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image | CVPR 2026 | project / github / dataset |
| 2026-03-31 | Town-Scale, Single Image, Latent Extension | Seoul National University | Extend3D: Town-Scale 3D Generation | CVPR 2026 | project |
| 2026-03-31 | Unbounded World, Flow Matching, Layouts | Princeton University | WorldFlow3D: Flowing Through 3D Distributions for Unbounded World Generation | arXiv | project |
| 2026-03-27 | Autoregressive 3DGS, Token Generation, Completion/Outpainting | Technical University of Munich | GaussianGPT: Towards Autoregressive 3D Gaussian Scene Generation | arXiv | project |
| 2026-03-12 | Multi-Floor, Language-to-3D, Long-Horizon Tasks | Tsinghua | MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks | CVPR 2026 | paper |
| 2026-03-06 | Compositional Scene, Panoramic Image, Feed-Forward | NTU | Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic Image | CVPR 2026 | paper |
| 2026-02-10 | Agentic Scene Generation, Sim-Ready, SAGE-10k | NVIDIA | SAGE: Scalable Agentic 3D Scene Generation for Embodied AI | CVPR 2026 | project / github |
| 2026-01-09 | Language-Guided, Infinite Worlds, Articulated Furniture | National Taiwan Univ | SceneFoundry: Generating Interactive Infinite 3D Worlds | arXiv | project |
| 2025-12-01 | Tabletop, Instance-Level, Interactive Scene, Text/Image | D-Robotics | TabletopGen: Instance-Level Interactive 3D Tabletop Scene Generation from Text or Single Image | arXiv | project / paper |
| 2025-11-18 | Single-Image Scene, Gaussian World, Scene Generation | Tsinghua University | GEN3D: Generating Domain-Free 3D Scenes from a Single Image | arXiv | paper |
| 2025-09-18 | Layout-Guided, Indoor Scenes, Decoupled Geometry/Appearance | Hong Kong University of Science and Technology | SPATIALGEN: Layout-guided 3D Indoor Scene Generation | 3DV 2026 | paper |
| 2025-08-21 | Single Image, Multi-Asset Scene, Feed-Forward | Shanghai Jiao Tong University | SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass | 3DV 2026 | project / github |
| 2025-08-11 | Panoramic, Explorable World, Matrix-Pano | Kunlun Wanwei | Matrix-3D: Omnidirectional Explorable 3D World Generation | arXiv | project / github |
| 2025-07-29 | Panoramic, Text/Image-to-World, Mesh Export | Tencent Hunyuan | HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels | arXiv | github |
| 2025-07-09 | VLA Scene Authoring, Simulation-Ready Worlds, Synthetic Data | NVIDIA / Stanford University | 3D-Generalist: Self-Improving Vision-Language-Action Models for Crafting 3D Worlds | 3DV 2026 | project / github |
| 2025-06-25 | Explorable Scene, Novel-View Restoration, Consistency | Beijing Academy of AI | WonderFree: Enhancing Novel View Quality and Cross-View Consistency for 3D Scene Exploration | arXiv | paper |
| 2025-03-13 | Scene Layout, Optimization, Generation | Tsinghua | HOG-Layout: Layout-Enhanced Scene Generation via Hierarchical Optimization | arXiv | project |
| 2025-02-20 | Octree, 3D Diffusion, Scene Generation | Zhejiang University | Octree Diffusion: Hierarchical Scene Generation via Octree Structures | arXiv | project |
| 2025-01-15 | Gaussian, GPT, Scene Generation | Shanghai AI Lab | GaussianGPT: Language-Driven Scene Generation with Gaussian Representation | arXiv | paper |
| 2024-12-19 | Scene Generation, Growing, Incremental | Tsinghua | WorldGrow: Incremental 3D Scene Generation | CVPR 2025 | project |
| 2024-11-05 | Splatting, Fluents, Scene Understanding | University of Cambridge | FluSplat: Fluent Scene Generation via Gaussian Splatting | CVPR 2025 | project |
| 2024-11-04 | GenXD, Any 3D and 4D, Scene Generation | NUS | GenXD: Generating Any 3D and 4D Scenes | ICLR 2025 | paper |
| 2024-06-17 | Procedural Scenes, Synthetic Data, Embodied AI | Princeton | Infinigen Indoors: Photorealistic Indoor Scenes using Procedural Generation | NeurIPS 2024 | project / github |
| 2024-06-13 | Interactive 3D Scene, FLAGS, Single Image | Stanford University | WonderWorld: Interactive 3D Scene Generation from a Single Image | CVPR 2025 | project |
| 2024-05-02 | EchoScene, Scene Graph, Diffusion, Indoor | TUM / JHU | EchoScene: Indoor Scene Generation via Information Echo | ECCV 2024 | paper |
| 2024-02-12 | SceneScape, Text-Driven, Consistent, Scene | Weizmann | SceneScape: Text-Driven Consistent Scene Generation | NeurIPS 2024 | paper |
| 2024-01-30 | BlockFusion, Expandable, Tri-plane, SIGGRAPH | Tencent / UTokyo | BlockFusion: Expandable 3D Scene Generation | SIGGRAPH 2024 (ACM TOG) | paper |
| 2023-12-01 | ControlRoom3D, Semantic Proxy, Room Generation | TUM / Meta | ControlRoom3D: Room Generation using Semantic Proxy Rooms | CVPR 2024 | paper |
| 2023-10-05 | Ctrl-Room, Text-to-3D, Layout Constraints | Simon Fraser | Ctrl-Room: Controllable Text-to-3D Room Meshes Generation | ECCV 2024 | paper |
| 2023-10-04 | MagicDrive, Street View, 3D Geometry Control | CUHK / HKUST | MagicDrive: Street View Generation with Diverse 3D Geometry Control | ICLR 2024 | project |
| 2023-06-15 | Procedural World, Synthetic Data, Simulation | Princeton | Infinite Photorealistic Worlds using Procedural Generation | CVPR 2023 | project |
| 2023-03-24 | DiffuScene, Diffusion, Indoor Scene Synthesis | TUM | DiffuScene: Denoising Diffusion for Generative Indoor Scene Synthesis | CVPR 2024 | paper |
| 2023-03-21 | Text-to-3D Room, Indoor Scenes, Mesh | LMU Munich | Text2Room: Extracting Textured 3D Meshes from 2D Text-to-Image Models | ICCV 2023 | project |
| 2023-02-02 | Unbounded 3D Scene, Generative Model, Driving | NVIDIA | SceneDreamer: Unbounded 3D Scene Generation from 2D Image Collections | CVPR 2023 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-18 | Multi-View Generation, Scene Assets, Training-Free | The University of Queensland | Scene-SAM3D: Multi-View Scene Asset Generation Without Fine-Tuning | arXiv | github |
| 2025-01-03 | Rover, Semantic, 3D Scene | Carnegie Mellon University | SEM-ROVER: Semantic Scene Exploration with Hierarchical Spatial Reasoning | ICLR 2025 | project |
| 2024-09-30 | Spatial, Generation, Language | Tsinghua | SpatialGen: Language-Driven Spatial Scene Generation | NeurIPS 2024 | project |
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-11 | Layout-Conditioned, 3DGS, Mixed Reality | Authors | SyncSpace: Layout-Conditioned 3D Gaussian Splatting for Space Reskinning in Mixed Reality | arXiv | paper |
| 2025-04-17 | Image-Pair-to-4D, Diffusion, Explicit 3D Motion | Technical University of Munich | TwoSquared: 4D Generation from 2D Image Pairs | 3DV 2026 Oral | paper |
| 2024-12-05 | 4Real-Video, Photo-Realistic, Video Diffusion, CVPR | Snap / KAUST | 4Real-Video: Generalizable Photo-Realistic 4D Video Diffusion | CVPR 2025 | paper |
| 2024-07-16 | Animate3D, Multi-View, Video Diffusion, NeurIPS | CASIA / Alibaba | Animate3D: Animating Any 3D Model with Multi-view Video Diffusion | NeurIPS 2024 | paper |
| 2024-05-31 | 4Diffusion, Multi-View Video, 4D, NeurIPS | CASIA / Shanghai AI Lab | 4Diffusion: Multi-view Video Diffusion Model for 4D Generation | NeurIPS 2024 | paper |
| 2024-05-26 | Diffusion4D, Video Diffusion, 4D, NeurIPS | Toronto / BJTU | Diffusion4D: Fast Spatial-temporal Consistent 4D Generation | NeurIPS 2024 | paper |
| 2024-05-03 | DreamScene4D, Multi-Object, Dynamic, NeurIPS | CMU | DreamScene4D: Dynamic Multi-Object Scene Generation | NeurIPS 2024 | paper |
| 2024-03-22 | STAG4D, Spatial-Temporal, 4D Gaussians, ECCV | Nanjing / CASIA | STAG4D: Spatial-Temporal Anchored Generative 4D Gaussians | ECCV 2024 | paper |
| 2023-11-29 | 4D-fy, Text-to-4D, Score Distillation, CVPR | KAUST / Snap | 4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling | CVPR 2024 | project |
| 2023-11-24 | Animate124, Image-to-4D, Animation, ICLR | NUS / Huawei | Animate124: Animating One Image to 4D Dynamic Scene | ICLR 2024 | paper |
| 2023-11-17 | Consistent4D, 360° Dynamic, Monocular Video, ICLR | CASIA / Nanjing | Consistent4D: Consistent 360° Dynamic Object Generation | ICLR 2024 | project |
3D editing covers methods that modify existing 3D assets, Gaussian / NeRF fields, meshes, voxel or latent states, and dynamic scenes. Entries are grouped by editable state: object-level, scene-level, and dynamic / 4D.
| Date | Keywords | Institute (first) | Paper / Resource | Publication | Others |
|---|---|---|---|---|---|
| 2026-07-27 | NeRF Editing, Object Removal, Robot Manipulation | Authors | NEO: NeRF It Once, Edit It Many Times for Continuous Object Manipulation | arXiv | paper |
| 2026-06-05 | Mesh Editing, Image-Guided, Local Morphing | Leiden University | 3DMorph: Single-Image-Guided Local 3D Shape Editing and Morphing | IJCNN 2026 | github |
| 2026-05-26 | PartFlow, Semantic-Part Transformation, Mask-Free | Nanyang Technological University | Feedforward 3D Editing Learns from Semantic-Part Transformation | arXiv | project / github / benchmark (Steer3D) |
| 2026-05-08 | VS3D, Velocity-Space, Mask-Free | Tsinghua University | Velocity-Space 3D Asset Editing | arXiv | paper |
| 2026-05-01 | Latent Editing, Object-Level, Structured 3D Latents | Seoul National University | InpaintSLat: Inpainting Structured 3D Latents via Initial Noise Optimization | arXiv | project |
| 2026-04-30 | MeshReGen, VecSet Regeneration, Image-Guided Editing | KAIST | MeshReGen: A Unified 3D Geometry Regeneration Framework | arXiv | project / benchmark (VoxHammer) |
| 2026-04-26 | Primitive Proxy, Shape Editing, Fine-Grained Control | Tel Aviv University | Prox-E: Fine-Grained 3D Shape Editing via Primitive-Based Abstractions | SIGGRAPH 2026 | project / github / benchmark (VoxHammer) |
| 2026-03-30 | 3DGS Editing, Object-Level, Single-View | Zhejiang Gongshang University | SVGS: Single-View to 3D Object Editing via Gaussian Splatting | ACM TOMM 2026 | project |
| 2026-02-25 | Voxel Editing, Object-Level, Rectified Voxel Flow | USTC | Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow | CVPR 2026 | project |
| 2026-02-05 | Native 3D Editing, Image-Conditioned, Latent-to-Latent | Aigency.ai / Tel Aviv University | ShapeUP: Scalable Image-Conditioned 3D Editing | SIGGRAPH 2026 | project / github |
| 2026-02-04 | Mesh Editing, Object-Level, Single-Image LRM | National Yang Ming Chiao Tung University | VecSet-Edit: Unleashing Pre-trained LRM for Mesh Editing from Single Image | arXiv | github / benchmark (VoxHammer) |
| 2025-12-15 | Steer3D, Text-Steerable, Edit3D-Bench (Ma) | Caltech | Feedforward 3D Editing via Text-Steerable Image-to-3D | arXiv | project / github / benchmark (Steer3D) |
| 2025-11-27 | Latent Anchor, Object-Level, Mask-Free | Zhejiang University | AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows | CVPR 2026 Oral | project / github / benchmark |
| 2025-11-21 | Native Editing, Object-Level, Full Attention | Fudan University / StepFun | Native 3D Editing with Full Attention | arXiv | paper |
| 2025-10-16 | FlowEdit, Object-Level, Mask-Free | Tsinghua University | NANO3D: A Training-Free Approach for Efficient 3D Editing Without Masks | ICLR 2026 | project / github / dataset |
| 2025-10-03 | 3DEditFormer, Paired Dataset, Mask-Free | East China Univ of Science & Technology | Towards Scalable and Consistent 3D Editing | arXiv | project / github |
| 2025-08-29 | 3D-LATTE, Text Instructions, 3D Diffusion Latent | University of Tübingen | 3D-LATTE: Latent Space 3D Editing from Textual Instructions | CVPR 2026 Oral | project / CVF |
| 2025-08-26 | VoxHammer, Edit3D-Bench (Li), Training-Free | Renmin University / Beihang University | VoxHammer: Training-Free Precise and Coherent 3D Editing in Native 3D Space | 3DV 2026 Oral | project / github / benchmark (VoxHammer) |
| 2025-07-15 | 3DGS Editing, Part-Level, Regularized SDS | Seoul National University | Robust 3D-Masked Part-level Editing in 3D Gaussian Splatting with Regularized Score Distillation Sampling | ICCV 2025 | CVF |
| 2025-06-25 | Image Prompt, Multi-View Propagation, Mask-Free | Tel Aviv University | EditP23: 3D Editing via Propagation of Image Prompts to Multi-View | ACM TOG 2025 | project / github |
| 2025-05-11 | CMD, Local Editing, Progressive Generation | HKUST | CMD: Controllable Multiview Diffusion for 3D Editing and Progressive Generation | SIGGRAPH 2025 | project |
| 2024-12-11 | Mesh Editing, Object-Level, Masked LR |
Truncated — view the full README on GitHub.
98 commits
53 commits
Python
100.0%