🌟 A curate list of papers, datasets, and projects for 3D Reconstruction and Generation.
95
43 commits
updated Mar 13, 2026
🌟 A curate list of papers, datasets, and projects for 3D Reconstruction and Generation.
:heart: Pull requests or issues for updates are very welcome.
| Title | Date | Link | Venue |
|---|---|---|---|
Emergent Extreme-View Geometry in 3D Foundation Models Abstract3D foundation models (3DFMs) have recently transformed 3D vision, enabling joint prediction of depths, poses, and point maps directly from images. Yet their ability to rea- son under extreme, non-overlapping views remains largely unexplored. In this work, we study their internal repre- sentations and find that 3DFMs exhibit an emergent un- derstanding of extreme-view geometry, despite never being trained for such conditions. To further enhance these capa- bilities, we introduce a lightweight alignment scheme that refines their internal 3D representation by tuning only a small subset of backbone bias terms, leaving all decoder heads frozen. This targeted adaptation substantially im- proves relative pose estimation under extreme viewpoints without degrading per-image depth or point quality. | 11/2025 | project | CVPR 2026 |
Emergent Outlier View Rejection in Visual Geometry Grounded Transformers AbstractReliable 3D reconstruction from in-the-wild image collec- tions is often hindered by “noisy” images—irrelevant in- puts with little or no view overlap with others. While tra- ditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed- forward 3D reconstruction models lack these explicit mech- anisms, leading to degraded performance under in-the-wild conditions. In this paper, we discover that the existing feed- forward reconstruction model, e.g., VGGT, despite lacking explicit outlier-rejection mechanisms or noise-aware train- ing, can inherently distinguish distractor images. | 12/2025 | project | CVPR 2026 |
AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend AbstractWe present AMB3R, a multi-view feed-forward model for dense 3D reconstruction on a metric-scale that addresses diverse 3D vision tasks. The key idea is to leverage a sparse, yet compact, volumetric scene representation as our backend, enabling geometric reasoning with spatial compactness. Although trained solely for multi-view re- construction, we demonstrate that AMB3R can be seam- lessly extended to uncalibrated visual odometry (online) or large-scale structure from motion without the need for task- specific fine-tuning or test-time optimization. | 11/2025 | project | CVPR 2026 |
Depth Anything 3: Recovering the Visual Space from Any Views AbstractWe present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. | 11/2025 | project | ICLR 2026 |
| Dens3R: A Foundation Model for 3D Geometry Prediction | 2025 | ICLR 2026 | |
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting AbstractWe present WorldMirror, an all-in-one, feed-forward model for versatile 3D geo- metric prediction tasks. Unlike existing methods constrained to image-only inputs or customized for a specific task, our framework flexibly integrates diverse geo- metric priors, including camera poses, intrinsics, and depth maps, while simulta- neously generating multiple 3D representations: dense point clouds, multi-view depth maps, camera parameters, surface normals, and 3D Gaussians. This elegant and unified architecture leverages available prior information to resolve structural ambiguities and delivers geometrically consistent 3D outputs in a single forward pass. | 10/2025 | arXiv | |
PAGE-4D: Disentangled Pose and Geometry for 4D Perception AbstractRecent 3D feed-forward models, such as the Visual Geometry Grounded Trans- former (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these mod- els often struggle in real-world scenarios involving complex dynamic elements, such as moving humans or deformable objects like umbrellas. To address this limitation, we introduce PAGE-4D, a feedforward model that extends VGGT to dynamic scenes, enabling camera pose estimation, depth prediction and point cloud reconstruction —all without post-processing. | 10/2025 | ICLR 2026 | |
DriveVGGT: Visual Geometry Transformer for Autonomous Driving AbstractFeed-forward reconstruction has recently gained significant attention, with VGGT being a notable example. However, directly applying VGGT to autonomous driving (AD) sys- tems leads to sub-optimal results due to the different priors between the two tasks. In AD systems, several important new priors need to be considered: (i) The overlap between camera views is minimal, as autonomous driving sensor se- tups are designed to achieve 360◦coverage at a low cost. (ii) The camera intrinsics and extrinsics are known, which introduces more constraints on the output and also enables the estimation of absolute scale. (iii) Relative positions of all cameras remain fixed though the ego vehicle is in motion. | 11/2025 | ECCV 2025 | |
π3: Scalable Permutation-Equivariant Visual Geometry Learning AbstractWe introduce π3, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, π3 em- ploys a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design makes our model inherently robust to input ordering and highly scalable. | 07/2025 | ICLR 2026 | |
BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection AbstractOpen-vocabulary 3D object detection has gained sig- nificant interest due to its critical applications in au- tonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial com- putational overhead and memory constraints, hindering real-time deployment in downstream tasks. To address this, we propose a novel reconstruction-free online framework tailored for memory-efficient and real-time 3D detection. Specifically, given streaming posed RGB-D video input, we leverage Cubify Anything as a pre-trained visual foundation model (VFM) for single-view 3D object detection by bound- ing boxes, coupled with CLIP to capture open-vocabulary semantics of detected objects. | 06/2025 | arXiv | |
Depth Anything with Any Prior AbstractThis work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap be- tween prior patterns, enhancing generalization across vary- ing scenarios. Second, we develop a conditioned monoc- ular depth estimation (MDE) model to refine the inherent noise of depth priors. | 05/2025 | project | ICML 2025 |
| MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models | 05/2025 | code | CVPR 2025 |
VGGT:Visual Geometry Grounded Transformer AbstractWe present VGGT, a feed-forward neural network that di- rectly infers all key 3D attributes of a scene, including cam- era parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and special- ized for single tasks. It is also simple and efficient, re- constructing images in under one second, and still out- performing alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, in- cluding camera parameter estimation, multi-view depth es- timation, dense point cloud reconstruction, and 3D point tracking. | 03/2025 | code | CVPR 2025 |
| MUSt3R: Multi-view Network for Stereo 3D Reconstruction | 03/2025 | CVPR 2025 | |
| Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | 03/2025 | CVPR 2025 | |
| FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | 02/2025 | code | CVPR 2025 |
EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose Estimation AbstractRelative camera pose estimation between two images is a fundamental task in 3D computer vision. Recently, many relative pose estimation networks have been explored for learning a mapping from two input images to their cor- responding relative pose, however, the estimated relative poses by these methods do not have the intrinsic Pose Per- mutation Equivariance (PPE) property: the estimated rel- ative pose from Image A to Image B should be the inverse of that from Image B to Image A. It means that permut- ing the input order of two images would cause these meth- ods to obtain inconsistent relative poses. To address this problem, we firstly introduce the concept of PPE mapping, which indicates such a mapping that captures the intrin- sic PPE property of relative poses. | 02/2025 | CVPR 2025 | |
RePoseD: Efficient Relative Pose Estimation With Known Depth Information AbstractRecent advances in monocular depth estimation meth- ods (MDE) and their improved accuracy open new possibil- ities for their applications. In this paper, we investigate how monocular depth estimates can be used for relative pose es- timation. In particular, we are interested in answering the question whether using MDEs improves results over tradi- tional point-based methods. We propose a novel framework for estimating the relative pose of two cameras from point correspondences with associated monocular depths. Since depth predictions are typically defined up to an unknown scale or even both unknown scale and shift parameters, our solvers jointly estimate the scale or both the scale and shift parameters along with the relative pose. | 01/2025 | ICCV 2025 | |
| Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | 01/2025 | code | CVPR 2025 |
Continuous 3D Perception Model with Persistent State AbstractWe present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolv- ing state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene re- construction that updates as new images arrive. | 01/2025 | code | CVPR 2025 (Oral) |
| Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization | 12/2024 | code | CVPR 2025 |
| MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | 12/2024 | code | CVPR 2025 |
Relative Pose Estimation through Affine Corrections of Monocular Depth Priors AbstractMonocular depth estimation (MDE) models have under- gone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metric (absolute) depth. However, effectively leveraging these predictions for geometric vision tasks, in particular relative pose estimation, remains rela- tively under explored. While depths provide rich constraints for cross-view image alignment, the intrinsic noise and am- biguity from the monocular depth priors present practical challenges to improving upon classic keypoint-based so- lutions. | 12/2024 | code | CVPR 2025 |
Understanding multi-view transformers AbstractMulti-view transformers such as DUSt3R [59] are revolu- tionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the inner mechanisms of multi-view transformers are unclear. Their black-box nature makes further improve- ments beyond data scaling challenging and complicates us- age in safety- and reliability-critical applications. Here, we present an approach for probing and visualizing 3D repre- sentations from the residual connections of the multi-view transformers’ layers. | 2025 | code | ICCV 2025 |
Peering into the Unknown: Active View Selection for 3D Reconstruction AbstractImagine trying to understand the shape of a teapot by viewing it from the front—you might see the spout, but completely miss the handle. Some perspectives naturally provide more information than others. How can an AI system determine which viewpoint offers the most valuable insight for accurate and efficient 3D object reconstruction? Active view selection (AVS) for 3D reconstruction remains a fundamental challenge in computer vision. The aim is to identify the minimal set of views that yields the most accurate 3D reconstruction. | 2025 | ICLR 2026 | |
| 3D Reconstruction with Spatial Memory | 08/2024 | code | 3DV 2025 |
Grounding Image Matching in 3D with MASt3R AbstractImage Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D problem. This makes sense as the goal of matching is to establish correspondences between 2D pixel fields, but also seems like a potentially hazardous choice. In this work, we take a different stance and propose to cast matching as a 3D task with DUSt3R, a recent and powerful 3D reconstruction framework based on Transformers. Based on pointmaps regression, this method displayed impressive robustness in matching views with extreme viewpoint changes, yet with limited accuracy. | 06/2024 | code | Arxiv 2024 |
| UniDepthV2: Universal Monocular Metric Depth Estimation | 06/2024 | CVPR 2024 | |
DUSt3R: Geometric 3D Vision Made Easy AbstractMulti-view stereo reconstruction (MVS) in the wild re- quires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangu- late corresponding pixels in 3D space, which is the core of all best performing MVS algorithms. In this work, we take an opposite stance and introduce DUSt3R1, a radi- cally novel paradigm for Dense and Unconstrained Stereo 3D Reconstruction of arbitrary image collections, i.e. oper- ating without prior information about camera calibration nor viewpoint poses. We cast the pairwise reconstruction problem as a regression of pointmaps, relaxing the hard constraints of usual projective camera models. We show 1https://dust3r.europe.naverlabs. | 12/2023 | code | CVPR 2024 |
RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation Abstract—Dense metric depth estimation using millimeter- wave radar typically requires dense LiDAR supervision, gen- erated via multi-frame projection and interpolation, to guide the learning of accurate depth from sparse radar measurements and RGB images. However, this paradigm is both costly and data- intensive. To address this, we propose RaCalNet, a novel frame- work that eliminates the need for dense supervision by using sparse LiDAR to supervise the learning of refined radar mea- surements, resulting in a supervision density of merely around 1% compared to dense-supervised methods. Unlike previous approaches that associate radar points with broad image regions and rely heavily on dense labels, RaCalNet first recalibrates and refines sparse radar points to construct accurate depth priors. | 06/2025 | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
GPA-VGGT: Adapting VGGT to Large scale Localization by Self-Supervised Learning Abstract—Transformer-based general visual geometry frame- works have shown promising performance in camera pose estima- tion and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise in camera pose estimation and 3D reconstruction. However, these models typically rely on ground truth labels for training, posing challenges when adapting to unlabeled and unseen scenes. In this paper, we propose a self-supervised framework to train VGGT with unlabeled data, thereby en- hancing its localization capability in large-scale environments. To achieve this, we extend conventional pair-wise relations to sequence-wise geometric constraints for self-supervised learning. | 01/2026 | code | arXiv |
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams AbstractThe grand vision of enabling persistent, large-scale 3D visual geometry understanding is shackled by the irrec- oncilable demands of scalability and long-term stability. While offline models like VGGT achieve inspiring geom- etry capability, their batch-based nature renders them ir- relevant for live systems. Streaming architectures, though the intended solution for live operation, have proven inade- quate. Existing methods either fail to support truly infinite- horizon inputs or suffer from catastrophic drift over long sequences. We shatter this long-standing dilemma with In- finiteVGGT, a causal visual geometry transformer that op- erationalizes the concept of a rolling memory through a bounded yet adaptive and perpetually expressive KV cache. | 01/2026 | code | arXiv |
AVGGT: Rethinking Global Attention for Accelerating VGGT AbstractSince DUSt3R, models such as VGGT and π3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention con- tributes to multi-view reasoning. In this paper, we first con- duct an in-depth investigation of the global attention mod- ules in VGGT and π3 to better understand their roles. Our analysis reveals a clear division of roles in the alternating *Equal contribution. †Corresponding author. global-frame architecture: early global layers do not form meaningful correspondences, middle layers perform cross- view alignment, and last layers provide only minor refine- ments. | 12/2025 | arXiv | |
FlashVGGT: Efficient and Scalable Visual Geometry Transformers Abstract3D reconstruction from multi-view images is a core chal- lenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to tra- ditional per-scene optimization techniques. Among them, state-of-the-art models like the Visual Geometry Ground- ing Transformer (VGGT) leverage full self-attention over all image tokens to capture global relationships. However, this approach suffers from poor scalability due to the quadratic complexity of self-attention and the large number of tokens generated in long image sequences. In this work, we in- troduce FlashVGGT, an efficient alternative that addresses this bottleneck through a descriptor-based attention mecha- nism. | 12/2025 | project | arXiv |
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging Abstract3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To ad- dress this, we propose LiteVGGT, achieving up to 10× speedup and substantial memory reduction, enabling effi- cient processing of 1000-image scenes. We derive two key insights for 3D reconstruction: 1) tokens from local im- age regions have inherent geometric correlations, leading to high similarity and computational redundancy; 2) token similarity acroses adjacent network layers remains stable, allowing for reusable merge decisions. | 12/2025 | project | arXiv |
HTTM: Head-wise Temporal Token Merging for Faster VGGT AbstractThe Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruc- tion, as it is the first model that directly infers all key 3D at- tributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism re- quires global attention layers that perform all-to-all atten- tion computation on tokens from all views. For reconstruc- tion of large scenes with long-sequence inputs, this causes a significant latency bottleneck. In this paper, we pro- pose head-wise temporal merging (HTTM), a training-free 3D token merging method for accelerating VGGT. | 11/2025 | arXiv | |
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer AbstractFoundation models for 3D vision have recently demon- strated remarkable capabilities in 3D perception. How- ever, scaling these models to long-sequence image inputs remains a significant challenge due to inference-time in- *Corresponding Author. efficiency. In this work, we present a detailed analysis of VGGT, a state-of-the-art feed-forward visual geometry model and identify its primary bottleneck. Visualization fur- ther reveals a token collapse phenomenon in the attention maps. Motivated by these findings, we explore the poten- tial of token merging in the feed-forward visual geometry model. Owing to the unique architectural and task-specific properties of 3D models, directly applying existing merg- arXiv:2509.02560v2 [cs.CV] 9 Nov 2025 ing techniques proves challenging. | 09/2025 | project | ICLR 2026 |
Faster VGGT with Block-Sparse Global Attention AbstractEfficient and accurate feed-forward multi-view recon- struction has long been an important task in computer vi- sion. Recent transformer-based models like VGGT and π3 have achieved impressive results with simple architectures, yet they face an inherent runtime bottleneck, due to the quadratic complexity of the global attention layers, that lim- its the scalability to large image sets. In this paper, we em- pirically analyze the global attention matrix of these models and observe that probability mass concentrates on a small subset of patch-patch interactions that correspond to cross- view geometric matches. | 09/2025 | arXiv | |
VGGT-Long: Chunk it, Loop it, Align it – Pushing VGGT's Limits on Long RGB Sequences AbstractFoundation models for 3D vision have recently demon- strated remarkable capabilities in 3D perception. How- ever, extending these models to large-scale RGB stream 3D reconstruction remains challenging due to memory lim- itations. In this work, we propose VGGT-Long, a sim- ple yet effective system that pushes the limits of monocu- 1† Corresponding author. lar 3D reconstruction to kilometer-scale, unbounded out- door environments. Our approach addresses the scalability bottlenecks of existing models through a chunk-based pro- cessing strategy combined with overlapping alignment and lightweight loop closure optimization. Without requiring camera calibration, depth supervision or model retraining, VGGT-Long achieves trajectory and reconstruction perfor- mance comparable to traditional methods. | 07/2025 | code | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
MIRROR: Make Your Object-Level Multi-View Generation Consistent AbstractMulti-view Diffusion has greatly advanced the de- velopment of 3D content creation by generating multiple images from distinct views, achieving remarkable photorealistic results. However, ex- isting works are still vulnerable to inconsistent 3D geometric structures (commonly known as Janus Problem) and severe artifacts. In this paper, we introduce MIRROR, a versatile plug-and-play method that rectifies such inconsistencies in a training-free manner, enabling the acquisition of high-fidelity, realistic structures without compro- mising diversity. Our key idea focuses on tracing the motion trajectory of physical points across adjacent viewpoints, enabling rectifications based on neighboring observations of the same region. | 2025 | ArXiV | |
Refine Any Object in Any Scene (RAISE) AbstractViewpoint missing of objects is common in scene reconstruction, as camera paths typically prioritize capturing the overall scene structure rather than individual objects. This makes it highly challenging to achieve high-fidelity object-level mod- eling while maintaining accurate scene-level representation. Addressing this issue is critical for advancing downstream tasks requiring detailed object understanding and appearance modeling. In this paper, we introduce Refine Any object In any ScenE (RAISE), a novel 3D enhancement framework that leverages 3D generative priors to recover fine-grained object geometry and appearance under missing views. | 06/2025 | code | arXiv |
| GenFusion: Closing the Loop between Reconstruction and Generation via Videos | 03/2025 | code | CVPR 2025 |
| Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models | 03/2025 | CVPR 2025 | |
| Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views | 03/2025 | CVPR 2025 | |
| LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors | 12/2024 | code | arXiv |
| MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views | 11/2024 | code | NeurIPS 2024 |
| 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion Priors | 10/2024 | NeurIPS 2024 | |
| ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis | 09/2024 | code | arXiv |
| ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model | 08/2024 | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
VGGT-SLAM 2.0: Real-time Dense Feed-forward SLAM Abstract—We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree- of-freedom drift and planar degeneracy from VGGT-SLAM by creating a new factor graph design while still addressing the reconstruction ambiguity of VGGT given unknown camera intrinsics. Secondly, by studying the attention layers of VGGT, we show that one of the layers is well suited to assist in image 1Laboratory for Information & Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA. {drmaggio, lcarlone}@mit.edu *Luca holds concurrent appointments as a faculty at the Massachusetts Institute of Technology and as an Amazon Scholar. | 01/2026 | CVPR 2026 | |
FUSER: Feed-Forward Multiview 3D Registration Transformer AbstractRegistration of multiview point clouds conventionally re- lies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expen- sive and inherently ill-posed without holistic geometric con- straints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves ab- solute translation cues, and performs efficient intra- and inter-scan reasoning through a Geometric Alternating At- tention module. | 12/2025 | code | arXiv |
VGGT4D: Mining Motion Cues in Visual Geometry Transformers AbstractReconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when moving objects dominate. Existing 4D approaches often rely on external priors, heavy post- optimization, or require fine-tuning on 4D datasets. In this paper, we propose VGGT4D, a training-free framework that extends the 3D foundation model VGGT for robust 4D scene reconstruction. Our approach is motivated by the key find- ing that VGGT’s global attention layers already implicitly encode rich, layer-wise dynamic cues. | 11/2025 | 3DV 2026 | |
| DropD-SLAM: RGB-D SLAM Without the Depth Sensor | 10/2025 | arXiv | |
PointSt3R: Point Tracking through 3D Grounded Correspondence AbstractRecent advances in foundational 3D reconstruction models, such as DUSt3R and MASt3R, have shown great potential in 2D and 3D correspondence in static scenes. In this pa- per, we propose to adapt them for the task of point track- ing through 3D grounded correspondence. We first demon- strate that these models are competitive point trackers when focusing on static points, present in current point tracking benchmarks (+33.5% on EgoPoints vs. CoTracker2). We propose to combine the reconstruction loss with training for dynamic correspondence along with a visibility head, and fine-tuning MASt3R for point tracking using a rela- tively small amount of synthetic data. | 10/2025 | project | arXiv |
| Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming | 09/2025 | arXiv | |
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool AbstractWe present WinT3R, a feed-forward reconstruction model capable of online pre- diction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures suffi- cient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. | 09/2025 | code | ICLR 2026 |
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization AbstractScene regression methods, such as VGGT [85], solve the Structure-from-Motion (SfM) problem by directly regress- ing camera poses and 3D scene structures from input im- ages. They demonstrate impressive performance in han- dling images under extreme viewpoint changes. However, these methods struggle to handle a large number of input images. To address this problem, we introduce SAIL-Recon, a feed-forward Transformer for large scale SfM, by aug- menting the scene regression network with visual localiza- tion capabilities. Specifically, our method first computes a neural scene representation from a subset of anchor im- ages. The regression network is then fine-tuned to recon- struct all input images conditioned on this neural scene representation. | 08/2025 | project | 3DV 2026 (Oral) |
StreamVGGT: Streaming 4D Visual Geometry Transformer AbstractPerceiving and reconstructing 4D spatial-temporal geometry from videos is a fun- damental yet challenging computer vision task. To facilitate interactive and real- time applications, we propose a streaming 4D visual geometry transformer that shares a similar philosophy with autoregressive large language models. We ex- plore a simple and efficient design and employ a causal transformer architecture to process the input sequence in an online manner. We use temporal causal at- tention and cache the historical keys and values as implicit memory to enable efficient streaming long-term 4D reconstruction. This design can handle real- time 4D reconstruction by incrementally integrating historical information while maintaining high-quality spatial consistency. | 07/2025 | code | ICLR 2026 |
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold AbstractWe present VGGT-SLAM, a dense RGB SLAM system constructed by incre- mentally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degrees-of-freedom projective transformation of the true geometry. | 05/2025 | NeurIPS 2025 | |
| MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion | 04/2025 | code | CVPR 2025 |
Easi3R: Estimating Disentangled Motion from DUSt3R Without Training AbstractRecent advances in DUSt3R have enabled robust estima- tion of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited scale and diversity of available 4D datasets present a major bottleneck for training a highly general- izable 4D model. This constraint has driven conventional 4D methods to fine-tune 3D models on scalable dynamic video data with additional geometric priors such as opti- cal flow and depths. In this work, we take an opposite path and introduce Easi3R, a simple yet efficient training-free method for 4D reconstruction. Our approach applies at- tention adaptation during inference, eliminating the need for from-scratch pre-training or network fine-tuning. | 2025 | ICCV 2025 | |
| RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation | 2025 | IROS 2025 | |
| MonST3R: Motion DUSt3R for Dynamic Scene Reconstruction | 2024 | project | ICLR 2025 |
MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors AbstractWe present a real-time monocular dense SLAM system de- signed bottom-up from MASt3R, a two-view 3D reconstruc- tion and matching prior. Equipped with this strong prior, our system is robust on in-the-wild video sequences despite making no assumption on a fixed or parametric camera model beyond a unique camera centre. We introduce ef- ficient methods for pointmap matching, camera tracking and local fusion, graph construction and loop closure, and second-order global optimisation. With known calibration, a simple modification to the system achieves state-of-the-art performance across various benchmarks. Altogether, we propose a plug-and-play monocular SLAM system capable of producing globally consistent poses and dense geometry while operating at 15 FPS. | 12/2024 | code | CVPR 2025 |
| SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | 12/2024 | code | CVPR 2025 |
| MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion | 09/2024 | code | CVPR 2025 |
| Title | Date | Link | Venue |
|---|---|---|---|
| ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation | 12/2024 | project | arXiv |
| Autonomous Character-Scene Interaction Synthesis from Text Instruction | 10/2024 | code | SIGGRAPH Aisa 2024 |
| Human-Object Interaction from Human-Level Instructions | 06/2024 | project | arXiv |
| Generating Human Motion in 3D Scenes from Text Descriptions | 05/2024 | code | CVPR 2024 |
| Generating Human Interaction Motions in Scenes with Text Control | 04/2024 | code | ECCV 2024 |
| Controllable Human-Object Interaction Synthesis | 12/2023 | code | ECCV 2024 oral |
| HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes | 10/2022 | project | NeurIPS 2022 |
| Title | Date | Link | Venue |
|---|---|---|---|
HUMAN3R: Everyone Everywhere All at Once AbstractWe present Human3R, a unified, feed-forward framework for online 4D human- scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact- aware refinement between humans and scenes, and heavy dependencies, e.g., human detection, depth estimation, and SLAM pre-processing, Human3R jointly recovers global multi-person SMPL-X bodies (“everyone”), dense 3D scene (“ev- erywhere”), and camera trajectories in a single forward pass (“all-at-once”). Our method builds upon the 4D online reconstruction model CUT3R, and uses parameter-efficient visual prompt tuning, to strive to preserve CUT3R’s rich spa- tiotemporal priors, while enabling direct readout of multiple SMPL-X bodies. | 10/2025 | project | ICLR 2026 |
| Title | Date | Link | Venue |
|---|---|---|---|
| CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image | 06/2025 | SIGGRAPH (Best paper) |
🌟 A curate list of papers, datasets, and projects for 3D Reconstruction and Generation.
95
43 commits
updated Mar 13, 2026
🌟 A curate list of papers, datasets, and projects for 3D Reconstruction and Generation.
:heart: Pull requests or issues for updates are very welcome.
| Title | Date | Link | Venue |
|---|---|---|---|
Emergent Extreme-View Geometry in 3D Foundation Models Abstract3D foundation models (3DFMs) have recently transformed 3D vision, enabling joint prediction of depths, poses, and point maps directly from images. Yet their ability to rea- son under extreme, non-overlapping views remains largely unexplored. In this work, we study their internal repre- sentations and find that 3DFMs exhibit an emergent un- derstanding of extreme-view geometry, despite never being trained for such conditions. To further enhance these capa- bilities, we introduce a lightweight alignment scheme that refines their internal 3D representation by tuning only a small subset of backbone bias terms, leaving all decoder heads frozen. This targeted adaptation substantially im- proves relative pose estimation under extreme viewpoints without degrading per-image depth or point quality. | 11/2025 | project | CVPR 2026 |
Emergent Outlier View Rejection in Visual Geometry Grounded Transformers AbstractReliable 3D reconstruction from in-the-wild image collec- tions is often hindered by “noisy” images—irrelevant in- puts with little or no view overlap with others. While tra- ditional Structure-from-Motion pipelines handle such cases through geometric verification and outlier rejection, feed- forward 3D reconstruction models lack these explicit mech- anisms, leading to degraded performance under in-the-wild conditions. In this paper, we discover that the existing feed- forward reconstruction model, e.g., VGGT, despite lacking explicit outlier-rejection mechanisms or noise-aware train- ing, can inherently distinguish distractor images. | 12/2025 | project | CVPR 2026 |
AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend AbstractWe present AMB3R, a multi-view feed-forward model for dense 3D reconstruction on a metric-scale that addresses diverse 3D vision tasks. The key idea is to leverage a sparse, yet compact, volumetric scene representation as our backend, enabling geometric reasoning with spatial compactness. Although trained solely for multi-view re- construction, we demonstrate that AMB3R can be seam- lessly extended to uncalibrated visual odometry (online) or large-scale structure from motion without the need for task- specific fine-tuning or test-time optimization. | 11/2025 | project | CVPR 2026 |
Depth Anything 3: Recovering the Visual Space from Any Views AbstractWe present Depth Anything 3 (DA3), a model that predicts spatially consistent geometry from an arbitrary number of visual inputs, with or without known camera poses. In pursuit of minimal modeling, DA3 yields two key insights: a single plain transformer (e.g., vanilla DINO encoder) is sufficient as a backbone without architectural specialization, and a singular depth-ray prediction target obviates the need for complex multi-task learning. Through our teacher-student training paradigm, the model achieves a level of detail and generalization on par with Depth Anything 2 (DA2). We establish a new visual geometry benchmark covering camera pose estimation, any-view geometry and visual rendering. | 11/2025 | project | ICLR 2026 |
| Dens3R: A Foundation Model for 3D Geometry Prediction | 2025 | ICLR 2026 | |
WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting AbstractWe present WorldMirror, an all-in-one, feed-forward model for versatile 3D geo- metric prediction tasks. Unlike existing methods constrained to image-only inputs or customized for a specific task, our framework flexibly integrates diverse geo- metric priors, including camera poses, intrinsics, and depth maps, while simulta- neously generating multiple 3D representations: dense point clouds, multi-view depth maps, camera parameters, surface normals, and 3D Gaussians. This elegant and unified architecture leverages available prior information to resolve structural ambiguities and delivers geometrically consistent 3D outputs in a single forward pass. | 10/2025 | arXiv | |
PAGE-4D: Disentangled Pose and Geometry for 4D Perception AbstractRecent 3D feed-forward models, such as the Visual Geometry Grounded Trans- former (VGGT), have shown strong capability in inferring 3D attributes of static scenes. However, since they are typically trained on static datasets, these mod- els often struggle in real-world scenarios involving complex dynamic elements, such as moving humans or deformable objects like umbrellas. To address this limitation, we introduce PAGE-4D, a feedforward model that extends VGGT to dynamic scenes, enabling camera pose estimation, depth prediction and point cloud reconstruction —all without post-processing. | 10/2025 | ICLR 2026 | |
DriveVGGT: Visual Geometry Transformer for Autonomous Driving AbstractFeed-forward reconstruction has recently gained significant attention, with VGGT being a notable example. However, directly applying VGGT to autonomous driving (AD) sys- tems leads to sub-optimal results due to the different priors between the two tasks. In AD systems, several important new priors need to be considered: (i) The overlap between camera views is minimal, as autonomous driving sensor se- tups are designed to achieve 360◦coverage at a low cost. (ii) The camera intrinsics and extrinsics are known, which introduces more constraints on the output and also enables the estimation of absolute scale. (iii) Relative positions of all cameras remain fixed though the ego vehicle is in motion. | 11/2025 | ECCV 2025 | |
π3: Scalable Permutation-Equivariant Visual Geometry Learning AbstractWe introduce π3, a feed-forward neural network that offers a novel approach to visual geometry reconstruction, breaking the reliance on a conventional fixed reference view. Previous methods often anchor their reconstructions to a designated viewpoint, an inductive bias that can lead to instability and failures if the reference is suboptimal. In contrast, π3 em- ploys a fully permutation-equivariant architecture to predict affine-invariant camera poses and scale-invariant local point maps without any reference frames. This design makes our model inherently robust to input ordering and highly scalable. | 07/2025 | ICLR 2026 | |
BoxFusion: Reconstruction-Free Open-Vocabulary 3D Object Detection AbstractOpen-vocabulary 3D object detection has gained sig- nificant interest due to its critical applications in au- tonomous driving and embodied AI. Existing detection methods, whether offline or online, typically rely on dense point cloud reconstruction, which imposes substantial com- putational overhead and memory constraints, hindering real-time deployment in downstream tasks. To address this, we propose a novel reconstruction-free online framework tailored for memory-efficient and real-time 3D detection. Specifically, given streaming posed RGB-D video input, we leverage Cubify Anything as a pre-trained visual foundation model (VFM) for single-view 3D object detection by bound- ing boxes, coupled with CLIP to capture open-vocabulary semantics of detected objects. | 06/2025 | arXiv | |
Depth Anything with Any Prior AbstractThis work presents Prior Depth Anything, a framework that combines incomplete but precise metric information in depth measurement with relative but complete geometric structures in depth prediction, generating accurate, dense, and detailed metric depth maps for any scene. To this end, we design a coarse-to-fine pipeline to progressively integrate the two complementary depth sources. First, we introduce pixel-level metric alignment and distance-aware weighting to pre-fill diverse metric priors by explicitly using depth prediction. It effectively narrows the domain gap be- tween prior patterns, enhancing generalization across vary- ing scenarios. Second, we develop a conditioned monoc- ular depth estimation (MDE) model to refine the inherent noise of depth priors. | 05/2025 | project | ICML 2025 |
| MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models | 05/2025 | code | CVPR 2025 |
VGGT:Visual Geometry Grounded Transformer AbstractWe present VGGT, a feed-forward neural network that di- rectly infers all key 3D attributes of a scene, including cam- era parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and special- ized for single tasks. It is also simple and efficient, re- constructing images in under one second, and still out- performing alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, in- cluding camera parameter estimation, multi-view depth es- timation, dense point cloud reconstruction, and 3D point tracking. | 03/2025 | code | CVPR 2025 |
| MUSt3R: Multi-view Network for Stereo 3D Reconstruction | 03/2025 | CVPR 2025 | |
| Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | 03/2025 | CVPR 2025 | |
| FLARE: Feed-forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | 02/2025 | code | CVPR 2025 |
EquiPose: Exploiting Permutation Equivariance for Relative Camera Pose Estimation AbstractRelative camera pose estimation between two images is a fundamental task in 3D computer vision. Recently, many relative pose estimation networks have been explored for learning a mapping from two input images to their cor- responding relative pose, however, the estimated relative poses by these methods do not have the intrinsic Pose Per- mutation Equivariance (PPE) property: the estimated rel- ative pose from Image A to Image B should be the inverse of that from Image B to Image A. It means that permut- ing the input order of two images would cause these meth- ods to obtain inconsistent relative poses. To address this problem, we firstly introduce the concept of PPE mapping, which indicates such a mapping that captures the intrin- sic PPE property of relative poses. | 02/2025 | CVPR 2025 | |
RePoseD: Efficient Relative Pose Estimation With Known Depth Information AbstractRecent advances in monocular depth estimation meth- ods (MDE) and their improved accuracy open new possibil- ities for their applications. In this paper, we investigate how monocular depth estimates can be used for relative pose es- timation. In particular, we are interested in answering the question whether using MDEs improves results over tradi- tional point-based methods. We propose a novel framework for estimating the relative pose of two cameras from point correspondences with associated monocular depths. Since depth predictions are typically defined up to an unknown scale or even both unknown scale and shift parameters, our solvers jointly estimate the scale or both the scale and shift parameters along with the relative pose. | 01/2025 | ICCV 2025 | |
| Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | 01/2025 | code | CVPR 2025 |
Continuous 3D Perception Model with Persistent State AbstractWe present a unified framework capable of solving a broad range of 3D tasks. Our approach features a stateful recurrent model that continuously updates its state representation with each new observation. Given a stream of images, this evolv- ing state can be used to generate metric-scale pointmaps (per-pixel 3D points) for each new input in an online fashion. These pointmaps reside within a common coordinate system, and can be accumulated into a coherent, dense scene re- construction that updates as new images arrive. | 01/2025 | code | CVPR 2025 (Oral) |
| Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization | 12/2024 | code | CVPR 2025 |
| MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | 12/2024 | code | CVPR 2025 |
Relative Pose Estimation through Affine Corrections of Monocular Depth Priors AbstractMonocular depth estimation (MDE) models have under- gone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metric (absolute) depth. However, effectively leveraging these predictions for geometric vision tasks, in particular relative pose estimation, remains rela- tively under explored. While depths provide rich constraints for cross-view image alignment, the intrinsic noise and am- biguity from the monocular depth priors present practical challenges to improving upon classic keypoint-based so- lutions. | 12/2024 | code | CVPR 2025 |
Understanding multi-view transformers AbstractMulti-view transformers such as DUSt3R [59] are revolu- tionizing 3D vision by solving 3D tasks in a feed-forward manner. However, contrary to previous optimization-based pipelines, the inner mechanisms of multi-view transformers are unclear. Their black-box nature makes further improve- ments beyond data scaling challenging and complicates us- age in safety- and reliability-critical applications. Here, we present an approach for probing and visualizing 3D repre- sentations from the residual connections of the multi-view transformers’ layers. | 2025 | code | ICCV 2025 |
Peering into the Unknown: Active View Selection for 3D Reconstruction AbstractImagine trying to understand the shape of a teapot by viewing it from the front—you might see the spout, but completely miss the handle. Some perspectives naturally provide more information than others. How can an AI system determine which viewpoint offers the most valuable insight for accurate and efficient 3D object reconstruction? Active view selection (AVS) for 3D reconstruction remains a fundamental challenge in computer vision. The aim is to identify the minimal set of views that yields the most accurate 3D reconstruction. | 2025 | ICLR 2026 | |
| 3D Reconstruction with Spatial Memory | 08/2024 | code | 3DV 2025 |
Grounding Image Matching in 3D with MASt3R AbstractImage Matching is a core component of all best-performing algorithms and pipelines in 3D vision. Yet despite matching being fundamentally a 3D problem, intrinsically linked to camera pose and scene geometry, it is typically treated as a 2D problem. This makes sense as the goal of matching is to establish correspondences between 2D pixel fields, but also seems like a potentially hazardous choice. In this work, we take a different stance and propose to cast matching as a 3D task with DUSt3R, a recent and powerful 3D reconstruction framework based on Transformers. Based on pointmaps regression, this method displayed impressive robustness in matching views with extreme viewpoint changes, yet with limited accuracy. | 06/2024 | code | Arxiv 2024 |
| UniDepthV2: Universal Monocular Metric Depth Estimation | 06/2024 | CVPR 2024 | |
DUSt3R: Geometric 3D Vision Made Easy AbstractMulti-view stereo reconstruction (MVS) in the wild re- quires to first estimate the camera parameters e.g. intrinsic and extrinsic parameters. These are usually tedious and cumbersome to obtain, yet they are mandatory to triangu- late corresponding pixels in 3D space, which is the core of all best performing MVS algorithms. In this work, we take an opposite stance and introduce DUSt3R1, a radi- cally novel paradigm for Dense and Unconstrained Stereo 3D Reconstruction of arbitrary image collections, i.e. oper- ating without prior information about camera calibration nor viewpoint poses. We cast the pairwise reconstruction problem as a regression of pointmaps, relaxing the hard constraints of usual projective camera models. We show 1https://dust3r.europe.naverlabs. | 12/2023 | code | CVPR 2024 |
RaCalNet: Radar Calibration Network for Sparse-Supervised Metric Depth Estimation Abstract—Dense metric depth estimation using millimeter- wave radar typically requires dense LiDAR supervision, gen- erated via multi-frame projection and interpolation, to guide the learning of accurate depth from sparse radar measurements and RGB images. However, this paradigm is both costly and data- intensive. To address this, we propose RaCalNet, a novel frame- work that eliminates the need for dense supervision by using sparse LiDAR to supervise the learning of refined radar mea- surements, resulting in a supervision density of merely around 1% compared to dense-supervised methods. Unlike previous approaches that associate radar points with broad image regions and rely heavily on dense labels, RaCalNet first recalibrates and refines sparse radar points to construct accurate depth priors. | 06/2025 | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
GPA-VGGT: Adapting VGGT to Large scale Localization by Self-Supervised Learning Abstract—Transformer-based general visual geometry frame- works have shown promising performance in camera pose estima- tion and 3D scene understanding. Recent advancements in Visual Geometry Grounded Transformer (VGGT) models have shown great promise in camera pose estimation and 3D reconstruction. However, these models typically rely on ground truth labels for training, posing challenges when adapting to unlabeled and unseen scenes. In this paper, we propose a self-supervised framework to train VGGT with unlabeled data, thereby en- hancing its localization capability in large-scale environments. To achieve this, we extend conventional pair-wise relations to sequence-wise geometric constraints for self-supervised learning. | 01/2026 | code | arXiv |
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams AbstractThe grand vision of enabling persistent, large-scale 3D visual geometry understanding is shackled by the irrec- oncilable demands of scalability and long-term stability. While offline models like VGGT achieve inspiring geom- etry capability, their batch-based nature renders them ir- relevant for live systems. Streaming architectures, though the intended solution for live operation, have proven inade- quate. Existing methods either fail to support truly infinite- horizon inputs or suffer from catastrophic drift over long sequences. We shatter this long-standing dilemma with In- finiteVGGT, a causal visual geometry transformer that op- erationalizes the concept of a rolling memory through a bounded yet adaptive and perpetually expressive KV cache. | 01/2026 | code | arXiv |
AVGGT: Rethinking Global Attention for Accelerating VGGT AbstractSince DUSt3R, models such as VGGT and π3 have shown strong multi-view 3D performance, but their heavy reliance on global self-attention results in high computational cost. Existing sparse-attention variants offer partial speedups, yet lack a systematic analysis of how global attention con- tributes to multi-view reasoning. In this paper, we first con- duct an in-depth investigation of the global attention mod- ules in VGGT and π3 to better understand their roles. Our analysis reveals a clear division of roles in the alternating *Equal contribution. †Corresponding author. global-frame architecture: early global layers do not form meaningful correspondences, middle layers perform cross- view alignment, and last layers provide only minor refine- ments. | 12/2025 | arXiv | |
FlashVGGT: Efficient and Scalable Visual Geometry Transformers Abstract3D reconstruction from multi-view images is a core chal- lenge in computer vision. Recently, feed-forward methods have emerged as efficient and robust alternatives to tra- ditional per-scene optimization techniques. Among them, state-of-the-art models like the Visual Geometry Ground- ing Transformer (VGGT) leverage full self-attention over all image tokens to capture global relationships. However, this approach suffers from poor scalability due to the quadratic complexity of self-attention and the large number of tokens generated in long image sequences. In this work, we in- troduce FlashVGGT, an efficient alternative that addresses this bottleneck through a descriptor-based attention mecha- nism. | 12/2025 | project | arXiv |
LiteVGGT: Boosting Vanilla VGGT via Geometry-aware Cached Token Merging Abstract3D vision foundation models like Visual Geometry Grounded Transformer (VGGT) have advanced greatly in geometric perception. However it is time-consuming and memory-intensive for long sequences, limiting application to large-scale scenes beyond hundreds of images. To ad- dress this, we propose LiteVGGT, achieving up to 10× speedup and substantial memory reduction, enabling effi- cient processing of 1000-image scenes. We derive two key insights for 3D reconstruction: 1) tokens from local im- age regions have inherent geometric correlations, leading to high similarity and computational redundancy; 2) token similarity acroses adjacent network layers remains stable, allowing for reusable merge decisions. | 12/2025 | project | arXiv |
HTTM: Head-wise Temporal Token Merging for Faster VGGT AbstractThe Visual Geometry Grounded Transformer (VGGT) marks a significant leap forward in 3D scene reconstruc- tion, as it is the first model that directly infers all key 3D at- tributes (camera poses, depths, and dense geometry) jointly in one pass. However, this joint inference mechanism re- quires global attention layers that perform all-to-all atten- tion computation on tokens from all views. For reconstruc- tion of large scenes with long-sequence inputs, this causes a significant latency bottleneck. In this paper, we pro- pose head-wise temporal merging (HTTM), a training-free 3D token merging method for accelerating VGGT. | 11/2025 | arXiv | |
FastVGGT: Training-Free Acceleration of Visual Geometry Transformer AbstractFoundation models for 3D vision have recently demon- strated remarkable capabilities in 3D perception. How- ever, scaling these models to long-sequence image inputs remains a significant challenge due to inference-time in- *Corresponding Author. efficiency. In this work, we present a detailed analysis of VGGT, a state-of-the-art feed-forward visual geometry model and identify its primary bottleneck. Visualization fur- ther reveals a token collapse phenomenon in the attention maps. Motivated by these findings, we explore the poten- tial of token merging in the feed-forward visual geometry model. Owing to the unique architectural and task-specific properties of 3D models, directly applying existing merg- arXiv:2509.02560v2 [cs.CV] 9 Nov 2025 ing techniques proves challenging. | 09/2025 | project | ICLR 2026 |
Faster VGGT with Block-Sparse Global Attention AbstractEfficient and accurate feed-forward multi-view recon- struction has long been an important task in computer vi- sion. Recent transformer-based models like VGGT and π3 have achieved impressive results with simple architectures, yet they face an inherent runtime bottleneck, due to the quadratic complexity of the global attention layers, that lim- its the scalability to large image sets. In this paper, we em- pirically analyze the global attention matrix of these models and observe that probability mass concentrates on a small subset of patch-patch interactions that correspond to cross- view geometric matches. | 09/2025 | arXiv | |
VGGT-Long: Chunk it, Loop it, Align it – Pushing VGGT's Limits on Long RGB Sequences AbstractFoundation models for 3D vision have recently demon- strated remarkable capabilities in 3D perception. How- ever, extending these models to large-scale RGB stream 3D reconstruction remains challenging due to memory lim- itations. In this work, we propose VGGT-Long, a sim- ple yet effective system that pushes the limits of monocu- 1† Corresponding author. lar 3D reconstruction to kilometer-scale, unbounded out- door environments. Our approach addresses the scalability bottlenecks of existing models through a chunk-based pro- cessing strategy combined with overlapping alignment and lightweight loop closure optimization. Without requiring camera calibration, depth supervision or model retraining, VGGT-Long achieves trajectory and reconstruction perfor- mance comparable to traditional methods. | 07/2025 | code | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
MIRROR: Make Your Object-Level Multi-View Generation Consistent AbstractMulti-view Diffusion has greatly advanced the de- velopment of 3D content creation by generating multiple images from distinct views, achieving remarkable photorealistic results. However, ex- isting works are still vulnerable to inconsistent 3D geometric structures (commonly known as Janus Problem) and severe artifacts. In this paper, we introduce MIRROR, a versatile plug-and-play method that rectifies such inconsistencies in a training-free manner, enabling the acquisition of high-fidelity, realistic structures without compro- mising diversity. Our key idea focuses on tracing the motion trajectory of physical points across adjacent viewpoints, enabling rectifications based on neighboring observations of the same region. | 2025 | ArXiV | |
Refine Any Object in Any Scene (RAISE) AbstractViewpoint missing of objects is common in scene reconstruction, as camera paths typically prioritize capturing the overall scene structure rather than individual objects. This makes it highly challenging to achieve high-fidelity object-level mod- eling while maintaining accurate scene-level representation. Addressing this issue is critical for advancing downstream tasks requiring detailed object understanding and appearance modeling. In this paper, we introduce Refine Any object In any ScenE (RAISE), a novel 3D enhancement framework that leverages 3D generative priors to recover fine-grained object geometry and appearance under missing views. | 06/2025 | code | arXiv |
| GenFusion: Closing the Loop between Reconstruction and Generation via Videos | 03/2025 | code | CVPR 2025 |
| Difix3D+: Improving 3D Reconstructions with Single-Step Diffusion Models | 03/2025 | CVPR 2025 | |
| Free360: Layered Gaussian Splatting for Unbounded 360-Degree View Synthesis from Extremely Sparse and Unposed Views | 03/2025 | CVPR 2025 | |
| LiftImage3D: Lifting Any Single Image to 3D Gaussians with Video Generation Priors | 12/2024 | code | arXiv |
| MVSplat360: Feed-Forward 360 Scene Synthesis from Sparse Views | 11/2024 | code | NeurIPS 2024 |
| 3DGS-Enhancer: Enhancing Unbounded 3D Gaussian Splatting with View-consistent 2D Diffusion Priors | 10/2024 | NeurIPS 2024 | |
| ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis | 09/2024 | code | arXiv |
| ReconX: Reconstruct Any Scene from Sparse Views with Video Diffusion Model | 08/2024 | arXiv |
| Title | Date | Link | Venue |
|---|---|---|---|
VGGT-SLAM 2.0: Real-time Dense Feed-forward SLAM Abstract—We present VGGT-SLAM 2.0, a real-time RGB feed-forward SLAM system which substantially improves upon VGGT-SLAM for incrementally aligning submaps created from VGGT. Firstly, we remove high-dimensional 15-degree- of-freedom drift and planar degeneracy from VGGT-SLAM by creating a new factor graph design while still addressing the reconstruction ambiguity of VGGT given unknown camera intrinsics. Secondly, by studying the attention layers of VGGT, we show that one of the layers is well suited to assist in image 1Laboratory for Information & Decision Systems, Massachusetts Institute of Technology, Cambridge, MA, USA. {drmaggio, lcarlone}@mit.edu *Luca holds concurrent appointments as a faculty at the Massachusetts Institute of Technology and as an Amazon Scholar. | 01/2026 | CVPR 2026 | |
FUSER: Feed-Forward Multiview 3D Registration Transformer AbstractRegistration of multiview point clouds conventionally re- lies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expen- sive and inherently ill-posed without holistic geometric con- straints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves ab- solute translation cues, and performs efficient intra- and inter-scan reasoning through a Geometric Alternating At- tention module. | 12/2025 | code | arXiv |
VGGT4D: Mining Motion Cues in Visual Geometry Transformers AbstractReconstructing dynamic 4D scenes is challenging, as it requires robust disentanglement of dynamic objects from the static background. While 3D foundation models like VGGT provide accurate 3D geometry, their performance drops markedly when moving objects dominate. Existing 4D approaches often rely on external priors, heavy post- optimization, or require fine-tuning on 4D datasets. In this paper, we propose VGGT4D, a training-free framework that extends the 3D foundation model VGGT for robust 4D scene reconstruction. Our approach is motivated by the key find- ing that VGGT’s global attention layers already implicitly encode rich, layer-wise dynamic cues. | 11/2025 | 3DV 2026 | |
| DropD-SLAM: RGB-D SLAM Without the Depth Sensor | 10/2025 | arXiv | |
PointSt3R: Point Tracking through 3D Grounded Correspondence AbstractRecent advances in foundational 3D reconstruction models, such as DUSt3R and MASt3R, have shown great potential in 2D and 3D correspondence in static scenes. In this pa- per, we propose to adapt them for the task of point track- ing through 3D grounded correspondence. We first demon- strate that these models are competitive point trackers when focusing on static points, present in current point tracking benchmarks (+33.5% on EgoPoints vs. CoTracker2). We propose to combine the reconstruction loss with training for dynamic correspondence along with a visibility head, and fine-tuning MASt3R for point tracking using a rela- tively small amount of synthetic data. | 10/2025 | project | arXiv |
| Evict3R: Training-Free Token Eviction for Memory-Bounded Streaming | 09/2025 | arXiv | |
WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool AbstractWe present WinT3R, a feed-forward reconstruction model capable of online pre- diction of precise camera poses and high-quality point maps. Previous methods suffer from a trade-off between reconstruction quality and real-time performance. To address this, we first introduce a sliding window mechanism that ensures suffi- cient information exchange among frames within the window, thereby improving the quality of geometric predictions without large computation. In addition, we leverage a compact representation of cameras and maintain a global camera token pool, which enhances the reliability of camera pose estimation without sacrificing efficiency. | 09/2025 | code | ICLR 2026 |
SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization AbstractScene regression methods, such as VGGT [85], solve the Structure-from-Motion (SfM) problem by directly regress- ing camera poses and 3D scene structures from input im- ages. They demonstrate impressive performance in han- dling images under extreme viewpoint changes. However, these methods struggle to handle a large number of input images. To address this problem, we introduce SAIL-Recon, a feed-forward Transformer for large scale SfM, by aug- menting the scene regression network with visual localiza- tion capabilities. Specifically, our method first computes a neural scene representation from a subset of anchor im- ages. The regression network is then fine-tuned to recon- struct all input images conditioned on this neural scene representation. | 08/2025 | project | 3DV 2026 (Oral) |
StreamVGGT: Streaming 4D Visual Geometry Transformer AbstractPerceiving and reconstructing 4D spatial-temporal geometry from videos is a fun- damental yet challenging computer vision task. To facilitate interactive and real- time applications, we propose a streaming 4D visual geometry transformer that shares a similar philosophy with autoregressive large language models. We ex- plore a simple and efficient design and employ a causal transformer architecture to process the input sequence in an online manner. We use temporal causal at- tention and cache the historical keys and values as implicit memory to enable efficient streaming long-term 4D reconstruction. This design can handle real- time 4D reconstruction by incrementally integrating historical information while maintaining high-quality spatial consistency. | 07/2025 | code | ICLR 2026 |
VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold AbstractWe present VGGT-SLAM, a dense RGB SLAM system constructed by incre- mentally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degrees-of-freedom projective transformation of the true geometry. | 05/2025 | NeurIPS 2025 | |
| MP-SfM: Monocular Surface Priors for Robust Structure-from-Motion | 04/2025 | code | CVPR 2025 |
Easi3R: Estimating Disentangled Motion from DUSt3R Without Training AbstractRecent advances in DUSt3R have enabled robust estima- tion of dense point clouds and camera parameters of static scenes, leveraging Transformer network architectures and direct supervision on large-scale 3D datasets. In contrast, the limited scale and diversity of available 4D datasets present a major bottleneck for training a highly general- izable 4D model. This constraint has driven conventional 4D methods to fine-tune 3D models on scalable dynamic video data with additional geometric priors such as opti- cal flow and depths. In this work, we take an opposite path and introduce Easi3R, a simple yet efficient training-free method for 4D reconstruction. Our approach applies at- tention adaptation during inference, eliminating the need for from-scratch pre-training or network fine-tuning. | 2025 | ICCV 2025 | |
| RA-NeRF: Robust Neural Radiance Field Reconstruction with Accurate Camera Pose Estimation | 2025 | IROS 2025 | |
| MonST3R: Motion DUSt3R for Dynamic Scene Reconstruction | 2024 | project | ICLR 2025 |
MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors AbstractWe present a real-time monocular dense SLAM system de- signed bottom-up from MASt3R, a two-view 3D reconstruc- tion and matching prior. Equipped with this strong prior, our system is robust on in-the-wild video sequences despite making no assumption on a fixed or parametric camera model beyond a unique camera centre. We introduce ef- ficient methods for pointmap matching, camera tracking and local fusion, graph construction and loop closure, and second-order global optimisation. With known calibration, a simple modification to the system achieves state-of-the-art performance across various benchmarks. Altogether, we propose a plug-and-play monocular SLAM system capable of producing globally consistent poses and dense geometry while operating at 15 FPS. | 12/2024 | code | CVPR 2025 |
| SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | 12/2024 | code | CVPR 2025 |
| MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion | 09/2024 | code | CVPR 2025 |
| Title | Date | Link | Venue |
|---|---|---|---|
| ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation | 12/2024 | project | arXiv |
| Autonomous Character-Scene Interaction Synthesis from Text Instruction | 10/2024 | code | SIGGRAPH Aisa 2024 |
| Human-Object Interaction from Human-Level Instructions | 06/2024 | project | arXiv |
| Generating Human Motion in 3D Scenes from Text Descriptions | 05/2024 | code | CVPR 2024 |
| Generating Human Interaction Motions in Scenes with Text Control | 04/2024 | code | ECCV 2024 |
| Controllable Human-Object Interaction Synthesis | 12/2023 | code | ECCV 2024 oral |
| HUMANISE: Language-conditioned Human Motion Generation in 3D Scenes | 10/2022 | project | NeurIPS 2022 |
| Title | Date | Link | Venue |
|---|---|---|---|
HUMAN3R: Everyone Everywhere All at Once AbstractWe present Human3R, a unified, feed-forward framework for online 4D human- scene reconstruction, in the world frame, from casually captured monocular videos. Unlike previous approaches that rely on multi-stage pipelines, iterative contact- aware refinement between humans and scenes, and heavy dependencies, e.g., human detection, depth estimation, and SLAM pre-processing, Human3R jointly recovers global multi-person SMPL-X bodies (“everyone”), dense 3D scene (“ev- erywhere”), and camera trajectories in a single forward pass (“all-at-once”). Our method builds upon the 4D online reconstruction model CUT3R, and uses parameter-efficient visual prompt tuning, to strive to preserve CUT3R’s rich spa- tiotemporal priors, while enabling direct readout of multiple SMPL-X bodies. | 10/2025 | project | ICLR 2026 |
| Title | Date | Link | Venue |
|---|---|---|---|
| CAST: Component-Aligned 3D Scene Reconstruction from an RGB Image | 06/2025 | SIGGRAPH (Best paper) |