3D-Vision-World/All-3R-SLAM-in-this-Repo

This is a list of relevant papers for 3D Geometric Foundation Models and Applications.

287

158 commits

updated Oct 3, 2026

See the code

README

3D Geometric Foundation Models (3R) with SLAM

Survey Paper

  • Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey, arXiv 2025. [Paper] [Website]
  • Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective, arXiv 2026. [Paper] [Website]

3D Geometric Foundation Models (3R)

  • DUSt3R: Geometric 3D Vision Made Easy, CVPR 2024. [Paper] [Code] [Website]
  • Monst3r: A simple approach for estimating geometry in the presence of motion, ICLR 2025. [Paper] [Code] [Website]
  • LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models, ICLR 2025. [Paper] [Website]
  • (CUT3R) Continuous 3D Perception Model with Persistent State, CVPR 2025. [Paper] [Code] [Website]
  • Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization, CVPR 2025. [Paper] [Code]
  • DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction, arXiv 2024. [Paper] [Code] [Website]
  • MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion, 3DV 2025. [Paper] [Code]
  • Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs, arXiv 2024. [Paper] [Project] [Code]
  • SAB3R: Semantic-Augmented Backbone in 3D Reconstruction, arXiv 2024. [Paper]
  • No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images, ICLR 2025. [Paper] [Code] [Website]
  • Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features, CVPR 2025. [Paper] [Code] [Project]
  • (Spann3R) 3D Reconstruction with Spatial Memory, 3DV 2025. [Paper] [Code] [Website]
  • Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass, CVPR 2025. [Paper] [Website] [Code]
  • InstantSplat: Unbounded Sparse-view Pose-free Gaussian Splatting in 40 Seconds, arXiv 2025. [Paper] [Website] [Code]
  • SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction, arXiv 2024. [Paper] [Code]
  • Align3R: Aligned Monocular Depth Estimation for Dynamic Videos, CVPR 2025. [Paper] [Code] [Website]
  • (MASt3R) Grounding Image Matching in 3D with MASt3R, ECCV 2024. [Paper] [Code] [Website]
  • VGGT: Visual Geometry Grounded Transformer, CVPR 2025. [Paper] [Code] [Website]
  • E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models, arXiv 2025. [Paper] [Code] [Website]
  • π^3: Scalable Permutation-Equivariant Visual Geometry Learning, arXiv 2025. [Paper] [Code] [Website]
  • Dens3R: A Foundation Model for 3D Geometry Prediction, ICCV 2025. [Paper] [Code] [Website]
  • LONG3R: Long Sequence Streaming 3D Reconstruction, ICCV 2025. [Paper] [Code] [Website]
  • PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction, ICCV 2025. [Paper] [Code] [Website]
  • Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos, CVPR 2026. [Paper]
  • G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration, arXiv 2025. [Paper]
  • ViPE: Video Pose Engine for 3D Geometric Perception, arXiv 2025. [Paper] [Code] [Website]
  • FastVGGT: Training-Free Acceleration of Visual Geometry Transformer, arXiv 2025. [Paper]
  • SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization, arXiv 2025. [Paper] [Code] [Website]
  • Streaming 4D Visual Geometry Transformer, arXiv 2025. [Paper] [Code] [Website]
  • MapAnything: Universal Feed-Forward Metric 3D Reconstruction, arXiv 2025. [Paper] [Code] [Website]
  • TTT3R: 3D Reconstruction as Test-Time Training, arXiv 2025. [Paper] [Code] [Website]
  • Co-Me: Confidence Guided Token Merging for Visual Geometric Transformers, arXiv 2025. [Paper] [Code] [Website]
  • MB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend, arXiv 2025. [Paper]
  • VG3T: Visual Geometry Grounded Gaussian Transformer, arXiv 2025. [Paper]
  • KV-Tracker: Real-Time Pose Tracking with Transformers, arXiv 2025. [Paper] [Website]
  • V-DPM: 4D Video Reconstruction with Dynamic Point Mapss, arXiv 2026. [Paper] [Website] [Code]
  • S-MUSt3R: Sliding Multi-view 3D Reconstruction, arXiv 2026. [Paper]
  • Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning, arXiv 2026. [Paper] [Website] [Code]
  • VGG-T3: Offline Feed-Forward 3D Reconstruction at Scale, CVPR 2026. [Paper]
  • OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution, CVPR Finding 2026 2026. [Paper] [Website]
  • ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time Training, CVPR 2026. [Paper] [Code] [Website]
  • Dark3R: Learning Structure from Motion in the Dark, CVPR 2026. [Paper] [Website]
  • GGPT: Geometry-Grounded Point Transformer, CVPR 2026. [Paper] [Code] [Website]
  • Repurposing Geometric Foundation Models for Multi-view Diffusion, arXiv 2026. [Paper] [Code] [Website]
  • NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction, ICLR 2026. [Paper] [Website
  • StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision, arXiv 2026. [Paper] [Code]
  • Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses, arXiv 2026. [Paper]
  • Learning 3D Reconstruction with Priors in Test Time, arXiv 2026. [Paper] [Website
  • Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction, CVPR 2026. [Paper] [Website [Code]
  • AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model, arXiv 2026. [Paper] [Website [Code]
  • Emergent Extreme-View Geometry in 3D Foundation Models, CVPR 2026. [Paper] [Website [Code]
  • VGGT-Ω, CVPR 2026. [Paper] [Website [Code]
  • Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory, arXiv 2026. [Paper]
  • UNIT: Unified Geometry Learning with Group Autoregressive Transformer, arXiv 2026. [Paper] [Website [Code]
  • Global Structure-from-Motion Meets Feedforward Reconstruction, CVPR 2026. [Paper] [Website [Code]
  • TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction, arXiv 2026. [Paper] [Website] [Code]
  • R3: 3D Reconstruction via Relative Regression, arXiv 2026. [Paper] [Website [Code]
  • EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization, ECCV 2026. [Paper]
  • RayTun3R: Online Camera Adaptation in 3D Foundation Models, arXiv 2026. [Paper]
  • VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues, arXiv 2026. [Paper] [Website
  • Hierarchical Structure-from-Motion Scales Feedforward Reconstruction, arXiv 2026. [Paper]
  • RoMa-Ω: What Feed-Forward 3D Models Know About Image Matching, ECCVw 2026. [Paper] [Code]
  • What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility, arXiv 2026. [Paper] [Code]
  • A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction Models, arXiv 2026. [Paper]

SLAM

  • Hier-SLAM++: Neuro-symbolic semantic slam with a hierarchically categorical Gaussian splatting, arXiv 2025. [Paper]
  • MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos, CVPR 2025. [Paper] [Website] [Code]
  • SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos, CVPR 2025. [Paper] [Code]
  • MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors, CVPR 2025. [Paper] [Website] [Code]
  • VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold, arXiv 2025. [Paper] [Code]
  • Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps, ICCV 2025. [Paper] [Website] [Code]
  • VGGT-Long: Chunk it, Loop it, Align it – Pushing VGGT’s Limits on Kilometer-scale Long RGB Sequences, ICRA, 2025. [Paper] [Code]
  • Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline, IROS, 2025. [Paper] [Code]
  • 3D Foundation Model-Based Loop Closing for Decentralized Collaborative SLAM, RAL, 2025. [Paper]
  • ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association, 3DV, 2026. [Paper] [Code]
  • SLAM-Former: Putting SLAM into One Transformer, arXiv 2025. [Paper] [Website] [Code]
  • MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM, arXiv, 2025. [Paper] [Code]
  • GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State, ICRA, 2026. [Paper]
  • EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction, ICRA, 2026. [Paper] [Website] [Code]
  • Visual Odometry with Transformers, arXiv 2025. [Paper] [Website] [Code]
  • ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation, arXiv, 2025. [Paper] [Website] [Code]
  • MASt3R-GS: Bridging 3D Reconstruction Priors with Gaussian Splatting for Real-Time Dense SLAM, IROSw, 2025. [Paper]
  • LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping, RAL, 2026. [Paper] [Code]
  • Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM, arXiv, 2025. [Paper]
  • SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Prior, arXiv, 2025. [Paper]
  • KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM, arXiv, 2025. [Paper] [Code]
  • Dynamic Visual SLAM using a General 3D Prior, CVPR, 2026. [Paper] [Code]
  • OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics, arXiv, 2025. [Paper]
  • Keyframe-Based Feed-Forward Visual Odometry, arXiv 2026. [Paper]
  • VGGT-SLAM 2.0: Real time Dense Feed-forward Scene Reconstruction, RSS 2026. [Paper]
  • VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency, arXiv 2026. [Paper]
  • VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation, arXiv 2026. [Paper]
  • VGGT-Geo: Probabilistic Geometric Fusion of Visual Geometry Grounded Transformer Priors for Robust Dense Indoor SLAM, ISPRS International Journal of Geo-Information 2026. [Paper]
  • IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping, arXiv 2026. [Paper]
  • AIM-SLAM: Dense Monocular SLAM via Adaptive and Informative Multi-View Keyframe Prioritization with Foundation Model, ICRA, 2026. [Paper] [Website]
  • Flash-Mono: Feed-Forward Accelerated Gaussian Splatting Monocular SLAM, ICLR, 2026. [Paper] [Website]
  • HyVGGT-VO: Tightly Coupled Hybrid Dense Visual Odometry with Feed-Forward Models, RAL, 2026. [Paper] [Website] [Code]
  • VGGT-SLAM++, arXiv, 2026. [Paper]
  • Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model, CVPR, 2026. [Paper] [Website]
  • Accelerating Transformer-Based Monocular SLAM via Geometric Utility Scoring, arXiv, 2026. [Paper] [Website]
  • Keep It CALM: Toward Calibration-Free Kilometer-Level SLAM with Visual Geometry Foundation Models via an Assistant Eye, arXiv, 2026. [Paper] [Code]
  • MR.ScaleMaster: Scale-Consistent Collaborative Mapping from Crowd-Sourced Monocular Videos, arXiv, 2026. [Paper] [Code]
  • RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments, arXiv, 2026. [Paper] [Website] [Code]
  • MASTD3R-SLAM: Monocular Adaptive Semantic Tracking and Dynamic Reconstruction SLAM, ICRA, 2026. Paper-todo
  • AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend, CVPR, 2026. [Paper] [Website] [Code]
  • TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction, CVPR, 2026. [Paper] [Website] [Code]
  • (LingBot-Map) Geometric Context Transformer for Streaming 3D Reconstruction, arXiv 2026. [Paper] [Website [Code]
  • TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference, CVPR, 2026. Paper-todo
  • LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction, arXiv, 2026. [Paper]
  • M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM, arXiv, 2026. [Paper] [Website] [Code]
  • Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using a Feed-Forward 3D Model, RSS, 2026. [Paper] [Code]
  • PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM, arXiv, 2026. [Paper] [Website]
  • CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems, arXiv, 2026. [Paper] [Website]
  • ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM, arXiv, 2026. [Paper] [Code]
  • Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping, arXiv, 2026. [Paper]
  • GeoGS-SLAM: Online Monocular Reconstruction Using Gaussian Splatting with Geometric Priors, ICRA, 2026. [Paper] [Website]
  • MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos, ECCV 2026. [Paper] [Website]
  • Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM, arXiv, 2026. [Paper]
  • UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization, ECCV, 2026. [Paper] [Website]
  • SLAMFormer-∞: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing, arXiv, 2026. [Paper] [Website]
  • VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction, ACM MM, 2026. [Paper] [Code]
  • On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios, arXiv, 2026. [Paper] [Code]
  • VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM, arXiv, 2026. [Paper]
  • RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos, IROS, 2026. [Paper]
  • FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM, AAAI, 2026. [Paper]
  • Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs, CoRL, 2026. [Paper] [Code]
  • FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry, arXiv, 2026. [Paper]
  • SURE-Map: Self-Correcting Streaming Geometric Foundation Models, arXiv, 2026. [Paper] [Website] [Code]
  • VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors, arXiv, 2026. [Paper]
  • AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend, arXiv, 2026. [Paper] [Website] [Code]
  • GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction, arXiv, 2026. [Paper]
  • Dense Monocular SLAM in Real-Time With Structured Gaussian Representation, RAL, 2026. [Paper]
  • Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering, arXiv, 2026. [Paper] [Code]
  • NOCTIF3R: Feed-Forward Monocular Real-Time SLAM for Photon-Limited Scenes on Embedded Hardware, arXiv, 2026. [Paper]
  • DAVIO: Dense Monocular–Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping, arXiv, 2026. [Paper] [Website] [Code]
  • PROFusion: Robust and Accurate Dense Reconstruction via Camera Pose Regression and Optimization, ICRA, 2026. [Paper] [Code]
  • World SLAM Model: Joint World Modeling for SLAM and Navigation, arXiv, 2026. [Paper] [Code] [Website]
  • StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry, arXiv, 2026. [Paper] [Code] [Website]
  • CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction, NeurIPS, 2026. [Paper] [Code]
  • Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors, arXiv, 2026. [Paper] [Code] [Website]

Calibration

  • Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction, arXiv 2025. [Paper]

Loop Closure Detection & Place Recognition

  • Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments, arXiv 2025. [Paper]
  • VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments, arXiv 2026. [Paper]
  • UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer, arXiv 2025. [Paper] [Code]
  • Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer, arXiv 2025. [Paper] [Code]

Downstream Tasks

  • SpatialLM: Training Large Language Models for Structured Indoor Modeling, NIPS 2025. [Paper] [Website] [Code]
  • SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation, arXiv 2025. [Paper] [Website] [Code]
  • Drone-Based Indoor Mapping for Augmented Indoor 3D Modeling Using Neural Simultaneous Localization And Mapping (SLAM), * IEEE International Conference on Consumer Electronics (ICCE) 2026*. [Paper]
3d-geometric-foundation-models
dust3r
mast3r
slam
vggt

Significant stargazers

Ryohei Sasaki

696 followers · starred May 2026

3D-Vision-World/All-3R-SLAM-in-this-Repo

This is a list of relevant papers for 3D Geometric Foundation Models and Applications.

287

158 commits

updated Oct 3, 2026

See the code

README

3D Geometric Foundation Models (3R) with SLAM

Survey Paper

  • Advances in Feed-Forward 3D Reconstruction and View Synthesis: A Survey, arXiv 2025. [Paper] [Website]
  • Feed-Forward 3D Scene Modeling: A Problem-Driven Perspective, arXiv 2026. [Paper] [Website]

3D Geometric Foundation Models (3R)

  • DUSt3R: Geometric 3D Vision Made Easy, CVPR 2024. [Paper] [Code] [Website]
  • Monst3r: A simple approach for estimating geometry in the presence of motion, ICLR 2025. [Paper] [Code] [Website]
  • LoRA3D: Low-Rank Self-Calibration of 3D Geometric Foundation Models, ICLR 2025. [Paper] [Website]
  • (CUT3R) Continuous 3D Perception Model with Persistent State, CVPR 2025. [Paper] [Code] [Website]
  • Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual Localization, CVPR 2025. [Paper] [Code]
  • DAS3R: Dynamics-Aware Gaussian Splatting for Static Scene Reconstruction, arXiv 2024. [Paper] [Code] [Website]
  • MASt3R-SfM: a Fully-Integrated Solution for Unconstrained Structure-from-Motion, 3DV 2025. [Paper] [Code]
  • Splatt3R: Zero-shot Gaussian Splatting from Uncalibrated Image Pairs, arXiv 2024. [Paper] [Project] [Code]
  • SAB3R: Semantic-Augmented Backbone in 3D Reconstruction, arXiv 2024. [Paper]
  • No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images, ICLR 2025. [Paper] [Code] [Website]
  • Doppelgangers++: Improved Visual Disambiguation with Geometric 3D Features, CVPR 2025. [Paper] [Code] [Project]
  • (Spann3R) 3D Reconstruction with Spatial Memory, 3DV 2025. [Paper] [Code] [Website]
  • Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass, CVPR 2025. [Paper] [Website] [Code]
  • InstantSplat: Unbounded Sparse-view Pose-free Gaussian Splatting in 40 Seconds, arXiv 2025. [Paper] [Website] [Code]
  • SPARS3R: Semantic Prior Alignment and Regularization for Sparse 3D Reconstruction, arXiv 2024. [Paper] [Code]
  • Align3R: Aligned Monocular Depth Estimation for Dynamic Videos, CVPR 2025. [Paper] [Code] [Website]
  • (MASt3R) Grounding Image Matching in 3D with MASt3R, ECCV 2024. [Paper] [Code] [Website]
  • VGGT: Visual Geometry Grounded Transformer, CVPR 2025. [Paper] [Code] [Website]
  • E3D-Bench: A Benchmark for End-to-End 3D Geometric Foundation Models, arXiv 2025. [Paper] [Code] [Website]
  • π^3: Scalable Permutation-Equivariant Visual Geometry Learning, arXiv 2025. [Paper] [Code] [Website]
  • Dens3R: A Foundation Model for 3D Geometry Prediction, ICCV 2025. [Paper] [Code] [Website]
  • LONG3R: Long Sequence Streaming 3D Reconstruction, ICCV 2025. [Paper] [Code] [Website]
  • PanoSplatt3R: Leveraging Perspective Pretraining for Generalized Unposed Wide-Baseline Panorama Reconstruction, ICCV 2025. [Paper] [Code] [Website]
  • Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos, CVPR 2026. [Paper]
  • G-CUT3R: Guided 3D Reconstruction with Camera and Depth Prior Integration, arXiv 2025. [Paper]
  • ViPE: Video Pose Engine for 3D Geometric Perception, arXiv 2025. [Paper] [Code] [Website]
  • FastVGGT: Training-Free Acceleration of Visual Geometry Transformer, arXiv 2025. [Paper]
  • SAIL-Recon: Large SfM by Augmenting Scene Regression with Localization, arXiv 2025. [Paper] [Code] [Website]
  • Streaming 4D Visual Geometry Transformer, arXiv 2025. [Paper] [Code] [Website]
  • MapAnything: Universal Feed-Forward Metric 3D Reconstruction, arXiv 2025. [Paper] [Code] [Website]
  • TTT3R: 3D Reconstruction as Test-Time Training, arXiv 2025. [Paper] [Code] [Website]
  • Co-Me: Confidence Guided Token Merging for Visual Geometric Transformers, arXiv 2025. [Paper] [Code] [Website]
  • MB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend, arXiv 2025. [Paper]
  • VG3T: Visual Geometry Grounded Gaussian Transformer, arXiv 2025. [Paper]
  • KV-Tracker: Real-Time Pose Tracking with Transformers, arXiv 2025. [Paper] [Website]
  • V-DPM: 4D Video Reconstruction with Dynamic Point Mapss, arXiv 2026. [Paper] [Website] [Code]
  • S-MUSt3R: Sliding Multi-view 3D Reconstruction, arXiv 2026. [Paper]
  • Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning, arXiv 2026. [Paper] [Website] [Code]
  • VGG-T3: Offline Feed-Forward 3D Reconstruction at Scale, CVPR 2026. [Paper]
  • OnlineX: Unified Online 3D Reconstruction and Understanding with Active-to-Stable State Evolution, CVPR Finding 2026 2026. [Paper] [Website]
  • ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time Training, CVPR 2026. [Paper] [Code] [Website]
  • Dark3R: Learning Structure from Motion in the Dark, CVPR 2026. [Paper] [Website]
  • GGPT: Geometry-Grounded Point Transformer, CVPR 2026. [Paper] [Code] [Website]
  • Repurposing Geometric Foundation Models for Multi-view Diffusion, arXiv 2026. [Paper] [Code] [Website]
  • NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction, ICLR 2026. [Paper] [Website
  • StereoVGGT: A Training-Free Visual Geometry Transformer for Stereo Vision, arXiv 2026. [Paper] [Code]
  • Fisheye3R: Adapting Unified 3D Feed-Forward Foundation Models to Fisheye Lenses, arXiv 2026. [Paper]
  • Learning 3D Reconstruction with Priors in Test Time, arXiv 2026. [Paper] [Website
  • Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction, CVPR 2026. [Paper] [Website [Code]
  • AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model, arXiv 2026. [Paper] [Website [Code]
  • Emergent Extreme-View Geometry in 3D Foundation Models, CVPR 2026. [Paper] [Website [Code]
  • VGGT-Ω, CVPR 2026. [Paper] [Website [Code]
  • Mamba-VGGT: Persistent Long-Sequence Video Geometry Grounded Transformer via External Sliding Window Mamba Memory, arXiv 2026. [Paper]
  • UNIT: Unified Geometry Learning with Group Autoregressive Transformer, arXiv 2026. [Paper] [Website [Code]
  • Global Structure-from-Motion Meets Feedforward Reconstruction, CVPR 2026. [Paper] [Website [Code]
  • TriSplat: Simulation-Ready Feed-Forward 3D Scene Reconstruction, arXiv 2026. [Paper] [Website] [Code]
  • R3: 3D Reconstruction via Relative Regression, arXiv 2026. [Paper] [Website [Code]
  • EPO: Boosting 3D Foundation Models with Edge-based Pose Optimization, ECCV 2026. [Paper]
  • RayTun3R: Online Camera Adaptation in 3D Foundation Models, arXiv 2026. [Paper]
  • VI3: Grounding Pretrained 3D Foundation Models with Inertial Cues, arXiv 2026. [Paper] [Website
  • Hierarchical Structure-from-Motion Scales Feedforward Reconstruction, arXiv 2026. [Paper]
  • RoMa-Ω: What Feed-Forward 3D Models Know About Image Matching, ECCVw 2026. [Paper] [Code]
  • What VGGT Knows About Overlap: Probing Geometric Foundation Models for Co-Visibility, arXiv 2026. [Paper] [Code]
  • A Calibration Audit of Confidence in Feed-Forward 3D Reconstruction Models, arXiv 2026. [Paper]

SLAM

  • Hier-SLAM++: Neuro-symbolic semantic slam with a hierarchically categorical Gaussian splatting, arXiv 2025. [Paper]
  • MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos, CVPR 2025. [Paper] [Website] [Code]
  • SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos, CVPR 2025. [Paper] [Code]
  • MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction Priors, CVPR 2025. [Paper] [Website] [Code]
  • VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold, arXiv 2025. [Paper] [Code]
  • Outdoor Monocular SLAM with Global Scale-Consistent 3D Gaussian Pointmaps, ICCV 2025. [Paper] [Website] [Code]
  • VGGT-Long: Chunk it, Loop it, Align it – Pushing VGGT’s Limits on Kilometer-scale Long RGB Sequences, ICRA, 2025. [Paper] [Code]
  • Pseudo Depth Meets Gaussian: A Feed-forward RGB SLAM Baseline, IROS, 2025. [Paper] [Code]
  • 3D Foundation Model-Based Loop Closing for Decentralized Collaborative SLAM, RAL, 2025. [Paper]
  • ViSTA-SLAM: Visual SLAM with Symmetric Two-view Association, 3DV, 2026. [Paper] [Code]
  • SLAM-Former: Putting SLAM into One Transformer, arXiv 2025. [Paper] [Website] [Code]
  • MASt3R-Fusion: Integrating Feed-Forward Visual Model with IMU, GNSS for High-Functionality SLAM, arXiv, 2025. [Paper] [Code]
  • GRS-SLAM3R: Real-Time Dense SLAM with Gated Recurrent State, ICRA, 2026. [Paper]
  • EC3R-SLAM: Efficient and Consistent Monocular Dense SLAM with Feed-Forward 3D Reconstruction, ICRA, 2026. [Paper] [Website] [Code]
  • Visual Odometry with Transformers, arXiv 2025. [Paper] [Website] [Code]
  • ARTDECO: Towards Efficient and High-Fidelity On-the-Fly 3D Reconstruction with Structured Scene Representation, arXiv, 2025. [Paper] [Website] [Code]
  • MASt3R-GS: Bridging 3D Reconstruction Priors with Gaussian Splatting for Real-Time Dense SLAM, IROSw, 2025. [Paper]
  • LiDAR-VGGT: Cross-Modal Coarse-to-Fine Fusion for Globally Consistent and Metric-Scale Dense Mapping, RAL, 2026. [Paper] [Code]
  • Building temporally coherent 3D maps with VGGT for memory-efficient Semantic SLAM, arXiv, 2025. [Paper]
  • SING3R-SLAM: Submap-based Indoor Monocular Gaussian SLAM with 3D Reconstruction Prior, arXiv, 2025. [Paper]
  • KM-ViPE: Online Tightly Coupled Vision-Language-Geometry Fusion for Open-Vocabulary Semantic SLAM, arXiv, 2025. [Paper] [Code]
  • Dynamic Visual SLAM using a General 3D Prior, CVPR, 2026. [Paper] [Code]
  • OpenMonoGS-SLAM: Monocular Gaussian Splatting SLAM with Open-set Semantics, arXiv, 2025. [Paper]
  • Keyframe-Based Feed-Forward Visual Odometry, arXiv 2026. [Paper]
  • VGGT-SLAM 2.0: Real time Dense Feed-forward Scene Reconstruction, RSS 2026. [Paper]
  • VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency, arXiv 2026. [Paper]
  • VGGT-based online 3D semantic SLAM for indoor scene understanding and navigation, arXiv 2026. [Paper]
  • VGGT-Geo: Probabilistic Geometric Fusion of Visual Geometry Grounded Transformer Priors for Robust Dense Indoor SLAM, ISPRS International Journal of Geo-Information 2026. [Paper]
  • IRIS-SLAM: Unified Geo-Instance Representations for Robust Semantic Localization and Mapping, arXiv 2026. [Paper]
  • AIM-SLAM: Dense Monocular SLAM via Adaptive and Informative Multi-View Keyframe Prioritization with Foundation Model, ICRA, 2026. [Paper] [Website]
  • Flash-Mono: Feed-Forward Accelerated Gaussian Splatting Monocular SLAM, ICLR, 2026. [Paper] [Website]
  • HyVGGT-VO: Tightly Coupled Hybrid Dense Visual Odometry with Feed-Forward Models, RAL, 2026. [Paper] [Website] [Code]
  • VGGT-SLAM++, arXiv, 2026. [Paper]
  • Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model, CVPR, 2026. [Paper] [Website]
  • Accelerating Transformer-Based Monocular SLAM via Geometric Utility Scoring, arXiv, 2026. [Paper] [Website]
  • Keep It CALM: Toward Calibration-Free Kilometer-Level SLAM with Visual Geometry Foundation Models via an Assistant Eye, arXiv, 2026. [Paper] [Code]
  • MR.ScaleMaster: Scale-Consistent Collaborative Mapping from Crowd-Sourced Monocular Videos, arXiv, 2026. [Paper] [Code]
  • RADIO-ViPE: Online Tightly Coupled Multi-Modal Fusion for Open-Vocabulary Semantic SLAM in Dynamic Environments, arXiv, 2026. [Paper] [Website] [Code]
  • MASTD3R-SLAM: Monocular Adaptive Semantic Tracking and Dynamic Reconstruction SLAM, ICRA, 2026. Paper-todo
  • AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend, CVPR, 2026. [Paper] [Website] [Code]
  • TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction, CVPR, 2026. [Paper] [Website] [Code]
  • (LingBot-Map) Geometric Context Transformer for Streaming 3D Reconstruction, arXiv 2026. [Paper] [Website [Code]
  • TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference, CVPR, 2026. Paper-todo
  • LEXI-SG: Monocular 3D Scene Graph Mapping with Room-Guided Feed-Forward Reconstruction, arXiv, 2026. [Paper]
  • M^3: Dense Matching Meets Multi-View Foundation Models for Monocular Gaussian Splatting SLAM, arXiv, 2026. [Paper] [Website] [Code]
  • Efficient Feature-Free Initialization for Monocular Visual-Inertial Systems Using a Feed-Forward 3D Model, RSS, 2026. [Paper] [Code]
  • PRISM-SLAM: Probabilistic Ray-Grounded Inference for Scale-aware Metric SLAM, arXiv, 2026. [Paper] [Website]
  • CoMo3R-SLAM: Collaborative Monocular Dense SLAM with Learned 3D Reconstruction Priors for Outdoor Multi-Agent Systems, arXiv, 2026. [Paper] [Website]
  • ScaRF-SLAM: Scale-Consistent Reconstruction with Feed-Forward Models and Classical Visual SLAM, arXiv, 2026. [Paper] [Code]
  • Anchor3R: Streaming 3D Reconstruction with Transient Anchors for Long-Horizon Visual Mapping, arXiv, 2026. [Paper]
  • GeoGS-SLAM: Online Monocular Reconstruction Using Gaussian Splatting with Geometric Priors, ICRA, 2026. [Paper] [Website]
  • MAGiSt3R: Multi-Agent Feed-forward 3D Reconstruction from Monocular RGB Videos, ECCV 2026. [Paper] [Website]
  • Look Up and Look Back: Hidden Attention and Latent Orientation in a Frozen Foundation Model for Panoramic SLAM, arXiv, 2026. [Paper]
  • UniSim-SLAM: Feed-Forward SLAM with Unified Sim(3) Optimization, ECCV, 2026. [Paper] [Website]
  • SLAMFormer-∞: Infinite SLAM Transformer for Unbounded Frontend and Backend Processing, arXiv, 2026. [Paper] [Website]
  • VGGT-Align: Bridging Local Reconstruction and Global Consistency for Long-Sequence 3D Reconstruction, ACM MM, 2026. [Paper] [Code]
  • On-the-Fly3R: Towards Robust Online 3D Reconstruction with Feed-Forward 3R Models for Large-Scale UAV Scenarios, arXiv, 2026. [Paper] [Code]
  • VOIM: Training-Free Open-Vocabulary 3D Instance Mapping for RGB-D and Monocular SLAM, arXiv, 2026. [Paper]
  • RoSe-SLAM: Robust Semantic-Aware Gaussian Splatting SLAM from Dynamic Monocular Videos, IROS, 2026. [Paper]
  • FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAM, AAAI, 2026. [Paper]
  • Functional-SLAM: Interaction-Aware Mapping with Online Functional Scene Graphs, CoRL, 2026. [Paper] [Code]
  • FFVO: A Feedforward Pose Decoder for Long-Horizon Visual Odometry, arXiv, 2026. [Paper]
  • SURE-Map: Self-Correcting Streaming Geometric Foundation Models, arXiv, 2026. [Paper] [Website] [Code]
  • VGGT-GS SLAM: Uncalibrated Monocular Gaussian Splatting SLAM with Feed-Forward Priors, arXiv, 2026. [Paper]
  • AMB3R-SLAM: Kilometer-scale SLAM with Hierarchical Backend, arXiv, 2026. [Paper] [Website] [Code]
  • GRF-Recon: Global Ray-Field Optimization for Long-Sequence Feed-forward Reconstruction, arXiv, 2026. [Paper]
  • Dense Monocular SLAM in Real-Time With Structured Gaussian Representation, RAL, 2026. [Paper]
  • Adaptive World Memory 3D Foundation Model for Scalable 3D Mapping, Localization, and Rendering, arXiv, 2026. [Paper] [Code]
  • NOCTIF3R: Feed-Forward Monocular Real-Time SLAM for Photon-Limited Scenes on Embedded Hardware, arXiv, 2026. [Paper]
  • DAVIO: Dense Monocular–Inertial SLAM with Feed-Forward Initialization and Pose-Conditioned Mapping, arXiv, 2026. [Paper] [Website] [Code]
  • PROFusion: Robust and Accurate Dense Reconstruction via Camera Pose Regression and Optimization, ICRA, 2026. [Paper] [Code]
  • World SLAM Model: Joint World Modeling for SLAM and Navigation, arXiv, 2026. [Paper] [Code] [Website]
  • StreamRig: Exploiting Intra-Rig Geometry for Streaming Multi-Camera Odometry, arXiv, 2026. [Paper] [Code] [Website]
  • CLoSeR: Closing the Loop for Long-Context Streaming Reconstruction, NeurIPS, 2026. [Paper] [Code]
  • Pow3R-SLAM: Real-Time RGB-D SLAM with 3D Reconstruction Priors, arXiv, 2026. [Paper] [Code] [Website]

Calibration

  • Calib3R: A 3D Foundation Model for Multi-Camera to Robot Calibration and 3D Metric-Scaled Scene Reconstruction, arXiv 2025. [Paper]

Loop Closure Detection & Place Recognition

  • Multi-modal Loop Closure Detection with Foundation Models in Severely Unstructured Environments, arXiv 2025. [Paper]
  • VGGT-MPR: VGGT-Enhanced Multimodal Place Recognition in Autonomous Driving Environments, arXiv 2026. [Paper]
  • UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer, arXiv 2025. [Paper] [Code]
  • Reloc-VGGT: Visual Re-localization with Geometry Grounded Transformer, arXiv 2025. [Paper] [Code]

Downstream Tasks

  • SpatialLM: Training Large Language Models for Structured Indoor Modeling, NIPS 2025. [Paper] [Website] [Code]
  • SpatialAnt: Autonomous Zero-Shot Robot Navigation via Active Scene Reconstruction and Visual Anticipation, arXiv 2025. [Paper] [Website] [Code]
  • Drone-Based Indoor Mapping for Augmented Indoor 3D Modeling Using Neural Simultaneous Localization And Mapping (SLAM), * IEEE International Conference on Consumer Electronics (ICCE) 2026*. [Paper]
3d-geometric-foundation-models
dust3r
mast3r
slam
vggt

Significant stargazers

Ryohei Sasaki

696 followers · starred May 2026