This is a collective repository for all 3D and 4D Reconstruction papers
44
15 commits
updated May 22, 2026
A curated collection of cutting-edge research papers on Feed-Forward 3D and 4D Reconstruction.
If you find this repo useful, please consider giving it a β!
| List | Description | |
|---|---|---|
| βοΈ | Awesome 3D/4D Editing | Papers on 3D and 4D scene/object editing |
| π¨ | Awesome 3D Generation | Papers on 3D/4D content generation |
| Section | Topic |
|---|---|
| ποΈ | 3D FeedForward Reconstruction |
| π | Dynamic FeedForward Reconstruction |
| π‘ | Streaming Reconstruction |
| π | Tracking |
End-to-end feed-forward models for static 3D scene reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion | π |
| 2 | CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow | π |
| 3 | VGGSfM: Visual Geometry Grounded Deep Structure From Motion | π |
| 4 | DUSt3R: Geometric 3D Vision Made Easy | π |
| 5 | Grounding Image Matching in 3D with MASt3R | π |
| 6 | 3D Reconstruction with Spatial Memory | π |
| 7 | MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion | π |
| 8 | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | π |
| 9 | MEt3R: Measuring Multi-View Consistency in Generated Images | π |
| 10 | Continuous 3D Perception Model with Persistent State | π |
| 11 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | π |
| 12 | FLARE: Feed-Forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | π |
| 13 | VGGT: Visual Geometry Grounded Transformer | π |
| 14 | Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | π |
| 15 | Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction | π |
| 16 | Test3R: Learning to Reconstruct 3D at Test Time | π |
| 17 | ΟΒ³: Scalable Permutation-Equivariant Visual Geometry Learning | π |
| 18 | Dens3R: A Foundation Model for 3D Geometry Prediction | π |
| 19 | VGGT-Long: Chunk It, Loop It, Align It β Pushing VGGT's Limits on Kilometer-Scale Long RGB Sequences | π |
| 20 | FastVGGT: Training-Free Acceleration of Visual Geometry Transformer | π |
| 21 | Faster VGGT with Block-Sparse Global Attention | π |
| 22 | MapAnything: Universal Feed-Forward Metric 3D Reconstruction | π |
| 23 | Quantized Visual Geometry Grounded Transformer | π |
| 24 | VGGT-X: When VGGT Meets Dense Novel View Synthesis | π |
| 25 | WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting | π |
| 26 | IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction | π |
| 27 | MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts | π |
| 28 | OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer | π |
| 29 | Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers | π |
| 30 | SwiftVGGT: A scalable visual geometry grounded transformer for large-scale scenes | π |
| 31 | HTTM: Head-Wise Temporal Token Merging for Faster VGGT | π |
| 32 | LiteVGGT: Boosting Vanilla VGGT via Geometry-Aware Cached Token Merging | π |
| 33 | MoE3D: A Mixture-of-Experts Module for 3D Reconstruction | π |
| 34 | Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video | π |
| 35 | Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning | π |
| 36 | VGG-TΒ³: Offline Feed-Forward 3D Reconstruction at Scale | π |
| 37 | DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation | π |
| 38 | ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time Training | π |
| 39 | HD-VGGT: High-Resolution Visual Geometry Transformer | π |
| 40 | Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model | π |
| 41 | Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself | π |
| 42 | Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation | π |
| 43 | Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction | π |
| 44 | PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers | π |
| 45 | Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction | π |
| 46 | TurboVGGT: Fast visual geometry reconstruction with adaptive alternating attention | π |
| 47 | VGGT-Ο | π |
| 48 | Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer | π |
| 49 | Trust it or not: Evidential uncertainty for feed-forward 3D reconstruction with Trust3R | π |
Feed-forward methods for dynamic / 4D scene reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion | π |
| 2 | Align3R: Aligned Monocular Depth Estimation for Dynamic Videos | π |
| 3 | MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos | π |
| 4 | Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving | π |
| 5 | Easi3R: Estimating Disentangled Motion from DUSt3R Without Training | π |
| 6 | POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction | π |
| 7 | DΒ²USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic Scenes | π |
| 8 | St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World | π |
| 9 | Human3R: Everyone Everywhere All at Once | π |
| 10 | PAGE-4D: DISENTANGLED POSE AND GEOMETRY ESTIMATION FOR VGGT-4D PERCEPTION | π |
| 11 | 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation | π |
| 12 | VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction | π |
| 13 | Efficiently Reconstructing Dynamic Scenes One D4RT at a Time | π |
| 14 | DVGT: Driving Visual Geometry Transformer | π |
| 15 | 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere | π |
| 16 | Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow | π |
| 17 | MoRe: Motion-Aware Feed-Forward 4D Reconstruction Transformer | π |
| 18 | DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving | π |
| 19 | Complet4R: Geometric Complete 4D Reconstruction | π |
| 20 | Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors | π |
| 21 | 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation | π |
Real-time and streaming approaches for sequential 3D/4D reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | π |
| 2 | Continuous 3D Perception Model with Persistent State | π |
| 3 | Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory | π |
| 4 | Streaming 4D Visual Geometry Transformer | π |
| 5 | LONG3R: Long Sequence Streaming 3D Reconstruction | π |
| 6 | Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction | π |
| 7 | STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer | π |
| 8 | WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool | π |
| 9 | TTT3R: 3D Reconstruction as Test-Time Training | π |
| 10 | MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction | π |
| 11 | InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams | π |
| 12 | TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction | π |
| 13 | LongStream: Long-Sequence Streaming Autoregressive Visual Geometry | π |
| 14 | XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression | π |
| 15 | LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory | π |
| 16 | FrameVGGT: Frame Evidence Rolling Memory for Streaming VGGT | π |
| 17 | MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction | π |
| 18 | RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction | π |
| 19 | PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video Sequences | π |
| 20 | Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training | π |
| 21 | Geometric Context Transformer for Streaming 3D Reconstruction | π |
| 22 | StreamCacheVGGT: Streaming visual geometry transformers with robust scoring and hybrid cache compression | π |
| 23 | Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D Reconstruction | π |
| 24 | Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval | π |
| 25 | GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction | π |
| 26 | Rethinking the state update gate for long-sequence recurrent 3D reconstruction | π |
| 27 | Mamba-VGGT: Persistent long-sequence video geometry grounded transformer via external sliding window mamba memory | π |
Point tracking methods in 2D and 3D space.
| # | Paper | Link |
|---|---|---|
| 1 | TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal Refinement | π |
| 2 | CoTracker: It Is Better to Track Together | π |
| 3 | SpatialTracker: Tracking Any 2D Pixels in 3D Space | π |
| 4 | CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos | π |
| 5 | DELTA: Dense Efficient Long-Range 3D Tracking for Any Video | π |
| 6 | TAPIP3D: Tracking Any Point in Persistent 3D Geometry | π |
| 7 | SpatialTrackerV2: 3D Point Tracking Made Easy | π |
If you find this repository useful, please consider giving it a β
Made with β€οΈ for the 3D Vision Community
15 commits
This is a collective repository for all 3D and 4D Reconstruction papers
44
15 commits
updated May 22, 2026
A curated collection of cutting-edge research papers on Feed-Forward 3D and 4D Reconstruction.
If you find this repo useful, please consider giving it a β!
| List | Description | |
|---|---|---|
| βοΈ | Awesome 3D/4D Editing | Papers on 3D and 4D scene/object editing |
| π¨ | Awesome 3D Generation | Papers on 3D/4D content generation |
| Section | Topic |
|---|---|
| ποΈ | 3D FeedForward Reconstruction |
| π | Dynamic FeedForward Reconstruction |
| π‘ | Streaming Reconstruction |
| π | Tracking |
End-to-end feed-forward models for static 3D scene reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View Completion | π |
| 2 | CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical Flow | π |
| 3 | VGGSfM: Visual Geometry Grounded Deep Structure From Motion | π |
| 4 | DUSt3R: Geometric 3D Vision Made Easy | π |
| 5 | Grounding Image Matching in 3D with MASt3R | π |
| 6 | 3D Reconstruction with Spatial Memory | π |
| 7 | MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion | π |
| 8 | MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 Seconds | π |
| 9 | MEt3R: Measuring Multi-View Consistency in Generated Images | π |
| 10 | Continuous 3D Perception Model with Persistent State | π |
| 11 | Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass | π |
| 12 | FLARE: Feed-Forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse Views | π |
| 13 | VGGT: Visual Geometry Grounded Transformer | π |
| 14 | Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene Priors | π |
| 15 | Mono3R: Exploiting Monocular Cues for Geometric 3D Reconstruction | π |
| 16 | Test3R: Learning to Reconstruct 3D at Test Time | π |
| 17 | ΟΒ³: Scalable Permutation-Equivariant Visual Geometry Learning | π |
| 18 | Dens3R: A Foundation Model for 3D Geometry Prediction | π |
| 19 | VGGT-Long: Chunk It, Loop It, Align It β Pushing VGGT's Limits on Kilometer-Scale Long RGB Sequences | π |
| 20 | FastVGGT: Training-Free Acceleration of Visual Geometry Transformer | π |
| 21 | Faster VGGT with Block-Sparse Global Attention | π |
| 22 | MapAnything: Universal Feed-Forward Metric 3D Reconstruction | π |
| 23 | Quantized Visual Geometry Grounded Transformer | π |
| 24 | VGGT-X: When VGGT Meets Dense Novel View Synthesis | π |
| 25 | WorldMirror: Universal 3D World Reconstruction with Any-Prior Prompting | π |
| 26 | IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction | π |
| 27 | MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts | π |
| 28 | OmniVGGT: Omni-Modality Driven Visual Geometry Grounded Transformer | π |
| 29 | Co-Me: Confidence-Guided Token Merging for Visual Geometric Transformers | π |
| 30 | SwiftVGGT: A scalable visual geometry grounded transformer for large-scale scenes | π |
| 31 | HTTM: Head-Wise Temporal Token Merging for Faster VGGT | π |
| 32 | LiteVGGT: Boosting Vanilla VGGT via Geometry-Aware Cached Token Merging | π |
| 33 | MoE3D: A Mixture-of-Experts Module for 3D Reconstruction | π |
| 34 | Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet Video | π |
| 35 | Flow3r: Factored Flow Prediction for Scalable Visual Geometry Learning | π |
| 36 | VGG-TΒ³: Offline Feed-Forward 3D Reconstruction at Scale | π |
| 37 | DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry Estimation | π |
| 38 | ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time Training | π |
| 39 | HD-VGGT: High-Resolution Visual Geometry Transformer | π |
| 40 | Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation Model | π |
| 41 | Free Geometry: Refining 3D Reconstruction from Longer Versions of Itself | π |
| 42 | Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation | π |
| 43 | Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D Reconstruction | π |
| 44 | PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry Transformers | π |
| 45 | Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D Reconstruction | π |
| 46 | TurboVGGT: Fast visual geometry reconstruction with adaptive alternating attention | π |
| 47 | VGGT-Ο | π |
| 48 | Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry Transformer | π |
| 49 | Trust it or not: Evidential uncertainty for feed-forward 3D reconstruction with Trust3R | π |
Feed-forward methods for dynamic / 4D scene reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion | π |
| 2 | Align3R: Aligned Monocular Depth Estimation for Dynamic Videos | π |
| 3 | MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic Videos | π |
| 4 | Driv3R: Learning Dense 4D Reconstruction for Autonomous Driving | π |
| 5 | Easi3R: Estimating Disentangled Motion from DUSt3R Without Training | π |
| 6 | POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D Reconstruction | π |
| 7 | DΒ²USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic Scenes | π |
| 8 | St4RTrack: Simultaneous 4D Reconstruction and Tracking in the World | π |
| 9 | Human3R: Everyone Everywhere All at Once | π |
| 10 | PAGE-4D: DISENTANGLED POSE AND GEOMETRY ESTIMATION FOR VGGT-4D PERCEPTION | π |
| 11 | 4D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry Estimation | π |
| 12 | VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction | π |
| 13 | Efficiently Reconstructing Dynamic Scenes One D4RT at a Time | π |
| 14 | DVGT: Driving Visual Geometry Transformer | π |
| 15 | 4RC: 4D Reconstruction via Conditional Querying Anytime and Anywhere | π |
| 16 | Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow | π |
| 17 | MoRe: Motion-Aware Feed-Forward 4D Reconstruction Transformer | π |
| 18 | DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous Driving | π |
| 19 | Complet4R: Geometric Complete 4D Reconstruction | π |
| 20 | Robust 4D Visual Geometry Transformer with Uncertainty-Aware Priors | π |
| 21 | 4DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth Estimation | π |
Real-time and streaming approaches for sequential 3D/4D reconstruction.
| # | Paper | Link |
|---|---|---|
| 1 | SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB Videos | π |
| 2 | Continuous 3D Perception Model with Persistent State | π |
| 3 | Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer Memory | π |
| 4 | Streaming 4D Visual Geometry Transformer | π |
| 5 | LONG3R: Long Sequence Streaming 3D Reconstruction | π |
| 6 | Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction | π |
| 7 | STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer | π |
| 8 | WinT3R: Window-Based Streaming Reconstruction with Camera Token Pool | π |
| 9 | TTT3R: 3D Reconstruction as Test-Time Training | π |
| 10 | MUT3R: Motion-aware Updating Transformer for Dynamic 3D Reconstruction | π |
| 11 | InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams | π |
| 12 | TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D Reconstruction | π |
| 13 | LongStream: Long-Sequence Streaming Autoregressive Visual Geometry | π |
| 14 | XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache Compression | π |
| 15 | LoGeR: Long-Context Geometric Reconstruction with Hybrid Memory | π |
| 16 | FrameVGGT: Frame Evidence Rolling Memory for Streaming VGGT | π |
| 17 | MeMix: Writing Less, Remembering More for Streaming 3D Reconstruction | π |
| 18 | RayMap3R: Inference-Time RayMap for Dynamic 3D Reconstruction | π |
| 19 | PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video Sequences | π |
| 20 | Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time Training | π |
| 21 | Geometric Context Transformer for Streaming 3D Reconstruction | π |
| 22 | StreamCacheVGGT: Streaming visual geometry transformers with robust scoring and hybrid cache compression | π |
| 23 | Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D Reconstruction | π |
| 24 | Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity Retrieval | π |
| 25 | GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D Reconstruction | π |
| 26 | Rethinking the state update gate for long-sequence recurrent 3D reconstruction | π |
| 27 | Mamba-VGGT: Persistent long-sequence video geometry grounded transformer via external sliding window mamba memory | π |
Point tracking methods in 2D and 3D space.
| # | Paper | Link |
|---|---|---|
| 1 | TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal Refinement | π |
| 2 | CoTracker: It Is Better to Track Together | π |
| 3 | SpatialTracker: Tracking Any 2D Pixels in 3D Space | π |
| 4 | CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real Videos | π |
| 5 | DELTA: Dense Efficient Long-Range 3D Tracking for Any Video | π |
| 6 | TAPIP3D: Tracking Any Point in Persistent 3D Geometry | π |
| 7 | SpatialTrackerV2: 3D Point Tracking Made Easy | π |
If you find this repository useful, please consider giving it a β
Made with β€οΈ for the 3D Vision Community
15 commits