2hiTee/awesome-feedforward-3D-4D-Reconstruction

This is a collective repository for all 3D and 4D Reconstruction papers

44

15 commits

updated May 22, 2026

See the code

README

πŸ”„ Awesome FeedForward 3D/4D Reconstruction πŸ”„

FeedForward 3D/4D Reconstruction

A curated collection of cutting-edge research papers on Feed-Forward 3D and 4D Reconstruction.

Awesome PRs Welcome If you find this repo useful, please consider giving it a ⭐!


πŸ”— Explore Our Other Curated Lists

ListDescription
✏️Awesome 3D/4D EditingPapers on 3D and 4D scene/object editing
🎨Awesome 3D GenerationPapers on 3D/4D content generation

πŸ“‹ Table of Contents


πŸ”οΈ 3D FeedForward Reconstruction

End-to-end feed-forward models for static 3D scene reconstruction.

#PaperLink
1CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionπŸ“„
2CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowπŸ“„
3VGGSfM: Visual Geometry Grounded Deep Structure From MotionπŸ“„
4DUSt3R: Geometric 3D Vision Made EasyπŸ“„
5Grounding Image Matching in 3D with MASt3RπŸ“„
63D Reconstruction with Spatial MemoryπŸ“„
7MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionπŸ“„
8MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsπŸ“„
9MEt3R: Measuring Multi-View Consistency in Generated ImagesπŸ“„
10Continuous 3D Perception Model with Persistent StateπŸ“„
11Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassπŸ“„
12FLARE: Feed-Forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsπŸ“„
13VGGT: Visual Geometry Grounded TransformerπŸ“„
14Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene PriorsπŸ“„
15Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionπŸ“„
16Test3R: Learning to Reconstruct 3D at Test TimeπŸ“„
17π³: Scalable Permutation-Equivariant Visual Geometry LearningπŸ“„
18Dens3R: A Foundation Model for 3D Geometry PredictionπŸ“„
19VGGT-Long: Chunk It, Loop It, Align It β€” Pushing VGGT's Limits on Kilometer-Scale Long RGB SequencesπŸ“„
20FastVGGT: Training-Free Acceleration of Visual Geometry TransformerπŸ“„
21Faster VGGT with Block-Sparse Global AttentionπŸ“„
22MapAnything: Universal Feed-Forward Metric 3D ReconstructionπŸ“„
23Quantized Visual Geometry Grounded TransformerπŸ“„
24VGGT-X: When VGGT Meets Dense Novel View SynthesisπŸ“„
25WorldMirror: Universal 3D World Reconstruction with Any-Prior PromptingπŸ“„
26IGGT: Instance-Grounded Geometry Transformer for Semantic 3D ReconstructionπŸ“„
27MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsπŸ“„
28OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerπŸ“„
29Co-Me: Confidence-Guided Token Merging for Visual Geometric TransformersπŸ“„
30SwiftVGGT: A scalable visual geometry grounded transformer for large-scale scenesπŸ“„
31HTTM: Head-Wise Temporal Token Merging for Faster VGGTπŸ“„
32LiteVGGT: Boosting Vanilla VGGT via Geometry-Aware Cached Token MergingπŸ“„
33MoE3D: A Mixture-of-Experts Module for 3D ReconstructionπŸ“„
34Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet VideoπŸ“„
35Flow3r: Factored Flow Prediction for Scalable Visual Geometry LearningπŸ“„
36VGG-TΒ³: Offline Feed-Forward 3D Reconstruction at ScaleπŸ“„
37DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationπŸ“„
38ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time TrainingπŸ“„
39HD-VGGT: High-Resolution Visual Geometry TransformerπŸ“„
40Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation ModelπŸ“„
41Free Geometry: Refining 3D Reconstruction from Longer Versions of ItselfπŸ“„
42Unlocking the Power of Critical Factors for 3D Visual Geometry EstimationπŸ“„
43Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D ReconstructionπŸ“„
44PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry TransformersπŸ“„
45Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D ReconstructionπŸ“„
46TurboVGGT: Fast visual geometry reconstruction with adaptive alternating attentionπŸ“„
47VGGT-Ο‰πŸ“„
48Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry TransformerπŸ“„
49Trust it or not: Evidential uncertainty for feed-forward 3D reconstruction with Trust3RπŸ“„

(⬆️ back to top)


🌊 Dynamic FeedForward Reconstruction

Feed-forward methods for dynamic / 4D scene reconstruction.

#PaperLink
1MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionπŸ“„
2Align3R: Aligned Monocular Depth Estimation for Dynamic VideosπŸ“„
3MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic VideosπŸ“„
4Driv3R: Learning Dense 4D Reconstruction for Autonomous DrivingπŸ“„
5Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingπŸ“„
6POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D ReconstructionπŸ“„
7DΒ²USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic ScenesπŸ“„
8St4RTrack: Simultaneous 4D Reconstruction and Tracking in the WorldπŸ“„
9Human3R: Everyone Everywhere All at OnceπŸ“„
10PAGE-4D: DISENTANGLED POSE AND GEOMETRY ESTIMATION FOR VGGT-4D PERCEPTIONπŸ“„
114D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry EstimationπŸ“„
12VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene ReconstructionπŸ“„
13Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeπŸ“„
14DVGT: Driving Visual Geometry TransformerπŸ“„
154RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereπŸ“„
16Flow4R: Unifying 4D Reconstruction and Tracking with Scene FlowπŸ“„
17MoRe: Motion-Aware Feed-Forward 4D Reconstruction TransformerπŸ“„
18DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingπŸ“„
19Complet4R: Geometric Complete 4D ReconstructionπŸ“„
20Robust 4D Visual Geometry Transformer with Uncertainty-Aware PriorsπŸ“„
214DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth EstimationπŸ“„

(⬆️ back to top)


πŸ“‘ Streaming Reconstruction

Real-time and streaming approaches for sequential 3D/4D reconstruction.

#PaperLink
1SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosπŸ“„
2Continuous 3D Perception Model with Persistent StateπŸ“„
3Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer MemoryπŸ“„
4Streaming 4D Visual Geometry TransformerπŸ“„
5LONG3R: Long Sequence Streaming 3D ReconstructionπŸ“„
6Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene ReconstructionπŸ“„
7STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerπŸ“„
8WinT3R: Window-Based Streaming Reconstruction with Camera Token PoolπŸ“„
9TTT3R: 3D Reconstruction as Test-Time TrainingπŸ“„
10MUT3R: Motion-aware Updating Transformer for Dynamic 3D ReconstructionπŸ“„
11InfiniteVGGT: Visual Geometry Grounded Transformer for Endless StreamsπŸ“„
12TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D ReconstructionπŸ“„
13LongStream: Long-Sequence Streaming Autoregressive Visual GeometryπŸ“„
14XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache CompressionπŸ“„
15LoGeR: Long-Context Geometric Reconstruction with Hybrid MemoryπŸ“„
16FrameVGGT: Frame Evidence Rolling Memory for Streaming VGGTπŸ“„
17MeMix: Writing Less, Remembering More for Streaming 3D ReconstructionπŸ“„
18RayMap3R: Inference-Time RayMap for Dynamic 3D ReconstructionπŸ“„
19PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video SequencesπŸ“„
20Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time TrainingπŸ“„
21Geometric Context Transformer for Streaming 3D ReconstructionπŸ“„
22StreamCacheVGGT: Streaming visual geometry transformers with robust scoring and hybrid cache compressionπŸ“„
23Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D ReconstructionπŸ“„
24Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity RetrievalπŸ“„
25GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D ReconstructionπŸ“„
26Rethinking the state update gate for long-sequence recurrent 3D reconstructionπŸ“„
27Mamba-VGGT: Persistent long-sequence video geometry grounded transformer via external sliding window mamba memoryπŸ“„

(⬆️ back to top)


πŸ“ Tracking

Point tracking methods in 2D and 3D space.

#PaperLink
1TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal RefinementπŸ“„
2CoTracker: It Is Better to Track TogetherπŸ“„
3SpatialTracker: Tracking Any 2D Pixels in 3D SpaceπŸ“„
4CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real VideosπŸ“„
5DELTA: Dense Efficient Long-Range 3D Tracking for Any VideoπŸ“„
6TAPIP3D: Tracking Any Point in Persistent 3D GeometryπŸ“„
7SpatialTrackerV2: 3D Point Tracking Made EasyπŸ“„

(⬆️ back to top)


If you find this repository useful, please consider giving it a ⭐

Made with ❀️ for the 3D Vision Community

Contributors

2hiTee

15 commits

2hiTee/awesome-feedforward-3D-4D-Reconstruction

This is a collective repository for all 3D and 4D Reconstruction papers

44

15 commits

updated May 22, 2026

See the code

README

πŸ”„ Awesome FeedForward 3D/4D Reconstruction πŸ”„

FeedForward 3D/4D Reconstruction

A curated collection of cutting-edge research papers on Feed-Forward 3D and 4D Reconstruction.

Awesome PRs Welcome If you find this repo useful, please consider giving it a ⭐!


πŸ”— Explore Our Other Curated Lists

ListDescription
✏️Awesome 3D/4D EditingPapers on 3D and 4D scene/object editing
🎨Awesome 3D GenerationPapers on 3D/4D content generation

πŸ“‹ Table of Contents


πŸ”οΈ 3D FeedForward Reconstruction

End-to-end feed-forward models for static 3D scene reconstruction.

#PaperLink
1CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionπŸ“„
2CroCo v2: Improved Cross-view Completion Pre-training for Stereo Matching and Optical FlowπŸ“„
3VGGSfM: Visual Geometry Grounded Deep Structure From MotionπŸ“„
4DUSt3R: Geometric 3D Vision Made EasyπŸ“„
5Grounding Image Matching in 3D with MASt3RπŸ“„
63D Reconstruction with Spatial MemoryπŸ“„
7MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionπŸ“„
8MV-DUSt3R+: Single-Stage Scene Reconstruction from Sparse Views In 2 SecondsπŸ“„
9MEt3R: Measuring Multi-View Consistency in Generated ImagesπŸ“„
10Continuous 3D Perception Model with Persistent StateπŸ“„
11Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward PassπŸ“„
12FLARE: Feed-Forward Geometry, Appearance and Camera Estimation from Uncalibrated Sparse ViewsπŸ“„
13VGGT: Visual Geometry Grounded TransformerπŸ“„
14Pow3R: Empowering Unconstrained 3D Reconstruction with Camera and Scene PriorsπŸ“„
15Mono3R: Exploiting Monocular Cues for Geometric 3D ReconstructionπŸ“„
16Test3R: Learning to Reconstruct 3D at Test TimeπŸ“„
17π³: Scalable Permutation-Equivariant Visual Geometry LearningπŸ“„
18Dens3R: A Foundation Model for 3D Geometry PredictionπŸ“„
19VGGT-Long: Chunk It, Loop It, Align It β€” Pushing VGGT's Limits on Kilometer-Scale Long RGB SequencesπŸ“„
20FastVGGT: Training-Free Acceleration of Visual Geometry TransformerπŸ“„
21Faster VGGT with Block-Sparse Global AttentionπŸ“„
22MapAnything: Universal Feed-Forward Metric 3D ReconstructionπŸ“„
23Quantized Visual Geometry Grounded TransformerπŸ“„
24VGGT-X: When VGGT Meets Dense Novel View SynthesisπŸ“„
25WorldMirror: Universal 3D World Reconstruction with Any-Prior PromptingπŸ“„
26IGGT: Instance-Grounded Geometry Transformer for Semantic 3D ReconstructionπŸ“„
27MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-ExpertsπŸ“„
28OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerπŸ“„
29Co-Me: Confidence-Guided Token Merging for Visual Geometric TransformersπŸ“„
30SwiftVGGT: A scalable visual geometry grounded transformer for large-scale scenesπŸ“„
31HTTM: Head-Wise Temporal Token Merging for Faster VGGTπŸ“„
32LiteVGGT: Boosting Vanilla VGGT via Geometry-Aware Cached Token MergingπŸ“„
33MoE3D: A Mixture-of-Experts Module for 3D ReconstructionπŸ“„
34Scalable Adaptation of 3D Geometric Foundation Models via Weak Supervision from Internet VideoπŸ“„
35Flow3r: Factored Flow Prediction for Scalable Visual Geometry LearningπŸ“„
36VGG-TΒ³: Offline Feed-Forward 3D Reconstruction at ScaleπŸ“„
37DAGE: Dual-Stream Architecture for Efficient and Fine-Grained Geometry EstimationπŸ“„
38ZipMap: Linear-Time Stateful 3D Reconstruction with Test-Time TrainingπŸ“„
39HD-VGGT: High-Resolution Visual Geometry TransformerπŸ“„
40Online3R: Online Learning for Consistent Sequential Reconstruction Based on Geometry Foundation ModelπŸ“„
41Free Geometry: Refining 3D Reconstruction from Longer Versions of ItselfπŸ“„
42Unlocking the Power of Critical Factors for 3D Visual Geometry EstimationπŸ“„
43Spark3R: Asymmetric Token Reduction Makes Fast Feed-Forward 3D ReconstructionπŸ“„
44PaceVGGT: Pre-Alternating-Attention Token Pruning for Visual Geometry TransformersπŸ“„
45Lite3R: A Model-Agnostic Framework for Efficient Feed-Forward 3D ReconstructionπŸ“„
46TurboVGGT: Fast visual geometry reconstruction with adaptive alternating attentionπŸ“„
47VGGT-Ο‰πŸ“„
48Not All Tasks Quantize Equally: Fisher-Guided Quantization for Visual Geometry TransformerπŸ“„
49Trust it or not: Evidential uncertainty for feed-forward 3D reconstruction with Trust3RπŸ“„

(⬆️ back to top)


🌊 Dynamic FeedForward Reconstruction

Feed-forward methods for dynamic / 4D scene reconstruction.

#PaperLink
1MonST3R: A Simple Approach for Estimating Geometry in the Presence of MotionπŸ“„
2Align3R: Aligned Monocular Depth Estimation for Dynamic VideosπŸ“„
3MegaSaM: Accurate, Fast, and Robust Structure and Motion from Casual Dynamic VideosπŸ“„
4Driv3R: Learning Dense 4D Reconstruction for Autonomous DrivingπŸ“„
5Easi3R: Estimating Disentangled Motion from DUSt3R Without TrainingπŸ“„
6POMATO: Marrying Pointmap Matching with Temporal Motion for Dynamic 3D ReconstructionπŸ“„
7DΒ²USt3R: Enhancing 3D Reconstruction with 4D Pointmaps for Dynamic ScenesπŸ“„
8St4RTrack: Simultaneous 4D Reconstruction and Tracking in the WorldπŸ“„
9Human3R: Everyone Everywhere All at OnceπŸ“„
10PAGE-4D: DISENTANGLED POSE AND GEOMETRY ESTIMATION FOR VGGT-4D PERCEPTIONπŸ“„
114D-VGGT: A General Foundation Model with SpatioTemporal Awareness for Dynamic Scene Geometry EstimationπŸ“„
12VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene ReconstructionπŸ“„
13Efficiently Reconstructing Dynamic Scenes One D4RT at a TimeπŸ“„
14DVGT: Driving Visual Geometry TransformerπŸ“„
154RC: 4D Reconstruction via Conditional Querying Anytime and AnywhereπŸ“„
16Flow4R: Unifying 4D Reconstruction and Tracking with Scene FlowπŸ“„
17MoRe: Motion-Aware Feed-Forward 4D Reconstruction TransformerπŸ“„
18DynamicVGGT: Learning Dynamic Point Maps for 4D Scene Reconstruction in Autonomous DrivingπŸ“„
19Complet4R: Geometric Complete 4D ReconstructionπŸ“„
20Robust 4D Visual Geometry Transformer with Uncertainty-Aware PriorsπŸ“„
214DVGGT-D: 4D Visual Geometry Transformer with Improved Dynamic Depth EstimationπŸ“„

(⬆️ back to top)


πŸ“‘ Streaming Reconstruction

Real-time and streaming approaches for sequential 3D/4D reconstruction.

#PaperLink
1SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosπŸ“„
2Continuous 3D Perception Model with Persistent StateπŸ“„
3Point3R: Streaming 3D Reconstruction with Explicit Spatial Pointer MemoryπŸ“„
4Streaming 4D Visual Geometry TransformerπŸ“„
5LONG3R: Long Sequence Streaming 3D ReconstructionπŸ“„
6Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene ReconstructionπŸ“„
7STream3R: Scalable Sequential 3D Reconstruction with Causal TransformerπŸ“„
8WinT3R: Window-Based Streaming Reconstruction with Camera Token PoolπŸ“„
9TTT3R: 3D Reconstruction as Test-Time TrainingπŸ“„
10MUT3R: Motion-aware Updating Transformer for Dynamic 3D ReconstructionπŸ“„
11InfiniteVGGT: Visual Geometry Grounded Transformer for Endless StreamsπŸ“„
12TTSA3R: Training-Free Temporal-Spatial Adaptive Persistent State for Streaming 3D ReconstructionπŸ“„
13LongStream: Long-Sequence Streaming Autoregressive Visual GeometryπŸ“„
14XStreamVGGT: Extremely Memory-Efficient Streaming Vision Geometry Grounded Transformer with KV Cache CompressionπŸ“„
15LoGeR: Long-Context Geometric Reconstruction with Hybrid MemoryπŸ“„
16FrameVGGT: Frame Evidence Rolling Memory for Streaming VGGTπŸ“„
17MeMix: Writing Less, Remembering More for Streaming 3D ReconstructionπŸ“„
18RayMap3R: Inference-Time RayMap for Dynamic 3D ReconstructionπŸ“„
19PAS3R: Pose-Adaptive Streaming 3D Reconstruction for Long Video SequencesπŸ“„
20Mem3R: Streaming 3D Reconstruction with Hybrid Memory via Test-Time TrainingπŸ“„
21Geometric Context Transformer for Streaming 3D ReconstructionπŸ“„
22StreamCacheVGGT: Streaming visual geometry transformers with robust scoring and hybrid cache compressionπŸ“„
23Ray-Aware Pointer Memory with Adaptive Updates for Streaming 3D ReconstructionπŸ“„
24Attention Itself Could Retrieve. RetrieveVGGT: Training-Free Long Context Streaming 3D Reconstruction via Query-Key Similarity RetrievalπŸ“„
25GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D ReconstructionπŸ“„
26Rethinking the state update gate for long-sequence recurrent 3D reconstructionπŸ“„
27Mamba-VGGT: Persistent long-sequence video geometry grounded transformer via external sliding window mamba memoryπŸ“„

(⬆️ back to top)


πŸ“ Tracking

Point tracking methods in 2D and 3D space.

#PaperLink
1TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal RefinementπŸ“„
2CoTracker: It Is Better to Track TogetherπŸ“„
3SpatialTracker: Tracking Any 2D Pixels in 3D SpaceπŸ“„
4CoTracker3: Simpler and Better Point Tracking by Pseudo-Labelling Real VideosπŸ“„
5DELTA: Dense Efficient Long-Range 3D Tracking for Any VideoπŸ“„
6TAPIP3D: Tracking Any Point in Persistent 3D GeometryπŸ“„
7SpatialTrackerV2: 3D Point Tracking Made EasyπŸ“„

(⬆️ back to top)


If you find this repository useful, please consider giving it a ⭐

Made with ❀️ for the 3D Vision Community

Contributors

2hiTee

15 commits