AIVFI/Video-Depth-Estimation-Rankings-and-Stereo-Video-Conversion-Rankings

Researchers, we look forward to models based on MiniMax H3. ChronoDepth Depth Any Video Depth Anything Depth Pro DepthCrafter DreamStereo DVD Elastic3D Eye2Eye FFN GemDepth GRT HairGuard InfiniDepth MoGe M2SVid NVDS Pixel-Perfect Video Depth Restereo StereoCrafter StereoPilot StereoWorld SVG UniDepth UniK3D Video Depth Anything ViGeo αDepth

262

257 commits

updated Sep 17, 2026

See the code

README

Video Depth Estimation Rankings
and Stereo Video Conversion Rankings

Table of Contents

Introduction

Stereo Video Conversion Rankings

  • New models based on MiniMax H3 will deliver a new level of quality and the new rankings will be based on them

Video Depth Estimation Rankings

Single Image Depth Estimation Rankings

Appendices


Purpose of this repository

"The 3D shows you a window into reality; the higher frame rate takes the glass out of the window"
— James Cameron

Researchers, if you have found your way here, please consider developing a new Stereo Video Conversion model based on MiniMax H3.

Back to Table of Contents

Awesome Stereo Video Conversion

The following list includes all Stereo Video Conversion methods from the last 17 months, from 1 April 2025 to 1 September 2026. This list was created because there is a significant problem with public access to the latest Stereo Video Conversion models, which makes it difficult for researchers to compare their work with the current state of the art and to use the same test set. Consequently, this also makes it difficult to present a single ranking that showcases all of the best models.

MethodBackboneSubmitted on
(arXiv)
     Venue     Official
  repository  
Website
code
GRTProPainter6 Jul 2026ICML-GitHub Stars
αDepth?29 May 2026arXiv--
StereoCrafter2Wan2.1-VACE-14B--GitHub Stars-
DreamStereoWan2.1-1.3B14 Apr 2026CVPR-GitHub Stars
HairGuardWan2.1-VACE-1.3B6 Jan 2026arXiv--
StereoPilotWan2.1-T2V-1.3B18 Dec 2025arXivGitHub StarsGitHub Stars
Elastic3DSVD16 Dec 2025CVPR-GitHub Stars
StereoWorldWan2.1-T2V-1.3B10 Dec 2025CVPRGitHub StarsGitHub Stars
RestereoStereoCrafter (based on SVD)6 Jun 2025CVPRW--
M2SVidSVD22 May 20253DVGitHub StarsGitHub Stars
Eye2EyeLumiere30 Apr 2025arXiv-GitHub Stars

Back to Table of Contents

Awesome Synthetic RGB-D Image Datasets for Training HD Video Depth Estimation Models

Although video depth estimation models should be trained mainly on synthetic RGB-D video datasets I decided to add two synthetic RGB-D image datasets because of their unique features.

Dataset     Venue     ResolutionUnique features
1SynthHuman
📌 Human faces 😍
ICCV384×512The dataset contains 98040 samples feature the face, 99976 sample feature the full body and 99992 samples feature the upper body. DAViD trained on this dataset alone achieved better depth estimation results than Depth Anything V2 Large, Depth Pro and even Sapiens-2B on the Goliath-Face test set. See the results in Table 1.
2MegaSynthCVPR512×512Huge size: 700K scenes and the incredible improvement in depth estimation results of the fine-tuned Depth Anything V2 ViT-B model on MegaSynth and evaluated on Hypersim. See the results in Table 6.

Back to Table of Contents

Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

The following list contains only synthetic RGB-D datasets in which at least some of the images can be composited into a video sequence of at least 32 frames. The minimum number of frames was chosen on the basis of the ablation studies shown in Table 5 by the Video Depth Anything researchers.

Most datasets contain ready-to-use video sequences of appropriately numbered images in individual folders, but in the case of the PLT-D3 dataset, images from at least two folders have to be combined to make a longer video sequence and in the case of the ClaraVid dataset, images have to be arranged in the correct order to make a 32-frame video sequence, for example in the order given in Appendix 4.

Researchers, if you are going to use the following list to select datasets to train your models check their quality very carefully and choose the best ones. I have only visually checked a few of them and have marked on the list 2 datasets to check particularly carefully and 2 datasets that in my opinion are not suitable for training video depth estimation models. I have given the reasons for such markings in the same Appendix 4.

In selecting the best datasets, comparisons of their quality can be very helpful, such as in Table 9, Table 6, another Table 6 for depth estimation models and TABLE V plus TABLE IV for stereo matching models, although a similar technique can also be used for depth estimation models.

Dataset       Venue       ResolutionV
G
D
A
3
T
M
o
3
M
o
2
D
P
U
D
2
G
D
V
D
A
D
V
D
D
C
R
D
1SynthVerse
📌 Human face 4K 😍

in the 164-frame battery-grabbing scene from "Charge"
SIGGRAPH4096×1716-----------
2PhysInOne
📌 Small particles 😍

in the PhysInOneP16
CVPR1120×1120-----------
3SDG-SynHuman
📌 Human poses 😍

631 million RGB frames!
arXiv1920×1080-----------
4BEDLAM2.0
📌 Human poses 😍
NeurIPS1280×720-----------
5C3I-SynFace
📌 Human faces 😍
DIB640×480-----------
6ClaraVidICCV4032x3024-----------
7StereoGenBencharXiv1280×1280-----------
8SpringCVPR1920×1080TTEEE------
9SDG-PhyxSimarXiv1920×1080-----------
10HorizonGSCVPR1920×1080-----------
11PLT-D3HD1920×1080-----------
12MVS-SynthCVPR1920×1080TTTTT-T----
13SYNTHIA-SFBMVC1920×1080-----------
14SynDrone
Check before use!
ICCVW1920×1080-----------
15Mid-AirCVPRW1024×1024--TT-------
16MatrixCityICCV1000×1000TTTT-T---T-
17StereoCarlaarXiv1600×900-----------
18LightwheelOcc-1600×900T----------
19SAIL-VOS 3DCVPR1280×800----T------
20SHIFTCVPR1280×800-----------
21SYNTHIA-Seqs
🚫 Do not use! 🚫
CVPR1280×760T-TT-------
22BEDLAMCVPR1280×720T---TT-----
23Dynamic ReplicaCVPR1280×720T---TTT--T-
24OmniWorld-GameICLR1280×720T-T--------
25WorldRoverarXiv1280×720-----------
26Syn4DECCV1280×720-----------
27InFlux++ SynthECCV1280×720-----------
28Infinigen SVTPAMI1280×720-----------
29InfinigenCVPR1280×720-----------
30DigiDogs
🚫 Do not use! 🚫
WACVW1280×720-----------
31Aria Synthetic Environments
Check before use!
-704×704T----------
32TartanGroundIROS640×640T----------
33TartanAir V2-640×640-----------
34BlinkVisionECCV960×540-----------
35PointOdysseyICCV960×540TT---TTT--E
36DyDToFCVPR960×540----------E
37IRSICME960×540-TTTT-TT---
38Scene FlowCVPR960×540-----------
39THUD++arXiv730×530-----------
40TAU AgentTCI1024×512-T---------
41TransPhy3DICRA512×512T----------
423D Ken BurnsTOG512×512-TTTT------
43SynPhoRest-848×480-----------
44TartanAirIROS640×480TTTTTTTTT-T
45ParallelDomain-4DECCV640×480-----------
46EDENWACV640×480-TTTTT-----
47GTA-SfMRAL640×480TTTT-------
48InteriorNetBMVC640×480-----------
49SYNTHIA-ALICCVW640×480-----------
50MPI SintelECCV1024×436EEEEEEEEEE-
51CarlaOccCVPR1408×376T----------
52Virtual KITTI 2arXiv1242×375-T--T-TT---
53Virtual KITTICVPR1242×375--------T--
54TartanAir ShibuyaICRA640×360-----------
Total: T (training)15111099664221
Total: E (testing)11222111112

Back to Table of Contents

170-frame ScanNet: TAE

📝 Note: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.

RKModel
Links:
         Venue   Repository    
  TAE ↓  
{Input fr.}
ECCV
Table 3
FFN
  TAE ↓  
{Input fr.}
ICML
Table 2&
GitHub Stars
GD
  TAE ↓  
{Input fr.}
CVPR
Table 1
VDA
1FFN
ECCV GitHub Stars
0.380 {MF}--
2GemDepth-VDA
ICML GitHub Stars
-0.47 {MF}-
3VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.570 {MF}
0.57 {MF}
VDA-L-Syn:
-
0.570 {MF}
VDA-L-Syn:
0.570 {MF}
4DVD v1.1
ICMLW GitHub Stars
-0.61 {MF}-
5DepthCrafter
CVPR GitHub Stars
0.639 {MF}-0.639 {MF}
6RollingDepth
CVPR GitHub Stars
-0.65 {MF}-
7Depth Any Video
ICLR GitHub Stars
0.967 {MF}-0.967 {MF}
8ChronoDepth
CVPR GitHub Stars
1.022 {MF}-1.022 {MF}
9Depth Anything V2 Large
NeurIPS GitHub Stars
1.140 {1}1.14 {1}1.140 {1}
10NVDS
ICCV GitHub Stars
2.176 {4}-2.176 {4}

Back to Table of Contents

500-frame Bonn RGB-D Dynamic: δ1

📝 Note 1: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.
📝 Note 2: A high rank of the Pixel-Perfect Video Depth model in this ranking does not guarantee that this model is suitable for practical applications, see https://github.com/gangweix/pixel-perfect-depth/issues/27. We are waiting for this model to be fixed and evaluated using temporal stability metrics.

RKModel
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
ECCV
Table 3
FFN
     δ1 ↑     
{Input fr.}
arXiv
TABLE II
PPVD
     δ1 ↑     
{Input fr.}
ICML
Table 1&
GitHub Stars
GD
     δ1 ↑     
{Input fr.}
CVPR
Table 1
VDA
1FFN
ECCV GitHub Stars
0.982 {MF}---
2Pixel-Perfect Video Depth
arXiv GitHub Stars
-0.979 {MF}--
3GemDepth-VDA
ICML GitHub Stars
--0.978 {MF}-
4VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.961 {MF}
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
0.961 {MF}
5DVD v1.1
ICMLW GitHub Stars
--0.948 {MF}-
6RollingDepth
CVPR GitHub Stars
-0.931 {MF}0.931 {MF}-
7Depth Anything V2 Large
NeurIPS GitHub Stars
0.864 {1}0.864 {1}0.864 {1}0.864 {1}
8DepthCrafter
CVPR GitHub Stars
0.803 {MF}0.803 {MF}0.803 {MF}0.803 {MF}
9NVDS
ICCV GitHub Stars
0.674 {4}0.674 {4}0.674 {4}0.674 {4}
10ChronoDepth
CVPR GitHub Stars
0.665 {MF}0.665 {MF}0.665 {MF}0.665 {MF}

Back to Table of Contents

Synth4K: δ1

RKModel
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
arXiv
Table C.1
MoGe-3
1MoGe-3 ViT-G Step 3
arXiv GitHub Stars
0.949 {1}
2MoGe-2
NeurIPS GitHub Stars
0.913 {1}
3Depth Anything 3
ICLR GitHub Stars
0.883 {1}
4InfiniDepth
CVPR GitHub Stars
0.878 {1}
5UniDepthV2
arXiv GitHub Stars
0.869 {1}
6Pixel-Perfect Depth
NeurIPS GitHub Stars
0.868 {1}
7Depth Pro
ICLR GitHub Stars
0.866 {1}
8UniK3D
CVPR GitHub Stars
0.862 {1}

Back to Table of Contents

Appendix 4: Notes on the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

📝 Note 1: Example of arranging images in the correct order to make a 32-frame video sequence for the ClaraVid dataset:

<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00360.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00320.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00280.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00240.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00200.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00160.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00120.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00080.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00040.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00000.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00001.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00002.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00003.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00004.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00005.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00006.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00007.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00008.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00009.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00010.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00011.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00012.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00013.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00014.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00015.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00016.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00017.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00018.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00019.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00059.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00099.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00139.jpg

📝 Note 2: Do not use the SYNTHIA-Seqs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. This is particularly evident in the example of tree leaves:
<your-data-path>/SYNTHIA-SEQS-01-SPRING/Depth/Stereo_Left/Omni_F/000071.png
<your-data-path>/SYNTHIA-SEQS-01-SPRING/RGB/Stereo_Left/Omni_F/000071.png.
📝 Note 3: Do not use the DigiDogs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. See the objects behind the campfire, the shifting position of the vegetation on the left and the clear banding on the depth map:
<your-data-path>/DigiDogs2024_full/09_22_2022/00054/images/img_00012.tiff.
📝 Note 4: Check before use the SynDrone dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/Town01_Opt_120_depth/Town01_Opt_120/ClearNoon/height20m/depth/00031.png.
📝 Note 5: Check before use the Aria Synthetic Environments dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/75/depth/depth0000109.png.

Back to Table of Contents

Appendix 5: List of all research papers from the above rankings

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
ChronoDepth-Learning Temporally Consistent Video Depth from Video Diffusion PriorsCVPRGitHub Stars
Depth Any VideoDAVDepth Any Video with Scalable Synthetic DataICLRGitHub Stars
Depth Anything 3DA3Depth Anything 3: Recovering the Visual Space from Any ViewsICLRGitHub Stars
Depth Anything V2DA V2Depth Anything V2NeurIPSGitHub Stars
Depth ProDPDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLRGitHub Stars
DepthCrafterDCDepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosCVPRGitHub Stars
DVD-DVD: Deterministic Video Depth Estimation with Generative PriorsICMLWGitHub Stars
FFN-Forget, Anticipate and Adapt: Test Time Training for Long VideosECCVGitHub Stars
GemDepthGDGemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthICMLGitHub Stars
InfiniDepth-InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsCVPRGitHub Stars
MoGe-2Mo2MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPSGitHub Stars
MoGe-3Mo3MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivGitHub Stars
NVDS-Neural Video Depth StabilizerICCVGitHub Stars
Pixel-Perfect DepthPPDPixel-Perfect Depth with Semantics-Prompted Diffusion TransformersNeurIPSGitHub Stars
Pixel-Perfect Video DepthPPVDPixel-Perfect Visual Geometry EstimationarXivGitHub Stars
RollingDepthRDVideo Depth without Video ModelsCVPRGitHub Stars
UniDepthV2UD2UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerarXivGitHub Stars
UniK3D-UniK3D: Universal Camera Monocular 3D EstimationCVPRGitHub Stars
Video Depth AnythingVDAVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPRGitHub Stars

Back to Table of Contents

Appendix 6: List of all research papers from the column headers of the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
Depth Anything 3DA3Depth Anything 3: Recovering the Visual Space from Any ViewsICLRGitHub Stars
Depth ProDPDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLRGitHub Stars
DepthCrafterDCDepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosCVPRGitHub Stars
DVD-DVD: Deterministic Video Depth Estimation with Generative PriorsICMLWGitHub Stars
GemDepthGDGemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthICMLGitHub Stars
MoGe-2Mo2MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPSGitHub Stars
MoGe-3Mo3MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivGitHub Stars
RollingDepthRDVideo Depth without Video ModelsCVPRGitHub Stars
UniDepthV2UD2UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerarXivGitHub Stars
Video Depth AnythingVDAVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPRGitHub Stars
ViGeoVGTowards Consistent Video Geometry EstimationarXivGitHub Stars

Back to Table of Contents

List of research papers to be added to the rankings

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
GRT-Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video GenerationICML-
αDepth-αDepth: Learning Single-Pass Soft Boundary Decomposition for Stereo ConversionarXiv-
DreamStereo-DreamStereo: Towards Real-Time Stereo Inpainting for HD VideosCVPR-
HairGuard-Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel ViewsarXiv-
StereoPilot-StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative PriorsarXivGitHub Stars
Elastic3D-Elastic3D: Controllable Stereo Video Conversion with Guided Latent DecodingCVPR-
StereoWorld-StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationCVPRGitHub Stars
Restereo-Restereo: Unifying diffusion stereo video generation and restorationCVPRW-
M2SVid-M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion3DVGitHub Stars
Eye2Eye-Eye2Eye: A Simple Approach for Monocular-to-Stereo Video SynthesisarXiv-
StereoCrafter-StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular VideosarXivGitHub Stars
SVG-SVG: 3D Stereoscopic Video Generation via Denoising Frame MatrixICLRGitHub Stars

Back to Table of Contents

3d
awesome
computer-vision
deep-learning
depth
depth-anything
depth-estimation
depth-map
depth-maps
depth-prediction
depth-pro
machine-learning
monocular-depth
monocular-depth-estimation
stereo
stereo-video
stereo-vision
video-depth
video-processing

AIVFI/Video-Depth-Estimation-Rankings-and-Stereo-Video-Conversion-Rankings

Researchers, we look forward to models based on MiniMax H3. ChronoDepth Depth Any Video Depth Anything Depth Pro DepthCrafter DreamStereo DVD Elastic3D Eye2Eye FFN GemDepth GRT HairGuard InfiniDepth MoGe M2SVid NVDS Pixel-Perfect Video Depth Restereo StereoCrafter StereoPilot StereoWorld SVG UniDepth UniK3D Video Depth Anything ViGeo αDepth

262

257 commits

updated Sep 17, 2026

See the code

README

Video Depth Estimation Rankings
and Stereo Video Conversion Rankings

Table of Contents

Introduction

Stereo Video Conversion Rankings

  • New models based on MiniMax H3 will deliver a new level of quality and the new rankings will be based on them

Video Depth Estimation Rankings

Single Image Depth Estimation Rankings

Appendices


Purpose of this repository

"The 3D shows you a window into reality; the higher frame rate takes the glass out of the window"
— James Cameron

Researchers, if you have found your way here, please consider developing a new Stereo Video Conversion model based on MiniMax H3.

Back to Table of Contents

Awesome Stereo Video Conversion

The following list includes all Stereo Video Conversion methods from the last 17 months, from 1 April 2025 to 1 September 2026. This list was created because there is a significant problem with public access to the latest Stereo Video Conversion models, which makes it difficult for researchers to compare their work with the current state of the art and to use the same test set. Consequently, this also makes it difficult to present a single ranking that showcases all of the best models.

MethodBackboneSubmitted on
(arXiv)
     Venue     Official
  repository  
Website
code
GRTProPainter6 Jul 2026ICML-GitHub Stars
αDepth?29 May 2026arXiv--
StereoCrafter2Wan2.1-VACE-14B--GitHub Stars-
DreamStereoWan2.1-1.3B14 Apr 2026CVPR-GitHub Stars
HairGuardWan2.1-VACE-1.3B6 Jan 2026arXiv--
StereoPilotWan2.1-T2V-1.3B18 Dec 2025arXivGitHub StarsGitHub Stars
Elastic3DSVD16 Dec 2025CVPR-GitHub Stars
StereoWorldWan2.1-T2V-1.3B10 Dec 2025CVPRGitHub StarsGitHub Stars
RestereoStereoCrafter (based on SVD)6 Jun 2025CVPRW--
M2SVidSVD22 May 20253DVGitHub StarsGitHub Stars
Eye2EyeLumiere30 Apr 2025arXiv-GitHub Stars

Back to Table of Contents

Awesome Synthetic RGB-D Image Datasets for Training HD Video Depth Estimation Models

Although video depth estimation models should be trained mainly on synthetic RGB-D video datasets I decided to add two synthetic RGB-D image datasets because of their unique features.

Dataset     Venue     ResolutionUnique features
1SynthHuman
📌 Human faces 😍
ICCV384×512The dataset contains 98040 samples feature the face, 99976 sample feature the full body and 99992 samples feature the upper body. DAViD trained on this dataset alone achieved better depth estimation results than Depth Anything V2 Large, Depth Pro and even Sapiens-2B on the Goliath-Face test set. See the results in Table 1.
2MegaSynthCVPR512×512Huge size: 700K scenes and the incredible improvement in depth estimation results of the fine-tuned Depth Anything V2 ViT-B model on MegaSynth and evaluated on Hypersim. See the results in Table 6.

Back to Table of Contents

Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

The following list contains only synthetic RGB-D datasets in which at least some of the images can be composited into a video sequence of at least 32 frames. The minimum number of frames was chosen on the basis of the ablation studies shown in Table 5 by the Video Depth Anything researchers.

Most datasets contain ready-to-use video sequences of appropriately numbered images in individual folders, but in the case of the PLT-D3 dataset, images from at least two folders have to be combined to make a longer video sequence and in the case of the ClaraVid dataset, images have to be arranged in the correct order to make a 32-frame video sequence, for example in the order given in Appendix 4.

Researchers, if you are going to use the following list to select datasets to train your models check their quality very carefully and choose the best ones. I have only visually checked a few of them and have marked on the list 2 datasets to check particularly carefully and 2 datasets that in my opinion are not suitable for training video depth estimation models. I have given the reasons for such markings in the same Appendix 4.

In selecting the best datasets, comparisons of their quality can be very helpful, such as in Table 9, Table 6, another Table 6 for depth estimation models and TABLE V plus TABLE IV for stereo matching models, although a similar technique can also be used for depth estimation models.

Dataset       Venue       ResolutionV
G
D
A
3
T
M
o
3
M
o
2
D
P
U
D
2
G
D
V
D
A
D
V
D
D
C
R
D
1SynthVerse
📌 Human face 4K 😍

in the 164-frame battery-grabbing scene from "Charge"
SIGGRAPH4096×1716-----------
2PhysInOne
📌 Small particles 😍

in the PhysInOneP16
CVPR1120×1120-----------
3SDG-SynHuman
📌 Human poses 😍

631 million RGB frames!
arXiv1920×1080-----------
4BEDLAM2.0
📌 Human poses 😍
NeurIPS1280×720-----------
5C3I-SynFace
📌 Human faces 😍
DIB640×480-----------
6ClaraVidICCV4032x3024-----------
7StereoGenBencharXiv1280×1280-----------
8SpringCVPR1920×1080TTEEE------
9SDG-PhyxSimarXiv1920×1080-----------
10HorizonGSCVPR1920×1080-----------
11PLT-D3HD1920×1080-----------
12MVS-SynthCVPR1920×1080TTTTT-T----
13SYNTHIA-SFBMVC1920×1080-----------
14SynDrone
Check before use!
ICCVW1920×1080-----------
15Mid-AirCVPRW1024×1024--TT-------
16MatrixCityICCV1000×1000TTTT-T---T-
17StereoCarlaarXiv1600×900-----------
18LightwheelOcc-1600×900T----------
19SAIL-VOS 3DCVPR1280×800----T------
20SHIFTCVPR1280×800-----------
21SYNTHIA-Seqs
🚫 Do not use! 🚫
CVPR1280×760T-TT-------
22BEDLAMCVPR1280×720T---TT-----
23Dynamic ReplicaCVPR1280×720T---TTT--T-
24OmniWorld-GameICLR1280×720T-T--------
25WorldRoverarXiv1280×720-----------
26Syn4DECCV1280×720-----------
27InFlux++ SynthECCV1280×720-----------
28Infinigen SVTPAMI1280×720-----------
29InfinigenCVPR1280×720-----------
30DigiDogs
🚫 Do not use! 🚫
WACVW1280×720-----------
31Aria Synthetic Environments
Check before use!
-704×704T----------
32TartanGroundIROS640×640T----------
33TartanAir V2-640×640-----------
34BlinkVisionECCV960×540-----------
35PointOdysseyICCV960×540TT---TTT--E
36DyDToFCVPR960×540----------E
37IRSICME960×540-TTTT-TT---
38Scene FlowCVPR960×540-----------
39THUD++arXiv730×530-----------
40TAU AgentTCI1024×512-T---------
41TransPhy3DICRA512×512T----------
423D Ken BurnsTOG512×512-TTTT------
43SynPhoRest-848×480-----------
44TartanAirIROS640×480TTTTTTTTT-T
45ParallelDomain-4DECCV640×480-----------
46EDENWACV640×480-TTTTT-----
47GTA-SfMRAL640×480TTTT-------
48InteriorNetBMVC640×480-----------
49SYNTHIA-ALICCVW640×480-----------
50MPI SintelECCV1024×436EEEEEEEEEE-
51CarlaOccCVPR1408×376T----------
52Virtual KITTI 2arXiv1242×375-T--T-TT---
53Virtual KITTICVPR1242×375--------T--
54TartanAir ShibuyaICRA640×360-----------
Total: T (training)15111099664221
Total: E (testing)11222111112

Back to Table of Contents

170-frame ScanNet: TAE

📝 Note: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.

RKModel
Links:
         Venue   Repository    
  TAE ↓  
{Input fr.}
ECCV
Table 3
FFN
  TAE ↓  
{Input fr.}
ICML
Table 2&
GitHub Stars
GD
  TAE ↓  
{Input fr.}
CVPR
Table 1
VDA
1FFN
ECCV GitHub Stars
0.380 {MF}--
2GemDepth-VDA
ICML GitHub Stars
-0.47 {MF}-
3VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.570 {MF}
0.57 {MF}
VDA-L-Syn:
-
0.570 {MF}
VDA-L-Syn:
0.570 {MF}
4DVD v1.1
ICMLW GitHub Stars
-0.61 {MF}-
5DepthCrafter
CVPR GitHub Stars
0.639 {MF}-0.639 {MF}
6RollingDepth
CVPR GitHub Stars
-0.65 {MF}-
7Depth Any Video
ICLR GitHub Stars
0.967 {MF}-0.967 {MF}
8ChronoDepth
CVPR GitHub Stars
1.022 {MF}-1.022 {MF}
9Depth Anything V2 Large
NeurIPS GitHub Stars
1.140 {1}1.14 {1}1.140 {1}
10NVDS
ICCV GitHub Stars
2.176 {4}-2.176 {4}

Back to Table of Contents

500-frame Bonn RGB-D Dynamic: δ1

📝 Note 1: This ranking is based on the evaluation protocol proposed by Video Depth Anything developers.
📝 Note 2: A high rank of the Pixel-Perfect Video Depth model in this ranking does not guarantee that this model is suitable for practical applications, see https://github.com/gangweix/pixel-perfect-depth/issues/27. We are waiting for this model to be fixed and evaluated using temporal stability metrics.

RKModel
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
ECCV
Table 3
FFN
     δ1 ↑     
{Input fr.}
arXiv
TABLE II
PPVD
     δ1 ↑     
{Input fr.}
ICML
Table 1&
GitHub Stars
GD
     δ1 ↑     
{Input fr.}
CVPR
Table 1
VDA
1FFN
ECCV GitHub Stars
0.982 {MF}---
2Pixel-Perfect Video Depth
arXiv GitHub Stars
-0.979 {MF}--
3GemDepth-VDA
ICML GitHub Stars
--0.978 {MF}-
4VDA-L
CVPR GitHub Stars
-
VDA-L-Syn:
0.961 {MF}
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
-
0.959 {MF}
VDA-L-Syn:
0.961 {MF}
5DVD v1.1
ICMLW GitHub Stars
--0.948 {MF}-
6RollingDepth
CVPR GitHub Stars
-0.931 {MF}0.931 {MF}-
7Depth Anything V2 Large
NeurIPS GitHub Stars
0.864 {1}0.864 {1}0.864 {1}0.864 {1}
8DepthCrafter
CVPR GitHub Stars
0.803 {MF}0.803 {MF}0.803 {MF}0.803 {MF}
9NVDS
ICCV GitHub Stars
0.674 {4}0.674 {4}0.674 {4}0.674 {4}
10ChronoDepth
CVPR GitHub Stars
0.665 {MF}0.665 {MF}0.665 {MF}0.665 {MF}

Back to Table of Contents

Synth4K: δ1

RKModel
Links:
         Venue   Repository    
     δ1 ↑     
{Input fr.}
arXiv
Table C.1
MoGe-3
1MoGe-3 ViT-G Step 3
arXiv GitHub Stars
0.949 {1}
2MoGe-2
NeurIPS GitHub Stars
0.913 {1}
3Depth Anything 3
ICLR GitHub Stars
0.883 {1}
4InfiniDepth
CVPR GitHub Stars
0.878 {1}
5UniDepthV2
arXiv GitHub Stars
0.869 {1}
6Pixel-Perfect Depth
NeurIPS GitHub Stars
0.868 {1}
7Depth Pro
ICLR GitHub Stars
0.866 {1}
8UniK3D
CVPR GitHub Stars
0.862 {1}

Back to Table of Contents

Appendix 4: Notes on the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

📝 Note 1: Example of arranging images in the correct order to make a 32-frame video sequence for the ClaraVid dataset:

<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00360.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00320.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00280.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00240.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00200.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00160.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00120.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00080.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00040.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00000.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00001.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00002.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00003.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00004.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00005.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00006.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00007.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00008.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00009.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00010.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00011.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00012.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00013.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00014.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00015.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00016.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00017.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00018.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00019.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00059.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00099.jpg
<your-data-path>/008_urban_dense_1/left_rgb/45deg_low_h/00139.jpg

📝 Note 2: Do not use the SYNTHIA-Seqs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. This is particularly evident in the example of tree leaves:
<your-data-path>/SYNTHIA-SEQS-01-SPRING/Depth/Stereo_Left/Omni_F/000071.png
<your-data-path>/SYNTHIA-SEQS-01-SPRING/RGB/Stereo_Left/Omni_F/000071.png.
📝 Note 3: Do not use the DigiDogs dataset for training HD video depth estimation models! The depth maps in this dataset do not match the corresponding RGB images. See the objects behind the campfire, the shifting position of the vegetation on the left and the clear banding on the depth map:
<your-data-path>/DigiDogs2024_full/09_22_2022/00054/images/img_00012.tiff.
📝 Note 4: Check before use the SynDrone dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/Town01_Opt_120_depth/Town01_Opt_120/ClearNoon/height20m/depth/00031.png.
📝 Note 5: Check before use the Aria Synthetic Environments dataset for training HD video depth estimation models! The depth maps in this dataset have large white areas of unknown depth, which should not happen with a synthetic dataset. Example depth map:
<your-data-path>/75/depth/depth0000109.png.

Back to Table of Contents

Appendix 5: List of all research papers from the above rankings

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
ChronoDepth-Learning Temporally Consistent Video Depth from Video Diffusion PriorsCVPRGitHub Stars
Depth Any VideoDAVDepth Any Video with Scalable Synthetic DataICLRGitHub Stars
Depth Anything 3DA3Depth Anything 3: Recovering the Visual Space from Any ViewsICLRGitHub Stars
Depth Anything V2DA V2Depth Anything V2NeurIPSGitHub Stars
Depth ProDPDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLRGitHub Stars
DepthCrafterDCDepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosCVPRGitHub Stars
DVD-DVD: Deterministic Video Depth Estimation with Generative PriorsICMLWGitHub Stars
FFN-Forget, Anticipate and Adapt: Test Time Training for Long VideosECCVGitHub Stars
GemDepthGDGemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthICMLGitHub Stars
InfiniDepth-InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit FieldsCVPRGitHub Stars
MoGe-2Mo2MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPSGitHub Stars
MoGe-3Mo3MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivGitHub Stars
NVDS-Neural Video Depth StabilizerICCVGitHub Stars
Pixel-Perfect DepthPPDPixel-Perfect Depth with Semantics-Prompted Diffusion TransformersNeurIPSGitHub Stars
Pixel-Perfect Video DepthPPVDPixel-Perfect Visual Geometry EstimationarXivGitHub Stars
RollingDepthRDVideo Depth without Video ModelsCVPRGitHub Stars
UniDepthV2UD2UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerarXivGitHub Stars
UniK3D-UniK3D: Universal Camera Monocular 3D EstimationCVPRGitHub Stars
Video Depth AnythingVDAVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPRGitHub Stars

Back to Table of Contents

Appendix 6: List of all research papers from the column headers of the table: Awesome Synthetic RGB-D Video Datasets for Training and Testing HD Video Depth Estimation Models

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
Depth Anything 3DA3Depth Anything 3: Recovering the Visual Space from Any ViewsICLRGitHub Stars
Depth ProDPDepth Pro: Sharp Monocular Metric Depth in Less Than a SecondICLRGitHub Stars
DepthCrafterDCDepthCrafter: Generating Consistent Long Depth Sequences for Open-world VideosCVPRGitHub Stars
DVD-DVD: Deterministic Video Depth Estimation with Generative PriorsICMLWGitHub Stars
GemDepthGDGemDepth: Geometry-Embedded Features for 3D-Consistent Video DepthICMLGitHub Stars
MoGe-2Mo2MoGe-2: Accurate Monocular Geometry with Metric Scale and Sharp DetailsNeurIPSGitHub Stars
MoGe-3Mo3MoGe-3: Fine-Detail Monocular Geometry Estimation with Self-Guided Sparse Volumetric RefinementarXivGitHub Stars
RollingDepthRDVideo Depth without Video ModelsCVPRGitHub Stars
UniDepthV2UD2UniDepthV2: Universal Monocular Metric Depth Estimation Made SimplerarXivGitHub Stars
Video Depth AnythingVDAVideo Depth Anything: Consistent Depth Estimation for Super-Long VideosCVPRGitHub Stars
ViGeoVGTowards Consistent Video Geometry EstimationarXivGitHub Stars

Back to Table of Contents

List of research papers to be added to the rankings

MethodAbbr.Paper     Venue     
(Alt link)
Official
  repository  
GRT-Geometric Reciprocity: Unlocking Self-Supervision for Stereoscopic Video GenerationICML-
αDepth-αDepth: Learning Single-Pass Soft Boundary Decomposition for Stereo ConversionarXiv-
DreamStereo-DreamStereo: Towards Real-Time Stereo Inpainting for HD VideosCVPR-
HairGuard-Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel ViewsarXiv-
StereoPilot-StereoPilot: Learning Unified and Efficient Stereo Conversion via Generative PriorsarXivGitHub Stars
Elastic3D-Elastic3D: Controllable Stereo Video Conversion with Guided Latent DecodingCVPR-
StereoWorld-StereoWorld: Geometry-Aware Monocular-to-Stereo Video GenerationCVPRGitHub Stars
Restereo-Restereo: Unifying diffusion stereo video generation and restorationCVPRW-
M2SVid-M2SVid: End-to-End Inpainting and Refinement for Monocular-to-Stereo Video Conversion3DVGitHub Stars
Eye2Eye-Eye2Eye: A Simple Approach for Monocular-to-Stereo Video SynthesisarXiv-
StereoCrafter-StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular VideosarXivGitHub Stars
SVG-SVG: 3D Stereoscopic Video Generation via Denoising Frame MatrixICLRGitHub Stars

Back to Table of Contents

3d
awesome
computer-vision
deep-learning
depth
depth-anything
depth-estimation
depth-map
depth-maps
depth-prediction
depth-pro
machine-learning
monocular-depth
monocular-depth-estimation
stereo
stereo-video
stereo-vision
video-depth
video-processing