git clone https://github.com/wjh892521292/DAR/
cd DAR
conda env create -f requirements.yml
conda activate dar
xformers for faster attention computation.pip install xformers==0.0.21
Prepare NYU Depth V2 and KITTI datasets from BTS. After downloading the datasets, change the data paths in the respective bash files to point to your dataset location where you have downloaded the datasets.
--data_path <path_to_nyu_dataset>\
Note the dataset structure inside the path you have given in the bash files should look like this:
NYU Depth V2
nyu
├── nyu_depth_v2
│ ├── official_splits
│ └── sync
KITTI:
kitti
├── KITTI
│ ├── 2011_09_26
│ ├── 2011_09_28
│ ├── 2011_09_29
│ ├── 2011_09_30
│ └── 2011_10_03
└── kitti_gt
├── 2011_09_26
├── 2011_09_28
├── 2011_09_29
├── 2011_09_30
└── 2011_10_03
Download the v1-5 checkpoint of stable-diffusion and put it in the <repo root>/checkpoints directory. Please create an empty directory if you find that such a path does not exist.
<repo root>/tools. Notebly, model weights are recommanded to be at <repo root>/checkpoints directory.We recommend using 8 NVIDIA A100 GPUs to train the model with a total batch size of 32. Inside the train_{kitti,nyu}.sh set the NPROC_PER_NODE variable and --batch_size argument to the desired values as per your system resources. For our method we set them as NPROC_PER_NODE=8 and --batch_size=4 (resulting in a total batch size of 32). Run the following instuction to train the DAR model:
Train on NYUv2 dataset:
bash train_nyu.sh
Train on KITTI dataset:
bash train_kitti.sh
If you have any questions, please feel free to contact wangjinhong@zju.edu.cn.
Thanks to the codebase of BTS, NeWCRFS, IEbins, EcoDepth, VAR.
If you find our work useful in your research, please consider citing using:
@InProceedings{wang2025scalable,
title={Scalable autoregressive monocular depth estimation},
author={Wang, Jinhong and Liu, Jian and Tang, Dongqi and Wang, Weiqiang and Li, Wentong and Chen, Danny and Chen, Jintai and Wu, Jian},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
pages={6262--6272},
year={2025}
}
Jupyter Notebook
78.1%
Python
21.7%
git clone https://github.com/wjh892521292/DAR/
cd DAR
conda env create -f requirements.yml
conda activate dar
xformers for faster attention computation.pip install xformers==0.0.21
Prepare NYU Depth V2 and KITTI datasets from BTS. After downloading the datasets, change the data paths in the respective bash files to point to your dataset location where you have downloaded the datasets.
--data_path <path_to_nyu_dataset>\
Note the dataset structure inside the path you have given in the bash files should look like this:
NYU Depth V2
nyu
├── nyu_depth_v2
│ ├── official_splits
│ └── sync
KITTI:
kitti
├── KITTI
│ ├── 2011_09_26
│ ├── 2011_09_28
│ ├── 2011_09_29
│ ├── 2011_09_30
│ └── 2011_10_03
└── kitti_gt
├── 2011_09_26
├── 2011_09_28
├── 2011_09_29
├── 2011_09_30
└── 2011_10_03
Download the v1-5 checkpoint of stable-diffusion and put it in the <repo root>/checkpoints directory. Please create an empty directory if you find that such a path does not exist.
<repo root>/tools. Notebly, model weights are recommanded to be at <repo root>/checkpoints directory.We recommend using 8 NVIDIA A100 GPUs to train the model with a total batch size of 32. Inside the train_{kitti,nyu}.sh set the NPROC_PER_NODE variable and --batch_size argument to the desired values as per your system resources. For our method we set them as NPROC_PER_NODE=8 and --batch_size=4 (resulting in a total batch size of 32). Run the following instuction to train the DAR model:
Train on NYUv2 dataset:
bash train_nyu.sh
Train on KITTI dataset:
bash train_kitti.sh
If you have any questions, please feel free to contact wangjinhong@zju.edu.cn.
Thanks to the codebase of BTS, NeWCRFS, IEbins, EcoDepth, VAR.
If you find our work useful in your research, please consider citing using:
@InProceedings{wang2025scalable,
title={Scalable autoregressive monocular depth estimation},
author={Wang, Jinhong and Liu, Jian and Tang, Dongqi and Wang, Weiqiang and Li, Wentong and Chen, Danny and Chen, Jintai and Wu, Jian},
booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
pages={6262--6272},
year={2025}
}
Jupyter Notebook
78.1%
Python
21.7%