wjh892521292/DAR

Jupyter Notebook

26

15 commits

updated Oct 9, 2025

See the code

README

Scalable Autoregressive Monocular Depth Estimation

CVPR 2025
Project Page arXiv page IEEE Xplore Paper IEEE Xplore Paper

News

  • [Jun 2025] Code released.
  • [Feb 2025] DAR accepted in CVPR'2025.

Installation

git clone https://github.com/wjh892521292/DAR/
cd DAR
conda env create -f requirements.yml
conda activate dar
  • (Optional) install and compile xformers for faster attention computation.
pip install xformers==0.0.21

Prepare Dataset

Prepare NYU Depth V2 and KITTI datasets from BTS. After downloading the datasets, change the data paths in the respective bash files to point to your dataset location where you have downloaded the datasets.

  • For example, when training on NYU Depth V2 dataset, the training bash file should contain:
    --data_path <path_to_nyu_dataset>\

Note the dataset structure inside the path you have given in the bash files should look like this:
NYU Depth V2

nyu
├── nyu_depth_v2
│   ├── official_splits
│   └── sync

KITTI:

kitti
├── KITTI
│   ├── 2011_09_26
│   ├── 2011_09_28
│   ├── 2011_09_29
│   ├── 2011_09_30
│   └── 2011_10_03
└── kitti_gt
    ├── 2011_09_26
    ├── 2011_09_28
    ├── 2011_09_29
    ├── 2011_09_30
    └── 2011_10_03

Pretrained Models

Download the v1-5 checkpoint of stable-diffusion and put it in the <repo root>/checkpoints directory. Please create an empty directory if you find that such a path does not exist.

  • (Optional) If you need to load model weights of pretrained visual backbones locally. You can download the backbone repository (e.g., dinov2 or ViT of CLIP) and put it in the <repo root>/tools. Notebly, model weights are recommanded to be at <repo root>/checkpoints directory.

Training

We recommend using 8 NVIDIA A100 GPUs to train the model with a total batch size of 32. Inside the train_{kitti,nyu}.sh set the NPROC_PER_NODE variable and --batch_size argument to the desired values as per your system resources. For our method we set them as NPROC_PER_NODE=8 and --batch_size=4 (resulting in a total batch size of 32). Run the following instuction to train the DAR model:

  1. Train on NYUv2 dataset:
    bash train_nyu.sh

  2. Train on KITTI dataset:
    bash train_kitti.sh

Contact

If you have any questions, please feel free to contact wangjinhong@zju.edu.cn.

Acknowledgment

Thanks to the codebase of BTS, NeWCRFS, IEbins, EcoDepth, VAR.

BibTeX (Citation)

If you find our work useful in your research, please consider citing using:

@InProceedings{wang2025scalable,
  title={Scalable autoregressive monocular depth estimation},
  author={Wang, Jinhong and Liu, Jian and Tang, Dongqi and Wang, Weiqiang and Li, Wentong and Chen, Danny and Chen, Jintai and Wu, Jian},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={6262--6272},
  year={2025}
}

wjh892521292/DAR

Jupyter Notebook

26

15 commits

updated Oct 9, 2025

See the code

README

Scalable Autoregressive Monocular Depth Estimation

CVPR 2025
Project Page arXiv page IEEE Xplore Paper IEEE Xplore Paper

News

  • [Jun 2025] Code released.
  • [Feb 2025] DAR accepted in CVPR'2025.

Installation

git clone https://github.com/wjh892521292/DAR/
cd DAR
conda env create -f requirements.yml
conda activate dar
  • (Optional) install and compile xformers for faster attention computation.
pip install xformers==0.0.21

Prepare Dataset

Prepare NYU Depth V2 and KITTI datasets from BTS. After downloading the datasets, change the data paths in the respective bash files to point to your dataset location where you have downloaded the datasets.

  • For example, when training on NYU Depth V2 dataset, the training bash file should contain:
    --data_path <path_to_nyu_dataset>\

Note the dataset structure inside the path you have given in the bash files should look like this:
NYU Depth V2

nyu
├── nyu_depth_v2
│   ├── official_splits
│   └── sync

KITTI:

kitti
├── KITTI
│   ├── 2011_09_26
│   ├── 2011_09_28
│   ├── 2011_09_29
│   ├── 2011_09_30
│   └── 2011_10_03
└── kitti_gt
    ├── 2011_09_26
    ├── 2011_09_28
    ├── 2011_09_29
    ├── 2011_09_30
    └── 2011_10_03

Pretrained Models

Download the v1-5 checkpoint of stable-diffusion and put it in the <repo root>/checkpoints directory. Please create an empty directory if you find that such a path does not exist.

  • (Optional) If you need to load model weights of pretrained visual backbones locally. You can download the backbone repository (e.g., dinov2 or ViT of CLIP) and put it in the <repo root>/tools. Notebly, model weights are recommanded to be at <repo root>/checkpoints directory.

Training

We recommend using 8 NVIDIA A100 GPUs to train the model with a total batch size of 32. Inside the train_{kitti,nyu}.sh set the NPROC_PER_NODE variable and --batch_size argument to the desired values as per your system resources. For our method we set them as NPROC_PER_NODE=8 and --batch_size=4 (resulting in a total batch size of 32). Run the following instuction to train the DAR model:

  1. Train on NYUv2 dataset:
    bash train_nyu.sh

  2. Train on KITTI dataset:
    bash train_kitti.sh

Contact

If you have any questions, please feel free to contact wangjinhong@zju.edu.cn.

Acknowledgment

Thanks to the codebase of BTS, NeWCRFS, IEbins, EcoDepth, VAR.

BibTeX (Citation)

If you find our work useful in your research, please consider citing using:

@InProceedings{wang2025scalable,
  title={Scalable autoregressive monocular depth estimation},
  author={Wang, Jinhong and Liu, Jian and Tang, Dongqi and Wang, Weiqiang and Li, Wentong and Chen, Danny and Chen, Jintai and Wu, Jian},
  booktitle={Proceedings of the Computer Vision and Pattern Recognition Conference},
  pages={6262--6272},
  year={2025}
}

Languages

Jupyter Notebook

78.1%

Python

21.7%