UVA-Computer-Vision-Lab/LabelAny3D

[NeurIPS 2025] LabelAny3D: Label Any Object 3D in the Wild

Python

135

3 commits

updated Aug 23, 2026

See the code

README

LabelAny3D: Label Any Object 3D in the Wild

Jin Yao, Radowan Mahmud Redoy, Sebastian Elbaum, Matthew B. Dwyer, Zezhou Cheng

Website Paper

Samples from COCO3D dataset COCO3D samples

COCO3D Dataset

The evaluation set of COCO3D and pseudo-labeled training set are available at Hugging Face.

3D BBox Human Refinement Interface

We release the source code for the refinement interface at https://github.com/UVA-Computer-Vision-Lab/3d_annotator.

Getting Started

πŸ“¦ Installation Guide - Setup instructions and external dependencies

πŸ“– COCO Pipeline Guide - Run the pipeline on COCO dataset

πŸ”§ OVMono3D Fine-tuning - Code for fine-tuning OVMono3D on LabelAny3D pseudo annotations

Citing

If you find this work useful for your research, please kindly cite:

@inproceedings{yao2025labelany3d,
  title={LabelAny3D: Label Any Object 3D in the Wild},
  author={Jin Yao and Radowan Mahmud Redoy and Sebastian Elbaum and Matthew B. Dwyer and Zezhou Cheng},
  booktitle={Neural Information Processing Systems (NeurIPS)},
  year={2025}
}

@inproceedings{yao2025open,
  title={Open Vocabulary Monocular 3D Object Detection},
  author={Yao, Jin and Gu, Hao and Chen, Xuweiyi and Wang, Jiayun and Cheng, Zezhou},
  booktitle={Proceedings of the International Conference on 3D Vision (3DV)},
  year={2026}
}

Acknowledgements

This work builds on many open-source projects:

  • Gen3DSR - 3D reconstruction framework
  • TRELLIS - 3D asset generation
  • MoGe - Monocular geometry estimation
  • DepthPro - Metric depth estimation
  • MASt3R - Dense matching
  • InvSR - Image super-resolution
  • COCONUT - COCO segmentation annotations
  • OVMono3D - Open vocabulary monocular 3D detection

Licenses

This repository is licensed under Apache-2.0; see NOTICE for third-party attribution.

External models keep their own terms. Unrestricted:

ComponentLicense
MoGeMIT
DepthProApple sample-code license
TRELLISMIT
One-2-3-45 / Zero123-XLApache-2.0 / MIT
RoMaMIT, DINOv2 backbone Apache-2.0
OneFormer, CLIPSegMIT, Apache-2.0

Non-commercial, optional in our pipeline:

ComponentLicense
MASt3R (bundles DUSt3R)CC BY-NC-SA 4.0
InvSRNTU S-Lab 1.0
Hunyuan3D-1Tencent Hunyuan Non-Commercial
UniDepthCC BY-NC 4.0
OVSAM, EntityV2S-Lab 1.0, CC BY-NC 4.0

COCO images keep their original Flickr terms; COCO3D annotations are CC BY 4.0.

3d-labelling
monocular-3d-detection

UVA-Computer-Vision-Lab/LabelAny3D

[NeurIPS 2025] LabelAny3D: Label Any Object 3D in the Wild

Python

135

3 commits

updated Aug 23, 2026

See the code

README

LabelAny3D: Label Any Object 3D in the Wild

Jin Yao, Radowan Mahmud Redoy, Sebastian Elbaum, Matthew B. Dwyer, Zezhou Cheng

Website Paper

Samples from COCO3D dataset COCO3D samples

COCO3D Dataset

The evaluation set of COCO3D and pseudo-labeled training set are available at Hugging Face.

3D BBox Human Refinement Interface

We release the source code for the refinement interface at https://github.com/UVA-Computer-Vision-Lab/3d_annotator.

Getting Started

πŸ“¦ Installation Guide - Setup instructions and external dependencies

πŸ“– COCO Pipeline Guide - Run the pipeline on COCO dataset

πŸ”§ OVMono3D Fine-tuning - Code for fine-tuning OVMono3D on LabelAny3D pseudo annotations

Citing

If you find this work useful for your research, please kindly cite:

@inproceedings{yao2025labelany3d,
  title={LabelAny3D: Label Any Object 3D in the Wild},
  author={Jin Yao and Radowan Mahmud Redoy and Sebastian Elbaum and Matthew B. Dwyer and Zezhou Cheng},
  booktitle={Neural Information Processing Systems (NeurIPS)},
  year={2025}
}

@inproceedings{yao2025open,
  title={Open Vocabulary Monocular 3D Object Detection},
  author={Yao, Jin and Gu, Hao and Chen, Xuweiyi and Wang, Jiayun and Cheng, Zezhou},
  booktitle={Proceedings of the International Conference on 3D Vision (3DV)},
  year={2026}
}

Acknowledgements

This work builds on many open-source projects:

  • Gen3DSR - 3D reconstruction framework
  • TRELLIS - 3D asset generation
  • MoGe - Monocular geometry estimation
  • DepthPro - Metric depth estimation
  • MASt3R - Dense matching
  • InvSR - Image super-resolution
  • COCONUT - COCO segmentation annotations
  • OVMono3D - Open vocabulary monocular 3D detection

Licenses

This repository is licensed under Apache-2.0; see NOTICE for third-party attribution.

External models keep their own terms. Unrestricted:

ComponentLicense
MoGeMIT
DepthProApple sample-code license
TRELLISMIT
One-2-3-45 / Zero123-XLApache-2.0 / MIT
RoMaMIT, DINOv2 backbone Apache-2.0
OneFormer, CLIPSegMIT, Apache-2.0

Non-commercial, optional in our pipeline:

ComponentLicense
MASt3R (bundles DUSt3R)CC BY-NC-SA 4.0
InvSRNTU S-Lab 1.0
Hunyuan3D-1Tencent Hunyuan Non-Commercial
UniDepthCC BY-NC 4.0
OVSAM, EntityV2S-Lab 1.0, CC BY-NC 4.0

COCO images keep their original Flickr terms; COCO3D annotations are CC BY 4.0.

3d-labelling
monocular-3d-detection

Languages

Python

98.6%

Shell

1.4%