PointWorld-DROID is the packaged DROID-derived annotation release used for training and evaluating the 3D world model, PointWorld. It contains precomputed 3D annotations derived from the official DROID dataset, including episode-level 3D point trajectories, optimized camera metadata, optional downsampled depth, and the released expert confidence artifact used by the PointWorld DROID evaluation pipeline.
This dataset is for research and development only.
This Hugging Face repository hosts the packaged release, not the original raw DROID scenes and not prebuilt WebDataset shards. After download, users should restore the packaged archives to the canonical HDF5/JSON layout. From there, they can either use the files directly as 3D annotations or follow the PointWorld repository workflow for integrity checking, conversion, training, and evaluation.
NVIDIA Corporation
04/15/2026
NVIDIA License
Research and development for world modeling, 3D vision, and robotic manipulation.
This release is intended for two common use cases:
Data Collection Method
Labeling Method
The release is distributed in two layers:
Restored canonical layout:
droid/
cameras/
*_cameras.json
confidence/
expert_confidence-seed=42.h5
depth_320x180/
*_depth.h5
flows-fs-optimized/
*_flows.h5
Local conversion to WebDataset train/test shards is supported by the PointWorld repository tooling.
Current release package counts inspected from the restored release tree:
Each *_flows.h5 file contains multiple clip groups keyed as "{start}:{end}", so the corpus contains substantially more clips than files.
Measurement of Total Data Storage: 3.91 TB
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with the applicable terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
Archive groups in the packaged release:
In the inspected packaged build, these resolved to:
huggingface-cli download nvidia/PointWorld-DROID \
--repo-type dataset \
--local-dir /path/to/PointWorld-DROID
cd /path/to/PointWorld-DROID
bash recover_dataset_from_parts.sh \
--packages . \
--out /path/to/pointworld_droid_restored \
--threads 0
After restoration, the canonical dataset root is:
/path/to/pointworld_droid_restored/droid
For end-to-end instructions on integrity checking, HDF5-to-WebDataset conversion, training, and evaluation, please use the PointWorld GitHub repository.
If you only want the annotations and do not plan to use the PointWorld training code, the restored files can be read directly.
Each DROID flow file is an episode-level HDF5 file. Clip groups are stored under keys of the form "{start}:{end}".
Per-clip robot-state datasets typically include:
Each camera group, such as camera_20103212_ext, contains:
Important notes:
Each depth file stores per-frame downsampled depth at 320x180. Camera groups typically contain:
The optional depth package is useful for analysis and downstream research. It is not intended to replace the full raw-to-annotation generation pipeline if your goal is to reproduce PointWorld's annotation generation from raw DROID data.
Each JSON sidecar stores optimized camera metadata for one episode and includes fields such as:
The released camera JSONs are retained for optimization-success episodes only.
This file is an auxiliary DROID evaluation artifact used by the released PointWorld DROID test pipeline. If you are not reproducing the PointWorld filtered DROID metrics, you can ignore it.
import io
import h5py
import numpy as np
from PIL import Image
def decode_jpeg_object(entry) -> np.ndarray:
if isinstance(entry, np.ndarray):
jpeg_bytes = entry.astype(np.uint8, copy=False).tobytes()
elif isinstance(entry, (bytes, bytearray, memoryview)):
jpeg_bytes = bytes(entry)
elif hasattr(entry, "tobytes"):
jpeg_bytes = entry.tobytes()
else:
jpeg_bytes = bytes(entry)
return np.array(Image.open(io.BytesIO(jpeg_bytes)).convert("RGB"))
path = "/path/to/pointworld_droid_restored/droid/flows-fs-optimized/example_flows.h5"
with h5py.File(path, "r") as f:
clip_key = sorted([k for k in f.keys() if ":" in k])[0]
clip = f[clip_key]
cam_key = sorted([k for k in clip.keys() if k.startswith("camera_")])[0]
cam = clip[cam_key]
rgb = decode_jpeg_object(cam["initial_rgb"][0])
depth_m = cam["initial_depth"][:].astype(np.float32) / 1000.0
scene_flows = cam["scene_flows"][:].astype(np.float32)
scene_colors = cam["scene_colors"][:]
scene_normals = cam["scene_normals"][:].astype(np.float32) / 127.0
scene_visibility = cam["scene_visibility"][:]
scene_depth_valid_mask = cam["scene_depth_valid_mask"][:]
If you use this dataset, please cite the PointWorld paper and the original DROID dataset.
@article{huang2026pointworld,
title={PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation},
author={Huang, Wenlong and Chao, Yu-Wei and Mousavian, Arsalan and Liu, Ming-Yu and Fox, Dieter and Mo, Kaichun and Li, Fei-Fei},
journal={arXiv preprint arXiv:2601.03782},
year={2026}
}
PointWorld-DROID is the packaged DROID-derived annotation release used for training and evaluating the 3D world model, PointWorld. It contains precomputed 3D annotations derived from the official DROID dataset, including episode-level 3D point trajectories, optimized camera metadata, optional downsampled depth, and the released expert confidence artifact used by the PointWorld DROID evaluation pipeline.
This dataset is for research and development only.
This Hugging Face repository hosts the packaged release, not the original raw DROID scenes and not prebuilt WebDataset shards. After download, users should restore the packaged archives to the canonical HDF5/JSON layout. From there, they can either use the files directly as 3D annotations or follow the PointWorld repository workflow for integrity checking, conversion, training, and evaluation.
NVIDIA Corporation
04/15/2026
NVIDIA License
Research and development for world modeling, 3D vision, and robotic manipulation.
This release is intended for two common use cases:
Data Collection Method
Labeling Method
The release is distributed in two layers:
Restored canonical layout:
droid/
cameras/
*_cameras.json
confidence/
expert_confidence-seed=42.h5
depth_320x180/
*_depth.h5
flows-fs-optimized/
*_flows.h5
Local conversion to WebDataset train/test shards is supported by the PointWorld repository tooling.
Current release package counts inspected from the restored release tree:
Each *_flows.h5 file contains multiple clip groups keyed as "{start}:{end}", so the corpus contains substantially more clips than files.
Measurement of Total Data Storage: 3.91 TB
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with the applicable terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
Archive groups in the packaged release:
In the inspected packaged build, these resolved to:
huggingface-cli download nvidia/PointWorld-DROID \
--repo-type dataset \
--local-dir /path/to/PointWorld-DROID
cd /path/to/PointWorld-DROID
bash recover_dataset_from_parts.sh \
--packages . \
--out /path/to/pointworld_droid_restored \
--threads 0
After restoration, the canonical dataset root is:
/path/to/pointworld_droid_restored/droid
For end-to-end instructions on integrity checking, HDF5-to-WebDataset conversion, training, and evaluation, please use the PointWorld GitHub repository.
If you only want the annotations and do not plan to use the PointWorld training code, the restored files can be read directly.
Each DROID flow file is an episode-level HDF5 file. Clip groups are stored under keys of the form "{start}:{end}".
Per-clip robot-state datasets typically include:
Each camera group, such as camera_20103212_ext, contains:
Important notes:
Each depth file stores per-frame downsampled depth at 320x180. Camera groups typically contain:
The optional depth package is useful for analysis and downstream research. It is not intended to replace the full raw-to-annotation generation pipeline if your goal is to reproduce PointWorld's annotation generation from raw DROID data.
Each JSON sidecar stores optimized camera metadata for one episode and includes fields such as:
The released camera JSONs are retained for optimization-success episodes only.
This file is an auxiliary DROID evaluation artifact used by the released PointWorld DROID test pipeline. If you are not reproducing the PointWorld filtered DROID metrics, you can ignore it.
import io
import h5py
import numpy as np
from PIL import Image
def decode_jpeg_object(entry) -> np.ndarray:
if isinstance(entry, np.ndarray):
jpeg_bytes = entry.astype(np.uint8, copy=False).tobytes()
elif isinstance(entry, (bytes, bytearray, memoryview)):
jpeg_bytes = bytes(entry)
elif hasattr(entry, "tobytes"):
jpeg_bytes = entry.tobytes()
else:
jpeg_bytes = bytes(entry)
return np.array(Image.open(io.BytesIO(jpeg_bytes)).convert("RGB"))
path = "/path/to/pointworld_droid_restored/droid/flows-fs-optimized/example_flows.h5"
with h5py.File(path, "r") as f:
clip_key = sorted([k for k in f.keys() if ":" in k])[0]
clip = f[clip_key]
cam_key = sorted([k for k in clip.keys() if k.startswith("camera_")])[0]
cam = clip[cam_key]
rgb = decode_jpeg_object(cam["initial_rgb"][0])
depth_m = cam["initial_depth"][:].astype(np.float32) / 1000.0
scene_flows = cam["scene_flows"][:].astype(np.float32)
scene_colors = cam["scene_colors"][:]
scene_normals = cam["scene_normals"][:].astype(np.float32) / 127.0
scene_visibility = cam["scene_visibility"][:]
scene_depth_valid_mask = cam["scene_depth_valid_mask"][:]
If you use this dataset, please cite the PointWorld paper and the original DROID dataset.
@article{huang2026pointworld,
title={PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation},
author={Huang, Wenlong and Chao, Yu-Wei and Mousavian, Arsalan and Liu, Ming-Yu and Fox, Dieter and Mo, Kaichun and Li, Fei-Fei},
journal={arXiv preprint arXiv:2601.03782},
year={2026}
}