PointWorld-BEHAVIOR is the packaged BEHAVIOR-derived annotation release used for training and evaluating the 3D world model PointWorld. It contains precomputed 3D annotations derived from BEHAVIOR simulation episodes, organized as episode-level HDF5 files that store robot state, camera parameters, initial RGB-D observations, and rigid-body scene geometry annotations.
This Hugging Face repository hosts the packaged release, not the original BEHAVIOR raw episodes nor the prebuilt WebDataset shards. After download, users should restore the packaged archives to the canonical HDF5 layout. From there, they can either use the files directly as 3D annotations or follow the PointWorld repository workflow for integrity checking, conversion, training, and evaluation.
This dataset is for research and development only.
NVIDIA Corporation
04/15/2026
NVIDIA License
This release is intended for research and development in world modeling, 3D vision, and robotic manipulation.
It is intended for two common use cases:
Data Collection Method: Synthetic
Labeling Method: Synthetic
The release is distributed in two layers:
Restored canonical layout:
behavior/
flows/
task-0000/
episode_*.hdf5
task-0001/
episode_*.hdf5
...
Local conversion to WebDataset train/test shards is supported by the PointWorld repository tooling.
Current release package counts inspected from the restored release tree:
Each episode file contains multiple clip groups keyed as "{start}:{end}", so the corpus contains substantially more clips than files.
Measurement of Total Data Storage: 718GB
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with the applicable terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
The packaged release is sharded by task. Typical archive paths are:
huggingface-cli download nvidia/PointWorld-BEHAVIOR \
--repo-type dataset \
--local-dir /path/to/PointWorld-BEHAVIOR
cd /path/to/PointWorld-BEHAVIOR
bash recover_dataset_from_parts.sh \
--packages . \
--out /path/to/pointworld_behavior_restored \
--threads 0
After restoration, the canonical dataset root is:
/path/to/pointworld_behavior_restored/behavior
For end-to-end instructions on integrity checking, HDF5-to-WebDataset conversion, training, and evaluation, please use the PointWorld GitHub repository.
If you only want the annotations and do not plan to use the PointWorld training code, the restored HDF5 files can be read directly.
Each BEHAVIOR file is an episode-level HDF5 file under:
behavior/flows/task-XXXX/episode_YYYYYYYY.hdf5
Clip groups are stored under keys of the form "{start}:{end}".
Typical robot-state fields include:
Each clip contains camera groups such as:
Each camera group contains:
Unlike PointWorld-DROID, PointWorld-BEHAVIOR does not store one dense scene_flows array for the whole scene. Instead, it stores mesh-local point samples plus a rigid pose trajectory for each mesh over time. This is more storage-efficient for simulated rigid-body scenes while still allowing exact reconstruction of per-frame scene geometry.
To recover per-frame points for a given camera and timestep:
The same mesh keys are shared across:
import io
import h5py
import numpy as np
from PIL import Image
def decode_jpeg_object(entry) -> np.ndarray:
if isinstance(entry, np.ndarray):
jpeg_bytes = entry.astype(np.uint8, copy=False).tobytes()
elif isinstance(entry, (bytes, bytearray, memoryview)):
jpeg_bytes = bytes(entry)
elif hasattr(entry, "tobytes"):
jpeg_bytes = entry.tobytes()
else:
jpeg_bytes = bytes(entry)
return np.array(Image.open(io.BytesIO(jpeg_bytes)).convert("RGB"))
path = "/path/to/pointworld_behavior_restored/behavior/flows/task-0000/episode_00000000.hdf5"
with h5py.File(path, "r") as f:
clip_key = sorted([k for k in f.keys() if ":" in k])[0]
clip = f[clip_key]
cam = clip["camera_head"]
rgb = decode_jpeg_object(cam["initial_rgb"][0])
depth_m = cam["initial_depth"][:].astype(np.float32) / 1000.0
mesh_name = sorted(cam["local_scene_points"].keys())[0]
local_points = cam["local_scene_points"][mesh_name][:].astype(np.float32)
local_colors = cam["local_scene_colors"][mesh_name][:]
local_normals = cam["local_scene_normals"][mesh_name][:].astype(np.float32) / 127.0
mesh_pose_traj = cam["scene_mesh_trajectories"][mesh_name][:].astype(np.float32)
If you use this dataset, please cite the PointWorld paper and the original BEHAVIOR dataset.
@article{huang2026pointworld,
title={PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation},
author={Huang, Wenlong and Chao, Yu-Wei and Mousavian, Arsalan and Liu, Ming-Yu and Fox, Dieter and Mo, Kaichun and Li, Fei-Fei},
journal={arXiv preprint arXiv:2601.03782},
year={2026}
}
PointWorld-BEHAVIOR is the packaged BEHAVIOR-derived annotation release used for training and evaluating the 3D world model PointWorld. It contains precomputed 3D annotations derived from BEHAVIOR simulation episodes, organized as episode-level HDF5 files that store robot state, camera parameters, initial RGB-D observations, and rigid-body scene geometry annotations.
This Hugging Face repository hosts the packaged release, not the original BEHAVIOR raw episodes nor the prebuilt WebDataset shards. After download, users should restore the packaged archives to the canonical HDF5 layout. From there, they can either use the files directly as 3D annotations or follow the PointWorld repository workflow for integrity checking, conversion, training, and evaluation.
This dataset is for research and development only.
NVIDIA Corporation
04/15/2026
NVIDIA License
This release is intended for research and development in world modeling, 3D vision, and robotic manipulation.
It is intended for two common use cases:
Data Collection Method: Synthetic
Labeling Method: Synthetic
The release is distributed in two layers:
Restored canonical layout:
behavior/
flows/
task-0000/
episode_*.hdf5
task-0001/
episode_*.hdf5
...
Local conversion to WebDataset train/test shards is supported by the PointWorld repository tooling.
Current release package counts inspected from the restored release tree:
Each episode file contains multiple clip groups keyed as "{start}:{end}", so the corpus contains substantially more clips than files.
Measurement of Total Data Storage: 718GB
NVIDIA believes Trustworthy AI is a shared responsibility and we have established policies and practices to enable development for a wide array of AI applications. When downloaded or used in accordance with the applicable terms of service, developers should work with their internal developer teams to ensure this dataset meets requirements for the relevant industry and use case and addresses unforeseen product misuse.
Please report quality, risk, security vulnerabilities or NVIDIA AI Concerns here.
The packaged release is sharded by task. Typical archive paths are:
huggingface-cli download nvidia/PointWorld-BEHAVIOR \
--repo-type dataset \
--local-dir /path/to/PointWorld-BEHAVIOR
cd /path/to/PointWorld-BEHAVIOR
bash recover_dataset_from_parts.sh \
--packages . \
--out /path/to/pointworld_behavior_restored \
--threads 0
After restoration, the canonical dataset root is:
/path/to/pointworld_behavior_restored/behavior
For end-to-end instructions on integrity checking, HDF5-to-WebDataset conversion, training, and evaluation, please use the PointWorld GitHub repository.
If you only want the annotations and do not plan to use the PointWorld training code, the restored HDF5 files can be read directly.
Each BEHAVIOR file is an episode-level HDF5 file under:
behavior/flows/task-XXXX/episode_YYYYYYYY.hdf5
Clip groups are stored under keys of the form "{start}:{end}".
Typical robot-state fields include:
Each clip contains camera groups such as:
Each camera group contains:
Unlike PointWorld-DROID, PointWorld-BEHAVIOR does not store one dense scene_flows array for the whole scene. Instead, it stores mesh-local point samples plus a rigid pose trajectory for each mesh over time. This is more storage-efficient for simulated rigid-body scenes while still allowing exact reconstruction of per-frame scene geometry.
To recover per-frame points for a given camera and timestep:
The same mesh keys are shared across:
import io
import h5py
import numpy as np
from PIL import Image
def decode_jpeg_object(entry) -> np.ndarray:
if isinstance(entry, np.ndarray):
jpeg_bytes = entry.astype(np.uint8, copy=False).tobytes()
elif isinstance(entry, (bytes, bytearray, memoryview)):
jpeg_bytes = bytes(entry)
elif hasattr(entry, "tobytes"):
jpeg_bytes = entry.tobytes()
else:
jpeg_bytes = bytes(entry)
return np.array(Image.open(io.BytesIO(jpeg_bytes)).convert("RGB"))
path = "/path/to/pointworld_behavior_restored/behavior/flows/task-0000/episode_00000000.hdf5"
with h5py.File(path, "r") as f:
clip_key = sorted([k for k in f.keys() if ":" in k])[0]
clip = f[clip_key]
cam = clip["camera_head"]
rgb = decode_jpeg_object(cam["initial_rgb"][0])
depth_m = cam["initial_depth"][:].astype(np.float32) / 1000.0
mesh_name = sorted(cam["local_scene_points"].keys())[0]
local_points = cam["local_scene_points"][mesh_name][:].astype(np.float32)
local_colors = cam["local_scene_colors"][mesh_name][:]
local_normals = cam["local_scene_normals"][mesh_name][:].astype(np.float32) / 127.0
mesh_pose_traj = cam["scene_mesh_trajectories"][mesh_name][:].astype(np.float32)
If you use this dataset, please cite the PointWorld paper and the original BEHAVIOR dataset.
@article{huang2026pointworld,
title={PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation},
author={Huang, Wenlong and Chao, Yu-Wei and Mousavian, Arsalan and Liu, Ming-Yu and Fox, Dieter and Mo, Kaichun and Li, Fei-Fei},
journal={arXiv preprint arXiv:2601.03782},
year={2026}
}