HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
55
224 commits
1 linked in READMEs
updated Jul 29, 2026
2,000 hours released ยท 6 synchronized camera views ยท 480+ scenes ยท 3 mm pose accuracy ยท <40 ยตs synchronization
๐ Project Website | ๐ฆ Dataset | ๐ Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations. HiFi-UMI-2K is the large-scale dataset produced by this system and is designed to provide action-grounded supervision for manipulation policy pre-training and post-training without requiring a robot during data collection.
The release contains a curated 2,000-hour subset of a source corpus exceeding 20,000 hours and 4.32 million episodes across more than 480 scenes. Each episode includes synchronized multi-view video, calibrated bimanual end-effector trajectories, gripper states, language annotations, task metadata, and quality-control information.
The HiFi-UMI system and dataset accompany the technical report:
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
The work asks whether increasing the fidelity of robot-free data can remove the remaining real-robot teleoperation "anchor" normally used during deployment-oriented post-training. The paper reports that policies trained only on HiFi-UMI task demonstrations can be deployed directly on a real bimanual robot, with no teleoperated robot demonstrations in the training loop.
| Property | Value |
|---|---|
| Public release | 2,000 hours |
| Source corpus | 20,000+ hours |
| Source-corpus episodes | 4.32M+ |
| Collection scenes | 480+ |
| Camera views per episode | 6 |
| Local end-effector error | 3 mm |
| Cross-sensor time offset | <40 ยตs |
| Dropped frames | <2 per hour |
| Gripper-state error | <0.1ยฐ |
| Trajectory reconstruction success | 98% |
| WBC replay validation success | 98% |
| Storage format | LeRobot v3-style Parquet + MP4 |
| License | CC BY 4.0 |
The 2,000-hour release is a curated subset of the larger source corpus. Statistics explicitly labeled as source-corpus statistics describe the full processed collection rather than the released subset alone.
The capture hardware is co-designed around four fidelity requirements:
Native bimanual capture hardware: two hands and four hand-camera views.
Raw captures pass through a closed-loop data-production pipeline:
Training and evaluation results feed back into later collection plans, allowing the corpus to be rebalanced toward missing tasks, objects, interaction dynamics, and recovery behaviors.
Trajectory reconstruction and six-view replay. Click the image to play the video.
The handwriting demonstration below provides a qualitative view of the local trajectory accuracy. The reconstructed bimanual trajectory preserves millimeter-scale pen motion and remains aligned with all six camera streams.
Click the image to play the handwriting reconstruction video.
The paper compares HiFi-UMI-only post-training with conventional in-domain real-robot teleoperation post-training. Architecture, initialization, optimization, action representation, and deployment protocol are held fixed within each model family.
| Policy family | HiFi-UMI post-training | Teleoperation post-training | Difference |
|---|---|---|---|
| StarVLA-QwenPI | 51.3% (82/160) | 53.8% (86/160) | -2.5 points |
| OpenPI-ฯ0.5 | 77.5% (124/160) | 74.4% (119/160) | +3.1 points |
| LingBot-VA | 56.9% (91/160) | 57.5% (92/160) | -0.6 points |
Additional findings:
The paper also pre-trains StarVLA-QwenPI on a 4,000-hour HiFi-UMI mixture:
These figures describe the research setting reported in the technical report. They are not universal performance guarantees for every policy, robot embodiment, or downstream task.
The repository is sharded at the top level:
chunk-XXXX/
โโโ part-0000/
โโโ data/
โ โโโ chunk-000/
โ โโโ file-000.parquet
โโโ videos/
โ โโโ observation.images.head_main/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.head_main_stereo_right/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.left_hand_up/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.left_hand_down/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.right_hand_up/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.right_hand_down/
โ โโโ chunk-000/file-000.mp4
โโโ meta/
โโโ info.json
โโโ modality.json
โโโ stats.json
โโโ tasks.parquet
โโโ episodes/
โโโ chunk-000/file-000.parquet
Path templates are recorded in meta/info.json:
{
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"
}
data/chunk-000/file-000.parquet is the frame-level training table. Each row corresponds to one synchronized frame.
| Column | Type | Description |
|---|---|---|
observation.state | float32[20] | Current bimanual end-effector and gripper state |
observation.state_valid | bool[20] | Per-dimension validity mask for observation.state |
action | float32[20] | Absolute next-state target action |
action_valid | bool[20] | Per-dimension validity mask for action |
timestamp | float32 | Time relative to the beginning of the episode, in seconds |
frame_index | int64 | Frame number within the episode, starting from 0 |
episode_index | int64 | Episode identifier within the current shard |
index | int64 | Continuous global frame index within the current shard |
task_index | int64 | Index into meta/tasks.parquet |
valid.frame | bool | Whether the frame is recommended for training |
Frames with valid.frame == false are intentionally retained in both Parquet and video files. This preserves strict alignment between data rows, video frames, and timestamps. Most training pipelines should filter to valid.frame == true.
The state and action vectors contain 20 dimensions:
observation.state = [right_10d, left_10d]
action = [right_action_10d, left_action_10d]
Each hand uses the following 10-dimensional layout:
[x, y, z, rot6d_0, rot6d_1, rot6d_2, rot6d_3, rot6d_4, rot6d_5, gripper]
| Slice | Meaning |
|---|---|
0:3 | End-effector position xyz, in meters (m) |
3:9 | 6D rotation representation using the first two rows of a rotation matrix |
9:10 | Gripper opening angle, in radians (rad) |
The stored action uses the same layout and units and represents an absolute next-state target. The final frame normally repeats the final available target so that action rows and video frames remain aligned.
The policy implementations in the paper may convert these stored targets into model-specific robot-centric relative actions during training. Users should not assume that the exported action is already expressed in the relative action convention used by a particular VLA, WAM, or robot controller.
meta/modality.json records the semantic slice of every state and action block, including right and left end-effector pose, gripper state, and the mapping between video keys and original camera streams.
Coordinate-frame conventions for the HiFi-UMI capture system.
The dataset uses the following coordinate-frame conventions:
| Coordinate frame | Origin | +X axis | +Y axis | +Z axis |
|---|---|---|---|---|
| Right-hand frame | At the fingertip | Forward, along the fingertip pointing direction | To the left | Upward |
| Left-hand frame | At the fingertip | Forward, along the fingertip pointing direction | To the left | Upward |
| Head-camera frame | At the optical center of the left camera in the head-mounted stereo pair | To the right in the image plane | Downward in the image plane | Forward, along the camera optical axis |
Both hand coordinate frames use the same axis convention. The head-camera frame follows the standard optical-camera convention, with +Z pointing into the observed scene.
The definitions above describe the local coordinate frames attached to the two hands and the head camera. Their time-varying trajectories are all expressed in the same shared world coordinate frame.
Z axis is aligned with the direction of gravity.X and Y axes lie in the plane perpendicular to gravity.Because the world origin is arbitrary, absolute world positions should not be compared directly across different recordings without additional alignment. Within a recording, the head and both hand trajectories are mutually aligned and can be transformed or compared in the shared world frame.
Videos are stored as:
videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Standard video keys:
observation.images.head_main
observation.images.head_main_stereo_right
observation.images.left_hand_up
observation.images.left_hand_down
observation.images.right_hand_up
observation.images.right_hand_down
Resolution, FPS, codec, pixel format, and channel count are defined in meta/info.json["features"].
Within each video shard, episode frames are concatenated in episode_index order. The position of each episode inside every video file is recorded in meta/episodes/chunk-000/file-000.parquet:
videos/{video_key}/chunk_index
videos/{video_key}/file_index
videos/{video_key}/from_timestamp
videos/{video_key}/to_timestamp
To locate the video frame corresponding to a data row:
episode_index.from_timestamp.timestamp.meta/info.json.meta/tasks.parquet maps natural-language task descriptions to task_index:
index.name = "task"
columns = ["task_index"]
The frame table references this mapping through data.task_index. Task strings associated with each episode are also stored in the tasks field of the episode metadata.
meta/episodes/chunk-000/file-000.parquet contains one row per episode.
| Column | Description |
|---|---|
episode_index | Episode identifier |
data/chunk_index | Data Parquet chunk identifier |
data/file_index | Data Parquet file identifier |
dataset_from_index | Inclusive start index in the frame table |
dataset_to_index | Exclusive end index in the frame table |
tasks | Task-text list associated with the episode |
length | Number of frames in the episode |
videos/{video_key}/... | Location of the episode in each video shard |
stats/{feature}/... | Episode-level statistics |
Expected invariants:
length = dataset_to_index - dataset_from_index
timestamp = frame_index / fps
meta/info.json is the primary entry point:
| Field | Description |
|---|---|
codebase_version | Dataset format version |
robot_type | Capture device or robot type |
dataset_id | Dataset identifier |
total_episodes | Number of episodes in the current part |
total_frames | Number of frames in the current part |
total_tasks | Number of tasks in the current part |
fps | Output frame rate |
splits | Train-split ranges |
data_path | Data Parquet path template |
video_path | Video path template |
features | Dtypes, shapes, frame rates, and video properties |
state_layout | Semantic layout of observation.state |
action_layout | Semantic layout of action |
video_processing | Cropping, scaling, encoding, and alignment strategy |
invalid_policy | Invalid-frame handling policy |
source | Export provenance and media mappings |
meta/stats.json contains part-level feature statistics:
min
max
mean
std
count
These values can be used for normalization. If a training job uses only valid.frame == true, we recommend recomputing normalization statistics after applying the same filtering rule.
If you use HiFi-UMI, please cite the technical report:
@article{simpleai2026hifiumi,
title = {HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone},
author = {{Simple AI} and Wei, Yuteng and Ma, Jinming and Wang, Jiawei and Zhou, Weitao and Zuo, Yushen and Rui, Ke and Li, Minglei and Zhang, Jinhao and Pan, Zhikang and Wang, Xiang and Jia, Haoran and Du, Huan and Zeng, Zicheng and Ma, Jun and Qin, Guiyu and Zhang, Di and Li, Xiaofei},
journal = {arXiv preprint arXiv:2607.25895},
year = {2026}
}
The dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
You may share and adapt the data, including for commercial use, provided that appropriate attribution is given, a link to the license is included, and modifications are indicated.
HiFi-UMI-2K is produced by Simple AI with contributions from the capture-system, data-engine, annotation, policy-learning, and real-robot evaluation teams, together with the operators and reviewers who collected and verified the demonstrations.
For project updates, paper release information, and additional videos, visit the HiFi-UMI project website.
HiFi-UMI-2K: High-Fidelity Robot-Free Manipulation Data
55
224 commits
1 linked in READMEs
updated Jul 29, 2026
2,000 hours released ยท 6 synchronized camera views ยท 480+ scenes ยท 3 mm pose accuracy ยท <40 ยตs synchronization
๐ Project Website | ๐ฆ Dataset | ๐ Paper: arXiv:2607.25895
Examples from the HiFi-UMI corpus. Click the image to play the video.
HiFi-UMI is a portable, high-fidelity bimanual capture system for collecting robot-free manipulation demonstrations. HiFi-UMI-2K is the large-scale dataset produced by this system and is designed to provide action-grounded supervision for manipulation policy pre-training and post-training without requiring a robot during data collection.
The release contains a curated 2,000-hour subset of a source corpus exceeding 20,000 hours and 4.32 million episodes across more than 480 scenes. Each episode includes synchronized multi-view video, calibrated bimanual end-effector trajectories, gripper states, language annotations, task metadata, and quality-control information.
The HiFi-UMI system and dataset accompany the technical report:
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone
The work asks whether increasing the fidelity of robot-free data can remove the remaining real-robot teleoperation "anchor" normally used during deployment-oriented post-training. The paper reports that policies trained only on HiFi-UMI task demonstrations can be deployed directly on a real bimanual robot, with no teleoperated robot demonstrations in the training loop.
| Property | Value |
|---|---|
| Public release | 2,000 hours |
| Source corpus | 20,000+ hours |
| Source-corpus episodes | 4.32M+ |
| Collection scenes | 480+ |
| Camera views per episode | 6 |
| Local end-effector error | 3 mm |
| Cross-sensor time offset | <40 ยตs |
| Dropped frames | <2 per hour |
| Gripper-state error | <0.1ยฐ |
| Trajectory reconstruction success | 98% |
| WBC replay validation success | 98% |
| Storage format | LeRobot v3-style Parquet + MP4 |
| License | CC BY 4.0 |
The 2,000-hour release is a curated subset of the larger source corpus. Statistics explicitly labeled as source-corpus statistics describe the full processed collection rather than the released subset alone.
The capture hardware is co-designed around four fidelity requirements:
Native bimanual capture hardware: two hands and four hand-camera views.
Raw captures pass through a closed-loop data-production pipeline:
Training and evaluation results feed back into later collection plans, allowing the corpus to be rebalanced toward missing tasks, objects, interaction dynamics, and recovery behaviors.
Trajectory reconstruction and six-view replay. Click the image to play the video.
The handwriting demonstration below provides a qualitative view of the local trajectory accuracy. The reconstructed bimanual trajectory preserves millimeter-scale pen motion and remains aligned with all six camera streams.
Click the image to play the handwriting reconstruction video.
The paper compares HiFi-UMI-only post-training with conventional in-domain real-robot teleoperation post-training. Architecture, initialization, optimization, action representation, and deployment protocol are held fixed within each model family.
| Policy family | HiFi-UMI post-training | Teleoperation post-training | Difference |
|---|---|---|---|
| StarVLA-QwenPI | 51.3% (82/160) | 53.8% (86/160) | -2.5 points |
| OpenPI-ฯ0.5 | 77.5% (124/160) | 74.4% (119/160) | +3.1 points |
| LingBot-VA | 56.9% (91/160) | 57.5% (92/160) | -0.6 points |
Additional findings:
The paper also pre-trains StarVLA-QwenPI on a 4,000-hour HiFi-UMI mixture:
These figures describe the research setting reported in the technical report. They are not universal performance guarantees for every policy, robot embodiment, or downstream task.
The repository is sharded at the top level:
chunk-XXXX/
โโโ part-0000/
โโโ data/
โ โโโ chunk-000/
โ โโโ file-000.parquet
โโโ videos/
โ โโโ observation.images.head_main/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.head_main_stereo_right/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.left_hand_up/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.left_hand_down/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.right_hand_up/
โ โ โโโ chunk-000/file-000.mp4
โ โโโ observation.images.right_hand_down/
โ โโโ chunk-000/file-000.mp4
โโโ meta/
โโโ info.json
โโโ modality.json
โโโ stats.json
โโโ tasks.parquet
โโโ episodes/
โโโ chunk-000/file-000.parquet
Path templates are recorded in meta/info.json:
{
"data_path": "data/chunk-{chunk_index:03d}/file-{file_index:03d}.parquet",
"video_path": "videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4"
}
data/chunk-000/file-000.parquet is the frame-level training table. Each row corresponds to one synchronized frame.
| Column | Type | Description |
|---|---|---|
observation.state | float32[20] | Current bimanual end-effector and gripper state |
observation.state_valid | bool[20] | Per-dimension validity mask for observation.state |
action | float32[20] | Absolute next-state target action |
action_valid | bool[20] | Per-dimension validity mask for action |
timestamp | float32 | Time relative to the beginning of the episode, in seconds |
frame_index | int64 | Frame number within the episode, starting from 0 |
episode_index | int64 | Episode identifier within the current shard |
index | int64 | Continuous global frame index within the current shard |
task_index | int64 | Index into meta/tasks.parquet |
valid.frame | bool | Whether the frame is recommended for training |
Frames with valid.frame == false are intentionally retained in both Parquet and video files. This preserves strict alignment between data rows, video frames, and timestamps. Most training pipelines should filter to valid.frame == true.
The state and action vectors contain 20 dimensions:
observation.state = [right_10d, left_10d]
action = [right_action_10d, left_action_10d]
Each hand uses the following 10-dimensional layout:
[x, y, z, rot6d_0, rot6d_1, rot6d_2, rot6d_3, rot6d_4, rot6d_5, gripper]
| Slice | Meaning |
|---|---|
0:3 | End-effector position xyz, in meters (m) |
3:9 | 6D rotation representation using the first two rows of a rotation matrix |
9:10 | Gripper opening angle, in radians (rad) |
The stored action uses the same layout and units and represents an absolute next-state target. The final frame normally repeats the final available target so that action rows and video frames remain aligned.
The policy implementations in the paper may convert these stored targets into model-specific robot-centric relative actions during training. Users should not assume that the exported action is already expressed in the relative action convention used by a particular VLA, WAM, or robot controller.
meta/modality.json records the semantic slice of every state and action block, including right and left end-effector pose, gripper state, and the mapping between video keys and original camera streams.
Coordinate-frame conventions for the HiFi-UMI capture system.
The dataset uses the following coordinate-frame conventions:
| Coordinate frame | Origin | +X axis | +Y axis | +Z axis |
|---|---|---|---|---|
| Right-hand frame | At the fingertip | Forward, along the fingertip pointing direction | To the left | Upward |
| Left-hand frame | At the fingertip | Forward, along the fingertip pointing direction | To the left | Upward |
| Head-camera frame | At the optical center of the left camera in the head-mounted stereo pair | To the right in the image plane | Downward in the image plane | Forward, along the camera optical axis |
Both hand coordinate frames use the same axis convention. The head-camera frame follows the standard optical-camera convention, with +Z pointing into the observed scene.
The definitions above describe the local coordinate frames attached to the two hands and the head camera. Their time-varying trajectories are all expressed in the same shared world coordinate frame.
Z axis is aligned with the direction of gravity.X and Y axes lie in the plane perpendicular to gravity.Because the world origin is arbitrary, absolute world positions should not be compared directly across different recordings without additional alignment. Within a recording, the head and both hand trajectories are mutually aligned and can be transformed or compared in the shared world frame.
Videos are stored as:
videos/{video_key}/chunk-{chunk_index:03d}/file-{file_index:03d}.mp4
Standard video keys:
observation.images.head_main
observation.images.head_main_stereo_right
observation.images.left_hand_up
observation.images.left_hand_down
observation.images.right_hand_up
observation.images.right_hand_down
Resolution, FPS, codec, pixel format, and channel count are defined in meta/info.json["features"].
Within each video shard, episode frames are concatenated in episode_index order. The position of each episode inside every video file is recorded in meta/episodes/chunk-000/file-000.parquet:
videos/{video_key}/chunk_index
videos/{video_key}/file_index
videos/{video_key}/from_timestamp
videos/{video_key}/to_timestamp
To locate the video frame corresponding to a data row:
episode_index.from_timestamp.timestamp.meta/info.json.meta/tasks.parquet maps natural-language task descriptions to task_index:
index.name = "task"
columns = ["task_index"]
The frame table references this mapping through data.task_index. Task strings associated with each episode are also stored in the tasks field of the episode metadata.
meta/episodes/chunk-000/file-000.parquet contains one row per episode.
| Column | Description |
|---|---|
episode_index | Episode identifier |
data/chunk_index | Data Parquet chunk identifier |
data/file_index | Data Parquet file identifier |
dataset_from_index | Inclusive start index in the frame table |
dataset_to_index | Exclusive end index in the frame table |
tasks | Task-text list associated with the episode |
length | Number of frames in the episode |
videos/{video_key}/... | Location of the episode in each video shard |
stats/{feature}/... | Episode-level statistics |
Expected invariants:
length = dataset_to_index - dataset_from_index
timestamp = frame_index / fps
meta/info.json is the primary entry point:
| Field | Description |
|---|---|
codebase_version | Dataset format version |
robot_type | Capture device or robot type |
dataset_id | Dataset identifier |
total_episodes | Number of episodes in the current part |
total_frames | Number of frames in the current part |
total_tasks | Number of tasks in the current part |
fps | Output frame rate |
splits | Train-split ranges |
data_path | Data Parquet path template |
video_path | Video path template |
features | Dtypes, shapes, frame rates, and video properties |
state_layout | Semantic layout of observation.state |
action_layout | Semantic layout of action |
video_processing | Cropping, scaling, encoding, and alignment strategy |
invalid_policy | Invalid-frame handling policy |
source | Export provenance and media mappings |
meta/stats.json contains part-level feature statistics:
min
max
mean
std
count
These values can be used for normalization. If a training job uses only valid.frame == true, we recommend recomputing normalization statistics after applying the same filtering rule.
If you use HiFi-UMI, please cite the technical report:
@article{simpleai2026hifiumi,
title = {HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone},
author = {{Simple AI} and Wei, Yuteng and Ma, Jinming and Wang, Jiawei and Zhou, Weitao and Zuo, Yushen and Rui, Ke and Li, Minglei and Zhang, Jinhao and Pan, Zhikang and Wang, Xiang and Jia, Haoran and Du, Huan and Zeng, Zicheng and Ma, Jun and Qin, Guiyu and Zhang, Di and Li, Xiaofei},
journal = {arXiv preprint arXiv:2607.25895},
year = {2026}
}
The dataset is released under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
You may share and adapt the data, including for commercial use, provided that appropriate attribution is given, a link to the license is included, and modifications are indicated.
HiFi-UMI-2K is produced by Simple AI with contributions from the capture-system, data-engine, annotation, policy-learning, and real-robot evaluation teams, together with the operators and reviewers who collected and verified the demonstrations.
For project updates, paper release information, and additional videos, visit the HiFi-UMI project website.