This dataset contains real-world robot teleoperation demonstrations collected using a 7-DoF robotic arm equipped with a dexterous hand and a head-mounted RGB camera. Each episode provides synchronized numerical state/action data and video recordings. The dataset is used for finetuning in the project VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Project page: https://microsoft.github.io/VITRA/
Each episode consists of two synchronized files:
<episode_id>.h5 β numerical data including robot states, actions, kinematics,
and metadata<episode_id>.mp4 β RGB video stream recorded from the head-mounted cameraThe two files correspond one-to-one and share the same episode identifier.
The dataset uses the following coordinate frames:

Figure 1. The hand_mount frame axes. Axis directions follow the human hand definition illustrated in the figure.
The dataset format is compatible with both right-arm-only episodes and dual-arm episodes. The currently released dataset contains only right-arm data.
/meta/has_left, /meta/has_right (episode-level)/mask/* (frame-level)Each .h5 file follows the structure below:
/
βββ meta/
β βββ instruction string
β βββ video_path string
β βββ frame_count int # T
β βββ fps float
β βββ has_left bool
β βββ has_right bool
β
βββ kinematics/
β βββ left_ee_urdf_to_hand_mount (4, 4) float64
β βββ right_ee_urdf_to_hand_mount (4, 4) float64
β βββ head_camera_to_left_arm_base (4, 4) float64
β βββ head_camera_to_right_arm_base (4, 4) float64
β
βββ observation/
β βββ camera/
β βββ intrinsics (3, 3) float64
β
βββ state/
β βββ left_arm_joint (T, Na) float64 # joint positions (rad)
β βββ right_arm_joint (T, Na) float64
β βββ left_hand_mount_pose (T, 6) float64 # hand_mount pose in arm_base: [x,y,z,rx,ry,rz]
β βββ right_hand_mount_pose (T, 6) float64 # hand_mount pose in arm_base: [x,y,z,rx,ry,rz]
| βββ left_hand_mount_pose_in_cam (T, 6) float64 # hand_mount pose in head_camera: [x,y,z,rx,ry,rz]
| βββ right_hand_mount_pose_in_cam (T, 6) float64 # hand_mount pose in head_camera: [x,y,z,rx,ry,rz]
β βββ left_hand_joint (T, Nh) float64
β βββ right_hand_joint (T, Nh) float64
β
βββ action/
β βββ left_arm_joint (T, Na) float64 # target joint positions (rad)
β βββ right_arm_joint (T, Na) float64 # target joint positions (rad)
β βββ left_hand_joint (T, Nh) float64 # target joint positions (rad)
β βββ right_hand_joint (T, Nh) float64 # target joint positions (rad)
β
βββ mask/
βββ left_arm (T,) bool
βββ right_arm (T,) bool
βββ left_hand (T,) bool
βββ right_hand (T,) bool
For all *_hand_mount_pose entries, poses are represented as:
[x, y, z, rx, ry, rz]
where:
(x, y, z) denotes the position of the hand_mount frame expressed in
arm_base (meters)(rx, ry, rz) denotes the rotation vector in axisβangle representation
(radians)A homogeneous transformation matrix is denoted by T (4Γ4).
All subscripts and superscripts are written on the right-hand side of T.
Example: T^{hand\_mount}_{arm\_base} represents the pose of hand_mount
expressed in the arm_base frame.
Different flange hardware or camera mounting configurations may be used across episodes or arms. As a result:
All kinematic and extrinsic transforms must be read from the current episode and must not be assumed constant.
The hand mounting pose expressed in arm_base is computed as:
T^{ee_urdf}{arm_base} \cdot T^{hand_mount}{ee_urdf} $$
where:
T^{ee\_urdf}_{arm\_base} is obtained via forward kinematics (FK) from the arm
joint positions, corresponding to the URDF end-effector frame (joint7).T^{hand\_mount}_{ee\_urdf} is a fixed, episode-specific transform provided under
/kinematics/*_ee_urdf_to_hand_mount.Camera extrinsics may also vary across episodes.
Transforms under /kinematics/head_camera_to_*_arm_base should likewise be
read from the current episode and must not be assumed constant.
The hand mounting pose expressed in head_camera frame (i.e. *_hand_mount_pose_in_cam) is:
(T^{head_camera}{arm_base})^{-1} \cdot T^{hand_mount}{arm_base} $$
where:
T^{head\_camera}_{arm\_base} is episode-specific transform provided under /kinematics/head_camera_to_*_arm_baseThis dataset contains real-world robot teleoperation demonstrations collected using a 7-DoF robotic arm equipped with a dexterous hand and a head-mounted RGB camera. Each episode provides synchronized numerical state/action data and video recordings. The dataset is used for finetuning in the project VITRA: Scalable Vision-Language-Action Model Pretraining for Robotic Manipulation with Real-Life Human Activity Videos
Project page: https://microsoft.github.io/VITRA/
Each episode consists of two synchronized files:
<episode_id>.h5 β numerical data including robot states, actions, kinematics,
and metadata<episode_id>.mp4 β RGB video stream recorded from the head-mounted cameraThe two files correspond one-to-one and share the same episode identifier.
The dataset uses the following coordinate frames:

Figure 1. The hand_mount frame axes. Axis directions follow the human hand definition illustrated in the figure.
The dataset format is compatible with both right-arm-only episodes and dual-arm episodes. The currently released dataset contains only right-arm data.
/meta/has_left, /meta/has_right (episode-level)/mask/* (frame-level)Each .h5 file follows the structure below:
/
βββ meta/
β βββ instruction string
β βββ video_path string
β βββ frame_count int # T
β βββ fps float
β βββ has_left bool
β βββ has_right bool
β
βββ kinematics/
β βββ left_ee_urdf_to_hand_mount (4, 4) float64
β βββ right_ee_urdf_to_hand_mount (4, 4) float64
β βββ head_camera_to_left_arm_base (4, 4) float64
β βββ head_camera_to_right_arm_base (4, 4) float64
β
βββ observation/
β βββ camera/
β βββ intrinsics (3, 3) float64
β
βββ state/
β βββ left_arm_joint (T, Na) float64 # joint positions (rad)
β βββ right_arm_joint (T, Na) float64
β βββ left_hand_mount_pose (T, 6) float64 # hand_mount pose in arm_base: [x,y,z,rx,ry,rz]
β βββ right_hand_mount_pose (T, 6) float64 # hand_mount pose in arm_base: [x,y,z,rx,ry,rz]
| βββ left_hand_mount_pose_in_cam (T, 6) float64 # hand_mount pose in head_camera: [x,y,z,rx,ry,rz]
| βββ right_hand_mount_pose_in_cam (T, 6) float64 # hand_mount pose in head_camera: [x,y,z,rx,ry,rz]
β βββ left_hand_joint (T, Nh) float64
β βββ right_hand_joint (T, Nh) float64
β
βββ action/
β βββ left_arm_joint (T, Na) float64 # target joint positions (rad)
β βββ right_arm_joint (T, Na) float64 # target joint positions (rad)
β βββ left_hand_joint (T, Nh) float64 # target joint positions (rad)
β βββ right_hand_joint (T, Nh) float64 # target joint positions (rad)
β
βββ mask/
βββ left_arm (T,) bool
βββ right_arm (T,) bool
βββ left_hand (T,) bool
βββ right_hand (T,) bool
For all *_hand_mount_pose entries, poses are represented as:
[x, y, z, rx, ry, rz]
where:
(x, y, z) denotes the position of the hand_mount frame expressed in
arm_base (meters)(rx, ry, rz) denotes the rotation vector in axisβangle representation
(radians)A homogeneous transformation matrix is denoted by T (4Γ4).
All subscripts and superscripts are written on the right-hand side of T.
Example: T^{hand\_mount}_{arm\_base} represents the pose of hand_mount
expressed in the arm_base frame.
Different flange hardware or camera mounting configurations may be used across episodes or arms. As a result:
All kinematic and extrinsic transforms must be read from the current episode and must not be assumed constant.
The hand mounting pose expressed in arm_base is computed as:
T^{ee_urdf}{arm_base} \cdot T^{hand_mount}{ee_urdf} $$
where:
T^{ee\_urdf}_{arm\_base} is obtained via forward kinematics (FK) from the arm
joint positions, corresponding to the URDF end-effector frame (joint7).T^{hand\_mount}_{ee\_urdf} is a fixed, episode-specific transform provided under
/kinematics/*_ee_urdf_to_hand_mount.Camera extrinsics may also vary across episodes.
Transforms under /kinematics/head_camera_to_*_arm_base should likewise be
read from the current episode and must not be assumed constant.
The hand mounting pose expressed in head_camera frame (i.e. *_hand_mount_pose_in_cam) is:
(T^{head_camera}{arm_base})^{-1} \cdot T^{hand_mount}{arm_base} $$
where:
T^{head\_camera}_{arm\_base} is episode-specific transform provided under /kinematics/head_camera_to_*_arm_base