Cirquar-Tech/wm_imagined

Dataset

Imagined Data

11

500 commits

updated Sep 29, 2026

See the code

README

Imagined Data

This repository hosts imagined interaction data generated by world models across different environments, tasks, and data sources. Data are organized into separate subdatasets, with additional types of imagined data to be added over time.

The repository currently contains only the RoboTwin2.0 subdataset. Storage formats, field definitions, and loading instructions are documented in the corresponding section for each subdataset.

Dataset Index

SubdatasetDirectoryContents
RoboTwin2.0RoboTwin2.0/Imagined interaction segments for 50 RoboTwin 2.0 tasks

RoboTwin2.0

The directory structure, field definitions, array shapes, and reading examples below apply to the data under RoboTwin2.0/.

The RoboTwin2.0 subdataset contains imagined interaction segments generated by a world model for 50 RoboTwin 2.0 tasks. Data are organized by task. Each HDF5 file stores one fixed-length chunk: 21 observation frames, 20 action steps, and the corresponding rewards and episode-ending flags.

Observations include RGB images from three camera viewsβ€”head, left wrist, and right wristβ€”along with joint states, gripper states, and end-effector poses for both arms. Frame 0 is the initial observation of the chunk; frames 1–20 are subsequent observations generated by the world model. Each file represents an interaction segment and should not be treated as a complete episode.

Directory Structure

RoboTwin2.0/
β”œβ”€β”€ adjust_bottle/
β”‚   β”œβ”€β”€ chunk_<uid>.hdf5
β”‚   └── ...
β”œβ”€β”€ handover_block/
β”‚   β”œβ”€β”€ chunk_<uid>.hdf5
β”‚   └── ...
β”œβ”€β”€ place_dual_shoes/
β”‚   └── ...
└── ...                         # 50 task directories in total

The task directory name identifies the task, whereas the task field inside each file contains the natural-language instruction for that sample.

File Schema

Each file contains 18 HDF5 datasets and a root attribute named source_identity. All dataset paths in the table below are relative to the file root. The root attribute is documented separately below.

FieldShapeDtypeDescription
task()UTF-8 stringNatural-language task instruction for the current chunk
obs/head_cam/rgb(21, 240, 320, 3)uint8RGB images from the head camera
obs/left_wrist_cam/rgb(21, 240, 320, 3)uint8RGB images from the left wrist camera
obs/right_wrist_cam/rgb(21, 240, 320, 3)uint8RGB images from the right wrist camera
obs/left_joint_pos(21, 6)float32Commanded positions of the six left-arm joints (rad)
obs/right_joint_pos(21, 6)float32Commanded positions of the six right-arm joints (rad)
obs/left_gripper(21,)float32Commanded opening fraction of the left gripper
obs/right_gripper(21,)float32Commanded opening fraction of the right gripper
obs/left_ee_pose(21, 7)float32Left-arm end-effector pose: [x, y, z, qw, qx, qy, qz]
obs/right_ee_pose(21, 7)float32Right-arm end-effector pose: [x, y, z, qw, qx, qy, qz]
action/left_joint_pos(20, 6)float64Raw target position commands for the six left-arm joints (rad)
action/right_joint_pos(20, 6)float64Raw target position commands for the six right-arm joints (rad)
action/left_gripper(20,)float64Raw action values for the left gripper
action/right_gripper(20,)float64Raw action values for the right gripper
rewards(20,)float32Rewards associated with the 20 action steps
terminations(20,)boolTask or environment termination flags
truncations(20,)boolTruncation flags, for example due to time limits
dones(20,)boolterminations OR truncations

Images

Image arrays use the (T, H, W, C) layout, with RGB channel order and pixel values in 0–255. The three views are aligned at each frame index, with a resolution of 240 pixels in height and 320 pixels in width.

RGB images are stored as pixel arrays using lossless HDF5 gzip compression at level 4, with an HDF5 storage chunk shape of (1, 240, 320, 3). Reading a dataset returns image arrays directly; no additional JPEG or PNG decoding is required.

Joint States, Gripper States, and End-Effector Poses

The joint and gripper states of both arms can be combined into a 14-dimensional vector in the following order:

[left_joint_1, ..., left_joint_6, left_gripper,
 right_joint_1, ..., right_joint_6, right_gripper]

obs/*_joint_pos and obs/*_gripper store the control targets used as the policy state. Row 0 is the initial state of the chunk. For the subsequent 20 rows, obs/*_joint_pos[t+1] is taken from the target joint angles in action/*_joint_pos[t], and obs/*_gripper[t+1] is obtained by clipping the corresponding gripper action to [0, 1]. Values are stored in the dtypes listed above and have not been normalized using the policy's normalization statistics. These fields do not represent measured joint or gripper feedback after robot motion.

For RoboTwin, the policy input state should use the joint positions and gripper states described here, not the end-effector poses.

Gripper values follow the convention 0 = closed, 1 = open. action/*_gripper retains the raw policy output, which may be slightly below 0 or above 1. Commands used for imagined execution clip gripper values to [0, 1], so a gripper action may differ from the gripper state in the next frame. Joint actions specify target positions, not joint increments.

The first three components of obs/*_ee_pose are positions in the world coordinate frame, expressed in meters. The remaining four components form a quaternion in wxyz order. The pose reference is the URDF link6 joint frame after RoboTwin's joint-coordinate correction, not the tool center point (TCP).

End-effector poses in generated frames come directly from the world model's predictions. During export, they are not replaced with forward kinematics computed from joint commands. Each end-effector pose field contains only the seven pose components; the separate obs/*_gripper fields come from the control targets described above.

Temporal Alignment

For t = 0, ..., 19:

obs[t]  -- action[t] -->  obs[t + 1]
                β”‚
                └── rewards[t], terminations[t], truncations[t], dones[t]
  • obs[0] is the last conditioning observation for the current chunk. For the first chunk of a branch, it comes from the offline data; for subsequent chunks, it is the final observation of the preceding chunk.
  • obs[1:21] contains the 20 subsequent observations generated for the current chunk.
  • action[t] corresponds to the transition from obs[t] to obs[t+1]. The reward array and all three flag arrays use the same action index.

Each file therefore contains one more observation than action. The policy generates an action sequence at each chunk boundary; 21 observation frames do not imply 21 policy calls. Files also do not include the world model's complete conditioning history across multiple frames.

Rewards and Episode-Ending Flags

This subdataset stores the sparse binary rewards used during data generation. The first 19 reward entries in each chunk are zero, and the final entry stores the chunk's 0/1 reward. A reward model assigns this reward based on the generated observations; it is not a human annotation for each frame.

terminations indicates termination, and truncations indicates truncation. At every step:

dones = terminations | truncations

The last step of a chunk does not necessarily have done=True: the same branch may continue with another chunk. Use the stored flags to determine episode boundaries rather than inferring them from file boundaries alone.

Root Attribute

The root attribute source_identity in each HDF5 file stores a JSON string containing three provenance fields:

{
  "branch_id": "<unique imagined branch identifier>",
  "chunk_index": 0,
  "transition_id": "<source record identifier>"
}
FieldDescription
branch_idIdentifier of the imagined rollout branch to which this chunk belongs
chunk_indexZero-based chunk index within the branch
transition_idTransition identifier in the original export record

Read this attribute with json.loads(f.attrs["source_identity"]). It is included in the distributed HDF5 files and is not counted among the 18 datasets listed above.

To link chunks, group them by branch_id and sort by chunk_index. Concatenate chunks only when both adjacent chunks are available and their indices are consecutive. Retain the shared boundary observation between adjacent chunks only once.

Cirquar-Tech/wm_imagined

Dataset

Imagined Data

11

500 commits

updated Sep 29, 2026

See the code

README

Imagined Data

This repository hosts imagined interaction data generated by world models across different environments, tasks, and data sources. Data are organized into separate subdatasets, with additional types of imagined data to be added over time.

The repository currently contains only the RoboTwin2.0 subdataset. Storage formats, field definitions, and loading instructions are documented in the corresponding section for each subdataset.

Dataset Index

SubdatasetDirectoryContents
RoboTwin2.0RoboTwin2.0/Imagined interaction segments for 50 RoboTwin 2.0 tasks

RoboTwin2.0

The directory structure, field definitions, array shapes, and reading examples below apply to the data under RoboTwin2.0/.

The RoboTwin2.0 subdataset contains imagined interaction segments generated by a world model for 50 RoboTwin 2.0 tasks. Data are organized by task. Each HDF5 file stores one fixed-length chunk: 21 observation frames, 20 action steps, and the corresponding rewards and episode-ending flags.

Observations include RGB images from three camera viewsβ€”head, left wrist, and right wristβ€”along with joint states, gripper states, and end-effector poses for both arms. Frame 0 is the initial observation of the chunk; frames 1–20 are subsequent observations generated by the world model. Each file represents an interaction segment and should not be treated as a complete episode.

Directory Structure

RoboTwin2.0/
β”œβ”€β”€ adjust_bottle/
β”‚   β”œβ”€β”€ chunk_<uid>.hdf5
β”‚   └── ...
β”œβ”€β”€ handover_block/
β”‚   β”œβ”€β”€ chunk_<uid>.hdf5
β”‚   └── ...
β”œβ”€β”€ place_dual_shoes/
β”‚   └── ...
└── ...                         # 50 task directories in total

The task directory name identifies the task, whereas the task field inside each file contains the natural-language instruction for that sample.

File Schema

Each file contains 18 HDF5 datasets and a root attribute named source_identity. All dataset paths in the table below are relative to the file root. The root attribute is documented separately below.

FieldShapeDtypeDescription
task()UTF-8 stringNatural-language task instruction for the current chunk
obs/head_cam/rgb(21, 240, 320, 3)uint8RGB images from the head camera
obs/left_wrist_cam/rgb(21, 240, 320, 3)uint8RGB images from the left wrist camera
obs/right_wrist_cam/rgb(21, 240, 320, 3)uint8RGB images from the right wrist camera
obs/left_joint_pos(21, 6)float32Commanded positions of the six left-arm joints (rad)
obs/right_joint_pos(21, 6)float32Commanded positions of the six right-arm joints (rad)
obs/left_gripper(21,)float32Commanded opening fraction of the left gripper
obs/right_gripper(21,)float32Commanded opening fraction of the right gripper
obs/left_ee_pose(21, 7)float32Left-arm end-effector pose: [x, y, z, qw, qx, qy, qz]
obs/right_ee_pose(21, 7)float32Right-arm end-effector pose: [x, y, z, qw, qx, qy, qz]
action/left_joint_pos(20, 6)float64Raw target position commands for the six left-arm joints (rad)
action/right_joint_pos(20, 6)float64Raw target position commands for the six right-arm joints (rad)
action/left_gripper(20,)float64Raw action values for the left gripper
action/right_gripper(20,)float64Raw action values for the right gripper
rewards(20,)float32Rewards associated with the 20 action steps
terminations(20,)boolTask or environment termination flags
truncations(20,)boolTruncation flags, for example due to time limits
dones(20,)boolterminations OR truncations

Images

Image arrays use the (T, H, W, C) layout, with RGB channel order and pixel values in 0–255. The three views are aligned at each frame index, with a resolution of 240 pixels in height and 320 pixels in width.

RGB images are stored as pixel arrays using lossless HDF5 gzip compression at level 4, with an HDF5 storage chunk shape of (1, 240, 320, 3). Reading a dataset returns image arrays directly; no additional JPEG or PNG decoding is required.

Joint States, Gripper States, and End-Effector Poses

The joint and gripper states of both arms can be combined into a 14-dimensional vector in the following order:

[left_joint_1, ..., left_joint_6, left_gripper,
 right_joint_1, ..., right_joint_6, right_gripper]

obs/*_joint_pos and obs/*_gripper store the control targets used as the policy state. Row 0 is the initial state of the chunk. For the subsequent 20 rows, obs/*_joint_pos[t+1] is taken from the target joint angles in action/*_joint_pos[t], and obs/*_gripper[t+1] is obtained by clipping the corresponding gripper action to [0, 1]. Values are stored in the dtypes listed above and have not been normalized using the policy's normalization statistics. These fields do not represent measured joint or gripper feedback after robot motion.

For RoboTwin, the policy input state should use the joint positions and gripper states described here, not the end-effector poses.

Gripper values follow the convention 0 = closed, 1 = open. action/*_gripper retains the raw policy output, which may be slightly below 0 or above 1. Commands used for imagined execution clip gripper values to [0, 1], so a gripper action may differ from the gripper state in the next frame. Joint actions specify target positions, not joint increments.

The first three components of obs/*_ee_pose are positions in the world coordinate frame, expressed in meters. The remaining four components form a quaternion in wxyz order. The pose reference is the URDF link6 joint frame after RoboTwin's joint-coordinate correction, not the tool center point (TCP).

End-effector poses in generated frames come directly from the world model's predictions. During export, they are not replaced with forward kinematics computed from joint commands. Each end-effector pose field contains only the seven pose components; the separate obs/*_gripper fields come from the control targets described above.

Temporal Alignment

For t = 0, ..., 19:

obs[t]  -- action[t] -->  obs[t + 1]
                β”‚
                └── rewards[t], terminations[t], truncations[t], dones[t]
  • obs[0] is the last conditioning observation for the current chunk. For the first chunk of a branch, it comes from the offline data; for subsequent chunks, it is the final observation of the preceding chunk.
  • obs[1:21] contains the 20 subsequent observations generated for the current chunk.
  • action[t] corresponds to the transition from obs[t] to obs[t+1]. The reward array and all three flag arrays use the same action index.

Each file therefore contains one more observation than action. The policy generates an action sequence at each chunk boundary; 21 observation frames do not imply 21 policy calls. Files also do not include the world model's complete conditioning history across multiple frames.

Rewards and Episode-Ending Flags

This subdataset stores the sparse binary rewards used during data generation. The first 19 reward entries in each chunk are zero, and the final entry stores the chunk's 0/1 reward. A reward model assigns this reward based on the generated observations; it is not a human annotation for each frame.

terminations indicates termination, and truncations indicates truncation. At every step:

dones = terminations | truncations

The last step of a chunk does not necessarily have done=True: the same branch may continue with another chunk. Use the stored flags to determine episode boundaries rather than inferring them from file boundaries alone.

Root Attribute

The root attribute source_identity in each HDF5 file stores a JSON string containing three provenance fields:

{
  "branch_id": "<unique imagined branch identifier>",
  "chunk_index": 0,
  "transition_id": "<source record identifier>"
}
FieldDescription
branch_idIdentifier of the imagined rollout branch to which this chunk belongs
chunk_indexZero-based chunk index within the branch
transition_idTransition identifier in the original export record

Read this attribute with json.loads(f.attrs["source_identity"]). It is included in the distributed HDF5 files and is not counted among the 18 datasets listed above.

To link chunks, group them by branch_id and sort by chunk_index. Concatenate chunks only when both adjacent chunks are available and their indices are consecutive. Retain the shared boundary observation between adjacent chunks only once.