π Project Page Β |Β π arXiv Β |Β π» Code Β |Β π§° EmbodiedEvalKit Β |Β π€ Models & Datasets
ποΈ Update β 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to
rft_datasets_json/. The complete JSON β media archive mapping is documented in the Dataset composition table below.
β οΈ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data used to train Embodied-R1.5. The corresponding model checkpoints, training scripts, and full evaluation suite are available at the project repository. For the Stage 1 SFT data, see the sibling dataset
Embodied-R1.5-SFT-Dataset.
π¦ ModelScope mirror. Some data files that could not be uploaded to HuggingFace are hosted on ModelScope instead: modelscope.cn/datasets/iffyuan/Embodied-R1.5-RFT-Dataset. Some datasets cannot be open-sourced due to institutional policy.
π JSON data location. All RFT training JSON files are stored in
rft_datasets_json/. Media archives referenced by these JSON files remain at the repository root or in their corresponding archive subdirectories.
This dataset is the Stage 2 reinforcement fine-tuning (RFT) corpus for Embodied-R1.5, a unified Embodied Foundation Model (EFM) built on Qwen3-VL-8B-Instruct. It is used after Stage 1 SFT to further sharpen embodied reasoning with a multi-task balanced RL recipe, and spans three core embodied capability dimensions:
RFT training uses the EasyR1 framework; see the GitHub repository for scripts and configs.
Each sample is a verifiable QA instance with a problem, multimodal references, and metadata used to compute task-specific rewards during RL:
{
"problem_id": "...",
"problem": "<image>\n...",
"images": ["path/to/image.jpg"],
"data_type": "image",
"problem_type": "point",
"options": [],
"data_source": "CoSyn-Point",
"answer": "..."
}
data_type is one of image | video.problem_type is one of point | bbox | trace | multiple choice | open-ended.point_2d) and boxes are normalized to the [0, 1000] range, regardless of original image resolution.depth value is in meters.<answer>...</answer> tags.The corpus is organized as per-task verifiable-QA JSON files under rft_datasets_json/, paired with LFS-tracked tar.gz / zip archives that contain the referenced images and videos. A small number of large archives are split into parts (handal_images/, qa/) for upload size limits.
| # | local json | hf archive |
|---|---|---|
| 1 | ER1.5_CoSyn-point_image_point.json | cosyn_images.tar.gz |
| 2 | ER1.5_Cosmos_video_qa.json | cosmos.tar.gz |
| 3 | ER1.5_Droid-Trace_image_trace.json | droid.tar.gz |
| 4 | ER1.5_EO_image_qa.json | eo-images.tar.gz |
| 5 | ER1.5_ER1-point_image_point.json | er1_point_images.tar.gz |
| 6 | ER1.5_ER1-trace_image_trace.json | er1-trace.tar.gz |
| 7 | ER1.5_ERQA2_image_qa.json | erqa.tar.gz |
| 8 | ER1.5_ERQA2_image_qa_sample3000.json | erqa.tar.gz |
| 9 | ER1.5_ERQA_Rush_image_qa.json | erqa.tar.gz |
| 10 | ER1.5_EmbSpatial_image_qa.json | embspatial_dataset.zip |
| 11 | ER1.5_HOI4D-Trace_image_trace.json | hoi4d-trace.tar.gz |
| 12 | ER1.5_HandAL_image_point.json | handal_images/handal_images.tar.gz.part00 + handal_images/handal_images.tar.gz.part01 + er1_point_images.tar.gz |
| 13 | ER1.5_InstructPart_image_point.json | instructpart.tar.gz |
| 14 | ER1.5_InternData-Trace_image_trace.json | interndata-trace.tar.gz |
| 15 | ER1.5_Ref_L4_image_point.json | ref_l4.tar.gz |
| 16 | ER1.5_Refspatial_image_point.json | refspatial.tar.gz |
| 17 | ER1.5_Robo2VLM_image_qa.json | robo2vlm_images.tar.gz |
| 18 | ER1.5_RoboVQA_image.json | robovqa_train.tar.gz |
| 19 | ER1.5_Roborefit_image_point.json | roborefit_images.tar.gz + er1_point_images.tar.gz |
| 20 | ER1.5_SAT_image_qa.json | sat-data.tar.gz |
| 21 | ER1.5_general_image_qa_filtered.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 22 | ER1.5_general_video_qa.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 23 | ER1.5_general_video_qa_50s_cleaned.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 24 | ER1.5_regular_simulation_image_point.json | regular-simulation.tar.gz |
| 25 | ER1.5_regular_synthetic_image_point.json | regular-synthetic.tar.gz |
| 26 | ER1.5_robocasa_partnet_2d_image_trace.json | robocasa_partnet.tar.gz |
| 27 | ER1.5_robocasa_partnet_3d_image_trace.json | robocasa_partnet.tar.gz |
| 28 | ER1.5_spatialssrl_image_qa.json | spatial-ssrl.tar.gz |
Additional archive er1.5_test.tar.gz at the repo root contains held-out evaluation samples and is not part of training.
RFT files always carry a modality suffix and may carry an optional cleanup suffix:
ER1.5_<dataset>_image_qa.json β image-level visual question answeringER1.5_<dataset>_image_point.json β image-level point / box groundingER1.5_<dataset>_image_trace.json β image-level visual trace generationER1.5_<dataset>_video_qa.json β video-level visual question answeringER1.5_<dataset>_image.json β image-only samples with mixed task typesOptional cleanup suffixes: _filtered.json (post-filtering), _sample3000.json (3k subsample), _50s_cleaned.json (β€50 s clips, cleaned).
To download only the RFT JSON files:
huggingface-cli download IffYuan/Embodied-R1.5-RFT-Dataset \
--repo-type dataset \
--include "rft_datasets_json/*.json" \
--local-dir Embodied-R1.5-RFT-Dataset
The Stage 1 SFT corpus (Embodied-R1.5-SFT-Dataset) is the union of (a) 33 open-source-derived ShareGPT files and (b) three automated generation pipelines. Stage 2 RFT narrows this to 28 task-specific verifiable-QA files and additionally includes:
ER1.5_ER1-point_image_point.json, ER1.5_ER1-trace_image_trace.json) β ER1 model's predictions used as reward-shaped rollouts.ERQA2, ERQA_Rush, EmbSpatial, spatialssrl) β not present in Stage 1 SFT.ER1.5_robocasa_partnet_2d_image_trace.json, ER1.5_robocasa_partnet_3d_image_trace.json) β finer-grained action traces than SFT.| Component | Status |
|---|---|
| Dataset card | β Available |
| Data JSON files | β
Available (28 verifiable-QA files in rft_datasets_json/, see table above) |
| Media archives | β
Available β er1.5_test.tar.gz is held-out and not for training |
| Full corpus | β Available |
If you find Embodied-R1.5 useful in your research, please cite our work:
@article{yuan2026embodied,
title={Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models},
author={Yuan, Yifu and Huang, Yaoting and Yao, Xianze and Li, Yutong and Zhang, Shuoheng and Han, Linqi and Li, Pengyi and Sun, Jiangeng and Jia, Wenting and Zhang, Zhao and Liu, Yuhao and Liao, Ruihao and Hu, Yucheng and Wu, Qiyu and Li, Yuxiao and Dong, Zibin and Ni, Fei and Zheng, Yan and Gu, Shuyang and Ma, Yi and Tang, Hongyao and Hu, Han and Hao, Jianye},
journal={arXiv preprint arXiv:2606.11324},
year={2026}
}
Released under the Apache 2.0 license.
π Project Page Β |Β π arXiv Β |Β π» Code Β |Β π§° EmbodiedEvalKit Β |Β π€ Models & Datasets
ποΈ Update β 2026-08-20 (20260820). All 28 Stage 2 RFT JSON annotation files have been uploaded to
rft_datasets_json/. The complete JSON β media archive mapping is documented in the Dataset composition table below.
β οΈ Partial release. This repository currently contains only a subset of the full Stage 2 RFT data used to train Embodied-R1.5. The corresponding model checkpoints, training scripts, and full evaluation suite are available at the project repository. For the Stage 1 SFT data, see the sibling dataset
Embodied-R1.5-SFT-Dataset.
π¦ ModelScope mirror. Some data files that could not be uploaded to HuggingFace are hosted on ModelScope instead: modelscope.cn/datasets/iffyuan/Embodied-R1.5-RFT-Dataset. Some datasets cannot be open-sourced due to institutional policy.
π JSON data location. All RFT training JSON files are stored in
rft_datasets_json/. Media archives referenced by these JSON files remain at the repository root or in their corresponding archive subdirectories.
This dataset is the Stage 2 reinforcement fine-tuning (RFT) corpus for Embodied-R1.5, a unified Embodied Foundation Model (EFM) built on Qwen3-VL-8B-Instruct. It is used after Stage 1 SFT to further sharpen embodied reasoning with a multi-task balanced RL recipe, and spans three core embodied capability dimensions:
RFT training uses the EasyR1 framework; see the GitHub repository for scripts and configs.
Each sample is a verifiable QA instance with a problem, multimodal references, and metadata used to compute task-specific rewards during RL:
{
"problem_id": "...",
"problem": "<image>\n...",
"images": ["path/to/image.jpg"],
"data_type": "image",
"problem_type": "point",
"options": [],
"data_source": "CoSyn-Point",
"answer": "..."
}
data_type is one of image | video.problem_type is one of point | bbox | trace | multiple choice | open-ended.point_2d) and boxes are normalized to the [0, 1000] range, regardless of original image resolution.depth value is in meters.<answer>...</answer> tags.The corpus is organized as per-task verifiable-QA JSON files under rft_datasets_json/, paired with LFS-tracked tar.gz / zip archives that contain the referenced images and videos. A small number of large archives are split into parts (handal_images/, qa/) for upload size limits.
| # | local json | hf archive |
|---|---|---|
| 1 | ER1.5_CoSyn-point_image_point.json | cosyn_images.tar.gz |
| 2 | ER1.5_Cosmos_video_qa.json | cosmos.tar.gz |
| 3 | ER1.5_Droid-Trace_image_trace.json | droid.tar.gz |
| 4 | ER1.5_EO_image_qa.json | eo-images.tar.gz |
| 5 | ER1.5_ER1-point_image_point.json | er1_point_images.tar.gz |
| 6 | ER1.5_ER1-trace_image_trace.json | er1-trace.tar.gz |
| 7 | ER1.5_ERQA2_image_qa.json | erqa.tar.gz |
| 8 | ER1.5_ERQA2_image_qa_sample3000.json | erqa.tar.gz |
| 9 | ER1.5_ERQA_Rush_image_qa.json | erqa.tar.gz |
| 10 | ER1.5_EmbSpatial_image_qa.json | embspatial_dataset.zip |
| 11 | ER1.5_HOI4D-Trace_image_trace.json | hoi4d-trace.tar.gz |
| 12 | ER1.5_HandAL_image_point.json | handal_images/handal_images.tar.gz.part00 + handal_images/handal_images.tar.gz.part01 + er1_point_images.tar.gz |
| 13 | ER1.5_InstructPart_image_point.json | instructpart.tar.gz |
| 14 | ER1.5_InternData-Trace_image_trace.json | interndata-trace.tar.gz |
| 15 | ER1.5_Ref_L4_image_point.json | ref_l4.tar.gz |
| 16 | ER1.5_Refspatial_image_point.json | refspatial.tar.gz |
| 17 | ER1.5_Robo2VLM_image_qa.json | robo2vlm_images.tar.gz |
| 18 | ER1.5_RoboVQA_image.json | robovqa_train.tar.gz |
| 19 | ER1.5_Roborefit_image_point.json | roborefit_images.tar.gz + er1_point_images.tar.gz |
| 20 | ER1.5_SAT_image_qa.json | sat-data.tar.gz |
| 21 | ER1.5_general_image_qa_filtered.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 22 | ER1.5_general_video_qa.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 23 | ER1.5_general_video_qa_50s_cleaned.json | qa/qa.tar.gz.part00 + qa/qa.tar.gz.part01 + qa/qa.tar.gz.part02 |
| 24 | ER1.5_regular_simulation_image_point.json | regular-simulation.tar.gz |
| 25 | ER1.5_regular_synthetic_image_point.json | regular-synthetic.tar.gz |
| 26 | ER1.5_robocasa_partnet_2d_image_trace.json | robocasa_partnet.tar.gz |
| 27 | ER1.5_robocasa_partnet_3d_image_trace.json | robocasa_partnet.tar.gz |
| 28 | ER1.5_spatialssrl_image_qa.json | spatial-ssrl.tar.gz |
Additional archive er1.5_test.tar.gz at the repo root contains held-out evaluation samples and is not part of training.
RFT files always carry a modality suffix and may carry an optional cleanup suffix:
ER1.5_<dataset>_image_qa.json β image-level visual question answeringER1.5_<dataset>_image_point.json β image-level point / box groundingER1.5_<dataset>_image_trace.json β image-level visual trace generationER1.5_<dataset>_video_qa.json β video-level visual question answeringER1.5_<dataset>_image.json β image-only samples with mixed task typesOptional cleanup suffixes: _filtered.json (post-filtering), _sample3000.json (3k subsample), _50s_cleaned.json (β€50 s clips, cleaned).
To download only the RFT JSON files:
huggingface-cli download IffYuan/Embodied-R1.5-RFT-Dataset \
--repo-type dataset \
--include "rft_datasets_json/*.json" \
--local-dir Embodied-R1.5-RFT-Dataset
The Stage 1 SFT corpus (Embodied-R1.5-SFT-Dataset) is the union of (a) 33 open-source-derived ShareGPT files and (b) three automated generation pipelines. Stage 2 RFT narrows this to 28 task-specific verifiable-QA files and additionally includes:
ER1.5_ER1-point_image_point.json, ER1.5_ER1-trace_image_trace.json) β ER1 model's predictions used as reward-shaped rollouts.ERQA2, ERQA_Rush, EmbSpatial, spatialssrl) β not present in Stage 1 SFT.ER1.5_robocasa_partnet_2d_image_trace.json, ER1.5_robocasa_partnet_3d_image_trace.json) β finer-grained action traces than SFT.| Component | Status |
|---|---|
| Dataset card | β Available |
| Data JSON files | β
Available (28 verifiable-QA files in rft_datasets_json/, see table above) |
| Media archives | β
Available β er1.5_test.tar.gz is held-out and not for training |
| Full corpus | β Available |
If you find Embodied-R1.5 useful in your research, please cite our work:
@article{yuan2026embodied,
title={Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models},
author={Yuan, Yifu and Huang, Yaoting and Yao, Xianze and Li, Yutong and Zhang, Shuoheng and Han, Linqi and Li, Pengyi and Sun, Jiangeng and Jia, Wenting and Zhang, Zhao and Liu, Yuhao and Liao, Ruihao and Hu, Yucheng and Wu, Qiyu and Li, Yuxiao and Dong, Zibin and Ni, Fei and Zheng, Yan and Gu, Shuyang and Ma, Yi and Tang, Hongyao and Hu, Han and Hao, Jianye},
journal={arXiv preprint arXiv:2606.11324},
year={2026}
}
Released under the Apache 2.0 license.