π Project Page Β |Β π arXiv Β |Β π» Code Β |Β π§° EmbodiedEvalKit Β |Β π€ Models & Datasets
ποΈ Update β 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to
sft_datasets_json/. The complete JSON β image/video data mapping is documented in the Dataset composition table below.
β οΈ Partial release. This repository currently contains only a subset of the full Stage 1 SFT data used to train Embodied-R1.5. The corresponding model checkpoints, training scripts, and full evaluation suite are available at the project repository.
π JSON file location. All SFT JSON files are stored in
sft_datasets_json/. They are not placed inside the individual image/video data folders.
π¦ ModelScope mirror. Some data files that could not be uploaded to HuggingFace are hosted on ModelScope instead: modelscope.cn/datasets/iffyuan/Embodied-R1.5-SFT-Dataset. Some datasets cannot be open-sourced due to institutional policy.
This dataset is the Stage 1 supervised fine-tuning (SFT) corpus for Embodied-R1.5, a unified Embodied Foundation Model (EFM) built on Qwen3-VL-8B-Instruct. The full corpus exceeds 15B tokens and spans three core embodied capability dimensions:
The data is built by integrating and restructuring open-source resources together with three automated data construction pipelines that target critical capability gaps.
Samples follow the ShareGPT conversation format, with multimodal references to images and videos:
{
"conversations": [
{"from": "human", "value": "<image>\n..."},
{"from": "gpt", "value": "...<answer>...</answer>"}
],
"image": ["path/to/image.jpg"]
}
point_2d) and boxes are normalized to the [0, 1000] range, regardless of original image resolution.depth value is in meters.<answer>...</answer> tags.All ShareGPT-format annotation files are centralized in sft_datasets_json/. The per-dataset folders contain the image/video data referenced by the JSON image / video fields.
| # | dataset_info key | Media folder | JSON file (under sft_datasets_json/) |
|---|---|---|---|
| 1 | RoboVQA | robovqa/ | ER1.5_robovqa_star.json |
| 2 | EgoPlan-IT | EgoPlan-Data/ | ER1.5_egoplan.json |
| 3 | euclid | euclid-30k/ | ER1.5_euclid.json |
| 4 | RoboPoint_object | RoboPoint-Data/ | ER1.5_robopoint_object.json |
| 5 | RoboPoint_region | RoboPoint-Data/ | ER1.5_robopoint_region.json |
| 6 | RoboRefit | er1-data/ | ER1.5_roborefit.json |
| 7 | HandAL | er1-data/ | ER1.5_handal_star.json |
| 8 | FSD-Point | er1-data/ | ER1.5_fsd_point_star.json |
| 9 | PACO-LVIS | PACO-LVIS/ | ER1.5_paco_lvis.json |
| 10 | Pixmo-Points | pixmo-points-images/ | ER1.5_pixmo_points.json |
| 11 | LVIS | VisualGenome_VG_100K_1_and_2/ | ER1.5_lvis.json |
| 12 | SAT | SAT-Data/ | ER1.5_sat.json |
| 13 | Ref-L4 | ref_l4/ | ER1.5_ref_l4.json |
| 14 | Robo2VLM_0111 | robo2vlm/ | ER1.5_robo2vlm.json |
| 15 | InstructPart | InstructPart/ | ER1.5_instructpart_star.json |
| 16 | PRISM | PRISM/ | ER1.5_prism_star.json |
| 17 | EO-1_0111 | EO-Data1.5M/ | ER1.5_eo.json |
| 18 | CoSyn-point | CoSyn-point/ | ER1.5_cosyn_star.json |
| 19 | RoboFAC | RoboFAC-dataset/ | ER1.5_robofac_star.json |
| 20 | Cosmos-Reasoning | cosmos/ | ER1.5_cosmos.json |
| 21 | VLM-3R | VSI-500K/ | ER1.5_vlm3r.json |
| 22 | MM-IF | MMIF-23k/ | ER1.5_mmif.json |
| 23 | RefSpatial | RefSpatial/ | ER1.5_refspatial.json |
| 24 | RefSpatial-3D | RefSpatial/ | ER1.5_refspatial_3d.json |
| 25 | FSD-Trace | er1-data/ | ER1.5_fsd_trace_star.json |
| 26 | PartNet-Maniskill | partnet_maniskill/ | ER1.5_partnet_maniskill_star.json |
| 27 | RoboFail | RoboFail/ | ER1.5_robofail_star.json |
| 28 | ManiskillFail | ManiskillFail/ | ER1.5_maniskill_fail_star.json |
| 29 | InternData-Trace | interndata/ | ER1.5_interndata_trace_star.json |
| 30 | HOI4D-Trace | hoi4d/ | ER1.5_hoi4d_trace_star.json |
| 31 | Droid-Trace | droid-trajectory/ | ER1.5_droid_trace_star.json |
| 32 | LLaVA-665K | llava_v1_5_mix665k/ | ER1.5_llava.json |
| 33 | BridgeDataFail | BridgeDataFail/ | ER1.5_bridgedatafail_star.json |
| 34 | OXE-trace | OXE-trace/ | ER1.5_oxe_trace_star.json |
ER1.5_<dataset>.json β direct use of open-source data without modificationER1.5_<dataset>_star.json β paper major modification or newly generated dataER1.5_<dataset>_trace_star.json β trajectory data with paper major modification| Component | Status |
|---|---|
| Dataset card | β Available |
| Data JSON files | β
Available (34 ShareGPT files in sft_datasets_json/, see table above) |
If you find Embodied-R1.5 useful in your research, please cite our work:
@article{yuan2026embodied,
title={Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models},
author={Yuan, Yifu and Huang, Yaoting and Yao, Xianze and Li, Yutong and Zhang, Shuoheng and Han, Linqi and Li, Pengyi and Sun, Jiangeng and Jia, Wenting and Zhang, Zhao and others},
journal={arXiv preprint arXiv:2606.11324},
year={2026}
}
Released under the Apache 2.0 license.
π Project Page Β |Β π arXiv Β |Β π» Code Β |Β π§° EmbodiedEvalKit Β |Β π€ Models & Datasets
ποΈ Update β 2026-08-20 (20260820). All 34 Stage 1 SFT JSON annotation files have been uploaded to
sft_datasets_json/. The complete JSON β image/video data mapping is documented in the Dataset composition table below.
β οΈ Partial release. This repository currently contains only a subset of the full Stage 1 SFT data used to train Embodied-R1.5. The corresponding model checkpoints, training scripts, and full evaluation suite are available at the project repository.
π JSON file location. All SFT JSON files are stored in
sft_datasets_json/. They are not placed inside the individual image/video data folders.
π¦ ModelScope mirror. Some data files that could not be uploaded to HuggingFace are hosted on ModelScope instead: modelscope.cn/datasets/iffyuan/Embodied-R1.5-SFT-Dataset. Some datasets cannot be open-sourced due to institutional policy.
This dataset is the Stage 1 supervised fine-tuning (SFT) corpus for Embodied-R1.5, a unified Embodied Foundation Model (EFM) built on Qwen3-VL-8B-Instruct. The full corpus exceeds 15B tokens and spans three core embodied capability dimensions:
The data is built by integrating and restructuring open-source resources together with three automated data construction pipelines that target critical capability gaps.
Samples follow the ShareGPT conversation format, with multimodal references to images and videos:
{
"conversations": [
{"from": "human", "value": "<image>\n..."},
{"from": "gpt", "value": "...<answer>...</answer>"}
],
"image": ["path/to/image.jpg"]
}
point_2d) and boxes are normalized to the [0, 1000] range, regardless of original image resolution.depth value is in meters.<answer>...</answer> tags.All ShareGPT-format annotation files are centralized in sft_datasets_json/. The per-dataset folders contain the image/video data referenced by the JSON image / video fields.
| # | dataset_info key | Media folder | JSON file (under sft_datasets_json/) |
|---|---|---|---|
| 1 | RoboVQA | robovqa/ | ER1.5_robovqa_star.json |
| 2 | EgoPlan-IT | EgoPlan-Data/ | ER1.5_egoplan.json |
| 3 | euclid | euclid-30k/ | ER1.5_euclid.json |
| 4 | RoboPoint_object | RoboPoint-Data/ | ER1.5_robopoint_object.json |
| 5 | RoboPoint_region | RoboPoint-Data/ | ER1.5_robopoint_region.json |
| 6 | RoboRefit | er1-data/ | ER1.5_roborefit.json |
| 7 | HandAL | er1-data/ | ER1.5_handal_star.json |
| 8 | FSD-Point | er1-data/ | ER1.5_fsd_point_star.json |
| 9 | PACO-LVIS | PACO-LVIS/ | ER1.5_paco_lvis.json |
| 10 | Pixmo-Points | pixmo-points-images/ | ER1.5_pixmo_points.json |
| 11 | LVIS | VisualGenome_VG_100K_1_and_2/ | ER1.5_lvis.json |
| 12 | SAT | SAT-Data/ | ER1.5_sat.json |
| 13 | Ref-L4 | ref_l4/ | ER1.5_ref_l4.json |
| 14 | Robo2VLM_0111 | robo2vlm/ | ER1.5_robo2vlm.json |
| 15 | InstructPart | InstructPart/ | ER1.5_instructpart_star.json |
| 16 | PRISM | PRISM/ | ER1.5_prism_star.json |
| 17 | EO-1_0111 | EO-Data1.5M/ | ER1.5_eo.json |
| 18 | CoSyn-point | CoSyn-point/ | ER1.5_cosyn_star.json |
| 19 | RoboFAC | RoboFAC-dataset/ | ER1.5_robofac_star.json |
| 20 | Cosmos-Reasoning | cosmos/ | ER1.5_cosmos.json |
| 21 | VLM-3R | VSI-500K/ | ER1.5_vlm3r.json |
| 22 | MM-IF | MMIF-23k/ | ER1.5_mmif.json |
| 23 | RefSpatial | RefSpatial/ | ER1.5_refspatial.json |
| 24 | RefSpatial-3D | RefSpatial/ | ER1.5_refspatial_3d.json |
| 25 | FSD-Trace | er1-data/ | ER1.5_fsd_trace_star.json |
| 26 | PartNet-Maniskill | partnet_maniskill/ | ER1.5_partnet_maniskill_star.json |
| 27 | RoboFail | RoboFail/ | ER1.5_robofail_star.json |
| 28 | ManiskillFail | ManiskillFail/ | ER1.5_maniskill_fail_star.json |
| 29 | InternData-Trace | interndata/ | ER1.5_interndata_trace_star.json |
| 30 | HOI4D-Trace | hoi4d/ | ER1.5_hoi4d_trace_star.json |
| 31 | Droid-Trace | droid-trajectory/ | ER1.5_droid_trace_star.json |
| 32 | LLaVA-665K | llava_v1_5_mix665k/ | ER1.5_llava.json |
| 33 | BridgeDataFail | BridgeDataFail/ | ER1.5_bridgedatafail_star.json |
| 34 | OXE-trace | OXE-trace/ | ER1.5_oxe_trace_star.json |
ER1.5_<dataset>.json β direct use of open-source data without modificationER1.5_<dataset>_star.json β paper major modification or newly generated dataER1.5_<dataset>_trace_star.json β trajectory data with paper major modification| Component | Status |
|---|---|
| Dataset card | β Available |
| Data JSON files | β
Available (34 ShareGPT files in sft_datasets_json/, see table above) |
If you find Embodied-R1.5 useful in your research, please cite our work:
@article{yuan2026embodied,
title={Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models},
author={Yuan, Yifu and Huang, Yaoting and Yao, Xianze and Li, Yutong and Zhang, Shuoheng and Han, Linqi and Li, Pengyi and Sun, Jiangeng and Jia, Wenting and Zhang, Zhao and others},
journal={arXiv preprint arXiv:2606.11324},
year={2026}
}
Released under the Apache 2.0 license.