[CVPR 2025] InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
Python
209
110 commits
updated Sep 28, 2026
Sirui Xu*
Dongting Li*
Yucheng Zhang*
Xiyan Xu*
Qi Long*
Ziyin Wang*
Yunzhi Lu
Shuchang Dong
Hezi Jiang
Akshat Gupta
Yu-Xiong Wang
Liang-Yan Gui
University of Illinois Urbana Champaign
*Equal contribution
CVPR 2025

We introduce InterAct, a comprehensive large-scale 3D human-object interaction (HOI) dataset, originally comprising 21.81 hours of HOI data consolidated from diverse sources, the dataset is meticulously refined by correcting contact artifacts and augmented with varied motion patterns to extend the total duration to approximately 30 hours. It includes 34.1K sequence-level detailed text descriptions.
The InterAct dataset is consolidated according to the licenses of its original data sources. For data approved for redistribution, direct download links are provided; for others, we supply processing code to convert the raw data into our standardized format.
Please follow the steps below to download, process, and organize the data. And make sure to review the license before use.
Please fill out this form to request non-commercial access to InterAct and InterAct-X. Once authorized, you'll receive the download links. Organize the data from NeuralDome, IMHD, CHAIRS, OMOMO, and its corrected and augmented data according to the following directory structure.
data
│── neuraldome
│ ├── objects
│ │ └── baseball
│ │ ├── baseball.obj # object mesh
│ │ └── sample_points.npy # sampled object pointcloud
│ └── ...
│ ├── objects_bps
│ │ └── baseball
│ │ └── baseball.npy # static bps representation
│ └── ...
│ ├── sequences
│ │ └── subject01_baseball_0
│ │ ├── action.npy
│ │ ├── action.txt
│ │ ├── human.npz
│ │ ├── markers.npy
│ │ ├── joints.npy
│ │ ├── motion.npy
│ │ ├── object.npz
│ │ └── text.txt
│ └── ...
│ └── sequences_canonical
│ └── subject01_baseball_0
│ ├── action.npy
│ ├── action.txt
│ ├── human.npz
│ ├── markers.npy
│ ├── joints.npy
│ ├── motion.npy
│ ├── object.npz
│ └── text.txt
│ └── ...
│── imhd
│── chairs
│── omomo
└── annotations
The GRAB, BEHAVE, INTERCAP datasets are available for academic research under custom licenses from the Max Planck Institute for Intelligent Systems. Note that we do not distribute the original motion data—instead, we provide the processing code and annotations. Besides, we support ParaHome and ARCTIC in addition to our original dataset. To download these datasets, please visit their respective websites and agree to the terms of their licenses:
Download SMPL+H, SMPLX, DMPLs.
Download SMPL+H mode from SMPL+H (choose Extended SMPL+H model used in the AMASS project), DMPL model from DMPL (choose DMPLs compatible with SMPL), and SMPL-X model from SMPL-X. Then, please place all the models under ./models/. The ./models/ folder tree should be:
models
│── smplh
│ ├── female
│ │ ├── model.npz
│ ├── male
│ │ ├── model.npz
│ ├── neutral
│ │ ├── model.npz
│ ├── SMPLH_FEMALE.pkl
│ ├── SMPLH_MALE.pkl
│ └── SMPLH_NEUTRAL.pkl
└── smplx
├── SMPLX_FEMALE.npz
├── SMPLX_FEMALE.pkl
├── SMPLX_MALE.npz
├── SMPLX_MALE.pkl
├── SMPLX_NEUTRAL.npz
└── SMPLX_NEUTRAL.pkl
Please follow smplx tools to merge SMPL-H and MANO parameters.
Prepare Environment
Create and activate a fresh environment:
conda create -n interact python=3.8
conda activate interact
pip install torch==2.0.0 torchvision==0.15.1 torchaudio==2.0.1 --index-url https://download.pytorch.org/whl/cu118
To install PyTorch3D, please follow the official instructions: Pytorch3D.
Install remaining packages:
pip install -r requirements.txt
python -m spacy download en_core_web_sm
bash install_human_body_prior.sh
BEHAVE
Download the motion data from this link, and put them into ./data/behave/sequences. Download object data from this link, and put them into ./data/behave/objects.
Expected File Structure:
data/behave/
├── sequences
│ ├── data_name
│ ├── object_fit_all.npz # object's pose sequences
│ └── smpl_fit_all.npz # human's pose sequences
└── objects
└── object_name
├── object_name.jpg # one photo of the object
├── object_name.obj # reconstructed 3D scan of the object
├── object_name.obj.mtl # mesh material property
├── object_name_tex.jpg # mesh texture
└── object_name_fxxx.ply # simplified object mesh
OMOMO
Download the dataset from this link, and download the text annotations from this link.
Expected File Structure:
data/omomo/raw
├── omomo_text_anno_json_data # Annotation JSON data
├── captured_objects
│ └── object_name_cleaned_simplified.obj # Simplified object mesh
├── test_diffusion_manip_seq_joints24.p # Test sequences
└── train_diffusion_manip_seq_joints24.p # Train sequences
InterCap
Dowload InterCap from the the project website. Please download the one with "new results via newly trained LEMO hand models"
Expected File Structure:
data/intercap/raw
└── 01
└── 01
└── Seg_id
├── res.pkl # Human and Object Motion
└── Mesh
└── 00000_second_obj.ply # Object mesh
...
GRAB
Download GRAB from the project website.
Expected File Structure:
data/grab/raw
├── grab
│ ├── s1
│ └── seq_name.npz # Human and Object Motion
...
└── tools
├── object_meshes # Object mesh
├── object_settings
├── subject_meshes # Subject mesh
└── subject_settings
ParaHome
Download ParaHome from the project website.
Download the annot2item.json from the project repository. In annot2item.json and the text_annotations.json for each sequence, there will be motions interacting with "cabinet", but there is no cabinet mesh in the dataset. After validating with visualization, all motion sequnces involving object "cabinet" is actually interacting with object "sink", so the process script will treat cabinet as sink.
Expected File Structure:
data/parahome/raw
├── seq
│ ├── s1
│ ├── text_annotations.json
│ ├── object_transformations.pkl
│ ├── object_in_scene.json
│ ├── joint_states.pkl
│ ├── joint_positions.pkl
│ ├── head_tips.pkl
│ ├── hand_joint_orientations.pkl
│ ├── bone_vectors.pkl
│ ├── body_joint_orientations.pkl
│ └── body_global_transform.pkl
...
├── scan
│ ├── book
│ └── simplified
│ └── base.obj
...
├── smplx_seq
│ ├── s1
│ ├── smplx_params.pkl
│ └── smplx_pose.pkl
└── annot2item.json
ARCTIC
Download raw sequences, and meta files from the project website
Download text annotations from this project website. The descriptions are labeled manually by this project. The motion sequences are split into sub-sequences with length 200-400 frames based on these descriptions. If a description is missing or unable to segment, the sequence will be processed but remain unsplit.
Expected File Structure:
data/arctic
├── description
│ ├── s01
│ ├── box_grab_01
│ └── description.txt
...
...
└── raw
├── meta
│ ├── object_vtemplates
│ ├── box
│ ├── bottom_keypoints_300.json
│ ├── bottom.obj
...
│ ├── mesh.obj
...
│ ├── parts.json
│ ├── top_keypoints_300.json
│ └── top.obj
...
│ ├── subject_vtemplates
│ ├── s01.obj
...
│ └── s10.obj
...
└── raw_seqs
├── s01
├── box_grab_01.smplx.npy
...
...
Data Processing
After organizing the raw data, execute the following steps to process the datasets into our standard representations.
Run the processing scripts for each dataset:
python process/process_behave.py
python process/process_grab.py
python process/process_intercap.py
python process/process_omomo.py
python process/process_parahome.py
python process/process_arctic.py
Canonicalize the object mesh:
python process/canonicalize_obj.py
Segment the sequences according to annotations and generate associated text files:
python process/process_text.py
python process/process_text_omomo.py
After processing, the directory structure under data/ should include all sub-datasets, including:
data
├── annotation
├── behave
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── omomo
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── intercap
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── grab
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── parahome
│ ├── objects
│ │ └── object_name
│ │ ├── base.obj
│ │ └── part1.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object_{object_name}_{part}.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object_{object_name}_{part}.npz
│ └── text.txt
└── arctic
├── objects
│ └── object_name
│ ├── top.obj
│ ├── bottom.obj
│ └── mesh.obj
├── sequences_seg
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
└── sequences_canonical
└── id
├── human.npz
├── object.npz
└── text.txt
For dataset parahome involving multiple objects with mixing rigid objects each with a single part and articulated objects each with multiple parts, the transformation of every part of the object is stored independently as a object_{object_name}_{part}.npz file with keys `angles` for rotation, `trans` for translation, and name.
For dataset arctic involving articulated objects with bottom and top parts, each object transformation is stored in only one objet.npz file with keys `angles` for rotation, `trans` for translation, name, and `arti` for the relative rotation of the top part with respect to the bottom part of the object. Refer to the visualization/visualize_arctic.py code for usage of arti.
Canonicalize the human data by running:
python process/canonicalize_human.py
# or multi_thread for speedup
python process/canonicalize_human_multi_thread.py
Sample object keypoints:
python process/sample_obj.py
Extract motion representations:
python process/motion_representation.py
Process the object bps for training:
python process/process_bps.py
To get the corrected OMOMO, please fill out this form to request non-commercial access, or process from scratch following the scripts below.
Step1: Correct the full-body hoi by:
python ./hoi_correction/optimize_fullbody.py --dataset behave
python ./hoi_correction/optimize_fullbody_intercap.py --dataset intercap
Step2: Correct the wrist by:
python ./hoi_correction/scan_diff.py --dataset omomo
python ./hoi_correction/correct_wrist.py --dataset omomo
Step3: Correct the hand by:
python ./hoi_correction/optimize.py --dataset omomo
python ./hoi_correction/optimize_hand_behave.py --dataset behave
Data
Register on the SMPL-X website, go to the
downloads section to get the correspondences and sample data,
by clicking on the Model correspondences button.
Create a folder
named transfer_data and extract the downloaded zip there. You should have the
following folder structure now:
process/smpl_conversion/transfer_data
├── meshes
│ ├── smpl
│ ├── smplx
├── smpl2smplh_def_transfer.pkl
├── smpl2smplx_deftrafo_setup.pkl
├── smplh2smpl_def_transfer.pkl
├── smplh2smplx_deftrafo_setup.pkl
├── smplx2smpl_deftrafo_setup.pkl
├── smplx2smplh_deftrafo_setup.pkl
├── smplx_mask_ids.npy
Unify the SMPL representation by:
cd ./process/smpl_conversion
python -m transfer_model --exp-cfg config_files/smplx2smplh.yaml --dataset grab
--dataset: dataset in [grab, omomo, chairs, intercap]
We adapt the smpl conversion code from https://github.com/vchoutas/smplx.git , special thanks to them! We have released the SMPL-H data for NeuralDome, IMHD, CHAIRS, and OMOMO at LIGHT.
python process/motion_representation_LIGHT.py
For the SMPL-H data for NeuralDome, IMHD, CHAIRS, and OMOMO, please refer to our release here.To load and explore our data, please refer to the demo notebook.
This pipeline depends on the requirements listed in the InterMimic project. Please make sure all dependencies are installed before running the script.
After completing the data preparation steps above, run the following to generate the simulation assets:
cd simulation
python interact2mimic.py --dataset_name [dataset]
After processing, the generated files will be organized as follows:
Motion files (.pt) are stored in
simulation/intermimic/InterAct/{dataset}
SMPL humanoid files (.xml) are stored in
simulation/intermimic/data/assets/{model_type}
Object files (.urdf) are stored in
simulation/intermimic/data/assets/objects/{dataset}
For details on data loading, replaying, and training with the processed data, please refer to the InterMimic repository. We adapt the conversion code from PHC, special thanks to them!
Additional dependency:
pointnet2_opsis required for this module. Runbash install_pointnet2_ops.shfrom the project root to install it.
Download pretrained model and evaluator models:
To train on our benchmark, execute the following steps:
cd text2interaction
python -m train.hoi_diff --save_dir ./save/t2m_interact --dataset interact
To evaluate on our benchmark, execute the following steps
Evaluate on the marker representation:
cd text2interaction
bash ./scripts/eval.sh
Evaluate on the marker representation with contact guidance used:
cd text2interaction
bash ./scripts/eval_wguide.sh
To inference with the trained model, execute the following steps
Inference with contact guidance:
cd text2interaction
bash ./scripts/run_sample_guide_contact.sh
Inference without contact guidance:
cd text2interaction
bash ./scripts/run_sample_nonguide.sh
To train on our benchmark, execute the following steps:
cd object2human
bash ./scripts/Train_markerContact_VecDist.sh
To evaluate on our benchmark, execute the following steps
cd object2human
bash ./scripts/Eval.sh
To train on our benchmark, execute the following steps:
cd human2object
bash ./scripts/train.sh
To evaluate on our benchmark, execute the following steps
cd human2object
bash ./scripts/eval_metrics.sh
To visualize the dataset, execute the following steps:
Run the visualization script:
python visualization/visualize.py [dataset_name]
Replace [dataset_name] with one of the following: behave, neuraldome, intercap, omomo, grab, imhd, chairs.
To visualize markers, run:
python visualization/visualize_markers.py
If you find this repository useful for your work, please cite:
@inproceedings{xu2025interact,
title = {{InterAct}: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation},
author = {Xu, Sirui and Li, Dongting and Zhang, Yucheng and Xu, Xiyan and Long, Qi and Wang, Ziyin and Lu, Yunzhi and Dong, Shuchang and Jiang, Hezi and Gupta, Akshat and Wang, Yu-Xiong and Gui, Liang-Yan},
booktitle = {CVPR},
year = {2025},
}
Please also consider citing the specific sub-dataset you used from InterAct as follows:
@inproceedings{taheri2020grab,
title = {{GRAB}: A Dataset of Whole-Body Human Grasping of Objects},
author = {Taheri, Omid and Ghorbani, Nima and Black, Michael J. and Tzionas, Dimitrios},
booktitle = {ECCV},
year = {2020},
}
@inproceedings{brahmbhatt2019contactdb,
title = {{ContactDB}: Analyzing and Predicting Grasp Contact via Thermal Imaging},
author = {Brahmbhatt, Samarth and Ham, Cusuh and Kemp, Charles C. and Hays, James},
booktitle = {CVPR},
year = {2019},
}
@inproceedings{bhatnagar2022behave,
title = {{BEHAVE}: Dataset and Method for Tracking Human Object Interactions},
author = {Bhatnagar, Bharat Lal and Xie, Xianghui and Petrov, Ilya and Sminchisescu, Cristian and Theobalt, Christian and Pons-Moll, Gerard},
booktitle = {CVPR},
year = {2022},
}
@article{huang2024intercap,
title = {{InterCap}: Joint Markerless {3D} Tracking of Humans and Objects in Interaction from Multi-view {RGB-D} Images},
author = {Huang, Yinghao and Taheri, Omid and Black, Michael J. and Tzionas, Dimitrios},
journal = {IJCV},
year = {2024}
}
@inproceedings{huang2022intercap,
title = {{InterCap}: {J}oint Markerless {3D} Tracking of Humans and Objects in Interaction},
author = {Huang, Yinghao and Taheri, Omid and Black, Michael J. and Tzionas, Dimitrios},
booktitle = {GCPR},
year = {2022},
}
@inproceedings{jiang2023full,
title = {Full-body articulated human-object interaction},
author = {Jiang, Nan and Liu, Tengyu and Cao, Zhexuan and Cui, Jieming and Zhang, Zhiyuan and Chen, Yixin and Wang, He and Zhu, Yixin and Huang, Siyuan},
booktitle = {ICCV},
year = {2023}
}
@inproceedings{zhang2023neuraldome,
title = {{NeuralDome}: A Neural Modeling Pipeline on Multi-View Human-Object Interactions},
author = {Juze Zhang and Haimin Luo and Hongdi Yang and Xinru Xu and Qianyang Wu and Ye Shi and Jingyi Yu and Lan Xu and Jingya Wang},
booktitle = {CVPR},
year = {2023},
}
@article{li2023object,
title = {Object Motion Guided Human Motion Synthesis},
author = {Li, Jiaman and Wu, Jiajun and Liu, C Karen},
journal = {ACM Trans. Graph.},
year = {2023}
}
@inproceedings{zhao2024imhoi,
author = {Zhao, Chengfeng and Zhang, Juze and Du, Jiashen and Shan, Ziwei and Wang, Junye and Yu, Jingyi and Wang, Jingya and Xu, Lan},
title = {{I'M HOI}: Inertia-aware Monocular Capture of 3D Human-Object Interactions},
booktitle = {CVPR},
year = {2024},
}
@inproceedings{kim2025parahome,
title = {Parahome: Parameterizing everyday home activities towards 3d generative modeling of human-object interactions},
author = {Kim, Jeonghwan and Kim, Jisoo and Na, Jeonghyeon and Joo, Hanbyul},
booktitle = {CVPR},
year = {2025}
}
@inproceedings{fan2023arctic,
title = {{ARCTIC}: A Dataset for Dexterous Bimanual Hand-Object Manipulation},
author = {Fan, Zicong and Taheri, Omid and Tzionas, Dimitrios and Kocabas, Muhammed and Kaufmann, Manuel and Black, Michael J. and Hilliges, Otmar},
booktitle = {CVPR},
year = {2023}
}
Python
99.4%
[CVPR 2025] InterAct: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation
Python
209
110 commits
updated Sep 28, 2026
Sirui Xu*
Dongting Li*
Yucheng Zhang*
Xiyan Xu*
Qi Long*
Ziyin Wang*
Yunzhi Lu
Shuchang Dong
Hezi Jiang
Akshat Gupta
Yu-Xiong Wang
Liang-Yan Gui
University of Illinois Urbana Champaign
*Equal contribution
CVPR 2025

We introduce InterAct, a comprehensive large-scale 3D human-object interaction (HOI) dataset, originally comprising 21.81 hours of HOI data consolidated from diverse sources, the dataset is meticulously refined by correcting contact artifacts and augmented with varied motion patterns to extend the total duration to approximately 30 hours. It includes 34.1K sequence-level detailed text descriptions.
The InterAct dataset is consolidated according to the licenses of its original data sources. For data approved for redistribution, direct download links are provided; for others, we supply processing code to convert the raw data into our standardized format.
Please follow the steps below to download, process, and organize the data. And make sure to review the license before use.
Please fill out this form to request non-commercial access to InterAct and InterAct-X. Once authorized, you'll receive the download links. Organize the data from NeuralDome, IMHD, CHAIRS, OMOMO, and its corrected and augmented data according to the following directory structure.
data
│── neuraldome
│ ├── objects
│ │ └── baseball
│ │ ├── baseball.obj # object mesh
│ │ └── sample_points.npy # sampled object pointcloud
│ └── ...
│ ├── objects_bps
│ │ └── baseball
│ │ └── baseball.npy # static bps representation
│ └── ...
│ ├── sequences
│ │ └── subject01_baseball_0
│ │ ├── action.npy
│ │ ├── action.txt
│ │ ├── human.npz
│ │ ├── markers.npy
│ │ ├── joints.npy
│ │ ├── motion.npy
│ │ ├── object.npz
│ │ └── text.txt
│ └── ...
│ └── sequences_canonical
│ └── subject01_baseball_0
│ ├── action.npy
│ ├── action.txt
│ ├── human.npz
│ ├── markers.npy
│ ├── joints.npy
│ ├── motion.npy
│ ├── object.npz
│ └── text.txt
│ └── ...
│── imhd
│── chairs
│── omomo
└── annotations
The GRAB, BEHAVE, INTERCAP datasets are available for academic research under custom licenses from the Max Planck Institute for Intelligent Systems. Note that we do not distribute the original motion data—instead, we provide the processing code and annotations. Besides, we support ParaHome and ARCTIC in addition to our original dataset. To download these datasets, please visit their respective websites and agree to the terms of their licenses:
Download SMPL+H, SMPLX, DMPLs.
Download SMPL+H mode from SMPL+H (choose Extended SMPL+H model used in the AMASS project), DMPL model from DMPL (choose DMPLs compatible with SMPL), and SMPL-X model from SMPL-X. Then, please place all the models under ./models/. The ./models/ folder tree should be:
models
│── smplh
│ ├── female
│ │ ├── model.npz
│ ├── male
│ │ ├── model.npz
│ ├── neutral
│ │ ├── model.npz
│ ├── SMPLH_FEMALE.pkl
│ ├── SMPLH_MALE.pkl
│ └── SMPLH_NEUTRAL.pkl
└── smplx
├── SMPLX_FEMALE.npz
├── SMPLX_FEMALE.pkl
├── SMPLX_MALE.npz
├── SMPLX_MALE.pkl
├── SMPLX_NEUTRAL.npz
└── SMPLX_NEUTRAL.pkl
Please follow smplx tools to merge SMPL-H and MANO parameters.
Prepare Environment
Create and activate a fresh environment:
conda create -n interact python=3.8
conda activate interact
pip install torch==2.0.0 torchvision==0.15.1 torchaudio==2.0.1 --index-url https://download.pytorch.org/whl/cu118
To install PyTorch3D, please follow the official instructions: Pytorch3D.
Install remaining packages:
pip install -r requirements.txt
python -m spacy download en_core_web_sm
bash install_human_body_prior.sh
BEHAVE
Download the motion data from this link, and put them into ./data/behave/sequences. Download object data from this link, and put them into ./data/behave/objects.
Expected File Structure:
data/behave/
├── sequences
│ ├── data_name
│ ├── object_fit_all.npz # object's pose sequences
│ └── smpl_fit_all.npz # human's pose sequences
└── objects
└── object_name
├── object_name.jpg # one photo of the object
├── object_name.obj # reconstructed 3D scan of the object
├── object_name.obj.mtl # mesh material property
├── object_name_tex.jpg # mesh texture
└── object_name_fxxx.ply # simplified object mesh
OMOMO
Download the dataset from this link, and download the text annotations from this link.
Expected File Structure:
data/omomo/raw
├── omomo_text_anno_json_data # Annotation JSON data
├── captured_objects
│ └── object_name_cleaned_simplified.obj # Simplified object mesh
├── test_diffusion_manip_seq_joints24.p # Test sequences
└── train_diffusion_manip_seq_joints24.p # Train sequences
InterCap
Dowload InterCap from the the project website. Please download the one with "new results via newly trained LEMO hand models"
Expected File Structure:
data/intercap/raw
└── 01
└── 01
└── Seg_id
├── res.pkl # Human and Object Motion
└── Mesh
└── 00000_second_obj.ply # Object mesh
...
GRAB
Download GRAB from the project website.
Expected File Structure:
data/grab/raw
├── grab
│ ├── s1
│ └── seq_name.npz # Human and Object Motion
...
└── tools
├── object_meshes # Object mesh
├── object_settings
├── subject_meshes # Subject mesh
└── subject_settings
ParaHome
Download ParaHome from the project website.
Download the annot2item.json from the project repository. In annot2item.json and the text_annotations.json for each sequence, there will be motions interacting with "cabinet", but there is no cabinet mesh in the dataset. After validating with visualization, all motion sequnces involving object "cabinet" is actually interacting with object "sink", so the process script will treat cabinet as sink.
Expected File Structure:
data/parahome/raw
├── seq
│ ├── s1
│ ├── text_annotations.json
│ ├── object_transformations.pkl
│ ├── object_in_scene.json
│ ├── joint_states.pkl
│ ├── joint_positions.pkl
│ ├── head_tips.pkl
│ ├── hand_joint_orientations.pkl
│ ├── bone_vectors.pkl
│ ├── body_joint_orientations.pkl
│ └── body_global_transform.pkl
...
├── scan
│ ├── book
│ └── simplified
│ └── base.obj
...
├── smplx_seq
│ ├── s1
│ ├── smplx_params.pkl
│ └── smplx_pose.pkl
└── annot2item.json
ARCTIC
Download raw sequences, and meta files from the project website
Download text annotations from this project website. The descriptions are labeled manually by this project. The motion sequences are split into sub-sequences with length 200-400 frames based on these descriptions. If a description is missing or unable to segment, the sequence will be processed but remain unsplit.
Expected File Structure:
data/arctic
├── description
│ ├── s01
│ ├── box_grab_01
│ └── description.txt
...
...
└── raw
├── meta
│ ├── object_vtemplates
│ ├── box
│ ├── bottom_keypoints_300.json
│ ├── bottom.obj
...
│ ├── mesh.obj
...
│ ├── parts.json
│ ├── top_keypoints_300.json
│ └── top.obj
...
│ ├── subject_vtemplates
│ ├── s01.obj
...
│ └── s10.obj
...
└── raw_seqs
├── s01
├── box_grab_01.smplx.npy
...
...
Data Processing
After organizing the raw data, execute the following steps to process the datasets into our standard representations.
Run the processing scripts for each dataset:
python process/process_behave.py
python process/process_grab.py
python process/process_intercap.py
python process/process_omomo.py
python process/process_parahome.py
python process/process_arctic.py
Canonicalize the object mesh:
python process/canonicalize_obj.py
Segment the sequences according to annotations and generate associated text files:
python process/process_text.py
python process/process_text_omomo.py
After processing, the directory structure under data/ should include all sub-datasets, including:
data
├── annotation
├── behave
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── omomo
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── intercap
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── grab
│ ├── objects
│ │ └── object_name
│ │ └── object_name.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
├── parahome
│ ├── objects
│ │ └── object_name
│ │ ├── base.obj
│ │ └── part1.obj
│ ├── sequences_seg
│ │ └── id
│ │ ├── human.npz
│ │ ├── object_{object_name}_{part}.npz
│ │ └── text.txt
│ └── sequences_canonical
│ └── id
│ ├── human.npz
│ ├── object_{object_name}_{part}.npz
│ └── text.txt
└── arctic
├── objects
│ └── object_name
│ ├── top.obj
│ ├── bottom.obj
│ └── mesh.obj
├── sequences_seg
│ └── id
│ ├── human.npz
│ ├── object.npz
│ └── text.txt
└── sequences_canonical
└── id
├── human.npz
├── object.npz
└── text.txt
For dataset parahome involving multiple objects with mixing rigid objects each with a single part and articulated objects each with multiple parts, the transformation of every part of the object is stored independently as a object_{object_name}_{part}.npz file with keys `angles` for rotation, `trans` for translation, and name.
For dataset arctic involving articulated objects with bottom and top parts, each object transformation is stored in only one objet.npz file with keys `angles` for rotation, `trans` for translation, name, and `arti` for the relative rotation of the top part with respect to the bottom part of the object. Refer to the visualization/visualize_arctic.py code for usage of arti.
Canonicalize the human data by running:
python process/canonicalize_human.py
# or multi_thread for speedup
python process/canonicalize_human_multi_thread.py
Sample object keypoints:
python process/sample_obj.py
Extract motion representations:
python process/motion_representation.py
Process the object bps for training:
python process/process_bps.py
To get the corrected OMOMO, please fill out this form to request non-commercial access, or process from scratch following the scripts below.
Step1: Correct the full-body hoi by:
python ./hoi_correction/optimize_fullbody.py --dataset behave
python ./hoi_correction/optimize_fullbody_intercap.py --dataset intercap
Step2: Correct the wrist by:
python ./hoi_correction/scan_diff.py --dataset omomo
python ./hoi_correction/correct_wrist.py --dataset omomo
Step3: Correct the hand by:
python ./hoi_correction/optimize.py --dataset omomo
python ./hoi_correction/optimize_hand_behave.py --dataset behave
Data
Register on the SMPL-X website, go to the
downloads section to get the correspondences and sample data,
by clicking on the Model correspondences button.
Create a folder
named transfer_data and extract the downloaded zip there. You should have the
following folder structure now:
process/smpl_conversion/transfer_data
├── meshes
│ ├── smpl
│ ├── smplx
├── smpl2smplh_def_transfer.pkl
├── smpl2smplx_deftrafo_setup.pkl
├── smplh2smpl_def_transfer.pkl
├── smplh2smplx_deftrafo_setup.pkl
├── smplx2smpl_deftrafo_setup.pkl
├── smplx2smplh_deftrafo_setup.pkl
├── smplx_mask_ids.npy
Unify the SMPL representation by:
cd ./process/smpl_conversion
python -m transfer_model --exp-cfg config_files/smplx2smplh.yaml --dataset grab
--dataset: dataset in [grab, omomo, chairs, intercap]
We adapt the smpl conversion code from https://github.com/vchoutas/smplx.git , special thanks to them! We have released the SMPL-H data for NeuralDome, IMHD, CHAIRS, and OMOMO at LIGHT.
python process/motion_representation_LIGHT.py
For the SMPL-H data for NeuralDome, IMHD, CHAIRS, and OMOMO, please refer to our release here.To load and explore our data, please refer to the demo notebook.
This pipeline depends on the requirements listed in the InterMimic project. Please make sure all dependencies are installed before running the script.
After completing the data preparation steps above, run the following to generate the simulation assets:
cd simulation
python interact2mimic.py --dataset_name [dataset]
After processing, the generated files will be organized as follows:
Motion files (.pt) are stored in
simulation/intermimic/InterAct/{dataset}
SMPL humanoid files (.xml) are stored in
simulation/intermimic/data/assets/{model_type}
Object files (.urdf) are stored in
simulation/intermimic/data/assets/objects/{dataset}
For details on data loading, replaying, and training with the processed data, please refer to the InterMimic repository. We adapt the conversion code from PHC, special thanks to them!
Additional dependency:
pointnet2_opsis required for this module. Runbash install_pointnet2_ops.shfrom the project root to install it.
Download pretrained model and evaluator models:
To train on our benchmark, execute the following steps:
cd text2interaction
python -m train.hoi_diff --save_dir ./save/t2m_interact --dataset interact
To evaluate on our benchmark, execute the following steps
Evaluate on the marker representation:
cd text2interaction
bash ./scripts/eval.sh
Evaluate on the marker representation with contact guidance used:
cd text2interaction
bash ./scripts/eval_wguide.sh
To inference with the trained model, execute the following steps
Inference with contact guidance:
cd text2interaction
bash ./scripts/run_sample_guide_contact.sh
Inference without contact guidance:
cd text2interaction
bash ./scripts/run_sample_nonguide.sh
To train on our benchmark, execute the following steps:
cd object2human
bash ./scripts/Train_markerContact_VecDist.sh
To evaluate on our benchmark, execute the following steps
cd object2human
bash ./scripts/Eval.sh
To train on our benchmark, execute the following steps:
cd human2object
bash ./scripts/train.sh
To evaluate on our benchmark, execute the following steps
cd human2object
bash ./scripts/eval_metrics.sh
To visualize the dataset, execute the following steps:
Run the visualization script:
python visualization/visualize.py [dataset_name]
Replace [dataset_name] with one of the following: behave, neuraldome, intercap, omomo, grab, imhd, chairs.
To visualize markers, run:
python visualization/visualize_markers.py
If you find this repository useful for your work, please cite:
@inproceedings{xu2025interact,
title = {{InterAct}: Advancing Large-Scale Versatile 3D Human-Object Interaction Generation},
author = {Xu, Sirui and Li, Dongting and Zhang, Yucheng and Xu, Xiyan and Long, Qi and Wang, Ziyin and Lu, Yunzhi and Dong, Shuchang and Jiang, Hezi and Gupta, Akshat and Wang, Yu-Xiong and Gui, Liang-Yan},
booktitle = {CVPR},
year = {2025},
}
Please also consider citing the specific sub-dataset you used from InterAct as follows:
@inproceedings{taheri2020grab,
title = {{GRAB}: A Dataset of Whole-Body Human Grasping of Objects},
author = {Taheri, Omid and Ghorbani, Nima and Black, Michael J. and Tzionas, Dimitrios},
booktitle = {ECCV},
year = {2020},
}
@inproceedings{brahmbhatt2019contactdb,
title = {{ContactDB}: Analyzing and Predicting Grasp Contact via Thermal Imaging},
author = {Brahmbhatt, Samarth and Ham, Cusuh and Kemp, Charles C. and Hays, James},
booktitle = {CVPR},
year = {2019},
}
@inproceedings{bhatnagar2022behave,
title = {{BEHAVE}: Dataset and Method for Tracking Human Object Interactions},
author = {Bhatnagar, Bharat Lal and Xie, Xianghui and Petrov, Ilya and Sminchisescu, Cristian and Theobalt, Christian and Pons-Moll, Gerard},
booktitle = {CVPR},
year = {2022},
}
@article{huang2024intercap,
title = {{InterCap}: Joint Markerless {3D} Tracking of Humans and Objects in Interaction from Multi-view {RGB-D} Images},
author = {Huang, Yinghao and Taheri, Omid and Black, Michael J. and Tzionas, Dimitrios},
journal = {IJCV},
year = {2024}
}
@inproceedings{huang2022intercap,
title = {{InterCap}: {J}oint Markerless {3D} Tracking of Humans and Objects in Interaction},
author = {Huang, Yinghao and Taheri, Omid and Black, Michael J. and Tzionas, Dimitrios},
booktitle = {GCPR},
year = {2022},
}
@inproceedings{jiang2023full,
title = {Full-body articulated human-object interaction},
author = {Jiang, Nan and Liu, Tengyu and Cao, Zhexuan and Cui, Jieming and Zhang, Zhiyuan and Chen, Yixin and Wang, He and Zhu, Yixin and Huang, Siyuan},
booktitle = {ICCV},
year = {2023}
}
@inproceedings{zhang2023neuraldome,
title = {{NeuralDome}: A Neural Modeling Pipeline on Multi-View Human-Object Interactions},
author = {Juze Zhang and Haimin Luo and Hongdi Yang and Xinru Xu and Qianyang Wu and Ye Shi and Jingyi Yu and Lan Xu and Jingya Wang},
booktitle = {CVPR},
year = {2023},
}
@article{li2023object,
title = {Object Motion Guided Human Motion Synthesis},
author = {Li, Jiaman and Wu, Jiajun and Liu, C Karen},
journal = {ACM Trans. Graph.},
year = {2023}
}
@inproceedings{zhao2024imhoi,
author = {Zhao, Chengfeng and Zhang, Juze and Du, Jiashen and Shan, Ziwei and Wang, Junye and Yu, Jingyi and Wang, Jingya and Xu, Lan},
title = {{I'M HOI}: Inertia-aware Monocular Capture of 3D Human-Object Interactions},
booktitle = {CVPR},
year = {2024},
}
@inproceedings{kim2025parahome,
title = {Parahome: Parameterizing everyday home activities towards 3d generative modeling of human-object interactions},
author = {Kim, Jeonghwan and Kim, Jisoo and Na, Jeonghyeon and Joo, Hanbyul},
booktitle = {CVPR},
year = {2025}
}
@inproceedings{fan2023arctic,
title = {{ARCTIC}: A Dataset for Dexterous Bimanual Hand-Object Manipulation},
author = {Fan, Zicong and Taheri, Omid and Tzionas, Dimitrios and Kocabas, Muhammed and Kaufmann, Manuel and Black, Michael J. and Hilliges, Otmar},
booktitle = {CVPR},
year = {2023}
}
Python
99.4%