RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
7
22 commits
1 linked in READMEs
updated Dec 13, 2024

For video data downloading, please have a look at this dataset.
We are excited to release a new video-instruction dataset for embodied navigation. It contains our ongoing efforts in collecting and annotating room tour videos for training an open-world navigation agent. This release contains our generated geometry-aware video-instruction data from 1847 room tour scenes. We also release all the intermediate products, such as 3D scene reconstruction using COLMAP, relative depth estimation, open-world object tags and localization.
The released data is organized in the following structure
We show an example file structure below, which follows the original COLMAP output structure to ease visualization and processing.
- -rrFiQpQmr0_0_100
- sparse # reconstructed sparse models
- 0 # original reconstructed sub-model 0
- 1+2 # merged models from sub-model 1 and sub-model 2
- -rrFiQpQmr0_90_190
- -rrFiQpQmr0_180_280
...
- -rrFiQpQmr0_360_460%450_522 # merged models from video clip 360_460 and 450_522
- sparse
- 0%0 # merged model from 360_460-0 and 450_522-0
We show the structure of the pickle file for each video, as {'frame_name': [boxes, detction_confidence, tags]}
{
...,
"output_frame_0514.png":[tensor([[0.6810, 0.4731, 0.1097, 0.1149],
...
[0.5800, 0.8618, 0.1217, 0.2756]]), tensor([0.5931, ..., 0.2586]),
['microwave(0.59)', ..., 'stool(0.25)']
],
...
}
We show the structure of the pickle file for each video, as
{
...,
"output_frame_0514.png": PIL.Image,
...
}
We show the structure of the collated trajectory JSON file, as
[
...,
{'annotation': [{'answers': ['Move forward from an initial position near a man, veering slightly to the right past a car and multiple flower beds, approaching a building and an entrance in the distance, while passing by various plants and following a path pavement, and finally nearing a house exterior and a stone before arriving at a doorway entrance with the building still in the far distance.'],
'question': 'Describe the camera movement by listing the objects that disappear from view as it pans in one direction.',
'question_id': 'train--0Z00G94Nl8-0'}],
'image_info': [{'image_id': 'output_frame_0001.png'},
{'image_id': 'output_frame_0201.png'},
{'image_id': 'output_frame_0207.png'},
{'image_id': 'output_frame_0213.png'},
{'image_id': 'output_frame_0219.png'},
{'image_id': 'output_frame_0225.png'}],
'seqence_id': '-0Z00G94Nl8_000',
'type': 'video_desc',
'videoId': '-0Z00G94Nl8'}
...
]
[
...,
{"path": # frames selected by geometry-aware approach
["-0Z00G94Nl8_output_frame_0373.png",
"-0Z00G94Nl8_output_frame_0378.png",
"-0Z00G94Nl8_output_frame_0407.png",
"-0Z00G94Nl8_output_frame_0416.png",
"-0Z00G94Nl8_output_frame_0430.png",
"-0Z00G94Nl8_output_frame_0442.png"],
"videoId": "-0Z00G94Nl8",
"path_id": "-0Z00G94Nl8_003",
"instructions": ["Navigate through the hallway, initially observing a clock in the distance, then passing a doorway to the right and approaching a staircase ahead. Continue past an armchair and under a chandelier, moving towards an archway and a balustrade, before reaching a more confined space with a cabinet and a closer doorway. Enter the bedroom, where you encounter a dresser and a vase, with a mirror to the right and a window with curtains ahead. Move deeper into the bedroom, drawing nearer to a bed adorned with pillows and a chandelier overhead, while a window with curtains is now to the right. Finally, align with the right side of the bedroom, where the bed, lamp, nightstand, and art are in the distance, and a closet comes"],
"longId": "-0Z00G94Nl8_90_190-0|389%393%407", # id of the navigable step, consisting of colmap model id (-0Z00G94Nl8_90_190-0) and navigable actions (389,393, 407)
"heading": 0.0,
"optView": "-0Z00G94Nl8_output_frame_0407.png" # optimizing target, the frame id (407)
}
...
]
We show the structure of the collated geometry information of each frame, which would be indexed via navigable action-instruction data path, as
[
...,
# indexed by colmap model id, {video_clip_id}-{sparse_model_id}, see https://huggingface.co/datasets/roomtour3d/roomtour3d#colmap_reconstruction
'-0Z00G94Nl8_0_100-0': {
...,
# indexed by frame id in integer
220:{'real_world_position':[0.798788158796955, 0.546947373343179, -7.9179695594989665], # projected camera location in meters, comparable to cameras in its specific colmap model, i.e.,
'pos': [0.798788158796955, 0.546947373343179, -7.9179695594989665], # same to real_world_position
'camera_world_position': array([ 0.4411149 , 0.30204082, -4.37254142]),
'yaw': 0.29741764247264585, # transformed into euc
'pitch': -1.3533569345478835
}
...,
},
...
]
We provide sampled and downscaled video frames here.
Or, you can download it on your own using the YouTube video id list here.
We uphold the rights of individuals and copyright holders. If you are featured in any of our video annotations or hold copyright to a video and wish to have its annotation removed from our dataset, please reach out to us. Send an email to hmf282@gmail.com with the subject line beginning with RoomTour3D-optout, or raise an issue with the same title format. We commit to reviewing your request promptly and taking suitable action.
Our text annotations are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License. They are available strictly for non-commercial research.
If you find our work useful for your research, please consider citing the paper
@article{han2024roomtour3d,
title={RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation},
author={Mingfei Han and Liang Ma and Kamila Zhumakhanova and Ekaterina Radionova and Jingyi Zhang and Xiaojun Chang and Xiaodan Liang and Ivan Laptev},
journal={arXiv preprint arXiv:2412.08591},
year={2024}
}
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
7
22 commits
1 linked in READMEs
updated Dec 13, 2024

For video data downloading, please have a look at this dataset.
We are excited to release a new video-instruction dataset for embodied navigation. It contains our ongoing efforts in collecting and annotating room tour videos for training an open-world navigation agent. This release contains our generated geometry-aware video-instruction data from 1847 room tour scenes. We also release all the intermediate products, such as 3D scene reconstruction using COLMAP, relative depth estimation, open-world object tags and localization.
The released data is organized in the following structure
We show an example file structure below, which follows the original COLMAP output structure to ease visualization and processing.
- -rrFiQpQmr0_0_100
- sparse # reconstructed sparse models
- 0 # original reconstructed sub-model 0
- 1+2 # merged models from sub-model 1 and sub-model 2
- -rrFiQpQmr0_90_190
- -rrFiQpQmr0_180_280
...
- -rrFiQpQmr0_360_460%450_522 # merged models from video clip 360_460 and 450_522
- sparse
- 0%0 # merged model from 360_460-0 and 450_522-0
We show the structure of the pickle file for each video, as {'frame_name': [boxes, detction_confidence, tags]}
{
...,
"output_frame_0514.png":[tensor([[0.6810, 0.4731, 0.1097, 0.1149],
...
[0.5800, 0.8618, 0.1217, 0.2756]]), tensor([0.5931, ..., 0.2586]),
['microwave(0.59)', ..., 'stool(0.25)']
],
...
}
We show the structure of the pickle file for each video, as
{
...,
"output_frame_0514.png": PIL.Image,
...
}
We show the structure of the collated trajectory JSON file, as
[
...,
{'annotation': [{'answers': ['Move forward from an initial position near a man, veering slightly to the right past a car and multiple flower beds, approaching a building and an entrance in the distance, while passing by various plants and following a path pavement, and finally nearing a house exterior and a stone before arriving at a doorway entrance with the building still in the far distance.'],
'question': 'Describe the camera movement by listing the objects that disappear from view as it pans in one direction.',
'question_id': 'train--0Z00G94Nl8-0'}],
'image_info': [{'image_id': 'output_frame_0001.png'},
{'image_id': 'output_frame_0201.png'},
{'image_id': 'output_frame_0207.png'},
{'image_id': 'output_frame_0213.png'},
{'image_id': 'output_frame_0219.png'},
{'image_id': 'output_frame_0225.png'}],
'seqence_id': '-0Z00G94Nl8_000',
'type': 'video_desc',
'videoId': '-0Z00G94Nl8'}
...
]
[
...,
{"path": # frames selected by geometry-aware approach
["-0Z00G94Nl8_output_frame_0373.png",
"-0Z00G94Nl8_output_frame_0378.png",
"-0Z00G94Nl8_output_frame_0407.png",
"-0Z00G94Nl8_output_frame_0416.png",
"-0Z00G94Nl8_output_frame_0430.png",
"-0Z00G94Nl8_output_frame_0442.png"],
"videoId": "-0Z00G94Nl8",
"path_id": "-0Z00G94Nl8_003",
"instructions": ["Navigate through the hallway, initially observing a clock in the distance, then passing a doorway to the right and approaching a staircase ahead. Continue past an armchair and under a chandelier, moving towards an archway and a balustrade, before reaching a more confined space with a cabinet and a closer doorway. Enter the bedroom, where you encounter a dresser and a vase, with a mirror to the right and a window with curtains ahead. Move deeper into the bedroom, drawing nearer to a bed adorned with pillows and a chandelier overhead, while a window with curtains is now to the right. Finally, align with the right side of the bedroom, where the bed, lamp, nightstand, and art are in the distance, and a closet comes"],
"longId": "-0Z00G94Nl8_90_190-0|389%393%407", # id of the navigable step, consisting of colmap model id (-0Z00G94Nl8_90_190-0) and navigable actions (389,393, 407)
"heading": 0.0,
"optView": "-0Z00G94Nl8_output_frame_0407.png" # optimizing target, the frame id (407)
}
...
]
We show the structure of the collated geometry information of each frame, which would be indexed via navigable action-instruction data path, as
[
...,
# indexed by colmap model id, {video_clip_id}-{sparse_model_id}, see https://huggingface.co/datasets/roomtour3d/roomtour3d#colmap_reconstruction
'-0Z00G94Nl8_0_100-0': {
...,
# indexed by frame id in integer
220:{'real_world_position':[0.798788158796955, 0.546947373343179, -7.9179695594989665], # projected camera location in meters, comparable to cameras in its specific colmap model, i.e.,
'pos': [0.798788158796955, 0.546947373343179, -7.9179695594989665], # same to real_world_position
'camera_world_position': array([ 0.4411149 , 0.30204082, -4.37254142]),
'yaw': 0.29741764247264585, # transformed into euc
'pitch': -1.3533569345478835
}
...,
},
...
]
We provide sampled and downscaled video frames here.
Or, you can download it on your own using the YouTube video id list here.
We uphold the rights of individuals and copyright holders. If you are featured in any of our video annotations or hold copyright to a video and wish to have its annotation removed from our dataset, please reach out to us. Send an email to hmf282@gmail.com with the subject line beginning with RoomTour3D-optout, or raise an issue with the same title format. We commit to reviewing your request promptly and taking suitable action.
Our text annotations are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0) License. They are available strictly for non-commercial research.
If you find our work useful for your research, please consider citing the paper
@article{han2024roomtour3d,
title={RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation},
author={Mingfei Han and Liang Ma and Kamila Zhumakhanova and Ekaterina Radionova and Jingyi Zhang and Xiaojun Chang and Xiaodan Liang and Ivan Laptev},
journal={arXiv preprint arXiv:2412.08591},
year={2024}
}