EgoExOR: An Egocentric–Exocentric Operating Room Dataset for Comprehensive Understanding of Surgical Activities
2
26 commits
1 linked in READMEs
updated Jun 1, 2026
Official code of the paper "EgoExOR: An Egocentric–Exocentric Operating Room Dataset for Comprehensive Understanding of Surgical Activities" submitted at NeurIPS 2025 Datasets & Benchmarks Track.
Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either provide partial egocentric views or sparse exocentric multi-view context, but do not explore the comprehensive combination of both. We introduce EgoExOR, the first OR dataset and accompanying benchmark to fuse first-person and third-person perspectives. Spanning 94 minutes (84,553 frames at 15 FPS) of two emulated spine procedures, Ultrasound-Guided Needle Insertion and Minimally Invasive Spine Surgery, EgoExOR integrates egocentric data (RGB, gaze, hand tracking, audio) from wearable glasses, exocentric RGB and depth from RGB-D cameras, and ultrasound imagery. Its detailed scene graph annotations, covering 36 entities and 22 relations (568,235 triplets), enable robust modeling of clinical interactions, supporting tasks like action recognition and human-centric perception. We evaluate the surgical scene graph generation performance of two adapted state-of-the-art models and offer a new baseline that explicitly leverages EgoExOR’s multimodal and multi-perspective signals. This new dataset and benchmark set a new foundation for OR perception, offering a rich, multimodal resource for next-generation clinical perception.
Figure: Overview of one timepoint from the EgoExoR dataset, showcasing synchronized multi-view egocentric RGB and exocentric RGB-D video streams, live ultrasound monitor feed, audio, a fused 3D point-cloud reconstruction, and gaze, hand‐pose and scene graph annotations.
Multiple Modalities: Each take includes RGB video, audio, eye gaze tracking, hand tracking, 3D point cloud data, and annotations, all captured simultaneously.
Time-Synchronized Streams: All modalities are aligned on a common timeline, enabling precise cross-modal correlation (e.g. each video frame has corresponding gaze coordinates, hand positions, etc.).
Research Applicability: EgoExOR aims to fill the gap in both egocentric and exocentric surgrical datasets, supporting development of AI assistants, skill assessment tools, and multimodal models in medical and augmented reality domains.
The dataset is available in two formats:
Individual files are organized hierarchically by surgery type, procedure, and take, with components like RGB frames, eye gaze, and annotations stored separately for efficiency. The splits.h5 file defines the train, validation, and test splits.
metadata/
vocabulary/
entity (Dataset: name, id)
relation (Dataset: name, id)
sources/
sources (Dataset: name, id)
eye_gaze/coordinates are mapped to this sources dataset for accurate source names. Do not use takes/<take_id>/sources/ for mapping camera IDs to get the source names, though the source names are listed in the same order.dataset/
version, creation_date, title
data/
<surgery_type>/
<procedure_id>/
takes/
<take_id>/
sources/
source_count (int), source_0 (e.g., 'head_surgeon'), source_1, ...
metadata/sources, but for camera/source ID mapping (in gaze), use metadata/sources to get accurate source names.frames/
rgb (Dataset: [num_frames, num_cameras, height, width, 3], uint8)
eye_gaze/
coordinates (Dataset: [num_frames, num_ego_cameras, 3], float32)
[-1., -1.].camera_id in the last dimension must be mapped to metadata/sources for the correct source name, not to takes/<take_id>/sources/.eye_gaze_depth/
values (Dataset: [num_frames, num_ego_cameras], float32)
eye_gaze/coordinates (can use camera/source ID from coordinates).hand_tracking/
positions (Dataset: [num_frames, num_ego_cameras, 17], float32)
NaN.audio/ (Optional)
waveform (Dataset: [num_samples, 2], float32)
snippets (Dataset: [num_frames, samples_per_snippet, 2], float32)
point_cloud/
coordinates (Dataset: [num_frames, num_points, 3], float32)
colors (Dataset: [num_frames, num_points, 3], float32)
annotations/
frame_idx
rel_annotations (Dataset: [n_annotations_per_frame, 3], object (byte string))
scene_graph (Dataset: [n_annotations_per_frame, 3], float32)
metadata/vocabulary, representing relationships in a structured format.splits.h5
train, validation, test).surgery_type, procedure_id, take_id, frame_id
surgery_type: Type of surgical procedure (e.g., "appendectomy").procedure_id: Unique identifier for a specific procedure.take_id: Identifier for a specific recording (subclip) of a procedure.frame_id: Identifier for individual frames within a take.The merged dataset file consolidates all data from the individual files into a single file, including the splits defined in splits.h5. This file follows the same structure as above, with an additional splits/ directory that organizes the data into train, validation, and test subsets.
splits/
train, validation, test
surgery_type, procedure_id, take_id, frame_id
data/ directory for easy access during machine learning tasks.gzip reduces file size, critical for video and point cloud data.data/surgery/procedure/take/modality) simplifies navigation.Released under the Apache 2.0 License, permitting free academic and commercial use with attribution.
@inproceedings{NEURIPS2025_5e3ffa2c,
author = {\"{O}zsoy, Ege and Mamur, Arda and Tristram, Felix and Pellegrini, Chantal and Wysocki, Magdalena and Busam, Benjamin and Navab, Nassir},
booktitle = {Advances in Neural Information Processing Systems},
editor = {D. Belgrave and C. Zhang and H. Lin and R. Pascanu and P. Koniusz and M. Ghassemi and N. Chen},
pages = {},
publisher = {Curran Associates, Inc.},
title = {EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding},
url = {https://proceedings.neurips.cc/paper_files/paper/2025/file/5e3ffa2c53dce23986ca0f8d1d2bbc7e-Paper-Datasets_and_Benchmarks_Track.pdf},
volume = {38},
year = {2025}
}
---
## 🤝 Contributing
Contributions are welcome! Submit pull requests to improve loaders, add visualizers, or share benchmark results.
---
*Dataset URL: [ardamamur/EgoExOR](https://huggingface.co/datasets/ardamamur/EgoExOR)*
*Last Updated: May 2025*
EgoExOR: An Egocentric–Exocentric Operating Room Dataset for Comprehensive Understanding of Surgical Activities
2
26 commits
1 linked in READMEs
updated Jun 1, 2026
Official code of the paper "EgoExOR: An Egocentric–Exocentric Operating Room Dataset for Comprehensive Understanding of Surgical Activities" submitted at NeurIPS 2025 Datasets & Benchmarks Track.
Operating rooms (ORs) demand precise coordination among surgeons, nurses, and equipment in a fast-paced, occlusion-heavy environment, necessitating advanced perception models to enhance safety and efficiency. Existing datasets either provide partial egocentric views or sparse exocentric multi-view context, but do not explore the comprehensive combination of both. We introduce EgoExOR, the first OR dataset and accompanying benchmark to fuse first-person and third-person perspectives. Spanning 94 minutes (84,553 frames at 15 FPS) of two emulated spine procedures, Ultrasound-Guided Needle Insertion and Minimally Invasive Spine Surgery, EgoExOR integrates egocentric data (RGB, gaze, hand tracking, audio) from wearable glasses, exocentric RGB and depth from RGB-D cameras, and ultrasound imagery. Its detailed scene graph annotations, covering 36 entities and 22 relations (568,235 triplets), enable robust modeling of clinical interactions, supporting tasks like action recognition and human-centric perception. We evaluate the surgical scene graph generation performance of two adapted state-of-the-art models and offer a new baseline that explicitly leverages EgoExOR’s multimodal and multi-perspective signals. This new dataset and benchmark set a new foundation for OR perception, offering a rich, multimodal resource for next-generation clinical perception.
Figure: Overview of one timepoint from the EgoExoR dataset, showcasing synchronized multi-view egocentric RGB and exocentric RGB-D video streams, live ultrasound monitor feed, audio, a fused 3D point-cloud reconstruction, and gaze, hand‐pose and scene graph annotations.
Multiple Modalities: Each take includes RGB video, audio, eye gaze tracking, hand tracking, 3D point cloud data, and annotations, all captured simultaneously.
Time-Synchronized Streams: All modalities are aligned on a common timeline, enabling precise cross-modal correlation (e.g. each video frame has corresponding gaze coordinates, hand positions, etc.).
Research Applicability: EgoExOR aims to fill the gap in both egocentric and exocentric surgrical datasets, supporting development of AI assistants, skill assessment tools, and multimodal models in medical and augmented reality domains.
The dataset is available in two formats:
Individual files are organized hierarchically by surgery type, procedure, and take, with components like RGB frames, eye gaze, and annotations stored separately for efficiency. The splits.h5 file defines the train, validation, and test splits.
metadata/
vocabulary/
entity (Dataset: name, id)
relation (Dataset: name, id)
sources/
sources (Dataset: name, id)
eye_gaze/coordinates are mapped to this sources dataset for accurate source names. Do not use takes/<take_id>/sources/ for mapping camera IDs to get the source names, though the source names are listed in the same order.dataset/
version, creation_date, title
data/
<surgery_type>/
<procedure_id>/
takes/
<take_id>/
sources/
source_count (int), source_0 (e.g., 'head_surgeon'), source_1, ...
metadata/sources, but for camera/source ID mapping (in gaze), use metadata/sources to get accurate source names.frames/
rgb (Dataset: [num_frames, num_cameras, height, width, 3], uint8)
eye_gaze/
coordinates (Dataset: [num_frames, num_ego_cameras, 3], float32)
[-1., -1.].camera_id in the last dimension must be mapped to metadata/sources for the correct source name, not to takes/<take_id>/sources/.eye_gaze_depth/
values (Dataset: [num_frames, num_ego_cameras], float32)
eye_gaze/coordinates (can use camera/source ID from coordinates).hand_tracking/
positions (Dataset: [num_frames, num_ego_cameras, 17], float32)
NaN.audio/ (Optional)
waveform (Dataset: [num_samples, 2], float32)
snippets (Dataset: [num_frames, samples_per_snippet, 2], float32)
point_cloud/
coordinates (Dataset: [num_frames, num_points, 3], float32)
colors (Dataset: [num_frames, num_points, 3], float32)
annotations/
frame_idx
rel_annotations (Dataset: [n_annotations_per_frame, 3], object (byte string))
scene_graph (Dataset: [n_annotations_per_frame, 3], float32)
metadata/vocabulary, representing relationships in a structured format.splits.h5
train, validation, test).surgery_type, procedure_id, take_id, frame_id
surgery_type: Type of surgical procedure (e.g., "appendectomy").procedure_id: Unique identifier for a specific procedure.take_id: Identifier for a specific recording (subclip) of a procedure.frame_id: Identifier for individual frames within a take.The merged dataset file consolidates all data from the individual files into a single file, including the splits defined in splits.h5. This file follows the same structure as above, with an additional splits/ directory that organizes the data into train, validation, and test subsets.
splits/
train, validation, test
surgery_type, procedure_id, take_id, frame_id
data/ directory for easy access during machine learning tasks.gzip reduces file size, critical for video and point cloud data.data/surgery/procedure/take/modality) simplifies navigation.Released under the Apache 2.0 License, permitting free academic and commercial use with attribution.
@inproceedings{NEURIPS2025_5e3ffa2c,
author = {\"{O}zsoy, Ege and Mamur, Arda and Tristram, Felix and Pellegrini, Chantal and Wysocki, Magdalena and Busam, Benjamin and Navab, Nassir},
booktitle = {Advances in Neural Information Processing Systems},
editor = {D. Belgrave and C. Zhang and H. Lin and R. Pascanu and P. Koniusz and M. Ghassemi and N. Chen},
pages = {},
publisher = {Curran Associates, Inc.},
title = {EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding},
url = {https://proceedings.neurips.cc/paper_files/paper/2025/file/5e3ffa2c53dce23986ca0f8d1d2bbc7e-Paper-Datasets_and_Benchmarks_Track.pdf},
volume = {38},
year = {2025}
}
---
## 🤝 Contributing
Contributions are welcome! Submit pull requests to improve loaders, add visualizers, or share benchmark results.
---
*Dataset URL: [ardamamur/EgoExOR](https://huggingface.co/datasets/ardamamur/EgoExOR)*
*Last Updated: May 2025*