ayiyayi/Awesome-Egocentric-and-Exocentric-Vision

42

12 commits

updated Nov 14, 2025

See the code

README


🌟 Awesome Egocentric-and-Exocentric Vision

Survey Paper Last Updated PRs Welcome Stars

A curated list of papers and datasets on Egocentric (first-person) and Exocentric (third-person) Vision. This repo complements our survey and is continuously updated.


📖 Contents


📑 Papers

🔹 Egocentric for Exocentric

✨ Video Generation

✨ Action Understanding

✨ View Birdification

✨ View Selection

🔹 Exocentric for Egocentric

✨ Video Generation

✨ Video Captioning

✨ Action Understanding

✨ Affordance Grounding

✨ Remote Drone Teleoperation

🔹 Joint Learning

✨ Video Captioning

✨ Cross-View Retrieval

✨ 3D Camera Localization

✨ Action Understanding

TitleVenueYearCode
Action recognition in the presence of one egocentric and multiple static camerasACCV2014-
An exocentric look at egocentric actions and vice versaCVIU2016-
Recognizing micro-actions and reactions from paired egocentric videosCVPR2016-
Actor and observer: Joint modeling of first and third-person videosCVPR2018Github
Holographic feature learning of egocentric-exocentric videos for multi-domain action recognitionTMM2021-
Holistic-guided disentangled learning with cross-video semantics mining for concurrent first-person and third-person activity recognitionTNNLS2022-
Look both ways: Self-supervising driver gaze estimation and road scene saliencyECCV2022Github
Aide: A vision-driven multi-view, multi-modal, multitasking dataset for assistive driving perceptionICCV2023Github
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignmentNeurIPS2023Github
Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view worldACMMM2023Github
Enhancing egocentric 3d pose estimation with third person viewsPR2023Github
Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instrumentsarxiv2023Github
Tell, don’t show: Language guidance eases transfer across domains in images and videosICML2024Github
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningCVPR2025Github
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsCVPR2025Github
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationACM MM2025-

✨ Egocentric Wearer Identification

✨ Cross-View Human Tracking and Association

✨ Robotic Manipulation

✨ Remote Drone Teleoperation

📊 Datasets

Action Understanding

DatasetYearPaperProjectCode
CMU-MMAC2009Guide to the carnegie mellon university multimodal activity (cmu-mmac) databaseProject-
Charades-Ego2018Actor and observer: Joint modeling of first and third-person videosProjectGithub
H2O2021H2o: Two hands manipulating objects for first person interaction recognitionProjectGithub
Assembly1012022Assembly101: A large-scale multi-view video dataset for understanding procedural activitiesProjectGithub
Homage2021Home action genome: Cooperative compositional action understandingProjectGithub
LEMMA2020Lemma: A multiview dataset for le arning m ulti-agent m ulti-task a ctivitiesProjectGithub
FT-HID2022Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis-Github
EgoPW2022Estimating egocentric 3d human pose in the wild with external weak supervisionProjectGithub
First2Third-Pose2022Enhancing egocentric 3d pose estimation with third person views-GitHub
ARCTIC2023ARCTIC: A dataset for dexterous bimanual hand-object manipulationProjectGithub
ECHP2023Egofish3d: Egocentric 3d pose estimation from a fisheye camera via selfsupervised learning-Github
AssemblyHands2023Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimationProjectGithub
EgoHumans2023Ego-humans: An ego-centric 3d multi-human benchmarkProjectGithub
-2023Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instrumentsProject-
ThermoHands2024Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal imageProjectGithub
OVR2024Ovr: A dataset for open vocabulary temporal repetition counting in videos-Github
Nymeria2024Nymeria: A massive collection of multimodal egocentric daily motion in the wildProjectGithub
OAKINK22024Oakink2: A dataset of bimanual hands-object manipulation in complex task completionProjectGithub
CORE4D2024Core4d: A 4d humanobject-human interaction dataset for collaborative object rearrangementProjectGithub
Ego-Exo4D2024Ego-exo4d: Understanding skilled human activity from first-and third-person perspectivesProjectGithub
EgoExoLearn2024Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world-Github
EgoExo-Fitness2024Egoexo-fitness: Towards egocentric and exocentric full-body action understanding-Github
EgoExOR2025EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding-Github
EPFL-Smart-Kitchen-302025EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models-Github

Driving

Affordance Grounding

Generation

Scene Understanding

Video Question Answering

Egocentric Wearer Identification

Cross-View Human Tracking and Association

Camera Registration

ayiyayi/Awesome-Egocentric-and-Exocentric-Vision

42

12 commits

updated Nov 14, 2025

See the code

README


🌟 Awesome Egocentric-and-Exocentric Vision

Survey Paper Last Updated PRs Welcome Stars

A curated list of papers and datasets on Egocentric (first-person) and Exocentric (third-person) Vision. This repo complements our survey and is continuously updated.


📖 Contents


📑 Papers

🔹 Egocentric for Exocentric

✨ Video Generation

✨ Action Understanding

✨ View Birdification

✨ View Selection

🔹 Exocentric for Egocentric

✨ Video Generation

✨ Video Captioning

✨ Action Understanding

✨ Affordance Grounding

✨ Remote Drone Teleoperation

🔹 Joint Learning

✨ Video Captioning

✨ Cross-View Retrieval

✨ 3D Camera Localization

✨ Action Understanding

TitleVenueYearCode
Action recognition in the presence of one egocentric and multiple static camerasACCV2014-
An exocentric look at egocentric actions and vice versaCVIU2016-
Recognizing micro-actions and reactions from paired egocentric videosCVPR2016-
Actor and observer: Joint modeling of first and third-person videosCVPR2018Github
Holographic feature learning of egocentric-exocentric videos for multi-domain action recognitionTMM2021-
Holistic-guided disentangled learning with cross-video semantics mining for concurrent first-person and third-person activity recognitionTNNLS2022-
Look both ways: Self-supervising driver gaze estimation and road scene saliencyECCV2022Github
Aide: A vision-driven multi-view, multi-modal, multitasking dataset for assistive driving perceptionICCV2023Github
Learning fine-grained view-invariant representations from unpaired ego-exo videos via temporal alignmentNeurIPS2023Github
Pov: Prompt-oriented view-agnostic learning for egocentric hand-object interaction in the multi-view worldACMMM2023Github
Enhancing egocentric 3d pose estimation with third person viewsPR2023Github
Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instrumentsarxiv2023Github
Tell, don’t show: Language guidance eases transfer across domains in images and videosICML2024Github
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation LearningCVPR2025Github
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video RepresentationsCVPR2025Github
Unsupervised Ego- and Exo-centric Dense Procedural Activity Captioning via Gaze Consensus AdaptationACM MM2025-

✨ Egocentric Wearer Identification

✨ Cross-View Human Tracking and Association

✨ Robotic Manipulation

✨ Remote Drone Teleoperation

📊 Datasets

Action Understanding

DatasetYearPaperProjectCode
CMU-MMAC2009Guide to the carnegie mellon university multimodal activity (cmu-mmac) databaseProject-
Charades-Ego2018Actor and observer: Joint modeling of first and third-person videosProjectGithub
H2O2021H2o: Two hands manipulating objects for first person interaction recognitionProjectGithub
Assembly1012022Assembly101: A large-scale multi-view video dataset for understanding procedural activitiesProjectGithub
Homage2021Home action genome: Cooperative compositional action understandingProjectGithub
LEMMA2020Lemma: A multiview dataset for le arning m ulti-agent m ulti-task a ctivitiesProjectGithub
FT-HID2022Ft-hid: a large-scale rgb-d dataset for first-and third-person human interaction analysis-Github
EgoPW2022Estimating egocentric 3d human pose in the wild with external weak supervisionProjectGithub
First2Third-Pose2022Enhancing egocentric 3d pose estimation with third person views-GitHub
ARCTIC2023ARCTIC: A dataset for dexterous bimanual hand-object manipulationProjectGithub
ECHP2023Egofish3d: Egocentric 3d pose estimation from a fisheye camera via selfsupervised learning-Github
AssemblyHands2023Assemblyhands: Towards egocentric activity understanding via 3d hand pose estimationProjectGithub
EgoHumans2023Ego-humans: An ego-centric 3d multi-human benchmarkProjectGithub
-2023Next-generation surgical navigation: Marker-less multi-view 6dof pose estimation of surgical instrumentsProject-
ThermoHands2024Thermohands: A benchmark for 3d hand pose estimation from egocentric thermal imageProjectGithub
OVR2024Ovr: A dataset for open vocabulary temporal repetition counting in videos-Github
Nymeria2024Nymeria: A massive collection of multimodal egocentric daily motion in the wildProjectGithub
OAKINK22024Oakink2: A dataset of bimanual hands-object manipulation in complex task completionProjectGithub
CORE4D2024Core4d: A 4d humanobject-human interaction dataset for collaborative object rearrangementProjectGithub
Ego-Exo4D2024Ego-exo4d: Understanding skilled human activity from first-and third-person perspectivesProjectGithub
EgoExoLearn2024Egoexolearn: A dataset for bridging asynchronous ego-and exo-centric view of procedural activities in real world-Github
EgoExo-Fitness2024Egoexo-fitness: Towards egocentric and exocentric full-body action understanding-Github
EgoExOR2025EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding-Github
EPFL-Smart-Kitchen-302025EPFL-Smart-Kitchen-30: Densely annotated cooking dataset with 3D kinematics to challenge video and language models-Github

Driving

Affordance Grounding

Generation

Scene Understanding

Video Question Answering

Egocentric Wearer Identification

Cross-View Human Tracking and Association

Camera Registration