A curated list of papers and datasets on Egocentric (first-person) and Exocentric (third-person) Vision. This repo complements our survey and is continuously updated.
| Title | Venue | Year | Code |
|---|---|---|---|
| Intention-driven Ego-to-Exo Video Generation | arxiv | 2024 | - |
| Ego-to-Exo: Interfacing Third Person Visuals from Egocentric Views in Real-time for Improved ROV Teleoperation | arxiv | 2024 | - |
| Title | Venue | Year | Code |
|---|---|---|---|
| From my view to yours: Egoaugmented learning in large vision language models for understanding exocentric daily living activities | arxiv | 2025 | Github |
| Title | Venue | Year | Code |
|---|---|---|---|
| Camera selection for occlusion-less surgery recording via training with an egocentric camera | Access | 2021 | - |
| Title | Venue | Year | Code |
|---|---|---|---|
| Exo2egodvc: Dense video captioning of egocentric procedural activities using web instructional videos | WACV | 2023 | Github |
| Retrieval-augmented egocentric video captioning | CVPR | 2024 | Github |
| Title | Venue | Year | Code |
|---|---|---|---|
| Spatial assisted human-drone collaborative navigation and interaction through immersive mixed reality | ICRA | 2024 | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| ThirdtoFirst | 2021 | Ego-exo: Transferring visual representations from third-person to first-person videos | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| 360 + x | 2024 | 360 + x: A panoptic multi-modal scene understanding dataset | Project | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| GazeVQA | 2023 | Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| TF2023 | 2024 | Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| CVMHT | 2020 | Complementary-view multiple human tracking | - | GitHub |
| DMHA | 2022 | Connecting the complementary-view videos: Joint camera identification and subject association | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| - | 2024 | From a bird’s eye view to see: Joint camera and subject registration without the camera calibration | - | Github |
A curated list of papers and datasets on Egocentric (first-person) and Exocentric (third-person) Vision. This repo complements our survey and is continuously updated.
| Title | Venue | Year | Code |
|---|---|---|---|
| Intention-driven Ego-to-Exo Video Generation | arxiv | 2024 | - |
| Ego-to-Exo: Interfacing Third Person Visuals from Egocentric Views in Real-time for Improved ROV Teleoperation | arxiv | 2024 | - |
| Title | Venue | Year | Code |
|---|---|---|---|
| From my view to yours: Egoaugmented learning in large vision language models for understanding exocentric daily living activities | arxiv | 2025 | Github |
| Title | Venue | Year | Code |
|---|---|---|---|
| Camera selection for occlusion-less surgery recording via training with an egocentric camera | Access | 2021 | - |
| Title | Venue | Year | Code |
|---|---|---|---|
| Exo2egodvc: Dense video captioning of egocentric procedural activities using web instructional videos | WACV | 2023 | Github |
| Retrieval-augmented egocentric video captioning | CVPR | 2024 | Github |
| Title | Venue | Year | Code |
|---|---|---|---|
| Spatial assisted human-drone collaborative navigation and interaction through immersive mixed reality | ICRA | 2024 | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| ThirdtoFirst | 2021 | Ego-exo: Transferring visual representations from third-person to first-person videos | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| 360 + x | 2024 | 360 + x: A panoptic multi-modal scene understanding dataset | Project | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| GazeVQA | 2023 | Gazevqa: A video question answering dataset for multiview eye-gaze task-oriented collaborations | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| TF2023 | 2024 | Fusing personal and environmental cues for identification and segmentation of first-person camera wearers in third-person views | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| CVMHT | 2020 | Complementary-view multiple human tracking | - | GitHub |
| DMHA | 2022 | Connecting the complementary-view videos: Joint camera identification and subject association | - | Github |
| Dataset | Year | Paper | Project | Code |
|---|---|---|---|---|
| - | 2024 | From a bird’s eye view to see: Joint camera and subject registration without the camera calibration | - | Github |