player0718/awesome-ego-video-datasets

πŸŽ₯ [Awesome] Egocentric / First-Person Video Datasets πŸ“š Papers, Benchmarks & Resources for Ego Vision

225

38 commits

updated Sep 21, 2026

See the code

README

πŸŽ₯ Awesome Egocentric Video Datasets

πŸ“œ A Curated List of Egocentric (First-Person) Video Datasets, Benchmarks, and Tools

Awesome Egocentric Video Datasets

Awesome Papers PRs Welcome License: CC0-1.0

Overview

Overview of egocentric video datasets

This repository tracks egocentric video datasets through a task-first view: every dataset appears once as a primary entry under one of seven research themes, with cross-links where it also matters. The goal is fast navigation for researchers who need scale, task fit, benchmark context, and official resources without bouncing across multiple index files.

Papers & Surveys

Sorted newest to oldest, with flagship surveys and corpus papers highlighted first.

  • Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI (2026) β€” Survey of egocentric VLMs spanning datasets, hand-object interaction, temporal reasoning, multimodal learning, wearable assistance, and human-to-robot transfer. arXiv

  • Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable (2026) β€” Position paper arguing that privacy leakage in always-on wearable video should be evaluated across the full data and model pipeline with standardized metrics and benchmarks. arXiv

  • Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges (2025) β€” Li et al., 2025 survey and benchmark paper of egocentric procedural activity understanding, focus on building egocentric procedural AI assistant. arXiv

  • [⭐️] Challenges and Trends in Egocentric Vision: A Survey (2025) β€” Li et al., 2025 survey of datasets, tasks, benchmarks, and open challenges in egocentric vision. arXiv

  • Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision (2025) β€” Cross-view survey covering ego-exo collaboration, paired capture, and collaborative perception. arXiv

  • HD-EPIC: A Highly-Detailed Egocentric Video Dataset (2025) β€” Dataset paper introducing fine-grained kitchen understanding with dense multimodal annotations. arXiv

  • Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives (2024) β€” Large-scale paired ego-exo dataset paper spanning skilled activity and multiview understanding. arXiv

  • EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World (2024) β€” Procedural ego-exo paper focused on asynchronous activity alignment in real environments. Paper

  • EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding (2023) β€” Benchmark paper targeting long-form memory and reasoning over egocentric video. arXiv

  • [⭐️] Ego4D: Around the World in 3,000 Hours of Egocentric Video (2022) β€” Flagship corpus paper introducing the large-scale Ego4D benchmark suite. arXiv

  • [⭐️] Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100 (2021) β€” Canonical kitchen benchmark paper for action recognition, detection, and anticipation. Paper

  • Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos (2018) β€” Early paired ego-exo benchmark for activity transfer and alignment. arXiv

🎬 Video Generation & World-Model Pretraining

Datasets here emphasize large-scale first-person pretraining corpora, video generation, editing, or world-model supervision from ego video.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ropedia Xperience-10M2026Large multi-stream experiencesMultimodal ego learningN/AHugging Face
WorldRover-10M20266,003 seq. / 21.9M frames / 202.7 hFirst-person world models, 3D explorationPaperN/A
H2R-Bench20266 manipulation families / 2 robot embodimentsHuman-to-robot video generationPaperN/A
Ego-OSCAR-550h2026~550 h/camera / 1,462 stereo sessionsStereo-inertial ego pretraining, capturePaperN/A
Ego2Robot202618,561 h / 15 robot morphologiesEgo-to-robot data synthesis, VLA pretrainingPaperSite
ACE-Data-02026150 h / 75K episodes / 200 tasksMultimodal embodied pretrainingPaperSite
EgoPlay2026106K event-triggered clip-prompt pairsEvent-triggered ego video editingPaperN/A
Open-AoE2026~2,000 h / 500+ contributorsManipulation pretraining, data toolchainPaperGitHub
EgoVid-Pro2026103K clips / ~12M framesHand-controlled ego video generationPaperN/A
RetailSMV202632,105 clips / 16.1K ego + 16.0K exoRetail world-model adaptationPaperSite
EgoCS-400K2026400K+ videos / 10K h gameplayAction-conditioned world modelsPaperN/A
WM-H (Wh0)202650K generated HOI episodesSynthetic dexterous VLA dataPaperSite
DreamDojo-HV2026Very large FP video (see paper)World models, pretrainingPaperN/A
Ego-1K2026Multiview clips (~1K takes)Neural 3D/4D synthesisPaperHugging Face
In-lab2026Lab tabletop trajectoriesSkills, world models (w/ DreamDojo)PaperN/A
HumanNet2026~1M h human-centric video (ego + exo)VLA / embodied pretrainingPaperN/A
MobileEgo Anywhere2026200 h smartphone-collected long-horizon egoLong-horizon ego data infra, VLAPaperN/A
EgoEdit2025100K editing pairsEgocentric video editingPaperSite
EgoVid-5M20245M clipsVideo generation, motion+textPaperSite

Entries

  • [⭐️] Ropedia Xperience-10M (2026) β€” 10M multimodal experiences with 6 RGB streams, stereo depth, pose/SLAM, hand-body mocap, audio, and IMU for large-scale ego pretraining. Site Code πŸ€—

  • WorldRover-10M (2026) β€” 6,003 synthetic exploration sequences from 32 environments (21.9M frames / 202.7 h, including 10.8M first-person frames) with first-person, third-person, and 360Β° views aligned to metric depth, trajectories, geometry, and action signals. arXiv

  • H2R-Bench (2026) β€” Cross-embodiment benchmark for transforming egocentric human demonstrations into robot manipulation videos, evaluated across six manipulation families, two target embodiments, and five dimensions covering execution, contact, embodiment, and visual quality. arXiv

  • Ego-OSCAR-550h (2026) β€” ~550 h per camera (1,462 stereo sessions) of calibrated everyday egocentric video with synchronized IMU, dense open-vocabulary action captions, per-frame 3D hand reconstructions, and an open sub-$200 capture stack. arXiv

  • Ego2Robot (2026) β€” 18,561 h of synthetic robot training data across 15 morphologies, generated from curated and in-the-wild egocentric manipulation video through action retargeting, robot-arm compositing, and multi-level quality curation. arXiv Site

  • ACE-Data-0 (2026) β€” 150 h and 75K interaction episodes across 200 household task categories, synchronizing ego/exo video, full-body and hand motion, object state, audio, and tactile signals at table and room scale. arXiv Site πŸ€—

  • EgoPlay (2026) β€” 106K event-triggered egocentric clip-prompt pairs, primarily derived from Ego4D, covering positive, fabricated-negative, and multi-event triggers for temporally restrained video editing and streaming evaluation. arXiv

  • Open-AoE (2026) β€” ~2,000 h of smartphone-collected manipulation video from 500+ contributors with bilingual text, MANO hand pose, camera trajectory, and atomic-action annotations, plus capture-to-training tools for VLA and world-model research. arXiv Code πŸ€—

  • EgoVid-Pro (2026) β€” 103K in-the-wild egocentric clips (~12M frames) with clean protagonist-only 3D hand trajectories, curated for HandsOnWorld and Plucker Hand Map conditioning in hand-controlled first-person video generation. arXiv

  • RetailSMV (2026) β€” 32,105 captioned retail clips from five supermarkets with synchronized staff-view egocentric and exocentric capture, predefined train/val/test splits, and a held-out protocol for video-world-model adaptation. arXiv Site

  • EgoCS-400K (2026) β€” 400K+ replay-grounded first-person Counter-Strike videos (10K h) aligned with player states, view directions, movements, keyboard/button inputs, events, and round context for action-conditioned rollout, captioning, and world-model training. arXiv

  • WM-H (Wh0) (2026) β€” 50K world-model-generated egocentric human-object manipulation episodes conditioned on language, objects, and scenes, then converted into robot-trainable supervision for dexterous VLA adaptation. arXiv Site Code

  • DreamDojo-HV (2026) β€” Very large FP video (see paper); World models, pretraining. arXiv Site

  • Ego-1K (2026) β€” Multiview clips (~1K takes); Neural 3D/4D synthesis. arXiv Site πŸ€—

  • In-lab (2026) β€” Lab tabletop trajectories; Skills, world models (w/ DreamDojo). arXiv Site

  • HumanNet (2026) β€” ~1M h of human-centric video (mix of ego and exo) with interaction-centric annotations; the authors report 1k h of ego human video outperforms 100 h of real-robot data for VLA training. arXiv

  • MobileEgo Anywhere (2026) β€” 200 h of hour-plus egocentric trajectories collected on commodity smartphones, released with an open-source mobile capture app and a processing pipeline aimed at VLA pretraining. arXiv

  • EgoEdit (2025) β€” 100K editing pairs; Egocentric video editing. arXiv Site

  • EgoVid-5M (2024) β€” 5M first-person clips curated for text-and-motion-conditioned video generation from wearable footage. arXiv Site Code

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
H2R-BenchCross-embodiment human-to-robot manipulation video generationH2R-BenchPaperStandalone
EgoPlayEvent-triggered editing, pre-trigger preservation, false-trigger robustnessEgoPlay / Ego4DPaperDataset+benchmark
ACE-Data-0Hierarchical signals-to-scenes-to-interactions evaluationACE-Data-0SiteDataset+benchmark
Ego2Robot / RoboTwin2.0 extensionOOD visual, spatial, embodiment, and semantic generalizationEgo2RobotSiteDataset+benchmark
HandsOnWorld / EgoVid-ProCamera-disentangled hand-controlled egocentric video generationEgoVid-ProPaperDataset+benchmark
RetailSMVRetail video-world-model adaptation and ego/exo viewpoint ablationsRetailSMVSiteDataset+benchmark
EgoCS-400KAction-conditioned future prediction, state/event-aware rollout, replay-grounded captioningEgoCS-400KPaperDataset+benchmark
Wh0 / WM-HSynthetic egocentric dexterous manipulation data for VLA adaptationWM-HSiteDataset+training resource
EgoEditEgocentric video editingEgoEdit / EgoEditDataProjectDataset+benchmark

🧠 Memory, Summarization & Long-form Understanding

This section collects long-horizon lifelog, summarization, and persistent-memory datasets where temporal continuity matters as much as recognition.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ EgoLife2025~266–300 h daily lifeLong-form assistants, memoryPaperSite
⭐ EgoSchema2023250+ h / 5K QALong-form video QAPaperSite
EgoMonth2026301 h / 738 clips / 1,443 QAMonth-level spatiotemporal memoryPaperHF
MEMORA-Bench202645 h / 18 participantsEmbodied action memory, planningPaperN/A
EgoServe20263K+ service instances / 4 horizonsProactive continuous-video assistancePaperSite
SuperMemory-VQA202652.9 h / 4,853 QALong-horizon memory VQAPaperHF
EgoMemReason2026500 MCQs over EgoLifeWeek-long memory reasoningPaperHF
EgoExoMem20262.6K MCQs / 390 videosCross-view memory reasoningPaperGitHub
EgoIntrospect2026180 h / 60 subjectsInternal-state reasoning, memoryPaperSite
MA-EgoQA20261,741 QA / 6 agents / 7 daysMulti-agent egocentric QAPaperSite
EgoStream20262,250 Q / 8,528 evals, streams up to 45.3 hStreaming episodic memoryPaperSite
VidChapters-7M2023817K videos / 7M chaptersChaptering (not ego-only)PaperSite
Multi-Ego2022~12 h / 41 seq.Multi-wearer, summarizationPaperGitHub
DoMSEV201880 h, 48 seq.Semantic fast-forward, first-person videoPaperSite
HUJI-EgoSeg201429 long egocentric videos (~1–5 h each), pixel-level temporal segmentation annotationsmemory, summarization & long-form understandingPaperSite
UT Ego2012~17 h, 4 long videosSummarization, long-form egoPaperSite
VINST / Visual Diaries201131 egocentric videos capturing daily commutes; used for temporal segmentation and video summarizationmemory, summarization & long-form understandingPaperSite

Entries

  • [⭐️] EgoLife (2025) β€” ~266-300 h of daily-life capture in EgoHouse with Meta Aria, third-person cameras, and mmWave sensors for persistent assistant memory. arXiv Site Code πŸ€—

  • [⭐️] EgoSchema (2023) β€” 250+ h of long-form Ego4D video with 5K QA pairs designed to probe memory and causal understanding over extended clips. arXiv Site Code

  • EgoMonth (2026) β€” 301 h across 738 wearable-camera clips from 20 participants recorded over 20–120 days, paired with 1,443 human-authored questions spanning schema consolidation, episodic indexing, and cascading reasoning. arXiv πŸ€—

  • MEMORA-Bench (2026) β€” 45 h of EPIC-KITCHENS-100 extension video from 18 participants for memory-grounded planning toward seen and unseen goals, plus structured assessment of environment, entity, activity, and inferred-knowledge memory. arXiv

  • EgoServe (2026) β€” 3K+ manually verified proactive-service instances over EgoLife, HoloAssist, and CaptainCook4D, organized into 10 assistance categories and four temporal horizons from instant alerts to multi-day habit coaching. arXiv Site Code πŸ€—

  • SuperMemory-VQA (2026) β€” 52.9 h of everyday Meta Aria recordings with RGB, processed gaze, IMU, SLAM trajectories, point clouds, redacted transcripts, and 4,853 human-verified long-horizon memory QA pairs. arXiv πŸ€—

  • EgoMemReason (2026) β€” 500 multiple-choice questions over week-long EgoLife video for entity, event, and behavior memory reasoning, with public questions and leaderboard evaluation. arXiv Site Code πŸ€—

  • EgoExoMem (2026) β€” 2.6K human-verified MCQs over 390 synchronized egocentric and exocentric videos from EgoExo4D and LEMMA for cross-view memory reasoning. arXiv Code

  • EgoIntrospect (2026) β€” 180 h of user-driven egocentric recordings from 60 subjects with synchronized video, audio, gaze, motion, and physiological signals for affective experience, request intent, and cognitive-memory reasoning; the paper states data will be made public. arXiv Site

  • MA-EgoQA (2026) β€” 1,741 QA pairs over six temporally aligned EgoLife egocentric streams spanning seven days, targeting multi-agent social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. arXiv Site Code

  • EgoStream (2026) β€” 2,250 curated questions expanded to 8,528 recall-conditioned evaluations via Answer Validity Windows over egocentric streams up to 45.3 h (curated from Ego4D, EgoLife, EgoTempo, Multi-Hop EgoQA, and HD-EPIC), spanning seven cognitive memory dimensions for diagnosing streaming episodic memory in video-language models. arXiv Site

  • VidChapters-7M (2023) β€” 817K videos / 7M chapters; Chaptering (not ego-only). arXiv Site Code

  • Multi-Ego (2022) β€” ~12 h / 41 seq; Multi-wearer, summarization. Paper arXiv Code

  • DoMSEV (2018) β€” 80 h, 48 seq; Semantic fast-forward, first-person video. Paper Site

  • HUJI-EgoSeg (2014) β€” 29 long egocentric videos (~1–5 h each), pixel-level temporal segmentation annotations; memory, summarization & long-form understanding. Paper Site

  • UT Ego (2012) β€” ~17 h, 4 long videos; Summarization, long-form ego. Paper Project

  • VINST / Visual Diaries (2011) β€” 31 egocentric videos capturing daily commutes; used for temporal segmentation and video summarization; memory, summarization & long-form understanding. Paper Site

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoMonthMonth-level schema, episodic, spatial, and cross-day memory reasoningEgoMonthHFDataset+benchmark
MEMORA-BenchEmbodied action memory formation, consolidation, retrieval, and planningEPIC-KITCHENS-100 extensionPaperStandalone
EgoServeProactive assistance over instant, short-term, episodic, and long-term contextEgoLife / HoloAssist / CaptainCook4DSiteStandalone
SuperMemory-VQALong-horizon egocentric memory VQA with answerability checksSuperMemory-VQAHFDataset+benchmark
EgoMemReasonWeek-long entity, event, and behavior memory reasoningEgoLifeHFStandalone
EgoExoMemCross-view memory reasoning over synchronized ego-exo videosEgoExo4D / LEMMAGitHubStandalone
EgoIntrospectUser internal-state reasoning over multimodal egocentric streamsEgoIntrospectSiteDataset+benchmark
MA-EgoQAMulti-agent egocentric video QA over week-long streamsEgoLifeSiteStandalone
EgoSchemaLong-form video-language understandingEgo4DSiteStandalone
EgoStreamStreaming episodic memory diagnosisEgo4D / EgoLife / EgoTempo / Multi-Hop EgoQA / HD-EPICSiteStandalone

πŸ’¬ VLMs, Instructions & QA

These datasets target egocentric video-language pretraining, instruction following, dialog, and question answering over first-person streams.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ HD-EPIC2025~41 h / dense labelsFine-grained kitchen, VQAPaperSite
⭐ EgoClip20223.8M clip–text pairsVideo-language pretrainingPaperGitHub
CrossView2026~6K questions / 4 domainsMulti-camera video QAPaperSite
EgoCross2026798 clips / 957 QACross-domain egocentric VideoQAPaperSite
HumanCLAW-Bench20261,218 episodes / 41 scenesClosed-loop embodied action intelligencePaperSite
EgoSafe-Bench20263K clips / 12K evaluation samplesVisual safety, forensic reasoningPaperN/A
VIABench2026761 videos / 46.9 h / 14.5K annotationsVisual-impairment assistancePaperGitHub
LongEgoRefer20261,498 refs / avg. 45 minLong-form video REC, groundingPaperGitHub
EgoGapBench20261,000 action-selection itemsEgocentric action selectionPaperGitHub
EgoSafetyBench20261,200 robot-view scenariosStreaming safety guardsPaperN/A
EgoSAT20261,997 videos / 165 h / 4.8K QAStreaming interaction understandingPaperSite
EgoTL2026100+ daily household tasksLong-horizon reasoning, spatial QAPaperSite
MM-Conv20266.7 h / 4,211 expressionsContext-aware 3D dialogue groundingPaperN/A
Minerva-Ego20261,160 QA / 156 videosSpatiotemporal reasoning tracesPaperGitHub
EgoEverything2026100+ h / 5K+ MCQsLong-context AR VideoQAPaperN/A
EgoEsportsQA20261,745 QA pairsEsports VideoQA, reasoningPaperN/A
GameplayQA2026~2.4K QA / 15 task categoriesPOV-synced multi-video gameplay QAPaperSite
MyEgo2026541 long videos / 5K QAPersonalized VideoQA, ego-groundingPaperGitHub
LifeDialBench / EgoMem2026EgoMem + LifeMem (coming soon)Lifelog memory, online evalPaperGitHub
Ego2Web2026500 video-instruction pairsEgocentric video-grounded web agentsPaperHF
NoRA20261,420 clips / support graphsNormative action reasoningPaperN/A
Causal-Plan-1M20261M QA / 22.2K clips / 770+ hPhysically grounded planning QAPaperHF
Pause-and-Think202610K QA clips + 300-sample benchmarkAssistive action suggestionPaperGitHub
EgoCoT-Bench2026351 videos / 3,172 QAGrounded operation-centric CoT QAPaperSite
EgoEMS202520+ h emergency scenariosEMS QA, multimodalPaperGitHub
HowToDIV2025~24 h instructionalDialog, procedural QAPaperGitHub
InterVLA202511.4 h interactionsInstruction, ego–exo mocapPaperSite
AssistQ2022100 long videos / 529 QAInstructional QAPaperGitHub
EgoTaskQA2022~2K videos / 40K QACausal & task QAPaperSite
EgoVQA2019600+ QAsVideo QAPaperN/A

Entries

  • [⭐️] HD-EPIC (2025) β€” ~41 h of densely labeled cooking video with recipe steps, audio events, gaze, 3D grounding, and VQA supervision. arXiv Site Code

  • [⭐️] EgoClip (2022) β€” 3.8M clip–text pairs; Video-language pretraining. arXiv Code

  • CrossView (2026) β€” ~6K multi-camera video questions across autonomous driving, surveillance, ego/exo activity, and robotics, including 4–7 synchronized Ego-Exo4D cameras per egocentric question. arXiv Site Code πŸ€—

  • EgoCross (2026) β€” 798 clips and 957 human-verified QA pairs across surgery, industrial assembly, extreme sports, and animal perspectives, designed to test cross-domain generalization beyond daily-life egocentric video. arXiv Site πŸ€—

  • HumanCLAW-Bench (2026) β€” 1,218 long-horizon find–navigate–interact episodes across 41 simulated indoor scenes for evaluating whether VLMs can select and sequence actions from a continuously updated egocentric body view. arXiv Site Code

  • EgoSafe-Bench (2026) β€” 12K evaluation samples formed from 3K mobile-captured first-person clips and hierarchical QA chains for feature anchoring, blind-spot deduction, intent inference, and logically consistent visual-safety reasoning. arXiv

  • VIABench (2026) β€” 761 first-person videos (46.9 h) recorded or shared by blind individuals with 14,526 curated annotations for proactive reminders, visual question answering, and vision-guided interaction in online and offline settings. arXiv Code

  • LongEgoRefer (2026) β€” 1,498 referring expressions over long-form Ego4D videos averaging 45 minutes, requiring temporal and spatial localization of sparse referred objects in untrimmed egocentric recordings. arXiv Code

  • EgoGapBench (2026) β€” 1,000 egocentric action-selection items from multi-agent scenes without first-person body cues, designed to test whether VLMs choose actions from the correct self/other perspective. arXiv Code

  • EgoSafetyBench (2026) β€” 1,200 egocentric robot-view scenarios annotated at half-second granularity, split into situational and visual-channel tracks for evaluating runtime VLM safety guards. arXiv

  • EgoSAT (2026) β€” 1,997 Ego4D videos (165 h) with ~4.8K QA pairs for streaming egocentric interaction understanding, covering retrospective, present, and prospective reasoning under partial observability. arXiv Site Code πŸ€—

  • EgoTL (2026) β€” 100+ daily household tasks with think-aloud chains, navigation and manipulation annotations, and metric spatial labels for long-horizon egocentric reasoning and QA. arXiv Site

  • MM-Conv (2026) β€” 6.7 h of egocentric VR interaction with synchronized speech, motion, gaze, 3D scene geometry, and 4,211 verified referring expressions for context-aware conversational grounding; full public release is described as following camera-ready. arXiv

  • Minerva-Ego (2026) β€” 1,160 hand-crafted multiple-choice questions over 156 HD-EPIC egocentric videos, paired with dense spatiotemporal human reasoning traces and object masks for diagnosable video reasoning. arXiv Code

  • EgoEverything (2026) β€” 100+ h of AR egocentric video with 5K+ multiple-choice questions for human-behavior-inspired long-context video understanding. arXiv

  • EgoEsportsQA (2026) β€” 1,745 expert QA pairs from first-person shooter esports videos for testing fast virtual-scene perception and tactical reasoning. arXiv

  • GameplayQA (2026) β€” ~2.4K QA pairs across 15 task categories and 3 cognitive levels for decision-dense, POV-synced multi-video understanding of multi-agent 3D gameplay, with densely labeled first-person player viewpoints (1.22 labels/sec); egocentric viewpoints are virtual/synthetic. ACL 2026. arXiv Site Code πŸ€—

  • MyEgo (2026) β€” 541 long egocentric videos with 5K personalized questions about the camera wearer, their activities, and past context. arXiv Code

  • LifeDialBench / EgoMem (2026) β€” Lifelog memory benchmark with EgoMem built from real-world egocentric videos and an online evaluation protocol; the project page currently marks the dataset and scripts as coming soon. arXiv Code

  • Ego2Web (2026) β€” 500 egocentric video-instruction pairs for evaluating web agents that must ground real-world first-person visual context before executing online tasks. arXiv Site Code πŸ€—

  • NoRA (2026) β€” 1,420 Ego4D-derived clips (190 human-verified HumanGold + 1,230 LLM-validated LLMSilver) with fact–reason–action support-graph annotations for benchmarking grounded normative action reasoning in VLMs. arXiv

  • Causal-Plan-1M (2026) β€” 1M QA pairs with four-stage causal reasoning-trace annotations over 22,201 egocentric clips (770+ h curated from Ego4D, EPIC-KITCHENS, HoloAssist, MECCANO, and others), plus the 1,200-instance Causal-Plan-Bench for physically grounded embodied planning. arXiv Code πŸ€—

  • Pause-and-Think (2026) β€” 10,051 reasoning-annotated training QA clips and a 300-sample benchmark curated from EPIC-KITCHENS, Ego4D, and Assembly101 with structured thinking/answer supervision for video-grounded assistive action suggestion and goal planning. arXiv Code

  • EgoCoT-Bench (2026) β€” 3,172 verifiable QA pairs over 351 first-person videos (sourced from Ego4D, EPIC-KITCHENS, Charades-Ego, MECCANO, and HD-EPIC) with step-by-step rationale annotations for evaluating grounded operation-centric chain-of-thought reasoning in MLLMs. arXiv Site Code πŸ€—

  • EgoEMS (2025) β€” 20+ h emergency scenarios; EMS QA, multimodal. arXiv Site Code

  • HowToDIV (2025) β€” ~24 h instructional; Dialog, procedural QA. arXiv Site Code

  • InterVLA (2025) β€” 11.4 h interactions; Instruction, ego–exo mocap. arXiv Site

  • AssistQ (2022) β€” 100 long videos / 529 QA; Instructional QA. arXiv Code

  • EgoTaskQA (2022) β€” ~2K videos / 40K QA; Causal & task QA. arXiv Site Code

  • EgoVQA (2019) β€” 600+ QAs; Video QA. Paper

  • EgoSchema β€” see Memory, Summarization & Long-form Understanding

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
CrossViewMulti-camera evidence integration across four domainsEgo-Exo4D / nuScenes / MEVA / AgiBotSiteStandalone
EgoCrossCross-domain egocentric VideoQA in surgery, industry, sports, and animal viewsEgoCrossSiteDataset+challenge
HumanCLAW-BenchClosed-loop find–navigate–interact action intelligenceHSSD simulated scenesSiteStandalone
EgoSafe-BenchHierarchical visual-safety and forensic reasoningEgoSafe-BenchPaperDataset+benchmark
VIABenchProactive reminder, VQA, and vision-guided assistanceVIABenchGitHubDataset+benchmark
LongEgoReferLong-form egocentric spatiotemporal referring expression comprehensionEgo4DGitHubStandalone
EgoGapBenchEgocentric action selection in multi-agent scenesEgoGapBenchGitHubStandalone
EgoSafetyBenchRuntime safety-guard evaluation over egocentric robot-view videoEgoSafetyBenchPaperStandalone
EgoSATRetrospective, present, and prospective streaming VideoQAEgo4DSiteStandalone
EgoTL-BenchLong-horizon planning, action reasoning, perceptual-metric understandingEgoTLSiteDataset+benchmark
MM-ConvContext-aware grounding in spontaneous 3D dialogueMM-ConvPaperDataset+benchmark
Minerva-EgoEgocentric multi-step VideoQA with spatiotemporal reasoning tracesHD-EPICGitHubStandalone
EgoEverythingLong-context AR VideoQAEgoEverythingPaperStandalone
EgoEsportsQAEsports VideoQA and tactical reasoningEgoEsportsQAPaperStandalone
GameplayQADecision-dense POV-synced multi-video gameplay QAGameplayQASiteDataset+benchmark
MyEgoPersonalized egocentric VideoQAMyEgoGitHubDataset+benchmark
LifeDialBench / EgoMemLifelog memory and online evaluationEgoMemGitHubDataset+benchmark
Ego2WebEgocentric video-grounded web-agent executionEgo2WebSiteDataset+benchmark
EgoExoBenchCross-perspective video understanding in MLLMsEgo-Exo4D / LEMMA / EgoExoLearn / TF2023 / EgoMe / CVMHATGitHubStandalone
EgoBabyVLM / Machine-DevBenchCross-modal learning from naturalistic egocentric videoNaturalistic infant/adult ego video corporaPaperStandalone
EgoEMSEMS QA, multimodal assessmentEgoEMSGitHubDataset+benchmark
HD-EPICFine-grained kitchen understanding, VQAHD-EPICSiteDataset+benchmark
HowToDIVMulti-turn instructional dialog & QAHowToDIVGitHubStandalone
InterVLAInstruction following, interaction understandingInterVLASiteDataset+benchmark
AssistQInstructional affordance-centric QAAssistQGitHubStandalone
EgoTaskQACausal, predictive, explanatory, counterfactual QAEgoTaskQASiteStandalone
EgoVQAEgocentric video QAEgoVQAOpen Access (ICCVW 2019 / EPIC)Standalone
NoRAGrounded normative action reasoningEgo4DPaperStandalone
Causal-Plan-BenchPhysically grounded embodied planning QACausal-Plan-1MHFDataset+benchmark
Pause-and-ThinkVideo-grounded assistive action suggestionEPIC-KITCHENS / Ego4D / Assembly101GitHubDataset+benchmark
EgoCoT-BenchGrounded operation-centric chain-of-thought QAEgo4D / EPIC-KITCHENS / Charades-Ego / MECCANO / HD-EPICSiteStandalone

πŸƒ Action & Activity Recognition

Canonical action-recognition, activity-analysis, affect, and interaction datasets built around first-person human behavior live here.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ego4D2022~3,670 hAR, VQA, forecasting, manyPaperSite
⭐ EPIC-KITCHENS-1002021100 h / 90K segmentsAction recognition, manyPaperSite
HUI360202671 h captured / 11 h filtered / 1M annotationsHuman-robot interaction anticipationPaperSite
EventKitchen20265.5 h / 10.8K action segmentsEvent-based action recognition, detectionPaperSite
InterPet4D20266.8M frames / 13 dogs / 23 peopleHuman-pet interaction, motion generationPaperSite
EgoPolice2026180+ h / 9 action classesBody-camera action recognition, VideoQAPaperGitHub
ChildLens2026108.58 h / 354 videosChild activity analysisPaperData
EgoScale202620k+ h labeled manipulation (paper)Action, dexterous transferPaperN/A
CogDrive (EyeCue)2026Multi-scenario driving clipsDriver cognitive-distraction detection, gaze + egoPaperN/A
Ego-METAS2026100+ h / 5 modalitiesOnline action segmentation, sensor routingPaperHF
Furhat Egocentric Dataset202620 seq. / ~25 minRobot-ego face/body tracking, re-IDPaperN/A
World In Your Hands20251000+ h labeled manipulation (paper)Action, dexterous transfer, VLA trainingPaperGitHub
EgoCampus2025~32 h / campus paths (paper)Gaze, pedestrian egoPaperGitHub
AEA2024143 seq. / ~7.3 hEveryday activities, AriaPaperSite
EgoSurgery (Phase / Tool / HTS)2024Open-surgery ego videoPhase, tools, segmentationPhase / Tool / HTSGitHub
EgoExo-Fitness202432 h / 1,276 seq.Full-body action, quality assessmentPaperGitHub
EΒ³ (Exploring Embodied Emotion)202450+ hEmotion, multimodal egoPaperGitHub
EGOFALLS2023Fall samples / AVFall detectionPaperSite
Epic-Sounding-Object20233.2K short clipsAudio-visual localizationPaperGitHub
HoloAssist2023169 hInteractive assistantsPaperSite
WEAR2023~19 h outdoor sportsActivity + IMUPaperSite
N-EPIC-KITCHENS2022Event + RGB subsetAction, neuromorphicPaperGitHub
Ego-Deliver20215,360 videosDelivery ego analysisN/ASite
HOMAGE202130 h / compositionalHome activitiesPaperSite
MECCANO2021~55 h industrialHOI, egoPaperSite
EGO-CH202027+ h, cultural sitesVisitor behavior, POI tasksPaperN/A
EgoCom202038.5 h conversationMultiperson ego dialogPaperGitHub
LEMMA2020Multi-view activitiesMulti-agent tasksPaperSite
Charades-Ego2018Paired ego / exoAlignment, actionsPaperSite
EGTEA Gaze+201828 h cookingGaze + action recognitionPaperSite
EgoGesture20172K+ videos / 24K samplesGesture recognitionPaperSite
Stanford ECM201731 hours, augmented with heart rate and accelerometer, 23–24 daily activity categoriesaction & activity recognitionPaperN/A
THU-READ20171,920 clips (8 subjects Γ— 40 actions Γ— 3 reps Γ— 2 modalities), RGB-D from helmet-mounted sensoraction & activity recognitionPaperSite
PEV (UTokyo Paired Ego-Video)20161,226 pairs of first-person clips, synchronous dyadic conversations, 8 interaction categories, 6 subjectsaction & activity recognitionPaperN/A
FPPA20155 subjects, 5 daily actions, egocentric video with hand and gaze cuesaction & activity recognitionPaperN/A
JPL-Interaction201384 videos, 7 activity types (4 positive, 1 neutral, 2 negative interactions), 320Γ—240@30 fpsaction & activity recognitionPaperSite
ADL2012~10 h, 20 participantsADL recognition, objectsPaperSite
Social Interactions20128 social events, ~60 hours, head-mounted cameras, multiple participants per eventaction & activity recognitionPaperSite
EgoAction2011First-person sports videos (skateboarding, skiing, cycling, etcaction & activity recognitionPaperN/A

Entries

  • [⭐️] Ego4D (2022) β€” ~3,670 h; AR, VQA, forecasting, many. arXiv Site Code

  • [⭐️] EPIC-KITCHENS-100 (2021) β€” 100 h of unscripted kitchen activity with 90K segments; the canonical egocentric action-recognition and anticipation benchmark. Paper Site Code

  • HUI360 (2026) β€” 71 h of in-the-wild 360Β° robot-egocentric capture (11 h retained for the benchmark) with more than 1M curated pose, face-keypoint, mask, tracking, and interaction annotations for human-robot interaction anticipation. arXiv Site

  • EventKitchen (2026) β€” 5.5 h of unscripted cooking from 10 participants in 13 kitchens, captured with stereo event cameras plus synchronized RGB, depth, and IMU and annotated with 10,762 action segments and 13,482 object boxes. arXiv Site

  • InterPet4D (2026) β€” 6.8M synchronized multi-view and egocentric frames from 13 dogs of 11 breeds interacting with 23 people, with audio, segmentation, 2D/3D keypoints, and human, hand, and pet meshes. arXiv Site πŸ€—

  • EgoPolice (2026) β€” 180+ h of real police body-worn camera footage with second-by-second annotations for nine high-stakes police/civilian action classes, classification folds, and multiple-choice VideoQA. arXiv Code

  • ChildLens (2026) β€” 108.58 h of child-worn egocentric video and audio from 62 children for activity analysis in everyday home behavior. Paper Site Data

  • EgoScale (2026) β€” 20k+ h labeled manipulation (paper); Action, dexterous transfer. arXiv Site

  • CogDrive (EyeCue) (2026) β€” Augmented multi-scenario driving dataset paired with the EyeCue gaze-empowered ego-video framework for driver cognitive-distraction detection (reported 74.38% accuracy). arXiv

  • Ego-METAS (2026) β€” 100+ h of untrimmed multimodal egocentric video (RGB, audio, gaze, IMU, monochrome) curated from Ego-Exo4D, CMU-MMAC, and CaptainCook4D with unified splits, pre-extracted features, and baseline sensor-routing policies for online energy-efficient temporal action segmentation. arXiv Site πŸ€—

  • Furhat Egocentric Dataset (2026) β€” 20 sequences (~25 min) of close-range multi-party interactions captured from a Furhat social robot's egocentric camera with amodal face/body boxes and consistent identities for multi-person tracking and re-identification in HRI; available on request under a Data Usage Agreement. arXiv

  • World In Your Hands (2025) β€” 1000+ h of labeled human manipulation data with video-language annotations for action understanding, dexterous transfer, and VLA training. arXiv Site Code

  • EgoCampus (2025) β€” ~32 h / campus paths (paper); Gaze, pedestrian ego. arXiv Site Code

  • AEA (2024) β€” 143 seq. / ~7.3 h; Everyday activities, Aria. arXiv Site πŸ€—

  • EgoSurgery (Phase / Tool / HTS) (2024) β€” Open-surgery ego video with companion phase, tool, and hand-tool segmentation releases for phase recognition, instrument analysis, and dense surgical interaction understanding. arXiv arXiv arXiv Code

  • EgoExo-Fitness (2024) β€” 32 h of synchronized egocentric and exocentric fitness video with temporal boundaries, sub-step annotations, action comments, and quality scores for full-body action understanding. arXiv Code πŸ€—

  • EΒ³ (Exploring Embodied Emotion) (2024) β€” 50+ h; Emotion, multimodal ego. Paper Code

  • EGOFALLS (2023) β€” Fall samples / AV; Fall detection. arXiv Site

  • Epic-Sounding-Object (2023) β€” 3.2K short clips; Audio-visual localization. Paper Code

  • HoloAssist (2023) β€” 169 h; Interactive assistants. Paper Site Code

  • WEAR (2023) β€” ~19 h outdoor sports; Activity + IMU. arXiv Site

  • N-EPIC-KITCHENS (2022) β€” Event + RGB subset; Action, neuromorphic. Paper Site Code

  • Ego-Deliver (2021) β€” 5,360 videos; Delivery ego analysis. Site

  • HOMAGE (2021) β€” 30 h / compositional; Home activities. Paper Site Code

  • MECCANO (2021) β€” ~55 h industrial; HOI, ego. Paper Site Code

  • EGO-CH (2020) β€” 27+ h, cultural sites; Visitor behavior, POI tasks. arXiv

  • EgoCom (2020) β€” 38.5 h conversation; Multiperson ego dialog. Paper Code

  • LEMMA (2020) β€” Multi-view activities; Multi-agent tasks. arXiv Site Code

  • Charades-Ego (2018) β€” Paired ego / exo; Alignment, actions. arXiv Site

  • EGTEA Gaze+ (2018) β€” 28 h cooking; Gaze + action recognition. Paper Site

  • EgoGesture (2017) β€” 2K+ videos / 24K samples; Gesture recognition. Paper Site

  • Stanford ECM (2017) β€” 31 hours, augmented with heart rate and accelerometer, 23–24 daily activity categories; action & activity recognition. Paper

  • THU-READ (2017) β€” 1,920 clips (8 subjects Γ— 40 actions Γ— 3 reps Γ— 2 modalities), RGB-D from helmet-mounted sensor; action & activity recognition. Paper Site

  • PEV (UTokyo Paired Ego-Video) (2016) β€” 1,226 pairs of first-person clips, synchronous dyadic conversations, 8 interaction categories, 6 subjects; action & activity recognition. Paper

  • FPPA (2015) β€” 5 subjects, 5 daily actions, egocentric video with hand and gaze cues; action & activity recognition. Paper

  • JPL-Interaction (2013) β€” 84 videos, 7 activity types (4 positive, 1 neutral, 2 negative interactions), 320Γ—240@30 fps; action & activity recognition. Paper Site

  • ADL (2012) β€” ~10 h, 20 participants; ADL recognition, objects. Paper Site

  • Social Interactions (2012) β€” 8 social events, ~60 hours, head-mounted cameras, multiple participants per event; action & activity recognition. Paper Site

  • EgoAction (2011) β€” First-person sports videos (skateboarding, skiing, cycling, etc; action & activity recognition. Paper

  • Ego-Exo4D β€” see 3D Scene Understanding & Localization

  • EgoExoLearn β€” see Procedural Activities & Skill Learning

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
HUI360In-the-wild human-robot interaction anticipation and cross-dataset transferHUI360 / SSUP-HRISiteDataset+benchmark
EventKitchenEvent-based action recognition, object detection, and stereo depthEventKitchenSiteDataset+benchmark
InterPet4DMultimodal human-pet interaction and pet-motion generationInterPet4DSiteDataset+benchmark
EgoPoliceHigh-stakes action classification and body-camera VideoQAEgoPoliceGitHubDataset+benchmark
EΒ³ (Exploring Embodied Emotion)Emotion recognition, classification, localization, reasoningEΒ³GitHubStandalone
EGOFALLSFall detection (visual + audio)EGOFALLSDataverseStandalone
EgoExo-FitnessAction localization, cross-view verification, skill determinationEgoExo-FitnessGitHubDataset+benchmark
Ego4DAction, forecasting, VQA, narration, …Ego4DSiteSuite
EPIC-KITCHENS-100Action recognition, detection, anticipation, …EPIC-KITCHENS-100SiteSuite
EgoGestureEgocentric hand gesture recognitionEgoGestureSiteStandalone
Ego-METASOnline energy-efficient temporal action segmentationEgo-Exo4D / CMU-MMAC / CaptainCook4DHFStandalone

βœ‹ Hand–Object Interaction, Dexterity & 3D

This section focuses on egocentric hands, dexterous manipulation, object interaction, tracking, and dense 3D understanding around the body and manipulated objects.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ HOI4D20222.4M frames / 4K seq.4D HOIPaperSite
⭐ EgoDex2025829 h / 30K trajectoriesDexterous manipulation, posePaperGitHub
EgoAffordance2026204K episodes / 17.2M affordancesVisual, grasp, trajectory affordancesPaperSite
H-Tac2026160 h / 135K episodesTactile-action pretrainingPaperN/A
EPIC-Contact20262.3K clips / 62.3K framesIn-the-wild 3D hand-object contactPaperSite
HT-Bench202610M RGB / 7.8M tactile framesFull-hand tactile representationPaperN/A
ForceBand202610 h multimodal force demossEMG-to-force, forceful manipulationPaperSite
EventEgoHands202648 clips / 129.6K framesRGB+event hand detectionPaperGitHub
EgoTactile2026~6 h / 768 clips / 63 objectsGrasp pressure from ego videoPaperSite
EgoDex-R20264.3M RGB-D frames / 5.6K seq.Dexterous manipulation, hand-object posePaperN/A
HA-Ego-1K2026~24 h / 484 multi-view clipsManipulation, 6-cam ego + IMUN/AHF
DexGloveHOI20263.5 h / 100K+ samplesVision-IMU 3D hand trackingPaperN/A
EgoTouch20261,891 episodes / 208 tasksTactile HOI, vision-to-touchPaperHF
EgoEVHands20265,419 annotated seq.Stereo event 3D hand pose, gesturePaperGitHub
EgoEMG202641 participants / 10+ hEMG + vision hand posePaperGitHub
HRDexDB20261.4K grasping trialsDexterous grasping, tactile, ego streamsPaperHF
TouchMoment20264,021 videos / 8,456 touch momentsContact moment detectionPaperN/A
EgoFun3D2026271 egocentric videosInteractive 3D objects, function templatesPaperSite
SHOW3D2026In-the-wild ego-exo HOI3D hand-object annotationsPaperN/A
FEEL2026Force-sync kitchen ego videoPhysical action understandingPaperSite
EgoPoints2025Point tracks + syntheticTracking in ego videoPaperGitHub
AssemblyHands20233M images / hands3D hand pose, assemblyPaperSite
EgoObjects20239.2K+ videosDetection, instance segPaperGitHub
ENIGMA-51202322 h industrialFine-grained behaviorPaperSite
POV-Surgery2023~88K frames, 53 seq. (synth.)Surgical hand–tool pose, segmentationPaperSite
VOST2023713 videosVOS, transforming objectsPaperSite
EgoBody2022125 seq. / multi-viewBody pose, interactionPaperSite
EgoHOS202211K+ imagesHand–object segmentationPaperGitHub
EgoPAT3D20221M+ frames RGB-D3D action target predictionPaperSite
Touch and Go202212K+ vis–tactile framesVision + touchPaperSite
VISOR2022EPIC + masks / relationsSegmentation, HOIPaperSite
H2O2021100K+ framesTwo-hand interactionPaperSite
TREK-1502021150 EPIC seq.Object trackingPaperSite
You2Me202014 seq., chest-mounted GoProBody pose via ego–exo interactionPaperGitHub
FPHA20181.2K seq. hand actionHand pose + actionPaperSite
EgoDexter2017~3.2K frames, 4 seq.Hand tracking under occlusionPaperSite
EgoHands20154.8K labeled framesHand detection / boxesPaperSite
BEOID201458 videos, 6 environments, 34 object interaction classes, ~30 fpshand–object interaction, dexterity & 3dPaperData
EDSH20132 videos (~5 min each), pixel-level hand segmentation, egocentric daily activitieshand–object interaction, dexterity & 3dPaperSite
Handled Objects200911 object categories, multiple grasp sequences, RGB + depth from wearable camerahand–object interaction, dexterity & 3dPaperN/A

Entries

  • [⭐️] HOI4D (2022) β€” 2.4M RGB-D frames with object poses, hand poses, interaction regions, and motion segmentation for category-level 4D HOI. Paper Site

  • [⭐️] EgoDex (2025) β€” 829 h / 30K trajectories; Dexterous manipulation, pose. arXiv Site Code

  • EgoAffordance (2026) β€” 204K egocentric manipulation episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances, automatically extracted in a shared 3D actionable representation for VLAff and robot transfer. arXiv Site

  • H-Tac (2026) β€” 160 h of egocentric human videos with tactile/action data across 300+ tasks and 135K episodes, introduced for human-centric transferable tactile-action pretraining and future tactile prediction. arXiv

  • EPIC-Contact (2026) β€” 2.3K in-the-wild EPIC-KITCHENS stable-grasp clips (62.3K frames) with dense bijective 3D hand-object contact correspondences and posed hand/object meshes for unconstrained 3D HOI pose estimation. arXiv Site Code πŸ€—

  • HT-Bench (2026) β€” Large-scale benchmark pairing egocentric vision with full-hand tactile sensing, comprising 10M RGB frames and 7.8M tactile frames across 226 tasks for tactile retrieval, inpainting, vision-to-touch synthesis, and multimodal prediction. arXiv

  • ForceBand (2026) β€” 10 h multimodal dataset with egocentric video, wrist sEMG, IMU, and fingertip force measurements across diverse everyday objects/actions, used to learn EMG-to-force labels for force-augmented robot demonstrations; public dataset release is marked as coming soon. arXiv Site

  • EventEgoHands (2026) β€” 48 egocentric clips (~1.2 h, 129.6K frames) pairing RGB with synthetic event streams (synthesized from EgoHands via v2e) and 393K hand bounding boxes for RGB-event hand detection under motion blur and low light. arXiv Code

  • EgoTactile (2026) β€” ~6 h (768 clips, 319K frames) of head- and neck-mounted egocentric video of 12 participants grasping 63 everyday objects with synchronized 162-taxel pressure-glove supervision and a bare-hand transfer subset for full-hand grasp pressure estimation; the dataset currently sits under an anonymous ICML-submission account. arXiv Site πŸ€—

  • EgoDex-R (2026) β€” 4.3M egocentric RGB-D frames across 5,600 manipulation sequences (1,000+ objects, 200+ daily task categories) with MANO hand poses, 6-DoF object trajectories, reconstructed meshes, and contact annotations, introduced in the EgoAERO paper; distinct from Apple's EgoDex. arXiv

  • HA-Ego-1K (2026) β€” ~24 h of privacy-redacted six-camera + IMU egocentric video (484 multi-view clips across 22 real-world work scenarios such as workshops, construction, and factories) captured with the head-worn Human Archive GSI Cap for dexterous-manipulation and long-horizon task research; gated access (CC BY-NC 4.0), no paper yet. Site πŸ€—

  • DexGloveHOI (2026) β€” 3.5 h / 100K+ synchronized egocentric vision-IMU samples with MoCap 3D hand-pose ground truth for dexterous hand-object interaction tracking; no official public data page was found. arXiv

  • EgoTouch (2026) β€” 1,891 bimanual hand-object interaction episodes across 208 manipulation tasks with synchronized egocentric and wrist RGB video, 3D hand pose, and dense tactile pressure maps. arXiv Site Code πŸ€—

  • EgoEVHands (2026) β€” 5,419 real-world stereo event-camera egocentric sequences with dense 2D/3D hand keypoints across 38 gesture classes; the official repository currently says code, models, and dataset links are to be uploaded. arXiv Code

  • EgoEMG (2026) β€” 10+ h of synchronized bilateral EMG, IMU, egocentric RGB, external RGB-D, and mocap-derived hand pose across 41 participants and 60 gesture classes. arXiv Code

  • HRDexDB (2026) β€” 1.4K dexterous human and robotic hand grasping trials with synchronized multi-view video, egocentric video streams, tactile signals, and 3D motion. arXiv πŸ€—

  • TouchMoment (2026) β€” 4,021 egocentric videos with 8,456 annotated hand-object contact moments for frame-precise touch detection. arXiv

  • EgoFun3D (2026) β€” 271 egocentric interaction videos with paired 3D geometry, 2D/3D segmentation, articulation labels, and function-template annotations. arXiv Site πŸ€—

  • SHOW3D (2026) β€” In-the-wild ego-exo capture of hands interacting with objects, with 3D hand-object annotations from a marker-less multi-camera system. arXiv

  • FEEL (2026) β€” Force-sync kitchen ego video; Physical action understanding. arXiv Site

  • EgoPoints (2025) β€” Point tracks + synthetic; Tracking in ego video. arXiv Site Code

  • AssemblyHands (2023) β€” 3M egocentric hand images on top of Assembly101 for detailed 3D hand pose estimation during assembly. Paper Site Code

  • EgoObjects (2023) β€” 9.2K+ videos; Detection, instance seg. Paper Site Code

  • ENIGMA-51 (2023) β€” 22 h industrial; Fine-grained behavior. arXiv Site Code

  • POV-Surgery (2023) β€” ~88K frames, 53 seq. (synth.); Surgical hand–tool pose, segmentation. arXiv Site Code

  • VOST (2023) β€” 713 videos; VOS, transforming objects. arXiv Site

  • EgoBody (2022) β€” 125 seq. / multi-view; Body pose, interaction. arXiv Site

  • EgoHOS (2022) β€” 11K+ images; Hand–object segmentation. arXiv Code

  • EgoPAT3D (2022) β€” 1M+ frames RGB-D; 3D action target prediction. Paper Site Code

  • Touch and Go (2022) β€” 12K+ vis–tactile frames; Vision + touch. arXiv Site Code

  • VISOR (2022) β€” EPIC + masks / relations; Segmentation, HOI. arXiv Site Code

  • H2O (2021) β€” 100K+ frames; Two-hand interaction. Paper Site Code

  • TREK-150 (2021) β€” 150 EPIC seq; Object tracking. arXiv Site Code

  • You2Me (2020) β€” 14 seq., chest-mounted GoPro; Body pose via ego–exo interaction. Paper arXiv Code

  • FPHA (2018) β€” 1,175 RGB-D sequences with 3D hand pose and action labels; a foundational first-person hand-action benchmark. Paper Site Code

  • EgoDexter (2017) β€” ~3.2K frames, 4 seq; Hand tracking under occlusion. arXiv Project

  • EgoHands (2015) β€” 4.8K labeled frames; Hand detection / boxes. Paper Project

  • BEOID (2014) β€” 58 videos, 6 environments, 34 object interaction classes, ~30 fps; hand–object interaction, dexterity & 3d. Paper Data

  • EDSH (2013) β€” 2 videos (~5 min each), pixel-level hand segmentation, egocentric daily activities; hand–object interaction, dexterity & 3d. Paper Site

  • Handled Objects (2009) β€” 11 object categories, multiple grasp sequences, RGB + depth from wearable camera; hand–object interaction, dexterity & 3d. Paper

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoAffordance / VLAffVisual, grasp, and trajectory affordance predictionEgoAffordanceSiteDataset+benchmark
H-TacHuman-to-robot tactile-action pretraining and future tactile predictionH-TacPaperDataset+pretraining resource
EPIC-Contact / HOPformerIn-the-wild egocentric 3D hand-object pose and contact estimationEPIC-ContactSiteDataset+benchmark
HT-BenchFull-hand tactile representation learning with egocentric visionHT-BenchPaperDataset+benchmark
ForceBand / EMG2ForcesEMG-to-fingertip-force prediction and forceful manipulation policy learningForceBandSiteDataset+benchmark
TouchMomentFrame-precise hand-object contact moment detectionTouchMomentPaperStandalone
EgoFun3DInteractive 3D object modeling and function-template inferenceEgoFun3DSiteDataset+benchmark
EgoEMGEMG-to-pose, vision-to-pose, and EMG+vision fusionEgoEMGGitHubDataset+benchmark
EgoTouch / TouchAnythingVision-to-touch prediction for bimanual HOIEgoTouchHFDataset+benchmark
DexGloveHOIVision-IMU 3D hand tracking under HOI occlusionDexGloveHOIPaperDataset+benchmark
EgoEVHandsStereo event 3D hand pose and gesture recognitionEgoEVHandsGitHubDataset+benchmark
AssemblyHandsEgocentric 3D hand poseAssembly101SiteStandalone
VISORVideo object segmentation, hand–object relationsEPIC-KITCHENSSiteStandalone
TREK-150Egocentric single-object trackingEPIC-KITCHENSSiteStandalone
EggHandEgocentric 3D hand pose forecastingEgoExo4DPaperMethod benchmark
FPHAHand action + 3D hand poseFPHASiteStandalone
EgoTactileFull-hand grasp pressure estimation from ego videoEgoTactileSiteDataset+benchmark
EventEgoHandsMultimodal RGB-event egocentric hand detectionEventEgoHands (from EgoHands)GitHubDataset+benchmark

πŸ“‹ Procedural Activities & Skill Learning

Datasets centered on step structure, instructional execution, assembly, or skill transfer from egocentric experience are grouped here.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ EgoExoLearn2024120 h ego+exoProcedural, async viewsPaperGitHub
⭐ Assembly1012022513 h multiviewAssembly, procedurePaperSite
EgoProceVQA20263,600 QA / 31 tasks / 4 scenariosKey-step procedural reasoningPaperSite
CoMind2026Dual ego + 2 exo views / 55 environmentsCollaborative activity, social reasoningPaperSite
VLK202648K synthetic paired trajectoriesHumanoid loco-manipulation, VLKPaperSite
EgoVerse20261,362 h / ~80K episodesRobot learning, manipulation skillsPaperSite
EgoLive2026Large-scale real-world task routinesRobot manipulation learningPaperN/A
EgoMAGIC20263,355 videos / 50 medical tasksField medicine, action detectionPaperZenodo
HumanEgo2026Minutes-per-task Aria demonstrationsHuman-to-robot policy learningPaperSite
EgoSPT202611,515 episodes / 112 task foldersSpatially prompted manipulation trajectoriesPaperHF
Ego-EXTRA202650 h / 15K+ VQAExpert-trainee assistancePaperSite
GM-1002026100+ tasks / 13K+ trajectoriesRobot manipulation, embodied evaluationPaperSite
SABER2026100+ h / 44.8K samplesRetail VLA adaptationPaperSite
EgoProactive / ProΒ²Bench2026700 recordings (22–55 min) / 42K eval instancesProactive procedural assistancePaperHF
EgoYC2 / Exo2EgoDVC2025~43 h cookingDense captioning, proceduralPaperGitHub
IndustReal2024~6 h industrialProcedure steps, errorsPaperSite
EgoProceL202262 videos / 16 tasksProcedure learningPaperSite
EPIC-Tent20197+ h, tent assemblyProcedural, dual HMD + gazePaperSite
CMU-MMAC201125 subjects, 5 cooking recipesprocedural activities & skill learningPaperSite
GTEA Gaze201117 meal preparation sessions, 7 cooking activities, gaze tracking annotationsprocedural activities & skill learningPaperSite

Entries

  • [⭐️] EgoExoLearn (2024) β€” 120 h ego+exo; Procedural, async views. Paper Site Code πŸ€—

  • [⭐️] Assembly101 (2022) β€” 513 h multiview; Assembly, procedure. Paper Site

  • EgoProceVQA (2026) β€” 3,600 key-step-centric questions across 31 everyday tasks and four procedural scenarios, covering six question types generated with EgoProceGen and human-checked for procedural reasoning evaluation. arXiv Site

  • CoMind (2026) β€” Collaborative cooking captured from two synchronized head-mounted cameras and two exocentric views, with audio, gaze, hand/object interactions, social cues, and aligned scans across 55 environments. arXiv Site

  • VLK (2026) β€” 48K synthetic vision-language-kinematics trajectories rendered as egocentric observations in reconstructed indoor 3DGS scenes, paired with language commands and whole-body humanoid kinematic trajectories for loco-manipulation. arXiv Site

  • EgoVerse (2026) β€” 1,362 h of egocentric human demonstrations spanning ~80K episodes and 1,965 tasks for robot learning from human manipulation experience. arXiv Site Code

  • EgoLive (2026) β€” Large-scale annotated egocentric recordings of real-world human task routines for robot manipulation learning. arXiv

  • EgoMAGIC (2026) β€” 3,355 egocentric field-medicine videos covering 50 medical tasks, with released medical training data and an action-detection challenge. arXiv Site

  • HumanEgo (2026) β€” Minutes-per-task human egocentric demonstrations collected with Aria glasses for zero-shot human-to-robot manipulation-policy learning via interaction-centric spatial representations. arXiv Site

  • EgoSPT (2026) β€” 11,515 processed egocentric manipulation episodes for spatially prompted visual trajectory prediction, with RGB video, end-effector poses, gripper widths, and valid-frame masks. arXiv πŸ€—

  • Ego-EXTRA (2026) β€” 50 h of unscripted expert-trainee egocentric procedural assistance across bike workshop, kitchen, bakery, and assembly scenarios, with dialogue transcripts and 15K+ VQA sets. Paper Site

  • GM-100 (2026) β€” 100+ detail-oriented robot manipulation tasks with 13K+ teleoperated trajectories and robot first-person camera views for embodied skill evaluation. arXiv Site Code

  • SABER (2026) β€” 100+ h of natural in-store retail activity with head-mounted egocentric video, 360-degree exocentric video, and 44.8K action samples for VLA adaptation. arXiv Site πŸ€—

  • EgoProactive / ProΒ²Bench (2026) β€” 700 Ray-Ban Meta smart-glasses recordings (22–55 min each) of cooking, crafts, DIY, and tutorial sessions with per-decision-point interrupt/silent labels and Out-of-Plan deviation-recovery annotations for proactive procedural assistance; ProΒ²Bench unifies five existing egocentric benchmarks into 42K evaluation and 250K training instances. arXiv πŸ€—

  • EgoYC2 / Exo2EgoDVC (2025) β€” ~43 h cooking; Dense captioning, procedural. arXiv Site Code

  • IndustReal (2024) β€” ~6 h industrial; Procedure steps, errors. Paper Site

  • EgoProceL (2022) β€” 62 videos / 16 tasks; Procedure learning. arXiv Site Code

  • EPIC-Tent (2019) β€” 7+ h, tent assembly; Procedural, dual HMD + gaze. Paper Site Code

  • CMU-MMAC (2011) β€” 25 subjects, 5 cooking recipes; procedural activities & skill learning. Paper Site

  • GTEA Gaze (2011) β€” 17 meal preparation sessions, 7 cooking activities, gaze tracking annotations; procedural activities & skill learning. Paper Site

  • Ego-Exo4D β€” see 3D Scene Understanding & Localization

  • HowToDIV β€” see VLMs, Instructions & QA

  • ADT (Aria Digital Twin) β€” see 3D Scene Understanding & Localization

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoProceVQAKey-step procedural understanding across six QA typesEgoProceVQASiteDataset+benchmark
CoMindJoint attention, socially conditioned interaction anticipation, collaborative handoverCoMindSiteDataset+benchmark
VLKVision-language-kinematics policy learning for humanoid navigation and object transportSynthetic 3DGS trajectoriesSiteDataset+benchmark
HumanEgoZero-shot human-to-robot manipulation from egocentric videoHumanEgoSiteDataset+benchmark
EgoSPT / SP-VTPSpatially prompted visual trajectory prediction for manipulationEgoSPTHFDataset+benchmark
Ego-EXTRAExpert-trainee procedural assistance and VQAEgo-EXTRASiteDataset+benchmark
EgoMAGICField-medicine action detectionEgoMAGICZenodoDataset+benchmark
GM-100Detail-oriented robot manipulation evaluationGM-100SiteDataset+benchmark
TAVISActive-vision imitation learning on humanoid robots (GR1T2, Reachy2) in IsaacLab; TAVIS-Head + TAVIS-Hands suites with the GALT anticipatory-gaze metricSimulation-only (no real-data release)PaperStandalone benchmark
EgoProactive / ProΒ²BenchProactive intervention timing and Out-of-Plan recovery guidanceEgoProactive + Ego4D / EPIC-KITCHENS / Ego-Exo4D / HoloAssist / HowTo100MHFDataset+benchmark

πŸ—ΊοΈ 3D Scene Understanding & Localization

These datasets emphasize geometry, localization, scene graphs, multiview capture, or machine-perception tasks grounded in ego video.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ego-Exo4D20241,286+ h ego+exoSkilled activity, many tasksPaperSite
⭐ ADT (Aria Digital Twin)2023200 seq., 2 scenesEgocentric 3D perceptionPaperSite
GST-Bench / GST-Train20266,790 min synthetic videoGlobal spatial awareness from ego videoPaperN/A
FloAff-Kitchen2026Cross-scene, multi-view kitchen benchmarkNavigation-to-manipulation affordancePaperSite
EgoHTR202655 seq. / 150K+ frames / 7 scenes4D human-terrain reconstructionPaperSite
SG-Ego20263.8M graphs / 7.3K Ego4D videosSpatio-temporal scene graphsPaperHF
PRISM2026270K samples / 11.8M framesRetail embodied VLM, spatial reasoningPaperHF
EgoTraj202610.7 h / 1.15M framesEgocentric trajectory predictionPaperGitHub
AIST-Living2026Egocentric video + GT motion in scanned env.Global pose, localizationPaperSite
OVO-S-Bench2026348 videos / 1,680 Q / 30 task typesStreaming spatial intelligence QAPaperSite
PVSG2023400 vids, ~150K framesPanoptic video scene graph (ego + third-person)PaperSite
DR(eye)VE2018~6 h driving, 555K framesGaze prediction, driving ego videoPaperSite
EgoCart2018Retail RGB-D, 9 videosIndoor / cart localizationPaperSite
IU ShareView20189 paired ego video setsPerson seg / ID across synchronized wearersPaperSite
OST201757 sequences, 55 subjects, ~15 min/video, egocentric object search tasks, eye-tracking ground truth3d scene understanding & localizationPaperGitHub

Entries

  • [⭐️] Ego-Exo4D (2024) β€” 1,286+ h of paired first- and third-person skilled activity with multiview geometry and a broad benchmark suite. arXiv Site

  • [⭐️] ADT (Aria Digital Twin) (2023) β€” 200 seq., 2 scenes; Egocentric 3D perception. Paper Site πŸ€—

  • GST-Bench / GST-Train (2026) β€” Human-verified global-spatial-temporal questions derived from 6,790 minutes of synthetic first-person exploration, requiring novel-view inference and mapping ego observations onto global top-down scenes, plus a companion training set. arXiv

  • FloAff-Kitchen (2026) β€” Cross-scene, multi-view benchmark for predicting where a mobile robot should stand to execute downstream manipulation, spanning varied skills, layouts, furniture styles, and egocentric viewpoints. arXiv Site

  • EgoHTR (2026) β€” 55 scene-aligned 4D human-terrain traversal sequences (150K+ frames across seven challenging scenes) with ego/exo Aria video, SLAM, IMU, 3D scans, and parametrized human motion for analysis, synthesis, and humanoid locomotion transfer. arXiv Site

  • SG-Ego (2026) β€” Large-scale spatio-temporal scene-graph annotations extending Ego4D: SG-Ego-Align provides ~3.8M graphs from 7,297 videos, while SG-Ego-Edit adds action-conditioned graph-edit forecasting samples for A-GEF. arXiv Site Code πŸ€—

  • PRISM (2026) β€” 270K-sample multi-view retail video SFT corpus with egocentric, exocentric, and 360-degree views for embodied VLM spatial, physical, and action reasoning. arXiv Site πŸ€—

  • EgoTraj (2026) β€” 10.7 h / 1.15M frames of Meta Quest Pro egocentric urban navigation with synchronized RGB, 6DoF head pose, gaze, and scene annotations for trajectory forecasting; the GitHub README says the dataset and dashboard will be released after publication. arXiv Code

  • AIST-Living (2026) β€” Dataset introduced with Map-Mono-Ego that pairs monocular egocentric video with ground-truth human motion in a pre-scanned 3D environment for globally consistent pose estimation. arXiv Site

  • OVO-S-Bench (2026) β€” 1,680 fully human-annotated questions over 348 continuous egocentric streams (indoor walkthroughs, daily activities, outdoor tours, and driving from nine sources) spanning 30 task types across four hierarchical levels, from instantaneous perception to allocentric mapping, for streaming spatial intelligence in multimodal LLMs. arXiv Site Code πŸ€—

  • PVSG (2023) β€” 400 vids, ~150K frames; Panoptic video scene graph (ego + third-person). Paper arXiv Site Code

  • DR(eye)VE (2018) β€” ~6 h driving, 555K frames; Gaze prediction, driving ego video. arXiv Site Code

  • EgoCart (2018) β€” Retail RGB-D, 9 videos; Indoor / cart localization. Paper Site

  • IU ShareView (2018) β€” 9 paired ego video sets; Person seg / ID across synchronized wearers. Paper arXiv Site

  • OST (2017) β€” 57 sequences, 55 subjects, ~15 min/video, egocentric object search tasks, eye-tracking ground truth; 3d scene understanding & localization. Paper Code

  • Ego-1K β€” see Video Generation & World-Model Pretraining

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
GST-BenchGlobal spatial-temporal VQA and allocentric mapping from ego streamsGST-BenchPaperDataset+benchmark
FloAff-KitchenTarget-conditioned floor-affordance prediction for mobile manipulationFloAff-KitchenSiteDataset+benchmark
EgoHTRScene-aligned 4D human motion reconstruction and terrain traversalEgoHTRSiteDataset+benchmark
A-GEF / SG-EgoAction-conditioned scene-graph edit forecasting and graph-text reasoningSG-EgoSiteDataset+benchmark
Ego-Exo4DEgo–exo skill understanding, many tasksEgo-Exo4DSiteSuite
EgoTrajEgocentric multimodal trajectory forecastingEgoTrajGitHubDataset+benchmark
EgoProxEgocentric 3D proximity reasoning VQAADT / EgoExo4DSiteStandalone
Map-Mono-EgoMap-grounded global human pose estimationAIST-LivingSiteDataset+benchmark
ADT (Aria Digital Twin)Egocentric 3D machine perceptionADTAriaDataset+benchmark
OVO-S-BenchStreaming spatial intelligence over continuous ego videoNine egocentric video sourcesSiteStandalone

πŸ› οΈ Tools & Libraries

NameDescriptionLink
Ego4D CLIOfficial downloader and tooling for accessing Ego4D releases.GitHub
HOMIE-toolkitToolkit released with Ropedia Xperience-10M for large-scale multimodal ego data.GitHub
Open-AoE ToolchainSmartphone capture, reconstruction, visualization, retargeting, and model-ready conversion for Open-AoE.GitHub
Ego-OSCAROpen-hardware stereo-inertial capture device and recording stack with a sub-$200 bill of materials.Paper
ego-stereo-cn-v1-toolsLoading + timing verification for the ego-stereo-cn-v1 LeRobot v3 stereo+IMU sample (hardware-synced).GitHub
AssemblyHands ToolkitOfficial toolkit for the AssemblyHands benchmark.GitHub
TREK-150 ToolkitToolkit for the TREK-150 egocentric tracking benchmark.GitHub

🀝 Contributing

  1. Add or update the dataset, benchmark, or survey directly in the matching section of README.md.
  2. Keep primary entries unique: one full entry under one topic, cross-links everywhere else.
  3. Preserve newest-to-oldest ordering inside each topic block, with flagship entries kept at the top.
  4. Follow the detailed checklist in CONTRIBUTING.md before opening a PR.

❀️ Contact

If you have suggestions, dataset updates, or find this project useful, feel free to contact Shen Yujiao at shenyujiao18@gmail.com.

License

CC0 1.0 Universal. See LICENSE.

awesome-list
computer-vision
datasets
egocentric-datasets
egocentric-vision
first-person-video
video-datasets

Contributors

player0718

32 commits

Jingkang50

3 commits

TateZhouSiu

1 commits

player0718/awesome-ego-video-datasets

πŸŽ₯ [Awesome] Egocentric / First-Person Video Datasets πŸ“š Papers, Benchmarks & Resources for Ego Vision

225

38 commits

updated Sep 21, 2026

See the code

README

πŸŽ₯ Awesome Egocentric Video Datasets

πŸ“œ A Curated List of Egocentric (First-Person) Video Datasets, Benchmarks, and Tools

Awesome Egocentric Video Datasets

Awesome Papers PRs Welcome License: CC0-1.0

Overview

Overview of egocentric video datasets

This repository tracks egocentric video datasets through a task-first view: every dataset appears once as a primary entry under one of seven research themes, with cross-links where it also matters. The goal is fast navigation for researchers who need scale, task fit, benchmark context, and official resources without bouncing across multiple index files.

Papers & Surveys

Sorted newest to oldest, with flagship surveys and corpus papers highlighted first.

  • Vision-Language Models for Egocentric Video: From Hand-Object Interaction to Embodied AI (2026) β€” Survey of egocentric VLMs spanning datasets, hand-object interaction, temporal reasoning, multimodal learning, wearable assistance, and human-to-robot transfer. arXiv

  • Position: Life-Logging Video Streams Make the Privacy-Utility Trade-off Inevitable (2026) β€” Position paper arguing that privacy leakage in always-on wearable video should be evaluated across the full data and model pipeline with standardized metrics and benchmarks. arXiv

  • Building Egocentric Procedural AI Assistant: Methods, Benchmarks, and Challenges (2025) β€” Li et al., 2025 survey and benchmark paper of egocentric procedural activity understanding, focus on building egocentric procedural AI assistant. arXiv

  • [⭐️] Challenges and Trends in Egocentric Vision: A Survey (2025) β€” Li et al., 2025 survey of datasets, tasks, benchmarks, and open challenges in egocentric vision. arXiv

  • Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision (2025) β€” Cross-view survey covering ego-exo collaboration, paired capture, and collaborative perception. arXiv

  • HD-EPIC: A Highly-Detailed Egocentric Video Dataset (2025) β€” Dataset paper introducing fine-grained kitchen understanding with dense multimodal annotations. arXiv

  • Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives (2024) β€” Large-scale paired ego-exo dataset paper spanning skilled activity and multiview understanding. arXiv

  • EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World (2024) β€” Procedural ego-exo paper focused on asynchronous activity alignment in real environments. Paper

  • EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding (2023) β€” Benchmark paper targeting long-form memory and reasoning over egocentric video. arXiv

  • [⭐️] Ego4D: Around the World in 3,000 Hours of Egocentric Video (2022) β€” Flagship corpus paper introducing the large-scale Ego4D benchmark suite. arXiv

  • [⭐️] Rescaling Egocentric Vision: Collection, Pipeline and Challenges for EPIC-KITCHENS-100 (2021) β€” Canonical kitchen benchmark paper for action recognition, detection, and anticipation. Paper

  • Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos (2018) β€” Early paired ego-exo benchmark for activity transfer and alignment. arXiv

🎬 Video Generation & World-Model Pretraining

Datasets here emphasize large-scale first-person pretraining corpora, video generation, editing, or world-model supervision from ego video.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ropedia Xperience-10M2026Large multi-stream experiencesMultimodal ego learningN/AHugging Face
WorldRover-10M20266,003 seq. / 21.9M frames / 202.7 hFirst-person world models, 3D explorationPaperN/A
H2R-Bench20266 manipulation families / 2 robot embodimentsHuman-to-robot video generationPaperN/A
Ego-OSCAR-550h2026~550 h/camera / 1,462 stereo sessionsStereo-inertial ego pretraining, capturePaperN/A
Ego2Robot202618,561 h / 15 robot morphologiesEgo-to-robot data synthesis, VLA pretrainingPaperSite
ACE-Data-02026150 h / 75K episodes / 200 tasksMultimodal embodied pretrainingPaperSite
EgoPlay2026106K event-triggered clip-prompt pairsEvent-triggered ego video editingPaperN/A
Open-AoE2026~2,000 h / 500+ contributorsManipulation pretraining, data toolchainPaperGitHub
EgoVid-Pro2026103K clips / ~12M framesHand-controlled ego video generationPaperN/A
RetailSMV202632,105 clips / 16.1K ego + 16.0K exoRetail world-model adaptationPaperSite
EgoCS-400K2026400K+ videos / 10K h gameplayAction-conditioned world modelsPaperN/A
WM-H (Wh0)202650K generated HOI episodesSynthetic dexterous VLA dataPaperSite
DreamDojo-HV2026Very large FP video (see paper)World models, pretrainingPaperN/A
Ego-1K2026Multiview clips (~1K takes)Neural 3D/4D synthesisPaperHugging Face
In-lab2026Lab tabletop trajectoriesSkills, world models (w/ DreamDojo)PaperN/A
HumanNet2026~1M h human-centric video (ego + exo)VLA / embodied pretrainingPaperN/A
MobileEgo Anywhere2026200 h smartphone-collected long-horizon egoLong-horizon ego data infra, VLAPaperN/A
EgoEdit2025100K editing pairsEgocentric video editingPaperSite
EgoVid-5M20245M clipsVideo generation, motion+textPaperSite

Entries

  • [⭐️] Ropedia Xperience-10M (2026) β€” 10M multimodal experiences with 6 RGB streams, stereo depth, pose/SLAM, hand-body mocap, audio, and IMU for large-scale ego pretraining. Site Code πŸ€—

  • WorldRover-10M (2026) β€” 6,003 synthetic exploration sequences from 32 environments (21.9M frames / 202.7 h, including 10.8M first-person frames) with first-person, third-person, and 360Β° views aligned to metric depth, trajectories, geometry, and action signals. arXiv

  • H2R-Bench (2026) β€” Cross-embodiment benchmark for transforming egocentric human demonstrations into robot manipulation videos, evaluated across six manipulation families, two target embodiments, and five dimensions covering execution, contact, embodiment, and visual quality. arXiv

  • Ego-OSCAR-550h (2026) β€” ~550 h per camera (1,462 stereo sessions) of calibrated everyday egocentric video with synchronized IMU, dense open-vocabulary action captions, per-frame 3D hand reconstructions, and an open sub-$200 capture stack. arXiv

  • Ego2Robot (2026) β€” 18,561 h of synthetic robot training data across 15 morphologies, generated from curated and in-the-wild egocentric manipulation video through action retargeting, robot-arm compositing, and multi-level quality curation. arXiv Site

  • ACE-Data-0 (2026) β€” 150 h and 75K interaction episodes across 200 household task categories, synchronizing ego/exo video, full-body and hand motion, object state, audio, and tactile signals at table and room scale. arXiv Site πŸ€—

  • EgoPlay (2026) β€” 106K event-triggered egocentric clip-prompt pairs, primarily derived from Ego4D, covering positive, fabricated-negative, and multi-event triggers for temporally restrained video editing and streaming evaluation. arXiv

  • Open-AoE (2026) β€” ~2,000 h of smartphone-collected manipulation video from 500+ contributors with bilingual text, MANO hand pose, camera trajectory, and atomic-action annotations, plus capture-to-training tools for VLA and world-model research. arXiv Code πŸ€—

  • EgoVid-Pro (2026) β€” 103K in-the-wild egocentric clips (~12M frames) with clean protagonist-only 3D hand trajectories, curated for HandsOnWorld and Plucker Hand Map conditioning in hand-controlled first-person video generation. arXiv

  • RetailSMV (2026) β€” 32,105 captioned retail clips from five supermarkets with synchronized staff-view egocentric and exocentric capture, predefined train/val/test splits, and a held-out protocol for video-world-model adaptation. arXiv Site

  • EgoCS-400K (2026) β€” 400K+ replay-grounded first-person Counter-Strike videos (10K h) aligned with player states, view directions, movements, keyboard/button inputs, events, and round context for action-conditioned rollout, captioning, and world-model training. arXiv

  • WM-H (Wh0) (2026) β€” 50K world-model-generated egocentric human-object manipulation episodes conditioned on language, objects, and scenes, then converted into robot-trainable supervision for dexterous VLA adaptation. arXiv Site Code

  • DreamDojo-HV (2026) β€” Very large FP video (see paper); World models, pretraining. arXiv Site

  • Ego-1K (2026) β€” Multiview clips (~1K takes); Neural 3D/4D synthesis. arXiv Site πŸ€—

  • In-lab (2026) β€” Lab tabletop trajectories; Skills, world models (w/ DreamDojo). arXiv Site

  • HumanNet (2026) β€” ~1M h of human-centric video (mix of ego and exo) with interaction-centric annotations; the authors report 1k h of ego human video outperforms 100 h of real-robot data for VLA training. arXiv

  • MobileEgo Anywhere (2026) β€” 200 h of hour-plus egocentric trajectories collected on commodity smartphones, released with an open-source mobile capture app and a processing pipeline aimed at VLA pretraining. arXiv

  • EgoEdit (2025) β€” 100K editing pairs; Egocentric video editing. arXiv Site

  • EgoVid-5M (2024) β€” 5M first-person clips curated for text-and-motion-conditioned video generation from wearable footage. arXiv Site Code

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
H2R-BenchCross-embodiment human-to-robot manipulation video generationH2R-BenchPaperStandalone
EgoPlayEvent-triggered editing, pre-trigger preservation, false-trigger robustnessEgoPlay / Ego4DPaperDataset+benchmark
ACE-Data-0Hierarchical signals-to-scenes-to-interactions evaluationACE-Data-0SiteDataset+benchmark
Ego2Robot / RoboTwin2.0 extensionOOD visual, spatial, embodiment, and semantic generalizationEgo2RobotSiteDataset+benchmark
HandsOnWorld / EgoVid-ProCamera-disentangled hand-controlled egocentric video generationEgoVid-ProPaperDataset+benchmark
RetailSMVRetail video-world-model adaptation and ego/exo viewpoint ablationsRetailSMVSiteDataset+benchmark
EgoCS-400KAction-conditioned future prediction, state/event-aware rollout, replay-grounded captioningEgoCS-400KPaperDataset+benchmark
Wh0 / WM-HSynthetic egocentric dexterous manipulation data for VLA adaptationWM-HSiteDataset+training resource
EgoEditEgocentric video editingEgoEdit / EgoEditDataProjectDataset+benchmark

🧠 Memory, Summarization & Long-form Understanding

This section collects long-horizon lifelog, summarization, and persistent-memory datasets where temporal continuity matters as much as recognition.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ EgoLife2025~266–300 h daily lifeLong-form assistants, memoryPaperSite
⭐ EgoSchema2023250+ h / 5K QALong-form video QAPaperSite
EgoMonth2026301 h / 738 clips / 1,443 QAMonth-level spatiotemporal memoryPaperHF
MEMORA-Bench202645 h / 18 participantsEmbodied action memory, planningPaperN/A
EgoServe20263K+ service instances / 4 horizonsProactive continuous-video assistancePaperSite
SuperMemory-VQA202652.9 h / 4,853 QALong-horizon memory VQAPaperHF
EgoMemReason2026500 MCQs over EgoLifeWeek-long memory reasoningPaperHF
EgoExoMem20262.6K MCQs / 390 videosCross-view memory reasoningPaperGitHub
EgoIntrospect2026180 h / 60 subjectsInternal-state reasoning, memoryPaperSite
MA-EgoQA20261,741 QA / 6 agents / 7 daysMulti-agent egocentric QAPaperSite
EgoStream20262,250 Q / 8,528 evals, streams up to 45.3 hStreaming episodic memoryPaperSite
VidChapters-7M2023817K videos / 7M chaptersChaptering (not ego-only)PaperSite
Multi-Ego2022~12 h / 41 seq.Multi-wearer, summarizationPaperGitHub
DoMSEV201880 h, 48 seq.Semantic fast-forward, first-person videoPaperSite
HUJI-EgoSeg201429 long egocentric videos (~1–5 h each), pixel-level temporal segmentation annotationsmemory, summarization & long-form understandingPaperSite
UT Ego2012~17 h, 4 long videosSummarization, long-form egoPaperSite
VINST / Visual Diaries201131 egocentric videos capturing daily commutes; used for temporal segmentation and video summarizationmemory, summarization & long-form understandingPaperSite

Entries

  • [⭐️] EgoLife (2025) β€” ~266-300 h of daily-life capture in EgoHouse with Meta Aria, third-person cameras, and mmWave sensors for persistent assistant memory. arXiv Site Code πŸ€—

  • [⭐️] EgoSchema (2023) β€” 250+ h of long-form Ego4D video with 5K QA pairs designed to probe memory and causal understanding over extended clips. arXiv Site Code

  • EgoMonth (2026) β€” 301 h across 738 wearable-camera clips from 20 participants recorded over 20–120 days, paired with 1,443 human-authored questions spanning schema consolidation, episodic indexing, and cascading reasoning. arXiv πŸ€—

  • MEMORA-Bench (2026) β€” 45 h of EPIC-KITCHENS-100 extension video from 18 participants for memory-grounded planning toward seen and unseen goals, plus structured assessment of environment, entity, activity, and inferred-knowledge memory. arXiv

  • EgoServe (2026) β€” 3K+ manually verified proactive-service instances over EgoLife, HoloAssist, and CaptainCook4D, organized into 10 assistance categories and four temporal horizons from instant alerts to multi-day habit coaching. arXiv Site Code πŸ€—

  • SuperMemory-VQA (2026) β€” 52.9 h of everyday Meta Aria recordings with RGB, processed gaze, IMU, SLAM trajectories, point clouds, redacted transcripts, and 4,853 human-verified long-horizon memory QA pairs. arXiv πŸ€—

  • EgoMemReason (2026) β€” 500 multiple-choice questions over week-long EgoLife video for entity, event, and behavior memory reasoning, with public questions and leaderboard evaluation. arXiv Site Code πŸ€—

  • EgoExoMem (2026) β€” 2.6K human-verified MCQs over 390 synchronized egocentric and exocentric videos from EgoExo4D and LEMMA for cross-view memory reasoning. arXiv Code

  • EgoIntrospect (2026) β€” 180 h of user-driven egocentric recordings from 60 subjects with synchronized video, audio, gaze, motion, and physiological signals for affective experience, request intent, and cognitive-memory reasoning; the paper states data will be made public. arXiv Site

  • MA-EgoQA (2026) β€” 1,741 QA pairs over six temporally aligned EgoLife egocentric streams spanning seven days, targeting multi-agent social interaction, task coordination, theory-of-mind, temporal reasoning, and environmental interaction. arXiv Site Code

  • EgoStream (2026) β€” 2,250 curated questions expanded to 8,528 recall-conditioned evaluations via Answer Validity Windows over egocentric streams up to 45.3 h (curated from Ego4D, EgoLife, EgoTempo, Multi-Hop EgoQA, and HD-EPIC), spanning seven cognitive memory dimensions for diagnosing streaming episodic memory in video-language models. arXiv Site

  • VidChapters-7M (2023) β€” 817K videos / 7M chapters; Chaptering (not ego-only). arXiv Site Code

  • Multi-Ego (2022) β€” ~12 h / 41 seq; Multi-wearer, summarization. Paper arXiv Code

  • DoMSEV (2018) β€” 80 h, 48 seq; Semantic fast-forward, first-person video. Paper Site

  • HUJI-EgoSeg (2014) β€” 29 long egocentric videos (~1–5 h each), pixel-level temporal segmentation annotations; memory, summarization & long-form understanding. Paper Site

  • UT Ego (2012) β€” ~17 h, 4 long videos; Summarization, long-form ego. Paper Project

  • VINST / Visual Diaries (2011) β€” 31 egocentric videos capturing daily commutes; used for temporal segmentation and video summarization; memory, summarization & long-form understanding. Paper Site

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoMonthMonth-level schema, episodic, spatial, and cross-day memory reasoningEgoMonthHFDataset+benchmark
MEMORA-BenchEmbodied action memory formation, consolidation, retrieval, and planningEPIC-KITCHENS-100 extensionPaperStandalone
EgoServeProactive assistance over instant, short-term, episodic, and long-term contextEgoLife / HoloAssist / CaptainCook4DSiteStandalone
SuperMemory-VQALong-horizon egocentric memory VQA with answerability checksSuperMemory-VQAHFDataset+benchmark
EgoMemReasonWeek-long entity, event, and behavior memory reasoningEgoLifeHFStandalone
EgoExoMemCross-view memory reasoning over synchronized ego-exo videosEgoExo4D / LEMMAGitHubStandalone
EgoIntrospectUser internal-state reasoning over multimodal egocentric streamsEgoIntrospectSiteDataset+benchmark
MA-EgoQAMulti-agent egocentric video QA over week-long streamsEgoLifeSiteStandalone
EgoSchemaLong-form video-language understandingEgo4DSiteStandalone
EgoStreamStreaming episodic memory diagnosisEgo4D / EgoLife / EgoTempo / Multi-Hop EgoQA / HD-EPICSiteStandalone

πŸ’¬ VLMs, Instructions & QA

These datasets target egocentric video-language pretraining, instruction following, dialog, and question answering over first-person streams.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ HD-EPIC2025~41 h / dense labelsFine-grained kitchen, VQAPaperSite
⭐ EgoClip20223.8M clip–text pairsVideo-language pretrainingPaperGitHub
CrossView2026~6K questions / 4 domainsMulti-camera video QAPaperSite
EgoCross2026798 clips / 957 QACross-domain egocentric VideoQAPaperSite
HumanCLAW-Bench20261,218 episodes / 41 scenesClosed-loop embodied action intelligencePaperSite
EgoSafe-Bench20263K clips / 12K evaluation samplesVisual safety, forensic reasoningPaperN/A
VIABench2026761 videos / 46.9 h / 14.5K annotationsVisual-impairment assistancePaperGitHub
LongEgoRefer20261,498 refs / avg. 45 minLong-form video REC, groundingPaperGitHub
EgoGapBench20261,000 action-selection itemsEgocentric action selectionPaperGitHub
EgoSafetyBench20261,200 robot-view scenariosStreaming safety guardsPaperN/A
EgoSAT20261,997 videos / 165 h / 4.8K QAStreaming interaction understandingPaperSite
EgoTL2026100+ daily household tasksLong-horizon reasoning, spatial QAPaperSite
MM-Conv20266.7 h / 4,211 expressionsContext-aware 3D dialogue groundingPaperN/A
Minerva-Ego20261,160 QA / 156 videosSpatiotemporal reasoning tracesPaperGitHub
EgoEverything2026100+ h / 5K+ MCQsLong-context AR VideoQAPaperN/A
EgoEsportsQA20261,745 QA pairsEsports VideoQA, reasoningPaperN/A
GameplayQA2026~2.4K QA / 15 task categoriesPOV-synced multi-video gameplay QAPaperSite
MyEgo2026541 long videos / 5K QAPersonalized VideoQA, ego-groundingPaperGitHub
LifeDialBench / EgoMem2026EgoMem + LifeMem (coming soon)Lifelog memory, online evalPaperGitHub
Ego2Web2026500 video-instruction pairsEgocentric video-grounded web agentsPaperHF
NoRA20261,420 clips / support graphsNormative action reasoningPaperN/A
Causal-Plan-1M20261M QA / 22.2K clips / 770+ hPhysically grounded planning QAPaperHF
Pause-and-Think202610K QA clips + 300-sample benchmarkAssistive action suggestionPaperGitHub
EgoCoT-Bench2026351 videos / 3,172 QAGrounded operation-centric CoT QAPaperSite
EgoEMS202520+ h emergency scenariosEMS QA, multimodalPaperGitHub
HowToDIV2025~24 h instructionalDialog, procedural QAPaperGitHub
InterVLA202511.4 h interactionsInstruction, ego–exo mocapPaperSite
AssistQ2022100 long videos / 529 QAInstructional QAPaperGitHub
EgoTaskQA2022~2K videos / 40K QACausal & task QAPaperSite
EgoVQA2019600+ QAsVideo QAPaperN/A

Entries

  • [⭐️] HD-EPIC (2025) β€” ~41 h of densely labeled cooking video with recipe steps, audio events, gaze, 3D grounding, and VQA supervision. arXiv Site Code

  • [⭐️] EgoClip (2022) β€” 3.8M clip–text pairs; Video-language pretraining. arXiv Code

  • CrossView (2026) β€” ~6K multi-camera video questions across autonomous driving, surveillance, ego/exo activity, and robotics, including 4–7 synchronized Ego-Exo4D cameras per egocentric question. arXiv Site Code πŸ€—

  • EgoCross (2026) β€” 798 clips and 957 human-verified QA pairs across surgery, industrial assembly, extreme sports, and animal perspectives, designed to test cross-domain generalization beyond daily-life egocentric video. arXiv Site πŸ€—

  • HumanCLAW-Bench (2026) β€” 1,218 long-horizon find–navigate–interact episodes across 41 simulated indoor scenes for evaluating whether VLMs can select and sequence actions from a continuously updated egocentric body view. arXiv Site Code

  • EgoSafe-Bench (2026) β€” 12K evaluation samples formed from 3K mobile-captured first-person clips and hierarchical QA chains for feature anchoring, blind-spot deduction, intent inference, and logically consistent visual-safety reasoning. arXiv

  • VIABench (2026) β€” 761 first-person videos (46.9 h) recorded or shared by blind individuals with 14,526 curated annotations for proactive reminders, visual question answering, and vision-guided interaction in online and offline settings. arXiv Code

  • LongEgoRefer (2026) β€” 1,498 referring expressions over long-form Ego4D videos averaging 45 minutes, requiring temporal and spatial localization of sparse referred objects in untrimmed egocentric recordings. arXiv Code

  • EgoGapBench (2026) β€” 1,000 egocentric action-selection items from multi-agent scenes without first-person body cues, designed to test whether VLMs choose actions from the correct self/other perspective. arXiv Code

  • EgoSafetyBench (2026) β€” 1,200 egocentric robot-view scenarios annotated at half-second granularity, split into situational and visual-channel tracks for evaluating runtime VLM safety guards. arXiv

  • EgoSAT (2026) β€” 1,997 Ego4D videos (165 h) with ~4.8K QA pairs for streaming egocentric interaction understanding, covering retrospective, present, and prospective reasoning under partial observability. arXiv Site Code πŸ€—

  • EgoTL (2026) β€” 100+ daily household tasks with think-aloud chains, navigation and manipulation annotations, and metric spatial labels for long-horizon egocentric reasoning and QA. arXiv Site

  • MM-Conv (2026) β€” 6.7 h of egocentric VR interaction with synchronized speech, motion, gaze, 3D scene geometry, and 4,211 verified referring expressions for context-aware conversational grounding; full public release is described as following camera-ready. arXiv

  • Minerva-Ego (2026) β€” 1,160 hand-crafted multiple-choice questions over 156 HD-EPIC egocentric videos, paired with dense spatiotemporal human reasoning traces and object masks for diagnosable video reasoning. arXiv Code

  • EgoEverything (2026) β€” 100+ h of AR egocentric video with 5K+ multiple-choice questions for human-behavior-inspired long-context video understanding. arXiv

  • EgoEsportsQA (2026) β€” 1,745 expert QA pairs from first-person shooter esports videos for testing fast virtual-scene perception and tactical reasoning. arXiv

  • GameplayQA (2026) β€” ~2.4K QA pairs across 15 task categories and 3 cognitive levels for decision-dense, POV-synced multi-video understanding of multi-agent 3D gameplay, with densely labeled first-person player viewpoints (1.22 labels/sec); egocentric viewpoints are virtual/synthetic. ACL 2026. arXiv Site Code πŸ€—

  • MyEgo (2026) β€” 541 long egocentric videos with 5K personalized questions about the camera wearer, their activities, and past context. arXiv Code

  • LifeDialBench / EgoMem (2026) β€” Lifelog memory benchmark with EgoMem built from real-world egocentric videos and an online evaluation protocol; the project page currently marks the dataset and scripts as coming soon. arXiv Code

  • Ego2Web (2026) β€” 500 egocentric video-instruction pairs for evaluating web agents that must ground real-world first-person visual context before executing online tasks. arXiv Site Code πŸ€—

  • NoRA (2026) β€” 1,420 Ego4D-derived clips (190 human-verified HumanGold + 1,230 LLM-validated LLMSilver) with fact–reason–action support-graph annotations for benchmarking grounded normative action reasoning in VLMs. arXiv

  • Causal-Plan-1M (2026) β€” 1M QA pairs with four-stage causal reasoning-trace annotations over 22,201 egocentric clips (770+ h curated from Ego4D, EPIC-KITCHENS, HoloAssist, MECCANO, and others), plus the 1,200-instance Causal-Plan-Bench for physically grounded embodied planning. arXiv Code πŸ€—

  • Pause-and-Think (2026) β€” 10,051 reasoning-annotated training QA clips and a 300-sample benchmark curated from EPIC-KITCHENS, Ego4D, and Assembly101 with structured thinking/answer supervision for video-grounded assistive action suggestion and goal planning. arXiv Code

  • EgoCoT-Bench (2026) β€” 3,172 verifiable QA pairs over 351 first-person videos (sourced from Ego4D, EPIC-KITCHENS, Charades-Ego, MECCANO, and HD-EPIC) with step-by-step rationale annotations for evaluating grounded operation-centric chain-of-thought reasoning in MLLMs. arXiv Site Code πŸ€—

  • EgoEMS (2025) β€” 20+ h emergency scenarios; EMS QA, multimodal. arXiv Site Code

  • HowToDIV (2025) β€” ~24 h instructional; Dialog, procedural QA. arXiv Site Code

  • InterVLA (2025) β€” 11.4 h interactions; Instruction, ego–exo mocap. arXiv Site

  • AssistQ (2022) β€” 100 long videos / 529 QA; Instructional QA. arXiv Code

  • EgoTaskQA (2022) β€” ~2K videos / 40K QA; Causal & task QA. arXiv Site Code

  • EgoVQA (2019) β€” 600+ QAs; Video QA. Paper

  • EgoSchema β€” see Memory, Summarization & Long-form Understanding

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
CrossViewMulti-camera evidence integration across four domainsEgo-Exo4D / nuScenes / MEVA / AgiBotSiteStandalone
EgoCrossCross-domain egocentric VideoQA in surgery, industry, sports, and animal viewsEgoCrossSiteDataset+challenge
HumanCLAW-BenchClosed-loop find–navigate–interact action intelligenceHSSD simulated scenesSiteStandalone
EgoSafe-BenchHierarchical visual-safety and forensic reasoningEgoSafe-BenchPaperDataset+benchmark
VIABenchProactive reminder, VQA, and vision-guided assistanceVIABenchGitHubDataset+benchmark
LongEgoReferLong-form egocentric spatiotemporal referring expression comprehensionEgo4DGitHubStandalone
EgoGapBenchEgocentric action selection in multi-agent scenesEgoGapBenchGitHubStandalone
EgoSafetyBenchRuntime safety-guard evaluation over egocentric robot-view videoEgoSafetyBenchPaperStandalone
EgoSATRetrospective, present, and prospective streaming VideoQAEgo4DSiteStandalone
EgoTL-BenchLong-horizon planning, action reasoning, perceptual-metric understandingEgoTLSiteDataset+benchmark
MM-ConvContext-aware grounding in spontaneous 3D dialogueMM-ConvPaperDataset+benchmark
Minerva-EgoEgocentric multi-step VideoQA with spatiotemporal reasoning tracesHD-EPICGitHubStandalone
EgoEverythingLong-context AR VideoQAEgoEverythingPaperStandalone
EgoEsportsQAEsports VideoQA and tactical reasoningEgoEsportsQAPaperStandalone
GameplayQADecision-dense POV-synced multi-video gameplay QAGameplayQASiteDataset+benchmark
MyEgoPersonalized egocentric VideoQAMyEgoGitHubDataset+benchmark
LifeDialBench / EgoMemLifelog memory and online evaluationEgoMemGitHubDataset+benchmark
Ego2WebEgocentric video-grounded web-agent executionEgo2WebSiteDataset+benchmark
EgoExoBenchCross-perspective video understanding in MLLMsEgo-Exo4D / LEMMA / EgoExoLearn / TF2023 / EgoMe / CVMHATGitHubStandalone
EgoBabyVLM / Machine-DevBenchCross-modal learning from naturalistic egocentric videoNaturalistic infant/adult ego video corporaPaperStandalone
EgoEMSEMS QA, multimodal assessmentEgoEMSGitHubDataset+benchmark
HD-EPICFine-grained kitchen understanding, VQAHD-EPICSiteDataset+benchmark
HowToDIVMulti-turn instructional dialog & QAHowToDIVGitHubStandalone
InterVLAInstruction following, interaction understandingInterVLASiteDataset+benchmark
AssistQInstructional affordance-centric QAAssistQGitHubStandalone
EgoTaskQACausal, predictive, explanatory, counterfactual QAEgoTaskQASiteStandalone
EgoVQAEgocentric video QAEgoVQAOpen Access (ICCVW 2019 / EPIC)Standalone
NoRAGrounded normative action reasoningEgo4DPaperStandalone
Causal-Plan-BenchPhysically grounded embodied planning QACausal-Plan-1MHFDataset+benchmark
Pause-and-ThinkVideo-grounded assistive action suggestionEPIC-KITCHENS / Ego4D / Assembly101GitHubDataset+benchmark
EgoCoT-BenchGrounded operation-centric chain-of-thought QAEgo4D / EPIC-KITCHENS / Charades-Ego / MECCANO / HD-EPICSiteStandalone

πŸƒ Action & Activity Recognition

Canonical action-recognition, activity-analysis, affect, and interaction datasets built around first-person human behavior live here.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ego4D2022~3,670 hAR, VQA, forecasting, manyPaperSite
⭐ EPIC-KITCHENS-1002021100 h / 90K segmentsAction recognition, manyPaperSite
HUI360202671 h captured / 11 h filtered / 1M annotationsHuman-robot interaction anticipationPaperSite
EventKitchen20265.5 h / 10.8K action segmentsEvent-based action recognition, detectionPaperSite
InterPet4D20266.8M frames / 13 dogs / 23 peopleHuman-pet interaction, motion generationPaperSite
EgoPolice2026180+ h / 9 action classesBody-camera action recognition, VideoQAPaperGitHub
ChildLens2026108.58 h / 354 videosChild activity analysisPaperData
EgoScale202620k+ h labeled manipulation (paper)Action, dexterous transferPaperN/A
CogDrive (EyeCue)2026Multi-scenario driving clipsDriver cognitive-distraction detection, gaze + egoPaperN/A
Ego-METAS2026100+ h / 5 modalitiesOnline action segmentation, sensor routingPaperHF
Furhat Egocentric Dataset202620 seq. / ~25 minRobot-ego face/body tracking, re-IDPaperN/A
World In Your Hands20251000+ h labeled manipulation (paper)Action, dexterous transfer, VLA trainingPaperGitHub
EgoCampus2025~32 h / campus paths (paper)Gaze, pedestrian egoPaperGitHub
AEA2024143 seq. / ~7.3 hEveryday activities, AriaPaperSite
EgoSurgery (Phase / Tool / HTS)2024Open-surgery ego videoPhase, tools, segmentationPhase / Tool / HTSGitHub
EgoExo-Fitness202432 h / 1,276 seq.Full-body action, quality assessmentPaperGitHub
EΒ³ (Exploring Embodied Emotion)202450+ hEmotion, multimodal egoPaperGitHub
EGOFALLS2023Fall samples / AVFall detectionPaperSite
Epic-Sounding-Object20233.2K short clipsAudio-visual localizationPaperGitHub
HoloAssist2023169 hInteractive assistantsPaperSite
WEAR2023~19 h outdoor sportsActivity + IMUPaperSite
N-EPIC-KITCHENS2022Event + RGB subsetAction, neuromorphicPaperGitHub
Ego-Deliver20215,360 videosDelivery ego analysisN/ASite
HOMAGE202130 h / compositionalHome activitiesPaperSite
MECCANO2021~55 h industrialHOI, egoPaperSite
EGO-CH202027+ h, cultural sitesVisitor behavior, POI tasksPaperN/A
EgoCom202038.5 h conversationMultiperson ego dialogPaperGitHub
LEMMA2020Multi-view activitiesMulti-agent tasksPaperSite
Charades-Ego2018Paired ego / exoAlignment, actionsPaperSite
EGTEA Gaze+201828 h cookingGaze + action recognitionPaperSite
EgoGesture20172K+ videos / 24K samplesGesture recognitionPaperSite
Stanford ECM201731 hours, augmented with heart rate and accelerometer, 23–24 daily activity categoriesaction & activity recognitionPaperN/A
THU-READ20171,920 clips (8 subjects Γ— 40 actions Γ— 3 reps Γ— 2 modalities), RGB-D from helmet-mounted sensoraction & activity recognitionPaperSite
PEV (UTokyo Paired Ego-Video)20161,226 pairs of first-person clips, synchronous dyadic conversations, 8 interaction categories, 6 subjectsaction & activity recognitionPaperN/A
FPPA20155 subjects, 5 daily actions, egocentric video with hand and gaze cuesaction & activity recognitionPaperN/A
JPL-Interaction201384 videos, 7 activity types (4 positive, 1 neutral, 2 negative interactions), 320Γ—240@30 fpsaction & activity recognitionPaperSite
ADL2012~10 h, 20 participantsADL recognition, objectsPaperSite
Social Interactions20128 social events, ~60 hours, head-mounted cameras, multiple participants per eventaction & activity recognitionPaperSite
EgoAction2011First-person sports videos (skateboarding, skiing, cycling, etcaction & activity recognitionPaperN/A

Entries

  • [⭐️] Ego4D (2022) β€” ~3,670 h; AR, VQA, forecasting, many. arXiv Site Code

  • [⭐️] EPIC-KITCHENS-100 (2021) β€” 100 h of unscripted kitchen activity with 90K segments; the canonical egocentric action-recognition and anticipation benchmark. Paper Site Code

  • HUI360 (2026) β€” 71 h of in-the-wild 360Β° robot-egocentric capture (11 h retained for the benchmark) with more than 1M curated pose, face-keypoint, mask, tracking, and interaction annotations for human-robot interaction anticipation. arXiv Site

  • EventKitchen (2026) β€” 5.5 h of unscripted cooking from 10 participants in 13 kitchens, captured with stereo event cameras plus synchronized RGB, depth, and IMU and annotated with 10,762 action segments and 13,482 object boxes. arXiv Site

  • InterPet4D (2026) β€” 6.8M synchronized multi-view and egocentric frames from 13 dogs of 11 breeds interacting with 23 people, with audio, segmentation, 2D/3D keypoints, and human, hand, and pet meshes. arXiv Site πŸ€—

  • EgoPolice (2026) β€” 180+ h of real police body-worn camera footage with second-by-second annotations for nine high-stakes police/civilian action classes, classification folds, and multiple-choice VideoQA. arXiv Code

  • ChildLens (2026) β€” 108.58 h of child-worn egocentric video and audio from 62 children for activity analysis in everyday home behavior. Paper Site Data

  • EgoScale (2026) β€” 20k+ h labeled manipulation (paper); Action, dexterous transfer. arXiv Site

  • CogDrive (EyeCue) (2026) β€” Augmented multi-scenario driving dataset paired with the EyeCue gaze-empowered ego-video framework for driver cognitive-distraction detection (reported 74.38% accuracy). arXiv

  • Ego-METAS (2026) β€” 100+ h of untrimmed multimodal egocentric video (RGB, audio, gaze, IMU, monochrome) curated from Ego-Exo4D, CMU-MMAC, and CaptainCook4D with unified splits, pre-extracted features, and baseline sensor-routing policies for online energy-efficient temporal action segmentation. arXiv Site πŸ€—

  • Furhat Egocentric Dataset (2026) β€” 20 sequences (~25 min) of close-range multi-party interactions captured from a Furhat social robot's egocentric camera with amodal face/body boxes and consistent identities for multi-person tracking and re-identification in HRI; available on request under a Data Usage Agreement. arXiv

  • World In Your Hands (2025) β€” 1000+ h of labeled human manipulation data with video-language annotations for action understanding, dexterous transfer, and VLA training. arXiv Site Code

  • EgoCampus (2025) β€” ~32 h / campus paths (paper); Gaze, pedestrian ego. arXiv Site Code

  • AEA (2024) β€” 143 seq. / ~7.3 h; Everyday activities, Aria. arXiv Site πŸ€—

  • EgoSurgery (Phase / Tool / HTS) (2024) β€” Open-surgery ego video with companion phase, tool, and hand-tool segmentation releases for phase recognition, instrument analysis, and dense surgical interaction understanding. arXiv arXiv arXiv Code

  • EgoExo-Fitness (2024) β€” 32 h of synchronized egocentric and exocentric fitness video with temporal boundaries, sub-step annotations, action comments, and quality scores for full-body action understanding. arXiv Code πŸ€—

  • EΒ³ (Exploring Embodied Emotion) (2024) β€” 50+ h; Emotion, multimodal ego. Paper Code

  • EGOFALLS (2023) β€” Fall samples / AV; Fall detection. arXiv Site

  • Epic-Sounding-Object (2023) β€” 3.2K short clips; Audio-visual localization. Paper Code

  • HoloAssist (2023) β€” 169 h; Interactive assistants. Paper Site Code

  • WEAR (2023) β€” ~19 h outdoor sports; Activity + IMU. arXiv Site

  • N-EPIC-KITCHENS (2022) β€” Event + RGB subset; Action, neuromorphic. Paper Site Code

  • Ego-Deliver (2021) β€” 5,360 videos; Delivery ego analysis. Site

  • HOMAGE (2021) β€” 30 h / compositional; Home activities. Paper Site Code

  • MECCANO (2021) β€” ~55 h industrial; HOI, ego. Paper Site Code

  • EGO-CH (2020) β€” 27+ h, cultural sites; Visitor behavior, POI tasks. arXiv

  • EgoCom (2020) β€” 38.5 h conversation; Multiperson ego dialog. Paper Code

  • LEMMA (2020) β€” Multi-view activities; Multi-agent tasks. arXiv Site Code

  • Charades-Ego (2018) β€” Paired ego / exo; Alignment, actions. arXiv Site

  • EGTEA Gaze+ (2018) β€” 28 h cooking; Gaze + action recognition. Paper Site

  • EgoGesture (2017) β€” 2K+ videos / 24K samples; Gesture recognition. Paper Site

  • Stanford ECM (2017) β€” 31 hours, augmented with heart rate and accelerometer, 23–24 daily activity categories; action & activity recognition. Paper

  • THU-READ (2017) β€” 1,920 clips (8 subjects Γ— 40 actions Γ— 3 reps Γ— 2 modalities), RGB-D from helmet-mounted sensor; action & activity recognition. Paper Site

  • PEV (UTokyo Paired Ego-Video) (2016) β€” 1,226 pairs of first-person clips, synchronous dyadic conversations, 8 interaction categories, 6 subjects; action & activity recognition. Paper

  • FPPA (2015) β€” 5 subjects, 5 daily actions, egocentric video with hand and gaze cues; action & activity recognition. Paper

  • JPL-Interaction (2013) β€” 84 videos, 7 activity types (4 positive, 1 neutral, 2 negative interactions), 320Γ—240@30 fps; action & activity recognition. Paper Site

  • ADL (2012) β€” ~10 h, 20 participants; ADL recognition, objects. Paper Site

  • Social Interactions (2012) β€” 8 social events, ~60 hours, head-mounted cameras, multiple participants per event; action & activity recognition. Paper Site

  • EgoAction (2011) β€” First-person sports videos (skateboarding, skiing, cycling, etc; action & activity recognition. Paper

  • Ego-Exo4D β€” see 3D Scene Understanding & Localization

  • EgoExoLearn β€” see Procedural Activities & Skill Learning

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
HUI360In-the-wild human-robot interaction anticipation and cross-dataset transferHUI360 / SSUP-HRISiteDataset+benchmark
EventKitchenEvent-based action recognition, object detection, and stereo depthEventKitchenSiteDataset+benchmark
InterPet4DMultimodal human-pet interaction and pet-motion generationInterPet4DSiteDataset+benchmark
EgoPoliceHigh-stakes action classification and body-camera VideoQAEgoPoliceGitHubDataset+benchmark
EΒ³ (Exploring Embodied Emotion)Emotion recognition, classification, localization, reasoningEΒ³GitHubStandalone
EGOFALLSFall detection (visual + audio)EGOFALLSDataverseStandalone
EgoExo-FitnessAction localization, cross-view verification, skill determinationEgoExo-FitnessGitHubDataset+benchmark
Ego4DAction, forecasting, VQA, narration, …Ego4DSiteSuite
EPIC-KITCHENS-100Action recognition, detection, anticipation, …EPIC-KITCHENS-100SiteSuite
EgoGestureEgocentric hand gesture recognitionEgoGestureSiteStandalone
Ego-METASOnline energy-efficient temporal action segmentationEgo-Exo4D / CMU-MMAC / CaptainCook4DHFStandalone

βœ‹ Hand–Object Interaction, Dexterity & 3D

This section focuses on egocentric hands, dexterous manipulation, object interaction, tracking, and dense 3D understanding around the body and manipulated objects.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ HOI4D20222.4M frames / 4K seq.4D HOIPaperSite
⭐ EgoDex2025829 h / 30K trajectoriesDexterous manipulation, posePaperGitHub
EgoAffordance2026204K episodes / 17.2M affordancesVisual, grasp, trajectory affordancesPaperSite
H-Tac2026160 h / 135K episodesTactile-action pretrainingPaperN/A
EPIC-Contact20262.3K clips / 62.3K framesIn-the-wild 3D hand-object contactPaperSite
HT-Bench202610M RGB / 7.8M tactile framesFull-hand tactile representationPaperN/A
ForceBand202610 h multimodal force demossEMG-to-force, forceful manipulationPaperSite
EventEgoHands202648 clips / 129.6K framesRGB+event hand detectionPaperGitHub
EgoTactile2026~6 h / 768 clips / 63 objectsGrasp pressure from ego videoPaperSite
EgoDex-R20264.3M RGB-D frames / 5.6K seq.Dexterous manipulation, hand-object posePaperN/A
HA-Ego-1K2026~24 h / 484 multi-view clipsManipulation, 6-cam ego + IMUN/AHF
DexGloveHOI20263.5 h / 100K+ samplesVision-IMU 3D hand trackingPaperN/A
EgoTouch20261,891 episodes / 208 tasksTactile HOI, vision-to-touchPaperHF
EgoEVHands20265,419 annotated seq.Stereo event 3D hand pose, gesturePaperGitHub
EgoEMG202641 participants / 10+ hEMG + vision hand posePaperGitHub
HRDexDB20261.4K grasping trialsDexterous grasping, tactile, ego streamsPaperHF
TouchMoment20264,021 videos / 8,456 touch momentsContact moment detectionPaperN/A
EgoFun3D2026271 egocentric videosInteractive 3D objects, function templatesPaperSite
SHOW3D2026In-the-wild ego-exo HOI3D hand-object annotationsPaperN/A
FEEL2026Force-sync kitchen ego videoPhysical action understandingPaperSite
EgoPoints2025Point tracks + syntheticTracking in ego videoPaperGitHub
AssemblyHands20233M images / hands3D hand pose, assemblyPaperSite
EgoObjects20239.2K+ videosDetection, instance segPaperGitHub
ENIGMA-51202322 h industrialFine-grained behaviorPaperSite
POV-Surgery2023~88K frames, 53 seq. (synth.)Surgical hand–tool pose, segmentationPaperSite
VOST2023713 videosVOS, transforming objectsPaperSite
EgoBody2022125 seq. / multi-viewBody pose, interactionPaperSite
EgoHOS202211K+ imagesHand–object segmentationPaperGitHub
EgoPAT3D20221M+ frames RGB-D3D action target predictionPaperSite
Touch and Go202212K+ vis–tactile framesVision + touchPaperSite
VISOR2022EPIC + masks / relationsSegmentation, HOIPaperSite
H2O2021100K+ framesTwo-hand interactionPaperSite
TREK-1502021150 EPIC seq.Object trackingPaperSite
You2Me202014 seq., chest-mounted GoProBody pose via ego–exo interactionPaperGitHub
FPHA20181.2K seq. hand actionHand pose + actionPaperSite
EgoDexter2017~3.2K frames, 4 seq.Hand tracking under occlusionPaperSite
EgoHands20154.8K labeled framesHand detection / boxesPaperSite
BEOID201458 videos, 6 environments, 34 object interaction classes, ~30 fpshand–object interaction, dexterity & 3dPaperData
EDSH20132 videos (~5 min each), pixel-level hand segmentation, egocentric daily activitieshand–object interaction, dexterity & 3dPaperSite
Handled Objects200911 object categories, multiple grasp sequences, RGB + depth from wearable camerahand–object interaction, dexterity & 3dPaperN/A

Entries

  • [⭐️] HOI4D (2022) β€” 2.4M RGB-D frames with object poses, hand poses, interaction regions, and motion segmentation for category-level 4D HOI. Paper Site

  • [⭐️] EgoDex (2025) β€” 829 h / 30K trajectories; Dexterous manipulation, pose. arXiv Site Code

  • EgoAffordance (2026) β€” 204K egocentric manipulation episodes with 5.6M visual affordances and 11.6M grasp and trajectory affordances, automatically extracted in a shared 3D actionable representation for VLAff and robot transfer. arXiv Site

  • H-Tac (2026) β€” 160 h of egocentric human videos with tactile/action data across 300+ tasks and 135K episodes, introduced for human-centric transferable tactile-action pretraining and future tactile prediction. arXiv

  • EPIC-Contact (2026) β€” 2.3K in-the-wild EPIC-KITCHENS stable-grasp clips (62.3K frames) with dense bijective 3D hand-object contact correspondences and posed hand/object meshes for unconstrained 3D HOI pose estimation. arXiv Site Code πŸ€—

  • HT-Bench (2026) β€” Large-scale benchmark pairing egocentric vision with full-hand tactile sensing, comprising 10M RGB frames and 7.8M tactile frames across 226 tasks for tactile retrieval, inpainting, vision-to-touch synthesis, and multimodal prediction. arXiv

  • ForceBand (2026) β€” 10 h multimodal dataset with egocentric video, wrist sEMG, IMU, and fingertip force measurements across diverse everyday objects/actions, used to learn EMG-to-force labels for force-augmented robot demonstrations; public dataset release is marked as coming soon. arXiv Site

  • EventEgoHands (2026) β€” 48 egocentric clips (~1.2 h, 129.6K frames) pairing RGB with synthetic event streams (synthesized from EgoHands via v2e) and 393K hand bounding boxes for RGB-event hand detection under motion blur and low light. arXiv Code

  • EgoTactile (2026) β€” ~6 h (768 clips, 319K frames) of head- and neck-mounted egocentric video of 12 participants grasping 63 everyday objects with synchronized 162-taxel pressure-glove supervision and a bare-hand transfer subset for full-hand grasp pressure estimation; the dataset currently sits under an anonymous ICML-submission account. arXiv Site πŸ€—

  • EgoDex-R (2026) β€” 4.3M egocentric RGB-D frames across 5,600 manipulation sequences (1,000+ objects, 200+ daily task categories) with MANO hand poses, 6-DoF object trajectories, reconstructed meshes, and contact annotations, introduced in the EgoAERO paper; distinct from Apple's EgoDex. arXiv

  • HA-Ego-1K (2026) β€” ~24 h of privacy-redacted six-camera + IMU egocentric video (484 multi-view clips across 22 real-world work scenarios such as workshops, construction, and factories) captured with the head-worn Human Archive GSI Cap for dexterous-manipulation and long-horizon task research; gated access (CC BY-NC 4.0), no paper yet. Site πŸ€—

  • DexGloveHOI (2026) β€” 3.5 h / 100K+ synchronized egocentric vision-IMU samples with MoCap 3D hand-pose ground truth for dexterous hand-object interaction tracking; no official public data page was found. arXiv

  • EgoTouch (2026) β€” 1,891 bimanual hand-object interaction episodes across 208 manipulation tasks with synchronized egocentric and wrist RGB video, 3D hand pose, and dense tactile pressure maps. arXiv Site Code πŸ€—

  • EgoEVHands (2026) β€” 5,419 real-world stereo event-camera egocentric sequences with dense 2D/3D hand keypoints across 38 gesture classes; the official repository currently says code, models, and dataset links are to be uploaded. arXiv Code

  • EgoEMG (2026) β€” 10+ h of synchronized bilateral EMG, IMU, egocentric RGB, external RGB-D, and mocap-derived hand pose across 41 participants and 60 gesture classes. arXiv Code

  • HRDexDB (2026) β€” 1.4K dexterous human and robotic hand grasping trials with synchronized multi-view video, egocentric video streams, tactile signals, and 3D motion. arXiv πŸ€—

  • TouchMoment (2026) β€” 4,021 egocentric videos with 8,456 annotated hand-object contact moments for frame-precise touch detection. arXiv

  • EgoFun3D (2026) β€” 271 egocentric interaction videos with paired 3D geometry, 2D/3D segmentation, articulation labels, and function-template annotations. arXiv Site πŸ€—

  • SHOW3D (2026) β€” In-the-wild ego-exo capture of hands interacting with objects, with 3D hand-object annotations from a marker-less multi-camera system. arXiv

  • FEEL (2026) β€” Force-sync kitchen ego video; Physical action understanding. arXiv Site

  • EgoPoints (2025) β€” Point tracks + synthetic; Tracking in ego video. arXiv Site Code

  • AssemblyHands (2023) β€” 3M egocentric hand images on top of Assembly101 for detailed 3D hand pose estimation during assembly. Paper Site Code

  • EgoObjects (2023) β€” 9.2K+ videos; Detection, instance seg. Paper Site Code

  • ENIGMA-51 (2023) β€” 22 h industrial; Fine-grained behavior. arXiv Site Code

  • POV-Surgery (2023) β€” ~88K frames, 53 seq. (synth.); Surgical hand–tool pose, segmentation. arXiv Site Code

  • VOST (2023) β€” 713 videos; VOS, transforming objects. arXiv Site

  • EgoBody (2022) β€” 125 seq. / multi-view; Body pose, interaction. arXiv Site

  • EgoHOS (2022) β€” 11K+ images; Hand–object segmentation. arXiv Code

  • EgoPAT3D (2022) β€” 1M+ frames RGB-D; 3D action target prediction. Paper Site Code

  • Touch and Go (2022) β€” 12K+ vis–tactile frames; Vision + touch. arXiv Site Code

  • VISOR (2022) β€” EPIC + masks / relations; Segmentation, HOI. arXiv Site Code

  • H2O (2021) β€” 100K+ frames; Two-hand interaction. Paper Site Code

  • TREK-150 (2021) β€” 150 EPIC seq; Object tracking. arXiv Site Code

  • You2Me (2020) β€” 14 seq., chest-mounted GoPro; Body pose via ego–exo interaction. Paper arXiv Code

  • FPHA (2018) β€” 1,175 RGB-D sequences with 3D hand pose and action labels; a foundational first-person hand-action benchmark. Paper Site Code

  • EgoDexter (2017) β€” ~3.2K frames, 4 seq; Hand tracking under occlusion. arXiv Project

  • EgoHands (2015) β€” 4.8K labeled frames; Hand detection / boxes. Paper Project

  • BEOID (2014) β€” 58 videos, 6 environments, 34 object interaction classes, ~30 fps; hand–object interaction, dexterity & 3d. Paper Data

  • EDSH (2013) β€” 2 videos (~5 min each), pixel-level hand segmentation, egocentric daily activities; hand–object interaction, dexterity & 3d. Paper Site

  • Handled Objects (2009) β€” 11 object categories, multiple grasp sequences, RGB + depth from wearable camera; hand–object interaction, dexterity & 3d. Paper

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoAffordance / VLAffVisual, grasp, and trajectory affordance predictionEgoAffordanceSiteDataset+benchmark
H-TacHuman-to-robot tactile-action pretraining and future tactile predictionH-TacPaperDataset+pretraining resource
EPIC-Contact / HOPformerIn-the-wild egocentric 3D hand-object pose and contact estimationEPIC-ContactSiteDataset+benchmark
HT-BenchFull-hand tactile representation learning with egocentric visionHT-BenchPaperDataset+benchmark
ForceBand / EMG2ForcesEMG-to-fingertip-force prediction and forceful manipulation policy learningForceBandSiteDataset+benchmark
TouchMomentFrame-precise hand-object contact moment detectionTouchMomentPaperStandalone
EgoFun3DInteractive 3D object modeling and function-template inferenceEgoFun3DSiteDataset+benchmark
EgoEMGEMG-to-pose, vision-to-pose, and EMG+vision fusionEgoEMGGitHubDataset+benchmark
EgoTouch / TouchAnythingVision-to-touch prediction for bimanual HOIEgoTouchHFDataset+benchmark
DexGloveHOIVision-IMU 3D hand tracking under HOI occlusionDexGloveHOIPaperDataset+benchmark
EgoEVHandsStereo event 3D hand pose and gesture recognitionEgoEVHandsGitHubDataset+benchmark
AssemblyHandsEgocentric 3D hand poseAssembly101SiteStandalone
VISORVideo object segmentation, hand–object relationsEPIC-KITCHENSSiteStandalone
TREK-150Egocentric single-object trackingEPIC-KITCHENSSiteStandalone
EggHandEgocentric 3D hand pose forecastingEgoExo4DPaperMethod benchmark
FPHAHand action + 3D hand poseFPHASiteStandalone
EgoTactileFull-hand grasp pressure estimation from ego videoEgoTactileSiteDataset+benchmark
EventEgoHandsMultimodal RGB-event egocentric hand detectionEventEgoHands (from EgoHands)GitHubDataset+benchmark

πŸ“‹ Procedural Activities & Skill Learning

Datasets centered on step structure, instructional execution, assembly, or skill transfer from egocentric experience are grouped here.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ EgoExoLearn2024120 h ego+exoProcedural, async viewsPaperGitHub
⭐ Assembly1012022513 h multiviewAssembly, procedurePaperSite
EgoProceVQA20263,600 QA / 31 tasks / 4 scenariosKey-step procedural reasoningPaperSite
CoMind2026Dual ego + 2 exo views / 55 environmentsCollaborative activity, social reasoningPaperSite
VLK202648K synthetic paired trajectoriesHumanoid loco-manipulation, VLKPaperSite
EgoVerse20261,362 h / ~80K episodesRobot learning, manipulation skillsPaperSite
EgoLive2026Large-scale real-world task routinesRobot manipulation learningPaperN/A
EgoMAGIC20263,355 videos / 50 medical tasksField medicine, action detectionPaperZenodo
HumanEgo2026Minutes-per-task Aria demonstrationsHuman-to-robot policy learningPaperSite
EgoSPT202611,515 episodes / 112 task foldersSpatially prompted manipulation trajectoriesPaperHF
Ego-EXTRA202650 h / 15K+ VQAExpert-trainee assistancePaperSite
GM-1002026100+ tasks / 13K+ trajectoriesRobot manipulation, embodied evaluationPaperSite
SABER2026100+ h / 44.8K samplesRetail VLA adaptationPaperSite
EgoProactive / ProΒ²Bench2026700 recordings (22–55 min) / 42K eval instancesProactive procedural assistancePaperHF
EgoYC2 / Exo2EgoDVC2025~43 h cookingDense captioning, proceduralPaperGitHub
IndustReal2024~6 h industrialProcedure steps, errorsPaperSite
EgoProceL202262 videos / 16 tasksProcedure learningPaperSite
EPIC-Tent20197+ h, tent assemblyProcedural, dual HMD + gazePaperSite
CMU-MMAC201125 subjects, 5 cooking recipesprocedural activities & skill learningPaperSite
GTEA Gaze201117 meal preparation sessions, 7 cooking activities, gaze tracking annotationsprocedural activities & skill learningPaperSite

Entries

  • [⭐️] EgoExoLearn (2024) β€” 120 h ego+exo; Procedural, async views. Paper Site Code πŸ€—

  • [⭐️] Assembly101 (2022) β€” 513 h multiview; Assembly, procedure. Paper Site

  • EgoProceVQA (2026) β€” 3,600 key-step-centric questions across 31 everyday tasks and four procedural scenarios, covering six question types generated with EgoProceGen and human-checked for procedural reasoning evaluation. arXiv Site

  • CoMind (2026) β€” Collaborative cooking captured from two synchronized head-mounted cameras and two exocentric views, with audio, gaze, hand/object interactions, social cues, and aligned scans across 55 environments. arXiv Site

  • VLK (2026) β€” 48K synthetic vision-language-kinematics trajectories rendered as egocentric observations in reconstructed indoor 3DGS scenes, paired with language commands and whole-body humanoid kinematic trajectories for loco-manipulation. arXiv Site

  • EgoVerse (2026) β€” 1,362 h of egocentric human demonstrations spanning ~80K episodes and 1,965 tasks for robot learning from human manipulation experience. arXiv Site Code

  • EgoLive (2026) β€” Large-scale annotated egocentric recordings of real-world human task routines for robot manipulation learning. arXiv

  • EgoMAGIC (2026) β€” 3,355 egocentric field-medicine videos covering 50 medical tasks, with released medical training data and an action-detection challenge. arXiv Site

  • HumanEgo (2026) β€” Minutes-per-task human egocentric demonstrations collected with Aria glasses for zero-shot human-to-robot manipulation-policy learning via interaction-centric spatial representations. arXiv Site

  • EgoSPT (2026) β€” 11,515 processed egocentric manipulation episodes for spatially prompted visual trajectory prediction, with RGB video, end-effector poses, gripper widths, and valid-frame masks. arXiv πŸ€—

  • Ego-EXTRA (2026) β€” 50 h of unscripted expert-trainee egocentric procedural assistance across bike workshop, kitchen, bakery, and assembly scenarios, with dialogue transcripts and 15K+ VQA sets. Paper Site

  • GM-100 (2026) β€” 100+ detail-oriented robot manipulation tasks with 13K+ teleoperated trajectories and robot first-person camera views for embodied skill evaluation. arXiv Site Code

  • SABER (2026) β€” 100+ h of natural in-store retail activity with head-mounted egocentric video, 360-degree exocentric video, and 44.8K action samples for VLA adaptation. arXiv Site πŸ€—

  • EgoProactive / ProΒ²Bench (2026) β€” 700 Ray-Ban Meta smart-glasses recordings (22–55 min each) of cooking, crafts, DIY, and tutorial sessions with per-decision-point interrupt/silent labels and Out-of-Plan deviation-recovery annotations for proactive procedural assistance; ProΒ²Bench unifies five existing egocentric benchmarks into 42K evaluation and 250K training instances. arXiv πŸ€—

  • EgoYC2 / Exo2EgoDVC (2025) β€” ~43 h cooking; Dense captioning, procedural. arXiv Site Code

  • IndustReal (2024) β€” ~6 h industrial; Procedure steps, errors. Paper Site

  • EgoProceL (2022) β€” 62 videos / 16 tasks; Procedure learning. arXiv Site Code

  • EPIC-Tent (2019) β€” 7+ h, tent assembly; Procedural, dual HMD + gaze. Paper Site Code

  • CMU-MMAC (2011) β€” 25 subjects, 5 cooking recipes; procedural activities & skill learning. Paper Site

  • GTEA Gaze (2011) β€” 17 meal preparation sessions, 7 cooking activities, gaze tracking annotations; procedural activities & skill learning. Paper Site

  • Ego-Exo4D β€” see 3D Scene Understanding & Localization

  • HowToDIV β€” see VLMs, Instructions & QA

  • ADT (Aria Digital Twin) β€” see 3D Scene Understanding & Localization

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
EgoProceVQAKey-step procedural understanding across six QA typesEgoProceVQASiteDataset+benchmark
CoMindJoint attention, socially conditioned interaction anticipation, collaborative handoverCoMindSiteDataset+benchmark
VLKVision-language-kinematics policy learning for humanoid navigation and object transportSynthetic 3DGS trajectoriesSiteDataset+benchmark
HumanEgoZero-shot human-to-robot manipulation from egocentric videoHumanEgoSiteDataset+benchmark
EgoSPT / SP-VTPSpatially prompted visual trajectory prediction for manipulationEgoSPTHFDataset+benchmark
Ego-EXTRAExpert-trainee procedural assistance and VQAEgo-EXTRASiteDataset+benchmark
EgoMAGICField-medicine action detectionEgoMAGICZenodoDataset+benchmark
GM-100Detail-oriented robot manipulation evaluationGM-100SiteDataset+benchmark
TAVISActive-vision imitation learning on humanoid robots (GR1T2, Reachy2) in IsaacLab; TAVIS-Head + TAVIS-Hands suites with the GALT anticipatory-gaze metricSimulation-only (no real-data release)PaperStandalone benchmark
EgoProactive / ProΒ²BenchProactive intervention timing and Out-of-Plan recovery guidanceEgoProactive + Ego4D / EPIC-KITCHENS / Ego-Exo4D / HoloAssist / HowTo100MHFDataset+benchmark

πŸ—ΊοΈ 3D Scene Understanding & Localization

These datasets emphasize geometry, localization, scene graphs, multiview capture, or machine-perception tasks grounded in ego video.

Datasets at a glance

NameYearScaleKey tasksPaperLink
⭐ Ego-Exo4D20241,286+ h ego+exoSkilled activity, many tasksPaperSite
⭐ ADT (Aria Digital Twin)2023200 seq., 2 scenesEgocentric 3D perceptionPaperSite
GST-Bench / GST-Train20266,790 min synthetic videoGlobal spatial awareness from ego videoPaperN/A
FloAff-Kitchen2026Cross-scene, multi-view kitchen benchmarkNavigation-to-manipulation affordancePaperSite
EgoHTR202655 seq. / 150K+ frames / 7 scenes4D human-terrain reconstructionPaperSite
SG-Ego20263.8M graphs / 7.3K Ego4D videosSpatio-temporal scene graphsPaperHF
PRISM2026270K samples / 11.8M framesRetail embodied VLM, spatial reasoningPaperHF
EgoTraj202610.7 h / 1.15M framesEgocentric trajectory predictionPaperGitHub
AIST-Living2026Egocentric video + GT motion in scanned env.Global pose, localizationPaperSite
OVO-S-Bench2026348 videos / 1,680 Q / 30 task typesStreaming spatial intelligence QAPaperSite
PVSG2023400 vids, ~150K framesPanoptic video scene graph (ego + third-person)PaperSite
DR(eye)VE2018~6 h driving, 555K framesGaze prediction, driving ego videoPaperSite
EgoCart2018Retail RGB-D, 9 videosIndoor / cart localizationPaperSite
IU ShareView20189 paired ego video setsPerson seg / ID across synchronized wearersPaperSite
OST201757 sequences, 55 subjects, ~15 min/video, egocentric object search tasks, eye-tracking ground truth3d scene understanding & localizationPaperGitHub

Entries

  • [⭐️] Ego-Exo4D (2024) β€” 1,286+ h of paired first- and third-person skilled activity with multiview geometry and a broad benchmark suite. arXiv Site

  • [⭐️] ADT (Aria Digital Twin) (2023) β€” 200 seq., 2 scenes; Egocentric 3D perception. Paper Site πŸ€—

  • GST-Bench / GST-Train (2026) β€” Human-verified global-spatial-temporal questions derived from 6,790 minutes of synthetic first-person exploration, requiring novel-view inference and mapping ego observations onto global top-down scenes, plus a companion training set. arXiv

  • FloAff-Kitchen (2026) β€” Cross-scene, multi-view benchmark for predicting where a mobile robot should stand to execute downstream manipulation, spanning varied skills, layouts, furniture styles, and egocentric viewpoints. arXiv Site

  • EgoHTR (2026) β€” 55 scene-aligned 4D human-terrain traversal sequences (150K+ frames across seven challenging scenes) with ego/exo Aria video, SLAM, IMU, 3D scans, and parametrized human motion for analysis, synthesis, and humanoid locomotion transfer. arXiv Site

  • SG-Ego (2026) β€” Large-scale spatio-temporal scene-graph annotations extending Ego4D: SG-Ego-Align provides ~3.8M graphs from 7,297 videos, while SG-Ego-Edit adds action-conditioned graph-edit forecasting samples for A-GEF. arXiv Site Code πŸ€—

  • PRISM (2026) β€” 270K-sample multi-view retail video SFT corpus with egocentric, exocentric, and 360-degree views for embodied VLM spatial, physical, and action reasoning. arXiv Site πŸ€—

  • EgoTraj (2026) β€” 10.7 h / 1.15M frames of Meta Quest Pro egocentric urban navigation with synchronized RGB, 6DoF head pose, gaze, and scene annotations for trajectory forecasting; the GitHub README says the dataset and dashboard will be released after publication. arXiv Code

  • AIST-Living (2026) β€” Dataset introduced with Map-Mono-Ego that pairs monocular egocentric video with ground-truth human motion in a pre-scanned 3D environment for globally consistent pose estimation. arXiv Site

  • OVO-S-Bench (2026) β€” 1,680 fully human-annotated questions over 348 continuous egocentric streams (indoor walkthroughs, daily activities, outdoor tours, and driving from nine sources) spanning 30 task types across four hierarchical levels, from instantaneous perception to allocentric mapping, for streaming spatial intelligence in multimodal LLMs. arXiv Site Code πŸ€—

  • PVSG (2023) β€” 400 vids, ~150K frames; Panoptic video scene graph (ego + third-person). Paper arXiv Site Code

  • DR(eye)VE (2018) β€” ~6 h driving, 555K frames; Gaze prediction, driving ego video. arXiv Site Code

  • EgoCart (2018) β€” Retail RGB-D, 9 videos; Indoor / cart localization. Paper Site

  • IU ShareView (2018) β€” 9 paired ego video sets; Person seg / ID across synchronized wearers. Paper arXiv Site

  • OST (2017) β€” 57 sequences, 55 subjects, ~15 min/video, egocentric object search tasks, eye-tracking ground truth; 3d scene understanding & localization. Paper Code

  • Ego-1K β€” see Video Generation & World-Model Pretraining

Benchmarks built on these datasets

BenchmarkCapabilityPrimary dataOfficial linkNotes
GST-BenchGlobal spatial-temporal VQA and allocentric mapping from ego streamsGST-BenchPaperDataset+benchmark
FloAff-KitchenTarget-conditioned floor-affordance prediction for mobile manipulationFloAff-KitchenSiteDataset+benchmark
EgoHTRScene-aligned 4D human motion reconstruction and terrain traversalEgoHTRSiteDataset+benchmark
A-GEF / SG-EgoAction-conditioned scene-graph edit forecasting and graph-text reasoningSG-EgoSiteDataset+benchmark
Ego-Exo4DEgo–exo skill understanding, many tasksEgo-Exo4DSiteSuite
EgoTrajEgocentric multimodal trajectory forecastingEgoTrajGitHubDataset+benchmark
EgoProxEgocentric 3D proximity reasoning VQAADT / EgoExo4DSiteStandalone
Map-Mono-EgoMap-grounded global human pose estimationAIST-LivingSiteDataset+benchmark
ADT (Aria Digital Twin)Egocentric 3D machine perceptionADTAriaDataset+benchmark
OVO-S-BenchStreaming spatial intelligence over continuous ego videoNine egocentric video sourcesSiteStandalone

πŸ› οΈ Tools & Libraries

NameDescriptionLink
Ego4D CLIOfficial downloader and tooling for accessing Ego4D releases.GitHub
HOMIE-toolkitToolkit released with Ropedia Xperience-10M for large-scale multimodal ego data.GitHub
Open-AoE ToolchainSmartphone capture, reconstruction, visualization, retargeting, and model-ready conversion for Open-AoE.GitHub
Ego-OSCAROpen-hardware stereo-inertial capture device and recording stack with a sub-$200 bill of materials.Paper
ego-stereo-cn-v1-toolsLoading + timing verification for the ego-stereo-cn-v1 LeRobot v3 stereo+IMU sample (hardware-synced).GitHub
AssemblyHands ToolkitOfficial toolkit for the AssemblyHands benchmark.GitHub
TREK-150 ToolkitToolkit for the TREK-150 egocentric tracking benchmark.GitHub

🀝 Contributing

  1. Add or update the dataset, benchmark, or survey directly in the matching section of README.md.
  2. Keep primary entries unique: one full entry under one topic, cross-links everywhere else.
  3. Preserve newest-to-oldest ordering inside each topic block, with flagship entries kept at the top.
  4. Follow the detailed checklist in CONTRIBUTING.md before opening a PR.

❀️ Contact

If you have suggestions, dataset updates, or find this project useful, feel free to contact Shen Yujiao at shenyujiao18@gmail.com.

License

CC0 1.0 Universal. See LICENSE.

awesome-list
computer-vision
datasets
egocentric-datasets
egocentric-vision
first-person-video
video-datasets

Contributors

player0718

32 commits

Jingkang50

3 commits

TateZhouSiu

1 commits