SuperMemory-VQA is an egocentric visual question answering benchmark for evaluating long-horizon memory in augmented reality assistant settings. The dataset is designed around practical questions a person might ask a wearable memory assistant, such as where an object was left, what someone said earlier, whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in 52.9 hours of everyday activities recorded by 10 participants wearing Gen 1 Meta Aria Glasses. Recordings include synchronized RGB video, processed gaze, IMU, SLAM trajectories, point clouds, and redacted audio transcripts. Raw audio is not released.
SuperMemory-VQA targets long-horizon, multimodal memory rather than short-clip video understanding. Questions may require retrieving evidence across hours, days, or multiple recording sessions, and many questions require linking more than one supporting moment.
Each question is represented as multiple choice. In addition to correct and incorrect answers, the benchmark includes calibrated unanswerable options so systems must decide when the available memory evidence is insufficient instead of hallucinating an answer.
The dataset covers six memory-oriented task categories:
Dataset entries are organized around individual QA examples. A typical example contains:
The released data is intended to support both end-to-end VQA evaluation and analysis of retrieval, grounding, temporal reasoning, and abstention behavior.
This dataset is intended for research on:
The primary benchmark setting is zero-shot evaluation on the released QA labels. Systems trained, fine-tuned, or otherwise optimized on SuperMemory-VQA labels should report that usage separately.
The paper evaluates systems using three complementary metrics:
These metrics are designed to separate safe abstention from grounded answer selection. A model can identify that a question is answerable while still selecting the wrong evidence-backed answer, so reporting all three metrics is recommended.
Data was collected under an IRB-approved protocol. Participants wore Gen 1 Meta Aria Glasses during loosely scripted everyday activities in a simulated home environment, including cooking, games, puzzles, exploration, outdoor walks, and errands. Each participant contributed 3 to 12 hours of recordings, and some participants contributed recordings spanning multiple days.
The glasses captured RGB video, grayscale SLAM streams, eye tracking, audio, IMU, magnetometer, and barometer data. The public release includes processed modalities needed for benchmark use, with privacy-preserving transformations as described below.
Question-answer pairs were generated with a human-in-the-loop pipeline:
The benchmark emphasizes questions whose answers are causally available from recorded evidence before the question time.
The dataset contains egocentric recordings from human participants and should be used with care. The release applies several privacy protections:
Although the dataset has been de-identified, egocentric video can still contain residual contextual information. Users should not attempt to identify participants or bystanders.
SuperMemory-VQA is an initial benchmark for long-horizon egocentric memory, not an exhaustive sample of all daily-life settings. The recordings come from 10 participants in loosely scripted indoor and outdoor activities centered on a simulated home environment. The dataset is English-only and may not reflect the full diversity of homes, cultures, languages, accessibility needs, privacy expectations, or unconstrained daily routines.
Because many examples involve personal activities and conversations, benchmark performance should not be interpreted as readiness for deployment in real AR memory assistants. Practical systems require additional safeguards for consent, privacy, user control, uncertainty communication, and secure data handling.
This dataset card declares the dataset license as CC BY-NC-SA 4.0.
SuperMemory-VQA is an egocentric visual question answering benchmark for evaluating long-horizon memory in augmented reality assistant settings. The dataset is designed around practical questions a person might ask a wearable memory assistant, such as where an object was left, what someone said earlier, whether a planned step was completed, or what happened next in a longer event.
The benchmark contains 4,853 human-verified question-answer pairs grounded in 52.9 hours of everyday activities recorded by 10 participants wearing Gen 1 Meta Aria Glasses. Recordings include synchronized RGB video, processed gaze, IMU, SLAM trajectories, point clouds, and redacted audio transcripts. Raw audio is not released.
SuperMemory-VQA targets long-horizon, multimodal memory rather than short-clip video understanding. Questions may require retrieving evidence across hours, days, or multiple recording sessions, and many questions require linking more than one supporting moment.
Each question is represented as multiple choice. In addition to correct and incorrect answers, the benchmark includes calibrated unanswerable options so systems must decide when the available memory evidence is insufficient instead of hallucinating an answer.
The dataset covers six memory-oriented task categories:
Dataset entries are organized around individual QA examples. A typical example contains:
The released data is intended to support both end-to-end VQA evaluation and analysis of retrieval, grounding, temporal reasoning, and abstention behavior.
This dataset is intended for research on:
The primary benchmark setting is zero-shot evaluation on the released QA labels. Systems trained, fine-tuned, or otherwise optimized on SuperMemory-VQA labels should report that usage separately.
The paper evaluates systems using three complementary metrics:
These metrics are designed to separate safe abstention from grounded answer selection. A model can identify that a question is answerable while still selecting the wrong evidence-backed answer, so reporting all three metrics is recommended.
Data was collected under an IRB-approved protocol. Participants wore Gen 1 Meta Aria Glasses during loosely scripted everyday activities in a simulated home environment, including cooking, games, puzzles, exploration, outdoor walks, and errands. Each participant contributed 3 to 12 hours of recordings, and some participants contributed recordings spanning multiple days.
The glasses captured RGB video, grayscale SLAM streams, eye tracking, audio, IMU, magnetometer, and barometer data. The public release includes processed modalities needed for benchmark use, with privacy-preserving transformations as described below.
Question-answer pairs were generated with a human-in-the-loop pipeline:
The benchmark emphasizes questions whose answers are causally available from recorded evidence before the question time.
The dataset contains egocentric recordings from human participants and should be used with care. The release applies several privacy protections:
Although the dataset has been de-identified, egocentric video can still contain residual contextual information. Users should not attempt to identify participants or bystanders.
SuperMemory-VQA is an initial benchmark for long-horizon egocentric memory, not an exhaustive sample of all daily-life settings. The recordings come from 10 participants in loosely scripted indoor and outdoor activities centered on a simulated home environment. The dataset is English-only and may not reflect the full diversity of homes, cultures, languages, accessibility needs, privacy expectations, or unconstrained daily routines.
Because many examples involve personal activities and conversations, benchmark performance should not be interpreted as readiness for deployment in real AR memory assistants. Practical systems require additional safeguards for consent, privacy, user control, uncertainty communication, and secure data handling.
This dataset card declares the dataset license as CC BY-NC-SA 4.0.