MobileMem is the dataset release for MobileMem: On-Device Memory for Continually Evolving Agents, a benchmark for evaluating long-term memory systems in realistic mobile-assistant scenarios.
MobileMem models a continually evolving personal assistant that receives heterogeneous, temporally ordered interactions from users, assistants, and mobile applications. The benchmark evaluates whether a memory system can preserve, update, retrieve, and reason over personal knowledge across long interaction horizons.
For details about the data synthesis pipeline, please refer to the MobileMem repository and the technical report.
MobileMem contains three complementary splits:
| Split | Modality | Description |
|---|---|---|
| text | Text | Long-horizon userβassistant conversations and structured mobile-app events for evaluating textual memory systems. |
| omni | Text and images | Multimodal mobile interactions with screenshots and photos. |
| struct | Structured and multimodal device data | App-native records distributed across calendars, notes, bills, documents, media, and other on-device sources. |
The text split corresponds to the MobileMem textual setting described in the paper. Mobile applications are connected to a system-level assistant through predefined templates. When a user creates or updates a record, the application converts the interaction into a structured textual event and forwards it to the memory layer.
text/
βββ mobilemem_data.json
βββ profiles/
βββ user_01.json
βββ user_02.json
mobilemem_data.json is a JSON array of post-processed trajectory records, with one element per synthesized user.profiles/*.json contains the initial persona inputs used to start trajectory synthesis. The person field in mobilemem_data.json is the final, trajectory-evolved profile and is not identical to the corresponding initial profile.Each trajectory record in mobilemem_data.json contains the following main fields:
| Field | Description |
|---|---|
person | Final evolved user profile at the end of trajectory synthesis, including attribute values, histories, evidence links, and update operations. |
sessions | Chronologically ordered leaf interaction sessions used as memory system input. |
graphs | Hierarchical temporal event graphs produced during trajectory synthesis. |
old_question_type_toolbook | Snapshot of the question taxonomy and question-answer (QA) pairs before post-processing quality control. |
question_type_toolbook | Canonical benchmark QA source. It contains the retained QA pairs after post-processing. These pairs are grouped by question type. |
revised_qa_pairs | Auxiliary flat list of QA pairs retained after post-processing, including pairs that require no textual revision. |
revision_results | Per-original-QA quality-control log containing revision and discard decisions with explanations. |
statistics | Post-processing counters such as the total, revised, discarded, skipped, and failed QA counts. |
Each session contains a list of messages:
{
"id": "session_...",
"event_id": "event_...",
"messages": [
{
"id": "message_...",
"name": "Calendar",
"content": "The textual interaction or app event.",
"role": "system",
"timestamp": "2024-12-04 19:53:00",
"metadata": {}
}
],
"started_at": "2024-12-04 19:53:00",
"ended_at": "2024-12-04 21:53:00"
}
Application events use the system role, while ordinary conversations use user and assistant. The name field identifies the speaker or source application.
Each finalized QA pair under question_type_toolbook.question_types[*].qa_pairs follows this structure:
{
"id": "qa_...",
"question": "Which event happened first, and on what dates?",
"question_type": "temporal-reasoning",
"question_form": "open_ended",
"golden_answers": [
"Event A occurred first on December 4, followed by Event B on December 6."
],
"difficulty": "hard",
"num_hops": 2,
"effective_timestamp": "2024-12-06 18:48:01",
"source_evidences": [
{
"id": "message_...",
"role": "system",
"timestamp": "2024-12-04 19:53:00",
"content": "..."
}
]
}
For benchmark evaluation, please feed sessions to the memory system in chronological order and flatten question_type_toolbook.question_types[*].qa_pairs to obtain the evaluation questions and their golden_answers. The source_evidences field is provided for evidence analysis and reference-grounded judging.
After cloning this dataset repository, run the following examples from the repository root: the directory containing this README and the text/ folder.
To load the raw JSON with Hugging Face Datasets:
from datasets import load_dataset
dataset = load_dataset(
"json",
data_files={"text": "text/mobilemem_data.json"},
)
text_split = dataset["text"]
To preserve the complete nested structure, load the file directly:
import json
from pathlib import Path
data_path = Path("text/mobilemem_data.json")
with data_path.open(encoding="utf-8") as data_file:
trajectories = json.load(data_file)
For memory construction and benchmark evaluation, MobileMem can also be loaded through MemBase. After setting up MemBase, run the following from the root of the cloned MemBase repository and replace the dataset path as needed:
from membase import DATASET_MAPPING
mobilemem_dataset = DATASET_MAPPING["MobileMem"].read_raw_data(
"/path/to/MobileMem/text/mobilemem_data.json"
)
print(repr(mobilemem_dataset))
The MemBase loader converts the raw records into its standardized trajectory, session, message, and QA data models.
The omni split extends MobileMem to multimodal mobile interactions. In addition to textual conversations and app events, each user trajectory includes mobile screenshots, camera photos, and other visual content.
omni/
βββ data.jsonl # Multimodal trajectory data with image references
βββ image.zip # All images referenced in data.jsonl
βββ questions.jsonl # Evaluation questions with image grounding
data.jsonl contains one JSON record per user, following the same session and message structure as the text split, with additional image_refs fields linking to visual content.image.zip contains all PNG/JPEG images (screenshots, photos, generated imagery) referenced in the trajectories.questions.jsonl contains the benchmark QA pairs for the omni split, including visual reasoning questions that require interpreting image content.Dialogue Structure contains:
| Field | Description |
|---|---|
turn | Turn index of the utterance |
role | Speaker role, either user or assistant |
content_type | Type of content, either text or image |
content | Text content, used when content_type is text |
image_inline | Image path, used when content_type is image |
Each question contains:
| Field | Description |
|---|---|
question_id | Unique question identifier |
question | Open-ended question text |
answer | Ground-truth answer |
question_format | Format of the question; all retained questions use open_ended |
question_type | Type of the question (e.g., single_hop, multi_hop, visual_reasoning) |
difficulty | Difficulty level of the question, one of easy, medium, or hard |
evidence | List of evidence supporting the answer |
image_refs | List of image paths involved in the question |
source_session_ids | List of source session identifiers |
source_event_ids | List of source event identifiers |
target | Target type for special questions; only a small number of records contain this field |
After downloading and extracting image.zip to the omni/ directory:
import json
from pathlib import Path
# Load trajectories
data_path = Path("omni/data.jsonl")
trajectories = []
with data_path.open(encoding="utf-8") as f:
for line in f:
trajectories.append(json.loads(line))
# Load questions
questions_path = Path("omni/questions.jsonl")
questions = []
with questions_path.open(encoding="utf-8") as f:
for line in f:
questions.append(json.loads(line))
For Hugging Face Datasets:
from datasets import load_dataset
dataset = load_dataset("json", data_files={
"data": "omni/data.jsonl",
"questions": "omni/questions.jsonl",
})
The struct split evaluates long-term memory in a device-native setting. It preserves information in separate app-like stores rather than flattening a user's history into a single conversation.
The complete split contains ten synthetic personas: a young professional, a young manager, a local business owner, a public-sector employee, a father, a mother, a truck driver, a delivery rider, a male university student, and a female university student.
struct/
βββ persona01/
βββ persona02/
βββ ...
βββ persona10/
Every persona directory follows the same organization. Using persona01/ to illustrate one persona:
persona01/
βββ bill/batch.json
βββ calendar/batch.json
βββ document/batch.json
βββ note/batch.json
βββ todo/batch.json
βββ voice/batch.json
βββ image.zip
βββ screen/<evidence_id>.html
βββ video/description.json
βββ event/user.json
βββ case.csv
| Source | Format | Description |
|---|---|---|
event | JSON | Persona profile data. |
bill, calendar, document, note, todo, voice | JSON | Unimodal structured application trajectory data. |
image | PNG | Multimodal image data grouped by category. |
screen | HTML | Multimodal screenshot data. |
video | MP4 | Prompts for generating multimodal video data. |
case.csv | CSV | Contains the scenario dimension, query ID, user scenario, query, evidence IDs, ground truth, gold answer, scoring criteria, passing threshold, and evaluated capabilities. |
Run the following from the dataset repository root, replacing persona01 with the desired persona directory:
import json
from pathlib import Path
persona_dir = Path("struct/persona01")
with (persona_dir / "event/user.json").open(encoding="utf-8") as f:
persona_profile = json.load(f)
with (persona_dir / "calendar/batch.json").open(encoding="utf-8") as f:
calendar_records = json.load(f)
MobileMem uses an end-to-end evaluation protocol. For the text and omni splits, a memory system processes trajectory messages in chronological order before answering the retained questions. For the struct split, it processes the app records and artifacts available by the query time, then answers the cases grounded in those sources. Predictions are compared with the corresponding reference answers.
Text-split evaluation code is maintained in the MemBase repository. Omni-split evaluation code is available in the MobileMem-Omni/eval directory. Struct-split reference cases and scoring rubrics are distributed with each persona directory.
If you use MobileMem, please cite:
@techreport{mobilemem,
title={MobileMem: Learning from a Year of Mobile Experiences},
author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
year={2026},
institution={OPPO and OpenKG},
eprint={2608.13606},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.13606},
}
This dataset is released under the MIT License.
MobileMem is the dataset release for MobileMem: On-Device Memory for Continually Evolving Agents, a benchmark for evaluating long-term memory systems in realistic mobile-assistant scenarios.
MobileMem models a continually evolving personal assistant that receives heterogeneous, temporally ordered interactions from users, assistants, and mobile applications. The benchmark evaluates whether a memory system can preserve, update, retrieve, and reason over personal knowledge across long interaction horizons.
For details about the data synthesis pipeline, please refer to the MobileMem repository and the technical report.
MobileMem contains three complementary splits:
| Split | Modality | Description |
|---|---|---|
| text | Text | Long-horizon userβassistant conversations and structured mobile-app events for evaluating textual memory systems. |
| omni | Text and images | Multimodal mobile interactions with screenshots and photos. |
| struct | Structured and multimodal device data | App-native records distributed across calendars, notes, bills, documents, media, and other on-device sources. |
The text split corresponds to the MobileMem textual setting described in the paper. Mobile applications are connected to a system-level assistant through predefined templates. When a user creates or updates a record, the application converts the interaction into a structured textual event and forwards it to the memory layer.
text/
βββ mobilemem_data.json
βββ profiles/
βββ user_01.json
βββ user_02.json
mobilemem_data.json is a JSON array of post-processed trajectory records, with one element per synthesized user.profiles/*.json contains the initial persona inputs used to start trajectory synthesis. The person field in mobilemem_data.json is the final, trajectory-evolved profile and is not identical to the corresponding initial profile.Each trajectory record in mobilemem_data.json contains the following main fields:
| Field | Description |
|---|---|
person | Final evolved user profile at the end of trajectory synthesis, including attribute values, histories, evidence links, and update operations. |
sessions | Chronologically ordered leaf interaction sessions used as memory system input. |
graphs | Hierarchical temporal event graphs produced during trajectory synthesis. |
old_question_type_toolbook | Snapshot of the question taxonomy and question-answer (QA) pairs before post-processing quality control. |
question_type_toolbook | Canonical benchmark QA source. It contains the retained QA pairs after post-processing. These pairs are grouped by question type. |
revised_qa_pairs | Auxiliary flat list of QA pairs retained after post-processing, including pairs that require no textual revision. |
revision_results | Per-original-QA quality-control log containing revision and discard decisions with explanations. |
statistics | Post-processing counters such as the total, revised, discarded, skipped, and failed QA counts. |
Each session contains a list of messages:
{
"id": "session_...",
"event_id": "event_...",
"messages": [
{
"id": "message_...",
"name": "Calendar",
"content": "The textual interaction or app event.",
"role": "system",
"timestamp": "2024-12-04 19:53:00",
"metadata": {}
}
],
"started_at": "2024-12-04 19:53:00",
"ended_at": "2024-12-04 21:53:00"
}
Application events use the system role, while ordinary conversations use user and assistant. The name field identifies the speaker or source application.
Each finalized QA pair under question_type_toolbook.question_types[*].qa_pairs follows this structure:
{
"id": "qa_...",
"question": "Which event happened first, and on what dates?",
"question_type": "temporal-reasoning",
"question_form": "open_ended",
"golden_answers": [
"Event A occurred first on December 4, followed by Event B on December 6."
],
"difficulty": "hard",
"num_hops": 2,
"effective_timestamp": "2024-12-06 18:48:01",
"source_evidences": [
{
"id": "message_...",
"role": "system",
"timestamp": "2024-12-04 19:53:00",
"content": "..."
}
]
}
For benchmark evaluation, please feed sessions to the memory system in chronological order and flatten question_type_toolbook.question_types[*].qa_pairs to obtain the evaluation questions and their golden_answers. The source_evidences field is provided for evidence analysis and reference-grounded judging.
After cloning this dataset repository, run the following examples from the repository root: the directory containing this README and the text/ folder.
To load the raw JSON with Hugging Face Datasets:
from datasets import load_dataset
dataset = load_dataset(
"json",
data_files={"text": "text/mobilemem_data.json"},
)
text_split = dataset["text"]
To preserve the complete nested structure, load the file directly:
import json
from pathlib import Path
data_path = Path("text/mobilemem_data.json")
with data_path.open(encoding="utf-8") as data_file:
trajectories = json.load(data_file)
For memory construction and benchmark evaluation, MobileMem can also be loaded through MemBase. After setting up MemBase, run the following from the root of the cloned MemBase repository and replace the dataset path as needed:
from membase import DATASET_MAPPING
mobilemem_dataset = DATASET_MAPPING["MobileMem"].read_raw_data(
"/path/to/MobileMem/text/mobilemem_data.json"
)
print(repr(mobilemem_dataset))
The MemBase loader converts the raw records into its standardized trajectory, session, message, and QA data models.
The omni split extends MobileMem to multimodal mobile interactions. In addition to textual conversations and app events, each user trajectory includes mobile screenshots, camera photos, and other visual content.
omni/
βββ data.jsonl # Multimodal trajectory data with image references
βββ image.zip # All images referenced in data.jsonl
βββ questions.jsonl # Evaluation questions with image grounding
data.jsonl contains one JSON record per user, following the same session and message structure as the text split, with additional image_refs fields linking to visual content.image.zip contains all PNG/JPEG images (screenshots, photos, generated imagery) referenced in the trajectories.questions.jsonl contains the benchmark QA pairs for the omni split, including visual reasoning questions that require interpreting image content.Dialogue Structure contains:
| Field | Description |
|---|---|
turn | Turn index of the utterance |
role | Speaker role, either user or assistant |
content_type | Type of content, either text or image |
content | Text content, used when content_type is text |
image_inline | Image path, used when content_type is image |
Each question contains:
| Field | Description |
|---|---|
question_id | Unique question identifier |
question | Open-ended question text |
answer | Ground-truth answer |
question_format | Format of the question; all retained questions use open_ended |
question_type | Type of the question (e.g., single_hop, multi_hop, visual_reasoning) |
difficulty | Difficulty level of the question, one of easy, medium, or hard |
evidence | List of evidence supporting the answer |
image_refs | List of image paths involved in the question |
source_session_ids | List of source session identifiers |
source_event_ids | List of source event identifiers |
target | Target type for special questions; only a small number of records contain this field |
After downloading and extracting image.zip to the omni/ directory:
import json
from pathlib import Path
# Load trajectories
data_path = Path("omni/data.jsonl")
trajectories = []
with data_path.open(encoding="utf-8") as f:
for line in f:
trajectories.append(json.loads(line))
# Load questions
questions_path = Path("omni/questions.jsonl")
questions = []
with questions_path.open(encoding="utf-8") as f:
for line in f:
questions.append(json.loads(line))
For Hugging Face Datasets:
from datasets import load_dataset
dataset = load_dataset("json", data_files={
"data": "omni/data.jsonl",
"questions": "omni/questions.jsonl",
})
The struct split evaluates long-term memory in a device-native setting. It preserves information in separate app-like stores rather than flattening a user's history into a single conversation.
The complete split contains ten synthetic personas: a young professional, a young manager, a local business owner, a public-sector employee, a father, a mother, a truck driver, a delivery rider, a male university student, and a female university student.
struct/
βββ persona01/
βββ persona02/
βββ ...
βββ persona10/
Every persona directory follows the same organization. Using persona01/ to illustrate one persona:
persona01/
βββ bill/batch.json
βββ calendar/batch.json
βββ document/batch.json
βββ note/batch.json
βββ todo/batch.json
βββ voice/batch.json
βββ image.zip
βββ screen/<evidence_id>.html
βββ video/description.json
βββ event/user.json
βββ case.csv
| Source | Format | Description |
|---|---|---|
event | JSON | Persona profile data. |
bill, calendar, document, note, todo, voice | JSON | Unimodal structured application trajectory data. |
image | PNG | Multimodal image data grouped by category. |
screen | HTML | Multimodal screenshot data. |
video | MP4 | Prompts for generating multimodal video data. |
case.csv | CSV | Contains the scenario dimension, query ID, user scenario, query, evidence IDs, ground truth, gold answer, scoring criteria, passing threshold, and evaluated capabilities. |
Run the following from the dataset repository root, replacing persona01 with the desired persona directory:
import json
from pathlib import Path
persona_dir = Path("struct/persona01")
with (persona_dir / "event/user.json").open(encoding="utf-8") as f:
persona_profile = json.load(f)
with (persona_dir / "calendar/batch.json").open(encoding="utf-8") as f:
calendar_records = json.load(f)
MobileMem uses an end-to-end evaluation protocol. For the text and omni splits, a memory system processes trajectory messages in chronological order before answering the retained questions. For the struct split, it processes the app records and artifacts available by the query time, then answers the cases grounded in those sources. Predictions are compared with the corresponding reference answers.
Text-split evaluation code is maintained in the MemBase repository. Omni-split evaluation code is available in the MobileMem-Omni/eval directory. Struct-split reference cases and scoring rubrics are distributed with each persona directory.
If you use MobileMem, please cite:
@techreport{mobilemem,
title={MobileMem: Learning from a Year of Mobile Experiences},
author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
year={2026},
institution={OPPO and OpenKG},
eprint={2608.13606},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2608.13606},
}
This dataset is released under the MIT License.