zjunlp/MobileMem

Dataset

MobileMem

3

33 commits

2 linked in READMEs

updated Sep 30, 2026

See the code

README

MobileMem

MobileMem is the dataset release for MobileMem: On-Device Memory for Continually Evolving Agents, a benchmark for evaluating long-term memory systems in realistic mobile-assistant scenarios.

MobileMem models a continually evolving personal assistant that receives heterogeneous, temporally ordered interactions from users, assistants, and mobile applications. The benchmark evaluates whether a memory system can preserve, update, retrieve, and reason over personal knowledge across long interaction horizons.

For details about the data synthesis pipeline, please refer to the MobileMem repository and the technical report.

Dataset Splits

MobileMem contains three complementary splits:

SplitModalityDescription
textTextLong-horizon user–assistant conversations and structured mobile-app events for evaluating textual memory systems.
omniText and imagesMultimodal mobile interactions with screenshots and photos.
structStructured and multimodal device dataApp-native records distributed across calendars, notes, bills, documents, media, and other on-device sources.

Text Split

The text split corresponds to the MobileMem textual setting described in the paper. Mobile applications are connected to a system-level assistant through predefined templates. When a user creates or updates a record, the application converts the interaction into a structured textual event and forwards it to the memory layer.

Files

text/
β”œβ”€β”€ mobilemem_data.json
└── profiles/
    β”œβ”€β”€ user_01.json
    └── user_02.json
  • mobilemem_data.json is a JSON array of post-processed trajectory records, with one element per synthesized user.
  • profiles/*.json contains the initial persona inputs used to start trajectory synthesis. The person field in mobilemem_data.json is the final, trajectory-evolved profile and is not identical to the corresponding initial profile.

Data Format

Each trajectory record in mobilemem_data.json contains the following main fields:

FieldDescription
personFinal evolved user profile at the end of trajectory synthesis, including attribute values, histories, evidence links, and update operations.
sessionsChronologically ordered leaf interaction sessions used as memory system input.
graphsHierarchical temporal event graphs produced during trajectory synthesis.
old_question_type_toolbookSnapshot of the question taxonomy and question-answer (QA) pairs before post-processing quality control.
question_type_toolbookCanonical benchmark QA source. It contains the retained QA pairs after post-processing. These pairs are grouped by question type.
revised_qa_pairsAuxiliary flat list of QA pairs retained after post-processing, including pairs that require no textual revision.
revision_resultsPer-original-QA quality-control log containing revision and discard decisions with explanations.
statisticsPost-processing counters such as the total, revised, discarded, skipped, and failed QA counts.

Each session contains a list of messages:

{
  "id": "session_...",
  "event_id": "event_...",
  "messages": [
    {
      "id": "message_...",
      "name": "Calendar",
      "content": "The textual interaction or app event.",
      "role": "system",
      "timestamp": "2024-12-04 19:53:00",
      "metadata": {}
    }
  ],
  "started_at": "2024-12-04 19:53:00",
  "ended_at": "2024-12-04 21:53:00"
}

Application events use the system role, while ordinary conversations use user and assistant. The name field identifies the speaker or source application.

Each finalized QA pair under question_type_toolbook.question_types[*].qa_pairs follows this structure:

{
  "id": "qa_...",
  "question": "Which event happened first, and on what dates?",
  "question_type": "temporal-reasoning",
  "question_form": "open_ended",
  "golden_answers": [
    "Event A occurred first on December 4, followed by Event B on December 6."
  ],
  "difficulty": "hard",
  "num_hops": 2,
  "effective_timestamp": "2024-12-06 18:48:01",
  "source_evidences": [
    {
      "id": "message_...",
      "role": "system",
      "timestamp": "2024-12-04 19:53:00",
      "content": "..."
    }
  ]
}

For benchmark evaluation, please feed sessions to the memory system in chronological order and flatten question_type_toolbook.question_types[*].qa_pairs to obtain the evaluation questions and their golden_answers. The source_evidences field is provided for evidence analysis and reference-grounded judging.

Loading the Text Split

After cloning this dataset repository, run the following examples from the repository root: the directory containing this README and the text/ folder.

To load the raw JSON with Hugging Face Datasets:

from datasets import load_dataset

dataset = load_dataset(
    "json",
    data_files={"text": "text/mobilemem_data.json"},
)
text_split = dataset["text"]

To preserve the complete nested structure, load the file directly:

import json
from pathlib import Path

data_path = Path("text/mobilemem_data.json")
with data_path.open(encoding="utf-8") as data_file:
    trajectories = json.load(data_file)

For memory construction and benchmark evaluation, MobileMem can also be loaded through MemBase. After setting up MemBase, run the following from the root of the cloned MemBase repository and replace the dataset path as needed:

from membase import DATASET_MAPPING

mobilemem_dataset = DATASET_MAPPING["MobileMem"].read_raw_data(
    "/path/to/MobileMem/text/mobilemem_data.json"
)
print(repr(mobilemem_dataset))

The MemBase loader converts the raw records into its standardized trajectory, session, message, and QA data models.

Omni Split

The omni split extends MobileMem to multimodal mobile interactions. In addition to textual conversations and app events, each user trajectory includes mobile screenshots, camera photos, and other visual content.

Files

omni/
β”œβ”€β”€ data.jsonl          # Multimodal trajectory data with image references
β”œβ”€β”€ image.zip           # All images referenced in data.jsonl
└── questions.jsonl     # Evaluation questions with image grounding
  • data.jsonl contains one JSON record per user, following the same session and message structure as the text split, with additional image_refs fields linking to visual content.
  • image.zip contains all PNG/JPEG images (screenshots, photos, generated imagery) referenced in the trajectories.
  • questions.jsonl contains the benchmark QA pairs for the omni split, including visual reasoning questions that require interpreting image content.

Data Format

Dialogue Structure contains:

FieldDescription
turnTurn index of the utterance
roleSpeaker role, either user or assistant
content_typeType of content, either text or image
contentText content, used when content_type is text
image_inlineImage path, used when content_type is image

Each question contains:

FieldDescription
question_idUnique question identifier
questionOpen-ended question text
answerGround-truth answer
question_formatFormat of the question; all retained questions use open_ended
question_typeType of the question (e.g., single_hop, multi_hop, visual_reasoning)
difficultyDifficulty level of the question, one of easy, medium, or hard
evidenceList of evidence supporting the answer
image_refsList of image paths involved in the question
source_session_idsList of source session identifiers
source_event_idsList of source event identifiers
targetTarget type for special questions; only a small number of records contain this field

Loading the Omni Split

After downloading and extracting image.zip to the omni/ directory:

import json
from pathlib import Path

# Load trajectories
data_path = Path("omni/data.jsonl")
trajectories = []
with data_path.open(encoding="utf-8") as f:
    for line in f:
        trajectories.append(json.loads(line))

# Load questions
questions_path = Path("omni/questions.jsonl")
questions = []
with questions_path.open(encoding="utf-8") as f:
    for line in f:
        questions.append(json.loads(line))

For Hugging Face Datasets:

from datasets import load_dataset

dataset = load_dataset("json", data_files={
    "data": "omni/data.jsonl",
    "questions": "omni/questions.jsonl",
})

Struct Split

The struct split evaluates long-term memory in a device-native setting. It preserves information in separate app-like stores rather than flattening a user's history into a single conversation.

The complete split contains ten synthetic personas: a young professional, a young manager, a local business owner, a public-sector employee, a father, a mother, a truck driver, a delivery rider, a male university student, and a female university student.

Files

struct/
β”œβ”€β”€ persona01/
β”œβ”€β”€ persona02/
β”œβ”€β”€ ...
└── persona10/

Every persona directory follows the same organization. Using persona01/ to illustrate one persona:

persona01/
    β”œβ”€β”€ bill/batch.json
    β”œβ”€β”€ calendar/batch.json
    β”œβ”€β”€ document/batch.json
    β”œβ”€β”€ note/batch.json
    β”œβ”€β”€ todo/batch.json
    β”œβ”€β”€ voice/batch.json
    β”œβ”€β”€ image.zip
    β”œβ”€β”€ screen/<evidence_id>.html
    β”œβ”€β”€ video/description.json
    β”œβ”€β”€ event/user.json
    └── case.csv

Data Format

SourceFormatDescription
eventJSONPersona profile data.
bill, calendar, document, note, todo, voiceJSONUnimodal structured application trajectory data.
imagePNGMultimodal image data grouped by category.
screenHTMLMultimodal screenshot data.
videoMP4Prompts for generating multimodal video data.
case.csvCSVContains the scenario dimension, query ID, user scenario, query, evidence IDs, ground truth, gold answer, scoring criteria, passing threshold, and evaluated capabilities.

Loading the Struct Split

Run the following from the dataset repository root, replacing persona01 with the desired persona directory:

import json
from pathlib import Path

persona_dir = Path("struct/persona01")

with (persona_dir / "event/user.json").open(encoding="utf-8") as f:
    persona_profile = json.load(f)

with (persona_dir / "calendar/batch.json").open(encoding="utf-8") as f:
    calendar_records = json.load(f)

Evaluation

MobileMem uses an end-to-end evaluation protocol. For the text and omni splits, a memory system processes trajectory messages in chronological order before answering the retained questions. For the struct split, it processes the app records and artifacts available by the query time, then answers the cases grounded in those sources. Predictions are compared with the corresponding reference answers.

Text-split evaluation code is maintained in the MemBase repository. Omni-split evaluation code is available in the MobileMem-Omni/eval directory. Struct-split reference cases and scoring rubrics are distributed with each persona directory.

Citation

If you use MobileMem, please cite:

@techreport{mobilemem,
      title={MobileMem: Learning from a Year of Mobile Experiences}, 
      author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
      year={2026},
      institution={OPPO and OpenKG},
      eprint={2608.13606},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.13606}, 
}

License

This dataset is released under the MIT License.

agent-memory
benchmark
long-term-memory
mobile-agent
multimodal
on-device
question-answering

zjunlp/MobileMem

Dataset

MobileMem

3

33 commits

2 linked in READMEs

updated Sep 30, 2026

See the code

README

MobileMem

MobileMem is the dataset release for MobileMem: On-Device Memory for Continually Evolving Agents, a benchmark for evaluating long-term memory systems in realistic mobile-assistant scenarios.

MobileMem models a continually evolving personal assistant that receives heterogeneous, temporally ordered interactions from users, assistants, and mobile applications. The benchmark evaluates whether a memory system can preserve, update, retrieve, and reason over personal knowledge across long interaction horizons.

For details about the data synthesis pipeline, please refer to the MobileMem repository and the technical report.

Dataset Splits

MobileMem contains three complementary splits:

SplitModalityDescription
textTextLong-horizon user–assistant conversations and structured mobile-app events for evaluating textual memory systems.
omniText and imagesMultimodal mobile interactions with screenshots and photos.
structStructured and multimodal device dataApp-native records distributed across calendars, notes, bills, documents, media, and other on-device sources.

Text Split

The text split corresponds to the MobileMem textual setting described in the paper. Mobile applications are connected to a system-level assistant through predefined templates. When a user creates or updates a record, the application converts the interaction into a structured textual event and forwards it to the memory layer.

Files

text/
β”œβ”€β”€ mobilemem_data.json
└── profiles/
    β”œβ”€β”€ user_01.json
    └── user_02.json
  • mobilemem_data.json is a JSON array of post-processed trajectory records, with one element per synthesized user.
  • profiles/*.json contains the initial persona inputs used to start trajectory synthesis. The person field in mobilemem_data.json is the final, trajectory-evolved profile and is not identical to the corresponding initial profile.

Data Format

Each trajectory record in mobilemem_data.json contains the following main fields:

FieldDescription
personFinal evolved user profile at the end of trajectory synthesis, including attribute values, histories, evidence links, and update operations.
sessionsChronologically ordered leaf interaction sessions used as memory system input.
graphsHierarchical temporal event graphs produced during trajectory synthesis.
old_question_type_toolbookSnapshot of the question taxonomy and question-answer (QA) pairs before post-processing quality control.
question_type_toolbookCanonical benchmark QA source. It contains the retained QA pairs after post-processing. These pairs are grouped by question type.
revised_qa_pairsAuxiliary flat list of QA pairs retained after post-processing, including pairs that require no textual revision.
revision_resultsPer-original-QA quality-control log containing revision and discard decisions with explanations.
statisticsPost-processing counters such as the total, revised, discarded, skipped, and failed QA counts.

Each session contains a list of messages:

{
  "id": "session_...",
  "event_id": "event_...",
  "messages": [
    {
      "id": "message_...",
      "name": "Calendar",
      "content": "The textual interaction or app event.",
      "role": "system",
      "timestamp": "2024-12-04 19:53:00",
      "metadata": {}
    }
  ],
  "started_at": "2024-12-04 19:53:00",
  "ended_at": "2024-12-04 21:53:00"
}

Application events use the system role, while ordinary conversations use user and assistant. The name field identifies the speaker or source application.

Each finalized QA pair under question_type_toolbook.question_types[*].qa_pairs follows this structure:

{
  "id": "qa_...",
  "question": "Which event happened first, and on what dates?",
  "question_type": "temporal-reasoning",
  "question_form": "open_ended",
  "golden_answers": [
    "Event A occurred first on December 4, followed by Event B on December 6."
  ],
  "difficulty": "hard",
  "num_hops": 2,
  "effective_timestamp": "2024-12-06 18:48:01",
  "source_evidences": [
    {
      "id": "message_...",
      "role": "system",
      "timestamp": "2024-12-04 19:53:00",
      "content": "..."
    }
  ]
}

For benchmark evaluation, please feed sessions to the memory system in chronological order and flatten question_type_toolbook.question_types[*].qa_pairs to obtain the evaluation questions and their golden_answers. The source_evidences field is provided for evidence analysis and reference-grounded judging.

Loading the Text Split

After cloning this dataset repository, run the following examples from the repository root: the directory containing this README and the text/ folder.

To load the raw JSON with Hugging Face Datasets:

from datasets import load_dataset

dataset = load_dataset(
    "json",
    data_files={"text": "text/mobilemem_data.json"},
)
text_split = dataset["text"]

To preserve the complete nested structure, load the file directly:

import json
from pathlib import Path

data_path = Path("text/mobilemem_data.json")
with data_path.open(encoding="utf-8") as data_file:
    trajectories = json.load(data_file)

For memory construction and benchmark evaluation, MobileMem can also be loaded through MemBase. After setting up MemBase, run the following from the root of the cloned MemBase repository and replace the dataset path as needed:

from membase import DATASET_MAPPING

mobilemem_dataset = DATASET_MAPPING["MobileMem"].read_raw_data(
    "/path/to/MobileMem/text/mobilemem_data.json"
)
print(repr(mobilemem_dataset))

The MemBase loader converts the raw records into its standardized trajectory, session, message, and QA data models.

Omni Split

The omni split extends MobileMem to multimodal mobile interactions. In addition to textual conversations and app events, each user trajectory includes mobile screenshots, camera photos, and other visual content.

Files

omni/
β”œβ”€β”€ data.jsonl          # Multimodal trajectory data with image references
β”œβ”€β”€ image.zip           # All images referenced in data.jsonl
└── questions.jsonl     # Evaluation questions with image grounding
  • data.jsonl contains one JSON record per user, following the same session and message structure as the text split, with additional image_refs fields linking to visual content.
  • image.zip contains all PNG/JPEG images (screenshots, photos, generated imagery) referenced in the trajectories.
  • questions.jsonl contains the benchmark QA pairs for the omni split, including visual reasoning questions that require interpreting image content.

Data Format

Dialogue Structure contains:

FieldDescription
turnTurn index of the utterance
roleSpeaker role, either user or assistant
content_typeType of content, either text or image
contentText content, used when content_type is text
image_inlineImage path, used when content_type is image

Each question contains:

FieldDescription
question_idUnique question identifier
questionOpen-ended question text
answerGround-truth answer
question_formatFormat of the question; all retained questions use open_ended
question_typeType of the question (e.g., single_hop, multi_hop, visual_reasoning)
difficultyDifficulty level of the question, one of easy, medium, or hard
evidenceList of evidence supporting the answer
image_refsList of image paths involved in the question
source_session_idsList of source session identifiers
source_event_idsList of source event identifiers
targetTarget type for special questions; only a small number of records contain this field

Loading the Omni Split

After downloading and extracting image.zip to the omni/ directory:

import json
from pathlib import Path

# Load trajectories
data_path = Path("omni/data.jsonl")
trajectories = []
with data_path.open(encoding="utf-8") as f:
    for line in f:
        trajectories.append(json.loads(line))

# Load questions
questions_path = Path("omni/questions.jsonl")
questions = []
with questions_path.open(encoding="utf-8") as f:
    for line in f:
        questions.append(json.loads(line))

For Hugging Face Datasets:

from datasets import load_dataset

dataset = load_dataset("json", data_files={
    "data": "omni/data.jsonl",
    "questions": "omni/questions.jsonl",
})

Struct Split

The struct split evaluates long-term memory in a device-native setting. It preserves information in separate app-like stores rather than flattening a user's history into a single conversation.

The complete split contains ten synthetic personas: a young professional, a young manager, a local business owner, a public-sector employee, a father, a mother, a truck driver, a delivery rider, a male university student, and a female university student.

Files

struct/
β”œβ”€β”€ persona01/
β”œβ”€β”€ persona02/
β”œβ”€β”€ ...
└── persona10/

Every persona directory follows the same organization. Using persona01/ to illustrate one persona:

persona01/
    β”œβ”€β”€ bill/batch.json
    β”œβ”€β”€ calendar/batch.json
    β”œβ”€β”€ document/batch.json
    β”œβ”€β”€ note/batch.json
    β”œβ”€β”€ todo/batch.json
    β”œβ”€β”€ voice/batch.json
    β”œβ”€β”€ image.zip
    β”œβ”€β”€ screen/<evidence_id>.html
    β”œβ”€β”€ video/description.json
    β”œβ”€β”€ event/user.json
    └── case.csv

Data Format

SourceFormatDescription
eventJSONPersona profile data.
bill, calendar, document, note, todo, voiceJSONUnimodal structured application trajectory data.
imagePNGMultimodal image data grouped by category.
screenHTMLMultimodal screenshot data.
videoMP4Prompts for generating multimodal video data.
case.csvCSVContains the scenario dimension, query ID, user scenario, query, evidence IDs, ground truth, gold answer, scoring criteria, passing threshold, and evaluated capabilities.

Loading the Struct Split

Run the following from the dataset repository root, replacing persona01 with the desired persona directory:

import json
from pathlib import Path

persona_dir = Path("struct/persona01")

with (persona_dir / "event/user.json").open(encoding="utf-8") as f:
    persona_profile = json.load(f)

with (persona_dir / "calendar/batch.json").open(encoding="utf-8") as f:
    calendar_records = json.load(f)

Evaluation

MobileMem uses an end-to-end evaluation protocol. For the text and omni splits, a memory system processes trajectory messages in chronological order before answering the retained questions. For the struct split, it processes the app records and artifacts available by the query time, then answers the cases grounded in those sources. Predictions are compared with the corresponding reference answers.

Text-split evaluation code is maintained in the MemBase repository. Omni-split evaluation code is available in the MobileMem-Omni/eval directory. Struct-split reference cases and scoring rubrics are distributed with each persona directory.

Citation

If you use MobileMem, please cite:

@techreport{mobilemem,
      title={MobileMem: Learning from a Year of Mobile Experiences}, 
      author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
      year={2026},
      institution={OPPO and OpenKG},
      eprint={2608.13606},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.13606}, 
}

License

This dataset is released under the MIT License.

agent-memory
benchmark
long-term-memory
mobile-agent
multimodal
on-device
question-answering