zjunlp/MobileMem

MobileMem: Learning from a Year of Mobile Experiences

Python

28

103 commits

updated Sep 30, 2026

See the code

README

MobileMem: Learning from a Year of Mobile Experiences

Technical Report Chinese Report Website HuggingFace License


MobileMem is a comprehensive benchmarking framework for evaluating on-device memory systems in realistic mobile environments.


📑 Table of Contents


🔥 News

  • 2026-09-30 — We released a structured Chinese version of the dataset and benchmark.
  • 2026-08-18 — We have officially released our technical report.
  • 2026-08-01 — We publicly release the MobileMem dataset.
  • 2026-05-16 — We launch the English version of the dataset and benchmark.
  • 2026-05-03 — We launch the Chinese version of the dataset and benchmark.

🎯 Applications

MobileMem is built from multiple heterogeneous sources to enable comprehensive on-device memory modeling.


📊 Dataset Structure

MobileMem contains three complementary tracks:

TrackModalityDescription
textTextLong-horizon user–assistant conversations and structured mobile-app events for evaluating textual memory systems.
omniText and imagesMultimodal mobile interactions with screenshots and photos.
structStructured dataSimulated structured data from on-device applications.

The dataset is available for download at HuggingFace.


🗂️ Project Structure

The repository is organized into three main tracks. Track-specific details remain in each subdirectory.

MobileMem/
│
├── text/                           # 📖 Text Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── keme/                       # 🛠️ KEME synthesis pipeline code
│   └── eval/                       # ⚙️ Evaluation scripts for text track
│
├── omni/                           # 🖼️ Omni Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── src/                        # 🛠️ Data construction pipeline code
│   └── eval/                       # ⚙️ Evaluation scripts for omni track
│
├── struct/                         # 🧩 Struct Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── construct/                  # 🛠️ Example generation and review (optional)
│   └── eval/                       # ⚙️ Evaluation for struct track
│
└── README.md                       # This file

🚀 Getting Started

MobileMem offers three benchmark tracks. Choose the path that fits your needs and navigate to the corresponding resources.

For an interactive visualization of the MobileMem data, visit the Dataset Explorer branch.

📖 Text Track

The textual benchmark for evaluating memory systems on long-term, knowledge-intensive mobile agent trajectories.

SectionDescriptionQuick Link
📥 Data AccessDownload the synthesized KEME trajectories and QA pairs from HuggingFace.Link
⚙️ How to EvaluateDetailed evaluation guide for reproducing leaderboard results is available in the MemBase repository.Link
🛠️ Data ConstructionReproduce the KEME synthesis pipeline from scratch.Link

🖼️ Omni Track

The multimodal benchmark for evaluating on-device memory with realistic mobile images and dialogues.

SectionDescriptionQuick Link
📥 Data AccessDownload the MobileMem-Omni dataset, including images and dialogues.Link
⚙️ How to EvaluateDetailed evaluation guide for reproducing leaderboard results is available in the MemBase repository.Link
🛠️ Data ConstructionRebuild the entire MobileMem-Omni dataset with the provided pipeline.Link

🧩 Struct Track

Simulated structured data from on-device applications.

SectionDescriptionQuick Link
📥 Data AccessDownload Struct cases and structured app records from HuggingFace.Link
⚙️ How to EvaluateEvaluate agent JSONL traces against Struct cases.Link
🛠️ Data ConstructionExample generation and review Skills.Link

🔍 Analyzing Failures with MemTrace

We recommend using MemTrace to perform an in-depth error analysis. MemTrace helps you visualize and diagnose where and why your memory system fails, making it easier to pinpoint areas for improvement. For an example of how to use MemTrace, please refer to the tutorial in MemBase.


🚩 Citation

If this work or datasets is helpful, please kindly cite as this:

@techreport{mobilemem,
      title={MobileMem: Learning from a Year of Mobile Experiences}, 
      author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
      year={2026},
      institution={OPPO and OpenKG},
      eprint={2608.13606},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.13606}, 
}
agent
artificial-intelligence
benchmark
dataset
large-language-models
long-horizon
memory
mobile
mobile-app
mobilemem
mobile-omni
natural-language-processing
phone

zjunlp/MobileMem

MobileMem: Learning from a Year of Mobile Experiences

Python

28

103 commits

updated Sep 30, 2026

See the code

README

MobileMem: Learning from a Year of Mobile Experiences

Technical Report Chinese Report Website HuggingFace License


MobileMem is a comprehensive benchmarking framework for evaluating on-device memory systems in realistic mobile environments.


📑 Table of Contents


🔥 News

  • 2026-09-30 — We released a structured Chinese version of the dataset and benchmark.
  • 2026-08-18 — We have officially released our technical report.
  • 2026-08-01 — We publicly release the MobileMem dataset.
  • 2026-05-16 — We launch the English version of the dataset and benchmark.
  • 2026-05-03 — We launch the Chinese version of the dataset and benchmark.

🎯 Applications

MobileMem is built from multiple heterogeneous sources to enable comprehensive on-device memory modeling.


📊 Dataset Structure

MobileMem contains three complementary tracks:

TrackModalityDescription
textTextLong-horizon user–assistant conversations and structured mobile-app events for evaluating textual memory systems.
omniText and imagesMultimodal mobile interactions with screenshots and photos.
structStructured dataSimulated structured data from on-device applications.

The dataset is available for download at HuggingFace.


🗂️ Project Structure

The repository is organized into three main tracks. Track-specific details remain in each subdirectory.

MobileMem/
│
├── text/                           # 📖 Text Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── keme/                       # 🛠️ KEME synthesis pipeline code
│   └── eval/                       # ⚙️ Evaluation scripts for text track
│
├── omni/                           # 🖼️ Omni Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── src/                        # 🛠️ Data construction pipeline code
│   └── eval/                       # ⚙️ Evaluation scripts for omni track
│
├── struct/                         # 🧩 Struct Track
│   ├── README.md                   # Track-specific guide and dataset download
│   ├── construct/                  # 🛠️ Example generation and review (optional)
│   └── eval/                       # ⚙️ Evaluation for struct track
│
└── README.md                       # This file

🚀 Getting Started

MobileMem offers three benchmark tracks. Choose the path that fits your needs and navigate to the corresponding resources.

For an interactive visualization of the MobileMem data, visit the Dataset Explorer branch.

📖 Text Track

The textual benchmark for evaluating memory systems on long-term, knowledge-intensive mobile agent trajectories.

SectionDescriptionQuick Link
📥 Data AccessDownload the synthesized KEME trajectories and QA pairs from HuggingFace.Link
⚙️ How to EvaluateDetailed evaluation guide for reproducing leaderboard results is available in the MemBase repository.Link
🛠️ Data ConstructionReproduce the KEME synthesis pipeline from scratch.Link

🖼️ Omni Track

The multimodal benchmark for evaluating on-device memory with realistic mobile images and dialogues.

SectionDescriptionQuick Link
📥 Data AccessDownload the MobileMem-Omni dataset, including images and dialogues.Link
⚙️ How to EvaluateDetailed evaluation guide for reproducing leaderboard results is available in the MemBase repository.Link
🛠️ Data ConstructionRebuild the entire MobileMem-Omni dataset with the provided pipeline.Link

🧩 Struct Track

Simulated structured data from on-device applications.

SectionDescriptionQuick Link
📥 Data AccessDownload Struct cases and structured app records from HuggingFace.Link
⚙️ How to EvaluateEvaluate agent JSONL traces against Struct cases.Link
🛠️ Data ConstructionExample generation and review Skills.Link

🔍 Analyzing Failures with MemTrace

We recommend using MemTrace to perform an in-depth error analysis. MemTrace helps you visualize and diagnose where and why your memory system fails, making it easier to pinpoint areas for improvement. For an example of how to use MemTrace, please refer to the tutorial in MemBase.


🚩 Citation

If this work or datasets is helpful, please kindly cite as this:

@techreport{mobilemem,
      title={MobileMem: Learning from a Year of Mobile Experiences}, 
      author={Xinle Deng and Yida Xue and Xiangyuan Ru and Yijun Chen and Buqiang Xu and Mingjun Mao and Xinjie Liu and Haoming Xu and Shuofei Qiao and Mengru Wang and Chen Jiang and Yuchen Eleanor Jiang and Lizhong Wang and Jason Wang and Li Zeng and Haofen Wang and Guilin Qi and Huajun Chen and Ningyu Zhang},
      year={2026},
      institution={OPPO and OpenKG},
      eprint={2608.13606},
      archivePrefix={arXiv},
      primaryClass={cs.AI},
      url={https://arxiv.org/abs/2608.13606}, 
}
agent
artificial-intelligence
benchmark
dataset
large-language-models
long-horizon
memory
mobile
mobile-app
mobilemem
mobile-omni
natural-language-processing
phone

Languages

Python

93.1%

HTML

6.5%