This repository introduce a comprehensive paper list, datasets, methods and tools for memory research.
354
53 commits
updated Dec 29, 2025
This repository introduce a comprehensive paper list, datasets, methods and tools for memory research.

Memory in the Age of AI Agents Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang, Boyang Liu, Fangyi Zhu, Jiahang Lin, Honglin Guo, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiejun Tan, Yanbin Yin, Jiongnan Liu, Zeyu Zhang, Zhongxiang Sun, Yutao Zhu, Hao Sun, Boci Peng, Zhenrong Cheng, Xuanbo Fan, Jiaxin Guo, Xinlei Yu, Zhenhong Zhou, Zewen Hu, Jiahao Huo, Junhao Wang, Yuwei Niu, Yu Wang, Zhenfei Yin, Xiaobin Hu, Yue Liao, Qiankun Li, Kun Wang, Wangchunshu Zhou, Yixin Liu, Dawei Cheng, Qi Zhang, Tao Gui, Shirui Pan, Yan Zhang, Philip Torr, Zhicheng Dou, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Shuicheng Yan. Arxiv 2025.
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models Jiahui Geng, Qing Li, Herbert Woisetschlaeger, Zongxiong Chen, Yuxia Wang, Preslav Nakov, Hans-Arno Jacobsen, Fakhri Karray. Arxiv 2025.
A Comprehensive Survey on Long Context Language Modeling Jiaheng Liu, Dawei Zhu, Zhiqi Bai, Yancheng He, Huanxuan Liao, Haoran Que, Zekun Wang, Chenchen Zhang, Ge Zhang, Jiebin Zhang, Yuanxing Zhang, Zhuo Chen, Hangyu Guo, Shilong Li, Ziqiang Liu, Yong Shan, Yifan Song, Jiayi Tian, Wenhao Wu, Zhejian Zhou, Ruijie Zhu, Junlan Feng, Yang Gao, Shizhu He, Zhoujun Li, Tianyu Liu, Fanyu Meng, Wenbo Su, Yingshui Tan, Zili Wang, Jian Yang, Wei Ye, Bo Zheng, Wangchunshu Zhou, Wenhao Huang, Sujian Li, Zhaoxiang Zhang Arxiv 2025.
A Survey of Personalized Large Language Models: Progress and Future Directions Jiahong Liu, Zexuan Qiu, Zhongyang Li, Quanyu Dai, Jieming Zhu, Minda Hu, Menglin Yang, Irwin King. Arxiv 2025.
Prompt Compression for Large Language Models: A Survey Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier NAACL 2025.
Cognitive Memory in Large Language Models Lianlei Shan, Shixian Luo, Zezhou Zhu, Yu Yuan, Yong Wu. Arxiv 2025.
Human-inspired Perspectives: A Survey on AI Long-term Memory Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W. Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, Junxiao Shen. Arxiv 2025.
Knowledge Conflicts for LLMs: A Survey Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, Wei Xu. EMNLP 2024.
A Survey on the Memory Mechanism of Large Language Model based Agents. Zhang, Zeyu and Bo, Xiaohe and Ma, Chen and Li, Rui and Chen, Xu and Dai, Quanyu and Zhu, Jieming and Dong, Zhenhua and Wen, Ji-Rong. Arxiv 2024.
Knowledge Editing for Large Language Models: A Survey Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, Jundong Li. Arxiv 2024.
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey Yunpeng Huang, Jingwei Xu, Junyu Lai, Zixu Jiang, Taolue Chen, Zenan Li, Yuan Yao, Xiaoxing Ma, Lijuan Yang, Hao Chen, Shupeng Li, Penghao Zhao. Arxiv 2024.
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Zijun Liu, Anirudh Ravi Vijjini, Jiaming He, and others. Arxiv 2025.
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents Piaohong Wang, Motong Tian, Jiaxian Li, Yuan Liang, Yuqing Wang, Qianben Chen, Tiannan Wang, Zhicong Lu, Jiawei Ma, Yuchen Eleanor Jiang, Wangchunshu Zhou. Arxiv 2025.
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, and others. Arxiv 2025.
Agent Learning via Early Experience Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, and others. Arxiv 2025.
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems Guibin Zhang, Muxin Fu, Guancheng Wan, Miao Yu, Kun Wang, Shuicheng Yan. Arxiv 2025.
HaluMem: Evaluating Hallucinations in Memory Systems of Agents Ding Chen, Simin Niu, Kehang Li, Peng Liu, Xiangping Zheng, Bo Tang, Xinchi Li, Feiyu Xiong, Zhiyu Li. Arxiv 2025.
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and others. Arxiv 2025.
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions Yuanzhe Hu, Yu Wang, Julian McAuley. Arxiv 2025.
Mem0: Building production-ready ai agents with scalable long-term memory Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav. Arxiv 2025.
MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu, Ding Chen, Jiawei Yang, Chenyang Xi, Huayi Lai, Jihao Zhao, Yezhaohui Wang, and others. Arxiv 2025.
LightMem: Lightweight and Efficient Memory-Augmented Generation Qingyang Zhang, Ningyu Zhang, and others. Arxiv 2025.
Memory OS of AI Agent Jiazheng Kang, Mingming Ji, Zhe Zhao, Ting Bai. Arxiv 2025.
MemU: An open-source memory framework for AI companions NevaMind-AI. GitHub 2025.
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents Haoran Sun, Shaoning Zeng. Arxiv 2025.
ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory Siru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long T. Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, Tomas Pfister. Arxiv 2025.
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, Urmish Thakker, James Zou, Kunle Olukotun. Arxiv 2025.
Coarse-to-Fine Grounded Memory for LLM Agent Planning Wei Yang, Jinwei Xiao, Hongming Zhang, Qingyang Zhang, Yanna Wang, Bo Xu. Arxiv 2025.
Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation Xinzge Gao, Chuanrui Hu, Bin Chen, Teng Li. Arxiv 2025.
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang. Arxiv 2025.
Tremu: Towards Neuro-Symbolic Temporal Reasoning for LLM-Agents with Memory in Multi-Session Dialogues Yubin Ge, Salvatore Romeo, Jason Cai, Raphael Shu, Monica Sunkara, Yassine Benajiba, Yi Zhang. Arxiv 2025.
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training Deven Mahesh Mistry, Anooshka Bajaj, Yash Aggarwal, Sahaj Singh Maini, Zoran Tiganj. Arxiv 2025.
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction Kaiqiao Han, Tianqing Fang, Zhaowei Wang, Yangqiu Song, Mark Steedman. NAACL 2025.
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework Zeyu Zhang, Quanyu Dai, Rui Li, Xiaohe Bo, Xu Chen, Zhenhua Dong. Arxiv 2025.
Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Hinrich Schütze, Volker Tresp, Yunpu Ma. Arxiv 2025.
Mem-$\alpha$: Learning Memory Construction via Reinforcement Learning Yu Wang, Ryuichi Takanobu, Zhiqi Liang, Yuzhen Mao, Yuanzhe Hu, Julian McAuley, Xiaojian Wu. Arxiv 2025.
AgentFly: Fine-tuning LLM Agents without Fine-tuning LLMs Huichi Zhou, Yihang Chen, Siyuan Guo, Xue Yan, Kin Hei Lee, Zihan Wang, Ka Yiu Lee, Guchun Zhang, Kun Shao, Linyi Yang, and others. Arxiv 2025.
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and others. Arxiv 2025.
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions Yuanzhe Hu, Yu Wang, Julian McAuley. Arxiv 2025.
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, Zhenhua Dong. Arxiv 2025.
MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents Yiming Du, Bingbing Wang, Yang He, Bin Liang, Baojun Wang, Zhongyang Li, Lin Gui, Jeff Z. Pan, Ruifeng Xu, Kam-Fai Wong. Arxiv 2025.
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, James A Burke. Arxiv 2025.
Memp: Exploring Agent Procedural Memory Runnan Fang, Yuan Liang, Xiaobin Wang, Jialong Wu, Shuofei Qiao, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Arxiv 2025.
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, Dong Yu. ICLR 2025.
Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations Nuo Chen, Hongguang Li, Jia Li, Yuxuan Li, Wei Wu. COLING 2025.
Towards Lifelong Dialogue Agents via Timeline-based Memory Management Kai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. NAACL 2025.
Zep: A Temporal Knowledge Graph Architecture for Agent Memory Preston Rasmussen, Daniel Chalef. Arxiv 2025.
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su. ICML 2025.
Disentangling Memory and Reasoning Ability in Large Language Models Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, Yongfeng Zhang. Arxiv 2025.
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation Hongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao, Defu Lian, Zhicheng Dou, Tiejun Huang. Arxiv 2025.
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue Hao Li, Chenghao Yang, An Zhang, Yang Deng, Xiang Wang, Tat-Seng Chua. NAACL 2025.
MemInsight: Autonomous Memory Augmentation for LLM Agents Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yi Zhang, Yassine Benajiba. Arxiv 2025.
Interpersonal Memory Matters: A New Task for Proactive Dialogue Utilizing Conversational History Bowen Wu, Wenqing Wang, Haoran Li, Ying Li, Jingsong Yu, Baoxun Wang. Arxiv 2025.
Echo: A Large Language Model with Temporal Episodic Memory WenTao Liu, Ruohua Zhang, Aimin Zhou, Feng Gao, JiaLi Liu. Arxiv 2025.
Improving Factuality with Explicit Working Memory Mingda Chen, Yang Li, Karthik Padthe, Rulin Shao, Alicia Sun, Luke Zettlemoyer, Gargi Ghosh, Wen-tau Yih. Arxiv 2025.
Memorization Over Reasoning? Exposing and Mitigating Verbatim Memorization in Large Language Models' Character Understanding Evaluation Yuxuan Jiang, Francis Ferraro. Arxiv 2025.
Self-Memory Alignment: Mitigating Factual Hallucinations with Generalized Improvement Siyuan Zhang, Yichi Zhang, Yinpeng Dong, Hang Su. Arxiv 2025.
Needle in the Haystack for Memory Based Large Language Models Elliot Nelson, Georgios Kollias, Payel Das, Subhajit Chaudhury, Soham Dan. ICLR 2025.
Evaluating Very Long-Term Conversational Memory of LLM Agents Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang. ACL 2024.
MemoryBank: Enhancing Large Language Models with Long-Term Memory Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang. AAAI 2024.
A-MEM: Agentic Memory for LLM Agents Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, Yongfeng Zhang. Arxiv 2024.
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, Liqiang Nie. NeurIPS 2024.
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, Yu Su. NeurIPS 2024.
"My agent understands me better": Integrating Dynamic Human-like Memory Recall and Consolidation in LLM-Based Agents Yu Hou, Tamoto Yuta, Mayur Tikundi. Arxiv 2024.
Mr.Steve: Instruction-Following Agents in Minecraft with What-Where-When Memory Yu Hou, Tamoto Yuta, Mayur Tikundi. ICLR 2024.
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization Shida Wang, Qianxiao Li. ICML 2024.
Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory Graphs Zheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu, Wei Shi. EMNLP 2024.
Towards Verifiable Text Generation with Evolving Memory and Self-Reflection Hao Sun, Hengyi Cai, Bo Wang, Yingyan Hou, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang, Dawei Yin. EMNLP 2024.
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering Yiming Du, Hongru Wang, Zhengyi Zhao, Bin Liang, Baojun Wang, Wanjun Zhong, Zezhong Wang, Kam-Fai Wong. Arxiv 2024.
An Iterative Associative Memory Model for Empathetic Response Generation Zhou Yang, Zhaochun Ren, Yufeng Wang, Haizhou Sun, Chao Chen, Xiaofei Zhu, Xiangwen Liao. ACL 2024.
COCOA: CBT-based Conversational Counseling Agent Using Memory Specialized in Cognitive Distortions and Dynamic Prompt Suyeon Lee, Jieun Kang, Harim Kim, Kyoung-Mee Chung, Dongha Lee, Jinyoung Yeo. Arxiv 2024.
Mixed-Session Conversation with Egocentric Memory Jihyoung Jang, Taeyoung Kim, Hyounghun Kim. EMNLP 2024.
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models Xihang Yue, Linchao Zhu, Yi Yang. ACL 2024.
Extractive Medical Entity Disambiguation with Memory Mechanism and Memorized Entity Information Guobiao Zhang, Xueping Peng, Tao Shen, Guodong Long, Jiasheng Si, Libo Qin, Wenpeng Lu. EMNLP 2024.
Ever-Evolving Memory by Blending and Refining the Past Seo Hyun Kim, Keummin Ka, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. Arxiv 2024.
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control Longtao Zheng, Rundong Wang, Xinrun Wang, Bo An. Arxiv 2024.
Moviechat: From dense token to sparse memory for long video understanding Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, Gaoang Wang. CVPR 2024.
lamp: when large language models meet personalization Alireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed Zamani. ACL 2024.
Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, Dong Wang. NAACL 2024.
IterCQR: Iterative Conversational Query Reformulation with Retrieval Guidance Yunah Jang, Kang-il Lee, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung. NAACL 2024.
Memory Layers at Scale Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh. Arxiv 2024.
MoT: Memory-of-Thought Enables ChatGPT to Self-Improve Sureman Lee, Yujie Qian, Yujia Xie, Yifan Hou, Xinyan Wang, Yiming Yang, Xiang Ren. EMNLP 2023.
Think-in-memory: Recalling and post-thinking enable llms with long-term memory Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, Guannan Zhang. Arxiv 2023.
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models Qingxue Wang, Ling Ding, Yaran Cao, Zhilang Tan, Shi Wang, Dacheng Tao, Liu Qiu. Arxiv 2023.
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination Yuwei Zhang, Yifan Hou, Xinyan Wang, Yiming Yang, Xiang Ren. Arxiv 2023.
SCM: Enhancing Large Language Model with Self-Controlled Memory Framework Bing Wang, Xinnian Liang, Jian Yang, Hui Huang, Shuangzhi Wu, Peihao Wu, Lu Lu, Zejun Ma, Zhoujun Li. Arxiv 2023.
LDM²: A Large Decision Model Imitating Human Cognition with Dynamic Memory Enhancement Xingjin Wang, Linjing Li, Dongfeng Zeng. EMMNLP 2023.
NarrativeXL: A Large-scale Dataset For Long-Term Memory Models Arseny Moskvichev, Ky-Vinh Mai. EMNLP 2023.
Who's Harry Potter? Approximate Unlearning in LLMs Ronen Eldan, Mark Russinovich. Arxiv 2023.
Active Retrieval Augmented Generation Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. EMNLP 2023.
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, Kangwook Lee. ACL 2023.
MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation Junru Lu, Siyu An, Mingbao Lin, Gabriele Pergola, Yulan He, Di Yin, Xing Sun, Yunsheng Wu. Arxiv 2023.
Learning Retrieval Augmentation for Personalized Dialogue Generation Qiushi Huang, Shuai Fu, Xubo Liu, Wenwu Wang, Tom Ko, Yu Zhang, Lilian Tang. EMNLP 2023.
Learning to Reason and Memorize with Self-Notes Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, Sainbayar Sukhbaatar. NeurIPS 2023.
RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation Shuai Liu, Hyundong Cho, Marjorie Freedman, Xuezhe Ma, Jonathan May. ACL 2023.
Enhancing Personalized Dialogue Generation with Contrastive Latent Variables: Combining Sparse and Dense Persona Yihong Tang, Bo Wang, Miao Fang, Dongming Zhao, Kun Huang, Ruifang He, Yuexian Hou. ACL 2023.
Transformer-based World Models Are Happy With 100k Interactions Jan Robine, Marc Höftmann, Tobias Uelwer, Stefan Harmeling. ICLR 2023.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation Jing Xu, Arthur Szlam, Jason Weston. ACL 2022.
Long Time No See! Open-Domain Conversation with Long-Term Persona Memory Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, Shihang Wang. ACL 2022.
Keep Me Updated! Memory Management in Long-term Conversations Sanghwan Bae, Donghyun Kwak, Soyoung Kang, Min Young Lee, Sungdong Kim, Yuin Jeong, Hyeri Kim, Sang-Woo Lee, Woomyoung Park, Nako Sung. EMNLP 2022.
Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Clark. EMNLP 2022.
Learning to Repair: Repairing Model Output Errors after Deployment Using a Dynamic Memory of Feedback Niket Tandon, Aman Madaan, Peter Clark, Yiming Yang. NAACL 2022.
There Are a Thousand Hamlets in a Thousand People's Eyes: Enhancing Knowledge-grounded Dialogue with Personal Memory Tingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen, Rui Yan. ACL 2022.
Training Language Models with Memory Augmentation Zexuan Zhong, Tao Lei, Danqi Chen. EMNLP 2022.
Improving Multi-turn Emotional Support Dialogue Generation with Lookahead Strategy Planning Yi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang, Ruihui Zhao, Bang Liu, Xiaodan Liang, Yefeng Zheng. EMNLP 2022.
Less is More: Learning to Refine Dialogue History for Personalized Dialogue Generation Hanxun Zhong, Zhicheng Dou, Yutao Zhu, Hongjin Qian, Ji-Rong Wen. NAACL 2022.
Leveraging Similar Users for Personalized Language Modeling with Limited Data Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. ACL 2022.
PerKGQA: Question Answering over Personalized Knowledge Graphs Ritam Dutt, Kasturi Bhattacharjee, Rashmi Gangadharaiah, Dan Roth, Carolyn Rose. NAACL 2022.
A Cooperative Memory Network for Personalized Task-oriented Dialogue Systems with Incomplete User Profiles Jiahuan Pei, Pengjie Ren, Maarten de Rijke. WWW 2021.
Episodic Memory in Lifelong Language Learning Cyprien de Masson d'Autume, Sebastian Ruder, Lingpeng Kong, Dani Yogatama. NeurIPS 2019.
Table-1: Datasets for Evaluating Long-Term Memory *(Continuously Updated)
| Dataset | Mo | Operations | DS Type | Per | TR | Metrics | Purpose | Year |
|---|---|---|---|---|---|---|---|---|
| PersonaMem-v2 | text | Updating, Retrieval | MS | ✓ | ✓ | Accuracy, Persona Score | Benchmark implicit user persona learning and agentic memory updates over long contexts. | 2025 |
| MemoryBench | text | Updating, Retrieval, Forgetting | QA | ✗ | ✓ | Accuracy, Retention Rate | Comprehensive benchmark for memory correctness, persistence, and continual learning. | 2025 |
| HaluMem | text | Retrieval | QA | ✗ | ✗ | Accuracy, Hallucination Rate, Omission Rate | Evaluate hallucinations in memory extraction, updating and retrieval. | 2025 |
| BFCL V4 | text (API/Code) | Updating, Retrieval, Reasoning | QA (API) | ✗ | ✗ | AST Accuracy, Execution Success | Benchmarking function-calling capabilities, specifically featuring a "Memory" category for CRUD tool usage. | 2025 |
| LongMemEval | text | Indexing, Retrieval, Compression | MS | ✗ | ✓ | Recall@K, NDCG@K, Accuracy | Benchmark chat assistants on long-term memory abilities, including temporal reasoning. | 2025 |
| LoCoMo | text + image | Indexing, Retrieval, Compression | MS | ✗ | ✓ | Accuracy, ROUGE, Precision, Recall, F1 | Evaluate long-term memory in LLMs across QA, event summarization, and multimodal dialogue tasks. | 2024 |
| MemoryBank | text | Updating, Retrieval | MS | ✓ | ✗ | Accuracy, Human Eval | Enhance LLMs with long-term memory capabilities, adapting to user personalities and contexts. | 2024 |
| PerLTQA | text | Retrieval | MS | ✓ | ✗ | MAP, Recall, Precision, F1, Accuracy, GPT4 score | To explore personal long-term memory question answering ability. | 2024 |
| MALP | text | Retrieval, Compression | QA | ✓ | ✗ | ROUGE, Accuracy, Win Rate | Preference-conditioned dialogue generation. Parameter-efficient fine-tuning (PEFT) for customization. | 2024 |
| DialSim | text | Retrieval | MS | ✓ | ✗ | Accuracy | To evaluate dialogue systems under realistic, real-time, and long-context multi-party conversation conditions. | 2024 |
| MovieChat-1K | text + video | question-answering + caption | QA | ✗ | ✓ | Accuracy | For long-term video understanding for Large Multimodal Models across video question-answering and video captioning tasks. | 2023 |
| CC | text | Retrieval | MS | ✗ | ✓ | BLEU, ROUGE | For long-term dialogue modeling with time and relationship context. | 2023 |
| LAMP | text | Consolidation, Retrieval, Compression | MS | ✓ | ✓ | Accuracy, F1, ROUGE | Multiple entries per user. Supports both user-based splits and time-based splits. | 2023 |
| MSC | text | Consolidation, Retrieval, Compression | MS | ✓ | ✗ | PPL | Evaluate and improve long-term dialogue models via multi-session chats with evolving knowledge. | 2022 |
| DuLeMon | text | Consolidation, Updating, Retrieval, Compression | MS | ✓ | ✗ | Accuracy, F1, Recall, Precision, PPL, BLEU, DISTINCT | For dynamic persona tracking and consistent long-term interaction. | 2022 |
| 2WikiMultiHopQA | table + knowledge base + text | Consolidation, Indexing, Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Multi-hop QA combining structured and unstructured data with reasoning paths. | 2020 |
| NQ | text | Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Open-domain QA based on real Google search queries. | 2019 |
| HotpotQA | text | Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Multi-hop QA with explainable reasoning and sentence-level supporting facts. | 2018 |
Note:
Table-2: Datasets for Long-Context Memory Evaluation *(Continuously Updated)*
| Dataset | Modality | Operations | Metrics | Purpose | Year |
|---|---|---|---|---|---|
| MMLongBench | text + image | compression, retrieval | SubEM, Accuracy, Rouge-L, Model-Based | 5 categories and 16 datasets for vision-language long-context evaluation | 2025 |
| WikiText-103 | text | compression | PPL | 100M-token Wikipedia corpus for long-context language modeling | 2016 |
| PG-19 | text | compression | PPL | Project Gutenberg books corpus for long-context language modeling | 2019 |
| LRA | text + image | compression, retrieval | Acc | Benchmark with 6 tasks for evaluating efficient long-context language models | 2020 |
| NarrativeQA | text | retrieval | Bleu-1, Bleu-4, Meteor, Rouge-L, MRR | QA dataset for evaluating long-context QA ability | 2017 |
| TriviaQA | text | retrieval | EM, F1 | QA dataset for evaluating long-context QA ability | 2017 |
| NaturalQuestions | text | retrieval | EM, F1 | QA dataset for evaluating long-context QA ability | 2019 |
| MusiQue | text | retrieval | F1 | Multi-hop QA dataset for evaluating long-context reasoning and QA | 2021 |
| CNN/DailyMail | text | compression | Rouge-1, Rouge-2, Rouge-L | News articles dataset for long document summarization | 2016 |
| GovReport | text | compression | Rouge-1, Rouge-2, Rouge-L, Bert Score | Government agency reports for long document summarization | 2021 |
| L-Eval | text | compression, retrieval | Rouge-L, F1, GPT4 | 20-subtask benchmark for diverse long-context language model evaluation | 2023 |
| LongBench | text | compression, retrieval | F1, Rouge-L, Accuracy, EM, Edit Sim | 14 English, 5 Chinese, 2 code tasks for long-context evaluation | 2023 |
| LongBench v2 | text + table + KG | compression, retrieval | Acc | Longer, more challenging tasks with consistent multi-choice format | 2024 |
| SWE-bench | text | compression, retrieval | Resolution rate (%Resolved) | 2,294 task instances from 12 popular python repositories from GitHub | 2023 |
| SWE-bench Multimodal | text + image | compression, retrieval | Resolution rate (%Resolved), Inference cost (Avg. $ Cost) | Extending the original benchmark with image modal with 517 task instances | 2024 |
| $\infty$Bench | text | compression, retrieval | F1, Acc, ROUGE-L-Sum | 12 sub-tasks specially designed for evaluating extreme long context language models | 2024 |
| LooGLE | text | compression, retrieval | Bleu-1, Bleu-4, Rouge-1, Rouge-4, Rouge-L, Meteor score, Bert score, GPT4 score | 7 major tasks specially designed for evaluating extreme long context language models | 2023 |
Table-3: Datasets for Parametric Memory Evaluation *(Continuously Updated)*
| Dataset | Modality | Operations | Metrics | Purpose | Year |
|---|---|---|---|---|---|
| KnowEdit | text | updating | Edit Success, Portability, Locality, Fluency | 6 datasets covering insertion, modification, and erasure | 2024 |
| MQUAKE-CF | text | updating | Edit-wise Success Rate, Instance-wise Accuracy, Multi-hop Accuracy | Counterfactual knowledge editing through multi-hop reasoning (up to 4 hops) | 2023 |
| MQUAKE-T | text | updating | Edit-wise Success Rate, Instance-wise Accuracy, Multi-hop Accuracy | Temporal knowledge editing with one edit per reasoning chain | 2023 |
| Counterfact | text | updating | Efficacy Score, Magnitude, Paraphrase & Neighborhood Scores | Tests substantial factual changes beyond superficial edits | 2022 |
| zsRE | text | updating | Success Rate, Retain Accuracy, Equivalence Accuracy, Perf. Deterioration | One of the earliest datasets for knowledge editing | 2021 |
| MUSE | text | forgetting | VerbMem, KnowMem, PrivLeak | Unlearning benchmark with 6 desirable properties | 2024 |
| KnowUnDo | text | forgetting | Unlearn Success, Retention Success, Perplexity, ROUGE-L | Test unlearning in copyrighted and privacy-sensitive domains | 2024 |
| RWKU | text | forgetting | ROUGE-L | Real-world unlearning under corpus-free, adversarial settings | 2024 |
| WMDP | text | forgetting | QA accuracy | Proxy for hazardous knowledge in bio/cyber/chemical domains | 2024 |
| TOFU | text | forgetting | Probability, ROUGE, Truth Ratio | Unlearning dataset of facts about 200 fictitious authors | 2024 |
| ABSA | text | consolidation | F1 | Aspect-based sentiment analysis for continual learning | 2024 |
| SGD | text | consolidation | JGA, FWT, BWT | Multi-turn task-oriented dialogue with evolving intents | 2020 |
| INSPIRED | text | consolidation | JGA, FWT, BWT | Task-oriented dialogue supporting user goal evolution | 2020 |
| Natural Question | text | consolidation | Indexing Accuracy, Hits@1 | Supports continual learning over evolving document corpora | 2019 |
Note:
Table-4: Datasets for Multi-Source Memory Evaluation *(Continuously Updated)*
| Dataset | Mo | Ops | Src# | Mod# | Task | Metrics | Purpose | Year |
|---|---|---|---|---|---|---|---|---|
| MultiChat | text + image | Retrieval | 2 | 2 | Retrieval | Precision, mAP, GPT-4 | Image-grounded sticker retrieval with cross-session image-text dialogue context. | 2025 |
| Context-conflicting | text | Compression | 2 | 1 | Conflict | DiffGR, EM, Similarity | Evaluates model handling of conflicting evidence across sources. | 2024 |
| EgoSchema | video + text | Retrieval, Compression | 3 | 2 | Fusion | Accuracy | Episodic video + social schema + conversation for long-term memory QA. | 2023 |
| Ego4D NLQ | video + text | Retrieval, Compression | 2 | 2 | Fusion | Recall@K | Natural language queries over egocentric video with temporal memory. | 2022 |
| 2WikiMultihopQA | text | Indexing, Retrieval, Compression | 2 | 1 | Reasoning | EM, F1 | Multi-hop QA across Wikipedia passages with sentence-level support. | 2020 |
| HybridQA | text | Retrieval, Compression | 2 | 1 | Reasoning | EM, F1 | Reasoning across structured tables and unstructured text. | 2020 |
| CommonsenseVQA | text + image | Retrieval, Compression | 2 | 2 | Fusion | Accuracy | Commonsense QA over visual scenes requiring visual-textual fusion. | 2019 |
| NaturalQuestions | text | Retrieval, Compression | >1* | 1 | Conflict | EM, F1 | QA over Google snippets; used for contradiction analysis. | 2019 |
| ComplexWebQuestions | text | Retrieval, Compression | >1* | 1 | Reasoning | EM, F1 | Compositional QA requiring multi-step reasoning over web snippets. | 2018 |
| HotpotQA | text | Retrieval, Compression | 2 | 1 | Conflict | EM, F1, Supporting Fact Accuracy | Multi-hop QA with paragraph- and sentence-level support. | 2018 |
| TriviaQA | text | Retrieval, Compression | ≥6 | 1 | Conflict | EM, F1 | QA with noisy web sources; useful for source disagreement analysis. | 2017 |
| WebQuestionsSP | text | Indexing, Retrieval, Compression | >1* | 1 | Reasoning | F1, Accuracy | Structured QA dataset with enhanced reasoning chains. | 2016 |
| Flickr30K | text + image | Retrieval, Compression | 2 | 2 | Retrieval | Similarity | Image-caption pairs for cross-modal retrieval and alignment. | 2014 |
Note:
Table-1: Component-Level Tools for Memory Management and Utilization. *(Continuously Updated)*
| Memory Tool | Function | Input/Output | Example Use |
|---|---|---|---|
| FAISS | Library for fast storage, indexing, and retrieval of high-dimensional vectors | Vector / Index, relevance score | Indexing large sets of text embeddings and retrieving relevant documents in RAG systems |
| Neo4j | Native graph database supporting ACID transactions and Cypher query language | Nodes and relationships with properties / Query results via Cypher | Modeling and retrieving complex relational data for use cases like fraud detection and recommendation engines |
| Chroma | AI-native embedding database for building LLM applications | Text / Embeddings | Managing knowledge, facts, and skills for LLMs |
| Milvus | Vector database for embedding similarity search and AI applications | Embeddings / Similar items | Unstructured data search and similarity matching |
| Qdrant | Vector similarity search engine and database | Embeddings / Similar items | Production-ready service with user-friendly API for vector search |
| Weaviate | Open-source vector database with built-in ML models | Data objects and vector embeddings / Search results | Scalable storage and retrieval for AI applications |
| BM25 | Probabilistic ranking function for estimating document relevance | Text queries / Ranked list of documents | Enhancing search engine results and document retrieval systems |
| Contriever | Unsupervised dense retriever trained with contrastive learning | Query text / List of similar documents | High-recall retrieval tasks in multilingual question-answering systems |
| Embedding Models (e.g., OpenAI) | Convert text, images, or audio into dense vector representations capturing semantic meaning | Raw data / Vector embeddings | Text similarity computation, recommendation systems, and clustering tasks |
Table-2: Framework-Level Tools for Memory Management and Utilization *(Continuously Updated)*
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Graphiti | Framework for building and querying temporally-aware knowledge graphs tailored for AI agents in dynamic environments | Multi-source data / Queryable knowledge graph | Constructing real-time knowledge graphs to enhance AI agent memory | Open |
| LlamaIndex | A flexible framework for building knowledge assistants using LLMs connected to enterprise data | Text / Context-augmented responses | Developing knowledge assistants that process complex data formats | Open |
| LangChain | Provides a framework for building context-aware, reasoning applications by connecting LLMs with external data sources | Input prompts / Multi-step reasoning outputs | Creating complex LLM applications like question-answering systems and chatbots | Open |
| LangGraph | Constructs controllable agent architectures supporting long-term memory and human-in-the-loop multi-agent systems | Graph state / State updates | Building complex task workflows with multiple AI agents | Open |
| EasyEdit | An easy-to-use knowledge editing framework for LLMs, enabling efficient behavior modification within specific domains | Edit instructions / Updated model behavior | Modifying LLM knowledge in specific domains, such as updating factual information | Open |
| CrewAI | A platform for building and deploying multi-agent systems, supporting automated workflows using any LLM and cloud platform | Multi-agent tasks / Collaborative results | Automating workflows across agents like project management and content generation | Open |
| Letta | Constructs stateful agents with long-term memory, advanced reasoning, and custom tools within a visual environment | User interactions / Improved response | Developing AI agents that learn and improve over time | Open |
| OpenHands | An open platform for autonomous software agents that maintains persistent context across file editing, command execution, and web browsing | Natural language tasks / Code patches, Terminal actions | Automating complex software engineering tasks like debugging and feature implementation with full project context | Open |
Table-3: Application Layer-Level Tools for Memory Management and Utilization (Continuously Updated)
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Mem0 | Provides a smart memory layer for LLMs, enabling direct addition, updating, and searching of memories in models | User interactions / Personalized responses | Enhancing AI systems with persistent context for customer support and personalized recommendations | Open |
| Zep | Integrates chat messages into a knowledge graph, offering accurate and relevant user information | Chat logs, business data / Knowledge graph query results | Augmenting AI agents with knowledge through continuous learning from user interactions | Open |
| Memary | An open memory layer that emulates human memory to help AI agents manage and utilize information effectively | Agent tasks / Memory management and utilization | Building AI agents with human-like memory characteristics | Open |
| Memobase | A user profile-based long-term memory system designed to provide personalized experiences in generative AI applications | User interactions / Personalized responses | Implementing virtual assistants, educational tools, and personalized AI companions | Open |
| O-Mem | An omni-memory system enabling agents to self-evolve and maintain long-horizon consistency through recursive memory consolidation | Long-term interaction logs / Evolved memory state | Creating self-evolving personal AI assistants that adapt to user growth over time | Open |
| MemOS | An operating system-like architecture that manages memory hierarchy (working/short/long-term) to optimize Memory-Augmented Generation | Agent queries, Complex contexts / Hierarchical memory blocks | Managing complex memory resources for agents handling multi-step reasoning tasks | Open |
Table-4: Product-Level Tools for Memory Utilization (Continuously Updated)
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Me.bot | AI-powered personal assistant that organizes notes, tasks, and memories, providing emotional support and productivity tools | User inputs (text, voice) / Organized notes, reminders, summaries | Personal productivity enhancement, emotional support, idea organization | Closed |
| ima.copilot | Intelligent workstation powered by Tencent's Mix Huang model, building a personal knowledge base for learning and work scenarios | User queries / Customized responses, knowledge retrieval | Enhancing learning efficiency, work productivity, knowledge management | Closed |
| Coze | Enables multi-agent collaboration across various platforms | User-defined workflows / Response | Deployed chatbots, AI agents | Closed |
| Grok | AI assistant developed by xAI, designed to provide truthful, useful, and curious responses, with real-time data access and image generation | Query / Informative answers, generated images | Answering questions, generating images, providing insights | Closed |
| ChatGPT | Conversational AI developed by OpenAI, capable of understanding and generating human-like text based on prompts | User prompts / Generated text responses | Answering questions, generating images, providing insights | Closed |
| Claude | AI assistant featuring "Projects" to ground answers in user-provided knowledge bases and massive context windows | Prompts, Files, Code / Text, Code, Artifacts | Analyzing large codebases, maintaining consistent style across documents via Projects | Closed |
| Doubao | A high-efficiency multimodal AI assistant capable of handling long-context interactions and diverse tasks | Text, Voice, Image / Answers, creative content | Daily conversation, writing assistance, coding, and role-playing | Closed |
| Siri | Intelligent voice assistant utilizing on-device personal semantic memory for cross-app actions and context understanding | Voice commands / Action execution, personal info retrieval | Device control, retrieving personal context ("When is Mom's flight?"), cross-app tasks | Closed |
| Xiaoyi | Huawei's smart assistant integrated into HarmonyOS, leveraging ecosystem memory for proactive services and document processing | Voice, Text, Documents / Summaries, suggestions, IoT control | Document summarization, smart home control, personalized travel planning | Closed |
| Zhixiaobao | Ant Group's financial AI agent that utilizes user financial history and market knowledge for personalized wealth management | Financial queries / Market analysis, investment advice | Financial planning, insurance analysis, market trend explanation | Closed |
Table-5: Key differences between human and agent memory across operational dimensions
| Aspect | Human Memory | Agent Memory |
|---|---|---|
| Storage | Distributed, interconnected neural systems across brain regions | Parametric, modular, and context-dependent (structured or unstructured) |
| Consolidation | Slow, biologically driven, passive | Fast, explicit, policy-driven and selective |
| Indexing | Implicit, associative, sparse codes via hippocampal circuits | Explicit, embedding-based, symbolic or key–value lookup |
| Updating | Indirect, reconsolidation-based, error-prone | Precise, programmable, supports rollback/unlearning |
| Forgetting | Passive decay or interference | Transparent, trackable, policy-controlled |
| Retrieval | Cue/context/emotion dependent, emotionally biased | Content-based, reproducible, similarity or query driven |
| Compression | Implicit, salience- and frequency-biased | Explicit, customizable (e.g., quantization, summarization) |
| Ownership | Individual and private | Shareable, replicable, and broadcastable |
| Volume | Biologically limited | Scalable, bounded only by storage and compute limits |
Please contact me if I miss your names in the list, I will add you back ASAP!
🤝🤝 Thanks for all the great contributors on GitHub!
If you find our repository and survey useful for your research, please consider citing the following paper:
@article{du2025rethinking,
title={Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions},
author={Du, Yiming and Huang, Wenyu and Zheng, Danna and Wang, Zhaowei and Montella, Sebastien and Lapata, Mirella and Wong, Kam-Fai and Pan, Jeff Z.},
journal={arXiv preprint arXiv:2505.00675},
year={2025},
url={https://arxiv.org/abs/2505.00675}
}
This repository introduce a comprehensive paper list, datasets, methods and tools for memory research.
354
53 commits
updated Dec 29, 2025
This repository introduce a comprehensive paper list, datasets, methods and tools for memory research.

Memory in the Age of AI Agents Yuyang Hu, Shichun Liu, Yanwei Yue, Guibin Zhang, Boyang Liu, Fangyi Zhu, Jiahang Lin, Honglin Guo, Shihan Dou, Zhiheng Xi, Senjie Jin, Jiejun Tan, Yanbin Yin, Jiongnan Liu, Zeyu Zhang, Zhongxiang Sun, Yutao Zhu, Hao Sun, Boci Peng, Zhenrong Cheng, Xuanbo Fan, Jiaxin Guo, Xinlei Yu, Zhenhong Zhou, Zewen Hu, Jiahao Huo, Junhao Wang, Yuwei Niu, Yu Wang, Zhenfei Yin, Xiaobin Hu, Yue Liao, Qiankun Li, Kun Wang, Wangchunshu Zhou, Yixin Liu, Dawei Cheng, Qi Zhang, Tao Gui, Shirui Pan, Yan Zhang, Philip Torr, Zhicheng Dou, Ji-Rong Wen, Xuanjing Huang, Yu-Gang Jiang, Shuicheng Yan. Arxiv 2025.
A Comprehensive Survey of Machine Unlearning Techniques for Large Language Models Jiahui Geng, Qing Li, Herbert Woisetschlaeger, Zongxiong Chen, Yuxia Wang, Preslav Nakov, Hans-Arno Jacobsen, Fakhri Karray. Arxiv 2025.
A Comprehensive Survey on Long Context Language Modeling Jiaheng Liu, Dawei Zhu, Zhiqi Bai, Yancheng He, Huanxuan Liao, Haoran Que, Zekun Wang, Chenchen Zhang, Ge Zhang, Jiebin Zhang, Yuanxing Zhang, Zhuo Chen, Hangyu Guo, Shilong Li, Ziqiang Liu, Yong Shan, Yifan Song, Jiayi Tian, Wenhao Wu, Zhejian Zhou, Ruijie Zhu, Junlan Feng, Yang Gao, Shizhu He, Zhoujun Li, Tianyu Liu, Fanyu Meng, Wenbo Su, Yingshui Tan, Zili Wang, Jian Yang, Wei Ye, Bo Zheng, Wangchunshu Zhou, Wenhao Huang, Sujian Li, Zhaoxiang Zhang Arxiv 2025.
A Survey of Personalized Large Language Models: Progress and Future Directions Jiahong Liu, Zexuan Qiu, Zhongyang Li, Quanyu Dai, Jieming Zhu, Minda Hu, Menglin Yang, Irwin King. Arxiv 2025.
Prompt Compression for Large Language Models: A Survey Zongqian Li, Yinhong Liu, Yixuan Su, Nigel Collier NAACL 2025.
Cognitive Memory in Large Language Models Lianlei Shan, Shixian Luo, Zezhou Zhu, Yu Yuan, Yong Wu. Arxiv 2025.
Human-inspired Perspectives: A Survey on AI Long-term Memory Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W. Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, Junxiao Shen. Arxiv 2025.
Knowledge Conflicts for LLMs: A Survey Rongwu Xu, Zehan Qi, Zhijiang Guo, Cunxiang Wang, Hongru Wang, Yue Zhang, Wei Xu. EMNLP 2024.
A Survey on the Memory Mechanism of Large Language Model based Agents. Zhang, Zeyu and Bo, Xiaohe and Ma, Chen and Li, Rui and Chen, Xu and Dai, Quanyu and Zhu, Jieming and Dong, Zhenhua and Wen, Ji-Rong. Arxiv 2024.
Knowledge Editing for Large Language Models: A Survey Song Wang, Yaochen Zhu, Haochen Liu, Zaiyi Zheng, Chen Chen, Jundong Li. Arxiv 2024.
Advancing Transformer Architecture in Long-Context Large Language Models: A Comprehensive Survey Yunpeng Huang, Jingwei Xu, Junyu Lai, Zixu Jiang, Taolue Chen, Zenan Li, Yuan Yao, Xiaoxing Ma, Lijuan Yang, Hao Chen, Shupeng Li, Penghao Zhao. Arxiv 2024.
PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory Bowen Jiang, Yuan Yuan, Maohao Shen, Zhuoqun Hao, Zhangchen Xu, Zichen Chen, Zijun Liu, Anirudh Ravi Vijjini, Jiaming He, and others. Arxiv 2025.
O-Mem: Omni Memory System for Personalized, Long Horizon, Self-Evolving Agents Piaohong Wang, Motong Tian, Jiaxian Li, Yuan Liang, Yuqing Wang, Qianben Chen, Tiannan Wang, Zhicong Lu, Jiawei Ma, Yuchen Eleanor Jiang, Wangchunshu Zhou. Arxiv 2025.
MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems Yifan Song, Weimin Xiong, Dawei Zhu, Cheng Li, Ke Wang, and others. Arxiv 2025.
Agent Learning via Early Experience Kai Zhang, Xiangchao Chen, Bo Liu, Tianci Xue, Zeyi Liao, Zhihan Liu, Xiyao Wang, Yuting Ning, Zhaorun Chen, Xiaohan Fu, and others. Arxiv 2025.
G-Memory: Tracing Hierarchical Memory for Multi-Agent Systems Guibin Zhang, Muxin Fu, Guancheng Wan, Miao Yu, Kun Wang, Shuicheng Yan. Arxiv 2025.
HaluMem: Evaluating Hallucinations in Memory Systems of Agents Ding Chen, Simin Niu, Kehang Li, Peng Liu, Xiangping Zheng, Bo Tang, Xinchi Li, Feiyu Xiong, Zhiyu Li. Arxiv 2025.
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and others. Arxiv 2025.
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions Yuanzhe Hu, Yu Wang, Julian McAuley. Arxiv 2025.
Mem0: Building production-ready ai agents with scalable long-term memory Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, Deshraj Yadav. Arxiv 2025.
MemOS: An Operating System for Memory-Augmented Generation (MAG) in Large Language Models Zhiyu Li, Shichao Song, Hanyu Wang, Simin Niu, Ding Chen, Jiawei Yang, Chenyang Xi, Huayi Lai, Jihao Zhao, Yezhaohui Wang, and others. Arxiv 2025.
LightMem: Lightweight and Efficient Memory-Augmented Generation Qingyang Zhang, Ningyu Zhang, and others. Arxiv 2025.
Memory OS of AI Agent Jiazheng Kang, Mingming Ji, Zhe Zhao, Ting Bai. Arxiv 2025.
MemU: An open-source memory framework for AI companions NevaMind-AI. GitHub 2025.
Hierarchical Memory for High-Efficiency Long-Term Reasoning in LLM Agents Haoran Sun, Shaoning Zeng. Arxiv 2025.
ReasoningBank: Scaling Agent Self-Evolving with Reasoning Memory Siru Ouyang, Jun Yan, I-Hung Hsu, Yanfei Chen, Ke Jiang, Zifeng Wang, Rujun Han, Long T. Le, Samira Daruki, Xiangru Tang, Vishy Tirumalashetty, George Lee, Mahsan Rofouei, Hangfei Lin, Jiawei Han, Chen-Yu Lee, Tomas Pfister. Arxiv 2025.
Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models Qizheng Zhang, Changran Hu, Shubhangi Upasani, Boyuan Ma, Fenglu Hong, Vamsidhar Kamanuru, Jay Rainton, Chen Wu, Mengmeng Ji, Hanchen Li, Urmish Thakker, James Zou, Kunle Olukotun. Arxiv 2025.
Coarse-to-Fine Grounded Memory for LLM Agent Planning Wei Yang, Jinwei Xiao, Hongming Zhang, Qingyang Zhang, Yanna Wang, Bo Xu. Arxiv 2025.
Chain-of-Memory: Enhancing GUI Agents for Cross-Application Navigation Xinzge Gao, Chuanrui Hu, Bin Chen, Teng Li. Arxiv 2025.
MEM1: Learning to Synergize Memory and Reasoning for Efficient Long-Horizon Agents Zijian Zhou, Ao Qu, Zhaoxuan Wu, Sunghwan Kim, Alok Prakash, Daniela Rus, Jinhua Zhao, Bryan Kian Hsiang Low, Paul Pu Liang. Arxiv 2025.
Tremu: Towards Neuro-Symbolic Temporal Reasoning for LLM-Agents with Memory in Multi-Session Dialogues Yubin Ge, Salvatore Romeo, Jason Cai, Raphael Shu, Monica Sunkara, Yassine Benajiba, Yi Zhang. Arxiv 2025.
Emergence of Episodic Memory in Transformers: Characterizing Changes in Temporal Structure of Attention Scores During Training Deven Mahesh Mistry, Anooshka Bajaj, Yash Aggarwal, Sahaj Singh Maini, Zoran Tiganj. Arxiv 2025.
Concept-Reversed Winograd Schema Challenge: Evaluating and Improving Robust Reasoning in Large Language Models via Abstraction Kaiqiao Han, Tianqing Fang, Zhaowei Wang, Yangqiu Song, Mark Steedman. NAACL 2025.
Learn to Memorize: Optimizing LLM-based Agents with Adaptive Memory Framework Zeyu Zhang, Quanyu Dai, Rui Li, Xiaohe Bo, Xu Chen, Zhenhua Dong. Arxiv 2025.
Memory-R1: Enhancing large language model agents to manage and utilize memories via reinforcement learning Sikuan Yan, Xiufeng Yang, Zuchao Huang, Ercong Nie, Zifeng Ding, Zonggen Li, Xiaowen Ma, Hinrich Schütze, Volker Tresp, Yunpu Ma. Arxiv 2025.
Mem-$\alpha$: Learning Memory Construction via Reinforcement Learning Yu Wang, Ryuichi Takanobu, Zhiqi Liang, Yuzhen Mao, Yuanzhe Hu, Julian McAuley, Xiaojian Wu. Arxiv 2025.
AgentFly: Fine-tuning LLM Agents without Fine-tuning LLMs Huichi Zhou, Yihang Chen, Siyuan Guo, Xue Yan, Kin Hei Lee, Zihan Wang, Ka Yiu Lee, Guchun Zhang, Kun Shao, Linyi Yang, and others. Arxiv 2025.
MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent Hongli Yu, Tinghong Chen, Jiangtao Feng, Jiangjie Chen, Weinan Dai, Qiying Yu, Ya-Qin Zhang, Wei-Ying Ma, Jingjing Liu, Mingxuan Wang, and others. Arxiv 2025.
Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions Yuanzhe Hu, Yu Wang, Julian McAuley. Arxiv 2025.
MemBench: Towards More Comprehensive Evaluation on the Memory of LLM-based Agents Haoran Tan, Zeyu Zhang, Chen Ma, Xu Chen, Quanyu Dai, Zhenhua Dong. Arxiv 2025.
MemGuide: Intent-Driven Memory Selection for Goal-Oriented Multi-Session LLM Agents Yiming Du, Bingbing Wang, Yang He, Bin Liang, Baojun Wang, Zhongyang Li, Lin Gui, Jeff Z. Pan, Ruifeng Xu, Kam-Fai Wong. Arxiv 2025.
MemTool: Optimizing Short-Term Memory Management for Dynamic Tool Calling in LLM Agent Multi-Turn Conversations Elias Lumer, Anmol Gulati, Vamse Kumar Subbiah, Pradeep Honaganahalli Basavaraju, James A Burke. Arxiv 2025.
Memp: Exploring Agent Procedural Memory Runnan Fang, Yuan Liang, Xiaobin Wang, Jialong Wu, Shuofei Qiao, Pengjun Xie, Fei Huang, Huajun Chen, Ningyu Zhang. Arxiv 2025.
LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory Di Wu, Hongwei Wang, Wenhao Yu, Yuwei Zhang, Kai-Wei Chang, Dong Yu. ICLR 2025.
Compress to Impress: Unleashing the Potential of Compressive Memory in Real-World Long-Term Conversations Nuo Chen, Hongguang Li, Jia Li, Yuxuan Li, Wei Wu. COLING 2025.
Towards Lifelong Dialogue Agents via Timeline-based Memory Management Kai Tzu-iunn Ong, Namyoung Kim, Minju Gwak, Hyungjoo Chae, Taeyoon Kwon, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. NAACL 2025.
Zep: A Temporal Knowledge Graph Architecture for Agent Memory Preston Rasmussen, Daniel Chalef. Arxiv 2025.
From RAG to Memory: Non-Parametric Continual Learning for Large Language Models Bernal Jiménez Gutiérrez, Yiheng Shu, Weijian Qi, Sizhe Zhou, Yu Su. ICML 2025.
Disentangling Memory and Reasoning Ability in Large Language Models Mingyu Jin, Weidi Luo, Sitao Cheng, Xinyi Wang, Wenyue Hua, Ruixiang Tang, William Yang Wang, Yongfeng Zhang. Arxiv 2025.
MemoRAG: Boosting Long Context Processing with Global Memory-Enhanced Retrieval Augmentation Hongjin Qian, Zheng Liu, Peitian Zhang, Kelong Mao, Defu Lian, Zhicheng Dou, Tiejun Huang. Arxiv 2025.
Hello Again! LLM-powered Personalized Agent for Long-term Dialogue Hao Li, Chenghao Yang, An Zhang, Yang Deng, Xiang Wang, Tat-Seng Chua. NAACL 2025.
MemInsight: Autonomous Memory Augmentation for LLM Agents Rana Salama, Jason Cai, Michelle Yuan, Anna Currey, Monica Sunkara, Yi Zhang, Yassine Benajiba. Arxiv 2025.
Interpersonal Memory Matters: A New Task for Proactive Dialogue Utilizing Conversational History Bowen Wu, Wenqing Wang, Haoran Li, Ying Li, Jingsong Yu, Baoxun Wang. Arxiv 2025.
Echo: A Large Language Model with Temporal Episodic Memory WenTao Liu, Ruohua Zhang, Aimin Zhou, Feng Gao, JiaLi Liu. Arxiv 2025.
Improving Factuality with Explicit Working Memory Mingda Chen, Yang Li, Karthik Padthe, Rulin Shao, Alicia Sun, Luke Zettlemoyer, Gargi Ghosh, Wen-tau Yih. Arxiv 2025.
Memorization Over Reasoning? Exposing and Mitigating Verbatim Memorization in Large Language Models' Character Understanding Evaluation Yuxuan Jiang, Francis Ferraro. Arxiv 2025.
Self-Memory Alignment: Mitigating Factual Hallucinations with Generalized Improvement Siyuan Zhang, Yichi Zhang, Yinpeng Dong, Hang Su. Arxiv 2025.
Needle in the Haystack for Memory Based Large Language Models Elliot Nelson, Georgios Kollias, Payel Das, Subhajit Chaudhury, Soham Dan. ICLR 2025.
Evaluating Very Long-Term Conversational Memory of LLM Agents Adyasha Maharana, Dong-Ho Lee, Sergey Tulyakov, Mohit Bansal, Francesco Barbieri, Yuwei Fang. ACL 2024.
MemoryBank: Enhancing Large Language Models with Long-Term Memory Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, Yanlin Wang. AAAI 2024.
A-MEM: Agentic Memory for LLM Agents Wujiang Xu, Kai Mei, Hang Gao, Juntao Tan, Zujie Liang, Yongfeng Zhang. Arxiv 2024.
Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks Zaijing Li, Yuquan Xie, Rui Shao, Gongwei Chen, Dongmei Jiang, Liqiang Nie. NeurIPS 2024.
HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models Bernal Jiménez Gutiérrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, Yu Su. NeurIPS 2024.
"My agent understands me better": Integrating Dynamic Human-like Memory Recall and Consolidation in LLM-Based Agents Yu Hou, Tamoto Yuta, Mayur Tikundi. Arxiv 2024.
Mr.Steve: Instruction-Following Agents in Minecraft with What-Where-When Memory Yu Hou, Tamoto Yuta, Mayur Tikundi. ICLR 2024.
StableSSM: Alleviating the Curse of Memory in State-space Models through Stable Reparameterization Shida Wang, Qianxiao Li. ICML 2024.
Crafting Personalized Agents through Retrieval-Augmented Generation on Editable Memory Graphs Zheng Wang, Zhongyang Li, Zeren Jiang, Dandan Tu, Wei Shi. EMNLP 2024.
Towards Verifiable Text Generation with Evolving Memory and Self-Reflection Hao Sun, Hengyi Cai, Bo Wang, Yingyan Hou, Xiaochi Wei, Shuaiqiang Wang, Yan Zhang, Dawei Yin. EMNLP 2024.
PerLTQA: A Personal Long-Term Memory Dataset for Memory Classification, Retrieval, and Synthesis in Question Answering Yiming Du, Hongru Wang, Zhengyi Zhao, Bin Liang, Baojun Wang, Wanjun Zhong, Zezhong Wang, Kam-Fai Wong. Arxiv 2024.
An Iterative Associative Memory Model for Empathetic Response Generation Zhou Yang, Zhaochun Ren, Yufeng Wang, Haizhou Sun, Chao Chen, Xiaofei Zhu, Xiangwen Liao. ACL 2024.
COCOA: CBT-based Conversational Counseling Agent Using Memory Specialized in Cognitive Distortions and Dynamic Prompt Suyeon Lee, Jieun Kang, Harim Kim, Kyoung-Mee Chung, Dongha Lee, Jinyoung Yeo. Arxiv 2024.
Mixed-Session Conversation with Egocentric Memory Jihyoung Jang, Taeyoung Kim, Hyounghun Kim. EMNLP 2024.
FragRel: Exploiting Fragment-level Relations in the External Memory of Large Language Models Xihang Yue, Linchao Zhu, Yi Yang. ACL 2024.
Extractive Medical Entity Disambiguation with Memory Mechanism and Memorized Entity Information Guobiao Zhang, Xueping Peng, Tao Shen, Guodong Long, Jiasheng Si, Libo Qin, Wenpeng Lu. EMNLP 2024.
Ever-Evolving Memory by Blending and Refining the Past Seo Hyun Kim, Keummin Ka, Yohan Jo, Seung-won Hwang, Dongha Lee, Jinyoung Yeo. Arxiv 2024.
Synapse: Trajectory-as-Exemplar Prompting with Memory for Computer Control Longtao Zheng, Rundong Wang, Xinrun Wang, Bo An. Arxiv 2024.
Moviechat: From dense token to sparse memory for long video understanding Enxin Song, Wenhao Chai, Guanhong Wang, Yucheng Zhang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo, Tian Ye, Yanting Zhang, Yan Lu, Jenq-Neng Hwang, Gaoang Wang. CVPR 2024.
lamp: when large language models meet personalization Alireza Salemi, Sheshera Mysore, Michael Bendersky, Hamed Zamani. ACL 2024.
Evidence-Driven Retrieval Augmented Response Generation for Online Misinformation Zhenrui Yue, Huimin Zeng, Yimeng Lu, Lanyu Shang, Yang Zhang, Dong Wang. NAACL 2024.
IterCQR: Iterative Conversational Query Reformulation with Retrieval Guidance Yunah Jang, Kang-il Lee, Hyunkyung Bae, Hwanhee Lee, Kyomin Jung. NAACL 2024.
Memory Layers at Scale Vincent-Pierre Berges, Barlas Oğuz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, Gargi Ghosh. Arxiv 2024.
MoT: Memory-of-Thought Enables ChatGPT to Self-Improve Sureman Lee, Yujie Qian, Yujia Xie, Yifan Hou, Xinyan Wang, Yiming Yang, Xiang Ren. EMNLP 2023.
Think-in-memory: Recalling and post-thinking enable llms with long-term memory Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, Guannan Zhang. Arxiv 2023.
Recursively Summarizing Enables Long-Term Dialogue Memory in Large Language Models Qingxue Wang, Ling Ding, Yaran Cao, Zhilang Tan, Shi Wang, Dacheng Tao, Liu Qiu. Arxiv 2023.
LLM-based Medical Assistant Personalization with Short- and Long-Term Memory Coordination Yuwei Zhang, Yifan Hou, Xinyan Wang, Yiming Yang, Xiang Ren. Arxiv 2023.
SCM: Enhancing Large Language Model with Self-Controlled Memory Framework Bing Wang, Xinnian Liang, Jian Yang, Hui Huang, Shuangzhi Wu, Peihao Wu, Lu Lu, Zejun Ma, Zhoujun Li. Arxiv 2023.
LDM²: A Large Decision Model Imitating Human Cognition with Dynamic Memory Enhancement Xingjin Wang, Linjing Li, Dongfeng Zeng. EMMNLP 2023.
NarrativeXL: A Large-scale Dataset For Long-Term Memory Models Arseny Moskvichev, Ky-Vinh Mai. EMNLP 2023.
Who's Harry Potter? Approximate Unlearning in LLMs Ronen Eldan, Mark Russinovich. Arxiv 2023.
Active Retrieval Augmented Generation Zhengbao Jiang, Frank Xu, Luyu Gao, Zhiqing Sun, Qian Liu, Jane Dwivedi-Yu, Yiming Yang, Jamie Callan, Graham Neubig. EMNLP 2023.
Prompted LLMs as Chatbot Modules for Long Open-domain Conversation Gibbeum Lee, Volker Hartmann, Jongho Park, Dimitris Papailiopoulos, Kangwook Lee. ACL 2023.
MemoChat: Tuning LLMs to Use Memos for Consistent Long-Range Open-Domain Conversation Junru Lu, Siyu An, Mingbao Lin, Gabriele Pergola, Yulan He, Di Yin, Xing Sun, Yunsheng Wu. Arxiv 2023.
Learning Retrieval Augmentation for Personalized Dialogue Generation Qiushi Huang, Shuai Fu, Xubo Liu, Wenwu Wang, Tom Ko, Yu Zhang, Lilian Tang. EMNLP 2023.
Learning to Reason and Memorize with Self-Notes Jack Lanchantin, Shubham Toshniwal, Jason Weston, Arthur Szlam, Sainbayar Sukhbaatar. NeurIPS 2023.
RECAP: Retrieval-Enhanced Context-Aware Prefix Encoder for Personalized Dialogue Response Generation Shuai Liu, Hyundong Cho, Marjorie Freedman, Xuezhe Ma, Jonathan May. ACL 2023.
Enhancing Personalized Dialogue Generation with Contrastive Latent Variables: Combining Sparse and Dense Persona Yihong Tang, Bo Wang, Miao Fang, Dongming Zhao, Kun Huang, Ruifang He, Yuexian Hou. ACL 2023.
Transformer-based World Models Are Happy With 100k Interactions Jan Robine, Marc Höftmann, Tobias Uelwer, Stefan Harmeling. ICLR 2023.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation Jing Xu, Arthur Szlam, Jason Weston. ACL 2022.
Long Time No See! Open-Domain Conversation with Long-Term Persona Memory Xinchao Xu, Zhibin Gou, Wenquan Wu, Zheng-Yu Niu, Hua Wu, Haifeng Wang, Shihang Wang. ACL 2022.
Keep Me Updated! Memory Management in Long-term Conversations Sanghwan Bae, Donghyun Kwak, Soyoung Kang, Min Young Lee, Sungdong Kim, Yuin Jeong, Hyeri Kim, Sang-Woo Lee, Woomyoung Park, Nako Sung. EMNLP 2022.
Towards Teachable Reasoning Systems: Using a Dynamic Memory of User Feedback for Continual System Improvement Bhavana Dalvi Mishra, Oyvind Tafjord, Peter Clark. EMNLP 2022.
Learning to Repair: Repairing Model Output Errors after Deployment Using a Dynamic Memory of Feedback Niket Tandon, Aman Madaan, Peter Clark, Yiming Yang. NAACL 2022.
There Are a Thousand Hamlets in a Thousand People's Eyes: Enhancing Knowledge-grounded Dialogue with Personal Memory Tingchen Fu, Xueliang Zhao, Chongyang Tao, Ji-Rong Wen, Rui Yan. ACL 2022.
Training Language Models with Memory Augmentation Zexuan Zhong, Tao Lei, Danqi Chen. EMNLP 2022.
Improving Multi-turn Emotional Support Dialogue Generation with Lookahead Strategy Planning Yi Cheng, Wenge Liu, Wenjie Li, Jiashuo Wang, Ruihui Zhao, Bang Liu, Xiaodan Liang, Yefeng Zheng. EMNLP 2022.
Less is More: Learning to Refine Dialogue History for Personalized Dialogue Generation Hanxun Zhong, Zhicheng Dou, Yutao Zhu, Hongjin Qian, Ji-Rong Wen. NAACL 2022.
Leveraging Similar Users for Personalized Language Modeling with Limited Data Charles Welch, Chenxi Gu, Jonathan K. Kummerfeld, Veronica Perez-Rosas, Rada Mihalcea. ACL 2022.
PerKGQA: Question Answering over Personalized Knowledge Graphs Ritam Dutt, Kasturi Bhattacharjee, Rashmi Gangadharaiah, Dan Roth, Carolyn Rose. NAACL 2022.
A Cooperative Memory Network for Personalized Task-oriented Dialogue Systems with Incomplete User Profiles Jiahuan Pei, Pengjie Ren, Maarten de Rijke. WWW 2021.
Episodic Memory in Lifelong Language Learning Cyprien de Masson d'Autume, Sebastian Ruder, Lingpeng Kong, Dani Yogatama. NeurIPS 2019.
Table-1: Datasets for Evaluating Long-Term Memory *(Continuously Updated)
| Dataset | Mo | Operations | DS Type | Per | TR | Metrics | Purpose | Year |
|---|---|---|---|---|---|---|---|---|
| PersonaMem-v2 | text | Updating, Retrieval | MS | ✓ | ✓ | Accuracy, Persona Score | Benchmark implicit user persona learning and agentic memory updates over long contexts. | 2025 |
| MemoryBench | text | Updating, Retrieval, Forgetting | QA | ✗ | ✓ | Accuracy, Retention Rate | Comprehensive benchmark for memory correctness, persistence, and continual learning. | 2025 |
| HaluMem | text | Retrieval | QA | ✗ | ✗ | Accuracy, Hallucination Rate, Omission Rate | Evaluate hallucinations in memory extraction, updating and retrieval. | 2025 |
| BFCL V4 | text (API/Code) | Updating, Retrieval, Reasoning | QA (API) | ✗ | ✗ | AST Accuracy, Execution Success | Benchmarking function-calling capabilities, specifically featuring a "Memory" category for CRUD tool usage. | 2025 |
| LongMemEval | text | Indexing, Retrieval, Compression | MS | ✗ | ✓ | Recall@K, NDCG@K, Accuracy | Benchmark chat assistants on long-term memory abilities, including temporal reasoning. | 2025 |
| LoCoMo | text + image | Indexing, Retrieval, Compression | MS | ✗ | ✓ | Accuracy, ROUGE, Precision, Recall, F1 | Evaluate long-term memory in LLMs across QA, event summarization, and multimodal dialogue tasks. | 2024 |
| MemoryBank | text | Updating, Retrieval | MS | ✓ | ✗ | Accuracy, Human Eval | Enhance LLMs with long-term memory capabilities, adapting to user personalities and contexts. | 2024 |
| PerLTQA | text | Retrieval | MS | ✓ | ✗ | MAP, Recall, Precision, F1, Accuracy, GPT4 score | To explore personal long-term memory question answering ability. | 2024 |
| MALP | text | Retrieval, Compression | QA | ✓ | ✗ | ROUGE, Accuracy, Win Rate | Preference-conditioned dialogue generation. Parameter-efficient fine-tuning (PEFT) for customization. | 2024 |
| DialSim | text | Retrieval | MS | ✓ | ✗ | Accuracy | To evaluate dialogue systems under realistic, real-time, and long-context multi-party conversation conditions. | 2024 |
| MovieChat-1K | text + video | question-answering + caption | QA | ✗ | ✓ | Accuracy | For long-term video understanding for Large Multimodal Models across video question-answering and video captioning tasks. | 2023 |
| CC | text | Retrieval | MS | ✗ | ✓ | BLEU, ROUGE | For long-term dialogue modeling with time and relationship context. | 2023 |
| LAMP | text | Consolidation, Retrieval, Compression | MS | ✓ | ✓ | Accuracy, F1, ROUGE | Multiple entries per user. Supports both user-based splits and time-based splits. | 2023 |
| MSC | text | Consolidation, Retrieval, Compression | MS | ✓ | ✗ | PPL | Evaluate and improve long-term dialogue models via multi-session chats with evolving knowledge. | 2022 |
| DuLeMon | text | Consolidation, Updating, Retrieval, Compression | MS | ✓ | ✗ | Accuracy, F1, Recall, Precision, PPL, BLEU, DISTINCT | For dynamic persona tracking and consistent long-term interaction. | 2022 |
| 2WikiMultiHopQA | table + knowledge base + text | Consolidation, Indexing, Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Multi-hop QA combining structured and unstructured data with reasoning paths. | 2020 |
| NQ | text | Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Open-domain QA based on real Google search queries. | 2019 |
| HotpotQA | text | Retrieval, Compression | QA | ✗ | ✗ | EM, F1 | Multi-hop QA with explainable reasoning and sentence-level supporting facts. | 2018 |
Note:
Table-2: Datasets for Long-Context Memory Evaluation *(Continuously Updated)*
| Dataset | Modality | Operations | Metrics | Purpose | Year |
|---|---|---|---|---|---|
| MMLongBench | text + image | compression, retrieval | SubEM, Accuracy, Rouge-L, Model-Based | 5 categories and 16 datasets for vision-language long-context evaluation | 2025 |
| WikiText-103 | text | compression | PPL | 100M-token Wikipedia corpus for long-context language modeling | 2016 |
| PG-19 | text | compression | PPL | Project Gutenberg books corpus for long-context language modeling | 2019 |
| LRA | text + image | compression, retrieval | Acc | Benchmark with 6 tasks for evaluating efficient long-context language models | 2020 |
| NarrativeQA | text | retrieval | Bleu-1, Bleu-4, Meteor, Rouge-L, MRR | QA dataset for evaluating long-context QA ability | 2017 |
| TriviaQA | text | retrieval | EM, F1 | QA dataset for evaluating long-context QA ability | 2017 |
| NaturalQuestions | text | retrieval | EM, F1 | QA dataset for evaluating long-context QA ability | 2019 |
| MusiQue | text | retrieval | F1 | Multi-hop QA dataset for evaluating long-context reasoning and QA | 2021 |
| CNN/DailyMail | text | compression | Rouge-1, Rouge-2, Rouge-L | News articles dataset for long document summarization | 2016 |
| GovReport | text | compression | Rouge-1, Rouge-2, Rouge-L, Bert Score | Government agency reports for long document summarization | 2021 |
| L-Eval | text | compression, retrieval | Rouge-L, F1, GPT4 | 20-subtask benchmark for diverse long-context language model evaluation | 2023 |
| LongBench | text | compression, retrieval | F1, Rouge-L, Accuracy, EM, Edit Sim | 14 English, 5 Chinese, 2 code tasks for long-context evaluation | 2023 |
| LongBench v2 | text + table + KG | compression, retrieval | Acc | Longer, more challenging tasks with consistent multi-choice format | 2024 |
| SWE-bench | text | compression, retrieval | Resolution rate (%Resolved) | 2,294 task instances from 12 popular python repositories from GitHub | 2023 |
| SWE-bench Multimodal | text + image | compression, retrieval | Resolution rate (%Resolved), Inference cost (Avg. $ Cost) | Extending the original benchmark with image modal with 517 task instances | 2024 |
| $\infty$Bench | text | compression, retrieval | F1, Acc, ROUGE-L-Sum | 12 sub-tasks specially designed for evaluating extreme long context language models | 2024 |
| LooGLE | text | compression, retrieval | Bleu-1, Bleu-4, Rouge-1, Rouge-4, Rouge-L, Meteor score, Bert score, GPT4 score | 7 major tasks specially designed for evaluating extreme long context language models | 2023 |
Table-3: Datasets for Parametric Memory Evaluation *(Continuously Updated)*
| Dataset | Modality | Operations | Metrics | Purpose | Year |
|---|---|---|---|---|---|
| KnowEdit | text | updating | Edit Success, Portability, Locality, Fluency | 6 datasets covering insertion, modification, and erasure | 2024 |
| MQUAKE-CF | text | updating | Edit-wise Success Rate, Instance-wise Accuracy, Multi-hop Accuracy | Counterfactual knowledge editing through multi-hop reasoning (up to 4 hops) | 2023 |
| MQUAKE-T | text | updating | Edit-wise Success Rate, Instance-wise Accuracy, Multi-hop Accuracy | Temporal knowledge editing with one edit per reasoning chain | 2023 |
| Counterfact | text | updating | Efficacy Score, Magnitude, Paraphrase & Neighborhood Scores | Tests substantial factual changes beyond superficial edits | 2022 |
| zsRE | text | updating | Success Rate, Retain Accuracy, Equivalence Accuracy, Perf. Deterioration | One of the earliest datasets for knowledge editing | 2021 |
| MUSE | text | forgetting | VerbMem, KnowMem, PrivLeak | Unlearning benchmark with 6 desirable properties | 2024 |
| KnowUnDo | text | forgetting | Unlearn Success, Retention Success, Perplexity, ROUGE-L | Test unlearning in copyrighted and privacy-sensitive domains | 2024 |
| RWKU | text | forgetting | ROUGE-L | Real-world unlearning under corpus-free, adversarial settings | 2024 |
| WMDP | text | forgetting | QA accuracy | Proxy for hazardous knowledge in bio/cyber/chemical domains | 2024 |
| TOFU | text | forgetting | Probability, ROUGE, Truth Ratio | Unlearning dataset of facts about 200 fictitious authors | 2024 |
| ABSA | text | consolidation | F1 | Aspect-based sentiment analysis for continual learning | 2024 |
| SGD | text | consolidation | JGA, FWT, BWT | Multi-turn task-oriented dialogue with evolving intents | 2020 |
| INSPIRED | text | consolidation | JGA, FWT, BWT | Task-oriented dialogue supporting user goal evolution | 2020 |
| Natural Question | text | consolidation | Indexing Accuracy, Hits@1 | Supports continual learning over evolving document corpora | 2019 |
Note:
Table-4: Datasets for Multi-Source Memory Evaluation *(Continuously Updated)*
| Dataset | Mo | Ops | Src# | Mod# | Task | Metrics | Purpose | Year |
|---|---|---|---|---|---|---|---|---|
| MultiChat | text + image | Retrieval | 2 | 2 | Retrieval | Precision, mAP, GPT-4 | Image-grounded sticker retrieval with cross-session image-text dialogue context. | 2025 |
| Context-conflicting | text | Compression | 2 | 1 | Conflict | DiffGR, EM, Similarity | Evaluates model handling of conflicting evidence across sources. | 2024 |
| EgoSchema | video + text | Retrieval, Compression | 3 | 2 | Fusion | Accuracy | Episodic video + social schema + conversation for long-term memory QA. | 2023 |
| Ego4D NLQ | video + text | Retrieval, Compression | 2 | 2 | Fusion | Recall@K | Natural language queries over egocentric video with temporal memory. | 2022 |
| 2WikiMultihopQA | text | Indexing, Retrieval, Compression | 2 | 1 | Reasoning | EM, F1 | Multi-hop QA across Wikipedia passages with sentence-level support. | 2020 |
| HybridQA | text | Retrieval, Compression | 2 | 1 | Reasoning | EM, F1 | Reasoning across structured tables and unstructured text. | 2020 |
| CommonsenseVQA | text + image | Retrieval, Compression | 2 | 2 | Fusion | Accuracy | Commonsense QA over visual scenes requiring visual-textual fusion. | 2019 |
| NaturalQuestions | text | Retrieval, Compression | >1* | 1 | Conflict | EM, F1 | QA over Google snippets; used for contradiction analysis. | 2019 |
| ComplexWebQuestions | text | Retrieval, Compression | >1* | 1 | Reasoning | EM, F1 | Compositional QA requiring multi-step reasoning over web snippets. | 2018 |
| HotpotQA | text | Retrieval, Compression | 2 | 1 | Conflict | EM, F1, Supporting Fact Accuracy | Multi-hop QA with paragraph- and sentence-level support. | 2018 |
| TriviaQA | text | Retrieval, Compression | ≥6 | 1 | Conflict | EM, F1 | QA with noisy web sources; useful for source disagreement analysis. | 2017 |
| WebQuestionsSP | text | Indexing, Retrieval, Compression | >1* | 1 | Reasoning | F1, Accuracy | Structured QA dataset with enhanced reasoning chains. | 2016 |
| Flickr30K | text + image | Retrieval, Compression | 2 | 2 | Retrieval | Similarity | Image-caption pairs for cross-modal retrieval and alignment. | 2014 |
Note:
Table-1: Component-Level Tools for Memory Management and Utilization. *(Continuously Updated)*
| Memory Tool | Function | Input/Output | Example Use |
|---|---|---|---|
| FAISS | Library for fast storage, indexing, and retrieval of high-dimensional vectors | Vector / Index, relevance score | Indexing large sets of text embeddings and retrieving relevant documents in RAG systems |
| Neo4j | Native graph database supporting ACID transactions and Cypher query language | Nodes and relationships with properties / Query results via Cypher | Modeling and retrieving complex relational data for use cases like fraud detection and recommendation engines |
| Chroma | AI-native embedding database for building LLM applications | Text / Embeddings | Managing knowledge, facts, and skills for LLMs |
| Milvus | Vector database for embedding similarity search and AI applications | Embeddings / Similar items | Unstructured data search and similarity matching |
| Qdrant | Vector similarity search engine and database | Embeddings / Similar items | Production-ready service with user-friendly API for vector search |
| Weaviate | Open-source vector database with built-in ML models | Data objects and vector embeddings / Search results | Scalable storage and retrieval for AI applications |
| BM25 | Probabilistic ranking function for estimating document relevance | Text queries / Ranked list of documents | Enhancing search engine results and document retrieval systems |
| Contriever | Unsupervised dense retriever trained with contrastive learning | Query text / List of similar documents | High-recall retrieval tasks in multilingual question-answering systems |
| Embedding Models (e.g., OpenAI) | Convert text, images, or audio into dense vector representations capturing semantic meaning | Raw data / Vector embeddings | Text similarity computation, recommendation systems, and clustering tasks |
Table-2: Framework-Level Tools for Memory Management and Utilization *(Continuously Updated)*
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Graphiti | Framework for building and querying temporally-aware knowledge graphs tailored for AI agents in dynamic environments | Multi-source data / Queryable knowledge graph | Constructing real-time knowledge graphs to enhance AI agent memory | Open |
| LlamaIndex | A flexible framework for building knowledge assistants using LLMs connected to enterprise data | Text / Context-augmented responses | Developing knowledge assistants that process complex data formats | Open |
| LangChain | Provides a framework for building context-aware, reasoning applications by connecting LLMs with external data sources | Input prompts / Multi-step reasoning outputs | Creating complex LLM applications like question-answering systems and chatbots | Open |
| LangGraph | Constructs controllable agent architectures supporting long-term memory and human-in-the-loop multi-agent systems | Graph state / State updates | Building complex task workflows with multiple AI agents | Open |
| EasyEdit | An easy-to-use knowledge editing framework for LLMs, enabling efficient behavior modification within specific domains | Edit instructions / Updated model behavior | Modifying LLM knowledge in specific domains, such as updating factual information | Open |
| CrewAI | A platform for building and deploying multi-agent systems, supporting automated workflows using any LLM and cloud platform | Multi-agent tasks / Collaborative results | Automating workflows across agents like project management and content generation | Open |
| Letta | Constructs stateful agents with long-term memory, advanced reasoning, and custom tools within a visual environment | User interactions / Improved response | Developing AI agents that learn and improve over time | Open |
| OpenHands | An open platform for autonomous software agents that maintains persistent context across file editing, command execution, and web browsing | Natural language tasks / Code patches, Terminal actions | Automating complex software engineering tasks like debugging and feature implementation with full project context | Open |
Table-3: Application Layer-Level Tools for Memory Management and Utilization (Continuously Updated)
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Mem0 | Provides a smart memory layer for LLMs, enabling direct addition, updating, and searching of memories in models | User interactions / Personalized responses | Enhancing AI systems with persistent context for customer support and personalized recommendations | Open |
| Zep | Integrates chat messages into a knowledge graph, offering accurate and relevant user information | Chat logs, business data / Knowledge graph query results | Augmenting AI agents with knowledge through continuous learning from user interactions | Open |
| Memary | An open memory layer that emulates human memory to help AI agents manage and utilize information effectively | Agent tasks / Memory management and utilization | Building AI agents with human-like memory characteristics | Open |
| Memobase | A user profile-based long-term memory system designed to provide personalized experiences in generative AI applications | User interactions / Personalized responses | Implementing virtual assistants, educational tools, and personalized AI companions | Open |
| O-Mem | An omni-memory system enabling agents to self-evolve and maintain long-horizon consistency through recursive memory consolidation | Long-term interaction logs / Evolved memory state | Creating self-evolving personal AI assistants that adapt to user growth over time | Open |
| MemOS | An operating system-like architecture that manages memory hierarchy (working/short/long-term) to optimize Memory-Augmented Generation | Agent queries, Complex contexts / Hierarchical memory blocks | Managing complex memory resources for agents handling multi-step reasoning tasks | Open |
Table-4: Product-Level Tools for Memory Utilization (Continuously Updated)
| Memory Tool | Function | Input/Output | Example Use | Source Type |
|---|---|---|---|---|
| Me.bot | AI-powered personal assistant that organizes notes, tasks, and memories, providing emotional support and productivity tools | User inputs (text, voice) / Organized notes, reminders, summaries | Personal productivity enhancement, emotional support, idea organization | Closed |
| ima.copilot | Intelligent workstation powered by Tencent's Mix Huang model, building a personal knowledge base for learning and work scenarios | User queries / Customized responses, knowledge retrieval | Enhancing learning efficiency, work productivity, knowledge management | Closed |
| Coze | Enables multi-agent collaboration across various platforms | User-defined workflows / Response | Deployed chatbots, AI agents | Closed |
| Grok | AI assistant developed by xAI, designed to provide truthful, useful, and curious responses, with real-time data access and image generation | Query / Informative answers, generated images | Answering questions, generating images, providing insights | Closed |
| ChatGPT | Conversational AI developed by OpenAI, capable of understanding and generating human-like text based on prompts | User prompts / Generated text responses | Answering questions, generating images, providing insights | Closed |
| Claude | AI assistant featuring "Projects" to ground answers in user-provided knowledge bases and massive context windows | Prompts, Files, Code / Text, Code, Artifacts | Analyzing large codebases, maintaining consistent style across documents via Projects | Closed |
| Doubao | A high-efficiency multimodal AI assistant capable of handling long-context interactions and diverse tasks | Text, Voice, Image / Answers, creative content | Daily conversation, writing assistance, coding, and role-playing | Closed |
| Siri | Intelligent voice assistant utilizing on-device personal semantic memory for cross-app actions and context understanding | Voice commands / Action execution, personal info retrieval | Device control, retrieving personal context ("When is Mom's flight?"), cross-app tasks | Closed |
| Xiaoyi | Huawei's smart assistant integrated into HarmonyOS, leveraging ecosystem memory for proactive services and document processing | Voice, Text, Documents / Summaries, suggestions, IoT control | Document summarization, smart home control, personalized travel planning | Closed |
| Zhixiaobao | Ant Group's financial AI agent that utilizes user financial history and market knowledge for personalized wealth management | Financial queries / Market analysis, investment advice | Financial planning, insurance analysis, market trend explanation | Closed |
Table-5: Key differences between human and agent memory across operational dimensions
| Aspect | Human Memory | Agent Memory |
|---|---|---|
| Storage | Distributed, interconnected neural systems across brain regions | Parametric, modular, and context-dependent (structured or unstructured) |
| Consolidation | Slow, biologically driven, passive | Fast, explicit, policy-driven and selective |
| Indexing | Implicit, associative, sparse codes via hippocampal circuits | Explicit, embedding-based, symbolic or key–value lookup |
| Updating | Indirect, reconsolidation-based, error-prone | Precise, programmable, supports rollback/unlearning |
| Forgetting | Passive decay or interference | Transparent, trackable, policy-controlled |
| Retrieval | Cue/context/emotion dependent, emotionally biased | Content-based, reproducible, similarity or query driven |
| Compression | Implicit, salience- and frequency-biased | Explicit, customizable (e.g., quantization, summarization) |
| Ownership | Individual and private | Shareable, replicable, and broadcastable |
| Volume | Biologically limited | Scalable, bounded only by storage and compute limits |
Please contact me if I miss your names in the list, I will add you back ASAP!
🤝🤝 Thanks for all the great contributors on GitHub!
If you find our repository and survey useful for your research, please consider citing the following paper:
@article{du2025rethinking,
title={Rethinking Memory in AI: Taxonomy, Operations, Topics, and Future Directions},
author={Du, Yiming and Huang, Wenyu and Zheng, Danna and Wang, Zhaowei and Montella, Sebastien and Lapata, Mirella and Wong, Kam-Fai and Pan, Jeff Z.},
journal={arXiv preprint arXiv:2505.00675},
year={2025},
url={https://arxiv.org/abs/2505.00675}
}