THUIR/MemoryBench-Full

Dataset

MemoryBench

3

2 commits

2 linked in READMEs

updated Dec 8, 2025

See the code

README

MemoryBench

MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.

Paper Link: https://arxiv.org/abs/2510.17281

Github: https://github.com/LittleDinoC/MemoryBench/

This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the paper) are, respectively, subsets of the training and test sets of this dataset and maintain the same proportion. However, please note that the sizes of different datasets vary significantly, which may require special handling when calculating averages to prevent any single dataset from disproportionately influencing the results.

For guidance on how to use this dataset, please refer to https://huggingface.co/datasets/THUIR/MemoryBench.

Citation

If you use MemoryBench in your research, please cite our paper:

@misc{ai2025memorybenchbenchmarkmemorycontinual,
      title={MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems}, 
      author={Qingyao Ai and Yichen Tang and Changyue Wang and Jianming Long and Weihang Su and Yiqun Liu},
      year={2025},
      eprint={2510.17281},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2510.17281}, 
}

Contributors

bebr2

2 commits

THUIR/MemoryBench-Full

Dataset

MemoryBench

3

2 commits

2 linked in READMEs

updated Dec 8, 2025

See the code

README

MemoryBench

MemoryBench aims to provide a standardized and extensible benchmark for evaluating memory and continual learning in LLM systems — encouraging future work toward more adaptive, feedback-driven, and efficient LLM systems.

Paper Link: https://arxiv.org/abs/2510.17281

Github: https://github.com/LittleDinoC/MemoryBench/

This is an extended version of MemoryBench. The training and test sets of THUIR/MemoryBench(the balanced version on which we conducted experiments in the paper) are, respectively, subsets of the training and test sets of this dataset and maintain the same proportion. However, please note that the sizes of different datasets vary significantly, which may require special handling when calculating averages to prevent any single dataset from disproportionately influencing the results.

For guidance on how to use this dataset, please refer to https://huggingface.co/datasets/THUIR/MemoryBench.

Citation

If you use MemoryBench in your research, please cite our paper:

@misc{ai2025memorybenchbenchmarkmemorycontinual,
      title={MemoryBench: A Benchmark for Memory and Continual Learning in LLM Systems}, 
      author={Qingyao Ai and Yichen Tang and Changyue Wang and Jianming Long and Weihang Su and Yiqun Liu},
      year={2025},
      eprint={2510.17281},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2510.17281}, 
}

Contributors

bebr2

2 commits