Synthetic data generation and benchmark implementation for "Episodic Memories Generation and Evaluation Benchmark for Large Language Models" [ICLR 2025]
Jupyter Notebook
69
36 commits
updated Oct 3, 2025
In honor of Endel Tulving (1927-2023), whose pioneering work on episodic memory inspired this work and fundamentally shaped our understanding of human memory systems.
Accepted at ICLR 2025
We evaluated different models focusing on two critical aspects of episodic memory:
| Model | ๐ฏ Simple Recall | โฑ๏ธ Chronological Awareness |
|---|---|---|
| gemini-2.5-pro๐ยฒ | 0.968 ๐ฅ | 0.796 ๐ฅ |
| gemini-2.5-flash๐ยฒ | 0.960 ๐ฅ | 0.817 ๐ฅ |
| gpt-5๐ยฒ | 0.942 ๐ฅ | 0.804 ๐ฅ |
| gpt-5-mini๐ยฒ | 0.830 | 0.442 |
| claude-sonnet-4๐ยฒ | 0.790 | 0.326 |
| grok-4-fast-reasoning๐ยฒ | 0.726 | 0.281 |
| gemini-2-pro๐ยน | 0.708 | 0.290 |
| gemini-2-flash-thinking๐ยน | 0.708 | 0.288 |
| gpt-4o | 0.670 | 0.204 |
| grok-4-fast-non-reasoning๐ยฒ | 0.602 | 0.122 |
| deepseek-v3๐ยน | 0.600 | 0.103 |
| gemini-2-flash๐ยน | 0.596 | 0.173 |
| deepseek-r1๐ยน | 0.572 | 0.147 |
| llama-3.1-405b | 0.504 | 0.129 |
| gpt-4o-mini | 0.492 | 0.077 |
| claude-3-haiku | 0.470 | 0.109 |
| claude-3-5-sonnet | 0.470 | 0.090 |
| o3-mini๐ยน | 0.424 | 0.044 |
| o1๐ยน | 0.384 | 0.052 |
| gpt-4.1-nano๐ยฒ | 0.356 | 0.090 |
| o1-mini | 0.300 | 0.033 |
๐ยน Evaluated after paper acceptance (February'25) ๐ยฒ Evaluated after paper acceptance (September'25)
Note: for deepseek-v3, deepseek-r1, and llama-3.1-405b, we used openrouter to run the experiments.
[Please refer to our paper for additional details.]
Our framework generates book-like narratives where each event includes a date, location, entity name, and detailed content - capturing the core elements of episodic memory.
Events are designed to have recurring elements. For example, in one of the produced dataset, the entity "Jackson Ramos" appears in five different chapters, enabling evaluation of the model's ability to track and recall multiple occurrences of the same entity across the narrative. Click to view tracking of Jackson Ramos over NYC map
The benchmark includes a comprehensive set of episodic memory questions based on cue composition and retrieval types: (full list of 36 question templates in the paper)
| Cueย ย ย ย ย ย ย ย ย ย ย ย ย ย | Description | Retrieved traceย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย | Template question for โ |
|---|---|---|---|
| (t, *, *, *) | Events at a specific time | - Spaces - Entities โ - Contents | List all protagonists involved in events on {t}. |
| (*, s, e, *) | Events involving entities at a specific location | - Times - Contents โ | List events involving {e} at {s}. |
| (*, s, e, c) | Events with specific location, entities, and content | - Times โ | When did events with {e} and {c} occur at {s}? |
| (t, s, e, c) | Events with specific time, location, entities, and content | - Event details โ | What happened with {e} and {c} at {s} on {t}? |
| (*, *, e, *) | Retrieves the most recent known location of an entity | - Times [latest] - Spaces [latest] โ - Contents [latest] | Where was {e} last seen? |
| (*, *, e, *) | Retrieves a chronological list of dates when an entity was observed | - Times [chrono.] โ
- Spaces [chrono.] - Contents [chrono.] | List all dates {e} was observed. |
Here's how we evaluate the episodic memory capabilities of a specific LLM (e.g., GPT-4 with in-context memory):
Question: Events dates for 'Jackson Ramos'?
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Model Answer: Ground Truth: โ
โ โข Sep 22, 2026 โข Sep 22, 2026 โ
โ โข Jun 14, 2025 โข Feb 27, 2026 โ
โ โข Aug 24, 2026 โ
โ โข Apr 09, 2026 โ
โ โข Jun 14, 2025 โ
โโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
F1 Score: 0.57
We release 11 synthetic episodic memory datasets, generated using our automated framework. These datasets vary in size and diversity.
Evaluation of LLM performance using these datasets is provided in the paper, and the raw data is available in the Installation section below.
| book style | length | generated with | variation | chapters | tokens | used in |
|---|---|---|---|---|---|---|
| default | short | claude-3-5-sonnet | / | 20 | 10k | main |
| default | long | claude-3-5-sonnet | / | 200 | 100k | main |
| default | very long | claude-3-5-sonnet | / | 2000 | 1M | / |
| default | short | claude-3-5-sonnet | ordered | 20 | 10k | ablation |
| default | long | claude-3-5-sonnet | ordered | 200 | 100k | ablation |
| default | short | gpt-4o | / | 20 | 14k | ablation |
| default | long | gpt-4o | / | 200 | 125k | ablation |
| world news | short | claude-3-5-sonnet | / | 20 | 7k | ablation |
| world news | long | claude-3-5-sonnet | / | 200 | 69k | ablation |
| sci-fi | short | claude-3-5-sonnet | / | 20 | 9k | ablation |
| sci-fi | long | claude-3-5-sonnet | / | 200 | 89k | ablation |
Below is a fictional chapter from our synthetically generated book "world news", demonstrating our methodology.
In a dramatic turn of events on May 11, 2026, Benjamin Green found himself documenting the rapid transformation of peaceful suburban streets into raging torrents of muddy water. The local meteorological station's emergency sirens blared through the rain-soaked air as Hamza Avila and Koa Berlin, emergency response coordinators, rushed to evacuate residents from the low-lying areas. Rising waters had already submerged vehicles to their windows, while the relentless downpour continued to intensify, creating treacherous conditions across the region.
As the situation in New South Wales deteriorated, Benjamin witnessed a flash flood emergency that would later be described as unprecedented in its ferocity. Water levels rose at an alarming rate of nearly one meter per hour, prompting Emilia Hooks, a veteran emergency services spokesperson, to declare it a "catastrophic event." The flood's destructive force was evident as debris-laden waters crashed through streets, uprooting trees and damaging infrastructure. Local authorities reported that over 300 residents were evacuated to emergency shelters, while rescue teams conducted more than 50 water rescues throughout the affected areas. The disaster response teams continue to monitor the situation as meteorologists predict additional rainfall in the coming hours.
git clone https://github.com/ahstat/episodic-memory-benchmark.git
conda create -n epbench
...
pip install -r requirements.txt
The data is available at this address: https://doi.org/10.6084/m9.figshare.28244480
It contains the generated documents, the question/answer pairs, and the evaluation results over all the models presented in the paper.
The downloaded zip file should be uncompressed into the data repository as follows:
episodic-memory-benchmark-main
โโโ epbench
โโโ data
โ โโโ Udefault_Sdefault_seed0
โ โโโ UdefaultOrdered_Sdefault_seed0
โ โโโ Unews_Snews_seed1
โ โโโ Uscifi_Sscifi_seed2
โโโ experiments
โโโ plots
โโโ src
.env file:For running new experiments with OpenAI, Anthropic, or OpenRouter models, the following variables are used:
PROXY = {"http": "xxxxxx", "https": "xxxxxx"}
OPENAI_API_KEY = 'sk-xxxxxx'
ANTHROPIC_API_KEY = 'sk-xxxxxx'
OPENROUTER_API_KEY = 'sk-xxxxxx'
To verify the setup, you can run a quick test using a single benchmark (short book generated from 20 events) with a single model (GPT-4o mini using in-context memory). Run the following command, which should complete within seconds if the data is already available. If not, it will automatically proceed through the generation, answering, and evaluation steps.
python ./epbench/experiments/quickstart.py --data_folder epbench/data --env_file .env --book_nb_events 20 --answering_kind prompting --answering_model_name gpt-4o-mini-2024-07-18
For reproducing the experiments and ablation studies presented in the paper, you can use the provided Jupyter notebooks, available in epbench/experiments/.
For the experiments in the main body of the paper:
step_1_generation.ipynb: generation of the documents with two rounds of verifications, and extraction of the selected question/answer pairsstep_2_answering.ipynb: predicting the answers given the document and the questions, using in-context, RAG, or fine-tuned models, and perform the evaluationsstep_3_results.ipynb: extract the results, including the CD plots and the summarizing tablesFor additional experiments and ablation studies available in the appendix:
rebuttal_ablation_on_news_and_scifi_books.ipynb (evaluating the world news and the scifi short books with gpt-4o)rebuttal_ablation_ordered_books.ipynb (building and evaluation of the ordered book)rebuttal_ablation_with_gpt_book.ipynb (applying the experiment on the GPT generated book)rebuttal_generating_book_variations.ipynb (building the world news, the scifi, and the very long default books)rebuttal_hallucinations_0_matching_events.ipynb (manual analysis of the hallucinations observed in the gpt-4o answers when there is 0 matching events)rebuttal_llama3.ipynb (evaluating the short default book with llama 3.1 405b and llama 3.2 3b)rebuttal_map.ipynb (illustration of the shared universe structure)rebuttal_realistic_partition_and_evaluation.ipynb (assessment of the degree of realism of each event and evaluation in the difference of performance between the realistic and the non-realistic events).@article{2025epmembench,
title={Episodic Memories Generation and Evaluation Benchmark for Large Language Models},
author={Huet, Alexis and Ben Houidi, Zied and Rossi, Dario},
journal={International Conference on Learning Representations},
year={2025}
}
MIT License - see LICENSE for details.
Jupyter Notebook
89.3%
Python
8.7%
HTML
2.0%
Synthetic data generation and benchmark implementation for "Episodic Memories Generation and Evaluation Benchmark for Large Language Models" [ICLR 2025]
Jupyter Notebook
69
36 commits
updated Oct 3, 2025
In honor of Endel Tulving (1927-2023), whose pioneering work on episodic memory inspired this work and fundamentally shaped our understanding of human memory systems.
Accepted at ICLR 2025
We evaluated different models focusing on two critical aspects of episodic memory:
| Model | ๐ฏ Simple Recall | โฑ๏ธ Chronological Awareness |
|---|---|---|
| gemini-2.5-pro๐ยฒ | 0.968 ๐ฅ | 0.796 ๐ฅ |
| gemini-2.5-flash๐ยฒ | 0.960 ๐ฅ | 0.817 ๐ฅ |
| gpt-5๐ยฒ | 0.942 ๐ฅ | 0.804 ๐ฅ |
| gpt-5-mini๐ยฒ | 0.830 | 0.442 |
| claude-sonnet-4๐ยฒ | 0.790 | 0.326 |
| grok-4-fast-reasoning๐ยฒ | 0.726 | 0.281 |
| gemini-2-pro๐ยน | 0.708 | 0.290 |
| gemini-2-flash-thinking๐ยน | 0.708 | 0.288 |
| gpt-4o | 0.670 | 0.204 |
| grok-4-fast-non-reasoning๐ยฒ | 0.602 | 0.122 |
| deepseek-v3๐ยน | 0.600 | 0.103 |
| gemini-2-flash๐ยน | 0.596 | 0.173 |
| deepseek-r1๐ยน | 0.572 | 0.147 |
| llama-3.1-405b | 0.504 | 0.129 |
| gpt-4o-mini | 0.492 | 0.077 |
| claude-3-haiku | 0.470 | 0.109 |
| claude-3-5-sonnet | 0.470 | 0.090 |
| o3-mini๐ยน | 0.424 | 0.044 |
| o1๐ยน | 0.384 | 0.052 |
| gpt-4.1-nano๐ยฒ | 0.356 | 0.090 |
| o1-mini | 0.300 | 0.033 |
๐ยน Evaluated after paper acceptance (February'25) ๐ยฒ Evaluated after paper acceptance (September'25)
Note: for deepseek-v3, deepseek-r1, and llama-3.1-405b, we used openrouter to run the experiments.
[Please refer to our paper for additional details.]
Our framework generates book-like narratives where each event includes a date, location, entity name, and detailed content - capturing the core elements of episodic memory.
Events are designed to have recurring elements. For example, in one of the produced dataset, the entity "Jackson Ramos" appears in five different chapters, enabling evaluation of the model's ability to track and recall multiple occurrences of the same entity across the narrative. Click to view tracking of Jackson Ramos over NYC map
The benchmark includes a comprehensive set of episodic memory questions based on cue composition and retrieval types: (full list of 36 question templates in the paper)
| Cueย ย ย ย ย ย ย ย ย ย ย ย ย ย | Description | Retrieved traceย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย ย | Template question for โ |
|---|---|---|---|
| (t, *, *, *) | Events at a specific time | - Spaces - Entities โ - Contents | List all protagonists involved in events on {t}. |
| (*, s, e, *) | Events involving entities at a specific location | - Times - Contents โ | List events involving {e} at {s}. |
| (*, s, e, c) | Events with specific location, entities, and content | - Times โ | When did events with {e} and {c} occur at {s}? |
| (t, s, e, c) | Events with specific time, location, entities, and content | - Event details โ | What happened with {e} and {c} at {s} on {t}? |
| (*, *, e, *) | Retrieves the most recent known location of an entity | - Times [latest] - Spaces [latest] โ - Contents [latest] | Where was {e} last seen? |
| (*, *, e, *) | Retrieves a chronological list of dates when an entity was observed | - Times [chrono.] โ
- Spaces [chrono.] - Contents [chrono.] | List all dates {e} was observed. |
Here's how we evaluate the episodic memory capabilities of a specific LLM (e.g., GPT-4 with in-context memory):
Question: Events dates for 'Jackson Ramos'?
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ Model Answer: Ground Truth: โ
โ โข Sep 22, 2026 โข Sep 22, 2026 โ
โ โข Jun 14, 2025 โข Feb 27, 2026 โ
โ โข Aug 24, 2026 โ
โ โข Apr 09, 2026 โ
โ โข Jun 14, 2025 โ
โโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโ
โ
โผ
F1 Score: 0.57
We release 11 synthetic episodic memory datasets, generated using our automated framework. These datasets vary in size and diversity.
Evaluation of LLM performance using these datasets is provided in the paper, and the raw data is available in the Installation section below.
| book style | length | generated with | variation | chapters | tokens | used in |
|---|---|---|---|---|---|---|
| default | short | claude-3-5-sonnet | / | 20 | 10k | main |
| default | long | claude-3-5-sonnet | / | 200 | 100k | main |
| default | very long | claude-3-5-sonnet | / | 2000 | 1M | / |
| default | short | claude-3-5-sonnet | ordered | 20 | 10k | ablation |
| default | long | claude-3-5-sonnet | ordered | 200 | 100k | ablation |
| default | short | gpt-4o | / | 20 | 14k | ablation |
| default | long | gpt-4o | / | 200 | 125k | ablation |
| world news | short | claude-3-5-sonnet | / | 20 | 7k | ablation |
| world news | long | claude-3-5-sonnet | / | 200 | 69k | ablation |
| sci-fi | short | claude-3-5-sonnet | / | 20 | 9k | ablation |
| sci-fi | long | claude-3-5-sonnet | / | 200 | 89k | ablation |
Below is a fictional chapter from our synthetically generated book "world news", demonstrating our methodology.
In a dramatic turn of events on May 11, 2026, Benjamin Green found himself documenting the rapid transformation of peaceful suburban streets into raging torrents of muddy water. The local meteorological station's emergency sirens blared through the rain-soaked air as Hamza Avila and Koa Berlin, emergency response coordinators, rushed to evacuate residents from the low-lying areas. Rising waters had already submerged vehicles to their windows, while the relentless downpour continued to intensify, creating treacherous conditions across the region.
As the situation in New South Wales deteriorated, Benjamin witnessed a flash flood emergency that would later be described as unprecedented in its ferocity. Water levels rose at an alarming rate of nearly one meter per hour, prompting Emilia Hooks, a veteran emergency services spokesperson, to declare it a "catastrophic event." The flood's destructive force was evident as debris-laden waters crashed through streets, uprooting trees and damaging infrastructure. Local authorities reported that over 300 residents were evacuated to emergency shelters, while rescue teams conducted more than 50 water rescues throughout the affected areas. The disaster response teams continue to monitor the situation as meteorologists predict additional rainfall in the coming hours.
git clone https://github.com/ahstat/episodic-memory-benchmark.git
conda create -n epbench
...
pip install -r requirements.txt
The data is available at this address: https://doi.org/10.6084/m9.figshare.28244480
It contains the generated documents, the question/answer pairs, and the evaluation results over all the models presented in the paper.
The downloaded zip file should be uncompressed into the data repository as follows:
episodic-memory-benchmark-main
โโโ epbench
โโโ data
โ โโโ Udefault_Sdefault_seed0
โ โโโ UdefaultOrdered_Sdefault_seed0
โ โโโ Unews_Snews_seed1
โ โโโ Uscifi_Sscifi_seed2
โโโ experiments
โโโ plots
โโโ src
.env file:For running new experiments with OpenAI, Anthropic, or OpenRouter models, the following variables are used:
PROXY = {"http": "xxxxxx", "https": "xxxxxx"}
OPENAI_API_KEY = 'sk-xxxxxx'
ANTHROPIC_API_KEY = 'sk-xxxxxx'
OPENROUTER_API_KEY = 'sk-xxxxxx'
To verify the setup, you can run a quick test using a single benchmark (short book generated from 20 events) with a single model (GPT-4o mini using in-context memory). Run the following command, which should complete within seconds if the data is already available. If not, it will automatically proceed through the generation, answering, and evaluation steps.
python ./epbench/experiments/quickstart.py --data_folder epbench/data --env_file .env --book_nb_events 20 --answering_kind prompting --answering_model_name gpt-4o-mini-2024-07-18
For reproducing the experiments and ablation studies presented in the paper, you can use the provided Jupyter notebooks, available in epbench/experiments/.
For the experiments in the main body of the paper:
step_1_generation.ipynb: generation of the documents with two rounds of verifications, and extraction of the selected question/answer pairsstep_2_answering.ipynb: predicting the answers given the document and the questions, using in-context, RAG, or fine-tuned models, and perform the evaluationsstep_3_results.ipynb: extract the results, including the CD plots and the summarizing tablesFor additional experiments and ablation studies available in the appendix:
rebuttal_ablation_on_news_and_scifi_books.ipynb (evaluating the world news and the scifi short books with gpt-4o)rebuttal_ablation_ordered_books.ipynb (building and evaluation of the ordered book)rebuttal_ablation_with_gpt_book.ipynb (applying the experiment on the GPT generated book)rebuttal_generating_book_variations.ipynb (building the world news, the scifi, and the very long default books)rebuttal_hallucinations_0_matching_events.ipynb (manual analysis of the hallucinations observed in the gpt-4o answers when there is 0 matching events)rebuttal_llama3.ipynb (evaluating the short default book with llama 3.1 405b and llama 3.2 3b)rebuttal_map.ipynb (illustration of the shared universe structure)rebuttal_realistic_partition_and_evaluation.ipynb (assessment of the degree of realism of each event and evaluation in the difference of performance between the realistic and the non-realistic events).@article{2025epmembench,
title={Episodic Memories Generation and Evaluation Benchmark for Large Language Models},
author={Huet, Alexis and Ben Houidi, Zied and Rossi, Dario},
journal={International Conference on Learning Representations},
year={2025}
}
MIT License - see LICENSE for details.
Jupyter Notebook
89.3%
Python
8.7%
HTML
2.0%