MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.
The benchmark contains 1,200 task templates, from which 12,000 examples are generated for each of nine entity-replacement settings. Together with the 1,200 factual examples, this yields 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.
| Split | Description |
|---|---|
factual | Original documents; no entity replacement (0%). |
fictional | All eligible annotated entities replaced (100%). |
fictional_named | Named identities replaced; numbers and dates retained. |
fictional_numtemp | Numbers and temporal values replaced; named identities retained. |
fictional_{10,20,30,50,80,90}pct | The indicated percentage of eligible entity IDs replaced, rounded to a whole number; the rest remain factual. |
There are 1,200 factual examples and 12,000 examples per replacement setting. These are paired experimental conditions, not independent train/test splits. API names retain fictional for compatibility; the replacement content is fictitious.
from datasets import load_dataset
data = load_dataset("zineddine/MemoReason", "paper_scoring")
factual = data["factual"]
fictitious = data["fictional"]
Each row contains id, document, question, answer, question_type, answer_type, and canary_guid (a contamination marker).
IDs follow {setting}:{document_id}_{variant_id}:{question_id}. Pair factual and fictitious examples by document and question, not row order; split the middle component at its last underscore.
answer_type records the original task group: variant, invariant, or refusal. For scoring, use the reference answers and accepted alternatives in evaluation/references.jsonl.gz, together with the paper's evaluation code; single-string exact matching alone is insufficient.
CC BY-SA 4.0, with upstream notices in LICENSE.md. Source attribution is in SOURCE_ATTRIBUTION.json. Fictitious passages are transformed texts, not factual claims.
MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.
The benchmark contains 1,200 task templates, from which 12,000 examples are generated for each of nine entity-replacement settings. Together with the 1,200 factual examples, this yields 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.
| Split | Description |
|---|---|
factual | Original documents; no entity replacement (0%). |
fictional | All eligible annotated entities replaced (100%). |
fictional_named | Named identities replaced; numbers and dates retained. |
fictional_numtemp | Numbers and temporal values replaced; named identities retained. |
fictional_{10,20,30,50,80,90}pct | The indicated percentage of eligible entity IDs replaced, rounded to a whole number; the rest remain factual. |
There are 1,200 factual examples and 12,000 examples per replacement setting. These are paired experimental conditions, not independent train/test splits. API names retain fictional for compatibility; the replacement content is fictitious.
from datasets import load_dataset
data = load_dataset("zineddine/MemoReason", "paper_scoring")
factual = data["factual"]
fictitious = data["fictional"]
Each row contains id, document, question, answer, question_type, answer_type, and canary_guid (a contamination marker).
IDs follow {setting}:{document_id}_{variant_id}:{question_id}. Pair factual and fictitious examples by document and question, not row order; split the middle component at its last underscore.
answer_type records the original task group: variant, invariant, or refusal. For scoring, use the reference answers and accepted alternatives in evaluation/references.jsonl.gz, together with the paper's evaluation code; single-string exact matching alone is insufficient.
CC BY-SA 4.0, with upstream notices in LICENSE.md. Source attribution is in SOURCE_ATTRIBUTION.json. Fictitious passages are transformed texts, not factual claims.