zineddine/MemoReason

Dataset

MemoReason

8

12 commits

updated Sep 28, 2026

See the code

README

MemoReason

MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.

The benchmark contains 1,200 task templates, from which 12,000 examples are generated for each of nine entity-replacement settings. Together with the 1,200 factual examples, this yields 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.

Subsets

SplitDescription
factualOriginal documents; no entity replacement (0%).
fictionalAll eligible annotated entities replaced (100%).
fictional_namedNamed identities replaced; numbers and dates retained.
fictional_numtempNumbers and temporal values replaced; named identities retained.
fictional_{10,20,30,50,80,90}pctThe indicated percentage of eligible entity IDs replaced, rounded to a whole number; the rest remain factual.

There are 1,200 factual examples and 12,000 examples per replacement setting. These are paired experimental conditions, not independent train/test splits. API names retain fictional for compatibility; the replacement content is fictitious.

Load

from datasets import load_dataset

data = load_dataset("zineddine/MemoReason", "paper_scoring")
factual = data["factual"]
fictitious = data["fictional"]

Each row contains id, document, question, answer, question_type, answer_type, and canary_guid (a contamination marker).

IDs follow {setting}:{document_id}_{variant_id}:{question_id}. Pair factual and fictitious examples by document and question, not row order; split the middle component at its last underscore.

answer_type records the original task group: variant, invariant, or refusal. For scoring, use the reference answers and accepted alternatives in evaluation/references.jsonl.gz, together with the paper's evaluation code; single-string exact matching alone is insufficient.

License

CC BY-SA 4.0, with upstream notices in LICENSE.md. Source attribution is in SOURCE_ATTRIBUTION.json. Fictitious passages are transformed texts, not factual claims.

zineddine/MemoReason

Dataset

MemoReason

8

12 commits

updated Sep 28, 2026

See the code

README

MemoReason

MemoReason evaluates how entity familiarity affects document-grounded reasoning in language models. It pairs factual passages with fictitious variants that preserve task structure and specified reasoning operations.

The benchmark contains 1,200 task templates, from which 12,000 examples are generated for each of nine entity-replacement settings. Together with the 1,200 factual examples, this yields 109,200 question-answer examples across ten settings. Questions cover extraction, arithmetic, temporal reasoning, and inference.

Subsets

SplitDescription
factualOriginal documents; no entity replacement (0%).
fictionalAll eligible annotated entities replaced (100%).
fictional_namedNamed identities replaced; numbers and dates retained.
fictional_numtempNumbers and temporal values replaced; named identities retained.
fictional_{10,20,30,50,80,90}pctThe indicated percentage of eligible entity IDs replaced, rounded to a whole number; the rest remain factual.

There are 1,200 factual examples and 12,000 examples per replacement setting. These are paired experimental conditions, not independent train/test splits. API names retain fictional for compatibility; the replacement content is fictitious.

Load

from datasets import load_dataset

data = load_dataset("zineddine/MemoReason", "paper_scoring")
factual = data["factual"]
fictitious = data["fictional"]

Each row contains id, document, question, answer, question_type, answer_type, and canary_guid (a contamination marker).

IDs follow {setting}:{document_id}_{variant_id}:{question_id}. Pair factual and fictitious examples by document and question, not row order; split the middle component at its last underscore.

answer_type records the original task group: variant, invariant, or refusal. For scoring, use the reference answers and accepted alternatives in evaluation/references.jsonl.gz, together with the paper's evaluation code; single-string exact matching alone is insufficient.

License

CC BY-SA 4.0, with upstream notices in LICENSE.md. Source attribution is in SOURCE_ATTRIBUTION.json. Fictitious passages are transformed texts, not factual claims.