Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants. Unlike benchmarks built around static dialogues, RHELM provides realistic, heterogeneous, and temporally evolving memory sources, together with challenging questions that require multi-hop reasoning, temporal synthesis, and hallucination detection.
β οΈ All characters, events, and personal details in this dataset are fully synthetic. Any resemblance to real individuals is coincidental.
| Item | Count |
|---|---|
| Characters (personas) | 10 |
| QA pairs | 1,305 |
Conversation sessions (.json) | 629 |
Emails (.txt) | 625 |
Attachments (.md / .html) | 1,053 |
| Type | Count |
|---|---|
| attachment | 249 |
| mixed | 210 |
| fact | 207 |
| hallucination | 197 |
| aggregation | 192 |
| temporal | 185 |
| misleading | 65 |
data/ (uploaded to repo root)
βββ conversations/<Character>/*.json # dated dialogue sessions
βββ emails/<Character>/*.txt # email threads
βββ attachments/<Character>/*.md|*.html# documents, notes, reports
βββ QA_final/low_score_qa_<Character>_all_validated.jsonl
Each line in a QA_final/*.jsonl file is a JSON object:
| Field | Description |
|---|---|
id | Unique question identifier |
question | The user query |
answer | Ground-truth answer |
question_date | Date the question is asked from |
question_type | One of: fact, temporal, hallucination, aggregation, misleading, attachment, mixed |
supporting_evidence | References to source items (e.g. "2024-10-13:1" or "56_report_task_*.md:Section") |
characteristics | Fine-grained challenge labels (see taxonomy) |
from datasets import load_dataset
qa = load_dataset("microsoft/RHELM", data_files="QA_final/*.jsonl", split="train")
print(qa[0])
To work with the full multi-source context (conversations, emails, attachments), download the repository snapshot:
from huggingface_hub import snapshot_download
local_dir = snapshot_download("microsoft/RHELM", repo_type="dataset")
RHELM organizes questions into 7 categories with 26 challenge characteristics across three QA domains: Dialogue History QA, External Source QA, and Hybrid Context QA. See the evaluation code repository for the full taxonomy and benchmark harness.
Released under CC BY 4.0.
1 commits
Benchmarking Realistic, Heterogeneous, and Evolving Long-Horizon Memory
RHELM is a benchmark for evaluating long-horizon memory capabilities in AI assistants. Unlike benchmarks built around static dialogues, RHELM provides realistic, heterogeneous, and temporally evolving memory sources, together with challenging questions that require multi-hop reasoning, temporal synthesis, and hallucination detection.
β οΈ All characters, events, and personal details in this dataset are fully synthetic. Any resemblance to real individuals is coincidental.
| Item | Count |
|---|---|
| Characters (personas) | 10 |
| QA pairs | 1,305 |
Conversation sessions (.json) | 629 |
Emails (.txt) | 625 |
Attachments (.md / .html) | 1,053 |
| Type | Count |
|---|---|
| attachment | 249 |
| mixed | 210 |
| fact | 207 |
| hallucination | 197 |
| aggregation | 192 |
| temporal | 185 |
| misleading | 65 |
data/ (uploaded to repo root)
βββ conversations/<Character>/*.json # dated dialogue sessions
βββ emails/<Character>/*.txt # email threads
βββ attachments/<Character>/*.md|*.html# documents, notes, reports
βββ QA_final/low_score_qa_<Character>_all_validated.jsonl
Each line in a QA_final/*.jsonl file is a JSON object:
| Field | Description |
|---|---|
id | Unique question identifier |
question | The user query |
answer | Ground-truth answer |
question_date | Date the question is asked from |
question_type | One of: fact, temporal, hallucination, aggregation, misleading, attachment, mixed |
supporting_evidence | References to source items (e.g. "2024-10-13:1" or "56_report_task_*.md:Section") |
characteristics | Fine-grained challenge labels (see taxonomy) |
from datasets import load_dataset
qa = load_dataset("microsoft/RHELM", data_files="QA_final/*.jsonl", split="train")
print(qa[0])
To work with the full multi-source context (conversations, emails, attachments), download the repository snapshot:
from huggingface_hub import snapshot_download
local_dir = snapshot_download("microsoft/RHELM", repo_type="dataset")
RHELM organizes questions into 7 categories with 26 challenge characteristics across three QA domains: Dialogue History QA, External Source QA, and Hybrid Context QA. See the evaluation code repository for the full taxonomy and benchmark harness.
Released under CC BY 4.0.
1 commits