Unified eval mix for Latent Context Language Models (LCLM). Four
benchmarks, one schema ({prompt, category, extra_info}), one repo.
| Config | Source | Rows | Notes |
|---|---|---|---|
ruler | tonychenxyz/ruler-full (memwrap, validation) | 39,000 | 13 tasks × 6 ctx lengths × 500 |
gsm8k | tonychenxyz/codellava-gsm8k-memwrap | 1,319 | grade-school math word problems |
longhealth5 | leonli66/longhealth5 (memwrap, test) | 400 | 5-doc patient-record QA |
longbench | nimitkalra/LongBench-v1 (memwrap, validation) | 4,750 | 21 English+Chinese long-context tasks |
Every row has three columns:
prompt (str): full chat-formatted prompt with
<|memory_start|>...<|memory_end|> markers wrapping the context to
be compressed by the LCLM encoder.category (str): task-and-length tag (e.g. niah_single_1_4096,
narrativeqa).extra_info (dict): per-task metadata including
ground_truth.answers (list of acceptable strings),
scoring_function (string-match flavor), and original-task fields.from datasets import load_dataset
ds = load_dataset("latent-context/lclm-eval", "ruler", split="test")
print(ds[0]["prompt"][:200])
print(ds[0]["category"])
print(ds[0]["extra_info"]["ground_truth"])
The LCLM eval pipeline reads extra_info.scoring_function per sample.
For RULER subtasks this is ruler_string_match_all /
ruler_string_match_part (official NVIDIA RULER reference impl,
case-insensitive substring match). For other benchmarks see the LCLM
benchmark code.
latent-context/0.6b-4b-LCLM-{4,8,16}x3 commits
Unified eval mix for Latent Context Language Models (LCLM). Four
benchmarks, one schema ({prompt, category, extra_info}), one repo.
| Config | Source | Rows | Notes |
|---|---|---|---|
ruler | tonychenxyz/ruler-full (memwrap, validation) | 39,000 | 13 tasks × 6 ctx lengths × 500 |
gsm8k | tonychenxyz/codellava-gsm8k-memwrap | 1,319 | grade-school math word problems |
longhealth5 | leonli66/longhealth5 (memwrap, test) | 400 | 5-doc patient-record QA |
longbench | nimitkalra/LongBench-v1 (memwrap, validation) | 4,750 | 21 English+Chinese long-context tasks |
Every row has three columns:
prompt (str): full chat-formatted prompt with
<|memory_start|>...<|memory_end|> markers wrapping the context to
be compressed by the LCLM encoder.category (str): task-and-length tag (e.g. niah_single_1_4096,
narrativeqa).extra_info (dict): per-task metadata including
ground_truth.answers (list of acceptable strings),
scoring_function (string-match flavor), and original-task fields.from datasets import load_dataset
ds = load_dataset("latent-context/lclm-eval", "ruler", split="test")
print(ds[0]["prompt"][:200])
print(ds[0]["category"])
print(ds[0]["extra_info"]["ground_truth"])
The LCLM eval pipeline reads extra_info.scoring_function per sample.
For RULER subtasks this is ruler_string_match_all /
ruler_string_match_part (official NVIDIA RULER reference impl,
case-insensitive substring match). For other benchmarks see the LCLM
benchmark code.
latent-context/0.6b-4b-LCLM-{4,8,16}x3 commits