15 repos
bigainlco/LooGLE
📜**Introduction**
17
106 commits
bigai-nlco/LooGLE
ACL 2024 | LooGLE: Long Context Evaluation for Long-Context Language Models
200
27 commits
booydar/babilong
BABILong is a benchmark for LLM evaluation using the needle-in-a-haystack approach.
257
86 commits
TIGER-AI-Lab/LongICLBench
Code and Data for "Long-context LLMs Struggle with Long In-context Learning" [TMLR2025]
113
latent-context/lclm-eval
LCLM evaluation datasets
1
3 commits
RMT-team/babilong
BABILong (100 samples) : a long-context needle-in-a-haystack benchmark for LLMs
21
99 commits
Lemoncoke/Marathon
Dataset Card for Marathon
4
8 commits
RMT-team/babilong-1k-samples
BABILong (1000 samples) : a long-context needle-in-a-haystack benchmark for LLMs
19 commits
Penguin-Scrolls/PenguinScrolls
No description
7
2 commits
princeton-nlp/HELMET
HELMET: How to Evaluate Long-context Language Models Effectively and Thoroughly
10
5 commits
RMT-team/babilong-train-5k-samples
BABILong (5k train samples) : a long-context needle-in-a-haystack benchmark for LLMs
21 commits
TIGER-Lab/LongICLBench
This is the benchmark we adopt in our TMLR2025 paper [Long-context LLMs Struggle with Long…
6
6 commits
Hambaobao/Marathon
Marathon: A Multiple-choice Long Context Evaluation Benchmark for Large Language Models.
The HELMET Benchmark
227
95 commits
PenguinScrolls: A User-Aligned Fine-Grained Benchmark for Long-Context Language Model Evaluation