Collection of models predictions on [BABILong dataset](https://huggingface.co/datasets/RMT-team/babilong).
0
stars
11
commits
1
linked in READMEs
May 3, 2025
updated
Collection of models predictions on BABILong dataset.
9 commits
1 commits
aybora/scorers_datasets_eval
princeton-nlp/HELMET
10
yale-nlp/MMVU-evaluation-results
LJ0815/DevEval
JUNJIE99/VISTA_Evaluation
3
AntResearchNLP/ViLaSR-eval
ASLP-lab/WSC-Eval
7
KwaiVGI/FullBench
booydar/babilong
BABILong is a benchmark for LLM evaluation using the needle-in-a-haystack approach.
257