A retrieval dataset that exposes fundamental theoretical limitations of embedding-based retrieval models. Despite using simple queries like "Who likes Apples?", state-of-the-art embedding models achieve less than 20% recall@100 on LIMIT full and cannot solve LIMIT-small (46 docs).
You can load the data using the datasets library from Huggingface (LIMIT, LIMIT-small):
from datasets import load_dataset
ds = load_dataset("orionweller/LIMIT-small", "corpus") # also available: queries, test (contains qrels).
Queries (1,000): Simple questions asking "Who likes [attribute]?"
Corpus (50k documents): Short biographical texts describing people and their preferences
Qrels (2,000): Each query has exactly 2 relevant documents (score=1), creating nearly all possible combinations of 2 documents from the 46 corpus documents (C(46,2) = 1,035 combinations).
The dataset follows standard MTEB format with three configurations:
default: Query-document relevance judgments (qrels), keys: corpus-id, query-id, score (1 for relevant)queries: Query texts with IDs , keys: _id, textcorpus: Document texts with IDs, keys: _id, title (empty), and textTests whether embedding models can represent all top-k combinations of relevant documents, based on theoretical results connecting embedding dimension to representational capacity. Despite the simple nature of queries, state-of-the-art models struggle due to fundamental dimensional limitations.
@misc{weller2025theoreticallimit,
title={On the Theoretical Limitations of Embedding-Based Retrieval},
author={Orion Weller and Michael Boratko and Iftekhar Naim and Jinhyuk Lee},
year={2025},
eprint={2508.21038},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2508.21038},
}
8 commits
2 commits
A retrieval dataset that exposes fundamental theoretical limitations of embedding-based retrieval models. Despite using simple queries like "Who likes Apples?", state-of-the-art embedding models achieve less than 20% recall@100 on LIMIT full and cannot solve LIMIT-small (46 docs).
You can load the data using the datasets library from Huggingface (LIMIT, LIMIT-small):
from datasets import load_dataset
ds = load_dataset("orionweller/LIMIT-small", "corpus") # also available: queries, test (contains qrels).
Queries (1,000): Simple questions asking "Who likes [attribute]?"
Corpus (50k documents): Short biographical texts describing people and their preferences
Qrels (2,000): Each query has exactly 2 relevant documents (score=1), creating nearly all possible combinations of 2 documents from the 46 corpus documents (C(46,2) = 1,035 combinations).
The dataset follows standard MTEB format with three configurations:
default: Query-document relevance judgments (qrels), keys: corpus-id, query-id, score (1 for relevant)queries: Query texts with IDs , keys: _id, textcorpus: Document texts with IDs, keys: _id, title (empty), and textTests whether embedding models can represent all top-k combinations of relevant documents, based on theoretical results connecting embedding dimension to representational capacity. Despite the simple nature of queries, state-of-the-art models struggle due to fundamental dimensional limitations.
@misc{weller2025theoreticallimit,
title={On the Theoretical Limitations of Embedding-Based Retrieval},
author={Orion Weller and Michael Boratko and Iftekhar Naim and Jinhyuk Lee},
year={2025},
eprint={2508.21038},
archivePrefix={arXiv},
primaryClass={cs.IR},
url={https://arxiv.org/abs/2508.21038},
}
8 commits
2 commits