Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with LateOn, in the TACHIOM multivector format.
ir_datasets (beir/fiqa/test); PyLate only did the encodingdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npylightonai/LateOn @ 62911e105059585d244384c7d17826e35f669c17config_sentence_transformers.json in the model repository. We did not override any of them| Token vectors (N) | 7,695,260 |
| Avg vectors per document | 133.5 (max 300) |
| Vectors per query | variable, 7 to 32 (no query expansion), zero-padded to 32 |
| Avg vectors per query | 16.7 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [7695260, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [57638] | Vectors per document; sum == N |
token_ids.npy | uint32 | [7695260] | Input token id of each row of documents.npy |
doc_ids.npy | string | [57638] | BEIR doc id of each document |
queries.npy | float32 | [648, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [648] | True number of vectors per query |
queries_ids.npy | string | [648] | BEIR query id of each query |
qrels.test.tsv | TREC | 1706 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 64800 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [648, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [648, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.5250 | 0.8353 |
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with LateOn, in the TACHIOM multivector format.
ir_datasets (beir/fiqa/test); PyLate only did the encodingdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npylightonai/LateOn @ 62911e105059585d244384c7d17826e35f669c17config_sentence_transformers.json in the model repository. We did not override any of them| Token vectors (N) | 7,695,260 |
| Avg vectors per document | 133.5 (max 300) |
| Vectors per query | variable, 7 to 32 (no query expansion), zero-padded to 32 |
| Avg vectors per query | 16.7 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [7695260, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [57638] | Vectors per document; sum == N |
token_ids.npy | uint32 | [7695260] | Input token id of each row of documents.npy |
doc_ids.npy | string | [57638] | BEIR doc id of each document |
queries.npy | float32 | [648, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [648] | True number of vectors per query |
queries_ids.npy | string | [648] | BEIR query id of each query |
qrels.test.tsv | TREC | 1706 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 64800 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [648, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [648, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.5250 | 0.8353 |