Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
ir_datasets (beir/fiqa/test); PyLate only did the encodingdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npycolbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acfartifact.metadata in the model repository. We did not override any of them| Token vectors (N) | 6,054,600 |
| Avg vectors per document | 105.0 (max 179) |
| Vectors per query | always 32 (padded with [MASK] expansion tokens, which are real embeddings) |
| Avg vectors per query | 32.0 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [6054600, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [57638] | Vectors per document; sum == N |
token_ids.npy | uint32 | [6054600] | Input token id of each row of documents.npy |
doc_ids.npy | string | [57638] | BEIR doc id of each document |
queries.npy | float32 | [648, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [648] | True number of vectors per query |
queries_ids.npy | string | [648] | BEIR query id of each query |
qrels.test.tsv | TREC | 1706 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 64800 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [648, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [648, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.3472 | 0.6277 |
Token-level (late-interaction) embeddings of the BEIR FiQA-2018 corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
ir_datasets (beir/fiqa/test); PyLate only did the encodingdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npycolbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acfartifact.metadata in the model repository. We did not override any of them| Token vectors (N) | 6,054,600 |
| Avg vectors per document | 105.0 (max 179) |
| Vectors per query | always 32 (padded with [MASK] expansion tokens, which are real embeddings) |
| Avg vectors per query | 32.0 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [6054600, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [57638] | Vectors per document; sum == N |
token_ids.npy | uint32 | [6054600] | Input token id of each row of documents.npy |
doc_ids.npy | string | [57638] | BEIR doc id of each document |
queries.npy | float32 | [648, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [648] | True number of vectors per query |
queries_ids.npy | string | [648] | BEIR query id of each query |
qrels.test.tsv | TREC | 1706 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 64800 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [648, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [648, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.3472 | 0.6277 |