Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
ir_datasets (beir/nfcorpus/test); PyLate only did the encodingtitle + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectorsdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npycolbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acfartifact.metadata in the model repository. We did not override any of them| Token vectors (N) | 561,120 |
| Avg vectors per document | 154.5 (max 173) |
| Vectors per query | always 32 (padded with [MASK] expansion tokens, which are real embeddings) |
| Avg vectors per query | 32.0 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [561120, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [3633] | Vectors per document; sum == N |
token_ids.npy | uint32 | [561120] | Input token id of each row of documents.npy |
doc_ids.npy | string | [3633] | BEIR doc id of each document |
queries.npy | float32 | [323, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [323] | True number of vectors per query |
queries_ids.npy | string | [323] | BEIR query id of each query |
qrels.test.tsv | TREC | 12334 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 32300 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [323, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [323, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.3299 | 0.2768 |
Token-level (late-interaction) embeddings of the BEIR NFCorpus corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.
ir_datasets (beir/nfcorpus/test); PyLate only did the encodingtitle + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectorsdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npycolbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acfartifact.metadata in the model repository. We did not override any of them| Token vectors (N) | 561,120 |
| Avg vectors per document | 154.5 (max 173) |
| Vectors per query | always 32 (padded with [MASK] expansion tokens, which are real embeddings) |
| Avg vectors per query | 32.0 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [561120, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [3633] | Vectors per document; sum == N |
token_ids.npy | uint32 | [561120] | Input token id of each row of documents.npy |
doc_ids.npy | string | [3633] | BEIR doc id of each document |
queries.npy | float32 | [323, 32, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [323] | True number of vectors per query |
queries_ids.npy | string | [323] | BEIR query id of each query |
qrels.test.tsv | TREC | 12334 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 32300 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [323, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [323, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.3299 | 0.2768 |