Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
ir_datasets (beir/scidocs); PyLate only did the encodingtitle + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectorsdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npylightonai/GTE-ModernColBERT-v1 @ 25f6f7bb8237b7ae25ae1d9b805ce17c0d1cc639config_sentence_transformers.json in the model repository. We did not override any of them| Token vectors (N) | 4,819,277 |
| Avg vectors per document | 187.8 (max 300) |
| Vectors per query | variable, 6 to 48 (no query expansion), zero-padded to 48 |
| Avg vectors per query | 17.5 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [4819277, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [25657] | Vectors per document; sum == N |
token_ids.npy | uint32 | [4819277] | Input token id of each row of documents.npy |
doc_ids.npy | string | [25657] | BEIR doc id of each document |
queries.npy | float32 | [1000, 48, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [1000] | True number of vectors per query |
queries_ids.npy | string | [1000] | BEIR query id of each query |
qrels.test.tsv | TREC | 29928 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 100000 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [1000, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [1000, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.1949 | 0.4407 |
Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with GTE-ModernColBERT, in the TACHIOM multivector format.
ir_datasets (beir/scidocs); PyLate only did the encodingtitle + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectorsdoc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npylightonai/GTE-ModernColBERT-v1 @ 25f6f7bb8237b7ae25ae1d9b805ce17c0d1cc639config_sentence_transformers.json in the model repository. We did not override any of them| Token vectors (N) | 4,819,277 |
| Avg vectors per document | 187.8 (max 300) |
| Vectors per query | variable, 6 to 48 (no query expansion), zero-padded to 48 |
| Avg vectors per query | 17.5 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | float16 (<f2) | [4819277, 128] | All document vectors, concatenated document by document |
doclens.npy | int32 | [25657] | Vectors per document; sum == N |
token_ids.npy | uint32 | [4819277] | Input token id of each row of documents.npy |
doc_ids.npy | string | [25657] | BEIR doc id of each document |
queries.npy | float32 | [1000, 48, 128] | Query vectors, zero-padded at the end |
query_lens.npy | int32 | [1000] | True number of vectors per query |
queries_ids.npy | string | [1000] | BEIR query id of each query |
qrels.test.tsv | TREC | 29928 lines | qid \t 0 \t docid \t relevance |
groundtruth/gt_top100.tsv | TSV | 100000 lines | Exhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions) |
groundtruth/gt_ids.npy | int32 | [1000, 100] | Same, as doc positions |
groundtruth/gt_scores.npy | float32 | [1000, 100] | Same, MaxSim scores |
Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.
Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.
| nDCG@10 | R@100 |
|---|---|
| 0.1949 | 0.4407 |