Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.
ir_datasets msmarco-passage), 8,841,823 passagesmsmarco-passage/dev/small)colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)[CLS] and the ColBERT [D] marker (token id 2)[MASK] expansion, no zero padding)| Token vectors (N) | 597,909,919 |
| Avg vectors per document | 67.6 (min 4, max 300) |
| Vectors per query | 32 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | uint16 (<u2) | [597909919, 128] | Raw float16 bit patterns stored as uint16. Read with .view(np.float16) |
doclens.npy | int32 | [8841823] | Vectors per document; sum == N |
token_ids_per_token.npy | int64 | [597909919] | Input token id of each row of documents.npy |
doc_ids.npy | int64 | [8841823] | MS MARCO pid (equals the row index) |
queries.npy | float32 | [6980, 32, 128] | Query vectors |
queries_ids.npy | int64 | [6980] | MS MARCO qid of each query |
Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.
documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16.
In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).
The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant
query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small).
Standard metric: MRR@10 (RR@10 in ir_measures).
Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.
ir_datasets msmarco-passage), 8,841,823 passagesmsmarco-passage/dev/small)colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)[CLS] and the ColBERT [D] marker (token id 2)[MASK] expansion, no zero padding)| Token vectors (N) | 597,909,919 |
| Avg vectors per document | 67.6 (min 4, max 300) |
| Vectors per query | 32 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | uint16 (<u2) | [597909919, 128] | Raw float16 bit patterns stored as uint16. Read with .view(np.float16) |
doclens.npy | int32 | [8841823] | Vectors per document; sum == N |
token_ids_per_token.npy | int64 | [597909919] | Input token id of each row of documents.npy |
doc_ids.npy | int64 | [8841823] | MS MARCO pid (equals the row index) |
queries.npy | float32 | [6980, 32, 128] | Query vectors |
queries_ids.npy | int64 | [6980] | MS MARCO qid of each query |
Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.
documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16.
In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).
The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant
query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small).
Standard metric: MRR@10 (RR@10 in ir_measures).