Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.
ir_datasets lotte/pooled/dev), 2,428,854 passageslotte/pooled/dev/search, 8,573 qrels)colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)[CLS] and the ColBERT [D] marker (token id 2)doc_maxlen=180[MASK] expansion, no zero padding)| Token vectors (N) | 266,205,513 |
| Avg vectors per document | 109.6 (min 3, max 180) |
| Vectors per query | 32 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | uint16 (<u2) | [266205513, 128] | Raw float16 bit patterns stored as uint16. Read with .view(np.float16) |
doclens.npy | int32 | [2428854] | Vectors per document; sum == N |
token_ids.npy | int64 | [266205513] | Input token id of each row of documents.npy |
queries.npy | float32 | [2931, 32, 128] | Query vectors |
queries_ids.npy | int64 | [2931] | LoTTE qid of each query (0 to 2930) |
Document ids are LoTTE passage ids and equal the row index: row i of doclens.npy is passage i.
documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16.
In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).
The official LoTTE pooled dev search qrels: 8,573 relevant query-passage pairs over the
2,931 queries, all with relevance 1 (ir_datasets lotte/pooled/dev/search).
Standard metric: Success@5.
Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.
ir_datasets lotte/pooled/dev), 2,428,854 passageslotte/pooled/dev/search, 8,573 qrels)colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)[CLS] and the ColBERT [D] marker (token id 2)doc_maxlen=180[MASK] expansion, no zero padding)| Token vectors (N) | 266,205,513 |
| Avg vectors per document | 109.6 (min 3, max 180) |
| Vectors per query | 32 |
| File | dtype | Shape | Content |
|---|---|---|---|
documents.npy | uint16 (<u2) | [266205513, 128] | Raw float16 bit patterns stored as uint16. Read with .view(np.float16) |
doclens.npy | int32 | [2428854] | Vectors per document; sum == N |
token_ids.npy | int64 | [266205513] | Input token id of each row of documents.npy |
queries.npy | float32 | [2931, 32, 128] | Query vectors |
queries_ids.npy | int64 | [2931] | LoTTE qid of each query (0 to 2930) |
Document ids are LoTTE passage ids and equal the row index: row i of doclens.npy is passage i.
documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16.
In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).
The official LoTTE pooled dev search qrels: 8,573 relevant query-passage pairs over the
2,931 queries, all with relevance 1 (ir_datasets lotte/pooled/dev/search).
Standard metric: Success@5.