tuskanny/ms_marco_colbertv2

Dataset

MS MARCO v1 Passage, ColBERTv2

0

6 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

MS MARCO v1 Passage, ColBERTv2

Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.

Source

  • Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages
  • Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small)
  • Document order: passage id order (row i is pid i)

Encoding

  • Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
  • Documents start with [CLS] and the ColBERT [D] marker (token id 2)
  • Queries: always 32 vectors ([MASK] expansion, no zero padding)
  • Vectors: 128-d, L2-normalized
  • Not recorded: exact checkpoint revision, encoding library, document length cap (the longest document has 300 vectors), and whether punctuation was dropped

Statistics

Token vectors (N)597,909,919
Avg vectors per document67.6 (min 4, max 300)
Vectors per query32

Files

FiledtypeShapeContent
documents.npyuint16 (<u2)[597909919, 128]Raw float16 bit patterns stored as uint16. Read with .view(np.float16)
doclens.npyint32[8841823]Vectors per document; sum == N
token_ids_per_token.npyint64[597909919]Input token id of each row of documents.npy
doc_ids.npyint64[8841823]MS MARCO pid (equals the row index)
queries.npyfloat32[6980, 32, 128]Query vectors
queries_ids.npyint64[6980]MS MARCO qid of each query

Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.

documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).

Relevance judgments

The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small). Standard metric: MRR@10 (RR@10 in ir_measures).

colbert
late-interaction
msmarco
multivector

tuskanny/ms_marco_colbertv2

Dataset

MS MARCO v1 Passage, ColBERTv2

0

6 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

MS MARCO v1 Passage, ColBERTv2

Token-level (late-interaction) ColBERTv2 embeddings of the MS MARCO v1 passage collection and the dev/small queries.

Source

  • Collection: MS MARCO v1 passage (ir_datasets msmarco-passage), 8,841,823 passages
  • Queries: dev/small, 6,980 queries and 7,437 qrels (msmarco-passage/dev/small)
  • Document order: passage id order (row i is pid i)

Encoding

  • Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
  • Documents start with [CLS] and the ColBERT [D] marker (token id 2)
  • Queries: always 32 vectors ([MASK] expansion, no zero padding)
  • Vectors: 128-d, L2-normalized
  • Not recorded: exact checkpoint revision, encoding library, document length cap (the longest document has 300 vectors), and whether punctuation was dropped

Statistics

Token vectors (N)597,909,919
Avg vectors per document67.6 (min 4, max 300)
Vectors per query32

Files

FiledtypeShapeContent
documents.npyuint16 (<u2)[597909919, 128]Raw float16 bit patterns stored as uint16. Read with .view(np.float16)
doclens.npyint32[8841823]Vectors per document; sum == N
token_ids_per_token.npyint64[597909919]Input token id of each row of documents.npy
doc_ids.npyint64[8841823]MS MARCO pid (equals the row index)
queries.npyfloat32[6980, 32, 128]Query vectors
queries_ids.npyint64[6980]MS MARCO qid of each query

Document ids are MS MARCO passage ids and equal the row index: row i of doclens.npy is passage i.

documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).

Relevance judgments

The official MS MARCO passage dev/small qrels (qrels.dev.small.tsv): 7,437 relevant query-passage pairs over the 6,980 queries, all with relevance 1 (ir_datasets msmarco-passage/dev/small). Standard metric: MRR@10 (RR@10 in ir_measures).

colbert
late-interaction
msmarco
multivector