tuskanny/scidocs_colbertv2

Dataset

SCIDOCS, ColBERTv2

0

3 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

SCIDOCS, ColBERTv2

Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.

Source

  • BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
  • 25,657 documents, 1,000 queries, 29,928 qrels
  • Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectors
  • Row order follows the BEIR corpus and query files; row i of doc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npy

Encoding

  • Model: colbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acf
  • Library: PyLate 1.6.0, CPU
  • Document length cap: 180 tokens (model default)
  • Query length: 32 tokens (model default)
  • Query expansion: yes (model default)
  • The model defaults come from artifact.metadata in the model repository. We did not override any of them
  • Documents: punctuation tokens are dropped (PyLate skiplist), as are padding tokens
  • Vectors: 128-d, L2-normalized

Statistics

Token vectors (N)3,770,957
Avg vectors per document147.0 (max 180)
Vectors per queryalways 32 (padded with [MASK] expansion tokens, which are real embeddings)
Avg vectors per query32.0

Files

FiledtypeShapeContent
documents.npyfloat16 (<f2)[3770957, 128]All document vectors, concatenated document by document
doclens.npyint32[25657]Vectors per document; sum == N
token_ids.npyuint32[3770957]Input token id of each row of documents.npy
doc_ids.npystring[25657]BEIR doc id of each document
queries.npyfloat32[1000, 32, 128]Query vectors, zero-padded at the end
query_lens.npyint32[1000]True number of vectors per query
queries_ids.npystring[1000]BEIR query id of each query
qrels.test.tsvTREC29928 linesqid \t 0 \t docid \t relevance
groundtruth/gt_top100.tsvTSV100000 linesExhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions)
groundtruth/gt_ids.npyint32[1000, 100]Same, as doc positions
groundtruth/gt_scores.npyfloat32[1000, 100]Same, MaxSim scores

Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.

Exhaustive-search effectiveness

Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.

nDCG@10R@100
0.15810.3578
beir
colbert
late-interaction
multivector
scidocs

tuskanny/scidocs_colbertv2

Dataset

SCIDOCS, ColBERTv2

0

3 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

SCIDOCS, ColBERTv2

Token-level (late-interaction) embeddings of the BEIR SCIDOCS corpus and queries, encoded with ColBERTv2, in the TACHIOM multivector format.

Source

  • BEIR SCIDOCS, test (its only split) split. Corpus, queries and qrels were read from the official BEIR files via ir_datasets (beir/scidocs); PyLate only did the encoding
  • 25,657 documents, 1,000 queries, 29,928 qrels
  • Text given to the encoder for each document: title + " " + text (BEIR title and body joined by a space). The text itself is not included, only its vectors
  • Row order follows the BEIR corpus and query files; row i of doc_ids.npy / queries_ids.npy identifies row i of doclens.npy / queries.npy

Encoding

  • Model: colbert-ir/colbertv2.0 @ c1e84128e85ef755c096a95bdb06b47793b13acf
  • Library: PyLate 1.6.0, CPU
  • Document length cap: 180 tokens (model default)
  • Query length: 32 tokens (model default)
  • Query expansion: yes (model default)
  • The model defaults come from artifact.metadata in the model repository. We did not override any of them
  • Documents: punctuation tokens are dropped (PyLate skiplist), as are padding tokens
  • Vectors: 128-d, L2-normalized

Statistics

Token vectors (N)3,770,957
Avg vectors per document147.0 (max 180)
Vectors per queryalways 32 (padded with [MASK] expansion tokens, which are real embeddings)
Avg vectors per query32.0

Files

FiledtypeShapeContent
documents.npyfloat16 (<f2)[3770957, 128]All document vectors, concatenated document by document
doclens.npyint32[25657]Vectors per document; sum == N
token_ids.npyuint32[3770957]Input token id of each row of documents.npy
doc_ids.npystring[25657]BEIR doc id of each document
queries.npyfloat32[1000, 32, 128]Query vectors, zero-padded at the end
query_lens.npyint32[1000]True number of vectors per query
queries_ids.npystring[1000]BEIR query id of each query
qrels.test.tsvTREC29928 linesqid \t 0 \t docid \t relevance
groundtruth/gt_top100.tsvTSV100000 linesExhaustive top-100: query_idx \t doc_idx \t rank \t score (0-based positions)
groundtruth/gt_ids.npyint32[1000, 100]Same, as doc positions
groundtruth/gt_scores.npyfloat32[1000, 100]Same, MaxSim scores

Zero padding does not change any score: a zero query vector adds 0 to the MaxSim of every document.

Exhaustive-search effectiveness

Exact MaxSim over the full collection (vectorium compute_groundtruth_multivec). These are the reference numbers for approximate search on this data.

nDCG@10R@100
0.15810.3578
beir
colbert
late-interaction
multivector
scidocs