tuskanny/lotte_pooled_colbertv2

Dataset

LoTTE pooled (dev), ColBERTv2

0

6 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

LoTTE pooled (dev), ColBERTv2

Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.

Source

  • Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages
  • Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels)
  • Document order: passage id order (row i is pid i)

Encoding

  • Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
  • Documents start with [CLS] and the ColBERT [D] marker (token id 2)
  • Document length cap: the longest document has 180 vectors, consistent with the ColBERTv2 default doc_maxlen=180
  • Queries: always 32 vectors ([MASK] expansion, no zero padding)
  • Vectors: 128-d, L2-normalized
  • Not recorded: exact checkpoint revision, encoding library, and whether punctuation was dropped

Statistics

Token vectors (N)266,205,513
Avg vectors per document109.6 (min 3, max 180)
Vectors per query32

Files

FiledtypeShapeContent
documents.npyuint16 (<u2)[266205513, 128]Raw float16 bit patterns stored as uint16. Read with .view(np.float16)
doclens.npyint32[2428854]Vectors per document; sum == N
token_ids.npyint64[266205513]Input token id of each row of documents.npy
queries.npyfloat32[2931, 32, 128]Query vectors
queries_ids.npyint64[2931]LoTTE qid of each query (0 to 2930)

Document ids are LoTTE passage ids and equal the row index: row i of doclens.npy is passage i.

documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).

Relevance judgments

The official LoTTE pooled dev search qrels: 8,573 relevant query-passage pairs over the 2,931 queries, all with relevance 1 (ir_datasets lotte/pooled/dev/search). Standard metric: Success@5.

colbert
late-interaction
lotte
multivector

tuskanny/lotte_pooled_colbertv2

Dataset

LoTTE pooled (dev), ColBERTv2

0

6 commits

1 linked in READMEs

updated Sep 23, 2026

See the code

README

LoTTE pooled (dev), ColBERTv2

Token-level (late-interaction) ColBERTv2 embeddings of the LoTTE pooled dev collection and its search queries.

Source

  • Collection: LoTTE pooled, dev split (ir_datasets lotte/pooled/dev), 2,428,854 passages
  • Queries: dev search queries, 2,931 (lotte/pooled/dev/search, 8,573 qrels)
  • Document order: passage id order (row i is pid i)

Encoding

  • Model: ColBERTv2 (colbert-ir/colbertv2.0, BERT-base-uncased tokenizer)
  • Documents start with [CLS] and the ColBERT [D] marker (token id 2)
  • Document length cap: the longest document has 180 vectors, consistent with the ColBERTv2 default doc_maxlen=180
  • Queries: always 32 vectors ([MASK] expansion, no zero padding)
  • Vectors: 128-d, L2-normalized
  • Not recorded: exact checkpoint revision, encoding library, and whether punctuation was dropped

Statistics

Token vectors (N)266,205,513
Avg vectors per document109.6 (min 3, max 180)
Vectors per query32

Files

FiledtypeShapeContent
documents.npyuint16 (<u2)[266205513, 128]Raw float16 bit patterns stored as uint16. Read with .view(np.float16)
doclens.npyint32[2428854]Vectors per document; sum == N
token_ids.npyint64[266205513]Input token id of each row of documents.npy
queries.npyfloat32[2931, 32, 128]Query vectors
queries_ids.npyint64[2931]LoTTE qid of each query (0 to 2930)

Document ids are LoTTE passage ids and equal the row index: row i of doclens.npy is passage i.

documents.npy stores float16 values as their raw 16-bit patterns, with dtype uint16. In numpy, read it with np.load("documents.npy", mmap_mode="r").view(np.float16).

Relevance judgments

The official LoTTE pooled dev search qrels: 8,573 relevant query-passage pairs over the 2,931 queries, all with relevance 1 (ir_datasets lotte/pooled/dev/search). Standard metric: Success@5.

colbert
late-interaction
lotte
multivector