robro612/scidocs_answerai_colbert_small

Dataset

scidocs_answerai_colbert_small

0

4 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

scidocs_answerai_colbert_small

Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.

Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that repo. Document, query and qrel ids are the source's own ids, unchanged.

Every document is one variable-length set of 96-d vectors; every query is one variable-length set of 96-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.

Files

filedtypeshapecontents
documents.npyfloat16 (<f2)[4,820,738, 96]every document vector, concatenated document by document (882.7 MiB)
doclens.npyint32[25,657]vectors per document; cumsum gives offsets
token_ids.npyuint32[4,820,738]tokenizer id of each documents.npy row, 1:1
doc_ids.npy<U40[25,657]original document ids
queries.npyfloat32 (<f4)[1,000, 32, 96]query vectors, zero-padded at the end (11.7 MiB)
query_lens.npyint32[1,000]true vectors per query, before padding
queries_ids.npy<U40[1,000]original query ids
qrels.test.tsvtext29,928 rowsTREC qrels, qid \t 0 \t docid \t relevance, no header
gt_top1000.tsvtext1,000,000 rowsexact MaxSim top-1000, see below
gt_top100.tsvtext100,000 rowsfirst 100 ranks of gt_top1000.tsv, same format

All positional indices (the gt_top*.tsv files, and the row order of every .npy file) refer to the order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates the ground truth.

Statistics

documents25,657
document vectors4,820,738
vectors per document (min / median / mean / max)5 / 189 / 187.9 / 300
queries1,000
vectors per query (min / median / mean / max)32 / 32 / 32.0 / 32
queries with at least one qrel1,000
qrels rows29,928
embedding dimension96

Encoding

modellightonai/answerai-colbert-small-v1
model revisione507cd12947a2b4b52201d150967df3c19a90590
librarysentence-transformers 6.1.0 MultiVectorEncoder (transformers 5.17.0, torch 2.13.0+cu126)
document compute dtypefloat16 (model weights loaded at this dtype for the document pass)
document storage dtypefp16
query compute dtypefloat32 (model weights loaded at this dtype for the query pass)
query storage dtypefp32
normalizationL2, by the model's own Normalize module, before the storage cast
document truncation300 tokens (the checkpoint's document_length), applied before the skiplist. 4,619 of 25,657 documents (18%) were longer and were cut to it; longest here 300 vectors
query truncationquery_length unset; fixed query expansion: every query is padded to 32 tokens with the tokenizer's mask token, which are not attended to, and those expansion vectors are kept
document skiplist32 words removed: ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\', ']', '^', '_', '`', '{', '
document inputtitle + "\n\n" + text when the corpus has a title, else text; stripped
query inputquery text, stripped of surrounding whitespace, formatted by the model's own query prompt/template
query vectorsevery vector the model emits for the query is kept, including any query-expansion tokens its template adds; query_lens counts them all
document paddingnone: documents.npy holds real vectors only, sum(doclens) == n_tokens
query paddingrows at or beyond query_lens[i] in queries.npy[i] are exactly zero
token_idstokenizer id of each kept document token (after the skiplist above), aligned 1:1 with documents.npy

Ground truth: gt_top1000.tsv and gt_top100.tsv

Exact brute-force MaxSim top-1000 per query over the full corpus, from the vectors in this repo. gt_top100.tsv holds the first 100 ranks per query of the same lists (the original layout of these exports).

No header; tab-separated qidx docidx rank score:

  • qidx: 0-based row into queries_ids.npy / queries.npy
  • docidx: 0-based position into doc_ids.npy / doclens.npy
  • rank: 1-based, descending score
  • score: sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors are included in the sum. Printed to 6 decimals.

Retrieval quality

Sanity check of the vectors, not a leaderboard number: gt_top1000.tsv (exact MaxSim over the full corpus) scored against qrels.test.tsv with ir_measures.

nDCG@10MRR@10Success@5Recall@100Recall@1000MAP@1000
0.18480.32020.45300.42320.67950.1314

Loading

import numpy as np

documents = np.load("documents.npy", mmap_mode="r")      # [n_tokens, 96] float16
doclens = np.load("doclens.npy")                         # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy")                         # [n_docs] str

def document(i):
    return documents[offsets[i]:offsets[i + 1]]           # [doclens[i], 96]

queries = np.load("queries.npy")                         # [n_queries, 32, 96] float32
query_lens = np.load("query_lens.npy")                   # [n_queries] int32
query_ids = np.load("queries_ids.npy")                   # [n_queries] str

def query(j):
    return queries[j, :query_lens[j]]                     # [query_lens[j], 96]

def maxsim(q, d):
    return (q @ d.astype(np.float32).T).max(axis=1).sum()

Validation

Checks run by the exporter on the files exactly as written here:

  • βœ… file set β€” missing=[] extra=[]
  • βœ… documents.npy dtype/shape β€” <f2 (4820738, 96)
  • βœ… doclens.npy dtype/shape β€” <i4 (25657,)
  • βœ… doc_ids.npy is a string array β€” <U40 (25657,)
  • βœ… queries.npy dtype/shape β€” <f4 (1000, 32, 96)
  • βœ… query_lens.npy dtype/shape β€” <i4 (1000,)
  • βœ… queries_ids.npy is a string array β€” <U40 (1000,)
  • βœ… sum(doclens) == n_tokens β€” 4820738 vs 4820738
  • βœ… no empty documents β€” min doclen 5
  • βœ… len(doc_ids) == len(doclens) == corpus size β€” 25657, 25657, 25657
  • βœ… doc_ids unique
  • βœ… query arrays aligned β€” 1000, 1000, 1000
  • βœ… doc and query dim agree β€” 96 / 96
  • βœ… token_ids.npy dtype/shape β€” <u4 (4820738,)
  • βœ… document vectors unit-norm (100k sample) β€” norm range [0.9995, 1.0005]
  • βœ… query vectors unit-norm β€” norm range [1.000000, 1.000000]
  • βœ… all vectors finite
  • βœ… gt_top100.tsv has k rows per query β€” 100000 rows, k=100
  • βœ… gt_top100.tsv rows grouped by qidx with ranks 1..k and descending scores
  • βœ… gt_top100.tsv indices in range
  • βœ… gt_top1000.tsv has k rows per query β€” 1000000 rows, k=1000
  • βœ… gt_top1000.tsv rows grouped by qidx with ranks 1..k and descending scores
  • βœ… gt_top1000.tsv indices in range
  • βœ… gt_top100.tsv is the first 100 ranks of gt_top1000.tsv

Provenance

exported2026-09-25
hardwareTesla V100S-PCIE-32GB
revised2026-09-29: ground truth extended to top-1000 (gt_top1000.tsv, exact MaxSim over this repo's vectors on Tesla V100S-PCIE-32GB); gt_top100.tsv rewritten as its first 100 ranks: 70 rows differ from the previous file, all of them documents with identical scores listed in a different order
colbert
embeddings
late-interaction
multi-vector
retrieval
text

robro612/scidocs_answerai_colbert_small

Dataset

scidocs_answerai_colbert_small

0

4 commits

1 linked in READMEs

updated Sep 29, 2026

See the code

README

scidocs_answerai_colbert_small

Multi-vector (late-interaction) embeddings of BEIR scidocs (beir/scidocs), encoded with lightonai/answerai-colbert-small-v1 at revision e507cd12947a2b4b52201d150967df3c19a90590.

Source data: ir_datasets beir/scidocs (ir_datasets 0.6.3), which downloads scidocs.zip (md5 38121350fc3a4d2f48850f6aff52e4a9). BEIR also publishes this corpus on the Hub as BeIR/scidocs, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that repo. Document, query and qrel ids are the source's own ids, unchanged.

Every document is one variable-length set of 96-d vectors; every query is one variable-length set of 96-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.

Files

filedtypeshapecontents
documents.npyfloat16 (<f2)[4,820,738, 96]every document vector, concatenated document by document (882.7 MiB)
doclens.npyint32[25,657]vectors per document; cumsum gives offsets
token_ids.npyuint32[4,820,738]tokenizer id of each documents.npy row, 1:1
doc_ids.npy<U40[25,657]original document ids
queries.npyfloat32 (<f4)[1,000, 32, 96]query vectors, zero-padded at the end (11.7 MiB)
query_lens.npyint32[1,000]true vectors per query, before padding
queries_ids.npy<U40[1,000]original query ids
qrels.test.tsvtext29,928 rowsTREC qrels, qid \t 0 \t docid \t relevance, no header
gt_top1000.tsvtext1,000,000 rowsexact MaxSim top-1000, see below
gt_top100.tsvtext100,000 rowsfirst 100 ranks of gt_top1000.tsv, same format

All positional indices (the gt_top*.tsv files, and the row order of every .npy file) refer to the order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates the ground truth.

Statistics

documents25,657
document vectors4,820,738
vectors per document (min / median / mean / max)5 / 189 / 187.9 / 300
queries1,000
vectors per query (min / median / mean / max)32 / 32 / 32.0 / 32
queries with at least one qrel1,000
qrels rows29,928
embedding dimension96

Encoding

modellightonai/answerai-colbert-small-v1
model revisione507cd12947a2b4b52201d150967df3c19a90590
librarysentence-transformers 6.1.0 MultiVectorEncoder (transformers 5.17.0, torch 2.13.0+cu126)
document compute dtypefloat16 (model weights loaded at this dtype for the document pass)
document storage dtypefp16
query compute dtypefloat32 (model weights loaded at this dtype for the query pass)
query storage dtypefp32
normalizationL2, by the model's own Normalize module, before the storage cast
document truncation300 tokens (the checkpoint's document_length), applied before the skiplist. 4,619 of 25,657 documents (18%) were longer and were cut to it; longest here 300 vectors
query truncationquery_length unset; fixed query expansion: every query is padded to 32 tokens with the tokenizer's mask token, which are not attended to, and those expansion vectors are kept
document skiplist32 words removed: ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\', ']', '^', '_', '`', '{', '
document inputtitle + "\n\n" + text when the corpus has a title, else text; stripped
query inputquery text, stripped of surrounding whitespace, formatted by the model's own query prompt/template
query vectorsevery vector the model emits for the query is kept, including any query-expansion tokens its template adds; query_lens counts them all
document paddingnone: documents.npy holds real vectors only, sum(doclens) == n_tokens
query paddingrows at or beyond query_lens[i] in queries.npy[i] are exactly zero
token_idstokenizer id of each kept document token (after the skiplist above), aligned 1:1 with documents.npy

Ground truth: gt_top1000.tsv and gt_top100.tsv

Exact brute-force MaxSim top-1000 per query over the full corpus, from the vectors in this repo. gt_top100.tsv holds the first 100 ranks per query of the same lists (the original layout of these exports).

No header; tab-separated qidx docidx rank score:

  • qidx: 0-based row into queries_ids.npy / queries.npy
  • docidx: 0-based position into doc_ids.npy / doclens.npy
  • rank: 1-based, descending score
  • score: sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors are included in the sum. Printed to 6 decimals.

Retrieval quality

Sanity check of the vectors, not a leaderboard number: gt_top1000.tsv (exact MaxSim over the full corpus) scored against qrels.test.tsv with ir_measures.

nDCG@10MRR@10Success@5Recall@100Recall@1000MAP@1000
0.18480.32020.45300.42320.67950.1314

Loading

import numpy as np

documents = np.load("documents.npy", mmap_mode="r")      # [n_tokens, 96] float16
doclens = np.load("doclens.npy")                         # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy")                         # [n_docs] str

def document(i):
    return documents[offsets[i]:offsets[i + 1]]           # [doclens[i], 96]

queries = np.load("queries.npy")                         # [n_queries, 32, 96] float32
query_lens = np.load("query_lens.npy")                   # [n_queries] int32
query_ids = np.load("queries_ids.npy")                   # [n_queries] str

def query(j):
    return queries[j, :query_lens[j]]                     # [query_lens[j], 96]

def maxsim(q, d):
    return (q @ d.astype(np.float32).T).max(axis=1).sum()

Validation

Checks run by the exporter on the files exactly as written here:

  • βœ… file set β€” missing=[] extra=[]
  • βœ… documents.npy dtype/shape β€” <f2 (4820738, 96)
  • βœ… doclens.npy dtype/shape β€” <i4 (25657,)
  • βœ… doc_ids.npy is a string array β€” <U40 (25657,)
  • βœ… queries.npy dtype/shape β€” <f4 (1000, 32, 96)
  • βœ… query_lens.npy dtype/shape β€” <i4 (1000,)
  • βœ… queries_ids.npy is a string array β€” <U40 (1000,)
  • βœ… sum(doclens) == n_tokens β€” 4820738 vs 4820738
  • βœ… no empty documents β€” min doclen 5
  • βœ… len(doc_ids) == len(doclens) == corpus size β€” 25657, 25657, 25657
  • βœ… doc_ids unique
  • βœ… query arrays aligned β€” 1000, 1000, 1000
  • βœ… doc and query dim agree β€” 96 / 96
  • βœ… token_ids.npy dtype/shape β€” <u4 (4820738,)
  • βœ… document vectors unit-norm (100k sample) β€” norm range [0.9995, 1.0005]
  • βœ… query vectors unit-norm β€” norm range [1.000000, 1.000000]
  • βœ… all vectors finite
  • βœ… gt_top100.tsv has k rows per query β€” 100000 rows, k=100
  • βœ… gt_top100.tsv rows grouped by qidx with ranks 1..k and descending scores
  • βœ… gt_top100.tsv indices in range
  • βœ… gt_top1000.tsv has k rows per query β€” 1000000 rows, k=1000
  • βœ… gt_top1000.tsv rows grouped by qidx with ranks 1..k and descending scores
  • βœ… gt_top1000.tsv indices in range
  • βœ… gt_top100.tsv is the first 100 ranks of gt_top1000.tsv

Provenance

exported2026-09-25
hardwareTesla V100S-PCIE-32GB
revised2026-09-29: ground truth extended to top-1000 (gt_top1000.tsv, exact MaxSim over this repo's vectors on Tesla V100S-PCIE-32GB); gt_top100.tsv rewritten as its first 100 ranks: 70 rows differ from the previous file, all of them documents with identical scores listed in a different order
colbert
embeddings
late-interaction
multi-vector
retrieval
text