Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
robro612/ModernBERT-XTR at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that repo. Document, query and qrel ids are the source's own ids,
unchanged.
Every document is one variable-length set of 128-d vectors; every query is one variable-length set of 128-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.
| file | dtype | shape | contents |
|---|---|---|---|
documents.npy | float16 (<f2) | [1,047,014, 128] | every document vector, concatenated document by document (255.6 MiB) |
doclens.npy | int32 | [3,633] | vectors per document; cumsum gives offsets |
token_ids.npy | uint32 | [1,047,014] | tokenizer id of each documents.npy row, 1:1 |
doc_ids.npy | <U8 | [3,633] | original document ids |
queries.npy | float32 (<f4) | [323, 32, 128] | query vectors, zero-padded at the end (5.0 MiB) |
query_lens.npy | int32 | [323] | true vectors per query, before padding |
queries_ids.npy | <U10 | [323] | original query ids |
qrels.test.tsv | text | 12,334 rows | TREC qrels, qid \t 0 \t docid \t relevance, no header |
gt_top1000.tsv | text | 323,000 rows | exact MaxSim top-1000, see below |
gt_top100.tsv | text | 32,300 rows | first 100 ranks of gt_top1000.tsv, same format |
All positional indices (the gt_top*.tsv files, and the row order of every .npy file) refer to the
order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates the ground truth.
| documents | 3,633 |
| document vectors | 1,047,014 |
| vectors per document (min / median / mean / max) | 24 / 295 / 288.2 / 491 |
| queries | 323 |
| vectors per query (min / median / mean / max) | 32 / 32 / 32.0 / 32 |
| queries with at least one qrel | 323 |
| qrels rows | 12,334 |
| embedding dimension | 128 |
| model | robro612/ModernBERT-XTR |
| model revision | 8c06fce0b8d2387be98183582eb9600af7bd1b8a |
| library | sentence-transformers 6.1.0 MultiVectorEncoder (transformers 5.17.0, torch 2.13.0+cu126) |
| document compute dtype | float16 (model weights loaded at this dtype for the document pass) |
| document storage dtype | fp16 |
| query compute dtype | float32 (model weights loaded at this dtype for the query pass) |
| query storage dtype | fp32 |
| normalization | L2, by the model's own Normalize module, before the storage cast |
| document truncation | 512 tokens (the checkpoint's document_length), applied before the skiplist. 204 of 3,633 documents (5.6%) were longer and were cut to it; longest here 491 vectors |
| query truncation | query_length unset; fixed query expansion: every query is padded to 32 tokens with the tokenizer's mask token, which are not attended to, and those expansion vectors are kept |
| document skiplist | 32 words removed: ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\', ']', '^', '_', '`', '{', ' |
| document input | title + "\n\n" + text when the corpus has a title, else text; stripped |
| query input | query text, stripped of surrounding whitespace, formatted by the model's own query prompt/template |
| query vectors | every vector the model emits for the query is kept, including any query-expansion tokens its template adds; query_lens counts them all |
| document padding | none: documents.npy holds real vectors only, sum(doclens) == n_tokens |
| query padding | rows at or beyond query_lens[i] in queries.npy[i] are exactly zero |
| token_ids | tokenizer id of each kept document token (after the skiplist above), aligned 1:1 with documents.npy |
gt_top1000.tsv and gt_top100.tsvExact brute-force MaxSim top-1000 per query over the full corpus, from the vectors in this repo. gt_top100.tsv holds the first 100 ranks per query of the same lists (the original layout of these exports).
No header; tab-separated qidx docidx rank score:
qidx: 0-based row into queries_ids.npy / queries.npydocidx: 0-based position into doc_ids.npy / doclens.npyrank: 1-based, descending scorescore: sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors
are included in the sum. Printed to 6 decimals.Sanity check of the vectors, not a leaderboard number: gt_top1000.tsv (exact MaxSim over the full
corpus) scored against qrels.test.tsv with ir_measures.
| nDCG@10 | MRR@10 | Success@5 | Recall@100 | Recall@1000 | MAP@1000 |
|---|---|---|---|---|---|
| 0.3430 | 0.5536 | 0.6594 | 0.2775 | 0.5897 | 0.1739 |
import numpy as np
documents = np.load("documents.npy", mmap_mode="r") # [n_tokens, 128] float16
doclens = np.load("doclens.npy") # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy") # [n_docs] str
def document(i):
return documents[offsets[i]:offsets[i + 1]] # [doclens[i], 128]
queries = np.load("queries.npy") # [n_queries, 32, 128] float32
query_lens = np.load("query_lens.npy") # [n_queries] int32
query_ids = np.load("queries_ids.npy") # [n_queries] str
def query(j):
return queries[j, :query_lens[j]] # [query_lens[j], 128]
def maxsim(q, d):
return (q @ d.astype(np.float32).T).max(axis=1).sum()
Checks run by the exporter on the files exactly as written here:
| exported | 2026-09-25 |
| hardware | Tesla V100S-PCIE-32GB |
| revised | 2026-09-29: ground truth extended to top-1000 (gt_top1000.tsv, exact MaxSim over this repo's vectors on Tesla V100S-PCIE-32GB); gt_top100.tsv rewritten as its first 100 ranks: 391 rows differ from the previous file, all of them documents with identical scores listed in a different order (1 tied pairs at rank 100 swapped in or out) |
Multi-vector (late-interaction) embeddings of BEIR nfcorpus (beir/nfcorpus/test), encoded with
robro612/ModernBERT-XTR at revision 8c06fce0b8d2387be98183582eb9600af7bd1b8a.
Source data: ir_datasets beir/nfcorpus/test (ir_datasets 0.6.3), which downloads nfcorpus.zip (md5 a89dba18a62ef92f7d323ec890a0d38d). BEIR also publishes this corpus on the Hub as BeIR/nfcorpus, whose card gives this dataset's license; the data here was loaded through ir_datasets, not from that repo. Document, query and qrel ids are the source's own ids,
unchanged.
Every document is one variable-length set of 128-d vectors; every query is one variable-length set of 128-d vectors. Documents and queries are stored at different precisions (fp16 and fp32 respectively), see Encoding.
| file | dtype | shape | contents |
|---|---|---|---|
documents.npy | float16 (<f2) | [1,047,014, 128] | every document vector, concatenated document by document (255.6 MiB) |
doclens.npy | int32 | [3,633] | vectors per document; cumsum gives offsets |
token_ids.npy | uint32 | [1,047,014] | tokenizer id of each documents.npy row, 1:1 |
doc_ids.npy | <U8 | [3,633] | original document ids |
queries.npy | float32 (<f4) | [323, 32, 128] | query vectors, zero-padded at the end (5.0 MiB) |
query_lens.npy | int32 | [323] | true vectors per query, before padding |
queries_ids.npy | <U10 | [323] | original query ids |
qrels.test.tsv | text | 12,334 rows | TREC qrels, qid \t 0 \t docid \t relevance, no header |
gt_top1000.tsv | text | 323,000 rows | exact MaxSim top-1000, see below |
gt_top100.tsv | text | 32,300 rows | first 100 ranks of gt_top1000.tsv, same format |
All positional indices (the gt_top*.tsv files, and the row order of every .npy file) refer to the
order of doc_ids.npy and queries_ids.npy. Reordering either file invalidates the ground truth.
| documents | 3,633 |
| document vectors | 1,047,014 |
| vectors per document (min / median / mean / max) | 24 / 295 / 288.2 / 491 |
| queries | 323 |
| vectors per query (min / median / mean / max) | 32 / 32 / 32.0 / 32 |
| queries with at least one qrel | 323 |
| qrels rows | 12,334 |
| embedding dimension | 128 |
| model | robro612/ModernBERT-XTR |
| model revision | 8c06fce0b8d2387be98183582eb9600af7bd1b8a |
| library | sentence-transformers 6.1.0 MultiVectorEncoder (transformers 5.17.0, torch 2.13.0+cu126) |
| document compute dtype | float16 (model weights loaded at this dtype for the document pass) |
| document storage dtype | fp16 |
| query compute dtype | float32 (model weights loaded at this dtype for the query pass) |
| query storage dtype | fp32 |
| normalization | L2, by the model's own Normalize module, before the storage cast |
| document truncation | 512 tokens (the checkpoint's document_length), applied before the skiplist. 204 of 3,633 documents (5.6%) were longer and were cut to it; longest here 491 vectors |
| query truncation | query_length unset; fixed query expansion: every query is padded to 32 tokens with the tokenizer's mask token, which are not attended to, and those expansion vectors are kept |
| document skiplist | 32 words removed: ['!', '"', '#', '$', '%', '&', "'", '(', ')', '*', '+', ',', '-', '.', '/', ':', ';', '<', '=', '>', '?', '@', '[', '\', ']', '^', '_', '`', '{', ' |
| document input | title + "\n\n" + text when the corpus has a title, else text; stripped |
| query input | query text, stripped of surrounding whitespace, formatted by the model's own query prompt/template |
| query vectors | every vector the model emits for the query is kept, including any query-expansion tokens its template adds; query_lens counts them all |
| document padding | none: documents.npy holds real vectors only, sum(doclens) == n_tokens |
| query padding | rows at or beyond query_lens[i] in queries.npy[i] are exactly zero |
| token_ids | tokenizer id of each kept document token (after the skiplist above), aligned 1:1 with documents.npy |
gt_top1000.tsv and gt_top100.tsvExact brute-force MaxSim top-1000 per query over the full corpus, from the vectors in this repo. gt_top100.tsv holds the first 100 ranks per query of the same lists (the original layout of these exports).
No header; tab-separated qidx docidx rank score:
qidx: 0-based row into queries_ids.npy / queries.npydocidx: 0-based position into doc_ids.npy / doclens.npyrank: 1-based, descending scorescore: sum over the query's query_lens[qidx] vectors of max over the document's vectors of the dot product, computed in fp32 with the fp16 document vectors upcast to fp32. Expansion vectors
are included in the sum. Printed to 6 decimals.Sanity check of the vectors, not a leaderboard number: gt_top1000.tsv (exact MaxSim over the full
corpus) scored against qrels.test.tsv with ir_measures.
| nDCG@10 | MRR@10 | Success@5 | Recall@100 | Recall@1000 | MAP@1000 |
|---|---|---|---|---|---|
| 0.3430 | 0.5536 | 0.6594 | 0.2775 | 0.5897 | 0.1739 |
import numpy as np
documents = np.load("documents.npy", mmap_mode="r") # [n_tokens, 128] float16
doclens = np.load("doclens.npy") # [n_docs] int32
offsets = np.concatenate([[0], np.cumsum(doclens)])
doc_ids = np.load("doc_ids.npy") # [n_docs] str
def document(i):
return documents[offsets[i]:offsets[i + 1]] # [doclens[i], 128]
queries = np.load("queries.npy") # [n_queries, 32, 128] float32
query_lens = np.load("query_lens.npy") # [n_queries] int32
query_ids = np.load("queries_ids.npy") # [n_queries] str
def query(j):
return queries[j, :query_lens[j]] # [query_lens[j], 128]
def maxsim(q, d):
return (q @ d.astype(np.float32).T).max(axis=1).sum()
Checks run by the exporter on the files exactly as written here:
| exported | 2026-09-25 |
| hardware | Tesla V100S-PCIE-32GB |
| revised | 2026-09-29: ground truth extended to top-1000 (gt_top1000.tsv, exact MaxSim over this repo's vectors on Tesla V100S-PCIE-32GB); gt_top100.tsv rewritten as its first 100 ranks: 391 rows differ from the previous file, all of them documents with identical scores listed in a different order (1 tied pairs at rank 100 swapped in or out) |