An extensive and commented list of resources on Late-Interaction Multivector Retrieval.
TeX
79
53 commits
updated Sep 28, 2026
An extensive and commented list of resources on late-interaction multivector retrieval.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab, Matei Zaharia
SIGIR, 2020
π paper | π οΈ code
COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List
Luyu Gao, Zhuyun Dai, Jamie Callan
NAACL, 2021
π paper
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia
NAACL, 2022
π paper | π οΈ code
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
Rohan Jha, Bo Wang, Michael GΓΌnther, Georgios Mastrapas, Saba Sturua, Isabelle Mohr, Andreas Koukounas, Mohammad Kalim Akram, Nan Wang, Han Xiao
MRL Workshop, 2024
π paper
PyLate: Flexible Training and Retrieval for Late Interaction Models
Antoine Chaffin, RaphaΓ«l Sourty
CIKM, 2025
π paper | π οΈ code
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
Antoine Chaffin, Luca Arnaboldi, AmΓ©lie Chatelain, Florent Krzakala
arXiv, 2026
π paper
Your Embedding Model is SMARTer Than You Think
Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee
arXiv, 2026
π paper | π οΈ code
Party is over: regularizing ColBERT models to fix efficient ANN methods
LightOn AI
Blog, 2026
π blog
NumColBERT: Non-Intrusive Numeracy Injection for Late-Interaction Retrieval Models
Haruki Fujimaki, Makoto P. Kato
arXiv, 2026
π paper
mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models
LightOn AI
Blog, 2026
π blog
GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval
Aarush Sinha
Blog, 2026
π blog
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Tom Aarsen, Antoine Chaffin, RaphaΓ«l Sourty
Blog, 2026
π blog
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Tom Aarsen
Blog, 2026
π blog
Exploring Static Embedding Retrieval
Logan Markewich
Blog, 2026
π blog
KURE-v2: A Korean-English Bilingual Late-Interaction Retrieval Model
Youngjoon Jang
Blog, 2026
π blog
SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning Intensive Retrieval
Abdelrahman Abdallah, Mohammed Ali, Adam Jatowt
EMNLP, 2026
π paper | π οΈ code
Compact Token Representations with Contextual Quantization for Efficient Document Re-ranking
Yingrui Yang, Yifan Qiao, Tao Yang
ACL, 2022
π paper
Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT)
Carlos Lassance, Maroua Maachou, Joohee Park, Stephane Clinchant
SIGIR, 2022
π paper
Introducing Neural Bag of Whole-Words with ColBERTer: Contextualized Late Interactions using Enhanced Reduction
Sebastian Hofstatter, Omar Khattab, Sophia Althammer, Mete Sertkan, Allan Hanbury
CIKM, 2022
π paper | π οΈ code
Joint Optimization of Multi-Vector Representation with Product Quantization
Yufan Fang, Jing Zhan, Yiqun Liu, Jiafeng Mao, Min Zhang, Shaoping Ma
NLPCC, 2022
π paper
Multi-Vector Retrieval as Sparse Alignment
Yujie Qian, Jinhyuk Lee, Sai Meher Karthik Duddu, Zhuyun Dai, Siddhartha Brahma, Iftekhar Naim, Tao Lei, Vincent Y. Zhao
arXiv, 2022
π paper
CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval
Minghan Li, Sean C. Lin, Barlas Oguz, Arnab Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen
ACL, 2023
π paper
SLIM: Sparsified Late Interaction for Multi-Vector Retrieval with Inverted Indexes
Minghan Li, Sheng-Chieh Lin, Xueguang Ma, Jimmy Lin
SIGIR, 2023
π paper
Static Pruning for Multi-Representation Dense Retrieval
Antonio Acquavia, Craig Macdonald, Nicola Tonellotto
DocEng, 2023
π paper | π οΈ code
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, Vincent Y. Zhao
NeurIPS, 2023
π paper
SPLATE: Sparse Late Interaction Retrieval
Thibault Formal, Stephane Clinchant, Herve Dejean, Carlos Lassance
SIGIR, 2024
π paper
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
Benjamin ClaviΓ©, Antoine Chaffin, Griffin Adams
arXiv, 2024
π paper
Muvera: Multi-Vector Retrieval via Fixed Dimensional Encodings
Laxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee, Vahab Mirrokni
NeurIPS, 2024
π paper
Enhancing ColBERT: A Method for Reducing Space Complexity and Accelerating Retrieval Speed
Hai Nguyen T., Huong Le T.
PACLIC, 2024
π paper
Token Pruning Optimization for Efficient Multi-vector Dense Retrieval
Shanxiu He, Mutasem Al-Darabsah, Suraj Nair, Jonathan May, Tarun Agarwal, Tao Yang, Choon Hui Teo
ECIR, 2025
π paper
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning
JoΓ£o Veneroso, Rajesh Jayaram, Jinmeng Rao, Gustavo HernΓ‘ndez Γbrego, Majid Hadian, Daniel Cer
arXiv, 2025
π paper
Towards Lossless Token Pruning in Late-Interaction Retrieval Models
Yuxuan Zong, Benjamin Piwowarski
SIGIR, 2025
π paper
ColPruner: Combining Complementary Pruning Approaches for ColBERT in Web Search
Wondo Rhee, Chan Lim, Taewon Yoon, Gyuhyeon Choi, Jooyoung Lee
ReNeuIR Workshop, 2025
π paper
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
Yubo Ma, Jinsong Li, Yuhang Zang, Xiaobao Wu, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Jiaqi Wang, Yixin Cao, Aixin Sun
ACL Findings, 2025
π paper
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
Yibo Yan, Mingdong Ou, Yi Cao, Xin Zou, Jiahao Huo, Shuliang Liu, James Kwok, Xuming Hu
arXiv, 2026
π paper
Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers
Ailar Mahdizadeh, Aria Salari, Sohail Rajabi, Shahriar Mirabbasi, Panos Nasiopoulos, Alireza Morsali
arXiv, 2026
π paper
AdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval
Jianxin You, Kun Ni
CIKM, 2026
π paper
Multi-Vector Index Compression in Any Modality
Hanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
SIGIR, 2026
π paper | π οΈ code
Learn to Pool: Lightweight Fine-Tuning for Flexible Multi-Vector Compression
Stefan Josef
LIR Workshop, 2026
π paper
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
Yash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski, Joseph Le Roux
SIGIR, 2026
π paper
CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction Retrieval
Rohit Kumar Salla, Manoj Saravanan, Ramya Manasa Amancherla
ICML, 2026
π paper
EigenLI: Spectral Approximations to Late Interaction
Archish S, Sabyasachi Basu, Ankit Garg, Ravishankar Krishnaswamy, Kirankumar Shiragur
arXiv, 2026
π paper
Generative Late-Interaction Embeddings For Visual Document Retrieval
Mohamed Eltahir, Talal Aloushan, Rose Khairoalsendi, Jana Shata, Mohammed Alhassan, Leen Alrehaili, Tanveer Hussain, Naeemullah Khan
arXiv, 2026
π paper
ColPali: Efficient Document Retrieval with Vision Language Models
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Celine Hudelot, Pierre Colombo
ICLR, 2025
π paper
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Arun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates, Kate Sanders, Kenton Murray, Reno Kriz, Celso M. de Melo, Benjamin Van Durme, Rama Chellappa
CVPR, 2025
π paper
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan, Rabiul Awal, Shambhavi Mishra, Akshay Kalkunte Suresh, Srivatsava Daruru, Enamul Hoque, Spandana Gella, Torsten Scholak, Sai Rajeswar
EMNLP, 2025
π paper
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
Zilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen, Xintao Chen, Vicente Ordonez, Vijai Mohan
ICLR, 2026
π paper | π οΈ code
AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval
Xinze Liu, Lei Yang, Dayan Wu, Hengjie Zhu, Zihao Zhang, Hanqi Wu, Tianzhu Hu, Peng Fu, Zheng Lin, Weiping Wang
arXiv, 2026
π paper
Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings
Rajesh Jayaram
arXiv, 2026
π paper
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models
Julian Killingback, Varad Ingale, Hamed Zamani, Cameron Musco
arXiv, 2026
π paper
Retrieval Needs Multivectors: An Exponential Separation
Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu, Ankit Garg, Kirankumar Shiragur
arXiv, 2026
π paper
Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval
Omar Khattab, Christopher Potts, Matei Zaharia
NeurIPS, 2021
π paper
PLAID: An Efficient Engine for Late Interaction Retrieval
Keshav Santhanam, Omar Khattab, Christopher Potts, Matei Zaharia
CIKM, 2022
π paper | π οΈ code
DESSERT: An Efficient Algorithm for Vector Set Search with Vector Set Queries
Joshua Engels, Benjamin Coleman, Vihan Lakshman, Anshumali Shrivastava
NeurIPS, 2023
π paper
Efficient Multi-Vector Dense Retrieval with Bit Vectors
Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
ECIR, 2024
π paper | π οΈ code
Efficient Constant-Space Multi-vector Retrieval
Sean MacAvaney, Antonio Mallia, Nicola Tonellotto
ECIR, 2025
π paper
IGP: Efficient Multi-Vector Retrieval via Proximity Graph Index
Ziyang Bian, Man Lung Yiu, Buzhou Tang
SIGIR, 2025
π paper | π οΈ code
WARP: An Efficient Engine for Multi-Vector Retrieval
Jan Luca Scheerer, Matei Zaharia, Christopher Potts, Gustavo Alonso, Omar Khattab
SIGIR, 2025
π paper | π οΈ code
Multivector Reranking in the Era of Strong First-Stage Retrievers
Silvio Martinico, Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
ECIR, 2026
π paper | π οΈ code | π οΈ code
SMVE: Sparse Multi-Vector Retrieval
Martin Spisak, Marek Galovic
LIR Workshop, 2026
π blog
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
Lixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng, Stefanie Jegelka, Chenyu You
ICML, 2026
π paper
LEMUR: Learned Multi-Vector Retrieval
Elias JÀÀsaari, Ville Hyvânen, Teemu Roos
ICML, 2026
π paper | π οΈ code
Efficient Multivector Retrieval with Token-Aware Clustering and Hierarchical Indexing
Silvio Martinico, Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
SIGIR, 2026
π paper | π οΈ code
ColBERTSaR: Sparsified ColBERT Index via Product Quantization
Eugene Yang, Andrew Yates, Dawn Lawrie, James Mayfield, Saron Samuel, Rohan Jha
SIGIR, 2026
π paper | π οΈ code
PLAID-PRF: Pseudo-Relevance Feedback with Centroid-like Tokens in PLAID
Xiao Wang, Sean MacAvaney, Craig Macdonald
SIGIR, 2026
π paper | π οΈ code
Chimera: Efficient Multi-Vector Retrieval via GPU-CPU Co-Processing
Yanqi Chen, Juelin Liu, Alexandra Meliou, Xiao Yan
arXiv, 2026
π paper
Col-Bandit: Query-Time Top-K Estimation for Late-Interaction Retrieval
Roi Pony, Adi Raz Goldfarb, Oshri Naparstek, Idan Friedman, Udi Barzelay, Eli Schwartz
arXiv, 2026
π paper
FLASH-MAXSIM: IO-Aware Fused Kernels for Late-Interaction Scoring
Roi Pony, Adi Raz Goldfarb, Idan Friedman, Daniel Ezer, Udi Barzelay
arXiv, 2026
π paper | π οΈ code
TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product Quantization
Ashutosh Sharma
arXiv, 2026
π paper | π οΈ code
A Reproducibility Study of PLAID
Sean MacAvaney, Nicola Tonellotto
SIGIR, 2024
π paper
A Replicability Study of XTR
Rohan Jha, Reno Kriz, Benjamin Van Durme
arXiv, 2026
π paper
Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions
Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
SIGIR, 2026
π paper
Comparing Token Pruning Approaches for Multi-Vector Retrieval
Ferdinand Schlatt, Hanno Barschel, Matthias Hagen
SIGIR, 2026
π paper
A Brief Comparison of Training-Free Multi-Vector Sequence Compression Methods
Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
LIR Workshop, 2026
π paper
A Survey of Late-Interaction Neural Retrieval: Paradigms, Systems, and Research Frontiers
Xiao Wang, Chuting Yu, Minghan Li, Binci Yang, Hang Li, Ben He
SSRN, 2026
π paper
ColBERT
Reference implementation for ColBERT and ColBERTv2, and includes PLAID support for efficient late-interaction retrieval.
RAGatouille
Python toolkit to train and serve ColBERT-based late-interaction retrievers.
PyLate
Python library for training, fine-tuning, inference, and retrieval with ColBERT-style late-interaction models on single and multi-GPU setups.
Sentence Transformers
Framework for dense, sparse, reranker, and (v6.0+) ColBERT-style multi-vector late-interaction models via the MultiVectorEncoder, with a unified training and inference API.
PyLate-rs
High-performance Rust inference engine for PyLate models, with Python bindings and optimized integration with FastPlaid for retrieval pipelines.
FastPlaid
GPU-optimized engine for ColBERT/PLAID-style late-interaction retrieval.
kANNolo
ANN library for dense, sparse, and multivector retrieval.
Vectorium ![]()
Rust library for compact storage/access of dense, sparse, and multivector embeddings.
Firn ![]()
Rust search engine for single-vector and late-interaction multivector namespaces, backed by LanceDB on object storage with RAM/NVMe result caching.
NextPlaid
CPU-oriented local-first multivector retrieval engine with memory-mapped storage.
EMVB
Reference implementation for Efficient Multi-Vector Dense Retrieval with Bit Vectors.
IGP
Official C++ implementation for IGP: proximity-graph indexing for multi-vector retrieval (with Python scripts for experiments).
WARP
Official implementation for WARP, an efficient multi-vector retrieval engine.
ColGrep
High-performance code search CLI tool powered by LateOn-Code and NextPlaid, enabling semantic + hybrid (regex + semantic) code retrieval locally with incremental indexing.
TACHIOM
Fast and scalable multivector retrieval system with Token-Aware Clustering (TAC) and hierarchical Product Quantization for efficient late-interaction search.
TopK
Managed retrieval engine with support for late-interaction search over billions of documents, online index updates, filtering, and more.
Flash-MaxSim
IO-aware Triton kernel for MaxSim scoring in ColBERT/ColPali pipelines: tile-by-tile on-chip computation with zero intermediate memory and INT8 quantization support.
maxsim
Ahead-of-time compiled MaxSim kernel with CUDA and Metal backends (NVIDIA + Apple Silicon), distributed as a HuggingFace kernels package.
late-interaction-kernels
Fused Triton kernels for MaxSim scoring with CUDA, Metal, and CPU backends, native PyLate/colpali-engine integration, and PLAID-style compressed-index support.
maxsim-cpu
CPU-only MaxSim kernel written in Rust (libxsmm on x86, Apple Accelerate on ARM) with Python bindings.
TileMaxSim
IO-aware Triton kernel for MaxSim scoring with dimension tiling for embeddings wider than 128 dims and fused product quantization, achieving 80%+ peak HBM bandwidth.
colbert-ir/colbertv2.0
Official ColBERTv2 checkpoint (MS MARCO-trained) from the ColBERT authors, widely used as the canonical baseline model.
lightonai/LateOn
State-of-the-art ColBERT model (149M, ModernBERT-based) achieving 57.22 NDCG@10 on BEIR with fully open training data and strong generalization under decontamination.
lightonai/mLateOn
Multilingual ColBERT model (307M, mmBERT-based) covering 9 languages with SOTA multilingual, long-document, and code retrieval; strong zero-shot generalization to unseen languages (57.56 NDCG@10 on BEIR).
jinaai/jina-colbert-v2
Multilingual late-interaction retriever (0.6B, JinaBERT-based) supporting 89 languages with Matryoshka token embeddings (128/96/64 dims) for flexible efficiency-precision tradeoffs.
chungimungi/GLInt
English ColBERT model (149M, LateOn-based) trained with geometry-matched hard negatives mined via MaxSim geometry, reporting 57.43 NDCG@10 on BEIR.
lightonai/LateOn-regularized
LateOn variant trained with STE-based regularization to fix compatibility with projection-based retrieval methods (MUVERA, SMVE).
lightonai/ColBERT-Zero
Large-scale fully pre-trained ColBERT checkpoint trained on public data and released with the ColBERT-Zero paper.
lightonai/GTE-ModernColBERT-v1
PyLate late-interaction checkpoint based on ModernBERT with 128-dimensional token embeddings and strong long-context retrieval behavior.
topk-io/Iso-ModernColBERT
Isotropically corrected version of GTE-ModernColBERT-v1 built for efficient inference and scalable retrieval.
sebastian-hofstaetter/colberter-128-32-msmarco / sebastian-hofstaetter/uni-colberter-128-1-msmarco
ColBERTer checkpoints trained on MS MARCO (128-dim, with 32 and 1 unique whole-word vectors per document respectively).
lightonai/LateOn-Code
Specialized ColBERT model (149M parameters) fine-tuned for code retrieval, achieving SOTA on MTEB Code benchmark.
lightonai/LateOn-Code-edge
Lightweight code retrieval model (17M parameters) for edge devices, matching larger models while running efficiently on CPU.
lightonai/Reason-ModernColBERT
Reasoning-focused late-interaction checkpoint fine-tuned on reasonir-hq, with strong BRIGHT benchmark performance for reasoning-intensive retrieval.
nlpai-lab/KURE-v2
Korean-English bilingual late-interaction model (154M, skt/A.X-Encoder-base) with 128-dim token vectors and 8,192-token context, reporting 0.8160 average nDCG@10 on MTEB(kor, v2).
nlpai-lab/KURE-v2-unsupervised
Stage-1 KURE-v2 checkpoint trained with weakly-supervised contrastive learning only (20.7M pairs, no relevance labels), reaching 0.7283 average nDCG@10 on MTEB(kor, v2).
DataScience-UIBK/SmallReason-ColBERT-32M
Ultra-small reasoning retriever (32M, mxbai-edge-colbert-v0-32m-based) with a query-side token-importance head, reporting 21.41 mean nDCG@10 on BRIGHT; load via WeightedColBERT.from_base(), as plain PyLate loading drops the head.
vidore/colpali-v1.3
Latest ColPali release (PaliGemma-3B + LoRA) for visual document retrieval, producing ColBERT-style multi-vector embeddings of page images.
vidore/colqwen2-v1.0
ColPali-style visual document retriever on Qwen2-VL-2B-Instruct, accepting dynamic image resolutions without aspect-ratio distortion (up to 768 patches).
vidore/colqwen2.5-v0.2
ColPali-style visual document retriever on Qwen2.5-VL-3B-Instruct, with dynamic image resolutions (up to 768 patches).
vidore/colSmol-256M / vidore/colSmol-500M
Lightweight ColPali-style visual document retrievers built on SmolVLM-256M-Instruct and SmolVLM-500M-Instruct.
NFCorpus3,633test]: 323nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 154.5 | 32.0 | 0.3299 |
modern_colbert | link | 128 | 237.4 | 8.6 | 0.3792 |
answerai_colbert_small | link | 96 | 235.5 | 32.0 | 0.3683 |
modernbert_xtr | link | 128 | 288.2 | 32.0 | 0.3430 |
lateon | link | 128 | 237.4 | 8.6 | 0.3809 |
lateon_hpool_regularized | link | 128 | 238.5 | 8.6 | 0.3803 |
mlateon | link | 128 | 343.7 | 7.4 | 0.3786 |
neomme_260m_li | link | 128 | 337.3 | 17.1 | 0.3081 |
SciFact5,183test]: 300nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
answerai_colbert_small | link | 96 | 235.7 | 32.0 | 0.7431 |
lateon | link | 128 | 231.2 | 21.0 | 0.7627 |
lateon_hpool_regularized | link | 128 | 231.2 | 21.0 | 0.7608 |
mlateon | link | 128 | 314.4 | 21.0 | 0.7605 |
neomme_260m_li | link | 128 | 321.2 | 31.0 | 0.7161 |
ArguAna8,674test]: 1,406nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
neomme_260m_li | link | 128 | 194.5 | 238.1 | 0.4163 |
SCIDOCS25,657test]: 1,000nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 147.0 | 32.0 | 0.1581 |
modern_colbert | link | 128 | 187.8 | 17.5 | 0.1949 |
answerai_colbert_small | link | 96 | 187.9 | 32.0 | 0.1848 |
lateon | link | 128 | 187.8 | 17.4 | 0.2190 |
lateon_hpool_regularized | link | 128 | 189.6 | 17.4 | 0.2057 |
mlateon | link | 128 | 227.7 | 15.3 | 0.2055 |
neomme_260m_li | link | 128 | 214.6 | 26.2 | 0.1531 |
FiQA-201857,638test]: 648nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 105.0 | 32.0 | 0.3472 |
modern_colbert | link | 128 | 133.5 | 16.7 | 0.4555 |
answerai_colbert_small | link | 96 | 126.9 | 32.0 | 0.4132 |
lateon | link | 128 | 133.5 | 16.7 | 0.5250 |
lateon_hpool_regularized | link | 128 | 133.5 | 16.7 | 0.5065 |
mlateon | link | 128 | 178.1 | 16.4 | 0.4999 |
neomme_260m_li | link | 128 | 151.8 | 24.4 | 0.3678 |
TREC-COVID171,332test]: 50nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
answerai_colbert_small | link | 96 | 171.0 | 32.0 | 0.8297 |
lateon | link | 128 | 171.1 | 17.4 | 0.8390 |
lateon_hpool_regularized | link | 128 | 171.1 | 17.4 | 0.8282 |
mlateon | link | 128 | 233.6 | 17.8 | 0.8194 |
neomme_260m_li | link | 128 | 231.3 | 24.5 | 0.7634 |
Quora522,931test]: 10,000nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
neomme_260m_li | link | 128 | 14.3 | 21.8 | 0.7014 |
LoTTE-pooled2,428,854dev/search]: 2,931Success@5| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | Success@5 | nDCG@10 |
|---|---|---|---|---|---|---|
colbertv2 | link | 128 | 109.6 | 32.0 | N/A | N/A |
answerai_colbert_small | link | 96 | 139.7 | 32.0 | N/A | 0.5154 |
lateon | link | 128 | 146.0 | 12.2 | N/A | 0.5895 |
lateon_hpool_regularized | link | 128 | 146.0 | 12.2 | N/A | 0.5784 |
mlateon | link | 128 | 218.5 | 11.9 | N/A | 0.5706 |
neomme_260m_li | link | 128 | 189.5 | 19.5 | N/A | 0.4291 |
MS MARCO v18,841,823dev.small]: 6,980MRR@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | MRR@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 67.6 | 32.0 | 0.397 |
answerai_colbert_small | link | 96 | 67.6 | 32.0 | 0.3692 |
lateon | link | 128 | 70.9 | 10.2 | 0.3922 |
lateon_hpool_regularized | link | 128 | 70.9 | 10.2 | 0.3793 |
mlateon | link | 128 | 79.6 | 9.7 | 0.3882 |
neomme_260m_li | link | 128 | 73.0 | 17.9 | 0.3267 |
ViDoRe v3nDCG@10| Subset | Encoding | Link | Documents | Queries [test] | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|---|---|---|
hr | neomme_260m_li | link | 1,110 | 1,908 | 128 | 2991.8 | 38.5 | 0.5520 |
computerscience | neomme_260m_li | link | 1,360 | 1,290 | 128 | 3266.0 | 34.8 | 0.6745 |
physics | neomme_260m_li | link | 1,674 | 1,812 | 128 | 2342.0 | 36.2 | 0.4261 |
energy | neomme_260m_li | link | 2,225 | 1,848 | 128 | 2906.5 | 37.0 | 0.5932 |
pharmaceuticals | neomme_260m_li | link | 2,313 | 2,184 | 128 | 2542.4 | 38.4 | 0.5958 |
financefr | neomme_260m_li | link | 2,384 | 1,920 | 128 | 3034.0 | 37.3 | 0.3771 |
finance | neomme_260m_li | link | 2,942 | 1,854 | 128 | 3266.0 | 37.8 | 0.5641 |
industrial | neomme_260m_li | link | 5,244 | 1,698 | 128 | 3224.6 | 39.7 | 0.3991 |
986 followers Β· starred Jun 2026
47 followers Β· starred Jun 2026
TeX
100.0%
An extensive and commented list of resources on Late-Interaction Multivector Retrieval.
TeX
79
53 commits
updated Sep 28, 2026
An extensive and commented list of resources on late-interaction multivector retrieval.
ColBERT: Efficient and Effective Passage Search via Contextualized Late Interaction over BERT
Omar Khattab, Matei Zaharia
SIGIR, 2020
π paper | π οΈ code
COIL: Revisit Exact Lexical Match in Information Retrieval with Contextualized Inverted List
Luyu Gao, Zhuyun Dai, Jamie Callan
NAACL, 2021
π paper
ColBERTv2: Effective and Efficient Retrieval via Lightweight Late Interaction
Keshav Santhanam, Omar Khattab, Jon Saad-Falcon, Christopher Potts, Matei Zaharia
NAACL, 2022
π paper | π οΈ code
Jina-ColBERT-v2: A General-Purpose Multilingual Late Interaction Retriever
Rohan Jha, Bo Wang, Michael GΓΌnther, Georgios Mastrapas, Saba Sturua, Isabelle Mohr, Andreas Koukounas, Mohammad Kalim Akram, Nan Wang, Han Xiao
MRL Workshop, 2024
π paper
PyLate: Flexible Training and Retrieval for Late Interaction Models
Antoine Chaffin, RaphaΓ«l Sourty
CIKM, 2025
π paper | π οΈ code
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models
Antoine Chaffin, Luca Arnaboldi, AmΓ©lie Chatelain, Florent Krzakala
arXiv, 2026
π paper
Your Embedding Model is SMARTer Than You Think
Jianrui Zhang, Hyun Jung Lee, Sukanta Ganguly, Tae-Eui Kam, Donghyun Kim, Yong Jae Lee
arXiv, 2026
π paper | π οΈ code
Party is over: regularizing ColBERT models to fix efficient ANN methods
LightOn AI
Blog, 2026
π blog
NumColBERT: Non-Intrusive Numeracy Injection for Late-Interaction Retrieval Models
Haruki Fujimaki, Makoto P. Kato
arXiv, 2026
π paper
mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models
LightOn AI
Blog, 2026
π blog
GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval
Aarush Sinha
Blog, 2026
π blog
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
Tom Aarsen, Antoine Chaffin, RaphaΓ«l Sourty
Blog, 2026
π blog
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Tom Aarsen
Blog, 2026
π blog
Exploring Static Embedding Retrieval
Logan Markewich
Blog, 2026
π blog
KURE-v2: A Korean-English Bilingual Late-Interaction Retrieval Model
Youngjoon Jang
Blog, 2026
π blog
SmallReason-ColBERT: An Ultra-Small Late-Interaction Retriever for Reasoning Intensive Retrieval
Abdelrahman Abdallah, Mohammed Ali, Adam Jatowt
EMNLP, 2026
π paper | π οΈ code
Compact Token Representations with Contextual Quantization for Efficient Document Re-ranking
Yingrui Yang, Yifan Qiao, Tao Yang
ACL, 2022
π paper
Learned Token Pruning in Contextualized Late Interaction over BERT (ColBERT)
Carlos Lassance, Maroua Maachou, Joohee Park, Stephane Clinchant
SIGIR, 2022
π paper
Introducing Neural Bag of Whole-Words with ColBERTer: Contextualized Late Interactions using Enhanced Reduction
Sebastian Hofstatter, Omar Khattab, Sophia Althammer, Mete Sertkan, Allan Hanbury
CIKM, 2022
π paper | π οΈ code
Joint Optimization of Multi-Vector Representation with Product Quantization
Yufan Fang, Jing Zhan, Yiqun Liu, Jiafeng Mao, Min Zhang, Shaoping Ma
NLPCC, 2022
π paper
Multi-Vector Retrieval as Sparse Alignment
Yujie Qian, Jinhyuk Lee, Sai Meher Karthik Duddu, Zhuyun Dai, Siddhartha Brahma, Iftekhar Naim, Tao Lei, Vincent Y. Zhao
arXiv, 2022
π paper
CITADEL: Conditional Token Interaction via Dynamic Lexical Routing for Efficient and Effective Multi-Vector Retrieval
Minghan Li, Sean C. Lin, Barlas Oguz, Arnab Ghoshal, Jimmy Lin, Yashar Mehdad, Wen-tau Yih, Xilun Chen
ACL, 2023
π paper
SLIM: Sparsified Late Interaction for Multi-Vector Retrieval with Inverted Indexes
Minghan Li, Sheng-Chieh Lin, Xueguang Ma, Jimmy Lin
SIGIR, 2023
π paper
Static Pruning for Multi-Representation Dense Retrieval
Antonio Acquavia, Craig Macdonald, Nicola Tonellotto
DocEng, 2023
π paper | π οΈ code
Rethinking the Role of Token Retrieval in Multi-Vector Retrieval
Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei, Iftekhar Naim, Ming-Wei Chang, Vincent Y. Zhao
NeurIPS, 2023
π paper
SPLATE: Sparse Late Interaction Retrieval
Thibault Formal, Stephane Clinchant, Herve Dejean, Carlos Lassance
SIGIR, 2024
π paper
Reducing the Footprint of Multi-Vector Retrieval with Minimal Performance Impact via Token Pooling
Benjamin ClaviΓ©, Antoine Chaffin, Griffin Adams
arXiv, 2024
π paper
Muvera: Multi-Vector Retrieval via Fixed Dimensional Encodings
Laxman Dhulipala, Majid Hadian, Rajesh Jayaram, Jason Lee, Vahab Mirrokni
NeurIPS, 2024
π paper
Enhancing ColBERT: A Method for Reducing Space Complexity and Accelerating Retrieval Speed
Hai Nguyen T., Huong Le T.
PACLIC, 2024
π paper
Token Pruning Optimization for Efficient Multi-vector Dense Retrieval
Shanxiu He, Mutasem Al-Darabsah, Suraj Nair, Jonathan May, Tarun Agarwal, Tao Yang, Choon Hui Teo
ECIR, 2025
π paper
CRISP: Clustering Multi-Vector Representations for Denoising and Pruning
JoΓ£o Veneroso, Rajesh Jayaram, Jinmeng Rao, Gustavo HernΓ‘ndez Γbrego, Majid Hadian, Daniel Cer
arXiv, 2025
π paper
Towards Lossless Token Pruning in Late-Interaction Retrieval Models
Yuxuan Zong, Benjamin Piwowarski
SIGIR, 2025
π paper
ColPruner: Combining Complementary Pruning Approaches for ColBERT in Web Search
Wondo Rhee, Chan Lim, Taewon Yoon, Gyuhyeon Choi, Jooyoung Lee
ReNeuIR Workshop, 2025
π paper
Towards Storage-Efficient Visual Document Retrieval: An Empirical Study on Reducing Patch-Level Embeddings
Yubo Ma, Jinsong Li, Yuhang Zang, Xiaobao Wu, Xiaoyi Dong, Pan Zhang, Yuhang Cao, Haodong Duan, Jiaqi Wang, Yixin Cao, Aixin Sun
ACL Findings, 2025
π paper
Sculpting the Vector Space: Towards Efficient Multi-Vector Visual Document Retrieval via Prune-then-Merge Framework
Yibo Yan, Mingdong Ou, Yi Cao, Xin Zou, Jiahao Huo, Shuliang Liu, James Kwok, Xuming Hu
arXiv, 2026
π paper
Coverage Matters: MarginMerge for Compressing Multi-Vector Visual Document Retrievers
Ailar Mahdizadeh, Aria Salari, Sohail Rajabi, Shahriar Mirabbasi, Panos Nasiopoulos, Alireza Morsali
arXiv, 2026
π paper
AdaMerge: Tuning-Free Patch Compression for Multi-Vector Visual Document Retrieval
Jianxin You, Kun Ni
CIKM, 2026
π paper
Multi-Vector Index Compression in Any Modality
Hanxiang Qin, Alexander Martin, Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
SIGIR, 2026
π paper | π οΈ code
Learn to Pool: Lightweight Fine-Tuning for Flexible Multi-Vector Compression
Stefan Josef
LIR Workshop, 2026
π paper
A Voronoi Cell Formulation for Principled Token Pruning in Late-Interaction Retrieval Models
Yash Kankanampati, Yuxuan Zong, Nadi Tomeh, Benjamin Piwowarski, Joseph Le Roux
SIGIR, 2026
π paper
CrossQ: Task-Aligned Cross-Token Conditional Quantization for Late Interaction Retrieval
Rohit Kumar Salla, Manoj Saravanan, Ramya Manasa Amancherla
ICML, 2026
π paper
EigenLI: Spectral Approximations to Late Interaction
Archish S, Sabyasachi Basu, Ankit Garg, Ravishankar Krishnaswamy, Kirankumar Shiragur
arXiv, 2026
π paper
Generative Late-Interaction Embeddings For Visual Document Retrieval
Mohamed Eltahir, Talal Aloushan, Rose Khairoalsendi, Jana Shata, Mohammed Alhassan, Leen Alrehaili, Tanveer Hussain, Naeemullah Khan
arXiv, 2026
π paper
ColPali: Efficient Document Retrieval with Vision Language Models
Manuel Faysse, Hugues Sibille, Tony Wu, Bilel Omrani, Gautier Viaud, Celine Hudelot, Pierre Colombo
ICLR, 2025
π paper
Video-ColBERT: Contextualized Late Interaction for Text-to-Video Retrieval
Arun V. Reddy, Alexander Martin, Eugene Yang, Andrew Yates, Kate Sanders, Kenton Murray, Reno Kriz, Celso M. de Melo, Benjamin Van Durme, Rama Chellappa
CVPR, 2025
π paper
ColMate: Contrastive Late Interaction and Masked Text for Multimodal Document Retrieval
Ahmed Masry, Megh Thakkar, Patrice Bechard, Sathwik Tejaswi Madhusudhan, Rabiul Awal, Shambhavi Mishra, Akshay Kalkunte Suresh, Srivatsava Daruru, Enamul Hoque, Spandana Gella, Torsten Scholak, Sai Rajeswar
EMNLP, 2025
π paper
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction
Zilin Xiao, Qi Ma, Mengting Gu, Chun-cheng Jason Chen, Xintao Chen, Vicente Ordonez, Vijai Mohan
ICLR, 2026
π paper | π οΈ code
AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval
Xinze Liu, Lei Yang, Dayan Wu, Hengjie Zhu, Zihao Zhang, Hanqi Wu, Tianzhu Hu, Peng Fu, Zheng Lin, Weiping Wang
arXiv, 2026
π paper
Multi-Vector Embeddings are Provably More Expressive than Single Vector Embeddings
Rajesh Jayaram
arXiv, 2026
π paper
Quantifying and Expanding the Theoretical Capacity of Late-Interaction Retrieval Models
Julian Killingback, Varad Ingale, Hamed Zamani, Cameron Musco
arXiv, 2026
π paper
Retrieval Needs Multivectors: An Exponential Separation
Mihir Agarwal, Viraj Agrawal, Sabyasachi Basu, Ankit Garg, Kirankumar Shiragur
arXiv, 2026
π paper
Baleen: Robust Multi-Hop Reasoning at Scale via Condensed Retrieval
Omar Khattab, Christopher Potts, Matei Zaharia
NeurIPS, 2021
π paper
PLAID: An Efficient Engine for Late Interaction Retrieval
Keshav Santhanam, Omar Khattab, Christopher Potts, Matei Zaharia
CIKM, 2022
π paper | π οΈ code
DESSERT: An Efficient Algorithm for Vector Set Search with Vector Set Queries
Joshua Engels, Benjamin Coleman, Vihan Lakshman, Anshumali Shrivastava
NeurIPS, 2023
π paper
Efficient Multi-Vector Dense Retrieval with Bit Vectors
Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
ECIR, 2024
π paper | π οΈ code
Efficient Constant-Space Multi-vector Retrieval
Sean MacAvaney, Antonio Mallia, Nicola Tonellotto
ECIR, 2025
π paper
IGP: Efficient Multi-Vector Retrieval via Proximity Graph Index
Ziyang Bian, Man Lung Yiu, Buzhou Tang
SIGIR, 2025
π paper | π οΈ code
WARP: An Efficient Engine for Multi-Vector Retrieval
Jan Luca Scheerer, Matei Zaharia, Christopher Potts, Gustavo Alonso, Omar Khattab
SIGIR, 2025
π paper | π οΈ code
Multivector Reranking in the Era of Strong First-Stage Retrievers
Silvio Martinico, Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
ECIR, 2026
π paper | π οΈ code | π οΈ code
SMVE: Sparse Multi-Vector Retrieval
Martin Spisak, Marek Galovic
LIR Workshop, 2026
π blog
No More K-means: Single-Stage Sparse Coding for Efficient Multi-Vector Retrieval
Lixuan Guo, Yifei Wang, Tiansheng Wen, Aosong Feng, Stefanie Jegelka, Chenyu You
ICML, 2026
π paper
LEMUR: Learned Multi-Vector Retrieval
Elias JÀÀsaari, Ville Hyvânen, Teemu Roos
ICML, 2026
π paper | π οΈ code
Efficient Multivector Retrieval with Token-Aware Clustering and Hierarchical Indexing
Silvio Martinico, Franco Maria Nardini, Cosimo Rulli, Rossano Venturini
SIGIR, 2026
π paper | π οΈ code
ColBERTSaR: Sparsified ColBERT Index via Product Quantization
Eugene Yang, Andrew Yates, Dawn Lawrie, James Mayfield, Saron Samuel, Rohan Jha
SIGIR, 2026
π paper | π οΈ code
PLAID-PRF: Pseudo-Relevance Feedback with Centroid-like Tokens in PLAID
Xiao Wang, Sean MacAvaney, Craig Macdonald
SIGIR, 2026
π paper | π οΈ code
Chimera: Efficient Multi-Vector Retrieval via GPU-CPU Co-Processing
Yanqi Chen, Juelin Liu, Alexandra Meliou, Xiao Yan
arXiv, 2026
π paper
Col-Bandit: Query-Time Top-K Estimation for Late-Interaction Retrieval
Roi Pony, Adi Raz Goldfarb, Oshri Naparstek, Idan Friedman, Udi Barzelay, Eli Schwartz
arXiv, 2026
π paper
FLASH-MAXSIM: IO-Aware Fused Kernels for Late-Interaction Scoring
Roi Pony, Adi Raz Goldfarb, Idan Friedman, Daniel Ezer, Udi Barzelay
arXiv, 2026
π paper | π οΈ code
TileMaxSim: IO-Aware GPU MaxSim Scoring with Dimension Tiling and Fused Product Quantization
Ashutosh Sharma
arXiv, 2026
π paper | π οΈ code
A Reproducibility Study of PLAID
Sean MacAvaney, Nicola Tonellotto
SIGIR, 2024
π paper
A Replicability Study of XTR
Rohan Jha, Reno Kriz, Benjamin Van Durme
arXiv, 2026
π paper
Reproduction Beyond Benchmarks: ConstBERT and ColBERT-v2 Across Backends and Query Distributions
Utshab Kumar Ghosh, Ashish David, Shubham Chatterjee
SIGIR, 2026
π paper
Comparing Token Pruning Approaches for Multi-Vector Retrieval
Ferdinand Schlatt, Hanno Barschel, Matthias Hagen
SIGIR, 2026
π paper
A Brief Comparison of Training-Free Multi-Vector Sequence Compression Methods
Rohan Jha, Chunsheng Zuo, Reno Kriz, Benjamin Van Durme
LIR Workshop, 2026
π paper
A Survey of Late-Interaction Neural Retrieval: Paradigms, Systems, and Research Frontiers
Xiao Wang, Chuting Yu, Minghan Li, Binci Yang, Hang Li, Ben He
SSRN, 2026
π paper
ColBERT
Reference implementation for ColBERT and ColBERTv2, and includes PLAID support for efficient late-interaction retrieval.
RAGatouille
Python toolkit to train and serve ColBERT-based late-interaction retrievers.
PyLate
Python library for training, fine-tuning, inference, and retrieval with ColBERT-style late-interaction models on single and multi-GPU setups.
Sentence Transformers
Framework for dense, sparse, reranker, and (v6.0+) ColBERT-style multi-vector late-interaction models via the MultiVectorEncoder, with a unified training and inference API.
PyLate-rs
High-performance Rust inference engine for PyLate models, with Python bindings and optimized integration with FastPlaid for retrieval pipelines.
FastPlaid
GPU-optimized engine for ColBERT/PLAID-style late-interaction retrieval.
kANNolo
ANN library for dense, sparse, and multivector retrieval.
Vectorium ![]()
Rust library for compact storage/access of dense, sparse, and multivector embeddings.
Firn ![]()
Rust search engine for single-vector and late-interaction multivector namespaces, backed by LanceDB on object storage with RAM/NVMe result caching.
NextPlaid
CPU-oriented local-first multivector retrieval engine with memory-mapped storage.
EMVB
Reference implementation for Efficient Multi-Vector Dense Retrieval with Bit Vectors.
IGP
Official C++ implementation for IGP: proximity-graph indexing for multi-vector retrieval (with Python scripts for experiments).
WARP
Official implementation for WARP, an efficient multi-vector retrieval engine.
ColGrep
High-performance code search CLI tool powered by LateOn-Code and NextPlaid, enabling semantic + hybrid (regex + semantic) code retrieval locally with incremental indexing.
TACHIOM
Fast and scalable multivector retrieval system with Token-Aware Clustering (TAC) and hierarchical Product Quantization for efficient late-interaction search.
TopK
Managed retrieval engine with support for late-interaction search over billions of documents, online index updates, filtering, and more.
Flash-MaxSim
IO-aware Triton kernel for MaxSim scoring in ColBERT/ColPali pipelines: tile-by-tile on-chip computation with zero intermediate memory and INT8 quantization support.
maxsim
Ahead-of-time compiled MaxSim kernel with CUDA and Metal backends (NVIDIA + Apple Silicon), distributed as a HuggingFace kernels package.
late-interaction-kernels
Fused Triton kernels for MaxSim scoring with CUDA, Metal, and CPU backends, native PyLate/colpali-engine integration, and PLAID-style compressed-index support.
maxsim-cpu
CPU-only MaxSim kernel written in Rust (libxsmm on x86, Apple Accelerate on ARM) with Python bindings.
TileMaxSim
IO-aware Triton kernel for MaxSim scoring with dimension tiling for embeddings wider than 128 dims and fused product quantization, achieving 80%+ peak HBM bandwidth.
colbert-ir/colbertv2.0
Official ColBERTv2 checkpoint (MS MARCO-trained) from the ColBERT authors, widely used as the canonical baseline model.
lightonai/LateOn
State-of-the-art ColBERT model (149M, ModernBERT-based) achieving 57.22 NDCG@10 on BEIR with fully open training data and strong generalization under decontamination.
lightonai/mLateOn
Multilingual ColBERT model (307M, mmBERT-based) covering 9 languages with SOTA multilingual, long-document, and code retrieval; strong zero-shot generalization to unseen languages (57.56 NDCG@10 on BEIR).
jinaai/jina-colbert-v2
Multilingual late-interaction retriever (0.6B, JinaBERT-based) supporting 89 languages with Matryoshka token embeddings (128/96/64 dims) for flexible efficiency-precision tradeoffs.
chungimungi/GLInt
English ColBERT model (149M, LateOn-based) trained with geometry-matched hard negatives mined via MaxSim geometry, reporting 57.43 NDCG@10 on BEIR.
lightonai/LateOn-regularized
LateOn variant trained with STE-based regularization to fix compatibility with projection-based retrieval methods (MUVERA, SMVE).
lightonai/ColBERT-Zero
Large-scale fully pre-trained ColBERT checkpoint trained on public data and released with the ColBERT-Zero paper.
lightonai/GTE-ModernColBERT-v1
PyLate late-interaction checkpoint based on ModernBERT with 128-dimensional token embeddings and strong long-context retrieval behavior.
topk-io/Iso-ModernColBERT
Isotropically corrected version of GTE-ModernColBERT-v1 built for efficient inference and scalable retrieval.
sebastian-hofstaetter/colberter-128-32-msmarco / sebastian-hofstaetter/uni-colberter-128-1-msmarco
ColBERTer checkpoints trained on MS MARCO (128-dim, with 32 and 1 unique whole-word vectors per document respectively).
lightonai/LateOn-Code
Specialized ColBERT model (149M parameters) fine-tuned for code retrieval, achieving SOTA on MTEB Code benchmark.
lightonai/LateOn-Code-edge
Lightweight code retrieval model (17M parameters) for edge devices, matching larger models while running efficiently on CPU.
lightonai/Reason-ModernColBERT
Reasoning-focused late-interaction checkpoint fine-tuned on reasonir-hq, with strong BRIGHT benchmark performance for reasoning-intensive retrieval.
nlpai-lab/KURE-v2
Korean-English bilingual late-interaction model (154M, skt/A.X-Encoder-base) with 128-dim token vectors and 8,192-token context, reporting 0.8160 average nDCG@10 on MTEB(kor, v2).
nlpai-lab/KURE-v2-unsupervised
Stage-1 KURE-v2 checkpoint trained with weakly-supervised contrastive learning only (20.7M pairs, no relevance labels), reaching 0.7283 average nDCG@10 on MTEB(kor, v2).
DataScience-UIBK/SmallReason-ColBERT-32M
Ultra-small reasoning retriever (32M, mxbai-edge-colbert-v0-32m-based) with a query-side token-importance head, reporting 21.41 mean nDCG@10 on BRIGHT; load via WeightedColBERT.from_base(), as plain PyLate loading drops the head.
vidore/colpali-v1.3
Latest ColPali release (PaliGemma-3B + LoRA) for visual document retrieval, producing ColBERT-style multi-vector embeddings of page images.
vidore/colqwen2-v1.0
ColPali-style visual document retriever on Qwen2-VL-2B-Instruct, accepting dynamic image resolutions without aspect-ratio distortion (up to 768 patches).
vidore/colqwen2.5-v0.2
ColPali-style visual document retriever on Qwen2.5-VL-3B-Instruct, with dynamic image resolutions (up to 768 patches).
vidore/colSmol-256M / vidore/colSmol-500M
Lightweight ColPali-style visual document retrievers built on SmolVLM-256M-Instruct and SmolVLM-500M-Instruct.
NFCorpus3,633test]: 323nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 154.5 | 32.0 | 0.3299 |
modern_colbert | link | 128 | 237.4 | 8.6 | 0.3792 |
answerai_colbert_small | link | 96 | 235.5 | 32.0 | 0.3683 |
modernbert_xtr | link | 128 | 288.2 | 32.0 | 0.3430 |
lateon | link | 128 | 237.4 | 8.6 | 0.3809 |
lateon_hpool_regularized | link | 128 | 238.5 | 8.6 | 0.3803 |
mlateon | link | 128 | 343.7 | 7.4 | 0.3786 |
neomme_260m_li | link | 128 | 337.3 | 17.1 | 0.3081 |
SciFact5,183test]: 300nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
answerai_colbert_small | link | 96 | 235.7 | 32.0 | 0.7431 |
lateon | link | 128 | 231.2 | 21.0 | 0.7627 |
lateon_hpool_regularized | link | 128 | 231.2 | 21.0 | 0.7608 |
mlateon | link | 128 | 314.4 | 21.0 | 0.7605 |
neomme_260m_li | link | 128 | 321.2 | 31.0 | 0.7161 |
ArguAna8,674test]: 1,406nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
neomme_260m_li | link | 128 | 194.5 | 238.1 | 0.4163 |
SCIDOCS25,657test]: 1,000nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 147.0 | 32.0 | 0.1581 |
modern_colbert | link | 128 | 187.8 | 17.5 | 0.1949 |
answerai_colbert_small | link | 96 | 187.9 | 32.0 | 0.1848 |
lateon | link | 128 | 187.8 | 17.4 | 0.2190 |
lateon_hpool_regularized | link | 128 | 189.6 | 17.4 | 0.2057 |
mlateon | link | 128 | 227.7 | 15.3 | 0.2055 |
neomme_260m_li | link | 128 | 214.6 | 26.2 | 0.1531 |
FiQA-201857,638test]: 648nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 105.0 | 32.0 | 0.3472 |
modern_colbert | link | 128 | 133.5 | 16.7 | 0.4555 |
answerai_colbert_small | link | 96 | 126.9 | 32.0 | 0.4132 |
lateon | link | 128 | 133.5 | 16.7 | 0.5250 |
lateon_hpool_regularized | link | 128 | 133.5 | 16.7 | 0.5065 |
mlateon | link | 128 | 178.1 | 16.4 | 0.4999 |
neomme_260m_li | link | 128 | 151.8 | 24.4 | 0.3678 |
TREC-COVID171,332test]: 50nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
answerai_colbert_small | link | 96 | 171.0 | 32.0 | 0.8297 |
lateon | link | 128 | 171.1 | 17.4 | 0.8390 |
lateon_hpool_regularized | link | 128 | 171.1 | 17.4 | 0.8282 |
mlateon | link | 128 | 233.6 | 17.8 | 0.8194 |
neomme_260m_li | link | 128 | 231.3 | 24.5 | 0.7634 |
Quora522,931test]: 10,000nDCG@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|
neomme_260m_li | link | 128 | 14.3 | 21.8 | 0.7014 |
LoTTE-pooled2,428,854dev/search]: 2,931Success@5| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | Success@5 | nDCG@10 |
|---|---|---|---|---|---|---|
colbertv2 | link | 128 | 109.6 | 32.0 | N/A | N/A |
answerai_colbert_small | link | 96 | 139.7 | 32.0 | N/A | 0.5154 |
lateon | link | 128 | 146.0 | 12.2 | N/A | 0.5895 |
lateon_hpool_regularized | link | 128 | 146.0 | 12.2 | N/A | 0.5784 |
mlateon | link | 128 | 218.5 | 11.9 | N/A | 0.5706 |
neomme_260m_li | link | 128 | 189.5 | 19.5 | N/A | 0.4291 |
MS MARCO v18,841,823dev.small]: 6,980MRR@10| Encoding | Link | Vector dim | Avg vectors per doc | Avg vectors per query | MRR@10 |
|---|---|---|---|---|---|
colbertv2 | link | 128 | 67.6 | 32.0 | 0.397 |
answerai_colbert_small | link | 96 | 67.6 | 32.0 | 0.3692 |
lateon | link | 128 | 70.9 | 10.2 | 0.3922 |
lateon_hpool_regularized | link | 128 | 70.9 | 10.2 | 0.3793 |
mlateon | link | 128 | 79.6 | 9.7 | 0.3882 |
neomme_260m_li | link | 128 | 73.0 | 17.9 | 0.3267 |
ViDoRe v3nDCG@10| Subset | Encoding | Link | Documents | Queries [test] | Vector dim | Avg vectors per doc | Avg vectors per query | nDCG@10 |
|---|---|---|---|---|---|---|---|---|
hr | neomme_260m_li | link | 1,110 | 1,908 | 128 | 2991.8 | 38.5 | 0.5520 |
computerscience | neomme_260m_li | link | 1,360 | 1,290 | 128 | 3266.0 | 34.8 | 0.6745 |
physics | neomme_260m_li | link | 1,674 | 1,812 | 128 | 2342.0 | 36.2 | 0.4261 |
energy | neomme_260m_li | link | 2,225 | 1,848 | 128 | 2906.5 | 37.0 | 0.5932 |
pharmaceuticals | neomme_260m_li | link | 2,313 | 2,184 | 128 | 2542.4 | 38.4 | 0.5958 |
financefr | neomme_260m_li | link | 2,384 | 1,920 | 128 | 3034.0 | 37.3 | 0.3771 |
finance | neomme_260m_li | link | 2,942 | 1,854 | 128 | 3266.0 | 37.8 | 0.5641 |
industrial | neomme_260m_li | link | 5,244 | 1,698 | 128 | 3224.6 | 39.7 | 0.3991 |
986 followers Β· starred Jun 2026
47 followers Β· starred Jun 2026
TeX
100.0%