chungimungi/GLInt

Model

GLInt

3

27 commits

1 linked in READMEs

updated Aug 27, 2026

See the code

README

GLInt

GLINT is a SOTA 149M-parameter English late-interaction retriever built from LateOn-unsupervised. It retains 128-dimensional token embeddings and uses MaxSim retrieval with 32 query tokens and 300 document tokens.

What is new in GLInt?

GLINT is designed around the mismatch between ordinary dense hard-negative mining and a late-interaction retriever. Dense mining selects documents that are difficult under one pooled vector; GLINT instead mines negatives under the same token-level MaxSim geometry used at retrieval time. This exposes lexical, compositional, and localized token matches that a single-vector miner can miss.

The training recipe has two stages:

  1. supervised fine-tuning with multi-vector (MaxSim) hard negatives;
  2. mixed listwise knowledge distillation over a diverse seven-source hard-negative mixture.

For the second stage, a frozen listwise teacher (jinaai/jina-reranker-v3.5) scores each 32-document candidate set jointly. GLINT distils that ordering with a sharpened listwise KL objective, while a false-negative-masked InfoNCE term preserves a direct retrieval signal.

Usage

Sentence Transformers

This model can be used with Sentence Transformers as a multi-vector (ColBERT-style late interaction) retriever via the MultiVectorEncoder:

pip install "sentence-transformers>=6.0.0"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("chungimungi/GLInt")

query = "Which planet is known as the Red Planet?"
documents = [
    "Venus is often called Earth's twin because of its similar size and proximity.",
    "Mars, known for its reddish appearance, is often referred to as the Red Planet.",
    "Jupiter, the largest planet in our solar system, has a prominent red spot.",
    "Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
]

query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# torch.Size([12, 128]) torch.Size([18, 128])

# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[11.6192, 11.7344, 11.6513, 11.7105]], device='cuda:0')

PyLate

from pylate import models

model = models.ColBERT("chungimungi/GLInt")
query_embeddings = model.encode(["what causes a lunar eclipse?"], is_query=True)
document_embeddings = model.encode(
    ["A lunar eclipse happens when Earth passes between the Sun and the Moon."],
    is_query=False,
)

Use a late-interaction backend such as PyLate/PLAID for corpus-scale retrieval. Scores are computed by summing, over query tokens, the maximum similarity to a document token.

Results

BEIR (15 datasets, NDCG@10)

ModelAverageSize (M)Embed dimArguAnaCQADupstackRetrievalClimateFEVERDBPediaFEVERFiQA2018HotpotQAMSMARCONFCorpusNQQuoraRetrievalSCIDOCSSciFactTRECCOVIDTouche2020
ColBERTv248.6311012846.5038.3017.6045.2078.5035.4067.5046.0033.7052.4085.5015.4068.9072.6026.00
Jina-ColBERT-v251.8560012836.6040.8023.9047.1080.5040.8076.6046.9034.6064.0088.7018.6067.8083.4027.40
ColBERT-small53.79339650.0938.7533.0745.5890.9641.1576.1143.5037.3059.1087.7218.4274.7784.5925.69
GTE-ModernColBERT-v154.7514912847.5241.0831.3347.5687.6745.2577.4845.6037.8361.6286.7119.2276.3384.8431.25
ColBERT-Zero55.3914912852.8241.4135.9047.4390.5242.5079.4545.9537.2161.8285.1919.8476.3378.2736.24
LateOn-unsupervised50.1114912843.1247.7118.7643.3665.7451.9468.1737.5137.1558.4189.4821.1376.8969.8122.53
LateOn57.2214912850.5247.3639.6745.9992.0253.1279.9845.6737.7963.9189.6721.9076.6183.6030.52
GLInt57.4314912852.3846.4934.1747.6892.4550.8582.5446.3837.5168.0390.0820.6577.1384.7830.26

BEIR-Decontaminated (14 datasets, NDCG@10)

ModelAverageArguAnaClimateFEVERDBPediaFEVERFiQA2018HotpotQAMS MARCONFCorpusNatural QuestionsQuoraSciDocsSciFactTREC-COVIDTouché-2020
GLInt62.5051.6736.3542.5092.8956.8881.1672.7026.2194.9792.0622.0289.0781.5134.97
LateOn61.452.242.131.792.757.978.970.327.093.191.515.188.980.936.8
DenseOn58.840.039.528.891.255.973.768.928.592.191.114.785.482.531.0
pplx-embed-v1-0.6b59.743.742.428.491.155.273.571.928.091.691.515.489.083.730.0
jina-v5-text-nano58.847.241.630.290.051.567.568.629.492.391.314.989.476.833.2
harrier-oss-v1-0.6b58.047.425.731.380.750.171.473.427.990.090.917.190.781.833.3
arctic-embed-l-v257.943.145.745.792.250.463.171.026.090.791.313.987.481.426.8
bge-large-en-v1.557.346.039.028.987.649.375.268.929.885.991.314.086.572.726.9
Qwen3-Embedding-0.6B57.048.438.025.386.449.162.263.625.888.390.015.385.587.931.8
GTE-ModernBERT56.652.547.525.994.155.565.564.826.184.590.811.688.662.423.1
bge-base-en-v1.556.245.632.926.786.844.572.766.827.485.691.113.887.676.628.1
Nomic v1.555.935.843.528.886.844.772.767.424.485.187.212.783.380.729.4
modernbert-embed-base55.636.537.824.787.846.062.765.324.389.389.912.985.582.733.1
ColBERT-Zero60.054.536.833.090.546.677.874.226.691.188.314.289.575.340.9
pplx-embed-v1-late-0.6b59.860.936.429.989.750.978.669.227.992.883.813.589.380.234.7
GTE-ModernColBERT59.348.833.533.288.150.277.371.627.393.189.113.687.781.435.3
colbert-small58.147.735.731.789.345.677.171.425.086.290.113.189.281.529.0

Training data and reproducibility

The corresponding private training artifacts are in GLINT-data. It contains the complete prepared SFT data, the 1,046,009-row seven-source KD mixture, and Jina teacher-score parquet shards. The repository contains no BEIR evaluation corpus or evaluation labels.

Citation

@misc{aarush2026glint,
  title={GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval},
  author={Aarush},
  year={2026},
  howpublished={\url{https://huggingface.co/blog/chungimungi/glint}},
}
colbert
endpoints_compatible
late-interaction
modernbert
multi-vector
pylate
retrieval
safetensors
sentence-similarity
sentence-transformers
text-embeddings-inference

chungimungi/GLInt

Model

GLInt

3

27 commits

1 linked in READMEs

updated Aug 27, 2026

See the code

README

GLInt

GLINT is a SOTA 149M-parameter English late-interaction retriever built from LateOn-unsupervised. It retains 128-dimensional token embeddings and uses MaxSim retrieval with 32 query tokens and 300 document tokens.

What is new in GLInt?

GLINT is designed around the mismatch between ordinary dense hard-negative mining and a late-interaction retriever. Dense mining selects documents that are difficult under one pooled vector; GLINT instead mines negatives under the same token-level MaxSim geometry used at retrieval time. This exposes lexical, compositional, and localized token matches that a single-vector miner can miss.

The training recipe has two stages:

  1. supervised fine-tuning with multi-vector (MaxSim) hard negatives;
  2. mixed listwise knowledge distillation over a diverse seven-source hard-negative mixture.

For the second stage, a frozen listwise teacher (jinaai/jina-reranker-v3.5) scores each 32-document candidate set jointly. GLINT distils that ordering with a sharpened listwise KL objective, while a false-negative-masked InfoNCE term preserves a direct retrieval signal.

Usage

Sentence Transformers

This model can be used with Sentence Transformers as a multi-vector (ColBERT-style late interaction) retriever via the MultiVectorEncoder:

pip install "sentence-transformers>=6.0.0"
from sentence_transformers import MultiVectorEncoder

model = MultiVectorEncoder("chungimungi/GLInt")

query = "Which planet is known as the Red Planet?"
documents = [
    "Venus is often called Earth's twin because of its similar size and proximity.",
    "Mars, known for its reddish appearance, is often referred to as the Red Planet.",
    "Jupiter, the largest planet in our solar system, has a prominent red spot.",
    "Saturn, famous for its rings, is sometimes mistaken for the Red Planet.",
]

query_embeddings = model.encode_query(query)
document_embeddings = model.encode_document(documents)
print(query_embeddings.shape, document_embeddings[0].shape)
# torch.Size([12, 128]) torch.Size([18, 128])

# MaxSim late-interaction scoring (higher is more relevant)
scores = model.similarity(query_embeddings, document_embeddings)
print(scores)
# tensor([[11.6192, 11.7344, 11.6513, 11.7105]], device='cuda:0')

PyLate

from pylate import models

model = models.ColBERT("chungimungi/GLInt")
query_embeddings = model.encode(["what causes a lunar eclipse?"], is_query=True)
document_embeddings = model.encode(
    ["A lunar eclipse happens when Earth passes between the Sun and the Moon."],
    is_query=False,
)

Use a late-interaction backend such as PyLate/PLAID for corpus-scale retrieval. Scores are computed by summing, over query tokens, the maximum similarity to a document token.

Results

BEIR (15 datasets, NDCG@10)

ModelAverageSize (M)Embed dimArguAnaCQADupstackRetrievalClimateFEVERDBPediaFEVERFiQA2018HotpotQAMSMARCONFCorpusNQQuoraRetrievalSCIDOCSSciFactTRECCOVIDTouche2020
ColBERTv248.6311012846.5038.3017.6045.2078.5035.4067.5046.0033.7052.4085.5015.4068.9072.6026.00
Jina-ColBERT-v251.8560012836.6040.8023.9047.1080.5040.8076.6046.9034.6064.0088.7018.6067.8083.4027.40
ColBERT-small53.79339650.0938.7533.0745.5890.9641.1576.1143.5037.3059.1087.7218.4274.7784.5925.69
GTE-ModernColBERT-v154.7514912847.5241.0831.3347.5687.6745.2577.4845.6037.8361.6286.7119.2276.3384.8431.25
ColBERT-Zero55.3914912852.8241.4135.9047.4390.5242.5079.4545.9537.2161.8285.1919.8476.3378.2736.24
LateOn-unsupervised50.1114912843.1247.7118.7643.3665.7451.9468.1737.5137.1558.4189.4821.1376.8969.8122.53
LateOn57.2214912850.5247.3639.6745.9992.0253.1279.9845.6737.7963.9189.6721.9076.6183.6030.52
GLInt57.4314912852.3846.4934.1747.6892.4550.8582.5446.3837.5168.0390.0820.6577.1384.7830.26

BEIR-Decontaminated (14 datasets, NDCG@10)

ModelAverageArguAnaClimateFEVERDBPediaFEVERFiQA2018HotpotQAMS MARCONFCorpusNatural QuestionsQuoraSciDocsSciFactTREC-COVIDTouché-2020
GLInt62.5051.6736.3542.5092.8956.8881.1672.7026.2194.9792.0622.0289.0781.5134.97
LateOn61.452.242.131.792.757.978.970.327.093.191.515.188.980.936.8
DenseOn58.840.039.528.891.255.973.768.928.592.191.114.785.482.531.0
pplx-embed-v1-0.6b59.743.742.428.491.155.273.571.928.091.691.515.489.083.730.0
jina-v5-text-nano58.847.241.630.290.051.567.568.629.492.391.314.989.476.833.2
harrier-oss-v1-0.6b58.047.425.731.380.750.171.473.427.990.090.917.190.781.833.3
arctic-embed-l-v257.943.145.745.792.250.463.171.026.090.791.313.987.481.426.8
bge-large-en-v1.557.346.039.028.987.649.375.268.929.885.991.314.086.572.726.9
Qwen3-Embedding-0.6B57.048.438.025.386.449.162.263.625.888.390.015.385.587.931.8
GTE-ModernBERT56.652.547.525.994.155.565.564.826.184.590.811.688.662.423.1
bge-base-en-v1.556.245.632.926.786.844.572.766.827.485.691.113.887.676.628.1
Nomic v1.555.935.843.528.886.844.772.767.424.485.187.212.783.380.729.4
modernbert-embed-base55.636.537.824.787.846.062.765.324.389.389.912.985.582.733.1
ColBERT-Zero60.054.536.833.090.546.677.874.226.691.188.314.289.575.340.9
pplx-embed-v1-late-0.6b59.860.936.429.989.750.978.669.227.992.883.813.589.380.234.7
GTE-ModernColBERT59.348.833.533.288.150.277.371.627.393.189.113.687.781.435.3
colbert-small58.147.735.731.789.345.677.171.425.086.290.113.189.281.529.0

Training data and reproducibility

The corresponding private training artifacts are in GLINT-data. It contains the complete prepared SFT data, the 1,046,009-row seven-source KD mixture, and Jina teacher-score parquet shards. The repository contains no BEIR evaluation corpus or evaluation labels.

Citation

@misc{aarush2026glint,
  title={GLInt: Geometry-Matched Hard Negatives for Late-Interaction Retrieval},
  author={Aarush},
  year={2026},
  howpublished={\url{https://huggingface.co/blog/chungimungi/glint}},
}
colbert
endpoints_compatible
late-interaction
modernbert
multi-vector
pylate
retrieval
safetensors
sentence-similarity
sentence-transformers
text-embeddings-inference