lightonai/MonoQwen2-VL-v0.1

Model

48

stars

33

commits

6

repos using this model

1

linked in READMEs

Jul 17, 2026

updated

peft
qwen2_vl
reranker
safetensors
sentence-transformers
vidore
visual-document-retrieval
Browse cluster: Visual Document Retrieval & ColPali

README

MonoQwen2-VL-v0.1

Model Overview

The MonoQwen2-VL-v0.1 is a multimodal reranker finetuned with LoRA from Qwen2-VL-2B, optimized for asserting pointwise image-query relevance using the MonoT5 objective. That is, given a couple of image and query fed into the prompt of the VLM, the model is tasked to generate "True" if the image is relevant to the query and "False" otherwise. During inference, a relevancy score can then be obtained by comparing the logits of the two tokens and this score can effectively be used to rerank the candidates generated by a first-stage retriever (such as DSE or ColPali) or filter them using a threshold.

The ColPali train set was used to train this model with negatives mined using DSE.

How to Use the Model

Using Sentence Transformers

Install Sentence Transformers (>= 5.6.0) with the image extra, plus peft for the LoRA adapter (requires transformers v5):

pip install "sentence-transformers[image]" peft

The reranking prompt used during training is baked into the bundled chat template, and the returned scores are the P("True") relevance probabilities:

from sentence_transformers import CrossEncoder

model = CrossEncoder("lightonai/MonoQwen2-VL-v0.1")

# Pass document page images as PIL.Image, a local file path, or an image URL
query = "What is the variable represented on the y-axis of the graph?"
pages = [
    "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc1.jpg",
    "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc2.jpg",
]

scores = model.predict([(query, page) for page in pages])
print(scores)
# [0.61328125 0.29492188]

rankings = model.rank(query, pages)
print(rankings)
# [{'corpus_id': 0, 'score': 0.61328125}, {'corpus_id': 1, 'score': 0.29492188}]

The LoRA adapter is loaded onto the Qwen/Qwen2-VL-2B-Instruct base model automatically. The model loads in bfloat16 by default, but you can pass model_kwargs={"dtype": torch.float32} to CrossEncoder(...) for full fp32 precision. Text passages can also be passed as documents in place of images.

Using Transformers

Below is a quick example to rerank a single image against a user query using this model:

import requests
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen2VLForConditionalGeneration

# Load processor and model
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-VL-2B-Instruct")
model = Qwen2VLForConditionalGeneration.from_pretrained(
    "lightonai/MonoQwen2-VL-v0.1",
    device_map="auto",
    # attn_implementation="flash_attention_2",
    # torch_dtype=torch.bfloat16,
)

# Define query and load image
query = "What is the variable represented on the y-axis of the graph?"
image_url = "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc1.jpg"
image = Image.open(requests.get(image_url, stream=True).raw)

# Construct the prompt and prepare input
prompt = (
    "Assert the relevance of the previous image document to the following query, "
    "answer True or False. The query is: {query}"
).format(query=query)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": prompt},
        ],
    }
]

# Apply chat template and tokenize
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=text, images=image, return_tensors="pt").to("cuda")

# Run inference to obtain logits
with torch.no_grad():
    outputs = model(**inputs)
    logits_for_last_token = outputs.logits[:, -1, :]

# Convert tokens and calculate relevance score
true_token_id = processor.tokenizer.convert_tokens_to_ids("True")
false_token_id = processor.tokenizer.convert_tokens_to_ids("False")
relevance_score = torch.softmax(logits_for_last_token[:, [true_token_id, false_token_id]], dim=-1)

# Extract and display probabilities
true_prob = relevance_score[0, 0].item()
false_prob = relevance_score[0, 1].item()

print(f"True probability: {true_prob:.4f}, False probability: {false_prob:.4f}")
# True probability: 0.6133, False probability: 0.3848

This example demonstrates how to use the model to assess the relevance of an image with respect to a query. It outputs the probability that the image is relevant ("True") or not relevant ("False").

Note: this example requires peft to be installed in your environment (pip install peft). If you don't want to use peft, you can use model.load_adapter on the original Qwen2-VL-2B model.

Performance Metrics

The model has been evaluated on ViDoRe Benchmark, by retrieving 10 elements with MrLight_dse-qwen2-2b-mrl-v1 and reranking them. The table below summarizes its ndcg@5 scores:

DatasetMrLight_dse-qwen2-2b-mrl-v1MonoQwen2-VL-v0.1 reranking
vidore/arxivqa_test_subsampled85.689.0
vidore/docvqa_test_subsampled57.159.7
vidore/infovqa_test_subsampled88.193.2
vidore/tabfquad_test_subsampled93.196.0
vidore/shiftproject_test82.093.0
vidore/syntheticDocQA_artificial_intelligence_test97.5100.0
vidore/syntheticDocQA_energy_test92.997.7
vidore/syntheticDocQA_government_reports_test96.098.0
vidore/syntheticDocQA_healthcare_industry_test96.499.3
vidore/tatdqa_test69.479.0
Mean85.890.5

License

This LoRA model is licensed under the Apache 2.0 license.

Citation

If you find the model useful, consider citing our work:

@misc{MonoQwen,
  title={MonoQwen: Visual Document Reranking},
  author={Chaffin, Antoine and Lac, Aurélien},
  url={https://huggingface.co/lightonai/MonoQwen2-VL-v0.1},
  year={2024}
}

Contributors

AL
Aurelien Lac

25 commits

NohTow

7 commits

merve

1 commits

lightonai/MonoQwen2-VL-v0.1

Model

48

stars

33

commits

6

repos using this model

1

linked in READMEs

Jul 17, 2026

updated

peft
qwen2_vl
reranker
safetensors
sentence-transformers
vidore
visual-document-retrieval
Browse cluster: Visual Document Retrieval & ColPali

README

MonoQwen2-VL-v0.1

Model Overview

The MonoQwen2-VL-v0.1 is a multimodal reranker finetuned with LoRA from Qwen2-VL-2B, optimized for asserting pointwise image-query relevance using the MonoT5 objective. That is, given a couple of image and query fed into the prompt of the VLM, the model is tasked to generate "True" if the image is relevant to the query and "False" otherwise. During inference, a relevancy score can then be obtained by comparing the logits of the two tokens and this score can effectively be used to rerank the candidates generated by a first-stage retriever (such as DSE or ColPali) or filter them using a threshold.

The ColPali train set was used to train this model with negatives mined using DSE.

How to Use the Model

Using Sentence Transformers

Install Sentence Transformers (>= 5.6.0) with the image extra, plus peft for the LoRA adapter (requires transformers v5):

pip install "sentence-transformers[image]" peft

The reranking prompt used during training is baked into the bundled chat template, and the returned scores are the P("True") relevance probabilities:

from sentence_transformers import CrossEncoder

model = CrossEncoder("lightonai/MonoQwen2-VL-v0.1")

# Pass document page images as PIL.Image, a local file path, or an image URL
query = "What is the variable represented on the y-axis of the graph?"
pages = [
    "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc1.jpg",
    "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc2.jpg",
]

scores = model.predict([(query, page) for page in pages])
print(scores)
# [0.61328125 0.29492188]

rankings = model.rank(query, pages)
print(rankings)
# [{'corpus_id': 0, 'score': 0.61328125}, {'corpus_id': 1, 'score': 0.29492188}]

The LoRA adapter is loaded onto the Qwen/Qwen2-VL-2B-Instruct base model automatically. The model loads in bfloat16 by default, but you can pass model_kwargs={"dtype": torch.float32} to CrossEncoder(...) for full fp32 precision. Text passages can also be passed as documents in place of images.

Using Transformers

Below is a quick example to rerank a single image against a user query using this model:

import requests
import torch
from PIL import Image
from transformers import AutoProcessor, Qwen2VLForConditionalGeneration

# Load processor and model
processor = AutoProcessor.from_pretrained("Qwen/Qwen2-VL-2B-Instruct")
model = Qwen2VLForConditionalGeneration.from_pretrained(
    "lightonai/MonoQwen2-VL-v0.1",
    device_map="auto",
    # attn_implementation="flash_attention_2",
    # torch_dtype=torch.bfloat16,
)

# Define query and load image
query = "What is the variable represented on the y-axis of the graph?"
image_url = "https://huggingface.co/lightonai/MonoQwen2-VL-v0.1/resolve/main/assets/doc1.jpg"
image = Image.open(requests.get(image_url, stream=True).raw)

# Construct the prompt and prepare input
prompt = (
    "Assert the relevance of the previous image document to the following query, "
    "answer True or False. The query is: {query}"
).format(query=query)

messages = [
    {
        "role": "user",
        "content": [
            {"type": "image", "image": image},
            {"type": "text", "text": prompt},
        ],
    }
]

# Apply chat template and tokenize
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = processor(text=text, images=image, return_tensors="pt").to("cuda")

# Run inference to obtain logits
with torch.no_grad():
    outputs = model(**inputs)
    logits_for_last_token = outputs.logits[:, -1, :]

# Convert tokens and calculate relevance score
true_token_id = processor.tokenizer.convert_tokens_to_ids("True")
false_token_id = processor.tokenizer.convert_tokens_to_ids("False")
relevance_score = torch.softmax(logits_for_last_token[:, [true_token_id, false_token_id]], dim=-1)

# Extract and display probabilities
true_prob = relevance_score[0, 0].item()
false_prob = relevance_score[0, 1].item()

print(f"True probability: {true_prob:.4f}, False probability: {false_prob:.4f}")
# True probability: 0.6133, False probability: 0.3848

This example demonstrates how to use the model to assess the relevance of an image with respect to a query. It outputs the probability that the image is relevant ("True") or not relevant ("False").

Note: this example requires peft to be installed in your environment (pip install peft). If you don't want to use peft, you can use model.load_adapter on the original Qwen2-VL-2B model.

Performance Metrics

The model has been evaluated on ViDoRe Benchmark, by retrieving 10 elements with MrLight_dse-qwen2-2b-mrl-v1 and reranking them. The table below summarizes its ndcg@5 scores:

DatasetMrLight_dse-qwen2-2b-mrl-v1MonoQwen2-VL-v0.1 reranking
vidore/arxivqa_test_subsampled85.689.0
vidore/docvqa_test_subsampled57.159.7
vidore/infovqa_test_subsampled88.193.2
vidore/tabfquad_test_subsampled93.196.0
vidore/shiftproject_test82.093.0
vidore/syntheticDocQA_artificial_intelligence_test97.5100.0
vidore/syntheticDocQA_energy_test92.997.7
vidore/syntheticDocQA_government_reports_test96.098.0
vidore/syntheticDocQA_healthcare_industry_test96.499.3
vidore/tatdqa_test69.479.0
Mean85.890.5

License

This LoRA model is licensed under the Apache 2.0 license.

Citation

If you find the model useful, consider citing our work:

@misc{MonoQwen,
  title={MonoQwen: Visual Document Reranking},
  author={Chaffin, Antoine and Lac, Aurélien},
  url={https://huggingface.co/lightonai/MonoQwen2-VL-v0.1},
  year={2024}
}

Contributors

AL
Aurelien Lac

25 commits

NohTow

7 commits

merve

1 commits