Universal multimodal retrieval (UMR) addresses complex retrieval tasks involving diverse modalities for both queries and candidates. Despite the success of state-of-the-art methods based on multimodal large language models (MLLMs) using contrastive learning principles, the mechanisms underlying their retrieval capabilities remain largely unexplored. This gap potentially leads to suboptimal performance and limited generalization ability.
In this study, we systematically analyze the key factors driving effective embedding learning for UMR using MLLMs. We implement a general MLLM-based embedding learning pipeline and investigate contributors to high-performing universal retrieval systems. Our analysis covers various aspects of embedding generation and training strategies, including progressive transition, hard negative mining, and re-ranker distillation. Our findings reveal that often-overlooked factors can significantly impact model performance.
Building on these insights, we introduce U-MARVEL (Universal Multimodal Retrieval via Embedding Learning), a unified framework that outperforms state-of-the-art competitors on the M-BEIR benchmark in supervised settings and demonstrates strong zero-shot performance on tasks such as composed image retrieval and text-to-video retrieval. These results highlight the generalization potential of our framework across various embedding-based retrieval tasks, providing valuable insights for future research.
├── checkpoints
│ ├── hf_models
│ │ └── Qwen2-VL-7B-Instruct
│ │ └── Qwen3-VL-4B-Instruct
│ └── U-MARVEL-Qwen2VL-7B-Instruct
│ └── U-MARVEL-Qwen3VL-4B-Instruct
To install requirements:
pip install -r requirements_qwen2_vl.txt
To train the model(s) in the paper, run this command: Shown here is the configuration for Qwen2-VL-7B. The training scripts for Qwen3-VL-4B and Qwen3-VL-8B are analogous and require similar modifications.
Download Qwen2-VL-7B and place it in ./checkpoints/hf_models/Qwen2-VL-7B-Instruct
For NLI dataset, please refer to link
For multimodal instruction tuning datset, please refer to M-BEIR
After downloading all of them, organize the data as follows in ./data
├── M-BEIR
├── nli_for_simcse.csv
├── rerank_data_for_training
├── flickr
├── coco
├── sharegpt4v
├── Urban1K
├── circo
├── genecis
├── vist
├── visdial
├── ccneg
├── sugar-crepe
├── MSVD
└── msrvtt
get_10percent_training_data.ipynb
get_training_data_local_format.ipynb
Run the
get_10percent_training_data.ipynbnotebook to extract 10% of the query data from the training set for subsequent distillation.Run the
get_training_data_local_format.ipynbnotebook to partition the M-BEIR dataset's queries and candidates into 16 tasks, which will be used for subsequent hard negative mining.
Fine-tuning with NLI dataset
python scripts/umarvel/nli/vtools_umarvel_progressive_transition_nli.py
sh scripts/merge_lora.sh
To fine-tune the model using the CC3M dataset, we executed the
scripts/umarvel/nli/vtools_umarvel_progressive_transition_nli.pyscript , resulting in the model namedqwen2-vl-7b_umarvel_progressive_transition_nli. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Fine-tuning with CC3M dataset
python scripts/umarvel/cc3m_sharegpt4v_laion/vtools_finetune_cc3m_llm.py
sh scripts/merge_lora.sh
To fine-tune the model using the CC3M dataset, we executed the
scripts/umarvel/cc3m/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_cc3m.pyscript. This resulted in the model namedqwen2-vl-7b_umarvel_progressive_transition_cc3m. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Fine-tuning with M-BEIR dataset
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_m-beir.py
sh scripts/merge_lora.sh
To fine-tune the model using the M-BEIR dataset, we executed the
scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_m-beir.pyscript. This resulted in the model namedqwen2-vl-7b_umarvel_progressive_transition_m-beir. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Get training data (full set) with local type negative samples for training point-wise model
python scripts/umarvel_rank/vtools_get_train_data_from_eval_train_data_local.py
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_local.py
Get training data (full set) with global type negative samples for train hard negative mining model
sh scripts/umarvel_rank/get_train_data_from_eval_train_data_global.sh
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_global.py
Fine-tuning with hard negative mining
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining.py
sh scripts/merge_lora.sh
Train point-wise rerank model
python scripts/umarvel_rank/vtools_train_rerank_multi_nodes_only_pointwise.py
Get top-100 negative samples for queries (local and global versions)
# Get top 100 negative samples (local version)
python scripts/umarvel_rank/vtools_get_train_data_from_eval_train_10percent_data_local.py
# Merge training data (local version)
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_local_10percent.py
# Get top 100 negative samples (global version)
sh scripts/umarvel_rank/get_train_data_from_eval_train_10percent_data_global.sh
# Merge training data (global version)
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_global_10percent.py
Process using the rerank model to obtain point-wise scores
# Process local version to obtain point-wise scores
python scripts/eval/rerank/vtools_eval_train_data_rerank_mbeir_pointwise_local_10percent.py
# Process global version to obtain point-wise scores
python scripts/eval/rerank/vtools_eval_train_data_rerank_mbeir_pointwise_global_10perncent.py
Obtain distillation scores
python scripts/eval/rerank/get_mbeir_base_rerank_pointwise_score_local.py
python scripts/eval/rerank/get_mbeir_base_rerank_pointwise_score_global.py
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_distillation.py
Compare distilled model with negative sample model
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.py # Compare distilled model with negative sample model
The script
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.pyis used to compare the performance of a distilled model with that of a model trained using hard negative samples.
To obtain the rerank+ model, execute the following commands:
python scripts/umarvel_rank/u-marvel+/vtools_get_train_data_from_eval_train_data_local-umarvel+.py
python scripts/umarvel_rank/u-marvel+/vtools_merge_train_data_from_eval_train_local_umarvel+.py
python scripts/umarvel_rank/u-marvel+/vtools_train_rerank_multi_nodes_only_pointwise-umarvel+.py
To evaluate our model on M-BEIR, run:
Evaluate models at each stage
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_progressive_transition_m-beir.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_hard_negative_mining.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_distillation.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.py
sh scripts/eval/bash/eval_mbeir_global.sh
Zero-shot evaluation
python scripts/eval/bash/vtools_eval_zero-shot.py
Evaluate rerank (local version)
python scripts/eval/rerank/umarvel+/vtools_eval_rerank_mbeir_pointwise_local_umarvel+.py
python eval/rerank/mbeir_rerank_pointwise_local_for_weight.py
Evaluate rerank (global version)
python scripts/eval/rerank/umarvel+/vtools_eval_rerank_mbeir_pointwise_global_umarvel+.py
python eval/rerank/mbeir_rerank_pointwise_global_for_weight.py
Zero-shot evaluation for umarvel+ model
sh scripts/eval/bash/eval_rerank_zeroshot.sh
python scripts/eval/bash/get_rerank_results_zeroshot.sh
The proposed U-MARVEL framework establishes new state-of-the-art performance across both single-model architectures and recall-then-rerank approaches on M-BEIR benchmark.
Many thanks to the code bases from LamRA .
If you use this code for your research or project, please cite:
@inproceedings{li2026umarvel,
title={U-{MARVEL}: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with {MLLM}s},
author={Xiaojie Li and Chu Li and Shi-Zhe Chen and Xi Chen},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}
7 commits
Python
92.0%
Jupyter Notebook
4.1%
Shell
3.9%
Universal multimodal retrieval (UMR) addresses complex retrieval tasks involving diverse modalities for both queries and candidates. Despite the success of state-of-the-art methods based on multimodal large language models (MLLMs) using contrastive learning principles, the mechanisms underlying their retrieval capabilities remain largely unexplored. This gap potentially leads to suboptimal performance and limited generalization ability.
In this study, we systematically analyze the key factors driving effective embedding learning for UMR using MLLMs. We implement a general MLLM-based embedding learning pipeline and investigate contributors to high-performing universal retrieval systems. Our analysis covers various aspects of embedding generation and training strategies, including progressive transition, hard negative mining, and re-ranker distillation. Our findings reveal that often-overlooked factors can significantly impact model performance.
Building on these insights, we introduce U-MARVEL (Universal Multimodal Retrieval via Embedding Learning), a unified framework that outperforms state-of-the-art competitors on the M-BEIR benchmark in supervised settings and demonstrates strong zero-shot performance on tasks such as composed image retrieval and text-to-video retrieval. These results highlight the generalization potential of our framework across various embedding-based retrieval tasks, providing valuable insights for future research.
├── checkpoints
│ ├── hf_models
│ │ └── Qwen2-VL-7B-Instruct
│ │ └── Qwen3-VL-4B-Instruct
│ └── U-MARVEL-Qwen2VL-7B-Instruct
│ └── U-MARVEL-Qwen3VL-4B-Instruct
To install requirements:
pip install -r requirements_qwen2_vl.txt
To train the model(s) in the paper, run this command: Shown here is the configuration for Qwen2-VL-7B. The training scripts for Qwen3-VL-4B and Qwen3-VL-8B are analogous and require similar modifications.
Download Qwen2-VL-7B and place it in ./checkpoints/hf_models/Qwen2-VL-7B-Instruct
For NLI dataset, please refer to link
For multimodal instruction tuning datset, please refer to M-BEIR
After downloading all of them, organize the data as follows in ./data
├── M-BEIR
├── nli_for_simcse.csv
├── rerank_data_for_training
├── flickr
├── coco
├── sharegpt4v
├── Urban1K
├── circo
├── genecis
├── vist
├── visdial
├── ccneg
├── sugar-crepe
├── MSVD
└── msrvtt
get_10percent_training_data.ipynb
get_training_data_local_format.ipynb
Run the
get_10percent_training_data.ipynbnotebook to extract 10% of the query data from the training set for subsequent distillation.Run the
get_training_data_local_format.ipynbnotebook to partition the M-BEIR dataset's queries and candidates into 16 tasks, which will be used for subsequent hard negative mining.
Fine-tuning with NLI dataset
python scripts/umarvel/nli/vtools_umarvel_progressive_transition_nli.py
sh scripts/merge_lora.sh
To fine-tune the model using the CC3M dataset, we executed the
scripts/umarvel/nli/vtools_umarvel_progressive_transition_nli.pyscript , resulting in the model namedqwen2-vl-7b_umarvel_progressive_transition_nli. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Fine-tuning with CC3M dataset
python scripts/umarvel/cc3m_sharegpt4v_laion/vtools_finetune_cc3m_llm.py
sh scripts/merge_lora.sh
To fine-tune the model using the CC3M dataset, we executed the
scripts/umarvel/cc3m/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_cc3m.pyscript. This resulted in the model namedqwen2-vl-7b_umarvel_progressive_transition_cc3m. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Fine-tuning with M-BEIR dataset
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_m-beir.py
sh scripts/merge_lora.sh
To fine-tune the model using the M-BEIR dataset, we executed the
scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_progressive_transition_m-beir.pyscript. This resulted in the model namedqwen2-vl-7b_umarvel_progressive_transition_m-beir. Subsequently, we adjusted the parameters and merged the LoRA with the base model by running themerge_lora.shscript.
Get training data (full set) with local type negative samples for training point-wise model
python scripts/umarvel_rank/vtools_get_train_data_from_eval_train_data_local.py
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_local.py
Get training data (full set) with global type negative samples for train hard negative mining model
sh scripts/umarvel_rank/get_train_data_from_eval_train_data_global.sh
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_global.py
Fine-tuning with hard negative mining
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining.py
sh scripts/merge_lora.sh
Train point-wise rerank model
python scripts/umarvel_rank/vtools_train_rerank_multi_nodes_only_pointwise.py
Get top-100 negative samples for queries (local and global versions)
# Get top 100 negative samples (local version)
python scripts/umarvel_rank/vtools_get_train_data_from_eval_train_10percent_data_local.py
# Merge training data (local version)
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_local_10percent.py
# Get top 100 negative samples (global version)
sh scripts/umarvel_rank/get_train_data_from_eval_train_10percent_data_global.sh
# Merge training data (global version)
python scripts/umarvel_rank/vtools_merge_train_data_from_eval_train_global_10percent.py
Process using the rerank model to obtain point-wise scores
# Process local version to obtain point-wise scores
python scripts/eval/rerank/vtools_eval_train_data_rerank_mbeir_pointwise_local_10percent.py
# Process global version to obtain point-wise scores
python scripts/eval/rerank/vtools_eval_train_data_rerank_mbeir_pointwise_global_10perncent.py
Obtain distillation scores
python scripts/eval/rerank/get_mbeir_base_rerank_pointwise_score_local.py
python scripts/eval/rerank/get_mbeir_base_rerank_pointwise_score_global.py
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_distillation.py
Compare distilled model with negative sample model
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.py # Compare distilled model with negative sample model
The script
python scripts/umarvel/mbeir/vtools_finetune_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.pyis used to compare the performance of a distilled model with that of a model trained using hard negative samples.
To obtain the rerank+ model, execute the following commands:
python scripts/umarvel_rank/u-marvel+/vtools_get_train_data_from_eval_train_data_local-umarvel+.py
python scripts/umarvel_rank/u-marvel+/vtools_merge_train_data_from_eval_train_local_umarvel+.py
python scripts/umarvel_rank/u-marvel+/vtools_train_rerank_multi_nodes_only_pointwise-umarvel+.py
To evaluate our model on M-BEIR, run:
Evaluate models at each stage
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_progressive_transition_m-beir.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_hard_negative_mining.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_distillation.py
python scripts/eval/fast_eval/vtools_eval_mbeir_local_fast_qwen2-vl-7b_umarvel_hard_negative_mining-continue_hard.py
sh scripts/eval/bash/eval_mbeir_global.sh
Zero-shot evaluation
python scripts/eval/bash/vtools_eval_zero-shot.py
Evaluate rerank (local version)
python scripts/eval/rerank/umarvel+/vtools_eval_rerank_mbeir_pointwise_local_umarvel+.py
python eval/rerank/mbeir_rerank_pointwise_local_for_weight.py
Evaluate rerank (global version)
python scripts/eval/rerank/umarvel+/vtools_eval_rerank_mbeir_pointwise_global_umarvel+.py
python eval/rerank/mbeir_rerank_pointwise_global_for_weight.py
Zero-shot evaluation for umarvel+ model
sh scripts/eval/bash/eval_rerank_zeroshot.sh
python scripts/eval/bash/get_rerank_results_zeroshot.sh
The proposed U-MARVEL framework establishes new state-of-the-art performance across both single-model architectures and recall-then-rerank approaches on M-BEIR benchmark.
Many thanks to the code bases from LamRA .
If you use this code for your research or project, please cite:
@inproceedings{li2026umarvel,
title={U-{MARVEL}: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with {MLLM}s},
author={Xiaojie Li and Chu Li and Shi-Zhe Chen and Xi Chen},
booktitle={The Fourteenth International Conference on Learning Representations},
year={2026}
}
7 commits
Python
92.0%
Jupyter Notebook
4.1%
Shell
3.9%