A curated collection of cutting-edge research papers on universal multimodal embedding learning.
| Title | Year | Paper | Github |
|---|---|---|---|
| E5-V: Universal embeddings with multimodal large language models | 2024 | Paper | Github |
| VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks | 2024 | Paper | Github |
| |MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs | 2024 | Paper | Huggingface |
| Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models | 2024 | Paper | Huggingface |
| VladVA: Discriminative Fine-tuning of LVLMs | 2024 | Paper | Unreleased |
| LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant | 2024 | Paper | Github |
| Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval | 2025 | Paper | Github |
| LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning | 2025 | Paper | Github |
| CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning | 2024 | Paper | Github |
| UniME: Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs | 2025 | Paper | Github |
| Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining | 2025 | Paper | Github |
| Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval | 2025 | Paper | Unreleased |
| mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data | 2025 | Paper | Github |
| Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval | 2025 | Paper | Github |
| Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying | 2025 | Paper | Github |
| Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models | 2025 | Paper | Unreleased |
| Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment | 2025 | Paper | Github |
| MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval | 2025 | Paper | Github |
| jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval | 2025 | Paper | Huggingface |
| PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning | 2025 | Paper | Unreleased |
| UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings | 2025 | Paper | Github |
| U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs | 2025 | Paper | Huggingface |
| Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment | 2025 | Paper | Unreleased |
| From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model | 2025 | Paper | Github |
| WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM | 2025 | Paper | Unreleased |
| Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval | 2025 | Paper | Github |
| VIRTUE: Visual-Interactive Text-Image Universal Embedder | 2025 | Paper | Unreleased |
| Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents | 2025 | Paper | Github |
| CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning | 2025 | Paper | Unreleased |
| SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model | 2025 | Paper | Unreleased |
| Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework | 2025 | Paper | Unreleased |
| Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation | 2025 | Paper | Github |
| Think Then Embed: Generative Context Improves Multimodal Embedding | 2025 | Paper | Unreleased |
| MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction | 2025 | Paper | Unreleased |
| UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning | 2025 | Paper | Github |
| MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval | 2025 | Paper | Unreleased |
| RzenEmbed: Towards Comprehensive Multimodal Retrieval | 2025 | Paper | Huggingface |
| Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation | 2025 | Paper | Unreleased |
| Scaling Language-Centric Omnimodal Representation Learning | 2025 | Paper | Github |
| Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding | 2025 | Paper | Unreleased |
| MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding | 2025 | Paper | Unreleased |
| MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding | 2025 | Paper | Unreleased |
| ReMatch: Boosting Representation through Matching for Multimodal Retrieval | 2025 | Paper | Github |
| FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning | 2025 | Paper | Unreleased |
| Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval | 2025 | Paper | Github |
| Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization | 2025 | Paper | Github |
| e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings | 2026 | Paper | Huggingface |
| Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking | 2026 | Paper | Huggingface |
| ObjEmbed: Towards Universal Multimodal Object Embeddings | 2026 | Paper | Github |
| |V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval | 2026 | Paper | Github |
| Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs | 2026 | Paper | Unreleased |
Continuously updating the list. Feel free to contribute by:
We are the UniME team, passionate about learning more robust representations through MLLM to empower various downstream tasks. If you are interest in universal multimodal embedding learning or would like to collaborate with us, feel free to reach out: kaichengyang0828@gmail.com
A curated collection of cutting-edge research papers on universal multimodal embedding learning.
| Title | Year | Paper | Github |
|---|---|---|---|
| E5-V: Universal embeddings with multimodal large language models | 2024 | Paper | Github |
| VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks | 2024 | Paper | Github |
| |MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs | 2024 | Paper | Huggingface |
| Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models | 2024 | Paper | Huggingface |
| VladVA: Discriminative Fine-tuning of LVLMs | 2024 | Paper | Unreleased |
| LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant | 2024 | Paper | Github |
| Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval | 2025 | Paper | Github |
| LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning | 2025 | Paper | Github |
| CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning | 2024 | Paper | Github |
| UniME: Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs | 2025 | Paper | Github |
| Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining | 2025 | Paper | Github |
| Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval | 2025 | Paper | Unreleased |
| mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data | 2025 | Paper | Github |
| Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval | 2025 | Paper | Github |
| Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying | 2025 | Paper | Github |
| Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models | 2025 | Paper | Unreleased |
| Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment | 2025 | Paper | Github |
| MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval | 2025 | Paper | Github |
| jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval | 2025 | Paper | Huggingface |
| PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning | 2025 | Paper | Unreleased |
| UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings | 2025 | Paper | Github |
| U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs | 2025 | Paper | Huggingface |
| Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment | 2025 | Paper | Unreleased |
| From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model | 2025 | Paper | Github |
| WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM | 2025 | Paper | Unreleased |
| Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval | 2025 | Paper | Github |
| VIRTUE: Visual-Interactive Text-Image Universal Embedder | 2025 | Paper | Unreleased |
| Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents | 2025 | Paper | Github |
| CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning | 2025 | Paper | Unreleased |
| SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model | 2025 | Paper | Unreleased |
| Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework | 2025 | Paper | Unreleased |
| Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation | 2025 | Paper | Github |
| Think Then Embed: Generative Context Improves Multimodal Embedding | 2025 | Paper | Unreleased |
| MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction | 2025 | Paper | Unreleased |
| UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning | 2025 | Paper | Github |
| MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval | 2025 | Paper | Unreleased |
| RzenEmbed: Towards Comprehensive Multimodal Retrieval | 2025 | Paper | Huggingface |
| Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation | 2025 | Paper | Unreleased |
| Scaling Language-Centric Omnimodal Representation Learning | 2025 | Paper | Github |
| Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding | 2025 | Paper | Unreleased |
| MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding | 2025 | Paper | Unreleased |
| MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding | 2025 | Paper | Unreleased |
| ReMatch: Boosting Representation through Matching for Multimodal Retrieval | 2025 | Paper | Github |
| FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning | 2025 | Paper | Unreleased |
| Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval | 2025 | Paper | Github |
| Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization | 2025 | Paper | Github |
| e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings | 2026 | Paper | Huggingface |
| Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking | 2026 | Paper | Huggingface |
| ObjEmbed: Towards Universal Multimodal Object Embeddings | 2026 | Paper | Github |
| |V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval | 2026 | Paper | Github |
| Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs | 2026 | Paper | Unreleased |
Continuously updating the list. Feel free to contribute by:
We are the UniME team, passionate about learning more robust representations through MLLM to empower various downstream tasks. If you are interest in universal multimodal embedding learning or would like to collaborate with us, feel free to reach out: kaichengyang0828@gmail.com