kaichengyang0828/MLLM-for-Universal-Embedding-Awesome

MLLM-based Universal Embedding Learning Paper List!

11

0 commits

updated Apr 3, 2026

See the code

README

📖 Universal Multimodal Embedding Papers

A curated collection of cutting-edge research papers on universal multimodal embedding learning.

📚 Paper List

TitleYearPaperGithub
E5-V: Universal embeddings with multimodal large language models2024PaperGithub
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks2024PaperGithub
|MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs2024PaperHuggingface
Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models2024PaperHuggingface
VladVA: Discriminative Fine-tuning of LVLMs2024PaperUnreleased
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant2024PaperGithub
Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval2025PaperGithub
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning2025PaperGithub
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning2024PaperGithub
UniME: Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs2025PaperGithub
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining2025PaperGithub
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval2025PaperUnreleased
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data2025PaperGithub
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval2025PaperGithub
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying2025PaperGithub
Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models2025PaperUnreleased
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment2025PaperGithub
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval2025PaperGithub
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval2025PaperHuggingface
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning2025PaperUnreleased
UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings2025PaperGithub
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs2025PaperHuggingface
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment2025PaperUnreleased
From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model2025PaperGithub
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM2025PaperUnreleased
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval2025PaperGithub
VIRTUE: Visual-Interactive Text-Image Universal Embedder2025PaperUnreleased
Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents2025PaperGithub
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning2025PaperUnreleased
SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model2025PaperUnreleased
Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework2025PaperUnreleased
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation2025PaperGithub
Think Then Embed: Generative Context Improves Multimodal Embedding2025PaperUnreleased
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction2025PaperUnreleased
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning2025PaperGithub
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval2025PaperUnreleased
RzenEmbed: Towards Comprehensive Multimodal Retrieval2025PaperHuggingface
Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation2025PaperUnreleased
Scaling Language-Centric Omnimodal Representation Learning2025PaperGithub
Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding2025PaperUnreleased
MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding2025PaperUnreleased
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding2025PaperUnreleased
ReMatch: Boosting Representation through Matching for Multimodal Retrieval2025PaperGithub
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning2025PaperUnreleased
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval2025PaperGithub
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization2025PaperGithub
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings2026PaperHuggingface
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking2026PaperHuggingface
ObjEmbed: Towards Universal Multimodal Object Embeddings2026PaperGithub
|V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval2026PaperGithub
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs2026PaperUnreleased

📝 Contributing

Continuously updating the list. Feel free to contribute by:

  • Adding new relevant papers
  • Reporting issues or broken links

🎙️ Who We Are?

We are the UniME team, passionate about learning more robust representations through MLLM to empower various downstream tasks. If you are interest in universal multimodal embedding learning or would like to collaborate with us, feel free to reach out: kaichengyang0828@gmail.com

kaichengyang0828/MLLM-for-Universal-Embedding-Awesome

MLLM-based Universal Embedding Learning Paper List!

11

0 commits

updated Apr 3, 2026

See the code

README

📖 Universal Multimodal Embedding Papers

A curated collection of cutting-edge research papers on universal multimodal embedding learning.

📚 Paper List

TitleYearPaperGithub
E5-V: Universal embeddings with multimodal large language models2024PaperGithub
VLM2Vec: Training Vision-Language Models for Massive Multimodal Embedding Tasks2024PaperGithub
|MM-Embed: Universal Multimodal Retrieval with Multimodal LLMs2024PaperHuggingface
Bridging Modalities: Improving Universal Multimodal Retrieval by Multimodal Large Language Models2024PaperHuggingface
VladVA: Discriminative Fine-tuning of LVLMs2024PaperUnreleased
LamRA: Large Multimodal Model as Your Advanced Retrieval Assistant2024PaperGithub
Any Information Is Just Worth One Single Screenshot: Unifying Search With Visualized Information Retrieval2025PaperGithub
LLaVE: Large Language and Vision Embedding Models with Hardness-Weighted Contrastive Learning2025PaperGithub
CAFe: Unifying Representation and Generation with Contrastive-Autoregressive Finetuning2024PaperGithub
UniME: Breaking the Modality Barrier: Universal Embedding Learning with Multimodal LLMs2025PaperGithub
Breaking the Batch Barrier (B3) of Contrastive Learning via Smart Batch Mining2025PaperGithub
Highlighting What Matters: Promptable Embeddings for Attribute-Focused Image Retrieval2025PaperUnreleased
mmE5: Improving Multimodal Multilingual Embeddings via High-quality Synthetic Data2025PaperGithub
Modality Curation: Building Universal Embeddings for Advanced Multimodal Information Retrieval2025PaperGithub
Improve Multi-Modal Embedding Learning via Explicit Hard Negative Gradient Amplifying2025PaperGithub
Incorporating Dense Knowledge Alignment into Unified Multimodal Representation Models2025PaperUnreleased
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment2025PaperGithub
MM-R5: MultiModal Reasoning-Enhanced ReRanker via Reinforcement Learning for Document Retrieval2025PaperGithub
jina-embeddings-v4: Universal Embeddings for Multimodal Multilingual Retrieval2025PaperHuggingface
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning2025PaperUnreleased
UniMoCo: Unified Modality Completion for Robust Multi-Modal Embeddings2025PaperGithub
U-MARVEL: Unveiling Key Factors for Universal Multimodal Retrieval via Embedding Learning with MLLMs2025PaperHuggingface
Context-Adaptive Multi-Prompt Embedding with Large Language Models for Vision-Language Alignment2025PaperUnreleased
From Generator to Embedder: Harnessing Innate Abilities of Multimodal LLMs via Building Zero-Shot Discriminative Embedding Model2025PaperGithub
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM2025PaperUnreleased
Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval2025PaperGithub
VIRTUE: Visual-Interactive Text-Image Universal Embedder2025PaperUnreleased
Vlm2vec-v2: Advancing multimodal embedding for videos, images, and visual documents2025PaperGithub
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning2025PaperUnreleased
SAIL-Embedding Technical Report: Omni-modal Embedding Foundation Model2025PaperUnreleased
Multi-Modal Multi-Task Unified Embedding Model (M3T-UEM): A Task-Adaptive Representation Learning Framework2025PaperUnreleased
Towards Mixed-Modal Retrieval for Universal Retrieval-Augmented Generation2025PaperGithub
Think Then Embed: Generative Context Improves Multimodal Embedding2025PaperUnreleased
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction2025PaperUnreleased
UniME-V2: MLLM-as-a-Judge for Universal Multimodal Embedding Learning2025PaperGithub
MCA: Modality Composition Awareness for Robust Composed Multimodal Retrieval2025PaperUnreleased
RzenEmbed: Towards Comprehensive Multimodal Retrieval2025PaperHuggingface
Let Multimodal Embedders Learn When to Augment Query via Adaptive Query Augmentation2025PaperUnreleased
Scaling Language-Centric Omnimodal Representation Learning2025PaperGithub
Compression then Matching: An Efficient Pre-training Paradigm for Multimodal Embedding2025PaperUnreleased
MOON: Generative MLLM-based Multimodal Representation Learning for E-commerce Product Understanding2025PaperUnreleased
MOON2.0: Dynamic Modality-balanced Multimodal Representation Learning for E-commerce Product Understanding2025PaperUnreleased
ReMatch: Boosting Representation through Matching for Multimodal Retrieval2025PaperGithub
FANoise: Singular Value-Adaptive Noise Modulation for Robust Multimodal Representation Learning2025PaperUnreleased
Reasoning Guided Embeddings: Leveraging MLLM Reasoning for Improved Multimodal Retrieval2025PaperGithub
Explore More, Learn Better: Parallel MLLM Embeddings under Mutual Information Minimization2025PaperGithub
e5-omni: Explicit Cross-modal Alignment for Omni-modal Embeddings2026PaperHuggingface
Qwen3-VL-Embedding and Qwen3-VL-Reranker: A Unified Framework for State-of-the-Art Multimodal Retrieval and Ranking2026PaperHuggingface
ObjEmbed: Towards Universal Multimodal Object Embeddings2026PaperGithub
|V-Retrver: Evidence-Driven Agentic Reasoning for Universal Multimodal Retrieval2026PaperGithub
Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs2026PaperUnreleased

📝 Contributing

Continuously updating the list. Feel free to contribute by:

  • Adding new relevant papers
  • Reporting issues or broken links

🎙️ Who We Are?

We are the UniME team, passionate about learning more robust representations through MLLM to empower various downstream tasks. If you are interest in universal multimodal embedding learning or would like to collaborate with us, feel free to reach out: kaichengyang0828@gmail.com