[ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding
Python
26
6 commits
updated Apr 29, 2026
π‘ Project Page | π Paper | π€ Model | π€ Training Data | π€ Eval Data
PLUME is a latent reasoning framework for universal multimodal embedding (UME). It replaces explicit chain-of-thought (CoT) generation with a short autoregressive rollout of continuous latent states, combined with a semantic-anchor-guided transition adapter (Latent MoE) and a progressive explicit-to-latent curriculum. Built on Qwen2-VL-2B, PLUME achieves 61.6 on the 78-task MMEB-v2 benchmark while delivering over 30x faster inference compared to explicit-CoT methods.
This repository is the official implementation of the paper PLUME: Latent Reasoning Based Universal Multimodal Embedding.
<ct>) tokens across training stages
Overview of PLUME. The bottom panel illustrates the latent rollout process. The top-left panel expands the semantic-anchor-guided transition adapter with shared and specialized experts. The top-right panel shows the progressive explicit-to-latent curriculum.
All methods share the same Qwen2-VL-2B backbone.
| Model | Image | Video | VisDoc | All |
|---|---|---|---|---|
| VLM2Vec-V2 | 64.9 | 34.9 | 65.4 | 58.0 |
| UME-R1 | 66.6 | 42.2 | 63.9 | 60.1 |
| PLUME | 66.3 | 44.1 | 67.5 | 61.6 |
Per-task performance comparison on MMEB-v2. PLUME consistently outperforms UME-R1 and single-pass baselines across most sub-tasks.
| Resource | Link | Description |
|---|---|---|
| Model weights | CUDAOUTOFMEMORY/PLUME-Qwen2-VL-2B | Pre-trained PLUME model (Qwen2-VL-2B + Latent MoE) |
| Training annotations | zhibinlan/UME-sft-train | JSONL annotations for training |
| Images & eval data | TIGER-Lab/MMEB-V2 | Multi-modal images for training and evaluation |
PLUME/
βββ README.md
βββ plume/ # Python package
β βββ train/ # Training
β β βββ train_plume.py # Core trainer (PlumeTrainer)
β β βββ train_plume_gc.py # Gradient-checkpointing variant
β β βββ latent_moe.py # Latent MoE transition module
β β βββ argument.py # Dataclass argument definitions
β βββ data/ # Data processing
β β βββ data_plume.py # Dataset, collator, curriculum sampler
β β βββ rope2d.py # 2D/3D RoPE index computation
β βββ eval/ # Evaluation analysis tools
β βββ compare_eval_results.py
β βββ compute_mmeb_image_hit1_avg.py
β βββ compute_mmeb_video_hit1_avg.py
β βββ analyze_max_tokens.py
βββ configs/
β βββ deepspeed/ # DeepSpeed configs (zero2/zero3/offload)
β βββ eval/ # MMEB dataset configs (image/video/visdoc)
βββ scripts/ # Shell launchers
β βββ train_singlenode.sh
β βββ train_multinode.sh
β βββ launch_multinode.sh
β βββ eval_plume_moe.sh
βββ VLM2Vec/ # Bundled MMEB eval engine with eval_twomode.py
βββ tools/
β βββ check_image.py
βββ docs/ # Documentation
βββ environment.md
βββ training.md
βββ evaluation.md
Many thanks to the code bases from VLM2Vec and Qwen2-VL.
If you use this code for your research or project, please cite:
@misc{he2026plumelatentreasoningbased,
title={PLUME: Latent Reasoning Based Universal Multimodal Embedding},
author={Chenwei He and Xiangzhao Hao and Tianyu Yang and Yuxiang Ma and Yuheng Jia and Lingxiang Wu and Chaoyang Zhao and Haiyun Guo and Jinqiao Wang},
year={2026},
eprint={2604.02073},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.02073},
}
See the repository root for license information.
Python
97.2%
Shell
2.8%
[ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding
Python
26
6 commits
updated Apr 29, 2026
π‘ Project Page | π Paper | π€ Model | π€ Training Data | π€ Eval Data
PLUME is a latent reasoning framework for universal multimodal embedding (UME). It replaces explicit chain-of-thought (CoT) generation with a short autoregressive rollout of continuous latent states, combined with a semantic-anchor-guided transition adapter (Latent MoE) and a progressive explicit-to-latent curriculum. Built on Qwen2-VL-2B, PLUME achieves 61.6 on the 78-task MMEB-v2 benchmark while delivering over 30x faster inference compared to explicit-CoT methods.
This repository is the official implementation of the paper PLUME: Latent Reasoning Based Universal Multimodal Embedding.
<ct>) tokens across training stages
Overview of PLUME. The bottom panel illustrates the latent rollout process. The top-left panel expands the semantic-anchor-guided transition adapter with shared and specialized experts. The top-right panel shows the progressive explicit-to-latent curriculum.
All methods share the same Qwen2-VL-2B backbone.
| Model | Image | Video | VisDoc | All |
|---|---|---|---|---|
| VLM2Vec-V2 | 64.9 | 34.9 | 65.4 | 58.0 |
| UME-R1 | 66.6 | 42.2 | 63.9 | 60.1 |
| PLUME | 66.3 | 44.1 | 67.5 | 61.6 |
Per-task performance comparison on MMEB-v2. PLUME consistently outperforms UME-R1 and single-pass baselines across most sub-tasks.
| Resource | Link | Description |
|---|---|---|
| Model weights | CUDAOUTOFMEMORY/PLUME-Qwen2-VL-2B | Pre-trained PLUME model (Qwen2-VL-2B + Latent MoE) |
| Training annotations | zhibinlan/UME-sft-train | JSONL annotations for training |
| Images & eval data | TIGER-Lab/MMEB-V2 | Multi-modal images for training and evaluation |
PLUME/
βββ README.md
βββ plume/ # Python package
β βββ train/ # Training
β β βββ train_plume.py # Core trainer (PlumeTrainer)
β β βββ train_plume_gc.py # Gradient-checkpointing variant
β β βββ latent_moe.py # Latent MoE transition module
β β βββ argument.py # Dataclass argument definitions
β βββ data/ # Data processing
β β βββ data_plume.py # Dataset, collator, curriculum sampler
β β βββ rope2d.py # 2D/3D RoPE index computation
β βββ eval/ # Evaluation analysis tools
β βββ compare_eval_results.py
β βββ compute_mmeb_image_hit1_avg.py
β βββ compute_mmeb_video_hit1_avg.py
β βββ analyze_max_tokens.py
βββ configs/
β βββ deepspeed/ # DeepSpeed configs (zero2/zero3/offload)
β βββ eval/ # MMEB dataset configs (image/video/visdoc)
βββ scripts/ # Shell launchers
β βββ train_singlenode.sh
β βββ train_multinode.sh
β βββ launch_multinode.sh
β βββ eval_plume_moe.sh
βββ VLM2Vec/ # Bundled MMEB eval engine with eval_twomode.py
βββ tools/
β βββ check_image.py
βββ docs/ # Documentation
βββ environment.md
βββ training.md
βββ evaluation.md
Many thanks to the code bases from VLM2Vec and Qwen2-VL.
If you use this code for your research or project, please cite:
@misc{he2026plumelatentreasoningbased,
title={PLUME: Latent Reasoning Based Universal Multimodal Embedding},
author={Chenwei He and Xiangzhao Hao and Tianyu Yang and Yuxiang Ma and Yuheng Jia and Lingxiang Wu and Chaoyang Zhao and Haiyun Guo and Jinqiao Wang},
year={2026},
eprint={2604.02073},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2604.02073},
}
See the repository root for license information.
Python
97.2%
Shell
2.8%