haoxiangzhao12138/PLUME

[ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding

Python

26

6 commits

updated Apr 29, 2026

See the code

README

PLUME: Latent Reasoning Based Universal Multimodal Embedding

🏑 Project Page | πŸ“„ Paper | πŸ€— Model | πŸ€— Training Data | πŸ€— Eval Data

PLUME is a latent reasoning framework for universal multimodal embedding (UME). It replaces explicit chain-of-thought (CoT) generation with a short autoregressive rollout of continuous latent states, combined with a semantic-anchor-guided transition adapter (Latent MoE) and a progressive explicit-to-latent curriculum. Built on Qwen2-VL-2B, PLUME achieves 61.6 on the 78-task MMEB-v2 benchmark while delivering over 30x faster inference compared to explicit-CoT methods.

Accuracy-Efficiency Tradeoff

This repository is the official implementation of the paper PLUME: Latent Reasoning Based Universal Multimodal Embedding.

πŸ“° News

  • [2026/04] Paper released on arXiv.
  • [2025] Code and model weights released.

πŸ“ TODO

  • Code Released: Training and evaluation pipeline.
  • Model Released: Pre-trained PLUME-Qwen2-VL-2B weights.
  • Paper: arXiv.

πŸ’‘ Highlights

  • Replaces hundreds of explicit reasoning tokens with only 8 latent steps, delivering 30.3x faster inference
  • 61.6 overall on the 78-task MMEB-v2 benchmark, surpassing UME-R1 (60.1) and VLM2Vec-V2 (58.0)
  • Curriculum-based latent reasoning -- gradually replaces chain-of-thought text with continuous thought (<ct>) tokens across training stages
  • Contrastive learning -- cross-device bidirectional contrastive loss for query/positive embedding alignment
  • Latent MoE -- Mixture-of-Experts transition layer with 4 routed experts + shared expert in the latent reasoning loop
  • Multi-modal evaluation -- MMEB image / video / visdoc evaluation with latent-MoE support

πŸ—οΈ Method

PLUME Method Overview

Overview of PLUME. The bottom panel illustrates the latent rollout process. The top-left panel expands the semantic-anchor-guided transition adapter with shared and specialized experts. The top-right panel shows the progressive explicit-to-latent curriculum.

🎞️ Results on MMEB-v2

All methods share the same Qwen2-VL-2B backbone.

ModelImageVideoVisDocAll
VLM2Vec-V264.934.965.458.0
UME-R166.642.263.960.1
PLUME66.344.167.561.6

Per-task Performance Comparison

Per-task performance comparison on MMEB-v2. PLUME consistently outperforms UME-R1 and single-pass baselines across most sub-tasks.

πŸ“¦ Model & Data

ResourceLinkDescription
Model weightsCUDAOUTOFMEMORY/PLUME-Qwen2-VL-2BPre-trained PLUME model (Qwen2-VL-2B + Latent MoE)
Training annotationszhibinlan/UME-sft-trainJSONL annotations for training
Images & eval dataTIGER-Lab/MMEB-V2Multi-modal images for training and evaluation

πŸ”§ Getting Started

πŸ“ Directory Structure

PLUME/
β”œβ”€β”€ README.md
β”œβ”€β”€ plume/                              # Python package
β”‚   β”œβ”€β”€ train/                          #   Training
β”‚   β”‚   β”œβ”€β”€ train_plume.py              #     Core trainer (PlumeTrainer)
β”‚   β”‚   β”œβ”€β”€ train_plume_gc.py           #     Gradient-checkpointing variant
β”‚   β”‚   β”œβ”€β”€ latent_moe.py              #     Latent MoE transition module
β”‚   β”‚   └── argument.py                #     Dataclass argument definitions
β”‚   β”œβ”€β”€ data/                           #   Data processing
β”‚   β”‚   β”œβ”€β”€ data_plume.py              #     Dataset, collator, curriculum sampler
β”‚   β”‚   └── rope2d.py                  #     2D/3D RoPE index computation
β”‚   └── eval/                           #   Evaluation analysis tools
β”‚       β”œβ”€β”€ compare_eval_results.py
β”‚       β”œβ”€β”€ compute_mmeb_image_hit1_avg.py
β”‚       β”œβ”€β”€ compute_mmeb_video_hit1_avg.py
β”‚       └── analyze_max_tokens.py
β”œβ”€β”€ configs/
β”‚   β”œβ”€β”€ deepspeed/                      #   DeepSpeed configs (zero2/zero3/offload)
β”‚   └── eval/                           #   MMEB dataset configs (image/video/visdoc)
β”œβ”€β”€ scripts/                            #   Shell launchers
β”‚   β”œβ”€β”€ train_singlenode.sh
β”‚   β”œβ”€β”€ train_multinode.sh
β”‚   β”œβ”€β”€ launch_multinode.sh
β”‚   └── eval_plume_moe.sh
β”œβ”€β”€ VLM2Vec/                            #   Bundled MMEB eval engine with eval_twomode.py
β”œβ”€β”€ tools/
β”‚   └── check_image.py
└── docs/                               #   Documentation
    β”œβ”€β”€ environment.md
    β”œβ”€β”€ training.md
    └── evaluation.md

🫑 Acknowledgements

Many thanks to the code bases from VLM2Vec and Qwen2-VL.

Citation

If you use this code for your research or project, please cite:

@misc{he2026plumelatentreasoningbased,
      title={PLUME: Latent Reasoning Based Universal Multimodal Embedding},
      author={Chenwei He and Xiangzhao Hao and Tianyu Yang and Yuxiang Ma and Yuheng Jia and Lingxiang Wu and Chaoyang Zhao and Haiyun Guo and Jinqiao Wang},
      year={2026},
      eprint={2604.02073},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2604.02073},
}

License

See the repository root for license information.

haoxiangzhao12138/PLUME

[ACMMM 2026] PLUME: Latent Reasoning Based Universal Multimodal Embedding

Python

26

6 commits

updated Apr 29, 2026

See the code

README

PLUME: Latent Reasoning Based Universal Multimodal Embedding

🏑 Project Page | πŸ“„ Paper | πŸ€— Model | πŸ€— Training Data | πŸ€— Eval Data

PLUME is a latent reasoning framework for universal multimodal embedding (UME). It replaces explicit chain-of-thought (CoT) generation with a short autoregressive rollout of continuous latent states, combined with a semantic-anchor-guided transition adapter (Latent MoE) and a progressive explicit-to-latent curriculum. Built on Qwen2-VL-2B, PLUME achieves 61.6 on the 78-task MMEB-v2 benchmark while delivering over 30x faster inference compared to explicit-CoT methods.

Accuracy-Efficiency Tradeoff

This repository is the official implementation of the paper PLUME: Latent Reasoning Based Universal Multimodal Embedding.

πŸ“° News

  • [2026/04] Paper released on arXiv.
  • [2025] Code and model weights released.

πŸ“ TODO

  • Code Released: Training and evaluation pipeline.
  • Model Released: Pre-trained PLUME-Qwen2-VL-2B weights.
  • Paper: arXiv.

πŸ’‘ Highlights

  • Replaces hundreds of explicit reasoning tokens with only 8 latent steps, delivering 30.3x faster inference
  • 61.6 overall on the 78-task MMEB-v2 benchmark, surpassing UME-R1 (60.1) and VLM2Vec-V2 (58.0)
  • Curriculum-based latent reasoning -- gradually replaces chain-of-thought text with continuous thought (<ct>) tokens across training stages
  • Contrastive learning -- cross-device bidirectional contrastive loss for query/positive embedding alignment
  • Latent MoE -- Mixture-of-Experts transition layer with 4 routed experts + shared expert in the latent reasoning loop
  • Multi-modal evaluation -- MMEB image / video / visdoc evaluation with latent-MoE support

πŸ—οΈ Method

PLUME Method Overview

Overview of PLUME. The bottom panel illustrates the latent rollout process. The top-left panel expands the semantic-anchor-guided transition adapter with shared and specialized experts. The top-right panel shows the progressive explicit-to-latent curriculum.

🎞️ Results on MMEB-v2

All methods share the same Qwen2-VL-2B backbone.

ModelImageVideoVisDocAll
VLM2Vec-V264.934.965.458.0
UME-R166.642.263.960.1
PLUME66.344.167.561.6

Per-task Performance Comparison

Per-task performance comparison on MMEB-v2. PLUME consistently outperforms UME-R1 and single-pass baselines across most sub-tasks.

πŸ“¦ Model & Data

ResourceLinkDescription
Model weightsCUDAOUTOFMEMORY/PLUME-Qwen2-VL-2BPre-trained PLUME model (Qwen2-VL-2B + Latent MoE)
Training annotationszhibinlan/UME-sft-trainJSONL annotations for training
Images & eval dataTIGER-Lab/MMEB-V2Multi-modal images for training and evaluation

πŸ”§ Getting Started

πŸ“ Directory Structure

PLUME/
β”œβ”€β”€ README.md
β”œβ”€β”€ plume/                              # Python package
β”‚   β”œβ”€β”€ train/                          #   Training
β”‚   β”‚   β”œβ”€β”€ train_plume.py              #     Core trainer (PlumeTrainer)
β”‚   β”‚   β”œβ”€β”€ train_plume_gc.py           #     Gradient-checkpointing variant
β”‚   β”‚   β”œβ”€β”€ latent_moe.py              #     Latent MoE transition module
β”‚   β”‚   └── argument.py                #     Dataclass argument definitions
β”‚   β”œβ”€β”€ data/                           #   Data processing
β”‚   β”‚   β”œβ”€β”€ data_plume.py              #     Dataset, collator, curriculum sampler
β”‚   β”‚   └── rope2d.py                  #     2D/3D RoPE index computation
β”‚   └── eval/                           #   Evaluation analysis tools
β”‚       β”œβ”€β”€ compare_eval_results.py
β”‚       β”œβ”€β”€ compute_mmeb_image_hit1_avg.py
β”‚       β”œβ”€β”€ compute_mmeb_video_hit1_avg.py
β”‚       └── analyze_max_tokens.py
β”œβ”€β”€ configs/
β”‚   β”œβ”€β”€ deepspeed/                      #   DeepSpeed configs (zero2/zero3/offload)
β”‚   └── eval/                           #   MMEB dataset configs (image/video/visdoc)
β”œβ”€β”€ scripts/                            #   Shell launchers
β”‚   β”œβ”€β”€ train_singlenode.sh
β”‚   β”œβ”€β”€ train_multinode.sh
β”‚   β”œβ”€β”€ launch_multinode.sh
β”‚   └── eval_plume_moe.sh
β”œβ”€β”€ VLM2Vec/                            #   Bundled MMEB eval engine with eval_twomode.py
β”œβ”€β”€ tools/
β”‚   └── check_image.py
└── docs/                               #   Documentation
    β”œβ”€β”€ environment.md
    β”œβ”€β”€ training.md
    └── evaluation.md

🫑 Acknowledgements

Many thanks to the code bases from VLM2Vec and Qwen2-VL.

Citation

If you use this code for your research or project, please cite:

@misc{he2026plumelatentreasoningbased,
      title={PLUME: Latent Reasoning Based Universal Multimodal Embedding},
      author={Chenwei He and Xiangzhao Hao and Tianyu Yang and Yuxiang Ma and Yuheng Jia and Lingxiang Wu and Chaoyang Zhao and Haiyun Guo and Jinqiao Wang},
      year={2026},
      eprint={2604.02073},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2604.02073},
}

License

See the repository root for license information.

Languages

Python

97.2%

Shell

2.8%