The dataset is constructed based on the training set of MMEB-V2. We employ the pure-thinking model GLM-4.1V-Thinking to generate chain-of-thought (CoT) rationales for both the query and the target in each pair. To ensure data quality, we filter out samples that meet any of the following criteria:
This filtering process results in a final set of 1.46 million cold-start SFT pairs.
If you find our work useful, please consider citing it.
@article{lan2025ume,
title={UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings},
author={Lan, Zhibin and Niu, Liqiang and Meng, Fandong and Zhou, Jie and Su, Jinsong},
journal={arXiv preprint arXiv:2511.00405},
year={2025}
}
The dataset is constructed based on the training set of MMEB-V2. We employ the pure-thinking model GLM-4.1V-Thinking to generate chain-of-thought (CoT) rationales for both the query and the target in each pair. To ensure data quality, we filter out samples that meet any of the following criteria:
This filtering process results in a final set of 1.46 million cold-start SFT pairs.
If you find our work useful, please consider citing it.
@article{lan2025ume,
title={UME-R1: Exploring Reasoning-Driven Generative Multimodal Embeddings},
author={Lan, Zhibin and Niu, Liqiang and Meng, Fandong and Zhou, Jie and Su, Jinsong},
journal={arXiv preprint arXiv:2511.00405},
year={2025}
}