Paper: MM-IFEngine: Towards Multimodal Instruction Following
Github: SYuan03/MM-IFEngine
Project Page: syuan03.github.io/MM-IFEngine/
MM-IFEval Evaluation: Using VLMEvalKit
😊 This is the official repo of MM-IFEngine datasets in MM-IFEngine: Towards Multimodal Instruction Following
🚀 We include both the SFT and DPO data in this repo as the v1 dataset (generated mainly by InternVL2.5-78B and Qwen2-VL-7B), which we used to train the model described in our paper.
💖 [2025.9.16 Update] We have released the v2 dataset (annotated mainly by GPT-4o), feel free to use it!
Using ShareGPT format from LLaMA-Factory
@article{ding2025mm,
title={MM-IFEngine: Towards Multimodal Instruction Following},
author={Ding, Shengyuan and Wu, Shenxi and Zhao, Xiangyu and Zang, Yuhang and Duan, Haodong and Dong, Xiaoyi and Zhang, Pan and Cao, Yuhang and Lin, Dahua and Wang, Jiaqi},
journal={arXiv preprint arXiv:2504.07957},
year={2025}
}
9 commits
Paper: MM-IFEngine: Towards Multimodal Instruction Following
Github: SYuan03/MM-IFEngine
Project Page: syuan03.github.io/MM-IFEngine/
MM-IFEval Evaluation: Using VLMEvalKit
😊 This is the official repo of MM-IFEngine datasets in MM-IFEngine: Towards Multimodal Instruction Following
🚀 We include both the SFT and DPO data in this repo as the v1 dataset (generated mainly by InternVL2.5-78B and Qwen2-VL-7B), which we used to train the model described in our paper.
💖 [2025.9.16 Update] We have released the v2 dataset (annotated mainly by GPT-4o), feel free to use it!
Using ShareGPT format from LLaMA-Factory
@article{ding2025mm,
title={MM-IFEngine: Towards Multimodal Instruction Following},
author={Ding, Shengyuan and Wu, Shenxi and Zhao, Xiangyu and Zang, Yuhang and Duan, Haodong and Dong, Xiaoyi and Zhang, Pan and Cao, Yuhang and Lin, Dahua and Wang, Jiaqi},
journal={arXiv preprint arXiv:2504.07957},
year={2025}
}
9 commits