[π GitHub] [π Paper] [π Blog] [π€ model] [π€ dataset] [π€ benchmark]
2025/04/11: We release a new version of VisualPRM400K (i.e., VisualPRM400K-v1.1), which includes additional data sources to enhance the data diversity.
VisualPRM400K is a dataset comprising approximately 400K multimodal process supervision data. We generate the data using an automatic data pipeline. The key idea is to estimate the expected accuracy \(mc_i\) of the given step \(s_{\leq i}\) based on Monte Carlo sampling and consider the step correct if \(mc_i>0\). Please see our paper or blog for more details.
NOTE: This dataset is formulated as multi-turn conversation and the expected accuracy \(mc_i\) has been converted into correctness token \(c_i \in {+,-}\). If you want to use the annotations for expected accuracy, please refer to this version.

This project is released under the MIT License. This project uses the pre-trained internlm2_5-7b-chat as a component, which is licensed under the Apache License 2.0.
If you find this project useful in your research, please consider citing:
@article{wang2025visualprm,
title={VisualPRM: An Effective Process Reward Model for Multimodal Reasoning},
author={Wang, Weiyun and Gao, Zhangwei and Chen, Lianjie and Chen, Zhe and Zhu, Jinguo and Zhao, Xiangyu and Liu, Yangzhou and Cao, Yue and Ye, Shenglong and Zhu, Xizhou and others},
journal={arXiv preprint arXiv:2503.10291},
year={2025}
}
7 commits
1 commits
[π GitHub] [π Paper] [π Blog] [π€ model] [π€ dataset] [π€ benchmark]
2025/04/11: We release a new version of VisualPRM400K (i.e., VisualPRM400K-v1.1), which includes additional data sources to enhance the data diversity.
VisualPRM400K is a dataset comprising approximately 400K multimodal process supervision data. We generate the data using an automatic data pipeline. The key idea is to estimate the expected accuracy \(mc_i\) of the given step \(s_{\leq i}\) based on Monte Carlo sampling and consider the step correct if \(mc_i>0\). Please see our paper or blog for more details.
NOTE: This dataset is formulated as multi-turn conversation and the expected accuracy \(mc_i\) has been converted into correctness token \(c_i \in {+,-}\). If you want to use the annotations for expected accuracy, please refer to this version.

This project is released under the MIT License. This project uses the pre-trained internlm2_5-7b-chat as a component, which is licensed under the Apache License 2.0.
If you find this project useful in your research, please consider citing:
@article{wang2025visualprm,
title={VisualPRM: An Effective Process Reward Model for Multimodal Reasoning},
author={Wang, Weiyun and Gao, Zhangwei and Chen, Lianjie and Chen, Zhe and Zhu, Jinguo and Zhao, Xiangyu and Liu, Yangzhou and Cao, Yue and Ye, Shenglong and Zhu, Xizhou and others},
journal={arXiv preprint arXiv:2503.10291},
year={2025}
}
7 commits
1 commits