CodeGoat24/VideoDPO

Dataset

0

stars

8

commits

2

linked in READMEs

Sep 3, 2025

updated

README

VideoDPO

Dataset Summary

This dataset is derived from VideoDPO for our UnifiedReward-7B training.

For further details, please refer to the following resources:

Citation

@article{unifiedreward-think,
  title={Unified multimodal chain-of-thought reward model through reinforcement fine-tuning},
  author={Wang, Yibin and Li, Zhimin and Zang, Yuhang and Wang, Chunyu and Lu, Qinglin and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2505.03318},
  year={2025}
}

@article{unifiedreward,
  title={Unified reward model for multimodal understanding and generation},
  author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2503.05236},
  year={2025}
}

Contributors

CodeGoat24

8 commits

CodeGoat24/VideoDPO

Dataset

0

stars

8

commits

2

linked in READMEs

Sep 3, 2025

updated

README

VideoDPO

Dataset Summary

This dataset is derived from VideoDPO for our UnifiedReward-7B training.

For further details, please refer to the following resources:

Citation

@article{unifiedreward-think,
  title={Unified multimodal chain-of-thought reward model through reinforcement fine-tuning},
  author={Wang, Yibin and Li, Zhimin and Zang, Yuhang and Wang, Chunyu and Lu, Qinglin and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2505.03318},
  year={2025}
}

@article{unifiedreward,
  title={Unified reward model for multimodal understanding and generation},
  author={Wang, Yibin and Zang, Yuhang and Li, Hao and Jin, Cheng and Wang, Jiaqi},
  journal={arXiv preprint arXiv:2503.05236},
  year={2025}
}

Contributors

CodeGoat24

8 commits