OpenMOSS-Team/Game-RL-Qwen2.5-VL-7B

Model

3

stars

16

commits

4

linked in READMEs

Jul 27, 2025

updated

conversational
endpoints_compatible
image-text-to-text
qwen2_5_vl
safetensors
text-generation-inference
transformers

README

This model (GameQA-Qwen2.5-VL-7B) results from training Qwen2.5-VL-7B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).

Evaluation Results on General Vision BenchMarks

(The inference and evaluation configurations were unified across both the original open-source models and our trained models.)

It's also found that getting trained on 5k samples from our GameQA dataset can lead to better results than on 8k samples from MAVIS and on multimodal-open-r1-8k-verified.

Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning

This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple cutting-edge open-source models exhibit significantly enhanced out-of-domain generalization.

[πŸ“– Paper] [πŸ’» Code] [πŸ€— GameQA-140K Dataset] [πŸ€— GameQA-5K Dataset] [πŸ€— GameQA-InternVL3-8B ] [πŸ€— GameQA-InternVL2.5-8B ] [πŸ€— GameQA-Qwen2.5-VL-7B] [πŸ€— GameQA-LLaVA-OV-7B ]

Code: https://github.com/tongjingqi/Code2Logic

News

  • We've open-sourced the three models trained with GRPO on GameQA on Huggingface.

Contributors

lkdhy

14 commits

Gabriel166

1 commits

nielsr

1 commits

OpenMOSS-Team/Game-RL-Qwen2.5-VL-7B

Model

3

stars

16

commits

4

linked in READMEs

Jul 27, 2025

updated

conversational
endpoints_compatible
image-text-to-text
qwen2_5_vl
safetensors
text-generation-inference
transformers

README

This model (GameQA-Qwen2.5-VL-7B) results from training Qwen2.5-VL-7B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).

Evaluation Results on General Vision BenchMarks

(The inference and evaluation configurations were unified across both the original open-source models and our trained models.)

It's also found that getting trained on 5k samples from our GameQA dataset can lead to better results than on 8k samples from MAVIS and on multimodal-open-r1-8k-verified.

Code2Logic: Game-Code-Driven Data Synthesis for Enhancing VLMs General Reasoning

This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple cutting-edge open-source models exhibit significantly enhanced out-of-domain generalization.

[πŸ“– Paper] [πŸ’» Code] [πŸ€— GameQA-140K Dataset] [πŸ€— GameQA-5K Dataset] [πŸ€— GameQA-InternVL3-8B ] [πŸ€— GameQA-InternVL2.5-8B ] [πŸ€— GameQA-Qwen2.5-VL-7B] [πŸ€— GameQA-LLaVA-OV-7B ]

Code: https://github.com/tongjingqi/Code2Logic

News

  • We've open-sourced the three models trained with GRPO on GameQA on Huggingface.

Contributors

lkdhy

14 commits

Gabriel166

1 commits

nielsr

1 commits