This model (GameQA-InternVL3-8B) results from training InternVL3-8B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).

(The inference and evaluation configurations were unified across both the original open-source models and our trained models.)
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple cutting-edge open-source models exhibit significantly enhanced out-of-domain generalization.
[📖 Paper] [🤗 GameQA-140K Dataset] [🤗 GameQA-5K Dataset] [🤗 GameQA-InternVL3-8B ] [🤗 GameQA-Qwen2.5-VL-7B] [🤗 GameQA-LLaVA-OV-7B ]

10 commits
1 commits
This model (GameQA-InternVL3-8B) results from training InternVL3-8B with GRPO solely on our GameQA-5K (sampled from the full GameQA-140K dataset).

(The inference and evaluation configurations were unified across both the original open-source models and our trained models.)
This is the first work, to the best of our knowledge, that leverages game code to synthesize multimodal reasoning data for training VLMs. Furthermore, when trained with a GRPO strategy solely on GameQA (synthesized via our proposed Code2Logic approach), multiple cutting-edge open-source models exhibit significantly enhanced out-of-domain generalization.
[📖 Paper] [🤗 GameQA-140K Dataset] [🤗 GameQA-5K Dataset] [🤗 GameQA-InternVL3-8B ] [🤗 GameQA-Qwen2.5-VL-7B] [🤗 GameQA-LLaVA-OV-7B ]

10 commits
1 commits