GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs). Generated using the innovative Code2Logic framework, it leverages game code to synthesize high-quality visual-language Chain-of-Thought (CoT) data. The dataset addresses the scarcity of multimodal reasoning data, critical for advancing complex multi-step reasoning in VLMs. Each sample includes visual game state, targeted question, original analysis, augmented step-by-step reasoning (refinement) and final answer, derived from the logical structures inherent in game code.
π Paper: Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMsβ General Reasoning
π Website: https://iclr26-game-rl.github.io
π» Code: https://github.com/tongjingqi/Game-RL
For a quick preview of the dataset, GameQA-data_studio-preview.parquet contains 300 sampled entries from the training set. This file is optimized for online viewing in tools like Data Studio and includes embedded image data.
For the full training dataset, please download games_data.json.
For the full test dataset, please download games_data_test.json.
Associated image files are available in games_images.zip and games_images_test.zip respectively.
| Attribute | Description |
|---|---|
| Size | ~140,000 question-answer pairs (126,760 training, 15,047 testing). |
| Diversity | 30 unique games, 158 distinct tasks covering various cognitive skills. |
| Game Categories | - 3D Spatial Perception and Understanding - Pattern Recognition and Matching - Multi-step Reasoning - Strategic Planning |
| Format | Visual Question Answering (VQA): - Game state image - Targeted question - Step-by-step reasoning - Final answer |
| Question Types | - Multiple-choice (typically 7-8 options) - Fill-in-the-blank (e.g., numbers, coordinates) |
| Challenging | Difficult for SOTA VLMs (<50% accuracy on test set). |
| Scalability & Cost | Code2Logic enables massive-scale generation with minimal cost after initial setup. |
| Difficulty Levels | - Plot Level (Image Complexity): Easy, Medium, Hard - QA Level (Task Complexity): Easy, Medium, Hard |
If you find our work or dataset useful, we would appreciate it if you could cite our work:
@article{tong2025game,
title={Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning},
author={Tong, Jingqi and Tang, Jixin and Li, Hangcheng and Mou, Yurong and Zhang, Ming and Zhao, Jun and Wen, Yanbo and Song, Fan and Zhan, Jiahao and Lu, Yuyang and others},
journal={arXiv preprint arXiv:2505.13886},
year={2025}
}
GameQA is a large-scale, diverse, and challenging multimodal reasoning dataset designed to enhance the general reasoning capabilities of Vision Language Models (VLMs). Generated using the innovative Code2Logic framework, it leverages game code to synthesize high-quality visual-language Chain-of-Thought (CoT) data. The dataset addresses the scarcity of multimodal reasoning data, critical for advancing complex multi-step reasoning in VLMs. Each sample includes visual game state, targeted question, original analysis, augmented step-by-step reasoning (refinement) and final answer, derived from the logical structures inherent in game code.
π Paper: Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMsβ General Reasoning
π Website: https://iclr26-game-rl.github.io
π» Code: https://github.com/tongjingqi/Game-RL
For a quick preview of the dataset, GameQA-data_studio-preview.parquet contains 300 sampled entries from the training set. This file is optimized for online viewing in tools like Data Studio and includes embedded image data.
For the full training dataset, please download games_data.json.
For the full test dataset, please download games_data_test.json.
Associated image files are available in games_images.zip and games_images_test.zip respectively.
| Attribute | Description |
|---|---|
| Size | ~140,000 question-answer pairs (126,760 training, 15,047 testing). |
| Diversity | 30 unique games, 158 distinct tasks covering various cognitive skills. |
| Game Categories | - 3D Spatial Perception and Understanding - Pattern Recognition and Matching - Multi-step Reasoning - Strategic Planning |
| Format | Visual Question Answering (VQA): - Game state image - Targeted question - Step-by-step reasoning - Final answer |
| Question Types | - Multiple-choice (typically 7-8 options) - Fill-in-the-blank (e.g., numbers, coordinates) |
| Challenging | Difficult for SOTA VLMs (<50% accuracy on test set). |
| Scalability & Cost | Code2Logic enables massive-scale generation with minimal cost after initial setup. |
| Difficulty Levels | - Plot Level (Image Complexity): Easy, Medium, Hard - QA Level (Task Complexity): Easy, Medium, Hard |
If you find our work or dataset useful, we would appreciate it if you could cite our work:
@article{tong2025game,
title={Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning},
author={Tong, Jingqi and Tang, Jixin and Li, Hangcheng and Mou, Yurong and Zhang, Ming and Zhao, Jun and Wen, Yanbo and Song, Fan and Zhan, Jiahao and Lu, Yuyang and others},
journal={arXiv preprint arXiv:2505.13886},
year={2025}
}