This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to improve the reasoning abilities of multi-modal large language models (MLLMs). It contains image-text pairs, where each pair consists of an image, a problem described in text, and the corresponding answer. The dataset is designed for unsupervised post-training of MLLMs.
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to improve the reasoning abilities of multi-modal large language models (MLLMs). It contains image-text pairs, where each pair consists of an image, a problem described in text, and the corresponding answer. The dataset is designed for unsupervised post-training of MLLMs.
The dataset contains 2101 examples in the training split. Each example includes:
images: An image related to the problem.problem: The problem statement (text).answer: The correct answer (text).4 commits
1 commits
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to improve the reasoning abilities of multi-modal large language models (MLLMs). It contains image-text pairs, where each pair consists of an image, a problem described in text, and the corresponding answer. The dataset is designed for unsupervised post-training of MLLMs.
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to improve the reasoning abilities of multi-modal large language models (MLLMs). It contains image-text pairs, where each pair consists of an image, a problem described in text, and the corresponding answer. The dataset is designed for unsupervised post-training of MLLMs.
The dataset contains 2101 examples in the training split. Each example includes:
images: An image related to the problem.problem: The problem statement (text).answer: The correct answer (text).4 commits
1 commits