This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista a
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista and We-Math.
The dataset consists of image-text pairs where images is the image input, problem is the problem description (text), and answer is the corresponding answer (text). The train split contains 8031 examples.
4 commits
1 commits
This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista a
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista and We-Math.
The dataset consists of image-text pairs where images is the image input, problem is the problem description (text), and answer is the corresponding answer (text). The train split contains 8031 examples.
4 commits
1 commits