WaltonFuture/GeoQA-8K-direct-synthesizing

Dataset

This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista a

1

5 commits

1 linked in READMEs

updated Jun 2, 2025

See the code

README

This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista and We-Math.

The dataset consists of image-text pairs where images is the image input, problem is the problem description (text), and answer is the corresponding answer (text). The train split contains 8031 examples.

Contributors

WaltonFuture

4 commits

nielsr

1 commits

WaltonFuture/GeoQA-8K-direct-synthesizing

Dataset

This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista a

1

5 commits

1 linked in READMEs

updated Jun 2, 2025

See the code

README

This dataset supports the unsupervised post-training of multi-modal large language models (MLLMs) as described in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO. It's designed to enable continual self-improvement without external supervision, using a self-rewarding mechanism based on majority voting over multiple sampled responses. The dataset is used to improve the reasoning ability of MLLMs, as demonstrated by significant improvements on benchmarks like MathVista and We-Math.

The dataset consists of image-text pairs where images is the image input, problem is the problem description (text), and answer is the corresponding answer (text). The train split contains 8031 examples.

Contributors

WaltonFuture

4 commits

nielsr

1 commits