WaltonFuture/MMR1-direct-synthesizing

Dataset

This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.

1

5 commits

1 linked in READMEs

updated Jun 2, 2025

See the code

README

This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.

The dataset contains the following:

  • images: Image data.
  • problem: Problem description (text).
  • answer: Answer (text).

The train split contains 5782 examples.

multimodal
reasoning
unsupervised-learning

Contributors

WaltonFuture

4 commits

nielsr

1 commits

WaltonFuture/MMR1-direct-synthesizing

Dataset

This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.

1

5 commits

1 linked in READMEs

updated Jun 2, 2025

See the code

README

This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.

The dataset contains the following:

  • images: Image data.
  • problem: Problem description (text).
  • answer: Answer (text).

The train split contains 5782 examples.

multimodal
reasoning
unsupervised-learning

Contributors

WaltonFuture

4 commits

nielsr

1 commits