This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.
The dataset contains the following:
images: Image data.problem: Problem description (text).answer: Answer (text).The train split contains 5782 examples.
4 commits
1 commits
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.
1
5 commits
1 linked in READMEs
updated Jun 2, 2025
This dataset is used in the paper Unsupervised Post-Training for Multi-Modal LLM Reasoning via GRPO to evaluate the performance of an unsupervised post-training framework (MM-UPT) for multi-modal LLMs. The dataset contains image-text pairs where the text represents a problem or question, and the corresponding answer. It is used to demonstrate the ability of MM-UPT to improve the reasoning capabilities of the Qwen2.5-VL-7B model without relying on any external supervised data.
The dataset contains the following:
images: Image data.problem: Problem description (text).answer: Answer (text).The train split contains 5782 examples.
4 commits
1 commits