CaraJ/MME-CoT

Dataset

MME-CoT 🔥🕵️: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

23

18 commits

6 linked in READMEs

updated Mar 19, 2025

See the code

README

MME-CoT 🔥🕵️: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Multimodal CoT Visual Reasoning MME-CoT

OpenAI o1 Kimi k1.5 GPT-4o

Official repository for "MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency".

🌟 For more details, please refer to the project page with dataset exploration and visualization tools.

[🍓Project Page] [📖 Paper] [🧑‍💻 Code] [📊 Huggingface Dataset] [🏆 Leaderboard] [👁️ Visualization]

💥 News

  • [2025.02.14] 🌟 We are very proud to launch MME-CoT, the first-ever comprehensive CoT evaluation benchmark of LMMs in Visual Reasoning! We release the arxiv paper and all data samples in huggingface dataset.

👀 About MME-CoT

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation.

In this paper, we introduce MME-CoT, a specialized benchmark evaluating the CoT reasoning performance of LMMs, spanning six domains: math, science, OCR, logic, space-time, and general scenes. As the first comprehensive study in this area, we propose a thorough evaluation suite incorporating three novel metrics that assess the reasoning quality, robustness, and efficiency at a fine-grained level.


Leveraging curated high-quality data and a unique evaluation strategy, we conduct an in-depth analysis of state-of-the-art LMMs, uncovering several key insights: (1) Models with reflection mechanism demonstrate a superior CoT quality, with Kimi k1.5 outperforming GPT-4o and demonstrating the highest quality results; (2) CoT prompting often degrades LMM performance on perception-heavy tasks, suggesting a potentially harmful overthinking behavior; (3) Although the CoT quality is high, LMMs with reflection exhibit significant inefficiency in both normal response and self-correction phases. We hope MME-CoT serves as a foundation for advancing multimodal reasoning in LMMs.


💡 Illustration of our CoT Quality Evaluation Strategy


💪 Illustration of our CoT Robustness Evaluation Strategy


⚡️ Illustration of our CoT Efficiency Evaluation Strategy


Evaluation

🏆 Leaderboard

Contributing to the Leaderboard

🚨 The Leaderboard is continuously being updated, welcoming the contribution of your excellent LMMs!

To contribute your model to the leaderboard, please email the prediction files of four tasks to 📫jdzcarr7@gmail.com.

Data Usage

We release the MME-CoT data and evaluation prompts for benchmarking on the leaderboard.

You can download the dataset from the 🤗 Huggingface by the following command (make sure that you have installed related packages):

from datasets import load_dataset

dataset = load_dataset("CaraJ/MME-CoT")

Explore our additional research on Vision-Language Large Models:

Contributors

CaraJ

17 commits

nielsr

1 commits

CaraJ/MME-CoT

Dataset

MME-CoT 🔥🕵️: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

23

18 commits

6 linked in READMEs

updated Mar 19, 2025

See the code

README

MME-CoT 🔥🕵️: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Multimodal CoT Visual Reasoning MME-CoT

OpenAI o1 Kimi k1.5 GPT-4o

Official repository for "MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency".

🌟 For more details, please refer to the project page with dataset exploration and visualization tools.

[🍓Project Page] [📖 Paper] [🧑‍💻 Code] [📊 Huggingface Dataset] [🏆 Leaderboard] [👁️ Visualization]

💥 News

  • [2025.02.14] 🌟 We are very proud to launch MME-CoT, the first-ever comprehensive CoT evaluation benchmark of LMMs in Visual Reasoning! We release the arxiv paper and all data samples in huggingface dataset.

👀 About MME-CoT

Answering questions with Chain-of-Thought (CoT) has significantly enhanced the reasoning capabilities of Large Language Models (LLMs), yet its impact on Large Multimodal Models (LMMs) still lacks a systematic assessment and in-depth investigation.

In this paper, we introduce MME-CoT, a specialized benchmark evaluating the CoT reasoning performance of LMMs, spanning six domains: math, science, OCR, logic, space-time, and general scenes. As the first comprehensive study in this area, we propose a thorough evaluation suite incorporating three novel metrics that assess the reasoning quality, robustness, and efficiency at a fine-grained level.


Leveraging curated high-quality data and a unique evaluation strategy, we conduct an in-depth analysis of state-of-the-art LMMs, uncovering several key insights: (1) Models with reflection mechanism demonstrate a superior CoT quality, with Kimi k1.5 outperforming GPT-4o and demonstrating the highest quality results; (2) CoT prompting often degrades LMM performance on perception-heavy tasks, suggesting a potentially harmful overthinking behavior; (3) Although the CoT quality is high, LMMs with reflection exhibit significant inefficiency in both normal response and self-correction phases. We hope MME-CoT serves as a foundation for advancing multimodal reasoning in LMMs.


💡 Illustration of our CoT Quality Evaluation Strategy


💪 Illustration of our CoT Robustness Evaluation Strategy


⚡️ Illustration of our CoT Efficiency Evaluation Strategy


Evaluation

🏆 Leaderboard

Contributing to the Leaderboard

🚨 The Leaderboard is continuously being updated, welcoming the contribution of your excellent LMMs!

To contribute your model to the leaderboard, please email the prediction files of four tasks to 📫jdzcarr7@gmail.com.

Data Usage

We release the MME-CoT data and evaluation prompts for benchmarking on the leaderboard.

You can download the dataset from the 🤗 Huggingface by the following command (make sure that you have installed related packages):

from datasets import load_dataset

dataset = load_dataset("CaraJ/MME-CoT")

Explore our additional research on Vision-Language Large Models:

Contributors

CaraJ

17 commits

nielsr

1 commits