Paper collections of multi-modal LLM for Math/STEM/Code.
146
71 commits
updated Aug 30, 2026

🔥 Collections of multi-modal LLM for Math/STEM/Code.
MAVIS: Mathematical Visual Instruction Tuning Preprint
Renrui Zhang, Xinyu Wei, Dongzhi Jiang, Yichi Zhang, Ziyu Guo,Chengzhuo Tong, Jiaming Liu, Aojun Zhou, Bin Wei, Shanghang Zhang, Peng Gao, Hongsheng Li.[Paper], 2024.7
COMET: “Cone of experience” enhanced large multimodal model for mathematical problem generation. Preprint
Sannyuya Liu, Jintian Feng, Zongkai Yang, Yawei Luo, Qian Wan, Xiaoxuan Shen, Jianwen Sun. [Paper], 2024.7
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B: A Technical Report. Preprint
Di Zhang, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, Wanli Ouyang. [Paper], 2024.6
Visual SKETCHPAD: Sketching as a Visual Chain of Thought for Multimodal Language Models. Preprint
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, Ranjay Krishna. [Paper], [Code], 2024.6
TextSquare: Scaling up Text-Centric Visual Instruction Tuning. Preprint
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, Can Huang. [Paper], 2024.4
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs. ACL 2024
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou. [Paper], 2024.3
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding. Preprint
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou. [Paper], 2024.3
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning. Preprint
Renqiu Xia, Bo Zhang, Hancheng Ye, Xiangchao Yan, Qi Liu, Hongbin Zhou, Zijun Chen, Min Dou, Botian Shi, Junchi Yan, Yu Qiao. [Paper], 2024.2
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions. Preprint
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, Jun Suzuki. [Paper], 2024.1
G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model. Preprint
Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, Lingpeng Kong. [Paper], [Code], 2023.12
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Models. Preprint
Anwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye, Qinghao Ye, Ming Yan, Chenliang Li, Qi Qian, Ji Zhang, Fei Huang. [Paper], 2023.11
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning. Preprint
Xingchen Zeng, Haichuan Lin, Yilin Ye, Wei Zeng. [Paper], [Code], 2024.7
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning. Preprint
Wenwen Zhuang, Xin Huang, Xiantao Zhang, Jin Zeng. [Paper], 2024.8
Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver. Preprint
Zeren Zhang, Jo-Ku Cheng, Jingyang Deng, Lu Tian, Jinwen Ma, Ziran Qin, Xiaokai Zhang, Na Zhu, and Tuo Leng. [Paper], 2024.9
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends. Preprint
Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, Mostafa M. Aref. [Paper], 2024.10
IMPROVE VISION LANGUAGE MODEL CHAIN-OFTHOUGHT REASONING. Preprint
Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, Yiming Yang. [Paper], 2024.10
R-COT : REVERSE CHAIN-OF-THOUGHT PROBLEM GENERATION FOR GEOMETRIC REASONING IN LARGE MULTIMODAL MODELS. Preprint
Linger Deng, Yuliang Liu, Bohan Li, Dongliang Luo, Liang Wu, Chengquan Zhang, Pengyuan Lyu, Ziyang Zhang, Gang Zhang, Errui Ding, Yingying Zhu, Xiang Bai. [Paper], 2024.10
GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models. Preprint
Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq, Christopher J. Pal. [Paper], 2024.10
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM. NeurIPS 2024 Workshop
Jeongwoo Lee, Kwangsuk Park, Jihyeon Park. [Paper], 2024.11
DIVING INTO SELF-EVOLVING TRAINING FOR MULTIMODAL REASONING. Preprint
Wei Liu, Junlong Li, Xiwen Zhang, Fan Zhou, Yu Cheng, Junxian He. [Paper], 2024.12
Slow Perception: Let’s Perceive Geometric Figures Step-by-step. Preprint
Haoran Wei, Youyang Yin, Yumeng Li, Jia Wang, Liang Zhao, Jianjian Sun, Zheng Ge, Xiangyu Zhang. [Paper], 2024.12
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM. Preprint
Yifan Du, Zikang Liu1, Yifan Li, Wayne Xin Zhao, Yuqi Huo, Bingning Wang, Weipeng ChenZheng Liu, Zhongyuan Wang, Ji-Rong Wen. [Paper], 2025.1
URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics. Preprint
Ruilin Luo, Zhuofan Zheng, Yifan Wang, Yiyao Yu, Xinzhe Ni, Zicheng Lin, Jin Zeng, Yujiu Yang. [Paper], 2025.1
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation. Preprint
Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark. [Paper], 2025.2
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning. Preprint
Ke Wang, Junting Pan, Linda Wei, Aojun Zhou, Weikang Shi, Zimu Lu, Han Xiao, Yunqiao Yang, Houxing Ren, Mingjie Zhan, Hongsheng Li. [Paper], 2025.5
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning. Preprint
Xinyan Chen, Renrui Zhang, Dongzhi Jiang, Aojun Zhou, Shilin Yan, Weifeng Lin, Hongsheng Li. [Paper], 2025.6
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction. Preprint
Chengzhi Xu, Yuyang Wang, Lai Wei, Lichao Sun, Weiran Huang. [Paper], 2025.6
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning. Preprint
Weikang Shi, etc. [Paper], 2025.10
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning. Preprint
Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang. [Paper], 2025.11
Multimodal OCR: Parse Anything from Documents Preprint
Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin, Guangwei Zhao, Hao Liu, et al. [Paper], 2026.3
Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing Preprint
Jiahe Song, Chuang Wang, Yinfan Wang, Hao Zheng, Rui Nie, Bowen Jiang, et al. [Paper], 2026.3
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation Preprint
Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu, Begum Demir, Nicu Sebe, Paolo Rota. [Paper], 2026.3
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models Preprint
Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, et al. [Paper], 2026.3
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Preprint
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, et al. [Paper], 2026.3
GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback Preprint
Giorgio Giannone, Anna Clare Doris, Amin Heyrani Nobari, Kai Xu, Akash Srivastava, Faez Ahmed. [Paper], 2026.3
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning Preprint
Jian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi, Zhiwen Fan. [Paper], 2026.3
When Choices Become Priors: Contrastive Decoding for Scientific Figure Multiple-Choice QA Preprint
Taeyun Roh, Eun-yeong Jo, Wonjune Jang, Jaewoo Kang. [Paper], 2026.3
MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction Preprint
Zitian Tang, Xu Zhang, Jianbo Yuan, Yang Zou, Varad Gunjal, Songyao Jiang, Davide Modolo. [Paper], 2026.4
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding Preprint
Xuanle Zhao, Xinyuan Cai, Xiang Cheng, Xiuyi Chen, Bo Xu. [Paper], 2026.4
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding Preprint
Hao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang, Minghui Liao, Jihao Wu, Wei Chen, Xiang Bai. [Paper], 2026.4
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Preprint
Yubo Jiang, Yitong An, Xin Yang, Abudukelimu Wuerkaixi, Xuxin Cheng, Fengying Xie, et al. [Paper], 2026.4
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images Preprint
Qingxiao Li, Lifeng Xu, QingLi Wang, Yudong Bai, Mingwei Ou, Shu Hu, Nan Xu. [Paper], 2026.4
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA Preprint
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri, Raviteja Bommireddy, Dinesh Manocha, Puneet Mathur, Vivek Gupta. [Paper], 2026.4
FT-RAG: A Fine-grained Retrieval-Augmented Generation Framework for Complex Table Reasoning Preprint
Zebin Guo, Weidong Geng, Ruichen Mao. [Paper], 2026.5
Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts Preprint
Hongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu, Qitong Yan, Yi Yang, et al. [Paper], 2026.5
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning Preprint
Qihua Dong, Ruozhen He, Junwen Chen, Yizhou Wang, Xu Ma, Songyao Jiang, Yun Fu. [Paper], 2026.5
ChartZero: Synthetic Priors Enable Zero Shot Chart Data Extraction Preprint
Md Touhidul Islam, Yasir Mahmud, Sujan Kumar Saha, Mark Tehranipoor, Farimah Farahmandi. [Paper], 2026.5
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding Preprint
Jiashun Zhu, Ronghao Fu, Jiasen Hu, Nachuan Xing, Xu Na, Xiao Yang, et al. [Paper], 2026.5
From Table to Cell: Attention for Better Reasoning with TABALIGN Preprint
Tung Sum Thomas Kwok, Zeyong Zhang, Xinyu Wang, Chunhe Wang, Xiaofeng Lin, Hanwei Wu, et al. [Paper], 2026.5
MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing Preprint
Chuang Tang, Chenhao Lin, Yin Xu, Hao Wang, Jinrui Zhou, Xin Li, et al. [Paper], [Code], 2026.5
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis Preprint
Jiachen Zhang, Junyi Lao, Chenghao Liu, Siyuan Liu, Shixin Wu, Linsen Zhang, et al. [Paper], 2026.5
BiNSGPS: Geometry Problem Solving via Bidirectional Neuro-Symbolic Interaction Preprint
Qi Wang, Peijie Wang, Fei Yin, Cheng-Lin Liu. [Paper], 2026.6
SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering CIKM 2026 Demo
Ayush Dwivedi, Qixin Wang, Ashvi Soni, Ruoteng Wang, Han Li, Animesh Mahapatra, et al. [Paper], 2026.6
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models Preprint
Chanik Kang, Raphael Pestourie, Haejun Chung. [Paper], 2026.6
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines Preprint
Xindian Ma, Xinyu Long, Yefei Zhang, Yanchen Liu, Xianghao Li, Yufu Wen, et al. [Paper], [Code], 2026.6
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding Preprint
Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, et al. [Paper], 2026.7
AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic Preprint
Zhe Xiao, Longfei Li, Xu He, Haoying Wu, Zixing Zhang, Mingyu Liu. [Paper], [Code], 2026.7
EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning Preprint
Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao. [Paper], [Project], 2026.8
Click2Poly: A VLM for Vector Mapping Buildings and Walls Preprint
Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin, Liuyun Duan, Sacha Lepretre. [Paper], 2026.8
Intern-S2-Preview: Scientific Agentic Foundation Model Technical Report
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, et al. [Paper], [Model], 2026.8
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Preprint
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu. [Paper], [Code], 2026.8
SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models Preprint
Kexin Ma, Jing Xiao, Bowen Xing, Liang Liao, Chia-Wen Lin. [Paper], 2026.8
Training-Free P&ID Information Extraction Through a Hybrid Computer Vision and Vision Language Model Pipeline for Digital Twins Advanced Engineering Informatics
Rafay Hayat Ali, Chan Young Park, Ashrant Aryal, H. David Jeong, Ghang Lee. [Paper], 2026.8
VortexChat: An Agentic Framework for Autonomous Multi-Objective Integrated Photonic Design Preprint
Faqian Chong, Yulun Wu, Shilong Li, Andrew Forbes, Hongsheng Chen, Song Han. [Paper], 2026.8
Unlocking Multimodal Protein Language Models at Inference Time Preprint
Yi Zhou, Qipeng Wang, Yunqing Liu, Jun Xia, Qing Li, Wenqi Fan. [Paper], [Code], 2026.8
| Name | Paper | Notes |
|---|---|---|
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | A benchmark consists of ∼21k multimodal multiple choice questions with diverse science topics. |
| CMM12K | COMET: “Cone of experience” enhanced large multimodal model for mathematical problem generation | A Chinese MM SFT dataset for math, not released |
| SPIQA | SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers | Designed to interpret complex figures and tables within the context of scientific research articles across various domains of computer science |
| InstructDoc | InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions | Collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format. |
| M-Paper | mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model | Built by parsing Latex source files of high-quality papers. |
| DocStruct4M/DocReason25K | mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding | Based on publicly available datasets. A high-quality instruction tuning dataset. |
| DocGenome | DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models | A structured document benchmark constructed by annotating 500K scientific documents from 153 disciplines in the arXiv. |
| ArXivCap/ArXivQA | Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models | A figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers. A QA dataset generated by prompting GPT-4V based on scientific figures. |
| FigureQA | FigureQA: An Annotated Figure Dataset for Visual Reasoning | A visual reasoning corpus of over one million QA pairs grounded in over 100,000 images. The images are synthetic, scientific-style figures from five classes: line plots, dotline plots, vertical and horizontal bar graphs, and pie charts |
| DVQA | DVQA: Understanding Data Visualizations via Question Answering | A dataset that tests many aspects of bar chart understanding in a question answering framework. |
| SciGraphQA | SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs | A synthetic multi-turn QA dataset related to academic graphs. |
| SciCap | SciCap: Generating Captions for Scientific Figures | A large-scale figure caption dataset based on Computer Science arXiv papers published between 2010 and 2020, contained over 416k figures that focused on graphplot. |
| FigCap | Figure Captioning with Reasoning and Sequence-Level Training | Generated based on FigureQA |
| FigureSeer | FigureSeer: Parsing Result-Figures in Research Papers | - |
| UniChart | UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning | A large-scale chart corpus for pretraining, covering a diverse range of visual styles and topics. |
| MapQA | MapQA: A Dataset for Question Answering on Choropleth Maps | A large-scale dataset of ~800K question-answer pairs over ~60K map images. |
| TabMWP | Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning | A dataset containing 38,431 open-domain grade-level problems that require mathematical reasoning on both textual and tabular data |
| CLEVR-Math | CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning | A multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction |
| GUICourse | GUICourse: From General Vision Language Model to Versatile GUI Agent | A suite of datasets to train visual-based GUI agents from general VLMs |
| PIN-14M | PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents | 14 million samples derived from Chinese and English sources, tailored to include complex web and scientific content. |
| MathV360K | Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models | 40K high-quality images with QA pairs from 24 existing datasets and synthesizing 320K new pairs. |
| MMSci | MMSci: A Multimodal Multi-Discipline Dataset for PhD-Level Scientific Comprehension | Collected a multimodal dataset from open-access scientific articles published in Nature Communications journals. |
| MAVIS-Caption/Instruct | MAVIS: Mathematical Visual Instruction Tuning | - |
| Geo170K | G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model | Utilize the geometry characteristic to construct a multi-modal geometry dataset, building upon existing datasets. |
| SciOL/MuLMS-Img | SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain | Pretraining corpus for multimodal models in the scientific domain. |
| PlotQA | PlotQA: Reasoning over Scientific Plots | With 28.9 million QA pairs over 224,377 plots on data from realworld sources and questions based on crowd-sourced question templates. |
| ChartInstructionData | Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning | A dataset of 467K, which includes 108K table-chart pairs and 359K chart-QA pairs. |
| MMTab | Multimodal Table Understanding | Dataset for multimodal table understanding problem, based on 14 publicly available table datasets of 8 domains. |
| Multimodal Self-Instruct | Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model | Instruction dataset for eight visual scenarios: charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles. |
| GeoGPT4V | GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation | Leverages GPT-4 and GPT-4V to generate relatively basic geometry problems with aligned text and images. |
| InfiMM-WebMath-40B | InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning | Interleaved image-text documents, comprises 24 million web pages, 85 million associated image URLs, and 40 billion text tokens, extracted and filtered from CommonCrawl. |
| MultiMath-300K | MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models | Spans K-12 levels with image captions and step-wise solutions. |
| MathVL | MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model | A fine-tuning dataset including both several public datasets and our curated Chinese dataset collected from K12 education levels. |
| AtomMATH | AtomThink: A Slow Thinking Framework for Multimodal Mathematical Reasoning | A large-scale multimodal dataset of long CoTs, and an atomic capability evaluation metric for mathematical tasks. |
| MAmmoTH-VL | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale | A dataset containing 12M instruction-response pairs to cover diverse, reasoning-intensive tasks with detailed and faithful rationales. |
| BIGDOCS | BIGDOCS: AN OPEN AND PERMISSIVELY-LICENSED DATASET FOR TRAINING MULTIMODAL MODELS ON DOCUMENT AND CODE TASKS | Comprising 7.5 million multimodal documents across 30 tasks. |
| 2.5 Years in Class | 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining | Collects over 2.5 years of instructional videos, totaling 22,000 class hours. |
| MM-PRM | MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision | A curated dataset of 10,000 multimodal math problems with verifiable answers, which serves as seed data. Leveraging a Monte Carlo Tree Search (MCTS)-based pipeline, generate over 700k step-level annotations without human labeling |
| BMMR | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset | a large-scale bilingual, multimodal, multidisciplinary reasoning dataset for the community to develop and evaluate large multimodal models |
| Zebra-CoT | Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning | a diverse large-scale dataset with 182,384 samples, containing logically coherent interleaved text-image reasoning traces |
| SldprtNet | SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design | A large-scale multimodal CAD dataset with over 242K industrial parts in STEP and SLDPRT formats for language-driven 3D design and CAD generation. |
| CharTide-2M | CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution | A 2M-sample chart-to-code training and alignment dataset built with tri-perspective tuning and inquiry-driven verification. |
| Chart2NCode | Aligned Multi-View Scripts for Universal Chart-to-Code Generation | 176K chart images paired with aligned executable scripts across multiple plotting languages for universal chart-to-code generation. |
| Zero-to-CAD | Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data | Around one million synthetic executable CAD construction sequences plus a curated 100K high-quality subset for CAD program generation. |
| CADFS | CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models | 450K real-world CAD models represented with FeatureScript and spanning 15 modeling operations for richer CAD generation. |
| DocAtlas | DocAtlas: Multilingual Document Understanding Across 80+ Languages | High-fidelity OCR datasets and benchmarks covering 82 languages and 9 document understanding evaluation tasks. |
| Sentinel2Cap | Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning | A human-annotated multimodal remote sensing captioning dataset pairing Sentinel-1 SAR and Sentinel-2 multi-spectral image patches with validated captions. |
| Ryze Evidence-Enriched Dataset | Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers | An automatically synthesized biomedical SFT corpus whose QA pairs retain the supporting visual element, caption, extracted structure, and referring prose; the pipeline produces millions of domain QA tokens without human annotation. |
| LabEmbodied-Data | LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories | Success-filtered laboratory robot demonstrations with multi-camera observations, language instructions, robot states, action trajectories, and structured annotations across 16 robot platforms and four task families. |
| FusionRS | FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning | 600K aligned RGB-infrared-style remote-sensing records with 599,992 source-text pairs and 45,913 IR-aware captions, organized into group-aware 580K/10K/10K train/validation/test splits. |
| Name | Paper | Note |
|---|---|---|
| GeoEval | GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving | An benchmark for evaluating MLLMs' capability in solving geometry math problems |
| Geometry3K | Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning | Consisting of 3,002 geometry problems with dense annotation in formal language. |
| GEOS | Solving Geometry Problems: Combining Text and Diagram Interpretation | - |
| GeoQA | GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning | 4,998 geometric problems with cor- responding annotated programs |
| GeoQA+ | An Augmented Benchmark Dataset for Geometric Question Answering through Dual Parallel Text Encoding | Based on GeoQA, newly annotate 2,518 geometric problems with richer types and greater difficulty |
| UniGeo | UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression | Contains 4,998 calculation problems and 9,543 proving problems |
| PGPS9K | A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram | Labeled with both fine-grained diagram annotation and interpretable solution program. |
| GeomVerse | GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning | A synthetic benchmark of geometry questions with controllable difficulty levels along multiple axes |
| MathVista | MATHVISTA: EVALUATING MATHEMATICAL REASONING OF FOUNDATION MODELS IN VISUAL CONTEXTS | A benchmark designed to combine challenges from diverse mathematical and visual tasks. |
| OlympiadBench | OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems | An Olympiad-level bilingual multimodal scientific benchmark, from mathematics and physics competitions |
| OlympicArena | OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI | Encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions. |
| SciBench | SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models | A benchmark for college-level scientific problems sourced from instructional textbooks. |
| MMMU | MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI | Designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. |
| CMMMU | CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark | A new Chinese Massive Multi-discipline Multimodal Understanding benchmark designed to evaluate LMMs on tasks demanding college-level subject knowledge and deliberate reasoning in a Chinese context. |
| MULTI | MULTI: Multimodal Understanding Leaderboard with Text and Images | Includes over 18,000 questions, and challenges MLLMs with a variety of tasks, ranging from formula derivation to image detail analysis and cross-modality reasoning. |
| M3GIA | M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark | Designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. |
| M3Exam | M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models | Sourced from real and official human exam questions for evaluating LLMs in a multilingual, multimodal, and multilevel context. |
| MathVerse | MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? | 2,612 high-quality, multi-subject math problems with diagrams from publicly available sources. |
| MATH-Vision | Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset | 3,040 high-quality mathe- matical problems with visual contexts sourced from real math competitions. |
| AI2D | A Diagram Is Worth A Dozen Images | A dataset of diagrams with annotations of constituents and relationships for over 5,000 diagrams and 15,000 QAs. |
| IconQA | IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning | A benchmark with the goal of answering a question in an icon image context. |
| TQA | Are You Smarter Than A Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension | Includes 1,076 lessons and 26,260 multi-modal questions, taken from middle school science curricula. |
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | A benchmark consists of ∼21k multimodal multiple choice questions with diverse science topics. |
| ChartX | ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning | A multi-modal evaluation set covering 18 chart types, 7 chart tasks, 22 disciplinary topics, and high-quality chart data |
| PlotQA | PlotQA: Reasoning over Scientific Plots | With 28.9 million question-answer pairs over 224,377 plots on data from realworld sources and questions based on crowd-sourced question templates. |
| Chart-to-text | Chart-to-Text: A Large-Scale Benchmark for Chart Summarization | A large-scale benchmark with two datasets and a total of 44,096 charts covering a wide range of topics and chart types. |
| ChartQA | ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning | A large-scale benchmark covering 9.6K human-written questions as well as 23.1K questions generated from human-written chart summaries. |
| OpenCQA | OpenCQA: Open-ended Question Answering with Charts | The goal is to answer an open-ended question about a chart with descriptive texts. |
| ChartBench | ChartBench: A Benchmark for Complex Visual Reasoning in Charts | A comprehensive benchmark designed to assess chart comprehension and data reliability through complex visual reasoning. |
| DocVQA | DocVQA: A Dataset for VQA on Document Images | Consists of 50,000 questions defined on 12,000+ document images |
| InfoVQA | InfographicVQA | Comprises a diverse collection of infographics along with question-answer annotations. |
| WTQ | Compositional Semantic Parsing on Semi-Structured Tables | A dataset of 22,033 complex questions on Wikipedia tables. |
| TableFact | TabFact : A Large-scale Dataset for Table-based Fact Verification | A large-scale dataset with 16k Wikipedia tables as the evidence for 118k human-annotated natural language statements. |
| MM-Math | MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification | Consists of 5,929 open-ended middle school math problems with visual contexts, with fine-grained classification. |
| MathCheck | Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist | A well-designed checklist for testing task generalization and reasoning robustness. |
| PuzzleVQA | PUZZLEVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns | A collection of 2000 puzzle instances based on abstract patterns. |
| SMART-101 | Are Deep Neural Networks SMARTer than Second Graders? | Evaluating the abstraction, deduction, and generalization abilities of neural networks in solving visul-linguistic puzzles. |
| AlgpPuzzleVQA | ARE LANGUAGE MODELS PUZZLE PRODIGIES? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning | Evaluate the capabilities in solving algorithmic puzzles. |
| ChartMimic | ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation | Aimed at assessing the visually-grounded code generation capabilities. |
| ChartSumm | ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries | - |
| MMCode | MMCode: Evaluating Multi-Modal Code Large Language Models with Visually Rich Programming Problems | Contains 3,548 questions and 6,620 images collected from real-world programming challenges harvested from 10 code competition websites. |
| Design2Code | Design2Code: How Far Are We From Automating Front-End Engineering | Manually curate a benchmark of 484 diverse real-world webpages |
| Plot2Code | Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots | A comprehensive visual coding benchmark. |
| CharXiv | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs | A comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. |
| We-Math | WE-MATH: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? | 6.5K visual math problems, spanning 67 hierarchical knowledge concepts and 5 layers of knowledge granularity. |
| SceMQA | SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark | A benchmark for scientific multimodal question answering at the college entrance leve. |
| TheoremQA | TheoremQA: A Theorem-driven Question Answering dataset | Curated by domain experts containing 800 high-quality questions covering 350 theorems from Math, Physics, EE&CS, and Finance. |
| NPHardEval4V | NPHardEval4V: A Dynamic Reasoning Benchmark of Multimodal Large Language Models | Built by converting textual description of questions from NPHardEval to image representations. |
| MathScape | MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark | Designed to evaluate photo-based math problem scenarios, assessing the theoretical understanding and application ability of MLLMs through a categorical hierarchical approach. |
| TableBench | TableBench: A Comprehensive and Complex Benchmark for Table Question Answering | Including 18 fields within four major categories of table question answering capabilities. |
| GRAB | GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models | Synthetic, comprised of 2170 questions, covering four tasks and 23 graph properties. |
| LogicVista | LogicVista: A Benchmark for Evaluating Multimodal Logical Reasoning | Evaluate general logical cognition abilities across 5 logical reasoning tasks encompassing 9 different capabilities, using a sample of 448 multiple-choice questions. |
| CMM-Math | CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models | Contains over 28,000 high-quality samples, featuring a variety of problem types with detailed solutions across 12 grade levels from elementary to high school in China. |
| SWE-bench Multimodal | SWE-BENCH MULTIMODAL: DO AI SYSTEMS GENERALIZE TO VISUAL SOFTWARE DOMAINS? | Contains 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. |
| MMIE | MMIE: MASSIVE MULTIMODAL INTERLEAVED COMPREHENSION BENCHMARK FOR LARGE VISIONLANGUAGE MODELS | Comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. |
| MultiChartQA | MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems | Multi-hop reasoning required to extract and integrate information from multiple charts, comprises 655 charts and 944 questions |
| Sketch2Code | Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping | Evaluating automating the conversion of rudimentary sketches into webpage prototypes, collected a total of 731 sketches for 484 webpage screenshots |
| PolyMath | POLYMATH: A CHALLENGING MULTI-MODAL MATHEMATICAL REASONING BENCHMARK | Comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning |
| VisAidMath | VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning | Includes 1,200 challenging problems from various mathematical branches, vision-aid formulations, and difficulty levels, collected from diverse sources such as textbooks, examination papers, and Olympiad problems |
| DYNAMATH | DYNAMATH: A DYNAMIC VISUAL BENCHMARK FOR EVALUATING MATHEMATICAL REASONING ROBUSTNESS OF VISION LANGUAGE MODELS | Includes 501 high-quality, multi-topic seed questions, each represented as a Python program |
| M3SCIQA | M3SCIQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models | Consists of 1,452 expert-annotated questions spanning 70 natural language processing paper clusters |
| M-LONGDOC | M-LONGDOC: A BENCHMARK FOR MULTIMODAL SUPER-LONG DOCUMENT UNDERSTANDING AND A RETRIEVAL-AWARE TUNING FRAMEWORK | A benchmark of 851 samples, and an automated framework to evaluate the performance of large multimodal models |
| VisOnlyQA | VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information | Includes 1,200 multiple-choice questions in 12 tasks on four categories of figures. Designed to directly evaluate the visual perception capabilities. |
| U-MATH | U-MATH: A UNIVERSITY-LEVEL BENCHMARK FOR EVALUATING MATHEMATICAL SKILLS IN LLMS | 1,100 unpublished open-ended university-level problems sourced from teaching materials. It is balanced across six core subjects, with 20% of multimodal problems. |
| DrawEduMath | DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images | An English-language dataset of 2,030 images of students' handwritten responses to K-12 math problems |
| MM-IQ | MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models | An evaluation framework comprising 2,710 meticulously curated test items spanning 8 distinct reasoning paradigms |
| LOST IN TIME | LOST IN TIME: CLOCK AND CALENDAR UNDERSTANDING CHALLENGES IN MULTIMODAL LLMS | Curated a structured dataset comprising two subsets: ClockQA and CalendarQA |
| ProJudge | ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges | Comprises 2,400 test cases and 50,118 step-level labels, spanning four scientific disciplines with diverse difficulty levels and multimodal content |
| MPBench | MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification | A comprehensive, multi-task, multimodal benchmark designed to systematically assess the effectiveness of PRMs in diverse scenarios |
| FlowVerse | MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems | A comprehensive benchmark that categorizes all information used during problem-solving into four components |
| ChartQAPRO | ChartQA PRO : A More Diverse and Challenging Benchmark for Chart Question Answering | A new benchmark that includes 1,341 charts from 157 diverse sources, spanning various chart types. 1,948 questions in various types, to better reflect real-world challenges |
| ChartMuseum | ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models | Chart Question Answering (QA) benchmark containing 1,162 expert-annotated questions spanning multiple reasoning types, curated from realworld charts across 184 sources |
| FullFront | FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow | Assesses three fundamental tasks that map directly to the front-end engineering pipeline: Webpage Design, Webpage Perception QA, and Webpage Code Generation |
| MMMR | MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks | 1,083 questions spanning six diverse reasoning types with symbolic depth and multi-hop demands |
| VideoMathQA | VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos | Spans 10 diverse mathematical domains, covering videos ranging from 10 seconds to over 1 hour. |
| Scientists’ First Exam | Scientists’ First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning | Comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. |
| SCIVER | SCIVER: Evaluating Foundation Models for Multimodal Scientific Claim Verification | Consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification. |
| MATHREAL | MATHREAL: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models | Comprising 2,000 mathematical questions with images captured by handheld mobile devices in authentic scenarios. |
| WE-MATH 2.0 | WE-MATH 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning | A comprehensive benchmark covering all 491 knowledge points with diverse reasoning step distributions. |
| Uni-MMMU | Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark | A benchmark of eight bidirectionally coupled tasks that enforce Gen–Und logical dependency. |
| GGBench | GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models | A benchmark designed specifically to evaluate geometric generative reasoning. |
| ScratchMath | Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math | 1,720 authentic handwritten math scratchwork samples for error cause explanation and classification across seven error types. |
| GeoAux-Bench | Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning | A benchmark for visual-text interleaved geometric reasoning where models must decide when and how to construct visual aids. |
| MultihopSpatial | MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model | Evaluates multi-hop compositional spatial reasoning and precise visual grounding for VLM and VLA-style agents. |
| AEC-Bench | AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction | Real-world AEC tasks requiring drawing understanding, cross-sheet reasoning, and project-level coordination. |
| GeoMMBench | GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing | 1,053 expert-level image-based multiple-choice questions spanning geoscience and remote sensing disciplines. |
| HM-Bench | HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing | A hyperspectral remote sensing benchmark for testing spectral-spatial perception and reasoning in MLLMs. |
| PaperScope | PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers | Evaluates agentic deep research over many scientific papers with evidence from text, tables, and figures. |
| MMCoIR | CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval | A multimodal code information retrieval benchmark across five visual domains, eight programming languages, and eleven libraries. |
| ReactBench | ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams | Probes topological reasoning over branching, converging, and cyclic chemical reaction diagrams. |
| ArXivDoc | Document-as-Image Representations Fall Short for Scientific Retrieval | A scientific document retrieval benchmark built from LaTeX sources to test text, tables, figures, and equations. |
| MathNet | MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval | 30K+ Olympiad-level math problems from 47 countries, with a retrieval benchmark for mathematical problem search. |
| STEP-STEM | Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks | Evaluates fine-grained visual traces in multimodal interleaved reasoning chains for STEM tasks. |
| OMIBench | OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model | An Olympiad-level benchmark where evidence is distributed across multiple images. |
| AstroVLBench | A Systematic Evaluation of Vision-Language Models for Observational Astronomical Reasoning Tasks | 4,100+ expert-verified instances across optical imaging, radio interferometry, photometry, time-domain signals, and spectroscopy. |
| SpecVQA | SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images | A scientific spectral image VQA benchmark covering seven representative spectrum types with expert annotations. |
| TopBench | TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering | Tests implicit prediction and latent-intent reasoning over table question answering rather than simple retrieval. |
| Text-to-CAD Retrieval | Text-to-CAD Retrieval: A Strong Baseline | Establishes a cross-modal retrieval benchmark for finding semantically relevant CAD models from natural-language queries. |
| AstroAlertBench | AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification | Evaluates multimodal LLM accuracy, reasoning, and honesty on astronomical alert classification. |
| TableVista | TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity | 3,000 table reasoning problems expanded into diverse visual variants for robustness and structural complexity testing. |
| ChartREG++ | ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse Referring Clues and Multi-Target Referring | A chart referring expression grounding benchmark with diverse referring clues and multi-target localization. |
| CADBench | CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation | 18,000 evaluation samples across six CAD benchmark families for multimodal CAD program generation. |
| Vision2Code | Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation | A reference-code-free benchmark and evaluation framework for image-to-code generation across multiple visual domains. |
| UHR-Micro | UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs | Diagnoses micro-target perception failures in ultra-high-resolution Earth observation VLMs. |
| CiteVQA | CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence | Tests whether document VQA models cite and ground the visual or textual evidence supporting their answers. |
| IndustryBench-MIPU | IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products | 4,559 industrial products, 27,652 images, and 103,703 attribute annotations across 18 categories for evaluating multi-image technical specification extraction. |
| DashboardMimic | Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards | 180 manually verified Plotly+Dash dashboard-code pairs spanning three difficulty levels, 20 visualization types, and eight interaction patterns, with dynamic tests for visual and behavioral fidelity. |
| SVLAT for MLLMs | Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy | Evaluates six MLLMs on a standardized 49-item scientific-visualization literacy test built from 18 visualizations, eight techniques, and 11 task types, with comparison data from 485 human participants. |
| TableParseMap | From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing | 916 real-world complex tables organized into five challenging scenarios and nine failure types, accompanied by a 1,977-table Consensus-Hard Set for cross-parser evaluation. |
| SonarBench | SonarLLM: A Native Sonar-Optical Multimodal Large Language Model for Underwater Perception | A paired sonar-optical benchmark covering recognition, counting, VQA, and captioning across 25 subsets, with controlled optical degradation to isolate cross-modal complementarity under turbidity. |
If you have any question about this opinionated list, do not hesitate to create an issue.
70 commits
1 commits
Paper collections of multi-modal LLM for Math/STEM/Code.
146
71 commits
updated Aug 30, 2026

🔥 Collections of multi-modal LLM for Math/STEM/Code.
MAVIS: Mathematical Visual Instruction Tuning Preprint
Renrui Zhang, Xinyu Wei, Dongzhi Jiang, Yichi Zhang, Ziyu Guo,Chengzhuo Tong, Jiaming Liu, Aojun Zhou, Bin Wei, Shanghang Zhang, Peng Gao, Hongsheng Li.[Paper], 2024.7
COMET: “Cone of experience” enhanced large multimodal model for mathematical problem generation. Preprint
Sannyuya Liu, Jintian Feng, Zongkai Yang, Yawei Luo, Qian Wan, Xiaoxuan Shen, Jianwen Sun. [Paper], 2024.7
Accessing GPT-4 level Mathematical Olympiad Solutions via Monte Carlo Tree Self-refine with LLaMa-3 8B: A Technical Report. Preprint
Di Zhang, Xiaoshui Huang, Dongzhan Zhou, Yuqiang Li, Wanli Ouyang. [Paper], 2024.6
Visual SKETCHPAD: Sketching as a Visual Chain of Thought for Multimodal Language Models. Preprint
Yushi Hu, Weijia Shi, Xingyu Fu, Dan Roth, Mari Ostendorf, Luke Zettlemoyer, Noah A Smith, Ranjay Krishna. [Paper], [Code], 2024.6
TextSquare: Scaling up Text-Centric Visual Instruction Tuning. Preprint
Jingqun Tang, Chunhui Lin, Zhen Zhao, Shu Wei, Binghong Wu, Qi Liu, Hao Feng, Yang Li, Siqi Wang, Lei Liao, Wei Shi, Yuliang Liu, Hao Liu, Yuan Xie, Xiang Bai, Can Huang. [Paper], 2024.4
Chart-based Reasoning: Transferring Capabilities from LLMs to VLMs. ACL 2024
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou. [Paper], 2024.3
mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding. Preprint
Anwen Hu, Haiyang Xu, Jiabo Ye, Ming Yan, Liang Zhang, Bo Zhang, Chen Li, Ji Zhang, Qin Jin, Fei Huang, Jingren Zhou. [Paper], 2024.3
ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning. Preprint
Renqiu Xia, Bo Zhang, Hancheng Ye, Xiangchao Yan, Qi Liu, Hongbin Zhou, Zijun Chen, Min Dou, Botian Shi, Junchi Yan, Yu Qiao. [Paper], 2024.2
InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions. Preprint
Ryota Tanaka, Taichi Iki, Kyosuke Nishida, Kuniko Saito, Jun Suzuki. [Paper], 2024.1
G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model. Preprint
Jiahui Gao, Renjie Pi, Jipeng Zhang, Jiacheng Ye, Wanjun Zhong, Yufei Wang, Lanqing Hong, Jianhua Han, Hang Xu, Zhenguo Li, Lingpeng Kong. [Paper], [Code], 2023.12
mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Models. Preprint
Anwen Hu, Yaya Shi, Haiyang Xu, Jiabo Ye, Qinghao Ye, Ming Yan, Chenliang Li, Qi Qian, Ji Zhang, Fei Huang. [Paper], 2023.11
Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning. Preprint
Xingchen Zeng, Haichuan Lin, Yilin Ye, Wei Zeng. [Paper], [Code], 2024.7
Math-PUMA: Progressive Upward Multimodal Alignment to Enhance Mathematical Reasoning. Preprint
Wenwen Zhuang, Xin Huang, Xiantao Zhang, Jin Zeng. [Paper], 2024.8
Diagram Formalization Enhanced Multi-Modal Geometry Problem Solver. Preprint
Zeren Zhang, Jo-Ku Cheng, Jingyang Deng, Lu Tian, Jinwen Ma, Ziran Qin, Xiaokai Zhang, Na Zhu, and Tuo Leng. [Paper], 2024.9
Transformers Utilization in Chart Understanding: A Review of Recent Advances & Future Trends. Preprint
Mirna Al-Shetairy, Hanan Hindy, Dina Khattab, Mostafa M. Aref. [Paper], 2024.10
IMPROVE VISION LANGUAGE MODEL CHAIN-OFTHOUGHT REASONING. Preprint
Ruohong Zhang, Bowen Zhang, Yanghao Li, Haotian Zhang, Zhiqing Sun, Zhe Gan, Yinfei Yang, Ruoming Pang, Yiming Yang. [Paper], 2024.10
R-COT : REVERSE CHAIN-OF-THOUGHT PROBLEM GENERATION FOR GEOMETRIC REASONING IN LARGE MULTIMODAL MODELS. Preprint
Linger Deng, Yuliang Liu, Bohan Li, Dongliang Luo, Liang Wu, Chengquan Zhang, Pengyuan Lyu, Ziyang Zhang, Gang Zhang, Errui Ding, Yingying Zhu, Xiang Bai. [Paper], 2024.10
GeoCoder: Solving Geometry Problems by Generating Modular Code through Vision-Language Models. Preprint
Aditya Sharma, Aman Dalmia, Mehran Kazemi, Amal Zouaq, Christopher J. Pal. [Paper], 2024.10
VISTA: Visual Integrated System for Tailored Automation in Math Problem Generation Using LLM. NeurIPS 2024 Workshop
Jeongwoo Lee, Kwangsuk Park, Jihyeon Park. [Paper], 2024.11
DIVING INTO SELF-EVOLVING TRAINING FOR MULTIMODAL REASONING. Preprint
Wei Liu, Junlong Li, Xiwen Zhang, Fan Zhou, Yu Cheng, Junxian He. [Paper], 2024.12
Slow Perception: Let’s Perceive Geometric Figures Step-by-step. Preprint
Haoran Wei, Youyang Yin, Yumeng Li, Jia Wang, Liang Zhao, Jianjian Sun, Zheng Ge, Xiangyu Zhang. [Paper], 2024.12
Virgo: A Preliminary Exploration on Reproducing o1-like MLLM. Preprint
Yifan Du, Zikang Liu1, Yifan Li, Wayne Xin Zhao, Yuqi Huo, Bingning Wang, Weipeng ChenZheng Liu, Zhongyuan Wang, Ji-Rong Wen. [Paper], 2025.1
URSA: Understanding and Verifying Chain-of-thought Reasoning in Multimodal Mathematics. Preprint
Ruilin Luo, Zhuofan Zheng, Yifan Wang, Yiyao Yu, Xinzhe Ni, Zicheng Lin, Jin Zeng, Yujiu Yang. [Paper], 2025.1
Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation. Preprint
Yue Yang, Ajay Patel, Matt Deitke, Tanmay Gupta, Luca Weihs, Andrew Head, Mark Yatskar, Chris Callison-Burch, Ranjay Krishna, Aniruddha Kembhavi, Christopher Clark. [Paper], 2025.2
MathCoder-VL: Bridging Vision and Code for Enhanced Multimodal Mathematical Reasoning. Preprint
Ke Wang, Junting Pan, Linda Wei, Aojun Zhou, Weikang Shi, Zimu Lu, Han Xiao, Yunqiao Yang, Houxing Ren, Mingjie Zhan, Hongsheng Li. [Paper], 2025.5
MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning. Preprint
Xinyan Chen, Renrui Zhang, Dongzhi Jiang, Aojun Zhou, Shilin Yan, Weifeng Lin, Hongsheng Li. [Paper], 2025.6
Improved Iterative Refinement for Chart-to-Code Generation via Structured Instruction. Preprint
Chengzhi Xu, Yuyang Wang, Lai Wei, Lichao Sun, Weiran Huang. [Paper], 2025.6
MathCanvas: Intrinsic Visual Chain-of-Thought for Multimodal Mathematical Reasoning. Preprint
Weikang Shi, etc. [Paper], 2025.10
MathSE: Improving Multimodal Mathematical Reasoning via Self-Evolving Iterative Reflection and Reward-Guided Fine-Tuning. Preprint
Jinhao Chen, Zhen Yang, Jianxin Shi, Tianyu Wo, Jie Tang. [Paper], 2025.11
Multimodal OCR: Parse Anything from Documents Preprint
Handong Zheng, Yumeng Li, Kaile Zhang, Liang Xin, Guangwei Zhao, Hao Liu, et al. [Paper], 2026.3
Molecular Identifier Visual Prompt and Verifiable Reinforcement Learning for Chemical Reaction Diagram Parsing Preprint
Jiahe Song, Chuang Wang, Yinfan Wang, Hao Zheng, Rui Nie, Bowen Jiang, et al. [Paper], 2026.3
TerraScope: Pixel-Grounded Visual Reasoning for Earth Observation Preprint
Yan Shu, Bin Ren, Zhitong Xiong, Xiao Xiang Zhu, Begum Demir, Nicu Sebe, Paolo Rota. [Paper], 2026.3
VFIG: Vectorizing Complex Figures in SVG with Vision-Language Models Preprint
Qijia He, Xunmei Liu, Hammaad Memon, Ziang Li, Zixian Ma, Jaemin Cho, et al. [Paper], 2026.3
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Preprint
Yicheng Zou, Dongsheng Zhu, Lin Zhu, Tong Zhu, Yunhua Zhou, Peiheng Zhou, et al. [Paper], 2026.3
GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric Feedback Preprint
Giorgio Giannone, Anna Clare Doris, Amin Heyrani Nobari, Kai Xu, Akash Srivastava, Faez Ahmed. [Paper], 2026.3
SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning Preprint
Jian Zhang, Shijie Zhou, Bangya Liu, Achuta Kadambi, Zhiwen Fan. [Paper], 2026.3
When Choices Become Priors: Contrastive Decoding for Scientific Figure Multiple-Choice QA Preprint
Taeyun Roh, Eun-yeong Jo, Wonjune Jang, Jaewoo Kang. [Paper], 2026.3
MM-ReCoder: Advancing Chart-to-Code Generation with Reinforcement Learning and Self-Correction Preprint
Zitian Tang, Xu Zhang, Jianbo Yuan, Yang Zou, Varad Gunjal, Songyao Jiang, Davide Modolo. [Paper], 2026.4
ChemVLR: Prioritizing Reasoning in Perception for Chemical Vision-Language Understanding Preprint
Xuanle Zhao, Xinyuan Cai, Xiang Cheng, Xiuyi Chen, Bo Xu. [Paper], 2026.4
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding Preprint
Hao Yan, Yuliang Liu, Xingchen Liu, Yuyi Zhang, Minghui Liao, Jihao Wu, Wei Chen, Xiang Bai. [Paper], 2026.4
V-tableR1: Process-Supervised Multimodal Table Reasoning with Critic-Guided Policy Optimization Preprint
Yubo Jiang, Yitong An, Xin Yang, Abudukelimu Wuerkaixi, Xuxin Cheng, Fengying Xie, et al. [Paper], 2026.4
S1-VL: Scientific Multimodal Reasoning Model with Thinking-with-Images Preprint
Qingxiao Li, Lifeng Xu, QingLi Wang, Yudong Bai, Mingwei Ou, Shu Hu, Nan Xu. [Paper], 2026.4
DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA Preprint
Anirudh Iyengar Kaniyar Narayana Iyengar, Tampu Ravi Kumar, Manan Suri, Raviteja Bommireddy, Dinesh Manocha, Puneet Mathur, Vivek Gupta. [Paper], 2026.4
FT-RAG: A Fine-grained Retrieval-Augmented Generation Framework for Complex Table Reasoning Preprint
Zebin Guo, Weidong Geng, Ruichen Mao. [Paper], 2026.5
Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense Charts Preprint
Hongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu, Qitong Yan, Yi Yang, et al. [Paper], 2026.5
Hierarchical Visual Agent: Managing Contexts in Joint Image-Text Space for Advanced Chart Reasoning Preprint
Qihua Dong, Ruozhen He, Junwen Chen, Yizhou Wang, Xu Ma, Songyao Jiang, Yun Fu. [Paper], 2026.5
ChartZero: Synthetic Priors Enable Zero Shot Chart Data Extraction Preprint
Md Touhidul Islam, Yasir Mahmud, Sujan Kumar Saha, Mark Tehranipoor, Farimah Farahmandi. [Paper], 2026.5
GeoVista: Visually Grounded Active Perception for Ultra-High-Resolution Remote Sensing Understanding Preprint
Jiashun Zhu, Ronghao Fu, Jiasen Hu, Nachuan Xing, Xu Na, Xiao Yang, et al. [Paper], 2026.5
From Table to Cell: Attention for Better Reasoning with TABALIGN Preprint
Tung Sum Thomas Kwok, Zeyong Zhang, Xinyu Wang, Chunhe Wang, Xiaofeng Lin, Hanwei Wu, et al. [Paper], 2026.5
MACReD: A Multi-Agent Collaborative Reasoning Framework for Reaction Diagram Parsing Preprint
Chuang Tang, Chenhao Lin, Yin Xu, Hao Wang, Jinrui Zhou, Xin Li, et al. [Paper], [Code], 2026.5
VFEAgent: A Multimodal Agent Framework for End-to-End Automated Finite Element Analysis Preprint
Jiachen Zhang, Junyi Lao, Chenghao Liu, Siyuan Liu, Shixin Wu, Linsen Zhang, et al. [Paper], 2026.5
BiNSGPS: Geometry Problem Solving via Bidirectional Neuro-Symbolic Interaction Preprint
Qi Wang, Peijie Wang, Fei Yin, Cheng-Lin Liu. [Paper], 2026.6
SAFE-Cascade: Cost-Adaptive Vision-Language Routing for Chart Question Answering CIKM 2026 Demo
Ayush Dwivedi, Qixin Wang, Ashvi Soni, Ruoteng Wang, Han Li, Animesh Mahapatra, et al. [Paper], 2026.6
VLM-Aware Meta-Optic Front-End Design for Frozen Vision-Language Models Preprint
Chanik Kang, Raphael Pestourie, Haejun Chung. [Paper], 2026.6
RiverONE: Generating Knowledge-Intensive VLM by Simulated Quantum Machines Preprint
Xindian Ma, Xinyu Long, Yefei Zhang, Yanchen Liu, Xianghao Li, Yufu Wen, et al. [Paper], [Code], 2026.6
MolSight: A Graph-Aware Vision-Language Model for Unified Chemical Image Understanding Preprint
Wenda Wang, Yihan Tong, Yuwei Hu, Xuchen Pan, Zhewei Wei, Yaliang Li, et al. [Paper], 2026.7
AutoVSR: Automatic Visual-to-Symbolic Reasoning for Symbolic Expression Generation from Circuit Schematic Preprint
Zhe Xiao, Longfei Li, Xu He, Haoying Wu, Zixing Zhang, Mingyu Liu. [Paper], [Code], 2026.7
EchoChange: A Diffusion Language Model with Dual Pass Remasking for Factual Remote Sensing Disaster Change Captioning Preprint
Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao. [Paper], [Project], 2026.8
Click2Poly: A VLM for Vector Mapping Buildings and Walls Preprint
Nicolas Girard, Jawher Ben Abdallah, Arno Gobbin, Liuyun Duan, Sacha Lepretre. [Paper], 2026.8
Intern-S2-Preview: Scientific Agentic Foundation Model Technical Report
Lei Bai, Jiaqi Cao, Chiyu Chen, Guanzhou Chen, Kai Chen, Guangran Cheng, et al. [Paper], [Model], 2026.8
OmniScientist: An Omni-Modal Omni-Discipline AI Scientist Preprint
Bobo Li, Hao Fei, Tianjie Ju, Mong-Li Lee, Wynne Hsu. [Paper], [Code], 2026.8
SA-GEM: Scale-Adaptive and Geospatial Evidence-Modulated Token Pruning for Efficient Remote Sensing Large Vision-Language Models Preprint
Kexin Ma, Jing Xiao, Bowen Xing, Liang Liao, Chia-Wen Lin. [Paper], 2026.8
Training-Free P&ID Information Extraction Through a Hybrid Computer Vision and Vision Language Model Pipeline for Digital Twins Advanced Engineering Informatics
Rafay Hayat Ali, Chan Young Park, Ashrant Aryal, H. David Jeong, Ghang Lee. [Paper], 2026.8
VortexChat: An Agentic Framework for Autonomous Multi-Objective Integrated Photonic Design Preprint
Faqian Chong, Yulun Wu, Shilong Li, Andrew Forbes, Hongsheng Chen, Song Han. [Paper], 2026.8
Unlocking Multimodal Protein Language Models at Inference Time Preprint
Yi Zhou, Qipeng Wang, Yunqing Liu, Jun Xia, Qing Li, Wenqi Fan. [Paper], [Code], 2026.8
| Name | Paper | Notes |
|---|---|---|
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | A benchmark consists of ∼21k multimodal multiple choice questions with diverse science topics. |
| CMM12K | COMET: “Cone of experience” enhanced large multimodal model for mathematical problem generation | A Chinese MM SFT dataset for math, not released |
| SPIQA | SPIQA: A Dataset for Multimodal Question Answering on Scientific Papers | Designed to interpret complex figures and tables within the context of scientific research articles across various domains of computer science |
| InstructDoc | InstructDoc: A Dataset for Zero-Shot Generalization of Visual Document Understanding with Instructions | Collection of 30 publicly available VDU datasets, each with diverse instructions in a unified format. |
| M-Paper | mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model | Built by parsing Latex source files of high-quality papers. |
| DocStruct4M/DocReason25K | mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding | Based on publicly available datasets. A high-quality instruction tuning dataset. |
| DocGenome | DocGenome: An Open Large-scale Scientific Document Benchmark for Training and Testing Multi-modal Large Language Models | A structured document benchmark constructed by annotating 500K scientific documents from 153 disciplines in the arXiv. |
| ArXivCap/ArXivQA | Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models | A figure-caption dataset comprising 6.4M images and 3.9M captions, sourced from 572K ArXiv papers. A QA dataset generated by prompting GPT-4V based on scientific figures. |
| FigureQA | FigureQA: An Annotated Figure Dataset for Visual Reasoning | A visual reasoning corpus of over one million QA pairs grounded in over 100,000 images. The images are synthetic, scientific-style figures from five classes: line plots, dotline plots, vertical and horizontal bar graphs, and pie charts |
| DVQA | DVQA: Understanding Data Visualizations via Question Answering | A dataset that tests many aspects of bar chart understanding in a question answering framework. |
| SciGraphQA | SciGraphQA: A Large-Scale Synthetic Multi-Turn Question-Answering Dataset for Scientific Graphs | A synthetic multi-turn QA dataset related to academic graphs. |
| SciCap | SciCap: Generating Captions for Scientific Figures | A large-scale figure caption dataset based on Computer Science arXiv papers published between 2010 and 2020, contained over 416k figures that focused on graphplot. |
| FigCap | Figure Captioning with Reasoning and Sequence-Level Training | Generated based on FigureQA |
| FigureSeer | FigureSeer: Parsing Result-Figures in Research Papers | - |
| UniChart | UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning | A large-scale chart corpus for pretraining, covering a diverse range of visual styles and topics. |
| MapQA | MapQA: A Dataset for Question Answering on Choropleth Maps | A large-scale dataset of ~800K question-answer pairs over ~60K map images. |
| TabMWP | Dynamic Prompt Learning via Policy Gradient for Semi-structured Mathematical Reasoning | A dataset containing 38,431 open-domain grade-level problems that require mathematical reasoning on both textual and tabular data |
| CLEVR-Math | CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning | A multi-modal math word problems dataset consisting of simple math word problems involving addition/subtraction |
| GUICourse | GUICourse: From General Vision Language Model to Versatile GUI Agent | A suite of datasets to train visual-based GUI agents from general VLMs |
| PIN-14M | PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents | 14 million samples derived from Chinese and English sources, tailored to include complex web and scientific content. |
| MathV360K | Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models | 40K high-quality images with QA pairs from 24 existing datasets and synthesizing 320K new pairs. |
| MMSci | MMSci: A Multimodal Multi-Discipline Dataset for PhD-Level Scientific Comprehension | Collected a multimodal dataset from open-access scientific articles published in Nature Communications journals. |
| MAVIS-Caption/Instruct | MAVIS: Mathematical Visual Instruction Tuning | - |
| Geo170K | G-LLaVA: Solving Geometric Problem with Multi-Modal Large Language Model | Utilize the geometry characteristic to construct a multi-modal geometry dataset, building upon existing datasets. |
| SciOL/MuLMS-Img | SciOL and MuLMS-Img: Introducing A Large-Scale Multimodal Scientific Dataset and Models for Image-Text Tasks in the Scientific Domain | Pretraining corpus for multimodal models in the scientific domain. |
| PlotQA | PlotQA: Reasoning over Scientific Plots | With 28.9 million QA pairs over 224,377 plots on data from realworld sources and questions based on crowd-sourced question templates. |
| ChartInstructionData | Advancing Multimodal Large Language Models in Chart Question Answering with Visualization-Referenced Instruction Tuning | A dataset of 467K, which includes 108K table-chart pairs and 359K chart-QA pairs. |
| MMTab | Multimodal Table Understanding | Dataset for multimodal table understanding problem, based on 14 publicly available table datasets of 8 domains. |
| Multimodal Self-Instruct | Multimodal Self-Instruct: Synthetic Abstract Image and Visual Reasoning Instruction Using Language Model | Instruction dataset for eight visual scenarios: charts, tables, simulated maps, dashboards, flowcharts, relation graphs, floor plans, and visual puzzles. |
| GeoGPT4V | GeoGPT4V: Towards Geometric Multi-modal Large Language Models with Geometric Image Generation | Leverages GPT-4 and GPT-4V to generate relatively basic geometry problems with aligned text and images. |
| InfiMM-WebMath-40B | InfiMM-WebMath-40B: Advancing Multimodal Pre-Training for Enhanced Mathematical Reasoning | Interleaved image-text documents, comprises 24 million web pages, 85 million associated image URLs, and 40 billion text tokens, extracted and filtered from CommonCrawl. |
| MultiMath-300K | MultiMath: Bridging Visual and Mathematical Reasoning for Large Language Models | Spans K-12 levels with image captions and step-wise solutions. |
| MathVL | MathGLM-Vision: Solving Mathematical Problems with Multi-Modal Large Language Model | A fine-tuning dataset including both several public datasets and our curated Chinese dataset collected from K12 education levels. |
| AtomMATH | AtomThink: A Slow Thinking Framework for Multimodal Mathematical Reasoning | A large-scale multimodal dataset of long CoTs, and an atomic capability evaluation metric for mathematical tasks. |
| MAmmoTH-VL | MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale | A dataset containing 12M instruction-response pairs to cover diverse, reasoning-intensive tasks with detailed and faithful rationales. |
| BIGDOCS | BIGDOCS: AN OPEN AND PERMISSIVELY-LICENSED DATASET FOR TRAINING MULTIMODAL MODELS ON DOCUMENT AND CODE TASKS | Comprising 7.5 million multimodal documents across 30 tasks. |
| 2.5 Years in Class | 2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining | Collects over 2.5 years of instructional videos, totaling 22,000 class hours. |
| MM-PRM | MM-PRM: Enhancing Multimodal Mathematical Reasoning with Scalable Step-Level Supervision | A curated dataset of 10,000 multimodal math problems with verifiable answers, which serves as seed data. Leveraging a Monte Carlo Tree Search (MCTS)-based pipeline, generate over 700k step-level annotations without human labeling |
| BMMR | BMMR: A Large-Scale Bilingual Multimodal Multi-Discipline Reasoning Dataset | a large-scale bilingual, multimodal, multidisciplinary reasoning dataset for the community to develop and evaluate large multimodal models |
| Zebra-CoT | Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning | a diverse large-scale dataset with 182,384 samples, containing logically coherent interleaved text-image reasoning traces |
| SldprtNet | SldprtNet: A Large-Scale Multimodal Dataset for CAD Generation in Language-Driven 3D Design | A large-scale multimodal CAD dataset with over 242K industrial parts in STEP and SLDPRT formats for language-driven 3D design and CAD generation. |
| CharTide-2M | CharTide: Data-Centric Chart-to-Code Generation via Tri-Perspective Tuning and Inquiry-Driven Evolution | A 2M-sample chart-to-code training and alignment dataset built with tri-perspective tuning and inquiry-driven verification. |
| Chart2NCode | Aligned Multi-View Scripts for Universal Chart-to-Code Generation | 176K chart images paired with aligned executable scripts across multiple plotting languages for universal chart-to-code generation. |
| Zero-to-CAD | Zero-to-CAD: Agentic Synthesis of Interpretable CAD Programs at Million-Scale Without Real Data | Around one million synthetic executable CAD construction sequences plus a curated 100K high-quality subset for CAD program generation. |
| CADFS | CADFS: A Big CAD Program Dataset and Framework for Computer-Aided Design with Large Language Models | 450K real-world CAD models represented with FeatureScript and spanning 15 modeling operations for richer CAD generation. |
| DocAtlas | DocAtlas: Multilingual Document Understanding Across 80+ Languages | High-fidelity OCR datasets and benchmarks covering 82 languages and 9 document understanding evaluation tasks. |
| Sentinel2Cap | Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning | A human-annotated multimodal remote sensing captioning dataset pairing Sentinel-1 SAR and Sentinel-2 multi-spectral image patches with validated captions. |
| Ryze Evidence-Enriched Dataset | Ryze: Evidence-Enriched Data Synthesis from Biomedical Papers | An automatically synthesized biomedical SFT corpus whose QA pairs retain the supporting visual element, caption, extracted structure, and referring prose; the pipeline produces millions of domain QA tokens without human annotation. |
| LabEmbodied-Data | LabVLA: Grounding Vision-Language-Action Models in Scientific Laboratories | Success-filtered laboratory robot demonstrations with multi-camera observations, language instructions, robot states, action trajectories, and structured annotations across 16 robot platforms and four task families. |
| FusionRS | FusionRS: A Large-Scale RGB-Infrared-Style Remote Sensing Dataset for Cross-Modal Vision-Language Learning | 600K aligned RGB-infrared-style remote-sensing records with 599,992 source-text pairs and 45,913 IR-aware captions, organized into group-aware 580K/10K/10K train/validation/test splits. |
| Name | Paper | Note |
|---|---|---|
| GeoEval | GeoEval: Benchmark for Evaluating LLMs and Multi-Modal Models on Geometry Problem-Solving | An benchmark for evaluating MLLMs' capability in solving geometry math problems |
| Geometry3K | Inter-GPS: Interpretable Geometry Problem Solving with Formal Language and Symbolic Reasoning | Consisting of 3,002 geometry problems with dense annotation in formal language. |
| GEOS | Solving Geometry Problems: Combining Text and Diagram Interpretation | - |
| GeoQA | GeoQA: A Geometric Question Answering Benchmark Towards Multimodal Numerical Reasoning | 4,998 geometric problems with cor- responding annotated programs |
| GeoQA+ | An Augmented Benchmark Dataset for Geometric Question Answering through Dual Parallel Text Encoding | Based on GeoQA, newly annotate 2,518 geometric problems with richer types and greater difficulty |
| UniGeo | UniGeo: Unifying Geometry Logical Reasoning via Reformulating Mathematical Expression | Contains 4,998 calculation problems and 9,543 proving problems |
| PGPS9K | A Multi-Modal Neural Geometric Solver with Textual Clauses Parsed from Diagram | Labeled with both fine-grained diagram annotation and interpretable solution program. |
| GeomVerse | GeomVerse: A Systematic Evaluation of Large Models for Geometric Reasoning | A synthetic benchmark of geometry questions with controllable difficulty levels along multiple axes |
| MathVista | MATHVISTA: EVALUATING MATHEMATICAL REASONING OF FOUNDATION MODELS IN VISUAL CONTEXTS | A benchmark designed to combine challenges from diverse mathematical and visual tasks. |
| OlympiadBench | OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems | An Olympiad-level bilingual multimodal scientific benchmark, from mathematics and physics competitions |
| OlympicArena | OlympicArena: Benchmarking Multi-discipline Cognitive Reasoning for Superintelligent AI | Encompass a wide range of disciplines spanning seven fields and 62 international Olympic competitions. |
| SciBench | SciBench: Evaluating College-Level Scientific Problem-Solving Abilities of Large Language Models | A benchmark for college-level scientific problems sourced from instructional textbooks. |
| MMMU | MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI | Designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. |
| CMMMU | CMMMU: A Chinese Massive Multi-discipline Multimodal Understanding Benchmark | A new Chinese Massive Multi-discipline Multimodal Understanding benchmark designed to evaluate LMMs on tasks demanding college-level subject knowledge and deliberate reasoning in a Chinese context. |
| MULTI | MULTI: Multimodal Understanding Leaderboard with Text and Images | Includes over 18,000 questions, and challenges MLLMs with a variety of tasks, ranging from formula derivation to image detail analysis and cross-modality reasoning. |
| M3GIA | M3GIA: A Cognition Inspired Multilingual and Multimodal General Intelligence Ability Benchmark | Designed to evaluate multimodal models on massive multi-discipline tasks demanding college-level subject knowledge and deliberate reasoning. |
| M3Exam | M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models | Sourced from real and official human exam questions for evaluating LLMs in a multilingual, multimodal, and multilevel context. |
| MathVerse | MATHVERSE: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? | 2,612 high-quality, multi-subject math problems with diagrams from publicly available sources. |
| MATH-Vision | Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset | 3,040 high-quality mathe- matical problems with visual contexts sourced from real math competitions. |
| AI2D | A Diagram Is Worth A Dozen Images | A dataset of diagrams with annotations of constituents and relationships for over 5,000 diagrams and 15,000 QAs. |
| IconQA | IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning | A benchmark with the goal of answering a question in an icon image context. |
| TQA | Are You Smarter Than A Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension | Includes 1,076 lessons and 26,260 multi-modal questions, taken from middle school science curricula. |
| ScienceQA | Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering | A benchmark consists of ∼21k multimodal multiple choice questions with diverse science topics. |
| ChartX | ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning | A multi-modal evaluation set covering 18 chart types, 7 chart tasks, 22 disciplinary topics, and high-quality chart data |
| PlotQA | PlotQA: Reasoning over Scientific Plots | With 28.9 million question-answer pairs over 224,377 plots on data from realworld sources and questions based on crowd-sourced question templates. |
| Chart-to-text | Chart-to-Text: A Large-Scale Benchmark for Chart Summarization | A large-scale benchmark with two datasets and a total of 44,096 charts covering a wide range of topics and chart types. |
| ChartQA | ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning | A large-scale benchmark covering 9.6K human-written questions as well as 23.1K questions generated from human-written chart summaries. |
| OpenCQA | OpenCQA: Open-ended Question Answering with Charts | The goal is to answer an open-ended question about a chart with descriptive texts. |
| ChartBench | ChartBench: A Benchmark for Complex Visual Reasoning in Charts | A comprehensive benchmark designed to assess chart comprehension and data reliability through complex visual reasoning. |
| DocVQA | DocVQA: A Dataset for VQA on Document Images | Consists of 50,000 questions defined on 12,000+ document images |
| InfoVQA | InfographicVQA | Comprises a diverse collection of infographics along with question-answer annotations. |
| WTQ | Compositional Semantic Parsing on Semi-Structured Tables | A dataset of 22,033 complex questions on Wikipedia tables. |
| TableFact | TabFact : A Large-scale Dataset for Table-based Fact Verification | A large-scale dataset with 16k Wikipedia tables as the evidence for 118k human-annotated natural language statements. |
| MM-Math | MM-MATH: Advancing Multimodal Math Evaluation with Process Evaluation and Fine-grained Classification | Consists of 5,929 open-ended middle school math problems with visual contexts, with fine-grained classification. |
| MathCheck | Is Your Model Really A Good Math Reasoner? Evaluating Mathematical Reasoning with Checklist | A well-designed checklist for testing task generalization and reasoning robustness. |
| PuzzleVQA | PUZZLEVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns | A collection of 2000 puzzle instances based on abstract patterns. |
| SMART-101 | Are Deep Neural Networks SMARTer than Second Graders? | Evaluating the abstraction, deduction, and generalization abilities of neural networks in solving visul-linguistic puzzles. |
| AlgpPuzzleVQA | ARE LANGUAGE MODELS PUZZLE PRODIGIES? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning | Evaluate the capabilities in solving algorithmic puzzles. |
| ChartMimic | ChartMimic: Evaluating LMM’s Cross-Modal Reasoning Capability via Chart-to-Code Generation | Aimed at assessing the visually-grounded code generation capabilities. |
| ChartSumm | ChartSumm: A Comprehensive Benchmark for Automatic Chart Summarization of Long and Short Summaries | - |
| MMCode | MMCode: Evaluating Multi-Modal Code Large Language Models with Visually Rich Programming Problems | Contains 3,548 questions and 6,620 images collected from real-world programming challenges harvested from 10 code competition websites. |
| Design2Code | Design2Code: How Far Are We From Automating Front-End Engineering | Manually curate a benchmark of 484 diverse real-world webpages |
| Plot2Code | Plot2Code: A Comprehensive Benchmark for Evaluating Multi-modal Large Language Models in Code Generation from Scientific Plots | A comprehensive visual coding benchmark. |
| CharXiv | CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs | A comprehensive evaluation suite involving 2,323 natural, challenging, and diverse charts from arXiv papers. |
| We-Math | WE-MATH: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning? | 6.5K visual math problems, spanning 67 hierarchical knowledge concepts and 5 layers of knowledge granularity. |
| SceMQA | SceMQA: A Scientific College Entrance Level Multimodal Question Answering Benchmark | A benchmark for scientific multimodal question answering at the college entrance leve. |
| TheoremQA | TheoremQA: A Theorem-driven Question Answering dataset | Curated by domain experts containing 800 high-quality questions covering 350 theorems from Math, Physics, EE&CS, and Finance. |
| NPHardEval4V | NPHardEval4V: A Dynamic Reasoning Benchmark of Multimodal Large Language Models | Built by converting textual description of questions from NPHardEval to image representations. |
| MathScape | MathScape: Evaluating MLLMs in multimodal Math Scenarios through a Hierarchical Benchmark | Designed to evaluate photo-based math problem scenarios, assessing the theoretical understanding and application ability of MLLMs through a categorical hierarchical approach. |
| TableBench | TableBench: A Comprehensive and Complex Benchmark for Table Question Answering | Including 18 fields within four major categories of table question answering capabilities. |
| GRAB | GRAB: A Challenging GRaph Analysis Benchmark for Large Multimodal Models | Synthetic, comprised of 2170 questions, covering four tasks and 23 graph properties. |
| LogicVista | LogicVista: A Benchmark for Evaluating Multimodal Logical Reasoning | Evaluate general logical cognition abilities across 5 logical reasoning tasks encompassing 9 different capabilities, using a sample of 448 multiple-choice questions. |
| CMM-Math | CMM-Math: A Chinese Multimodal Math Dataset To Evaluate and Enhance the Mathematics Reasoning of Large Multimodal Models | Contains over 28,000 high-quality samples, featuring a variety of problem types with detailed solutions across 12 grade levels from elementary to high school in China. |
| SWE-bench Multimodal | SWE-BENCH MULTIMODAL: DO AI SYSTEMS GENERALIZE TO VISUAL SOFTWARE DOMAINS? | Contains 617 task instances collected from 17 JavaScript libraries used for web interface design, diagramming, data visualization, syntax highlighting, and interactive mapping. |
| MMIE | MMIE: MASSIVE MULTIMODAL INTERLEAVED COMPREHENSION BENCHMARK FOR LARGE VISIONLANGUAGE MODELS | Comprises 20K meticulously curated multimodal queries, spanning 3 categories, 12 fields, and 102 subfields, including mathematics, coding, physics, literature, health, and arts. |
| MultiChartQA | MultiChartQA: Benchmarking Vision-Language Models on Multi-Chart Problems | Multi-hop reasoning required to extract and integrate information from multiple charts, comprises 655 charts and 944 questions |
| Sketch2Code | Sketch2Code: Evaluating Vision-Language Models for Interactive Web Design Prototyping | Evaluating automating the conversion of rudimentary sketches into webpage prototypes, collected a total of 731 sketches for 484 webpage screenshots |
| PolyMath | POLYMATH: A CHALLENGING MULTI-MODAL MATHEMATICAL REASONING BENCHMARK | Comprises 5,000 manually collected high-quality images of cognitive textual and visual challenges across 10 distinct categories, including pattern recognition, spatial reasoning, and relative reasoning |
| VisAidMath | VisAidMath: Benchmarking Visual-Aided Mathematical Reasoning | Includes 1,200 challenging problems from various mathematical branches, vision-aid formulations, and difficulty levels, collected from diverse sources such as textbooks, examination papers, and Olympiad problems |
| DYNAMATH | DYNAMATH: A DYNAMIC VISUAL BENCHMARK FOR EVALUATING MATHEMATICAL REASONING ROBUSTNESS OF VISION LANGUAGE MODELS | Includes 501 high-quality, multi-topic seed questions, each represented as a Python program |
| M3SCIQA | M3SCIQA: A Multi-Modal Multi-Document Scientific QA Benchmark for Evaluating Foundation Models | Consists of 1,452 expert-annotated questions spanning 70 natural language processing paper clusters |
| M-LONGDOC | M-LONGDOC: A BENCHMARK FOR MULTIMODAL SUPER-LONG DOCUMENT UNDERSTANDING AND A RETRIEVAL-AWARE TUNING FRAMEWORK | A benchmark of 851 samples, and an automated framework to evaluate the performance of large multimodal models |
| VisOnlyQA | VisOnlyQA: Large Vision Language Models Still Struggle with Visual Perception of Geometric Information | Includes 1,200 multiple-choice questions in 12 tasks on four categories of figures. Designed to directly evaluate the visual perception capabilities. |
| U-MATH | U-MATH: A UNIVERSITY-LEVEL BENCHMARK FOR EVALUATING MATHEMATICAL SKILLS IN LLMS | 1,100 unpublished open-ended university-level problems sourced from teaching materials. It is balanced across six core subjects, with 20% of multimodal problems. |
| DrawEduMath | DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images | An English-language dataset of 2,030 images of students' handwritten responses to K-12 math problems |
| MM-IQ | MM-IQ: Benchmarking Human-Like Abstraction and Reasoning in Multimodal Models | An evaluation framework comprising 2,710 meticulously curated test items spanning 8 distinct reasoning paradigms |
| LOST IN TIME | LOST IN TIME: CLOCK AND CALENDAR UNDERSTANDING CHALLENGES IN MULTIMODAL LLMS | Curated a structured dataset comprising two subsets: ClockQA and CalendarQA |
| ProJudge | ProJudge: A Multi-Modal Multi-Discipline Benchmark and Instruction-Tuning Dataset for MLLM-based Process Judges | Comprises 2,400 test cases and 50,118 step-level labels, spanning four scientific disciplines with diverse difficulty levels and multimodal content |
| MPBench | MPBench: A Comprehensive Multimodal Reasoning Benchmark for Process Errors Identification | A comprehensive, multi-task, multimodal benchmark designed to systematically assess the effectiveness of PRMs in diverse scenarios |
| FlowVerse | MathFlow: Enhancing the Perceptual Flow of MLLMs for Visual Mathematical Problems | A comprehensive benchmark that categorizes all information used during problem-solving into four components |
| ChartQAPRO | ChartQA PRO : A More Diverse and Challenging Benchmark for Chart Question Answering | A new benchmark that includes 1,341 charts from 157 diverse sources, spanning various chart types. 1,948 questions in various types, to better reflect real-world challenges |
| ChartMuseum | ChartMuseum: Testing Visual Reasoning Capabilities of Large Vision-Language Models | Chart Question Answering (QA) benchmark containing 1,162 expert-annotated questions spanning multiple reasoning types, curated from realworld charts across 184 sources |
| FullFront | FullFront: Benchmarking MLLMs Across the Full Front-End Engineering Workflow | Assesses three fundamental tasks that map directly to the front-end engineering pipeline: Webpage Design, Webpage Perception QA, and Webpage Code Generation |
| MMMR | MMMR: Benchmarking Massive Multi-Modal Reasoning Tasks | 1,083 questions spanning six diverse reasoning types with symbolic depth and multi-hop demands |
| VideoMathQA | VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos | Spans 10 diverse mathematical domains, covering videos ranging from 10 seconds to over 1 hour. |
| Scientists’ First Exam | Scientists’ First Exam: Probing Cognitive Abilities of MLLM via Perception, Understanding, and Reasoning | Comprises 830 expert-verified VQA pairs across three question types, spanning 66 multimodal tasks across five high-value disciplines. |
| SCIVER | SCIVER: Evaluating Foundation Models for Multimodal Scientific Claim Verification | Consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification. |
| MATHREAL | MATHREAL: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models | Comprising 2,000 mathematical questions with images captured by handheld mobile devices in authentic scenarios. |
| WE-MATH 2.0 | WE-MATH 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning | A comprehensive benchmark covering all 491 knowledge points with diverse reasoning step distributions. |
| Uni-MMMU | Uni-MMMU: A Massive Multi-discipline Multimodal Unified Benchmark | A benchmark of eight bidirectionally coupled tasks that enforce Gen–Und logical dependency. |
| GGBench | GGBench: A Geometric Generative Reasoning Benchmark for Unified Multimodal Models | A benchmark designed specifically to evaluate geometric generative reasoning. |
| ScratchMath | Can MLLMs Read Students' Minds? Unpacking Multimodal Error Analysis in Handwritten Math | 1,720 authentic handwritten math scratchwork samples for error cause explanation and classification across seven error types. |
| GeoAux-Bench | Thinking with Constructions: A Benchmark and Policy Optimization for Visual-Text Interleaved Geometric Reasoning | A benchmark for visual-text interleaved geometric reasoning where models must decide when and how to construct visual aids. |
| MultihopSpatial | MultihopSpatial: Multi-hop Compositional Spatial Reasoning Benchmark for Vision-Language Model | Evaluates multi-hop compositional spatial reasoning and precise visual grounding for VLM and VLA-style agents. |
| AEC-Bench | AEC-Bench: A Multimodal Benchmark for Agentic Systems in Architecture, Engineering, and Construction | Real-world AEC tasks requiring drawing understanding, cross-sheet reasoning, and project-level coordination. |
| GeoMMBench | GeoMMBench and GeoMMAgent: Toward Expert-Level Multimodal Intelligence in Geoscience and Remote Sensing | 1,053 expert-level image-based multiple-choice questions spanning geoscience and remote sensing disciplines. |
| HM-Bench | HM-Bench: A Comprehensive Benchmark for Multimodal Large Language Models in Hyperspectral Remote Sensing | A hyperspectral remote sensing benchmark for testing spectral-spatial perception and reasoning in MLLMs. |
| PaperScope | PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers | Evaluates agentic deep research over many scientific papers with evidence from text, tables, and figures. |
| MMCoIR | CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval | A multimodal code information retrieval benchmark across five visual domains, eight programming languages, and eleven libraries. |
| ReactBench | ReactBench: A Benchmark for Topological Reasoning in MLLMs on Chemical Reaction Diagrams | Probes topological reasoning over branching, converging, and cyclic chemical reaction diagrams. |
| ArXivDoc | Document-as-Image Representations Fall Short for Scientific Retrieval | A scientific document retrieval benchmark built from LaTeX sources to test text, tables, figures, and equations. |
| MathNet | MathNet: A Global Multimodal Benchmark for Mathematical Reasoning and Retrieval | 30K+ Olympiad-level math problems from 47 countries, with a retrieval benchmark for mathematical problem search. |
| STEP-STEM | Unveiling Fine-Grained Visual Traces: Evaluating Multimodal Interleaved Reasoning Chains in Multimodal STEM Tasks | Evaluates fine-grained visual traces in multimodal interleaved reasoning chains for STEM tasks. |
| OMIBench | OMIBench: Benchmarking Olympiad-Level Multi-Image Reasoning in Large Vision-Language Model | An Olympiad-level benchmark where evidence is distributed across multiple images. |
| AstroVLBench | A Systematic Evaluation of Vision-Language Models for Observational Astronomical Reasoning Tasks | 4,100+ expert-verified instances across optical imaging, radio interferometry, photometry, time-domain signals, and spectroscopy. |
| SpecVQA | SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images | A scientific spectral image VQA benchmark covering seven representative spectrum types with expert annotations. |
| TopBench | TopBench: A Benchmark for Implicit Prediction and Reasoning over Tabular Question Answering | Tests implicit prediction and latent-intent reasoning over table question answering rather than simple retrieval. |
| Text-to-CAD Retrieval | Text-to-CAD Retrieval: A Strong Baseline | Establishes a cross-modal retrieval benchmark for finding semantically relevant CAD models from natural-language queries. |
| AstroAlertBench | AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification | Evaluates multimodal LLM accuracy, reasoning, and honesty on astronomical alert classification. |
| TableVista | TableVista: Benchmarking Multimodal Table Reasoning under Visual and Structural Complexity | 3,000 table reasoning problems expanded into diverse visual variants for robustness and structural complexity testing. |
| ChartREG++ | ChartREG++: Towards Benchmarking and Improving Chart Referring Expression Grounding under Diverse Referring Clues and Multi-Target Referring | A chart referring expression grounding benchmark with diverse referring clues and multi-target localization. |
| CADBench | CADBench: A Multimodal Benchmark for AI-Assisted CAD Program Generation | 18,000 evaluation samples across six CAD benchmark families for multimodal CAD program generation. |
| Vision2Code | Vision2Code: A Multi-Domain Benchmark for Evaluating Image-to-Code Generation | A reference-code-free benchmark and evaluation framework for image-to-code generation across multiple visual domains. |
| UHR-Micro | UHR-Micro: Diagnosing and Mitigating the Resolution Illusion in Earth Observation VLMs | Diagnoses micro-target perception failures in ultra-high-resolution Earth observation VLMs. |
| CiteVQA | CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence | Tests whether document VQA models cite and ground the visual or textual evidence supporting their answers. |
| IndustryBench-MIPU | IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products | 4,559 industrial products, 27,652 images, and 103,703 attribute annotations across 18 categories for evaluating multi-image technical specification extraction. |
| DashboardMimic | Dashboard2Code: Evaluating Multimodal Models on Reconstructing Interactive Dashboards | 180 manually verified Plotly+Dash dashboard-code pairs spanning three difficulty levels, 20 visualization types, and eight interaction patterns, with dynamic tests for visual and behavioral fidelity. |
| SVLAT for MLLMs | Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy | Evaluates six MLLMs on a standardized 49-item scientific-visualization literacy test built from 18 visualizations, eight techniques, and 11 task types, with comparison data from 485 human participants. |
| TableParseMap | From Diagnosis to Correction: Benchmarking and Improving Real-World Table Parsing | 916 real-world complex tables organized into five challenging scenarios and nine failure types, accompanied by a 1,977-table Consensus-Hard Set for cross-parser evaluation. |
| SonarBench | SonarLLM: A Native Sonar-Optical Multimodal Large Language Model for Underwater Perception | A paired sonar-optical benchmark covering recognition, counting, VQA, and captioning across 25 subsets, with controlled optical degradation to isolate cross-modal complementarity under turbidity. |
If you have any question about this opinionated list, do not hesitate to create an issue.
70 commits
1 commits