dxzxy12138/PhysReason

PhysReason Becnhmark

Python

19

27 commits

updated Jul 8, 2025

See the code

README

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

arXiv Dataset Project Page

PhysReason is accepted by ACL-2025-main

πŸ“‹ Overview

PhysReason is a comprehensive physics-based reasoning benchmark consisting of 1,200 physics problems spanning multiple domains, with a focus on both knowledge-based (25%) and reasoning-based (75%) questions. This benchmark addresses the critical gap in evaluating large language models' capabilities in physics-based reasoning, which requires applying physics theorems and constraints in complex problem-solving scenarios.

✨ Key Features

  • πŸ“Š Dataset Size: 1,200 carefully curated physics problems
  • 🎯 Problem Types: Strategic mix of knowledge-based (25%) and reasoning-based (75%) questions
  • πŸ“š Theorem Coverage: Comprehensive coverage of 147 physics theorems
  • 🎨 Visual Content: 81% of problems include diagrams and visual elements
  • πŸ“ˆ Difficulty Levels: Four distinct levels - Knowledge, Easy, Medium, Hard
  • πŸ”„ Step-by-step Solutions: Average of 8.1 solution steps per problem (15.6 for hard problems)
  • 🌍 Multi-modal: Supports both text and image inputs

πŸ”§ Data Collection

Our rigorous data collection process ensures high-quality, challenging problems:

  • πŸ“– Sources: Global college entrance exams and international physics competitions
  • βš™οΈ Process: Standardized using MinerU framework for consistent formatting
  • βœ… Quality Control: Two-phase translation process with expert verification
  • πŸ” Filtering: Systematically excluded easily searchable problems to prevent data leakage
  • πŸ“Š Classification: Difficulty levels based on solving time and theorem complexity analysis

πŸ“Š Benchmark Comparison

BenchmarkMulti-modalSizeKnowledgeQuestion TypeAvg. TStep-by-stepAvg. TAvg. S
JEEBench❌123CEEOE,MC169.7---
MMLU-Pro❌1299COLMC52.1---
GPQA❌227PH.D.OE111.4❌197.23.6
SciEval❌1657-OE,MC154.5---
SciBenchβœ…295COLOE80.5❌315.92.8
MMMUβœ…443COLOE,MC53.8---
ScienceQAβœ…617K1-K12MC13.3❌63.02.4
OlympiadBenchβœ…2334COMPOE222.0❌199.83.7
EMMAβœ…156-MC109.5---
Ours-Knowledgeβœ…300CEE+COMPOE163.7βœ…196.53.3
Ours-Easyβœ…300CEE+COMPOE171.2βœ…241.55.0
Ours-Mediumβœ…300CEE+COMPOE229.2βœ…391.38.4
Ours-Hardβœ…300CEE+COMPOE340.9βœ…936.115.6
Ours-Fullβœ…1200CEE+COMPOE226.3βœ…441.38.1

πŸ” Evaluation Framework

We introduce the Physics Solution Auto Scoring (PSAS) framework with two complementary evaluation approaches:

PSAS-A (Answer Level Evaluation)

  • Sub-question Assessment: Evaluates answers for each sub-question independently
  • LLM-based Extraction: Uses advanced language models for answer extraction
  • Semantic Verification: Ensures semantic consistency between extracted and ground truth answers
  • Weighted Scoring: Considers solution step lengths as weights for different sub-questions

PSAS-S (Step Level Evaluation)

Provides detailed step-by-step assessment through four phases:

  1. Data Extraction: Parses model responses and reference solutions
  2. Scoring: Evaluates correctness of each reasoning step
  3. First Error Detection: Identifies where models first deviate from correct reasoning
  4. Error Analysis: Classifies error types into four key bottlenecks:
    • Physics Theorem Application
    • Physics Process Understanding
    • Calculation
    • Physics Condition Analysis

πŸš€ Usage

Core Evaluation Files

  • answer_evaluation_with_ds_ch_prompt.py: Answer-level evaluation using Chinese prompts
  • answer_evaluation_with_ds_en_prompt.py: Answer-level evaluation using English prompts
  • format_result_ds.py: Optimizes unstable outputs into stable, consistent formats
  • step_evaluation_with_ds_ch_prompt.py: Step-level evaluation using Chinese prompts
  • step_evaluation_with_ds_en_prompt.py: Step-level evaluation using English prompts

πŸ“ˆ Experimental Results

Non-O-like Models Performance

ModelInputKnowledgeEasyMediumHardAvg.
Qwen2VL-72BQ, I41.92/62.4724.04/45.2615.97/36.134.83/24.2316.96/42.88
InternVL2.5-78BQ, I28.34/64.7124.16/50.6917.72/38.569.71/25.9519.98/45.89
GPT-4oQ, I50.71/65.8233.87/51.9822.73/42.3611.03/24.7129.58/47.23
Deepseek-V3-671BQ, IC55.86/66.1440.06/52.7726.63/44.0213.73/26.8734.07/48.42
Claude-3.5-SonnetQ, I54.14/66.4541.35/55.8528.14/44.8615.11/28.5134.69/49.88
Gemini-2.0-FlashQ, I65.08/75.0454.84/68.6039.79/55.6721.99/38.3945.20/60.40
Gemini-2.0-ProQ, I67.99/79.0155.43/71.4744.29/57.7423.81/42.6647.88/62.74

O-like Models Performance

ModelInputKnowledgeEasyMediumHardAvg.
o1-miniQ, IC53.90/65.7435.21/52.2622.24/40.1910.61/26.8030.49/47.18
QvQ-72BQ, I62.44/70.9253.74/64.6528.18/54.8814.30/36.4732.67/57.66
Gemini-2.0-Flash-Thinking-1206Q, I65.35/77.2051.89/67.4944.43/58.9527.14/45.4847.20/63.07
QwQ-32BQ, IC62.03/76.2854.92/71.0843.64/62.1422.99/42.1945.89/63.87
GLM-ZeroQ, IC64.95/80.3654.11/71.5441.32/63.6723.04/47.4646.52/65.76
o3-mini-highQ, IC70.67/83.6167.20/81.9545.31/64.5730.12/47.2353.32/69.34
Gemini-2.0-Flash-Thinking-0121Q, I73.44/84.1563.17/75.9450.41/66.6031.90/48.4754.73/69.73
Deepseek-R1Q, IC75.11/85.9165.08/79.8154.84/72.0231.95/51.5056.75/73.26

PhysReason-mini Results

ModelK.E.M.H.Avg.
o1-mini54.8030.3315.417.9227.11
QvQ-72B51.1737.1029.8322.1335.06
QwQ-32B64.4050.0738.8827.4545.20
Gemini-2.0-Flash-Thinking-120671.4749.9736.8322.9745.42
GLM-Zero72.7050.1743.4224.7047.75
o172.4753.3749.3125.3250.12
o3-mini-high71.1063.2047.0231.9353.31
Gemini-2.0-Flash-Thinking-012176.3356.8751.8532.6154.42
Deepseek-R185.1760.7747.2433.2356.60

πŸ”‘ Key Findings

  • Performance Gap: Even top-performing models achieve less than 60% on answer-level evaluation
  • Difficulty Scaling: Performance drops significantly from knowledge questions (75.11%) to hard problems (31.95%)
  • O-like Model Advantage: Models with enhanced reasoning capabilities show superior performance
  • Multi-modal Benefits: Visual content significantly enhances model understanding and performance
  • Four Critical Bottlenecks identified through step-level evaluation:
    1. Physics Theorem Application
    2. Physics Process Understanding
    3. Calculation Accuracy
    4. Physics Condition Analysis

πŸ“ Citation

If you find PhysReason useful in your research, please cite our paper:

@article{zhang2025physreason,
  title={Physreason: A comprehensive benchmark towards physics-based reasoning},
  author={Zhang, Xinyu and Dong, Yuxuan and Wu, Yanrui and Huang, Jiaxing and Jia, Chengyou and Fernando, Basura and Shou, Mike Zheng and Zhang, Lingling and Liu, Jun},
  journal={arXiv preprint arXiv:2502.12054},
  year={2025}
}

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ“§ Contact

We welcome contributions to PhysReason! Please contact us for more details.


πŸ”— Quick Links:

dxzxy12138/PhysReason

PhysReason Becnhmark

Python

19

27 commits

updated Jul 8, 2025

See the code

README

PhysReason: A Comprehensive Benchmark towards Physics-Based Reasoning

arXiv Dataset Project Page

PhysReason is accepted by ACL-2025-main

πŸ“‹ Overview

PhysReason is a comprehensive physics-based reasoning benchmark consisting of 1,200 physics problems spanning multiple domains, with a focus on both knowledge-based (25%) and reasoning-based (75%) questions. This benchmark addresses the critical gap in evaluating large language models' capabilities in physics-based reasoning, which requires applying physics theorems and constraints in complex problem-solving scenarios.

✨ Key Features

  • πŸ“Š Dataset Size: 1,200 carefully curated physics problems
  • 🎯 Problem Types: Strategic mix of knowledge-based (25%) and reasoning-based (75%) questions
  • πŸ“š Theorem Coverage: Comprehensive coverage of 147 physics theorems
  • 🎨 Visual Content: 81% of problems include diagrams and visual elements
  • πŸ“ˆ Difficulty Levels: Four distinct levels - Knowledge, Easy, Medium, Hard
  • πŸ”„ Step-by-step Solutions: Average of 8.1 solution steps per problem (15.6 for hard problems)
  • 🌍 Multi-modal: Supports both text and image inputs

πŸ”§ Data Collection

Our rigorous data collection process ensures high-quality, challenging problems:

  • πŸ“– Sources: Global college entrance exams and international physics competitions
  • βš™οΈ Process: Standardized using MinerU framework for consistent formatting
  • βœ… Quality Control: Two-phase translation process with expert verification
  • πŸ” Filtering: Systematically excluded easily searchable problems to prevent data leakage
  • πŸ“Š Classification: Difficulty levels based on solving time and theorem complexity analysis

πŸ“Š Benchmark Comparison

BenchmarkMulti-modalSizeKnowledgeQuestion TypeAvg. TStep-by-stepAvg. TAvg. S
JEEBench❌123CEEOE,MC169.7---
MMLU-Pro❌1299COLMC52.1---
GPQA❌227PH.D.OE111.4❌197.23.6
SciEval❌1657-OE,MC154.5---
SciBenchβœ…295COLOE80.5❌315.92.8
MMMUβœ…443COLOE,MC53.8---
ScienceQAβœ…617K1-K12MC13.3❌63.02.4
OlympiadBenchβœ…2334COMPOE222.0❌199.83.7
EMMAβœ…156-MC109.5---
Ours-Knowledgeβœ…300CEE+COMPOE163.7βœ…196.53.3
Ours-Easyβœ…300CEE+COMPOE171.2βœ…241.55.0
Ours-Mediumβœ…300CEE+COMPOE229.2βœ…391.38.4
Ours-Hardβœ…300CEE+COMPOE340.9βœ…936.115.6
Ours-Fullβœ…1200CEE+COMPOE226.3βœ…441.38.1

πŸ” Evaluation Framework

We introduce the Physics Solution Auto Scoring (PSAS) framework with two complementary evaluation approaches:

PSAS-A (Answer Level Evaluation)

  • Sub-question Assessment: Evaluates answers for each sub-question independently
  • LLM-based Extraction: Uses advanced language models for answer extraction
  • Semantic Verification: Ensures semantic consistency between extracted and ground truth answers
  • Weighted Scoring: Considers solution step lengths as weights for different sub-questions

PSAS-S (Step Level Evaluation)

Provides detailed step-by-step assessment through four phases:

  1. Data Extraction: Parses model responses and reference solutions
  2. Scoring: Evaluates correctness of each reasoning step
  3. First Error Detection: Identifies where models first deviate from correct reasoning
  4. Error Analysis: Classifies error types into four key bottlenecks:
    • Physics Theorem Application
    • Physics Process Understanding
    • Calculation
    • Physics Condition Analysis

πŸš€ Usage

Core Evaluation Files

  • answer_evaluation_with_ds_ch_prompt.py: Answer-level evaluation using Chinese prompts
  • answer_evaluation_with_ds_en_prompt.py: Answer-level evaluation using English prompts
  • format_result_ds.py: Optimizes unstable outputs into stable, consistent formats
  • step_evaluation_with_ds_ch_prompt.py: Step-level evaluation using Chinese prompts
  • step_evaluation_with_ds_en_prompt.py: Step-level evaluation using English prompts

πŸ“ˆ Experimental Results

Non-O-like Models Performance

ModelInputKnowledgeEasyMediumHardAvg.
Qwen2VL-72BQ, I41.92/62.4724.04/45.2615.97/36.134.83/24.2316.96/42.88
InternVL2.5-78BQ, I28.34/64.7124.16/50.6917.72/38.569.71/25.9519.98/45.89
GPT-4oQ, I50.71/65.8233.87/51.9822.73/42.3611.03/24.7129.58/47.23
Deepseek-V3-671BQ, IC55.86/66.1440.06/52.7726.63/44.0213.73/26.8734.07/48.42
Claude-3.5-SonnetQ, I54.14/66.4541.35/55.8528.14/44.8615.11/28.5134.69/49.88
Gemini-2.0-FlashQ, I65.08/75.0454.84/68.6039.79/55.6721.99/38.3945.20/60.40
Gemini-2.0-ProQ, I67.99/79.0155.43/71.4744.29/57.7423.81/42.6647.88/62.74

O-like Models Performance

ModelInputKnowledgeEasyMediumHardAvg.
o1-miniQ, IC53.90/65.7435.21/52.2622.24/40.1910.61/26.8030.49/47.18
QvQ-72BQ, I62.44/70.9253.74/64.6528.18/54.8814.30/36.4732.67/57.66
Gemini-2.0-Flash-Thinking-1206Q, I65.35/77.2051.89/67.4944.43/58.9527.14/45.4847.20/63.07
QwQ-32BQ, IC62.03/76.2854.92/71.0843.64/62.1422.99/42.1945.89/63.87
GLM-ZeroQ, IC64.95/80.3654.11/71.5441.32/63.6723.04/47.4646.52/65.76
o3-mini-highQ, IC70.67/83.6167.20/81.9545.31/64.5730.12/47.2353.32/69.34
Gemini-2.0-Flash-Thinking-0121Q, I73.44/84.1563.17/75.9450.41/66.6031.90/48.4754.73/69.73
Deepseek-R1Q, IC75.11/85.9165.08/79.8154.84/72.0231.95/51.5056.75/73.26

PhysReason-mini Results

ModelK.E.M.H.Avg.
o1-mini54.8030.3315.417.9227.11
QvQ-72B51.1737.1029.8322.1335.06
QwQ-32B64.4050.0738.8827.4545.20
Gemini-2.0-Flash-Thinking-120671.4749.9736.8322.9745.42
GLM-Zero72.7050.1743.4224.7047.75
o172.4753.3749.3125.3250.12
o3-mini-high71.1063.2047.0231.9353.31
Gemini-2.0-Flash-Thinking-012176.3356.8751.8532.6154.42
Deepseek-R185.1760.7747.2433.2356.60

πŸ”‘ Key Findings

  • Performance Gap: Even top-performing models achieve less than 60% on answer-level evaluation
  • Difficulty Scaling: Performance drops significantly from knowledge questions (75.11%) to hard problems (31.95%)
  • O-like Model Advantage: Models with enhanced reasoning capabilities show superior performance
  • Multi-modal Benefits: Visual content significantly enhances model understanding and performance
  • Four Critical Bottlenecks identified through step-level evaluation:
    1. Physics Theorem Application
    2. Physics Process Understanding
    3. Calculation Accuracy
    4. Physics Condition Analysis

πŸ“ Citation

If you find PhysReason useful in your research, please cite our paper:

@article{zhang2025physreason,
  title={Physreason: A comprehensive benchmark towards physics-based reasoning},
  author={Zhang, Xinyu and Dong, Yuxuan and Wu, Yanrui and Huang, Jiaxing and Jia, Chengyou and Fernando, Basura and Shou, Mike Zheng and Zhang, Lingling and Liu, Jun},
  journal={arXiv preprint arXiv:2502.12054},
  year={2025}
}

πŸ“„ License

This project is licensed under the MIT License - see the LICENSE file for details.

πŸ“§ Contact

We welcome contributions to PhysReason! Please contact us for more details.


πŸ”— Quick Links:

Languages

Python

71.9%

HTML

28.1%