A curated list of 100+ AI benchmarks across various domains including Agent Capabilities, Reasoning, Translation, Code Generation, Multimodal, and other AI domains.
Visit our automatically generated website for a better browsing experience with search, filtering, and detailed benchmark information.
We welcome contributions! Please read our CONTRIBUTING.md for guidelines on how to add new benchmarks or improve existing entries.
LiveCodeBench - A comprehensive benchmark for evaluating code generation capabilities of large language models on real-world programming tasks
HumanEval - Evaluating Large Language Models Trained on Code
BigCodeBench - Comprehensive benchmark for code generation and understanding
SciCode - A Research Coding Benchmark Curated by Scientists
Aider Polyglot - Multilingual coding benchmark for AI assistants
can-ai-code - Simple evaluation framework for code generation models with multiple programming languages
Kotlin-bench - Benchmark specifically designed for evaluating Kotlin programming capabilities
Gorilla - Large Language Model Connected with Massive APIs for function calling evaluation
ToolSandbox - A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Ο-bench - A Benchmark for Tool-Agent-User Interaction in Real-World Domains
ΟΒ²-bench - Evaluating Conversational Agents in a Dual-Control Environment
Sudoku-Bench - Benchmark testing logical reasoning capabilities through Sudoku puzzle solving
NonoBench - A benchmark suite for evaluating LLM performance on Nonogram (Picross) puzzle solving across different grid sizes
Tests as Prompt - A Test-Driven-Development Benchmark for LLM Code Generation
Web-Bench - A LLM Code Benchmark Based on Web Standards and Frameworks
KernelBench - Can LLMs Write GPU Kernels?
CypherBench - Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
Spider 2.0 - Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
BIRD - A comprehensive benchmark for large-scale database grounded text-to-SQL evaluation with 12,751+ question-SQL pairs across 95 databases spanning 37 professional domains
LongVideoBench - A Benchmark for Long-context Interleaved Video-Language Understanding
Video-MME - The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
MLVU - Multi-task Long Video Understanding Benchmark
Physics IQ Benchmark - Do generative video models understand physical principles?
VBench - Comprehensive Benchmark Suite for Video Generative Models
VideoScore2 - Interpretable generated-video evaluation across visual quality, text-video alignment, and physical/common-sense consistency
VBench-2.0 - Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
FAVOR-Bench - A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
VideoPhy - Evaluating Physical Commonsense for Video Generation
VideoPhy 2 - Challenging Action-Centric Physical Commonsense Evaluation of Video Generation
TempCompass - Do Video LLMs Really Understand Videos?
GenVidBench - A Challenging Benchmark for Detecting AI-Generated Video
PhyGenBench - Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
IPV-Bench - Impossible Videos
VidCapBench - A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
StructEval - Benchmarking LLMs' structured-output generation and conversion across 18 text and visual formats; StructEval-V evaluates renderable HTML, React, SVG, Canvas, and visualization-code outputs with structural and VQA checks
MLLM-as-a-Judge - Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Vision Language Models are Biased - Benchmark examining bias in vision-language models
ManipBench - Benchmarking Vision-Language Models for Low-Level Robot Manipulation
ENIGMAEVAL - A Benchmark of Long Multimodal Reasoning Challenges
VisuLogic - A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
PHYBench - Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
EmbodiedBench - Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
SEED-Bench - Benchmarking Multimodal Large Language Models
LOKI - A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
CapArena - Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
Intelligent Document Processing (IDP) Leaderboard - Comprehensive evaluation of document processing capabilities
olmOCR-Bench - Benchmark for evaluating optical character recognition capabilities
Benchmark: Testing AI Models On Engineering Drawings - An evaluation of leading AI models on their ability to extract dimensional and tolerance data from real-world mechanical engineering drawings.
Evaluating the Translation Performance of Large Language Models Based on Euas-20 - Comprehensive evaluation of LLM translation capabilities using the Euas-20 dataset
DiBiMT - A Gold Evaluation Benchmark for Studying Lexical Ambiguity in Machine Translation
TransBench - Multilingual Translation Leaderboard for Industrial-Scale Applications
FRMT - A Benchmark for Few-Shot Region-Aware Machine Translation
VNTL Leaderboard - Japanese Visual Novels into English translating leaderboard for evaluating translation quality
SwiLTra-Bench - The Swiss Legal Translation Benchmark for evaluating legal document translation
DATETIME - A new benchmark to measure LLM translation and reasoning capabilities with temporal expressions
MultiNRC - A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
LEXam - Benchmarking Legal Reasoning on 340 Law Exams from Swiss, EU, and international law examinations
CaseLaw - Legal case law analysis and reasoning benchmark
SwarmBench - Benchmarking LLMs' Swarm Intelligence
Realistic Evaluations for Agents Leaderboard (REAL) - Comprehensive agent evaluation framework
Vending-Bench - A Benchmark for Long-Term Coherence of Autonomous Agents
SUPER - Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
TravelPlanner - A Benchmark for Real-World Planning with Language Agents
LongProc - Benchmarking Long-Context Language Models on Long Procedural Generation
OlympiadBench - A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
ZebraLogic - Benchmarking the Logical Reasoning Ability of Language Models
TruthfulQA - Measuring How Models Imitate Human Falsehoods (New version)
MathChat - Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
MMLU-Pro - Enhanced version of the Massive Multitask Language Understanding benchmark
SuperGPQA - Scaling LLM Evaluation across 285 Graduate Disciplines
PhysBench - Benchmarking and Enhancing VLMs for Physical World Understanding
MATH-Perturb - Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
PhD Knowledge Not Required - A Reasoning Challenge for Large Language Models
Gravity-Bench-v1 - A Benchmark on Gravitational Physics Discovery for Agents
MMLU - Measuring Massive Multitask Language Understanding
MATH - Measuring Mathematical Problem Solving With the MATH Dataset
GPQA - A Graduate-Level Google-Proof Q&A Benchmark
DROP - A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
MGSM - Multilingual Grade School Math Benchmark (MGSM), Language Models are Multilingual Chain-of-Thought Reasoners
FrontierMath - A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
MuSR - Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
AIME Benchmark - American Invitational Mathematics Examination benchmark
Humanity's Last Exam - Comprehensive evaluation benchmark for advanced AI capabilities
ProcessBench - Identifying Process Errors in Mathematical Reasoning
SimpleQA - Measuring short-form factuality in large language models
BrowseComp - A Simple Yet Challenging Benchmark for Browsing Agents
HealthBench - Evaluating Large Language Models Towards Improved Human Health
QuestBench - Can LLMs ask the right question to acquire information in reasoning tasks?
MedAgentsBench - Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
Can Language Models Falsify? - Evaluating Algorithmic Reasoning with Counterexample Creation
BIG-Bench Extra Hard - Enhanced version of the BIG-Bench benchmark with more challenging tasks
JailbreakBench - An Open Robustness Benchmark for Jailbreaking Large Language Models
SnitchBench - Benchmark for evaluating model safety and information leakage
DIF - A Framework for Benchmarking and Verifying Implicit Bias in LLMs
BenchClaw - 10-dimension benchmark suite for evaluating AI agent capabilities including reasoning, knowledge, creativity, coding, ethics, efficiency, robustness, communication, collaboration, and adaptability
CRMArena-Pro - Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
CXMArena - Unified Dataset to benchmark performance in realistic CXM Scenarios
LLMs Battle Snake - Game-based evaluation using Snake gameplay
ARC-AGI-1 - Abstraction and Reasoning Corpus for Artificial General Intelligence
ARC-AGI-2 - A New Challenge for Frontier AI Reasoning Systems
Factorio Learning Environment - Complex game environment for AI agent training and evaluation
Balrog - Benchmarking Agentic LLM and VLM Reasoning On Games
WeirdML Benchmark v1 (Archived) - Unconventional machine learning challenges
WeirdML Benchmark v2 - Unconventional machine learning challenges
PokerBench - Training Large Language Models to become Professional Poker Players
GameArena - Evaluating LLM Reasoning through Live Computer Games
Longform Creative Writing - Benchmark for evaluating long-form creative writing capabilities
Creative Writing v3 - Enhanced creative writing evaluation benchmark
EQ-Bench 3 - Emotional Intelligence Benchmarks for LLMs
Judgemark v2 - Benchmark for evaluating judgment and decision-making capabilities
BuzzBench - Humor analysis benchmark for evaluating understanding of comedy and wit
RealCritic - Towards Effectiveness-Driven Evaluation of Language Model Critiques
PRMBench - A Fine-grained and Challenging Benchmark for Process-Level Reward Models
AHA Leaderboard - AI - Human alignment
This project is licensed under the MIT License - see the LICENSE file for details.
TypeScript
74.8%
JavaScript
17.6%
CSS
7.6%
A curated list of 100+ AI benchmarks across various domains including Agent Capabilities, Reasoning, Translation, Code Generation, Multimodal, and other AI domains.
Visit our automatically generated website for a better browsing experience with search, filtering, and detailed benchmark information.
We welcome contributions! Please read our CONTRIBUTING.md for guidelines on how to add new benchmarks or improve existing entries.
LiveCodeBench - A comprehensive benchmark for evaluating code generation capabilities of large language models on real-world programming tasks
HumanEval - Evaluating Large Language Models Trained on Code
BigCodeBench - Comprehensive benchmark for code generation and understanding
SciCode - A Research Coding Benchmark Curated by Scientists
Aider Polyglot - Multilingual coding benchmark for AI assistants
can-ai-code - Simple evaluation framework for code generation models with multiple programming languages
Kotlin-bench - Benchmark specifically designed for evaluating Kotlin programming capabilities
Gorilla - Large Language Model Connected with Massive APIs for function calling evaluation
ToolSandbox - A Stateful, Conversational, Interactive Evaluation Benchmark for LLM Tool Use Capabilities
Ο-bench - A Benchmark for Tool-Agent-User Interaction in Real-World Domains
ΟΒ²-bench - Evaluating Conversational Agents in a Dual-Control Environment
Sudoku-Bench - Benchmark testing logical reasoning capabilities through Sudoku puzzle solving
NonoBench - A benchmark suite for evaluating LLM performance on Nonogram (Picross) puzzle solving across different grid sizes
Tests as Prompt - A Test-Driven-Development Benchmark for LLM Code Generation
Web-Bench - A LLM Code Benchmark Based on Web Standards and Frameworks
KernelBench - Can LLMs Write GPU Kernels?
CypherBench - Towards Precise Retrieval over Full-scale Modern Knowledge Graphs in the LLM Era
Spider 2.0 - Evaluating Language Models on Real-World Enterprise Text-to-SQL Workflows
BIRD - A comprehensive benchmark for large-scale database grounded text-to-SQL evaluation with 12,751+ question-SQL pairs across 95 databases spanning 37 professional domains
LongVideoBench - A Benchmark for Long-context Interleaved Video-Language Understanding
Video-MME - The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
MLVU - Multi-task Long Video Understanding Benchmark
Physics IQ Benchmark - Do generative video models understand physical principles?
VBench - Comprehensive Benchmark Suite for Video Generative Models
VideoScore2 - Interpretable generated-video evaluation across visual quality, text-video alignment, and physical/common-sense consistency
VBench-2.0 - Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness
FAVOR-Bench - A Comprehensive Benchmark for Fine-Grained Video Motion Understanding
VideoPhy - Evaluating Physical Commonsense for Video Generation
VideoPhy 2 - Challenging Action-Centric Physical Commonsense Evaluation of Video Generation
TempCompass - Do Video LLMs Really Understand Videos?
GenVidBench - A Challenging Benchmark for Detecting AI-Generated Video
PhyGenBench - Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation
IPV-Bench - Impossible Videos
VidCapBench - A Comprehensive Benchmark of Video Captioning for Controllable Text-to-Video Generation
StructEval - Benchmarking LLMs' structured-output generation and conversion across 18 text and visual formats; StructEval-V evaluates renderable HTML, React, SVG, Canvas, and visualization-code outputs with structural and VQA checks
MLLM-as-a-Judge - Assessing Multimodal LLM-as-a-Judge with Vision-Language Benchmark
Vision Language Models are Biased - Benchmark examining bias in vision-language models
ManipBench - Benchmarking Vision-Language Models for Low-Level Robot Manipulation
ENIGMAEVAL - A Benchmark of Long Multimodal Reasoning Challenges
VisuLogic - A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
PHYBench - Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
EmbodiedBench - Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
SEED-Bench - Benchmarking Multimodal Large Language Models
LOKI - A Comprehensive Synthetic Data Detection Benchmark using Large Multimodal Models
CapArena - Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
Intelligent Document Processing (IDP) Leaderboard - Comprehensive evaluation of document processing capabilities
olmOCR-Bench - Benchmark for evaluating optical character recognition capabilities
Benchmark: Testing AI Models On Engineering Drawings - An evaluation of leading AI models on their ability to extract dimensional and tolerance data from real-world mechanical engineering drawings.
Evaluating the Translation Performance of Large Language Models Based on Euas-20 - Comprehensive evaluation of LLM translation capabilities using the Euas-20 dataset
DiBiMT - A Gold Evaluation Benchmark for Studying Lexical Ambiguity in Machine Translation
TransBench - Multilingual Translation Leaderboard for Industrial-Scale Applications
FRMT - A Benchmark for Few-Shot Region-Aware Machine Translation
VNTL Leaderboard - Japanese Visual Novels into English translating leaderboard for evaluating translation quality
SwiLTra-Bench - The Swiss Legal Translation Benchmark for evaluating legal document translation
DATETIME - A new benchmark to measure LLM translation and reasoning capabilities with temporal expressions
MultiNRC - A Challenging and Native Multilingual Reasoning Evaluation Benchmark for LLMs
LEXam - Benchmarking Legal Reasoning on 340 Law Exams from Swiss, EU, and international law examinations
CaseLaw - Legal case law analysis and reasoning benchmark
SwarmBench - Benchmarking LLMs' Swarm Intelligence
Realistic Evaluations for Agents Leaderboard (REAL) - Comprehensive agent evaluation framework
Vending-Bench - A Benchmark for Long-Term Coherence of Autonomous Agents
SUPER - Evaluating Agents on Setting Up and Executing Tasks from Research Repositories
TravelPlanner - A Benchmark for Real-World Planning with Language Agents
LongProc - Benchmarking Long-Context Language Models on Long Procedural Generation
OlympiadBench - A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
ZebraLogic - Benchmarking the Logical Reasoning Ability of Language Models
TruthfulQA - Measuring How Models Imitate Human Falsehoods (New version)
MathChat - Benchmarking Mathematical Reasoning and Instruction Following in Multi-Turn Interactions
MMLU-Pro - Enhanced version of the Massive Multitask Language Understanding benchmark
SuperGPQA - Scaling LLM Evaluation across 285 Graduate Disciplines
PhysBench - Benchmarking and Enhancing VLMs for Physical World Understanding
MATH-Perturb - Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations
PhD Knowledge Not Required - A Reasoning Challenge for Large Language Models
Gravity-Bench-v1 - A Benchmark on Gravitational Physics Discovery for Agents
MMLU - Measuring Massive Multitask Language Understanding
MATH - Measuring Mathematical Problem Solving With the MATH Dataset
GPQA - A Graduate-Level Google-Proof Q&A Benchmark
DROP - A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs
MGSM - Multilingual Grade School Math Benchmark (MGSM), Language Models are Multilingual Chain-of-Thought Reasoners
FrontierMath - A Benchmark for Evaluating Advanced Mathematical Reasoning in AI
MuSR - Testing the Limits of Chain-of-thought with Multistep Soft Reasoning
AIME Benchmark - American Invitational Mathematics Examination benchmark
Humanity's Last Exam - Comprehensive evaluation benchmark for advanced AI capabilities
ProcessBench - Identifying Process Errors in Mathematical Reasoning
SimpleQA - Measuring short-form factuality in large language models
BrowseComp - A Simple Yet Challenging Benchmark for Browsing Agents
HealthBench - Evaluating Large Language Models Towards Improved Human Health
QuestBench - Can LLMs ask the right question to acquire information in reasoning tasks?
MedAgentsBench - Benchmarking Thinking Models and Agent Frameworks for Complex Medical Reasoning
Can Language Models Falsify? - Evaluating Algorithmic Reasoning with Counterexample Creation
BIG-Bench Extra Hard - Enhanced version of the BIG-Bench benchmark with more challenging tasks
JailbreakBench - An Open Robustness Benchmark for Jailbreaking Large Language Models
SnitchBench - Benchmark for evaluating model safety and information leakage
DIF - A Framework for Benchmarking and Verifying Implicit Bias in LLMs
BenchClaw - 10-dimension benchmark suite for evaluating AI agent capabilities including reasoning, knowledge, creativity, coding, ethics, efficiency, robustness, communication, collaboration, and adaptability
CRMArena-Pro - Holistic Assessment of LLM Agents Across Diverse Business Scenarios and Interactions
CXMArena - Unified Dataset to benchmark performance in realistic CXM Scenarios
LLMs Battle Snake - Game-based evaluation using Snake gameplay
ARC-AGI-1 - Abstraction and Reasoning Corpus for Artificial General Intelligence
ARC-AGI-2 - A New Challenge for Frontier AI Reasoning Systems
Factorio Learning Environment - Complex game environment for AI agent training and evaluation
Balrog - Benchmarking Agentic LLM and VLM Reasoning On Games
WeirdML Benchmark v1 (Archived) - Unconventional machine learning challenges
WeirdML Benchmark v2 - Unconventional machine learning challenges
PokerBench - Training Large Language Models to become Professional Poker Players
GameArena - Evaluating LLM Reasoning through Live Computer Games
Longform Creative Writing - Benchmark for evaluating long-form creative writing capabilities
Creative Writing v3 - Enhanced creative writing evaluation benchmark
EQ-Bench 3 - Emotional Intelligence Benchmarks for LLMs
Judgemark v2 - Benchmark for evaluating judgment and decision-making capabilities
BuzzBench - Humor analysis benchmark for evaluating understanding of comedy and wit
RealCritic - Towards Effectiveness-Driven Evaluation of Language Model Critiques
PRMBench - A Fine-grained and Challenging Benchmark for Process-Level Reward Models
AHA Leaderboard - AI - Human alignment
This project is licensed under the MIT License - see the LICENSE file for details.
TypeScript
74.8%
JavaScript
17.6%
CSS
7.6%