panilya/awesome-ai-benchmarks

Awesome AI Benchmarks

TypeScript

52

99 commits

updated Aug 28, 2026

See the code

README

Awesome AI Benchmarks Awesome

A curated list of 100+ AI benchmarks across various domains including Agent Capabilities, Reasoning, Translation, Code Generation, Multimodal, and other AI domains.

πŸ“‹ Table of Contents

πŸ“Š Statistics

  • Total Benchmarks: 114
  • Categories: 6
  • Subcategories: 24
  • Last Updated: 2026-07-28

πŸš€ Website

Visit our automatically generated website for a better browsing experience with search, filtering, and detailed benchmark information.

πŸ“ Contributing

We welcome contributions! Please read our CONTRIBUTING.md for guidelines on how to add new benchmarks or improve existing entries.

πŸ“š Benchmarks

Programming & Code Generation

Code Generation & Evaluation

API & Tool Usage

Terminal & Environment

Logic & Reasoning

Database & Query

Multimodal & Vision

Video Understanding

Multimodal Evaluation

OCR & Document Understanding

Translation & Multilingual

General Translation Evaluation

Domain-Specific Translation

Multilingual Reasoning

Domain-Specific

Agriculture

Cultural & Social Understanding

  • WorldView-Bench - A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

Industry 4.0 Applications

  • AssetOpsBench - A Benchmark for developing, orchestrating, and evaluating domain-specific AI agents in industrial asset operations and maintenance

Agent Capabilities & Reasoning

Web Agents

Swarm Intelligence

Long-term Coherence

Scientific & Academic Reasoning

Security & Robustness

Business & CRM

World Modeling & Simulation

Game & Interactive

Creative & Evaluation

Memory & Episodic Tasks

Creative Writing

  • Longform Creative Writing - Benchmark for evaluating long-form creative writing capabilities

  • Creative Writing v3 - Enhanced creative writing evaluation benchmark

Judgment & Analysis

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

  • Thanks to all the researchers and organizations who created these benchmarks
  • Inspired by other "awesome" lists in the open source community

Contributors

panilya

59 commits

actions-user

31 commits

mauricekleine

2 commits

panilya/awesome-ai-benchmarks

Awesome AI Benchmarks

TypeScript

52

99 commits

updated Aug 28, 2026

See the code

README

Awesome AI Benchmarks Awesome

A curated list of 100+ AI benchmarks across various domains including Agent Capabilities, Reasoning, Translation, Code Generation, Multimodal, and other AI domains.

πŸ“‹ Table of Contents

πŸ“Š Statistics

  • Total Benchmarks: 114
  • Categories: 6
  • Subcategories: 24
  • Last Updated: 2026-07-28

πŸš€ Website

Visit our automatically generated website for a better browsing experience with search, filtering, and detailed benchmark information.

πŸ“ Contributing

We welcome contributions! Please read our CONTRIBUTING.md for guidelines on how to add new benchmarks or improve existing entries.

πŸ“š Benchmarks

Programming & Code Generation

Code Generation & Evaluation

API & Tool Usage

Terminal & Environment

Logic & Reasoning

Database & Query

Multimodal & Vision

Video Understanding

Multimodal Evaluation

OCR & Document Understanding

Translation & Multilingual

General Translation Evaluation

Domain-Specific Translation

Multilingual Reasoning

Domain-Specific

Agriculture

Cultural & Social Understanding

  • WorldView-Bench - A Benchmark for Evaluating Global Cultural Perspectives in Large Language Models

Industry 4.0 Applications

  • AssetOpsBench - A Benchmark for developing, orchestrating, and evaluating domain-specific AI agents in industrial asset operations and maintenance

Agent Capabilities & Reasoning

Web Agents

Swarm Intelligence

Long-term Coherence

Scientific & Academic Reasoning

Security & Robustness

Business & CRM

World Modeling & Simulation

Game & Interactive

Creative & Evaluation

Memory & Episodic Tasks

Creative Writing

  • Longform Creative Writing - Benchmark for evaluating long-form creative writing capabilities

  • Creative Writing v3 - Enhanced creative writing evaluation benchmark

Judgment & Analysis

License

This project is licensed under the MIT License - see the LICENSE file for details.

Acknowledgments

  • Thanks to all the researchers and organizations who created these benchmarks
  • Inspired by other "awesome" lists in the open source community

Contributors

panilya

59 commits

actions-user

31 commits

mauricekleine

2 commits

Languages

TypeScript

74.8%

JavaScript

17.6%

CSS

7.6%