tongye98/Awesome-Code-Benchmark

A comprehensive code domain benchmark review of LLM researches.

245

141 commits

updated Sep 16, 2026

See the code

README

👨‍💻 Awesome Code Benchmark

Awesome PRs Welcome

A comprehensive code domain benchmark review of LLM researches.

Oryx Video-ChatGPT

Table of Contents

Taxonomy

This list organizes code benchmarks by primary capability and software-engineering workflow. Each benchmark entry includes compact metadata for task type, granularity, interaction pattern, and evaluation method.

DimensionValues
GranularityFunction, API, Query / Database, File, Project, UI, Repository, Workflow
InteractionSingle-turn, Multi-turn, Agentic, Async, Multi-agent
EvaluationUnit Tests, Execution, Performance, Human, LLM-as-Judge, Security Exploit, Economic
EnvironmentNone, Sandbox, Terminal, Browser, IDE, CI
FreshnessStatic, Dynamic, Held-out, Contamination-resistant

Surveys

  1. Software Development Life Cycle Perspective A Survey of Benchmarks for Code Large Language Models and Agents from Xi’an Jiaotong University

  2. Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks from Zhejiang University

  3. A Survey on Large Language Model Benchmarks from Shenzhen Key Laboratory for High Performance Data Mining

🚀 Benchmark Categories

Repository & Agentic Software Engineering

Program Repair, Testing & Debugging

Security, Reliability & Robustness

Code Understanding, Search & Review

Performance Optimization

Frontend, UI & Visual-Interactive Development

Code Generation & Completion

awesome
benchmarks
bug-fixing
code-completion
code-efficiency
code-generation
codellm
codellms
data-science
llm-benchmarking
multimodal
reasoning

Contributors

tongye98

91 commits

luferw

33 commits

miko99jh

6 commits

SeekingDream

4 commits

tongye98/Awesome-Code-Benchmark

A comprehensive code domain benchmark review of LLM researches.

245

141 commits

updated Sep 16, 2026

See the code

README

👨‍💻 Awesome Code Benchmark

Awesome PRs Welcome

A comprehensive code domain benchmark review of LLM researches.

Oryx Video-ChatGPT

Table of Contents

Taxonomy

This list organizes code benchmarks by primary capability and software-engineering workflow. Each benchmark entry includes compact metadata for task type, granularity, interaction pattern, and evaluation method.

DimensionValues
GranularityFunction, API, Query / Database, File, Project, UI, Repository, Workflow
InteractionSingle-turn, Multi-turn, Agentic, Async, Multi-agent
EvaluationUnit Tests, Execution, Performance, Human, LLM-as-Judge, Security Exploit, Economic
EnvironmentNone, Sandbox, Terminal, Browser, IDE, CI
FreshnessStatic, Dynamic, Held-out, Contamination-resistant

Surveys

  1. Software Development Life Cycle Perspective A Survey of Benchmarks for Code Large Language Models and Agents from Xi’an Jiaotong University

  2. Assessing and Advancing Benchmarks for Evaluating Large Language Models in Software Engineering Tasks from Zhejiang University

  3. A Survey on Large Language Model Benchmarks from Shenzhen Key Laboratory for High Performance Data Mining

🚀 Benchmark Categories

Repository & Agentic Software Engineering

Program Repair, Testing & Debugging

Security, Reliability & Robustness

Code Understanding, Search & Review

Performance Optimization

Frontend, UI & Visual-Interactive Development

Code Generation & Completion

awesome
benchmarks
bug-fixing
code-completion
code-efficiency
code-generation
codellm
codellms
data-science
llm-benchmarking
multimodal
reasoning

Contributors

tongye98

91 commits

luferw

33 commits

miko99jh

6 commits

SeekingDream

4 commits