EvoCodeBench is an evolutionary code generation benchmark aligned with real-world code repositories.
Details of EvoCodeBench can be found in our paper "EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-world Code Repositories".
Paper: https://arxiv.org/pdf/2404.00599.pdf
Data and Source Code: https://github.com/seketeam/EvoCodeBench
Compared to existing benchmarks (e.g., HumanEval), EvoCodeBench has the following features:
A sample in EvoCodeBench is shown as follows.

Based on EvoCodeBench, we propose repository-level code generation and evaluate 10 popular LLMs (e.g., gpt-4, gpt-3.5, DeepSeek Coder, StarCoder 2, CodeLLaMa, Gemma, and Qwen 1.5). All prompts and models' completions have been released for further community analysis.
The evaluation results are shown as follows.

9 commits
EvoCodeBench is an evolutionary code generation benchmark aligned with real-world code repositories.
Details of EvoCodeBench can be found in our paper "EvoCodeBench: An Evolving Code Generation Benchmark Aligned with Real-world Code Repositories".
Paper: https://arxiv.org/pdf/2404.00599.pdf
Data and Source Code: https://github.com/seketeam/EvoCodeBench
Compared to existing benchmarks (e.g., HumanEval), EvoCodeBench has the following features:
A sample in EvoCodeBench is shown as follows.

Based on EvoCodeBench, we propose repository-level code generation and evaluate 10 popular LLMs (e.g., gpt-4, gpt-3.5, DeepSeek Coder, StarCoder 2, CodeLLaMa, Gemma, and Qwen 1.5). All prompts and models' completions have been released for further community analysis.
The evaluation results are shown as follows.

9 commits