LLM-based Code Understanding and Benchmarking

1 repo

Tools and benchmarks for evaluating language models on software engineering tasks, particularly code generation, bug fixing, and repository-level problem solving. The cluster centers on SWE-bench, a large-scale benchmark for evaluating LLMs on real GitHub issues, along with agents and frameworks that use LLMs to automatically resolve software engineering problems. These repositories focus on measuring and improving how well AI systems can understand, analyze, and modify real-world codebases.

Python · 1
ai ·643
code-gen ·643
code-search ·643
gpt-4 ·643
python ·643
tree-sitter ·643
LLM-based Code Understanding and Benchmarking — Shadowgraph