1 repo
Tools and benchmarks for evaluating language models on software engineering tasks, particularly code generation, bug fixing, and repository-level problem solving. The cluster centers on SWE-bench, a large-scale benchmark for evaluating LLMs on real GitHub issues, along with agents and frameworks that use LLMs to automatically resolve software engineering problems. These repositories focus on measuring and improving how well AI systems can understand, analyze, and modify real-world codebases.