12 repos
SWE-bench/SWE-bench
SWE-bench: Can Language Models Resolve Real-world Github Issues?
5,868
709 commits
princeton-nlp/SWE-bench
727 commits
SWE-rebench/SWE-bench-fork
Fork to run instances from SWE-rebench
30
586 commits
Elfsong/Mercury
Code Efficiency Benchmark
87
64 commits
nebius/SWE-rebench
Dataset Summary
73
20 commits
ScalingIntelligence/codemonkeys
No description
58
1 commits
SWE-agent/SWE-agent
SWE-agent takes a GitHub issue and tries to automatically fix it, using your LM of choice. It can…
20,350
2,087 commits
microsoft/OpenRCA
[ICLR'25] OpenRCA: Can Large Language Models Locate the Root Cause of Software Failures?
423
23 commits
amazon-science/SWE-PolyBench
SWE-PolyBench: A multi-language benchmark for repository level evaluation of coding agents
89
83 commits
RUCAIBox/SWE-World
52
6 commits
Vaibhavi1707/unibench_code
SWEAgents and CodeLLM Evaluation Framework
0
2 commits
KoYejune0302/bob-swe
16 commits