facebookresearch/swe-sweep

How many bugs can LMs find & fix in large codebases?

Python

10

4 commits

updated Oct 1, 2026

See the code

See what people are saying

README

SWE-sweep logo
SWE-sweep

How many bugs can LMs find & fix in large codebases?

Given a real repository, an agent must discover & repair as many bugs as they can.
Agents are not given any hint about the type of bug or its location.

Quickstart

[!note] The code in this repo is only a thin wrapper around Harbor, however see the warning below regarding harbor version

[!warning]

  • Harbor version: Harbor 0.23 does not yet support the separate verifier environments and collect hooks used by these tasks. The project therefore pins a compatible revision from Harbor's source repository until Harbor 0.24 is released.
  • Final score: We build a micro-average of the bugs resolved (equivalent to a weighted average of the individual task scores). See evaluation notes.

We recommend uv for managing Python environments.

git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync   # or pip install .

Verify your setup:

sweep infra doctor
Development setup

Clone the repository and install the editable package with its dev dependencies:

git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync --extra test

Run the test suite (matches CI, which tests on Python 3.12 and 3.13):

uv run pytest

Run the linting/formatting hooks:

uvx pre-commit run --all-files

Evaluate a single task

Pass a task name and a unified diff against that task's base commit:

sweep eval <task> <path/to/submission.diff>  # see raw harbor command below
Raw harbor command
uv run harbor run \
--path tasks/pandora-bench__dateutil \
--agent swesweep.patch_agent:PatchAgent \
--agent-kwarg "patch_path=$(realpath /path/to/model.patch)" \
--jobs-dir jobs \
--n-concurrent 1 \
--yes

The command runs the corresponding Harbor task with a small patch-applying agent. Harbor then evaluates the resulting checkout in the task's separate verifier environment. Results are written under jobs/ by default.

How evaluation works

Evaluation follows the following pseudo-code:

reset_to_base_commit()
apply(agent_patch)
reset(test_files)
build_if_needed()
base_commit_results = run_visible_suite()  # test -> pass/fail

subtask_results = {}   # subtask -> {test -> pass/fail}
for subtask in subtasks:
  apply(subtask.test_patch)
  subtask_results[subtask.id] = run(subtask.hidden_tests)
  revert(subtask.test_patch)

if any_new_failures(base_commit_results):
  task_score = 0
else:
  task_score = sum(passed_subtasks) / len(subtasks)

Calculate the final score

The benchmark score is calculated as

total number of bugs solved across tasks / total number of bugs =
    = mean(number of bugs in task * task score)

Summarize one or more directories containing graded Harbor trials:

sweep info path/to/graded-solutions --per-task
sweep info path/to/graded-solutions --json

Citation

@misc{lieret2026swesweep,
  title  = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
  author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
            Yuxiang Wei and Jeremy Ma and Sten Sootla and
            Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
            Rui Hou and Ofir Press and John Yang},
  year   = {2026},
  note   = {Preprint},
  url    = {https://github.com/facebookresearch/swe-sweep}
}

License

SWE-sweep is licensed under the terms of the license found in LICENSE.

ai
ai-agents
benchmark
benchmarking
harbor
harbor-framework
llm
llm-benchmark
llm-benchmarking
llm-benchmarks
swe-bench

facebookresearch/swe-sweep

How many bugs can LMs find & fix in large codebases?

Python

10

4 commits

updated Oct 1, 2026

See the code

See what people are saying

README

SWE-sweep logo
SWE-sweep

How many bugs can LMs find & fix in large codebases?

Given a real repository, an agent must discover & repair as many bugs as they can.
Agents are not given any hint about the type of bug or its location.

Quickstart

[!note] The code in this repo is only a thin wrapper around Harbor, however see the warning below regarding harbor version

[!warning]

  • Harbor version: Harbor 0.23 does not yet support the separate verifier environments and collect hooks used by these tasks. The project therefore pins a compatible revision from Harbor's source repository until Harbor 0.24 is released.
  • Final score: We build a micro-average of the bugs resolved (equivalent to a weighted average of the individual task scores). See evaluation notes.

We recommend uv for managing Python environments.

git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync   # or pip install .

Verify your setup:

sweep infra doctor
Development setup

Clone the repository and install the editable package with its dev dependencies:

git clone https://github.com/facebookresearch/swe-sweep.git
cd swe-sweep
uv sync --extra test

Run the test suite (matches CI, which tests on Python 3.12 and 3.13):

uv run pytest

Run the linting/formatting hooks:

uvx pre-commit run --all-files

Evaluate a single task

Pass a task name and a unified diff against that task's base commit:

sweep eval <task> <path/to/submission.diff>  # see raw harbor command below
Raw harbor command
uv run harbor run \
--path tasks/pandora-bench__dateutil \
--agent swesweep.patch_agent:PatchAgent \
--agent-kwarg "patch_path=$(realpath /path/to/model.patch)" \
--jobs-dir jobs \
--n-concurrent 1 \
--yes

The command runs the corresponding Harbor task with a small patch-applying agent. Harbor then evaluates the resulting checkout in the task's separate verifier environment. Results are written under jobs/ by default.

How evaluation works

Evaluation follows the following pseudo-code:

reset_to_base_commit()
apply(agent_patch)
reset(test_files)
build_if_needed()
base_commit_results = run_visible_suite()  # test -> pass/fail

subtask_results = {}   # subtask -> {test -> pass/fail}
for subtask in subtasks:
  apply(subtask.test_patch)
  subtask_results[subtask.id] = run(subtask.hidden_tests)
  revert(subtask.test_patch)

if any_new_failures(base_commit_results):
  task_score = 0
else:
  task_score = sum(passed_subtasks) / len(subtasks)

Calculate the final score

The benchmark score is calculated as

total number of bugs solved across tasks / total number of bugs =
    = mean(number of bugs in task * task score)

Summarize one or more directories containing graded Harbor trials:

sweep info path/to/graded-solutions --per-task
sweep info path/to/graded-solutions --json

Citation

@misc{lieret2026swesweep,
  title  = {{SWE-sweep}: Can Agents Autonomously Find and Fix Bugs?},
  author = {Kilian Lieret and Jeffrey Jian Ma and Rahul Kindi and
            Yuxiang Wei and Jeremy Ma and Sten Sootla and
            Parth Thakkar and Chao Beyond Zhou and Pengcheng Yin and
            Rui Hou and Ofir Press and John Yang},
  year   = {2026},
  note   = {Preprint},
  url    = {https://github.com/facebookresearch/swe-sweep}
}

License

SWE-sweep is licensed under the terms of the license found in LICENSE.

ai
ai-agents
benchmark
benchmarking
harbor
harbor-framework
llm
llm-benchmark
llm-benchmarking
llm-benchmarks
swe-bench

Languages

Python

88.1%

Dockerfile

10.0%

Shell

1.8%