nebius/SWE-rebench

Dataset

73

stars

20

commits

2

linked in READMEs

Dec 23, 2025

updated

README

Dataset Summary

SWE-rebench is a large-scale dataset designed to support training and evaluation of LLM-based software engineering (SWE) agents, building upon and expanding our earlier release, SWE-bench-extra. It is constructed using a fully automated pipeline that continuously extracts real-world interactive SWE tasks from GitHub repositories at scale, as detailed in our paper SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. The dataset currently comprises over 21,000 issue–pull request pairs from 3,400+ Python repositories, each validated for correctness through automated environment setup and test execution. A curated subset of these tasks also forms the basis of our continuously updated SWE-rebench leaderboard. SWE-rebench builds upon and extends the methodology of SWE-bench by incorporating several key enhancements detailed in our paper, including:

  • A fully automated pipeline for continuous task collection.
  • LLM-driven extraction and validation of environment installation instructions.
  • An automated LLM-based task quality assessment pipeline that annotates tasks with labels such as clarity, complexity, or test patch validity.

We’ve released 7,500 pre-built Docker images used in our RL pipeline. They’re publicly available on Docker Hub. You do not need to build them yourself.

News

[2025/08/05] Uploaded the corresponding Docker images for 7,500 tasks to Docker Hub.

How to Use

from datasets import load_dataset
ds = load_dataset('nebius/SWE-rebench')

Dataset Structure

The SWE-rebench dataset schema extends the original SWE-bench schema with additional fields to support richer analysis. The complete schema is detailed in the table below. For more information about this data and methodology behind collecting it, please refer to our paper.

Field nameTypeDescription
instance_idstrA formatted instance identifier, usually as repo_owner__repo_name-PR-number.
patchstrThe gold patch, the patch generated by the PR (minus test-related code), that resolved the issue.
repostrThe repository owner/name identifier from GitHub.
base_commitstrThe commit hash of the repository representing the HEAD of the repository before the solution PR is applied.
hints_textstrComments made on the issue prior to the creation of the solution PR’s first commit creation date.
created_atstrThe creation date of the pull request.
test_patchstrA test-file patch that was contributed by the solution PR.
problem_statementstrThe issue title and body.
versionstrInstallation version to use for running evaluation.
environment_setup_commitstrCommit hash to use for environment setup and installation.
FAIL_TO_PASSstrA JSON list of strings that represent the set of tests resolved by the PR and tied to the issue resolution.
PASS_TO_PASSstrA JSON list of strings that represent tests that should pass before and after the PR application.
metastrA JSON dictionary indicating whether the instance is lite, along with a list of failed lite validators if it is not.
license_namestrThe type of license of the repository.
install_configstrInstallation configuration for setting up the repository.
requirementsstrFreezed requirements for the repository.
environmentstrEnvironment configuration for the repository.

To execute tasks from SWE-rebench (i.e., set up their environments, apply patches, and run tests), we provide a fork of the original SWE-bench execution framework, adapted for our dataset's structure and features. Our fork is based on the SWE-bench framework, specifically from its Release 4.0.3. The primary modification introduces functionality to source environment installation constants directly from the install_config field present in each task instance within SWE-rebench. This allows for more flexible and task-specific environment setups.

You can find the details of this modification in the following commit:

To build the necessary Docker images and run agents on SWE-rebench tasks, you have two main options:

  1. Use our SWE-bench fork directly: Clone the fork and utilize its scripts for building images and executing tasks. The framework will automatically use the install_config from each task.
  2. Integrate similar functionality into your existing codebase: If you have your own execution framework based on SWE-bench or a different system, you can adapt it by implementing a similar mechanism to parse and utilize the install_config field from the SWE-rebench task instances. The aforementioned commit can serve as a reference for this integration.

License

The dataset is licensed under the Creative Commons Attribution 4.0 license. However, please respect the license of each specific repository on which a particular instance is based. To facilitate this, the license of each repository at the time of the commit is provided for every instance.

Citation

@misc{badertdinov2025swerebenchautomatedpipelinetask,
      title={SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents}, 
      author={Ibragim Badertdinov and Alexander Golubev and Maksim Nekrashevich and Anton Shevtsov and Simon Karasik and Andrei Andriushchenko and Maria Trofimova and Daria Litvintseva and Boris Yangel},
      year={2025},
      eprint={2505.20411},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2505.20411}
}

Contributors

ibragim-bad

19 commits

nielsr

1 commits

nebius/SWE-rebench

Dataset

73

stars

20

commits

2

linked in READMEs

Dec 23, 2025

updated

README

Dataset Summary

SWE-rebench is a large-scale dataset designed to support training and evaluation of LLM-based software engineering (SWE) agents, building upon and expanding our earlier release, SWE-bench-extra. It is constructed using a fully automated pipeline that continuously extracts real-world interactive SWE tasks from GitHub repositories at scale, as detailed in our paper SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. The dataset currently comprises over 21,000 issue–pull request pairs from 3,400+ Python repositories, each validated for correctness through automated environment setup and test execution. A curated subset of these tasks also forms the basis of our continuously updated SWE-rebench leaderboard. SWE-rebench builds upon and extends the methodology of SWE-bench by incorporating several key enhancements detailed in our paper, including:

  • A fully automated pipeline for continuous task collection.
  • LLM-driven extraction and validation of environment installation instructions.
  • An automated LLM-based task quality assessment pipeline that annotates tasks with labels such as clarity, complexity, or test patch validity.

We’ve released 7,500 pre-built Docker images used in our RL pipeline. They’re publicly available on Docker Hub. You do not need to build them yourself.

News

[2025/08/05] Uploaded the corresponding Docker images for 7,500 tasks to Docker Hub.

How to Use

from datasets import load_dataset
ds = load_dataset('nebius/SWE-rebench')

Dataset Structure

The SWE-rebench dataset schema extends the original SWE-bench schema with additional fields to support richer analysis. The complete schema is detailed in the table below. For more information about this data and methodology behind collecting it, please refer to our paper.

Field nameTypeDescription
instance_idstrA formatted instance identifier, usually as repo_owner__repo_name-PR-number.
patchstrThe gold patch, the patch generated by the PR (minus test-related code), that resolved the issue.
repostrThe repository owner/name identifier from GitHub.
base_commitstrThe commit hash of the repository representing the HEAD of the repository before the solution PR is applied.
hints_textstrComments made on the issue prior to the creation of the solution PR’s first commit creation date.
created_atstrThe creation date of the pull request.
test_patchstrA test-file patch that was contributed by the solution PR.
problem_statementstrThe issue title and body.
versionstrInstallation version to use for running evaluation.
environment_setup_commitstrCommit hash to use for environment setup and installation.
FAIL_TO_PASSstrA JSON list of strings that represent the set of tests resolved by the PR and tied to the issue resolution.
PASS_TO_PASSstrA JSON list of strings that represent tests that should pass before and after the PR application.
metastrA JSON dictionary indicating whether the instance is lite, along with a list of failed lite validators if it is not.
license_namestrThe type of license of the repository.
install_configstrInstallation configuration for setting up the repository.
requirementsstrFreezed requirements for the repository.
environmentstrEnvironment configuration for the repository.

To execute tasks from SWE-rebench (i.e., set up their environments, apply patches, and run tests), we provide a fork of the original SWE-bench execution framework, adapted for our dataset's structure and features. Our fork is based on the SWE-bench framework, specifically from its Release 4.0.3. The primary modification introduces functionality to source environment installation constants directly from the install_config field present in each task instance within SWE-rebench. This allows for more flexible and task-specific environment setups.

You can find the details of this modification in the following commit:

To build the necessary Docker images and run agents on SWE-rebench tasks, you have two main options:

  1. Use our SWE-bench fork directly: Clone the fork and utilize its scripts for building images and executing tasks. The framework will automatically use the install_config from each task.
  2. Integrate similar functionality into your existing codebase: If you have your own execution framework based on SWE-bench or a different system, you can adapt it by implementing a similar mechanism to parse and utilize the install_config field from the SWE-rebench task instances. The aforementioned commit can serve as a reference for this integration.

License

The dataset is licensed under the Creative Commons Attribution 4.0 license. However, please respect the license of each specific repository on which a particular instance is based. To facilitate this, the license of each repository at the time of the commit is provided for every instance.

Citation

@misc{badertdinov2025swerebenchautomatedpipelinetask,
      title={SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents}, 
      author={Ibragim Badertdinov and Alexander Golubev and Maksim Nekrashevich and Anton Shevtsov and Simon Karasik and Andrei Andriushchenko and Maria Trofimova and Daria Litvintseva and Boris Yangel},
      year={2025},
      eprint={2505.20411},
      archivePrefix={arXiv},
      primaryClass={cs.SE},
      url={https://arxiv.org/abs/2505.20411}
}

Contributors

ibragim-bad

19 commits

nielsr

1 commits