SWE-bench/SWE-smith

Dataset

[12/14/2025] NOTE: We will no longer actively update this dataset.

57

28 commits

1 linked in READMEs

updated Dec 14, 2025

See the code

README

CodePaperSite

[12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo.

The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents.

All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation.

agents
code
software-engineering

Contributors

john-b-yang

25 commits

JO
John

2 commits

nielsr

1 commits

SWE-bench/SWE-smith

Dataset

[12/14/2025] NOTE: We will no longer actively update this dataset.

57

28 commits

1 linked in READMEs

updated Dec 14, 2025

See the code

README

CodePaperSite

[12/14/2025] NOTE: We will no longer actively update this dataset. While this dataset is still functional and usable, we recommend you use the `SWE-bench/SWE-smith-[lang]` datasets. For better maintainability and ease-of-use, we are maintaining language-specific datasets in lieu of this mono-repo.

The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents.

All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation.

agents
code
software-engineering

Contributors

john-b-yang

25 commits

JO
John

2 commits

nielsr

1 commits