[12/14/2025] NOTE: We will no longer actively update this dataset.
57
28 commits
1 linked in READMEs
updated Dec 14, 2025
The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents.
All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation.
[12/14/2025] NOTE: We will no longer actively update this dataset.
57
28 commits
1 linked in READMEs
updated Dec 14, 2025
The SWE-smith Dataset is a training dataset of 50137 task instances from 128 GitHub repositories, collected using the SWE-smith toolkit. It is the largest dataset to date for training software engineering agents.
All SWE-smith task instances come with an executable environment. To learn more about how to use this dataset to train Language Models for Software Engineering, please refer to the documentation.