Gold-patch-validated subset of
PrimeIntellect/SWE-Lego-Real-Data
(itself a fixed fork of SWE-Lego's real-data split). The
resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply
test_patch, apply the gold patch, run the row's test_cmd in its image, require every
F2P/P2P test to report PASSED.
dropped split iff 0/10 passes on retry; flaky rows stay in resolved.cloud-custodian, pytorch/ignite, sympy
lead). Schema and row content are otherwise unchanged.License mirrors upstream: Apache-2.0.
| Split | Rows |
|---|---|
resolved | 4,323 |
dropped | 109 |
Install the swelego_v1 taskset from
research-environments, then run it
end-to-end with verifiers:
uv pip install --prerelease=allow "git+https://github.com/PrimeIntellect-ai/research-environments.git#subdirectory=environments/swe/swelego_v1"
uv run eval --taskset.id swelego_v1 -m <your-model> -n 100 -r 4
Snapshot of the SWE-Lego/SWE-Lego-Real-Data
card at card-build time — see the live card for updates.
SWE-Lego/SWE-Lego-Real-Data dataset cardPaper | Github | HF Collection
SWE-Lego-Real-Data contains 18k real github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training.
.
└── data
├── resolved-00000-of-00001.parquet (5k github issues with resolved trajectories)
└── unresolved-00000-of-00001.parquet (13k github issues with unresolved trajectories)
The effectiveness of the dataset has been demonstrated by training exclusively with SFT from Qwen3-8B and Qwen3-32B, and evaluated on SWE-Bench-Verified:
We’ve open-sourced everything—our dataset, code, and training scripts, for everyone to progress on scaling and improving software engineering agents. For more details, please refer to our Github and Paper.
import json
from datasets import load_dataset
# Load dataset
ds = load_dataset("SWE-Lego/SWE-Lego-Real-Data", split="resolved")
# Select required columns
processed_ds = ds.select_columns(["instance_id", "messages"])
# Convert to list format
data_list = processed_ds.to_list()
# Save as JSON file
filename = "swe_lego_real_data_resolved_trajectories.json"
with open(filename, "w", encoding="utf-8") as f:
json.dump(data_list, f, ensure_ascii=False, indent=4)
print(f"Saved {len(data_list)} records to {filename}")
SWE-Lego-Real-Data is built upon SWE-rebench, a large-scale dataset comprising 21k issue–pull request pairs.
Please cite our paper if you find the repo helpful in your work:
@misc{swelego,
title={SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving},
author={Chaofan Tao and Jierun Chen and Yuxin Jiang and Kaiqi Kou and Shaowei Wang and Ruoyu Wang and Xiaohui Li and Sidi Yang and Yiming Du and Jianbo Dai and Zhiming Mao and Xinyu Wang and Lifeng Shang and Haoli Bai},
year={2026},
eprint={2601.01426},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2601.01426},
}
Gold-patch-validated subset of
PrimeIntellect/SWE-Lego-Real-Data
(itself a fixed fork of SWE-Lego's real-data split). The
resolved split contains 4,323 / 4,432 rows (97.54%) verified scoreable end-to-end: apply
test_patch, apply the gold patch, run the row's test_cmd in its image, require every
F2P/P2P test to report PASSED.
dropped split iff 0/10 passes on retry; flaky rows stay in resolved.cloud-custodian, pytorch/ignite, sympy
lead). Schema and row content are otherwise unchanged.License mirrors upstream: Apache-2.0.
| Split | Rows |
|---|---|
resolved | 4,323 |
dropped | 109 |
Install the swelego_v1 taskset from
research-environments, then run it
end-to-end with verifiers:
uv pip install --prerelease=allow "git+https://github.com/PrimeIntellect-ai/research-environments.git#subdirectory=environments/swe/swelego_v1"
uv run eval --taskset.id swelego_v1 -m <your-model> -n 100 -r 4
Snapshot of the SWE-Lego/SWE-Lego-Real-Data
card at card-build time — see the live card for updates.
SWE-Lego/SWE-Lego-Real-Data dataset cardPaper | Github | HF Collection
SWE-Lego-Real-Data contains 18k real github issues (Python language) and their multi-turn agent trajectories. The column named messages is collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.53.0) agent scaffolding, which can be directly used for SFT training.
.
└── data
├── resolved-00000-of-00001.parquet (5k github issues with resolved trajectories)
└── unresolved-00000-of-00001.parquet (13k github issues with unresolved trajectories)
The effectiveness of the dataset has been demonstrated by training exclusively with SFT from Qwen3-8B and Qwen3-32B, and evaluated on SWE-Bench-Verified:
We’ve open-sourced everything—our dataset, code, and training scripts, for everyone to progress on scaling and improving software engineering agents. For more details, please refer to our Github and Paper.
import json
from datasets import load_dataset
# Load dataset
ds = load_dataset("SWE-Lego/SWE-Lego-Real-Data", split="resolved")
# Select required columns
processed_ds = ds.select_columns(["instance_id", "messages"])
# Convert to list format
data_list = processed_ds.to_list()
# Save as JSON file
filename = "swe_lego_real_data_resolved_trajectories.json"
with open(filename, "w", encoding="utf-8") as f:
json.dump(data_list, f, ensure_ascii=False, indent=4)
print(f"Saved {len(data_list)} records to {filename}")
SWE-Lego-Real-Data is built upon SWE-rebench, a large-scale dataset comprising 21k issue–pull request pairs.
Please cite our paper if you find the repo helpful in your work:
@misc{swelego,
title={SWE-Lego: Pushing the Limits of Supervised Fine-tuning for Software Issue Resolving},
author={Chaofan Tao and Jierun Chen and Yuxin Jiang and Kaiqi Kou and Shaowei Wang and Ruoyu Wang and Xiaohui Li and Sidi Yang and Yiming Du and Jianbo Dai and Zhiming Mao and Xinyu Wang and Lifeng Shang and Haoli Bai},
year={2026},
eprint={2601.01426},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2601.01426},
}