yfwu/SWE-Together

Dataset

SWE-Together

0

3 commits

1 linked in READMEs

updated Jul 7, 2026

See the code

README

SWE-Together

SWE-Together: Evaluating Coding Agents in Interactive User Sessions.

SWE-Together reconstructs the multi-turn loop from real userโ€“agent coding sessions, replaying each with a reactive user simulator that asks questions, adds requirements, and pushes back โ€” preserving the original user's intent. This dataset holds the 109 discriminating tasks of the canonical suite as one metadata row per task.

from datasets import load_dataset

ds = load_dataset("yfwu/SWE-Together", split="test")
print(ds[0]["instruction"])   # the real user's first message, verbatim

What each row is

A task is a first user message plus a replayable, containerized interaction. The heavy artifacts โ€” the environment image, the deterministic verifier, and the user-simulator prompts โ€” live in the GitHub repo under tasks/<task_id>/. Each row here distills that task's metadata so it is browsable and loadable; docker_image and task_id point back to the full task.

Fields

FieldTypeDescription
task_idstringTask folder name; key into tasks/<task_id>/ in the GitHub repo.
instructionstringThe real user's first message, verbatim (what the agent reads).
repostringUpstream GitHub repo the session was on (owner/name), when known.
repo_urlstringURL of that repo, when known.
base_commitstringCommit the environment is built from, when known.
languagestringPrimary language (SWE-rebench-tier tasks).
difficultystringeasy / medium / hard.
categorystringfeature / bugfix / refactor.
tagslist[string]Free-form task tags.
scoring_tierstringswerebench (log-parser + FAIL_TO_PASS) or legacy (weighted F2P/P2P gates).
num_user_intentsintDistinct user intents across the session (turns that carry a request/question).
expert_time_estimate_minfloatHuman expert time estimate.
junior_time_estimate_minfloatHuman junior time estimate.
agent_timeout_secfloatPer-trial agent timeout.
docker_imagestringThe task's environment image (ghcr.io/togetherbench/...).
allow_internet, cpus, memoryโ€”Sandbox resource policy.
fail_to_pass, pass_to_passlist[string]Test targets (SWE-rebench tier).
test_cmd, log_parserstringVerifier command + parser (SWE-rebench tier).
source_fileslist[string]Files the reference change touches (SWE-rebench tier).
reference_patchstringGold patch (unified diff) reconstructed from the session, when available.
patch_files_changed, patch_additions, patch_deletionsintDiff stats for the reference patch.
oracle_intentsstring (JSON)Ordered user intents [{intent_id, source_turn, intent_kind, text, verbatim_excerpt}] driving the multi-turn loop.
completeness_goalsstring (JSON)Judge rubric goals for the agentic correctness score.
test_manifeststring (YAML)Legacy-tier F2P/P2P scoring gates and weights.

Running the benchmark

This table is for browsing and programmatic access. To actually run a coding agent against a task (with the user simulator and the deterministic verifier), use the harness in the GitHub repo:

git clone https://github.com/Togetherbench/SWE-Together
cd SWE-Together && uv sync
.venv/bin/python launch.py canonical_full109.json --stage run --models opencode_opus48 --execute

Citation

If you use SWE-Together, please cite our paper:

@article{wu2026swetogether,
  title   = {SWE-Together: Evaluating Coding Agents in Interactive User Sessions},
  author  = {Wu, Yifan and Zhao, Zhuokai and Li, Songlin and Lee, Ho Hin and Zhu, Jiacheng and Wu, Shirley and Yu, Tianhe and Li, Serena and Zhang, Lizhu and Fan, Xiangjun and Li, Shengzhi},
  year    = {2026},
  journal = {arXiv preprint arXiv:2606.29957},
  url     = {https://arxiv.org/pdf/2606.29957}
}

License

Released under the Apache-2.0 license, matching the source repository. Task content derives from public repositories; each task records its upstream repo/repo_url.

agents
benchmark
code
coding-agents
multi-turn

yfwu/SWE-Together

Dataset

SWE-Together

0

3 commits

1 linked in READMEs

updated Jul 7, 2026

See the code

README

SWE-Together

SWE-Together: Evaluating Coding Agents in Interactive User Sessions.

SWE-Together reconstructs the multi-turn loop from real userโ€“agent coding sessions, replaying each with a reactive user simulator that asks questions, adds requirements, and pushes back โ€” preserving the original user's intent. This dataset holds the 109 discriminating tasks of the canonical suite as one metadata row per task.

from datasets import load_dataset

ds = load_dataset("yfwu/SWE-Together", split="test")
print(ds[0]["instruction"])   # the real user's first message, verbatim

What each row is

A task is a first user message plus a replayable, containerized interaction. The heavy artifacts โ€” the environment image, the deterministic verifier, and the user-simulator prompts โ€” live in the GitHub repo under tasks/<task_id>/. Each row here distills that task's metadata so it is browsable and loadable; docker_image and task_id point back to the full task.

Fields

FieldTypeDescription
task_idstringTask folder name; key into tasks/<task_id>/ in the GitHub repo.
instructionstringThe real user's first message, verbatim (what the agent reads).
repostringUpstream GitHub repo the session was on (owner/name), when known.
repo_urlstringURL of that repo, when known.
base_commitstringCommit the environment is built from, when known.
languagestringPrimary language (SWE-rebench-tier tasks).
difficultystringeasy / medium / hard.
categorystringfeature / bugfix / refactor.
tagslist[string]Free-form task tags.
scoring_tierstringswerebench (log-parser + FAIL_TO_PASS) or legacy (weighted F2P/P2P gates).
num_user_intentsintDistinct user intents across the session (turns that carry a request/question).
expert_time_estimate_minfloatHuman expert time estimate.
junior_time_estimate_minfloatHuman junior time estimate.
agent_timeout_secfloatPer-trial agent timeout.
docker_imagestringThe task's environment image (ghcr.io/togetherbench/...).
allow_internet, cpus, memoryโ€”Sandbox resource policy.
fail_to_pass, pass_to_passlist[string]Test targets (SWE-rebench tier).
test_cmd, log_parserstringVerifier command + parser (SWE-rebench tier).
source_fileslist[string]Files the reference change touches (SWE-rebench tier).
reference_patchstringGold patch (unified diff) reconstructed from the session, when available.
patch_files_changed, patch_additions, patch_deletionsintDiff stats for the reference patch.
oracle_intentsstring (JSON)Ordered user intents [{intent_id, source_turn, intent_kind, text, verbatim_excerpt}] driving the multi-turn loop.
completeness_goalsstring (JSON)Judge rubric goals for the agentic correctness score.
test_manifeststring (YAML)Legacy-tier F2P/P2P scoring gates and weights.

Running the benchmark

This table is for browsing and programmatic access. To actually run a coding agent against a task (with the user simulator and the deterministic verifier), use the harness in the GitHub repo:

git clone https://github.com/Togetherbench/SWE-Together
cd SWE-Together && uv sync
.venv/bin/python launch.py canonical_full109.json --stage run --models opencode_opus48 --execute

Citation

If you use SWE-Together, please cite our paper:

@article{wu2026swetogether,
  title   = {SWE-Together: Evaluating Coding Agents in Interactive User Sessions},
  author  = {Wu, Yifan and Zhao, Zhuokai and Li, Songlin and Lee, Ho Hin and Zhu, Jiacheng and Wu, Shirley and Yu, Tianhe and Li, Serena and Zhang, Lizhu and Fan, Xiangjun and Li, Shengzhi},
  year    = {2026},
  journal = {arXiv preprint arXiv:2606.29957},
  url     = {https://arxiv.org/pdf/2606.29957}
}

License

Released under the Apache-2.0 license, matching the source repository. Task content derives from public repositories; each task records its upstream repo/repo_url.

agents
benchmark
code
coding-agents
multi-turn