SWE-Together: Evaluating Coding Agents in Interactive User Sessions.
SWE-Together reconstructs the multi-turn loop from real userโagent coding sessions, replaying each with a reactive user simulator that asks questions, adds requirements, and pushes back โ preserving the original user's intent. This dataset holds the 109 discriminating tasks of the canonical suite as one metadata row per task.
from datasets import load_dataset
ds = load_dataset("yfwu/SWE-Together", split="test")
print(ds[0]["instruction"]) # the real user's first message, verbatim
A task is a first user message plus a replayable, containerized interaction.
The heavy artifacts โ the environment image, the deterministic verifier, and the
user-simulator prompts โ live in the GitHub repo
under tasks/<task_id>/. Each row here distills that task's metadata so it is
browsable and loadable; docker_image and task_id point back to the full task.
| Field | Type | Description |
|---|---|---|
task_id | string | Task folder name; key into tasks/<task_id>/ in the GitHub repo. |
instruction | string | The real user's first message, verbatim (what the agent reads). |
repo | string | Upstream GitHub repo the session was on (owner/name), when known. |
repo_url | string | URL of that repo, when known. |
base_commit | string | Commit the environment is built from, when known. |
language | string | Primary language (SWE-rebench-tier tasks). |
difficulty | string | easy / medium / hard. |
category | string | feature / bugfix / refactor. |
tags | list[string] | Free-form task tags. |
scoring_tier | string | swerebench (log-parser + FAIL_TO_PASS) or legacy (weighted F2P/P2P gates). |
num_user_intents | int | Distinct user intents across the session (turns that carry a request/question). |
expert_time_estimate_min | float | Human expert time estimate. |
junior_time_estimate_min | float | Human junior time estimate. |
agent_timeout_sec | float | Per-trial agent timeout. |
docker_image | string | The task's environment image (ghcr.io/togetherbench/...). |
allow_internet, cpus, memory | โ | Sandbox resource policy. |
fail_to_pass, pass_to_pass | list[string] | Test targets (SWE-rebench tier). |
test_cmd, log_parser | string | Verifier command + parser (SWE-rebench tier). |
source_files | list[string] | Files the reference change touches (SWE-rebench tier). |
reference_patch | string | Gold patch (unified diff) reconstructed from the session, when available. |
patch_files_changed, patch_additions, patch_deletions | int | Diff stats for the reference patch. |
oracle_intents | string (JSON) | Ordered user intents [{intent_id, source_turn, intent_kind, text, verbatim_excerpt}] driving the multi-turn loop. |
completeness_goals | string (JSON) | Judge rubric goals for the agentic correctness score. |
test_manifest | string (YAML) | Legacy-tier F2P/P2P scoring gates and weights. |
This table is for browsing and programmatic access. To actually run a coding agent against a task (with the user simulator and the deterministic verifier), use the harness in the GitHub repo:
git clone https://github.com/Togetherbench/SWE-Together
cd SWE-Together && uv sync
.venv/bin/python launch.py canonical_full109.json --stage run --models opencode_opus48 --execute
If you use SWE-Together, please cite our paper:
@article{wu2026swetogether,
title = {SWE-Together: Evaluating Coding Agents in Interactive User Sessions},
author = {Wu, Yifan and Zhao, Zhuokai and Li, Songlin and Lee, Ho Hin and Zhu, Jiacheng and Wu, Shirley and Yu, Tianhe and Li, Serena and Zhang, Lizhu and Fan, Xiangjun and Li, Shengzhi},
year = {2026},
journal = {arXiv preprint arXiv:2606.29957},
url = {https://arxiv.org/pdf/2606.29957}
}
Released under the Apache-2.0 license, matching the
source repository. Task content
derives from public repositories; each task records its upstream repo/repo_url.
SWE-Together: Evaluating Coding Agents in Interactive User Sessions.
SWE-Together reconstructs the multi-turn loop from real userโagent coding sessions, replaying each with a reactive user simulator that asks questions, adds requirements, and pushes back โ preserving the original user's intent. This dataset holds the 109 discriminating tasks of the canonical suite as one metadata row per task.
from datasets import load_dataset
ds = load_dataset("yfwu/SWE-Together", split="test")
print(ds[0]["instruction"]) # the real user's first message, verbatim
A task is a first user message plus a replayable, containerized interaction.
The heavy artifacts โ the environment image, the deterministic verifier, and the
user-simulator prompts โ live in the GitHub repo
under tasks/<task_id>/. Each row here distills that task's metadata so it is
browsable and loadable; docker_image and task_id point back to the full task.
| Field | Type | Description |
|---|---|---|
task_id | string | Task folder name; key into tasks/<task_id>/ in the GitHub repo. |
instruction | string | The real user's first message, verbatim (what the agent reads). |
repo | string | Upstream GitHub repo the session was on (owner/name), when known. |
repo_url | string | URL of that repo, when known. |
base_commit | string | Commit the environment is built from, when known. |
language | string | Primary language (SWE-rebench-tier tasks). |
difficulty | string | easy / medium / hard. |
category | string | feature / bugfix / refactor. |
tags | list[string] | Free-form task tags. |
scoring_tier | string | swerebench (log-parser + FAIL_TO_PASS) or legacy (weighted F2P/P2P gates). |
num_user_intents | int | Distinct user intents across the session (turns that carry a request/question). |
expert_time_estimate_min | float | Human expert time estimate. |
junior_time_estimate_min | float | Human junior time estimate. |
agent_timeout_sec | float | Per-trial agent timeout. |
docker_image | string | The task's environment image (ghcr.io/togetherbench/...). |
allow_internet, cpus, memory | โ | Sandbox resource policy. |
fail_to_pass, pass_to_pass | list[string] | Test targets (SWE-rebench tier). |
test_cmd, log_parser | string | Verifier command + parser (SWE-rebench tier). |
source_files | list[string] | Files the reference change touches (SWE-rebench tier). |
reference_patch | string | Gold patch (unified diff) reconstructed from the session, when available. |
patch_files_changed, patch_additions, patch_deletions | int | Diff stats for the reference patch. |
oracle_intents | string (JSON) | Ordered user intents [{intent_id, source_turn, intent_kind, text, verbatim_excerpt}] driving the multi-turn loop. |
completeness_goals | string (JSON) | Judge rubric goals for the agentic correctness score. |
test_manifest | string (YAML) | Legacy-tier F2P/P2P scoring gates and weights. |
This table is for browsing and programmatic access. To actually run a coding agent against a task (with the user simulator and the deterministic verifier), use the harness in the GitHub repo:
git clone https://github.com/Togetherbench/SWE-Together
cd SWE-Together && uv sync
.venv/bin/python launch.py canonical_full109.json --stage run --models opencode_opus48 --execute
If you use SWE-Together, please cite our paper:
@article{wu2026swetogether,
title = {SWE-Together: Evaluating Coding Agents in Interactive User Sessions},
author = {Wu, Yifan and Zhao, Zhuokai and Li, Songlin and Lee, Ho Hin and Zhu, Jiacheng and Wu, Shirley and Yu, Tianhe and Li, Serena and Zhang, Lizhu and Fan, Xiangjun and Li, Shengzhi},
year = {2026},
journal = {arXiv preprint arXiv:2606.29957},
url = {https://arxiv.org/pdf/2606.29957}
}
Released under the Apache-2.0 license, matching the
source repository. Task content
derives from public repositories; each task records its upstream repo/repo_url.