nebius/SWE-rebench-openhands-trajectories

Dataset

Dataset Summary

147

13 commits

2 linked in READMEs

updated Dec 27, 2025

See the code

README

Dataset Summary

SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding. This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from nebius/SWE-rebench. Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations.

MetricSWE-bench/SWE-smith-trajectoriesKwai-Klear/SWE-smith-mini_swe_agent_plus-trajectories-66knebius/SWE-agent-trajectoriesSWE-Gym/OpenHands-Sampled-TrajectoriesR2E-Gym/R2EGym-SFT-TrajectoriesOurs
ScaffoldingSWE-agentmini-swe-agent-plusClosed-sourceOpenHandsOpenHandsOpenHands (v0.54.0)
Bootstrapping ModelNameclaude-3-7-sonnet-20250219
claude-3-5-sonnet-20241022
gpt-4o-2024-08-06
unknown*Qwen2.5-72B-Instruct
Llama3-70B-Instruct
gpt-4o-2024-08-06
claude-3-5-sonnet-20241022
Claude-Sonnet-3.5-v2Qwen3-Coder-480B-A35B-Instruct
Uses function calling
Repositories1291231,20211unknown*1,823
IssuesResolved Count7,27010,8948382942,0483,792
Real-world/SyntheticSyntheticSyntheticReal-worldReal-worldReal-worldReal-world
TrajectoriesTotal Count49,89765,99480,0366,0553,23167,074
Successful Count21,51365,99413,3894913,23132,161
TurnsMax Count1511574085042100
Average Count30.234.326.418.916.164.3

Table 1: Comparison of statistics across different datasets containing multi-turn trajectories of agent’s interactions with executable SWE environments.

Note: Entries marked with asterisk (*) indicate statistics whose values couldn't be derived from available data.

For more details see our report in Nebius blog.


How to use

import json

from datasets import load_dataset


ds = load_dataset("nebius/SWE-rebench-openhands-trajectories", split="train")

role2field_names = {
 "system": ["role", "content"],
 "assistant": ["role", "content", "tool_calls"],
 "user": ["role", "content"],
 "tool": ["role", "content", "name", "tool_call_id"],
}

def filter_and_deserialize(row):
    trajectory = []
    for msg in row["trajectory"]:
        msg = {
            field_name: msg[field_name] for field_name in role2field_names[msg["role"]]
        }
        if (msg["role"] == "assistant") and (msg["tool_calls"] is not None):
            for i, tool_call in enumerate(msg["tool_calls"]):
                if "arguments" in tool_call.get("function", {}):
                    msg["tool_calls"][i]["function"]["arguments"] = json.loads(
                        tool_call["function"]["arguments"]
                    )
        trajectory.append(msg)
    return row | {"trajectory": trajectory}

first_trajectory = filter_and_deserialize(ds[0])["trajectory"]

for msg in first_trajectory:
    print(msg)

Dataset Structure

Each row contains the following information about trajectory:

Field NameTypeDescription
trajectory_idstrThe identifier unique for each collected trajectory.
instance_idstrGitHub issue identifier consisting of repository name and issue number. Can be joined with corresponding Docker images from nebius/SWE-rebench.
repostrThe repository identifier.
trajectorylistComplete conversation history with roles: 'system' (initial prompt), 'assistant' (model reasoning/actions), 'user' and 'tool' (environment observations).
model_patchstrFinal code modifications produced by the agent in unified diff format.
exit_statusstrContains 'submit' in case the agent completes trajectory with a terminating action, or an error message of the OpenHands agent otherwise.
resolvedintBinary indicator of task success: 1 if the agent solved the issue, 0 otherwise.
gen_tests_correctintNumber of agent-generated tests that correctly validate the solution (fail before applying the golden patch, pass after). null if no tests were generated. This metric validates agent's test writing ability.
pred_passes_gen_testintNumber of agent-generated tests passed by the agent's own solution (model_patch). null if no tests were generated. This metric evaluates predicted solution correctness against the agent's own tests.

Table 2: Dataset field descriptions and schema.

To our knowledge, no other publicly available agent trajectory dataset includes evaluation of agent-generated tests.

Important Note: arguments field inside tool calls present in assistant steps of trajectory is serialized to string format for storage efficiency. When training on this data, you may want to deserialize it first to ensure chat templates apply the same formatting during training and inference.


Citation

@article{trofimova2025openhandstrajs,
  title={OpenHands Trajectories with Qwen3-Coder-480B-A35B-Instruct},
  author={Trofimova, Maria and Shevtsov, Anton and Ibragim, Badertdinov and Pyaev, Konstantin and Karasik, Simon and Golubev, Alexander},
  year={2025},
  journal={Nebius blog},
  note={}
}
agents
code
software
synthetic
tools

Contributors

vim-ary

7 commits

ibragim-bad

6 commits

nebius/SWE-rebench-openhands-trajectories

Dataset

Dataset Summary

147

13 commits

2 linked in READMEs

updated Dec 27, 2025

See the code

README

Dataset Summary

SWE-rebench-OpenHands-Trajectories is a dataset of multi-turn agent trajectories for software engineering tasks, collected using Qwen/Qwen3-Coder-480B-A35B-Instruct with OpenHands (v0.54.0) agent scaffolding. This dataset captures complete agent execution traces as they attempt to resolve real GitHub issues from nebius/SWE-rebench. Each trajectory contains the agent's step-by-step reasoning, actions, and environmental observations.

MetricSWE-bench/SWE-smith-trajectoriesKwai-Klear/SWE-smith-mini_swe_agent_plus-trajectories-66knebius/SWE-agent-trajectoriesSWE-Gym/OpenHands-Sampled-TrajectoriesR2E-Gym/R2EGym-SFT-TrajectoriesOurs
ScaffoldingSWE-agentmini-swe-agent-plusClosed-sourceOpenHandsOpenHandsOpenHands (v0.54.0)
Bootstrapping ModelNameclaude-3-7-sonnet-20250219
claude-3-5-sonnet-20241022
gpt-4o-2024-08-06
unknown*Qwen2.5-72B-Instruct
Llama3-70B-Instruct
gpt-4o-2024-08-06
claude-3-5-sonnet-20241022
Claude-Sonnet-3.5-v2Qwen3-Coder-480B-A35B-Instruct
Uses function calling
Repositories1291231,20211unknown*1,823
IssuesResolved Count7,27010,8948382942,0483,792
Real-world/SyntheticSyntheticSyntheticReal-worldReal-worldReal-worldReal-world
TrajectoriesTotal Count49,89765,99480,0366,0553,23167,074
Successful Count21,51365,99413,3894913,23132,161
TurnsMax Count1511574085042100
Average Count30.234.326.418.916.164.3

Table 1: Comparison of statistics across different datasets containing multi-turn trajectories of agent’s interactions with executable SWE environments.

Note: Entries marked with asterisk (*) indicate statistics whose values couldn't be derived from available data.

For more details see our report in Nebius blog.


How to use

import json

from datasets import load_dataset


ds = load_dataset("nebius/SWE-rebench-openhands-trajectories", split="train")

role2field_names = {
 "system": ["role", "content"],
 "assistant": ["role", "content", "tool_calls"],
 "user": ["role", "content"],
 "tool": ["role", "content", "name", "tool_call_id"],
}

def filter_and_deserialize(row):
    trajectory = []
    for msg in row["trajectory"]:
        msg = {
            field_name: msg[field_name] for field_name in role2field_names[msg["role"]]
        }
        if (msg["role"] == "assistant") and (msg["tool_calls"] is not None):
            for i, tool_call in enumerate(msg["tool_calls"]):
                if "arguments" in tool_call.get("function", {}):
                    msg["tool_calls"][i]["function"]["arguments"] = json.loads(
                        tool_call["function"]["arguments"]
                    )
        trajectory.append(msg)
    return row | {"trajectory": trajectory}

first_trajectory = filter_and_deserialize(ds[0])["trajectory"]

for msg in first_trajectory:
    print(msg)

Dataset Structure

Each row contains the following information about trajectory:

Field NameTypeDescription
trajectory_idstrThe identifier unique for each collected trajectory.
instance_idstrGitHub issue identifier consisting of repository name and issue number. Can be joined with corresponding Docker images from nebius/SWE-rebench.
repostrThe repository identifier.
trajectorylistComplete conversation history with roles: 'system' (initial prompt), 'assistant' (model reasoning/actions), 'user' and 'tool' (environment observations).
model_patchstrFinal code modifications produced by the agent in unified diff format.
exit_statusstrContains 'submit' in case the agent completes trajectory with a terminating action, or an error message of the OpenHands agent otherwise.
resolvedintBinary indicator of task success: 1 if the agent solved the issue, 0 otherwise.
gen_tests_correctintNumber of agent-generated tests that correctly validate the solution (fail before applying the golden patch, pass after). null if no tests were generated. This metric validates agent's test writing ability.
pred_passes_gen_testintNumber of agent-generated tests passed by the agent's own solution (model_patch). null if no tests were generated. This metric evaluates predicted solution correctness against the agent's own tests.

Table 2: Dataset field descriptions and schema.

To our knowledge, no other publicly available agent trajectory dataset includes evaluation of agent-generated tests.

Important Note: arguments field inside tool calls present in assistant steps of trajectory is serialized to string format for storage efficiency. When training on this data, you may want to deserialize it first to ensure chat templates apply the same formatting during training and inference.


Citation

@article{trofimova2025openhandstrajs,
  title={OpenHands Trajectories with Qwen3-Coder-480B-A35B-Instruct},
  author={Trofimova, Maria and Shevtsov, Anton and Ibragim, Badertdinov and Pyaev, Konstantin and Karasik, Simon and Golubev, Alexander},
  year={2025},
  journal={Nebius blog},
  note={}
}
agents
code
software
synthetic
tools

Contributors

vim-ary

7 commits

ibragim-bad

6 commits