[ICML 2026] Evaluating AI Agents on Continuous Software Evolution
72
stars
272
commits
Python
primary language
Sep 3, 2026
updated
Evaluating AI agents on continuous software evolution, structured as a milestone DAG, across real release histories.
Most existing benchmarks evaluate agents on isolated, one-shot tasks. But real-world workflows are not a bag of independent missions, they are continuous processes where tasks build on each other, dependencies interleave, and context accumulates over a long session.
SWE-Milestone is a general-purpose evaluation harness for continuous tasks. It drops an AI agent into a working environment and challenges it to complete an ordered sequence of milestones. As the agent works, SWE-Milestone silently extracts checkpoints, evaluates each milestone, and asynchronously unlocks downstream tasks, enabling fine-grained, per-milestone analysis without interrupting the agent's session.
A real milestone DAG: the scikit-learn 1.5.2 → 1.6.0 itinerary — nodes are milestones, edges are dependencies that gate when downstream tasks unlock.
Currently focused on software evolution, SWE-Milestone's architecture is designed to extend to other domains.
Test Your Model: Out of the box, SWE-Milestone ships with the SWE-Milestone Benchmark (long-horizon software evolution itineraries from 7 real-world repos) and 4 pre-configured agent frameworks (Claude Code, Codex, Gemini CLI, OpenHands). Provide a model API key and start evaluating.
Bring Your Own Agent: The agent layer is decoupled from the evaluation engine (see below). Plug in your own agent by implementing a lightweight adapter. SWE-Milestone also provides a per-milestone analysis framework for detailed performance breakdowns.
Bring Your Own Data: Supply your own task descriptions, test environments (Docker), test list for scoring, and task dependencies. SWE-Milestone handles orchestration, checkpoint-based evaluation, and reporting, enabling continuous task evaluation beyond coding.
Each evaluation trial works as follows:
agent-impl-milestone_001).0. Prerequisites
UNIFIED_API_KEY and UNIFIED_BASE_URL1. Installation
git clone https://github.com/DeepCommit-ai/SWE-Milestone.git
cd SWE-Milestone
uv sync
2. Data & Docker Images
Workspace data is hosted on HuggingFace. Docker images are hosted on DockerHub.
# Download workspace data
git lfs install
git clone https://huggingface.co/datasets/DeepCommit-ai/SWE-Milestone-data
# Point the harness at it — set once in .env_private (gitignored, auto-loaded on every launch)
cp .env .env_private
# edit .env_private → SWE_MILESTONE_DATA_ROOT=$PWD/SWE-Milestone-data
# Align the data checkout to the pinned benchmark version (manifests/BENCHMARK_VERSION)
./scripts/pull_data.sh
# Pull all images by content digest (login first — anonymous pulls are rate-limited, ~115 images)
docker login
./scripts/pull_images.sh
See docs/setup.md for the full data layout and image naming scheme, and docs/versioning.md for the versioning policy.
Trial configs then use data_root: ${SWE_MILESTONE_DATA_ROOT} — no host path to repeat, and the same variable drives pull_data.sh and the launch-time version gate. (Anti-cheat quarantine is auto-on per repo and needs no extra host paths; see docs/quarantine.md.)
Hand
docs/running-trials.mdto your agent — it has everything needed to launch trials, monitor progress, recover stuck repos autonomously, and manage trial IDs.
1. Configure — copy the template and edit:
cp trial_config.example.yaml trial_config.yaml
# modify trial_config.yaml
data_root: ${SWE_MILESTONE_DATA_ROOT} # set once in .env_private (or hardcode a path)
trial_name: my_experiment # name for this evaluation run
agent: claude-code # agent: claude-code | codex | gemini-cli | openhands
model: claude-opus-4-7 # model identifier (use claude-opus-4-7[1m] for 1M context)
timeout: 18000 # optional: max agent runtime per repo (seconds)
# reasoning_effort: high # optional: low | medium | high | xhigh | max
# repos: [navidrome, ripgrep] # optional: run only these repos (default: all)
# agent_version: 2.1.210 # optional: pin the agent CLI version (claude-code | codex | gemini-cli)
Naming convention: a trial name without a suffix runs the 200K context regime in Claude Code (pinned via
auto_compact_window: 200000); only a-1msuffix means 1M context (e.g.claude-code_opus-4.7-1m, modelclaude-opus-4-7[1m]).
New model or Vertex AI? See docs/adding-a-model.md; for Vertex (ADC auth, no API key) set
vertex_ai: true— see docs/vertex-ai.md.
2. Run — evaluate across all repos:
export UNIFIED_API_KEY=sk-...
export UNIFIED_BASE_URL=https://... # optional, for proxy or custom endpoints
# NOTE: if UNIFIED_BASE_URL is a custom domain, add it to WHITELISTED_DOMAINS
# in harness/e2e/container_setup.py — agent containers block all other outbound traffic.
python scripts/run_all.py --config trial_config.yaml
See docs/running-trials.md for the day-to-day operational runbook (launch, monitor, recover from stuck repos), and docs/advanced.md for single-repo / single-milestone debugging, result collection,
e2e_config.yaml, and lock internals.The full docs index is docs/README.md; documents for maintaining the benchmark itself (re-evaluation, image patching, dataset contracts) live under docs/maintainers/.
3. Monitor — check progress in another terminal:
./scripts/monitor.sh # auto-detects trial, compact view
./scripts/monitor.sh my_experiment --detail # per-milestone breakdown
./scripts/monitor.sh my_experiment --full # full table with all columns
Agent containers enforce an iptables outbound whitelist — only API and package-manager domains are allowed (e.g. api.anthropic.com, registry.npmjs.org, pypi.org); code-hosting sites (GitHub, GitLab, …) are blocked to prevent data leakage. Routing through a custom proxy? Add its domain to WHITELISTED_DOMAINS in harness/e2e/container_setup.py.
Plain-HTTP (port 80) blocked by your host? No action needed — the harness rewrites apt sources to HTTPS automatically.
Re-run the same command; it resumes from where the worker left off:
python scripts/run_all.py --config trial_config.yaml
Evaluation protocol: reported SWE-Milestone results resume trials until every milestone is submitted and evaluated. Each resume reuses the agent session (preserving the model's memory of prior work); after three consecutive resumes with no new submissions, a fresh session is rotated in automatically. Reproducibility studies should follow the same setting.
⚠️ Not every "no progress" is the agent's fault: rate-limit (429), quota, or auth (401/403) errors also stop workers, and session rotation won't fix them. Check the session jsonl for api_error_status: 429 / "reached your usage limit" / 401 and fix the key before resuming.
Never delete a running trial's container — all in-progress work (code, git history, conversation memory) lives inside it; once it's gone only --force (restart from scratch) remains. Snapshots of already-evaluated milestones survive on the host under evaluation/.
Containers start with --pull=never: a missing local image fails loudly instead of silently fetching. Align your machine with the release (digest-exact, near-free thanks to layer dedup):
./scripts/pull_images.sh # pulls by content digest from manifests/digests-<version>.tsv (version = manifests/BENCHMARK_VERSION)
python3 scripts/verify_image_digests.py --local # confirm local bytes match the frozen manifest
The launcher also checks that the data repo is on the matching version tag (default = manifests/BENCHMARK_VERSION, overridable with SWE_MILESTONE_IMAGE_TAG); if it refuses with a version mismatch, update the data checkout: git -C $SWE_MILESTONE_DATA_ROOT fetch --tags && git -C $SWE_MILESTONE_DATA_ROOT checkout $(cat manifests/BENCHMARK_VERSION).
We welcome contributions! Whether it's adding support for new agents, new task domains, new datasets, bug fixes, or documentation improvements.
Welcome to cite our paper if you find SWE-Milestone useful!
@misc{deng2026swemilestoneevaluatingaiagents,
title={SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution},
author={Gangda Deng and Zhaoling Chen and Zhongming Yu and Haoyang Fan and Yuhong Liu and Yuxin Yang and Dhruv Parikh and Rajgopal Kannan and Le Cong and Mengdi Wang and Qian Zhang and Viktor Prasanna and Xiangru Tang and Xingyao Wang},
year={2026},
eprint={2603.13428},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2603.13428},
}
This project is licensed under the MIT License.
Python
98.9%
Shell
1.0%
[ICML 2026] Evaluating AI Agents on Continuous Software Evolution
72
stars
272
commits
Python
primary language
Sep 3, 2026
updated
Evaluating AI agents on continuous software evolution, structured as a milestone DAG, across real release histories.
Most existing benchmarks evaluate agents on isolated, one-shot tasks. But real-world workflows are not a bag of independent missions, they are continuous processes where tasks build on each other, dependencies interleave, and context accumulates over a long session.
SWE-Milestone is a general-purpose evaluation harness for continuous tasks. It drops an AI agent into a working environment and challenges it to complete an ordered sequence of milestones. As the agent works, SWE-Milestone silently extracts checkpoints, evaluates each milestone, and asynchronously unlocks downstream tasks, enabling fine-grained, per-milestone analysis without interrupting the agent's session.
A real milestone DAG: the scikit-learn 1.5.2 → 1.6.0 itinerary — nodes are milestones, edges are dependencies that gate when downstream tasks unlock.
Currently focused on software evolution, SWE-Milestone's architecture is designed to extend to other domains.
Test Your Model: Out of the box, SWE-Milestone ships with the SWE-Milestone Benchmark (long-horizon software evolution itineraries from 7 real-world repos) and 4 pre-configured agent frameworks (Claude Code, Codex, Gemini CLI, OpenHands). Provide a model API key and start evaluating.
Bring Your Own Agent: The agent layer is decoupled from the evaluation engine (see below). Plug in your own agent by implementing a lightweight adapter. SWE-Milestone also provides a per-milestone analysis framework for detailed performance breakdowns.
Bring Your Own Data: Supply your own task descriptions, test environments (Docker), test list for scoring, and task dependencies. SWE-Milestone handles orchestration, checkpoint-based evaluation, and reporting, enabling continuous task evaluation beyond coding.
Each evaluation trial works as follows:
agent-impl-milestone_001).0. Prerequisites
UNIFIED_API_KEY and UNIFIED_BASE_URL1. Installation
git clone https://github.com/DeepCommit-ai/SWE-Milestone.git
cd SWE-Milestone
uv sync
2. Data & Docker Images
Workspace data is hosted on HuggingFace. Docker images are hosted on DockerHub.
# Download workspace data
git lfs install
git clone https://huggingface.co/datasets/DeepCommit-ai/SWE-Milestone-data
# Point the harness at it — set once in .env_private (gitignored, auto-loaded on every launch)
cp .env .env_private
# edit .env_private → SWE_MILESTONE_DATA_ROOT=$PWD/SWE-Milestone-data
# Align the data checkout to the pinned benchmark version (manifests/BENCHMARK_VERSION)
./scripts/pull_data.sh
# Pull all images by content digest (login first — anonymous pulls are rate-limited, ~115 images)
docker login
./scripts/pull_images.sh
See docs/setup.md for the full data layout and image naming scheme, and docs/versioning.md for the versioning policy.
Trial configs then use data_root: ${SWE_MILESTONE_DATA_ROOT} — no host path to repeat, and the same variable drives pull_data.sh and the launch-time version gate. (Anti-cheat quarantine is auto-on per repo and needs no extra host paths; see docs/quarantine.md.)
Hand
docs/running-trials.mdto your agent — it has everything needed to launch trials, monitor progress, recover stuck repos autonomously, and manage trial IDs.
1. Configure — copy the template and edit:
cp trial_config.example.yaml trial_config.yaml
# modify trial_config.yaml
data_root: ${SWE_MILESTONE_DATA_ROOT} # set once in .env_private (or hardcode a path)
trial_name: my_experiment # name for this evaluation run
agent: claude-code # agent: claude-code | codex | gemini-cli | openhands
model: claude-opus-4-7 # model identifier (use claude-opus-4-7[1m] for 1M context)
timeout: 18000 # optional: max agent runtime per repo (seconds)
# reasoning_effort: high # optional: low | medium | high | xhigh | max
# repos: [navidrome, ripgrep] # optional: run only these repos (default: all)
# agent_version: 2.1.210 # optional: pin the agent CLI version (claude-code | codex | gemini-cli)
Naming convention: a trial name without a suffix runs the 200K context regime in Claude Code (pinned via
auto_compact_window: 200000); only a-1msuffix means 1M context (e.g.claude-code_opus-4.7-1m, modelclaude-opus-4-7[1m]).
New model or Vertex AI? See docs/adding-a-model.md; for Vertex (ADC auth, no API key) set
vertex_ai: true— see docs/vertex-ai.md.
2. Run — evaluate across all repos:
export UNIFIED_API_KEY=sk-...
export UNIFIED_BASE_URL=https://... # optional, for proxy or custom endpoints
# NOTE: if UNIFIED_BASE_URL is a custom domain, add it to WHITELISTED_DOMAINS
# in harness/e2e/container_setup.py — agent containers block all other outbound traffic.
python scripts/run_all.py --config trial_config.yaml
See docs/running-trials.md for the day-to-day operational runbook (launch, monitor, recover from stuck repos), and docs/advanced.md for single-repo / single-milestone debugging, result collection,
e2e_config.yaml, and lock internals.The full docs index is docs/README.md; documents for maintaining the benchmark itself (re-evaluation, image patching, dataset contracts) live under docs/maintainers/.
3. Monitor — check progress in another terminal:
./scripts/monitor.sh # auto-detects trial, compact view
./scripts/monitor.sh my_experiment --detail # per-milestone breakdown
./scripts/monitor.sh my_experiment --full # full table with all columns
Agent containers enforce an iptables outbound whitelist — only API and package-manager domains are allowed (e.g. api.anthropic.com, registry.npmjs.org, pypi.org); code-hosting sites (GitHub, GitLab, …) are blocked to prevent data leakage. Routing through a custom proxy? Add its domain to WHITELISTED_DOMAINS in harness/e2e/container_setup.py.
Plain-HTTP (port 80) blocked by your host? No action needed — the harness rewrites apt sources to HTTPS automatically.
Re-run the same command; it resumes from where the worker left off:
python scripts/run_all.py --config trial_config.yaml
Evaluation protocol: reported SWE-Milestone results resume trials until every milestone is submitted and evaluated. Each resume reuses the agent session (preserving the model's memory of prior work); after three consecutive resumes with no new submissions, a fresh session is rotated in automatically. Reproducibility studies should follow the same setting.
⚠️ Not every "no progress" is the agent's fault: rate-limit (429), quota, or auth (401/403) errors also stop workers, and session rotation won't fix them. Check the session jsonl for api_error_status: 429 / "reached your usage limit" / 401 and fix the key before resuming.
Never delete a running trial's container — all in-progress work (code, git history, conversation memory) lives inside it; once it's gone only --force (restart from scratch) remains. Snapshots of already-evaluated milestones survive on the host under evaluation/.
Containers start with --pull=never: a missing local image fails loudly instead of silently fetching. Align your machine with the release (digest-exact, near-free thanks to layer dedup):
./scripts/pull_images.sh # pulls by content digest from manifests/digests-<version>.tsv (version = manifests/BENCHMARK_VERSION)
python3 scripts/verify_image_digests.py --local # confirm local bytes match the frozen manifest
The launcher also checks that the data repo is on the matching version tag (default = manifests/BENCHMARK_VERSION, overridable with SWE_MILESTONE_IMAGE_TAG); if it refuses with a version mismatch, update the data checkout: git -C $SWE_MILESTONE_DATA_ROOT fetch --tags && git -C $SWE_MILESTONE_DATA_ROOT checkout $(cat manifests/BENCHMARK_VERSION).
We welcome contributions! Whether it's adding support for new agents, new task domains, new datasets, bug fixes, or documentation improvements.
Welcome to cite our paper if you find SWE-Milestone useful!
@misc{deng2026swemilestoneevaluatingaiagents,
title={SWE-Milestone: Evaluating AI Agents on Continuous Software Evolution},
author={Gangda Deng and Zhaoling Chen and Zhongming Yu and Haoyang Fan and Yuhong Liu and Yuxin Yang and Dhruv Parikh and Rajgopal Kannan and Le Cong and Mengdi Wang and Qian Zhang and Viktor Prasanna and Xiangru Tang and Xingyao Wang},
year={2026},
eprint={2603.13428},
archivePrefix={arXiv},
primaryClass={cs.SE},
url={https://arxiv.org/abs/2603.13428},
}
This project is licensed under the MIT License.
Python
98.9%
Shell
1.0%