Incident Arena: a 20-task SRE incident repair benchmark for Harbor
Python
0
2 commits
updated Sep 30, 2026
Can an agent diagnose and repair a live production-style incident?
Incident Arena has 20 SRE incidents across Frappe, Saleor, and a Slack-like service. Each Harbor task starts a broken service stack. An agent investigates and repairs the incident, then the benchmark checks whether the service is healthy.
Use a Linux machine with Docker. Each task needs 8 CPUs, 16 GB of memory, and 40–77 GB of storage. Clone the repo and install Harbor:
git clone https://github.com/abundant-ai/incident-arena.git
cd incident-arena
uv tool install 'harbor @ git+https://github.com/abundant-ai/harbor.git@main'
Set the API key for the model you plan to run, then try a task:
export ANTHROPIC_API_KEY=...
./scripts/run-benchmark.sh 005 -a claude-code -m claude-sonnet-5-5 --ak reasoning_effort=high
For an OpenAI model:
export OPENAI_API_KEY=...
./scripts/run-benchmark.sh 006 -a codex -m gpt-6.1-sol --ak reasoning_effort=medium
Use ./scripts/run-benchmark.sh --list to list tasks, or all in place of a task number to run all 20 sequentially. You can pass other harbor run options after the task number. To check that a task can be loaded without running it:
./scripts/run-benchmark.sh 000 -a nop --dry-run
Harbor writes results under jobs/. A full trial can take more than an hour.
Apache 2.0. Vendored upstream components retain their own notices in the task directories.
Python
94.3%
Shell
3.1%
Go Template
2.4%
Incident Arena: a 20-task SRE incident repair benchmark for Harbor
Python
0
2 commits
updated Sep 30, 2026
Can an agent diagnose and repair a live production-style incident?
Incident Arena has 20 SRE incidents across Frappe, Saleor, and a Slack-like service. Each Harbor task starts a broken service stack. An agent investigates and repairs the incident, then the benchmark checks whether the service is healthy.
Use a Linux machine with Docker. Each task needs 8 CPUs, 16 GB of memory, and 40–77 GB of storage. Clone the repo and install Harbor:
git clone https://github.com/abundant-ai/incident-arena.git
cd incident-arena
uv tool install 'harbor @ git+https://github.com/abundant-ai/harbor.git@main'
Set the API key for the model you plan to run, then try a task:
export ANTHROPIC_API_KEY=...
./scripts/run-benchmark.sh 005 -a claude-code -m claude-sonnet-5-5 --ak reasoning_effort=high
For an OpenAI model:
export OPENAI_API_KEY=...
./scripts/run-benchmark.sh 006 -a codex -m gpt-6.1-sol --ak reasoning_effort=medium
Use ./scripts/run-benchmark.sh --list to list tasks, or all in place of a task number to run all 20 sequentially. You can pass other harbor run options after the task number. To check that a task can be loaded without running it:
./scripts/run-benchmark.sh 000 -a nop --dry-run
Harbor writes results under jobs/. A full trial can take more than an hour.
Apache 2.0. Vendored upstream components retain their own notices in the task directories.
Python
94.3%
Shell
3.1%
Go Template
2.4%