abundant-ai/incident-arena

Incident Arena: a 20-task SRE incident repair benchmark for Harbor

Python

0

2 commits

updated Sep 30, 2026

See the code

README

Incident Arena

Can an agent diagnose and repair a live production-style incident?

Incident Arena has 20 SRE incidents across Frappe, Saleor, and a Slack-like service. Each Harbor task starts a broken service stack. An agent investigates and repairs the incident, then the benchmark checks whether the service is healthy.

Getting started

Use a Linux machine with Docker. Each task needs 8 CPUs, 16 GB of memory, and 40–77 GB of storage. Clone the repo and install Harbor:

git clone https://github.com/abundant-ai/incident-arena.git
cd incident-arena
uv tool install 'harbor @ git+https://github.com/abundant-ai/harbor.git@main'

Set the API key for the model you plan to run, then try a task:

export ANTHROPIC_API_KEY=...
./scripts/run-benchmark.sh 005 -a claude-code -m claude-sonnet-5-5 --ak reasoning_effort=high

For an OpenAI model:

export OPENAI_API_KEY=...
./scripts/run-benchmark.sh 006 -a codex -m gpt-6.1-sol --ak reasoning_effort=medium

Use ./scripts/run-benchmark.sh --list to list tasks, or all in place of a task number to run all 20 sequentially. You can pass other harbor run options after the task number. To check that a task can be loaded without running it:

./scripts/run-benchmark.sh 000 -a nop --dry-run

Harbor writes results under jobs/. A full trial can take more than an hour.

Tasks

License

Apache 2.0. Vendored upstream components retain their own notices in the task directories.

benchmark
harbor
incident-response
kubernetes
sre

abundant-ai/incident-arena

Incident Arena: a 20-task SRE incident repair benchmark for Harbor

Python

0

2 commits

updated Sep 30, 2026

See the code

README

Incident Arena

Can an agent diagnose and repair a live production-style incident?

Incident Arena has 20 SRE incidents across Frappe, Saleor, and a Slack-like service. Each Harbor task starts a broken service stack. An agent investigates and repairs the incident, then the benchmark checks whether the service is healthy.

Getting started

Use a Linux machine with Docker. Each task needs 8 CPUs, 16 GB of memory, and 40–77 GB of storage. Clone the repo and install Harbor:

git clone https://github.com/abundant-ai/incident-arena.git
cd incident-arena
uv tool install 'harbor @ git+https://github.com/abundant-ai/harbor.git@main'

Set the API key for the model you plan to run, then try a task:

export ANTHROPIC_API_KEY=...
./scripts/run-benchmark.sh 005 -a claude-code -m claude-sonnet-5-5 --ak reasoning_effort=high

For an OpenAI model:

export OPENAI_API_KEY=...
./scripts/run-benchmark.sh 006 -a codex -m gpt-6.1-sol --ak reasoning_effort=medium

Use ./scripts/run-benchmark.sh --list to list tasks, or all in place of a task number to run all 20 sequentially. You can pass other harbor run options after the task number. To check that a task can be loaded without running it:

./scripts/run-benchmark.sh 000 -a nop --dry-run

Harbor writes results under jobs/. A full trial can take more than an hour.

Tasks

License

Apache 2.0. Vendored upstream components retain their own notices in the task directories.

benchmark
harbor
incident-response
kubernetes
sre

Languages

Python

94.3%

Shell

3.1%

Go Template

2.4%