Claude Code skill for delegating development and verification tasks to pi coding agent — with adversarial review loop
Shell
0
39 commits
updated Oct 2, 2026
Claude Code plans and checks. A cheaper model does the heavy lifting.
Say "delegate to pi" and the heavy part of a coding task (reading files, editing, running tests) runs on pi, an open-source coding agent that can drive any model you pick, even one on your own machine. Claude only writes the brief and reads a short result, and your tests decide whether it worked.
[!TIP] On multi-file tasks, delegating cut Claude's cost by a third (−33%) with every hidden test still passing.
| Plain Claude | With pi-delegate | Verdict | |
|---|---|---|---|
| 🧩 Multi-file features | $0.082 | $0.055 | ✅ −33% Claude cost |
| ✏️ Tiny edits (ten-line fixes) | $0.042 | $0.047 | ⚠️ +11%: do these yourself |
| 🧪 Hidden checks passing | 28 / 28 | 28 / 28 | ✅ same quality |
[!NOTE] The honest fine print. The numbers count Claude's cost only: pi's own spend comes on top (nothing if you self-host, cents on a hosted cheap model). Delegated runs are also slower, several times in our runs on a self-hosted pi (about 25 s plain vs 2-3 minutes). Samples are small (2 runs per task); the benchmark takes about five minutes to rerun yourself (how).
| ✅ A good fit | ❌ Not a fit |
|---|---|
| You use Claude Code and your tasks touch several files | Quick one-file edits (Claude alone is cheaper) |
| You want fewer Claude tokens, or less of your usage limit, spent on routine implementation | Tasks with no way to check the result |
| You have a test command that can say whether the work is right | Speed matters more than cost |
1. Add the plugin
/plugin marketplace add randomm/pi-delegate
/plugin install pi-delegate@pi-delegate
2. Set up pi once
npm install -g @earendil-works/pi-coding-agent # then run `pi`, /login (or export an API key), /model
Cheap and local models work; the benchmark used a self-hosted Qwen. Also install jq, and on macOS
brew install coreutils for the timeout that bounds each pi call.
delegate to pi: add a
--jsonflag to cli.py, verify withpython3 -m unittest
Claude makes one call, pi does the work, and your verify command decides whether it worked:
EXIT CODE: 0
<pi's summary>
VERIFY: PASS (retries=0)
cli.py | 12 +++++++++---
If verification fails, pi gets one more attempt with the failure output. If it still fails, Claude says so instead of claiming success.
--verify "<your tests>" runs after pi. No model reviews the
work: a same-model reviewer approves most changes, tests don't..env/*.pem/*.key files, and disables
git push for pi. This guards against mistakes, not a malicious model. For stronger isolation set
PI_DELEGATE_WRAP (sandbox recipes in docs/configuration.md) or use a
disposable clone or container.skills/delegate/run.sh in this repo is plain bash. Any agent that can run a shell command can use it:
git clone https://github.com/randomm/pi-delegate.git
bash pi-delegate/skills/delegate/run.sh --verify "pytest -q" <<'TASK'
Fix the failing date parsing in utils.py; do not commit.
TASK
bench/quick.sh -n 3 # about five minutes: plain Claude vs Claude + pi-delegate, prints a REWARD score
Method, tasks and the slower real-repo protocol: docs/benchmark.md · results. Settings (timeouts, safety, sandbox): docs/configuration.md.
Apache License 2.0, see LICENSE. Copyright 2026 Janni Turunen.
Shell
92.2%
Python
7.8%
Claude Code skill for delegating development and verification tasks to pi coding agent — with adversarial review loop
Shell
0
39 commits
updated Oct 2, 2026
Claude Code plans and checks. A cheaper model does the heavy lifting.
Say "delegate to pi" and the heavy part of a coding task (reading files, editing, running tests) runs on pi, an open-source coding agent that can drive any model you pick, even one on your own machine. Claude only writes the brief and reads a short result, and your tests decide whether it worked.
[!TIP] On multi-file tasks, delegating cut Claude's cost by a third (−33%) with every hidden test still passing.
| Plain Claude | With pi-delegate | Verdict | |
|---|---|---|---|
| 🧩 Multi-file features | $0.082 | $0.055 | ✅ −33% Claude cost |
| ✏️ Tiny edits (ten-line fixes) | $0.042 | $0.047 | ⚠️ +11%: do these yourself |
| 🧪 Hidden checks passing | 28 / 28 | 28 / 28 | ✅ same quality |
[!NOTE] The honest fine print. The numbers count Claude's cost only: pi's own spend comes on top (nothing if you self-host, cents on a hosted cheap model). Delegated runs are also slower, several times in our runs on a self-hosted pi (about 25 s plain vs 2-3 minutes). Samples are small (2 runs per task); the benchmark takes about five minutes to rerun yourself (how).
| ✅ A good fit | ❌ Not a fit |
|---|---|
| You use Claude Code and your tasks touch several files | Quick one-file edits (Claude alone is cheaper) |
| You want fewer Claude tokens, or less of your usage limit, spent on routine implementation | Tasks with no way to check the result |
| You have a test command that can say whether the work is right | Speed matters more than cost |
1. Add the plugin
/plugin marketplace add randomm/pi-delegate
/plugin install pi-delegate@pi-delegate
2. Set up pi once
npm install -g @earendil-works/pi-coding-agent # then run `pi`, /login (or export an API key), /model
Cheap and local models work; the benchmark used a self-hosted Qwen. Also install jq, and on macOS
brew install coreutils for the timeout that bounds each pi call.
delegate to pi: add a
--jsonflag to cli.py, verify withpython3 -m unittest
Claude makes one call, pi does the work, and your verify command decides whether it worked:
EXIT CODE: 0
<pi's summary>
VERIFY: PASS (retries=0)
cli.py | 12 +++++++++---
If verification fails, pi gets one more attempt with the failure output. If it still fails, Claude says so instead of claiming success.
--verify "<your tests>" runs after pi. No model reviews the
work: a same-model reviewer approves most changes, tests don't..env/*.pem/*.key files, and disables
git push for pi. This guards against mistakes, not a malicious model. For stronger isolation set
PI_DELEGATE_WRAP (sandbox recipes in docs/configuration.md) or use a
disposable clone or container.skills/delegate/run.sh in this repo is plain bash. Any agent that can run a shell command can use it:
git clone https://github.com/randomm/pi-delegate.git
bash pi-delegate/skills/delegate/run.sh --verify "pytest -q" <<'TASK'
Fix the failing date parsing in utils.py; do not commit.
TASK
bench/quick.sh -n 3 # about five minutes: plain Claude vs Claude + pi-delegate, prints a REWARD score
Method, tasks and the slower real-repo protocol: docs/benchmark.md · results. Settings (timeouts, safety, sandbox): docs/configuration.md.
Apache License 2.0, see LICENSE. Copyright 2026 Janni Turunen.
Shell
92.2%
Python
7.8%