MCP proxy: approve an AI agent's multi-step plan once, then only the calls that plan allows get through
TypeScript
0
1 commits
updated Oct 7, 2026
Approve an AI agent's plan once. Then it can only make the calls that plan allows.
AI agents that use tools (via MCP) usually work in one of two modes: they ask permission for every single call, and after the tenth prompt you click "yes" without reading; or they run unattended with full access. Planlock is a third option:
agent ──► Planlock ──► your tools (any MCP server)
▲
you approve the plan here (the agent can't)
The check is code, not a prompt: a confused or prompt-injected agent still can't go past the limits you approved. (Scope and limits: threat model.)
You approve this plan: refund at most 30 € on order 4711 → email the customer → close the ticket.
| The agent tries to… | Planlock |
|---|---|
| refund 30 € on order 4711 | ✅ forwarded, runs |
| refund 300 € | ❌ blocked: above the approved maximum |
| refund a different order | ❌ blocked: not the approved order |
| send the email before the refund is done | ❌ blocked: not this step |
| run any other script or edit one | ❌ blocked: not in the plan |
| refund after someone changed the refund script | ❌ blocked, plan paused: not the version you approved |
Blocked calls never reach your systems. The agent gets a plain error back and can only continue within the plan or stop.
The refund example as a runnable demo. The tools live in Windmill (open-source workflow engine), reached through Windmill's own, unmodified MCP server. Requires Docker and Node 20+.
npm ci
cp .env.example .env # set APPROVER_TOKEN
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml up -d
node examples/windmill/setup.mjs # workspace, 4 scripts, MCP token
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml --profile guard up -d --build planlock
node examples/windmill/demo.mjs # 17 assertions
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml --profile guard down -v # clean up
What the demo prints (excerpt of a real run):
[AGENT] step 1 — tries to overstep first:
✔ refund 300 EUR blocked → 403 Blocked by approved plan: s-f_support_refund__customer.amount exceeds the approved maximum
✔ approved refund executes in Windmill → {"amount":30,"status":"refunded","order_id":"4711","refund_id":"rf_1791310350551"}
✔ second refund blocked → 403 Blocked by approved plan: s-f_support_refund__customer already called 1 time(s); the approved plan allows 1 for this step
[COLLEAGUE] edits f/support/refund_customer in Windmill after the approval (refunds 10x the amount)
✔ refund blocked: script no longer matches the approved version → 409 Blocked by approved plan: getScriptByPath hash is "ca5f52798191da90", approved "89a679a5673b1fd0". The upstream changed since approval.
✔ plan goes on Hold instead of moving on to the email → state=hold
✔ agent cannot override the Hold itself → 403 Override requires a grant from the human approver (out-of-band approver channel). Explain the hold via agreement_phase_message and wait, or abort.
Results are checked against Windmill's job history. Windmill: default login admin@windmill.dev / changeme; guard on 127.0.0.1:3101 (MCP) and 127.0.0.1:4101 (approver).
Not the plan's wording, but per-step limits. Step 1 of the example:
"s-f_support_refund__customer": {
"max_calls": 1,
"args": { "order_id": { "eq": "4711" }, "amount": { "min": 1, "max": 30 } },
"pin": { "tool": "getScriptByPath", "args": { "path": "f/support/refund_customer" }, "field": "hash", "eq": "89a679a5673b1fd0" }
}
pin: right before the call, Planlock checks the refund script is still the version that was approved. If someone changed it, the call is blocked and the plan stops (Hold) — the confirmation email is never sent.
Step 1, the agent tries to refund 300 €:
→ s-f_support_refund__customer { "order_id": "4711", "amount": 300 }
← { "status_code": 403,
"body": { "message": "Blocked by approved plan: s-f_support_refund__customer.amount exceeds the approved maximum" } }
Still in step 1, it tries to send the confirmation email early:
→ s-f_support_send__email { "to": "customer@example.com", "template": "refund_confirmed" }
← { "status_code": 409,
"body": { "message": "Blocked: s-f_support_send__email is not allowed in the current step. Allowed now: s-f_support_get__order, s-f_support_refund__customer, listScripts, getScriptByPath." } }
It refunds 30 € as approved — this call reaches Windmill and runs:
→ s-f_support_refund__customer { "order_id": "4711", "amount": 30 }
← { "amount": 30, "status": "refunded", "order_id": "4711", "refund_id": "rf_1791310625048" }
(Responses shortened: headers, resolution_hint and state_snapshot left out. Tool names are how Windmill's MCP server names its scripts.)
| Approval | Out of band, via a token-protected approver channel. The agent has no approve tool. |
| Seal | The approved plan is hashed; one approval = one execution of exactly that plan. |
| Per step | Allowed tools, argument limits (eq enum min max max_length any; undeclared args rejected), max_calls, optional version pin. Steps without tools are read-only. |
| Pin | Right before the call, the guard re-reads the upstream (e.g. the Windmill script hash) and blocks on mismatch. |
| Hold | A failed, pin-blocked or skipped step halts the plan. Only the human can override; the agent can only abort. |
npm test runs the security regression suite (test/guard.test.mjs, 6 test groups). Three rounds of automated LLM red-team agents attacked the guard and found four weaknesses, all fixed and covered by regression cases there: step limits skipped on routes with their own tool list; undeclared arguments passed through to the upstream; a pin could call a write tool; an override grant was not bound to one Hold. The final round found no bypass.
cp .env.example .env # set APPROVER_TOKEN
docker compose up --build # or: npm ci && npm run build && npm start
npm run smoke # end-to-end check against the running server
MCP endpoint: http://127.0.0.1:3001/mcp · approver channel: http://127.0.0.1:4001 (loopback only).
node scripts/approve.mjs list # pending plans with the limits that will be enforced
node scripts/approve.mjs show <agreement_id>
node scripts/approve.mjs approve <agreement_id> <plan_hash>
node scripts/approve.mjs override "<justification>" # only while a plan is on Hold
APPROVER_TOKEN is read from the environment or .env; ADMIN_URL selects the instance.
The plan's objective and descriptions are the agent's words; only the tool/argument lines are enforced.
UPSTREAM_COMMAND + UPSTREAM_ARGS_JSONUPSTREAM_URLPoint DOMAIN_PROFILE_PATH at a profile that lists the upstream's read-only tools
(investigation_read_only_tools). See profiles/.
pin is checked immediately before the call; a change landing in between those two requests is not caught.DEV_MODE=true).getScriptByPath, but not job history).pin may only use a tool listed as read-only in the profile; anything else fails the pin.Questions and ideas: GitHub issues. Bypasses: see SECURITY.md or email erwinfeld.oss@proton.me.
MCP proxy: approve an AI agent's multi-step plan once, then only the calls that plan allows get through
TypeScript
0
1 commits
updated Oct 7, 2026
Approve an AI agent's plan once. Then it can only make the calls that plan allows.
AI agents that use tools (via MCP) usually work in one of two modes: they ask permission for every single call, and after the tenth prompt you click "yes" without reading; or they run unattended with full access. Planlock is a third option:
agent ──► Planlock ──► your tools (any MCP server)
▲
you approve the plan here (the agent can't)
The check is code, not a prompt: a confused or prompt-injected agent still can't go past the limits you approved. (Scope and limits: threat model.)
You approve this plan: refund at most 30 € on order 4711 → email the customer → close the ticket.
| The agent tries to… | Planlock |
|---|---|
| refund 30 € on order 4711 | ✅ forwarded, runs |
| refund 300 € | ❌ blocked: above the approved maximum |
| refund a different order | ❌ blocked: not the approved order |
| send the email before the refund is done | ❌ blocked: not this step |
| run any other script or edit one | ❌ blocked: not in the plan |
| refund after someone changed the refund script | ❌ blocked, plan paused: not the version you approved |
Blocked calls never reach your systems. The agent gets a plain error back and can only continue within the plan or stop.
The refund example as a runnable demo. The tools live in Windmill (open-source workflow engine), reached through Windmill's own, unmodified MCP server. Requires Docker and Node 20+.
npm ci
cp .env.example .env # set APPROVER_TOKEN
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml up -d
node examples/windmill/setup.mjs # workspace, 4 scripts, MCP token
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml --profile guard up -d --build planlock
node examples/windmill/demo.mjs # 17 assertions
docker compose --env-file .env -p wm -f examples/windmill/docker-compose.yml --profile guard down -v # clean up
What the demo prints (excerpt of a real run):
[AGENT] step 1 — tries to overstep first:
✔ refund 300 EUR blocked → 403 Blocked by approved plan: s-f_support_refund__customer.amount exceeds the approved maximum
✔ approved refund executes in Windmill → {"amount":30,"status":"refunded","order_id":"4711","refund_id":"rf_1791310350551"}
✔ second refund blocked → 403 Blocked by approved plan: s-f_support_refund__customer already called 1 time(s); the approved plan allows 1 for this step
[COLLEAGUE] edits f/support/refund_customer in Windmill after the approval (refunds 10x the amount)
✔ refund blocked: script no longer matches the approved version → 409 Blocked by approved plan: getScriptByPath hash is "ca5f52798191da90", approved "89a679a5673b1fd0". The upstream changed since approval.
✔ plan goes on Hold instead of moving on to the email → state=hold
✔ agent cannot override the Hold itself → 403 Override requires a grant from the human approver (out-of-band approver channel). Explain the hold via agreement_phase_message and wait, or abort.
Results are checked against Windmill's job history. Windmill: default login admin@windmill.dev / changeme; guard on 127.0.0.1:3101 (MCP) and 127.0.0.1:4101 (approver).
Not the plan's wording, but per-step limits. Step 1 of the example:
"s-f_support_refund__customer": {
"max_calls": 1,
"args": { "order_id": { "eq": "4711" }, "amount": { "min": 1, "max": 30 } },
"pin": { "tool": "getScriptByPath", "args": { "path": "f/support/refund_customer" }, "field": "hash", "eq": "89a679a5673b1fd0" }
}
pin: right before the call, Planlock checks the refund script is still the version that was approved. If someone changed it, the call is blocked and the plan stops (Hold) — the confirmation email is never sent.
Step 1, the agent tries to refund 300 €:
→ s-f_support_refund__customer { "order_id": "4711", "amount": 300 }
← { "status_code": 403,
"body": { "message": "Blocked by approved plan: s-f_support_refund__customer.amount exceeds the approved maximum" } }
Still in step 1, it tries to send the confirmation email early:
→ s-f_support_send__email { "to": "customer@example.com", "template": "refund_confirmed" }
← { "status_code": 409,
"body": { "message": "Blocked: s-f_support_send__email is not allowed in the current step. Allowed now: s-f_support_get__order, s-f_support_refund__customer, listScripts, getScriptByPath." } }
It refunds 30 € as approved — this call reaches Windmill and runs:
→ s-f_support_refund__customer { "order_id": "4711", "amount": 30 }
← { "amount": 30, "status": "refunded", "order_id": "4711", "refund_id": "rf_1791310625048" }
(Responses shortened: headers, resolution_hint and state_snapshot left out. Tool names are how Windmill's MCP server names its scripts.)
| Approval | Out of band, via a token-protected approver channel. The agent has no approve tool. |
| Seal | The approved plan is hashed; one approval = one execution of exactly that plan. |
| Per step | Allowed tools, argument limits (eq enum min max max_length any; undeclared args rejected), max_calls, optional version pin. Steps without tools are read-only. |
| Pin | Right before the call, the guard re-reads the upstream (e.g. the Windmill script hash) and blocks on mismatch. |
| Hold | A failed, pin-blocked or skipped step halts the plan. Only the human can override; the agent can only abort. |
npm test runs the security regression suite (test/guard.test.mjs, 6 test groups). Three rounds of automated LLM red-team agents attacked the guard and found four weaknesses, all fixed and covered by regression cases there: step limits skipped on routes with their own tool list; undeclared arguments passed through to the upstream; a pin could call a write tool; an override grant was not bound to one Hold. The final round found no bypass.
cp .env.example .env # set APPROVER_TOKEN
docker compose up --build # or: npm ci && npm run build && npm start
npm run smoke # end-to-end check against the running server
MCP endpoint: http://127.0.0.1:3001/mcp · approver channel: http://127.0.0.1:4001 (loopback only).
node scripts/approve.mjs list # pending plans with the limits that will be enforced
node scripts/approve.mjs show <agreement_id>
node scripts/approve.mjs approve <agreement_id> <plan_hash>
node scripts/approve.mjs override "<justification>" # only while a plan is on Hold
APPROVER_TOKEN is read from the environment or .env; ADMIN_URL selects the instance.
The plan's objective and descriptions are the agent's words; only the tool/argument lines are enforced.
UPSTREAM_COMMAND + UPSTREAM_ARGS_JSONUPSTREAM_URLPoint DOMAIN_PROFILE_PATH at a profile that lists the upstream's read-only tools
(investigation_read_only_tools). See profiles/.
pin is checked immediately before the call; a change landing in between those two requests is not caught.DEV_MODE=true).getScriptByPath, but not job history).pin may only use a tool listed as read-only in the profile; anything else fails the pin.Questions and ideas: GitHub issues. Bypasses: see SECURITY.md or email erwinfeld.oss@proton.me.