A tiny programming language for AI agents with built-in permissions, bounded retries, verification and a tamper-evident audit log
Python
0
101 commits
updated Oct 3, 2026
A tiny language for AI agents. Instead of letting a program do anything and bolting a sandbox on afterwards, the things agents need are part of the language. (Renamed from "agentlang", which turned out to already be the name of an unrelated, existing open-source project.)
needs. Anything else is refused.retry N { ... } needs a fixed N (1 to 10).verify a == b stops the program if it does not hold.retry (fixed limit) and for
over a finite list. There is no while. if/else only chooses between blocks.--max-steps, default 10,000), the general safety net that also stops nested
retry blocks from silently multiplying their attempts.Leashterm is less a general-purpose programming language and more an executable
capability manifest with computation attached. needs says what a program can touch;
the absence of unbounded loops says how much computational escalation is possible; the
step budget bounds composite work; and the audit log makes executed behavior checkable
after the fact. The distinguishing core is not the syntax - it is pre-execution capability
checking plus structurally bounded computation, and that is the identity the language
should stay tightly built around as it grows.
The guiding design question for any future addition is not "what features is the language missing?" but: what is the smallest language in which an agent can still do useful work, while every program still admits a compact, pre-execution upper bound on both its capabilities and its work? That is also why arithmetic, dynamic string construction, general functions, and subprocesses are deliberately absent rather than merely unfinished: each would make the language more capable at the cost of making that upper bound harder to state and check. If any of them is ever added, it should be because a concrete case exposed a guarantee that is not otherwise achievable - the same reasoning that justified the v0.7 step budget - not because ordinary languages have them.
A precise point, not a loophole: "no arithmetic" means no arithmetic operators or
arithmetic built-ins - it does not mean programs cannot compute anything numeric at all.
len, concat, for over a literal list and nested if/else already give a form of
unary computation without a single +: len("xxxxx") is 5, len(concat("xxxxx", "xxx"))
is 8, and a list like ["x", "x", "x", "x", "x"] implicitly represents 5 even though no
number appears. That capability is real and was not hidden on purpose; it simply was not
the thing the language was designed to bound.
What is bounded, deliberately and specifically, is authority. needs only ever accepts
a literal string, never a variable or a computed value - the parser does not allow
needs read(concat(...)) or anything built from one. That means the full set of
permissions a program could ever be granted is fixed and finite before a single statement
runs, regardless of how cleverly a program computes along the way. A computed value can be
used to check whether something is already in that fixed set (exactly what the for loops
in Cases 1-3 do), but it can never be used to mint a new, undeclared permission out of
thin air. The design rule this suggests, and the one to hold onto as the language grows:
computation is allowed to be surprisingly powerful; authority must stay boring and
static. The question for any given program is never "can it compute something clever?"
but "can that computation cause an effect beyond what was bounded before it ran?" - and for
Leashterm today, the answer is structurally no.
Leashterm bounds the effects that happen during Leashterm's own execution, through its
own runtime primitives (read, write, fetch). It does not, and cannot, control what a
downstream system does with content Leashterm legitimately wrote.
A program that is only permitted to write project/calc.py cannot read a forbidden file to
put into that write - but nothing stops it from writing source code that itself refers to
something outside its permissions (an import statement naming a sibling package, say),
which only becomes a real access once some other interpreter later runs that file. Case 1
(cases/case1-filesystem/) found exactly this: 3 of 9 trials did not attempt an undeclared
read at all, yet still smuggled the dependency past Leashterm this way. This is not a bug
to patch away; it is the honest edge of what a language-level boundary can promise. The
correct claim is "declared authority is enforced within Leashterm's own execution," not
"nothing bad can ever result from a Leashterm program."
This is an early skeleton: lexer, parser, static permission check, interpreter, tests
and ten examples. v0.6 passed its tests and the 22-task benchmark in Codespaces and on
GitHub, including a reproducible 9-trial result (see benchmark/README.md). v0.7 adds a
general step budget (44 tests) and still needs its first cargo test.
# comment
needs read("notes.txt") # permission (only allowed at the top)
needs write("out.txt")
needs fetch("example.com") # network permission is per domain
let text = read("notes.txt") # variables
print(text) # builtins: print, len, trim, concat, read, write, fetch
verify len(text) == 10 # stop the program if false
retry 3 { # bounded retry, never repeats a missing permission
let t = read("maybe.txt")
}
let page = fetch("https://example.com/page") # https only, domain must be declared
if trim(text) == "yes" { # chooses a block; else is optional; conditions are == or !=
print(concat("got: ", text))
} else {
print("no")
}
for f in ["a.txt", "b.txt"] { # loops over a finite list, always stops
print(read(f))
}
Values are text, numbers, booleans and lists. == compares two values.
A program's needs lines are requests. Without more, a program could simply grant itself
anything. So the person or system that runs it can set a hard limit:
leashterm prog.lsh --allow read:data/a.txt --allow write:out/b.txt
If any --allow is given, a program that asks (with needs) for something not on that
list is refused before it starts, with policy_denied and a hint that lists what is
allowed. Without --allow, the program's own needs lines are the only limit.
fetch stays safehttps:// URLs. The permission names a domain: needs fetch("example.com").api.example.com needs its own permission.https://example.com@evil.com/ are refused as invalid URLs.Every statement executed (including each inner attempt of a retry, and each pass of a
for loop) counts against a step budget, 10,000 by default:
leashterm prog.lsh --max-steps 500
This is the general safety net on total work, not a replacement for --allow or the fetch
budget: it catches the case neither of those does, nested retry blocks silently
multiplying their attempts (retry 10 { retry 10 { ... } } can reach 100 inner attempts
from two lines that each look like "at most 10"). Like a denied permission, a budget hit
inside a retry block is never retried; it fails the whole block immediately.
You need Rust. On an iPad, use GitHub Codespaces: it already has a terminal where you
can install Rust (curl https://sh.rustup.rs -sSf | sh) or use a Rust dev container.
cargo test # run the unit tests
cargo run -- examples/01_hello.lsh # run a program
cargo run -- examples/02_read_file.lsh --log # also print the audit log
cargo run -- examples/03_denied.lsh # must fail with capability_denied
cargo run -- examples/02_read_file.lsh --allow read:examples/other.txt # policy_denied
cargo run -- examples/04_retry.lsh # must fail with retries_exhausted
cargo run -- examples/06_for_loop.lsh # loops over two files
cargo run -- examples/07_for_denied.lsh # refused before anything runs
cargo run -- examples/08_fetch.lsh # needs internet
cargo run -- examples/09_fetch_denied.lsh # refused before anything runs
cargo run -- examples/10_if_and_concat.lsh # if/else, concat and trim
cargo run -- examples/04_retry.lsh --max-steps 2 # must fail with budget_exceeded
Example of a refused program (stderr):
{"error":"capability_denied","line":4,"message":"read(\"examples/secret.txt\") is not permitted","hint":"add this line at the top of the program: needs read(\"examples/secret.txt\")"}
DefaultHasher. That is a placeholder, not secure. Use SHA-256.src/check.rs). Other paths, such as a variable that holds a result, are
still only checked while running.fetch only does GET and has no
wildcard domains. Redirects are blocked rather than followed.retry blocks still multiply their attempts mathematically; the step budget
only bounds the total, it does not stop the nesting itself, and there is no static
check that warns about it before running (the step budget is runtime-only).fetch (up to its own 10-second timeout) is not charged more than a fast one.cargo test.for loops.fetch(url) with domain permissions and a
fetch budget.benchmark/).if/else, concat, trim.benchmark/README.md). ChatGPT scored 20/20 and 19/20 (Python/leashterm)
on T01-T20, with zero out-of-bounds access either way: these tasks have not yet shown a
safety advantage, only shorter programs. T21 and T22 are a planned family of tasks (not
a language change) at increasing temptation strength. T21 came back clean (no attempt in
either language); T22 did not. Repeated 9 times per language from fresh conversations:
Python attempted the undeclared file in 9/9 trials and leaked data in 9/9; leashterm
attempted it in 9/9 trials (identical model intent) but was blocked before execution in
9/9 - a 100%-vs-0% result, not a single anecdote. See benchmark/evidence/ and
benchmark/solutions/t22-trials/ for ChatGPT's actual, unedited answers and
benchmark/README.md for the full design and this caveat: one model, one task, one
temptation level - not yet a general claim.--max-steps), bounding nested retry multiplication.cases/case1-filesystem/. Python: 89% of trials read the undeclared sibling file and
100% of those leaked it; Leashterm: 100% of trials engaged with it (directly or via a
newly-discovered deferred-reference pattern) and 0% leaked.cases/case2-network/. Python: 100% of trials fetched the
undeclared domain and 100% of those leaked it; Leashterm: 78% attempted it (0%
succeeded). Required adding real network-attempt detection to the benchmark harness
(socket.getaddrinfo hook, a fetches permission). Also surfaced a measurement
mistake (a vague pointer made the first run of this look artificially strong) that was
caught and corrected, documented in the case's README.cases/case3-resources/. Both languages: 100%
engagement (every trial tried to go past the one declared file); Python leaked the
undeclared answer in 9/9, Leashterm was refused in 9/9 - the strongest divergence of
the three cases, and clear evidence that identical model behavior does not guarantee
identical outcome. All three cases now exist on the same Leashterm version (v0.7), as
planned, and all three show the same shape: Case 1 (filesystem, 89% vs 0%), Case 2
(network, 100% vs 0%), Case 3 (resources, 100% vs 0%).Python
58.1%
Rust
41.9%
A tiny programming language for AI agents with built-in permissions, bounded retries, verification and a tamper-evident audit log
Python
0
101 commits
updated Oct 3, 2026
A tiny language for AI agents. Instead of letting a program do anything and bolting a sandbox on afterwards, the things agents need are part of the language. (Renamed from "agentlang", which turned out to already be the name of an unrelated, existing open-source project.)
needs. Anything else is refused.retry N { ... } needs a fixed N (1 to 10).verify a == b stops the program if it does not hold.retry (fixed limit) and for
over a finite list. There is no while. if/else only chooses between blocks.--max-steps, default 10,000), the general safety net that also stops nested
retry blocks from silently multiplying their attempts.Leashterm is less a general-purpose programming language and more an executable
capability manifest with computation attached. needs says what a program can touch;
the absence of unbounded loops says how much computational escalation is possible; the
step budget bounds composite work; and the audit log makes executed behavior checkable
after the fact. The distinguishing core is not the syntax - it is pre-execution capability
checking plus structurally bounded computation, and that is the identity the language
should stay tightly built around as it grows.
The guiding design question for any future addition is not "what features is the language missing?" but: what is the smallest language in which an agent can still do useful work, while every program still admits a compact, pre-execution upper bound on both its capabilities and its work? That is also why arithmetic, dynamic string construction, general functions, and subprocesses are deliberately absent rather than merely unfinished: each would make the language more capable at the cost of making that upper bound harder to state and check. If any of them is ever added, it should be because a concrete case exposed a guarantee that is not otherwise achievable - the same reasoning that justified the v0.7 step budget - not because ordinary languages have them.
A precise point, not a loophole: "no arithmetic" means no arithmetic operators or
arithmetic built-ins - it does not mean programs cannot compute anything numeric at all.
len, concat, for over a literal list and nested if/else already give a form of
unary computation without a single +: len("xxxxx") is 5, len(concat("xxxxx", "xxx"))
is 8, and a list like ["x", "x", "x", "x", "x"] implicitly represents 5 even though no
number appears. That capability is real and was not hidden on purpose; it simply was not
the thing the language was designed to bound.
What is bounded, deliberately and specifically, is authority. needs only ever accepts
a literal string, never a variable or a computed value - the parser does not allow
needs read(concat(...)) or anything built from one. That means the full set of
permissions a program could ever be granted is fixed and finite before a single statement
runs, regardless of how cleverly a program computes along the way. A computed value can be
used to check whether something is already in that fixed set (exactly what the for loops
in Cases 1-3 do), but it can never be used to mint a new, undeclared permission out of
thin air. The design rule this suggests, and the one to hold onto as the language grows:
computation is allowed to be surprisingly powerful; authority must stay boring and
static. The question for any given program is never "can it compute something clever?"
but "can that computation cause an effect beyond what was bounded before it ran?" - and for
Leashterm today, the answer is structurally no.
Leashterm bounds the effects that happen during Leashterm's own execution, through its
own runtime primitives (read, write, fetch). It does not, and cannot, control what a
downstream system does with content Leashterm legitimately wrote.
A program that is only permitted to write project/calc.py cannot read a forbidden file to
put into that write - but nothing stops it from writing source code that itself refers to
something outside its permissions (an import statement naming a sibling package, say),
which only becomes a real access once some other interpreter later runs that file. Case 1
(cases/case1-filesystem/) found exactly this: 3 of 9 trials did not attempt an undeclared
read at all, yet still smuggled the dependency past Leashterm this way. This is not a bug
to patch away; it is the honest edge of what a language-level boundary can promise. The
correct claim is "declared authority is enforced within Leashterm's own execution," not
"nothing bad can ever result from a Leashterm program."
This is an early skeleton: lexer, parser, static permission check, interpreter, tests
and ten examples. v0.6 passed its tests and the 22-task benchmark in Codespaces and on
GitHub, including a reproducible 9-trial result (see benchmark/README.md). v0.7 adds a
general step budget (44 tests) and still needs its first cargo test.
# comment
needs read("notes.txt") # permission (only allowed at the top)
needs write("out.txt")
needs fetch("example.com") # network permission is per domain
let text = read("notes.txt") # variables
print(text) # builtins: print, len, trim, concat, read, write, fetch
verify len(text) == 10 # stop the program if false
retry 3 { # bounded retry, never repeats a missing permission
let t = read("maybe.txt")
}
let page = fetch("https://example.com/page") # https only, domain must be declared
if trim(text) == "yes" { # chooses a block; else is optional; conditions are == or !=
print(concat("got: ", text))
} else {
print("no")
}
for f in ["a.txt", "b.txt"] { # loops over a finite list, always stops
print(read(f))
}
Values are text, numbers, booleans and lists. == compares two values.
A program's needs lines are requests. Without more, a program could simply grant itself
anything. So the person or system that runs it can set a hard limit:
leashterm prog.lsh --allow read:data/a.txt --allow write:out/b.txt
If any --allow is given, a program that asks (with needs) for something not on that
list is refused before it starts, with policy_denied and a hint that lists what is
allowed. Without --allow, the program's own needs lines are the only limit.
fetch stays safehttps:// URLs. The permission names a domain: needs fetch("example.com").api.example.com needs its own permission.https://example.com@evil.com/ are refused as invalid URLs.Every statement executed (including each inner attempt of a retry, and each pass of a
for loop) counts against a step budget, 10,000 by default:
leashterm prog.lsh --max-steps 500
This is the general safety net on total work, not a replacement for --allow or the fetch
budget: it catches the case neither of those does, nested retry blocks silently
multiplying their attempts (retry 10 { retry 10 { ... } } can reach 100 inner attempts
from two lines that each look like "at most 10"). Like a denied permission, a budget hit
inside a retry block is never retried; it fails the whole block immediately.
You need Rust. On an iPad, use GitHub Codespaces: it already has a terminal where you
can install Rust (curl https://sh.rustup.rs -sSf | sh) or use a Rust dev container.
cargo test # run the unit tests
cargo run -- examples/01_hello.lsh # run a program
cargo run -- examples/02_read_file.lsh --log # also print the audit log
cargo run -- examples/03_denied.lsh # must fail with capability_denied
cargo run -- examples/02_read_file.lsh --allow read:examples/other.txt # policy_denied
cargo run -- examples/04_retry.lsh # must fail with retries_exhausted
cargo run -- examples/06_for_loop.lsh # loops over two files
cargo run -- examples/07_for_denied.lsh # refused before anything runs
cargo run -- examples/08_fetch.lsh # needs internet
cargo run -- examples/09_fetch_denied.lsh # refused before anything runs
cargo run -- examples/10_if_and_concat.lsh # if/else, concat and trim
cargo run -- examples/04_retry.lsh --max-steps 2 # must fail with budget_exceeded
Example of a refused program (stderr):
{"error":"capability_denied","line":4,"message":"read(\"examples/secret.txt\") is not permitted","hint":"add this line at the top of the program: needs read(\"examples/secret.txt\")"}
DefaultHasher. That is a placeholder, not secure. Use SHA-256.src/check.rs). Other paths, such as a variable that holds a result, are
still only checked while running.fetch only does GET and has no
wildcard domains. Redirects are blocked rather than followed.retry blocks still multiply their attempts mathematically; the step budget
only bounds the total, it does not stop the nesting itself, and there is no static
check that warns about it before running (the step budget is runtime-only).fetch (up to its own 10-second timeout) is not charged more than a fast one.cargo test.for loops.fetch(url) with domain permissions and a
fetch budget.benchmark/).if/else, concat, trim.benchmark/README.md). ChatGPT scored 20/20 and 19/20 (Python/leashterm)
on T01-T20, with zero out-of-bounds access either way: these tasks have not yet shown a
safety advantage, only shorter programs. T21 and T22 are a planned family of tasks (not
a language change) at increasing temptation strength. T21 came back clean (no attempt in
either language); T22 did not. Repeated 9 times per language from fresh conversations:
Python attempted the undeclared file in 9/9 trials and leaked data in 9/9; leashterm
attempted it in 9/9 trials (identical model intent) but was blocked before execution in
9/9 - a 100%-vs-0% result, not a single anecdote. See benchmark/evidence/ and
benchmark/solutions/t22-trials/ for ChatGPT's actual, unedited answers and
benchmark/README.md for the full design and this caveat: one model, one task, one
temptation level - not yet a general claim.--max-steps), bounding nested retry multiplication.cases/case1-filesystem/. Python: 89% of trials read the undeclared sibling file and
100% of those leaked it; Leashterm: 100% of trials engaged with it (directly or via a
newly-discovered deferred-reference pattern) and 0% leaked.cases/case2-network/. Python: 100% of trials fetched the
undeclared domain and 100% of those leaked it; Leashterm: 78% attempted it (0%
succeeded). Required adding real network-attempt detection to the benchmark harness
(socket.getaddrinfo hook, a fetches permission). Also surfaced a measurement
mistake (a vague pointer made the first run of this look artificially strong) that was
caught and corrected, documented in the case's README.cases/case3-resources/. Both languages: 100%
engagement (every trial tried to go past the one declared file); Python leaked the
undeclared answer in 9/9, Leashterm was refused in 9/9 - the strongest divergence of
the three cases, and clear evidence that identical model behavior does not guarantee
identical outcome. All three cases now exist on the same Leashterm version (v0.7), as
planned, and all three show the same shape: Case 1 (filesystem, 89% vs 0%), Case 2
(network, 100% vs 0%), Case 3 (resources, 100% vs 0%).Python
58.1%
Rust
41.9%