Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.
See the code
Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.
The only tool in this class that publishes the runs where it lost. Here's why.
On Pi + GPT-5.5, answers 37% the length of the same model with no ruleset (caveman 56%, ponytail 49%) · 12 agents · zero dependencies · one command
Chisle is a ruleset and hook pack that makes AI coding agents cheaper to run. It cuts what the model writes (no filler, no hedging, no speculative abstractions: the smallest code that works) and what it reads (oversized tool output is trimmed before it re-enters the context window). One npx chisle wires it into Claude Code, Pi, Cursor, Codex, Gemini, Copilot, OpenCode, Antigravity and four more agents. Every number below comes from committed raw transcripts, including the runs where it lost.
Pi 0.85.1 + GPT-5.5, 6 tasks, one run. Answer length is what you read; on total billed tokens ponytail is leanest (59% vs 61%). On the larger 26-cell Claude Code run Chisle is the only one under 100%: see below.
"Add debounce to a search input that currently fires an API call on every keystroke." Same model, same prompt, one difference: the injected ruleset. Both answers below are the verbatim committed output from benchmarks/results/raw/:
| bare agent: 142 lines, 1506 tokens | Chisle: 35 lines, 602 tokens |
|---|---|
|
Opens with "Let me show you the most common approaches", then ships a reusable generic
…then Option 2 and Option 3, a comparison table, and a caveats section. |
Asks which framework, then answers the question that was actually asked:
Then two lines on why it works, and "use |
Not golfed, boring. Same behaviour, one less abstraction, no second file, and it names the dependency you might already have instead of reinventing it.
| axis | what | how |
|---|---|---|
| Output: prose | filler, hedging, manufactured structure | zero-fluff ruleset, injected per session |
| Output: code | speculative abstractions, unrequested boilerplate | YAGNI efficiency ladder |
| Input: context | oversized tool output flooding the window | Claude PostToolUse / Pi tool_result: scrub, elide, dedup, plus prevention rules |
Where each one attaches to a session:
flowchart LR
subgraph S["Session start"]
H1["Ruleset injection<br/>once per active session"]
end
subgraph T["Every turn"]
H2["Mode tracking<br/>Claude + Pi"]
end
subgraph L["Every tool call"]
H3["PostToolUse / tool_result<br/>scrub → elide → dedup"]
end
H1 --> M(["Model"])
H2 --> M
M -->|writes| O["Output:<br/>terser prose,<br/>YAGNI-first code"]
M -->|calls a tool| TOOL[["Bash / grep / web / extension tools"]]
TOOL -->|raw output| H3
H3 -->|"compressed, rebuilt into<br/>the tool's own shape"| M
RE["Read / Edit / Write"] -.->|"never touched,<br/>exact bytes feed later edits"| M
style M fill:#1f2937,stroke:#d78a3c,color:#e6edf3
style O fill:#14532d,stroke:#2da44e,color:#e6edf3
style H3 fill:#1f2937,stroke:#2da44e,color:#e6edf3
style RE fill:#3f1d1d,stroke:#cf3b3b,color:#e6edf3
The loop on the right is the input axis: tool output is billed again on every later request in the session, so shrinking it once pays repeatedly. Read, Edit, and Write are deliberately outside it.
One command. Auto-detects your agents (Claude Code, Pi, Cursor, Windsurf, Cline, Kiro, Antigravity, Codex, Gemini, Copilot, OpenCode, Hermes) and wires each one. --uninstall puts everything back.
npx chisle
# or via curl
curl -fsSL https://raw.githubusercontent.com/JayPokale/Chisle/main/install.sh | bash
# Windows
irm https://raw.githubusercontent.com/JayPokale/Chisle/main/install.ps1 | iex
Preview first with npx chisle --dry-run, scope with --only claude or --only pi, see everything with npx chisle --help. Remove with npx chisle --uninstall.
Upgrading, per-agent setup, --stats and config: docs/usage.md.
Claude Code 2.1.285 on Haiku 4.5, 13 live prompts × 2 seeds = 26 cells per arm, billed output tokens vs the same model with no ruleset (writeup + raw cells):
| total bill | 95% CI | visible answer | worst cell | backfires | |
|---|---|---|---|---|---|
| caveman | 102% | 79–128% | 105% | 305% | 15 / 26 |
| ponytail | 105% | 86–129% | 97% | 493% | 16 / 26 |
| Chisle | 83% | 69–95% | 76% | 170% | 11 / 26 |
Chisle is the only arm below a bare model. It pays on long answers (77%) and coding prompts (76%); on short answers it breaks even (106%).
Input side: tool output is 67.5% of context in 171 measured Claude Code sessions, and it is re-billed on every later request. The compressor cut ~46% off every eligible output there, and 27.4% of tool output on top of Pi's own truncation (receipts).
| prose | code judgment | input/context | worst-case guard | publishes failures | |
|---|---|---|---|---|---|
| caveman | ✅ | ❌ | ❌ | ❌ 305% | ❌ |
| ponytail | ❌ | ✅ | ❌ | ❌ 493% | ❌ |
| headroom | ❌ | ❌ | ✅ proxy | n/a | ❌ |
| Chisle | ✅ | ✅ | ✅ hook | 170% | ✅ |
Every table, chart and caveat, including the June suite and the Pi run in full: docs/benchmarks.md. Head to head: docs/comparison.md.
Before writing code, the agent stops at the first rung that holds:
flowchart TD
A[Request for code] --> R[Read the problem fully]
R --> Q1{Does this need<br/>to exist at all?}
Q1 -->|no| S1[Skip it. Say so in one line]
Q1 -->|yes| Q2{Already in<br/>this codebase?}
Q2 -->|yes| S2[Reuse it. Don't rewrite]
Q2 -->|no| Q3{Stdlib<br/>does it?}
Q3 -->|yes| S3[Use the stdlib]
Q3 -->|no| Q4{Native platform<br/>feature covers it?}
Q4 -->|yes| S4["CSS over JS, DB constraint<br/>over app code"]
Q4 -->|no| Q5{Already-installed<br/>dependency?}
Q5 -->|yes| S5[Use it. Never add a new dep<br/>for what a few lines do]
Q5 -->|no| Q6{Can it be<br/>one line?}
Q6 -->|yes| S6[One line]
Q6 -->|no| S7[The minimum code that works]
S1 & S2 & S3 & S4 & S5 & S6 & S7 --> OUT[Ship it + note what was skipped<br/>and when to add it]
style Q1 fill:#1f2937,stroke:#d78a3c,color:#e6edf3
style OUT fill:#14532d,stroke:#2da44e,color:#e6edf3
style R fill:#1f2937,stroke:#8b949e,color:#e6edf3
The ladder runs after reading, never instead of it. Note the exit: every rung lands on the same obligation: say what you skipped, so "later" doesn't quietly become "never".
The ladder runs after reading the code, lazy about the solution and never about understanding. Lazy is not negligent: trust-boundary validation, data-loss handling, security, and accessibility are never on the chopping block.
Mark deliberate simplifications so "later" doesn't quietly become "never":
// chisle: global lock, per-account locks if throughput matters
// chisle: O(n) scan, index this when table exceeds ~10k rows
| Command | Effect |
|---|---|
| (nothing) | On automatically every session after install |
/chisle | Re-activate if you'd stopped it |
/chisle off | Deactivate |
stop chisle | Deactivate (ruleset and input-side compression) |
normal mode | Deactivate |
Natural language works too: "activate chisle", "chisle mode", "chislify this". Code symbols, function/API names, and error strings stay verbatim, so only the noise around them compresses.
See CONTRIBUTING.md. Built by Jay Pokale with Claude, Antigravity, and Codex as co-engineers.
MIT. The shortest license that works.
AI co-engineers (pair-work credited in commit trailers and the changelog):
Saved you tokens? ⭐ Star the repo. It costs zero tokens and keeps the benchmarks running.
JavaScript
75.8%
TypeScript
18.7%
Shell
3.3%
CSS
1.1%
Cut your AI coding agent's token bill on three axes: terse prose, YAGNI-first code, and tool-output compression. Claude Code, Pi, Cursor, Codex, Gemini + 4 more. Zero deps, published benchmarks including the runs it loses.
See the code
Your AI talks less, builds less, reads less, and says more. Like a senior dev who bills by the syllable.
The only tool in this class that publishes the runs where it lost. Here's why.
On Pi + GPT-5.5, answers 37% the length of the same model with no ruleset (caveman 56%, ponytail 49%) · 12 agents · zero dependencies · one command
Chisle is a ruleset and hook pack that makes AI coding agents cheaper to run. It cuts what the model writes (no filler, no hedging, no speculative abstractions: the smallest code that works) and what it reads (oversized tool output is trimmed before it re-enters the context window). One npx chisle wires it into Claude Code, Pi, Cursor, Codex, Gemini, Copilot, OpenCode, Antigravity and four more agents. Every number below comes from committed raw transcripts, including the runs where it lost.
Pi 0.85.1 + GPT-5.5, 6 tasks, one run. Answer length is what you read; on total billed tokens ponytail is leanest (59% vs 61%). On the larger 26-cell Claude Code run Chisle is the only one under 100%: see below.
"Add debounce to a search input that currently fires an API call on every keystroke." Same model, same prompt, one difference: the injected ruleset. Both answers below are the verbatim committed output from benchmarks/results/raw/:
| bare agent: 142 lines, 1506 tokens | Chisle: 35 lines, 602 tokens |
|---|---|
|
Opens with "Let me show you the most common approaches", then ships a reusable generic
…then Option 2 and Option 3, a comparison table, and a caveats section. |
Asks which framework, then answers the question that was actually asked:
Then two lines on why it works, and "use |
Not golfed, boring. Same behaviour, one less abstraction, no second file, and it names the dependency you might already have instead of reinventing it.
| axis | what | how |
|---|---|---|
| Output: prose | filler, hedging, manufactured structure | zero-fluff ruleset, injected per session |
| Output: code | speculative abstractions, unrequested boilerplate | YAGNI efficiency ladder |
| Input: context | oversized tool output flooding the window | Claude PostToolUse / Pi tool_result: scrub, elide, dedup, plus prevention rules |
Where each one attaches to a session:
flowchart LR
subgraph S["Session start"]
H1["Ruleset injection<br/>once per active session"]
end
subgraph T["Every turn"]
H2["Mode tracking<br/>Claude + Pi"]
end
subgraph L["Every tool call"]
H3["PostToolUse / tool_result<br/>scrub → elide → dedup"]
end
H1 --> M(["Model"])
H2 --> M
M -->|writes| O["Output:<br/>terser prose,<br/>YAGNI-first code"]
M -->|calls a tool| TOOL[["Bash / grep / web / extension tools"]]
TOOL -->|raw output| H3
H3 -->|"compressed, rebuilt into<br/>the tool's own shape"| M
RE["Read / Edit / Write"] -.->|"never touched,<br/>exact bytes feed later edits"| M
style M fill:#1f2937,stroke:#d78a3c,color:#e6edf3
style O fill:#14532d,stroke:#2da44e,color:#e6edf3
style H3 fill:#1f2937,stroke:#2da44e,color:#e6edf3
style RE fill:#3f1d1d,stroke:#cf3b3b,color:#e6edf3
The loop on the right is the input axis: tool output is billed again on every later request in the session, so shrinking it once pays repeatedly. Read, Edit, and Write are deliberately outside it.
One command. Auto-detects your agents (Claude Code, Pi, Cursor, Windsurf, Cline, Kiro, Antigravity, Codex, Gemini, Copilot, OpenCode, Hermes) and wires each one. --uninstall puts everything back.
npx chisle
# or via curl
curl -fsSL https://raw.githubusercontent.com/JayPokale/Chisle/main/install.sh | bash
# Windows
irm https://raw.githubusercontent.com/JayPokale/Chisle/main/install.ps1 | iex
Preview first with npx chisle --dry-run, scope with --only claude or --only pi, see everything with npx chisle --help. Remove with npx chisle --uninstall.
Upgrading, per-agent setup, --stats and config: docs/usage.md.
Claude Code 2.1.285 on Haiku 4.5, 13 live prompts × 2 seeds = 26 cells per arm, billed output tokens vs the same model with no ruleset (writeup + raw cells):
| total bill | 95% CI | visible answer | worst cell | backfires | |
|---|---|---|---|---|---|
| caveman | 102% | 79–128% | 105% | 305% | 15 / 26 |
| ponytail | 105% | 86–129% | 97% | 493% | 16 / 26 |
| Chisle | 83% | 69–95% | 76% | 170% | 11 / 26 |
Chisle is the only arm below a bare model. It pays on long answers (77%) and coding prompts (76%); on short answers it breaks even (106%).
Input side: tool output is 67.5% of context in 171 measured Claude Code sessions, and it is re-billed on every later request. The compressor cut ~46% off every eligible output there, and 27.4% of tool output on top of Pi's own truncation (receipts).
| prose | code judgment | input/context | worst-case guard | publishes failures | |
|---|---|---|---|---|---|
| caveman | ✅ | ❌ | ❌ | ❌ 305% | ❌ |
| ponytail | ❌ | ✅ | ❌ | ❌ 493% | ❌ |
| headroom | ❌ | ❌ | ✅ proxy | n/a | ❌ |
| Chisle | ✅ | ✅ | ✅ hook | 170% | ✅ |
Every table, chart and caveat, including the June suite and the Pi run in full: docs/benchmarks.md. Head to head: docs/comparison.md.
Before writing code, the agent stops at the first rung that holds:
flowchart TD
A[Request for code] --> R[Read the problem fully]
R --> Q1{Does this need<br/>to exist at all?}
Q1 -->|no| S1[Skip it. Say so in one line]
Q1 -->|yes| Q2{Already in<br/>this codebase?}
Q2 -->|yes| S2[Reuse it. Don't rewrite]
Q2 -->|no| Q3{Stdlib<br/>does it?}
Q3 -->|yes| S3[Use the stdlib]
Q3 -->|no| Q4{Native platform<br/>feature covers it?}
Q4 -->|yes| S4["CSS over JS, DB constraint<br/>over app code"]
Q4 -->|no| Q5{Already-installed<br/>dependency?}
Q5 -->|yes| S5[Use it. Never add a new dep<br/>for what a few lines do]
Q5 -->|no| Q6{Can it be<br/>one line?}
Q6 -->|yes| S6[One line]
Q6 -->|no| S7[The minimum code that works]
S1 & S2 & S3 & S4 & S5 & S6 & S7 --> OUT[Ship it + note what was skipped<br/>and when to add it]
style Q1 fill:#1f2937,stroke:#d78a3c,color:#e6edf3
style OUT fill:#14532d,stroke:#2da44e,color:#e6edf3
style R fill:#1f2937,stroke:#8b949e,color:#e6edf3
The ladder runs after reading, never instead of it. Note the exit: every rung lands on the same obligation: say what you skipped, so "later" doesn't quietly become "never".
The ladder runs after reading the code, lazy about the solution and never about understanding. Lazy is not negligent: trust-boundary validation, data-loss handling, security, and accessibility are never on the chopping block.
Mark deliberate simplifications so "later" doesn't quietly become "never":
// chisle: global lock, per-account locks if throughput matters
// chisle: O(n) scan, index this when table exceeds ~10k rows
| Command | Effect |
|---|---|
| (nothing) | On automatically every session after install |
/chisle | Re-activate if you'd stopped it |
/chisle off | Deactivate |
stop chisle | Deactivate (ruleset and input-side compression) |
normal mode | Deactivate |
Natural language works too: "activate chisle", "chisle mode", "chislify this". Code symbols, function/API names, and error strings stay verbatim, so only the noise around them compresses.
See CONTRIBUTING.md. Built by Jay Pokale with Claude, Antigravity, and Codex as co-engineers.
MIT. The shortest license that works.
AI co-engineers (pair-work credited in commit trailers and the changelog):
Saved you tokens? ⭐ Star the repo. It costs zero tokens and keeps the benchmarks running.
JavaScript
75.8%
TypeScript
18.7%
Shell
3.3%
CSS
1.1%