Your coding agent says it's done. WTF checks.
TypeScript
0
0 commits
updated Sep 20, 2026
Any agent. Any Git repo. Local. No account. No AI required.
WTF works with any coding agent (Claude Code, Cursor, Copilot, Codex, Aider) because it inspects the resulting software change, not the agent.
npx agent-wtf
$ npx agent-wtf
WTF — what just happened?
2 files changed · +7 / -5
VERIFIED
○ Tests not yet run · Run wtf verify to validate
tests (npm test)
PAY ATTENTION
1. TESTS
Test skipped or disabled
test/charge.test.js:9
> test.skip('handles VIP coupon cap calculation', () => {
ALSO
⚠ 1 skipped test
⚠ 1 debug statement (console.log)
Review surface:
12 lines to review across 2 files
Then:
$ npx agent-wtf verify
WTF — what just happened?
2 files changed · +6 / -5
VERIFIED
✓ tests (254ms)
Review surface:
11 lines to review across 2 files
Agent finishes
↓
WTF
↓
Evidence
↓
Agent fixes
↓
WTF verify
↓
Human gets receipt
Coding agents can generate more code in two minutes than you can review in an afternoon.
When an agent claims: "Done! Refactored the billing module and all tests pass." — humans are left with three questions:
WTF gives you the answers in under a second.
Most agent diffs are dominated by lockfiles, minified bundles, snapshots, and generated boilerplate.
WTF separates mechanical churn from code that actually deserves human attention:
3,812 changed lines
↓
WTF
↓
94 meaningful lines to review (97.5% compressed)
You review what matters. WTF accounts for the rest.
git diff?| Standard Tooling | What Happens with Coding Agents | How WTF Solves It |
|---|---|---|
git status | Lists modified files, but treats a 2,000-line lockfile the same as an auth timeout modification. | Separates signal from churn: Classifies mechanical lines vs. meaningful lines deserving human review. |
git diff | Floods your terminal with generated boilerplate, snapshots, and minified bundles. | Focuses human attention: Automatically highlights high-risk patterns (auth, DB migrations, env vars, debug leftovers). |
| Agent Claims | Believes the agent when it claims "refactored billing module and all tests pass". | Verifies independently: Flags skipped or disabled tests (test.skip) and generates an unforgeable local execution receipt (wtf verify). |
No installation required:
npx agent-wtf
Or install globally:
npm install -g agent-wtf
| Command | What it does |
|---|---|
wtf | See what changed and what needs attention in your working tree (~50ms). |
wtf init-agent | Automatically configure your repository for autonomous agent self-auditing. |
wtf verify | Discover and run your tests/builds to produce an independently verified receipt. |
wtf show | View exact diff snippets and line evidence for every finding. |
wtf --json | Machine-readable evidence schema (wtf/0.1) for coding agents. |
Run this once in any repository:
npx agent-wtf init-agent
This automatically configures your repository's agent rules (AGENTS.md, CLAUDE.md, .cursorrules, and .github/copilot-instructions.md).
From that moment on, whenever Claude Code, Cursor, Copilot, or Cline works in your repo, the agent autonomously:
test.skip), debug leftovers (console.log), and schema risks.wtf verify to independently validate your test suite.(To view the markdown template without modifying files, pass npx agent-wtf init-agent --print).
WTF is designed to inspect machine-generated changes, so it treats repository content as untrusted input.
shell: false) with baseline Git configuration overrides.Normal wtf analysis does not intentionally execute project code.
wtf verify is different: it runs your project’s verification commands locally with your user permissions and is not sandboxed. Only use it on code you trust to execute.
See SECURITY.md for details.
WTF is an evidence ledger, not an oracle.
We strictly avoid fabricated confidence scores (e.g. "87% safe" or "clean code guarantee"). Instead, WTF categorizes facts into four strict evidence tiers:
REPORTED: What something claims happened (e.g., an agent summary).OBSERVED: What WTF directly confirmed in the Git diff (e.g., session timeout altered, .env introduced).VERIFIED: What WTF independently executed and validated (e.g., test runner exited code 0).UNKNOWN: What available evidence cannot prove (e.g., tests exist but have not been run).WTF does not claim to catch every bug or replace human judgment. It eliminates the blind spots between what the machine claimed and what the machine actually did.
git clone https://github.com/LinusInnovator/wtf.git
cd wtf
npm install
npm run build
npm test
npm run gauntlet
Built by @LinusInnovator. Explored in depth at Great Delights.
TypeScript
96.4%
Shell
3.5%
Your coding agent says it's done. WTF checks.
TypeScript
0
0 commits
updated Sep 20, 2026
Any agent. Any Git repo. Local. No account. No AI required.
WTF works with any coding agent (Claude Code, Cursor, Copilot, Codex, Aider) because it inspects the resulting software change, not the agent.
npx agent-wtf
$ npx agent-wtf
WTF — what just happened?
2 files changed · +7 / -5
VERIFIED
○ Tests not yet run · Run wtf verify to validate
tests (npm test)
PAY ATTENTION
1. TESTS
Test skipped or disabled
test/charge.test.js:9
> test.skip('handles VIP coupon cap calculation', () => {
ALSO
⚠ 1 skipped test
⚠ 1 debug statement (console.log)
Review surface:
12 lines to review across 2 files
Then:
$ npx agent-wtf verify
WTF — what just happened?
2 files changed · +6 / -5
VERIFIED
✓ tests (254ms)
Review surface:
11 lines to review across 2 files
Agent finishes
↓
WTF
↓
Evidence
↓
Agent fixes
↓
WTF verify
↓
Human gets receipt
Coding agents can generate more code in two minutes than you can review in an afternoon.
When an agent claims: "Done! Refactored the billing module and all tests pass." — humans are left with three questions:
WTF gives you the answers in under a second.
Most agent diffs are dominated by lockfiles, minified bundles, snapshots, and generated boilerplate.
WTF separates mechanical churn from code that actually deserves human attention:
3,812 changed lines
↓
WTF
↓
94 meaningful lines to review (97.5% compressed)
You review what matters. WTF accounts for the rest.
git diff?| Standard Tooling | What Happens with Coding Agents | How WTF Solves It |
|---|---|---|
git status | Lists modified files, but treats a 2,000-line lockfile the same as an auth timeout modification. | Separates signal from churn: Classifies mechanical lines vs. meaningful lines deserving human review. |
git diff | Floods your terminal with generated boilerplate, snapshots, and minified bundles. | Focuses human attention: Automatically highlights high-risk patterns (auth, DB migrations, env vars, debug leftovers). |
| Agent Claims | Believes the agent when it claims "refactored billing module and all tests pass". | Verifies independently: Flags skipped or disabled tests (test.skip) and generates an unforgeable local execution receipt (wtf verify). |
No installation required:
npx agent-wtf
Or install globally:
npm install -g agent-wtf
| Command | What it does |
|---|---|
wtf | See what changed and what needs attention in your working tree (~50ms). |
wtf init-agent | Automatically configure your repository for autonomous agent self-auditing. |
wtf verify | Discover and run your tests/builds to produce an independently verified receipt. |
wtf show | View exact diff snippets and line evidence for every finding. |
wtf --json | Machine-readable evidence schema (wtf/0.1) for coding agents. |
Run this once in any repository:
npx agent-wtf init-agent
This automatically configures your repository's agent rules (AGENTS.md, CLAUDE.md, .cursorrules, and .github/copilot-instructions.md).
From that moment on, whenever Claude Code, Cursor, Copilot, or Cline works in your repo, the agent autonomously:
test.skip), debug leftovers (console.log), and schema risks.wtf verify to independently validate your test suite.(To view the markdown template without modifying files, pass npx agent-wtf init-agent --print).
WTF is designed to inspect machine-generated changes, so it treats repository content as untrusted input.
shell: false) with baseline Git configuration overrides.Normal wtf analysis does not intentionally execute project code.
wtf verify is different: it runs your project’s verification commands locally with your user permissions and is not sandboxed. Only use it on code you trust to execute.
See SECURITY.md for details.
WTF is an evidence ledger, not an oracle.
We strictly avoid fabricated confidence scores (e.g. "87% safe" or "clean code guarantee"). Instead, WTF categorizes facts into four strict evidence tiers:
REPORTED: What something claims happened (e.g., an agent summary).OBSERVED: What WTF directly confirmed in the Git diff (e.g., session timeout altered, .env introduced).VERIFIED: What WTF independently executed and validated (e.g., test runner exited code 0).UNKNOWN: What available evidence cannot prove (e.g., tests exist but have not been run).WTF does not claim to catch every bug or replace human judgment. It eliminates the blind spots between what the machine claimed and what the machine actually did.
git clone https://github.com/LinusInnovator/wtf.git
cd wtf
npm install
npm run build
npm test
npm run gauntlet
Built by @LinusInnovator. Explored in depth at Great Delights.
TypeScript
96.4%
Shell
3.5%