Deterministic scanner that deletes tests that can never fail — asks the type checker instead of an LLM. TypeScript (Jest/Vitest/bun:test) + Python (pytest/mypy).
Python
13
81 commits
updated Sep 30, 2026
AI agents write tests that can never fail. This skill deletes them.
Every AI-generated test suite has them:
assert isinstance(add(1, 2), int) # the function is typed `-> int`. mypy knew.
expect(true).toBe(true); // bold claim
Tests like these can't catch a regression — they re-assert what the compiler already proved, or assert nothing at all. They burn CI minutes and inflate coverage into false confidence. Empirical studies find test smells in 38–100% of LLM-generated test suites.
Detection is fully deterministic — no LLM in the loop, one script invocation per repo. Instead of asking an AI "is this test useful?", it asks the tools that already know:
typeof result === 'number' is a fact of the
type, the assertion is not a test. Plain-JS test files (.test.js etc.) get
the syntactic categories; type-guaranteed needs TypeScript (or stays off
without checkJs).mypy reveal_type probes (one mypy run for the whole
repo) answer the same for isinstance(...) and is not None.return, assertions swallowed by empty catch/except: pass, duplicate
bodies, len(x) >= 0.Every finding is one of two tiers, and each tier is handled by whoever can actually decide it:
--fix, no LLM
involved. The scripts guard every known way types lie: any/unknown,
as casts, non-null !, index signatures, unchecked index access, structural
instanceof (only nominal classes are provable), cast()/type: ignore.instanceof, mock-echo variants, pytest.raises(Exception), unawaited async
assertions, rotten-green conditional asserts). The script writes down exactly
why it's uncertain, and the coding agent running the skill adjudicates each
one against the surrounding code — delete, keep, or rewrite — then
proposes the calls for your approval before touching anything.It also knows what not to flag: toBeDefined() on .find() results
(T | undefined — a real check), enum contract locks, custom assertion
helpers, deliberate determinism tests, and "must not raise" contract tests.
The skill now asks four questions the scanner cannot answer: What observable contract does the test protect? What regression would make it fail? Is that contract already covered at a stronger boundary? Does the test require a production seam used only by tests? For bug regressions, run the test against the pre-fix behavior when safe to confirm it fails for the intended reason.
This review catches leads such as copied implementation expectations, source
greps, and mocks that supply the asserted outcome. They require reading the
production path and existing tests; none become automatic deletions. See the
authoring gate and
prevention examples.
.rejects.toThrow(...) never awaited, so it
silently never ran.The honest headline isn't a line count — it's that --fix only removes what it
can prove, so whatever it deletes could never have caught a regression. Clone
it and run it on your own repo.
npx skills add shmulc8/captain-obvious@captain-obvious
Also ships a .claude-plugin/plugin.json manifest, or copy
skills/captain-obvious/ into ~/.claude/skills/. Then ask your agent to
"clean up the useless tests" — or run the scanners directly:
# report-only (add --json out.json to save)
node skills/captain-obvious/scripts/captain_obvious_ts.mjs --project <repo>
python3 skills/captain-obvious/scripts/captain_obvious_py.py --path <repo> --mypy "uv run mypy"
# delete proven findings (needs a clean git tree — review the diff after)
node ... --fix
# confirm rotten-green asserts against real coverage (lcov / istanbul / coverage.py)
node ... --coverage coverage/lcov.info
python3 ... --coverage coverage.json
With --coverage, a conditional-assert whose line never executed is promoted
to proven rotten; one that did execute is a confirmed false positive and is
dropped — the dynamic half of the ICSE'19 analysis, using coverage your runner
already emits.
Trust boundary: the scan runs the project's own toolchain — mypy (with whatever plugins the repo's config loads),
uv/poetryenvironments, the repo's owntypescriptpackage. Treat scanning like running the repo's type checker: only do it on code you trust.
Requires Python ≥ 3.9 for the Python scanner and hook; the TS scanner uses the target repo's own typescript and requires ≥ 4.0.
Removing dead tests after the fact pays for them twice — once to write them,
once to clean them up. Installed as a plugin, captain-obvious also ships a
PreToolUse hook that catches them before they land: whenever the agent
writes or edits a test file, the pending content runs through a fast
syntactic-only single-file scan (--file --stdin; no mypy, no tsc, no
side effects), and the call is denied with a per-finding reason if it
would introduce proven can-never-fail patterns. The agent fixes the test on
the spot.
CAPTAIN_OBVIOUS_HOOK=off (default) | block | warn.npx skills add) don't get hooks — paste the
test-writing rules from
skills/captain-obvious/references/prevention.md
into your CLAUDE.md instead.--check --base <ref>)The write-time hook only fires inside an agent session. --check --base <ref>
is the CI twin of that hook — a correctness gate on new dead tests, not a
CI-time coverage optimizer. It is report-only (never writes; --check --fix is
rejected) and exits 1 only when a proven syntactic finding is newly
introduced vs <ref> in a changed test file — pre-existing findings never
gate, and it excludes type-guaranteed findings (the base-side scan is
syntactic-only). Any git trouble other than "file absent in base" (bad ref,
shallow clone, diff failure) fails open: a captain-obvious: note to stderr
and exit 0, so the tool can never invent a CI failure.
# .github/workflows/captain-obvious.yml
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # need history to diff against the base
- uses: <owner>/captain-obvious@main
with:
base: ${{ github.event.pull_request.base.sha }} # default: origin/main
A .pre-commit-hooks.yaml (id: captain-obvious-check, --base HEAD) ships
for pre-commit users. Both wrappers run --no-types for speed and syntactic
parity with the comparison.
| Category | Example | Level |
|---|---|---|
type-guaranteed | expect(typeof f()).toBe('number') when f(): number; assert x is not None on non-Optional | proven |
constant-assert | expect(true).toBe(true), assert x == x (incl. self.assertEqual(1, 1)-style unittest assert-methods, report-only) | proven |
boundary-tautology | expect(arr.length).toBeGreaterThanOrEqual(0) | proven |
local-const-echo | const expected = 5; expect(expected).toBe(5) | proven |
mock-echo | stub returns 5 → assert it returns 5 | proven / advisory |
dead-assert / swallowed-assert / never-asserts | assertion after return, or inside try {} catch {} | proven |
missed-fail | a forced-fail marker (pytest.fail() / throw new Error) stuck in dead code — can never fire | proven |
missed-skip | a conditional early return/skip above the asserts — if it fires, they never run | advisory |
duplicate-test | identical body in the same suite | proven |
silent-smoke | assertion-free test whose every call is silently try/caught — or that contains no calls at all — it can do nothing and can never fail | proven |
no-assert | no assertion anywhere — a smoke test (legit by design, per ICSE '19), surfaced not deleted | advisory |
conditional-assert | assertion gated behind if — rotten green (ICSE '19) | advisory |
floating-async-assert | unawaited expect(p).rejects... (silent pass under Jest) | advisory |
smoke-only | expect(fn).not.toThrow() as the only check | advisory |
self-compare-call | expect(f(a)).toEqual(f(a)) | advisory |
broad-raises | pytest.raises(Exception) as the only check | advisory |
skipped-test | it.skip / xit / @pytest.mark.skip — never runs | advisory |
Full semantics, escape hatches, and false-positive guards — plus the academic
grounding (the rotten-green lineage: Delplanque ICSE '19 → RTj Java '19 →
Robinson ESEC/FSE '23 Google Test, which this tool extends to TS + Python) —
are in references/detectors.md.
result.length >= 0 is caught,
result.length < 1e9 is not) — only mutation testing proves those useless.Python
73.2%
JavaScript
26.8%
Deterministic scanner that deletes tests that can never fail — asks the type checker instead of an LLM. TypeScript (Jest/Vitest/bun:test) + Python (pytest/mypy).
Python
13
81 commits
updated Sep 30, 2026
AI agents write tests that can never fail. This skill deletes them.
Every AI-generated test suite has them:
assert isinstance(add(1, 2), int) # the function is typed `-> int`. mypy knew.
expect(true).toBe(true); // bold claim
Tests like these can't catch a regression — they re-assert what the compiler already proved, or assert nothing at all. They burn CI minutes and inflate coverage into false confidence. Empirical studies find test smells in 38–100% of LLM-generated test suites.
Detection is fully deterministic — no LLM in the loop, one script invocation per repo. Instead of asking an AI "is this test useful?", it asks the tools that already know:
typeof result === 'number' is a fact of the
type, the assertion is not a test. Plain-JS test files (.test.js etc.) get
the syntactic categories; type-guaranteed needs TypeScript (or stays off
without checkJs).mypy reveal_type probes (one mypy run for the whole
repo) answer the same for isinstance(...) and is not None.return, assertions swallowed by empty catch/except: pass, duplicate
bodies, len(x) >= 0.Every finding is one of two tiers, and each tier is handled by whoever can actually decide it:
--fix, no LLM
involved. The scripts guard every known way types lie: any/unknown,
as casts, non-null !, index signatures, unchecked index access, structural
instanceof (only nominal classes are provable), cast()/type: ignore.instanceof, mock-echo variants, pytest.raises(Exception), unawaited async
assertions, rotten-green conditional asserts). The script writes down exactly
why it's uncertain, and the coding agent running the skill adjudicates each
one against the surrounding code — delete, keep, or rewrite — then
proposes the calls for your approval before touching anything.It also knows what not to flag: toBeDefined() on .find() results
(T | undefined — a real check), enum contract locks, custom assertion
helpers, deliberate determinism tests, and "must not raise" contract tests.
The skill now asks four questions the scanner cannot answer: What observable contract does the test protect? What regression would make it fail? Is that contract already covered at a stronger boundary? Does the test require a production seam used only by tests? For bug regressions, run the test against the pre-fix behavior when safe to confirm it fails for the intended reason.
This review catches leads such as copied implementation expectations, source
greps, and mocks that supply the asserted outcome. They require reading the
production path and existing tests; none become automatic deletions. See the
authoring gate and
prevention examples.
.rejects.toThrow(...) never awaited, so it
silently never ran.The honest headline isn't a line count — it's that --fix only removes what it
can prove, so whatever it deletes could never have caught a regression. Clone
it and run it on your own repo.
npx skills add shmulc8/captain-obvious@captain-obvious
Also ships a .claude-plugin/plugin.json manifest, or copy
skills/captain-obvious/ into ~/.claude/skills/. Then ask your agent to
"clean up the useless tests" — or run the scanners directly:
# report-only (add --json out.json to save)
node skills/captain-obvious/scripts/captain_obvious_ts.mjs --project <repo>
python3 skills/captain-obvious/scripts/captain_obvious_py.py --path <repo> --mypy "uv run mypy"
# delete proven findings (needs a clean git tree — review the diff after)
node ... --fix
# confirm rotten-green asserts against real coverage (lcov / istanbul / coverage.py)
node ... --coverage coverage/lcov.info
python3 ... --coverage coverage.json
With --coverage, a conditional-assert whose line never executed is promoted
to proven rotten; one that did execute is a confirmed false positive and is
dropped — the dynamic half of the ICSE'19 analysis, using coverage your runner
already emits.
Trust boundary: the scan runs the project's own toolchain — mypy (with whatever plugins the repo's config loads),
uv/poetryenvironments, the repo's owntypescriptpackage. Treat scanning like running the repo's type checker: only do it on code you trust.
Requires Python ≥ 3.9 for the Python scanner and hook; the TS scanner uses the target repo's own typescript and requires ≥ 4.0.
Removing dead tests after the fact pays for them twice — once to write them,
once to clean them up. Installed as a plugin, captain-obvious also ships a
PreToolUse hook that catches them before they land: whenever the agent
writes or edits a test file, the pending content runs through a fast
syntactic-only single-file scan (--file --stdin; no mypy, no tsc, no
side effects), and the call is denied with a per-finding reason if it
would introduce proven can-never-fail patterns. The agent fixes the test on
the spot.
CAPTAIN_OBVIOUS_HOOK=off (default) | block | warn.npx skills add) don't get hooks — paste the
test-writing rules from
skills/captain-obvious/references/prevention.md
into your CLAUDE.md instead.--check --base <ref>)The write-time hook only fires inside an agent session. --check --base <ref>
is the CI twin of that hook — a correctness gate on new dead tests, not a
CI-time coverage optimizer. It is report-only (never writes; --check --fix is
rejected) and exits 1 only when a proven syntactic finding is newly
introduced vs <ref> in a changed test file — pre-existing findings never
gate, and it excludes type-guaranteed findings (the base-side scan is
syntactic-only). Any git trouble other than "file absent in base" (bad ref,
shallow clone, diff failure) fails open: a captain-obvious: note to stderr
and exit 0, so the tool can never invent a CI failure.
# .github/workflows/captain-obvious.yml
- uses: actions/checkout@v4
with: { fetch-depth: 0 } # need history to diff against the base
- uses: <owner>/captain-obvious@main
with:
base: ${{ github.event.pull_request.base.sha }} # default: origin/main
A .pre-commit-hooks.yaml (id: captain-obvious-check, --base HEAD) ships
for pre-commit users. Both wrappers run --no-types for speed and syntactic
parity with the comparison.
| Category | Example | Level |
|---|---|---|
type-guaranteed | expect(typeof f()).toBe('number') when f(): number; assert x is not None on non-Optional | proven |
constant-assert | expect(true).toBe(true), assert x == x (incl. self.assertEqual(1, 1)-style unittest assert-methods, report-only) | proven |
boundary-tautology | expect(arr.length).toBeGreaterThanOrEqual(0) | proven |
local-const-echo | const expected = 5; expect(expected).toBe(5) | proven |
mock-echo | stub returns 5 → assert it returns 5 | proven / advisory |
dead-assert / swallowed-assert / never-asserts | assertion after return, or inside try {} catch {} | proven |
missed-fail | a forced-fail marker (pytest.fail() / throw new Error) stuck in dead code — can never fire | proven |
missed-skip | a conditional early return/skip above the asserts — if it fires, they never run | advisory |
duplicate-test | identical body in the same suite | proven |
silent-smoke | assertion-free test whose every call is silently try/caught — or that contains no calls at all — it can do nothing and can never fail | proven |
no-assert | no assertion anywhere — a smoke test (legit by design, per ICSE '19), surfaced not deleted | advisory |
conditional-assert | assertion gated behind if — rotten green (ICSE '19) | advisory |
floating-async-assert | unawaited expect(p).rejects... (silent pass under Jest) | advisory |
smoke-only | expect(fn).not.toThrow() as the only check | advisory |
self-compare-call | expect(f(a)).toEqual(f(a)) | advisory |
broad-raises | pytest.raises(Exception) as the only check | advisory |
skipped-test | it.skip / xit / @pytest.mark.skip — never runs | advisory |
Full semantics, escape hatches, and false-positive guards — plus the academic
grounding (the rotten-green lineage: Delplanque ICSE '19 → RTj Java '19 →
Robinson ESEC/FSE '23 Google Test, which this tool extends to TS + Python) —
are in references/detectors.md.
result.length >= 0 is caught,
result.length < 1e9 is not) — only mutation testing proves those useless.Python
73.2%
JavaScript
26.8%