shmulc8/captain-obvious

Deterministic scanner that deletes tests that can never fail — asks the type checker instead of an LLM. TypeScript (Jest/Vitest/bun:test) + Python (pytest/mypy).

Python

13

81 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a tool that finds tests which can never fail (r/SideProject)

Captain Obvious checks TypeScript and Python tests for tautologies, assertions guaranteed by types, and tests with no real checks. The scanners are deterministic and make no model calls. It also includes a reusable agent skill. Free, MIT licensed:…

4

Sep 30, 2026

README

Captain Obvious 🫡

Captain Obvious Logo

AI agents write tests that can never fail. This skill deletes them.

captain-obvious scanning a repo: proven findings are auto-deletable, advisory findings are surfaced only

Every AI-generated test suite has them:

assert isinstance(add(1, 2), int)   # the function is typed `-> int`. mypy knew.
expect(true).toBe(true);            // bold claim

Tests like these can't catch a regression — they re-assert what the compiler already proved, or assert nothing at all. They burn CI minutes and inflate coverage into false confidence. Empirical studies find test smells in 38–100% of LLM-generated test suites.

🧠 The idea: ask the type checker, not another LLM

Detection is fully deterministic — no LLM in the loop, one script invocation per repo. Instead of asking an AI "is this test useful?", it asks the tools that already know:

  • TypeScript: the project's own compiler API answers "what is the static type of this expression?" — if typeof result === 'number' is a fact of the type, the assertion is not a test. Plain-JS test files (.test.js etc.) get the syntactic categories; type-guaranteed needs TypeScript (or stays off without checkJs).
  • Python: batched mypy reveal_type probes (one mypy run for the whole repo) answer the same for isinstance(...) and is not None.
  • AST passes catch the rest: tautologies, mock-echo tests, assertions after a return, assertions swallowed by empty catch/except: pass, duplicate bodies, len(x) >= 0.

✅ Proven vs ⚠️ advisory

Every finding is one of two tiers, and each tier is handled by whoever can actually decide it:

  • proven — cannot fail, by construction. Auto-deleted by --fix, no LLM involved. The scripts guard every known way types lie: any/unknown, as casts, non-null !, index signatures, unchecked index access, structural instanceof (only nominal classes are provable), cast()/type: ignore.
  • advisory — almost certainly useless but not provable (structural instanceof, mock-echo variants, pytest.raises(Exception), unawaited async assertions, rotten-green conditional asserts). The script writes down exactly why it's uncertain, and the coding agent running the skill adjudicates each one against the surrounding code — delete, keep, or rewrite — then proposes the calls for your approval before touching anything.

It also knows what not to flag: toBeDefined() on .find() results (T | undefined — a real check), enum contract locks, custom assertion helpers, deliberate determinism tests, and "must not raise" contract tests.

✍️ Before writing or removing a test

The skill now asks four questions the scanner cannot answer: What observable contract does the test protect? What regression would make it fail? Is that contract already covered at a stronger boundary? Does the test require a production seam used only by tests? For bug regressions, run the test against the pre-fix behavior when safe to confirm it fails for the intended reason.

This review catches leads such as copied implementation expectations, source greps, and mocks that supply the asserted outcome. They require reading the production path and existing tests; none become automatic deletions. See the authoring gate and prevention examples.

📊 Real-world results

  • Internal repo, ~1,800 pytest tests: 124 lines deleted across 21 files — 111 assertions re-checking what mypy already guaranteed, plus one all-constant test. Type-checker and full CI green before and after; change auto-approved.
  • TypeScript CLI repo, ~730 tests: zero deletions (it doesn't invent work) — but it caught a real bug: a .rejects.toThrow(...) never awaited, so it silently never ran.

The honest headline isn't a line count — it's that --fix only removes what it can prove, so whatever it deletes could never have caught a regression. Clone it and run it on your own repo.

🚀 Install & use

npx skills add shmulc8/captain-obvious@captain-obvious

Also ships a .claude-plugin/plugin.json manifest, or copy skills/captain-obvious/ into ~/.claude/skills/. Then ask your agent to "clean up the useless tests" — or run the scanners directly:

# report-only (add --json out.json to save)
node    skills/captain-obvious/scripts/captain_obvious_ts.mjs --project <repo>
python3 skills/captain-obvious/scripts/captain_obvious_py.py  --path <repo> --mypy "uv run mypy"

# delete proven findings (needs a clean git tree — review the diff after)
node ... --fix

# confirm rotten-green asserts against real coverage (lcov / istanbul / coverage.py)
node    ... --coverage coverage/lcov.info
python3 ... --coverage coverage.json

With --coverage, a conditional-assert whose line never executed is promoted to proven rotten; one that did execute is a confirmed false positive and is dropped — the dynamic half of the ICSE'19 analysis, using coverage your runner already emits.

Trust boundary: the scan runs the project's own toolchain — mypy (with whatever plugins the repo's config loads), uv/poetry environments, the repo's own typescript package. Treat scanning like running the repo's type checker: only do it on code you trust.

Requires Python ≥ 3.9 for the Python scanner and hook; the TS scanner uses the target repo's own typescript and requires ≥ 4.0.

🛡️ Prevention (write-time hook)

Removing dead tests after the fact pays for them twice — once to write them, once to clean them up. Installed as a plugin, captain-obvious also ships a PreToolUse hook that catches them before they land: whenever the agent writes or edits a test file, the pending content runs through a fast syntactic-only single-file scan (--file --stdin; no mypy, no tsc, no side effects), and the call is denied with a per-finding reason if it would introduce proven can-never-fail patterns. The agent fixes the test on the spot.

  • Proven, newly-introduced findings only. Advisories never block, and a pre-existing finding elsewhere in the file never blocks an unrelated edit. TDD red is safe: a deliberately failing test is not a can-never-fail test.
  • Fails open, always. Missing node, scanner crash, timeout, huge or syntactically-broken content — the write goes through.
  • Configurable: CAPTAIN_OBVIOUS_HOOK=off (default) | block | warn.
  • Skill-only installs (npx skills add) don't get hooks — paste the test-writing rules from skills/captain-obvious/references/prevention.md into your CLAUDE.md instead.

CI gate (--check --base <ref>)

The write-time hook only fires inside an agent session. --check --base <ref> is the CI twin of that hook — a correctness gate on new dead tests, not a CI-time coverage optimizer. It is report-only (never writes; --check --fix is rejected) and exits 1 only when a proven syntactic finding is newly introduced vs <ref> in a changed test file — pre-existing findings never gate, and it excludes type-guaranteed findings (the base-side scan is syntactic-only). Any git trouble other than "file absent in base" (bad ref, shallow clone, diff failure) fails open: a captain-obvious: note to stderr and exit 0, so the tool can never invent a CI failure.

# .github/workflows/captain-obvious.yml
- uses: actions/checkout@v4
  with: { fetch-depth: 0 }        # need history to diff against the base
- uses: <owner>/captain-obvious@main
  with:
    base: ${{ github.event.pull_request.base.sha }}   # default: origin/main

A .pre-commit-hooks.yaml (id: captain-obvious-check, --base HEAD) ships for pre-commit users. Both wrappers run --no-types for speed and syntactic parity with the comparison.

🔍 Detector catalog

CategoryExampleLevel
type-guaranteedexpect(typeof f()).toBe('number') when f(): number; assert x is not None on non-Optionalproven
constant-assertexpect(true).toBe(true), assert x == x (incl. self.assertEqual(1, 1)-style unittest assert-methods, report-only)proven
boundary-tautologyexpect(arr.length).toBeGreaterThanOrEqual(0)proven
local-const-echoconst expected = 5; expect(expected).toBe(5)proven
mock-echostub returns 5 → assert it returns 5proven / advisory
dead-assert / swallowed-assert / never-assertsassertion after return, or inside try {} catch {}proven
missed-faila forced-fail marker (pytest.fail() / throw new Error) stuck in dead code — can never fireproven
missed-skipa conditional early return/skip above the asserts — if it fires, they never runadvisory
duplicate-testidentical body in the same suiteproven
silent-smokeassertion-free test whose every call is silently try/caught — or that contains no calls at all — it can do nothing and can never failproven
no-assertno assertion anywhere — a smoke test (legit by design, per ICSE '19), surfaced not deletedadvisory
conditional-assertassertion gated behind if — rotten green (ICSE '19)advisory
floating-async-assertunawaited expect(p).rejects... (silent pass under Jest)advisory
smoke-onlyexpect(fn).not.toThrow() as the only checkadvisory
self-compare-callexpect(f(a)).toEqual(f(a))advisory
broad-raisespytest.raises(Exception) as the only checkadvisory
skipped-testit.skip / xit / @pytest.mark.skip — never runsadvisory

Full semantics, escape hatches, and false-positive guards — plus the academic grounding (the rotten-green lineage: Delplanque ICSE '19 → RTj Java '19 → Robinson ESEC/FSE '23 Google Test, which this tool extends to TS + Python) — are in references/detectors.md.

⚠️ Honest limitations

  • Cannot catch weak-but-executing assertions (result.length >= 0 is caught, result.length < 1e9 is not) — only mutation testing proves those useless.
  • Cross-file duplicates are advisory-only; coverage-subsumption is out of scope.
  • Dynamically-built assertions are invisible.
  • "Deleted nothing" does not mean "your suite is sound."

shmulc8/captain-obvious

Deterministic scanner that deletes tests that can never fail — asks the type checker instead of an LLM. TypeScript (Jest/Vitest/bun:test) + Python (pytest/mypy).

Python

13

81 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a tool that finds tests which can never fail (r/SideProject)

Captain Obvious checks TypeScript and Python tests for tautologies, assertions guaranteed by types, and tests with no real checks. The scanners are deterministic and make no model calls. It also includes a reusable agent skill. Free, MIT licensed:…

4

Sep 30, 2026

README

Captain Obvious 🫡

Captain Obvious Logo

AI agents write tests that can never fail. This skill deletes them.

captain-obvious scanning a repo: proven findings are auto-deletable, advisory findings are surfaced only

Every AI-generated test suite has them:

assert isinstance(add(1, 2), int)   # the function is typed `-> int`. mypy knew.
expect(true).toBe(true);            // bold claim

Tests like these can't catch a regression — they re-assert what the compiler already proved, or assert nothing at all. They burn CI minutes and inflate coverage into false confidence. Empirical studies find test smells in 38–100% of LLM-generated test suites.

🧠 The idea: ask the type checker, not another LLM

Detection is fully deterministic — no LLM in the loop, one script invocation per repo. Instead of asking an AI "is this test useful?", it asks the tools that already know:

  • TypeScript: the project's own compiler API answers "what is the static type of this expression?" — if typeof result === 'number' is a fact of the type, the assertion is not a test. Plain-JS test files (.test.js etc.) get the syntactic categories; type-guaranteed needs TypeScript (or stays off without checkJs).
  • Python: batched mypy reveal_type probes (one mypy run for the whole repo) answer the same for isinstance(...) and is not None.
  • AST passes catch the rest: tautologies, mock-echo tests, assertions after a return, assertions swallowed by empty catch/except: pass, duplicate bodies, len(x) >= 0.

✅ Proven vs ⚠️ advisory

Every finding is one of two tiers, and each tier is handled by whoever can actually decide it:

  • proven — cannot fail, by construction. Auto-deleted by --fix, no LLM involved. The scripts guard every known way types lie: any/unknown, as casts, non-null !, index signatures, unchecked index access, structural instanceof (only nominal classes are provable), cast()/type: ignore.
  • advisory — almost certainly useless but not provable (structural instanceof, mock-echo variants, pytest.raises(Exception), unawaited async assertions, rotten-green conditional asserts). The script writes down exactly why it's uncertain, and the coding agent running the skill adjudicates each one against the surrounding code — delete, keep, or rewrite — then proposes the calls for your approval before touching anything.

It also knows what not to flag: toBeDefined() on .find() results (T | undefined — a real check), enum contract locks, custom assertion helpers, deliberate determinism tests, and "must not raise" contract tests.

✍️ Before writing or removing a test

The skill now asks four questions the scanner cannot answer: What observable contract does the test protect? What regression would make it fail? Is that contract already covered at a stronger boundary? Does the test require a production seam used only by tests? For bug regressions, run the test against the pre-fix behavior when safe to confirm it fails for the intended reason.

This review catches leads such as copied implementation expectations, source greps, and mocks that supply the asserted outcome. They require reading the production path and existing tests; none become automatic deletions. See the authoring gate and prevention examples.

📊 Real-world results

  • Internal repo, ~1,800 pytest tests: 124 lines deleted across 21 files — 111 assertions re-checking what mypy already guaranteed, plus one all-constant test. Type-checker and full CI green before and after; change auto-approved.
  • TypeScript CLI repo, ~730 tests: zero deletions (it doesn't invent work) — but it caught a real bug: a .rejects.toThrow(...) never awaited, so it silently never ran.

The honest headline isn't a line count — it's that --fix only removes what it can prove, so whatever it deletes could never have caught a regression. Clone it and run it on your own repo.

🚀 Install & use

npx skills add shmulc8/captain-obvious@captain-obvious

Also ships a .claude-plugin/plugin.json manifest, or copy skills/captain-obvious/ into ~/.claude/skills/. Then ask your agent to "clean up the useless tests" — or run the scanners directly:

# report-only (add --json out.json to save)
node    skills/captain-obvious/scripts/captain_obvious_ts.mjs --project <repo>
python3 skills/captain-obvious/scripts/captain_obvious_py.py  --path <repo> --mypy "uv run mypy"

# delete proven findings (needs a clean git tree — review the diff after)
node ... --fix

# confirm rotten-green asserts against real coverage (lcov / istanbul / coverage.py)
node    ... --coverage coverage/lcov.info
python3 ... --coverage coverage.json

With --coverage, a conditional-assert whose line never executed is promoted to proven rotten; one that did execute is a confirmed false positive and is dropped — the dynamic half of the ICSE'19 analysis, using coverage your runner already emits.

Trust boundary: the scan runs the project's own toolchain — mypy (with whatever plugins the repo's config loads), uv/poetry environments, the repo's own typescript package. Treat scanning like running the repo's type checker: only do it on code you trust.

Requires Python ≥ 3.9 for the Python scanner and hook; the TS scanner uses the target repo's own typescript and requires ≥ 4.0.

🛡️ Prevention (write-time hook)

Removing dead tests after the fact pays for them twice — once to write them, once to clean them up. Installed as a plugin, captain-obvious also ships a PreToolUse hook that catches them before they land: whenever the agent writes or edits a test file, the pending content runs through a fast syntactic-only single-file scan (--file --stdin; no mypy, no tsc, no side effects), and the call is denied with a per-finding reason if it would introduce proven can-never-fail patterns. The agent fixes the test on the spot.

  • Proven, newly-introduced findings only. Advisories never block, and a pre-existing finding elsewhere in the file never blocks an unrelated edit. TDD red is safe: a deliberately failing test is not a can-never-fail test.
  • Fails open, always. Missing node, scanner crash, timeout, huge or syntactically-broken content — the write goes through.
  • Configurable: CAPTAIN_OBVIOUS_HOOK=off (default) | block | warn.
  • Skill-only installs (npx skills add) don't get hooks — paste the test-writing rules from skills/captain-obvious/references/prevention.md into your CLAUDE.md instead.

CI gate (--check --base <ref>)

The write-time hook only fires inside an agent session. --check --base <ref> is the CI twin of that hook — a correctness gate on new dead tests, not a CI-time coverage optimizer. It is report-only (never writes; --check --fix is rejected) and exits 1 only when a proven syntactic finding is newly introduced vs <ref> in a changed test file — pre-existing findings never gate, and it excludes type-guaranteed findings (the base-side scan is syntactic-only). Any git trouble other than "file absent in base" (bad ref, shallow clone, diff failure) fails open: a captain-obvious: note to stderr and exit 0, so the tool can never invent a CI failure.

# .github/workflows/captain-obvious.yml
- uses: actions/checkout@v4
  with: { fetch-depth: 0 }        # need history to diff against the base
- uses: <owner>/captain-obvious@main
  with:
    base: ${{ github.event.pull_request.base.sha }}   # default: origin/main

A .pre-commit-hooks.yaml (id: captain-obvious-check, --base HEAD) ships for pre-commit users. Both wrappers run --no-types for speed and syntactic parity with the comparison.

🔍 Detector catalog

CategoryExampleLevel
type-guaranteedexpect(typeof f()).toBe('number') when f(): number; assert x is not None on non-Optionalproven
constant-assertexpect(true).toBe(true), assert x == x (incl. self.assertEqual(1, 1)-style unittest assert-methods, report-only)proven
boundary-tautologyexpect(arr.length).toBeGreaterThanOrEqual(0)proven
local-const-echoconst expected = 5; expect(expected).toBe(5)proven
mock-echostub returns 5 → assert it returns 5proven / advisory
dead-assert / swallowed-assert / never-assertsassertion after return, or inside try {} catch {}proven
missed-faila forced-fail marker (pytest.fail() / throw new Error) stuck in dead code — can never fireproven
missed-skipa conditional early return/skip above the asserts — if it fires, they never runadvisory
duplicate-testidentical body in the same suiteproven
silent-smokeassertion-free test whose every call is silently try/caught — or that contains no calls at all — it can do nothing and can never failproven
no-assertno assertion anywhere — a smoke test (legit by design, per ICSE '19), surfaced not deletedadvisory
conditional-assertassertion gated behind if — rotten green (ICSE '19)advisory
floating-async-assertunawaited expect(p).rejects... (silent pass under Jest)advisory
smoke-onlyexpect(fn).not.toThrow() as the only checkadvisory
self-compare-callexpect(f(a)).toEqual(f(a))advisory
broad-raisespytest.raises(Exception) as the only checkadvisory
skipped-testit.skip / xit / @pytest.mark.skip — never runsadvisory

Full semantics, escape hatches, and false-positive guards — plus the academic grounding (the rotten-green lineage: Delplanque ICSE '19 → RTj Java '19 → Robinson ESEC/FSE '23 Google Test, which this tool extends to TS + Python) — are in references/detectors.md.

⚠️ Honest limitations

  • Cannot catch weak-but-executing assertions (result.length >= 0 is caught, result.length < 1e9 is not) — only mutation testing proves those useless.
  • Cross-file duplicates are advisory-only; coverage-subsumption is out of scope.
  • Dynamically-built assertions are invisible.
  • "Deleted nothing" does not mean "your suite is sound."

Languages

Python

73.2%

JavaScript

26.8%