Claude Code skill + GitHub Action that proves your tests would have caught the bug: every changed test runs with the fix and without it. Works in Codex, Cursor and other agents too.
TypeScript
1
11 commits
updated Sep 26, 2026
A green CI says your tests pass. Receipts says whether they would have caught the bug.
Install · How it works · Verdicts · Study · По-русски
Coding agents now write most of the tests in their pull requests, and CI turns green. But a test that passes with the fix may pass without it too: then it checks nothing about the bug. Receipts runs every test a change adds or edits twice, once with the change and once with the code reverted to the base branch, and tells you which tests actually prove the change.
it/test blocks (vitest, jest, including describe nesting, parametrize rows and .each tables) are matched against the changed lines.| Verdict | With the change | Without it | What it means | |
|---|---|---|---|---|
| ✅ | PROVEN | passes | fails | The test catches what the change fixes |
| 🛡️ | GUARD | passes | passes | Guards neighboring behavior while another test proves the change. Healthy |
| ⚠️ | THEATER | passes | passes | No test proves the change: they would all pass anyway |
| 🟡 | WEAK | passes | fails on import | The code it calls did not exist yet. Normal for a feature, thin for a fix |
| ❌ | BROKEN | fails | The change does not pass its own test | |
| 🔁 | FLAKY | mixed | mixed | Different results on the same code |
| ⏭️ | SKIPPED | skipped | The test did not run |
For a refactor the expectation flips: every test must pass on both sides (✅ PRESERVED), and a test that fails on the old code means the refactor changed behavior (⚠️ CHANGED). Receipts reads the kind of change from PR labels, a fix: / feat: / refactor: title or commit messages. It also lists tests a change removed, and source changes that came with no test at all.
/plugin marketplace add syntaxixr/receipts
/plugin install receipts-check@receipts
The plugin brings the prove-fix skill: after Claude fixes a bug it runs Receipts on its own tests, and when a test is THEATER it rewrites the test, not the fix, until the change is proven. You can also call it yourself with /receipts-check:prove-fix.
It also has a Stop hook, off by default, that won't let Claude end a turn while its changed tests prove nothing. Turn it on in a repository you trust with git config receipts.hook true. More in docs/agents.md.
A re-enactment of a real session: the verdicts are real output from Receipts on this change.
The same skill works in Claude Code, Codex, Cursor, Gemini CLI, Copilot and other agents that read SKILL.md. The CLI is bundled inside it.
npx skills add syntaxixr/receipts
Claude apps (claude.ai, Claude Desktop): download prove-fix.zip from the release page and upload it in Settings → Capabilities → Skills. It runs where code execution can see your repository.
Prefer plain instructions? docs/agents.md has an AGENTS.md block that does the same with npx.
# .github/workflows/receipts.yml
name: Receipts
on: pull_request # never pull_request_target
permissions:
contents: read
pull-requests: write # for the report comment
jobs:
receipts:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
persist-credentials: false # keep the token away from the tests
# Set the project up exactly as for your normal test job, e.g.:
- uses: actions/setup-python@v7
with: { python-version: "3.12" }
- run: pip install -e . pytest
- uses: syntaxixr/receipts@main # or pin a commit SHA
The action posts one report on the PR and keeps it updated, writes the job summary, and fails the check on broken,theater,changed by default.
| Input | Default | |
|---|---|---|
fail-on | broken,theater,changed | Verdicts that fail the check. Also: weak, flaky, skipped, no-tests |
mode | auto | fix, feat or refactor |
base | PR base branch | Ref to compare against |
reruns | 2 | Extra runs for failing tests, to tell flakes from failures |
timeout | 300 | Seconds per test run; a test still running counts as failed (on the old code: the hang the change fixed) |
js-runner | from package.json | vitest or jest |
python | python | Interpreter for pytest |
working-directory | . | Where to run |
comment | true | Post the PR comment |
github-token | github.token | Token for the comment; only the comment step sees it |
comment-author | github-actions[bot] | Login that owns the report comment; change it only with a custom token |
Outputs: passed ("true" or "false") and receipt (path to the JSON receipt).
cd your-project
npx github:syntaxixr/receipts check
It checks committed and uncommitted work against the merge base with origin/HEAD, main or master, and puts every file back when it is done. Nothing to install: the build is committed and has no dependencies. Without npx, clone this repository and run node plugin/skills/prove-fix/scripts/cli.js check from your project.
receipts check [--base <ref>] [--mode fix|feat|refactor] [--fail-on <verdicts>]
[--reruns <n>] [--timeout <sec>] [--python <path>] [--js-runner vitest|jest]
[--committed] [--json <file>] [--markdown <file>]
receipts restore
Exit codes: 0 pass, 1 a --fail-on verdict was found, 2 Receipts itself failed.
We ran Receipts over 181 real changes: 81 maintainer fix commits in twelve libraries (click, sqlparse, marshmallow, dateutil, dayjs…) and 100 pull requests with coding-agent fingerprints in five agent-heavy repositories (Claude Agent SDK, OpenAI Agents SDK, MCP Python SDK, fastmcp, simonw/llm).
| Proven | Mixed | Unproven | Weak only | |
|---|---|---|---|---|
| Maintainer fix commits | 90% | 3% | 7% | 0 |
| Agent pull requests | 82% | 5% | 2% | 10% |
Most tests do their job. The gap that shows up only in agent work is WEAK: in 10% of agent PRs every test failed on the old code only because it imported a name the change added, so nothing ever ran against the old behavior. Method, caveats and every change with a link: docs/study.md. The data is on Hugging Face: syntaxixr/receipts-study.
pip install -e .) keep importing from the original checkout. Every original byte is saved under .git/receipts-backup first and restored afterwards, on Ctrl-C too. After a hard kill, receipts restore puts everything back.pytest, vitest run or jest directly, not your npm test. Set what the tests depend on (a TZ, service URLs, feature flags) on the CI step or in your shell.--timeout counts as failed, so a fix for a hang is PROVEN.Limits in this version: one JS runner per repository, detected from the root package.json (use working-directory for a sub-package); pnpm catalog: versions are not supported yet; tests are located by pattern, not by a full parser. CI runs on Linux and Windows.
Receipts runs a change's tests, so it runs the change's code, exactly as your CI's test job does. Run it where you would run those tests anyway: in Actions with pull_request (never pull_request_target), with persist-credentials: false.
What Receipts itself guarantees:
src/ into a link to ~/.ssh is refused), no file/directory swaps.Details and the threat model: SECURITY.md. Report a vulnerability privately via Security → Report a vulnerability.
npm ci
pip install pytest pytest-xdist
npm test # builds, then runs the suite against real pytest, vitest and jest projects
node docs/media/make-media.mjs # redraws the README images from real runs (needs ffmpeg and Edge or Chrome)
The build in plugin/skills/prove-fix/scripts/ is committed, because the plugin, the skill, the CLI and the Action all run it as is. Rebuild and commit it with every change to src/; CI fails when they drift.
MIT
TypeScript
52.4%
JavaScript
44.1%
Python
3.5%
Claude Code skill + GitHub Action that proves your tests would have caught the bug: every changed test runs with the fix and without it. Works in Codex, Cursor and other agents too.
TypeScript
1
11 commits
updated Sep 26, 2026
A green CI says your tests pass. Receipts says whether they would have caught the bug.
Install · How it works · Verdicts · Study · По-русски
Coding agents now write most of the tests in their pull requests, and CI turns green. But a test that passes with the fix may pass without it too: then it checks nothing about the bug. Receipts runs every test a change adds or edits twice, once with the change and once with the code reverted to the base branch, and tells you which tests actually prove the change.
it/test blocks (vitest, jest, including describe nesting, parametrize rows and .each tables) are matched against the changed lines.| Verdict | With the change | Without it | What it means | |
|---|---|---|---|---|
| ✅ | PROVEN | passes | fails | The test catches what the change fixes |
| 🛡️ | GUARD | passes | passes | Guards neighboring behavior while another test proves the change. Healthy |
| ⚠️ | THEATER | passes | passes | No test proves the change: they would all pass anyway |
| 🟡 | WEAK | passes | fails on import | The code it calls did not exist yet. Normal for a feature, thin for a fix |
| ❌ | BROKEN | fails | The change does not pass its own test | |
| 🔁 | FLAKY | mixed | mixed | Different results on the same code |
| ⏭️ | SKIPPED | skipped | The test did not run |
For a refactor the expectation flips: every test must pass on both sides (✅ PRESERVED), and a test that fails on the old code means the refactor changed behavior (⚠️ CHANGED). Receipts reads the kind of change from PR labels, a fix: / feat: / refactor: title or commit messages. It also lists tests a change removed, and source changes that came with no test at all.
/plugin marketplace add syntaxixr/receipts
/plugin install receipts-check@receipts
The plugin brings the prove-fix skill: after Claude fixes a bug it runs Receipts on its own tests, and when a test is THEATER it rewrites the test, not the fix, until the change is proven. You can also call it yourself with /receipts-check:prove-fix.
It also has a Stop hook, off by default, that won't let Claude end a turn while its changed tests prove nothing. Turn it on in a repository you trust with git config receipts.hook true. More in docs/agents.md.
A re-enactment of a real session: the verdicts are real output from Receipts on this change.
The same skill works in Claude Code, Codex, Cursor, Gemini CLI, Copilot and other agents that read SKILL.md. The CLI is bundled inside it.
npx skills add syntaxixr/receipts
Claude apps (claude.ai, Claude Desktop): download prove-fix.zip from the release page and upload it in Settings → Capabilities → Skills. It runs where code execution can see your repository.
Prefer plain instructions? docs/agents.md has an AGENTS.md block that does the same with npx.
# .github/workflows/receipts.yml
name: Receipts
on: pull_request # never pull_request_target
permissions:
contents: read
pull-requests: write # for the report comment
jobs:
receipts:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
with:
fetch-depth: 0
persist-credentials: false # keep the token away from the tests
# Set the project up exactly as for your normal test job, e.g.:
- uses: actions/setup-python@v7
with: { python-version: "3.12" }
- run: pip install -e . pytest
- uses: syntaxixr/receipts@main # or pin a commit SHA
The action posts one report on the PR and keeps it updated, writes the job summary, and fails the check on broken,theater,changed by default.
| Input | Default | |
|---|---|---|
fail-on | broken,theater,changed | Verdicts that fail the check. Also: weak, flaky, skipped, no-tests |
mode | auto | fix, feat or refactor |
base | PR base branch | Ref to compare against |
reruns | 2 | Extra runs for failing tests, to tell flakes from failures |
timeout | 300 | Seconds per test run; a test still running counts as failed (on the old code: the hang the change fixed) |
js-runner | from package.json | vitest or jest |
python | python | Interpreter for pytest |
working-directory | . | Where to run |
comment | true | Post the PR comment |
github-token | github.token | Token for the comment; only the comment step sees it |
comment-author | github-actions[bot] | Login that owns the report comment; change it only with a custom token |
Outputs: passed ("true" or "false") and receipt (path to the JSON receipt).
cd your-project
npx github:syntaxixr/receipts check
It checks committed and uncommitted work against the merge base with origin/HEAD, main or master, and puts every file back when it is done. Nothing to install: the build is committed and has no dependencies. Without npx, clone this repository and run node plugin/skills/prove-fix/scripts/cli.js check from your project.
receipts check [--base <ref>] [--mode fix|feat|refactor] [--fail-on <verdicts>]
[--reruns <n>] [--timeout <sec>] [--python <path>] [--js-runner vitest|jest]
[--committed] [--json <file>] [--markdown <file>]
receipts restore
Exit codes: 0 pass, 1 a --fail-on verdict was found, 2 Receipts itself failed.
We ran Receipts over 181 real changes: 81 maintainer fix commits in twelve libraries (click, sqlparse, marshmallow, dateutil, dayjs…) and 100 pull requests with coding-agent fingerprints in five agent-heavy repositories (Claude Agent SDK, OpenAI Agents SDK, MCP Python SDK, fastmcp, simonw/llm).
| Proven | Mixed | Unproven | Weak only | |
|---|---|---|---|---|
| Maintainer fix commits | 90% | 3% | 7% | 0 |
| Agent pull requests | 82% | 5% | 2% | 10% |
Most tests do their job. The gap that shows up only in agent work is WEAK: in 10% of agent PRs every test failed on the old code only because it imported a name the change added, so nothing ever ran against the old behavior. Method, caveats and every change with a link: docs/study.md. The data is on Hugging Face: syntaxixr/receipts-study.
pip install -e .) keep importing from the original checkout. Every original byte is saved under .git/receipts-backup first and restored afterwards, on Ctrl-C too. After a hard kill, receipts restore puts everything back.pytest, vitest run or jest directly, not your npm test. Set what the tests depend on (a TZ, service URLs, feature flags) on the CI step or in your shell.--timeout counts as failed, so a fix for a hang is PROVEN.Limits in this version: one JS runner per repository, detected from the root package.json (use working-directory for a sub-package); pnpm catalog: versions are not supported yet; tests are located by pattern, not by a full parser. CI runs on Linux and Windows.
Receipts runs a change's tests, so it runs the change's code, exactly as your CI's test job does. Run it where you would run those tests anyway: in Actions with pull_request (never pull_request_target), with persist-credentials: false.
What Receipts itself guarantees:
src/ into a link to ~/.ssh is refused), no file/directory swaps.Details and the threat model: SECURITY.md. Report a vulnerability privately via Security → Report a vulnerability.
npm ci
pip install pytest pytest-xdist
npm test # builds, then runs the suite against real pytest, vitest and jest projects
node docs/media/make-media.mjs # redraws the README images from real runs (needs ffmpeg and Edge or Chrome)
The build in plugin/skills/prove-fix/scripts/ is committed, because the plugin, the skill, the CLI and the Action all run it as is. Rebuild and commit it with every change to src/; CI fails when they drift.
MIT
TypeScript
52.4%
JavaScript
44.1%
Python
3.5%