heyosseus/sloppy

Static analysis for the debt AI agents leave in PHP: 26 rules, Claude Code hooks, git-diff review, a Rector and Pint fix pass, Pest expectations, CI annotations and an MCP server. Deterministic, local, no LLM.

PHP

146

36 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Are AI coding assistants creating a new category of technical debt? I built a tool to measure some of it (r/SideProject)

AI agents write PHP fast, and they make the same mistakes fast. I've been seeing patterns like: * `catch (Throwable)` blocks that silently hide failures * 200-line controller actions * queries inside loops * `// ... existing code ...` left where a method body should be * tests "fixed" with…

1

Oct 4, 2026

README

Sloppy

Your AI agent writes the PHP. Sloppy makes it clean up after itself.
It catches the code people regret — god methods, swallowed exceptions, N+1 queries — inside Claude Code, your CI and your Pest suite.
Deterministic and local: no model, no API key, no network.

tests packagist downloads php license ko-fi

What Claude Code runs after the agent edits a PHP file

Real output. Claude Code runs this after every edit and hands it to the model, which fixes the catch block before you ever see it.

Quick start

composer require --dev heyosseus/sloppy

vendor/bin/sloppy                   # scan the project (php artisan sloppy in Laravel)
vendor/bin/sloppy agents install    # make Claude Code check its own work

That's it: no configuration file, no account. Sloppy finds your source roots from composer.json. It runs on any PHP 8.3+ project, and adds ten Eloquent-aware rules when it finds Laravel 12 or 13.

A first scan of a codebase with history does not hand you hundreds of findings to work through. It lists the defects worth fixing first, sums up the rest per file, and offers to baseline what is already there.

Trying it before adding a dependency? Run composer global require heyosseus/sloppy, or download the phar. See Getting started.

How the agent loop works

Agents produce code fast, and they produce the same mistakes fast: the catch (Throwable) that hides a failure, the 200-line controller action, the query inside a loop. Code review catches them late. Sloppy catches them while the agent still has the code in hand.

  1. The rules go in first. CLAUDE.md (or AGENTS.md, .cursorrules, Copilot, Windsurf, Laravel Boost) gets this project's rules, each written as an instruction the agent can follow.
  2. Every edit is checked. After each change to a PHP file, Claude Code runs Sloppy and hands the model anything that edit introduced, then the model fixes it.
  3. It can't finish dirty. When the agent tries to call the task done, new findings at or above your threshold send it back to fix them.

It's built never to get in the way. Findings a file already had are never reported, so the agent doesn't wander off "fixing" code nobody asked it to touch. The finish check blocks once, so a false positive costs one round trip, not the session. And if Sloppy can't run (no git, a broken config), the agent carries on.

There's also an MCP server for agents that should scan on demand. See Coding agents.

What it catches

26 rules, plus two checks that compare your change with its base. Each one is tested to fire on the pattern and to stay silent on ordinary Laravel code.

Rules
ComplexitySL101 God Method · SL102 God Class · SL103 Excessive Nesting
DuplicationSL104 Duplicate Logic · SL111 Copy-Paste Drift, a near-copy whose one difference looks like an unfinished edit
Dead code & dependenciesSL105 Dead Private Method · SL106 Unused Constructor Dependency · SL112 Placeholder Implementation, the // ... existing code ... or "not implemented" left where a body should be
Error handlingSL107 Swallowed Exception
ReadabilitySL108 Redundant Condition · SL109 Narrative Comment · SL110 Defensive Programming Noise
LaravelSL201 Business Logic In Controller · SL202 Inline Validation · SL208 Direct External API Call · SL209 Model Doing Too Much
PerformanceSL203 Possible N+1 · SL204 Query Inside Loop · SL205 Collection Instead Of Query · SL210 Suspicious Model::all()
DependenciesSL206 Excessive Controller Dependencies · SL207 Excessive Service Dependencies
Architecture (advisory)SL301 Abstraction Inflation · SL302 Empty Wrapper Class · SL303 Single-Use Abstraction
SuppressionSL501 Unexplained Suppression · SL502 Baseline Growth, new entries in your PHPStan or Psalm baseline · SL503 Weakened Test, a test skipped, stripped of assertions or given assertTrue(true) to make it pass

Every finding says where it is, what was measured, how sure the analyser is, why the pattern costs you, and what to do instead. See Rules, or write your own.

Beyond the agent

The same analyser, wherever else you want the answer.

Start with what matters. A first scan of a medium-sized application finds hundreds of things, and nobody reads hundreds of things. So sloppy lists the defects first (swallowed exceptions, queries in loops, unfinished bodies), highest risk first. It sums up the long methods and narrating comments per file, putting the files that keep changing at the top. Then it tells you which commands shorten the list without reading it: sloppy fix for the mechanical findings, sloppy baseline for the debt that is already there. --all lists everything.

Review a change. sloppy diff main reports only what your branch introduced, never what it inherited. sloppy review main ranks the same findings by risk, so you know which file to read first and where to stop.

php artisan sloppy:diff main

Adopt it on a codebase with history. sloppy baseline accepts today's debt, so only new findings fail the build. Entries are keyed on class and member, not line numbers, so the baseline survives ordinary editing.

Gate it in CI. sloppy ci reads the pipeline it runs in: it annotates the diff on a GitHub pull request and fills the Code Quality widget on GitLab. The GitHub Action is three lines:

- uses: heyosseus/sloppy-action@v1
  with:
    diff-branch: main

Already on PHPStan? Get the same findings inside the run you already have:

composer require --dev heyosseus/phpstan-sloppy

Each one is a PHPStan error with its own identifier (sloppy.SL107), so @phpstan-ignore, ignoreErrors and PHPStan baselines work on it. It reports exactly what sloppy ci would fail on. See heyosseus/phpstan-sloppy.

Fail your tests on new debt. A Pest plugin ships with the package:

it('has no new slop on this branch', function (): void {
    expectCleanSloppyDiff('main');
});

Fix what a machine can fix. sloppy fix hands the mechanical findings to Rector, scoped to the files that have them, deletes the comments that only restate their code, then formats with Pint. It tells you which findings still need a person, and why no tool should touch them.

Keep score while you work. sloppy watch keeps the score and what to read first on screen, redrawing on every save. It's made to sit beside an agent that is writing code.

See it where you already look. Use SARIF for GitHub code scanning and your IDE, --format=github for inline annotations, and Markdown for a PR comment. There's also a Filament widget, a NativePHP menu-bar label, and JSON output with a published schema.

See Everyday workflow and CI and code scanning.

Why trust the numbers

It is not an AI detector. Nobody can reliably tell from source code who or what wrote it, and Sloppy never tries. It detects patterns that turn into maintenance cost, whoever wrote them. A 200-line controller action that swallows a Throwable is a problem whether a person or an agent wrote it at 3am.

Sloppy saysIt meansIt does not mean
Confidence: 88%How sure the analyser is that the pattern is really thereAny probability that the code was AI-generated
Score: 67/100A code-quality risk measure for the analysed paths"67% of this code is AI-generated"

Every number shows its working. Add --explain-risk and each score and risk prints the arithmetic that produced it. The same code always produces the same report, which is what makes Sloppy usable as a gate. See Score, severity and risk.

It is tuned for silence. Rules need several signals, not one; they understand framework conventions; and they back off wherever code could be reached indirectly. Sloppy runs over its own source in CI, and a test asserts that ordinary Laravel code produces zero findings. The rules and the score are checked against eight open-source Laravel applications, 782,578 lines, where the four healthiest score 87 to 90.

It complements your tools rather than replacing them:

ToolAnswers
PHPStan / PsalmIs this type-correct?
Pint / PHP_CodeSnifferIs this formatted consistently?
Pest / PHPUnitDoes this behave correctly?
RectorCan this be transformed mechanically?
SloppyIs this shaped like code somebody will regret?

If PHPStan can prove it, Sloppy stays out of it.

Documentation

Getting startedInstall options, every command, your first scan and how a long one is triaged, adopting on an existing codebase
Coding agentsClaude Code hooks, rulesets for every agent, Laravel Boost, the MCP server
RulesEvery rule in detail, and how false positives are kept down
Everyday workflowDiff and review, sloppy fix over Rector, Pint and narrating comments, Pest expectations, watch, Filament and NativePHP
CI and code scanningsloppy ci, the GitHub Action, GitLab, SARIF, annotations, exit codes
Score, severity and riskHow every number is calculated, and what coverage, git history and PHPStan baselines feed
Configurationconfig/sloppy.php, and taming a noisy first run
JSON outputThe machine-readable report and its contract
Custom rulesWriting your own rule, and where it shows up

Roadmap

Still ahead: inline pull-request review comments, sloppy explain for a longer write-up of one finding, HTML reports, project architecture policies, and more rules for the shortcuts agents take, such as configuration keys and routes that do not exist. Anything AI-assisted will be opt-in and separate: the analyser will always work with no API key, no network and no model.

Contributing

Found a false positive? That's a bug worth reporting, not a threshold to work around. Open an issue. To work on Sloppy itself, see CONTRIBUTING.md. composer test runs Rector, Pint, PHPStan at level 8, 100% type coverage and the suite.

Security issues: see SECURITY.md.

License

MIT. See LICENSE.md.

ai
code-quality
filament
laravel
linter
mcp
pest
php
pint
rectorphp
static-analysis
technical-debt

heyosseus/sloppy

Static analysis for the debt AI agents leave in PHP: 26 rules, Claude Code hooks, git-diff review, a Rector and Pint fix pass, Pest expectations, CI annotations and an MCP server. Deterministic, local, no LLM.

PHP

146

36 commits

updated Oct 4, 2026

See the code

See what people are saying

SourceMessageScoreDate

Are AI coding assistants creating a new category of technical debt? I built a tool to measure some of it (r/SideProject)

AI agents write PHP fast, and they make the same mistakes fast. I've been seeing patterns like: * `catch (Throwable)` blocks that silently hide failures * 200-line controller actions * queries inside loops * `// ... existing code ...` left where a method body should be * tests "fixed" with…

1

Oct 4, 2026

README

Sloppy

Your AI agent writes the PHP. Sloppy makes it clean up after itself.
It catches the code people regret — god methods, swallowed exceptions, N+1 queries — inside Claude Code, your CI and your Pest suite.
Deterministic and local: no model, no API key, no network.

tests packagist downloads php license ko-fi

What Claude Code runs after the agent edits a PHP file

Real output. Claude Code runs this after every edit and hands it to the model, which fixes the catch block before you ever see it.

Quick start

composer require --dev heyosseus/sloppy

vendor/bin/sloppy                   # scan the project (php artisan sloppy in Laravel)
vendor/bin/sloppy agents install    # make Claude Code check its own work

That's it: no configuration file, no account. Sloppy finds your source roots from composer.json. It runs on any PHP 8.3+ project, and adds ten Eloquent-aware rules when it finds Laravel 12 or 13.

A first scan of a codebase with history does not hand you hundreds of findings to work through. It lists the defects worth fixing first, sums up the rest per file, and offers to baseline what is already there.

Trying it before adding a dependency? Run composer global require heyosseus/sloppy, or download the phar. See Getting started.

How the agent loop works

Agents produce code fast, and they produce the same mistakes fast: the catch (Throwable) that hides a failure, the 200-line controller action, the query inside a loop. Code review catches them late. Sloppy catches them while the agent still has the code in hand.

  1. The rules go in first. CLAUDE.md (or AGENTS.md, .cursorrules, Copilot, Windsurf, Laravel Boost) gets this project's rules, each written as an instruction the agent can follow.
  2. Every edit is checked. After each change to a PHP file, Claude Code runs Sloppy and hands the model anything that edit introduced, then the model fixes it.
  3. It can't finish dirty. When the agent tries to call the task done, new findings at or above your threshold send it back to fix them.

It's built never to get in the way. Findings a file already had are never reported, so the agent doesn't wander off "fixing" code nobody asked it to touch. The finish check blocks once, so a false positive costs one round trip, not the session. And if Sloppy can't run (no git, a broken config), the agent carries on.

There's also an MCP server for agents that should scan on demand. See Coding agents.

What it catches

26 rules, plus two checks that compare your change with its base. Each one is tested to fire on the pattern and to stay silent on ordinary Laravel code.

Rules
ComplexitySL101 God Method · SL102 God Class · SL103 Excessive Nesting
DuplicationSL104 Duplicate Logic · SL111 Copy-Paste Drift, a near-copy whose one difference looks like an unfinished edit
Dead code & dependenciesSL105 Dead Private Method · SL106 Unused Constructor Dependency · SL112 Placeholder Implementation, the // ... existing code ... or "not implemented" left where a body should be
Error handlingSL107 Swallowed Exception
ReadabilitySL108 Redundant Condition · SL109 Narrative Comment · SL110 Defensive Programming Noise
LaravelSL201 Business Logic In Controller · SL202 Inline Validation · SL208 Direct External API Call · SL209 Model Doing Too Much
PerformanceSL203 Possible N+1 · SL204 Query Inside Loop · SL205 Collection Instead Of Query · SL210 Suspicious Model::all()
DependenciesSL206 Excessive Controller Dependencies · SL207 Excessive Service Dependencies
Architecture (advisory)SL301 Abstraction Inflation · SL302 Empty Wrapper Class · SL303 Single-Use Abstraction
SuppressionSL501 Unexplained Suppression · SL502 Baseline Growth, new entries in your PHPStan or Psalm baseline · SL503 Weakened Test, a test skipped, stripped of assertions or given assertTrue(true) to make it pass

Every finding says where it is, what was measured, how sure the analyser is, why the pattern costs you, and what to do instead. See Rules, or write your own.

Beyond the agent

The same analyser, wherever else you want the answer.

Start with what matters. A first scan of a medium-sized application finds hundreds of things, and nobody reads hundreds of things. So sloppy lists the defects first (swallowed exceptions, queries in loops, unfinished bodies), highest risk first. It sums up the long methods and narrating comments per file, putting the files that keep changing at the top. Then it tells you which commands shorten the list without reading it: sloppy fix for the mechanical findings, sloppy baseline for the debt that is already there. --all lists everything.

Review a change. sloppy diff main reports only what your branch introduced, never what it inherited. sloppy review main ranks the same findings by risk, so you know which file to read first and where to stop.

php artisan sloppy:diff main

Adopt it on a codebase with history. sloppy baseline accepts today's debt, so only new findings fail the build. Entries are keyed on class and member, not line numbers, so the baseline survives ordinary editing.

Gate it in CI. sloppy ci reads the pipeline it runs in: it annotates the diff on a GitHub pull request and fills the Code Quality widget on GitLab. The GitHub Action is three lines:

- uses: heyosseus/sloppy-action@v1
  with:
    diff-branch: main

Already on PHPStan? Get the same findings inside the run you already have:

composer require --dev heyosseus/phpstan-sloppy

Each one is a PHPStan error with its own identifier (sloppy.SL107), so @phpstan-ignore, ignoreErrors and PHPStan baselines work on it. It reports exactly what sloppy ci would fail on. See heyosseus/phpstan-sloppy.

Fail your tests on new debt. A Pest plugin ships with the package:

it('has no new slop on this branch', function (): void {
    expectCleanSloppyDiff('main');
});

Fix what a machine can fix. sloppy fix hands the mechanical findings to Rector, scoped to the files that have them, deletes the comments that only restate their code, then formats with Pint. It tells you which findings still need a person, and why no tool should touch them.

Keep score while you work. sloppy watch keeps the score and what to read first on screen, redrawing on every save. It's made to sit beside an agent that is writing code.

See it where you already look. Use SARIF for GitHub code scanning and your IDE, --format=github for inline annotations, and Markdown for a PR comment. There's also a Filament widget, a NativePHP menu-bar label, and JSON output with a published schema.

See Everyday workflow and CI and code scanning.

Why trust the numbers

It is not an AI detector. Nobody can reliably tell from source code who or what wrote it, and Sloppy never tries. It detects patterns that turn into maintenance cost, whoever wrote them. A 200-line controller action that swallows a Throwable is a problem whether a person or an agent wrote it at 3am.

Sloppy saysIt meansIt does not mean
Confidence: 88%How sure the analyser is that the pattern is really thereAny probability that the code was AI-generated
Score: 67/100A code-quality risk measure for the analysed paths"67% of this code is AI-generated"

Every number shows its working. Add --explain-risk and each score and risk prints the arithmetic that produced it. The same code always produces the same report, which is what makes Sloppy usable as a gate. See Score, severity and risk.

It is tuned for silence. Rules need several signals, not one; they understand framework conventions; and they back off wherever code could be reached indirectly. Sloppy runs over its own source in CI, and a test asserts that ordinary Laravel code produces zero findings. The rules and the score are checked against eight open-source Laravel applications, 782,578 lines, where the four healthiest score 87 to 90.

It complements your tools rather than replacing them:

ToolAnswers
PHPStan / PsalmIs this type-correct?
Pint / PHP_CodeSnifferIs this formatted consistently?
Pest / PHPUnitDoes this behave correctly?
RectorCan this be transformed mechanically?
SloppyIs this shaped like code somebody will regret?

If PHPStan can prove it, Sloppy stays out of it.

Documentation

Getting startedInstall options, every command, your first scan and how a long one is triaged, adopting on an existing codebase
Coding agentsClaude Code hooks, rulesets for every agent, Laravel Boost, the MCP server
RulesEvery rule in detail, and how false positives are kept down
Everyday workflowDiff and review, sloppy fix over Rector, Pint and narrating comments, Pest expectations, watch, Filament and NativePHP
CI and code scanningsloppy ci, the GitHub Action, GitLab, SARIF, annotations, exit codes
Score, severity and riskHow every number is calculated, and what coverage, git history and PHPStan baselines feed
Configurationconfig/sloppy.php, and taming a noisy first run
JSON outputThe machine-readable report and its contract
Custom rulesWriting your own rule, and where it shows up

Roadmap

Still ahead: inline pull-request review comments, sloppy explain for a longer write-up of one finding, HTML reports, project architecture policies, and more rules for the shortcuts agents take, such as configuration keys and routes that do not exist. Anything AI-assisted will be opt-in and separate: the analyser will always work with no API key, no network and no model.

Contributing

Found a false positive? That's a bug worth reporting, not a threshold to work around. Open an issue. To work on Sloppy itself, see CONTRIBUTING.md. composer test runs Rector, Pint, PHPStan at level 8, 100% type coverage and the suite.

Security issues: see SECURITY.md.

License

MIT. See LICENSE.md.

ai
code-quality
filament
laravel
linter
mcp
pest
php
pint
rectorphp
static-analysis
technical-debt

Languages

PHP

99.6%