Placebo-controlled benchmark of the most-starred coding agent skills
Shell
1
101 commits
updated Oct 6, 2026
A placebo-controlled test of the most-starred skills for coding agents: does a skill do better than the same amount of neutral text?
Drug trials give the control group a sugar pill. We gave coding agents one: neutral instructions of the same length, installed the same way as each skill.
2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (1 worse, 6 no better). Control: a same-length neutral placebo, installed like the skill.
Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.
Measured 2026-09-30 on 15 public tasks (SWE-bench Verified, Terminal-Bench 2.1, OpenThoughts-TBLite) with claude-opus-5-5 in Claude Code, 30 trials per arm, 450 trials in total. Method registered before the first run: METHOD.md. Per-trial records: results/; the agents' full logs: release v0.1.0 (sha256 in results/*/*/AGENT_LOGS.json). Reproduce a row: uvx skill-placebo run <owner/repo>.
| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
|---|---|---|---|---|---|---|
| superpowers † | 0.98 [0.91, 1.06] | −2% | 83% / 87% | −3 [−10, +0] | no better than placebo | 30/30 |
| mattpocock | 1.05 [0.99, 1.12] | +5% | 83% / 87% | −3 [−10, +0] | no better than placebo | 30/30 |
| karpathy | 1.09 [0.93, 1.23] | +9% | 83% / 87% | −3 [−13, +7] | no better than placebo | 30/30 |
| ponytail † | 0.88 [0.76, 0.97] | −12% | 90% / 87% | +3 [−7, +13] | beats placebo | 30/30 |
| caveman | 0.99 [0.91, 1.06] | −1% | 90% / 83% | +7 [+0, +17] | no better than placebo | 30/30 |
| agent-skills | 0.95 [0.90, 0.99] | −5% | 83% / 83% | +0 [−10, +10] | beats placebo | 30/30 |
| i-have-adhd † | 0.92 [0.86, 0.98] | −8% | 87% / 87% | +0 [−10, +10] | no better than placebo | 30/30 |
| planning-with-files | 1.05 [0.91, 1.19] | +5% | 80% / 100% | −20 [−37, −7] | worse than placebo | 30/30 |
| compound-engineering | 1.04 [0.91, 1.23] | +4% | 87% / 87% | +0 [−10, +10] | no better than placebo | 30/30 |
CIs are unadjusted; verdicts use Holm-adjusted p across 9 skills (i-have-adhd: cost CI excludes 1, Holm-adjusted p = 0.095).
† Scripted approval, not an author-documented mode (METHOD.md 5.1): the scripted turn fired in 4 of 448 recorded trials (ponytail 1, i-have-adhd 2, placebo cc-3 1), never for superpowers.
Corrected on 2026-10-06 (METHOD.md, amendment 19): two agent timeouts whose tests passed afterwards now count as failed trials, as sections 5 and 7 require; planning-with-files moves from "no better" to "worse than placebo". Costs are unchanged.
R is the skill's mean cost divided by its placebo's; below 1 the skill is cheaper. D is the pass-rate difference. Verdicts follow METHOD.md 9.1, with Holm correction across the 9 skills. D stays out of the headline: with this many trials its 95% CI is about ±10 points, too wide to rank skills by. This is v1: 30 trials per arm (METHOD.md, amendments 16-17); more runs come as updates.
| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
|---|---|---|---|---|---|---|
| ponytail | 1.09 [0.90, 1.23] | +9% | 90% / 100% | −10 [−30, +0] | no better than placebo | 10/10 |
| agent-skills | 1.24 [1.10, 1.42] | +24% | 100% / 80% | +20 [+0, +60] | worse than placebo | 10/10 |
| compound-engineering | 1.42 [1.23, 1.64] | +42% | 100% / 100% | +0 [+0, +0] | worse than placebo | 10/10 |
These are the pilot's kill test on Codex: 3 skills on 5 tasks, 10 trials per arm. The Codex main run did not take place for v1 (METHOD.md, amendment 17); it may come as an update.
| Skill | Claimed (quote, source) | Their setup | Our closest measure | Measured |
|---|---|---|---|---|
| superpowers | no numeric claim in README at 8ca22db | verdict only | ||
| mattpocock | no numeric claim in README at c55ee46 | verdict only | ||
| karpathy | no numeric claim in README at 2c60614 | verdict only | ||
| ponytail | ~54% less code (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | final diff size vs baseline: lines added + deleted + untracked, SWE-bench tasks only, trials after the measure was switched on (METHOD.md amendment 12a) | −13%, 95% CI [−27%, +12%]; n = 11 / 10 trials on 7 tasks (few trials: read with care) |
| ponytail | ~20% cheaper (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | cost vs baseline (and vs placebo) | −1%, 95% CI [−19%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | ~27% faster (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | agent wall time vs baseline | −11%, 95% CI [−30%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | -22% tokens (README.md:85) | same benchmark; the -22% column is tokens | total tokens vs baseline | +4%, 95% CI [−11%, +18%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (talking like a caveman: the agent's output) (GitHub repository description) | repository description, read with gh api on 2026-09-30 | output tokens vs baseline (comparable to the claim) | −2%, 95% CI [−9%, +6%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (GitHub repository description) | repository description, read with gh api on 2026-09-30 | total tokens vs baseline (input and cache included; not the claim's metric) | +18%, 95% CI [+8%, +28%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (GitHub repository description) | repository description, read with gh api on 2026-09-30 | cost vs baseline (not the claim's metric) | +14%, 95% CI [+8%, +21%]; n = 30 / 30 trials on 15 tasks |
| caveman | -50% output tokens vs a terse control (README.md:206) | ten dev questions, skill vs a plain 'Answer concisely.' control, claude-opus-4-6 | output tokens vs placebo (our placebo is neutral, not terse) | −5%, 95% CI [−14%, +3%] |
| agent-skills | no numeric claim in README at 2686b62 | verdict only | ||
| i-have-adhd | no numeric claim in README at 839872f | verdict only | ||
| planning-with-files | 96.7% assertion pass rate (README.md:33) | v2.21.0 eval on claude-sonnet-4-6, 30 assertions of file-pattern fidelity, not task success (README.md:702) | not comparable: our pass is the task's own tests; pass rate vs baseline and placebo is shown | 80% vs 87% without the skill, D −7 [−20, +7] pp |
| planning-with-files | 13.3 → 5.0 turns after a context wipe (README.md:75) | turns to resume after a context wipe, internal benchmark v1 | not measured: single-session tasks, no context wipe (METHOD.md section 13) | not measured by this design |
| compound-engineering | no numeric claim in README at e80c5c4 | verdict only |
The claims were measured by their authors on other tasks, models and baselines (the "their setup" column), so a gap between the two columns does not mean the claim was wrong. It shows how much the number changes in a different setup.
If your skill was installed wrong or its placebo is unfair to it, open an issue: skill, commit, what is wrong, how to check. Every trial is public, so you can point at the exact runs. A confirmed installation error means your skill's arms are rerun in full, the result is updated, an amendment in METHOD.md records it, and the issue is linked here (METHOD.md, section 14).
uvx skill-placebo run DietrichGebert/ponytail # one skill: baseline, placebo, skill on the same tasks
good first issue)MIT, see LICENSE.
[](https://github.com/simonether/skill-placebo#results-claude-code-claude-opus-5-5)
Made by Simon (@simonether) · I take AI-built apps from demo to production at Keelfast.
Placebo-controlled benchmark of the most-starred coding agent skills
Shell
1
101 commits
updated Oct 6, 2026
A placebo-controlled test of the most-starred skills for coding agents: does a skill do better than the same amount of neutral text?
Drug trials give the control group a sugar pill. We gave coding agents one: neutral instructions of the same length, installed the same way as each skill.
2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (1 worse, 6 no better). Control: a same-length neutral placebo, installed like the skill.
Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.
Measured 2026-09-30 on 15 public tasks (SWE-bench Verified, Terminal-Bench 2.1, OpenThoughts-TBLite) with claude-opus-5-5 in Claude Code, 30 trials per arm, 450 trials in total. Method registered before the first run: METHOD.md. Per-trial records: results/; the agents' full logs: release v0.1.0 (sha256 in results/*/*/AGENT_LOGS.json). Reproduce a row: uvx skill-placebo run <owner/repo>.
| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
|---|---|---|---|---|---|---|
| superpowers † | 0.98 [0.91, 1.06] | −2% | 83% / 87% | −3 [−10, +0] | no better than placebo | 30/30 |
| mattpocock | 1.05 [0.99, 1.12] | +5% | 83% / 87% | −3 [−10, +0] | no better than placebo | 30/30 |
| karpathy | 1.09 [0.93, 1.23] | +9% | 83% / 87% | −3 [−13, +7] | no better than placebo | 30/30 |
| ponytail † | 0.88 [0.76, 0.97] | −12% | 90% / 87% | +3 [−7, +13] | beats placebo | 30/30 |
| caveman | 0.99 [0.91, 1.06] | −1% | 90% / 83% | +7 [+0, +17] | no better than placebo | 30/30 |
| agent-skills | 0.95 [0.90, 0.99] | −5% | 83% / 83% | +0 [−10, +10] | beats placebo | 30/30 |
| i-have-adhd † | 0.92 [0.86, 0.98] | −8% | 87% / 87% | +0 [−10, +10] | no better than placebo | 30/30 |
| planning-with-files | 1.05 [0.91, 1.19] | +5% | 80% / 100% | −20 [−37, −7] | worse than placebo | 30/30 |
| compound-engineering | 1.04 [0.91, 1.23] | +4% | 87% / 87% | +0 [−10, +10] | no better than placebo | 30/30 |
CIs are unadjusted; verdicts use Holm-adjusted p across 9 skills (i-have-adhd: cost CI excludes 1, Holm-adjusted p = 0.095).
† Scripted approval, not an author-documented mode (METHOD.md 5.1): the scripted turn fired in 4 of 448 recorded trials (ponytail 1, i-have-adhd 2, placebo cc-3 1), never for superpowers.
Corrected on 2026-10-06 (METHOD.md, amendment 19): two agent timeouts whose tests passed afterwards now count as failed trials, as sections 5 and 7 require; planning-with-files moves from "no better" to "worse than placebo". Costs are unchanged.
R is the skill's mean cost divided by its placebo's; below 1 the skill is cheaper. D is the pass-rate difference. Verdicts follow METHOD.md 9.1, with Holm correction across the 9 skills. D stays out of the headline: with this many trials its 95% CI is about ±10 points, too wide to rank skills by. This is v1: 30 trials per arm (METHOD.md, amendments 16-17); more runs come as updates.
| Skill | Cost vs placebo R [95% CI] | Change | Pass skill / placebo | D, pp [95% CI] | Verdict | n |
|---|---|---|---|---|---|---|
| ponytail | 1.09 [0.90, 1.23] | +9% | 90% / 100% | −10 [−30, +0] | no better than placebo | 10/10 |
| agent-skills | 1.24 [1.10, 1.42] | +24% | 100% / 80% | +20 [+0, +60] | worse than placebo | 10/10 |
| compound-engineering | 1.42 [1.23, 1.64] | +42% | 100% / 100% | +0 [+0, +0] | worse than placebo | 10/10 |
These are the pilot's kill test on Codex: 3 skills on 5 tasks, 10 trials per arm. The Codex main run did not take place for v1 (METHOD.md, amendment 17); it may come as an update.
| Skill | Claimed (quote, source) | Their setup | Our closest measure | Measured |
|---|---|---|---|---|
| superpowers | no numeric claim in README at 8ca22db | verdict only | ||
| mattpocock | no numeric claim in README at c55ee46 | verdict only | ||
| karpathy | no numeric claim in README at 2c60614 | verdict only | ||
| ponytail | ~54% less code (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | final diff size vs baseline: lines added + deleted + untracked, SWE-bench tasks only, trials after the measure was switched on (METHOD.md amendment 12a) | −13%, 95% CI [−27%, +12%]; n = 11 / 10 trials on 7 tasks (few trials: read with care) |
| ponytail | ~20% cheaper (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | cost vs baseline (and vs placebo) | −1%, 95% CI [−19%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | ~27% faster (README.md:33) | mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34) | agent wall time vs baseline | −11%, 95% CI [−30%, +13%]; n = 30 / 30 trials on 15 tasks |
| ponytail | -22% tokens (README.md:85) | same benchmark; the -22% column is tokens | total tokens vs baseline | +4%, 95% CI [−11%, +18%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (talking like a caveman: the agent's output) (GitHub repository description) | repository description, read with gh api on 2026-09-30 | output tokens vs baseline (comparable to the claim) | −2%, 95% CI [−9%, +6%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (GitHub repository description) | repository description, read with gh api on 2026-09-30 | total tokens vs baseline (input and cache included; not the claim's metric) | +18%, 95% CI [+8%, +28%]; n = 30 / 30 trials on 15 tasks |
| caveman | 65% fewer tokens (GitHub repository description) | repository description, read with gh api on 2026-09-30 | cost vs baseline (not the claim's metric) | +14%, 95% CI [+8%, +21%]; n = 30 / 30 trials on 15 tasks |
| caveman | -50% output tokens vs a terse control (README.md:206) | ten dev questions, skill vs a plain 'Answer concisely.' control, claude-opus-4-6 | output tokens vs placebo (our placebo is neutral, not terse) | −5%, 95% CI [−14%, +3%] |
| agent-skills | no numeric claim in README at 2686b62 | verdict only | ||
| i-have-adhd | no numeric claim in README at 839872f | verdict only | ||
| planning-with-files | 96.7% assertion pass rate (README.md:33) | v2.21.0 eval on claude-sonnet-4-6, 30 assertions of file-pattern fidelity, not task success (README.md:702) | not comparable: our pass is the task's own tests; pass rate vs baseline and placebo is shown | 80% vs 87% without the skill, D −7 [−20, +7] pp |
| planning-with-files | 13.3 → 5.0 turns after a context wipe (README.md:75) | turns to resume after a context wipe, internal benchmark v1 | not measured: single-session tasks, no context wipe (METHOD.md section 13) | not measured by this design |
| compound-engineering | no numeric claim in README at e80c5c4 | verdict only |
The claims were measured by their authors on other tasks, models and baselines (the "their setup" column), so a gap between the two columns does not mean the claim was wrong. It shows how much the number changes in a different setup.
If your skill was installed wrong or its placebo is unfair to it, open an issue: skill, commit, what is wrong, how to check. Every trial is public, so you can point at the exact runs. A confirmed installation error means your skill's arms are rerun in full, the result is updated, an amendment in METHOD.md records it, and the issue is linked here (METHOD.md, section 14).
uvx skill-placebo run DietrichGebert/ponytail # one skill: baseline, placebo, skill on the same tasks
good first issue)MIT, see LICENSE.
[](https://github.com/simonether/skill-placebo#results-claude-code-claude-opus-5-5)
Made by Simon (@simonether) · I take AI-built apps from demo to production at Keelfast.