simonether/skill-placebo

Placebo-controlled benchmark of the most-starred coding agent skills

Shell

1

101 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I gave the 9 most popular Claude Code skills a sugar pill. 2 beat it, 1 did worse than the pill. (r/ClaudeAI)

[Cost with each skill vs its same-length placebo \(95% CI\). Below 1 = the skill is cheaper. planning-with-files is \\"worse\\" on pass rate: 80% vs 100%.](https://preview.redd.it/9uhp4kqe2vth1.png?width=1280&format=png&auto=webp&s=65e83709a9b3923d8474ec5a01c49278…

9

Oct 6, 2026

Show HN: Popular Claude Code skills vs. a placebo: 2 beat it, 1 did worse

1

Oct 6, 2026

README

Forest plot: cost ratio of each skill vs its same-length placebo with 95% confidence intervals

skill-placebo

A placebo-controlled test of the most-starred skills for coding agents: does a skill do better than the same amount of neutral text?

Drug trials give the control group a sugar pill. We gave coding agents one: neutral instructions of the same length, installed the same way as each skill.

2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (1 worse, 6 no better). Control: a same-length neutral placebo, installed like the skill.
Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.
Measured 2026-09-30 on 15 public tasks (SWE-bench Verified, Terminal-Bench 2.1, OpenThoughts-TBLite) with claude-opus-5-5 in Claude Code, 30 trials per arm, 450 trials in total. Method registered before the first run: METHOD.md. Per-trial records: results/; the agents' full logs: release v0.1.0 (sha256 in results/*/*/AGENT_LOGS.json). Reproduce a row: uvx skill-placebo run <owner/repo>.

Results: Claude Code (claude-opus-5-5)

SkillCost vs placebo R [95% CI]ChangePass skill / placeboD, pp [95% CI]Verdictn
superpowers †0.98 [0.91, 1.06]−2%83% / 87%−3 [−10, +0]no better than placebo30/30
mattpocock1.05 [0.99, 1.12]+5%83% / 87%−3 [−10, +0]no better than placebo30/30
karpathy1.09 [0.93, 1.23]+9%83% / 87%−3 [−13, +7]no better than placebo30/30
ponytail †0.88 [0.76, 0.97]−12%90% / 87%+3 [−7, +13]beats placebo30/30
caveman0.99 [0.91, 1.06]−1%90% / 83%+7 [+0, +17]no better than placebo30/30
agent-skills0.95 [0.90, 0.99]−5%83% / 83%+0 [−10, +10]beats placebo30/30
i-have-adhd †0.92 [0.86, 0.98]−8%87% / 87%+0 [−10, +10]no better than placebo30/30
planning-with-files1.05 [0.91, 1.19]+5%80% / 100%−20 [−37, −7]worse than placebo30/30
compound-engineering1.04 [0.91, 1.23]+4%87% / 87%+0 [−10, +10]no better than placebo30/30

CIs are unadjusted; verdicts use Holm-adjusted p across 9 skills (i-have-adhd: cost CI excludes 1, Holm-adjusted p = 0.095).
† Scripted approval, not an author-documented mode (METHOD.md 5.1): the scripted turn fired in 4 of 448 recorded trials (ponytail 1, i-have-adhd 2, placebo cc-3 1), never for superpowers.
Corrected on 2026-10-06 (METHOD.md, amendment 19): two agent timeouts whose tests passed afterwards now count as failed trials, as sections 5 and 7 require; planning-with-files moves from "no better" to "worse than placebo". Costs are unchanged.

R is the skill's mean cost divided by its placebo's; below 1 the skill is cheaper. D is the pass-rate difference. Verdicts follow METHOD.md 9.1, with Holm correction across the 9 skills. D stays out of the headline: with this many trials its 95% CI is about ±10 points, too wide to rank skills by. This is v1: 30 trials per arm (METHOD.md, amendments 16-17); more runs come as updates.

Codex (gpt-6-sol): pilot only, secondary

SkillCost vs placebo R [95% CI]ChangePass skill / placeboD, pp [95% CI]Verdictn
ponytail1.09 [0.90, 1.23]+9%90% / 100%−10 [−30, +0]no better than placebo10/10
agent-skills1.24 [1.10, 1.42]+24%100% / 80%+20 [+0, +60]worse than placebo10/10
compound-engineering1.42 [1.23, 1.64]+42%100% / 100%+0 [+0, +0]worse than placebo10/10

These are the pilot's kill test on Codex: 3 skills on 5 tasks, 10 trials per arm. The Codex main run did not take place for v1 (METHOD.md, amendment 17); it may come as an update.

What the READMEs claim, and what we measured

SkillClaimed (quote, source)Their setupOur closest measureMeasured
superpowersno numeric claim in README at 8ca22dbverdict only
mattpocockno numeric claim in README at c55ee46verdict only
karpathyno numeric claim in README at 2c60614verdict only
ponytail~54% less code (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)final diff size vs baseline: lines added + deleted + untracked, SWE-bench tasks only, trials after the measure was switched on (METHOD.md amendment 12a)−13%, 95% CI [−27%, +12%]; n = 11 / 10 trials on 7 tasks (few trials: read with care)
ponytail~20% cheaper (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)cost vs baseline (and vs placebo)−1%, 95% CI [−19%, +13%]; n = 30 / 30 trials on 15 tasks
ponytail~27% faster (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)agent wall time vs baseline−11%, 95% CI [−30%, +13%]; n = 30 / 30 trials on 15 tasks
ponytail-22% tokens (README.md:85)same benchmark; the -22% column is tokenstotal tokens vs baseline+4%, 95% CI [−11%, +18%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (talking like a caveman: the agent's output) (GitHub repository description)repository description, read with gh api on 2026-09-30output tokens vs baseline (comparable to the claim)−2%, 95% CI [−9%, +6%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (GitHub repository description)repository description, read with gh api on 2026-09-30total tokens vs baseline (input and cache included; not the claim's metric)+18%, 95% CI [+8%, +28%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (GitHub repository description)repository description, read with gh api on 2026-09-30cost vs baseline (not the claim's metric)+14%, 95% CI [+8%, +21%]; n = 30 / 30 trials on 15 tasks
caveman-50% output tokens vs a terse control (README.md:206)ten dev questions, skill vs a plain 'Answer concisely.' control, claude-opus-4-6output tokens vs placebo (our placebo is neutral, not terse)−5%, 95% CI [−14%, +3%]
agent-skillsno numeric claim in README at 2686b62verdict only
i-have-adhdno numeric claim in README at 839872fverdict only
planning-with-files96.7% assertion pass rate (README.md:33)v2.21.0 eval on claude-sonnet-4-6, 30 assertions of file-pattern fidelity, not task success (README.md:702)not comparable: our pass is the task's own tests; pass rate vs baseline and placebo is shown80% vs 87% without the skill, D −7 [−20, +7] pp
planning-with-files13.3 → 5.0 turns after a context wipe (README.md:75)turns to resume after a context wipe, internal benchmark v1not measured: single-session tasks, no context wipe (METHOD.md section 13)not measured by this design
compound-engineeringno numeric claim in README at e80c5c4verdict only

The claims were measured by their authors on other tasks, models and baselines (the "their setup" column), so a gap between the two columns does not mean the claim was wrong. It shows how much the number changes in a different setup.

How it works

  • Every skill runs in three arms on the same tasks: no skill, placebo, skill.
  • The placebo is neutral text sized to the skill's always-on token footprint (within ±10%, measured), delivered through the same mechanism: plugin, hook or memory file (METHOD.md 4.1).
  • Cost is recorded tokens times public list prices: an estimate by tokens, not an invoice. The 95% CIs come from a cluster bootstrap over tasks.
  • Skills are installed as their authors document, at a pinned commit (skills.lock.json). Every trial is published, including the failed and interrupted ones.

For skill authors

If your skill was installed wrong or its placebo is unfair to it, open an issue: skill, commit, what is wrong, how to check. Every trial is public, so you can point at the exact runs. A confirmed installation error means your skill's arms are rerun in full, the result is updated, an amendment in METHOD.md records it, and the issue is linked here (METHOD.md, section 14).

Reproduce

uvx skill-placebo run DietrichGebert/ponytail   # one skill: baseline, placebo, skill on the same tasks

Limits

  • One model in the main run: claude-opus-5-5 in Claude Code at medium effort, the harness default. Codex (gpt-6-sol) ran only in a small pilot. Other models can behave differently.
  • The tasks are public and probably in the models' training data. That affects every arm equally, but absolute pass rates say little about new work.
  • Most tasks sat at the ceiling for claude-opus-5-5, so the pass rate carries little information here and the headline is cost.
  • Tasks take minutes. Skills that pay off over long sessions (memory, multi-day plans) are not measured.

Port me

  • Run a skill on OpenCode or Gemini CLI (good first issue)
  • Propose a skill: open an issue with the repository
  • Rerun on fresh tasks (a recent SWE-rebench slice)

License

MIT, see LICENSE.

Badge for tested skills

[![placebo-tested](https://raw.githubusercontent.com/simonether/skill-placebo/main/docs/badges/<skill>-cc.svg)](https://github.com/simonether/skill-placebo#results-claude-code-claude-opus-5-5)


Made by Simon (@simonether) · I take AI-built apps from demo to production at Keelfast.

agent-skills
ai-agents
benchmark
claude-code
codex
llm
placebo

simonether/skill-placebo

Placebo-controlled benchmark of the most-starred coding agent skills

Shell

1

101 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I gave the 9 most popular Claude Code skills a sugar pill. 2 beat it, 1 did worse than the pill. (r/ClaudeAI)

[Cost with each skill vs its same-length placebo \(95&amp;#37; CI\). Below 1 = the skill is cheaper. planning-with-files is \\"worse\\" on pass rate: 80&amp;#37; vs 100&amp;#37;.](https://preview.redd.it/9uhp4kqe2vth1.png?width=1280&amp;format=png&amp;auto=webp&amp;s=65e83709a9b3923d8474ec5a01c49278…

9

Oct 6, 2026

Show HN: Popular Claude Code skills vs. a placebo: 2 beat it, 1 did worse

1

Oct 6, 2026

README

Forest plot: cost ratio of each skill vs its same-length placebo with 95% confidence intervals

skill-placebo

A placebo-controlled test of the most-starred skills for coding agents: does a skill do better than the same amount of neutral text?

Drug trials give the control group a sugar pill. We gave coding agents one: neutral instructions of the same length, installed the same way as each skill.

2 of 9 skills beat a same-length placebo, both at Holm-adjusted p = 0.049 (1 worse, 6 no better). Control: a same-length neutral placebo, installed like the skill.
Secondary, against no skill: none of the 9 skills was measurably cheaper than running without a skill; a same-length placebo alone changed cost by +2% to +16%.
Measured 2026-09-30 on 15 public tasks (SWE-bench Verified, Terminal-Bench 2.1, OpenThoughts-TBLite) with claude-opus-5-5 in Claude Code, 30 trials per arm, 450 trials in total. Method registered before the first run: METHOD.md. Per-trial records: results/; the agents' full logs: release v0.1.0 (sha256 in results/*/*/AGENT_LOGS.json). Reproduce a row: uvx skill-placebo run <owner/repo>.

Results: Claude Code (claude-opus-5-5)

SkillCost vs placebo R [95% CI]ChangePass skill / placeboD, pp [95% CI]Verdictn
superpowers †0.98 [0.91, 1.06]−2%83% / 87%−3 [−10, +0]no better than placebo30/30
mattpocock1.05 [0.99, 1.12]+5%83% / 87%−3 [−10, +0]no better than placebo30/30
karpathy1.09 [0.93, 1.23]+9%83% / 87%−3 [−13, +7]no better than placebo30/30
ponytail †0.88 [0.76, 0.97]−12%90% / 87%+3 [−7, +13]beats placebo30/30
caveman0.99 [0.91, 1.06]−1%90% / 83%+7 [+0, +17]no better than placebo30/30
agent-skills0.95 [0.90, 0.99]−5%83% / 83%+0 [−10, +10]beats placebo30/30
i-have-adhd †0.92 [0.86, 0.98]−8%87% / 87%+0 [−10, +10]no better than placebo30/30
planning-with-files1.05 [0.91, 1.19]+5%80% / 100%−20 [−37, −7]worse than placebo30/30
compound-engineering1.04 [0.91, 1.23]+4%87% / 87%+0 [−10, +10]no better than placebo30/30

CIs are unadjusted; verdicts use Holm-adjusted p across 9 skills (i-have-adhd: cost CI excludes 1, Holm-adjusted p = 0.095).
† Scripted approval, not an author-documented mode (METHOD.md 5.1): the scripted turn fired in 4 of 448 recorded trials (ponytail 1, i-have-adhd 2, placebo cc-3 1), never for superpowers.
Corrected on 2026-10-06 (METHOD.md, amendment 19): two agent timeouts whose tests passed afterwards now count as failed trials, as sections 5 and 7 require; planning-with-files moves from "no better" to "worse than placebo". Costs are unchanged.

R is the skill's mean cost divided by its placebo's; below 1 the skill is cheaper. D is the pass-rate difference. Verdicts follow METHOD.md 9.1, with Holm correction across the 9 skills. D stays out of the headline: with this many trials its 95% CI is about ±10 points, too wide to rank skills by. This is v1: 30 trials per arm (METHOD.md, amendments 16-17); more runs come as updates.

Codex (gpt-6-sol): pilot only, secondary

SkillCost vs placebo R [95% CI]ChangePass skill / placeboD, pp [95% CI]Verdictn
ponytail1.09 [0.90, 1.23]+9%90% / 100%−10 [−30, +0]no better than placebo10/10
agent-skills1.24 [1.10, 1.42]+24%100% / 80%+20 [+0, +60]worse than placebo10/10
compound-engineering1.42 [1.23, 1.64]+42%100% / 100%+0 [+0, +0]worse than placebo10/10

These are the pilot's kill test on Codex: 3 skills on 5 tasks, 10 trials per arm. The Codex main run did not take place for v1 (METHOD.md, amendment 17); it may come as an update.

What the READMEs claim, and what we measured

SkillClaimed (quote, source)Their setupOur closest measureMeasured
superpowersno numeric claim in README at 8ca22dbverdict only
mattpocockno numeric claim in README at c55ee46verdict only
karpathyno numeric claim in README at 2c60614verdict only
ponytail~54% less code (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)final diff size vs baseline: lines added + deleted + untracked, SWE-bench tasks only, trials after the measure was switched on (METHOD.md amendment 12a)−13%, 95% CI [−27%, +12%]; n = 11 / 10 trials on 7 tasks (few trials: read with care)
ponytail~20% cheaper (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)cost vs baseline (and vs placebo)−1%, 95% CI [−19%, +13%]; n = 30 / 30 trials on 15 tasks
ponytail~27% faster (README.md:33)mean over 12 feature tasks, Haiku 4.5, n=4, vs the same agent with no skill (README.md:34)agent wall time vs baseline−11%, 95% CI [−30%, +13%]; n = 30 / 30 trials on 15 tasks
ponytail-22% tokens (README.md:85)same benchmark; the -22% column is tokenstotal tokens vs baseline+4%, 95% CI [−11%, +18%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (talking like a caveman: the agent's output) (GitHub repository description)repository description, read with gh api on 2026-09-30output tokens vs baseline (comparable to the claim)−2%, 95% CI [−9%, +6%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (GitHub repository description)repository description, read with gh api on 2026-09-30total tokens vs baseline (input and cache included; not the claim's metric)+18%, 95% CI [+8%, +28%]; n = 30 / 30 trials on 15 tasks
caveman65% fewer tokens (GitHub repository description)repository description, read with gh api on 2026-09-30cost vs baseline (not the claim's metric)+14%, 95% CI [+8%, +21%]; n = 30 / 30 trials on 15 tasks
caveman-50% output tokens vs a terse control (README.md:206)ten dev questions, skill vs a plain 'Answer concisely.' control, claude-opus-4-6output tokens vs placebo (our placebo is neutral, not terse)−5%, 95% CI [−14%, +3%]
agent-skillsno numeric claim in README at 2686b62verdict only
i-have-adhdno numeric claim in README at 839872fverdict only
planning-with-files96.7% assertion pass rate (README.md:33)v2.21.0 eval on claude-sonnet-4-6, 30 assertions of file-pattern fidelity, not task success (README.md:702)not comparable: our pass is the task's own tests; pass rate vs baseline and placebo is shown80% vs 87% without the skill, D −7 [−20, +7] pp
planning-with-files13.3 → 5.0 turns after a context wipe (README.md:75)turns to resume after a context wipe, internal benchmark v1not measured: single-session tasks, no context wipe (METHOD.md section 13)not measured by this design
compound-engineeringno numeric claim in README at e80c5c4verdict only

The claims were measured by their authors on other tasks, models and baselines (the "their setup" column), so a gap between the two columns does not mean the claim was wrong. It shows how much the number changes in a different setup.

How it works

  • Every skill runs in three arms on the same tasks: no skill, placebo, skill.
  • The placebo is neutral text sized to the skill's always-on token footprint (within ±10%, measured), delivered through the same mechanism: plugin, hook or memory file (METHOD.md 4.1).
  • Cost is recorded tokens times public list prices: an estimate by tokens, not an invoice. The 95% CIs come from a cluster bootstrap over tasks.
  • Skills are installed as their authors document, at a pinned commit (skills.lock.json). Every trial is published, including the failed and interrupted ones.

For skill authors

If your skill was installed wrong or its placebo is unfair to it, open an issue: skill, commit, what is wrong, how to check. Every trial is public, so you can point at the exact runs. A confirmed installation error means your skill's arms are rerun in full, the result is updated, an amendment in METHOD.md records it, and the issue is linked here (METHOD.md, section 14).

Reproduce

uvx skill-placebo run DietrichGebert/ponytail   # one skill: baseline, placebo, skill on the same tasks

Limits

  • One model in the main run: claude-opus-5-5 in Claude Code at medium effort, the harness default. Codex (gpt-6-sol) ran only in a small pilot. Other models can behave differently.
  • The tasks are public and probably in the models' training data. That affects every arm equally, but absolute pass rates say little about new work.
  • Most tasks sat at the ceiling for claude-opus-5-5, so the pass rate carries little information here and the headline is cost.
  • Tasks take minutes. Skills that pay off over long sessions (memory, multi-day plans) are not measured.

Port me

  • Run a skill on OpenCode or Gemini CLI (good first issue)
  • Propose a skill: open an issue with the repository
  • Rerun on fresh tasks (a recent SWE-rebench slice)

License

MIT, see LICENSE.

Badge for tested skills

[![placebo-tested](https://raw.githubusercontent.com/simonether/skill-placebo/main/docs/badges/<skill>-cc.svg)](https://github.com/simonether/skill-placebo#results-claude-code-claude-opus-5-5)


Made by Simon (@simonether) · I take AI-built apps from demo to production at Keelfast.

agent-skills
ai-agents
benchmark
claude-code
codex
llm
placebo