
Does Claude Code's built-in /goal know when the work is really done? This dataset has every valid run from the benchmark behind goalpost, a hooks plugin that makes /goal pass real checks before Claude is allowed to stop.
33 runs: Sonnet 5.5 and Haiku 4.5, four small tasks (a ledger, a text toolkit, a job queue, a spreadsheet engine), each run with plain /goal and with goalpost. Every task has hidden checks the agent never sees, and every run was graded three times.

plain /goal | with goalpost | |
|---|---|---|
| Sonnet 5.5, runs fully correct | 8 of 9 | 9 of 9 |
| Haiku 4.5, hidden checks passed (mean of the 4 task means) | 79.8% | 89.7% |
| Haiku 4.5, runs fully correct | 1 of 8 | 1 of 7 |
| Time per run (Sonnet / Haiku) | 2.7 / 5.7 min | 5.2 / 15.1 min |
| Cost per run (Sonnet / Haiku) | $0.46 / $0.64 | $0.91 / $1.25 |
The built-in judge said "met" in all 17 plain runs, and 8 of those were broken (1 Sonnet, 7 Haiku). It reads the conversation and never runs anything. goalpost never ended a Sonnet run with failing checks, and it moved Haiku the right way, but it did not make Haiku reliable, and it costs about twice the time and money.
Runs vary a lot. With two runs per task, plain Haiku scored 82% and 4% on two runs of the same spreadsheet prompt. Without that task the Haiku means are 92.0% (plain) and 94.4% (goalpost), so most of the Haiku gap comes from one task.
The samples are small. My account's rate limit cut several runs, which is why the Haiku arms have 8 and 7 runs. Every run that was thrown out, and why, is listed in the full report.
The pair in the gif (Haiku, text toolkit task) is the most striking of the 33, picked for the video. The table above is the honest summary.
One row per run in data/runs.jsonl:
run_id, round, model, task, arm (bare or goalpost), rep, claude_code_version, started_athidden_checks_passed, hidden_checks_total, requirements_passed, requirements_total, score, fully_correct, graded_timesbuiltin_judge_final_verdict: the last verdict of the built-in /goal judge (met in 32 runs, none in one run where it never gave a final verdict)done_while_checks_failed: the judge said met but hidden checks failedgoalpost_end, goalpost_stop_blocks, goalpost_denies, goalpost_audits: what goalpost did (empty for plain runs)wall_seconds, cost_usd, turns, tool_calls, subagents, diff_statrequirements: every hidden check with its resultagent_last_message: what the agent said at the end, cut to 1500 charactersfrom datasets import load_dataset
runs = load_dataset("syntaxixr/goalpost-benchmark", split="train")
The tasks, the hidden graders and the harness are in the repository, so every number here can be re-run.

Does Claude Code's built-in /goal know when the work is really done? This dataset has every valid run from the benchmark behind goalpost, a hooks plugin that makes /goal pass real checks before Claude is allowed to stop.
33 runs: Sonnet 5.5 and Haiku 4.5, four small tasks (a ledger, a text toolkit, a job queue, a spreadsheet engine), each run with plain /goal and with goalpost. Every task has hidden checks the agent never sees, and every run was graded three times.

plain /goal | with goalpost | |
|---|---|---|
| Sonnet 5.5, runs fully correct | 8 of 9 | 9 of 9 |
| Haiku 4.5, hidden checks passed (mean of the 4 task means) | 79.8% | 89.7% |
| Haiku 4.5, runs fully correct | 1 of 8 | 1 of 7 |
| Time per run (Sonnet / Haiku) | 2.7 / 5.7 min | 5.2 / 15.1 min |
| Cost per run (Sonnet / Haiku) | $0.46 / $0.64 | $0.91 / $1.25 |
The built-in judge said "met" in all 17 plain runs, and 8 of those were broken (1 Sonnet, 7 Haiku). It reads the conversation and never runs anything. goalpost never ended a Sonnet run with failing checks, and it moved Haiku the right way, but it did not make Haiku reliable, and it costs about twice the time and money.
Runs vary a lot. With two runs per task, plain Haiku scored 82% and 4% on two runs of the same spreadsheet prompt. Without that task the Haiku means are 92.0% (plain) and 94.4% (goalpost), so most of the Haiku gap comes from one task.
The samples are small. My account's rate limit cut several runs, which is why the Haiku arms have 8 and 7 runs. Every run that was thrown out, and why, is listed in the full report.
The pair in the gif (Haiku, text toolkit task) is the most striking of the 33, picked for the video. The table above is the honest summary.
One row per run in data/runs.jsonl:
run_id, round, model, task, arm (bare or goalpost), rep, claude_code_version, started_athidden_checks_passed, hidden_checks_total, requirements_passed, requirements_total, score, fully_correct, graded_timesbuiltin_judge_final_verdict: the last verdict of the built-in /goal judge (met in 32 runs, none in one run where it never gave a final verdict)done_while_checks_failed: the judge said met but hidden checks failedgoalpost_end, goalpost_stop_blocks, goalpost_denies, goalpost_audits: what goalpost did (empty for plain runs)wall_seconds, cost_usd, turns, tool_calls, subagents, diff_statrequirements: every hidden check with its resultagent_last_message: what the agent said at the end, cut to 1500 charactersfrom datasets import load_dataset
runs = load_dataset("syntaxixr/goalpost-benchmark", split="train")
The tasks, the hidden graders and the harness are in the repository, so every number here can be re-run.