syntaxixr/goalpost-benchmark

Dataset

goalpost benchmark

0

4 commits

1 linked in READMEs

updated Oct 7, 2026

See the code

README

goalpost: makes Claude Code's /goal actually finish the job

goalpost benchmark

Does Claude Code's built-in /goal know when the work is really done? This dataset has every valid run from the benchmark behind goalpost, a hooks plugin that makes /goal pass real checks before Claude is allowed to stop.

33 runs: Sonnet 5.5 and Haiku 4.5, four small tasks (a ledger, a text toolkit, a job queue, a spreadsheet engine), each run with plain /goal and with goalpost. Every task has hidden checks the agent never sees, and every run was graded three times.

Same model, same task: plain /goal ends at 28/34 hidden checks, with goalpost at 34/34

What it shows

plain /goalwith goalpost
Sonnet 5.5, runs fully correct8 of 99 of 9
Haiku 4.5, hidden checks passed (mean of the 4 task means)79.8%89.7%
Haiku 4.5, runs fully correct1 of 81 of 7
Time per run (Sonnet / Haiku)2.7 / 5.7 min5.2 / 15.1 min
Cost per run (Sonnet / Haiku)$0.46 / $0.64$0.91 / $1.25

The built-in judge said "met" in all 17 plain runs, and 8 of those were broken (1 Sonnet, 7 Haiku). It reads the conversation and never runs anything. goalpost never ended a Sonnet run with failing checks, and it moved Haiku the right way, but it did not make Haiku reliable, and it costs about twice the time and money.

Runs vary a lot. With two runs per task, plain Haiku scored 82% and 4% on two runs of the same spreadsheet prompt. Without that task the Haiku means are 92.0% (plain) and 94.4% (goalpost), so most of the Haiku gap comes from one task.

The samples are small. My account's rate limit cut several runs, which is why the Haiku arms have 8 and 7 runs. Every run that was thrown out, and why, is listed in the full report.

The pair in the gif (Haiku, text toolkit task) is the most striking of the 33, picked for the video. The table above is the honest summary.

Fields

One row per run in data/runs.jsonl:

  • run_id, round, model, task, arm (bare or goalpost), rep, claude_code_version, started_at
  • hidden_checks_passed, hidden_checks_total, requirements_passed, requirements_total, score, fully_correct, graded_times
  • builtin_judge_final_verdict: the last verdict of the built-in /goal judge (met in 32 runs, none in one run where it never gave a final verdict)
  • done_while_checks_failed: the judge said met but hidden checks failed
  • goalpost_end, goalpost_stop_blocks, goalpost_denies, goalpost_audits: what goalpost did (empty for plain runs)
  • wall_seconds, cost_usd, turns, tool_calls, subagents, diff_stat
  • requirements: every hidden check with its result
  • agent_last_message: what the agent said at the end, cut to 1500 characters
from datasets import load_dataset
runs = load_dataset("syntaxixr/goalpost-benchmark", split="train")

The tasks, the hidden graders and the harness are in the repository, so every number here can be re-run.

agents
benchmark
claude-code
coding-agents
evaluation
verification

syntaxixr/goalpost-benchmark

Dataset

goalpost benchmark

0

4 commits

1 linked in READMEs

updated Oct 7, 2026

See the code

README

goalpost: makes Claude Code's /goal actually finish the job

goalpost benchmark

Does Claude Code's built-in /goal know when the work is really done? This dataset has every valid run from the benchmark behind goalpost, a hooks plugin that makes /goal pass real checks before Claude is allowed to stop.

33 runs: Sonnet 5.5 and Haiku 4.5, four small tasks (a ledger, a text toolkit, a job queue, a spreadsheet engine), each run with plain /goal and with goalpost. Every task has hidden checks the agent never sees, and every run was graded three times.

Same model, same task: plain /goal ends at 28/34 hidden checks, with goalpost at 34/34

What it shows

plain /goalwith goalpost
Sonnet 5.5, runs fully correct8 of 99 of 9
Haiku 4.5, hidden checks passed (mean of the 4 task means)79.8%89.7%
Haiku 4.5, runs fully correct1 of 81 of 7
Time per run (Sonnet / Haiku)2.7 / 5.7 min5.2 / 15.1 min
Cost per run (Sonnet / Haiku)$0.46 / $0.64$0.91 / $1.25

The built-in judge said "met" in all 17 plain runs, and 8 of those were broken (1 Sonnet, 7 Haiku). It reads the conversation and never runs anything. goalpost never ended a Sonnet run with failing checks, and it moved Haiku the right way, but it did not make Haiku reliable, and it costs about twice the time and money.

Runs vary a lot. With two runs per task, plain Haiku scored 82% and 4% on two runs of the same spreadsheet prompt. Without that task the Haiku means are 92.0% (plain) and 94.4% (goalpost), so most of the Haiku gap comes from one task.

The samples are small. My account's rate limit cut several runs, which is why the Haiku arms have 8 and 7 runs. Every run that was thrown out, and why, is listed in the full report.

The pair in the gif (Haiku, text toolkit task) is the most striking of the 33, picked for the video. The table above is the honest summary.

Fields

One row per run in data/runs.jsonl:

  • run_id, round, model, task, arm (bare or goalpost), rep, claude_code_version, started_at
  • hidden_checks_passed, hidden_checks_total, requirements_passed, requirements_total, score, fully_correct, graded_times
  • builtin_judge_final_verdict: the last verdict of the built-in /goal judge (met in 32 runs, none in one run where it never gave a final verdict)
  • done_while_checks_failed: the judge said met but hidden checks failed
  • goalpost_end, goalpost_stop_blocks, goalpost_denies, goalpost_audits: what goalpost did (empty for plain runs)
  • wall_seconds, cost_usd, turns, tool_calls, subagents, diff_stat
  • requirements: every hidden check with its result
  • agent_last_message: what the agent said at the end, cut to 1500 characters
from datasets import load_dataset
runs = load_dataset("syntaxixr/goalpost-benchmark", split="train")

The tasks, the hidden graders and the harness are in the repository, so every number here can be re-run.

agents
benchmark
claude-code
coding-agents
evaluation
verification