vhsgreed/cost-predictor-spike

Pre-registered study: what an AI coding-agent run will cost before vs during the run. Code, plan, outputs.

HTML

0

40 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

This study of 424k agent runs showed you can't predict what a Claude agent run will cost up front, but you can tell when one is going wrong (r/ClaudeAI)

I spent the past week testing whether the cost of a coding-agent run can be predicted before it starts, using 424,108 public sessions from the AgentLogs dataset (mostly Claude Sonnet, triggered from GitHub) and writing down every test in git before running it. Knowing the model and what started the…

2

Oct 7, 2026

README

Agent cost prediction: lab record

Code, pre-registration and outputs behind the article draft journal/cost-prediction-draft.md ("What an AI coding-agent run will cost: what can be known before it starts, and what can only be seen while it runs").

  • agentlogs/PLAN.md: every hypothesis and pass bar, with the date it was written; check commit history to see that each bar predates the data it was tested on. Also the change log of errors and amendments.
  • agentlogs/*.py: analysis scripts. agentlogs/*.out / *.json: their committed outputs.
  • agentlogs/figures/: figures 1 to 5 and groups_full.csv (175 model x trigger x prompt-length groups: n, p10, p50, p90, mean).

No raw dataset rows are committed (agentlogs/handcheck.html holds derived tool-call patterns for 60 AgentLogs sessions, CC BY 4.0, used for the H2 hand-check). All inputs are public and downloaded at pinned revisions (see PLAN.md).

Environment

Python 3.14. pip install -r requirements.txt. About 60 GB of free disk for the AgentLogs pass (shards are deleted after parsing), plus about 10 GB for the external datasets. H5 also needs Ollama with nomic-embed-text.

Reproduce

Run from agentlogs/, in this order:

python test_lab.py                   # unit tests (must pass first)
python build_dataset.py ...          # dev/val/test sets from AgentLogs shards (see PLAN.md for shard lists)
python validate_h1.py                # H1 validation
python locked_test.py                # H1 + H4 locked test (one look)
python run_full.py                   # all 276 shards -> full_rows/
python full_tables.py && python figures.py
python h5_final.py                   # H5
python explore_r3.py && python confirm_r3.py        # H6, H7
python replicate_r4.py               # R4 (needs external/ downloads, see PLAN.md round 4)
python outcome_r5.py                 # O1 (needs external/ downloads, see PLAN.md round 5)
python rescore_dev.py && python revision_v02.py && python alarm_runtime_r5.py   # v0.2 supplementary analyses

Seeds are fixed in each script. Bootstrap CIs may differ in the last digit across platforms.

Not included

Box A of the article (one long conversation, Fig. 6) uses the author's own agent logs. Those logs and the scripts that read them are kept on local hardware and are not published; Box A is illustrative only.

Licence

Code: MIT. Data licences are those of the source datasets (AgentLogs CC BY 4.0; SWE-smith MIT; Nebius datasets CC BY 4.0; melissapan CC BY 4.0; Exgentic v2 no licence listed and pi sessions "other": only aggregate statistics from these two are reported).

vhsgreed/cost-predictor-spike

Pre-registered study: what an AI coding-agent run will cost before vs during the run. Code, plan, outputs.

HTML

0

40 commits

updated Oct 7, 2026

See the code

See what people are saying

SourceMessageScoreDate

This study of 424k agent runs showed you can't predict what a Claude agent run will cost up front, but you can tell when one is going wrong (r/ClaudeAI)

I spent the past week testing whether the cost of a coding-agent run can be predicted before it starts, using 424,108 public sessions from the AgentLogs dataset (mostly Claude Sonnet, triggered from GitHub) and writing down every test in git before running it. Knowing the model and what started the…

2

Oct 7, 2026

README

Agent cost prediction: lab record

Code, pre-registration and outputs behind the article draft journal/cost-prediction-draft.md ("What an AI coding-agent run will cost: what can be known before it starts, and what can only be seen while it runs").

  • agentlogs/PLAN.md: every hypothesis and pass bar, with the date it was written; check commit history to see that each bar predates the data it was tested on. Also the change log of errors and amendments.
  • agentlogs/*.py: analysis scripts. agentlogs/*.out / *.json: their committed outputs.
  • agentlogs/figures/: figures 1 to 5 and groups_full.csv (175 model x trigger x prompt-length groups: n, p10, p50, p90, mean).

No raw dataset rows are committed (agentlogs/handcheck.html holds derived tool-call patterns for 60 AgentLogs sessions, CC BY 4.0, used for the H2 hand-check). All inputs are public and downloaded at pinned revisions (see PLAN.md).

Environment

Python 3.14. pip install -r requirements.txt. About 60 GB of free disk for the AgentLogs pass (shards are deleted after parsing), plus about 10 GB for the external datasets. H5 also needs Ollama with nomic-embed-text.

Reproduce

Run from agentlogs/, in this order:

python test_lab.py                   # unit tests (must pass first)
python build_dataset.py ...          # dev/val/test sets from AgentLogs shards (see PLAN.md for shard lists)
python validate_h1.py                # H1 validation
python locked_test.py                # H1 + H4 locked test (one look)
python run_full.py                   # all 276 shards -> full_rows/
python full_tables.py && python figures.py
python h5_final.py                   # H5
python explore_r3.py && python confirm_r3.py        # H6, H7
python replicate_r4.py               # R4 (needs external/ downloads, see PLAN.md round 4)
python outcome_r5.py                 # O1 (needs external/ downloads, see PLAN.md round 5)
python rescore_dev.py && python revision_v02.py && python alarm_runtime_r5.py   # v0.2 supplementary analyses

Seeds are fixed in each script. Bootstrap CIs may differ in the last digit across platforms.

Not included

Box A of the article (one long conversation, Fig. 6) uses the author's own agent logs. Those logs and the scripts that read them are kept on local hardware and are not published; Box A is illustrative only.

Licence

Code: MIT. Data licences are those of the source datasets (AgentLogs CC BY 4.0; SWE-smith MIT; Nebius datasets CC BY 4.0; melissapan CC BY 4.0; Exgentic v2 no licence listed and pi sessions "other": only aggregate statistics from these two are reported).