An exploratory longitudinal N-of-1 study of human-AI coordination
A long-running interaction with a frontier model became much easier for one user to coordinate than fresh instances. We tested several simple explanations instead of assuming a new mechanism.
| Test | Result |
|---|---|
| Different state compression strategies | Different representations, but utility advantage not established |
| Compact state vs richer/full trajectory continuation | High quality across conditions; no useful directional advantage |
| Heavy Claim Lifecycle governance vs lightweight baseline | 7/7 vs 7/7 — NO_MEASURABLE_ADVANTAGE |
| 496-word static coordination packet vs no packet | Practical coordination benefit NOT ESTABLISHED (1/5 burden dimensions improved) |
The packet session felt substantially smoother to the participant, yet blind coding found more wrong-branch episodes (5 vs 4).
That leaves a narrower open question:
What, if anything, does ongoing human-AI calibration contribute that static state representations fail to preserve?
The participant-author reports that they did not have enough programming experience to implement Claim Lifecycle independently. Yet the human-AI-tool workflow produced executable prototypes, regression/adversarial tests, reviewed repairs, frozen evaluation contracts, and machine evidence.
Approximate workflow:
This is not proof that programming expertise is unnecessary. It is an N-of-1 capability-access case showing why the human-AI pair itself became part of the research question.
Raw private transcripts are intentionally excluded from v0.3.
Release intent: criticism, literature connections, and methodological feedback — not endorsement or proof of novelty.
40 commits
Python
100.0%
An exploratory longitudinal N-of-1 study of human-AI coordination
A long-running interaction with a frontier model became much easier for one user to coordinate than fresh instances. We tested several simple explanations instead of assuming a new mechanism.
| Test | Result |
|---|---|
| Different state compression strategies | Different representations, but utility advantage not established |
| Compact state vs richer/full trajectory continuation | High quality across conditions; no useful directional advantage |
| Heavy Claim Lifecycle governance vs lightweight baseline | 7/7 vs 7/7 — NO_MEASURABLE_ADVANTAGE |
| 496-word static coordination packet vs no packet | Practical coordination benefit NOT ESTABLISHED (1/5 burden dimensions improved) |
The packet session felt substantially smoother to the participant, yet blind coding found more wrong-branch episodes (5 vs 4).
That leaves a narrower open question:
What, if anything, does ongoing human-AI calibration contribute that static state representations fail to preserve?
The participant-author reports that they did not have enough programming experience to implement Claim Lifecycle independently. Yet the human-AI-tool workflow produced executable prototypes, regression/adversarial tests, reviewed repairs, frozen evaluation contracts, and machine evidence.
Approximate workflow:
This is not proof that programming expertise is unnecessary. It is an N-of-1 capability-access case showing why the human-AI pair itself became part of the research question.
Raw private transcripts are intentionally excluded from v0.3.
Release intent: criticism, literature connections, and methodological feedback — not endorsement or proof of novelty.
40 commits
Python
100.0%