Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.
0
5 commits
updated Sep 30, 2026
Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.
Here's a question that ruined one of my evenings:
If GitHub, your database, and every third-party dashboard you use all vanished tonight — code, data, configs, all of it — how long would it take you to rebuild?
Not "restore from backup." Rebuild. From memory. From scratch.
I sat with that for a while, and the honest answer was: I have no idea. Weeks, maybe. Some things I'd never get back at all — not because they're technically hard, but because I no longer remember why they were set up that way, or in what order, or what the magic values were in some vendor dashboard I clicked through six months ago.
Git would save my code. Git would not save my business.
That night I started writing a file. I call the practice Respawn: a one-file checkpoint for your entire system. The file itself is the single most valuable document I own that contains zero lines of product code.
Your repo versions your code. It does not version:
git log shows you what changed; it never shows you what you decided not to do and why.I call the gap between "the code" and "the thing that actually runs" the reconstruction gap, and for most solo founders and small teams, it's enormous. We just don't notice because nothing has exploded yet.
A Respawn file is a single markdown document with one job: a competent engineer (or a competent AI) should be able to reconstruct the entire system from this file alone, without talking to me.
It's not documentation in the "please read our wiki" sense. It's a requirements manifest — closer to a Dockerfile for your whole operation than to a README. The positioning line at the top of mine reads:
This file is the complete implementation manifest for the current version of the system — think of it like the requirements checklist for provisioning a server. Purpose: disaster rebuilds, environment migrations, and handoffs. Feed it to an AI or an engineer and they can reproduce a functionally identical system from zero, no redesign, no re-debugging required.
Here's the skeleton. Steal it, adapt it, delete what doesn't apply. (Fourteen numbered sections plus §0 below; my live spec has since grown to fifteen — starters stay small, living documents grow.)
# <Project> — Respawn Spec
> Snapshot: <commit> (<date>) · Build: <status>
> Runtime: <exact versions>
## 0. How to read this
## 1. System overview & topology (ASCII diagram)
## 2. Tech stack (exact versions, not "latest")
## 3. Repo structure
## 4. Environment variables (names + purpose, NEVER values)
## 5. Database (every table, every column, every RPC, idempotency design)
## 6. Billing lifecycle (products, prices, webhook events, ledger invariants)
## 7. Auth & anti-abuse (every gate, every threshold, every policy)
## 8. API routes (complete table)
## 9. Pages / surfaces (complete table)
## 10. Brand assets (where the logos live, naming conventions)
## 11. Deployment (exact build commands, platform settings, launch checklist IN ORDER)
## 12. Disaster recovery runbook (phases 0–N, worst case first)
## 13. Known limitations & TODOs (what's held together with duct tape)
## 14. E2E acceptance matrix (the checklist you run after any rebuild)
A few notes on sections people get wrong:
§4 (env vars): list every variable's name and purpose. Never values. The spec tells you what to fill in; your password manager tells you what to fill in it with. If a variable's purpose needs more than one sentence, that's a smell — simplify the variable.
§6 (billing): this is the section most devs skip and most businesses die without. Prices, billing cycles, webhook event → system action mapping, ledger invariants ("refunds never touch X", "credits expire under condition Y"). If money moves through your system, the money rules are load-bearing documentation.
§12 (runbook): start from the worst case — everything is gone — and work forward in phases: prepare → data layer → third parties → application → verification. Each step must be executable by someone who has never seen your system. If a step says "configure the thing," rewrite it until it says exactly which dashboard, which page, which field, which value.
§13 (limitations): be brutally honest here. "The webhook signature verification has never been tested end-to-end." "This table stores a secret in plaintext; encrypt on migration." Future-you will thank present-you. This section is also where you park every "we'll fix it later" so "later" actually happens.
§14 (acceptance matrix): a numbered table of scenarios and expected outcomes. After any rebuild or migration, you run this top to bottom. It turns "I think it's working" into "I verified 30 specific behaviors."
A disaster recovery doc that isn't maintained is worse than none at all — it's confidently wrong instructions during a crisis.
The standard failure mode: you write it once, feel good, and it rots within a month. Wikis are where runbooks go to die.
Here's the system that actually works for me — two copies, two jobs:
Live copy lives in the repo (docs/ENGINEERING_SPEC.md). It gets updated in the same session as any code change that affects it. Changed pricing? The spec updates before you close the laptop. New webhook event? Same session. No "I'll update the docs later" — later never comes.
Respawn points live outside the repo (ENGINEERING_SPEC_6.05.16.md, 6.05.17.md, ...). Every time the system reaches a stable state, you set a respawn point — freeze a versioned snapshot. Old ones are kept, never overwritten. This is your time machine: "what did the billing model look like in March?" (Mine is at 6.05.17 as of this writing — thirteen snapshots in and still maintained, which is the entire point.)
And the real trick: make updating it someone else's job. I have a standing rule with my AI assistant — every time the project reaches a stable state, it regenerates the snapshot automatically. The cost of maintenance drops to near zero, which is the only cost at which documentation actually gets maintained. If you have CI, a linter that fails when a new env var appears without a spec entry works too. The principle is the same: the update must be cheaper than skipping it.
A recent scar that sharpened this system: I once declared a UI fix "shipped" because the source code said so — then production testing proved me wrong. The root cause was structural, not a stale deploy: a conditionally rendered panel plus tab state initialized from window.location.hash during render meant the statically prerendered HTML never contained the anchor at all, and hash visits triggered a hydration mismatch instead of a clean tab switch. My spec had been describing the source; the build artifact was the truth. My acceptance check now greps the build output, not the repo. The lesson generalizes: your respawn point must describe the thing that actually ships, not the thing you think you wrote — and "the repo says so" is never the final verification.
There's exactly one acceptance test for this file, and you should run it mentally every quarter:
Hand this file — and nothing else — to a competent engineer who has never seen your project. Could they rebuild it?
Not "understand it." Rebuild it. End to end. Database, third-party configs, billing, auth, deployment. If the answer is "mostly, except they'd have to ask me about..." — every one of those exceptions is a bug in your spec. File it, fix it.
(I've never actually run this test with a human. One day I might. The AI version — feeding it to a fresh model instance and seeing what questions it asks — is a decent proxy and deeply humbling.)
This isn't magic, and you should know what it doesn't do:
Below is the full skeleton I start from for every project. Copy it, fill it in, and never think about the 3 AM scenario again — or rather, think about it once, write it down, and then sleep well.
# <Project Name> — Respawn Spec
> **Nature:** Complete implementation manifest for the current version.
> Think: requirements checklist for provisioning a server.
> **Purpose:** Disaster rebuilds, migrations, handoffs. Feed to an AI
> or engineer → functionally identical system from zero.
> **Maintenance rule:** Any change to pricing, ledger math, webhook
> events, or schema MUST update the corresponding section same-session.
>
> - Snapshot: commit `<sha>` (<date>)
> - Build: `<command>` <status>
> - Runtime: <exact versions>
---
## 0. How to read this
- §1–3: 10-minute overview for decision makers / new owners.
- §4–10: exact reproduction details for engineers.
- §11–12: deployment & disaster runbooks (execute in order).
- §13–14: boundaries & acceptance criteria (check before launch).
---
## 1. System overview & topology
<ASCII diagram: every component, every arrow, every trust boundary.>
## 2. Tech stack
<Exact versions. Not "latest". If it matters, pin it.>
## 3. Repo structure
<Directory tree with one-line purpose per directory.>
## 4. Environment variables
| Variable | Purpose | Required | Default |
|---|---|---|---|
| ... | ... | ... | ... |
**NEVER commit values.** Names and purpose only.
## 5. Database
### 5.1 Tables
<Table: every column, type, constraint, and WHY it exists.>
### 5.2 Functions / RPCs
<Signature, behavior, atomicity guarantees.>
### 5.3 Idempotency design
<How you prevent double-processing. Key formats.>
## 6. Billing lifecycle
### 6.1 Products & prices
<Every product, price, cycle. Where each was created (dashboard/API).>
### 6.2 Webhook event map
| Event | System action | Idempotency key |
|---|---|---|
### 6.3 Ledger invariants
<The rules money obeys. What can never happen.>
## 7. Auth & anti-abuse
<Every gate: what triggers it, what it checks, what happens on failure.
Thresholds as numbers, not adjectives.>
## 8. API routes
| Route | Method | Auth | Purpose |
|---|---|---|---|
## 9. Pages / surfaces
| Route | Purpose | Auth required |
|---|---|---|
## 10. Brand assets
<Where files live, naming conventions, what NOT to use.>
## 11. Deployment
### 11.1 Build & local run
<Exact commands. What "success" looks like.>
### 11.2 Platform settings
<Build command, output dir, flags, quirks you discovered the hard way.>
### 11.3 Launch checklist (IN ORDER)
<Numbered. Each step executable by a stranger.>
## 12. Disaster recovery runbook
> Assume worst case: code, database, and all third-party configs are gone.
**Phase 0 — Prepare**
**Phase 1 — Data layer**
**Phase 2 — Third parties**
**Phase 3 — Application**
**Phase 4 — Verification (run §14)**
## 13. Known limitations & TODOs
<Brutal honesty. What's untested, what's duct-taped, what's next.>
## 14. E2E acceptance matrix
| # | Scenario | Expected |
|---|---|---|
| 1 | ... | ... |
---
*End. Last updated: <date>.*
That's it. One file. Set your respawn point. The cheapest insurance policy in software.
The question isn't whether you can afford the afternoon it takes to write it. It's whether you can afford the month it would take to rebuild without it.
If you build your own version of this, I'd genuinely like to hear about it — especially the parts where my skeleton didn't fit. The best disaster recovery system is the one that survives contact with your actual disaster.
Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.
0
5 commits
updated Sep 30, 2026
Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.
Here's a question that ruined one of my evenings:
If GitHub, your database, and every third-party dashboard you use all vanished tonight — code, data, configs, all of it — how long would it take you to rebuild?
Not "restore from backup." Rebuild. From memory. From scratch.
I sat with that for a while, and the honest answer was: I have no idea. Weeks, maybe. Some things I'd never get back at all — not because they're technically hard, but because I no longer remember why they were set up that way, or in what order, or what the magic values were in some vendor dashboard I clicked through six months ago.
Git would save my code. Git would not save my business.
That night I started writing a file. I call the practice Respawn: a one-file checkpoint for your entire system. The file itself is the single most valuable document I own that contains zero lines of product code.
Your repo versions your code. It does not version:
git log shows you what changed; it never shows you what you decided not to do and why.I call the gap between "the code" and "the thing that actually runs" the reconstruction gap, and for most solo founders and small teams, it's enormous. We just don't notice because nothing has exploded yet.
A Respawn file is a single markdown document with one job: a competent engineer (or a competent AI) should be able to reconstruct the entire system from this file alone, without talking to me.
It's not documentation in the "please read our wiki" sense. It's a requirements manifest — closer to a Dockerfile for your whole operation than to a README. The positioning line at the top of mine reads:
This file is the complete implementation manifest for the current version of the system — think of it like the requirements checklist for provisioning a server. Purpose: disaster rebuilds, environment migrations, and handoffs. Feed it to an AI or an engineer and they can reproduce a functionally identical system from zero, no redesign, no re-debugging required.
Here's the skeleton. Steal it, adapt it, delete what doesn't apply. (Fourteen numbered sections plus §0 below; my live spec has since grown to fifteen — starters stay small, living documents grow.)
# <Project> — Respawn Spec
> Snapshot: <commit> (<date>) · Build: <status>
> Runtime: <exact versions>
## 0. How to read this
## 1. System overview & topology (ASCII diagram)
## 2. Tech stack (exact versions, not "latest")
## 3. Repo structure
## 4. Environment variables (names + purpose, NEVER values)
## 5. Database (every table, every column, every RPC, idempotency design)
## 6. Billing lifecycle (products, prices, webhook events, ledger invariants)
## 7. Auth & anti-abuse (every gate, every threshold, every policy)
## 8. API routes (complete table)
## 9. Pages / surfaces (complete table)
## 10. Brand assets (where the logos live, naming conventions)
## 11. Deployment (exact build commands, platform settings, launch checklist IN ORDER)
## 12. Disaster recovery runbook (phases 0–N, worst case first)
## 13. Known limitations & TODOs (what's held together with duct tape)
## 14. E2E acceptance matrix (the checklist you run after any rebuild)
A few notes on sections people get wrong:
§4 (env vars): list every variable's name and purpose. Never values. The spec tells you what to fill in; your password manager tells you what to fill in it with. If a variable's purpose needs more than one sentence, that's a smell — simplify the variable.
§6 (billing): this is the section most devs skip and most businesses die without. Prices, billing cycles, webhook event → system action mapping, ledger invariants ("refunds never touch X", "credits expire under condition Y"). If money moves through your system, the money rules are load-bearing documentation.
§12 (runbook): start from the worst case — everything is gone — and work forward in phases: prepare → data layer → third parties → application → verification. Each step must be executable by someone who has never seen your system. If a step says "configure the thing," rewrite it until it says exactly which dashboard, which page, which field, which value.
§13 (limitations): be brutally honest here. "The webhook signature verification has never been tested end-to-end." "This table stores a secret in plaintext; encrypt on migration." Future-you will thank present-you. This section is also where you park every "we'll fix it later" so "later" actually happens.
§14 (acceptance matrix): a numbered table of scenarios and expected outcomes. After any rebuild or migration, you run this top to bottom. It turns "I think it's working" into "I verified 30 specific behaviors."
A disaster recovery doc that isn't maintained is worse than none at all — it's confidently wrong instructions during a crisis.
The standard failure mode: you write it once, feel good, and it rots within a month. Wikis are where runbooks go to die.
Here's the system that actually works for me — two copies, two jobs:
Live copy lives in the repo (docs/ENGINEERING_SPEC.md). It gets updated in the same session as any code change that affects it. Changed pricing? The spec updates before you close the laptop. New webhook event? Same session. No "I'll update the docs later" — later never comes.
Respawn points live outside the repo (ENGINEERING_SPEC_6.05.16.md, 6.05.17.md, ...). Every time the system reaches a stable state, you set a respawn point — freeze a versioned snapshot. Old ones are kept, never overwritten. This is your time machine: "what did the billing model look like in March?" (Mine is at 6.05.17 as of this writing — thirteen snapshots in and still maintained, which is the entire point.)
And the real trick: make updating it someone else's job. I have a standing rule with my AI assistant — every time the project reaches a stable state, it regenerates the snapshot automatically. The cost of maintenance drops to near zero, which is the only cost at which documentation actually gets maintained. If you have CI, a linter that fails when a new env var appears without a spec entry works too. The principle is the same: the update must be cheaper than skipping it.
A recent scar that sharpened this system: I once declared a UI fix "shipped" because the source code said so — then production testing proved me wrong. The root cause was structural, not a stale deploy: a conditionally rendered panel plus tab state initialized from window.location.hash during render meant the statically prerendered HTML never contained the anchor at all, and hash visits triggered a hydration mismatch instead of a clean tab switch. My spec had been describing the source; the build artifact was the truth. My acceptance check now greps the build output, not the repo. The lesson generalizes: your respawn point must describe the thing that actually ships, not the thing you think you wrote — and "the repo says so" is never the final verification.
There's exactly one acceptance test for this file, and you should run it mentally every quarter:
Hand this file — and nothing else — to a competent engineer who has never seen your project. Could they rebuild it?
Not "understand it." Rebuild it. End to end. Database, third-party configs, billing, auth, deployment. If the answer is "mostly, except they'd have to ask me about..." — every one of those exceptions is a bug in your spec. File it, fix it.
(I've never actually run this test with a human. One day I might. The AI version — feeding it to a fresh model instance and seeing what questions it asks — is a decent proxy and deeply humbling.)
This isn't magic, and you should know what it doesn't do:
Below is the full skeleton I start from for every project. Copy it, fill it in, and never think about the 3 AM scenario again — or rather, think about it once, write it down, and then sleep well.
# <Project Name> — Respawn Spec
> **Nature:** Complete implementation manifest for the current version.
> Think: requirements checklist for provisioning a server.
> **Purpose:** Disaster rebuilds, migrations, handoffs. Feed to an AI
> or engineer → functionally identical system from zero.
> **Maintenance rule:** Any change to pricing, ledger math, webhook
> events, or schema MUST update the corresponding section same-session.
>
> - Snapshot: commit `<sha>` (<date>)
> - Build: `<command>` <status>
> - Runtime: <exact versions>
---
## 0. How to read this
- §1–3: 10-minute overview for decision makers / new owners.
- §4–10: exact reproduction details for engineers.
- §11–12: deployment & disaster runbooks (execute in order).
- §13–14: boundaries & acceptance criteria (check before launch).
---
## 1. System overview & topology
<ASCII diagram: every component, every arrow, every trust boundary.>
## 2. Tech stack
<Exact versions. Not "latest". If it matters, pin it.>
## 3. Repo structure
<Directory tree with one-line purpose per directory.>
## 4. Environment variables
| Variable | Purpose | Required | Default |
|---|---|---|---|
| ... | ... | ... | ... |
**NEVER commit values.** Names and purpose only.
## 5. Database
### 5.1 Tables
<Table: every column, type, constraint, and WHY it exists.>
### 5.2 Functions / RPCs
<Signature, behavior, atomicity guarantees.>
### 5.3 Idempotency design
<How you prevent double-processing. Key formats.>
## 6. Billing lifecycle
### 6.1 Products & prices
<Every product, price, cycle. Where each was created (dashboard/API).>
### 6.2 Webhook event map
| Event | System action | Idempotency key |
|---|---|---|
### 6.3 Ledger invariants
<The rules money obeys. What can never happen.>
## 7. Auth & anti-abuse
<Every gate: what triggers it, what it checks, what happens on failure.
Thresholds as numbers, not adjectives.>
## 8. API routes
| Route | Method | Auth | Purpose |
|---|---|---|---|
## 9. Pages / surfaces
| Route | Purpose | Auth required |
|---|---|---|
## 10. Brand assets
<Where files live, naming conventions, what NOT to use.>
## 11. Deployment
### 11.1 Build & local run
<Exact commands. What "success" looks like.>
### 11.2 Platform settings
<Build command, output dir, flags, quirks you discovered the hard way.>
### 11.3 Launch checklist (IN ORDER)
<Numbered. Each step executable by a stranger.>
## 12. Disaster recovery runbook
> Assume worst case: code, database, and all third-party configs are gone.
**Phase 0 — Prepare**
**Phase 1 — Data layer**
**Phase 2 — Third parties**
**Phase 3 — Application**
**Phase 4 — Verification (run §14)**
## 13. Known limitations & TODOs
<Brutal honesty. What's untested, what's duct-taped, what's next.>
## 14. E2E acceptance matrix
| # | Scenario | Expected |
|---|---|---|
| 1 | ... | ... |
---
*End. Last updated: <date>.*
That's it. One file. Set your respawn point. The cheapest insurance policy in software.
The question isn't whether you can afford the afternoon it takes to write it. It's whether you can afford the month it would take to rebuild without it.
If you build your own version of this, I'd genuinely like to hear about it — especially the parts where my skeleton didn't fit. The best disaster recovery system is the one that survives contact with your actual disaster.