Domsect/respawn

Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.

0

5 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

Git Is Not a Backup Plan

1

Sep 30, 2026

README

Git Is Not a Backup Plan

Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.


The 3 AM thought experiment

Here's a question that ruined one of my evenings:

If GitHub, your database, and every third-party dashboard you use all vanished tonight — code, data, configs, all of it — how long would it take you to rebuild?

Not "restore from backup." Rebuild. From memory. From scratch.

I sat with that for a while, and the honest answer was: I have no idea. Weeks, maybe. Some things I'd never get back at all — not because they're technically hard, but because I no longer remember why they were set up that way, or in what order, or what the magic values were in some vendor dashboard I clicked through six months ago.

Git would save my code. Git would not save my business.

That night I started writing a file. I call the practice Respawn: a one-file checkpoint for your entire system. The file itself is the single most valuable document I own that contains zero lines of product code.

What git can't save

Your repo versions your code. It does not version:

  • Everything outside the repo. The payment provider dashboard where you hand-created six price objects. The auth provider's OAuth app config and callback URLs. The hosting platform's build command, output directory, and compatibility flags. The DNS records. The environment variables and the order they need to be set in.
  • Decisions. Why the billing ledger works the way it does. Why a certain policy exists. Why you chose constraint X over the "obvious" alternative. git log shows you what changed; it never shows you what you decided not to do and why.
  • Sequence. Disaster recovery is an ordered procedure, not a pile of facts. "Create the database, then run the schema, then configure the OAuth app pointing at the new domain, then..." — that ordering lives in someone's head, usually yours, usually at 3 AM when you need it most.
  • Business constraints. Pricing freezes. Legal red lines ("never claim X on the website"). Policies that look arbitrary but exist because of a painful lesson. These aren't code, but rebuilding without them means rebuilding wrong.

I call the gap between "the code" and "the thing that actually runs" the reconstruction gap, and for most solo founders and small teams, it's enormous. We just don't notice because nothing has exploded yet.

The file: your respawn point

A Respawn file is a single markdown document with one job: a competent engineer (or a competent AI) should be able to reconstruct the entire system from this file alone, without talking to me.

It's not documentation in the "please read our wiki" sense. It's a requirements manifest — closer to a Dockerfile for your whole operation than to a README. The positioning line at the top of mine reads:

This file is the complete implementation manifest for the current version of the system — think of it like the requirements checklist for provisioning a server. Purpose: disaster rebuilds, environment migrations, and handoffs. Feed it to an AI or an engineer and they can reproduce a functionally identical system from zero, no redesign, no re-debugging required.

Anatomy of the spec

Here's the skeleton. Steal it, adapt it, delete what doesn't apply. (Fourteen numbered sections plus §0 below; my live spec has since grown to fifteen — starters stay small, living documents grow.)

# <Project> — Respawn Spec

> Snapshot: <commit> (<date>) · Build: <status>
> Runtime: <exact versions>

## 0. How to read this
## 1. System overview & topology (ASCII diagram)
## 2. Tech stack (exact versions, not "latest")
## 3. Repo structure
## 4. Environment variables (names + purpose, NEVER values)
## 5. Database (every table, every column, every RPC, idempotency design)
## 6. Billing lifecycle (products, prices, webhook events, ledger invariants)
## 7. Auth & anti-abuse (every gate, every threshold, every policy)
## 8. API routes (complete table)
## 9. Pages / surfaces (complete table)
## 10. Brand assets (where the logos live, naming conventions)
## 11. Deployment (exact build commands, platform settings, launch checklist IN ORDER)
## 12. Disaster recovery runbook (phases 0–N, worst case first)
## 13. Known limitations & TODOs (what's held together with duct tape)
## 14. E2E acceptance matrix (the checklist you run after any rebuild)

A few notes on sections people get wrong:

§4 (env vars): list every variable's name and purpose. Never values. The spec tells you what to fill in; your password manager tells you what to fill in it with. If a variable's purpose needs more than one sentence, that's a smell — simplify the variable.

§6 (billing): this is the section most devs skip and most businesses die without. Prices, billing cycles, webhook event → system action mapping, ledger invariants ("refunds never touch X", "credits expire under condition Y"). If money moves through your system, the money rules are load-bearing documentation.

§12 (runbook): start from the worst case — everything is gone — and work forward in phases: prepare → data layer → third parties → application → verification. Each step must be executable by someone who has never seen your system. If a step says "configure the thing," rewrite it until it says exactly which dashboard, which page, which field, which value.

§13 (limitations): be brutally honest here. "The webhook signature verification has never been tested end-to-end." "This table stores a secret in plaintext; encrypt on migration." Future-you will thank present-you. This section is also where you park every "we'll fix it later" so "later" actually happens.

§14 (acceptance matrix): a numbered table of scenarios and expected outcomes. After any rebuild or migration, you run this top to bottom. It turns "I think it's working" into "I verified 30 specific behaviors."

The maintenance trick (the part everyone gets wrong)

A disaster recovery doc that isn't maintained is worse than none at all — it's confidently wrong instructions during a crisis.

The standard failure mode: you write it once, feel good, and it rots within a month. Wikis are where runbooks go to die.

Here's the system that actually works for me — two copies, two jobs:

  1. Live copy lives in the repo (docs/ENGINEERING_SPEC.md). It gets updated in the same session as any code change that affects it. Changed pricing? The spec updates before you close the laptop. New webhook event? Same session. No "I'll update the docs later" — later never comes.

  2. Respawn points live outside the repo (ENGINEERING_SPEC_6.05.16.md, 6.05.17.md, ...). Every time the system reaches a stable state, you set a respawn point — freeze a versioned snapshot. Old ones are kept, never overwritten. This is your time machine: "what did the billing model look like in March?" (Mine is at 6.05.17 as of this writing — thirteen snapshots in and still maintained, which is the entire point.)

And the real trick: make updating it someone else's job. I have a standing rule with my AI assistant — every time the project reaches a stable state, it regenerates the snapshot automatically. The cost of maintenance drops to near zero, which is the only cost at which documentation actually gets maintained. If you have CI, a linter that fails when a new env var appears without a spec entry works too. The principle is the same: the update must be cheaper than skipping it.

A recent scar that sharpened this system: I once declared a UI fix "shipped" because the source code said so — then production testing proved me wrong. The root cause was structural, not a stale deploy: a conditionally rendered panel plus tab state initialized from window.location.hash during render meant the statically prerendered HTML never contained the anchor at all, and hash visits triggered a hydration mismatch instead of a clean tab switch. My spec had been describing the source; the build artifact was the truth. My acceptance check now greps the build output, not the repo. The lesson generalizes: your respawn point must describe the thing that actually ships, not the thing you think you wrote — and "the repo says so" is never the final verification.

The test

There's exactly one acceptance test for this file, and you should run it mentally every quarter:

Hand this file — and nothing else — to a competent engineer who has never seen your project. Could they rebuild it?

Not "understand it." Rebuild it. End to end. Database, third-party configs, billing, auth, deployment. If the answer is "mostly, except they'd have to ask me about..." — every one of those exceptions is a bug in your spec. File it, fix it.

(I've never actually run this test with a human. One day I might. The AI version — feeding it to a fresh model instance and seeing what questions it asks — is a decent proxy and deeply humbling.)

Honest limitations

This isn't magic, and you should know what it doesn't do:

  • It doesn't replace backups. This file tells you how to rebuild; it doesn't contain your data. Back up your database like an adult.
  • It rots if you let it. The two-copy system plus automatic updates is load-bearing. A stale spec is a liability.
  • It doesn't capture everything. Vendor dashboards change their UI. APIs deprecate. The spec is a map, not the territory — but a slightly outdated map beats no map.
  • It's overkill for weekend projects. If you can rebuild it from memory in an afternoon, you don't need this. This is for systems where the reconstruction gap is measured in weeks.

Who should steal this

  • Solo founders. Your bus factor is 1. This is the cheapest bus-factor insurance that exists — it's a markdown file.
  • Small teams. Onboarding becomes "read the spec, then ask questions" instead of "shadow someone for two weeks and hope."
  • Anyone shipping AI-generated code. If an AI built your app and you couldn't rebuild it yourself, you don't have a codebase — you have a hostage situation. The spec file is the ransom note in reverse.
  • Agencies and freelancers. Hand the client a rebuild kit, not a zip file and a prayer.

The template

Below is the full skeleton I start from for every project. Copy it, fill it in, and never think about the 3 AM scenario again — or rather, think about it once, write it down, and then sleep well.

# <Project Name> — Respawn Spec

> **Nature:** Complete implementation manifest for the current version.
> Think: requirements checklist for provisioning a server.
> **Purpose:** Disaster rebuilds, migrations, handoffs. Feed to an AI
> or engineer → functionally identical system from zero.
> **Maintenance rule:** Any change to pricing, ledger math, webhook
> events, or schema MUST update the corresponding section same-session.
>
> - Snapshot: commit `<sha>` (<date>)
> - Build: `<command>` <status>
> - Runtime: <exact versions>

---

## 0. How to read this

- §1–3: 10-minute overview for decision makers / new owners.
- §4–10: exact reproduction details for engineers.
- §11–12: deployment & disaster runbooks (execute in order).
- §13–14: boundaries & acceptance criteria (check before launch).

---

## 1. System overview & topology

<ASCII diagram: every component, every arrow, every trust boundary.>

## 2. Tech stack

<Exact versions. Not "latest". If it matters, pin it.>

## 3. Repo structure

<Directory tree with one-line purpose per directory.>

## 4. Environment variables

| Variable | Purpose | Required | Default |
|---|---|---|---|
| ... | ... | ... | ... |

**NEVER commit values.** Names and purpose only.

## 5. Database

### 5.1 Tables
<Table: every column, type, constraint, and WHY it exists.>

### 5.2 Functions / RPCs
<Signature, behavior, atomicity guarantees.>

### 5.3 Idempotency design
<How you prevent double-processing. Key formats.>

## 6. Billing lifecycle

### 6.1 Products & prices
<Every product, price, cycle. Where each was created (dashboard/API).>

### 6.2 Webhook event map
| Event | System action | Idempotency key |
|---|---|---|

### 6.3 Ledger invariants
<The rules money obeys. What can never happen.>

## 7. Auth & anti-abuse

<Every gate: what triggers it, what it checks, what happens on failure.
Thresholds as numbers, not adjectives.>

## 8. API routes

| Route | Method | Auth | Purpose |
|---|---|---|---|

## 9. Pages / surfaces

| Route | Purpose | Auth required |
|---|---|---|

## 10. Brand assets

<Where files live, naming conventions, what NOT to use.>

## 11. Deployment

### 11.1 Build & local run
<Exact commands. What "success" looks like.>

### 11.2 Platform settings
<Build command, output dir, flags, quirks you discovered the hard way.>

### 11.3 Launch checklist (IN ORDER)
<Numbered. Each step executable by a stranger.>

## 12. Disaster recovery runbook

> Assume worst case: code, database, and all third-party configs are gone.

**Phase 0 — Prepare**
**Phase 1 — Data layer**
**Phase 2 — Third parties**
**Phase 3 — Application**
**Phase 4 — Verification (run §14)**

## 13. Known limitations & TODOs

<Brutal honesty. What's untested, what's duct-taped, what's next.>

## 14. E2E acceptance matrix

| # | Scenario | Expected |
|---|---|---|
| 1 | ... | ... |

---

*End. Last updated: <date>.*

That's it. One file. Set your respawn point. The cheapest insurance policy in software.

The question isn't whether you can afford the afternoon it takes to write it. It's whether you can afford the month it would take to rebuild without it.


If you build your own version of this, I'd genuinely like to hear about it — especially the parts where my skeleton didn't fit. The best disaster recovery system is the one that survives contact with your actual disaster.

Domsect/respawn

Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.

0

5 commits

updated Sep 30, 2026

See the code

See what people are saying

SourceMessageScoreDate

Git Is Not a Backup Plan

1

Sep 30, 2026

README

Git Is Not a Backup Plan

Respawn is a one-file checkpoint for your entire system — when everything dies, you come back whole.


The 3 AM thought experiment

Here's a question that ruined one of my evenings:

If GitHub, your database, and every third-party dashboard you use all vanished tonight — code, data, configs, all of it — how long would it take you to rebuild?

Not "restore from backup." Rebuild. From memory. From scratch.

I sat with that for a while, and the honest answer was: I have no idea. Weeks, maybe. Some things I'd never get back at all — not because they're technically hard, but because I no longer remember why they were set up that way, or in what order, or what the magic values were in some vendor dashboard I clicked through six months ago.

Git would save my code. Git would not save my business.

That night I started writing a file. I call the practice Respawn: a one-file checkpoint for your entire system. The file itself is the single most valuable document I own that contains zero lines of product code.

What git can't save

Your repo versions your code. It does not version:

  • Everything outside the repo. The payment provider dashboard where you hand-created six price objects. The auth provider's OAuth app config and callback URLs. The hosting platform's build command, output directory, and compatibility flags. The DNS records. The environment variables and the order they need to be set in.
  • Decisions. Why the billing ledger works the way it does. Why a certain policy exists. Why you chose constraint X over the "obvious" alternative. git log shows you what changed; it never shows you what you decided not to do and why.
  • Sequence. Disaster recovery is an ordered procedure, not a pile of facts. "Create the database, then run the schema, then configure the OAuth app pointing at the new domain, then..." — that ordering lives in someone's head, usually yours, usually at 3 AM when you need it most.
  • Business constraints. Pricing freezes. Legal red lines ("never claim X on the website"). Policies that look arbitrary but exist because of a painful lesson. These aren't code, but rebuilding without them means rebuilding wrong.

I call the gap between "the code" and "the thing that actually runs" the reconstruction gap, and for most solo founders and small teams, it's enormous. We just don't notice because nothing has exploded yet.

The file: your respawn point

A Respawn file is a single markdown document with one job: a competent engineer (or a competent AI) should be able to reconstruct the entire system from this file alone, without talking to me.

It's not documentation in the "please read our wiki" sense. It's a requirements manifest — closer to a Dockerfile for your whole operation than to a README. The positioning line at the top of mine reads:

This file is the complete implementation manifest for the current version of the system — think of it like the requirements checklist for provisioning a server. Purpose: disaster rebuilds, environment migrations, and handoffs. Feed it to an AI or an engineer and they can reproduce a functionally identical system from zero, no redesign, no re-debugging required.

Anatomy of the spec

Here's the skeleton. Steal it, adapt it, delete what doesn't apply. (Fourteen numbered sections plus §0 below; my live spec has since grown to fifteen — starters stay small, living documents grow.)

# <Project> — Respawn Spec

> Snapshot: <commit> (<date>) · Build: <status>
> Runtime: <exact versions>

## 0. How to read this
## 1. System overview & topology (ASCII diagram)
## 2. Tech stack (exact versions, not "latest")
## 3. Repo structure
## 4. Environment variables (names + purpose, NEVER values)
## 5. Database (every table, every column, every RPC, idempotency design)
## 6. Billing lifecycle (products, prices, webhook events, ledger invariants)
## 7. Auth & anti-abuse (every gate, every threshold, every policy)
## 8. API routes (complete table)
## 9. Pages / surfaces (complete table)
## 10. Brand assets (where the logos live, naming conventions)
## 11. Deployment (exact build commands, platform settings, launch checklist IN ORDER)
## 12. Disaster recovery runbook (phases 0–N, worst case first)
## 13. Known limitations & TODOs (what's held together with duct tape)
## 14. E2E acceptance matrix (the checklist you run after any rebuild)

A few notes on sections people get wrong:

§4 (env vars): list every variable's name and purpose. Never values. The spec tells you what to fill in; your password manager tells you what to fill in it with. If a variable's purpose needs more than one sentence, that's a smell — simplify the variable.

§6 (billing): this is the section most devs skip and most businesses die without. Prices, billing cycles, webhook event → system action mapping, ledger invariants ("refunds never touch X", "credits expire under condition Y"). If money moves through your system, the money rules are load-bearing documentation.

§12 (runbook): start from the worst case — everything is gone — and work forward in phases: prepare → data layer → third parties → application → verification. Each step must be executable by someone who has never seen your system. If a step says "configure the thing," rewrite it until it says exactly which dashboard, which page, which field, which value.

§13 (limitations): be brutally honest here. "The webhook signature verification has never been tested end-to-end." "This table stores a secret in plaintext; encrypt on migration." Future-you will thank present-you. This section is also where you park every "we'll fix it later" so "later" actually happens.

§14 (acceptance matrix): a numbered table of scenarios and expected outcomes. After any rebuild or migration, you run this top to bottom. It turns "I think it's working" into "I verified 30 specific behaviors."

The maintenance trick (the part everyone gets wrong)

A disaster recovery doc that isn't maintained is worse than none at all — it's confidently wrong instructions during a crisis.

The standard failure mode: you write it once, feel good, and it rots within a month. Wikis are where runbooks go to die.

Here's the system that actually works for me — two copies, two jobs:

  1. Live copy lives in the repo (docs/ENGINEERING_SPEC.md). It gets updated in the same session as any code change that affects it. Changed pricing? The spec updates before you close the laptop. New webhook event? Same session. No "I'll update the docs later" — later never comes.

  2. Respawn points live outside the repo (ENGINEERING_SPEC_6.05.16.md, 6.05.17.md, ...). Every time the system reaches a stable state, you set a respawn point — freeze a versioned snapshot. Old ones are kept, never overwritten. This is your time machine: "what did the billing model look like in March?" (Mine is at 6.05.17 as of this writing — thirteen snapshots in and still maintained, which is the entire point.)

And the real trick: make updating it someone else's job. I have a standing rule with my AI assistant — every time the project reaches a stable state, it regenerates the snapshot automatically. The cost of maintenance drops to near zero, which is the only cost at which documentation actually gets maintained. If you have CI, a linter that fails when a new env var appears without a spec entry works too. The principle is the same: the update must be cheaper than skipping it.

A recent scar that sharpened this system: I once declared a UI fix "shipped" because the source code said so — then production testing proved me wrong. The root cause was structural, not a stale deploy: a conditionally rendered panel plus tab state initialized from window.location.hash during render meant the statically prerendered HTML never contained the anchor at all, and hash visits triggered a hydration mismatch instead of a clean tab switch. My spec had been describing the source; the build artifact was the truth. My acceptance check now greps the build output, not the repo. The lesson generalizes: your respawn point must describe the thing that actually ships, not the thing you think you wrote — and "the repo says so" is never the final verification.

The test

There's exactly one acceptance test for this file, and you should run it mentally every quarter:

Hand this file — and nothing else — to a competent engineer who has never seen your project. Could they rebuild it?

Not "understand it." Rebuild it. End to end. Database, third-party configs, billing, auth, deployment. If the answer is "mostly, except they'd have to ask me about..." — every one of those exceptions is a bug in your spec. File it, fix it.

(I've never actually run this test with a human. One day I might. The AI version — feeding it to a fresh model instance and seeing what questions it asks — is a decent proxy and deeply humbling.)

Honest limitations

This isn't magic, and you should know what it doesn't do:

  • It doesn't replace backups. This file tells you how to rebuild; it doesn't contain your data. Back up your database like an adult.
  • It rots if you let it. The two-copy system plus automatic updates is load-bearing. A stale spec is a liability.
  • It doesn't capture everything. Vendor dashboards change their UI. APIs deprecate. The spec is a map, not the territory — but a slightly outdated map beats no map.
  • It's overkill for weekend projects. If you can rebuild it from memory in an afternoon, you don't need this. This is for systems where the reconstruction gap is measured in weeks.

Who should steal this

  • Solo founders. Your bus factor is 1. This is the cheapest bus-factor insurance that exists — it's a markdown file.
  • Small teams. Onboarding becomes "read the spec, then ask questions" instead of "shadow someone for two weeks and hope."
  • Anyone shipping AI-generated code. If an AI built your app and you couldn't rebuild it yourself, you don't have a codebase — you have a hostage situation. The spec file is the ransom note in reverse.
  • Agencies and freelancers. Hand the client a rebuild kit, not a zip file and a prayer.

The template

Below is the full skeleton I start from for every project. Copy it, fill it in, and never think about the 3 AM scenario again — or rather, think about it once, write it down, and then sleep well.

# <Project Name> — Respawn Spec

> **Nature:** Complete implementation manifest for the current version.
> Think: requirements checklist for provisioning a server.
> **Purpose:** Disaster rebuilds, migrations, handoffs. Feed to an AI
> or engineer → functionally identical system from zero.
> **Maintenance rule:** Any change to pricing, ledger math, webhook
> events, or schema MUST update the corresponding section same-session.
>
> - Snapshot: commit `<sha>` (<date>)
> - Build: `<command>` <status>
> - Runtime: <exact versions>

---

## 0. How to read this

- §1–3: 10-minute overview for decision makers / new owners.
- §4–10: exact reproduction details for engineers.
- §11–12: deployment & disaster runbooks (execute in order).
- §13–14: boundaries & acceptance criteria (check before launch).

---

## 1. System overview & topology

<ASCII diagram: every component, every arrow, every trust boundary.>

## 2. Tech stack

<Exact versions. Not "latest". If it matters, pin it.>

## 3. Repo structure

<Directory tree with one-line purpose per directory.>

## 4. Environment variables

| Variable | Purpose | Required | Default |
|---|---|---|---|
| ... | ... | ... | ... |

**NEVER commit values.** Names and purpose only.

## 5. Database

### 5.1 Tables
<Table: every column, type, constraint, and WHY it exists.>

### 5.2 Functions / RPCs
<Signature, behavior, atomicity guarantees.>

### 5.3 Idempotency design
<How you prevent double-processing. Key formats.>

## 6. Billing lifecycle

### 6.1 Products & prices
<Every product, price, cycle. Where each was created (dashboard/API).>

### 6.2 Webhook event map
| Event | System action | Idempotency key |
|---|---|---|

### 6.3 Ledger invariants
<The rules money obeys. What can never happen.>

## 7. Auth & anti-abuse

<Every gate: what triggers it, what it checks, what happens on failure.
Thresholds as numbers, not adjectives.>

## 8. API routes

| Route | Method | Auth | Purpose |
|---|---|---|---|

## 9. Pages / surfaces

| Route | Purpose | Auth required |
|---|---|---|

## 10. Brand assets

<Where files live, naming conventions, what NOT to use.>

## 11. Deployment

### 11.1 Build & local run
<Exact commands. What "success" looks like.>

### 11.2 Platform settings
<Build command, output dir, flags, quirks you discovered the hard way.>

### 11.3 Launch checklist (IN ORDER)
<Numbered. Each step executable by a stranger.>

## 12. Disaster recovery runbook

> Assume worst case: code, database, and all third-party configs are gone.

**Phase 0 — Prepare**
**Phase 1 — Data layer**
**Phase 2 — Third parties**
**Phase 3 — Application**
**Phase 4 — Verification (run §14)**

## 13. Known limitations & TODOs

<Brutal honesty. What's untested, what's duct-taped, what's next.>

## 14. E2E acceptance matrix

| # | Scenario | Expected |
|---|---|---|
| 1 | ... | ... |

---

*End. Last updated: <date>.*

That's it. One file. Set your respawn point. The cheapest insurance policy in software.

The question isn't whether you can afford the afternoon it takes to write it. It's whether you can afford the month it would take to rebuild without it.


If you build your own version of this, I'd genuinely like to hear about it — especially the parts where my skeleton didn't fit. The best disaster recovery system is the one that survives contact with your actual disaster.