Product demo videos from a script. Write the flow once as Playwright steps with a line of narration on each; DemoForge drives the app, records it, zooms in on every click, draws a smooth cursor, speaks the narration and exports an MP4. When the UI changes, run the same command again.

Needs Node, pnpm and ffmpeg. espeak-ng (dnf install espeak-ng /
apt install espeak-ng) gives you a free local voice; no API key needed.
pnpm install
pnpm -C packages/core build
pnpm exec playwright-core install chromium
node scripts/demoforge.mjs record demos/todomvc/flow.json # drive the app, record
node scripts/demoforge.mjs open demos/todomvc # review in the editor
node scripts/demoforge.mjs export demos/todomvc # -> demos/todomvc/demo.mp4
demos/todomvc/flow.json records Playwright's public TodoMVC demo — copy it
and point url at your own app.
packages/core @demoforge/core — types, zoom planner, evaluator, cursor path
apps/editor Vite + React + Tailwind — player, timeline, compositor, export
apps/extension MV3 Chrome extension — capture by hand + click log
scripts/ demoforge.mjs — the record / open / export CLI
pnpm -r test runs the tests. demoforge-docs/ has the product vision and
architecture notes.
Narration needs a voice provider; nothing else does, and the editor tells you which ones are available.
espeak-ng (local) — free, offline, robotic.
Gemini AI — copy .env.example to .env and put a key in
GEMINI_API_KEY. The same key writes the script. Sounds like a person. .env at the repo root or in
apps/editor/ both work; the key is read by the dev server only and never
reaches the browser.
A free-tier key allows only a few requests a minute, so generating a long script pauses when the quota says to and picks up again — the button tells you how long it is waiting.
ElevenLabs — ELEVENLABS_API_KEY, same two locations. The best voices.
Metered per character: the free tier is 10,000 characters a month, personal
use only, and asks you to credit ElevenLabs. The panel shows what is left
and what the next generate will cost, and warns before a run that would run
out partway.
Mistral Voxtral — MISTRAL_API_KEY, same two locations.
pnpm -C apps/extension build, load apps/extension/dist
unpacked in Chrome, open any http(s) page, click the DemoForge action →
Start. It captures the tab with chrome.tabCapture and logs every click,
input, scroll and navigation against the recorder's own clock.~/Downloads/demoforge/<timestamp>/:
recording.webm and demo.json (a DemoRecording).pnpm -C apps/editor dev and drop both into the editor. Zooms are
planned from the click log; drag the pills to move or resize them, add and
delete, restyle the frame.apps/editor/vite-export.ts): the recording is decoded straight
through, each frame drawn by the editor's own compose() on a Skia canvas,
and piped into x264 — an 86 s demo in about 2 minutes. Without a server it
falls back to ffmpeg.wasm in the browser, several times slower.Tell Claude Code "record a demo of /login" and it does the rest. The
demoforge skill (.claude/skills/demoforge/, symlink it into
~/.claude/skills/ to use it from any repo) has the agent read the page,
write demos/<name>/flow.json — Playwright actions with a say line of
narration on each — and run:
node scripts/demoforge.mjs record demos/<name>/flow.json # video + demo.json + project with script
node scripts/demoforge.mjs steps demos/<name> # frames the agent checks
node scripts/demoforge.mjs open demos/<name> # editor at ?demo=<name>
The recorder paces itself to the narration: each line starts a beat before
its action and the step holds until the line has been said. setup steps
(signing in) run off camera. Capture is 2× device pixels, encoded as VP9 from the
screencast's own frames — Playwright's built-in recorder is capped at 1 Mbps. ${NAME} in a flow is read from the environment
or .env, so credentials stay out of the file. The flow is the only file in a
take that is committed — re-recording after a UI change is the same command.
An icon rail on the right opens six panels:
| Panel | What it does |
|---|---|
| Script & voice | Narration lines on the timeline, drafted from the click log or written by hand, spoken by a local TTS |
| Background | Image / Colour / Gradient tabs — 18 generated wallpapers, custom upload, gradient presets with editable stops and angle |
| Zoom | Auto-zoom toggle, per-zoom or global scale, re-plan from the click log |
| Captions | Add at the playhead, draft a set from the click log, edit text, position / size / colour |
| Effects | Padding, corner radius, shadow blur / offset / strength |
| Layout | Output aspect — Original, 16:9, 9:16, 1:1, 4:3, 4:5 |
| Cursor | Show, click pulse, size, smoothing |
Aiming a zoom. Select a pill and the preview drops back to the unzoomed frame with a rectangle showing exactly what that zoom will crop. Click or drag anywhere on the frame to move it, and a floating inspector gives you the zoom level, focus mode, reset and delete.
A zoom's focus is auto by default: it points at the nearest click, and re-aims itself if you drag the pill somewhere else on the timeline. Placing a point by hand switches it to manual, and nothing moves it again until you reset it.
Captions are drawn over the frame but outside the zoom transform — a caption belongs to the viewer, not to the picture, so it does not slide or grow when the camera moves. From clicks drafts one cue per click out of the element text the extension already recorded; that is string formatting, not AI.
Cutting. Press T to drop a cut at the playhead, then drag its edges;
I and O cut everything before or after the playhead. Cut spans are shaded
across every lane, playback jumps over them, and they are gone from the export
— video and audio both.
Cuts are the one place timeline time and source time come apart. The edit list
in packages/core/src/edits.ts owns that conversion, and the rule that keeps
it cheap is that everything else stays in source time: zoom keyframes,
captions, the cursor path and the event log are never remapped. The exporter
walks edited time, maps each frame back through editedToSource(), and
composites at a source timestamp exactly as the preview does.
The timeline has a scrubbable ruler with amber marks at every logged click, a cut lane, a zoom lane, a caption lane, and a clip lane. Ctrl+Scroll zooms the view about the pointer, Shift+Scroll pans, and the window follows the playhead.
| Key | |
|---|---|
Space | play / pause |
Z | add a zoom at the playhead |
C | add a caption at the playhead |
N | add a narration line at the playhead |
S | save the project |
T | cut a section out at the playhead |
I O | cut everything before / after the playhead |
Delete | remove the selection |
Esc | deselect |
← → | step one frame (hold Shift for a second) |
Home End | jump to start / end |
Wallpapers are generated, not shipped — a base colour plus soft radial blobs, painted by one function used for both the picker swatch and the full frame, so the swatch cannot lie and there are no binary assets in the repo.
A script is a list of lines, each anchored to a source timestamp on the same clock as everything else. There are two ways to get one.
Write the script with AI is the good one. The editor breaks the demo into steps — an opening, then one per click — grabs a frame of the screen at each, rings the spot that was clicked, and sends the lot to a model along with how many seconds it has to talk at each step. It writes to that budget, naming what is actually on screen. Give it a sentence about what the demo is for and it gets markedly better; that brief is saved with the project.
Two things the model is deliberately not trusted with:
Frames of your recording go to Google when you press it. Nothing else in the editor sends anything anywhere.
From clicks is the offline fallback: one line per step from the element text already in the log, string templates, no network. Rough, but instant.
Generate voiceover then speaks every line that has changed, measures how long it actually took, and lays the results onto one track.
The mixdown plays in the preview (the captured tab audio stays muted there) and is muxed into the export with the recording ducked underneath the voice.
A few things follow from how it is wired:
<name>.dfp.json and, when there is a voiceover,
<name>.narration.wav. Drop both back in with the video and the narration
plays immediately — nothing is respoken, which matters when the voice is
metered. The JSON stays a few readable kilobytes rather than megabytes of
base64, and the .wav is an ordinary file you can listen to or edit
elsewhere..wav and everything still works — it just costs a full
regenerate.est.. If lines start
talking over each other the panel says so and offers to space them out.GET /api/tts says who can speak
and with which voices; POST /api/tts returns WAV
(apps/editor/vite-tts.ts). The panel renders whatever the server reports,
so adding a provider is a server-side change. The mixdown, the timeline and
the exporter never learn who spoke.pcm_24000 — so the WAV header is
written server-side before the audio ever reaches the browser. (44.1 kHz
from ElevenLabs needs a Pro subscription; 24 kHz does not.)Save project (or S) writes <name>.dfp.json — the recording plus every
edit (zooms, captions, cuts, the narration script and the style) as plain
readable JSON — and <name>.narration.wav alongside it when there is a
voiceover. Drop the JSON, the video, and the .wav back in to carry on
exactly where you were. A raw demo.json still opens too; it just gets
freshly planned zooms.
The schema lives in packages/core/src/project.ts, not in the editor, because
the point is that the editor is not the only thing that can write one. A
script, a CI job, or Claude Code can open a project, change the zooms or
captions, write it back, and the editor will render exactly that.
{
"format": "demoforge-project",
"version": 1,
"mediaName": "recording.webm", // referenced, not embedded
"narrationName": "recording.narration.wav", // ditto; "" when there is none
"brief": "AirSense is an air-quality dashboard for facilities teams.",
"recording": { /* the DemoRecording from capture */ },
"zooms": [
{ "tStart": 1500, "tEnd": 3700, "targetXNorm": 0.42, "targetYNorm": 0.31,
"scale": 1.8, "easing": "easeInOutCubic", "focus": "auto" }
],
"captions": [
{ "tStart": 2000, "tEnd": 4200, "text": "Click \"Add Widget\"" }
],
"script": [ // narration; text only, audio is regenerated
{ "tStart": 800, "text": "Start by clicking Add Widget.", "audioMs": 2100 }
],
"style": { "background": { "kind": "wallpaper", "id": "cobalt" }, "aspect": null,
"padding": 0.05, "radius": 0.02, "shadow": { "blur": 0.05, "y": 0.018, "alpha": 0.5 },
"cursor": { "show": true, "size": 0.045, "smoothing": 0.4, "clicks": true },
"captions": { "size": 0.045, "position": "bottom", "color": "#ffffff",
"background": "rgba(2,6,23,0.72)" },
"voice": { "provider": "local", "voice": "en-us+f3", "rate": 170,
"direction": "", "gain": 1, "duck": 0.25 } }
}
Notes for anything editing one by hand:
zooms and captions are time-ordered, so "the third zoom" is stable.
Neither list may overlap itself.focus: "auto" means the zoom is aimed at the nearest click and will re-aim
if moved; "manual" pins it.parseProject() is a trust boundary: it sorts and de-overlaps the lists,
clamps every number into range, drops zero-length spans, and falls back to
defaults rather than letting NaN reach the renderer. It throws only on a
missing recording or a format version it does not understand — so a
roughly-right file loads rather than failing.cuts are spans of source video the demo skips, in source time. They are
sorted, clamped, merged when they overlap, and dropped when shorter than
100 ms. A legacy trim: {startMs, endMs} (a span to keep) is migrated
into the equivalent head and tail cuts.script lines carry audioMs only as a cached measurement. Change text
and drop it — a stale length lays the timeline out for audio that no longer
exists. Lines may overlap; that is reported, not prevented.style.voice.duck is what the captured recording drops to while the voice
is talking, gain is the voice's own level.brief is what the demo is about, in your words. It is context for whoever
writes the narration — the AI writer reads it — and it is worth keeping so
the next rewrite starts from the same understanding.narrationName names the rendered voiceover sitting next to the project.
The editor takes any dropped .wav as the narration, so the name is a hint
rather than a requirement — files get renamed.style.voice.provider is "local", "gemini" or "elevenlabs"; voice is that
voice is that provider's own id (en-us+f3, Iapetus, or an ElevenLabs
voice id). rate is used by the local provider, direction by Gemini —
both are always stored, so switching provider and back keeps your settings.zooms, captions, cuts, script or style entirely is fine;
they default.Everything flows through one type, DemoRecording, defined once in
packages/core. The rules that make Phase 1's editor reusable unchanged in
Phase 2 (AI voiceover) and Phase 3 (Playwright re-record):
source. Not the planner, editor, compositor or
exporter. That branch is the seam that would break Phase 3.render/geometry.ts, against the video's real decoded size.t0 is stamped at MediaRecorder.start(); every
DemoEvent.t is ms since it. The editor reconciles against the decoded
duration on import and warns on drift over 100 ms../scripts/make-fixture.sh # markers at exactly known coordinates
pnpm -C apps/editor dev # drop fixture/demo.json + recording.webm
Every zoom must land dead centre on its marker; the double-click at 5.0/5.12 s
must produce one zoom, and the giant panel at 13 s none. pnpm -r test asserts
all of that headlessly.
For recording-side timing, see apps/extension/README.md.
pnpm dev; the fix is the same
backend the Phase 2 TODO already calls for.Product demo videos from a script. Write the flow once as Playwright steps with a line of narration on each; DemoForge drives the app, records it, zooms in on every click, draws a smooth cursor, speaks the narration and exports an MP4. When the UI changes, run the same command again.

Needs Node, pnpm and ffmpeg. espeak-ng (dnf install espeak-ng /
apt install espeak-ng) gives you a free local voice; no API key needed.
pnpm install
pnpm -C packages/core build
pnpm exec playwright-core install chromium
node scripts/demoforge.mjs record demos/todomvc/flow.json # drive the app, record
node scripts/demoforge.mjs open demos/todomvc # review in the editor
node scripts/demoforge.mjs export demos/todomvc # -> demos/todomvc/demo.mp4
demos/todomvc/flow.json records Playwright's public TodoMVC demo — copy it
and point url at your own app.
packages/core @demoforge/core — types, zoom planner, evaluator, cursor path
apps/editor Vite + React + Tailwind — player, timeline, compositor, export
apps/extension MV3 Chrome extension — capture by hand + click log
scripts/ demoforge.mjs — the record / open / export CLI
pnpm -r test runs the tests. demoforge-docs/ has the product vision and
architecture notes.
Narration needs a voice provider; nothing else does, and the editor tells you which ones are available.
espeak-ng (local) — free, offline, robotic.
Gemini AI — copy .env.example to .env and put a key in
GEMINI_API_KEY. The same key writes the script. Sounds like a person. .env at the repo root or in
apps/editor/ both work; the key is read by the dev server only and never
reaches the browser.
A free-tier key allows only a few requests a minute, so generating a long script pauses when the quota says to and picks up again — the button tells you how long it is waiting.
ElevenLabs — ELEVENLABS_API_KEY, same two locations. The best voices.
Metered per character: the free tier is 10,000 characters a month, personal
use only, and asks you to credit ElevenLabs. The panel shows what is left
and what the next generate will cost, and warns before a run that would run
out partway.
Mistral Voxtral — MISTRAL_API_KEY, same two locations.
pnpm -C apps/extension build, load apps/extension/dist
unpacked in Chrome, open any http(s) page, click the DemoForge action →
Start. It captures the tab with chrome.tabCapture and logs every click,
input, scroll and navigation against the recorder's own clock.~/Downloads/demoforge/<timestamp>/:
recording.webm and demo.json (a DemoRecording).pnpm -C apps/editor dev and drop both into the editor. Zooms are
planned from the click log; drag the pills to move or resize them, add and
delete, restyle the frame.apps/editor/vite-export.ts): the recording is decoded straight
through, each frame drawn by the editor's own compose() on a Skia canvas,
and piped into x264 — an 86 s demo in about 2 minutes. Without a server it
falls back to ffmpeg.wasm in the browser, several times slower.Tell Claude Code "record a demo of /login" and it does the rest. The
demoforge skill (.claude/skills/demoforge/, symlink it into
~/.claude/skills/ to use it from any repo) has the agent read the page,
write demos/<name>/flow.json — Playwright actions with a say line of
narration on each — and run:
node scripts/demoforge.mjs record demos/<name>/flow.json # video + demo.json + project with script
node scripts/demoforge.mjs steps demos/<name> # frames the agent checks
node scripts/demoforge.mjs open demos/<name> # editor at ?demo=<name>
The recorder paces itself to the narration: each line starts a beat before
its action and the step holds until the line has been said. setup steps
(signing in) run off camera. Capture is 2× device pixels, encoded as VP9 from the
screencast's own frames — Playwright's built-in recorder is capped at 1 Mbps. ${NAME} in a flow is read from the environment
or .env, so credentials stay out of the file. The flow is the only file in a
take that is committed — re-recording after a UI change is the same command.
An icon rail on the right opens six panels:
| Panel | What it does |
|---|---|
| Script & voice | Narration lines on the timeline, drafted from the click log or written by hand, spoken by a local TTS |
| Background | Image / Colour / Gradient tabs — 18 generated wallpapers, custom upload, gradient presets with editable stops and angle |
| Zoom | Auto-zoom toggle, per-zoom or global scale, re-plan from the click log |
| Captions | Add at the playhead, draft a set from the click log, edit text, position / size / colour |
| Effects | Padding, corner radius, shadow blur / offset / strength |
| Layout | Output aspect — Original, 16:9, 9:16, 1:1, 4:3, 4:5 |
| Cursor | Show, click pulse, size, smoothing |
Aiming a zoom. Select a pill and the preview drops back to the unzoomed frame with a rectangle showing exactly what that zoom will crop. Click or drag anywhere on the frame to move it, and a floating inspector gives you the zoom level, focus mode, reset and delete.
A zoom's focus is auto by default: it points at the nearest click, and re-aims itself if you drag the pill somewhere else on the timeline. Placing a point by hand switches it to manual, and nothing moves it again until you reset it.
Captions are drawn over the frame but outside the zoom transform — a caption belongs to the viewer, not to the picture, so it does not slide or grow when the camera moves. From clicks drafts one cue per click out of the element text the extension already recorded; that is string formatting, not AI.
Cutting. Press T to drop a cut at the playhead, then drag its edges;
I and O cut everything before or after the playhead. Cut spans are shaded
across every lane, playback jumps over them, and they are gone from the export
— video and audio both.
Cuts are the one place timeline time and source time come apart. The edit list
in packages/core/src/edits.ts owns that conversion, and the rule that keeps
it cheap is that everything else stays in source time: zoom keyframes,
captions, the cursor path and the event log are never remapped. The exporter
walks edited time, maps each frame back through editedToSource(), and
composites at a source timestamp exactly as the preview does.
The timeline has a scrubbable ruler with amber marks at every logged click, a cut lane, a zoom lane, a caption lane, and a clip lane. Ctrl+Scroll zooms the view about the pointer, Shift+Scroll pans, and the window follows the playhead.
| Key | |
|---|---|
Space | play / pause |
Z | add a zoom at the playhead |
C | add a caption at the playhead |
N | add a narration line at the playhead |
S | save the project |
T | cut a section out at the playhead |
I O | cut everything before / after the playhead |
Delete | remove the selection |
Esc | deselect |
← → | step one frame (hold Shift for a second) |
Home End | jump to start / end |
Wallpapers are generated, not shipped — a base colour plus soft radial blobs, painted by one function used for both the picker swatch and the full frame, so the swatch cannot lie and there are no binary assets in the repo.
A script is a list of lines, each anchored to a source timestamp on the same clock as everything else. There are two ways to get one.
Write the script with AI is the good one. The editor breaks the demo into steps — an opening, then one per click — grabs a frame of the screen at each, rings the spot that was clicked, and sends the lot to a model along with how many seconds it has to talk at each step. It writes to that budget, naming what is actually on screen. Give it a sentence about what the demo is for and it gets markedly better; that brief is saved with the project.
Two things the model is deliberately not trusted with:
Frames of your recording go to Google when you press it. Nothing else in the editor sends anything anywhere.
From clicks is the offline fallback: one line per step from the element text already in the log, string templates, no network. Rough, but instant.
Generate voiceover then speaks every line that has changed, measures how long it actually took, and lays the results onto one track.
The mixdown plays in the preview (the captured tab audio stays muted there) and is muxed into the export with the recording ducked underneath the voice.
A few things follow from how it is wired:
<name>.dfp.json and, when there is a voiceover,
<name>.narration.wav. Drop both back in with the video and the narration
plays immediately — nothing is respoken, which matters when the voice is
metered. The JSON stays a few readable kilobytes rather than megabytes of
base64, and the .wav is an ordinary file you can listen to or edit
elsewhere..wav and everything still works — it just costs a full
regenerate.est.. If lines start
talking over each other the panel says so and offers to space them out.GET /api/tts says who can speak
and with which voices; POST /api/tts returns WAV
(apps/editor/vite-tts.ts). The panel renders whatever the server reports,
so adding a provider is a server-side change. The mixdown, the timeline and
the exporter never learn who spoke.pcm_24000 — so the WAV header is
written server-side before the audio ever reaches the browser. (44.1 kHz
from ElevenLabs needs a Pro subscription; 24 kHz does not.)Save project (or S) writes <name>.dfp.json — the recording plus every
edit (zooms, captions, cuts, the narration script and the style) as plain
readable JSON — and <name>.narration.wav alongside it when there is a
voiceover. Drop the JSON, the video, and the .wav back in to carry on
exactly where you were. A raw demo.json still opens too; it just gets
freshly planned zooms.
The schema lives in packages/core/src/project.ts, not in the editor, because
the point is that the editor is not the only thing that can write one. A
script, a CI job, or Claude Code can open a project, change the zooms or
captions, write it back, and the editor will render exactly that.
{
"format": "demoforge-project",
"version": 1,
"mediaName": "recording.webm", // referenced, not embedded
"narrationName": "recording.narration.wav", // ditto; "" when there is none
"brief": "AirSense is an air-quality dashboard for facilities teams.",
"recording": { /* the DemoRecording from capture */ },
"zooms": [
{ "tStart": 1500, "tEnd": 3700, "targetXNorm": 0.42, "targetYNorm": 0.31,
"scale": 1.8, "easing": "easeInOutCubic", "focus": "auto" }
],
"captions": [
{ "tStart": 2000, "tEnd": 4200, "text": "Click \"Add Widget\"" }
],
"script": [ // narration; text only, audio is regenerated
{ "tStart": 800, "text": "Start by clicking Add Widget.", "audioMs": 2100 }
],
"style": { "background": { "kind": "wallpaper", "id": "cobalt" }, "aspect": null,
"padding": 0.05, "radius": 0.02, "shadow": { "blur": 0.05, "y": 0.018, "alpha": 0.5 },
"cursor": { "show": true, "size": 0.045, "smoothing": 0.4, "clicks": true },
"captions": { "size": 0.045, "position": "bottom", "color": "#ffffff",
"background": "rgba(2,6,23,0.72)" },
"voice": { "provider": "local", "voice": "en-us+f3", "rate": 170,
"direction": "", "gain": 1, "duck": 0.25 } }
}
Notes for anything editing one by hand:
zooms and captions are time-ordered, so "the third zoom" is stable.
Neither list may overlap itself.focus: "auto" means the zoom is aimed at the nearest click and will re-aim
if moved; "manual" pins it.parseProject() is a trust boundary: it sorts and de-overlaps the lists,
clamps every number into range, drops zero-length spans, and falls back to
defaults rather than letting NaN reach the renderer. It throws only on a
missing recording or a format version it does not understand — so a
roughly-right file loads rather than failing.cuts are spans of source video the demo skips, in source time. They are
sorted, clamped, merged when they overlap, and dropped when shorter than
100 ms. A legacy trim: {startMs, endMs} (a span to keep) is migrated
into the equivalent head and tail cuts.script lines carry audioMs only as a cached measurement. Change text
and drop it — a stale length lays the timeline out for audio that no longer
exists. Lines may overlap; that is reported, not prevented.style.voice.duck is what the captured recording drops to while the voice
is talking, gain is the voice's own level.brief is what the demo is about, in your words. It is context for whoever
writes the narration — the AI writer reads it — and it is worth keeping so
the next rewrite starts from the same understanding.narrationName names the rendered voiceover sitting next to the project.
The editor takes any dropped .wav as the narration, so the name is a hint
rather than a requirement — files get renamed.style.voice.provider is "local", "gemini" or "elevenlabs"; voice is that
voice is that provider's own id (en-us+f3, Iapetus, or an ElevenLabs
voice id). rate is used by the local provider, direction by Gemini —
both are always stored, so switching provider and back keeps your settings.zooms, captions, cuts, script or style entirely is fine;
they default.Everything flows through one type, DemoRecording, defined once in
packages/core. The rules that make Phase 1's editor reusable unchanged in
Phase 2 (AI voiceover) and Phase 3 (Playwright re-record):
source. Not the planner, editor, compositor or
exporter. That branch is the seam that would break Phase 3.render/geometry.ts, against the video's real decoded size.t0 is stamped at MediaRecorder.start(); every
DemoEvent.t is ms since it. The editor reconciles against the decoded
duration on import and warns on drift over 100 ms../scripts/make-fixture.sh # markers at exactly known coordinates
pnpm -C apps/editor dev # drop fixture/demo.json + recording.webm
Every zoom must land dead centre on its marker; the double-click at 5.0/5.12 s
must produce one zoom, and the giant panel at 13 s none. pnpm -r test asserts
all of that headlessly.
For recording-side timing, see apps/extension/README.md.
pnpm dev; the fix is the same
backend the Phase 2 TODO already calls for.