ghywel/ai-storyboard

Import this repo for a toolchain and methodology to generate your own complex music videos with Claude Opus 5.5 or whatever comes next.

TypeScript

0

1 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

The Stone at the Bottom of the Sea - Extended Remix (r/ClaudeAI)

**Lyrics and video by Claude Opus 5.5 - music by Suno - concept and art direction by me.** Alternative remix 1 - [https://www.youtube.com/watch?v=Wby2MKhZqUw](https://www.youtube.com/watch?v=Wby2MKhZqUw) Alternative remix 2 -…

0

Oct 3, 2026

README

drawn in code

A toolchain and a method for making music videos in which every frame is drawn in code. The song comes from an AI music generator; the words, timing analysis, design and code are written with Claude; a human directs.

The output is a 3840×2160, 60 fps film with motion blur, where:

  • every sung word is on screen at the moment it is sung;
  • every cut lands on the beat;
  • a character with a face acts every line.

It was built for one brief and is written down here so the next film starts further on.

A frame of the demo take

What is here

PathWhat it is
METHOD.mdThe method end to end: from a brief to words, takes, timing, treatment, a shared kit, parallel plates, gates, checks and the final render
DIRECTION.mdHow to set and hold an artistic direction: palettes, type, the readable lyric, environments with clues, a character who acts (anime faces, chibi pops, lines of force), silhouettes that emote, motifs, camera grammar, the compositor's restraint, variants
MUSIC.mdThe link between the lyric (Claude), the song (Suno, or another generator) and the timing
LESSONS.mdWhat broke, how it was found and the fix, with numbers
PRIOR-ART.mdWhose work this stands on, the nearest prior art to each part, the licences that matter, and what is not new; the verified surveys are in research/
memory/The working rules in Claude Code memory format, to copy into a project's memory
templates/A treatment, a variants plan, a scene guide, a brief for a plate author, and plates.json
engine/The renderer: TypeScript scenes on Canvas2D and three.js, rendered offline in headless Chromium and encoded in the page with WebCodecs. It includes the shared kit (src/scenes/_motifs.ts, _manga.ts, _post.ts) and a demo scene.
analysis/The timing pipeline: Demucs stems, CTC forced alignment, a whisper cross-check, the beat grid, bars and sections
tools/Measuring tools (palette.py, song-analyze.py, dynamics.py, srt-lines.py, srt-check.py), soundtrack.py, make-demo-take.py, the final render (render-final.sh) and its memory guard (memwatch.py)
takes/One folder per take: the master WAV and its analysis. takes/demo/ is made by tools/make-demo-take.py.
examples/stone/The worked example: one brief, three films, every treatment, guide, timeline and scene file

Try it

You need bun and FFmpeg. The renderer drives Playwright's Chromium:

python3 tools/make-demo-take.py
cd engine
bun install
bunx playwright-core install chromium chromium-headless-shell
bun scripts/render.ts stills --take demo --t 3.6,10.1 --out ../out/demo
bun scripts/render.ts video --take demo --out ../out/demo.mp4
cd ..
TAKE=demo tools/render-final.sh

The last command writes out/final/demo-4k60.mp4.

To preview in a browser, run bunx vite in engine/ and open http://localhost:5173/?take=demo.

Making a film

Read METHOD.md, then DIRECTION.md, then examples/stone/README.md. The short version:

  1. Write. Write the lyric for the generator (MUSIC.md). Generate takes, and choose by ear.
  2. Time. Decode the take to takes/<take>/audio.wav, write the sung lines to lyrics.src.json, then run TAKE=<take> analysis/run-take.sh and analysis/analyze.py.
  3. Script. Write the treatment and its shot-by-shot script, and commit them (templates/TREATMENT.md).
  4. Build the shared kit. Build the character, sets, cast and lyric presenter, look-test them, then freeze the kit.
  5. Plates. List the plates in takes/<take>/plates.json (or a TypeScript timeline). Build them in parallel from a scene guide (templates/SCENE-GUIDE.md).
  6. Gates. Storyboard gate, then a 1080p preview, then TAKE=<take> tools/render-final.sh. Verify the result before you deliver.

Requirements and platform

  • Engine: bun, Playwright's Chromium (WebCodecs H.264 with hardware encoding) and FFmpeg.
  • Analysis: Python 3.12+ with uv, and a GPU for Demucs and the acoustic models.
    • The alignment fuses two acoustic models. MMS_FA's weights are licensed CC-BY-NC 4.0 (non-commercial); wav2vec2 LV60K is MIT. For commercial work, set ALIGN_MODELS=lv60k to use the MIT model alone. On one take this put 560 of 571 words within 0.1 s of the two-model result; the worst word was off by 0.74 s.
    • whisper_run.py uses mlx-whisper, which runs on Apple silicon; swap in another whisper elsewhere.
    • run-take.sh asks Demucs for -d mps; change it to cuda or cpu as needed.
  • Platform: developed and measured on a 16 GB Apple-silicon laptop running macOS. tools/memwatch.py reads macOS memory statistics (vm_stat, sysctl, top).

Credits

  • The engine and analysis are forked from pdoom-video by Giacomo Magnanini (MIT; engine/LICENSE-pdoom-video.txt, analysis/LICENSE-pdoom-video.txt). That repository had already solved the hard parts: offline rendering with motion blur, word-level forced alignment for sung vocals, and a scene API built for music.
  • Fonts: Archivo, Cormorant Garamond and IBM Plex Mono, under the SIL Open Font License (engine/public/fonts/FONTS.md).
  • Made by Knight Commander Gareth (direction) and Claude (Anthropic): words, analysis, design and code.

ghywel/ai-storyboard

Import this repo for a toolchain and methodology to generate your own complex music videos with Claude Opus 5.5 or whatever comes next.

TypeScript

0

1 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

The Stone at the Bottom of the Sea - Extended Remix (r/ClaudeAI)

**Lyrics and video by Claude Opus 5.5 - music by Suno - concept and art direction by me.** Alternative remix 1 - [https://www.youtube.com/watch?v=Wby2MKhZqUw](https://www.youtube.com/watch?v=Wby2MKhZqUw) Alternative remix 2 -…

0

Oct 3, 2026

README

drawn in code

A toolchain and a method for making music videos in which every frame is drawn in code. The song comes from an AI music generator; the words, timing analysis, design and code are written with Claude; a human directs.

The output is a 3840×2160, 60 fps film with motion blur, where:

  • every sung word is on screen at the moment it is sung;
  • every cut lands on the beat;
  • a character with a face acts every line.

It was built for one brief and is written down here so the next film starts further on.

A frame of the demo take

What is here

PathWhat it is
METHOD.mdThe method end to end: from a brief to words, takes, timing, treatment, a shared kit, parallel plates, gates, checks and the final render
DIRECTION.mdHow to set and hold an artistic direction: palettes, type, the readable lyric, environments with clues, a character who acts (anime faces, chibi pops, lines of force), silhouettes that emote, motifs, camera grammar, the compositor's restraint, variants
MUSIC.mdThe link between the lyric (Claude), the song (Suno, or another generator) and the timing
LESSONS.mdWhat broke, how it was found and the fix, with numbers
PRIOR-ART.mdWhose work this stands on, the nearest prior art to each part, the licences that matter, and what is not new; the verified surveys are in research/
memory/The working rules in Claude Code memory format, to copy into a project's memory
templates/A treatment, a variants plan, a scene guide, a brief for a plate author, and plates.json
engine/The renderer: TypeScript scenes on Canvas2D and three.js, rendered offline in headless Chromium and encoded in the page with WebCodecs. It includes the shared kit (src/scenes/_motifs.ts, _manga.ts, _post.ts) and a demo scene.
analysis/The timing pipeline: Demucs stems, CTC forced alignment, a whisper cross-check, the beat grid, bars and sections
tools/Measuring tools (palette.py, song-analyze.py, dynamics.py, srt-lines.py, srt-check.py), soundtrack.py, make-demo-take.py, the final render (render-final.sh) and its memory guard (memwatch.py)
takes/One folder per take: the master WAV and its analysis. takes/demo/ is made by tools/make-demo-take.py.
examples/stone/The worked example: one brief, three films, every treatment, guide, timeline and scene file

Try it

You need bun and FFmpeg. The renderer drives Playwright's Chromium:

python3 tools/make-demo-take.py
cd engine
bun install
bunx playwright-core install chromium chromium-headless-shell
bun scripts/render.ts stills --take demo --t 3.6,10.1 --out ../out/demo
bun scripts/render.ts video --take demo --out ../out/demo.mp4
cd ..
TAKE=demo tools/render-final.sh

The last command writes out/final/demo-4k60.mp4.

To preview in a browser, run bunx vite in engine/ and open http://localhost:5173/?take=demo.

Making a film

Read METHOD.md, then DIRECTION.md, then examples/stone/README.md. The short version:

  1. Write. Write the lyric for the generator (MUSIC.md). Generate takes, and choose by ear.
  2. Time. Decode the take to takes/<take>/audio.wav, write the sung lines to lyrics.src.json, then run TAKE=<take> analysis/run-take.sh and analysis/analyze.py.
  3. Script. Write the treatment and its shot-by-shot script, and commit them (templates/TREATMENT.md).
  4. Build the shared kit. Build the character, sets, cast and lyric presenter, look-test them, then freeze the kit.
  5. Plates. List the plates in takes/<take>/plates.json (or a TypeScript timeline). Build them in parallel from a scene guide (templates/SCENE-GUIDE.md).
  6. Gates. Storyboard gate, then a 1080p preview, then TAKE=<take> tools/render-final.sh. Verify the result before you deliver.

Requirements and platform

  • Engine: bun, Playwright's Chromium (WebCodecs H.264 with hardware encoding) and FFmpeg.
  • Analysis: Python 3.12+ with uv, and a GPU for Demucs and the acoustic models.
    • The alignment fuses two acoustic models. MMS_FA's weights are licensed CC-BY-NC 4.0 (non-commercial); wav2vec2 LV60K is MIT. For commercial work, set ALIGN_MODELS=lv60k to use the MIT model alone. On one take this put 560 of 571 words within 0.1 s of the two-model result; the worst word was off by 0.74 s.
    • whisper_run.py uses mlx-whisper, which runs on Apple silicon; swap in another whisper elsewhere.
    • run-take.sh asks Demucs for -d mps; change it to cuda or cpu as needed.
  • Platform: developed and measured on a 16 GB Apple-silicon laptop running macOS. tools/memwatch.py reads macOS memory statistics (vm_stat, sysctl, top).

Credits

  • The engine and analysis are forked from pdoom-video by Giacomo Magnanini (MIT; engine/LICENSE-pdoom-video.txt, analysis/LICENSE-pdoom-video.txt). That repository had already solved the hard parts: offline rendering with motion blur, word-level forced alignment for sung vocals, and a scene API built for music.
  • Fonts: Archivo, Cormorant Garamond and IBM Plex Mono, under the SIL Open Font License (engine/public/fonts/FONTS.md).
  • Made by Knight Commander Gareth (direction) and Claude (Anthropic): words, analysis, design and code.

Languages

TypeScript

66.6%

Python

31.4%

Shell

1.4%