Import this repo for a toolchain and methodology to generate your own complex music videos with Claude Opus 5.5 or whatever comes next.
TypeScript
0
1 commits
updated Oct 3, 2026
A toolchain and a method for making music videos in which every frame is drawn in code. The song comes from an AI music generator; the words, timing analysis, design and code are written with Claude; a human directs.
The output is a 3840×2160, 60 fps film with motion blur, where:
It was built for one brief and is written down here so the next film starts further on.

| Path | What it is |
|---|---|
METHOD.md | The method end to end: from a brief to words, takes, timing, treatment, a shared kit, parallel plates, gates, checks and the final render |
DIRECTION.md | How to set and hold an artistic direction: palettes, type, the readable lyric, environments with clues, a character who acts (anime faces, chibi pops, lines of force), silhouettes that emote, motifs, camera grammar, the compositor's restraint, variants |
MUSIC.md | The link between the lyric (Claude), the song (Suno, or another generator) and the timing |
LESSONS.md | What broke, how it was found and the fix, with numbers |
PRIOR-ART.md | Whose work this stands on, the nearest prior art to each part, the licences that matter, and what is not new; the verified surveys are in research/ |
memory/ | The working rules in Claude Code memory format, to copy into a project's memory |
templates/ | A treatment, a variants plan, a scene guide, a brief for a plate author, and plates.json |
engine/ | The renderer: TypeScript scenes on Canvas2D and three.js, rendered offline in headless Chromium and encoded in the page with WebCodecs. It includes the shared kit (src/scenes/_motifs.ts, _manga.ts, _post.ts) and a demo scene. |
analysis/ | The timing pipeline: Demucs stems, CTC forced alignment, a whisper cross-check, the beat grid, bars and sections |
tools/ | Measuring tools (palette.py, song-analyze.py, dynamics.py, srt-lines.py, srt-check.py), soundtrack.py, make-demo-take.py, the final render (render-final.sh) and its memory guard (memwatch.py) |
takes/ | One folder per take: the master WAV and its analysis. takes/demo/ is made by tools/make-demo-take.py. |
examples/stone/ | The worked example: one brief, three films, every treatment, guide, timeline and scene file |
You need bun and FFmpeg. The renderer drives Playwright's Chromium:
python3 tools/make-demo-take.py
cd engine
bun install
bunx playwright-core install chromium chromium-headless-shell
bun scripts/render.ts stills --take demo --t 3.6,10.1 --out ../out/demo
bun scripts/render.ts video --take demo --out ../out/demo.mp4
cd ..
TAKE=demo tools/render-final.sh
The last command writes out/final/demo-4k60.mp4.
To preview in a browser, run bunx vite in engine/ and open http://localhost:5173/?take=demo.
Read METHOD.md, then DIRECTION.md, then examples/stone/README.md. The short version:
MUSIC.md). Generate takes, and choose by ear.takes/<take>/audio.wav, write the sung lines to lyrics.src.json, then run
TAKE=<take> analysis/run-take.sh and analysis/analyze.py.templates/TREATMENT.md).takes/<take>/plates.json (or a TypeScript timeline). Build them in parallel from
a scene guide (templates/SCENE-GUIDE.md).TAKE=<take> tools/render-final.sh. Verify the result
before you deliver.ALIGN_MODELS=lv60k to use the MIT model alone. On one take this
put 560 of 571 words within 0.1 s of the two-model result; the worst word was off by 0.74 s.whisper_run.py uses mlx-whisper, which runs on Apple silicon; swap in another whisper elsewhere.run-take.sh asks Demucs for -d mps; change it to cuda or cpu as needed.tools/memwatch.py reads macOS
memory statistics (vm_stat, sysctl, top).engine/LICENSE-pdoom-video.txt, analysis/LICENSE-pdoom-video.txt). That repository had already
solved the hard parts: offline rendering with motion blur, word-level forced alignment for sung vocals, and a scene
API built for music.engine/public/fonts/FONTS.md).TypeScript
66.6%
Python
31.4%
Shell
1.4%
Import this repo for a toolchain and methodology to generate your own complex music videos with Claude Opus 5.5 or whatever comes next.
TypeScript
0
1 commits
updated Oct 3, 2026
A toolchain and a method for making music videos in which every frame is drawn in code. The song comes from an AI music generator; the words, timing analysis, design and code are written with Claude; a human directs.
The output is a 3840×2160, 60 fps film with motion blur, where:
It was built for one brief and is written down here so the next film starts further on.

| Path | What it is |
|---|---|
METHOD.md | The method end to end: from a brief to words, takes, timing, treatment, a shared kit, parallel plates, gates, checks and the final render |
DIRECTION.md | How to set and hold an artistic direction: palettes, type, the readable lyric, environments with clues, a character who acts (anime faces, chibi pops, lines of force), silhouettes that emote, motifs, camera grammar, the compositor's restraint, variants |
MUSIC.md | The link between the lyric (Claude), the song (Suno, or another generator) and the timing |
LESSONS.md | What broke, how it was found and the fix, with numbers |
PRIOR-ART.md | Whose work this stands on, the nearest prior art to each part, the licences that matter, and what is not new; the verified surveys are in research/ |
memory/ | The working rules in Claude Code memory format, to copy into a project's memory |
templates/ | A treatment, a variants plan, a scene guide, a brief for a plate author, and plates.json |
engine/ | The renderer: TypeScript scenes on Canvas2D and three.js, rendered offline in headless Chromium and encoded in the page with WebCodecs. It includes the shared kit (src/scenes/_motifs.ts, _manga.ts, _post.ts) and a demo scene. |
analysis/ | The timing pipeline: Demucs stems, CTC forced alignment, a whisper cross-check, the beat grid, bars and sections |
tools/ | Measuring tools (palette.py, song-analyze.py, dynamics.py, srt-lines.py, srt-check.py), soundtrack.py, make-demo-take.py, the final render (render-final.sh) and its memory guard (memwatch.py) |
takes/ | One folder per take: the master WAV and its analysis. takes/demo/ is made by tools/make-demo-take.py. |
examples/stone/ | The worked example: one brief, three films, every treatment, guide, timeline and scene file |
You need bun and FFmpeg. The renderer drives Playwright's Chromium:
python3 tools/make-demo-take.py
cd engine
bun install
bunx playwright-core install chromium chromium-headless-shell
bun scripts/render.ts stills --take demo --t 3.6,10.1 --out ../out/demo
bun scripts/render.ts video --take demo --out ../out/demo.mp4
cd ..
TAKE=demo tools/render-final.sh
The last command writes out/final/demo-4k60.mp4.
To preview in a browser, run bunx vite in engine/ and open http://localhost:5173/?take=demo.
Read METHOD.md, then DIRECTION.md, then examples/stone/README.md. The short version:
MUSIC.md). Generate takes, and choose by ear.takes/<take>/audio.wav, write the sung lines to lyrics.src.json, then run
TAKE=<take> analysis/run-take.sh and analysis/analyze.py.templates/TREATMENT.md).takes/<take>/plates.json (or a TypeScript timeline). Build them in parallel from
a scene guide (templates/SCENE-GUIDE.md).TAKE=<take> tools/render-final.sh. Verify the result
before you deliver.ALIGN_MODELS=lv60k to use the MIT model alone. On one take this
put 560 of 571 words within 0.1 s of the two-model result; the worst word was off by 0.74 s.whisper_run.py uses mlx-whisper, which runs on Apple silicon; swap in another whisper elsewhere.run-take.sh asks Demucs for -d mps; change it to cuda or cpu as needed.tools/memwatch.py reads macOS
memory statistics (vm_stat, sysctl, top).engine/LICENSE-pdoom-video.txt, analysis/LICENSE-pdoom-video.txt). That repository had already
solved the hard parts: offline rendering with motion blur, word-level forced alignment for sung vocals, and a scene
API built for music.engine/public/fonts/FONTS.md).TypeScript
66.6%
Python
31.4%
Shell
1.4%