Private, local AI workstation — LLM, image, music, TTS/STT, and voice cloning on your own GPU. Zero API cost, nothing leaves your machine.
15
8 commits
updated Oct 6, 2026
A private, local, zero-API-cost AI workstation — language, image, music, speech, and voice models, all running on your own GPU, in one browser tab.
No API keys. No subscriptions. No data leaving your machine. Just your GPU and a browser at
http://127.0.0.1:8800.
Local AI Studio is a self-hosted creative workstation that brings together large-language, image, music, speech-to-text, text-to-speech, and voice-cloning models behind a single tabbed web app — all running locally on an NVIDIA GPU, with nothing billed and nothing sent to the cloud.
It didn't start out this ambitious. The project began as a much smaller problem: I wanted Claude to be able to call on local models to help build a graphic novel — generating reference art, iterating on panels, and keeping the whole pipeline private and free to run as many times as I wanted. That meant giving Claude a reliable way to talk to an image model running on my own hardware.
Once that worked, it was hard to stop. A local LLM followed, then text-to-speech and voice cloning for narrating scripts, then music generation for soundtracking scenes, then speech-to-text for turning spoken notes back into text, then a proper multi-scene Story Maker for long-form fiction, and an Audiobook pipeline to turn any of it into chaptered narration. Somewhere along the way it stopped being "a tool that helps Claude make a graphic novel" and became a full studio — with its own UI, its own control panel, and its own setup runbook so it can be rebuilt on a fresh machine in one pass.
That's still the spirit of the project: give an AI (or yourself) a local set of creative tools with no per-call cost, no rate limits, and no vendor lock-in — you own the weights, the pipeline, and the output.
The studio used to make you manage the card yourself. Every tab had its own Load button, its controls stayed dead until that model was resident, and starting anything on a second tab meant stopping the first. Only one heavy model fits in 16 GB — that part is physics — but the scheduling was pushed onto you, and it was the single most tiring thing about using the place.
That is now the software's problem. Every tab submits to one job queue, and a runner drains it, loading and unloading models on the way. Line up a video, four images, a song and a 3D mesh in any order, walk away, and come back to all of them.
Interactive work — a chat turn, a transcription, one spoken line — is submitted as priority and jumps the line, so the studio still answers immediately while a long render grinds away behind it.
The visible result is a Queue button in the topbar with a live count and a countdown, and a drawer listing every job: what it is, which model it needs, how long it has been running, how long is left, and a ▲ to run one sooner.
Two changes, and they work together.
It writes files. The local LLM can create and edit them, not just answer. Ask for a
page, a script or a document and it emits the whole file; the server writes it into a
sandboxed workspace/ folder and an artifact panel appears beside the chat — list,
preview, edit by hand, save, rename, delete. HTML previews render live in an iframe.
Every model-supplied path is resolved inside workspace/ and rejected if it tries to
escape, so a hallucinated path can't reach the studio's own code.
It remembers. Conversations are kept — a Conversations list beside the chat holds every thread, newest first, each named after the question that started it, with its turn count, its age and the files it produced. Click one to pick it up exactly where you left it. Rename it, delete it, or start a new one; deleting a thread never touches the files it made.
The transcript lives on the server, in chats/<id>.json, not in the browser tab. That
is the part that matters: the page holds an id, not a history. Reload, open a second tab,
restart the studio, come back tomorrow — same conversation. Close the browser mid-reply
and the turn is still saved when it lands, because the queue job that is writing it does
not care whether anyone is watching. It also means the model's context is assembled
server-side from the real transcript rather than from whatever the page happened to still
have in memory.
.glb mesh (Hunyuan3D 2.1),
from the same kind of single image Sprite Studio already takes.Under the hood, the local LLM's context window was raised from Ollama's 4096-token default to 32k. That default had been silently truncating long prompts — a 32,000-token input was arriving as 4,096 tokens with the front discarded. See Language model settings below.
Everything Before: a 12-minute narrated sci-fi film. Claude Opus 5.5 wrote, planned and directed it in Claude Code, and every shot was rendered on one RTX 4080 SUPER with this studio: keyframes from Image, picture and ambient sound from Video, and the score from Music Generation. Nothing in it was filmed or licensed from stock.
To make your own, follow the Film Recipe, the same flow written for any subject: brief, script, timed plan, voice, keyframes, clips, sound and assembly.
One prompt, and Claude Code builds the whole studio — the clone, the runbook, 29 source files, five conda environments, 91.7 GB of weights, and then what the finished studio actually makes.
The second half of that film stands on its own as the Video tab capability demo.
Every tab below submits its work to the same job queue — there is nothing to load first, and jobs from different tabs line up together.
| Tab | What it does | Backend |
|---|---|---|
| 🧠 Language | Code / research / vision prompts to a local LLM, with a 🔓 Unlocked (uncensored) option. The model can also write real files — ask for a page, a script or a document and it lands in a sandboxed workspace/ folder with an artifact panel beside the chat: preview, hand-edit, save, rename, delete, with HTML rendered live. Conversations are kept and listed beside the chat — named, dated, resumable, and stored server-side in chats/, so a reload or a restart doesn't lose the thread. | Ollama |
| 🎨 Image — Generate | Text → image (FLUX.2 Klein) | ComfyUI |
| ✂️ Image — Edit | Reference-guided edit / remove / reframe / outpaint | ComfyUI |
| 🕹️ Sprite Studio | One reference image → style-matched 2D game sprites: single actions or a full animation set (idle/walk/run/jump/fall/crouch/attack/hurt/death), true transparent backgrounds, per-action strips + combined sprite sheet with engine-ready JSON metadata, per-frame re-roll | ComfyUI + rembg |
| 🧊 Image → 3D | One reference image → a 3D mesh (Hunyuan3D 2.1), saved as .glb. Takes the same kind of single reference image Sprite Studio does, so there is no new concept to learn. Geometry only — the mesh arrives untextured, ready to paint in Blender or have the reference baked onto it. | ComfyUI + mesh3d.py |
| 🎬 Video | Text → video with native stereo audio, generated together in one sampler pass (MiniMax H3) — nothing to mux afterwards. Three input modes: Text → video; First / last frame (drop one keyframe for image→video, or both and H3 interpolates between them); and References, which accepts up to 9 images, 3 video clips (each with its own optional soundtrack) and 3 audio clips — addressed in the prompt as <Picture 1>, <Video 1>, <Audio 1> — to carry identity, motion, camera style or a cloned voice into the result. Up to 15s at 24fps, rendered at H3's native 768p (running below native makes it stop animating — see the performance notes below). Streams the official 21GB int8 unets through 16GB of VRAM via ComfyUI's per-module offload, with a measured, self-correcting ETA rather than a spinner. See the licensing note below before sharing anything made here. | ComfyUI + h3gen.py |
| 📊 Composer | A text brief → a fully arranged, mixed, multitrack instrumental. Two engines: Arrange (a local LLM plans, the studio engine writes every note) and Orchestrate (SymphonyGen — a real 211M-parameter orchestral model writes 32 bars of multitrack score itself, no LLM involved, ~37s). Orchestrate merges up to 27 generated desks into orchestral families, seats them across the stereo field, and puts dynamics back into a model output that arrives at a flat velocity; the existing mixer, DAW grid, piano roll, stems and master chain then run unchanged. It can also re-orchestrate a MIDI you hand it. In Arrange, a local LLM plans the musical direction (instruments, key, tempo, structure, chords, mix, automation) from a fixed "studio" menu — General MIDI patches played by FluidSynth — then a deterministic engine (no note-level AI) writes every part, mixes each track through its own FX chain (saturation, tone shelves, tempo-synced delay, convolution reverb, automated sends), and adds production moves (risers, impacts, downlifters, drops, sidechain ducking). A 3-step wizard — Set up (brief/style/length/instrument count) → Arrange (DAW-style clip grid, per-track mix/automation, piano-roll note editing, re-render without a new LLM call) → Export (master MP3/WAV, multitrack MIDI, per-instrument FLAC stems). 8 genre templates mean it never fails even on a bad LLM response, and an "LLM off" mode composes from the template alone with no model load at all | Ollama (planning) + SymphonyGen (symphgen.py) + composerkit.py (CPU-only render) |
| 🎵 Music Generation | Full songs & instrumentals from style tags + lyrics (ACE-Step 1.5 XL) — structure/vocal/energy lyric tags, BPM/Key/time-signature control, remix mode | ComfyUI |
| 🔊 Sound Effects | Foley, impacts and ambience from a description (Stable Audio 3.0 Small SFX) — the third leg of game audio next to ACE-Step's music and the TTS engines' voice. Renders several takes at once, because picking between takes is the whole workflow for sound effects. Distilled to 8 steps, so a take takes about a second. The tab asks for the source, material, space and — the one that decides whether you get a smash or a tap — how the sound evolves over time. | ComfyUI + sfxgen.py |
| 🎹 Lullaby | Any song → soft lullaby instrumental. A workbench splits the song into 6 tracks (vocals/guitar/piano/other/bass/drums) with scrubbable waveform players so you pick exactly what carries into the result, then three engines: Remix (default — the selected tracks are cleaned, dynamics flattened so it stays soft throughout, then ACE-Step audio-to-audio re-imagines it with lullaby tags; closely resembles the original, with denoise/softness/slowdown controls), Piano (melody transcribed directly from the selected tracks, key/chords detected, rebuilt as a rocking piano + music-box arrangement at 55-88bpm on the Salamander sampled grand), and Melody Match (traces each sung note's continuous pitch curve via FCPE — real note boundaries, no scale-snap or quantization — onto a single portamento-capable instrument: cello/violin/flute/synth voice/music box; a per-track Route selector lets some ticked stems go through Melody Match while others get a full Piano-style rebuilt arrangement in the same render, mixed together, with an optional ACE-Step polish pass afterward) | lullabykit (2-pass Demucs + basic-pitch/FCPE + librosa + FluidSynth) + ACE-Step |
| ✂️ Track Splitter | Any song → its 6 individual instrument tracks (vocals/guitar/piano/other/bass/drums), each with a scrubbable player and its own download, plus a "download all" zip and a persistent library of past splits — shares its separation cache with the Lullaby tab | lullabykit (Demucs) |
| 🎙️ Speech → Text | Transcribe audio | NeMo Parakeet |
| 🔊 Text → Speech | Fast narration (Kokoro) and voice cloning (Chatterbox), with a one-click clean up this recording pass (DeepFilterNet denoise + de-reverb) on the cloning reference clip — cloning quality is capped by whatever mic and room the reference came from. Runs on the CPU in its own env, so it never unloads the voice model. | conda envs |
| 🗣️ Voice Studio | Fine-tune & reuse a personal voice | XTTS-v2 |
| 📖 Story Maker | Timeline-driven multi-scene story / novel generation | koboldcpp (Cydonia-24B) / Ollama |
| 📚 Audiobook | Story project or pasted text → chaptered MP3s with natural pacing and loudness normalization | TTS worker + ffmpeg |
Plus a CLI (studioctl.ps1) and a visual control panel (studio_gui.pyw, with a one-click Desktop shortcut) to start, stop, and monitor the whole stack — Ollama, ComfyUI, and the studio server — from one place.
The job queue. Seven jobs from four different tabs, lined up on one card. The running job says what it is doing (loading image:base4b); each pending job says how long it takes and which model it needs; one has never run before and says so instead of guessing. ▲ runs a job sooner, × cancels it.

Home — model status, pinning, and the queue countdown in the topbar ![]() | Language — four kept conversations on the left, the file the model just wrote previewing live on the right ![]() |
Story Maker ![]() | Image — Generate ![]() |
Image — Edit ![]() | Music — ACE-Step ![]() |
Speech → Text ![]() | Text → Speech ![]() |
Voice Studio — create a voice ![]() | Audiobook ![]() |
Sound Effects — four takes of one prompt, because picking between takes is the workflow ![]() | Image → 3D — one reference image, 508,068 triangles, in a viewer written into the page ![]() |
| Sprite Studio — single action | Sprite Studio — full sprite set (all 9 built-in actions selectable) |
Lullaby — Remix engine (default) ![]() | Lullaby — Piano engine ![]() |
Lullaby — Melody Match engine ![]() | Melody Match — per-track routing (vocals → instrument, piano → rebuilt arrangement, mixed together) ![]() |
Track Splitter ![]() | Track Splitter — results (per-track players + Play all selected) ![]() |
MiniMax H3 — text → video with native audio, generated in one pass. ⚠️ Licence restricts local deployment in the UK, EU, US, and South Korea, and anything you share must be disclosed as AI-generated — see the notice below.

Recommended length: 5–15 seconds. Generation time scales with length, not with a fixed startup cost — measured across four real generations on this rig (RTX 4080, 768p native, 20 steps), a 10.125s clip consistently took ~31.4 minutes end-to-end (model load included), which works out to roughly 3 minutes of generation time per second of video. As a rule of thumb: a 5s clip ≈ 15 min, a 15s clip ≈ 45 min. Going shorter than 5s wastes most of that time on model loading rather than the clip itself; going much past 15s just multiplies an already-long wait for diminishing narrative payoff.
Step 1 — Set up (brief, style, length, sound options) ![]() | Step 2 — Arrange (clip grid: tracks × sections, real generated song) ![]() |
Step 2 — track inspector (instrument/mix/automation + piano-roll notes) ![]() | Step 3 — Export (master, full mix table, arrangement structure, per-instrument stems) ![]() |
The entire install is captured in a single, self-contained runbook — SETUP_FOR_CLAUDE.md — designed to be handed to Claude Code (or followed by hand) on a fresh machine. It carries the full source of every component embedded inline, so nothing needs to be fetched from a separate repo to bootstrap the app itself.
SETUP_FOR_CLAUDE.md to Claude Code, or work through its phases manually:
For a deeper dive, see the companion docs in /docs:
A single Python stdlib HTTP server (server.py) serves the UI (index.html) and brokers every request to a backend.
One job queue sits between the UI and all of them. Tabs don't call backends directly any more — they POST /api/queue/add with a job kind, and a runner per lane (one GPU, one CPU) takes jobs off the list, makes sure the right model is resident, runs the job, and records how long it took. A single-threaded runner is also what makes the existing one-at-a-time job classes safe to line up: their "already running" rejection is now unreachable. The whole UI polls one endpoint, /api/queue, which carries every job's progress, the resident model and the pin state in a single response.
The backends:
spritekit.py under ComfyUI's venv)lullabykit/) — the Lullaby pipeline: its own venv (torch/CUDA, Demucs, basic-pitch/FCPE) plus bundled FluidSynth binaries and the FluidR3 GM soundfont; runs as a transient subprocess job, not a resident workerThree folders are written to on your behalf rather than by you: workspace/ (the Language tab's files, sandboxed — no path from the model can resolve outside it), chats/ (one JSON per conversation, which is why the Language tab survives a reload) and _tmp/ (the persisted queue, staged uploads, and the measured-timing ledger that every ETA is drawn from).
The whole stack is controlled from studioctl.ps1 (CLI) or studio_gui.pyw (visual control panel), which start, stop, and health-check every service.
Local AI Studio is glue code and a UI around excellent open-weight models and tools built by other people:
Check each model's own license before any commercial use — several of the above are personal/research use only. This project itself adds no additional restriction beyond what each model's license already requires.
Three of the obvious choices here turned out to be wrong, all in the same direction: the settings that look right for a small GPU quietly destroy the output. All numbers below are same-prompt, same-seed, nothing else on the GPU.
1. Run at native 768p. Below native, H3 stops animating.
| Short edge | Near-static frames | Audio mean | Time |
|---|---|---|---|
| 480p | 89 / 100 | −40.7 dB | 261s |
| 768p (native) | 0 / 100 | −36.8 dB | 779s |
This is the big one. At 480p the model composes a handsome scene and then barely moves it — four frames spanning five seconds are near-identical. At native 768 short edge the camera actually moves. ComfyUI's template ships a ~480p selector so it runs on modest cards; that's an accessibility default, not a quality one. The tab now sizes by short edge rather than megapixels, because a megapixel target lands 32px under native at 16:9 (1344×736) on exactly the axis that decides whether the clip moves.
2. Use the official int8 weights, not a Q3 GGUF — it is better and faster.
| Model | Result | Time |
|---|---|---|
MiniMax-H3-FL2VA-Q3_K_M.gguf (15.6GB) | subject collapses into abstract dark geometry | 779s |
minimax_h3_fl2va_pruned_int8_convrot (21GB) | renders the actual subject | 573s |
Counter-intuitive on a 16GB card, but the bigger file wins twice over. comfy-kitchen has
a fused dequantize_int8_convrot CUDA kernel (live once you're on cu130) while GGUF
Q3_K dequantises the slow way — and ~3.4 bits/weight visibly costs a 20B transformer its
prompt adherence. Both unets stay selectable via --unet.
3. EasyCache is fast but it eats the audio. SageAttention does nothing at all.
| Configuration | Wall | vs baseline | Audio mean |
|---|---|---|---|
| Baseline | 261s | — | −40.7 dB |
| + SageAttention 2.2 (FP8) | 260s | 1.00× — no gain | — |
| + EasyCache | 160s | 1.63× | −49.8 dB |
EasyCache skips 8 of 20 steps, which explains the speedup exactly — but picture and sound share one latent and the skip decision is dominated by the video channels, so the audio branch is starved of steps it needed. 9 dB quieter. Fine for silent drafts, off by default otherwise.
SageAttention speeds up attention arithmetic, but the GPU here is waiting on PCIe, not
on maths: roughly half the model streams from system RAM every step
(7269 MB loaded, 8073 MB offloaded). There is no stall for a faster kernel to fill. On
a card that fits the whole model in VRAM the balance would likely flip back.
Practical default: 768p, 20 steps, int8 weights, EasyCache and SageAttention off — about 9½ minutes for a 5-second clip. Drop to 480p only for checking composition, and never judge motion from a draft.
Ollama's default context window is 4096 tokens, and its OpenAI-compatible
/v1/chat/completions endpoint — which is what this studio calls — will not accept
num_ctx per request. So the window can only be set on the server, and left alone it fails
silently rather than loudly:
| prompt sent | tokens the model saw | answer | |
|---|---|---|---|
| Ollama default | ~32,800 | 4,096 | (empty) |
| After the fix | ~32,800 | 25,265 | correct |
studioctl.ps1 now starts Ollama with:
OLLAMA_CONTEXT_LENGTH=32768
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
The last two are what make the first affordable — quantising the KV cache roughly halves its memory for a perplexity change too small to notice, which is what buys the larger window inside 16 GB. Loaded with a 32k context, gpt-oss:20b sits at about 14.0 GB.
Two settings that turned out to matter more than expected, both measured rather than assumed:
low. It is gpt-oss's single largest quality knob. The
original reason was real — deep reasoning could consume the entire token budget and
return an empty answer — but that cause had already been fixed by raising max_tokens.
Code and research tasks now run at high, with a fallback that retries once at shallow
effort if the model still runs itself out of budget.reasoning=high, temperature
0.2 returned an empty answer 1 run in 3, and raising the budget from 8192 to 16384
tokens did not help — still 1 in 3. The same prompt at temperature 1.0 answered 3 for
3 and averaged 14s against 36s. Over-constrained, the chain of thought loops instead of
terminating. OpenAI's recommendation of 1.0 for gpt-oss is a reliability setting, not a
style preference.If Ollama is already running when the studio starts, these apply only after it restarts —
studioctl says so rather than reporting a healthy service that is quietly truncating.
The Video tab is the one part of this studio that is not freely usable everywhere, and the restriction is unusual, so it is worth stating plainly.
MiniMax H3 is released under the MiniMax H3 Community License Agreement, which defines an "Applicable Territory" of worldwide, excluding the Excluded Territories. The excluded territories are:
the European Union · the United Kingdom · the United States · the Republic of Korea
Two things that are easy to miss:
If you live in an Excluded Territory you can apply for individual authorisation:
Licence request form → https://platform.minimax.io/h3-license
MiniMax have confirmed that individuals may apply — enter Personal/None where the form asks
for a company name. Their official licence Q&A thread, which is the source for the
clarifications above, is at
https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/12.
Status for this repository's maintainer: individual authorisation for local deployment has been granted by MiniMax under a separate agreement. That authorisation is personal to the maintainer and does not transfer to anyone who clones this repository.
Until you hold your own, treat the Video tab as local and personal only — do not redistribute, publish or commercially use its
output. Every other tab in this studio is unaffected. The Video tab carries this same notice
in-app so it cannot be missed. The setup package now installs it (SETUP_FOR_CLAUDE.md
PHASE 5h), which is why that phase opens by putting this notice in front of the user and asking
before it spends the ~63 GB: skipping it leaves the rest of the studio fully working.
If you are redistributing Local AI Studio, you must pass this notice on to your users; their obligations depend on where they are, not where you are.
Local AI Studio is released under the Apache License 2.0: use it, change it and build on it, commercially or not, as long as you keep the licence and NOTICE with it and mark the files you changed.
That covers this project's own code and docs only. The models and tools the setup runbook downloads are not part of it and keep their own licences, several of which are personal/research-use only and one (MiniMax H3) territory-restricted. See Credits & licensing and the MiniMax H3 notice above before using any output commercially.
Private, local AI workstation — LLM, image, music, TTS/STT, and voice cloning on your own GPU. Zero API cost, nothing leaves your machine.
15
8 commits
updated Oct 6, 2026
A private, local, zero-API-cost AI workstation — language, image, music, speech, and voice models, all running on your own GPU, in one browser tab.
No API keys. No subscriptions. No data leaving your machine. Just your GPU and a browser at
http://127.0.0.1:8800.
Local AI Studio is a self-hosted creative workstation that brings together large-language, image, music, speech-to-text, text-to-speech, and voice-cloning models behind a single tabbed web app — all running locally on an NVIDIA GPU, with nothing billed and nothing sent to the cloud.
It didn't start out this ambitious. The project began as a much smaller problem: I wanted Claude to be able to call on local models to help build a graphic novel — generating reference art, iterating on panels, and keeping the whole pipeline private and free to run as many times as I wanted. That meant giving Claude a reliable way to talk to an image model running on my own hardware.
Once that worked, it was hard to stop. A local LLM followed, then text-to-speech and voice cloning for narrating scripts, then music generation for soundtracking scenes, then speech-to-text for turning spoken notes back into text, then a proper multi-scene Story Maker for long-form fiction, and an Audiobook pipeline to turn any of it into chaptered narration. Somewhere along the way it stopped being "a tool that helps Claude make a graphic novel" and became a full studio — with its own UI, its own control panel, and its own setup runbook so it can be rebuilt on a fresh machine in one pass.
That's still the spirit of the project: give an AI (or yourself) a local set of creative tools with no per-call cost, no rate limits, and no vendor lock-in — you own the weights, the pipeline, and the output.
The studio used to make you manage the card yourself. Every tab had its own Load button, its controls stayed dead until that model was resident, and starting anything on a second tab meant stopping the first. Only one heavy model fits in 16 GB — that part is physics — but the scheduling was pushed onto you, and it was the single most tiring thing about using the place.
That is now the software's problem. Every tab submits to one job queue, and a runner drains it, loading and unloading models on the way. Line up a video, four images, a song and a 3D mesh in any order, walk away, and come back to all of them.
Interactive work — a chat turn, a transcription, one spoken line — is submitted as priority and jumps the line, so the studio still answers immediately while a long render grinds away behind it.
The visible result is a Queue button in the topbar with a live count and a countdown, and a drawer listing every job: what it is, which model it needs, how long it has been running, how long is left, and a ▲ to run one sooner.
Two changes, and they work together.
It writes files. The local LLM can create and edit them, not just answer. Ask for a
page, a script or a document and it emits the whole file; the server writes it into a
sandboxed workspace/ folder and an artifact panel appears beside the chat — list,
preview, edit by hand, save, rename, delete. HTML previews render live in an iframe.
Every model-supplied path is resolved inside workspace/ and rejected if it tries to
escape, so a hallucinated path can't reach the studio's own code.
It remembers. Conversations are kept — a Conversations list beside the chat holds every thread, newest first, each named after the question that started it, with its turn count, its age and the files it produced. Click one to pick it up exactly where you left it. Rename it, delete it, or start a new one; deleting a thread never touches the files it made.
The transcript lives on the server, in chats/<id>.json, not in the browser tab. That
is the part that matters: the page holds an id, not a history. Reload, open a second tab,
restart the studio, come back tomorrow — same conversation. Close the browser mid-reply
and the turn is still saved when it lands, because the queue job that is writing it does
not care whether anyone is watching. It also means the model's context is assembled
server-side from the real transcript rather than from whatever the page happened to still
have in memory.
.glb mesh (Hunyuan3D 2.1),
from the same kind of single image Sprite Studio already takes.Under the hood, the local LLM's context window was raised from Ollama's 4096-token default to 32k. That default had been silently truncating long prompts — a 32,000-token input was arriving as 4,096 tokens with the front discarded. See Language model settings below.
Everything Before: a 12-minute narrated sci-fi film. Claude Opus 5.5 wrote, planned and directed it in Claude Code, and every shot was rendered on one RTX 4080 SUPER with this studio: keyframes from Image, picture and ambient sound from Video, and the score from Music Generation. Nothing in it was filmed or licensed from stock.
To make your own, follow the Film Recipe, the same flow written for any subject: brief, script, timed plan, voice, keyframes, clips, sound and assembly.
One prompt, and Claude Code builds the whole studio — the clone, the runbook, 29 source files, five conda environments, 91.7 GB of weights, and then what the finished studio actually makes.
The second half of that film stands on its own as the Video tab capability demo.
Every tab below submits its work to the same job queue — there is nothing to load first, and jobs from different tabs line up together.
| Tab | What it does | Backend |
|---|---|---|
| 🧠 Language | Code / research / vision prompts to a local LLM, with a 🔓 Unlocked (uncensored) option. The model can also write real files — ask for a page, a script or a document and it lands in a sandboxed workspace/ folder with an artifact panel beside the chat: preview, hand-edit, save, rename, delete, with HTML rendered live. Conversations are kept and listed beside the chat — named, dated, resumable, and stored server-side in chats/, so a reload or a restart doesn't lose the thread. | Ollama |
| 🎨 Image — Generate | Text → image (FLUX.2 Klein) | ComfyUI |
| ✂️ Image — Edit | Reference-guided edit / remove / reframe / outpaint | ComfyUI |
| 🕹️ Sprite Studio | One reference image → style-matched 2D game sprites: single actions or a full animation set (idle/walk/run/jump/fall/crouch/attack/hurt/death), true transparent backgrounds, per-action strips + combined sprite sheet with engine-ready JSON metadata, per-frame re-roll | ComfyUI + rembg |
| 🧊 Image → 3D | One reference image → a 3D mesh (Hunyuan3D 2.1), saved as .glb. Takes the same kind of single reference image Sprite Studio does, so there is no new concept to learn. Geometry only — the mesh arrives untextured, ready to paint in Blender or have the reference baked onto it. | ComfyUI + mesh3d.py |
| 🎬 Video | Text → video with native stereo audio, generated together in one sampler pass (MiniMax H3) — nothing to mux afterwards. Three input modes: Text → video; First / last frame (drop one keyframe for image→video, or both and H3 interpolates between them); and References, which accepts up to 9 images, 3 video clips (each with its own optional soundtrack) and 3 audio clips — addressed in the prompt as <Picture 1>, <Video 1>, <Audio 1> — to carry identity, motion, camera style or a cloned voice into the result. Up to 15s at 24fps, rendered at H3's native 768p (running below native makes it stop animating — see the performance notes below). Streams the official 21GB int8 unets through 16GB of VRAM via ComfyUI's per-module offload, with a measured, self-correcting ETA rather than a spinner. See the licensing note below before sharing anything made here. | ComfyUI + h3gen.py |
| 📊 Composer | A text brief → a fully arranged, mixed, multitrack instrumental. Two engines: Arrange (a local LLM plans, the studio engine writes every note) and Orchestrate (SymphonyGen — a real 211M-parameter orchestral model writes 32 bars of multitrack score itself, no LLM involved, ~37s). Orchestrate merges up to 27 generated desks into orchestral families, seats them across the stereo field, and puts dynamics back into a model output that arrives at a flat velocity; the existing mixer, DAW grid, piano roll, stems and master chain then run unchanged. It can also re-orchestrate a MIDI you hand it. In Arrange, a local LLM plans the musical direction (instruments, key, tempo, structure, chords, mix, automation) from a fixed "studio" menu — General MIDI patches played by FluidSynth — then a deterministic engine (no note-level AI) writes every part, mixes each track through its own FX chain (saturation, tone shelves, tempo-synced delay, convolution reverb, automated sends), and adds production moves (risers, impacts, downlifters, drops, sidechain ducking). A 3-step wizard — Set up (brief/style/length/instrument count) → Arrange (DAW-style clip grid, per-track mix/automation, piano-roll note editing, re-render without a new LLM call) → Export (master MP3/WAV, multitrack MIDI, per-instrument FLAC stems). 8 genre templates mean it never fails even on a bad LLM response, and an "LLM off" mode composes from the template alone with no model load at all | Ollama (planning) + SymphonyGen (symphgen.py) + composerkit.py (CPU-only render) |
| 🎵 Music Generation | Full songs & instrumentals from style tags + lyrics (ACE-Step 1.5 XL) — structure/vocal/energy lyric tags, BPM/Key/time-signature control, remix mode | ComfyUI |
| 🔊 Sound Effects | Foley, impacts and ambience from a description (Stable Audio 3.0 Small SFX) — the third leg of game audio next to ACE-Step's music and the TTS engines' voice. Renders several takes at once, because picking between takes is the whole workflow for sound effects. Distilled to 8 steps, so a take takes about a second. The tab asks for the source, material, space and — the one that decides whether you get a smash or a tap — how the sound evolves over time. | ComfyUI + sfxgen.py |
| 🎹 Lullaby | Any song → soft lullaby instrumental. A workbench splits the song into 6 tracks (vocals/guitar/piano/other/bass/drums) with scrubbable waveform players so you pick exactly what carries into the result, then three engines: Remix (default — the selected tracks are cleaned, dynamics flattened so it stays soft throughout, then ACE-Step audio-to-audio re-imagines it with lullaby tags; closely resembles the original, with denoise/softness/slowdown controls), Piano (melody transcribed directly from the selected tracks, key/chords detected, rebuilt as a rocking piano + music-box arrangement at 55-88bpm on the Salamander sampled grand), and Melody Match (traces each sung note's continuous pitch curve via FCPE — real note boundaries, no scale-snap or quantization — onto a single portamento-capable instrument: cello/violin/flute/synth voice/music box; a per-track Route selector lets some ticked stems go through Melody Match while others get a full Piano-style rebuilt arrangement in the same render, mixed together, with an optional ACE-Step polish pass afterward) | lullabykit (2-pass Demucs + basic-pitch/FCPE + librosa + FluidSynth) + ACE-Step |
| ✂️ Track Splitter | Any song → its 6 individual instrument tracks (vocals/guitar/piano/other/bass/drums), each with a scrubbable player and its own download, plus a "download all" zip and a persistent library of past splits — shares its separation cache with the Lullaby tab | lullabykit (Demucs) |
| 🎙️ Speech → Text | Transcribe audio | NeMo Parakeet |
| 🔊 Text → Speech | Fast narration (Kokoro) and voice cloning (Chatterbox), with a one-click clean up this recording pass (DeepFilterNet denoise + de-reverb) on the cloning reference clip — cloning quality is capped by whatever mic and room the reference came from. Runs on the CPU in its own env, so it never unloads the voice model. | conda envs |
| 🗣️ Voice Studio | Fine-tune & reuse a personal voice | XTTS-v2 |
| 📖 Story Maker | Timeline-driven multi-scene story / novel generation | koboldcpp (Cydonia-24B) / Ollama |
| 📚 Audiobook | Story project or pasted text → chaptered MP3s with natural pacing and loudness normalization | TTS worker + ffmpeg |
Plus a CLI (studioctl.ps1) and a visual control panel (studio_gui.pyw, with a one-click Desktop shortcut) to start, stop, and monitor the whole stack — Ollama, ComfyUI, and the studio server — from one place.
The job queue. Seven jobs from four different tabs, lined up on one card. The running job says what it is doing (loading image:base4b); each pending job says how long it takes and which model it needs; one has never run before and says so instead of guessing. ▲ runs a job sooner, × cancels it.

Home — model status, pinning, and the queue countdown in the topbar ![]() | Language — four kept conversations on the left, the file the model just wrote previewing live on the right ![]() |
Story Maker ![]() | Image — Generate ![]() |
Image — Edit ![]() | Music — ACE-Step ![]() |
Speech → Text ![]() | Text → Speech ![]() |
Voice Studio — create a voice ![]() | Audiobook ![]() |
Sound Effects — four takes of one prompt, because picking between takes is the workflow ![]() | Image → 3D — one reference image, 508,068 triangles, in a viewer written into the page ![]() |
| Sprite Studio — single action | Sprite Studio — full sprite set (all 9 built-in actions selectable) |
Lullaby — Remix engine (default) ![]() | Lullaby — Piano engine ![]() |
Lullaby — Melody Match engine ![]() | Melody Match — per-track routing (vocals → instrument, piano → rebuilt arrangement, mixed together) ![]() |
Track Splitter ![]() | Track Splitter — results (per-track players + Play all selected) ![]() |
MiniMax H3 — text → video with native audio, generated in one pass. ⚠️ Licence restricts local deployment in the UK, EU, US, and South Korea, and anything you share must be disclosed as AI-generated — see the notice below.

Recommended length: 5–15 seconds. Generation time scales with length, not with a fixed startup cost — measured across four real generations on this rig (RTX 4080, 768p native, 20 steps), a 10.125s clip consistently took ~31.4 minutes end-to-end (model load included), which works out to roughly 3 minutes of generation time per second of video. As a rule of thumb: a 5s clip ≈ 15 min, a 15s clip ≈ 45 min. Going shorter than 5s wastes most of that time on model loading rather than the clip itself; going much past 15s just multiplies an already-long wait for diminishing narrative payoff.
Step 1 — Set up (brief, style, length, sound options) ![]() | Step 2 — Arrange (clip grid: tracks × sections, real generated song) ![]() |
Step 2 — track inspector (instrument/mix/automation + piano-roll notes) ![]() | Step 3 — Export (master, full mix table, arrangement structure, per-instrument stems) ![]() |
The entire install is captured in a single, self-contained runbook — SETUP_FOR_CLAUDE.md — designed to be handed to Claude Code (or followed by hand) on a fresh machine. It carries the full source of every component embedded inline, so nothing needs to be fetched from a separate repo to bootstrap the app itself.
SETUP_FOR_CLAUDE.md to Claude Code, or work through its phases manually:
For a deeper dive, see the companion docs in /docs:
A single Python stdlib HTTP server (server.py) serves the UI (index.html) and brokers every request to a backend.
One job queue sits between the UI and all of them. Tabs don't call backends directly any more — they POST /api/queue/add with a job kind, and a runner per lane (one GPU, one CPU) takes jobs off the list, makes sure the right model is resident, runs the job, and records how long it took. A single-threaded runner is also what makes the existing one-at-a-time job classes safe to line up: their "already running" rejection is now unreachable. The whole UI polls one endpoint, /api/queue, which carries every job's progress, the resident model and the pin state in a single response.
The backends:
spritekit.py under ComfyUI's venv)lullabykit/) — the Lullaby pipeline: its own venv (torch/CUDA, Demucs, basic-pitch/FCPE) plus bundled FluidSynth binaries and the FluidR3 GM soundfont; runs as a transient subprocess job, not a resident workerThree folders are written to on your behalf rather than by you: workspace/ (the Language tab's files, sandboxed — no path from the model can resolve outside it), chats/ (one JSON per conversation, which is why the Language tab survives a reload) and _tmp/ (the persisted queue, staged uploads, and the measured-timing ledger that every ETA is drawn from).
The whole stack is controlled from studioctl.ps1 (CLI) or studio_gui.pyw (visual control panel), which start, stop, and health-check every service.
Local AI Studio is glue code and a UI around excellent open-weight models and tools built by other people:
Check each model's own license before any commercial use — several of the above are personal/research use only. This project itself adds no additional restriction beyond what each model's license already requires.
Three of the obvious choices here turned out to be wrong, all in the same direction: the settings that look right for a small GPU quietly destroy the output. All numbers below are same-prompt, same-seed, nothing else on the GPU.
1. Run at native 768p. Below native, H3 stops animating.
| Short edge | Near-static frames | Audio mean | Time |
|---|---|---|---|
| 480p | 89 / 100 | −40.7 dB | 261s |
| 768p (native) | 0 / 100 | −36.8 dB | 779s |
This is the big one. At 480p the model composes a handsome scene and then barely moves it — four frames spanning five seconds are near-identical. At native 768 short edge the camera actually moves. ComfyUI's template ships a ~480p selector so it runs on modest cards; that's an accessibility default, not a quality one. The tab now sizes by short edge rather than megapixels, because a megapixel target lands 32px under native at 16:9 (1344×736) on exactly the axis that decides whether the clip moves.
2. Use the official int8 weights, not a Q3 GGUF — it is better and faster.
| Model | Result | Time |
|---|---|---|
MiniMax-H3-FL2VA-Q3_K_M.gguf (15.6GB) | subject collapses into abstract dark geometry | 779s |
minimax_h3_fl2va_pruned_int8_convrot (21GB) | renders the actual subject | 573s |
Counter-intuitive on a 16GB card, but the bigger file wins twice over. comfy-kitchen has
a fused dequantize_int8_convrot CUDA kernel (live once you're on cu130) while GGUF
Q3_K dequantises the slow way — and ~3.4 bits/weight visibly costs a 20B transformer its
prompt adherence. Both unets stay selectable via --unet.
3. EasyCache is fast but it eats the audio. SageAttention does nothing at all.
| Configuration | Wall | vs baseline | Audio mean |
|---|---|---|---|
| Baseline | 261s | — | −40.7 dB |
| + SageAttention 2.2 (FP8) | 260s | 1.00× — no gain | — |
| + EasyCache | 160s | 1.63× | −49.8 dB |
EasyCache skips 8 of 20 steps, which explains the speedup exactly — but picture and sound share one latent and the skip decision is dominated by the video channels, so the audio branch is starved of steps it needed. 9 dB quieter. Fine for silent drafts, off by default otherwise.
SageAttention speeds up attention arithmetic, but the GPU here is waiting on PCIe, not
on maths: roughly half the model streams from system RAM every step
(7269 MB loaded, 8073 MB offloaded). There is no stall for a faster kernel to fill. On
a card that fits the whole model in VRAM the balance would likely flip back.
Practical default: 768p, 20 steps, int8 weights, EasyCache and SageAttention off — about 9½ minutes for a 5-second clip. Drop to 480p only for checking composition, and never judge motion from a draft.
Ollama's default context window is 4096 tokens, and its OpenAI-compatible
/v1/chat/completions endpoint — which is what this studio calls — will not accept
num_ctx per request. So the window can only be set on the server, and left alone it fails
silently rather than loudly:
| prompt sent | tokens the model saw | answer | |
|---|---|---|---|
| Ollama default | ~32,800 | 4,096 | (empty) |
| After the fix | ~32,800 | 25,265 | correct |
studioctl.ps1 now starts Ollama with:
OLLAMA_CONTEXT_LENGTH=32768
OLLAMA_FLASH_ATTENTION=1
OLLAMA_KV_CACHE_TYPE=q8_0
The last two are what make the first affordable — quantising the KV cache roughly halves its memory for a perplexity change too small to notice, which is what buys the larger window inside 16 GB. Loaded with a 32k context, gpt-oss:20b sits at about 14.0 GB.
Two settings that turned out to matter more than expected, both measured rather than assumed:
low. It is gpt-oss's single largest quality knob. The
original reason was real — deep reasoning could consume the entire token budget and
return an empty answer — but that cause had already been fixed by raising max_tokens.
Code and research tasks now run at high, with a fallback that retries once at shallow
effort if the model still runs itself out of budget.reasoning=high, temperature
0.2 returned an empty answer 1 run in 3, and raising the budget from 8192 to 16384
tokens did not help — still 1 in 3. The same prompt at temperature 1.0 answered 3 for
3 and averaged 14s against 36s. Over-constrained, the chain of thought loops instead of
terminating. OpenAI's recommendation of 1.0 for gpt-oss is a reliability setting, not a
style preference.If Ollama is already running when the studio starts, these apply only after it restarts —
studioctl says so rather than reporting a healthy service that is quietly truncating.
The Video tab is the one part of this studio that is not freely usable everywhere, and the restriction is unusual, so it is worth stating plainly.
MiniMax H3 is released under the MiniMax H3 Community License Agreement, which defines an "Applicable Territory" of worldwide, excluding the Excluded Territories. The excluded territories are:
the European Union · the United Kingdom · the United States · the Republic of Korea
Two things that are easy to miss:
If you live in an Excluded Territory you can apply for individual authorisation:
Licence request form → https://platform.minimax.io/h3-license
MiniMax have confirmed that individuals may apply — enter Personal/None where the form asks
for a company name. Their official licence Q&A thread, which is the source for the
clarifications above, is at
https://huggingface.co/MiniMaxAI/MiniMax-H3/discussions/12.
Status for this repository's maintainer: individual authorisation for local deployment has been granted by MiniMax under a separate agreement. That authorisation is personal to the maintainer and does not transfer to anyone who clones this repository.
Until you hold your own, treat the Video tab as local and personal only — do not redistribute, publish or commercially use its
output. Every other tab in this studio is unaffected. The Video tab carries this same notice
in-app so it cannot be missed. The setup package now installs it (SETUP_FOR_CLAUDE.md
PHASE 5h), which is why that phase opens by putting this notice in front of the user and asking
before it spends the ~63 GB: skipping it leaves the rest of the studio fully working.
If you are redistributing Local AI Studio, you must pass this notice on to your users; their obligations depend on where they are, not where you are.
Local AI Studio is released under the Apache License 2.0: use it, change it and build on it, commercially or not, as long as you keep the licence and NOTICE with it and mark the files you changed.
That covers this project's own code and docs only. The models and tools the setup runbook downloads are not part of it and keep their own licences, several of which are personal/research-use only and one (MiniMax H3) territory-restricted. See Credits & licensing and the MiniMax H3 notice above before using any output commercially.