Krea 2, MiniMax & Klein 9B LoRA - LoKR Studio — train, profile, repair, and extract Krea 2, Flux 2 Klein 9B & MiniMax LoRAs & LoKRs
341
stars
1,492
commits
Python
primary language
Sep 5, 2026
updated
Fine-tune base models on consumer GPUs — down to 16 GB. Fix broken LoRAs without retraining. Remix any LoRA into new variations in seconds.
A train · fine-tune · repair · explore workbench built end-to-end for Flux 2 Klein 9B, Krea 2 and MiniMax H3 — training on photos, video, sound and voices, from quick LoRAs to the full base model.
No GPU, or want a bigger one? Fizgig runs on rented hardware — one click, nothing to install.
Deploying through that link supports Fizgig's development at no extra cost to you.
Start-to-finish walkthrough — install, prep, caption, train, and the workbench tools
📰 Latest news
- 🧪 Fizgig 5.0 — Full fine-tuning graduates: train the MiniMax H3 and Krea 2 base models themselves, on consumer GPUs down to 16 GB. No adapter, no rank bottleneck — full-rank updates that change how the model represents a concept instead of filtering its output. One checkbox applies the whole recipe and the planner sizes the run to your card; photos, voice and video clips all fine-tune (2.3 s clips confirmed by measured runs on every tier, longer with video on the likeness blocks), and the built-in Checkpoint to LoRA utility turns the result into an ordinary shareable file — rank 64 was perceptually indistinguishable from the full checkpoint. Experimental; NVIDIA only for now. Details ↓ · Release notes
- Fizgig 4.3.1 — 12 GB cards confirmed training MiniMax H3 — a community field report on an RTX 5070 proved H3 LoRA training runs stable at 12 GB, and the two crashes in its way are fixed: checkpoint saves no longer die on low memory, and previews no longer fragment VRAM into a next-step OOM. Also in this maintenance release: captioning no longer slows down your next training run (it runs in its own process now — built by @scryptio). Release notes
- Fizgig 4.3 — AMD Radeon support arrives — Fizgig now trains on AMD with ROCm (RDNA1 through RDNA4, Strix Point / Halo, Instinct). Windows is the supported path with its own one-click installer; Linux is experimental. Built by @scryptio and tested in the open by the community. Also in the release: identity distillation now fits 16 GB cards — the 32B text encoder streams layer by layer, contributed by @rintic-13 — the Repair Studio gains a side-by-side compare view with likeness and quality metrics, and Fizgig speaks Korean via a community add-on by @ssain3d-lgtm. Details ↓ · Release notes
- Fizgig 4.2 — the workbench opens to MiniMax H3, and what it found ships as features — all five post-training tools now work on H3 LoRAs, with previews rendered as 22-frame clips judged by their middle frame. Using those tools on real LoRAs produced the first H3 block map — and its biggest finding is now Optimised Likeness Learning, a default-on checkbox that trains photos on the identity blocks only: sharper, more prompt-responsive, better sound, fewer epochs. Plus a ✨ MiniMax H3 Style preset, an Append Transcription button that Whispers a clip's speech into its caption, and fully offline transcription. Details ↓ · Release notes
- Fizgig 4.0 — video, sound and voices — MiniMax H3 now trains on video clips, on their sound, and on voice recordings alone: photos, clips and voice files in one folder train one LoRA in one run. Gizmo, a new bundled prep tool, cuts to-spec clips from any footage, auto-chops long videos at scene cuts, and records a voice dataset from nothing but a mic and ten minutes of reading. Training previews render in 6 steps with the Turbo LoRA and can carry their generated sound, opening in the gallery as playable clips. And 16 GB / 24 GB cards now train on the accurate int8 base — block swap streams one-way, ~6× faster, contributed by @rintic-13. Details ↓ · Release notes
- One-click cloud training on RunPod — no GPU, or want a 5090 for the afternoon? The official Fizgig template deploys the full app to a rented GPU in your browser: nothing to install, your files persist until you terminate the pod, and the in-app RunPod panel can even auto-stop the pod when your run finishes so an idle GPU never bills overnight. ⚡ Deploy → · Guide
Every trainer makes LoRAs. Fizgig is built around what you do with them afterwards — and that's the part nobody else has.
.safetensors.Under the workbench sits a fast, light trainer tuned to fit your GPU: a full Klein 9B LoRA trains on 16 GB, the 12.9B Krea 2 on 8 GB, and the 33B MiniMax H3 on 16 GB — block swap, quantisation and previews all size themselves to your VRAM automatically, and if a preview can't fit, training keeps running and saving. It loads kohya / PEFT / OneTrainer / AI-Toolkit / LyCORIS LoRAs, auto-converted, and saves kohya .safetensors that drop straight into ComfyUI.
Free and open source. A good first run: pick a ✨ built-in preset on the Training tab and go.
Each tool works on a trained run's output or any LoRA you've downloaded — and they hand off to each other (profile → repair → explore → compare, one closed loop). All three families: Klein, Krea 2 and MiniMax H3 (H3 previews render a short clip, judged by its middle frame — the model's native regime).
A live slider per transformer block (32 on Klein, up to 50 + the token refiners on MiniMax H3) with a side-by-side preview that updates as you drag. Turbo Preview caches per-block activations so late-block edits redraw up to 97% faster; the baked save is always exact. Blend blocks from a second donor LoRA, balance the pair per block, condition previews on a reference photo, and save a .safetensors that works in ComfyUI at strength 1.0.
Evolutionary discovery: the app mutates blocks and shows four variants — pick a favourite and it becomes the new baseline. Freeze what you like, set how far composition drifts, cycle seeds — and send any baseline to Repair Studio (and back) with one click.
Point it at a training run and it renders every epoch on one fixed seed, with a crossfade slider — drag until it looks best and stop. An optional likeness score (ArcFace, CPU) rates each epoch against a training photo and jumps you to the best. Then make it shareable: epoch-morph clips, seed / prompt / strength travels, a comparison sheet (with/without-LoRA grid, same seed per row), all exportable as looping MP4/GIF with an optional deflicker pass. Works on any folder of LoRAs, or a single file.
A per-block activation profile as a colour-coded HTML report — which blocks carry style, identity, and detail, and where they overlap. Repair Studio reads its sidecar automatically and shows the findings inline when you load the same LoRA.
Distil any Klein, Krea 2 or MiniMax H3 LoRA to a lower rank — Fast presets run weight-only SVD with no models loaded; Klein's activation-weighted presets add block and timestep targeting. PEFT and LyCORIS sources supported.
A from-scratch native port: 12.9B single-stream MMDiT, Qwen-Image VAE, Qwen3-VL-4B text encoder. Train on the RAW model; previews render on the training model itself with the official Turbo LoRA (auto-downloads) applied for the render only. Pick Krea 2 from the Base Model selector on the Training tab and the ✨ Krea 2 Defaults preset applies itself.
Everything works on Krea 2: all five workbench tools, Pause/Resume, Context LoRA, Adaptive LR, reference images, the live sample override — and LoKR training (pick it from Network Type; factor 8 or below for the quality edge, standard LoRA is ~20% faster). Output is ComfyUI-ready.
8 GB is enough. Users train full Krea 2 LoRAs on 8 GB with everything on Auto and batch size 1. Auto reads your free VRAM and picks INT8, NF4 or fp8 plus the right block swap — the console explains its choice. On longer runs the transformer blocks torch.compile automatically for roughly 2× faster steps.
Four Training-tab toggles no other trainer has:
Edit any caption yourself mid-run from the Problem Images window — no restart. When nothing is improving any more, a plateau banner names the best-checkpoint window to scrub in LoRA Royale. Pause, resume, restart: a resumed run replays its own loss log and loses nothing.
📣 Help map Krea 2's blocks — open an issue. Krea 2's per-block roles aren't charted yet, which is why the colour-coded sliders and layer targeting are Klein-only for now. The Profiler's weight-only report is the instrument — share what you find and it drives the presets and Repair Studio colour-coding to come.
Fizgig trains LoRAs for MiniMax H3, MiniMax's open-weight ~33B video model, from ordinary still-image datasets — and from short video clips, their sound, and voice recordings (details ↓) — on a single consumer GPU. Output loads straight into ComfyUI's H3 workflows, including the pruned inference builds.
The full studio, as of 4.2. H3 trains, previews and pauses/resumes like the other families — and all five workbench tools now work on H3 LoRAs too, with previews rendered as short clips judged by their middle frame. It was those tools, on real LoRAs, that produced the block map behind Optimised Likeness Learning below.
How it works: pick MiniMax H3 from the Base Model selector and the usual flow applies — Start-tab folder, Captions, Samples, Training. Leave Blocks Swap and Base Precision on Auto: at launch the trainer reads your free VRAM (close ComfyUI first) and picks the base precision and block-swap count together:
| Free VRAM | What Auto does |
|---|---|
| ~30 GB | int8, no block swap, up to 1 MP |
| ~22 GB | int8, ~14 blocks streamed |
| ~15 GB | int8, ~36 blocks streamed |
| ≤12 GB | 4-bit, as before |
int8 is the checkpoint's own storage and the most accurate base (~0.17% error). Block swap streams one way only — ~6.4× faster than round-trip swap, which is what lets 16 and 24 GB cards keep the accurate base (design contributed by @rintic-13, #73). Hit an OOM anyway? Set Blocks Swap to a number to override the planner.
Three built-in presets ship; Fast applies the moment you pick the family:
| Preset | Settings |
|---|---|
| ✨ MiniMax H3 Fast | LoRA dim/alpha 8, 50 epochs, flat 2e-4, 0.25 MP, Training Structure Likeness and Style, adamw. Reaches likeness in a few hundred steps, and the lower rank tends to come out more flexible |
| ✨ MiniMax H3 (Lower LR - slower) | The same at rank 16, 60 epochs, flat 1e-4 — more suitable for larger datasets with longer trains |
| ✨ MiniMax H3 Style | The Fast recipe on the measured style blocks, 0-3, 6-47 — style lives almost everywhere in H3 except the few blocks that only do identity and voice |

Optimised Likeness Learning ships ticked (Fast and Lower LR; Style unticks it): photo steps train only the identity blocks (20-49) while video and audio clips train the full model. Measured against full-model photo training: sharper, much better prompt following, better sound, fewer epochs — and the occasional deformed preview of full-model photo runs is gone. Untick it for style or scene training; while it's on, Blocks to Train is disabled with a note.
0.25 MP is the default, and it holds up — four times cheaper per step than 1 MP, and the extra resolution has not paid for itself in testing. Raise it if a specific dataset asks for it.
Previews default to 768×768, 56-frame clips with sound — a short watchable clip with the model's generated audio, opened in the gallery as a playable video (never autoplay). Without the audio VAE set, clips render silent; stills and other lengths stay in the dropdown. Set the Turbo LoRA in Preferences and previews render in 6 steps instead of 20 — previews only, never the saved LoRA. A preview that outgrows VRAM steps itself down a ladder rather than dying — a shorter clip first, then resolution to a 512×512 floor — and the size that fit is saved as the new default.
…train on video clips? Cut them with Gizmo (launch it from the Image Prep tab, or the Launch Gizmo .bat) — it exports clips already on H3's spec — drop them into the training folder next to your images, and caption them on the Captions tab like a photo. Photos, clips and voice recordings all train together in the same folder — no settings, no separate runs.
…make clips from my footage? Open Gizmo, drop a video on it, scrub to a moment, pick a length, Add to queue — repeat, then Export queue.
…chop a long video automatically? Gizmo's ✂ Auto-chop scene-detects the whole source and offers every segment as a thumbnail — click to keep or skip, and the keepers join the queue.
…train a voice from a recording? Gizmo's Voice tab: open any audio file (or a video, for its soundtrack), mark segments on the waveform, caption the sound, export — segments come out training-ready with their captions beside them.
…record a voice dataset from scratch? Voice tab → 🎙 Record: read the prompted sentences while holding the button (or the R key). Every take arrives trimmed and captioned; ten minutes of reading is a usable dataset.
…keep a clip's sound out of training? Mute it in Gizmo — it adds _mute to the filename, reversible by renaming. The video still trains.
…train photos, clips and a voice into one LoRA? Same folder, one trigger word, one run, any mix. If one category is much smaller, Finish one category early on the Training tab lets it finish at a chosen epoch while the rest trains on.
…get fast previews while training? Set the Turbo LoRA (~780 MB, its own Preferences row): 6-step previews with the Turbo at 75% on top of your training LoRA. Adjustable on the Samples tab.
…hear what it's generating while training? Pick a "with sound" Sample length on the Samples tab. Each preview carries its generated soundtrack, playable in the gallery.
…get a clip's spoken words into its caption? Open it in the caption editor (Captions tab → click the clip): any non-muted video shows an 🎤 Append Transcription button that Whispers the speech into the caption as saying "…" — Gizmo's grammar, without leaving the tab.
…set it up? One extra model file for sound: the audio VAE (~605 MB), on its own Preferences row. Blank = clips train silent; required only once the folder has voice recordings. Fizgig points out both new files once at startup if your H3 paths are set.
Stills teach H3 a look; clips teach it motion, and clips with sound teach it a voice. Clips cost far more per step than stills — start with a handful. Drop .mp4 clips into the training folder alongside your images and caption them like photos. A clip has to be on spec, and Fizgig refuses one that isn't rather than quietly fixing it:
| Requirement | |
|---|---|
| Container | .mp4 |
| Frame rate | exactly 24 fps |
| Frame count | 5, 22, 39, 56, 73, 90, 107 or 124 frames |
| Dimensions | multiples of 32 |
| Audio | 32 kHz stereo, or no track at all |

Gizmo makes clips that hit it — mark every section you want (frame-accurate stepping, first/last-frame previews, a ▶ Play of the exact clip), then export the lot in one go. Crop to the subject: a clip's cost is its pixels, so drag a rectangle and every token goes on what you want learned — with shape locks (1:1, 16:9, 9:16…) when you want consistent framing. High-frame-rate footage can keep extra frames as slow motion, offered as a choice. Clips are cut at native resolution and resized to your Target Megapixels at training time, so cutting large keeps the choice open.
What it costs: 22 frames is the shortest that shows real movement at ~7× a still per step; 124 frames is ~37×. Gizmo says which lengths your card can train, at which megapixels, before you cut anything:
| Clip | 16 GB | 24 GB | 32 GB |
|---|---|---|---|
| up to 56 frames | up to 0.25 MP | up to 0.5 MP | up to 0.5 MP |
| 73–90 frames | — | up to 0.25 MP | up to 0.5 MP |
| 107–124 frames | — | up to 0.25 MP | up to 0.25 MP |
Drop .wav / .mp3 / .flac / .m4a files into the training folder — alone or mixed with stills and clips. Rate and channels are converted for you; duration is the strict part:
| Requirement | |
|---|---|
| Formats | .wav .mp3 .flac .m4a — any rate or channel count |
| Duration | exactly 0.917, 1.625, 2.333, 3.042, 3.750, 4.458 or 5.167 s (±25 ms) |
| Content | actual sound — digital silence is refused |
| Caption | a .txt beside the file, or it silently won't train |
| Audio VAE | required — the ~605 MB Preferences row |

Gizmo's Voice tab cuts them for you — open a recording (or a video, for its soundtrack), mark segments on the waveform, pick a length, caption, export sample-exact. Caption the voice, not a picture — "a man speaking calmly, low pitch, unhurried" — with your trigger word leading; the Transcribe button (Whisper) appends the spoken words. Or record the dataset from scratch: 🎙 Record prompts sentences across every length and five tonal flavours, rolls a delivery style per take, and every hold-and-release lands trimmed, captioned and ready to queue. Set Training Structure to Likeness and Style for voices — tested head-to-head, it converges much faster; Fizgig reminds you when it sees voice files.
Each has a Download link on its row in Preferences:
| Model | Size | Notes |
|---|---|---|
| DiT — pruned int8 | ~21 GB | The training base — minimax_h3_fl2va_pruned_int8_convrot.safetensors, the same file ComfyUI runs. (The ~66 GB bf16 file also works for LoRA training, NF4 at load — but full fine-tuning needs this int8 file) |
| Qwen3-VL-32B text encoder | ~15.7 GB | The nvfp4 file — same one ComfyUI uses. Loaded once for caching, then freed |
| Video VAE | ~4.9 GB | Caching and preview decode |
| Audio VAE (optional) | ~605 MB | Sound training and previews with sound |
| Turbo LoRA (optional) | ~780 MB | 6-step previews — minimax_h3_turbo_v4_step600.safetensors; you may have it in ComfyUI's loras folder |
| DiT — reference (optional) | ~21 GB | Only for reference distillation (ref2va) |
Yes, you train on the pruned file. "Pruned" here swaps the AdaLN modulation MLP for a curve table — that branch only sees the timestep, so nothing a LoRA learns lives there. You train against the exact weights you deploy on.
Every control has a hint in the app; the highlights:
20-49 for likeness, 0-3, 6-47 for style (the Style preset sets it), voice core 38-48. Type ranges (3-12, 22, 31-33) to experiment beyond them.minimax_h3_turbo_v4_step600_ema the strongest checkpoint.Settings are read at launch; Pause → Resume relaunches with your current settings, so a pause is the moment to change them mid-run.
Everything above trains a LoRA. This trains the base model itself — no adapter, no rank bottleneck — on a single consumer GPU. Tick ⚗ Fine-tune the BASE MODEL instead of training a LoRA on the Training tab.
A note on where this is at. I first got fine-tuning working on Krea 2 shortly after its release, and I've been deliberately cautious about shipping it — first proving it to myself, then refining it through the MiniMax H3 work. This is the point where it needs the community to develop further. I don't expect every scenario to work perfectly yet — but it works, the numbers below are measured, and there's a solid foundation here to build on. Field reports genuinely shape what gets built next. I'm also aware this technique is model-agnostic at heart — it opens the door to fine-tuning other models, and I'm open to going there. But for that to happen it needs practical community support around those models — code, PRs, testing, that kind of thing — so I have the time necessary to make it happen. — Peter
New to fine-tuning? The extended "How do I…?" guide answers everything this section can't fit — including five-minute recipes for both families: tick Fine-tune, let the settings switch themselves, and change almost nothing.
One idea makes everything else here make sense: an "epoch" trains one slice of the model. The trainable window rotates each epoch, so it takes a full cycle — typically 4 epochs — for every part of the model to train once. Rule of thumb: 4 fine-tune epochs ≈ 1 true epoch of the whole model. That's why the epoch defaults look high, and why saves land on cycle boundaries — each saved checkpoint is a whole, evenly trained model.
Note on VRAM: the "trains on 8 GB" figures elsewhere in this README are for LoRA training. Full fine-tuning is a different animal — but it now tiers itself to your card, and fine-tuning defaults to a 4-bit NF4 frozen base that halves the model held on the card: on 32 GB and 24 GB the classic full-depth windows stay resident at full speed, and on 16 GB the frozen blocks stream from system RAM — slower steps, but the same component-mode learning. The planner measures your free VRAM at launch and prints the plan it chose.
What can my card fine-tune? The short answer, at the default training resolution:
| Your card | Krea 2 — photos | MiniMax H3 — photos | H3 — voice | H3 — video, confirmed | H3 — video on likeness blocks, expected |
|---|---|---|---|---|---|
| 16 GB | ✅ | ✅ | ✅ | ✅ up to 2.3 s | up to 3.8 s |
| 24 GB | ✅ | ✅ | ✅ | ✅ up to 2.3 s | up to 5.2 s |
| 32 GB | ✅ | ✅ | ✅ | ✅ up to 3.8 s | up to 5.2 s |
A few things worth knowing about that table: clip lengths follow Gizmo's grid, so 2.3 s means the 56-frame slot — cut your clips there and everything fits, confirmed by measured runs on every tier. On 32 GB, 3.8 s is also confirmed, even with video training the whole model. Beyond that, the Restrict video to likeness blocks tickbox (on by default with Optimised Likeness Learning — in our tests it trains video just as well, and it makes clips far lighter) extends the expected range: up to 5.2 s on 24 GB and 32 GB, and 3.8 s on 16 GB — conservative arithmetic from the measured constants, not yet individually measured, so treat those as expected rather than promised. Whole-model 5.2 s clips need more than 32 GB (measured). With the restriction unticked, one clip anywhere in your folder trains the whole model, so a mixed photos + clips dataset uses the clip column. And 12 GB cards train LoRAs, not fine-tunes — 16 GB is the fine-tune floor.
"A full fine-tune of a 12.9B–33B model on 16 GB" sounds like a trick, so here's the arithmetic. Only one slice of the model is ever trainable at a time — the trainable window rotates each epoch, so gradients and optimizer state exist for that slice alone. The frozen rest is held 4-bit (NF4) at half size and, on 16 GB, streamed from system RAM. The bf16 master copy lives in CPU RAM, never on the card. Those three together are the whole magic, and the numbers are measured, not projected: 8.8–12.3 GB peaks on a 16 GB card for H3, 8.4–11.0 GB for Krea 2 — and the console prints your own run's peak every epoch, so you can watch the claim hold live. Mechanism, tiers and every "how do I" in the extended guide: docs/FINETUNE_HOWDOI.md.
Which model files. Fine-tuning uses the same training bases you already have — nothing new to download:
krea2_raw_bf16.safetensors, ~26 GB), the same
file LoRA training uses. The fp8 Turbo is the preview model and can't be fine-tuned.minimax_h3_fl2va_pruned_int8_convrot.safetensors, ~21 GB) — again the same file the LoRA
path trains against and ComfyUI runs. The ~66 GB bf16 file, which LoRA training accepts, does
not work for fine-tuning; the trainer refuses it with a clear message.A finished fine-tune checkpoint is itself a valid base for either family — point the model path at it to train further (the console prints the exact continuation settings at every save). And — easy to miss — you can set it as the family's base in Preferences and train LoRAs on top of your own fine-tuned model: teach the base your world or cast once, then quick LoRAs for individual subjects ride on it. Deploy those LoRAs with the same fine-tuned base in ComfyUI. And Pause / Resume works on a fine-tune: Pause saves a full checkpoint even between the regular save epochs, and Resume continues it — rotation window, checkpoint numbering and the remaining epoch count all carry over.
Why bother. A LoRA constrains every update to a low-rank subspace, so concepts compete for the same handful of directions. That's why LoRAs tend to drag pose, framing and lighting toward the training set along with the likeness — they behave a bit like a filter over the model's output. A full-rank update can change how the model represents a concept, so it composes with what the model already knows. In our own tests, multi-character and concept teaching seemed to land at a much deeper level than LoRA training, with much better results — and the built-in Checkpoint to LoRA converter turns the result into a shareable file, and works very well. Beyond that, we're deliberately letting the community find the ceiling.
How it fits. A naive full fine-tune of Krea 2 (12.9B) needs roughly 78 GB — bf16 weights, gradients and optimizer state at once. Rotating windows make only part of the model trainable at a time, advancing each epoch, so gradients and optimizer state only ever exist for the active slice. Over a full cycle every weight trains. Around that sit three decisions that do the heavy lifting: a CPU-resident bf16 master copy is the source of truth, so training never round-trips through fp8 and quantisation can't erase the small updates being learned; optimizer-in-backward consumes and frees each gradient the moment it lands (worth 5.2 GB); and Adafactor's factored state is ~10× smaller than AdamW's.
It sizes itself to your card. Leave Window on Auto (by VRAM) and Fizgig measures the memory actually free at launch, picks the largest window that fits, and prints what it chose and why. Measured Krea 2 peaks (RTX 5090):
| Window mode | Peak VRAM | Speed | Fits |
|---|---|---|---|
| component + 4-bit NF4 (the default) — full-depth windows, resident | ~16 GB (24 GB budget) / ~21–23 GB (32 GB, more headroom held) | ~1.0 s/it | 24 GB and up |
| component + 4-bit NF4 + streaming | 8.4–11.0 GB | ~2.8 s/it | 16 GB |
| component on the fp8 base (explicit Base-precision pick) — depth-split + streamed | 15.6–17.6 GB | ~3.0 s/it | 24 GB |
4-bit NF4 is the fine-tune default, and you don't have to do anything to get it. It halves the frozen base, which on a 24 GB card is enough to keep the classic full-depth component windows resident instead of depth-splitting and streaming them: 4 windows instead of 8 — a full pass over every weight in 4 epochs rather than 8 — at roughly 3× the step speed (measured ~1.0 s/it against ~3.0 s/it for the fp8 base, same dataset, same 24 GB budget). On 16 GB it is the only base that fits at all.
The trade is that the frozen part of the model is held more coarsely while the trainable window learns against it. Your saved checkpoint is unaffected either way — it's written in bf16 from a master copy that never passes through a quantiser. If you want the more accurate frozen context and have the VRAM, pick fp8 under Base precision and it will be used.
Component is the best mode — and Auto now stays in it at every depth. Every window spans the model's full depth — attention across all 28 blocks, then each MLP matrix in turn — so a concept is learned by every layer at once rather than one depth slice at a time. The text-fusion stack stays trainable throughout: rotation would never reach it, and it's where prompt-to-concept binding happens. Where the budget used to force a mode change, the planner now depth-splits the windows instead (a fat window trains in slices — more windows per cycle, still full speed), and below that the frozen out-of-window blocks stream from system RAM — slower steps, but still component-mode learning. The console prints the chosen plan and why.
Block mode remains an explicit Window-dropdown choice — contiguous depth slices with frozen blocks streamed, slower than component at every budget. It's not yet quality-tested; every good result so far came from component runs.
The same checkbox under the MiniMax H3 family fine-tunes the 33B model, with the recipe adapted to it: component windows only — each window trains one attention or MLP matrix across all 50 blocks (4 windows per cycle), with the token refiner trainable throughout, so every window spans the model's full depth from the very first epoch.
mlp.fc1 trains in two slices, a 5-window cycle,
still full speed, no offloading; measured peaks 19.1–21.5 GB. On 16 GB the frozen
out-of-window blocks also stream from system RAM (~7 GB staged, a 9-window cycle):
measured peaks 8.8–12.3 GB at ~1.5× the step time — a full fine-tune of a 33B video model
on a 16 GB card. The console prints the chosen plan and why.Saves, previews and numbering run on the rotation cycle, not the Samples tab. The save
cadence snaps to cycle boundaries — the Save-every box follows the FT controls live in the GUI,
and the trainer snaps it again at launch — so every checkpoint compares like-for-like, with each
window trained equally. Previews ride the saves: one render per saved checkpoint plus the final
one, overriding the Samples tab's "every N epochs" (prompts, resolution, seed and the live
sample override still come from the Samples tab and status bar as usual — every sample in the
gallery maps to a file you can deploy). Checkpoints are numbered by epoch (-000004,
-000008, …) and the numbering continues across Pause/Resume, so a resumed run never overwrites
an earlier save. Krea 2 fine-tunes behave exactly the same way — saves snap to the cycle,
previews ride them (rendered on the training DiT with the Turbo LoRA), numbering carries over.
The output is a normal H3 checkpoint: load it in ComfyUI directly, or run Checkpoint to LoRA
(run_diff_to_lora.bat in your Fizgig folder) on it (the extractor decodes the int8 format natively) for a shareable LoRA.
If you're coming from LoRA training, recalibrate before anything else: fine-tuning wants much lower learning rates than LoRAs. A LoRA nudges a small adapter riding on a frozen model; a fine-tune moves the model's own weights, so the rates you're used to typing land very differently here — what's a normal LoRA rate can wreck a fine-tune outright.
Full fine-tuning moves every weight, so a long run on a handful of subjects drifts the model's whole notion of people — there's no low-rank bound to limit it the way there is with a LoRA. Point Regularisation images at a folder of ordinary photos of the broader class and they train at a reduced learning rate (LR ×, default 0.2) as an anchor rather than a lesson. That multiplier is a real dial, not a set-and-forget: 0.1–0.3 tethers the model's prior while your subject trains; push it toward 1.0 and the reg set trains like a second subject set — class-balanced training rather than a light anchor, which is a different (valid) thing. If a fine-tune drifts the broader class, raise it a step; if the subject learns too slowly, lower it. Worth a little experimentation per dataset.
Use real photos, not model output — anchoring a fine-tune to its own samples distils its artifacts back in, and there's nothing bounding that drift. Caption them normally: anything you leave unsaid gets attributed to the class word itself. Leave the folder empty to train without one.
A fine-tune produces a ~26 GB checkpoint, which is not what anyone wants to share. The
Checkpoint to LoRA utility — run_diff_to_lora.bat in your Fizgig folder, which opens
its own small window separate from the main app (Linux/pods: ./run_diff_to_lora.sh) — takes the base model
you started from and the checkpoint you produced, and extracts the difference as an ordinary
kohya .safetensors — at several ranks at once, since one SVD per layer serves them all.
This turned out to work far better than expected: rank 64 was perceptually indistinguishable from the full 26 GB checkpoint at ~0.5 GB, and quality degrades smoothly at lower ranks rather than falling off a cliff.
The result worth knowing: in our testing, a LoRA extracted from a fine-tune came out better than a LoRA trained directly at the same or higher rank on the same dataset. A low-rank file can hold a solution that low-rank training struggles to find — so fine-tune-then-extract isn't a workaround; the full-rank phase is the mechanism, and the extraction is nearly free.
Being straight about the trade-offs, because they're real:
The foundation: fast, light, and tuned for one model.
Loads kohya, PEFT, OneTrainer (OMI + legacy), AI-Toolkit, and LyCORIS (LoKR / LoHa) — auto-converted, and LoKR/LoHa run natively everywhere: Repair Studio, Profiler, Extract, Context LoRA. Repair Studio and Explorer save LoKR as LoKR, losslessly. Output is .safetensors that drops straight into ComfyUI.
Fizgig ships as a ready-made cloud image — the whole app in a browser tab, not a cut-down web version. Drag datasets in and LoRAs out with a built-in file manager, download models in one click, and optionally have the pod shut itself down when training finishes. Your models and datasets persist between sessions.
⚡ Deploy on RunPod → · Read the guide first
install_fizgig_rocm.bat (supported path). Linux: ./install_fizgig_rocm.sh — highly experimental (newer gfx like RDNA4, desktop compositor + training on the same GPU, and driver resets are common; use Windows ROCm or NVIDIA Linux for production training). Optional system amdrocm-amdsmi for accurate status-bar VRAM via amd-smi.Clone the repo:
git clone https://github.com/shootthesound/Fizgig.git
cd Fizgig
Clone it rather than downloading the ZIP — update_fizgig.bat updates by pulling with git, and a ZIP can't.
Open a terminal in your Fizgig folder and run:
git init
git remote add origin https://github.com/shootthesound/Fizgig.git
git fetch --depth 1 origin master
git reset --hard FETCH_HEAD
git branch -M master
git branch --set-upstream-to=origin/master master
Your model paths, output LoRAs, caches, presets and the venv are all left alone. update_fizgig.bat works normally from then on.
Windows (NVIDIA, one-click) — double-click install_fizgig.bat. It creates a venv, installs CUDA 12.8 PyTorch and all dependencies, pre-downloads the InsightFace models, and verifies CUDA is visible to PyTorch. Launch with run_fizgig.bat; update later with update_fizgig.bat.
Windows (AMD ROCm) — needs a full Python 3.12 install first (the ROCm bitsandbytes wheel is cp312-only; Fizgig's GUI needs Tkinter). Do not use the embeddable zip. Install from Windows downloads:
py install 3.12.Then double-click install_fizgig_rocm.bat (NVIDIA users never run this). It picks 3.12 via py -3.12 / python3.12 (not whatever python defaults to — e.g. 3.14). GPU detection follows, then pinned multi-arch wheels from AMD ROCm nightlies (https://rocm.nightlies.amd.com/whl-multi-arch/ — not built by Fizgig):
torch==2.12.0+rocm7.15.0a20260728torchvision==0.27.0+rocm7.15.0a20260728rocm-sdk-devel==7.15.0a20260728Override with TORCH_PIN / TORCHVISION_PIN / ROCM_SDK_DEVEL_PIN if needed. bitsandbytes is a pinned community Windows ROCm wheel from 0xDELUXA/bitsandbytes_win_rocm — built by neither AMD nor Fizgig. Shared deps come from requirements.txt with CUDA torch/bitsandbytes and NVIDIA-only nvidia-ml-py filtered out (filter_requirements_rocm.py). Launch with run_fizgig_rocm.bat; update later with update_fizgig_rocm.bat (not update_fizgig.bat — that script installs CUDA torch and would wipe the ROCm stack).
--experimental (unsupported): install_fizgig_rocm.bat --experimental installs unpinned torch[device-ARCH] / torchvision[device-ARCH] / rocm-sdk-devel from the same multi-arch index and leaves BNB_ROCM_VERSION unset so bitsandbytes auto-selects its highest matching DLL (no fallback warning while the resolved torch stays inside the wheel's HIP 7.13-7.16 range). This is not the same as Linux ROCM_CHANNEL=nightly (which stays on the constrained 7.14 / bitsandbytes 714 lane). Windows already installs from AMD nightlies with pinned versions by default; --experimental only drops those pins. Local experimentation only. Do not open GitHub issues for crashes, install failures, or training problems when --experimental was used — those reports will not be supported. Use the pinned install (no flag) for anything you expect help with.
Linux (AMD ROCm — highly experimental) — expect crashes, GPU resets, and incomplete model support on many setups. Best-effort only; Windows ROCm or NVIDIA Linux are the supported training paths. Prerequisites: amdgpu driver loaded (/dev/kfd), user in render/video groups. See Install ROCm and PyTorch for ROCm. Then:
chmod +x install_fizgig_rocm.sh
./install_fizgig_rocm.sh
./run_fizgig_rocm.sh
The script detects your gfx target (detect_gpu_linux.py). Nightly is the Linux default — TheRock multi-arch RELEASES.md index plus a [device-gfx*] extra for your GPU (e.g. gfx1201 → device-gfx1201). Unpinned nightly resolves the latest torch 2.12 + ROCm 7.14.0a* stack (matches libbitsandbytes_rocm714.so). Override with TORCH_PIN=…, ROCM_META_PIN=…, or TORCH_NIGHTLY_MINOR=….
Stable (repo.amd.com, no nightly alphas): pin torch==2.12.0+rocm7.14.0 + rocm-sdk==7.14.0 (cp310–cp314):
ROCM_CHANNEL=stable ./install_fizgig_rocm.sh
Try torch 2.14 (nightly only today — can increase sampling VRAM pressure vs 2.12):
ROCM_CHANNEL=nightly TORCH_NIGHTLY_MINOR=2.14 ./install_fizgig_rocm.sh
# or an explicit pin, e.g.:
# TORCH_PIN=2.14.0a0+rocm7.14.0a20260625 ROCM_CHANNEL=nightly ./install_fizgig_rocm.sh
# (paired torchvision ~0.29.0a0+rocm7.14.0a… — installer resolves the match)
Linux ROCm cache scripts import fizgig.rocm.cache_exit only when FIZGIG_GPU_BACKEND=rocm (set by run_fizgig_rocm.sh); NVIDIA and other platforms call main() unchanged. Opt out: FIZGIG_ROCM_NO_FAST_EXIT=1 ./run_fizgig_rocm.sh.
Then shared deps from requirements.txt (filtered) and bitsandbytes>=0.50.0 for ROCm.
Linux / macOS (NVIDIA CUDA path) — install_fizgig.py is CUDA-only (captioning / image prep on macOS; training needs a CUDA or ROCm GPU). On AMD-only Linux hosts it prints a hand-off to the ROCm installer and exits:
python install_fizgig.py
chmod +x run_fizgig.sh
./run_fizgig.sh
VRAM status bar on AMD: the existing NVIDIA pynvml / nvidia-smi path is unchanged; AMD readers (vram_monitor.read_amd_gpu_vram) run only as a fallback. Windows ROCm uses typeperf; Linux ROCm uses the amd-smi CLI when available (AMD SMI / ROCm Core SDK, e.g. sudo apt install amdrocm-amdsmi). Fizgig picks the GPU with the largest VRAM total (skips empty iGPU entries). Legacy rocm-smi is a fallback. Do not pip install amdsmi — the PyPI package is outdated.
buffalo_l (~300 MB, during install), Florence-2 (~500 MB–1.5 GB, first AI caption), and Helsinki-NLP opus-mt-en-zh (~300 MB, first bilingual translation).Fizgig doesn't bundle weights. You only need the family you're using — and Preferences has a ⬇ Download models for me button under each model card that downloads, verifies, and fills in the paths (Klein needs a free HuggingFace token for BFL's licence; Krea 2 needs no account). Every row also has a manual Download link. CLI:
python -m fizgig.scripts.fetch_models --family krea2 # ~32 GB, no account needed
python -m fizgig.scripts.fetch_models --family klein # ~34 GB, needs a token
python -m fizgig.scripts.fetch_models --family tools # Florence-2, face model, translator
| Model | File | Size | Source |
|---|---|---|---|
| Base DiT (fp8) — recommended | flux-2-klein-base-9b-fp8.safetensors | ~9.5 GB | black-forest-labs/FLUX.2-klein-base-9b-fp8 |
| Base DiT (bf16) | flux-2-klein-base-9b.safetensors | ~17 GB | black-forest-labs/FLUX.2-klein-base-9B |
| Distilled DiT | flux-2-klein-9b-fp8.safetensors | ~9 GB | black-forest-labs/FLUX.2-klein-9b-fp8 |
| VAE / AE | ae.safetensors | ~320 MB | black-forest-labs/FLUX.2-dev (from root, not the vae/ subfolder) |
| Text Encoder | qwen_3_8b.safetensors | ~15 GB | Comfy-Org/vae-text-encorder-for-flux-klein-9b |
Training runs on the Base DiT — the fp8 version is recommended on every GPU (same quality, half the VRAM). The Distilled DiT powers the 4-step previews and the workbench.
All files live in the one Comfy-Org/Krea-2 repo.
| Model | File | Size |
|---|---|---|
| RAW DiT (bf16) — training | krea2_raw_bf16.safetensors | ~26 GB |
| Turbo DiT (fp8) — workbench | krea2_turbo_fp8_scaled.safetensors | ~13 GB |
| Turbo LoRA (auto-downloads) | krea2_turbo_lora_rank_64_bf16.safetensors | ~470 MB |
| Qwen-Image VAE | qwen_image_vae.safetensors | ~250 MB |
| Text Encoder — recommended | qwen3vl_4b_fp8_scaled.safetensors | ~5.2 GB |
| Text Encoder — full precision | qwen3vl_4b_bf16.safetensors | ~8.9 GB |
The text-encoder slot is open: any Qwen3-VL-4B in the ComfyUI layout loads — fp8_scaled (recommended, captions we couldn't tell apart), bf16, or a community fine-tune/abliterated build, which changes how your dataset gets captioned.
MiniMax H3's files are listed in its section above.
Training — the fp8 Base stays resident at ~9.6 GB, so a 9B LoRA fits 16 GB (~14 GB observed). Smaller cards: the 4-bit (NF4) base toggle drops the base to ~5.6 GB — a full LoRA trains in ~7.5 GB, fitting 10–12 GB cards with no swap.
Workbench (Distilled 4-step):
| Block Swap | Min VRAM |
|---|---|
| 0 | 24 GB+ |
| 8 | 16 GB |
| 12 | 14 GB |
| 16 | 12 GB |
On first launch Fizgig auto-detects your VRAM and picks the default; your own choice sticks.
| Your card | What to do |
|---|---|
| 8 GB | Everything on Auto, batch size 1, stock preset defaults |
| 10–12 GB | Same — headroom to raise batch size or resolution |
| 16 GB+ | Same — Auto will usually pick the faster INT8 path |
Auto budgets from your free VRAM and the console explains its choice. If a preview can't fit, previews auto-disable and training keeps running and saving.
See the Auto table in its section — 16 GB and up trains on the accurate int8 base with streamed block swap; ≤12 GB falls back to 4-bit. On 16 GB-class cards, previews cap themselves at 768×640 and 22 frames (sound kept) — larger picks in the menus simply clamp, with a console note.
Turn off Hardware-accelerated GPU scheduling (Settings → System → Display → Graphics → Default graphics settings), then reboot. With it off, Fizgig runs training at low priority so your desktop stays smooth — training speed is unaffected.
Launch Fizgig and work left-to-right through the numbered tabs:
The unnumbered tabs are the post-training workbench: Profiler, Repair Studio, LoRA the Explorer, LoRA Royale, Extract, and Preferences.
One tool lives outside the main window: Checkpoint to LoRA (run_diff_to_lora.bat, or
python diff_to_lora_gui.py) — point it at a base model and a fine-tuned checkpoint and it writes
an ordinary LoRA at whichever ranks you tick. Only needed if you use the experimental full
fine-tune above.
Headless? Everything the trainer does is also available from the command line — see docs/CLI.md.
Community translations — Korean (한국어): Fizgig-Korean-Translated-Ver by @ssain3d-lgtm — an unofficial add-on that translates the UI at runtime without touching any Fizgig files, with a one-script uninstall. If you hit a bug while it's installed, uninstall and reproduce before reporting here; layer issues go to that repo.
If Fizgig saves you time or helps you make better LoRAs, consider supporting development:
Fizgig is open source under the Apache License 2.0 — free to use, modify, and redistribute, including commercially, with attribution and no warranty. Third-party components under compatible permissive licenses (and other terms where noted) are listed in THIRD_PARTY_NOTICES.md.
Copyright © 2026 Peter Neill.
Model weights are not covered by this license — each model carries its own terms from its publisher (see the Download links in Preferences).
I'm available for consulting on local AI training pipelines, custom workflow tooling, and private model work — the same engineering that's in Fizgig, applied to your studio's hardware and IP. Get in touch: peter@shootthesound.com.
Python
98.0%
Shell
1.2%
Krea 2, MiniMax & Klein 9B LoRA - LoKR Studio — train, profile, repair, and extract Krea 2, Flux 2 Klein 9B & MiniMax LoRAs & LoKRs
341
stars
1,492
commits
Python
primary language
Sep 5, 2026
updated
Fine-tune base models on consumer GPUs — down to 16 GB. Fix broken LoRAs without retraining. Remix any LoRA into new variations in seconds.
A train · fine-tune · repair · explore workbench built end-to-end for Flux 2 Klein 9B, Krea 2 and MiniMax H3 — training on photos, video, sound and voices, from quick LoRAs to the full base model.
No GPU, or want a bigger one? Fizgig runs on rented hardware — one click, nothing to install.
Deploying through that link supports Fizgig's development at no extra cost to you.
Start-to-finish walkthrough — install, prep, caption, train, and the workbench tools
📰 Latest news
- 🧪 Fizgig 5.0 — Full fine-tuning graduates: train the MiniMax H3 and Krea 2 base models themselves, on consumer GPUs down to 16 GB. No adapter, no rank bottleneck — full-rank updates that change how the model represents a concept instead of filtering its output. One checkbox applies the whole recipe and the planner sizes the run to your card; photos, voice and video clips all fine-tune (2.3 s clips confirmed by measured runs on every tier, longer with video on the likeness blocks), and the built-in Checkpoint to LoRA utility turns the result into an ordinary shareable file — rank 64 was perceptually indistinguishable from the full checkpoint. Experimental; NVIDIA only for now. Details ↓ · Release notes
- Fizgig 4.3.1 — 12 GB cards confirmed training MiniMax H3 — a community field report on an RTX 5070 proved H3 LoRA training runs stable at 12 GB, and the two crashes in its way are fixed: checkpoint saves no longer die on low memory, and previews no longer fragment VRAM into a next-step OOM. Also in this maintenance release: captioning no longer slows down your next training run (it runs in its own process now — built by @scryptio). Release notes
- Fizgig 4.3 — AMD Radeon support arrives — Fizgig now trains on AMD with ROCm (RDNA1 through RDNA4, Strix Point / Halo, Instinct). Windows is the supported path with its own one-click installer; Linux is experimental. Built by @scryptio and tested in the open by the community. Also in the release: identity distillation now fits 16 GB cards — the 32B text encoder streams layer by layer, contributed by @rintic-13 — the Repair Studio gains a side-by-side compare view with likeness and quality metrics, and Fizgig speaks Korean via a community add-on by @ssain3d-lgtm. Details ↓ · Release notes
- Fizgig 4.2 — the workbench opens to MiniMax H3, and what it found ships as features — all five post-training tools now work on H3 LoRAs, with previews rendered as 22-frame clips judged by their middle frame. Using those tools on real LoRAs produced the first H3 block map — and its biggest finding is now Optimised Likeness Learning, a default-on checkbox that trains photos on the identity blocks only: sharper, more prompt-responsive, better sound, fewer epochs. Plus a ✨ MiniMax H3 Style preset, an Append Transcription button that Whispers a clip's speech into its caption, and fully offline transcription. Details ↓ · Release notes
- Fizgig 4.0 — video, sound and voices — MiniMax H3 now trains on video clips, on their sound, and on voice recordings alone: photos, clips and voice files in one folder train one LoRA in one run. Gizmo, a new bundled prep tool, cuts to-spec clips from any footage, auto-chops long videos at scene cuts, and records a voice dataset from nothing but a mic and ten minutes of reading. Training previews render in 6 steps with the Turbo LoRA and can carry their generated sound, opening in the gallery as playable clips. And 16 GB / 24 GB cards now train on the accurate int8 base — block swap streams one-way, ~6× faster, contributed by @rintic-13. Details ↓ · Release notes
- One-click cloud training on RunPod — no GPU, or want a 5090 for the afternoon? The official Fizgig template deploys the full app to a rented GPU in your browser: nothing to install, your files persist until you terminate the pod, and the in-app RunPod panel can even auto-stop the pod when your run finishes so an idle GPU never bills overnight. ⚡ Deploy → · Guide
Every trainer makes LoRAs. Fizgig is built around what you do with them afterwards — and that's the part nobody else has.
.safetensors.Under the workbench sits a fast, light trainer tuned to fit your GPU: a full Klein 9B LoRA trains on 16 GB, the 12.9B Krea 2 on 8 GB, and the 33B MiniMax H3 on 16 GB — block swap, quantisation and previews all size themselves to your VRAM automatically, and if a preview can't fit, training keeps running and saving. It loads kohya / PEFT / OneTrainer / AI-Toolkit / LyCORIS LoRAs, auto-converted, and saves kohya .safetensors that drop straight into ComfyUI.
Free and open source. A good first run: pick a ✨ built-in preset on the Training tab and go.
Each tool works on a trained run's output or any LoRA you've downloaded — and they hand off to each other (profile → repair → explore → compare, one closed loop). All three families: Klein, Krea 2 and MiniMax H3 (H3 previews render a short clip, judged by its middle frame — the model's native regime).
A live slider per transformer block (32 on Klein, up to 50 + the token refiners on MiniMax H3) with a side-by-side preview that updates as you drag. Turbo Preview caches per-block activations so late-block edits redraw up to 97% faster; the baked save is always exact. Blend blocks from a second donor LoRA, balance the pair per block, condition previews on a reference photo, and save a .safetensors that works in ComfyUI at strength 1.0.
Evolutionary discovery: the app mutates blocks and shows four variants — pick a favourite and it becomes the new baseline. Freeze what you like, set how far composition drifts, cycle seeds — and send any baseline to Repair Studio (and back) with one click.
Point it at a training run and it renders every epoch on one fixed seed, with a crossfade slider — drag until it looks best and stop. An optional likeness score (ArcFace, CPU) rates each epoch against a training photo and jumps you to the best. Then make it shareable: epoch-morph clips, seed / prompt / strength travels, a comparison sheet (with/without-LoRA grid, same seed per row), all exportable as looping MP4/GIF with an optional deflicker pass. Works on any folder of LoRAs, or a single file.
A per-block activation profile as a colour-coded HTML report — which blocks carry style, identity, and detail, and where they overlap. Repair Studio reads its sidecar automatically and shows the findings inline when you load the same LoRA.
Distil any Klein, Krea 2 or MiniMax H3 LoRA to a lower rank — Fast presets run weight-only SVD with no models loaded; Klein's activation-weighted presets add block and timestep targeting. PEFT and LyCORIS sources supported.
A from-scratch native port: 12.9B single-stream MMDiT, Qwen-Image VAE, Qwen3-VL-4B text encoder. Train on the RAW model; previews render on the training model itself with the official Turbo LoRA (auto-downloads) applied for the render only. Pick Krea 2 from the Base Model selector on the Training tab and the ✨ Krea 2 Defaults preset applies itself.
Everything works on Krea 2: all five workbench tools, Pause/Resume, Context LoRA, Adaptive LR, reference images, the live sample override — and LoKR training (pick it from Network Type; factor 8 or below for the quality edge, standard LoRA is ~20% faster). Output is ComfyUI-ready.
8 GB is enough. Users train full Krea 2 LoRAs on 8 GB with everything on Auto and batch size 1. Auto reads your free VRAM and picks INT8, NF4 or fp8 plus the right block swap — the console explains its choice. On longer runs the transformer blocks torch.compile automatically for roughly 2× faster steps.
Four Training-tab toggles no other trainer has:
Edit any caption yourself mid-run from the Problem Images window — no restart. When nothing is improving any more, a plateau banner names the best-checkpoint window to scrub in LoRA Royale. Pause, resume, restart: a resumed run replays its own loss log and loses nothing.
📣 Help map Krea 2's blocks — open an issue. Krea 2's per-block roles aren't charted yet, which is why the colour-coded sliders and layer targeting are Klein-only for now. The Profiler's weight-only report is the instrument — share what you find and it drives the presets and Repair Studio colour-coding to come.
Fizgig trains LoRAs for MiniMax H3, MiniMax's open-weight ~33B video model, from ordinary still-image datasets — and from short video clips, their sound, and voice recordings (details ↓) — on a single consumer GPU. Output loads straight into ComfyUI's H3 workflows, including the pruned inference builds.
The full studio, as of 4.2. H3 trains, previews and pauses/resumes like the other families — and all five workbench tools now work on H3 LoRAs too, with previews rendered as short clips judged by their middle frame. It was those tools, on real LoRAs, that produced the block map behind Optimised Likeness Learning below.
How it works: pick MiniMax H3 from the Base Model selector and the usual flow applies — Start-tab folder, Captions, Samples, Training. Leave Blocks Swap and Base Precision on Auto: at launch the trainer reads your free VRAM (close ComfyUI first) and picks the base precision and block-swap count together:
| Free VRAM | What Auto does |
|---|---|
| ~30 GB | int8, no block swap, up to 1 MP |
| ~22 GB | int8, ~14 blocks streamed |
| ~15 GB | int8, ~36 blocks streamed |
| ≤12 GB | 4-bit, as before |
int8 is the checkpoint's own storage and the most accurate base (~0.17% error). Block swap streams one way only — ~6.4× faster than round-trip swap, which is what lets 16 and 24 GB cards keep the accurate base (design contributed by @rintic-13, #73). Hit an OOM anyway? Set Blocks Swap to a number to override the planner.
Three built-in presets ship; Fast applies the moment you pick the family:
| Preset | Settings |
|---|---|
| ✨ MiniMax H3 Fast | LoRA dim/alpha 8, 50 epochs, flat 2e-4, 0.25 MP, Training Structure Likeness and Style, adamw. Reaches likeness in a few hundred steps, and the lower rank tends to come out more flexible |
| ✨ MiniMax H3 (Lower LR - slower) | The same at rank 16, 60 epochs, flat 1e-4 — more suitable for larger datasets with longer trains |
| ✨ MiniMax H3 Style | The Fast recipe on the measured style blocks, 0-3, 6-47 — style lives almost everywhere in H3 except the few blocks that only do identity and voice |

Optimised Likeness Learning ships ticked (Fast and Lower LR; Style unticks it): photo steps train only the identity blocks (20-49) while video and audio clips train the full model. Measured against full-model photo training: sharper, much better prompt following, better sound, fewer epochs — and the occasional deformed preview of full-model photo runs is gone. Untick it for style or scene training; while it's on, Blocks to Train is disabled with a note.
0.25 MP is the default, and it holds up — four times cheaper per step than 1 MP, and the extra resolution has not paid for itself in testing. Raise it if a specific dataset asks for it.
Previews default to 768×768, 56-frame clips with sound — a short watchable clip with the model's generated audio, opened in the gallery as a playable video (never autoplay). Without the audio VAE set, clips render silent; stills and other lengths stay in the dropdown. Set the Turbo LoRA in Preferences and previews render in 6 steps instead of 20 — previews only, never the saved LoRA. A preview that outgrows VRAM steps itself down a ladder rather than dying — a shorter clip first, then resolution to a 512×512 floor — and the size that fit is saved as the new default.
…train on video clips? Cut them with Gizmo (launch it from the Image Prep tab, or the Launch Gizmo .bat) — it exports clips already on H3's spec — drop them into the training folder next to your images, and caption them on the Captions tab like a photo. Photos, clips and voice recordings all train together in the same folder — no settings, no separate runs.
…make clips from my footage? Open Gizmo, drop a video on it, scrub to a moment, pick a length, Add to queue — repeat, then Export queue.
…chop a long video automatically? Gizmo's ✂ Auto-chop scene-detects the whole source and offers every segment as a thumbnail — click to keep or skip, and the keepers join the queue.
…train a voice from a recording? Gizmo's Voice tab: open any audio file (or a video, for its soundtrack), mark segments on the waveform, caption the sound, export — segments come out training-ready with their captions beside them.
…record a voice dataset from scratch? Voice tab → 🎙 Record: read the prompted sentences while holding the button (or the R key). Every take arrives trimmed and captioned; ten minutes of reading is a usable dataset.
…keep a clip's sound out of training? Mute it in Gizmo — it adds _mute to the filename, reversible by renaming. The video still trains.
…train photos, clips and a voice into one LoRA? Same folder, one trigger word, one run, any mix. If one category is much smaller, Finish one category early on the Training tab lets it finish at a chosen epoch while the rest trains on.
…get fast previews while training? Set the Turbo LoRA (~780 MB, its own Preferences row): 6-step previews with the Turbo at 75% on top of your training LoRA. Adjustable on the Samples tab.
…hear what it's generating while training? Pick a "with sound" Sample length on the Samples tab. Each preview carries its generated soundtrack, playable in the gallery.
…get a clip's spoken words into its caption? Open it in the caption editor (Captions tab → click the clip): any non-muted video shows an 🎤 Append Transcription button that Whispers the speech into the caption as saying "…" — Gizmo's grammar, without leaving the tab.
…set it up? One extra model file for sound: the audio VAE (~605 MB), on its own Preferences row. Blank = clips train silent; required only once the folder has voice recordings. Fizgig points out both new files once at startup if your H3 paths are set.
Stills teach H3 a look; clips teach it motion, and clips with sound teach it a voice. Clips cost far more per step than stills — start with a handful. Drop .mp4 clips into the training folder alongside your images and caption them like photos. A clip has to be on spec, and Fizgig refuses one that isn't rather than quietly fixing it:
| Requirement | |
|---|---|
| Container | .mp4 |
| Frame rate | exactly 24 fps |
| Frame count | 5, 22, 39, 56, 73, 90, 107 or 124 frames |
| Dimensions | multiples of 32 |
| Audio | 32 kHz stereo, or no track at all |

Gizmo makes clips that hit it — mark every section you want (frame-accurate stepping, first/last-frame previews, a ▶ Play of the exact clip), then export the lot in one go. Crop to the subject: a clip's cost is its pixels, so drag a rectangle and every token goes on what you want learned — with shape locks (1:1, 16:9, 9:16…) when you want consistent framing. High-frame-rate footage can keep extra frames as slow motion, offered as a choice. Clips are cut at native resolution and resized to your Target Megapixels at training time, so cutting large keeps the choice open.
What it costs: 22 frames is the shortest that shows real movement at ~7× a still per step; 124 frames is ~37×. Gizmo says which lengths your card can train, at which megapixels, before you cut anything:
| Clip | 16 GB | 24 GB | 32 GB |
|---|---|---|---|
| up to 56 frames | up to 0.25 MP | up to 0.5 MP | up to 0.5 MP |
| 73–90 frames | — | up to 0.25 MP | up to 0.5 MP |
| 107–124 frames | — | up to 0.25 MP | up to 0.25 MP |
Drop .wav / .mp3 / .flac / .m4a files into the training folder — alone or mixed with stills and clips. Rate and channels are converted for you; duration is the strict part:
| Requirement | |
|---|---|
| Formats | .wav .mp3 .flac .m4a — any rate or channel count |
| Duration | exactly 0.917, 1.625, 2.333, 3.042, 3.750, 4.458 or 5.167 s (±25 ms) |
| Content | actual sound — digital silence is refused |
| Caption | a .txt beside the file, or it silently won't train |
| Audio VAE | required — the ~605 MB Preferences row |

Gizmo's Voice tab cuts them for you — open a recording (or a video, for its soundtrack), mark segments on the waveform, pick a length, caption, export sample-exact. Caption the voice, not a picture — "a man speaking calmly, low pitch, unhurried" — with your trigger word leading; the Transcribe button (Whisper) appends the spoken words. Or record the dataset from scratch: 🎙 Record prompts sentences across every length and five tonal flavours, rolls a delivery style per take, and every hold-and-release lands trimmed, captioned and ready to queue. Set Training Structure to Likeness and Style for voices — tested head-to-head, it converges much faster; Fizgig reminds you when it sees voice files.
Each has a Download link on its row in Preferences:
| Model | Size | Notes |
|---|---|---|
| DiT — pruned int8 | ~21 GB | The training base — minimax_h3_fl2va_pruned_int8_convrot.safetensors, the same file ComfyUI runs. (The ~66 GB bf16 file also works for LoRA training, NF4 at load — but full fine-tuning needs this int8 file) |
| Qwen3-VL-32B text encoder | ~15.7 GB | The nvfp4 file — same one ComfyUI uses. Loaded once for caching, then freed |
| Video VAE | ~4.9 GB | Caching and preview decode |
| Audio VAE (optional) | ~605 MB | Sound training and previews with sound |
| Turbo LoRA (optional) | ~780 MB | 6-step previews — minimax_h3_turbo_v4_step600.safetensors; you may have it in ComfyUI's loras folder |
| DiT — reference (optional) | ~21 GB | Only for reference distillation (ref2va) |
Yes, you train on the pruned file. "Pruned" here swaps the AdaLN modulation MLP for a curve table — that branch only sees the timestep, so nothing a LoRA learns lives there. You train against the exact weights you deploy on.
Every control has a hint in the app; the highlights:
20-49 for likeness, 0-3, 6-47 for style (the Style preset sets it), voice core 38-48. Type ranges (3-12, 22, 31-33) to experiment beyond them.minimax_h3_turbo_v4_step600_ema the strongest checkpoint.Settings are read at launch; Pause → Resume relaunches with your current settings, so a pause is the moment to change them mid-run.
Everything above trains a LoRA. This trains the base model itself — no adapter, no rank bottleneck — on a single consumer GPU. Tick ⚗ Fine-tune the BASE MODEL instead of training a LoRA on the Training tab.
A note on where this is at. I first got fine-tuning working on Krea 2 shortly after its release, and I've been deliberately cautious about shipping it — first proving it to myself, then refining it through the MiniMax H3 work. This is the point where it needs the community to develop further. I don't expect every scenario to work perfectly yet — but it works, the numbers below are measured, and there's a solid foundation here to build on. Field reports genuinely shape what gets built next. I'm also aware this technique is model-agnostic at heart — it opens the door to fine-tuning other models, and I'm open to going there. But for that to happen it needs practical community support around those models — code, PRs, testing, that kind of thing — so I have the time necessary to make it happen. — Peter
New to fine-tuning? The extended "How do I…?" guide answers everything this section can't fit — including five-minute recipes for both families: tick Fine-tune, let the settings switch themselves, and change almost nothing.
One idea makes everything else here make sense: an "epoch" trains one slice of the model. The trainable window rotates each epoch, so it takes a full cycle — typically 4 epochs — for every part of the model to train once. Rule of thumb: 4 fine-tune epochs ≈ 1 true epoch of the whole model. That's why the epoch defaults look high, and why saves land on cycle boundaries — each saved checkpoint is a whole, evenly trained model.
Note on VRAM: the "trains on 8 GB" figures elsewhere in this README are for LoRA training. Full fine-tuning is a different animal — but it now tiers itself to your card, and fine-tuning defaults to a 4-bit NF4 frozen base that halves the model held on the card: on 32 GB and 24 GB the classic full-depth windows stay resident at full speed, and on 16 GB the frozen blocks stream from system RAM — slower steps, but the same component-mode learning. The planner measures your free VRAM at launch and prints the plan it chose.
What can my card fine-tune? The short answer, at the default training resolution:
| Your card | Krea 2 — photos | MiniMax H3 — photos | H3 — voice | H3 — video, confirmed | H3 — video on likeness blocks, expected |
|---|---|---|---|---|---|
| 16 GB | ✅ | ✅ | ✅ | ✅ up to 2.3 s | up to 3.8 s |
| 24 GB | ✅ | ✅ | ✅ | ✅ up to 2.3 s | up to 5.2 s |
| 32 GB | ✅ | ✅ | ✅ | ✅ up to 3.8 s | up to 5.2 s |
A few things worth knowing about that table: clip lengths follow Gizmo's grid, so 2.3 s means the 56-frame slot — cut your clips there and everything fits, confirmed by measured runs on every tier. On 32 GB, 3.8 s is also confirmed, even with video training the whole model. Beyond that, the Restrict video to likeness blocks tickbox (on by default with Optimised Likeness Learning — in our tests it trains video just as well, and it makes clips far lighter) extends the expected range: up to 5.2 s on 24 GB and 32 GB, and 3.8 s on 16 GB — conservative arithmetic from the measured constants, not yet individually measured, so treat those as expected rather than promised. Whole-model 5.2 s clips need more than 32 GB (measured). With the restriction unticked, one clip anywhere in your folder trains the whole model, so a mixed photos + clips dataset uses the clip column. And 12 GB cards train LoRAs, not fine-tunes — 16 GB is the fine-tune floor.
"A full fine-tune of a 12.9B–33B model on 16 GB" sounds like a trick, so here's the arithmetic. Only one slice of the model is ever trainable at a time — the trainable window rotates each epoch, so gradients and optimizer state exist for that slice alone. The frozen rest is held 4-bit (NF4) at half size and, on 16 GB, streamed from system RAM. The bf16 master copy lives in CPU RAM, never on the card. Those three together are the whole magic, and the numbers are measured, not projected: 8.8–12.3 GB peaks on a 16 GB card for H3, 8.4–11.0 GB for Krea 2 — and the console prints your own run's peak every epoch, so you can watch the claim hold live. Mechanism, tiers and every "how do I" in the extended guide: docs/FINETUNE_HOWDOI.md.
Which model files. Fine-tuning uses the same training bases you already have — nothing new to download:
krea2_raw_bf16.safetensors, ~26 GB), the same
file LoRA training uses. The fp8 Turbo is the preview model and can't be fine-tuned.minimax_h3_fl2va_pruned_int8_convrot.safetensors, ~21 GB) — again the same file the LoRA
path trains against and ComfyUI runs. The ~66 GB bf16 file, which LoRA training accepts, does
not work for fine-tuning; the trainer refuses it with a clear message.A finished fine-tune checkpoint is itself a valid base for either family — point the model path at it to train further (the console prints the exact continuation settings at every save). And — easy to miss — you can set it as the family's base in Preferences and train LoRAs on top of your own fine-tuned model: teach the base your world or cast once, then quick LoRAs for individual subjects ride on it. Deploy those LoRAs with the same fine-tuned base in ComfyUI. And Pause / Resume works on a fine-tune: Pause saves a full checkpoint even between the regular save epochs, and Resume continues it — rotation window, checkpoint numbering and the remaining epoch count all carry over.
Why bother. A LoRA constrains every update to a low-rank subspace, so concepts compete for the same handful of directions. That's why LoRAs tend to drag pose, framing and lighting toward the training set along with the likeness — they behave a bit like a filter over the model's output. A full-rank update can change how the model represents a concept, so it composes with what the model already knows. In our own tests, multi-character and concept teaching seemed to land at a much deeper level than LoRA training, with much better results — and the built-in Checkpoint to LoRA converter turns the result into a shareable file, and works very well. Beyond that, we're deliberately letting the community find the ceiling.
How it fits. A naive full fine-tune of Krea 2 (12.9B) needs roughly 78 GB — bf16 weights, gradients and optimizer state at once. Rotating windows make only part of the model trainable at a time, advancing each epoch, so gradients and optimizer state only ever exist for the active slice. Over a full cycle every weight trains. Around that sit three decisions that do the heavy lifting: a CPU-resident bf16 master copy is the source of truth, so training never round-trips through fp8 and quantisation can't erase the small updates being learned; optimizer-in-backward consumes and frees each gradient the moment it lands (worth 5.2 GB); and Adafactor's factored state is ~10× smaller than AdamW's.
It sizes itself to your card. Leave Window on Auto (by VRAM) and Fizgig measures the memory actually free at launch, picks the largest window that fits, and prints what it chose and why. Measured Krea 2 peaks (RTX 5090):
| Window mode | Peak VRAM | Speed | Fits |
|---|---|---|---|
| component + 4-bit NF4 (the default) — full-depth windows, resident | ~16 GB (24 GB budget) / ~21–23 GB (32 GB, more headroom held) | ~1.0 s/it | 24 GB and up |
| component + 4-bit NF4 + streaming | 8.4–11.0 GB | ~2.8 s/it | 16 GB |
| component on the fp8 base (explicit Base-precision pick) — depth-split + streamed | 15.6–17.6 GB | ~3.0 s/it | 24 GB |
4-bit NF4 is the fine-tune default, and you don't have to do anything to get it. It halves the frozen base, which on a 24 GB card is enough to keep the classic full-depth component windows resident instead of depth-splitting and streaming them: 4 windows instead of 8 — a full pass over every weight in 4 epochs rather than 8 — at roughly 3× the step speed (measured ~1.0 s/it against ~3.0 s/it for the fp8 base, same dataset, same 24 GB budget). On 16 GB it is the only base that fits at all.
The trade is that the frozen part of the model is held more coarsely while the trainable window learns against it. Your saved checkpoint is unaffected either way — it's written in bf16 from a master copy that never passes through a quantiser. If you want the more accurate frozen context and have the VRAM, pick fp8 under Base precision and it will be used.
Component is the best mode — and Auto now stays in it at every depth. Every window spans the model's full depth — attention across all 28 blocks, then each MLP matrix in turn — so a concept is learned by every layer at once rather than one depth slice at a time. The text-fusion stack stays trainable throughout: rotation would never reach it, and it's where prompt-to-concept binding happens. Where the budget used to force a mode change, the planner now depth-splits the windows instead (a fat window trains in slices — more windows per cycle, still full speed), and below that the frozen out-of-window blocks stream from system RAM — slower steps, but still component-mode learning. The console prints the chosen plan and why.
Block mode remains an explicit Window-dropdown choice — contiguous depth slices with frozen blocks streamed, slower than component at every budget. It's not yet quality-tested; every good result so far came from component runs.
The same checkbox under the MiniMax H3 family fine-tunes the 33B model, with the recipe adapted to it: component windows only — each window trains one attention or MLP matrix across all 50 blocks (4 windows per cycle), with the token refiner trainable throughout, so every window spans the model's full depth from the very first epoch.
mlp.fc1 trains in two slices, a 5-window cycle,
still full speed, no offloading; measured peaks 19.1–21.5 GB. On 16 GB the frozen
out-of-window blocks also stream from system RAM (~7 GB staged, a 9-window cycle):
measured peaks 8.8–12.3 GB at ~1.5× the step time — a full fine-tune of a 33B video model
on a 16 GB card. The console prints the chosen plan and why.Saves, previews and numbering run on the rotation cycle, not the Samples tab. The save
cadence snaps to cycle boundaries — the Save-every box follows the FT controls live in the GUI,
and the trainer snaps it again at launch — so every checkpoint compares like-for-like, with each
window trained equally. Previews ride the saves: one render per saved checkpoint plus the final
one, overriding the Samples tab's "every N epochs" (prompts, resolution, seed and the live
sample override still come from the Samples tab and status bar as usual — every sample in the
gallery maps to a file you can deploy). Checkpoints are numbered by epoch (-000004,
-000008, …) and the numbering continues across Pause/Resume, so a resumed run never overwrites
an earlier save. Krea 2 fine-tunes behave exactly the same way — saves snap to the cycle,
previews ride them (rendered on the training DiT with the Turbo LoRA), numbering carries over.
The output is a normal H3 checkpoint: load it in ComfyUI directly, or run Checkpoint to LoRA
(run_diff_to_lora.bat in your Fizgig folder) on it (the extractor decodes the int8 format natively) for a shareable LoRA.
If you're coming from LoRA training, recalibrate before anything else: fine-tuning wants much lower learning rates than LoRAs. A LoRA nudges a small adapter riding on a frozen model; a fine-tune moves the model's own weights, so the rates you're used to typing land very differently here — what's a normal LoRA rate can wreck a fine-tune outright.
Full fine-tuning moves every weight, so a long run on a handful of subjects drifts the model's whole notion of people — there's no low-rank bound to limit it the way there is with a LoRA. Point Regularisation images at a folder of ordinary photos of the broader class and they train at a reduced learning rate (LR ×, default 0.2) as an anchor rather than a lesson. That multiplier is a real dial, not a set-and-forget: 0.1–0.3 tethers the model's prior while your subject trains; push it toward 1.0 and the reg set trains like a second subject set — class-balanced training rather than a light anchor, which is a different (valid) thing. If a fine-tune drifts the broader class, raise it a step; if the subject learns too slowly, lower it. Worth a little experimentation per dataset.
Use real photos, not model output — anchoring a fine-tune to its own samples distils its artifacts back in, and there's nothing bounding that drift. Caption them normally: anything you leave unsaid gets attributed to the class word itself. Leave the folder empty to train without one.
A fine-tune produces a ~26 GB checkpoint, which is not what anyone wants to share. The
Checkpoint to LoRA utility — run_diff_to_lora.bat in your Fizgig folder, which opens
its own small window separate from the main app (Linux/pods: ./run_diff_to_lora.sh) — takes the base model
you started from and the checkpoint you produced, and extracts the difference as an ordinary
kohya .safetensors — at several ranks at once, since one SVD per layer serves them all.
This turned out to work far better than expected: rank 64 was perceptually indistinguishable from the full 26 GB checkpoint at ~0.5 GB, and quality degrades smoothly at lower ranks rather than falling off a cliff.
The result worth knowing: in our testing, a LoRA extracted from a fine-tune came out better than a LoRA trained directly at the same or higher rank on the same dataset. A low-rank file can hold a solution that low-rank training struggles to find — so fine-tune-then-extract isn't a workaround; the full-rank phase is the mechanism, and the extraction is nearly free.
Being straight about the trade-offs, because they're real:
The foundation: fast, light, and tuned for one model.
Loads kohya, PEFT, OneTrainer (OMI + legacy), AI-Toolkit, and LyCORIS (LoKR / LoHa) — auto-converted, and LoKR/LoHa run natively everywhere: Repair Studio, Profiler, Extract, Context LoRA. Repair Studio and Explorer save LoKR as LoKR, losslessly. Output is .safetensors that drops straight into ComfyUI.
Fizgig ships as a ready-made cloud image — the whole app in a browser tab, not a cut-down web version. Drag datasets in and LoRAs out with a built-in file manager, download models in one click, and optionally have the pod shut itself down when training finishes. Your models and datasets persist between sessions.
⚡ Deploy on RunPod → · Read the guide first
install_fizgig_rocm.bat (supported path). Linux: ./install_fizgig_rocm.sh — highly experimental (newer gfx like RDNA4, desktop compositor + training on the same GPU, and driver resets are common; use Windows ROCm or NVIDIA Linux for production training). Optional system amdrocm-amdsmi for accurate status-bar VRAM via amd-smi.Clone the repo:
git clone https://github.com/shootthesound/Fizgig.git
cd Fizgig
Clone it rather than downloading the ZIP — update_fizgig.bat updates by pulling with git, and a ZIP can't.
Open a terminal in your Fizgig folder and run:
git init
git remote add origin https://github.com/shootthesound/Fizgig.git
git fetch --depth 1 origin master
git reset --hard FETCH_HEAD
git branch -M master
git branch --set-upstream-to=origin/master master
Your model paths, output LoRAs, caches, presets and the venv are all left alone. update_fizgig.bat works normally from then on.
Windows (NVIDIA, one-click) — double-click install_fizgig.bat. It creates a venv, installs CUDA 12.8 PyTorch and all dependencies, pre-downloads the InsightFace models, and verifies CUDA is visible to PyTorch. Launch with run_fizgig.bat; update later with update_fizgig.bat.
Windows (AMD ROCm) — needs a full Python 3.12 install first (the ROCm bitsandbytes wheel is cp312-only; Fizgig's GUI needs Tkinter). Do not use the embeddable zip. Install from Windows downloads:
py install 3.12.Then double-click install_fizgig_rocm.bat (NVIDIA users never run this). It picks 3.12 via py -3.12 / python3.12 (not whatever python defaults to — e.g. 3.14). GPU detection follows, then pinned multi-arch wheels from AMD ROCm nightlies (https://rocm.nightlies.amd.com/whl-multi-arch/ — not built by Fizgig):
torch==2.12.0+rocm7.15.0a20260728torchvision==0.27.0+rocm7.15.0a20260728rocm-sdk-devel==7.15.0a20260728Override with TORCH_PIN / TORCHVISION_PIN / ROCM_SDK_DEVEL_PIN if needed. bitsandbytes is a pinned community Windows ROCm wheel from 0xDELUXA/bitsandbytes_win_rocm — built by neither AMD nor Fizgig. Shared deps come from requirements.txt with CUDA torch/bitsandbytes and NVIDIA-only nvidia-ml-py filtered out (filter_requirements_rocm.py). Launch with run_fizgig_rocm.bat; update later with update_fizgig_rocm.bat (not update_fizgig.bat — that script installs CUDA torch and would wipe the ROCm stack).
--experimental (unsupported): install_fizgig_rocm.bat --experimental installs unpinned torch[device-ARCH] / torchvision[device-ARCH] / rocm-sdk-devel from the same multi-arch index and leaves BNB_ROCM_VERSION unset so bitsandbytes auto-selects its highest matching DLL (no fallback warning while the resolved torch stays inside the wheel's HIP 7.13-7.16 range). This is not the same as Linux ROCM_CHANNEL=nightly (which stays on the constrained 7.14 / bitsandbytes 714 lane). Windows already installs from AMD nightlies with pinned versions by default; --experimental only drops those pins. Local experimentation only. Do not open GitHub issues for crashes, install failures, or training problems when --experimental was used — those reports will not be supported. Use the pinned install (no flag) for anything you expect help with.
Linux (AMD ROCm — highly experimental) — expect crashes, GPU resets, and incomplete model support on many setups. Best-effort only; Windows ROCm or NVIDIA Linux are the supported training paths. Prerequisites: amdgpu driver loaded (/dev/kfd), user in render/video groups. See Install ROCm and PyTorch for ROCm. Then:
chmod +x install_fizgig_rocm.sh
./install_fizgig_rocm.sh
./run_fizgig_rocm.sh
The script detects your gfx target (detect_gpu_linux.py). Nightly is the Linux default — TheRock multi-arch RELEASES.md index plus a [device-gfx*] extra for your GPU (e.g. gfx1201 → device-gfx1201). Unpinned nightly resolves the latest torch 2.12 + ROCm 7.14.0a* stack (matches libbitsandbytes_rocm714.so). Override with TORCH_PIN=…, ROCM_META_PIN=…, or TORCH_NIGHTLY_MINOR=….
Stable (repo.amd.com, no nightly alphas): pin torch==2.12.0+rocm7.14.0 + rocm-sdk==7.14.0 (cp310–cp314):
ROCM_CHANNEL=stable ./install_fizgig_rocm.sh
Try torch 2.14 (nightly only today — can increase sampling VRAM pressure vs 2.12):
ROCM_CHANNEL=nightly TORCH_NIGHTLY_MINOR=2.14 ./install_fizgig_rocm.sh
# or an explicit pin, e.g.:
# TORCH_PIN=2.14.0a0+rocm7.14.0a20260625 ROCM_CHANNEL=nightly ./install_fizgig_rocm.sh
# (paired torchvision ~0.29.0a0+rocm7.14.0a… — installer resolves the match)
Linux ROCm cache scripts import fizgig.rocm.cache_exit only when FIZGIG_GPU_BACKEND=rocm (set by run_fizgig_rocm.sh); NVIDIA and other platforms call main() unchanged. Opt out: FIZGIG_ROCM_NO_FAST_EXIT=1 ./run_fizgig_rocm.sh.
Then shared deps from requirements.txt (filtered) and bitsandbytes>=0.50.0 for ROCm.
Linux / macOS (NVIDIA CUDA path) — install_fizgig.py is CUDA-only (captioning / image prep on macOS; training needs a CUDA or ROCm GPU). On AMD-only Linux hosts it prints a hand-off to the ROCm installer and exits:
python install_fizgig.py
chmod +x run_fizgig.sh
./run_fizgig.sh
VRAM status bar on AMD: the existing NVIDIA pynvml / nvidia-smi path is unchanged; AMD readers (vram_monitor.read_amd_gpu_vram) run only as a fallback. Windows ROCm uses typeperf; Linux ROCm uses the amd-smi CLI when available (AMD SMI / ROCm Core SDK, e.g. sudo apt install amdrocm-amdsmi). Fizgig picks the GPU with the largest VRAM total (skips empty iGPU entries). Legacy rocm-smi is a fallback. Do not pip install amdsmi — the PyPI package is outdated.
buffalo_l (~300 MB, during install), Florence-2 (~500 MB–1.5 GB, first AI caption), and Helsinki-NLP opus-mt-en-zh (~300 MB, first bilingual translation).Fizgig doesn't bundle weights. You only need the family you're using — and Preferences has a ⬇ Download models for me button under each model card that downloads, verifies, and fills in the paths (Klein needs a free HuggingFace token for BFL's licence; Krea 2 needs no account). Every row also has a manual Download link. CLI:
python -m fizgig.scripts.fetch_models --family krea2 # ~32 GB, no account needed
python -m fizgig.scripts.fetch_models --family klein # ~34 GB, needs a token
python -m fizgig.scripts.fetch_models --family tools # Florence-2, face model, translator
| Model | File | Size | Source |
|---|---|---|---|
| Base DiT (fp8) — recommended | flux-2-klein-base-9b-fp8.safetensors | ~9.5 GB | black-forest-labs/FLUX.2-klein-base-9b-fp8 |
| Base DiT (bf16) | flux-2-klein-base-9b.safetensors | ~17 GB | black-forest-labs/FLUX.2-klein-base-9B |
| Distilled DiT | flux-2-klein-9b-fp8.safetensors | ~9 GB | black-forest-labs/FLUX.2-klein-9b-fp8 |
| VAE / AE | ae.safetensors | ~320 MB | black-forest-labs/FLUX.2-dev (from root, not the vae/ subfolder) |
| Text Encoder | qwen_3_8b.safetensors | ~15 GB | Comfy-Org/vae-text-encorder-for-flux-klein-9b |
Training runs on the Base DiT — the fp8 version is recommended on every GPU (same quality, half the VRAM). The Distilled DiT powers the 4-step previews and the workbench.
All files live in the one Comfy-Org/Krea-2 repo.
| Model | File | Size |
|---|---|---|
| RAW DiT (bf16) — training | krea2_raw_bf16.safetensors | ~26 GB |
| Turbo DiT (fp8) — workbench | krea2_turbo_fp8_scaled.safetensors | ~13 GB |
| Turbo LoRA (auto-downloads) | krea2_turbo_lora_rank_64_bf16.safetensors | ~470 MB |
| Qwen-Image VAE | qwen_image_vae.safetensors | ~250 MB |
| Text Encoder — recommended | qwen3vl_4b_fp8_scaled.safetensors | ~5.2 GB |
| Text Encoder — full precision | qwen3vl_4b_bf16.safetensors | ~8.9 GB |
The text-encoder slot is open: any Qwen3-VL-4B in the ComfyUI layout loads — fp8_scaled (recommended, captions we couldn't tell apart), bf16, or a community fine-tune/abliterated build, which changes how your dataset gets captioned.
MiniMax H3's files are listed in its section above.
Training — the fp8 Base stays resident at ~9.6 GB, so a 9B LoRA fits 16 GB (~14 GB observed). Smaller cards: the 4-bit (NF4) base toggle drops the base to ~5.6 GB — a full LoRA trains in ~7.5 GB, fitting 10–12 GB cards with no swap.
Workbench (Distilled 4-step):
| Block Swap | Min VRAM |
|---|---|
| 0 | 24 GB+ |
| 8 | 16 GB |
| 12 | 14 GB |
| 16 | 12 GB |
On first launch Fizgig auto-detects your VRAM and picks the default; your own choice sticks.
| Your card | What to do |
|---|---|
| 8 GB | Everything on Auto, batch size 1, stock preset defaults |
| 10–12 GB | Same — headroom to raise batch size or resolution |
| 16 GB+ | Same — Auto will usually pick the faster INT8 path |
Auto budgets from your free VRAM and the console explains its choice. If a preview can't fit, previews auto-disable and training keeps running and saving.
See the Auto table in its section — 16 GB and up trains on the accurate int8 base with streamed block swap; ≤12 GB falls back to 4-bit. On 16 GB-class cards, previews cap themselves at 768×640 and 22 frames (sound kept) — larger picks in the menus simply clamp, with a console note.
Turn off Hardware-accelerated GPU scheduling (Settings → System → Display → Graphics → Default graphics settings), then reboot. With it off, Fizgig runs training at low priority so your desktop stays smooth — training speed is unaffected.
Launch Fizgig and work left-to-right through the numbered tabs:
The unnumbered tabs are the post-training workbench: Profiler, Repair Studio, LoRA the Explorer, LoRA Royale, Extract, and Preferences.
One tool lives outside the main window: Checkpoint to LoRA (run_diff_to_lora.bat, or
python diff_to_lora_gui.py) — point it at a base model and a fine-tuned checkpoint and it writes
an ordinary LoRA at whichever ranks you tick. Only needed if you use the experimental full
fine-tune above.
Headless? Everything the trainer does is also available from the command line — see docs/CLI.md.
Community translations — Korean (한국어): Fizgig-Korean-Translated-Ver by @ssain3d-lgtm — an unofficial add-on that translates the UI at runtime without touching any Fizgig files, with a one-script uninstall. If you hit a bug while it's installed, uninstall and reproduce before reporting here; layer issues go to that repo.
If Fizgig saves you time or helps you make better LoRAs, consider supporting development:
Fizgig is open source under the Apache License 2.0 — free to use, modify, and redistribute, including commercially, with attribution and no warranty. Third-party components under compatible permissive licenses (and other terms where noted) are listed in THIRD_PARTY_NOTICES.md.
Copyright © 2026 Peter Neill.
Model weights are not covered by this license — each model carries its own terms from its publisher (see the Download links in Preferences).
I'm available for consulting on local AI training pipelines, custom workflow tooling, and private model work — the same engineering that's in Fizgig, applied to your studio's hardware and IP. Get in touch: peter@shootthesound.com.
Python
98.0%
Shell
1.2%