Extension for Stable Diffusion WebUI Forge - Neo
Generate videos with sound with MiniMax H3 inside Forge Neo, using the checkpoint, preset and Generate button you already know - describe the scene and its sounds, start from a picture, or give it pictures of the people and things to show, and get an MP4 with picture and audio made together.
"You shall run... on sixteen gigabytes of VRAM... on Forge Neo!" - made by this extension on a 16 GB card with 32 GB of RAM (first frame from Krea 2, 8 seconds with sound). Click for the MP4.
[!Important] Work in progress. Text-to-video with sound, first/last frame and reference pictures work. Tested on a 48 GB A40, a 24 GB RTX 4090 and a 16 GB RTX 2000 Ada with 32 GB of system RAM, using the smaller files. Cards under 16 GB have not been tested yet.
This extension requires an up-to-date Forge Neo (the
neobranch, revisiond70373eof 3 October 2026 or later, with comfy-kitchen 0.2.37). On an older version it stays disabled and tells you so in the console - Forge Neo itself keeps working as usual.
--use-ck-attention - Forge Neo's INT8 attention now runs H3 (comfy-kitchen 0.2.37 or later), 3-10% faster<lora:name:weight> syntaxMade inside Forge Neo, with sound. The previews are silent - click one to download the MP4 with sound.
| Bus stop, 15 s | Neon alley, 15 s | Bicycle, from a picture |
|---|---|---|
![]() | ![]() | ![]() |
| Turbo LoRA, 8 steps | 20 steps | img2img, first frame |
| Motorcycle, 8 s | Night train, from a picture | Cooking dinner, with speech |
|---|---|---|
![]() | ![]() | ![]() |
| Synthwave soundtrack and engine | img2img, first frame | Dialogue and kitchen sounds |
| Puppy, 5 s | Two friends, in Portuguese | Village festival, 10 s |
|---|---|---|
![]() | ![]() | ![]() |
| RTX 4090 (24 GB), smaller files | RTX 4090, 10 s, dialogue | RTX 4090, 960Γ544 |
Every prompt, the exact settings and the generation times - plus more clips, FastH3 and a community checkpoint - are in the Examples page of the wiki.
[!Tip] Step-by-step guides, every setting explained and measured times are in the wiki.
<Picture 1>, <Picture 2>...)https://github.com/eduardoabreu81/minimax-h3-forge-neo| Part | File | Folder |
|---|---|---|
| H3 checkpoint | minimax_h3_fl2va_pruned_int8_convrot | models/Stable-diffusion |
| Text encoder | qwen3vl_32b_minimax_h3_int8_convrot | models/text_encoder |
| Video VAE | minimax_h3_video_vae_fp16 | models/VAE |
| Audio VAE | minimax_h3_audio_vae_fp32 | models/VAE |
| Turbo LoRA (optional) | minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16 | models/Lora |
For a 24 GB or 16 GB card, or to use less system RAM, swap the checkpoint and the text encoder for these smaller files (same folders, same VAEs):
| Part | File | Size |
|---|---|---|
| H3 checkpoint | minimax_h3_fl2va_pruned_w4a8_mixed (Kijai) | 12.5 GB |
| Text encoder | qwen3vl_32b_minimax_h3_int4_convrot (Merserk) | 14.9 GB |
[!Tip] On a 16 GB card, start Forge Neo with
--cuda-malloc(for example inwebui-user.bat). Without it, the memory left behind by the text encoder is too fragmented for the checkpoint and the first step runs out of VRAM. With it, the smaller set ran a 3-second 640Γ384 clip without Never OOM; for longer or larger clips turn on Never OOM Integrated. See the 16 GB page.
[!Note] H3 is a large model and Forge keeps its files in system RAM. With the smaller set, a 16 GB card ran with 32 GB of RAM; on a 24 GB RTX 4090, Forge used about 35 GB of RAM with the smaller set and about 67 GB with the INT8 checkpoint and the INT4 text encoder; on a 48 GB card the INT8 set peaked at about 46 GiB. Other files that work - FastH3, GGUF checkpoints, Kijai's INT8 video VAE, community checkpoints - and the measurements are in the Models and Performance and Memory pages.
| Setup | Sampler | Schedule type | Steps | CFG | Shift |
|---|---|---|---|---|---|
| Regular (the h3 preset) | Res Multistep | Simple | 20 | 1 | 12 |
| Turbo LoRA at weight 1 | Res Multistep | Simple | 8, or 12 with speech | 1 | 6 |
| FastH3 checkpoint | Res Multistep | Simple | 8 | 1 | 10 |
Start from the h3 preset and change the steps and Shift by hand for a turbo LoRA or FastH3.
Launch options - two Forge Neo flags worth adding for H3. They go in COMMANDLINE_ARGS, in webui-user.bat on Windows or webui-user.sh on Linux, for example set COMMANDLINE_ARGS=--cuda-malloc --use-ck-attention:
| Flag | What it does | When to use it |
|---|---|---|
--cuda-malloc | Lets the NVIDIA driver organize the graphics card's memory. H3 swaps very large parts in and out of the card (the text encoder, then the model), and the default organizer can leave the free memory in pieces too small to use | Always on 16 GB cards - without it the first step can run out of memory. Harmless elsewhere |
--use-ck-attention | A faster way to compute attention, the heaviest part of every step, using 8-bit math | Any time - 3-10% faster, same picture. Needs the comfy-kitchen that comes with Forge Neo from 3 October 2026 |
More about both on the wiki: 16 GB Cards and Speed Options.
Sparse Attention Integrated - speeds up long clips:
| Scene | Timestep Range | Extra Tokens | Saves |
|---|---|---|---|
| Simple motion (walking, running, talking) | 0.15 - 0.85 (the default) | 0 | about 25% |
| Turns, spins, flips, vehicles cornering | 0.50 - 1.00 | 256 | about 15% |
With the default range, a person or object turning on itself may "morph" instead of rotating; starting at 0.50 keeps the motion of the regular result. With FastH3 the same switch turns on its own sparse attention.
Width and Height must be multiples of 32. The model looks best at 768 on the short side - 1152x768, 768x1152, 1024x576 or 576x1024. Use at least 544 with a turbo LoRA.
Frames sets the length, at 24 frames per second:
| Frames | Length |
|---|---|
| 22 | about 1 second |
| 73 | about 3 seconds |
| 124 | about 5 seconds |
| 192 | 8 seconds |
| 362 | about 15 seconds (the maximum) |
<d>[English] ...</d> after its speakerref2va in the checkpoint's file name, which is how the extension recognizes itAGPL-3.0 - see LICENSE
Made with β€οΈ for the Stable Diffusion community
Report Bug β’ Request Feature β’ β Ko-fi
Extension for Stable Diffusion WebUI Forge - Neo
Generate videos with sound with MiniMax H3 inside Forge Neo, using the checkpoint, preset and Generate button you already know - describe the scene and its sounds, start from a picture, or give it pictures of the people and things to show, and get an MP4 with picture and audio made together.
"You shall run... on sixteen gigabytes of VRAM... on Forge Neo!" - made by this extension on a 16 GB card with 32 GB of RAM (first frame from Krea 2, 8 seconds with sound). Click for the MP4.
[!Important] Work in progress. Text-to-video with sound, first/last frame and reference pictures work. Tested on a 48 GB A40, a 24 GB RTX 4090 and a 16 GB RTX 2000 Ada with 32 GB of system RAM, using the smaller files. Cards under 16 GB have not been tested yet.
This extension requires an up-to-date Forge Neo (the
neobranch, revisiond70373eof 3 October 2026 or later, with comfy-kitchen 0.2.37). On an older version it stays disabled and tells you so in the console - Forge Neo itself keeps working as usual.
--use-ck-attention - Forge Neo's INT8 attention now runs H3 (comfy-kitchen 0.2.37 or later), 3-10% faster<lora:name:weight> syntaxMade inside Forge Neo, with sound. The previews are silent - click one to download the MP4 with sound.
| Bus stop, 15 s | Neon alley, 15 s | Bicycle, from a picture |
|---|---|---|
![]() | ![]() | ![]() |
| Turbo LoRA, 8 steps | 20 steps | img2img, first frame |
| Motorcycle, 8 s | Night train, from a picture | Cooking dinner, with speech |
|---|---|---|
![]() | ![]() | ![]() |
| Synthwave soundtrack and engine | img2img, first frame | Dialogue and kitchen sounds |
| Puppy, 5 s | Two friends, in Portuguese | Village festival, 10 s |
|---|---|---|
![]() | ![]() | ![]() |
| RTX 4090 (24 GB), smaller files | RTX 4090, 10 s, dialogue | RTX 4090, 960Γ544 |
Every prompt, the exact settings and the generation times - plus more clips, FastH3 and a community checkpoint - are in the Examples page of the wiki.
[!Tip] Step-by-step guides, every setting explained and measured times are in the wiki.
<Picture 1>, <Picture 2>...)https://github.com/eduardoabreu81/minimax-h3-forge-neo| Part | File | Folder |
|---|---|---|
| H3 checkpoint | minimax_h3_fl2va_pruned_int8_convrot | models/Stable-diffusion |
| Text encoder | qwen3vl_32b_minimax_h3_int8_convrot | models/text_encoder |
| Video VAE | minimax_h3_video_vae_fp16 | models/VAE |
| Audio VAE | minimax_h3_audio_vae_fp32 | models/VAE |
| Turbo LoRA (optional) | minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16 | models/Lora |
For a 24 GB or 16 GB card, or to use less system RAM, swap the checkpoint and the text encoder for these smaller files (same folders, same VAEs):
| Part | File | Size |
|---|---|---|
| H3 checkpoint | minimax_h3_fl2va_pruned_w4a8_mixed (Kijai) | 12.5 GB |
| Text encoder | qwen3vl_32b_minimax_h3_int4_convrot (Merserk) | 14.9 GB |
[!Tip] On a 16 GB card, start Forge Neo with
--cuda-malloc(for example inwebui-user.bat). Without it, the memory left behind by the text encoder is too fragmented for the checkpoint and the first step runs out of VRAM. With it, the smaller set ran a 3-second 640Γ384 clip without Never OOM; for longer or larger clips turn on Never OOM Integrated. See the 16 GB page.
[!Note] H3 is a large model and Forge keeps its files in system RAM. With the smaller set, a 16 GB card ran with 32 GB of RAM; on a 24 GB RTX 4090, Forge used about 35 GB of RAM with the smaller set and about 67 GB with the INT8 checkpoint and the INT4 text encoder; on a 48 GB card the INT8 set peaked at about 46 GiB. Other files that work - FastH3, GGUF checkpoints, Kijai's INT8 video VAE, community checkpoints - and the measurements are in the Models and Performance and Memory pages.
| Setup | Sampler | Schedule type | Steps | CFG | Shift |
|---|---|---|---|---|---|
| Regular (the h3 preset) | Res Multistep | Simple | 20 | 1 | 12 |
| Turbo LoRA at weight 1 | Res Multistep | Simple | 8, or 12 with speech | 1 | 6 |
| FastH3 checkpoint | Res Multistep | Simple | 8 | 1 | 10 |
Start from the h3 preset and change the steps and Shift by hand for a turbo LoRA or FastH3.
Launch options - two Forge Neo flags worth adding for H3. They go in COMMANDLINE_ARGS, in webui-user.bat on Windows or webui-user.sh on Linux, for example set COMMANDLINE_ARGS=--cuda-malloc --use-ck-attention:
| Flag | What it does | When to use it |
|---|---|---|
--cuda-malloc | Lets the NVIDIA driver organize the graphics card's memory. H3 swaps very large parts in and out of the card (the text encoder, then the model), and the default organizer can leave the free memory in pieces too small to use | Always on 16 GB cards - without it the first step can run out of memory. Harmless elsewhere |
--use-ck-attention | A faster way to compute attention, the heaviest part of every step, using 8-bit math | Any time - 3-10% faster, same picture. Needs the comfy-kitchen that comes with Forge Neo from 3 October 2026 |
More about both on the wiki: 16 GB Cards and Speed Options.
Sparse Attention Integrated - speeds up long clips:
| Scene | Timestep Range | Extra Tokens | Saves |
|---|---|---|---|
| Simple motion (walking, running, talking) | 0.15 - 0.85 (the default) | 0 | about 25% |
| Turns, spins, flips, vehicles cornering | 0.50 - 1.00 | 256 | about 15% |
With the default range, a person or object turning on itself may "morph" instead of rotating; starting at 0.50 keeps the motion of the regular result. With FastH3 the same switch turns on its own sparse attention.
Width and Height must be multiples of 32. The model looks best at 768 on the short side - 1152x768, 768x1152, 1024x576 or 576x1024. Use at least 544 with a turbo LoRA.
Frames sets the length, at 24 frames per second:
| Frames | Length |
|---|---|
| 22 | about 1 second |
| 73 | about 3 seconds |
| 124 | about 5 seconds |
| 192 | 8 seconds |
| 362 | about 15 seconds (the maximum) |
<d>[English] ...</d> after its speakerref2va in the checkpoint's file name, which is how the extension recognizes itAGPL-3.0 - see LICENSE
Made with β€οΈ for the Stable Diffusion community
Report Bug β’ Request Feature β’ β Ko-fi