Seven nodes that assemble spec-compliant MiniMax H3 prompts, so you write shot content and the scaffolding is generated for you.
Specs implemented:
VIDEO_PROMPT_WRITING_GUIDE_base_en.md — T2VA / I2VA / FL2VA / L2VAVIDEO_PROMPT_WRITING_GUIDE_ref_en.md — full-reference (r2v)Unzip so you end up with:
ComfyUI\custom_nodes\comfyui-minimax-h3\__init__.py
Restart ComfyUI. No dependencies — pure stdlib.
On startup the console prints:
[MiniMax H3] wrote ...\comfyui-minimax-h3\web\minimax_h3_buttons.js
[MiniMax H3] serving at /extensions/comfyui-minimax-h3/minimax_h3_buttons.js
[MiniMax H3 Prompt] registered 7 nodes: ...
The web folder is written by the Python on every start, so there is only one
file to install and the script can't fall out of sync. If the buttons don't
appear, hard-refresh the browser (Ctrl+Shift+R) — the old script gets cached.
All nodes appear under the MiniMax H3 category.
Base modes — one Shot node per shot:
Shot → Shot → Shot → Assemble → prompt / total_seconds / frames
Full reference (r2v) — Subject chain plus the same Shot chain:
Subject → Subject ─┐
├→ Ref Prompt Builder r2v → prompt / total_seconds / frames
Shot → Shot ───────┘
cut_verb, seconds, text, and a shots chain socket (leave empty on the
first). Cut timestamps are computed cumulatively — set 5 / 2 / 3.5 and you get
[Shot 2] At 00:05.0, and [Shot 3] At 00:07.0, automatically. Inserting a
node mid-chain renumbers everything downstream.
Two buttons under the text box wrap a highlighted selection in <d>[English] ...</d> or in double quotes for on-screen text.
Terminates a Shot chain. mode picks T2VA / I2VA / FL2VA / L2VA and the right
alignment instruction is prepended. FL2VA and L2VA take their duration and last
shot index from the chain, so they can't drift from what you actually wrote.
One subject_definitions line. Two phrasings, chosen by the role:
<Subject 1> is <description>, whose <role> comes from <reference>.<Picture 1> is the <role> of <reference>, showing <description>.role_2 / reference_2 handle a subject drawing on two assets — appearance
from an image, motion from a video.
The six sections in fixed spec order. subjects and shots chain inputs
override their text boxes when connected. Three ordered task_type dropdowns
build the bracketed summary prefix.
Fixed eight-slot alternative to the Shot chain, with one global cut verb.
Single-body builder — you write the whole description including shot headers.
Durations in (one per line), formatted shot headers out. Useful when typing a body by hand instead of chaining Shot nodes.
24fps is hardcoded. MiniMax H3 is trained at 24fps, so it's a module
constant rather than a widget. The frames output snaps to the nearest valid
17n+5 length — the lattice isn't evenly spaced, so a shot can shift by up to
~0.35s.
Cut times use one decimal (00:05.0). The FL2VA/L2VA alignment line keeps
two decimals, which the spec requires.
FL2VA bracket style is reproduced verbatim from the official guide, which is
internally inconsistent — I2VA and L2VA use <Picture 1> / [Shot 1] while
FL2VA uses bare Picture 1 / Shot 1. Set NORMALISE_BRACKETS = True near the
top of __init__.py to make them match, if FL2VA misbehaves.
Seven nodes that assemble spec-compliant MiniMax H3 prompts, so you write shot content and the scaffolding is generated for you.
Specs implemented:
VIDEO_PROMPT_WRITING_GUIDE_base_en.md — T2VA / I2VA / FL2VA / L2VAVIDEO_PROMPT_WRITING_GUIDE_ref_en.md — full-reference (r2v)Unzip so you end up with:
ComfyUI\custom_nodes\comfyui-minimax-h3\__init__.py
Restart ComfyUI. No dependencies — pure stdlib.
On startup the console prints:
[MiniMax H3] wrote ...\comfyui-minimax-h3\web\minimax_h3_buttons.js
[MiniMax H3] serving at /extensions/comfyui-minimax-h3/minimax_h3_buttons.js
[MiniMax H3 Prompt] registered 7 nodes: ...
The web folder is written by the Python on every start, so there is only one
file to install and the script can't fall out of sync. If the buttons don't
appear, hard-refresh the browser (Ctrl+Shift+R) — the old script gets cached.
All nodes appear under the MiniMax H3 category.
Base modes — one Shot node per shot:
Shot → Shot → Shot → Assemble → prompt / total_seconds / frames
Full reference (r2v) — Subject chain plus the same Shot chain:
Subject → Subject ─┐
├→ Ref Prompt Builder r2v → prompt / total_seconds / frames
Shot → Shot ───────┘
cut_verb, seconds, text, and a shots chain socket (leave empty on the
first). Cut timestamps are computed cumulatively — set 5 / 2 / 3.5 and you get
[Shot 2] At 00:05.0, and [Shot 3] At 00:07.0, automatically. Inserting a
node mid-chain renumbers everything downstream.
Two buttons under the text box wrap a highlighted selection in <d>[English] ...</d> or in double quotes for on-screen text.
Terminates a Shot chain. mode picks T2VA / I2VA / FL2VA / L2VA and the right
alignment instruction is prepended. FL2VA and L2VA take their duration and last
shot index from the chain, so they can't drift from what you actually wrote.
One subject_definitions line. Two phrasings, chosen by the role:
<Subject 1> is <description>, whose <role> comes from <reference>.<Picture 1> is the <role> of <reference>, showing <description>.role_2 / reference_2 handle a subject drawing on two assets — appearance
from an image, motion from a video.
The six sections in fixed spec order. subjects and shots chain inputs
override their text boxes when connected. Three ordered task_type dropdowns
build the bracketed summary prefix.
Fixed eight-slot alternative to the Shot chain, with one global cut verb.
Single-body builder — you write the whole description including shot headers.
Durations in (one per line), formatted shot headers out. Useful when typing a body by hand instead of chaining Shot nodes.
24fps is hardcoded. MiniMax H3 is trained at 24fps, so it's a module
constant rather than a widget. The frames output snaps to the nearest valid
17n+5 length — the lattice isn't evenly spaced, so a shot can shift by up to
~0.35s.
Cut times use one decimal (00:05.0). The FL2VA/L2VA alignment line keeps
two decimals, which the spec requires.
FL2VA bracket style is reproduced verbatim from the official guide, which is
internally inconsistent — I2VA and L2VA use <Picture 1> / [Shot 1] while
FL2VA uses bare Picture 1 / Shot 1. Set NORMALISE_BRACKETS = True near the
top of __init__.py to make them match, if FL2VA misbehaves.