This repository provides scripts for training LoRA (Low-Rank Adaptation) models with HunyuanVideo, Wan2.1/2.2, FramePack, FLUX.1 Kontext, FLUX.2 dev/klein, Qwen-Image, Z-Image, and MiniMax-H3 architectures.
This repository is unofficial and not affiliated with the official repositories of these architectures.
This repository is under development.
We are grateful to the following companies for their generous sponsorship:
If you find this project helpful, please consider supporting its development via GitHub Sponsors. Your support is greatly appreciated!
GitHub Discussions Enabled: We've enabled GitHub Discussions for community Q&A, knowledge sharing, and technical information exchange. Please use Issues for bug reports and feature requests, and Discussions for questions and sharing experiences. Join the conversation →
September 16, 2026
--convrot_int8), as an alternative to --fp8_base --fp8_scaled. See PR #1008.
triton for the fused kernels. See the Krea 2 documentation for details.audio_path field for audio-capable architectures (currently MiniMax-H3); a same-stem audio sidecar file or the embedded audio track is used when omitted. PR #1020, PR #1021--attn_mode sdpa raising an error in the shared attention backends; it is now an alias of torch. Thank you rossnot PR #1092.enable_bucket and bucket_no_upscale when caching latents; the video caching path always bucketed regardless of the setting. Thank you christopher5106 PR #1100.
enable_bucket = true are now cached at the single configured resolution (resized and center-cropped), as image datasets always were. If you relied on bucketing without setting it, add enable_bucket = true to the dataset. Otherwise, re-run latent caching (and text encoder output caching for MiniMax-H3 fl2va / ref2va, whose caches embed the resized control images) so the caches match the configured resolution.--output_dir or --output_name is missing, instead of failing at the first save. Thank you rossnot PR #1070.--gradient_checkpointing_cpu_offload is now honored (activation CPU offloading during gradient checkpointing). Thank you rockerBOO PR #1101.--turbo_lora to compose a Turbo LoRA on top of the RAW model for sample generation during training, as an alternative to --turbo_dit. It can be combined with block swap, fp8 and ConvRot int8. See the Krea 2 documentation for details. Thank you rockerBOO PR #1103.July 14, 2026
--log_grad_metrics option to log gradient norm diagnostics (grad/norm, grad/mean_norm, grad/max, measured before gradient clipping) to the tracker. Thank you rockerBOO PR #988.
--max_grad_norm value. Disabled by default. See the advanced configuration documentation for details.We are grateful to everyone who has been contributing to the Musubi Tuner ecosystem through documentation and third-party tools. To support these valuable contributions, we recommend working with our releases as stable reference points, as this project is under active development and breaking changes may occur.
You can find the latest release and version history in our releases page.
This repository provides recommended instructions to help AI agents like Claude and Gemini understand our project context and coding standards.
To use them, you need to opt-in by creating your own configuration file in the project root.
Quick Setup:
Create a CLAUDE.md, GEMINI.md, and/or AGENTS.md file in the project root.
Add the following line to your CLAUDE.md to import the repository's recommended prompt (currently they are the almost same):
@./.ai/claude.prompt.md
or for Gemini:
@./.ai/gemini.prompt.md
You may be also import the prompt depending on the agent you are using with the custom prompt file such as AGENTS.md.
You can now add your own personal instructions below the import line (e.g., Always include a short summary of the change before diving into details.).
This approach ensures that you have full control over the instructions given to your agent while benefiting from the shared project context. Your CLAUDE.md, GEMINI.md and AGENTS.md (as well as Claude's .mcp.json) are already listed in .gitignore, so they won't be committed to the repository.
--blocks_to_swap, --fp8_llm, etc.For detailed information on specific architectures, configurations, and advanced features, please refer to the documentation below.
Architecture-specific:
Common Configuration & Usage:
Python 3.10 or later is required (verified with 3.10).
Create a virtual environment and install PyTorch and torchvision matching your CUDA version.
PyTorch 2.5.1 or later is required (see note).
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124
Install the required dependencies using the following command.
pip install -e .
Optionally, you can use FlashAttention and SageAttention (for inference only; see SageAttention Installation for installation instructions).
Optional dependencies for additional features:
ascii-magic: Used for dataset verificationmatplotlib: Used for timestep visualizationtensorboard: Used for logging training progressprompt-toolkit: Used for interactive prompt editing in Wan2.1 and FramePack inference scripts. If installed, it will be automatically used in interactive mode. Especially useful in Linux environments for easier prompt editing.pip install ascii-magic matplotlib tensorboard prompt-toolkit
You can also install using uv, but installation with uv is experimental. Feedback is welcome.
curl -LsSf https://astral.sh/uv/install.sh | sh
Follow the instructions to add the uv path manually until you restart your session...
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
Follow the instructions to add the uv path manually until you reboot your system... or just reboot your system at this point.
Model download procedures vary by architecture. Please refer to the architecture-specific documents in the Documentation section for instructions.
Please refer to here.
Pre-caching procedures vary by architecture. Please refer to the architecture-specific documents in the Documentation section for instructions.
Run accelerate config to configure Accelerate. Choose appropriate values for each question based on your environment (either input values directly or use arrow keys and enter to select; uppercase is default, so if the default value is fine, just press enter without inputting anything). For training with a single GPU, answer the questions as follows:
- In which compute environment are you running?: This machine
- Which type of machine are you using?: No distributed training
- Do you want to run your training on CPU only (even if a GPU / Apple Silicon / Ascend NPU device is available)?[yes/NO]: NO
- Do you wish to optimize your script with torch dynamo?[yes/NO]: NO
- Do you want to use DeepSpeed? [yes/NO]: NO
- What GPU(s) (by id) should be used for training on this machine as a comma-seperated list? [all]: all
- Would you like to enable numa efficiency? (Currently only supported on NVIDIA hardware). [yes/NO]: NO
- Do you wish to use mixed precision?: bf16
Note: In some cases, you may encounter the error ValueError: fp16 mixed precision requires a GPU. If this happens, answer "0" to the sixth question (What GPU(s) (by id) should be used for training on this machine as a comma-separated list? [all]:). This means that only the first GPU (id 0) will be used.
Training and inference procedures vary significantly by architecture. Please refer to the architecture-specific documents in the Documentation section and the various configuration documents for detailed instructions.
sdbsd has provided a Windows-compatible SageAttention implementation and pre-built wheels here: https://github.com/sdbds/SageAttention-for-windows. After installing triton, if your Python, PyTorch, and CUDA versions match, you can download and install the pre-built wheel from the Releases page. Thanks to sdbsd for this contribution.
For reference, the build and installation instructions are as follows. You may need to update Microsoft Visual C++ Redistributable to the latest version.
Download and install triton 3.1.0 wheel matching your Python version from here.
Install Microsoft Visual Studio 2022 or Build Tools for Visual Studio 2022, configured for C++ builds.
Clone the SageAttention repository in your preferred directory:
git clone https://github.com/thu-ml/SageAttention.git
Open x64 Native Tools Command Prompt for VS 2022 from the Start menu under Visual Studio 2022.
Activate your venv, navigate to the SageAttention folder, and run the following command. If you get a DISTUTILS not configured error, set set DISTUTILS_USE_SDK=1 and try again:
python setup.py install
This completes the SageAttention installation.
If you specify torch for --attn_mode, use PyTorch 2.5.1 or later (earlier versions may result in black videos).
If you use an earlier version, use xformers or SageAttention.
This repository is unofficial and not affiliated with the official repositories of the supported architectures.
This repository is experimental and under active development. While we welcome community usage and feedback, please note:
If you encounter any issues or bugs, please create an Issue in this repository with:
We welcome contributions! Please see CONTRIBUTING.md for details.
Code under the hunyuan_model directory is modified from HunyuanVideo and follows their license.
Code under the hunyuan_video_1_5 directory is modified from HunyuanVideo 1.5 and follows their license.
Code under the wan directory is modified from Wan2.1. The license is under the Apache License 2.0.
Code under the frame_pack directory is modified from FramePack. The license is under the Apache License 2.0.
Code in modules/convrot_int8_kernels.py is modified from comfy-kitchen (in turn derived from dxqb/OneTrainer and ComfyUI-Flux2-INT8). The license is under the Apache License 2.0.
Other code is under the Apache License 2.0. Some code is copied and modified from Diffusers.
(top 30 of 36)
Python
100.0%
This repository provides scripts for training LoRA (Low-Rank Adaptation) models with HunyuanVideo, Wan2.1/2.2, FramePack, FLUX.1 Kontext, FLUX.2 dev/klein, Qwen-Image, Z-Image, and MiniMax-H3 architectures.
This repository is unofficial and not affiliated with the official repositories of these architectures.
This repository is under development.
We are grateful to the following companies for their generous sponsorship:
If you find this project helpful, please consider supporting its development via GitHub Sponsors. Your support is greatly appreciated!
GitHub Discussions Enabled: We've enabled GitHub Discussions for community Q&A, knowledge sharing, and technical information exchange. Please use Issues for bug reports and feature requests, and Discussions for questions and sharing experiences. Join the conversation →
September 16, 2026
--convrot_int8), as an alternative to --fp8_base --fp8_scaled. See PR #1008.
triton for the fused kernels. See the Krea 2 documentation for details.audio_path field for audio-capable architectures (currently MiniMax-H3); a same-stem audio sidecar file or the embedded audio track is used when omitted. PR #1020, PR #1021--attn_mode sdpa raising an error in the shared attention backends; it is now an alias of torch. Thank you rossnot PR #1092.enable_bucket and bucket_no_upscale when caching latents; the video caching path always bucketed regardless of the setting. Thank you christopher5106 PR #1100.
enable_bucket = true are now cached at the single configured resolution (resized and center-cropped), as image datasets always were. If you relied on bucketing without setting it, add enable_bucket = true to the dataset. Otherwise, re-run latent caching (and text encoder output caching for MiniMax-H3 fl2va / ref2va, whose caches embed the resized control images) so the caches match the configured resolution.--output_dir or --output_name is missing, instead of failing at the first save. Thank you rossnot PR #1070.--gradient_checkpointing_cpu_offload is now honored (activation CPU offloading during gradient checkpointing). Thank you rockerBOO PR #1101.--turbo_lora to compose a Turbo LoRA on top of the RAW model for sample generation during training, as an alternative to --turbo_dit. It can be combined with block swap, fp8 and ConvRot int8. See the Krea 2 documentation for details. Thank you rockerBOO PR #1103.July 14, 2026
--log_grad_metrics option to log gradient norm diagnostics (grad/norm, grad/mean_norm, grad/max, measured before gradient clipping) to the tracker. Thank you rockerBOO PR #988.
--max_grad_norm value. Disabled by default. See the advanced configuration documentation for details.We are grateful to everyone who has been contributing to the Musubi Tuner ecosystem through documentation and third-party tools. To support these valuable contributions, we recommend working with our releases as stable reference points, as this project is under active development and breaking changes may occur.
You can find the latest release and version history in our releases page.
This repository provides recommended instructions to help AI agents like Claude and Gemini understand our project context and coding standards.
To use them, you need to opt-in by creating your own configuration file in the project root.
Quick Setup:
Create a CLAUDE.md, GEMINI.md, and/or AGENTS.md file in the project root.
Add the following line to your CLAUDE.md to import the repository's recommended prompt (currently they are the almost same):
@./.ai/claude.prompt.md
or for Gemini:
@./.ai/gemini.prompt.md
You may be also import the prompt depending on the agent you are using with the custom prompt file such as AGENTS.md.
You can now add your own personal instructions below the import line (e.g., Always include a short summary of the change before diving into details.).
This approach ensures that you have full control over the instructions given to your agent while benefiting from the shared project context. Your CLAUDE.md, GEMINI.md and AGENTS.md (as well as Claude's .mcp.json) are already listed in .gitignore, so they won't be committed to the repository.
--blocks_to_swap, --fp8_llm, etc.For detailed information on specific architectures, configurations, and advanced features, please refer to the documentation below.
Architecture-specific:
Common Configuration & Usage:
Python 3.10 or later is required (verified with 3.10).
Create a virtual environment and install PyTorch and torchvision matching your CUDA version.
PyTorch 2.5.1 or later is required (see note).
pip install torch torchvision --index-url https://download.pytorch.org/whl/cu124
Install the required dependencies using the following command.
pip install -e .
Optionally, you can use FlashAttention and SageAttention (for inference only; see SageAttention Installation for installation instructions).
Optional dependencies for additional features:
ascii-magic: Used for dataset verificationmatplotlib: Used for timestep visualizationtensorboard: Used for logging training progressprompt-toolkit: Used for interactive prompt editing in Wan2.1 and FramePack inference scripts. If installed, it will be automatically used in interactive mode. Especially useful in Linux environments for easier prompt editing.pip install ascii-magic matplotlib tensorboard prompt-toolkit
You can also install using uv, but installation with uv is experimental. Feedback is welcome.
curl -LsSf https://astral.sh/uv/install.sh | sh
Follow the instructions to add the uv path manually until you restart your session...
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
Follow the instructions to add the uv path manually until you reboot your system... or just reboot your system at this point.
Model download procedures vary by architecture. Please refer to the architecture-specific documents in the Documentation section for instructions.
Please refer to here.
Pre-caching procedures vary by architecture. Please refer to the architecture-specific documents in the Documentation section for instructions.
Run accelerate config to configure Accelerate. Choose appropriate values for each question based on your environment (either input values directly or use arrow keys and enter to select; uppercase is default, so if the default value is fine, just press enter without inputting anything). For training with a single GPU, answer the questions as follows:
- In which compute environment are you running?: This machine
- Which type of machine are you using?: No distributed training
- Do you want to run your training on CPU only (even if a GPU / Apple Silicon / Ascend NPU device is available)?[yes/NO]: NO
- Do you wish to optimize your script with torch dynamo?[yes/NO]: NO
- Do you want to use DeepSpeed? [yes/NO]: NO
- What GPU(s) (by id) should be used for training on this machine as a comma-seperated list? [all]: all
- Would you like to enable numa efficiency? (Currently only supported on NVIDIA hardware). [yes/NO]: NO
- Do you wish to use mixed precision?: bf16
Note: In some cases, you may encounter the error ValueError: fp16 mixed precision requires a GPU. If this happens, answer "0" to the sixth question (What GPU(s) (by id) should be used for training on this machine as a comma-separated list? [all]:). This means that only the first GPU (id 0) will be used.
Training and inference procedures vary significantly by architecture. Please refer to the architecture-specific documents in the Documentation section and the various configuration documents for detailed instructions.
sdbsd has provided a Windows-compatible SageAttention implementation and pre-built wheels here: https://github.com/sdbds/SageAttention-for-windows. After installing triton, if your Python, PyTorch, and CUDA versions match, you can download and install the pre-built wheel from the Releases page. Thanks to sdbsd for this contribution.
For reference, the build and installation instructions are as follows. You may need to update Microsoft Visual C++ Redistributable to the latest version.
Download and install triton 3.1.0 wheel matching your Python version from here.
Install Microsoft Visual Studio 2022 or Build Tools for Visual Studio 2022, configured for C++ builds.
Clone the SageAttention repository in your preferred directory:
git clone https://github.com/thu-ml/SageAttention.git
Open x64 Native Tools Command Prompt for VS 2022 from the Start menu under Visual Studio 2022.
Activate your venv, navigate to the SageAttention folder, and run the following command. If you get a DISTUTILS not configured error, set set DISTUTILS_USE_SDK=1 and try again:
python setup.py install
This completes the SageAttention installation.
If you specify torch for --attn_mode, use PyTorch 2.5.1 or later (earlier versions may result in black videos).
If you use an earlier version, use xformers or SageAttention.
This repository is unofficial and not affiliated with the official repositories of the supported architectures.
This repository is experimental and under active development. While we welcome community usage and feedback, please note:
If you encounter any issues or bugs, please create an Issue in this repository with:
We welcome contributions! Please see CONTRIBUTING.md for details.
Code under the hunyuan_model directory is modified from HunyuanVideo and follows their license.
Code under the hunyuan_video_1_5 directory is modified from HunyuanVideo 1.5 and follows their license.
Code under the wan directory is modified from Wan2.1. The license is under the Apache License 2.0.
Code under the frame_pack directory is modified from FramePack. The license is under the Apache License 2.0.
Code in modules/convrot_int8_kernels.py is modified from comfy-kitchen (in turn derived from dxqb/OneTrainer and ComfyUI-Flux2-INT8). The license is under the Apache License 2.0.
Other code is under the Apache License 2.0. Some code is copied and modified from Diffusers.
(top 30 of 36)
Python
100.0%