A UI made in Pyside6 to make training LoRA/LoCon and other LoRA type models in sd-scripts easy
Python
90
1,529 commits
updated Sep 21, 2026
After these changes, use_ramtorch_network training as it relates to glora via lycoris seems more effective. I still need to review other network types for any issues.
A larger change I made was incorporating EXPERIMENTAL support for sdxl flow matching (Based on Bluvoll et alls work: https://github.com/bluvoll/sd-scripts).
Model: https://civitai.com/models/2071356/experimental-noobai-with-rectified-flow-eq-vae
Args as pulled from Bluvoll's readme (pass as extra training args in easy scripts)
--flow_model Required to train into Rectified Flow target
--flow_use_ot Not needed, want info? ask lodestone
--flow_timestep_distribution = uniform or logit_normal, just run uniform
--flow_uniform_static_ratio allows static values for shift, default 2.5
--contrastive_flow_matching Magic thingie that makes training results a bit sharper
--cfm_lambda needed by the one above, default 0.05, I prefer 0.02
--flow_logit_mean needed by logit_normal timestep_distribution
--flow_logit_std needed by logit_normal timestep_distribution
--flow_uniform_base_pixels default 1048576 or 1024x1024 allows dynamic shifting based on resolutions useful for higher than 1024x1024 training.
--flow_uniform_shift allows dynamic shifting
--vae_custom_scale suggested for Anzhc's eq-vae put 0.1406
--vae_custom_shift suggested for Anzhc's eq-vae put -0.4743
--vae_reflection_padding suggested to use with Anzhc's eq-vae, my shitty experiment wasn't trained with this.
I did regression to make sure that SDXL non-flow matching still worked and that flow matching will run without error, I did not do complete regression of all model and configuration permutations.
Please open issues if any are observed.
Pushed some fixes and enhancements for ramtorch's helper for applying it to modules generally and lycoris specifically. use_ramtorch_network wasn't working as intended, now it appears to be working.
RamTorch
lycoris
sd_scripts
So overall, fixed use_ramtorch_network for lycoris, improved ramtorch application to models so it should be faster and use less RAM.
Reimplemented is_val subset support, also reimplemented making all subset parameters static for is_val or val split, this was missed when I reimplemented validation loss in the new branch.
As a result, even if you change nothing, after updating, validation loss metrics will change given the same training settings due to the enforcement of static parameters (e.x. cropping, shuffling captions, etc, anything that introduces randomization is disabled for is_val subsets OR dynamic subsets created by val split). Ultimately this will lead to a more consistent measure of validation loss going forward.
There was a bug with the ramtorch fork pyproject where it wasn't properly pulling in the modules directory. This has been corrected. Running update.bat or update.sh should get everything updated correctly.
I also added experimental ramtorch support to Kohya's loras, not fully tested, let me know if there are issues.
I am currently working on a new branch to rebase off the latest sd3 from upstream sd_scripts, the new branch chain (refresh/refresh/sd-upstream) doesn't have all the features of current default branch chain (flux/flux/sd3), but runs leaner and supports newer models. I am still working on adding back features deemed useful. RamTorch is not working correctly in the existing default branch (flux), please switch to the new branch (refresh)
Right now some of the significant things that aren't present are:
Ramtorch (Shout out to lodestone for his great work! https://github.com/lodestone-rock/RamTorch) is supported for network training, it can be applied to the base model and network/lora, in addition, optimizer state offload is implemented to two optimizers currently.
To use ramtorch:
The vram savings can be absurdly massive, for example, someone I have testing was able to train a 512 linear dim, 256 conv dim locon for SDXL, BS 7, and still didn't fully fill their VRAM, ~22GB out of 24GB, ofc, filled more system RAM. Others have been running as high as BS 18 with smaller dims.
With any reasonable batch size (4+), the overhead of ramtorch and CPU offloading appears to be negligible, in fact, it may actually speed up things due to use of streams, asynchronous operations, non blocking, etc.
The quality of outputs is not degraded in anyway as far as I have observed.
I havent nesscarily tested every lycoris setting or type fully, there may be some gaps. We have a discussion thread going around RamTorch in this context at https://github.com/67372a/LoRA_Easy_Training_Scripts/discussions/51, but feel free to open issues for any encountered.
It came to my attention when I was looking through lycoris code to make enhancements, that the bypass_mode arg was not being properly handled, if you did not explicitly pass false, it would be None, and end up resolving to TRUE. This would end up bypassing weight decomposition for DoRA, as there exists no logic, despite original author's documentation, that it should be overridden to false. Now, i have fixed this in my fork so that bypass mode defaults to FALSE, and if dora is enabled, it will also be forced to false. This may have significant effects how training behaves in cases where bypass_mode was being erroneously being enabled, especially DoRA!
In essence, during the forward pass, weight decomposition for DoRA was not being applied, the bypass route would also not apply DoRA style network dropout if set.
So, all DoRAs trained without explictly setting bypass_mode to false (which I doubt most people would do) were not trained correctly / fully, as such, the full benefits were not realized.
A set of training scripts written in python for use in Kohya's SD-Scripts. It has a UI written in pyside6 to help streamline the process of training models.
This project uses uv (from Astral) for fast package installation and virtual environment management. uv is installed automatically by the installer scripts if it is not already present — you do not need to install it manually. After installation, virtual environments are created with uv venv --python 3.11 --seed, which ensures pip is available inside each venv for compatibility.
Open up a command line within the folder that you want to install to, then run:
git clone https://github.com/67372a/LoRA_Easy_Training_Scripts -b refresh
cd LoRA_Easy_Training_Scripts
install.bat
The installer will:
./venv using uv venv.uv pip install.backend/sd_scripts/venv and installs all ML/training dependencies (PyTorch, xformers, etc.) via uv pip install.A few questions will be asked along the way — just make sure to answer them.
git clone --recurse-submodules https://github.com/67372a/LoRA_Easy_Training_Scripts -b refresh
cd LoRA_Easy_Training_Scripts
git submodule update --init --recursive
./install.sh
The installer will automatically download and install uv (via the standalone curl installer) and Python 3.11 if needed, then create venvs and install all dependencies using uv pip install.
If you specifically need to target Python 3.11 explicitly, you can use ./install311.sh instead.
The backend can be installed independently from the frontend UI (useful for headless/remote training setups):
cd backend
./install.sh
The backend has its own copy of install_uv.py and is fully self-contained — it does not require the frontend repo to be present.
If you are using one of the installers, one of the questions it will ask you is "Are you using this locally? (y/n):" — make sure you say y if you are going to be training on the computer you are using. This is very important to get correctly because the backend will not install otherwise, and you will be stuck wondering why it is not doing anything.
If you wish to train LoRAs but you lack the hardware, you can use this Google Colab created by Jelosus2 to be able to train them. Additionally, you can check the guide if you have trouble setting all up. Note: You still need to git clone the repo and install the UI on your local machine. Be sure to answer the prompt "Are you using this locally? (y/n):" with "n".
For Colab, the installer uses backend/colab_install.sh which runs the backend installer with the colab flag. This installs uv, creates the venv with Python 3.11, installs all ML dependencies via uv pip install, and configures accelerate automatically.
You can launch the UI using the run.bat file if you are on windows, or run.sh file if you are on linux.
The UI looks like this:

and has a bunch of features to it to make using it as easy as I could. So lets start with the basics. The UI is divided into two parts, the "args list" and the "subset list", this replaces the old naming scheme of <number>_<name> to try and reduce confusion. The subset list allows you to add and remove subsets to have however many you want!

You are also able to collapse and expand the sections of the "args list", so that way you can have open only the section you are working on at the moment.

Block Weight training is possible through setting the weights, dims, and alpha in the network args.

Pretty much every file selector has two ways to add a file without having to type it all in, a proper file dialog, and a way to drag and drop the value in!

TOML saving and loading are available so that you don't have to put in every variable every time you launch the program. All you need to do is either use the menu on the top right, or the keybind listed.
NOTE: This change is entirely different from the old system, so unfortunately the JSON files of the old scripts are no longer valid.
I have added a custom scheduler, CosineAnnealingWarmupRestarts. This scheduler allows restarts which restart with a decay, so that each restart has a bit less lr than the last, A few things to note about it though, warmup steps are not all applied at the beginning, but rather per epoch, I have set it up so that the warmup steps get divided evenly among them, decay is settable, and it uniquely has a minimum lr, which is set to 0 instead if the lr provided is smaller.
The Queue System is intuitive and easy to use, allowing you to save a config into a little button on the bottom left then allowing you to pull it back up for editing if you need to. Additionally you can use the arrow keys to change the positions of the queue items. A cool thing about this is that you can still edit args and even add or remove queue items while something else is training.

And finally, we have the ability to switch themes. These themes are only possible because of the great repo that adds in some material design and the ability to apply them on the fly called qt-material, give them a look as the work they've done is amazing.
The themes also save between boots

I'd like to take a moment and look at what the output of the TOML saving and loading system looks like so that people can change it if they want outside of the UI.
here is an example of what a config file looks like:
[[subsets]]
num_repeats = 10
keep_tokens = 1
caption_extension = ".txt"
shuffle_caption = true
flip_aug = false
color_aug = false
random_crop = true
is_reg = false
image_dir = "F:/Desktop/stable diffusion/LoRA/lora_datasets/atla_data/10_atla"
[noise_args]
[sample_args]
[logging_args]
[general_args.args]
pretrained_model_name_or_path = "F:/Desktop/stable diffusion/LoRA/nai-fp16.safetensors"
mixed_precision = "bf16"
seed = 23
clip_skip = 2
xformers = true
max_data_loader_n_workers = 1
persistent_data_loader_workers = true
max_token_length = 225
prior_loss_weight = 1.0
max_train_epochs = 2
[general_args.dataset_args]
resolution = 768
batch_size = 2
[network_args.args]
network_dim = 8
network_alpha = 1.0
[optimizer_args.args]
optimizer_type = "AdamW8bit"
lr_scheduler = "cosine"
learning_rate = 0.0001
lr_scheduler_num_cycles = 2
warmup_ratio = 0.05
[saving_args.args]
output_dir = "F:/Desktop"
save_precision = "fp16"
save_model_as = "safetensors"
output_name = "test"
save_every_n_epochs = 1
[bucket_args.dataset_args]
enable_bucket = true
min_bucket_reso = 256
max_bucket_reso = 1024
bucket_reso_steps = 64
[optimizer_args.args.optimizer_args]
weight_decay = 0.1
betas = "0.9,0.99"
As you can see everything is sectioned off into their own sections. Generally they are seperated into two groups, args, and dataset_args, this is because of the nature of the config and dataset_confg files within sd-scripts. Generally speaking, the only section that you might want to edit that doesn't correspond to a UI element (for now) is the [optimizer_args.args.optimizer_args] section, which you can add, delete, or change options for the optimizer, A proper UI for it will come later, once I figure out how I want to set it up.
ip Noise gamma to the UI, as it looked like it was useful.scale_weight_norms to the optimizer argsnetwork_dropout, rank_dropout, and module_dropout to the network argsscale_v_pred_loss_like_noise_pred to the general args, as that is directly tied to using a V Param based SD2.x model.Python
99.8%
A UI made in Pyside6 to make training LoRA/LoCon and other LoRA type models in sd-scripts easy
Python
90
1,529 commits
updated Sep 21, 2026
After these changes, use_ramtorch_network training as it relates to glora via lycoris seems more effective. I still need to review other network types for any issues.
A larger change I made was incorporating EXPERIMENTAL support for sdxl flow matching (Based on Bluvoll et alls work: https://github.com/bluvoll/sd-scripts).
Model: https://civitai.com/models/2071356/experimental-noobai-with-rectified-flow-eq-vae
Args as pulled from Bluvoll's readme (pass as extra training args in easy scripts)
--flow_model Required to train into Rectified Flow target
--flow_use_ot Not needed, want info? ask lodestone
--flow_timestep_distribution = uniform or logit_normal, just run uniform
--flow_uniform_static_ratio allows static values for shift, default 2.5
--contrastive_flow_matching Magic thingie that makes training results a bit sharper
--cfm_lambda needed by the one above, default 0.05, I prefer 0.02
--flow_logit_mean needed by logit_normal timestep_distribution
--flow_logit_std needed by logit_normal timestep_distribution
--flow_uniform_base_pixels default 1048576 or 1024x1024 allows dynamic shifting based on resolutions useful for higher than 1024x1024 training.
--flow_uniform_shift allows dynamic shifting
--vae_custom_scale suggested for Anzhc's eq-vae put 0.1406
--vae_custom_shift suggested for Anzhc's eq-vae put -0.4743
--vae_reflection_padding suggested to use with Anzhc's eq-vae, my shitty experiment wasn't trained with this.
I did regression to make sure that SDXL non-flow matching still worked and that flow matching will run without error, I did not do complete regression of all model and configuration permutations.
Please open issues if any are observed.
Pushed some fixes and enhancements for ramtorch's helper for applying it to modules generally and lycoris specifically. use_ramtorch_network wasn't working as intended, now it appears to be working.
RamTorch
lycoris
sd_scripts
So overall, fixed use_ramtorch_network for lycoris, improved ramtorch application to models so it should be faster and use less RAM.
Reimplemented is_val subset support, also reimplemented making all subset parameters static for is_val or val split, this was missed when I reimplemented validation loss in the new branch.
As a result, even if you change nothing, after updating, validation loss metrics will change given the same training settings due to the enforcement of static parameters (e.x. cropping, shuffling captions, etc, anything that introduces randomization is disabled for is_val subsets OR dynamic subsets created by val split). Ultimately this will lead to a more consistent measure of validation loss going forward.
There was a bug with the ramtorch fork pyproject where it wasn't properly pulling in the modules directory. This has been corrected. Running update.bat or update.sh should get everything updated correctly.
I also added experimental ramtorch support to Kohya's loras, not fully tested, let me know if there are issues.
I am currently working on a new branch to rebase off the latest sd3 from upstream sd_scripts, the new branch chain (refresh/refresh/sd-upstream) doesn't have all the features of current default branch chain (flux/flux/sd3), but runs leaner and supports newer models. I am still working on adding back features deemed useful. RamTorch is not working correctly in the existing default branch (flux), please switch to the new branch (refresh)
Right now some of the significant things that aren't present are:
Ramtorch (Shout out to lodestone for his great work! https://github.com/lodestone-rock/RamTorch) is supported for network training, it can be applied to the base model and network/lora, in addition, optimizer state offload is implemented to two optimizers currently.
To use ramtorch:
The vram savings can be absurdly massive, for example, someone I have testing was able to train a 512 linear dim, 256 conv dim locon for SDXL, BS 7, and still didn't fully fill their VRAM, ~22GB out of 24GB, ofc, filled more system RAM. Others have been running as high as BS 18 with smaller dims.
With any reasonable batch size (4+), the overhead of ramtorch and CPU offloading appears to be negligible, in fact, it may actually speed up things due to use of streams, asynchronous operations, non blocking, etc.
The quality of outputs is not degraded in anyway as far as I have observed.
I havent nesscarily tested every lycoris setting or type fully, there may be some gaps. We have a discussion thread going around RamTorch in this context at https://github.com/67372a/LoRA_Easy_Training_Scripts/discussions/51, but feel free to open issues for any encountered.
It came to my attention when I was looking through lycoris code to make enhancements, that the bypass_mode arg was not being properly handled, if you did not explicitly pass false, it would be None, and end up resolving to TRUE. This would end up bypassing weight decomposition for DoRA, as there exists no logic, despite original author's documentation, that it should be overridden to false. Now, i have fixed this in my fork so that bypass mode defaults to FALSE, and if dora is enabled, it will also be forced to false. This may have significant effects how training behaves in cases where bypass_mode was being erroneously being enabled, especially DoRA!
In essence, during the forward pass, weight decomposition for DoRA was not being applied, the bypass route would also not apply DoRA style network dropout if set.
So, all DoRAs trained without explictly setting bypass_mode to false (which I doubt most people would do) were not trained correctly / fully, as such, the full benefits were not realized.
A set of training scripts written in python for use in Kohya's SD-Scripts. It has a UI written in pyside6 to help streamline the process of training models.
This project uses uv (from Astral) for fast package installation and virtual environment management. uv is installed automatically by the installer scripts if it is not already present — you do not need to install it manually. After installation, virtual environments are created with uv venv --python 3.11 --seed, which ensures pip is available inside each venv for compatibility.
Open up a command line within the folder that you want to install to, then run:
git clone https://github.com/67372a/LoRA_Easy_Training_Scripts -b refresh
cd LoRA_Easy_Training_Scripts
install.bat
The installer will:
./venv using uv venv.uv pip install.backend/sd_scripts/venv and installs all ML/training dependencies (PyTorch, xformers, etc.) via uv pip install.A few questions will be asked along the way — just make sure to answer them.
git clone --recurse-submodules https://github.com/67372a/LoRA_Easy_Training_Scripts -b refresh
cd LoRA_Easy_Training_Scripts
git submodule update --init --recursive
./install.sh
The installer will automatically download and install uv (via the standalone curl installer) and Python 3.11 if needed, then create venvs and install all dependencies using uv pip install.
If you specifically need to target Python 3.11 explicitly, you can use ./install311.sh instead.
The backend can be installed independently from the frontend UI (useful for headless/remote training setups):
cd backend
./install.sh
The backend has its own copy of install_uv.py and is fully self-contained — it does not require the frontend repo to be present.
If you are using one of the installers, one of the questions it will ask you is "Are you using this locally? (y/n):" — make sure you say y if you are going to be training on the computer you are using. This is very important to get correctly because the backend will not install otherwise, and you will be stuck wondering why it is not doing anything.
If you wish to train LoRAs but you lack the hardware, you can use this Google Colab created by Jelosus2 to be able to train them. Additionally, you can check the guide if you have trouble setting all up. Note: You still need to git clone the repo and install the UI on your local machine. Be sure to answer the prompt "Are you using this locally? (y/n):" with "n".
For Colab, the installer uses backend/colab_install.sh which runs the backend installer with the colab flag. This installs uv, creates the venv with Python 3.11, installs all ML dependencies via uv pip install, and configures accelerate automatically.
You can launch the UI using the run.bat file if you are on windows, or run.sh file if you are on linux.
The UI looks like this:

and has a bunch of features to it to make using it as easy as I could. So lets start with the basics. The UI is divided into two parts, the "args list" and the "subset list", this replaces the old naming scheme of <number>_<name> to try and reduce confusion. The subset list allows you to add and remove subsets to have however many you want!

You are also able to collapse and expand the sections of the "args list", so that way you can have open only the section you are working on at the moment.

Block Weight training is possible through setting the weights, dims, and alpha in the network args.

Pretty much every file selector has two ways to add a file without having to type it all in, a proper file dialog, and a way to drag and drop the value in!

TOML saving and loading are available so that you don't have to put in every variable every time you launch the program. All you need to do is either use the menu on the top right, or the keybind listed.
NOTE: This change is entirely different from the old system, so unfortunately the JSON files of the old scripts are no longer valid.
I have added a custom scheduler, CosineAnnealingWarmupRestarts. This scheduler allows restarts which restart with a decay, so that each restart has a bit less lr than the last, A few things to note about it though, warmup steps are not all applied at the beginning, but rather per epoch, I have set it up so that the warmup steps get divided evenly among them, decay is settable, and it uniquely has a minimum lr, which is set to 0 instead if the lr provided is smaller.
The Queue System is intuitive and easy to use, allowing you to save a config into a little button on the bottom left then allowing you to pull it back up for editing if you need to. Additionally you can use the arrow keys to change the positions of the queue items. A cool thing about this is that you can still edit args and even add or remove queue items while something else is training.

And finally, we have the ability to switch themes. These themes are only possible because of the great repo that adds in some material design and the ability to apply them on the fly called qt-material, give them a look as the work they've done is amazing.
The themes also save between boots

I'd like to take a moment and look at what the output of the TOML saving and loading system looks like so that people can change it if they want outside of the UI.
here is an example of what a config file looks like:
[[subsets]]
num_repeats = 10
keep_tokens = 1
caption_extension = ".txt"
shuffle_caption = true
flip_aug = false
color_aug = false
random_crop = true
is_reg = false
image_dir = "F:/Desktop/stable diffusion/LoRA/lora_datasets/atla_data/10_atla"
[noise_args]
[sample_args]
[logging_args]
[general_args.args]
pretrained_model_name_or_path = "F:/Desktop/stable diffusion/LoRA/nai-fp16.safetensors"
mixed_precision = "bf16"
seed = 23
clip_skip = 2
xformers = true
max_data_loader_n_workers = 1
persistent_data_loader_workers = true
max_token_length = 225
prior_loss_weight = 1.0
max_train_epochs = 2
[general_args.dataset_args]
resolution = 768
batch_size = 2
[network_args.args]
network_dim = 8
network_alpha = 1.0
[optimizer_args.args]
optimizer_type = "AdamW8bit"
lr_scheduler = "cosine"
learning_rate = 0.0001
lr_scheduler_num_cycles = 2
warmup_ratio = 0.05
[saving_args.args]
output_dir = "F:/Desktop"
save_precision = "fp16"
save_model_as = "safetensors"
output_name = "test"
save_every_n_epochs = 1
[bucket_args.dataset_args]
enable_bucket = true
min_bucket_reso = 256
max_bucket_reso = 1024
bucket_reso_steps = 64
[optimizer_args.args.optimizer_args]
weight_decay = 0.1
betas = "0.9,0.99"
As you can see everything is sectioned off into their own sections. Generally they are seperated into two groups, args, and dataset_args, this is because of the nature of the config and dataset_confg files within sd-scripts. Generally speaking, the only section that you might want to edit that doesn't correspond to a UI element (for now) is the [optimizer_args.args.optimizer_args] section, which you can add, delete, or change options for the optimizer, A proper UI for it will come later, once I figure out how I want to set it up.
ip Noise gamma to the UI, as it looked like it was useful.scale_weight_norms to the optimizer argsnetwork_dropout, rank_dropout, and module_dropout to the network argsscale_v_pred_loss_like_noise_pred to the general args, as that is directly tied to using a V Param based SD2.x model.Python
99.8%