WanGP by DeepBeepMeep : The best Open Source Video Generative Models Accessible to the GPU Poor
WanGP supports the Wan (and derived models) but also Hunyuan Video, Flux, Qwen, Z-Image, LongCat, Kandinsky, LTXV, LTX-2, Qwen3 TTS, Chatterbox, HearMula, ... with:
Discord Server to get Help from the WanGP Community and show your Best Gens: https://discord.gg/g7efUW9jGV
Follow DeepBeepMeep on Twitter/X to get the Latest News: https://x.com/deepbeepmeep
Qwen3.5 VL Abliterated Prompt Enhancer: new choice of Prompt Enhancer
Also you can now expand or override a System Prompt prompt Enhancer with add @ or @@ (check new doc PROMPTS.md)
GGUF CUDA Kernels: 15% speed gain when using GGUF on Diffusion Video Models & x3 speed with GGUF LLM (Qwen 3.5 VL GGUF for instance). GGUF Kernels are for the moment only available for Windows (please check docs/INSTALLATION.md).
LTX2.3 Improvements
WanGP API: rejoice developers (or agents) among you ! WanGP offers now an internal API that allows you to use WanGP as a backend for your apps. It is subject to compliance to the terms & conditions of WanGP license and more specifically to inform the users of your app that WanGP is working behind the scene.
LTX Desktop WanGP: as a sample app (made just for fun) that uses WanGP API, you may try LTX Desktop. This app offers Video / Audio nice editing capabilities but will require 32+ VRAM to run. As now it uses WanGP as its core engine, VRAM requirements are much smaller. It will use LTX 2.3 for Video Gen & Z Image turbo fo Image gen. You can reuse (in theory) your current WanGP install with LTX Destop WanGP. https://github.com/deepbeepmeep/LTX-Desktop-WanGP
New Audio Ouput formats in mp4: audio stored in video file can now be of higher quality (AAC192 - AAC320) or ALAC (lossless). Please note that you wont be able listen to ALAC audio track directly in the webapp.
Also note as people preferred mataynone v1 over v2 I have added an option to select matanyone version in the Config / Extension tab
update 10.9871: Improved Qwen3.5 GGUF Prompt Enhancer Output Quality & added Think mode
update 10.9872: Added LTX 2.0/2.3 frames injection
update 10.9873: Fixed low fidelity LTX2 injected frames + added Image Strength slider for end & injected frames
Control Video Support (Ic lora Union Control) will let you transfer Human Motion, Edges, ... in your new video.
For expert users, Dev finetune offers extra new configurable settings (modality guidance, audio guidance, *STG pertubation/skip self attention *, guidance rescaling). LTX team suggests: Cfg=3, Audio cfg=7, Modality Cfg=3, Rescale=0.7, STG Perturbation Skip Attention on all steps.
I recommend to stick to the Distilled finetune for higher resolutions (see sample video below) as it seems to have been distilled from a higher quality model (pro model?).
Kiwi Edit: a great model that lets you edit video and / or inject objects in a video. It exists in 3 flavours depending on what you want to do
SVI PRO2 End Frames: this should allow in theory to generate very long shots by splitting one shot into sub shots (sliding windows) by inserting key frames (the End Frames). This is an alternative to the Infinitalk references frames method (see my old release notes). I am waiting for your feedback to know which method is the best one.
Upgraded Models Selector with already Downloaded indicator: Next to each model or finetune, you will find a colored square: Blue = fully downloaded & available, Yellow = partially downloaded & Black = not downloaded at all. Please note that the square color will depend on your current choices of requested model quantization.
Upgraded Models Manager: colors squares have also been added so that you can see in glance what has already been downloaded. New filter for a quick model lookout. List of missing files per finetune.
Matanyone 2: everyone favorite Mask extractor has been been updated and is now more precise
update 10.981: LTX2.3 Ic Lora Support & expert settings, Matanyone 2, SVI Pro end frames
Here comes the (last ?) missing bit in WanGP of the Text To Speech offering: emotions
There isnt many TTS models around that let you express emotions, so I hope you will forgive me for adding an old TTS model (6 months old!) in WanGP: Index TTS 2.
But in WanGP, you wont just get the vanilla version of Index TTS:
Here is how to use it: By default Index TTS, will detect automatically the emotion to apply to a Text Prompt based on the text itself. However, it will apply the same emotion for the whole prompt. If you want a different emotion per sentence, just insert empty lines between each sentence.
You can also set manually which emotion you expect with [] tags, here is one example for one speaker:
[fear] At the very beginning I was so afraid to speak.
[sadness] Nobody would talk to me. I felt so alone.
[disgust] They would just ignore me and pretend that I didnt exist
[happy] By chance I discovered this wonderful App, and now everything is different.
[anger] I have a new voice and now everybody will have no choice but to listen to my words !!!
You can mix emotions [sadness,disgust] or if you want to precise the weight of one or several emotions [sadness=0.7,disgust] (in any case total of weights is 1)
Remember two speakers mode requires to insert "Speaker 1:" & "Speaker 2:" to indicate who is talking.
There is only one snag: Index TTS 2 supports only English & Chinese. But dont' panic ! not all is lost. There is a workaround:
With this new release of WanGP you should have the best TTS (Text To Speech) experience you can find:
Qwen3 TTS Powered Up:
Heart Mula Powered Up:
Ace Step 1.5 Powered Up:
Also you now have a choice of multiple Prompt Enhancements for Qwen3 TTS & Kugel Audio: Prompt Enhancer can now generate for you either a Monologue or a Dialogue between two Speakers
Please note that to use the new Cuda Graph, mode you will need to select either vllm or cuda graph in Configuration / Performance / Language Models Decoder Engine. Profiles 1,3 or 3+ will need to be enabled for the corresponding Model. vllm is a powered up version of cuda graph that may not always work with all GPUs. But don't worry if it is not available for your GPU there will be an automatic fallback to cuda graph.
Ace Step 1.5 Turbo Super Charged: all the best features of Ace Step 1.5 are now in WanGP and are Fast & Easy to use:
LoKr support: this "Lora" like format has been tested with Flux Klein 9B
Optimized Int8 Kernels: all the Quantized INT8 checkpoints (most of the quantized checkpoints) used with WanGP should be now 10% faster !!!. You will need to install Triton. It is experimental, so for the moment it needs to be enabled manually in the Config / Performance tab. Please share your feedback on discord by mentioning your GPU so that I know if it works properly.
Auto Queue Saved if Gen Error: if for whatever reason you have got an error during a Gen, the queue will now be automatically saved. So you can try again this queue later (with a different config or when the related bug is fixed, if ever ...).
UI Updates (thx Tophness!): Updated the Self-Refiner UI to a dynamic, slider-based interface (no more manual text input). Improved queue reordering: items can now be dragged and dropped directly onto the Top and Bottom buttons while rearranging the queue in order to snap scroll to the top and bottom.
Kugel Audio Audio Split: Kugel Audio is a great model but strangely it tends to accelerate with long speeches. In order to avoid this effect, we need to split audio speeches. You can either do that manually by inserting an Empty Line or by specifiying an Auto Audio Split Duration (don't worry WanGP will try to split between lines or sentences).
update 10.81: Fixes
update 10.82: UI update
update 10.83: Kugel Audio Split
update 10.84: Ace Step RAM optimizations (fixed memory leak & reduce RAM requirements)
Note to RTX 50xx owners: you will need to upgrade to pytorch 2.10 (see upgrade procedure below) to be able to use Triton
The competition between Open Source & Close Source has never been that hot !
Please note that when using the Ace Step LM variants, this may get very slow with Memory Profiles 2 or 4 since the LM is an Autoregressive Model. It is why I recommed to stick to Memory Profiles 1/3/3+ unless you have very little VRAM.
Kugel Audio is entirely an Autoregressive Model and quite VRAM Hungry. So either you've got 16GB VRAM and you can run it with Memory Profile 1/3/3+ or you will have to go the slow way with other Profiles.
LTX-2 Base Tweaks: new Quality features if you found the base model was too fast :
Flux Klein 4B & 9B Base Models: Z Image has its base model in WanGP, so it was fair that Flux Klein would have its base model too. Base Models require more steps (up 50) and guidance > 1 but are good starting points for finetunes
The real novelty about this new release is that is has been tested and tuned to work with more recent versions of Python, Pytorch & Cuda. My end goal is to have everbody upgrade to Python 3.11, Pytorch 2.10, Cuda 13/13.1. Once we are all there it will be much easier to provide precompiled kernels for Nunchaku NVPF4, Sage Attention, Flash Attention, ... So please follow the manual upgrade instructions below (no Pinokio auto upgrade for the moment) and let me know on Discord if it works with all generations of GPUs (starting from GTX10xx to RTX50xx). You will find the kernels for this new setup in the guides/INSTALLATION.md.
Note that PyTorch 2.10 represents at last a decent upgrade, no memory leak when switching models (pytorch 2.8) and bad perfs / VRAM peaks with VAE decoding (pytorch 2.9).
Update: It seems GTX10xx doesnt support Cuda 13.0. Dont't worry I will keep WanGP compatibility with Pytorch 2.7.1 / Cuda 12.8.
Update 10.61: added Self Refiner
WanGP Special TTS (Text To Speech) Release:
Heart Mula: Suno quality song with lyrics on your local PC. You can generate up to 4 min of music.
Ace Step v1: while waiting for Ace Step v1.5 (which should be released very soon), enjoy this oldie (2025!) but goodie song generatpr as an appetizer. Ace Step v1 is a very fast Song generator. It is a Diffusion based, so dont hesitate to turn on Profile 4 to go as low as 4B VRAM while remaining fast.
Qwen 3 TTS: you can either do Voice Cloning, Generate a Custom Voice based on a Prompt or use a Predefined Voice
TTS Features:
Z Image Base: try it if you are into the Z Image hype but it will be probably useless for you unless you are a researcher and / or want to build a finetune out of it. This model requires from 35 to 50 steps (4x to 6x slower than Z Image turbo) and cfg > 1 (an additional 2x slower) and there is no Reinforcement Learning so Output Images wont be as good. The plus side is a higher diversity and Native Negative Prompt (versus Z Image virtual Negative Prompt using NAG).
Note that Z Image Base is very sensitive to the Attention Mode: it is not compatible with Sage 1 as it produces black frames. So I have disabled Sage for RTX 30xx. Also there are reports it produces some vertical banding artifacts with Sage 2
Flux 1/2 NAG : Flux 2 Klein is your new best friend but you miss Negative Prompts, NAG support for Distilled models will make you best buddies forever as NAG simulates Negative prompts.
Various Improvements:
update 10.51: new Heart Mula Finetune better at following instructions, Extra settings (cfg, top k) for TTS models, Rife v4
update 10.52: updated plugin list and added version tracking
update 10.53: video/audio galleries now support deletions
update 10.54: added Z Image Base, prompt enhancers improvements, configurable loras root folder
update 10.55: blocked Sage with Z Image on RTX30xx and added override attention mode settings, allowed changing config during generation
update 10.56: added NAG for Flux 1/2 & Ace Step v1
GPUs are expensive, RAM is expensive, SSD are expensive, sadly we live now in a GPU & RAM poor.
WanGP comes again to the rescue:
GGUF support: as some of you know, I am not a big fan of this format because when used with image / video generative models we don't get any speed boost (matrices multiplications are still done at 16 bits), VRAM savings are small and quality is worse than with int8/fp8. Still gguf has one advantage: it consumes less RAM and harddrive space. So enjoy gguf support. I have added ready to use Kijai gguf finetunes for LTX-2.
Models Manager PlugIn: use this Plugin to identify how much space is taken by each model / finetune and delete the ones you no longer use. Try to avoid deleting shared files otherwise they will be downloaded again.
LTX-2 Dual Video & Audio Control: you no longer need to extract the audio track of a Control Video if you want to use it as well to drive the video generation. New mode will allow you to use both motion and audio from Video Control.
LTX-2 - Custom VAE URL: some users have asked if they could use the old Distiller VAE instead of the new one. To do that, create a finetune def based on an existing model definition and save it in the finetunes/ folder with this entry (check the docs/FINETUNES.md doc):
"VAE_URLs": ["https://huggingface.co/DeepBeepMeep/LTX-2/resolve/main/ltx-2-19b_vae_old.safetensors"]
Flux 2 Klein 4B & 9B: try these distilled models as fast as Z_Image if not faster but with out of the box image edition capabiltities
Flux 2 & Qwen Outpainting + Lanpaint: the inpaint mode of these models support now outpainting + more combination possible with Lanpaint
RAM Optimizations for multi minutes Videos: processing, saving, spatial & Temporal upsampling very long videos should require much less RAM.
Text Encoder Cache: if you are asking a Text prompt already used recently with the current model, it will be taken straight from a cache. The cache is optimized to consume little RAM. It wont work with certain models such as Qwen where the Text Prompt is combined internally with an Image.
update 10.41: added Flux 2 klein
update 10.42: added RAM optimizations & Text Encoder Cache
update 10.43: added outpainting for Qwen & Flux 2, Lanpaint for Flux 2
So dont be surprised if the old checkpoints are deleted and new are downloaded !!!.
LTX-2 Multi Passes Loras multipliers: LTX-2 supports now loras multiplier that depend on the Pass No. For instance "1;0.5" means 1 will the strength for the first LTX-2 pass and 0.5 will be the strength for the second pass.
New Profile 3.5: here is the lost kid of Profile 3 & Profile 5, you got tons of VRAM, but little RAM ? Profile 3.5 will be your new friend as it will no longer use Reserved RAM to accelerate transfers. Use Profile 3.5 only if you can fit entirely a Diffusion / Transformer model in VRAM, otherwise the gen may be much slower.
NVFP4 Quantization for LTX-2 & Flux 2: you will now be able to load NV FP4 model checkpoints in WanGP. On top of Wan NV4 which was added recently, we now have LTX-2 (non distilled) & Flux 2 support. NV FP4 uses slightly less VRAM and up to 30% less RAM.
To enjoy fully the NV FP4 checkpoints (at least 30% faster gens), you will need a RTX 50xx and to upgrade to Pytorch 2.9.1 / Cuda 13 with the latest version of lightx2v kernels (check docs/INSTALLATION.md). To observe the speed gain, you have to make sure the workload is quite high (high res, long video).
With WanGP 10.21 HD 720p Video Gens of 10s just need now 8GB of VRAM!
LTX Team said this video gen was for 4k. So I had no choice but to squeeze more VRAM with further optimizations.
After much suffering I have managed to reduce by at least 1/3 the VRAM requirements of LTX-2, which means:
3K/4K resolutions will be available only if you enable them in the Config / General tab.
Ic Loras support: Use a Control Video to transfer Pose, Depth, Canny Edges. I have added some extra tweaks: with WanGP you can restrict the transfer to a masked area, define a denoising strength (how much the control video is going to be followed) and a masking strength (how much unmasked area is impacted)
Start Image Strength: This new slider will appear below a Start Image or Source Video. If you set it to values lower than 1 you may to reduce the static image effect, you get sometime with LTX-2 i2v
Custom Gemma Text Encoder for LTX-2: As a practical case, the Heretic text encoder is now supported by WanGP. Check the finetune doc, but in short create a finetune that has a text_encoder_URLS key that contains a list of one or more file paths or URLs.
Experimental Auto Recovery Failed Lora Pin: Some users (with usually PC with less than 64 GB of RAM) have reported Out Of Memory although a model seemed to load just fine when starting a gen with Loras. This is sometime related to WanGP attempting (and failing due to unsufficient reserved RAM) to pin the Loras to Reserved Memory for faster gen. I have experimented a recovery mode that should release sufficient ressources to continue the Video Gen. This may solve the oom crashes with LTX-2 Default (non distilled)
Max Loras Pinned Slider: If the Auto Recovery Mode is still not sufficient, I have added a Slider at the bottom of the Configuration / Performance tab that you can use to prevent WanGP from Pinning Loras (to do so set it to 0). As if there is no loading attempt there wont be any crash...
update 10.21: added slider Loras Max Pinning slider
update 10.22: added support for custom LTX-2 Text Encoder + Auto Recovery mode if Lora Pinning failed
update 10.23: Fixed text prompt ignore in profile 1 & 2 (this created random output videos)
With WanGP v10.11 you can now force your soundtrack, it works like Multitalk / Avatar except in theory it should work with any kind of sound (not just vocals). Thanks to Kijai for showing it was possible.
Z Image Twin Folder Turbo: Z Image even faster as this variant can generate images with as little as 1 step (3 steps recommend)
Qwen LanPaint: very precise In Painting, offers a better integration of the inpainted area in the rest of the image. Beware it is up to 5x slower as it "searches" for the best replacement.
Optimized Pytorch Compiler : Patience is the Mother of Virtue. Finally I may (or may not) have fixed the PyTorch compiler with the Wan models. It should work in much diverse situations and takes much less time.
LongCat Video: experimental support which includes LongCat Avatar a talking head model. For the moment it is mostly for models collectors as it is very slow. It needs 40+ steps and each step contains up 3 passes.
MMaudio NSFW: for alternative audio background
update v10.11: LTX-2, use your own soundtrack
See full changelog: Changelog
One-click installation:
Get started instantly with Pinokio App
It is recommended to use in Pinokio the Community Scripts wan2gp or wan2gp-amd by Morpheus rather than the official Pinokio install.
Manual installation: (old python 3.10, to be deprecated)
git clone https://github.com/deepbeepmeep/Wan2GP.git
cd Wan2GP
conda create -n wan2gp python=3.10.9
conda activate wan2gp
pip install torch==2.7.1 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/cu128
pip install -r requirements.txt
Manual installation: (new python 3.11 setup)
git clone https://github.com/deepbeepmeep/Wan2GP.git
cd Wan2GP
conda create -n wan2gp python=3.11.14
conda activate wan2gp
pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
Run the application:
python wgp.py
First time using WanGP ? Just check the Guides tab, and you will find a selection of recommended models to use.
Update the application (stay in the old pyton / pytorch version): If using Pinokio use Pinokio to update otherwise: Get in the directory where WanGP is installed and:
git pull
conda activate wan2gp
pip install -r requirements.txt
Upgrade to 3.11, Pytorch 2.10, Cuda 13/13.1 (for non GTX10xx users) I recommend creating a new conda env for the Python 3.11 to avoid bad surprises. Let's call the new conda env wangp (instead of wan2gp the old name of this project) Get in the directory where WanGP is installed and:
git pull
conda create -n wangp python=3.11.9
conda activate wangp
pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
Git Errors Once you are done you will have to reinstall Sage Attention, Triton, Flash Attention. Check the Installation Guide -
if you get some error messages related to git, you may try the following (beware this will overwrite local changes made to the source code of WanGP):
git fetch origin && git reset --hard origin/main
conda activate wangp
pip install -r requirements.txt
When you have the confirmation it works well you can then delete the old conda env:
conda uninstall -n wan2gp --all
Run headless (batch processing):
Process saved queues without launching the web UI:
# Process a saved queue
python wgp.py --process my_queue.zip
Create your queue in the web UI, save it with "Save Queue", then process it headless. See CLI Documentation for details.
For Debian-based systems (Ubuntu, Debian, etc.):
./run-docker-cuda-deb.sh
This automated script will:
Docker environment includes:
Supported GPUs: RTX 40XX, RTX 30XX, RTX 20XX, GTX 16XX, GTX 10XX, Tesla V100, A100, H100, and more.
For detailed installation instructions for different GPU generations:
For detailed installation instructions for different GPU generations:
Made with β€οΈ by DeepBeepMeep
Python
97.7%
JavaScript
1.2%
WanGP by DeepBeepMeep : The best Open Source Video Generative Models Accessible to the GPU Poor
WanGP supports the Wan (and derived models) but also Hunyuan Video, Flux, Qwen, Z-Image, LongCat, Kandinsky, LTXV, LTX-2, Qwen3 TTS, Chatterbox, HearMula, ... with:
Discord Server to get Help from the WanGP Community and show your Best Gens: https://discord.gg/g7efUW9jGV
Follow DeepBeepMeep on Twitter/X to get the Latest News: https://x.com/deepbeepmeep
Qwen3.5 VL Abliterated Prompt Enhancer: new choice of Prompt Enhancer
Also you can now expand or override a System Prompt prompt Enhancer with add @ or @@ (check new doc PROMPTS.md)
GGUF CUDA Kernels: 15% speed gain when using GGUF on Diffusion Video Models & x3 speed with GGUF LLM (Qwen 3.5 VL GGUF for instance). GGUF Kernels are for the moment only available for Windows (please check docs/INSTALLATION.md).
LTX2.3 Improvements
WanGP API: rejoice developers (or agents) among you ! WanGP offers now an internal API that allows you to use WanGP as a backend for your apps. It is subject to compliance to the terms & conditions of WanGP license and more specifically to inform the users of your app that WanGP is working behind the scene.
LTX Desktop WanGP: as a sample app (made just for fun) that uses WanGP API, you may try LTX Desktop. This app offers Video / Audio nice editing capabilities but will require 32+ VRAM to run. As now it uses WanGP as its core engine, VRAM requirements are much smaller. It will use LTX 2.3 for Video Gen & Z Image turbo fo Image gen. You can reuse (in theory) your current WanGP install with LTX Destop WanGP. https://github.com/deepbeepmeep/LTX-Desktop-WanGP
New Audio Ouput formats in mp4: audio stored in video file can now be of higher quality (AAC192 - AAC320) or ALAC (lossless). Please note that you wont be able listen to ALAC audio track directly in the webapp.
Also note as people preferred mataynone v1 over v2 I have added an option to select matanyone version in the Config / Extension tab
update 10.9871: Improved Qwen3.5 GGUF Prompt Enhancer Output Quality & added Think mode
update 10.9872: Added LTX 2.0/2.3 frames injection
update 10.9873: Fixed low fidelity LTX2 injected frames + added Image Strength slider for end & injected frames
Control Video Support (Ic lora Union Control) will let you transfer Human Motion, Edges, ... in your new video.
For expert users, Dev finetune offers extra new configurable settings (modality guidance, audio guidance, *STG pertubation/skip self attention *, guidance rescaling). LTX team suggests: Cfg=3, Audio cfg=7, Modality Cfg=3, Rescale=0.7, STG Perturbation Skip Attention on all steps.
I recommend to stick to the Distilled finetune for higher resolutions (see sample video below) as it seems to have been distilled from a higher quality model (pro model?).
Kiwi Edit: a great model that lets you edit video and / or inject objects in a video. It exists in 3 flavours depending on what you want to do
SVI PRO2 End Frames: this should allow in theory to generate very long shots by splitting one shot into sub shots (sliding windows) by inserting key frames (the End Frames). This is an alternative to the Infinitalk references frames method (see my old release notes). I am waiting for your feedback to know which method is the best one.
Upgraded Models Selector with already Downloaded indicator: Next to each model or finetune, you will find a colored square: Blue = fully downloaded & available, Yellow = partially downloaded & Black = not downloaded at all. Please note that the square color will depend on your current choices of requested model quantization.
Upgraded Models Manager: colors squares have also been added so that you can see in glance what has already been downloaded. New filter for a quick model lookout. List of missing files per finetune.
Matanyone 2: everyone favorite Mask extractor has been been updated and is now more precise
update 10.981: LTX2.3 Ic Lora Support & expert settings, Matanyone 2, SVI Pro end frames
Here comes the (last ?) missing bit in WanGP of the Text To Speech offering: emotions
There isnt many TTS models around that let you express emotions, so I hope you will forgive me for adding an old TTS model (6 months old!) in WanGP: Index TTS 2.
But in WanGP, you wont just get the vanilla version of Index TTS:
Here is how to use it: By default Index TTS, will detect automatically the emotion to apply to a Text Prompt based on the text itself. However, it will apply the same emotion for the whole prompt. If you want a different emotion per sentence, just insert empty lines between each sentence.
You can also set manually which emotion you expect with [] tags, here is one example for one speaker:
[fear] At the very beginning I was so afraid to speak.
[sadness] Nobody would talk to me. I felt so alone.
[disgust] They would just ignore me and pretend that I didnt exist
[happy] By chance I discovered this wonderful App, and now everything is different.
[anger] I have a new voice and now everybody will have no choice but to listen to my words !!!
You can mix emotions [sadness,disgust] or if you want to precise the weight of one or several emotions [sadness=0.7,disgust] (in any case total of weights is 1)
Remember two speakers mode requires to insert "Speaker 1:" & "Speaker 2:" to indicate who is talking.
There is only one snag: Index TTS 2 supports only English & Chinese. But dont' panic ! not all is lost. There is a workaround:
With this new release of WanGP you should have the best TTS (Text To Speech) experience you can find:
Qwen3 TTS Powered Up:
Heart Mula Powered Up:
Ace Step 1.5 Powered Up:
Also you now have a choice of multiple Prompt Enhancements for Qwen3 TTS & Kugel Audio: Prompt Enhancer can now generate for you either a Monologue or a Dialogue between two Speakers
Please note that to use the new Cuda Graph, mode you will need to select either vllm or cuda graph in Configuration / Performance / Language Models Decoder Engine. Profiles 1,3 or 3+ will need to be enabled for the corresponding Model. vllm is a powered up version of cuda graph that may not always work with all GPUs. But don't worry if it is not available for your GPU there will be an automatic fallback to cuda graph.
Ace Step 1.5 Turbo Super Charged: all the best features of Ace Step 1.5 are now in WanGP and are Fast & Easy to use:
LoKr support: this "Lora" like format has been tested with Flux Klein 9B
Optimized Int8 Kernels: all the Quantized INT8 checkpoints (most of the quantized checkpoints) used with WanGP should be now 10% faster !!!. You will need to install Triton. It is experimental, so for the moment it needs to be enabled manually in the Config / Performance tab. Please share your feedback on discord by mentioning your GPU so that I know if it works properly.
Auto Queue Saved if Gen Error: if for whatever reason you have got an error during a Gen, the queue will now be automatically saved. So you can try again this queue later (with a different config or when the related bug is fixed, if ever ...).
UI Updates (thx Tophness!): Updated the Self-Refiner UI to a dynamic, slider-based interface (no more manual text input). Improved queue reordering: items can now be dragged and dropped directly onto the Top and Bottom buttons while rearranging the queue in order to snap scroll to the top and bottom.
Kugel Audio Audio Split: Kugel Audio is a great model but strangely it tends to accelerate with long speeches. In order to avoid this effect, we need to split audio speeches. You can either do that manually by inserting an Empty Line or by specifiying an Auto Audio Split Duration (don't worry WanGP will try to split between lines or sentences).
update 10.81: Fixes
update 10.82: UI update
update 10.83: Kugel Audio Split
update 10.84: Ace Step RAM optimizations (fixed memory leak & reduce RAM requirements)
Note to RTX 50xx owners: you will need to upgrade to pytorch 2.10 (see upgrade procedure below) to be able to use Triton
The competition between Open Source & Close Source has never been that hot !
Please note that when using the Ace Step LM variants, this may get very slow with Memory Profiles 2 or 4 since the LM is an Autoregressive Model. It is why I recommed to stick to Memory Profiles 1/3/3+ unless you have very little VRAM.
Kugel Audio is entirely an Autoregressive Model and quite VRAM Hungry. So either you've got 16GB VRAM and you can run it with Memory Profile 1/3/3+ or you will have to go the slow way with other Profiles.
LTX-2 Base Tweaks: new Quality features if you found the base model was too fast :
Flux Klein 4B & 9B Base Models: Z Image has its base model in WanGP, so it was fair that Flux Klein would have its base model too. Base Models require more steps (up 50) and guidance > 1 but are good starting points for finetunes
The real novelty about this new release is that is has been tested and tuned to work with more recent versions of Python, Pytorch & Cuda. My end goal is to have everbody upgrade to Python 3.11, Pytorch 2.10, Cuda 13/13.1. Once we are all there it will be much easier to provide precompiled kernels for Nunchaku NVPF4, Sage Attention, Flash Attention, ... So please follow the manual upgrade instructions below (no Pinokio auto upgrade for the moment) and let me know on Discord if it works with all generations of GPUs (starting from GTX10xx to RTX50xx). You will find the kernels for this new setup in the guides/INSTALLATION.md.
Note that PyTorch 2.10 represents at last a decent upgrade, no memory leak when switching models (pytorch 2.8) and bad perfs / VRAM peaks with VAE decoding (pytorch 2.9).
Update: It seems GTX10xx doesnt support Cuda 13.0. Dont't worry I will keep WanGP compatibility with Pytorch 2.7.1 / Cuda 12.8.
Update 10.61: added Self Refiner
WanGP Special TTS (Text To Speech) Release:
Heart Mula: Suno quality song with lyrics on your local PC. You can generate up to 4 min of music.
Ace Step v1: while waiting for Ace Step v1.5 (which should be released very soon), enjoy this oldie (2025!) but goodie song generatpr as an appetizer. Ace Step v1 is a very fast Song generator. It is a Diffusion based, so dont hesitate to turn on Profile 4 to go as low as 4B VRAM while remaining fast.
Qwen 3 TTS: you can either do Voice Cloning, Generate a Custom Voice based on a Prompt or use a Predefined Voice
TTS Features:
Z Image Base: try it if you are into the Z Image hype but it will be probably useless for you unless you are a researcher and / or want to build a finetune out of it. This model requires from 35 to 50 steps (4x to 6x slower than Z Image turbo) and cfg > 1 (an additional 2x slower) and there is no Reinforcement Learning so Output Images wont be as good. The plus side is a higher diversity and Native Negative Prompt (versus Z Image virtual Negative Prompt using NAG).
Note that Z Image Base is very sensitive to the Attention Mode: it is not compatible with Sage 1 as it produces black frames. So I have disabled Sage for RTX 30xx. Also there are reports it produces some vertical banding artifacts with Sage 2
Flux 1/2 NAG : Flux 2 Klein is your new best friend but you miss Negative Prompts, NAG support for Distilled models will make you best buddies forever as NAG simulates Negative prompts.
Various Improvements:
update 10.51: new Heart Mula Finetune better at following instructions, Extra settings (cfg, top k) for TTS models, Rife v4
update 10.52: updated plugin list and added version tracking
update 10.53: video/audio galleries now support deletions
update 10.54: added Z Image Base, prompt enhancers improvements, configurable loras root folder
update 10.55: blocked Sage with Z Image on RTX30xx and added override attention mode settings, allowed changing config during generation
update 10.56: added NAG for Flux 1/2 & Ace Step v1
GPUs are expensive, RAM is expensive, SSD are expensive, sadly we live now in a GPU & RAM poor.
WanGP comes again to the rescue:
GGUF support: as some of you know, I am not a big fan of this format because when used with image / video generative models we don't get any speed boost (matrices multiplications are still done at 16 bits), VRAM savings are small and quality is worse than with int8/fp8. Still gguf has one advantage: it consumes less RAM and harddrive space. So enjoy gguf support. I have added ready to use Kijai gguf finetunes for LTX-2.
Models Manager PlugIn: use this Plugin to identify how much space is taken by each model / finetune and delete the ones you no longer use. Try to avoid deleting shared files otherwise they will be downloaded again.
LTX-2 Dual Video & Audio Control: you no longer need to extract the audio track of a Control Video if you want to use it as well to drive the video generation. New mode will allow you to use both motion and audio from Video Control.
LTX-2 - Custom VAE URL: some users have asked if they could use the old Distiller VAE instead of the new one. To do that, create a finetune def based on an existing model definition and save it in the finetunes/ folder with this entry (check the docs/FINETUNES.md doc):
"VAE_URLs": ["https://huggingface.co/DeepBeepMeep/LTX-2/resolve/main/ltx-2-19b_vae_old.safetensors"]
Flux 2 Klein 4B & 9B: try these distilled models as fast as Z_Image if not faster but with out of the box image edition capabiltities
Flux 2 & Qwen Outpainting + Lanpaint: the inpaint mode of these models support now outpainting + more combination possible with Lanpaint
RAM Optimizations for multi minutes Videos: processing, saving, spatial & Temporal upsampling very long videos should require much less RAM.
Text Encoder Cache: if you are asking a Text prompt already used recently with the current model, it will be taken straight from a cache. The cache is optimized to consume little RAM. It wont work with certain models such as Qwen where the Text Prompt is combined internally with an Image.
update 10.41: added Flux 2 klein
update 10.42: added RAM optimizations & Text Encoder Cache
update 10.43: added outpainting for Qwen & Flux 2, Lanpaint for Flux 2
So dont be surprised if the old checkpoints are deleted and new are downloaded !!!.
LTX-2 Multi Passes Loras multipliers: LTX-2 supports now loras multiplier that depend on the Pass No. For instance "1;0.5" means 1 will the strength for the first LTX-2 pass and 0.5 will be the strength for the second pass.
New Profile 3.5: here is the lost kid of Profile 3 & Profile 5, you got tons of VRAM, but little RAM ? Profile 3.5 will be your new friend as it will no longer use Reserved RAM to accelerate transfers. Use Profile 3.5 only if you can fit entirely a Diffusion / Transformer model in VRAM, otherwise the gen may be much slower.
NVFP4 Quantization for LTX-2 & Flux 2: you will now be able to load NV FP4 model checkpoints in WanGP. On top of Wan NV4 which was added recently, we now have LTX-2 (non distilled) & Flux 2 support. NV FP4 uses slightly less VRAM and up to 30% less RAM.
To enjoy fully the NV FP4 checkpoints (at least 30% faster gens), you will need a RTX 50xx and to upgrade to Pytorch 2.9.1 / Cuda 13 with the latest version of lightx2v kernels (check docs/INSTALLATION.md). To observe the speed gain, you have to make sure the workload is quite high (high res, long video).
With WanGP 10.21 HD 720p Video Gens of 10s just need now 8GB of VRAM!
LTX Team said this video gen was for 4k. So I had no choice but to squeeze more VRAM with further optimizations.
After much suffering I have managed to reduce by at least 1/3 the VRAM requirements of LTX-2, which means:
3K/4K resolutions will be available only if you enable them in the Config / General tab.
Ic Loras support: Use a Control Video to transfer Pose, Depth, Canny Edges. I have added some extra tweaks: with WanGP you can restrict the transfer to a masked area, define a denoising strength (how much the control video is going to be followed) and a masking strength (how much unmasked area is impacted)
Start Image Strength: This new slider will appear below a Start Image or Source Video. If you set it to values lower than 1 you may to reduce the static image effect, you get sometime with LTX-2 i2v
Custom Gemma Text Encoder for LTX-2: As a practical case, the Heretic text encoder is now supported by WanGP. Check the finetune doc, but in short create a finetune that has a text_encoder_URLS key that contains a list of one or more file paths or URLs.
Experimental Auto Recovery Failed Lora Pin: Some users (with usually PC with less than 64 GB of RAM) have reported Out Of Memory although a model seemed to load just fine when starting a gen with Loras. This is sometime related to WanGP attempting (and failing due to unsufficient reserved RAM) to pin the Loras to Reserved Memory for faster gen. I have experimented a recovery mode that should release sufficient ressources to continue the Video Gen. This may solve the oom crashes with LTX-2 Default (non distilled)
Max Loras Pinned Slider: If the Auto Recovery Mode is still not sufficient, I have added a Slider at the bottom of the Configuration / Performance tab that you can use to prevent WanGP from Pinning Loras (to do so set it to 0). As if there is no loading attempt there wont be any crash...
update 10.21: added slider Loras Max Pinning slider
update 10.22: added support for custom LTX-2 Text Encoder + Auto Recovery mode if Lora Pinning failed
update 10.23: Fixed text prompt ignore in profile 1 & 2 (this created random output videos)
With WanGP v10.11 you can now force your soundtrack, it works like Multitalk / Avatar except in theory it should work with any kind of sound (not just vocals). Thanks to Kijai for showing it was possible.
Z Image Twin Folder Turbo: Z Image even faster as this variant can generate images with as little as 1 step (3 steps recommend)
Qwen LanPaint: very precise In Painting, offers a better integration of the inpainted area in the rest of the image. Beware it is up to 5x slower as it "searches" for the best replacement.
Optimized Pytorch Compiler : Patience is the Mother of Virtue. Finally I may (or may not) have fixed the PyTorch compiler with the Wan models. It should work in much diverse situations and takes much less time.
LongCat Video: experimental support which includes LongCat Avatar a talking head model. For the moment it is mostly for models collectors as it is very slow. It needs 40+ steps and each step contains up 3 passes.
MMaudio NSFW: for alternative audio background
update v10.11: LTX-2, use your own soundtrack
See full changelog: Changelog
One-click installation:
Get started instantly with Pinokio App
It is recommended to use in Pinokio the Community Scripts wan2gp or wan2gp-amd by Morpheus rather than the official Pinokio install.
Manual installation: (old python 3.10, to be deprecated)
git clone https://github.com/deepbeepmeep/Wan2GP.git
cd Wan2GP
conda create -n wan2gp python=3.10.9
conda activate wan2gp
pip install torch==2.7.1 torchvision torchaudio --index-url https://download.pytorch.org/whl/test/cu128
pip install -r requirements.txt
Manual installation: (new python 3.11 setup)
git clone https://github.com/deepbeepmeep/Wan2GP.git
cd Wan2GP
conda create -n wan2gp python=3.11.14
conda activate wan2gp
pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
Run the application:
python wgp.py
First time using WanGP ? Just check the Guides tab, and you will find a selection of recommended models to use.
Update the application (stay in the old pyton / pytorch version): If using Pinokio use Pinokio to update otherwise: Get in the directory where WanGP is installed and:
git pull
conda activate wan2gp
pip install -r requirements.txt
Upgrade to 3.11, Pytorch 2.10, Cuda 13/13.1 (for non GTX10xx users) I recommend creating a new conda env for the Python 3.11 to avoid bad surprises. Let's call the new conda env wangp (instead of wan2gp the old name of this project) Get in the directory where WanGP is installed and:
git pull
conda create -n wangp python=3.11.9
conda activate wangp
pip install torch==2.10.0 torchvision torchaudio --index-url https://download.pytorch.org/whl/cu130
pip install -r requirements.txt
Git Errors Once you are done you will have to reinstall Sage Attention, Triton, Flash Attention. Check the Installation Guide -
if you get some error messages related to git, you may try the following (beware this will overwrite local changes made to the source code of WanGP):
git fetch origin && git reset --hard origin/main
conda activate wangp
pip install -r requirements.txt
When you have the confirmation it works well you can then delete the old conda env:
conda uninstall -n wan2gp --all
Run headless (batch processing):
Process saved queues without launching the web UI:
# Process a saved queue
python wgp.py --process my_queue.zip
Create your queue in the web UI, save it with "Save Queue", then process it headless. See CLI Documentation for details.
For Debian-based systems (Ubuntu, Debian, etc.):
./run-docker-cuda-deb.sh
This automated script will:
Docker environment includes:
Supported GPUs: RTX 40XX, RTX 30XX, RTX 20XX, GTX 16XX, GTX 10XX, Tesla V100, A100, H100, and more.
For detailed installation instructions for different GPU generations:
For detailed installation instructions for different GPU generations:
Made with β€οΈ by DeepBeepMeep
Python
97.7%
JavaScript
1.2%