madebyollin/texture-fix-vae-for-qwen-image-2.1

Model

Texture-Fix-VAE-for-Qwen-Image-2.1

114

19 commits

2 linked in READMEs

updated Oct 6, 2026

See the code

README

Texture-Fix-VAE-for-Qwen-Image-2.1

Built with Qwen* (an unofficial finetune of the Qwen-Image-2.1 VAE)

Texture-Fix-VAE-for-Qwen-Image-2.1 is the Qwen-Image-2.1 VAE, but finetuned to produce cleaner textures with no checkerboard artifacts (and no NaNs when running in fp16).

Comparison Settings

Texture-Fix-VAE-for-Qwen-Image-2.1's improved decoding is most noticeable in detailed, photo-style images. The latents for the VAE comparison image below were generated by Qwen-Image-2.1 from the photo-style prompt:

Landscape photograph of a subalpine wildflower meadow in the Pacific Northwest in midsummer: a clear mountain stream winding over mossy boulders through purple lupine and red paintbrush, dense old-growth Douglas fir and western red cedar forest behind, a snow-capped volcano in the distance, golden late-afternoon light, highly detailed

Qwen-Image-2.1-VAE (full-res)πŸͺ„ Texture-Fix-VAE-for-Qwen-Image-2.1 πŸͺ„ (full-res)

Usage

ComfyUI

Download texture_fix_vae_for_qwen_image_2.1_bf16.safetensors into ComfyUI/models/vae/ and select it in the Load VAE node, in place of qwen_image_2.1_vae_bf16.safetensors. It also works with --fp16-vae.

🧨 Diffusers

import torch
from diffusers import QwenImage21Pipeline, AutoencoderKLQwenImage21

vae = AutoencoderKLQwenImage21.from_pretrained("madebyollin/texture-fix-vae-for-qwen-image-2.1", torch_dtype=torch.bfloat16)
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", vae=vae, torch_dtype=torch.bfloat16).to("cuda")

Mechanism

Fixing Checkerboard Artifacts

Texture-Fix-VAE-for-Qwen-Image-2.1 was cured of checkerboard artifacts by finetuning the highest-resolution blocks briefly using the adversarial recipe developed for TAESD.

The TAESD recipe, like most image autoencoder training recipes, uses a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial (GAN) loss terms. Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (among other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts. This figure from DC-AE shows the importance of including adversarial (GAN) loss:

Demo of the effects of adversarial loss, courtesy of the DC-AE paper

Figure: from Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models (Chen et al., 2024, arXiv:2410.10733), licensed under CC BY 4.0; cropped to the first two rows.

I suspect the original Qwen-Image-2.1-VAE was trained without a working adversarial loss term.

Fixing NaNs in FP16

Texture-Fix-VAE-for-Qwen-Image-2.1 runs fine in fp16, whereas the original Qwen-Image-2.1 VAE decoder often produces NaNs (the πŸŸͺ magenta-highlighted regions below):

Qwen-Image-2.1-VAE (bf16)⚠️ Qwen-Image-2.1-VAE (fp16)Texture-Fix-VAE-for-Qwen-Image-2.1 (fp16)

The original VAE decoder mostly produces NaNs in fully-transparent regions, but they bleed into surrounding content. You can reproduce this visualization with fp16_demo.py.

Texture-Fix-VAE-for-Qwen-Image-2.1 was cured of the NaNs-in-FP16 issue by:

  1. Finetuning the decoder weights to keep outputs the same while reducing activation magnitudes (like SDXL-VAE-FP16-Fix)
  2. Manually shrinking certain weights to further reduce the scale of the residual stream (every residual branch starts with an RMSNorm, so this doesn't change the outputs)

Metrics

Texture-Fix-VAE-for-Qwen-Image-2.1 makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) slightly worse.

MetricQwen-Image-2.1-VAE (bf16)Texture-Fix-VAE-for-Qwen-Image-2.1 (bf16)Texture-Fix-VAE-for-Qwen-Image-2.1 (fp16)
rFID ↓ (COCO val2017, 5000 images @ 256Β²)3.382.072.06
PSNR ↑ (COCO val2017 @ 256Β²)33.2932.7332.79
LPIPS ↓ (COCO val2017 @ 256Β²)0.03570.03840.0380
PSNR ↑ (DIV2K valid, native 1024Β² crops)32.8532.3532.40
LPIPS ↓ (DIV2K valid, native 1024Β² crops)0.04600.04900.0486

Texture-Fix-VAE-for-Qwen-Image-2.1 also reduces the internal activation magnitude, preventing NaNs/overflows in fp16.

fp16 Metric (encoder + decoder in fp16)Qwen-Image-2.1-VAETexture-Fix-VAE-for-Qwen-Image-2.1
Largest decoder activation ↓ (fp16 max is 65504)~3.5M (overflows)~1.3k
Inputs with NaN outputs ↓ (185 stress-test inputs: transparent / opaque / synthetic images and random latents, 256Β² to 2048Β²)750

Version History

  • 2026-10-05 Updated release; now also fixes decoder NaNs in fp16, metrics are similar
    • Finetuned the decoder like SDXL-VAE-FP16-Fix (scale + bias of each conv and the RMSNorm gains, 61k parameters) to match the initial release while penalizing large activations, with extra transparent training crops (~60k steps)
    • Re-ran the initial release's finetune (8000 steps) to restore texture detail
    • Rescaled the weights by exact powers of 2 (decoder /512, encoder /8)
  • 2026-09-24 Initial release; finetuned decoder to fix the checkerboard artifacts
    • ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters)

Attribution Notice

This fine-tuned VAE is based on Qwen/Qwen-Image-2.1; original materials Β© 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd., licensed under the Qwen RESEARCH LICENSE AGREEMENT (see LICENSE) for non-commercial/research use only. Built with Qwen*.

* In the sense that the initial VAE weights are from Qwen-Image. The decoder fine-tuning work was performed by madebyollin and Claude Opus.

diffusers
qwen-image
safetensors

madebyollin/texture-fix-vae-for-qwen-image-2.1

Model

Texture-Fix-VAE-for-Qwen-Image-2.1

114

19 commits

2 linked in READMEs

updated Oct 6, 2026

See the code

README

Texture-Fix-VAE-for-Qwen-Image-2.1

Built with Qwen* (an unofficial finetune of the Qwen-Image-2.1 VAE)

Texture-Fix-VAE-for-Qwen-Image-2.1 is the Qwen-Image-2.1 VAE, but finetuned to produce cleaner textures with no checkerboard artifacts (and no NaNs when running in fp16).

Comparison Settings

Texture-Fix-VAE-for-Qwen-Image-2.1's improved decoding is most noticeable in detailed, photo-style images. The latents for the VAE comparison image below were generated by Qwen-Image-2.1 from the photo-style prompt:

Landscape photograph of a subalpine wildflower meadow in the Pacific Northwest in midsummer: a clear mountain stream winding over mossy boulders through purple lupine and red paintbrush, dense old-growth Douglas fir and western red cedar forest behind, a snow-capped volcano in the distance, golden late-afternoon light, highly detailed

Qwen-Image-2.1-VAE (full-res)πŸͺ„ Texture-Fix-VAE-for-Qwen-Image-2.1 πŸͺ„ (full-res)

Usage

ComfyUI

Download texture_fix_vae_for_qwen_image_2.1_bf16.safetensors into ComfyUI/models/vae/ and select it in the Load VAE node, in place of qwen_image_2.1_vae_bf16.safetensors. It also works with --fp16-vae.

🧨 Diffusers

import torch
from diffusers import QwenImage21Pipeline, AutoencoderKLQwenImage21

vae = AutoencoderKLQwenImage21.from_pretrained("madebyollin/texture-fix-vae-for-qwen-image-2.1", torch_dtype=torch.bfloat16)
pipe = QwenImage21Pipeline.from_pretrained("Qwen/Qwen-Image-2.1", vae=vae, torch_dtype=torch.bfloat16).to("cuda")

Mechanism

Fixing Checkerboard Artifacts

Texture-Fix-VAE-for-Qwen-Image-2.1 was cured of checkerboard artifacts by finetuning the highest-resolution blocks briefly using the adversarial recipe developed for TAESD.

The TAESD recipe, like most image autoencoder training recipes, uses a mix of PSNR-focused (MSE/MAE), LPIPS, and adversarial (GAN) loss terms. Whenever precise details can't be reconstructed, MSE/MAE loss encourages blurring, LPIPS loss encourages blurring+checkerboarding (among other artifacts), and adversarial loss encourages generating sharp/plausible (but fake) detail without obvious artifacts. This figure from DC-AE shows the importance of including adversarial (GAN) loss:

Demo of the effects of adversarial loss, courtesy of the DC-AE paper

Figure: from Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models (Chen et al., 2024, arXiv:2410.10733), licensed under CC BY 4.0; cropped to the first two rows.

I suspect the original Qwen-Image-2.1-VAE was trained without a working adversarial loss term.

Fixing NaNs in FP16

Texture-Fix-VAE-for-Qwen-Image-2.1 runs fine in fp16, whereas the original Qwen-Image-2.1 VAE decoder often produces NaNs (the πŸŸͺ magenta-highlighted regions below):

Qwen-Image-2.1-VAE (bf16)⚠️ Qwen-Image-2.1-VAE (fp16)Texture-Fix-VAE-for-Qwen-Image-2.1 (fp16)

The original VAE decoder mostly produces NaNs in fully-transparent regions, but they bleed into surrounding content. You can reproduce this visualization with fp16_demo.py.

Texture-Fix-VAE-for-Qwen-Image-2.1 was cured of the NaNs-in-FP16 issue by:

  1. Finetuning the decoder weights to keep outputs the same while reducing activation magnitudes (like SDXL-VAE-FP16-Fix)
  2. Manually shrinking certain weights to further reduce the scale of the residual stream (every residual branch starts with an RMSNorm, so this doesn't change the outputs)

Metrics

Texture-Fix-VAE-for-Qwen-Image-2.1 makes perceptual quality metrics (rFID) better and reconstruction accuracy metrics (LPIPS/PSNR) slightly worse.

MetricQwen-Image-2.1-VAE (bf16)Texture-Fix-VAE-for-Qwen-Image-2.1 (bf16)Texture-Fix-VAE-for-Qwen-Image-2.1 (fp16)
rFID ↓ (COCO val2017, 5000 images @ 256Β²)3.382.072.06
PSNR ↑ (COCO val2017 @ 256Β²)33.2932.7332.79
LPIPS ↓ (COCO val2017 @ 256Β²)0.03570.03840.0380
PSNR ↑ (DIV2K valid, native 1024Β² crops)32.8532.3532.40
LPIPS ↓ (DIV2K valid, native 1024Β² crops)0.04600.04900.0486

Texture-Fix-VAE-for-Qwen-Image-2.1 also reduces the internal activation magnitude, preventing NaNs/overflows in fp16.

fp16 Metric (encoder + decoder in fp16)Qwen-Image-2.1-VAETexture-Fix-VAE-for-Qwen-Image-2.1
Largest decoder activation ↓ (fp16 max is 65504)~3.5M (overflows)~1.3k
Inputs with NaN outputs ↓ (185 stress-test inputs: transparent / opaque / synthetic images and random latents, 256Β² to 2048Β²)750

Version History

  • 2026-10-05 Updated release; now also fixes decoder NaNs in fp16, metrics are similar
    • Finetuned the decoder like SDXL-VAE-FP16-Fix (scale + bias of each conv and the RMSNorm gains, 61k parameters) to match the initial release while penalizing large activations, with extra transparent training crops (~60k steps)
    • Re-ran the initial release's finetune (8000 steps) to restore texture detail
    • Rescaled the weights by exact powers of 2 (decoder /512, encoder /8)
  • 2026-09-24 Initial release; finetuned decoder to fix the checkerboard artifacts
    • ~5000 steps at learning rate 3e-5, with only the two highest-resolution decoder stages and output head unfrozen (7.5M trainable parameters)

Attribution Notice

This fine-tuned VAE is based on Qwen/Qwen-Image-2.1; original materials Β© 2026 Hangzhou Tongyi Laboratory Technology Co., Ltd., licensed under the Qwen RESEARCH LICENSE AGREEMENT (see LICENSE) for non-commercial/research use only. Built with Qwen*.

* In the sense that the initial VAE weights are from Qwen-Image. The decoder fine-tuning work was performed by madebyollin and Claude Opus.

diffusers
qwen-image
safetensors