Read our How to Run Qwen-Image-2.1 Guide! 💜
107
15 commits
2 linked in READMEs
updated Sep 22, 2026
This is a GGUF quantized version of Qwen-Image-2.1.
unsloth/Qwen-Image-2.1-GGUF uses Unsloth Dynamic 2.0 methodology for SOTA performance.
vae/qwen_image_2.1_vae_bf16.safetensors. Text encoder: unsloth/Qwen3-VL-8B-Instruct-GGUF Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf, the Dynamic 2.0 4-bit rung rather than the uniform Q4_K_M. Measured against the Q4_K_M encoder at a shared seed, with the denoiser and VAE held fixed: LPIPS 0.029, SSIM 0.959, 5.15 GB vs 5.03 GB, 36.5 s vs 39.0 s.
sd-cli --diffusion-model qwen-image-2.1-Q4_K_M.gguf \
--vae qwen_image_2.1_vae_bf16.safetensors \
--llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf \
-p "a cartoon sloth mascot waving, flat vector illustration, bright colours" \
--steps 20 --cfg-scale 6.0 --sampling-method euler -W 1024 -H 1024 --diffusion-fa \
-o out.png
Rendered with the Q4_K_M denoiser and the Q4_K_M text encoder, 1024x1024, 20 steps, cfg 6.0, euler.
![]() | ![]() |
![]() | ![]() |
🤖 ModelScope | 🤗 HuggingFace | 📑 Blog | 🖥️ Demo | 🫨 Discord | 💬 WeChat
We are excited to open-source Qwen-Image-2.1, a unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Four key improvements define this release:
For more details, see the GitHub repo and Blog.
pip install torch>=2.4.0
pip install transformers>=5.17
pip install git+https://github.com/huggingface/diffusers
pip install accelerate pillow
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_example.png")
import torch
from PIL import Image
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
input_image = Image.open("input.png")
image = pipe(
prompt="Change the background to a sunset beach",
image=input_image,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("edit_example.png")
Use the recommended prompt format for transparent images:
image = pipe(
prompt="This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("transparent_example.png")
aspect_ratios = {
"1:1": (2048, 2048),
"4:3": (2400, 1792),
"3:4": (1792, 2400),
"3:2": (2528, 1696),
"2:3": (1696, 2528),
"16:9": (2752, 1536),
"9:16": (1536, 2752),
}
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
Native transparent image generation
Group photograph generated from six portrait references
Text rendering
This model is licensed under the Qwen Research License Agreement.
15 commits
Read our How to Run Qwen-Image-2.1 Guide! 💜
107
15 commits
2 linked in READMEs
updated Sep 22, 2026
This is a GGUF quantized version of Qwen-Image-2.1.
unsloth/Qwen-Image-2.1-GGUF uses Unsloth Dynamic 2.0 methodology for SOTA performance.
vae/qwen_image_2.1_vae_bf16.safetensors. Text encoder: unsloth/Qwen3-VL-8B-Instruct-GGUF Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf, the Dynamic 2.0 4-bit rung rather than the uniform Q4_K_M. Measured against the Q4_K_M encoder at a shared seed, with the denoiser and VAE held fixed: LPIPS 0.029, SSIM 0.959, 5.15 GB vs 5.03 GB, 36.5 s vs 39.0 s.
sd-cli --diffusion-model qwen-image-2.1-Q4_K_M.gguf \
--vae qwen_image_2.1_vae_bf16.safetensors \
--llm Qwen3-VL-8B-Instruct-UD-Q4_K_XL.gguf \
-p "a cartoon sloth mascot waving, flat vector illustration, bright colours" \
--steps 20 --cfg-scale 6.0 --sampling-method euler -W 1024 -H 1024 --diffusion-fa \
-o out.png
Rendered with the Q4_K_M denoiser and the Q4_K_M text encoder, 1024x1024, 20 steps, cfg 6.0, euler.
![]() | ![]() |
![]() | ![]() |
🤖 ModelScope | 🤗 HuggingFace | 📑 Blog | 🖥️ Demo | 🫨 Discord | 💬 WeChat
We are excited to open-source Qwen-Image-2.1, a unified text-to-image generation and image editing model in the Qwen family. With just 7B parameters in its visual generation component (32 Single-Stream DiT layers), Qwen-Image-2.1 balances generation quality, inference efficiency, and versatility.
Four key improvements define this release:
For more details, see the GitHub repo and Blog.
pip install torch>=2.4.0
pip install transformers>=5.17
pip install git+https://github.com/huggingface/diffusers
pip install accelerate pillow
import torch
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
image = pipe(
prompt="A neon shop sign that reads \"QWEN IMAGE 2.1\", rainy night, reflections on wet pavement",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("t2i_example.png")
import torch
from PIL import Image
from diffusers import QwenImage21Pipeline
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
).to("cuda")
input_image = Image.open("input.png")
image = pipe(
prompt="Change the background to a sunset beach",
image=input_image,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("edit_example.png")
Use the recommended prompt format for transparent images:
image = pipe(
prompt="This is an RGBA image with transparency. A cute cartoon dragon sticker. The image has alpha channel and the background is transparent.",
width=2048, height=2048,
num_inference_steps=40,
generator=torch.Generator("cuda").manual_seed(42),
).images[0]
image.save("transparent_example.png")
aspect_ratios = {
"1:1": (2048, 2048),
"4:3": (2400, 1792),
"3:4": (1792, 2400),
"3:2": (2528, 1696),
"2:3": (1696, 2528),
"16:9": (2752, 1536),
"9:16": (1536, 2752),
}
pipe = QwenImage21Pipeline.from_pretrained(
"Qwen/Qwen-Image-2.1", torch_dtype=torch.bfloat16
)
pipe.enable_model_cpu_offload()
Native transparent image generation
Group photograph generated from six portrait references
Text rendering
This model is licensed under the Qwen Research License Agreement.
15 commits