hypersniper05/MCP-Image-Generator-Uncensored

Self-hosted MCP server for image generation and editing with the uncensored Qwen-Image-2.1 model

Python

0

1 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen Image 2.1 Uncensored MCP (r/StableDiffusion)

[https://github.com/hypersniper05/MCP-Image-Generator-Uncensored](https://github.com/hypersniper05/MCP-Image-Generator-Uncensored) Just wanted to share with you guys my workflow converted into an MCP. It's a Qwen Image 2.1 Uncensored MCP with a few extra models I found helpful: a watermark-removal…

4

Oct 6, 2026

README

MCP Image Generator (Uncensored)

A self-hosted Model Context Protocol server that creates and edits images with the uncensored Qwen-Image-2.1 model on your own computer. It runs in Docker, on an NVIDIA GPU or on the CPU, and any MCP client on your network can use it over Streamable HTTP.

The default model is an uncensored build of Qwen-Image-2.1 with no built-in content filter. To use the standard model instead, set model.variant: base in config.yaml.

  • Text to image in any size or aspect ratio, up to about 4 megapixels (2048x2048)
  • Image editing: change, add or remove things, restyle, or combine up to 10 images
  • Transparent backgrounds (RGBA PNG) and background removal
  • Seamless tileable textures, including height and normal maps that stay seamless
  • 360 panoramas with a built-in 360 viewer
  • Upscaling 2x or 4x, and watermark removal
  • Models download automatically on the first start

Examples

All images below were made by this server with the default quality settings. The prompts are listed under the gallery.

A lighthouse on a rocky coast at sunset
Text to image
The same scene edited to a snowy winter night
Edit: the image on the left, made into a snowy winter night
A vintage travel poster that reads SEE THE ALPS
Text in images: text in quotes is written as given
Iron Man standing on a rooftop at night
Uncensored: famous characters, which many image services refuse
A cartoon rocket icon on a transparent background
Icon with no background (transparent PNG)
A red fox on a transparent background
Transparent background
The lighthouse photo covered with a tiled SAMPLE watermark and a corner badge
Watermark removal: before (a test watermark added to the first image)
The same photo with the watermark removed
Watermark removal: after remove_watermark

An equirectangular 360 panorama of a mountain meadow
360 panorama, opened in the built-in 360 viewer from the link in the result

Seamless textures

Each texture below is one tile (top). Repeated 3x3 (bottom), it shows no seams.

A red brick wall tileAn oak wood planks tileA blue and white ceramic tile
The brick tile repeated three by threeThe wood tile repeated three by threeThe ceramic tile repeated three by three

A stone tile, its height map and its normal map, each repeated two by two
Texture maps: a stone tile, then a height map and a normal map made from it with an edit. Each is shown repeated 2x2: they all stay seamless.

A detail enlarged four times: plain resize on the left, AI upscaler on the right
Upscaling 4x. Left: plain enlargement. Right: the upscaler.

Prompts and settings
ImageTool and settingsPrompt
Text to imagegenerate_image, 1344x768, seed 20261006A lighthouse on a rocky coast at sunset, waves breaking on the rocks, warm golden light, a small fishing boat in the distance, dramatic clouds, photorealistic
Editedit_image with the image above, seed 7Make it a snowy winter night with the lighthouse beam switched on and snow on the rocks, keep everything else unchanged
Text in imagesgenerate_image, 768x1024, seed 31A vintage travel poster, flat screen-print illustration of snowy mountains above a lake with a red train on a bridge, bold title text at the top that reads "SEE THE ALPS", smaller text at the bottom that reads "By rail, every season"
Uncensoredgenerate_image, 768x1024, seed 505Iron Man in his red and gold armor standing on a rooftop at night, city lights below, cinematic lighting, photorealistic
Icongenerate_image, 1024x1024, transparent: true, seed 404A cute cartoon rocket ship, flat vector illustration, bold outlines, bright colors
Transparentgenerate_image, 1024x1024, transparent: true, seed 11A red fox sitting, full body, soft studio light
360 panoramagenerate_panorama, 2048x1024, seed 2026A mountain meadow with wildflowers, a stone cabin, a wooden boardwalk and a lake
Brickgenerate_image, 512x512, tileable: true, seed 111Large red bricks with light grey mortar, close-up, photorealistic, even lighting
Woodgenerate_image, 512x512, tileable: true, seed 222Wide oak floor planks with visible wood grain and knots, top-down, natural light
Ceramicgenerate_image, 512x512, tileable: true, seed 303Blue and white ceramic tiles with a floral pattern
Stone tilegenerate_image, 512x512, tileable: true, seed 512A moss-covered stone floor, top-down
Height mapedit_image with the stone tile, seed 1Convert <image1> into a grayscale height map for a game material: white = high stone tops, black = low grout and gaps, keep the exact layout of every stone
Normal mapedit_image with the stone tile, seed 2Convert <image1> into a normal map for a game material, keep the exact layout of every stone
Watermark removalremove_watermark, seed 9, on the text-to-image result with a tiled "SAMPLE" text and a corner badge added- (no prompt needed)
Upscalingupscale_image, scale: 4 on the text-to-image result-

Characters and brands shown belong to their owners.

Requirements

GPU modeCPU mode
DockerDocker Desktop (Windows, macOS) or Docker Engine with Compose (Linux)same
HardwareNVIDIA GPU, RTX 20xx or newer, 12 GB VRAM or more (see GPU memory), 32 GB RAM recommendedAny modern 64-bit CPU, 24 GB RAM or more
DriverNVIDIA driver 570 or newer-
DiskAbout 12 GB for models and 6 GB for the Docker imageAbout 12 GB and 1 GB

GPU setup

Check that Docker can see your GPU:

docker run --rm --gpus all nvidia/cuda:12.8.2-base-ubuntu24.04 nvidia-smi

GPU memory

Every feature, including 2048x2048 images, edits with 10 input images, 2880x1440 panoramas and upscaling to 8192 px, works on a single 12 GB card:

GPU memorySettingsNotes
24 GB or moredefaultEverything stays on the GPU, the fastest setup. Peak use is about 18 GB
12-16 GBthe settings belowPeak use is about 11 GB. The text encoder's weights stay in system RAM, so the server uses up to about 15 GB of RAM
8-10 GBoffload: cpu and smaller sizesNot tested. All weights stream from RAM: slower, and the largest sizes may not fit

Settings for a 12 GB card, in config.yaml:

gpu:
  max_vram_gb: {0: 10}
  offload: {text_encoder: cpu}
generation:
  prefix_cache_type: q8_0

Quick start

git clone https://github.com/hypersniper05/MCP-Image-Generator-Uncensored.git
cd MCP-Image-Generator-Uncensored
./start.sh            # Linux / macOS
.\start.cmd           # Windows (or double-click start.cmd)

The first start builds the Docker image (10-30 minutes) and downloads about 12 GB of models. When it is ready, the script prints the address:

Ready. MCP endpoint (Streamable HTTP, no auth):
    http://localhost:5005/mcp
  • Stop: ./stop.sh or stop.cmd. Logs: docker compose logs -f.
  • Web page: open http://localhost:5005/ to see the status, upload images and browse recent results.
  • Settings: the first start creates config.yaml from config.example.yaml. To run on the CPU, set device: cpu in config.yaml and run the start script again.
  • Updates: after pulling new code, run ./start.sh --build (or start.cmd -Build) to rebuild the image.
Start without the scripts
cp config.example.yaml config.yaml
cp .env.example .env                  # for CPU mode, set COMPOSE_PROFILES=cpu in .env
docker compose up -d --build

The server has no authentication and listens on all network interfaces, so other computers on your network can reach it at http://<this-computer's-ip>:5005/mcp. Only run it on a network you trust.

Connect an MCP client

The endpoint is http://<host>:5005/mcp (Streamable HTTP).

MCP Inspector (quick test in a browser): run npx @modelcontextprotocol/inspector, choose Streamable HTTP, enter http://localhost:5005/mcp and click Connect.

VS Code (.vscode/mcp.json):

{ "servers": { "imagegen": { "type": "http", "url": "http://localhost:5005/mcp" } } }

Cursor (~/.cursor/mcp.json) and most other clients:

{ "mcpServers": { "imagegen": { "url": "http://localhost:5005/mcp" } } }

Clients that only support stdio can connect through mcp-remote:

{ "mcpServers": { "imagegen": { "command": "npx", "args": ["-y", "mcp-remote", "http://localhost:5005/mcp", "--allow-http"] } } }

Large images take minutes. If an image is not ready within 50 seconds, the tool returns a job_id and the client gets the image later with get_job, so clients with short timeouts still work.

llama.cpp web UI and optional request headers

For the llama.cpp web UI (llama-server started with --ui-mcp-proxy), add the server in the web UI's MCP settings, or for every browser through the file passed with --ui-config-file:

{
  "mcpServers": "[{\"id\": \"imagegen\", \"name\": \"Image Gen\", \"url\": \"http://127.0.0.1:5005/mcp\", \"enabled\": true, \"useProxy\": true, \"headers\": \"{\\\"X-Imagegen-Max-Wait\\\": \\\"20\\\", \\\"X-Inline-Max-Bytes\\\": \\\"16000000\\\"}\"}]"
}

Optional headers a client can send:

HeaderEffect
X-Imagegen-Max-Wait: 20Wait at most this many seconds before returning a job_id
X-Inline-Max-Bytes: 16000000Send the full image inline instead of a preview (for clients without a message size limit)
X-Imagegen-Inline: data-uri-textReturn images as data-URI text, for clients that drop MCP image content
X-Forwarded-Host: 192.168.1.50:5005Host name to use in returned links when the client connects through a proxy

Tools

ToolWhat it does
generate_imageText to image. Options: size (small, medium, large, xl) or width and height, aspect_ratio, transparent, tileable, seed, steps, negative_prompt
edit_imageEdit one image or combine up to 10 (refer to them as <image1>, <image2>, ...). Optional mask. Edits of a seamless tile stay seamless and keep its size
generate_panorama360 panorama (2:1). Optional image to turn a photo into a full 360
remove_backgroundCut out the subject into a transparent PNG
upscale_imageEnlarge 2x or 4x (up to 8192 px per side). Panoramas and tiles stay seamless
remove_watermarkRemove watermarks, logos and overlaid text; the rest of the image is kept as it was
get_job, cancel_jobGet the result of, or cancel, a job that was still running
list_images, view_imageRecent results and uploads, and a way for the model to look at one
server_statusModel, GPU placement, download progress and running jobs

Input images can be a data URL or base64, an http(s) URL, a file name or link of an image this server made, or an image uploaded on the server's web page (http://<host>:5005/upload).

Results include the image (or a preview of a large one) and a link to the full file, which is saved in ./outputs. Panoramas also get a link to the 360 viewer.

Tips

  • For the best quality use size: "xl". Leave steps and cfg_scale unset.
  • Write full sentences: subject, setting, lighting, style. Put text that should appear in quotes: a neon sign that says "OPEN 24/7".
  • For edits, say what to change and add "keep everything else unchanged".
  • For tiles, describe a surface or pattern that fills the whole picture, e.g. "moss-covered cobblestones, top-down".

Configuration

All settings are in config.yaml (created from the commented config.example.yaml). Restart after a change: docker compose restart. The most useful ones:

SettingDefaultWhat it does
devicegpugpu or cpu
model.quantQ4_K_MModel size: Q4_K_M (4.6 GB), Q6_K (5.9 GB) or Q8_0 (7.6 GB)
generation.steps40Quality vs. speed. 25 is a faster draft
generation.default_size1024x1024Size when a request gives none
generation.panorama_size2048x1024Default panorama size (2880x1440 at most)
generation.wait_seconds50How long a tool waits before returning a job_id
generation.idle_unload_seconds0Free the GPU memory after this many idle seconds (0 = keep loaded)
outputs.keep_days7Delete old results after this many days (0 = never)
server.port5005Port of the server
server.public_urlemptyAddress used in returned links, e.g. http://192.168.1.50:5005

.env (created from .env.example) holds Docker settings such as PUBLIC_URL, HF_TOKEN (only needed if Hugging Face rate-limits your downloads) and the build options.

config.yaml and .env are your local files and are not part of the repository.

Choosing GPUs

GPU numbers are the ones nvidia-smi shows. Everything runs on GPU 0 by default. The model has three parts, and each can go on a different GPU:

gpu:
  diffusion: [0]     # the main model, runs every step: use the fastest GPU
  text_encoder: 1    # reads the prompt once per request
  vae: 0             # turns the result into pixels

On a shared computer, only the GPUs you list are used. For cards with less memory, see GPU memory.

Troubleshooting

  • No GPU found: run the nvidia-smi check from Requirements. On Linux, install the NVIDIA Container Toolkit. Or use device: cpu.
  • Out of GPU memory: use the 12 GB settings, a smaller size, or move the text encoder to another GPU.
  • The client times out: lower generation.wait_seconds below the client's timeout.
  • Other computers cannot connect: use this computer's IP address instead of localhost, allow port 5005 in the firewall, and set server.public_url so image links work.
  • CPU mode is very slow or crashes on Windows: Docker Desktop gives WSL only half your RAM. Raise it in %UserProfile%\.wslconfig ([wsl2] then memory=24GB), run wsl --shutdown and restart Docker Desktop.
  • The server shows an error: read docker compose logs --tail 100, fix the cause, and run the start script again.

Development

python -m venv .venv && . .venv/bin/activate
pip install -e ".[test]"
pytest

python -m imagegen_mcp --config config.yaml --check validates a config file without starting anything. python scripts/smoke_test.py http://localhost:5005/mcp tests a running server.

Images are made by stable-diffusion.cpp, built inside the Docker image from a pinned release with one small patch (docker/patches/sdcpp-circular-json.patch) that turns on its wrap-around mode per request for tileable images.

Credits and licenses

The code in this repository is MIT licensed (see LICENSE). It builds on the work of others; the model files are downloaded from their original sources and keep their own licenses:

PartByLicense
Qwen-Image-2.1 image modelQwen team, AlibabaQwen Research License (non-commercial)
Uncensored Qwen-Image-2.1 GGUF (the default model files)abenzerpsQwen Research License (non-commercial)
Texture-fix VAE (image decoder)madebyollinQwen Research License (non-commercial)
Qwen3-VL-8B-Instruct GGUF (prompt and image encoder)Qwen team, AlibabaApache-2.0
Watermark removal LoRA (v1.0 for Qwen 2.1)saladinCivitai license: no commercial use
BiRefNet background removal, via rembgPeng Zheng et al.; Daniel GatisMIT (the optional isnet-general-use model: Apache-2.0)
4xNomos2_otf_esrgan upscaler (models)Philip HofmannCC BY 4.0
Real-ESRGAN (RealESRGAN_x4plus, optional upscaler)Xintao WangBSD-3-Clause
stable-diffusion.cpp and ggml (the inference engine)leejet and contributorsMIT
Pannellum (the 360 viewer)Matthew PetroffMIT

The Qwen-Image-2.1 files allow non-commercial use only: read their license before using images commercially. Only remove watermarks from images you have the rights to edit. The :gpu Docker image is based on NVIDIA's CUDA image (license).

The uncensored model has no built-in content filter. You are responsible for how you use it and what you make with it.

hypersniper05/MCP-Image-Generator-Uncensored

Self-hosted MCP server for image generation and editing with the uncensored Qwen-Image-2.1 model

Python

0

1 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

Qwen Image 2.1 Uncensored MCP (r/StableDiffusion)

[https://github.com/hypersniper05/MCP-Image-Generator-Uncensored](https://github.com/hypersniper05/MCP-Image-Generator-Uncensored) Just wanted to share with you guys my workflow converted into an MCP. It's a Qwen Image 2.1 Uncensored MCP with a few extra models I found helpful: a watermark-removal…

4

Oct 6, 2026

README

MCP Image Generator (Uncensored)

A self-hosted Model Context Protocol server that creates and edits images with the uncensored Qwen-Image-2.1 model on your own computer. It runs in Docker, on an NVIDIA GPU or on the CPU, and any MCP client on your network can use it over Streamable HTTP.

The default model is an uncensored build of Qwen-Image-2.1 with no built-in content filter. To use the standard model instead, set model.variant: base in config.yaml.

  • Text to image in any size or aspect ratio, up to about 4 megapixels (2048x2048)
  • Image editing: change, add or remove things, restyle, or combine up to 10 images
  • Transparent backgrounds (RGBA PNG) and background removal
  • Seamless tileable textures, including height and normal maps that stay seamless
  • 360 panoramas with a built-in 360 viewer
  • Upscaling 2x or 4x, and watermark removal
  • Models download automatically on the first start

Examples

All images below were made by this server with the default quality settings. The prompts are listed under the gallery.

A lighthouse on a rocky coast at sunset
Text to image
The same scene edited to a snowy winter night
Edit: the image on the left, made into a snowy winter night
A vintage travel poster that reads SEE THE ALPS
Text in images: text in quotes is written as given
Iron Man standing on a rooftop at night
Uncensored: famous characters, which many image services refuse
A cartoon rocket icon on a transparent background
Icon with no background (transparent PNG)
A red fox on a transparent background
Transparent background
The lighthouse photo covered with a tiled SAMPLE watermark and a corner badge
Watermark removal: before (a test watermark added to the first image)
The same photo with the watermark removed
Watermark removal: after remove_watermark

An equirectangular 360 panorama of a mountain meadow
360 panorama, opened in the built-in 360 viewer from the link in the result

Seamless textures

Each texture below is one tile (top). Repeated 3x3 (bottom), it shows no seams.

A red brick wall tileAn oak wood planks tileA blue and white ceramic tile
The brick tile repeated three by threeThe wood tile repeated three by threeThe ceramic tile repeated three by three

A stone tile, its height map and its normal map, each repeated two by two
Texture maps: a stone tile, then a height map and a normal map made from it with an edit. Each is shown repeated 2x2: they all stay seamless.

A detail enlarged four times: plain resize on the left, AI upscaler on the right
Upscaling 4x. Left: plain enlargement. Right: the upscaler.

Prompts and settings
ImageTool and settingsPrompt
Text to imagegenerate_image, 1344x768, seed 20261006A lighthouse on a rocky coast at sunset, waves breaking on the rocks, warm golden light, a small fishing boat in the distance, dramatic clouds, photorealistic
Editedit_image with the image above, seed 7Make it a snowy winter night with the lighthouse beam switched on and snow on the rocks, keep everything else unchanged
Text in imagesgenerate_image, 768x1024, seed 31A vintage travel poster, flat screen-print illustration of snowy mountains above a lake with a red train on a bridge, bold title text at the top that reads "SEE THE ALPS", smaller text at the bottom that reads "By rail, every season"
Uncensoredgenerate_image, 768x1024, seed 505Iron Man in his red and gold armor standing on a rooftop at night, city lights below, cinematic lighting, photorealistic
Icongenerate_image, 1024x1024, transparent: true, seed 404A cute cartoon rocket ship, flat vector illustration, bold outlines, bright colors
Transparentgenerate_image, 1024x1024, transparent: true, seed 11A red fox sitting, full body, soft studio light
360 panoramagenerate_panorama, 2048x1024, seed 2026A mountain meadow with wildflowers, a stone cabin, a wooden boardwalk and a lake
Brickgenerate_image, 512x512, tileable: true, seed 111Large red bricks with light grey mortar, close-up, photorealistic, even lighting
Woodgenerate_image, 512x512, tileable: true, seed 222Wide oak floor planks with visible wood grain and knots, top-down, natural light
Ceramicgenerate_image, 512x512, tileable: true, seed 303Blue and white ceramic tiles with a floral pattern
Stone tilegenerate_image, 512x512, tileable: true, seed 512A moss-covered stone floor, top-down
Height mapedit_image with the stone tile, seed 1Convert <image1> into a grayscale height map for a game material: white = high stone tops, black = low grout and gaps, keep the exact layout of every stone
Normal mapedit_image with the stone tile, seed 2Convert <image1> into a normal map for a game material, keep the exact layout of every stone
Watermark removalremove_watermark, seed 9, on the text-to-image result with a tiled "SAMPLE" text and a corner badge added- (no prompt needed)
Upscalingupscale_image, scale: 4 on the text-to-image result-

Characters and brands shown belong to their owners.

Requirements

GPU modeCPU mode
DockerDocker Desktop (Windows, macOS) or Docker Engine with Compose (Linux)same
HardwareNVIDIA GPU, RTX 20xx or newer, 12 GB VRAM or more (see GPU memory), 32 GB RAM recommendedAny modern 64-bit CPU, 24 GB RAM or more
DriverNVIDIA driver 570 or newer-
DiskAbout 12 GB for models and 6 GB for the Docker imageAbout 12 GB and 1 GB

GPU setup

Check that Docker can see your GPU:

docker run --rm --gpus all nvidia/cuda:12.8.2-base-ubuntu24.04 nvidia-smi

GPU memory

Every feature, including 2048x2048 images, edits with 10 input images, 2880x1440 panoramas and upscaling to 8192 px, works on a single 12 GB card:

GPU memorySettingsNotes
24 GB or moredefaultEverything stays on the GPU, the fastest setup. Peak use is about 18 GB
12-16 GBthe settings belowPeak use is about 11 GB. The text encoder's weights stay in system RAM, so the server uses up to about 15 GB of RAM
8-10 GBoffload: cpu and smaller sizesNot tested. All weights stream from RAM: slower, and the largest sizes may not fit

Settings for a 12 GB card, in config.yaml:

gpu:
  max_vram_gb: {0: 10}
  offload: {text_encoder: cpu}
generation:
  prefix_cache_type: q8_0

Quick start

git clone https://github.com/hypersniper05/MCP-Image-Generator-Uncensored.git
cd MCP-Image-Generator-Uncensored
./start.sh            # Linux / macOS
.\start.cmd           # Windows (or double-click start.cmd)

The first start builds the Docker image (10-30 minutes) and downloads about 12 GB of models. When it is ready, the script prints the address:

Ready. MCP endpoint (Streamable HTTP, no auth):
    http://localhost:5005/mcp
  • Stop: ./stop.sh or stop.cmd. Logs: docker compose logs -f.
  • Web page: open http://localhost:5005/ to see the status, upload images and browse recent results.
  • Settings: the first start creates config.yaml from config.example.yaml. To run on the CPU, set device: cpu in config.yaml and run the start script again.
  • Updates: after pulling new code, run ./start.sh --build (or start.cmd -Build) to rebuild the image.
Start without the scripts
cp config.example.yaml config.yaml
cp .env.example .env                  # for CPU mode, set COMPOSE_PROFILES=cpu in .env
docker compose up -d --build

The server has no authentication and listens on all network interfaces, so other computers on your network can reach it at http://<this-computer's-ip>:5005/mcp. Only run it on a network you trust.

Connect an MCP client

The endpoint is http://<host>:5005/mcp (Streamable HTTP).

MCP Inspector (quick test in a browser): run npx @modelcontextprotocol/inspector, choose Streamable HTTP, enter http://localhost:5005/mcp and click Connect.

VS Code (.vscode/mcp.json):

{ "servers": { "imagegen": { "type": "http", "url": "http://localhost:5005/mcp" } } }

Cursor (~/.cursor/mcp.json) and most other clients:

{ "mcpServers": { "imagegen": { "url": "http://localhost:5005/mcp" } } }

Clients that only support stdio can connect through mcp-remote:

{ "mcpServers": { "imagegen": { "command": "npx", "args": ["-y", "mcp-remote", "http://localhost:5005/mcp", "--allow-http"] } } }

Large images take minutes. If an image is not ready within 50 seconds, the tool returns a job_id and the client gets the image later with get_job, so clients with short timeouts still work.

llama.cpp web UI and optional request headers

For the llama.cpp web UI (llama-server started with --ui-mcp-proxy), add the server in the web UI's MCP settings, or for every browser through the file passed with --ui-config-file:

{
  "mcpServers": "[{\"id\": \"imagegen\", \"name\": \"Image Gen\", \"url\": \"http://127.0.0.1:5005/mcp\", \"enabled\": true, \"useProxy\": true, \"headers\": \"{\\\"X-Imagegen-Max-Wait\\\": \\\"20\\\", \\\"X-Inline-Max-Bytes\\\": \\\"16000000\\\"}\"}]"
}

Optional headers a client can send:

HeaderEffect
X-Imagegen-Max-Wait: 20Wait at most this many seconds before returning a job_id
X-Inline-Max-Bytes: 16000000Send the full image inline instead of a preview (for clients without a message size limit)
X-Imagegen-Inline: data-uri-textReturn images as data-URI text, for clients that drop MCP image content
X-Forwarded-Host: 192.168.1.50:5005Host name to use in returned links when the client connects through a proxy

Tools

ToolWhat it does
generate_imageText to image. Options: size (small, medium, large, xl) or width and height, aspect_ratio, transparent, tileable, seed, steps, negative_prompt
edit_imageEdit one image or combine up to 10 (refer to them as <image1>, <image2>, ...). Optional mask. Edits of a seamless tile stay seamless and keep its size
generate_panorama360 panorama (2:1). Optional image to turn a photo into a full 360
remove_backgroundCut out the subject into a transparent PNG
upscale_imageEnlarge 2x or 4x (up to 8192 px per side). Panoramas and tiles stay seamless
remove_watermarkRemove watermarks, logos and overlaid text; the rest of the image is kept as it was
get_job, cancel_jobGet the result of, or cancel, a job that was still running
list_images, view_imageRecent results and uploads, and a way for the model to look at one
server_statusModel, GPU placement, download progress and running jobs

Input images can be a data URL or base64, an http(s) URL, a file name or link of an image this server made, or an image uploaded on the server's web page (http://<host>:5005/upload).

Results include the image (or a preview of a large one) and a link to the full file, which is saved in ./outputs. Panoramas also get a link to the 360 viewer.

Tips

  • For the best quality use size: "xl". Leave steps and cfg_scale unset.
  • Write full sentences: subject, setting, lighting, style. Put text that should appear in quotes: a neon sign that says "OPEN 24/7".
  • For edits, say what to change and add "keep everything else unchanged".
  • For tiles, describe a surface or pattern that fills the whole picture, e.g. "moss-covered cobblestones, top-down".

Configuration

All settings are in config.yaml (created from the commented config.example.yaml). Restart after a change: docker compose restart. The most useful ones:

SettingDefaultWhat it does
devicegpugpu or cpu
model.quantQ4_K_MModel size: Q4_K_M (4.6 GB), Q6_K (5.9 GB) or Q8_0 (7.6 GB)
generation.steps40Quality vs. speed. 25 is a faster draft
generation.default_size1024x1024Size when a request gives none
generation.panorama_size2048x1024Default panorama size (2880x1440 at most)
generation.wait_seconds50How long a tool waits before returning a job_id
generation.idle_unload_seconds0Free the GPU memory after this many idle seconds (0 = keep loaded)
outputs.keep_days7Delete old results after this many days (0 = never)
server.port5005Port of the server
server.public_urlemptyAddress used in returned links, e.g. http://192.168.1.50:5005

.env (created from .env.example) holds Docker settings such as PUBLIC_URL, HF_TOKEN (only needed if Hugging Face rate-limits your downloads) and the build options.

config.yaml and .env are your local files and are not part of the repository.

Choosing GPUs

GPU numbers are the ones nvidia-smi shows. Everything runs on GPU 0 by default. The model has three parts, and each can go on a different GPU:

gpu:
  diffusion: [0]     # the main model, runs every step: use the fastest GPU
  text_encoder: 1    # reads the prompt once per request
  vae: 0             # turns the result into pixels

On a shared computer, only the GPUs you list are used. For cards with less memory, see GPU memory.

Troubleshooting

  • No GPU found: run the nvidia-smi check from Requirements. On Linux, install the NVIDIA Container Toolkit. Or use device: cpu.
  • Out of GPU memory: use the 12 GB settings, a smaller size, or move the text encoder to another GPU.
  • The client times out: lower generation.wait_seconds below the client's timeout.
  • Other computers cannot connect: use this computer's IP address instead of localhost, allow port 5005 in the firewall, and set server.public_url so image links work.
  • CPU mode is very slow or crashes on Windows: Docker Desktop gives WSL only half your RAM. Raise it in %UserProfile%\.wslconfig ([wsl2] then memory=24GB), run wsl --shutdown and restart Docker Desktop.
  • The server shows an error: read docker compose logs --tail 100, fix the cause, and run the start script again.

Development

python -m venv .venv && . .venv/bin/activate
pip install -e ".[test]"
pytest

python -m imagegen_mcp --config config.yaml --check validates a config file without starting anything. python scripts/smoke_test.py http://localhost:5005/mcp tests a running server.

Images are made by stable-diffusion.cpp, built inside the Docker image from a pinned release with one small patch (docker/patches/sdcpp-circular-json.patch) that turns on its wrap-around mode per request for tileable images.

Credits and licenses

The code in this repository is MIT licensed (see LICENSE). It builds on the work of others; the model files are downloaded from their original sources and keep their own licenses:

PartByLicense
Qwen-Image-2.1 image modelQwen team, AlibabaQwen Research License (non-commercial)
Uncensored Qwen-Image-2.1 GGUF (the default model files)abenzerpsQwen Research License (non-commercial)
Texture-fix VAE (image decoder)madebyollinQwen Research License (non-commercial)
Qwen3-VL-8B-Instruct GGUF (prompt and image encoder)Qwen team, AlibabaApache-2.0
Watermark removal LoRA (v1.0 for Qwen 2.1)saladinCivitai license: no commercial use
BiRefNet background removal, via rembgPeng Zheng et al.; Daniel GatisMIT (the optional isnet-general-use model: Apache-2.0)
4xNomos2_otf_esrgan upscaler (models)Philip HofmannCC BY 4.0
Real-ESRGAN (RealESRGAN_x4plus, optional upscaler)Xintao WangBSD-3-Clause
stable-diffusion.cpp and ggml (the inference engine)leejet and contributorsMIT
Pannellum (the 360 viewer)Matthew PetroffMIT

The Qwen-Image-2.1 files allow non-commercial use only: read their license before using images commercially. Only remove watermarks from images you have the rights to edit. The :gpu Docker image is based on NVIDIA's CUDA image (license).

The uncensored model has no built-in content filter. You are responsible for how you use it and what you make with it.