newideas99/ultra-fast-image-gen

2-8B parameter image gen that actually runs fast on your Mac. 14 seconds. No cloud. No GPU rental.

297

stars

53

commits

Python

primary language

Jun 11, 2026

updated

README

Ultra Fast Image Gen

AI image generation and editing on Mac Silicon and CUDA. Generate images from text or transform existing images with state-of-the-art diffusion models.

Ultra Fast Image Gen web UI

Features

  • Image Generation: Create images from text prompts
  • Image Editing: Upload up to 6 reference images and transform them with natural language
  • Multiple Models: FLUX.2-klein and Z-Image Turbo
  • Quantized Models: Low memory usage with 4bit/int8 quantization
  • Anima Turbo AIO: Local patched Metal runner with Turbo LoRA baked into the GGUF
  • Uncensored FLUX.2 2K lanes: MFLUX/MLX fast path and PyTorch SDNQ fallback with MPS optimizations
  • Bonsai Image 4B: Ternary FLUX.2 Klein, in-process MLX on Apple Silicon
  • LoRA Support: Load custom LoRA adapters with Z-Image Full model
  • Cross-Platform: Apple Silicon (MPS) and NVIDIA GPUs (CUDA)

Supported Models

ModelVRAMFeaturesSpeed
FLUX.2-klein-4B Uncensored MFLUX HS~7GB RSS @ 2KText-to-image, uncensored GGUF TE, validated to 2KFastest current 2K FLUX lane
FLUX.2-klein-4B Uncensored SDNQ HS~8GB MPS / low RAM @ 2KText-to-image, uncensored GGUF TE, exact chunked MPS attention + HSPyTorch 2K fallback
FLUX.2-klein-4B (4bit SDNQ)<8GB @ 512px, <16GB @ 1024pxText-to-image + Image editingFast
FLUX.2-klein-9B (4bit SDNQ)~12GB @ 512px, ~20GB @ 1024pxText-to-image + Image editing (Higher Quality)Fast
FLUX.2-klein-4B (Int8)~16GBText-to-image + Image editingFast
Z-Image Turbo (Quantized)~8GBText-to-imageFastest
Anima Turbo AIO Q4 (Metal)~3GB model + unified memoryText-to-image, baked Turbo LoRA~16s internal @ 512x768 / 8 steps
Bonsai Image 4B (Ternary MLX)~3.7GB, Apple Silicon onlyText-to-image, 4 steps~15s @ 512x512 / 4 steps
Z-Image Turbo (Full)~24GBText-to-image + LoRASlower

The uncensored variants do not re-download a separate base model: the SDNQ HS lane and the plain uncensored model reuse the FLUX.2-klein-4B (4bit SDNQ) backbone, so the only extra download is the uncensored Qwen3 text encoder (~2.5GB GGUF at the default q4_k_m quant).

Note: the uncensored text encoder repo is gated on Hugging Face (instant auto-approval). Accept the terms on the model page once, then paste a token in the app's ⋯ menu (it's validated and stored in .env for you — or create .env yourself with HF_TOKEN=hf_...).

Quick Start (1-Click)

  1. Download/clone the repo
  2. Double-click Launch.command (if macOS says it's from an unidentified developer, right-click the file → Open → Open)
  3. First run will auto-install dependencies (~5 min); later runs reinstall if requirements.txt changed
  4. The launcher installs/builds the patched Anima Metal runner and the MFLUX 2K runtime if needed — if either fails, the app still starts and the rest of the models work
  5. Browser opens automatically to the UI
  6. Models download from inside the app: each one in the model picker has a Download button with live progress (or they fetch automatically on first generate)
  7. For the Uncensored models, paste your Hugging Face token in the ⋯ menu (gated text encoder — accept the terms on the model page once)

Bonsai Image 4B is not part of this default install — it's an opt-in extra (uv sync --extra bonsai, Apple Silicon + python 3.11+). See Bonsai Image 4B (optional).

Manual Installation

git clone https://github.com/newideas99/ultra-fast-image-gen.git
cd ultra-fast-image-gen

python3.11 -m venv venv
source venv/bin/activate

pip install -r requirements.txt

# For the uncensored models (gated text-encoder repo):
echo "HF_TOKEN=hf_your_token_here" > .env

Anima Fresh Install

Launch.command runs this automatically when the Anima runner is missing:

scripts/setup_anima_metal_runner.sh

The setup script clones stable-diffusion.cpp, checks out the tested revision, applies the bundled ggml Metal patch for Anima VAE ops, builds sd-cli with Metal enabled, and writes ~/anima-comfyui/run_anima_aio_metal.sh.

The downloaded Anima-P3-Turbo-AIO-Q4_K.gguf already has the Anima Turbo LoRA merged in. The app does not need a separate anima-turbo-lora-v0.1.safetensors file for this model.

MFLUX HS 2K Lane Fresh Install

Launch.command runs this automatically when the MFLUX runtime is missing:

scripts/setup_mflux_hs.sh

The setup script clones mflux at the tested 0.17.5 revision into ~/.cache/ultra-fast-image-gen/mflux, applies the bundled hidden-state-compression + uncensored-GGUF-text-encoder patch (patches/mflux_hs_uncensored_gguf.patch), and pre-builds the uv environment. Only the "FLUX.2-klein-4B Uncensored MFLUX HS" model needs this; the SDNQ HS 2K lane runs on the normal Python dependencies. Override the install location with ULTRA_FAST_MFLUX_HS_DIR.

Bonsai Image 4B (optional)

Bonsai (ternary) is not installed by default — it's an opt-in extra. It runs in-process on MLX and is Apple Silicon + python 3.11+ only. Enable it with uv (the extra carries a mlx override that plain pip won't honour, so uv is required here):

uv sync --extra bonsai

That pulls prism-image-studio + mflux-prism and stock mlx (prebuilt wheel — no Xcode/Metal build needed). Weights (~3.7 GB) auto-download from Hugging Face on first use.

Note: only the ternary arm is supported. The 1-bit/binary arm needed a source-built mlx fork whose Metal kernels don't compile on current macOS, so it's dropped — ternary uses standard 2-bit affine quant on stock mlx instead.

Usage

Web UI

python server.py

Then open http://localhost:7860 in your browser.

The UI is a dependency-free HTML/CSS/JS frontend served by a small FastAPI backend. It shows real per-step progress, supports batch generation (up to 8 images per run), resolution presets, drag-and-drop image editing, a session gallery, storage management, and remembers your settings between visits.

Each model in the picker has its own Download / Delete button with live progress, so you can pre-fetch weights before generating; deleting a model keeps any base files another downloaded model still shares. Generating with a model you haven't downloaded also just works — the download progress shows up in the status card first. Switching models unloads the previous one from memory before the new one loads.

Model Selection

  • FLUX.2-klein-4B Uncensored MFLUX HS: Default. Fastest 2K text-to-image lane (~100s @ 2048x2048). Uses the patched MFLUX runtime (scripts/setup_mflux_hs.sh) plus the uncensored Qwen GGUF text encoder
  • FLUX.2-klein-4B Uncensored SDNQ HS: PyTorch SDNQ 2K text-to-image lane with exact MPS query chunking and hidden-state compression; no extra setup
  • FLUX.2-klein-4B (4bit SDNQ): Lowest memory, supports image editing
  • FLUX.2-klein-9B (4bit SDNQ): Higher quality 9B model, more memory
  • FLUX.2-klein-4B (Int8): Alternative quantization, more memory
  • Z-Image Turbo (Quantized): Fastest text-to-image, no image editing
  • Anima Turbo AIO Q4 (Metal): Uses ~/anima-comfyui/run_anima_aio_metal.sh; auto-downloads the Turbo AIO GGUF if missing and defaults to 512x768, 8 steps, CFG 1
  • Bonsai Image 4B (Ternary MLX): In-process MLX on Apple Silicon, 4-step distilled. Weights auto-download on first use (opt-in extra)
  • Z-Image Turbo (Full): Use when you need LoRA support

Image Editing (FLUX.2-klein)

  1. Select a classic FLUX.2-klein model from the dropdown (4bit SDNQ, 9B, or Int8 — the 2K uncensored lanes are text-to-image in the UI)
  2. Upload up to 6 images in the gallery
  3. Write a prompt describing the changes you want
  4. Select output resolution (1024px, 1280px, or 1536px)
  5. Click Generate

Command Line

Each model has its own sub-command with the options it needs:

# Z-Image Turbo (quantized) — fastest, ~3.5 GB
python generate.py zimage-quant a beautiful sunset over mountains

# Z-Image Turbo (full precision) — with optional LoRA
python generate.py zimage-full a beautiful sunset --lora my.safetensors --lora-strength 0.8

# FLUX.2-klein-4B (4bit SDNQ)
python generate.py flux2-4b-sdnq a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-4B (Int8)
python generate.py flux2-4b-int8 a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-9B (4bit SDNQ) — higher quality
python generate.py flux2-9b-sdnq a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-4B Uncensored MFLUX/MLX HS — fastest current 2K path
python generate.py flux2-4b-uncensored-mflux-hs a beautiful sunset --width 2048 --height 2048 --steps 4

# FLUX.2-klein-4B Uncensored PyTorch SDNQ HS — optimized PyTorch/MPS 2K path
python generate.py flux2-4b-uncensored-sdnq-hs a beautiful sunset --width 2048 --height 2048 --steps 4

# Image-to-image editing (FLUX.2-klein models only)
python generate.py flux2-4b-sdnq transform the fox into a wolf --input-images ref.png

# Anima Turbo AIO (Metal runner, baked Turbo LoRA)
python generate.py anima anime portrait, detailed eyes --anima-preset Balanced

# Bonsai Image 4B ternary (MLX, Apple Silicon) — needs the bonsai extra
python generate.py bonsai-ternary a red fox in snow --steps 4

Quotes around the prompt are optional — all words before the first --flag are joined into the prompt.

Common options (all sub-commands):

OptionDefaultDescription
--height512 (Anima: 768)Image height in pixels
--width512Image width in pixels
--seedrandomFixed seed for reproducibility
--outputoutput.pngOutput file path
--devicempsmps, cuda, or cpu

Z-Image options:

OptionDefaultDescription
--steps5Inference steps
--loraPath to LoRA .safetensors (zimage-full only)
--lora-strength1.0LoRA weight (zimage-full only)

FLUX.2-klein options:

OptionDefaultDescription
--steps28Inference steps
--guidance3.5Classifier-free guidance scale
--input-imagesUp to 6 reference images for editing

Uncensored 2K speed-lane options:

OptionDefaultDescription
--steps4Distilled klein inference steps
--guidance0.0Guidance scale
--gguf-quantq4_k_mUncensored Qwen GGUF text encoder quant
--qchunk1024PyTorch MPS attention query chunk (sdnq-hs only)
--hs-stride2Hidden-state compression stride
--hs-max-transformer-forwardsteps - 1Leave the final transformer forward exact
--mflux-dir~/.cache/ultra-fast-image-gen/mfluxPatched MFLUX checkout (mflux-hs only; see scripts/setup_mflux_hs.sh)

Anima options:

OptionDefaultDescription
--stepsfrom presetInference steps (overrides the preset)
--cfg-scale1.0Anima CFG scale
--anima-presetBalancedFast (3 steps), Balanced (8), or Quality (16)

Bonsai options (Apple Silicon only; --device is accepted but ignored — always MLX):

OptionDefaultDescription
--steps4Inference steps

MCP Server (AI coding agents)

The repo doubles as a local MCP server, so AI coding assistants (OpenCode, Claude Desktop, Cursor, Claude Code) can generate and edit project assets — hero banners, icons, backgrounds — without cloud APIs. It exposes two tools, generate_image and edit_image (1-6 reference images), which run the same generate.py CLI under the hood.

Install the mcp package into the app's environment first:

venv/bin/pip install mcp        # or: uv sync --extra mcp

OpenCode (1-click)

python3 scripts/install-opencode-mcp.py

The script registers the server in ~/.config/opencode/opencode.json and installs a "website-visual-assets" skill. It reuses HF_TOKEN from your environment or the repo .env (set via the web UI's ⋯ menu), prompting only if neither exists.

Manual configuration

OpenCode (~/.config/opencode/opencode.json):

{
  "mcp": {
    "ultra-fast-image-gen": {
      "type": "local",
      "command": ["/path/to/ultra-fast-image-gen/venv/bin/python", "/path/to/ultra-fast-image-gen/mcp_server.py"],
      "environment": { "HF_TOKEN": "hf_..." },
      "timeout": 3600000
    }
  }
}

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "ultra-fast-image-gen": {
      "command": "/path/to/ultra-fast-image-gen/venv/bin/python",
      "args": ["/path/to/ultra-fast-image-gen/mcp_server.py"],
      "env": { "HF_TOKEN": "hf_..." }
    }
  }
}

Agent instructions

Drop something like this into your project's CLAUDE.md, .opencode/agents.md, or .cursorrules so the agent reaches for the tools:

When asked to create or modify visual assets (e.g. "generate a hero banner",
"make the logo background dark"), use the `generate_image` / `edit_image`
tools from the `ultra-fast-image-gen` MCP server. The default model
"zimage-quant" is ultra-fast and lowest-memory; use "flux2-4b-sdnq" or
"flux2-9b-sdnq" only when higher quality is explicitly requested. Save
outputs to logical project paths (e.g. public/images/hero.png).

No HF_TOKEN is needed for the standard models (Z-Image, FLUX.2-klein, Anima) — weights download to the local Hugging Face cache on first use. The token only matters for the gated uncensored text encoder.

Heads-up: an MCP generation and a web-UI generation each load their own model copy, so running both at once doubles memory use.

Benchmarks

FLUX.2-klein-4B

HardwareResolutionStepsTime
Apple Silicon512x5124~8s
CUDA (RTX 3090)512x5124~3s

FLUX.2-klein-4B Uncensored 2K Speed Lanes

BackendResolutionStepsTimeNotes
MFLUX/MLX HS2048x20484100.2s fresh-process wall / ~69s denoiseIncludes uncensored GGUF TE load in fresh CLI process
PyTorch SDNQ HS2048x20484~110s generation wallExact query-chunked MPS attention + HS compression

Z-Image Turbo (Quantized)

MacResolutionStepsTime
M2 Max512x512714s
M2 Max768x768731s
M1 Max512x512723s

Anima Turbo AIO Q4 (Metal)

Recommended settings:

  • Fast: 3 steps with Spectrum cache
  • Balanced/default: 8 steps with Spectrum cache
  • Quality/model-card default: 16 steps with cache disabled
MacResolutionStepsTime
M2 Max512x76838.82s internal / 11.63s wall
M2 Max512x768411.13s internal / 13.65s wall
M2 Max512x768815.62s internal / 18.69s wall

Memory Requirements

ModelRAM/VRAM Required
FLUX.2-klein-4B Uncensored MFLUX HS~7GB RSS @ 2048px
FLUX.2-klein-4B Uncensored SDNQ HSLow MPS memory @ 2048px; slower PyTorch fallback
FLUX.2-klein-4B (4bit SDNQ)8GB @ 512px, 16GB @ 1024px
FLUX.2-klein-9B (4bit SDNQ)12GB @ 512px, 20GB @ 1024px
FLUX.2-klein-4B (Int8)16GB
Z-Image (Quantized)8GB
Z-Image (Full)24GB+
Anima Turbo AIO Q4 (Metal)32GB recommended for local setup
Bonsai Image 4B (MLX)~4GB unified memory, Apple Silicon

Credits

License

See the original model licenses for usage terms.

Contributors

newideas99

39 commits

lczyk

8 commits

John-K

2 commits

newideas99/ultra-fast-image-gen

2-8B parameter image gen that actually runs fast on your Mac. 14 seconds. No cloud. No GPU rental.

297

stars

53

commits

Python

primary language

Jun 11, 2026

updated

README

Ultra Fast Image Gen

AI image generation and editing on Mac Silicon and CUDA. Generate images from text or transform existing images with state-of-the-art diffusion models.

Ultra Fast Image Gen web UI

Features

  • Image Generation: Create images from text prompts
  • Image Editing: Upload up to 6 reference images and transform them with natural language
  • Multiple Models: FLUX.2-klein and Z-Image Turbo
  • Quantized Models: Low memory usage with 4bit/int8 quantization
  • Anima Turbo AIO: Local patched Metal runner with Turbo LoRA baked into the GGUF
  • Uncensored FLUX.2 2K lanes: MFLUX/MLX fast path and PyTorch SDNQ fallback with MPS optimizations
  • Bonsai Image 4B: Ternary FLUX.2 Klein, in-process MLX on Apple Silicon
  • LoRA Support: Load custom LoRA adapters with Z-Image Full model
  • Cross-Platform: Apple Silicon (MPS) and NVIDIA GPUs (CUDA)

Supported Models

ModelVRAMFeaturesSpeed
FLUX.2-klein-4B Uncensored MFLUX HS~7GB RSS @ 2KText-to-image, uncensored GGUF TE, validated to 2KFastest current 2K FLUX lane
FLUX.2-klein-4B Uncensored SDNQ HS~8GB MPS / low RAM @ 2KText-to-image, uncensored GGUF TE, exact chunked MPS attention + HSPyTorch 2K fallback
FLUX.2-klein-4B (4bit SDNQ)<8GB @ 512px, <16GB @ 1024pxText-to-image + Image editingFast
FLUX.2-klein-9B (4bit SDNQ)~12GB @ 512px, ~20GB @ 1024pxText-to-image + Image editing (Higher Quality)Fast
FLUX.2-klein-4B (Int8)~16GBText-to-image + Image editingFast
Z-Image Turbo (Quantized)~8GBText-to-imageFastest
Anima Turbo AIO Q4 (Metal)~3GB model + unified memoryText-to-image, baked Turbo LoRA~16s internal @ 512x768 / 8 steps
Bonsai Image 4B (Ternary MLX)~3.7GB, Apple Silicon onlyText-to-image, 4 steps~15s @ 512x512 / 4 steps
Z-Image Turbo (Full)~24GBText-to-image + LoRASlower

The uncensored variants do not re-download a separate base model: the SDNQ HS lane and the plain uncensored model reuse the FLUX.2-klein-4B (4bit SDNQ) backbone, so the only extra download is the uncensored Qwen3 text encoder (~2.5GB GGUF at the default q4_k_m quant).

Note: the uncensored text encoder repo is gated on Hugging Face (instant auto-approval). Accept the terms on the model page once, then paste a token in the app's ⋯ menu (it's validated and stored in .env for you — or create .env yourself with HF_TOKEN=hf_...).

Quick Start (1-Click)

  1. Download/clone the repo
  2. Double-click Launch.command (if macOS says it's from an unidentified developer, right-click the file → Open → Open)
  3. First run will auto-install dependencies (~5 min); later runs reinstall if requirements.txt changed
  4. The launcher installs/builds the patched Anima Metal runner and the MFLUX 2K runtime if needed — if either fails, the app still starts and the rest of the models work
  5. Browser opens automatically to the UI
  6. Models download from inside the app: each one in the model picker has a Download button with live progress (or they fetch automatically on first generate)
  7. For the Uncensored models, paste your Hugging Face token in the ⋯ menu (gated text encoder — accept the terms on the model page once)

Bonsai Image 4B is not part of this default install — it's an opt-in extra (uv sync --extra bonsai, Apple Silicon + python 3.11+). See Bonsai Image 4B (optional).

Manual Installation

git clone https://github.com/newideas99/ultra-fast-image-gen.git
cd ultra-fast-image-gen

python3.11 -m venv venv
source venv/bin/activate

pip install -r requirements.txt

# For the uncensored models (gated text-encoder repo):
echo "HF_TOKEN=hf_your_token_here" > .env

Anima Fresh Install

Launch.command runs this automatically when the Anima runner is missing:

scripts/setup_anima_metal_runner.sh

The setup script clones stable-diffusion.cpp, checks out the tested revision, applies the bundled ggml Metal patch for Anima VAE ops, builds sd-cli with Metal enabled, and writes ~/anima-comfyui/run_anima_aio_metal.sh.

The downloaded Anima-P3-Turbo-AIO-Q4_K.gguf already has the Anima Turbo LoRA merged in. The app does not need a separate anima-turbo-lora-v0.1.safetensors file for this model.

MFLUX HS 2K Lane Fresh Install

Launch.command runs this automatically when the MFLUX runtime is missing:

scripts/setup_mflux_hs.sh

The setup script clones mflux at the tested 0.17.5 revision into ~/.cache/ultra-fast-image-gen/mflux, applies the bundled hidden-state-compression + uncensored-GGUF-text-encoder patch (patches/mflux_hs_uncensored_gguf.patch), and pre-builds the uv environment. Only the "FLUX.2-klein-4B Uncensored MFLUX HS" model needs this; the SDNQ HS 2K lane runs on the normal Python dependencies. Override the install location with ULTRA_FAST_MFLUX_HS_DIR.

Bonsai Image 4B (optional)

Bonsai (ternary) is not installed by default — it's an opt-in extra. It runs in-process on MLX and is Apple Silicon + python 3.11+ only. Enable it with uv (the extra carries a mlx override that plain pip won't honour, so uv is required here):

uv sync --extra bonsai

That pulls prism-image-studio + mflux-prism and stock mlx (prebuilt wheel — no Xcode/Metal build needed). Weights (~3.7 GB) auto-download from Hugging Face on first use.

Note: only the ternary arm is supported. The 1-bit/binary arm needed a source-built mlx fork whose Metal kernels don't compile on current macOS, so it's dropped — ternary uses standard 2-bit affine quant on stock mlx instead.

Usage

Web UI

python server.py

Then open http://localhost:7860 in your browser.

The UI is a dependency-free HTML/CSS/JS frontend served by a small FastAPI backend. It shows real per-step progress, supports batch generation (up to 8 images per run), resolution presets, drag-and-drop image editing, a session gallery, storage management, and remembers your settings between visits.

Each model in the picker has its own Download / Delete button with live progress, so you can pre-fetch weights before generating; deleting a model keeps any base files another downloaded model still shares. Generating with a model you haven't downloaded also just works — the download progress shows up in the status card first. Switching models unloads the previous one from memory before the new one loads.

Model Selection

  • FLUX.2-klein-4B Uncensored MFLUX HS: Default. Fastest 2K text-to-image lane (~100s @ 2048x2048). Uses the patched MFLUX runtime (scripts/setup_mflux_hs.sh) plus the uncensored Qwen GGUF text encoder
  • FLUX.2-klein-4B Uncensored SDNQ HS: PyTorch SDNQ 2K text-to-image lane with exact MPS query chunking and hidden-state compression; no extra setup
  • FLUX.2-klein-4B (4bit SDNQ): Lowest memory, supports image editing
  • FLUX.2-klein-9B (4bit SDNQ): Higher quality 9B model, more memory
  • FLUX.2-klein-4B (Int8): Alternative quantization, more memory
  • Z-Image Turbo (Quantized): Fastest text-to-image, no image editing
  • Anima Turbo AIO Q4 (Metal): Uses ~/anima-comfyui/run_anima_aio_metal.sh; auto-downloads the Turbo AIO GGUF if missing and defaults to 512x768, 8 steps, CFG 1
  • Bonsai Image 4B (Ternary MLX): In-process MLX on Apple Silicon, 4-step distilled. Weights auto-download on first use (opt-in extra)
  • Z-Image Turbo (Full): Use when you need LoRA support

Image Editing (FLUX.2-klein)

  1. Select a classic FLUX.2-klein model from the dropdown (4bit SDNQ, 9B, or Int8 — the 2K uncensored lanes are text-to-image in the UI)
  2. Upload up to 6 images in the gallery
  3. Write a prompt describing the changes you want
  4. Select output resolution (1024px, 1280px, or 1536px)
  5. Click Generate

Command Line

Each model has its own sub-command with the options it needs:

# Z-Image Turbo (quantized) — fastest, ~3.5 GB
python generate.py zimage-quant a beautiful sunset over mountains

# Z-Image Turbo (full precision) — with optional LoRA
python generate.py zimage-full a beautiful sunset --lora my.safetensors --lora-strength 0.8

# FLUX.2-klein-4B (4bit SDNQ)
python generate.py flux2-4b-sdnq a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-4B (Int8)
python generate.py flux2-4b-int8 a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-9B (4bit SDNQ) — higher quality
python generate.py flux2-9b-sdnq a beautiful sunset --guidance 3.5 --steps 28

# FLUX.2-klein-4B Uncensored MFLUX/MLX HS — fastest current 2K path
python generate.py flux2-4b-uncensored-mflux-hs a beautiful sunset --width 2048 --height 2048 --steps 4

# FLUX.2-klein-4B Uncensored PyTorch SDNQ HS — optimized PyTorch/MPS 2K path
python generate.py flux2-4b-uncensored-sdnq-hs a beautiful sunset --width 2048 --height 2048 --steps 4

# Image-to-image editing (FLUX.2-klein models only)
python generate.py flux2-4b-sdnq transform the fox into a wolf --input-images ref.png

# Anima Turbo AIO (Metal runner, baked Turbo LoRA)
python generate.py anima anime portrait, detailed eyes --anima-preset Balanced

# Bonsai Image 4B ternary (MLX, Apple Silicon) — needs the bonsai extra
python generate.py bonsai-ternary a red fox in snow --steps 4

Quotes around the prompt are optional — all words before the first --flag are joined into the prompt.

Common options (all sub-commands):

OptionDefaultDescription
--height512 (Anima: 768)Image height in pixels
--width512Image width in pixels
--seedrandomFixed seed for reproducibility
--outputoutput.pngOutput file path
--devicempsmps, cuda, or cpu

Z-Image options:

OptionDefaultDescription
--steps5Inference steps
--loraPath to LoRA .safetensors (zimage-full only)
--lora-strength1.0LoRA weight (zimage-full only)

FLUX.2-klein options:

OptionDefaultDescription
--steps28Inference steps
--guidance3.5Classifier-free guidance scale
--input-imagesUp to 6 reference images for editing

Uncensored 2K speed-lane options:

OptionDefaultDescription
--steps4Distilled klein inference steps
--guidance0.0Guidance scale
--gguf-quantq4_k_mUncensored Qwen GGUF text encoder quant
--qchunk1024PyTorch MPS attention query chunk (sdnq-hs only)
--hs-stride2Hidden-state compression stride
--hs-max-transformer-forwardsteps - 1Leave the final transformer forward exact
--mflux-dir~/.cache/ultra-fast-image-gen/mfluxPatched MFLUX checkout (mflux-hs only; see scripts/setup_mflux_hs.sh)

Anima options:

OptionDefaultDescription
--stepsfrom presetInference steps (overrides the preset)
--cfg-scale1.0Anima CFG scale
--anima-presetBalancedFast (3 steps), Balanced (8), or Quality (16)

Bonsai options (Apple Silicon only; --device is accepted but ignored — always MLX):

OptionDefaultDescription
--steps4Inference steps

MCP Server (AI coding agents)

The repo doubles as a local MCP server, so AI coding assistants (OpenCode, Claude Desktop, Cursor, Claude Code) can generate and edit project assets — hero banners, icons, backgrounds — without cloud APIs. It exposes two tools, generate_image and edit_image (1-6 reference images), which run the same generate.py CLI under the hood.

Install the mcp package into the app's environment first:

venv/bin/pip install mcp        # or: uv sync --extra mcp

OpenCode (1-click)

python3 scripts/install-opencode-mcp.py

The script registers the server in ~/.config/opencode/opencode.json and installs a "website-visual-assets" skill. It reuses HF_TOKEN from your environment or the repo .env (set via the web UI's ⋯ menu), prompting only if neither exists.

Manual configuration

OpenCode (~/.config/opencode/opencode.json):

{
  "mcp": {
    "ultra-fast-image-gen": {
      "type": "local",
      "command": ["/path/to/ultra-fast-image-gen/venv/bin/python", "/path/to/ultra-fast-image-gen/mcp_server.py"],
      "environment": { "HF_TOKEN": "hf_..." },
      "timeout": 3600000
    }
  }
}

Claude Desktop (~/Library/Application Support/Claude/claude_desktop_config.json):

{
  "mcpServers": {
    "ultra-fast-image-gen": {
      "command": "/path/to/ultra-fast-image-gen/venv/bin/python",
      "args": ["/path/to/ultra-fast-image-gen/mcp_server.py"],
      "env": { "HF_TOKEN": "hf_..." }
    }
  }
}

Agent instructions

Drop something like this into your project's CLAUDE.md, .opencode/agents.md, or .cursorrules so the agent reaches for the tools:

When asked to create or modify visual assets (e.g. "generate a hero banner",
"make the logo background dark"), use the `generate_image` / `edit_image`
tools from the `ultra-fast-image-gen` MCP server. The default model
"zimage-quant" is ultra-fast and lowest-memory; use "flux2-4b-sdnq" or
"flux2-9b-sdnq" only when higher quality is explicitly requested. Save
outputs to logical project paths (e.g. public/images/hero.png).

No HF_TOKEN is needed for the standard models (Z-Image, FLUX.2-klein, Anima) — weights download to the local Hugging Face cache on first use. The token only matters for the gated uncensored text encoder.

Heads-up: an MCP generation and a web-UI generation each load their own model copy, so running both at once doubles memory use.

Benchmarks

FLUX.2-klein-4B

HardwareResolutionStepsTime
Apple Silicon512x5124~8s
CUDA (RTX 3090)512x5124~3s

FLUX.2-klein-4B Uncensored 2K Speed Lanes

BackendResolutionStepsTimeNotes
MFLUX/MLX HS2048x20484100.2s fresh-process wall / ~69s denoiseIncludes uncensored GGUF TE load in fresh CLI process
PyTorch SDNQ HS2048x20484~110s generation wallExact query-chunked MPS attention + HS compression

Z-Image Turbo (Quantized)

MacResolutionStepsTime
M2 Max512x512714s
M2 Max768x768731s
M1 Max512x512723s

Anima Turbo AIO Q4 (Metal)

Recommended settings:

  • Fast: 3 steps with Spectrum cache
  • Balanced/default: 8 steps with Spectrum cache
  • Quality/model-card default: 16 steps with cache disabled
MacResolutionStepsTime
M2 Max512x76838.82s internal / 11.63s wall
M2 Max512x768411.13s internal / 13.65s wall
M2 Max512x768815.62s internal / 18.69s wall

Memory Requirements

ModelRAM/VRAM Required
FLUX.2-klein-4B Uncensored MFLUX HS~7GB RSS @ 2048px
FLUX.2-klein-4B Uncensored SDNQ HSLow MPS memory @ 2048px; slower PyTorch fallback
FLUX.2-klein-4B (4bit SDNQ)8GB @ 512px, 16GB @ 1024px
FLUX.2-klein-9B (4bit SDNQ)12GB @ 512px, 20GB @ 1024px
FLUX.2-klein-4B (Int8)16GB
Z-Image (Quantized)8GB
Z-Image (Full)24GB+
Anima Turbo AIO Q4 (Metal)32GB recommended for local setup
Bonsai Image 4B (MLX)~4GB unified memory, Apple Silicon

Credits

License

See the original model licenses for usage terms.

Contributors

newideas99

39 commits

lczyk

8 commits

John-K

2 commits

Languages

Python

72.3%

JavaScript

10.9%

Shell

7.5%

CSS

5.7%

HTML

3.6%