Local image and video generation on your own machine. No account with a generation service, no per-image credits, no prompts leaving your computer.
Vison is a desktop app around a C++/Vulkan backend built on stable-diffusion.cpp. You pick a model, it downloads once, and everything after that runs on your GPU.
Status: early. One developer, one machine, Windows only so far. It works — the models below were run on a 6 GB laptop GPU to produce real output — but it has not been through many hands yet. Expect rough edges, and please report them.
Vison reads your GPU's actual VRAM and marks models you cannot run, rather than letting you download 20 GB and then fail.
Grab the installer from Releases. It is currently unsigned, so Windows SmartScreen will warn you — "More info" then "Run anyway", or build from source below if you would rather not take that on trust.
Full detail in docs/BUILD.md. The short version:
# The vendored dependencies are not in this repo - they are upstream clones
# carrying local patches. This fetches them at the pinned commits and applies
# the patches.
sh scripts/bootstrap-third-party.sh
cmake -S . -B build -A x64 -DCMAKE_BUILD_TYPE=Release -DVISON_VULKAN=ON
cmake --build build --config Release --target vison_server --parallel
# Optional, but without it your installer ships with no video muxer, and
# most people who install it won't have ffmpeg on PATH either - video
# generation will silently fall back to numbered PNG frames. This fetches
# the same licence-clean (LGPL) ffmpeg build the CI release pipeline bundles.
# See docs/BUILD.md#ffmpeg.
sh scripts/fetch-ffmpeg.sh
cd app && npm ci && npm run build
You need the Vulkan SDK, not just the runtime — glslc compiles ggml's
compute shaders and only ships in the SDK.
Everything is pulled from Hugging Face on demand. The registry lives in
server/src/server.cpp.
Size is every file the model needs, not just the transformer — a model is useless without its text encoders and VAE, so that is what the app downloads and what it reports. Several models share files (the Wan models share one umt5 encoder; Qwen-Image and HunyuanVideo share one Qwen2.5-VL; Z-Image and FLUX share one VAE), so the second one you download costs less than its row suggests.
| Task | Model | Size | Advisory VRAM |
|---|---|---|---|
| Image | SDXL Turbo | 3.8 GB | 4 GB |
| Image | Z-Image Turbo | 7.5 GB | 4 GB |
| Image | FLUX.1 Schnell | 11.4 GB | 8 GB |
| Image | Qwen-Image | 16.8 GB | 8 GB |
| Video | Wan 2.1 T2V 1.3B | 6.7 GB | 6 GB |
| Video | Wan 2.2 TI2V 5B | 8.7 GB | 6 GB |
| Video | Wan 2.2 T2V A14B | 22.1 GB | 16 GB |
| Video | HunyuanVideo 1.5 | 22.6 GB | 16 GB |
| Video | MiniMax H3 (video + audio) | 33.0 GB | 24 GB |
| Upscale | ESRGAN 4x Remacri (photographic) | 32 MB | low |
| Upscale | Real-ESRGAN x4 Anime 6B | 9 MB | low |
The VRAM column is advisory only. What the app actually does is measure your
RAM and your card and tell you, per model, whether it fits, is tight, or will
not finish — see check_compatibility in core/src/device.cpp.
Two upscalers, and which one you want: Remacri is the default and the right choice for anything meant to look real — it keeps skin, hair, foliage and grain. The Anime 6B model is trained to flatten exactly that texture in favour of clean line work, which is what you want on drawn or animated output and not what you want on a photoreal render.
Being straight about coverage: SDXL Turbo, Z-Image Turbo, FLUX.1 Schnell, the two smaller Wan video models and the Anime 6B upscaler have all been run here and produce output. Qwen-Image, Wan 2.2 A14B, HunyuanVideo 1.5 and MiniMax H3 are registered and wired up but have never been run — they do not fit on the development machine. The Remacri upscaler has not been run here either; it is the same ESRGAN architecture and loader as the Anime 6B model that has, and its weights come from the author of the vision.cpp we use to run them, but that is an argument rather than a test. If you have the hardware, running any of these is the single most useful thing you could report back.
Everything runs locally. The backend listens on 127.0.0.1, the weights sit on
your disk, and generated images and video are written to your own folders. No
prompt, image, or video is uploaded anywhere.
Two things do reach the network, both only when you ask: downloading model weights from Hugging Face, and Google sign-in.
Sign-in is optional. Vison is fully usable without an account — generation, upscaling and history all work signed out, and nothing is withheld behind it. It exists for people who want it; there is no paid tier for it to unlock. A build made without OAuth credentials, which is what you get from a fresh clone, simply has no Sign in item in its account menu.
Electron + React -> IPC -> main process -> HTTP 127.0.0.1:11439
|
vison_server.exe (C++)
|
stable-diffusion.cpp + ggml (Vulkan)
Chat history is SQLite in your user data directory, with an FTS5 index so you can search conversations by what you asked for rather than by title.
Windows only, for now, and honestly rather than by principle: the packaging
config copies vison_server.exe and *.dll, which matches nothing a macOS or
Linux build produces. The backend itself is portable C++ and Vulkan — most of
the work is packaging and testing, not porting. Help welcome.
See CONTRIBUTING.md. The one thing to know before you start:
third_party/ is not vendored into this repo, and one of the local patches is
load-bearing — without it FLUX is rejected as an unknown architecture. Run
scripts/bootstrap-third-party.sh first.
Bug reports are as valuable as patches, especially on hardware other than a 6 GB NVIDIA laptop GPU. There are issue templates for both.
Found something exploitable? SECURITY.md — it also lists the two things that look like vulnerabilities here and are not.
Vison is free and MIT-licensed, and it will stay that way. There is no paid tier, nothing held back, and no plan to add either.
If it is useful to you and you would like to chip in, there is a Sponsor button at the top of this repository, and the same links live in the app under the account menu → Support Vison. Entirely optional — a bug report on hardware I do not own is worth just as much, and the app says so too.
What money goes to, concretely: hardware Vison has never been tested on (three of the models above have never been run at all), and a code signing certificate, so the installer stops tripping SmartScreen.
The links shown in the app come from app/src/support-links.ts; the ones on
the repository page come from .github/FUNDING.yml. Keep the two in step.
MIT.
Vison builds on other people's work, all of it permissively licensed:
stable-diffusion.cpp and
ggml (MIT),
vision.cpp (MIT), and others listed in
app/build-resources/licenses/THIRD-PARTY-NOTICES.txt, which is generated from
the licences actually present in the tree and shipped inside the app.
The bundled ffmpeg.exe is an LGPL build with libvpx and no GPL
components; the build refuses to package a GPL or non-free one. Video is
encoded as VP9 in WebM, which is royalty-free — that is a deliberate licensing
choice, not a technical one.
Model weights are not covered by this licence. Each carries its own terms from whoever published it; FLUX.1 Schnell, the Wan models and the rest are all different. Check them before using output commercially.
C++
50.6%
TypeScript
34.8%
Python
6.8%
JavaScript
6.3%
Local image and video generation on your own machine. No account with a generation service, no per-image credits, no prompts leaving your computer.
Vison is a desktop app around a C++/Vulkan backend built on stable-diffusion.cpp. You pick a model, it downloads once, and everything after that runs on your GPU.
Status: early. One developer, one machine, Windows only so far. It works — the models below were run on a 6 GB laptop GPU to produce real output — but it has not been through many hands yet. Expect rough edges, and please report them.
Vison reads your GPU's actual VRAM and marks models you cannot run, rather than letting you download 20 GB and then fail.
Grab the installer from Releases. It is currently unsigned, so Windows SmartScreen will warn you — "More info" then "Run anyway", or build from source below if you would rather not take that on trust.
Full detail in docs/BUILD.md. The short version:
# The vendored dependencies are not in this repo - they are upstream clones
# carrying local patches. This fetches them at the pinned commits and applies
# the patches.
sh scripts/bootstrap-third-party.sh
cmake -S . -B build -A x64 -DCMAKE_BUILD_TYPE=Release -DVISON_VULKAN=ON
cmake --build build --config Release --target vison_server --parallel
# Optional, but without it your installer ships with no video muxer, and
# most people who install it won't have ffmpeg on PATH either - video
# generation will silently fall back to numbered PNG frames. This fetches
# the same licence-clean (LGPL) ffmpeg build the CI release pipeline bundles.
# See docs/BUILD.md#ffmpeg.
sh scripts/fetch-ffmpeg.sh
cd app && npm ci && npm run build
You need the Vulkan SDK, not just the runtime — glslc compiles ggml's
compute shaders and only ships in the SDK.
Everything is pulled from Hugging Face on demand. The registry lives in
server/src/server.cpp.
Size is every file the model needs, not just the transformer — a model is useless without its text encoders and VAE, so that is what the app downloads and what it reports. Several models share files (the Wan models share one umt5 encoder; Qwen-Image and HunyuanVideo share one Qwen2.5-VL; Z-Image and FLUX share one VAE), so the second one you download costs less than its row suggests.
| Task | Model | Size | Advisory VRAM |
|---|---|---|---|
| Image | SDXL Turbo | 3.8 GB | 4 GB |
| Image | Z-Image Turbo | 7.5 GB | 4 GB |
| Image | FLUX.1 Schnell | 11.4 GB | 8 GB |
| Image | Qwen-Image | 16.8 GB | 8 GB |
| Video | Wan 2.1 T2V 1.3B | 6.7 GB | 6 GB |
| Video | Wan 2.2 TI2V 5B | 8.7 GB | 6 GB |
| Video | Wan 2.2 T2V A14B | 22.1 GB | 16 GB |
| Video | HunyuanVideo 1.5 | 22.6 GB | 16 GB |
| Video | MiniMax H3 (video + audio) | 33.0 GB | 24 GB |
| Upscale | ESRGAN 4x Remacri (photographic) | 32 MB | low |
| Upscale | Real-ESRGAN x4 Anime 6B | 9 MB | low |
The VRAM column is advisory only. What the app actually does is measure your
RAM and your card and tell you, per model, whether it fits, is tight, or will
not finish — see check_compatibility in core/src/device.cpp.
Two upscalers, and which one you want: Remacri is the default and the right choice for anything meant to look real — it keeps skin, hair, foliage and grain. The Anime 6B model is trained to flatten exactly that texture in favour of clean line work, which is what you want on drawn or animated output and not what you want on a photoreal render.
Being straight about coverage: SDXL Turbo, Z-Image Turbo, FLUX.1 Schnell, the two smaller Wan video models and the Anime 6B upscaler have all been run here and produce output. Qwen-Image, Wan 2.2 A14B, HunyuanVideo 1.5 and MiniMax H3 are registered and wired up but have never been run — they do not fit on the development machine. The Remacri upscaler has not been run here either; it is the same ESRGAN architecture and loader as the Anime 6B model that has, and its weights come from the author of the vision.cpp we use to run them, but that is an argument rather than a test. If you have the hardware, running any of these is the single most useful thing you could report back.
Everything runs locally. The backend listens on 127.0.0.1, the weights sit on
your disk, and generated images and video are written to your own folders. No
prompt, image, or video is uploaded anywhere.
Two things do reach the network, both only when you ask: downloading model weights from Hugging Face, and Google sign-in.
Sign-in is optional. Vison is fully usable without an account — generation, upscaling and history all work signed out, and nothing is withheld behind it. It exists for people who want it; there is no paid tier for it to unlock. A build made without OAuth credentials, which is what you get from a fresh clone, simply has no Sign in item in its account menu.
Electron + React -> IPC -> main process -> HTTP 127.0.0.1:11439
|
vison_server.exe (C++)
|
stable-diffusion.cpp + ggml (Vulkan)
Chat history is SQLite in your user data directory, with an FTS5 index so you can search conversations by what you asked for rather than by title.
Windows only, for now, and honestly rather than by principle: the packaging
config copies vison_server.exe and *.dll, which matches nothing a macOS or
Linux build produces. The backend itself is portable C++ and Vulkan — most of
the work is packaging and testing, not porting. Help welcome.
See CONTRIBUTING.md. The one thing to know before you start:
third_party/ is not vendored into this repo, and one of the local patches is
load-bearing — without it FLUX is rejected as an unknown architecture. Run
scripts/bootstrap-third-party.sh first.
Bug reports are as valuable as patches, especially on hardware other than a 6 GB NVIDIA laptop GPU. There are issue templates for both.
Found something exploitable? SECURITY.md — it also lists the two things that look like vulnerabilities here and are not.
Vison is free and MIT-licensed, and it will stay that way. There is no paid tier, nothing held back, and no plan to add either.
If it is useful to you and you would like to chip in, there is a Sponsor button at the top of this repository, and the same links live in the app under the account menu → Support Vison. Entirely optional — a bug report on hardware I do not own is worth just as much, and the app says so too.
What money goes to, concretely: hardware Vison has never been tested on (three of the models above have never been run at all), and a code signing certificate, so the installer stops tripping SmartScreen.
The links shown in the app come from app/src/support-links.ts; the ones on
the repository page come from .github/FUNDING.yml. Keep the two in step.
MIT.
Vison builds on other people's work, all of it permissively licensed:
stable-diffusion.cpp and
ggml (MIT),
vision.cpp (MIT), and others listed in
app/build-resources/licenses/THIRD-PARTY-NOTICES.txt, which is generated from
the licences actually present in the tree and shipped inside the app.
The bundled ffmpeg.exe is an LGPL build with libvpx and no GPL
components; the build refuses to package a GPL or non-free one. Video is
encoded as VP9 in WebM, which is royalty-free — that is a deliberate licensing
choice, not a technical one.
Model weights are not covered by this licence. Each carries its own terms from whoever published it; FLUX.1 Schnell, the Wan models and the rest are all different. Check them before using output commercially.
C++
50.6%
TypeScript
34.8%
Python
6.8%
JavaScript
6.3%