A friendly GUI for running local LLMs with vLLM and llama.cpp — hardware-aware traffic-light guidance for every setting. Built on ideas from vllm-cli (Chen-zexi).
Python
41
66 commits
updated Oct 6, 2026
This app was developed off excellent work by Chen-zexi on vllm-cli: many thanks for the CLI version, which the author recommends for use when VRAM is tight (booting into a TTY without a display server can free over 1 GB of GPU memory). The idea behind this project was the frustration that comes with attempting to squeeze LLMs of various types onto individual hardware. The flags are cryptic, loading fails frequently, and documentation is not always helpful. This is an effort to help the local LLM community more easily run models — measuring available memory, making suggestions, and adjusting settings so a model can run without hours of trial and error. Or at least, fewer hours of trial and error.
Read the full story (and a shoutout to the AI that built this) in docs/ABOUT.md.
A friendly, browser-based control panel for downloading and running large language models ("LLMs" — the AI models behind chatbots like ChatGPT, except running on your own computer) using vLLM or llama.cpp. Built for people who don't want to memorize cryptic command-line flags or spend an afternoon guessing why a model won't load.
Every setting is rated 🟢 / 🟡 / 🔴 (green / yellow / red) against your actual computer and the model you picked, with a one-or-two-sentence plain-English explanation. A live "fuel gauge" shows whether the model will fit before you click launch. If a launch fails anyway, the app translates the error into something you can actually act on.

Experimental: SGLang support. An
experimental/sglang-integrationbranch adds SGLang as a fourth engine. It is unstable — upstream bugs inapache-tvm-fficause processor loading failures and JIT crashes, especially on NVIDIA Blackwell GPUs (RTX 5060 Ti). Testing only; vLLM and llama.cpp remain the reliable options.
A computer that can run local AI models. In practice this means:
Python 3.10 or newer. Most Macs and Linux systems already have this. On Windows, install it from python.org (tick "Add Python to PATH" during install).
At least one "engine" — the program that actually runs the model:
.gguf model files. Recommended for
most people starting out.pip install vllm or use their
Docker image.Don't worry too much about choosing — the app detects what you have installed and tells you what's missing, with instructions, on the Settings page.
Launch → Advanced settings includes linear, MoE (mixture-of-experts), and attention backends (alternative implementations of the model's calculations). Leave them automatic unless comparing a particular implementation. vLLM 0.30 already improves automatic NVFP4 selection (4-bit model computation) on compatible RTX 50-series models; selecting a backend does not convert a model to that format.
The optional b12x implementations target SM120/121 (GPU capability versions,
including the RTX 5060 Ti). Install the vllm[b12x] extra in the same Python
environment as the configured vLLM executable, preserving your chosen vLLM
version. Docker users need the dependency inside their image. The launcher
does not install it automatically.
For vLLM 0.30 the attention selector is B12X, despite earlier release
discussion using B12X_ATTN. It requires BF16 (16-bit brain floating point)
model computation and compatible conversation memory, and cannot split context
processing across GPUs. A model stored in float32 qualifies when precision is
left on automatic, because vLLM runs it as BF16 on these GPUs. B12X MoE also
has model-format restrictions and cannot use expert parallelism (splitting
experts among separate workers).
Native checks inspect the selected executable, not the launcher's Python.
Unavailable checks are shown as unverified; Docker runtime/package support is
always unverified here. The first check of a vLLM installation can take up to
a minute while vLLM loads its libraries: the Launch page shows "still being
checked" in the meantime, and Launch waits for the answer. Results are
reused for 10 minutes and rechecked as soon as you install or upgrade packages
in that environment. Options typed in Extra arguments are checked the way
vLLM reads them, including underscore spellings such as --linear_backend.
Package presence and recognized flags do not guarantee
a model will run or be faster. Compare performance on your own GPU before
keeping an override. See the versioned B12X documentation.
Open a terminal (on Windows: search for "Command Prompt" or "PowerShell"; on Mac: search for "Terminal"; on Linux: you know where it is) and run:
# 1. Get the code
git clone https://github.com/jimdawdy-hub/Local-LLM-Launcher-GUI
cd Local-LLM-Launcher-GUI
# 2. (Recommended) create an isolated Python environment
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 3. Install the app
pip install .
# 4. Run it
local-llm-launcher
Your web browser should open automatically to http://127.0.0.1:8765. If it
doesn't, open that address yourself. This only runs on your own computer —
nothing is sent anywhere else, and no other device on your network can reach
it. (Model servers and Open WebUI follow the same loopback-only rule unless
you turn on the LAN access toggle in Settings, which binds them to
0.0.0.0 for other devices on your network.) If port 8765 is already in use,
the app automatically picks the next free port and tells you in the terminal.
To stop the app, go back to the terminal window and press Ctrl+C.
Reaching it from your other devices through Tailscale (optional). The app
still listens only on this computer; Tailscale's serve feature forwards your
private tailnet to it. Start the app with the address Tailscale gives you:
local-llm-launcher --no-browser --allow-host mybox.tailnet-name.ts.net
tailscale serve --bg --https=8765 http://127.0.0.1:8765
Then open https://mybox.tailnet-name.ts.net:8765 on any device in your
tailnet. --allow-host accepts exact names only (no wildcards) and can be
repeated. Engine source updates stay available only from the computer itself.
.bin files) is downloaded and
counted.


llama-server if it wasn't found
automatically.Which GPUs to use takes the GPU numbers shown by nvidia-smi, separated by
commas: 0,1 for both of two cards, 1 for just the second. Entries such as
0 1 (a space instead of a comma) are refused with an explanation rather than
silently misread. For native vLLM the launcher sets
CUDA_DEVICE_ORDER=PCI_BUS_ID, so the numbers always match nvidia-smi. That
includes a CUDA_VISIBLE_DEVICES you set before starting the launcher, which
CUDA would otherwise number fastest card first. The memory advice budgets the
GPUs selected here, or those in CUDA_VISIBLE_DEVICES when this is empty. If
you add vLLM's own --device-ids under Extra arguments, it picks positions
within the GPUs selected here.
For llama.cpp, the advanced launch settings include:
--tensor-split, for example 3,1 for unequal cards or 40,40,40 for equal
shares on three cards. These are proportions, not exact layer counts.
--split-mode layer distributes layers and their conversation memory (KV
cache) across the cards. The older row mode keeps KV on the main GPU;
experimental tensor mode needs Flash Attention and remains model/backend
dependent. Recent llama.cpp builds support tensor with compressed q8_0
or q4_0 K/V cache. Older builds may need updating; switching to layer or
f16 cache is another option. NVFP4 describes model weights, not the K/V
cache: an NVFP4 GGUF can request Q8 conversation memory independently.
Safetensors-format NVFP4 models still require an engine that loads that
format, such as vLLM. See upstream tensor/quantized-KV support.--cpu-moe keeps all experts (the selectively used
parts of a mixture-of-experts model) in RAM. --n-cpu-moe N keeps experts
from the first N layers in RAM; reducing N lets more live on the GPU.
Transfers between RAM and GPU can reduce speed. These controls cannot promise
a precise memory fit without knowing the model's individual tensor sizes.--ubatch-size limits how many tokens are processed
together within a batch. Smaller values may reduce working memory at a speed
cost.--numa distribute spreads llama.cpp threads across nodes; isolate uses the
startup node; numactl respects CPU placement arranged externally. The
separate Interleave memory option wraps native llama.cpp or vLLM with
numactl --interleave=all. It spreads memory allocations across allowed
nodes and does not itself bind CPU threads. It requires Linux and numactl
and does not apply to Docker launches.NUMA options are off by default. They support ordinary multi-CPU computers as well as VMs; recommendations use the nodes visible and allowed to the app. A VM exposing one node cannot use these options to distribute work across hidden host nodes. There is no automatic change to host configuration or system page caches. Benchmark your workload before retaining a NUMA setting.
The existing memory-mapping and RAM-lock controls are translated to the newer
--load-mode syntax when the selected llama.cpp binary supports it. Older
binaries keep their legacy flags. Other advanced flags still depend on your
installed engine version; failed launches retain their logs for diagnosis.
The video linked in issue #15 also uses a modified llama.cpp fork. Its expert prefetch/pinning optimizations are not stock flags, and this launcher does not promise that video's speedup.
In Settings → Build current engine source, check requirements, inspect the target commit (an exact source revision), then explicitly choose to build it. Nothing is downloaded or upgraded automatically. Supported recipes are Linux llama.cpp CPU/CUDA and Linux native vLLM with NVIDIA CUDA prerequisites. Other platforms and Docker installations keep their manual installation path.
Builds download official upstream source and dependencies into separate folders under the app data directory. Current source may contain unreleased changes. Prerequisites and free disk space are checked before starting; compiler jobs are limited to reduce CPU and memory pressure. Progress, recent output, and failures appear in Settings. Only a validated executable is selected for future launches. Running servers keep their existing executable, and older copies are retained. You can restore an older executable path in Settings.
Use a browser directly on the launcher computer for these update controls. A failed build keeps the previous engine selected, removes its downloaded source and build environment to free disk space, and keeps its log. An interrupted build (for example, if the launcher is closed) also leaves the previous engine selected, but its partial files may remain in the app data directory. The launcher must stay open to finish a build and select its result.
The app adds up how much memory (VRAM, on a GPU) a model needs — its weights, plus a working memory area for the conversation — and compares it to how much your hardware actually has free right now:
These numbers update live, because how much memory is "free" changes as you open and close other programs.
Behind the scenes, vLLM checks memory in two separate steps when starting up: first it loads the model's weights, then it reserves room for the conversation (the "KV cache"). A model can pass the first check and still fail the second — loading successfully and then refusing to start. The app checks both steps before you launch, so a "should load but then fail at the last second" scenario shows up as a yellow warning with a fix (shorter context, more memory headroom, or compression) instead of a 10-15 minute wait followed by a cryptic error.
MIT — see LICENSE. Portions adapted from vllm-cli, © Chen-zexi, MIT license.
A friendly GUI for running local LLMs with vLLM and llama.cpp — hardware-aware traffic-light guidance for every setting. Built on ideas from vllm-cli (Chen-zexi).
Python
41
66 commits
updated Oct 6, 2026
This app was developed off excellent work by Chen-zexi on vllm-cli: many thanks for the CLI version, which the author recommends for use when VRAM is tight (booting into a TTY without a display server can free over 1 GB of GPU memory). The idea behind this project was the frustration that comes with attempting to squeeze LLMs of various types onto individual hardware. The flags are cryptic, loading fails frequently, and documentation is not always helpful. This is an effort to help the local LLM community more easily run models — measuring available memory, making suggestions, and adjusting settings so a model can run without hours of trial and error. Or at least, fewer hours of trial and error.
Read the full story (and a shoutout to the AI that built this) in docs/ABOUT.md.
A friendly, browser-based control panel for downloading and running large language models ("LLMs" — the AI models behind chatbots like ChatGPT, except running on your own computer) using vLLM or llama.cpp. Built for people who don't want to memorize cryptic command-line flags or spend an afternoon guessing why a model won't load.
Every setting is rated 🟢 / 🟡 / 🔴 (green / yellow / red) against your actual computer and the model you picked, with a one-or-two-sentence plain-English explanation. A live "fuel gauge" shows whether the model will fit before you click launch. If a launch fails anyway, the app translates the error into something you can actually act on.

Experimental: SGLang support. An
experimental/sglang-integrationbranch adds SGLang as a fourth engine. It is unstable — upstream bugs inapache-tvm-fficause processor loading failures and JIT crashes, especially on NVIDIA Blackwell GPUs (RTX 5060 Ti). Testing only; vLLM and llama.cpp remain the reliable options.
A computer that can run local AI models. In practice this means:
Python 3.10 or newer. Most Macs and Linux systems already have this. On Windows, install it from python.org (tick "Add Python to PATH" during install).
At least one "engine" — the program that actually runs the model:
.gguf model files. Recommended for
most people starting out.pip install vllm or use their
Docker image.Don't worry too much about choosing — the app detects what you have installed and tells you what's missing, with instructions, on the Settings page.
Launch → Advanced settings includes linear, MoE (mixture-of-experts), and attention backends (alternative implementations of the model's calculations). Leave them automatic unless comparing a particular implementation. vLLM 0.30 already improves automatic NVFP4 selection (4-bit model computation) on compatible RTX 50-series models; selecting a backend does not convert a model to that format.
The optional b12x implementations target SM120/121 (GPU capability versions,
including the RTX 5060 Ti). Install the vllm[b12x] extra in the same Python
environment as the configured vLLM executable, preserving your chosen vLLM
version. Docker users need the dependency inside their image. The launcher
does not install it automatically.
For vLLM 0.30 the attention selector is B12X, despite earlier release
discussion using B12X_ATTN. It requires BF16 (16-bit brain floating point)
model computation and compatible conversation memory, and cannot split context
processing across GPUs. A model stored in float32 qualifies when precision is
left on automatic, because vLLM runs it as BF16 on these GPUs. B12X MoE also
has model-format restrictions and cannot use expert parallelism (splitting
experts among separate workers).
Native checks inspect the selected executable, not the launcher's Python.
Unavailable checks are shown as unverified; Docker runtime/package support is
always unverified here. The first check of a vLLM installation can take up to
a minute while vLLM loads its libraries: the Launch page shows "still being
checked" in the meantime, and Launch waits for the answer. Results are
reused for 10 minutes and rechecked as soon as you install or upgrade packages
in that environment. Options typed in Extra arguments are checked the way
vLLM reads them, including underscore spellings such as --linear_backend.
Package presence and recognized flags do not guarantee
a model will run or be faster. Compare performance on your own GPU before
keeping an override. See the versioned B12X documentation.
Open a terminal (on Windows: search for "Command Prompt" or "PowerShell"; on Mac: search for "Terminal"; on Linux: you know where it is) and run:
# 1. Get the code
git clone https://github.com/jimdawdy-hub/Local-LLM-Launcher-GUI
cd Local-LLM-Launcher-GUI
# 2. (Recommended) create an isolated Python environment
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
# 3. Install the app
pip install .
# 4. Run it
local-llm-launcher
Your web browser should open automatically to http://127.0.0.1:8765. If it
doesn't, open that address yourself. This only runs on your own computer —
nothing is sent anywhere else, and no other device on your network can reach
it. (Model servers and Open WebUI follow the same loopback-only rule unless
you turn on the LAN access toggle in Settings, which binds them to
0.0.0.0 for other devices on your network.) If port 8765 is already in use,
the app automatically picks the next free port and tells you in the terminal.
To stop the app, go back to the terminal window and press Ctrl+C.
Reaching it from your other devices through Tailscale (optional). The app
still listens only on this computer; Tailscale's serve feature forwards your
private tailnet to it. Start the app with the address Tailscale gives you:
local-llm-launcher --no-browser --allow-host mybox.tailnet-name.ts.net
tailscale serve --bg --https=8765 http://127.0.0.1:8765
Then open https://mybox.tailnet-name.ts.net:8765 on any device in your
tailnet. --allow-host accepts exact names only (no wildcards) and can be
repeated. Engine source updates stay available only from the computer itself.
.bin files) is downloaded and
counted.


llama-server if it wasn't found
automatically.Which GPUs to use takes the GPU numbers shown by nvidia-smi, separated by
commas: 0,1 for both of two cards, 1 for just the second. Entries such as
0 1 (a space instead of a comma) are refused with an explanation rather than
silently misread. For native vLLM the launcher sets
CUDA_DEVICE_ORDER=PCI_BUS_ID, so the numbers always match nvidia-smi. That
includes a CUDA_VISIBLE_DEVICES you set before starting the launcher, which
CUDA would otherwise number fastest card first. The memory advice budgets the
GPUs selected here, or those in CUDA_VISIBLE_DEVICES when this is empty. If
you add vLLM's own --device-ids under Extra arguments, it picks positions
within the GPUs selected here.
For llama.cpp, the advanced launch settings include:
--tensor-split, for example 3,1 for unequal cards or 40,40,40 for equal
shares on three cards. These are proportions, not exact layer counts.
--split-mode layer distributes layers and their conversation memory (KV
cache) across the cards. The older row mode keeps KV on the main GPU;
experimental tensor mode needs Flash Attention and remains model/backend
dependent. Recent llama.cpp builds support tensor with compressed q8_0
or q4_0 K/V cache. Older builds may need updating; switching to layer or
f16 cache is another option. NVFP4 describes model weights, not the K/V
cache: an NVFP4 GGUF can request Q8 conversation memory independently.
Safetensors-format NVFP4 models still require an engine that loads that
format, such as vLLM. See upstream tensor/quantized-KV support.--cpu-moe keeps all experts (the selectively used
parts of a mixture-of-experts model) in RAM. --n-cpu-moe N keeps experts
from the first N layers in RAM; reducing N lets more live on the GPU.
Transfers between RAM and GPU can reduce speed. These controls cannot promise
a precise memory fit without knowing the model's individual tensor sizes.--ubatch-size limits how many tokens are processed
together within a batch. Smaller values may reduce working memory at a speed
cost.--numa distribute spreads llama.cpp threads across nodes; isolate uses the
startup node; numactl respects CPU placement arranged externally. The
separate Interleave memory option wraps native llama.cpp or vLLM with
numactl --interleave=all. It spreads memory allocations across allowed
nodes and does not itself bind CPU threads. It requires Linux and numactl
and does not apply to Docker launches.NUMA options are off by default. They support ordinary multi-CPU computers as well as VMs; recommendations use the nodes visible and allowed to the app. A VM exposing one node cannot use these options to distribute work across hidden host nodes. There is no automatic change to host configuration or system page caches. Benchmark your workload before retaining a NUMA setting.
The existing memory-mapping and RAM-lock controls are translated to the newer
--load-mode syntax when the selected llama.cpp binary supports it. Older
binaries keep their legacy flags. Other advanced flags still depend on your
installed engine version; failed launches retain their logs for diagnosis.
The video linked in issue #15 also uses a modified llama.cpp fork. Its expert prefetch/pinning optimizations are not stock flags, and this launcher does not promise that video's speedup.
In Settings → Build current engine source, check requirements, inspect the target commit (an exact source revision), then explicitly choose to build it. Nothing is downloaded or upgraded automatically. Supported recipes are Linux llama.cpp CPU/CUDA and Linux native vLLM with NVIDIA CUDA prerequisites. Other platforms and Docker installations keep their manual installation path.
Builds download official upstream source and dependencies into separate folders under the app data directory. Current source may contain unreleased changes. Prerequisites and free disk space are checked before starting; compiler jobs are limited to reduce CPU and memory pressure. Progress, recent output, and failures appear in Settings. Only a validated executable is selected for future launches. Running servers keep their existing executable, and older copies are retained. You can restore an older executable path in Settings.
Use a browser directly on the launcher computer for these update controls. A failed build keeps the previous engine selected, removes its downloaded source and build environment to free disk space, and keeps its log. An interrupted build (for example, if the launcher is closed) also leaves the previous engine selected, but its partial files may remain in the app data directory. The launcher must stay open to finish a build and select its result.
The app adds up how much memory (VRAM, on a GPU) a model needs — its weights, plus a working memory area for the conversation — and compares it to how much your hardware actually has free right now:
These numbers update live, because how much memory is "free" changes as you open and close other programs.
Behind the scenes, vLLM checks memory in two separate steps when starting up: first it loads the model's weights, then it reserves room for the conversation (the "KV cache"). A model can pass the first check and still fail the second — loading successfully and then refusing to start. The app checks both steps before you launch, so a "should load but then fail at the last second" scenario shows up as a yellow warning with a fix (shorter context, more memory headroom, or compression) instead of a 10-15 minute wait followed by a cryptic error.
MIT — see LICENSE. Portions adapted from vllm-cli, © Chen-zexi, MIT license.