Single-file installer for a club-3090 webserver providing an admin control panel, a reverse proxy that automatically routes requests to the correct containers via a single endpoint with configurable vLLM setting presets, live log management, GPU power and fan speed controls, multi-instance GPU orchestration and API access control for multiple users
Python
0
241 commits
updated Sep 29, 2026
Single-file installer for a club-3090 inference host, providing a full-featured browser admin control panel, distro-aware install/update logic, systemd service setup, a reverse proxy that automatically routes requests to the correct containers via a single endpoint with optional configurable vLLM setting presets, live log management, GPU power and fan speed controls, multi-instance GPU orchestration, and API access control for multiple users.
This repository is the server-management layer. It integrates with the upstream club-3090 runtime and, on a fresh install clones that upstream repo automatically into /opt/ai/club-3090 before wiring up the control plane.
v0.10.1* release family. The admin panel shows an unsupported-checkout warning when the local upstream club-3090 checkout falls outside the release patterns baked into the installer. You may let --migrate pull the latest upstream checkout, but if upstream introduced breaking changes, use that warning as a cue to audit the affected presets or AI Studio lanes before treating the host as production-ready./opt/club3090-control and an admin web panel running on localhost:8008/admin:8009 so you can chat with the LLM on a unified port regardless of what docker containers are in use.club-3090 checkoutGLOBAL, GPU<n>, and PAIRx_y targets, including built-in sequential auto-pairs and custom user-defined pairspull.shActive after the backend fully finishes booting, with boot/error browser notifications and fast failure reporting when a compose never actually starts a container/admin routes so refreshes can show the last known panel state while the remote service is temporarily unreachable and reconnect automatically when service returns--online mode that opens and forwards only the admin/proxy ports to the internet--online use-https mode that enables a Caddy-backed self-signed HTTPS frontend for the admin panel and proxy--skip-temps installer switch for operators who do not want the extra temperature helper, dependency install, bootloader iomem=relaxed staging, or reboot noticeVideo coming soon. Click any thumbnail to open the full-size screenshot.
This section is written for someone who just wants the server working and may not be so familiar with the complexities of LLM configuration.
Make sure your Linux machine already has:
sudoIf you plan to download gated Hugging Face models, also have your Hugging Face token ready.
By default the installer also prepares optional GDDR6/GDDR6X junction and VRAM temperature telemetry for supported NVIDIA cards. It vendors the helper source and NVIDIA NVML header inside the installer, compiles the helper locally against the host NVIDIA/libpci libraries, stages the required iomem=relaxed kernel command-line flag when it is not already active, and tells you to reboot only when that reboot is needed for the new readings to work.
Run the installer directly:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash
If you already downloaded this repo locally, you can also run:
bash install-club3090-server.sh
If you need gated model downloads and already have a Hugging Face token, use:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
HF_TOKEN=hf_xxx bash
On a fresh machine, the installer will create /opt/ai/ if needed, clone the upstream club-3090 repo into /opt/ai/club-3090, fix script permissions, and then continue with the normal setup.
When the installer finishes, open a browser and go to:
http://YOUR-SERVER-IP:8008/admin
If you installed with custom ports, replace 8008 with your chosen admin port.
When the login screen appears, sign in with your normal Linux user credentials.
Use:
In other words, log in with the same account you would normally use for sudo or for signing into that Linux machine. The admin panel does not create a separate default password for you.
After logging in, open the Presets tab.
You will see discovered model presets from the local club-3090 checkout. Some presets may show that downloads are still required.
If a preset is not ready:
Download button.Audit Logs.If you prefer to prepare the model manually in the terminal first, a common example is:
cd /opt/ai/club-3090
bash scripts/setup.sh qwen3.6-27b
Some advanced presets, such as DFlash variants, may require extra model files. The admin panel will usually tell you what command is needed.
Still in the Presets tab:
Apply.Active, the card shows how long launch took, and the action button becomes StopIf something fails, the preset will show Error and the logs will show what happened.
For most beginners:
Once a preset is active, the proxy is usually available on:
http://YOUR-SERVER-IP:8009/v1
That proxy gives you one stable API endpoint even if you switch between different backends later.
You can test it with:
curl http://YOUR-SERVER-IP:8009/v1/models
Or with a chat request:
curl -s http://YOUR-SERVER-IP:8009/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.6-27b-autoround","messages":[{"role":"user","content":"Hello"}],"max_tokens":128}'
Here is the simple mental model for the main tabs:
Main: quick status, running containers, uptime, and GPU overviewSystem: service controls, power/fan controls, and machine informationPresets: download models, start/stop presets, and manage which runtime is activeChat: built-in local streaming chat client with server-backed conversation storage, lazy detail loading, per-conversation stats caching, share/export links, configuration, and local or remote MCP serversAI Studio: multimodal setup, model-resource checks, gallery browsing, and Plan/Interactive generation lanes from the Presets and Chat surfacesUsers: create API users and keys if you want to share the server safelyMetrics: request counts, usage history, runtime performance data, storage browsing, and detachable monitoringLogs: Docker logs and audit logs for troubleshootingAfter installation, most people will want to do these things:
Presets.Launch on a recommended preset.Active.http://YOUR-SERVER-IP:8009/v1 into their AI client.Most OpenAI-compatible apps need:
http://YOUR-SERVER-IP:8009/v1If the client asks for an endpoint and does not mention presets, start with the plain proxy URL above.
Try these checks in order:
Active, not just Booting.Docker Logs and Audit Logs for the exact failure message.If you changed the upstream repo manually or pulled a new upstream version, run update or migrate again so the control layer and upstream runtime stay in sync.
The script reads /etc/os-release and installs the right package names for the detected family.
systemddocker composeclub-3090 repo into /opt/ai/club-3090 on first installAs of v0.5, the server no longer ships a hardcoded preset catalog. Instead it scans the local upstream repo under /opt/ai/club-3090, parses compose headers and compose files directly, and builds /opt/club3090-control/runtime_inventory.json.
That means the Presets tab now reflects whatever models and variants exist in the checked-out upstream repo, including:
Upstream switch tags such as vllm/default, vllm/dual, vllm/gemma-mtp, and llamacpp/default are still recognized when present, but the control layer can also launch compose variants that are only discoverable by scanning the repo tree.
Since the broader v0.8.x, v0.9.x, and v0.10.x migrations, the inventory path is also registry/profile-aware and understands the newer nested upstream compose hierarchy, so the admin panel can surface richer upstream metadata such as:
autoround-int4, awq, unsloth-q4km, ubergarm-iq4ks, and ex0bit-prismThe older single-mode/legacy-dual distinction has gradually been replaced by a clearer scope-oriented control model.
The admin panel now thinks in terms of:
GLOBAL: fan a compatible preset out across every eligible target automaticallyGPU0, GPU1, GPU2, ...: explicit single-card targetsPAIR0_1, PAIR2_3, ...: dual-card targetsImportant behaviors added across the v0.7.x and v0.8.x series include:
If you use many presets often, the Summary view is also worth calling out. It keeps a relaunch-oriented cache of recent and active presets per model, including temporary boot entries and bulk stop/restart actions.
One of the biggest v0.8.x additions is the first-class Custom Model flow in the Presets tab.
Instead of manually treating every non-curated model as a one-off local hack, the panel now supports:
[+] Custom Model entry pointscripts/pull.shThis keeps the UI model-first while still deferring the actual import/evaluation logic to upstream Club-3090.
The admin panel now includes a first-class AI Studio surface for upstream Club-3090's unified ai-studio scene. It is meant to make image, video, audio, and speech workflows usable from the same browser panel as inference presets.
Current AI Studio support includes:
The control layer keeps AI Studio integration in its own wrapper/setup logic instead of patching files inside the upstream checkout. Legacy /admin/image-studio/... routes remain as backend compatibility aliases, but the browser and current API surface use /admin/ai-studio/....
Model Scores are the built-in preset validation and scoring system. They are designed to answer two different questions:
The Benchmarks modal and preset score badges are backed by saved artifacts under /opt/club3090-control/benchmarks. The current score reader only treats a Quick or Full result as valid when every required stage for that mode has a saved zero return code. Partial selected-stage shells, stale failures, and superseded sidecar JSON are ignored or archived so old artifacts do not masquerade as current scores.
Benchmark behavior includes:
The Metrics tab now uses lightweight live status payloads and bounded chart history so charts can refresh at foreground cadence even while benchmark history or model-resource scans are large. Peak markers are persisted separately from the visible rolling window, and each chart draws a single dashed peak guide for the recorded max.
Metrics also provides:
When the admin panel is served from a secure HTTPS /admin route, a Service Worker caches the admin shell and last-known status snapshot. A normal open panel labels itself disconnected only after continuous status failures exceed the configured timeout, while a refreshed page falls back to the cached shell if the remote service cannot be contacted quickly. Any successful live status response automatically exits disconnected mode.
Default install:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash
Custom upstream checkout path and default mode:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
CLUB3090_DIR=/opt/ai/club-3090 DEFAULT_MODE=vllm/dual-dflash bash
Pass a Hugging Face token inline when gated downloads are needed:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
HF_TOKEN=hf_xxx bash
Running a local copy:
bash install-club3090-server.sh
The installer reads CLUB3090_DIR from a .env file in the current working directory before falling back to /opt/ai/club-3090. Use an absolute path:
CLUB3090_DIR=/home/gchamon/services/club-3090
An exported CLUB3090_DIR takes precedence over this bootstrap file. Set CLUB3090_INSTALLER_ENV_FILE to read the bootstrap value from another file. After resolving the checkout, the installer separately reads ${CLUB3090_DIR}/.env for repository settings such as MODEL_DIR, Hugging Face cache paths, and HF_TOKEN.
Custom admin and proxy ports:
bash install-club3090-server.sh --ports 18008:18009
Enable the local automation API:
bash install-club3090-server.sh --local-automation
Persist a Hugging Face token into the local club-3090 repo .env so setup jobs and later admin-panel model downloads can reuse it automatically:
bash install-club3090-server.sh --hf-token hf_xxx
Public internet exposure:
bash install-club3090-server.sh --online
Public internet exposure with TLS layer:
bash install-club3090-server.sh --online use-https
Skip optional junction/VRAM temperature telemetry setup and keep the old install/update behavior:
bash install-club3090-server.sh --skip-temps
--skip-temps bypasses helper unpacking, compiler/libpci dependency installation, bootloader iomem=relaxed staging, and the related reboot notice.
Updates the control layer without intentionally stopping an already-running backend:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash -s -- --update
Or from a local copy:
bash install-club3090-server.sh --update
Public-mode update:
bash install-club3090-server.sh --update --online
Update with custom ports:
bash install-club3090-server.sh --update --ports 18008:18009
Use --migrate when the upstream repo checkout itself needs to be replaced with a fresh clone while preserving downloaded model assets and runtime state:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash -s -- --migrate
Or from a local copy:
bash install-club3090-server.sh --migrate
If an interrupted migration needs to be discarded and restarted from scratch instead of resumed:
bash install-club3090-server.sh --migrate restart
--migrate:
club-3090 repo into the original path.env, models-cache, AI Studio assets, generated results, LMCache KV data, safe compose cache directories, and the effective MODEL_DIRMODEL_DIR paths during migration and mirrors preserved weights into the effective post-merge model directory when it differs from the repo-local defaultsetup.sh reclones themOLD, OLD (2), and later generations when upstream presets changed againOLD presets while preserving genuinely different generations for comparisonOLD rows can be audited against their current counterpart/opt/club3090-control/migration-state.env--migrateHF_TOKEN=... or --hf-token ..., then stores that token in the repo .env so later setup and admin-panel download jobs can reuse itIf DEFAULT_MODE is not set, the installer counts GPUs with nvidia-smi and chooses:
vllm/default when 0 or 1 GPU is detectedvllm/dual when 2 or more GPUs are detectedFor testing or overrides, you can force the detected count:
CLUB3090_GPU_COUNT_OVERRIDE=1 bash install-club3090-server.sh
You can still hard-set a default mode:
DEFAULT_MODE=vllm/long-text bash install-club3090-server.sh
When the server is using single-card presets, the control service can manage one backend instance per GPU. Each instance has:
The instance catalog is stored in:
/opt/club3090-control/instances.json
On multi-GPU systems, this lets you run:
This is intentionally separate from the legacy dual-GPU presets. If no per-GPU single-card instances are enabled, the control service can still fall back to the old upstream dual-mode flow.
In newer releases this model also extends cleanly into the scope-based control plane:
v0.8.4x cycleWhen you install or update with --online, the script treats the control service as the only public surface.
That means:
8010-8020 and per-instance ports like 8200+ are intentionally left privateThis prevents security bypass around API-key checks, user limits, and per-instance access controls.
If you also pass use-https, the installer binds the Python control service to loopback-only internal ports and places Caddy in front of it as the public HTTPS endpoint for both admin and proxy traffic.
When HTTPS mode is enabled, the installer places Caddy in front of the admin UI and proxy on the public admin/proxy ports. Direct IP access uses the self-signed certificate path under /opt/club3090-control/tls.crt and /opt/club3090-control/tls.key by default. For a browser-trusted certificate, point a DNS or DDNS hostname at the server and set CLUB3090_HTTPS_HOST or CLUB3090_ONLINE_HOST; Caddy will then request and renew a normal Let's Encrypt certificate for that hostname. Direct IP Let's Encrypt certificates are only attempted with the explicit --online use-https:cert-ip mode, or when you explicitly set an IP host and CLUB3090_ENABLE_IP_LE=true.
If the server is already on Tailscale, --online use-https:tailscale detects the node's MagicDNS name, such as aardwolf-halfmoon.ts.net, and writes Caddy routes for that hostname. This avoids router port-forwarding for certificate issuance and is intended for access from devices inside the tailnet. Public internet access still requires Tailscale Funnel or a separately configured public hostname/DDNS path.
--online use-https:tailscale:enable-funnel additionally configures Tailscale Funnel. Because Funnel only exposes HTTPS on selected public ports, the installer maps the admin UI to https://YOUR-NODE.ts.net/admin on public port 443 and the proxy API to https://YOUR-NODE.ts.net:8443/v1/chat/completions. Funnel may require interactive approval in Tailscale or a tailnet policy update the first time it is used.
Tailscale HTTPS modes skip persistent router exposure because tailnet/Funnel routing does not depend on NAT mappings. Caddy binds the .ts.net hostname to the Tailscale interface and obtains its certificate directly from the local Tailscale daemon, which keeps that certificate renewed automatically. A separate LAN-bound public-IP route retains the self-signed certificate as a fallback while a managed on-boot, on-update, and daily Certbot job attempts to replace it with a short-lived Let's Encrypt IP certificate. For each attempt, the refresh job requests a temporary 10-minute TCP/80 lease through UPnP or NAT-PMP and lets that lease expire automatically. Issuance failures are logged and leave the working fallback untouched; the router or upstream DMZ must still make public port 80 reach the server for the ACME challenge to succeed. Renewed certificates are activated with a graceful Caddy reload so active admin streams are not interrupted. Private Tailscale IP addresses cannot receive publicly trusted certificates, so use the detected .ts.net hostname for tailnet access.
The admin and proxy ports are configurable at install/update time with:
--ports ADMIN:PROXY
Example:
bash install-club3090-server.sh --ports 18008:18009
Defaults remain:
80088009The chosen ports are propagated into the generated control service and, when --online is used, those same ports are the ones opened in the firewall and requested through UPnP. In use-https mode, those public ports are owned by Caddy while the control service moves behind it onto internal loopback-only ports. The installer also opens ports 80 and 443 when needed so Caddy can complete ACME validation for automatic certificates.
http://SERVER:8008/adminhttp://SERVER:8009/v1http://SERVER:8009/v1/chat/completionsPer-GPU routes are also supported:
http://SERVER:8009/GPU0/v1/chat/completionshttp://SERVER:8009/GPU1/v1/chat/completionshttp://SERVER:8009/GPU2/v1/chat/completionsCustom preset routes also work globally and per GPU:
http://SERVER:8009/v1/<preset>/chat/completionshttp://SERVER:8009/<preset>/v1/chat/completionshttp://SERVER:8009/GPU0/v1/<preset>/chat/completionshttp://SERVER:8009/GPU1/<preset>/v1/chat/completionsLength-capped variants remain available with short- and concise- prefixes.
The admin web panel uses Linux account credentials through pamtester and is served from:
http://SERVER:8008/admin
With --online use-https, the public entrypoint becomes:
https://SERVER:8008/admin
The inference proxy can run in either mode:
:8009 are allowed without a per-user API keyThis is controlled from the Users tab in the admin panel. In --online installs, authenticated proxy mode is enabled by default.
Per-user API entries support:
legacy, GPU0, GPU1, or *The proxy enforces those controls before forwarding a request to the backend.
User groups/plans can define shared allowed-target rules and quota defaults, so multiple users can inherit the same service tier without repeating the whole policy by hand.
If installed or updated with --local-automation, the control service also exposes a loopback-only API:
http://127.0.0.1:10881
This API is intentionally not part of the online/public surface. It is:
/opt/club3090-control/local_api_tokenTypical uses:
The current implementation exposes local user-management and server-config endpoints for same-machine tools.
--online, use-https, and --local-automation are explicit opt-in flags on each install or update run. If you omit them on a later --update, the installer turns those surfaces back off and removes the tracked firewall and UPnP exposure it previously created.
The admin UI is designed to control the whole server from one place. It exposes:
/opt/club3090-control/conversations/state.json, with the conversation list loading first and individual transcripts fetched on demand to keep large histories responsiveShift while using the archive button for permanent deletion instead/admin routes, with automatic reconnection when live status returnsmetadata.json/opt/club3090-control/audit.logAuthentication uses Linux usernames and passwords through pamtester.
On Ubuntu and Debian-family systems the installer pulls pamtester from apt. On Arch-family systems it first tries the normal package path and then falls back to yay or paru if needed, failing safely with a clear manual-install message if it still cannot be installed.
Successful admin logins are also cached in a private control-owned session file, so browser cookies can survive controller restarts and normal --update runs without forcing constant PAM reauthentication. Failed PAM checks are coalesced briefly to avoid stale browser credentials hammering the host account lockout policy.
The console log follower and browser log stream both understand the newer instance-aware docker naming scheme.
That means:
The browser-side logging stack is also much richer now than the early v0.7.x path:
The built-in local chat client has grown significantly since early v0.7.x and the v0.9.x stabilization pass. In addition to basic local chat, it now includes:
The chat renderer and syntax pipeline also received a long series of fixes across the v0.7.x, v0.8.x, and v0.9.x lines:
code_syntax.json theme/config instead of brittle embedded fallbacks/opt/club3090-control/active_mode/opt/club3090-control/last_good_mode/opt/club3090-control/instances.json/opt/club3090-control/runtime_inventory.json/opt/club3090-control/server_config.json/opt/club3090-control/admin_sessions.json/opt/club3090-control/groups.json/opt/club3090-control/network_state.json/opt/club3090-control/conversations/state.json/opt/club3090-control/users.json/opt/club3090-control/custom_models.json and /opt/club3090-control/custom-models//opt/club3090-control/benchmarks//opt/club3090-control/code_syntax.json/opt/club3090-control/local_api_tokengputemps helper under /opt/club3090-control/bin/gputemps for junction/hotspot and VRAM temperature telemetryThe script writes these systemd units:
club3090-vllm.service
club3090-control.service
club3090-benchmarks.service
club3090-caddy.service
--online use-https is enabledclub3090-console-log.service
tty1club3090-headless-x.service
These services are gated by the kernel command-line flag club3090.server=1, so they can be installed without forcing server mode on every boot.
/opt/club3090-control/control.py/opt/club3090-control/start-vllm-last-mode.sh/opt/club3090-control/follow-vllm-log.sh/opt/club3090-control/prepare-headless-x.sh/opt/club3090-control/active_mode/opt/club3090-control/last_good_mode/opt/club3090-control/instances.json/opt/club3090-control/server_config.json/opt/club3090-control/admin_sessions.json/opt/club3090-control/groups.json/opt/club3090-control/users.json/opt/club3090-control/custom_models.json/opt/club3090-control/custom-models//opt/club3090-control/benchmarks//opt/club3090-control/code_syntax.json/opt/club3090-control/network_state.json/opt/club3090-control/Caddyfile/opt/club3090-control/local_api_token/opt/club3090-control/bin/gputemps/opt/club3090-control/include/nvml.h/opt/club3090-control/src/gputemps/gputemps.c/opt/club3090-control/src/gputemps/nvml.h/opt/club3090-control/control.log/opt/club3090-control/audit.log/etc/systemd/system/club3090-vllm.service/etc/systemd/system/club3090-control.service/etc/systemd/system/club3090-benchmarks.service/etc/systemd/system/club3090-caddy.service/etc/systemd/system/club3090-console-log.service/etc/systemd/system/club3090-headless-x.serviceclub-3090; it does not replace itnvidia-settings and a private Xorg display--skip-tempsiomem=relaxed kernel switch is required for that extra telemetry on systems that otherwise block GPU MMIO mapping. It is useful on trusted single-user inference hosts, but it does relax a kernel I/O-memory safety boundary, so security-sensitive or shared systems may prefer --skip-tempsnvidia-settings, Xorg, pamtester, OpenSSL, Caddy, miniupnpc, and firewall tooling as needed, and exits safely with a clear error if a required dependency still cannot be installedinstall-club3090-server.sh: full installer/updaterREADME.md: feature and usage referencesrc/control/: split-source Python backend modules used to build the shipped control planesrc/web/: split-source HTML/CSS/JS for the admin panelsrc/build/: build pipeline, smoke tests, and vendored payload inputs used to compose the shipped single-file installersrc/build/vendor/: vendored helper source/header payloads embedded into the installer for optional junction/VRAM temperature telemetrymetadata.json: root release metadata consumed by the build and updater flowsThe project has used a split-source build pipeline since the v0.7.0 refactor, but still ships a single integrated installer artifact. Day-to-day development happens in src/control/, src/web/, and src/build/, then the root build.py wrapper regenerates the monolithic script and bundled assets from those sources. Shipped control, updater, admin UI, code-syntax, and optional temperature-helper payloads are compressed into the installer so release artifacts stay deterministic and do not depend on helper repositories being reachable during install.
Build usage:
python build.py --changes "...": normal release build; advances the numeric patch version such as v0.9.32 -> v0.9.33python build.py --iterative --changes "...": iterative rebuild on the same numeric release; advances only the letter suffix such as v0.9.32 -> v0.9.32a -> v0.9.32bchange_log_latest; only the non-iterative mode rolls the previous numeric release notes down into change_log_releasepython build.py --list-smoke-tests prints numbered smoke checks, and --smoke-tests=ID,ID can run only the relevant targeted checks during narrow fix iterations--change remains accepted as a compatibility alias, but new build notes should use --changesPython
50.0%
Shell
24.0%
JavaScript
22.7%
CSS
2.6%
Single-file installer for a club-3090 webserver providing an admin control panel, a reverse proxy that automatically routes requests to the correct containers via a single endpoint with configurable vLLM setting presets, live log management, GPU power and fan speed controls, multi-instance GPU orchestration and API access control for multiple users
Python
0
241 commits
updated Sep 29, 2026
Single-file installer for a club-3090 inference host, providing a full-featured browser admin control panel, distro-aware install/update logic, systemd service setup, a reverse proxy that automatically routes requests to the correct containers via a single endpoint with optional configurable vLLM setting presets, live log management, GPU power and fan speed controls, multi-instance GPU orchestration, and API access control for multiple users.
This repository is the server-management layer. It integrates with the upstream club-3090 runtime and, on a fresh install clones that upstream repo automatically into /opt/ai/club-3090 before wiring up the control plane.
v0.10.1* release family. The admin panel shows an unsupported-checkout warning when the local upstream club-3090 checkout falls outside the release patterns baked into the installer. You may let --migrate pull the latest upstream checkout, but if upstream introduced breaking changes, use that warning as a cue to audit the affected presets or AI Studio lanes before treating the host as production-ready./opt/club3090-control and an admin web panel running on localhost:8008/admin:8009 so you can chat with the LLM on a unified port regardless of what docker containers are in use.club-3090 checkoutGLOBAL, GPU<n>, and PAIRx_y targets, including built-in sequential auto-pairs and custom user-defined pairspull.shActive after the backend fully finishes booting, with boot/error browser notifications and fast failure reporting when a compose never actually starts a container/admin routes so refreshes can show the last known panel state while the remote service is temporarily unreachable and reconnect automatically when service returns--online mode that opens and forwards only the admin/proxy ports to the internet--online use-https mode that enables a Caddy-backed self-signed HTTPS frontend for the admin panel and proxy--skip-temps installer switch for operators who do not want the extra temperature helper, dependency install, bootloader iomem=relaxed staging, or reboot noticeVideo coming soon. Click any thumbnail to open the full-size screenshot.
This section is written for someone who just wants the server working and may not be so familiar with the complexities of LLM configuration.
Make sure your Linux machine already has:
sudoIf you plan to download gated Hugging Face models, also have your Hugging Face token ready.
By default the installer also prepares optional GDDR6/GDDR6X junction and VRAM temperature telemetry for supported NVIDIA cards. It vendors the helper source and NVIDIA NVML header inside the installer, compiles the helper locally against the host NVIDIA/libpci libraries, stages the required iomem=relaxed kernel command-line flag when it is not already active, and tells you to reboot only when that reboot is needed for the new readings to work.
Run the installer directly:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash
If you already downloaded this repo locally, you can also run:
bash install-club3090-server.sh
If you need gated model downloads and already have a Hugging Face token, use:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
HF_TOKEN=hf_xxx bash
On a fresh machine, the installer will create /opt/ai/ if needed, clone the upstream club-3090 repo into /opt/ai/club-3090, fix script permissions, and then continue with the normal setup.
When the installer finishes, open a browser and go to:
http://YOUR-SERVER-IP:8008/admin
If you installed with custom ports, replace 8008 with your chosen admin port.
When the login screen appears, sign in with your normal Linux user credentials.
Use:
In other words, log in with the same account you would normally use for sudo or for signing into that Linux machine. The admin panel does not create a separate default password for you.
After logging in, open the Presets tab.
You will see discovered model presets from the local club-3090 checkout. Some presets may show that downloads are still required.
If a preset is not ready:
Download button.Audit Logs.If you prefer to prepare the model manually in the terminal first, a common example is:
cd /opt/ai/club-3090
bash scripts/setup.sh qwen3.6-27b
Some advanced presets, such as DFlash variants, may require extra model files. The admin panel will usually tell you what command is needed.
Still in the Presets tab:
Apply.Active, the card shows how long launch took, and the action button becomes StopIf something fails, the preset will show Error and the logs will show what happened.
For most beginners:
Once a preset is active, the proxy is usually available on:
http://YOUR-SERVER-IP:8009/v1
That proxy gives you one stable API endpoint even if you switch between different backends later.
You can test it with:
curl http://YOUR-SERVER-IP:8009/v1/models
Or with a chat request:
curl -s http://YOUR-SERVER-IP:8009/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model":"qwen3.6-27b-autoround","messages":[{"role":"user","content":"Hello"}],"max_tokens":128}'
Here is the simple mental model for the main tabs:
Main: quick status, running containers, uptime, and GPU overviewSystem: service controls, power/fan controls, and machine informationPresets: download models, start/stop presets, and manage which runtime is activeChat: built-in local streaming chat client with server-backed conversation storage, lazy detail loading, per-conversation stats caching, share/export links, configuration, and local or remote MCP serversAI Studio: multimodal setup, model-resource checks, gallery browsing, and Plan/Interactive generation lanes from the Presets and Chat surfacesUsers: create API users and keys if you want to share the server safelyMetrics: request counts, usage history, runtime performance data, storage browsing, and detachable monitoringLogs: Docker logs and audit logs for troubleshootingAfter installation, most people will want to do these things:
Presets.Launch on a recommended preset.Active.http://YOUR-SERVER-IP:8009/v1 into their AI client.Most OpenAI-compatible apps need:
http://YOUR-SERVER-IP:8009/v1If the client asks for an endpoint and does not mention presets, start with the plain proxy URL above.
Try these checks in order:
Active, not just Booting.Docker Logs and Audit Logs for the exact failure message.If you changed the upstream repo manually or pulled a new upstream version, run update or migrate again so the control layer and upstream runtime stay in sync.
The script reads /etc/os-release and installs the right package names for the detected family.
systemddocker composeclub-3090 repo into /opt/ai/club-3090 on first installAs of v0.5, the server no longer ships a hardcoded preset catalog. Instead it scans the local upstream repo under /opt/ai/club-3090, parses compose headers and compose files directly, and builds /opt/club3090-control/runtime_inventory.json.
That means the Presets tab now reflects whatever models and variants exist in the checked-out upstream repo, including:
Upstream switch tags such as vllm/default, vllm/dual, vllm/gemma-mtp, and llamacpp/default are still recognized when present, but the control layer can also launch compose variants that are only discoverable by scanning the repo tree.
Since the broader v0.8.x, v0.9.x, and v0.10.x migrations, the inventory path is also registry/profile-aware and understands the newer nested upstream compose hierarchy, so the admin panel can surface richer upstream metadata such as:
autoround-int4, awq, unsloth-q4km, ubergarm-iq4ks, and ex0bit-prismThe older single-mode/legacy-dual distinction has gradually been replaced by a clearer scope-oriented control model.
The admin panel now thinks in terms of:
GLOBAL: fan a compatible preset out across every eligible target automaticallyGPU0, GPU1, GPU2, ...: explicit single-card targetsPAIR0_1, PAIR2_3, ...: dual-card targetsImportant behaviors added across the v0.7.x and v0.8.x series include:
If you use many presets often, the Summary view is also worth calling out. It keeps a relaunch-oriented cache of recent and active presets per model, including temporary boot entries and bulk stop/restart actions.
One of the biggest v0.8.x additions is the first-class Custom Model flow in the Presets tab.
Instead of manually treating every non-curated model as a one-off local hack, the panel now supports:
[+] Custom Model entry pointscripts/pull.shThis keeps the UI model-first while still deferring the actual import/evaluation logic to upstream Club-3090.
The admin panel now includes a first-class AI Studio surface for upstream Club-3090's unified ai-studio scene. It is meant to make image, video, audio, and speech workflows usable from the same browser panel as inference presets.
Current AI Studio support includes:
The control layer keeps AI Studio integration in its own wrapper/setup logic instead of patching files inside the upstream checkout. Legacy /admin/image-studio/... routes remain as backend compatibility aliases, but the browser and current API surface use /admin/ai-studio/....
Model Scores are the built-in preset validation and scoring system. They are designed to answer two different questions:
The Benchmarks modal and preset score badges are backed by saved artifacts under /opt/club3090-control/benchmarks. The current score reader only treats a Quick or Full result as valid when every required stage for that mode has a saved zero return code. Partial selected-stage shells, stale failures, and superseded sidecar JSON are ignored or archived so old artifacts do not masquerade as current scores.
Benchmark behavior includes:
The Metrics tab now uses lightweight live status payloads and bounded chart history so charts can refresh at foreground cadence even while benchmark history or model-resource scans are large. Peak markers are persisted separately from the visible rolling window, and each chart draws a single dashed peak guide for the recorded max.
Metrics also provides:
When the admin panel is served from a secure HTTPS /admin route, a Service Worker caches the admin shell and last-known status snapshot. A normal open panel labels itself disconnected only after continuous status failures exceed the configured timeout, while a refreshed page falls back to the cached shell if the remote service cannot be contacted quickly. Any successful live status response automatically exits disconnected mode.
Default install:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash
Custom upstream checkout path and default mode:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
CLUB3090_DIR=/opt/ai/club-3090 DEFAULT_MODE=vllm/dual-dflash bash
Pass a Hugging Face token inline when gated downloads are needed:
curl -fsSL https://tinyurl.com/club-3090-webserver | \
HF_TOKEN=hf_xxx bash
Running a local copy:
bash install-club3090-server.sh
The installer reads CLUB3090_DIR from a .env file in the current working directory before falling back to /opt/ai/club-3090. Use an absolute path:
CLUB3090_DIR=/home/gchamon/services/club-3090
An exported CLUB3090_DIR takes precedence over this bootstrap file. Set CLUB3090_INSTALLER_ENV_FILE to read the bootstrap value from another file. After resolving the checkout, the installer separately reads ${CLUB3090_DIR}/.env for repository settings such as MODEL_DIR, Hugging Face cache paths, and HF_TOKEN.
Custom admin and proxy ports:
bash install-club3090-server.sh --ports 18008:18009
Enable the local automation API:
bash install-club3090-server.sh --local-automation
Persist a Hugging Face token into the local club-3090 repo .env so setup jobs and later admin-panel model downloads can reuse it automatically:
bash install-club3090-server.sh --hf-token hf_xxx
Public internet exposure:
bash install-club3090-server.sh --online
Public internet exposure with TLS layer:
bash install-club3090-server.sh --online use-https
Skip optional junction/VRAM temperature telemetry setup and keep the old install/update behavior:
bash install-club3090-server.sh --skip-temps
--skip-temps bypasses helper unpacking, compiler/libpci dependency installation, bootloader iomem=relaxed staging, and the related reboot notice.
Updates the control layer without intentionally stopping an already-running backend:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash -s -- --update
Or from a local copy:
bash install-club3090-server.sh --update
Public-mode update:
bash install-club3090-server.sh --update --online
Update with custom ports:
bash install-club3090-server.sh --update --ports 18008:18009
Use --migrate when the upstream repo checkout itself needs to be replaced with a fresh clone while preserving downloaded model assets and runtime state:
curl -fsSL https://tinyurl.com/club-3090-webserver | bash -s -- --migrate
Or from a local copy:
bash install-club3090-server.sh --migrate
If an interrupted migration needs to be discarded and restarted from scratch instead of resumed:
bash install-club3090-server.sh --migrate restart
--migrate:
club-3090 repo into the original path.env, models-cache, AI Studio assets, generated results, LMCache KV data, safe compose cache directories, and the effective MODEL_DIRMODEL_DIR paths during migration and mirrors preserved weights into the effective post-merge model directory when it differs from the repo-local defaultsetup.sh reclones themOLD, OLD (2), and later generations when upstream presets changed againOLD presets while preserving genuinely different generations for comparisonOLD rows can be audited against their current counterpart/opt/club3090-control/migration-state.env--migrateHF_TOKEN=... or --hf-token ..., then stores that token in the repo .env so later setup and admin-panel download jobs can reuse itIf DEFAULT_MODE is not set, the installer counts GPUs with nvidia-smi and chooses:
vllm/default when 0 or 1 GPU is detectedvllm/dual when 2 or more GPUs are detectedFor testing or overrides, you can force the detected count:
CLUB3090_GPU_COUNT_OVERRIDE=1 bash install-club3090-server.sh
You can still hard-set a default mode:
DEFAULT_MODE=vllm/long-text bash install-club3090-server.sh
When the server is using single-card presets, the control service can manage one backend instance per GPU. Each instance has:
The instance catalog is stored in:
/opt/club3090-control/instances.json
On multi-GPU systems, this lets you run:
This is intentionally separate from the legacy dual-GPU presets. If no per-GPU single-card instances are enabled, the control service can still fall back to the old upstream dual-mode flow.
In newer releases this model also extends cleanly into the scope-based control plane:
v0.8.4x cycleWhen you install or update with --online, the script treats the control service as the only public surface.
That means:
8010-8020 and per-instance ports like 8200+ are intentionally left privateThis prevents security bypass around API-key checks, user limits, and per-instance access controls.
If you also pass use-https, the installer binds the Python control service to loopback-only internal ports and places Caddy in front of it as the public HTTPS endpoint for both admin and proxy traffic.
When HTTPS mode is enabled, the installer places Caddy in front of the admin UI and proxy on the public admin/proxy ports. Direct IP access uses the self-signed certificate path under /opt/club3090-control/tls.crt and /opt/club3090-control/tls.key by default. For a browser-trusted certificate, point a DNS or DDNS hostname at the server and set CLUB3090_HTTPS_HOST or CLUB3090_ONLINE_HOST; Caddy will then request and renew a normal Let's Encrypt certificate for that hostname. Direct IP Let's Encrypt certificates are only attempted with the explicit --online use-https:cert-ip mode, or when you explicitly set an IP host and CLUB3090_ENABLE_IP_LE=true.
If the server is already on Tailscale, --online use-https:tailscale detects the node's MagicDNS name, such as aardwolf-halfmoon.ts.net, and writes Caddy routes for that hostname. This avoids router port-forwarding for certificate issuance and is intended for access from devices inside the tailnet. Public internet access still requires Tailscale Funnel or a separately configured public hostname/DDNS path.
--online use-https:tailscale:enable-funnel additionally configures Tailscale Funnel. Because Funnel only exposes HTTPS on selected public ports, the installer maps the admin UI to https://YOUR-NODE.ts.net/admin on public port 443 and the proxy API to https://YOUR-NODE.ts.net:8443/v1/chat/completions. Funnel may require interactive approval in Tailscale or a tailnet policy update the first time it is used.
Tailscale HTTPS modes skip persistent router exposure because tailnet/Funnel routing does not depend on NAT mappings. Caddy binds the .ts.net hostname to the Tailscale interface and obtains its certificate directly from the local Tailscale daemon, which keeps that certificate renewed automatically. A separate LAN-bound public-IP route retains the self-signed certificate as a fallback while a managed on-boot, on-update, and daily Certbot job attempts to replace it with a short-lived Let's Encrypt IP certificate. For each attempt, the refresh job requests a temporary 10-minute TCP/80 lease through UPnP or NAT-PMP and lets that lease expire automatically. Issuance failures are logged and leave the working fallback untouched; the router or upstream DMZ must still make public port 80 reach the server for the ACME challenge to succeed. Renewed certificates are activated with a graceful Caddy reload so active admin streams are not interrupted. Private Tailscale IP addresses cannot receive publicly trusted certificates, so use the detected .ts.net hostname for tailnet access.
The admin and proxy ports are configurable at install/update time with:
--ports ADMIN:PROXY
Example:
bash install-club3090-server.sh --ports 18008:18009
Defaults remain:
80088009The chosen ports are propagated into the generated control service and, when --online is used, those same ports are the ones opened in the firewall and requested through UPnP. In use-https mode, those public ports are owned by Caddy while the control service moves behind it onto internal loopback-only ports. The installer also opens ports 80 and 443 when needed so Caddy can complete ACME validation for automatic certificates.
http://SERVER:8008/adminhttp://SERVER:8009/v1http://SERVER:8009/v1/chat/completionsPer-GPU routes are also supported:
http://SERVER:8009/GPU0/v1/chat/completionshttp://SERVER:8009/GPU1/v1/chat/completionshttp://SERVER:8009/GPU2/v1/chat/completionsCustom preset routes also work globally and per GPU:
http://SERVER:8009/v1/<preset>/chat/completionshttp://SERVER:8009/<preset>/v1/chat/completionshttp://SERVER:8009/GPU0/v1/<preset>/chat/completionshttp://SERVER:8009/GPU1/<preset>/v1/chat/completionsLength-capped variants remain available with short- and concise- prefixes.
The admin web panel uses Linux account credentials through pamtester and is served from:
http://SERVER:8008/admin
With --online use-https, the public entrypoint becomes:
https://SERVER:8008/admin
The inference proxy can run in either mode:
:8009 are allowed without a per-user API keyThis is controlled from the Users tab in the admin panel. In --online installs, authenticated proxy mode is enabled by default.
Per-user API entries support:
legacy, GPU0, GPU1, or *The proxy enforces those controls before forwarding a request to the backend.
User groups/plans can define shared allowed-target rules and quota defaults, so multiple users can inherit the same service tier without repeating the whole policy by hand.
If installed or updated with --local-automation, the control service also exposes a loopback-only API:
http://127.0.0.1:10881
This API is intentionally not part of the online/public surface. It is:
/opt/club3090-control/local_api_tokenTypical uses:
The current implementation exposes local user-management and server-config endpoints for same-machine tools.
--online, use-https, and --local-automation are explicit opt-in flags on each install or update run. If you omit them on a later --update, the installer turns those surfaces back off and removes the tracked firewall and UPnP exposure it previously created.
The admin UI is designed to control the whole server from one place. It exposes:
/opt/club3090-control/conversations/state.json, with the conversation list loading first and individual transcripts fetched on demand to keep large histories responsiveShift while using the archive button for permanent deletion instead/admin routes, with automatic reconnection when live status returnsmetadata.json/opt/club3090-control/audit.logAuthentication uses Linux usernames and passwords through pamtester.
On Ubuntu and Debian-family systems the installer pulls pamtester from apt. On Arch-family systems it first tries the normal package path and then falls back to yay or paru if needed, failing safely with a clear manual-install message if it still cannot be installed.
Successful admin logins are also cached in a private control-owned session file, so browser cookies can survive controller restarts and normal --update runs without forcing constant PAM reauthentication. Failed PAM checks are coalesced briefly to avoid stale browser credentials hammering the host account lockout policy.
The console log follower and browser log stream both understand the newer instance-aware docker naming scheme.
That means:
The browser-side logging stack is also much richer now than the early v0.7.x path:
The built-in local chat client has grown significantly since early v0.7.x and the v0.9.x stabilization pass. In addition to basic local chat, it now includes:
The chat renderer and syntax pipeline also received a long series of fixes across the v0.7.x, v0.8.x, and v0.9.x lines:
code_syntax.json theme/config instead of brittle embedded fallbacks/opt/club3090-control/active_mode/opt/club3090-control/last_good_mode/opt/club3090-control/instances.json/opt/club3090-control/runtime_inventory.json/opt/club3090-control/server_config.json/opt/club3090-control/admin_sessions.json/opt/club3090-control/groups.json/opt/club3090-control/network_state.json/opt/club3090-control/conversations/state.json/opt/club3090-control/users.json/opt/club3090-control/custom_models.json and /opt/club3090-control/custom-models//opt/club3090-control/benchmarks//opt/club3090-control/code_syntax.json/opt/club3090-control/local_api_tokengputemps helper under /opt/club3090-control/bin/gputemps for junction/hotspot and VRAM temperature telemetryThe script writes these systemd units:
club3090-vllm.service
club3090-control.service
club3090-benchmarks.service
club3090-caddy.service
--online use-https is enabledclub3090-console-log.service
tty1club3090-headless-x.service
These services are gated by the kernel command-line flag club3090.server=1, so they can be installed without forcing server mode on every boot.
/opt/club3090-control/control.py/opt/club3090-control/start-vllm-last-mode.sh/opt/club3090-control/follow-vllm-log.sh/opt/club3090-control/prepare-headless-x.sh/opt/club3090-control/active_mode/opt/club3090-control/last_good_mode/opt/club3090-control/instances.json/opt/club3090-control/server_config.json/opt/club3090-control/admin_sessions.json/opt/club3090-control/groups.json/opt/club3090-control/users.json/opt/club3090-control/custom_models.json/opt/club3090-control/custom-models//opt/club3090-control/benchmarks//opt/club3090-control/code_syntax.json/opt/club3090-control/network_state.json/opt/club3090-control/Caddyfile/opt/club3090-control/local_api_token/opt/club3090-control/bin/gputemps/opt/club3090-control/include/nvml.h/opt/club3090-control/src/gputemps/gputemps.c/opt/club3090-control/src/gputemps/nvml.h/opt/club3090-control/control.log/opt/club3090-control/audit.log/etc/systemd/system/club3090-vllm.service/etc/systemd/system/club3090-control.service/etc/systemd/system/club3090-benchmarks.service/etc/systemd/system/club3090-caddy.service/etc/systemd/system/club3090-console-log.service/etc/systemd/system/club3090-headless-x.serviceclub-3090; it does not replace itnvidia-settings and a private Xorg display--skip-tempsiomem=relaxed kernel switch is required for that extra telemetry on systems that otherwise block GPU MMIO mapping. It is useful on trusted single-user inference hosts, but it does relax a kernel I/O-memory safety boundary, so security-sensitive or shared systems may prefer --skip-tempsnvidia-settings, Xorg, pamtester, OpenSSL, Caddy, miniupnpc, and firewall tooling as needed, and exits safely with a clear error if a required dependency still cannot be installedinstall-club3090-server.sh: full installer/updaterREADME.md: feature and usage referencesrc/control/: split-source Python backend modules used to build the shipped control planesrc/web/: split-source HTML/CSS/JS for the admin panelsrc/build/: build pipeline, smoke tests, and vendored payload inputs used to compose the shipped single-file installersrc/build/vendor/: vendored helper source/header payloads embedded into the installer for optional junction/VRAM temperature telemetrymetadata.json: root release metadata consumed by the build and updater flowsThe project has used a split-source build pipeline since the v0.7.0 refactor, but still ships a single integrated installer artifact. Day-to-day development happens in src/control/, src/web/, and src/build/, then the root build.py wrapper regenerates the monolithic script and bundled assets from those sources. Shipped control, updater, admin UI, code-syntax, and optional temperature-helper payloads are compressed into the installer so release artifacts stay deterministic and do not depend on helper repositories being reachable during install.
Build usage:
python build.py --changes "...": normal release build; advances the numeric patch version such as v0.9.32 -> v0.9.33python build.py --iterative --changes "...": iterative rebuild on the same numeric release; advances only the letter suffix such as v0.9.32 -> v0.9.32a -> v0.9.32bchange_log_latest; only the non-iterative mode rolls the previous numeric release notes down into change_log_releasepython build.py --list-smoke-tests prints numbered smoke checks, and --smoke-tests=ID,ID can run only the relevant targeted checks during narrow fix iterations--change remains accepted as a compatibility alias, but new build notes should use --changesPython
50.0%
Shell
24.0%
JavaScript
22.7%
CSS
2.6%