moooff/HermitUI

Keep your AI conversations private. A zero-dependency, single-file local chat UI that connects to any AI endpoint.

HTML

7

145 commits

updated Sep 28, 2026

See the code

README

hermitui-logo

A lightweight, modern, and ephemeral single-page web interface for local AI models.

License: AGPL v3 Vanilla JS Zero Install

🌐 Online Demo • Try it • In-Browser AI • Benchmarks • Connect a server • Features

HermitUI demo: a GGUF model is downloaded from Hugging Face and answers in real time — WebGPU-accelerated, fully in-browser

In-browser inference with WebGPU acceleration: the model downloads straight into memory (timelapsed) and the answer streams in real time.

Try the Live Online Demo Try the In-Browser AI Demo

Left: connect it to your local AI server. Right: one click downloads a small but capable model (Qwen3-0.6B, ~380 MB) and chats fully inside your browser — no server at all.

HermitUI is a chat interface that is one .html file. No install, no server, no build step, no npm — double-click it and it opens.

Two things set it apart, and the combination is the point:

  • 🧠 It runs models itself. GGUF models execute entirely in your browser via llama.cpp compiled to WebAssembly, with WebGPU acceleration. A 12.1 GB model loads in a tab and decodes at 43 tok/s — and we measured it properly.
  • 🔒 It stores absolutely nothing. No localStorage, no IndexedDB, no cookies, no model cache, no telemetry. Close the tab and the conversation and the model are gone — and you can check that in a minute.

Or ignore all of that and point it at LM Studio, Ollama, llama.cpp, or vLLM as a normal client.

Built for the machines where nothing else fits: air-gapped boxes, locked-down corporate and government networks, shared kiosks and hot desks.

⚡ Try it in 60 seconds

One click: open the 🧠 In-Browser AI Demo — it pre-fills Qwen3-0.6B (~380 MB) via the #gguf= hash parameter; confirm the banner and chat.

Or do it by hand:

  1. Save dist/hermit-ui-wllama.html to disk (right-click → Save link as…, since GitHub serves raw .html as plain text), then open it in your browser.
  2. Settings → Backend Mode → True Offline (Wllama GGUF), then paste into the URL field: hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
  3. Hit ⬇️ Load (~380 MB download) and chat — no server, no install, and nothing persisted.

🧠 In-browser inference

HermitUI can run true offline inference entirely in the browser — no local server or OpenAI-compatible endpoint required. It's powered by wllama (llama.cpp compiled to WebAssembly, with optional WebGPU acceleration): you load a .gguf model file and chat with it directly on the page.

This ships as a dedicated build output, dist/hermit-ui-wllama.html — the regular standalone app plus a Backend Mode switch in Settings (Remote / Local API ↔ True Offline (Wllama GGUF)). The main builds stay lean: the feature is stripped out of them at build time.

  • 🔌 The app needs no network: The wllama engine (JS + WASM) is embedded directly into the file at build time (gzipped, decompressed in-browser via the native DecompressionStream API), so the ~6 MB file is complete on its own — perfect for USB-stick distribution to air-gapped machines. Pair it with a .gguf from disk and the whole stack is offline; only the optional download-by-URL path touches the network.
  • 📂 Local GGUF loading: Pick a .gguf file from disk and run it fully client-side, with an optional WebGPU toggle for hardware acceleration. Download a model once, keep it, and re-pick it every session — no network involved, so this is also the fastest way to use HermitUI repeatedly (and the only way on an air-gapped machine).
  • 🔗 Load by URL / Hugging Face: Paste a direct .gguf link, a Hugging Face /blob/ page URL (auto-rewritten to /resolve/), or the hf:user/repo/file.gguf shorthand, then hit Load. The model streams straight into memory with a live progress bar — true to the ephemerality promise, nothing is written to browser storage. A URL-loaded model is therefore fetched again next session — unless you hit 💾 Save a copy to disk under the URL field, which hands the same file to your browser's download flow so the file picker can load it from then on. A model can also be baked into a shareable link: hermit-ui-wllama.html#gguf=hf:user/repo/file.gguf (see Configuration via URL).
  • 🎚️ Configurable inference: Adjustable context window (n_ctx, default 32k — automatically halved until it fits in memory, with the effective size shown in the status line) and max output tokens per reply (default 4096); temperature, top-p, and seed from the regular settings apply too.
  • 🧩 Layered chat-template handling: Uses the model's own embedded tokenizer.chat_template when present, otherwise auto-detects a sane format from the model architecture (ChatML, Llama 3, Mistral, Gemma, Phi-3, Zephyr, Alpaca, …), with a manual override.
  • 🐛 Quake-style debug console: A drop-down console with graduated verbosity levels (Off → Errors → Warnings → Info → Debug) that surfaces engine init, download/load progress, model metadata, the exact prompt sent, and native llama.cpp logs.
  • ⏱️ Live tokens/s: A real-time generation-speed readout while the model streams.

Every link below opens the wllama build with that model pre-filled via #gguf= — confirm the banner and it streams straight into memory. Nothing is written to browser storage, so a link-loaded model is fetched again next session unless you keep a copy of the .gguf. Start small; the bigger rungs need a modern Chrome/Edge (see Browser support).

ModelDownloadTry it
Qwen3-0.6B0.4 GB▶ Run in browser
Qwen3-1.7B1.1 GB▶ Run in browser
Qwen3-4B ⭐2.5 GB▶ Run in browser
Qwen3-8B5.0 GB▶ Run in browser
Gemma-4-E2B3.1 GB▶ Run in browser
Gemma-4-E4B5.0 GB▶ Run in browser
Gemma-4-12B7.1 GB▶ Run in browser
gpt-oss-20b ⚠️12.1 GB▶ Run in browser

⭐ = best speed/quality trade-off on a WebGPU machine. ⚠️ = the fastest model above 7 GB measured here (~43 t/s), but a 12 GB download that needs Chrome/Edge and a mostly-free GPU — see the note under the results table. Quants are Q4_K_M from unsloth except gpt-oss-20b, which is MXFP4 from ggml-org. On a CPU-only machine, use Qwen3-1.7B or smaller.

Download once, reuse offline

You only pay for the transfer once. The links above re-fetch the model each session, which is fine for a first try but wasteful afterwards. Keep the .gguf on disk instead:

  1. Put the model URL in Settings → Backend Mode → True Offline, then click 💾 Save a copy to disk under the URL field. That downloads the same file through your browser's normal download flow, to a location you pick — the app itself still writes nothing.
  2. Next session, load that file with the file picker in the same panel. No transfer, no network at all — and it's the only way to work on an air-gapped machine.

(Downloading the .gguf straight from Hugging Face works identically; the save link just points at the same file.)

Browsers don't allow a file path to be pre-filled from a link, so it stays a manual pick each session, and you still pay the model's load time — just not the download. The table below puts that at ≈6 s for Qwen3-0.6B up to ≈80 s for gpt-oss-20b, measured over a local HTTP server, so a disk pick is if anything quicker.

📊 Benchmarks

A Playwright harness drives the unmodified dist/hermit-ui-wllama.html — via #gguf= and the app's own buttons, exactly as a user would — and has every model answer the same 10 questions, scored for correctness by hand. 16 threads, RTX 5070 Ti, Edge (WebGPU), 3+ runs per model on a verified-idle GPU (the harness refuses to start otherwise).

ModelSizeLoadavg TTFTdecode t/send-to-end t/s
Qwen3-0.6B0.4 GB5.7s0.66s7956
Qwen3-1.7B1.1 GB11.3s0.91s7354
Qwen3-4B2.5 GB21.4s1.03s6443
Qwen3-8B5.0 GB41.5s1.17s5535
Gemma-4-E2B3.1 GB26.5s9.3s406.1
Gemma-4-E4B5.0 GB36.6s11.7s325.6
Gemma-4-12B7.1 GB58.6s15.2s364.0
gpt-oss-20b †12.1 GB~80s3.4s4315

Load = engine init + model transfer from a local HTTP server (not your Hugging Face download time). Decode = pure generation speed with TTFT excluded; end-to-end = what the app's own stats readout shows, prompt processing included.

† gpt-oss-20b is MXFP4, not Q4_K_M, and it is a reasoning model. Its hidden thinking trace inflates the end-to-end denominator — one terse-but-correct logic answer read 2.9 t/s end-to-end while decoding at 40.5 t/s — so that column is not comparable to the rest. Judge it on the decode column, which is excellent — and the answers are genuinely good. It is the strongest large rung here if you have the VRAM to spare.

CPU-only (WebGPU off, same machine): Qwen3-0.6B ≈ 16 t/s, Qwen3-1.7B ≈ 10 t/s. Everything larger is unusable — Qwen3-4B ≈ 3 t/s, Gemma-4-E2B ≈ 1.6 t/s.

What the numbers say:

  • Qwen3 is the sweet spot in a browser. Even 8B stays interactive on WebGPU, and TTFT is ~1 s across the whole family.
  • Gemma-4 decodes fine but prompt-processes slowly under wllama — 9–15 s before the first token drags the end-to-end figure into single digits even though tokens then arrive at 30–40 t/s. All its answers were correct: an engine-side prompt-eval gap, not a model-quality one.
  • 12.1 GB runs, and runs well — and 13.0 GB does too. Qwen3.8-27B at 4 bpw loads and holds ~8 t/s, so the practical ceiling is your VRAM rather than any engine limit. gpt-oss-20b out-decodes the 7.1 GB Gemma-4-12B on a 16 GB card — a sparse MoE activating only ~3.6B params per token beats a dense model despite being 70% larger on disk. Architecture predicts throughput, not file size. WASM Memory64 (Chrome/Edge) is required: without it, anything above ~4 GB fails outright.
  • The GPU matters more than the model. 0.6B → 8B costs ~30% of decode throughput; dropping to CPU costs ~80%. Free VRAM matters most of all — a contended GPU understated these same runs by 23×, so if your numbers look nothing like these, check what else is on your card first.

Reproduce it yourself — one pip install, no Node. Each run writes review.md with the timings and every answer in full, so quality is reviewable and not just asserted, plus a machine-readable run.json.

Qwen3.8-27B — verified working in the browser

Qwen3.8-27B at 4 bpw runs in the web build. At 13.0 GB it is the largest model verified here, ahead of gpt-oss-20b.

Modelbyteshape/Qwen3.8-27B-GGUF → Qwen3.8-27B-IQ4_XS-4.00bpw.gguf
Size / quant13.0 GB, IQ4_XS (4.00 bpw)
Architectureqwen35 — hybrid attention + SSM, 65 layers, 248k vocab
Load time18–23 s from a local file (engine init + weights)
VRAM14.2–14.9 GB of a 16 GB card, at n_ctx 8192
Speed (app readout)≈8 tok/s — 7.7–8.5 across seven replies of 550–920 tokens

The speed figure is exactly what the app's own live tokens/s readout shows, i.e. end-to-end with prompt processing included — the same column as end-to-end t/s in the table above, not a decode-only number. Short replies read lower (≈4–6 t/s at 50–170 tokens, and ≈2.5 t/s on the very first reply after a load) purely because a fixed TTFT dominates a small denominator. Judge it on the longer generations.

Two caveats worth stating plainly:

  • Verified via the file picker, not #gguf=. A 13 GB one-click link was not tested, and at this size the download-once-and-re-pick flow is the sensible route anyway. Grab the .gguf from the repo above, then load it with the file picker.
  • Chrome/Edge only, and it needs the VRAM. WASM Memory64 is required (as for every rung above ~4 GB), and 13 GB of weights leaves little headroom on a 16 GB card — keep n_ctx modest (8192 was used here; this model's KV runs ~64 KB per token).

It is a reasoning model that thinks at xhigh by default, which is slow enough to matter on a 27B: use the 🧠 Think control in the chatbox to drop to low or Off for routine questions. Measured on three trivial prompts, Off finished in 24.9 s with zero reasoning versus 57.7 s and ~935 characters of reasoning at the default — see Flexible thinking control.

This rung was verified by hand rather than through the benchmark harness, so it has no avg TTFT / decode t/s entry in the table above.

For comparison, the same file under native llama-server on the same GPU decodes at 77–101 t/s — about 10× faster. The browser build's appeal is that it needs no install and nothing leaves the tab, not throughput.

Browser support & model size limits

How large a model you can load — and how fast it runs — depends on two WebAssembly/GPU features of your browser, which wllama detects at load time:

CapabilityWhat it enablesChrome / EdgeFirefoxSafari
JSPI (WebAssembly.Suspending)Streams the GGUF straight into the engine instead of copying it whole into the WASM heap → model size limited only by your RAM/VRAM✅ Chrome 137+⚠️ 153+ only❌ none → ~3 GB cap
WebGPU (in workers)Hardware-accelerated inference✅ mature⚠️ new / may fail to initialize → CPU fallback⚠️ present in recent versions, untested here

HermitUI probes both at load time (by capability, not user-agent) and warns you in the panel before you start a download that can't succeed.

In practice:

  • Chrome / Edge: Multi-GB models (7B+ quants) load and run fine, with WebGPU acceleration. The limit is your actual RAM/VRAM.
  • Firefox before 153: Without JSPI, wllama falls back to copying the entire model file into the 4 GiB WASM heap. Models larger than roughly 3 GB fail with the cryptic error source array is too long (an unchecked allocation failure inside wllama). Fix: update to Firefox 153+, which enables JSPI by default. You can verify support by typing !!WebAssembly.Suspending into the DevTools console — it must print true.
  • Firefox speed: Even with JSPI, Firefox's WebGPU support is much newer than Chrome's and may not initialize inside the wllama worker, dropping inference to single-threaded CPU WASM — noticeably slower than Chrome on the same machine. Check the debug console (verbosity Debug, then reload the model) to see whether a WebGPU device or the CPU backend was picked. If WebGPU misbehaves, try unchecking the WebGPU toggle — a clean CPU run can beat a broken GPU path.
  • Safari: no JSPI at all, so the same ~3 GB ceiling applies with no version to upgrade to — Qwen3-0.6B and 1.7B are fine, Qwen3-4B is the realistic top rung, and anything larger fails. The benchmarks above were not run on Safari. The rest of HermitUI (the normal client mode) works everywhere.

🔌 Connect to your own endpoint

Prefer to run the model outside the browser? The standalone build (index.html / dist/hermit-ui-standalone.html) talks to anything that speaks the OpenAI chat completions API.

  1. Start your local AI server — LM Studio, Ollama, llama.cpp, or vLLM.
  2. Open HermitUI — double-click index.html; it opens in any modern browser.
  3. Configure — click ⚙️ Settings in the top right to set the API URL, model name, API key, or system prompt.
Configuration examples for popular servers

LM Studio (the default)

  1. Launch LM Studio and start the Local Server.
  2. API URL: http://localhost:1234/v1/chat/completions
  3. Model Name: leave blank, or set to the specific model identifier you loaded.
  4. Tip: ensure CORS is enabled in the LM Studio settings.

Ollama

  1. Start your Ollama server from the terminal, making sure to enable CORS:
    OLLAMA_ORIGINS="*" ollama serve
    
  2. API URL: http://localhost:11434/v1/chat/completions
  3. Model Name: the name of the model you pulled (e.g., llama3, mistral, deepseek-coder).

vLLM

  1. Start your vLLM server — it serves an OpenAI-compatible endpoint and already allows every origin by default:
    vllm serve meta-llama/Llama-3.1-8B-Instruct
    
    To pin the allowlist instead, pass --allowed-origins a JSON array: --allowed-origins '["https://example.com"]'.
  2. API URL: http://localhost:8000/v1/chat/completions
  3. Model Name: the model you served (e.g., meta-llama/Llama-3.1-8B-Instruct).

llama.cpp (llama-server)

  1. Start llama-server with a GGUF. --jinja is the important flag — it applies the model's own chat template, which is what makes the 🧠 Think control work:
    llama-server --model model.gguf --jinja --host 0.0.0.0 --port 8080
    
  2. API URL: http://localhost:8080/v1/chat/completions
  3. Model Name: anything — llama-server serves whichever model it loaded.
  4. Tip: add --reasoning-format deepseek for a reasoning model. It returns the trace in reasoning_content, which HermitUI renders in the collapsible think block. Settings → Check Reasoning Support reads this server's /props and will report reasoning support exactly.

Reference config: a 27B on a 16 GB card, measured

If you have a 16 GB GPU, this is a config that fits and what each flag costs. Qwen3.8-27B at 4 bpw, one RTX 5070 Ti, using the model's own MTP layer for speculative decoding. Nothing here is card-specific except the numbers — the same reasoning applies to any 16 GB card.

llama-server --model Qwen3.8-27B-IQ4_XS-4.00bpw.gguf \
  --ctx-size 40000 --jinja --reasoning-format deepseek --reasoning-preserve \
  --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
  --spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
  --cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \
  --n-gpu-layers all --parallel 1 --threads 32 --batch-size 1024 --ubatch-size 1024 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0 \
  --host 0.0.0.0 --port 8080

Speed (llama-server's own timings, so decode-only — TTFT excluded):

WorkloadWith --spec-typeWithoutGain
Prose (200-word explainer, ×3)77.1 t/s50.4 t/s1.53×
Reasoning prompt91–96 t/s—~1.85×
Code generation100.7 t/s50.2 t/s2.01×

Prompt processing reached 1,489 t/s on a 6,742-token prompt. The no-speculation baseline was flat to within 0.3 t/s across four runs, so those gains are signal rather than noise.

Draft acceptance explains the spread — 98.7% on code, 82–88% on the reasoning prompt, 50–65% on prose. MTP predicts structured text far better than free prose, so the speedup is largest exactly where you generate code.

⚠️ --parallel 1 is not optional — it is worth 3.6×. Omit it and the server defaults to 4 slots, which flips llama.cpp to a unified KV cache (kv_unified = 'true') sized across all of them, so every decoded token does far more attention work. Measured on this machine, same GPU, same everything else: 20.4 t/s prose / 29.2 t/s code with 4 slots, versus 73.8 / 97.1 with one. VRAM barely moves (~15.5 GB either way), so nothing looks wrong — you just silently lose three quarters of your throughput. Check total_slots at http://localhost:8080/props; it must read 1.

Where the VRAM goes, at --ctx-size 40000 — as nvidia-smi and the server's own load log report it:

VRAM
Model weights on GPU (4 bpw)12,376 MiB
Compute buffers (--ubatch-size 1024)396 MiB
KV cache @ 40k q8_0, plus CUDA context overhead≈2,260 MiB
MTP draft context (what --spec-type costs)≈860 MiB
Total resident15,893 of 16,303 MiB

That leaves roughly 410 MiB spare — it runs, but with no room for anything else on the card. Don't run a WebGPU browser (including HermitUI's own in-browser mode) against the same GPU at this context size.

If it doesn't fit on your card, cut in this order:

  1. Lower --ctx-size. The KV cache scales linearly with it, so halving 40k → 20k returns roughly 1,000 MiB. This is the cheapest VRAM you will find.
  2. Keep --cache-type-k q8_0 --cache-type-v q8_0 (and --flash-attn on, which the quantized-KV path wants). q8_0 stores one byte per element where f16 stores two, so reverting to f16 would roughly double the KV figure above — at 40k context that is the difference between fitting and not.
  3. Drop --spec-type last. It returns ~860 MiB but costs the 1.5–2.0× decode speedup above, so trade it away only once the context is as small as you can live with.

For scale: this is roughly 10× the ~8 t/s the same model reaches in the in-browser build. Native is far faster — the browser build buys zero-install and privacy, not speed.

Cloud models (OpenRouter, OpenAI, Groq, …)

  1. API URL: the provider's chat completions endpoint (e.g., https://openrouter.ai/api/v1/chat/completions).
  2. Model Name: the model you want (e.g., anthropic/claude-opus-4.8) — check your provider's model list for the exact slug.
  3. API Key: enter your provider's key in the settings menu.

[!WARNING] Privacy note: using cloud models is generally not advised if you require strict privacy. Your data leaves your machine, and it is unclear how these providers handle, store, or train on it. For true ephemerality, stick to local models.

Troubleshooting (CORS)

If HermitUI fails to connect to your local AI server (e.g., a "Network Error"), it is most likely CORS. Because HermitUI runs as a local file (file://), browsers block its requests to http://localhost unless the server explicitly allows it.

  • LM Studio: "Local Server" tab → find the CORS toggle → turn it ON.
  • Ollama: set OLLAMA_ORIGINS before starting, e.g. OLLAMA_ORIGINS="*" ollama serve.
  • vLLM: allows all origins out of the box — if you narrowed it, widen --allowed-origins (it takes a JSON array, e.g. '["*"]').

✨ Features

  • 📦 Zero-dependency setup: all external libraries (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid) and the Inter font are bundled directly into the file. No installation, no build step. (A CDN-linked developer version lives in dist/hermit-ui-cdn.html.)
  • 🔒 Privacy first & ephemeral: no localStorage, IndexedDB, or cookies — nothing survives the tab.
  • 🧠 Thinking-model support: built-in parser formats <think>, <thought>, and <reasoning> tags as they stream from reasoning models.
  • 🎚️ Flexible thinking control: a 🧠 Think selector beside the persona picker sets reasoning depth — Off / Low / Medium / High — for models that support it. Off genuinely stops the reasoning trace rather than shortening it, which is the difference between a 25 s and a 58 s answer on a 27B. It only appears where it can actually do something: for a local GGUF the model's own chat template is scanned for the reasoning variables, and for a remote endpoint Settings → Check Reasoning Support reads llama.cpp's /props capability flags (or Ollama's template). Levels are validated against the set the template accepts, so a value it would reject is never sent.
  • ⚡ Real-time streaming with 📊 live performance stats — prompt tokens, completion tokens, tokens/second, and total duration.
  • 🖼️ Image & vision support: upload, drag-and-drop, or paste (Ctrl+V) images for vision-capable models, sent as image_url content per the OpenAI schema with automatic vision-model detection.
  • 📝 Rich rendering: Markdown with per-block copy buttons, syntax highlighting, 🧮 LaTeX math ($…$, $$…$$, \(…\), \[…\]) rendered via KaTeX to native MathML — no webfonts, works mid-stream and offline — and 📈 Mermaid diagrams from ```mermaid fences.
  • 🎭 Personas: switch between preset system prompts (technical, general, writing, tutor) on the fly.
  • ✏️ Edit & regenerate any previous message without restarting the conversation.
More features
  • 📎 Context attachments: drag-and-drop or upload text files to inject their contents into your prompt.
  • 🎛️ Advanced sampling controls: temperature, max_tokens, top_p, presence_penalty, frequency_penalty, and seed from a collapsible Settings panel. Params are only sent when set, keeping payloads compatible with minimal backends.
  • 🎨 Modern UI/UX: clean, responsive design with smooth micro-animations, comprehensive CSS variables for theming, and a glassmorphism feel. Light and dark themes.
  • 💾 Chat export & import: download the entire conversation as a formatted Markdown file (📤, Ctrl+E), and load one back in later with 📂 (Ctrl+I) — messages, system prompt and attached context are restored, and the model picks the conversation up where you left it. The exported file is the only copy: nothing is stored in the browser. Images can't come back, since the export records them as [N images attached].
  • ⚙️ Customizable settings: API URL, model name, API key, and system prompt via the on-page settings overlay.

🔗 Configuration via URL

You can pre-configure HermitUI through the URL fragment (the part after #), so a single link or bookmark carries the whole connection setup:

hermit-ui-standalone.html#api=http://localhost:8080/v1&model=qwen3-8b
hermit-ui-standalone.html#api=https://api.groq.com/openai/v1&key=gsk_...&model=llama-3.3-70b
hermit-ui-wllama.html#gguf=hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
ParameterEffect
apiAPI base URL (same as the Settings field)
modelModel name
keyAPI key
personaPreset persona: technical, general, writing, or tutor
gguf(wllama build only) GGUF model to load in-browser — direct URL, Hugging Face link, or hf:user/repo/file.gguf shorthand. Shows a one-click confirmation banner before downloading.

Why the fragment and not ?query: the part after # never leaves your browser — it is not sent in any HTTP request — and nothing is stored, so this stays true to the ephemerality promise (the URL is the config; refresh keeps your setup). Applied settings are always announced in a toast, so a shared link can't reconfigure the app invisibly. Free-text system prompts are deliberately not supported as a parameter, since a link could smuggle a malicious prompt.

[!NOTE] A key in the URL is never transmitted, but it does end up in your browser history (and any bookmark you save). Prefer entering keys in Settings on shared machines.

🎯 Ideal use cases

  • Heavily regulated environments: enterprise or government networks where software installation is restricted, but a secure local or remote inference endpoint is accessible.
  • Air-gapped systems: distribute on a USB stick and run on disconnected machines — either against a local network LLM server, or with the wllama build and a .gguf, against nothing at all.
  • Ephemeral kiosks & shared terminals: no chat history is ever written, making it safe for public workstations and desk-sharing environments.

🏗️ Architecture & philosophy

HermitUI enforces strict architectural constraints to remain lightweight and accessible:

  • Single file constraint: the final product is always a single, standalone .html file. The src/ directory is a blueprint only — its split into index.html, style.css, and script.js exists for maintainability, and build.py assembles them back into one file.
  • Vanilla only: no React, Vue, Angular, or other frontend frameworks.
  • No build tools: no package.json, npm, Webpack, or Vite.
  • No CSS frameworks: pure vanilla CSS, no Tailwind or Bootstrap.
  • Security: all rendered AI responses are sanitized with DOMPurify to prevent XSS.

The live build runs on GitHub Pages at moooff.github.io/HermitUI.

Verify the privacy claim

"Stores nothing" is the whole pitch, so don't take it on trust — checking takes about a minute.

  • At runtime: open DevTools → Application → Storage and use the app normally. Local Storage, Session Storage, IndexedDB, Cookies and Cache Storage stay empty for the entire session, including after a model has loaded. This is the authoritative check: it covers the bundled libraries and the embedded wllama engine, not just HermitUI's own code.
  • In the source: grep the file you downloaded.
    grep -o "localStorage\|sessionStorage\|indexedDB\|document\.cookie" dist/hermit-ui-standalone.html | wc -l
    
    The answer is 2, and both are false positives — Highlight.js's list of JavaScript keywords contains the strings localStorage and sessionStorage. There is not one call site. (In the wllama build the engine ships gzipped, so grep sees the app but not the engine; use the runtime check above to cover it.)
  • On the network: the Network tab shows requests only to endpoints you configured yourself — your API server, or the model URL if you chose to load a GGUF that way. No analytics, no phone-home, and no CDN at runtime in the standalone builds.

Model weights get the same treatment. A downloaded GGUF is streamed into an in-memory Blob instead of going through wllama's own URL loader, specifically because that one would persist the model to OPFS — see downloadGgufToBlob in src/script.js.

📦 Building & development

The root index.html (a copy of dist/hermit-ui-standalone.html) is a completely offline, standalone build: web fonts and images are base64-encoded and external JS/CSS libraries are injected directly into the file, which is what makes it work in air-gapped environments.

To modify it, edit the modular sources in src/ — index.html, style.css, and script.js, which reference libraries via CDN for convenient local development — then run:

python build.py        # or python3 build.py

Prerequisites: Python 3 and, on the first run, an internet connection. That is the whole list — build.py uses only the standard library, so there is no pip install, no package.json, and no Node.

The network requirement is worth spelling out, since it cuts against the rest of the project: the output runs offline, but producing it does not. On the first build, build.py downloads the pinned library versions (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid, the Inter font, and the wllama engine — ~14 MB total) into libs/, verifying each one against the SRI hash pinned in src/index.html. Everything after that is cached and offline; pass --refresh to force a re-download. So if you need to build on an air-gapped machine, copy a populated libs/ directory across with the repo.

This generates the standalone build at dist/hermit-ui-standalone.html, copies it to the root index.html for GitHub Pages, and creates the alternative builds in dist/. The standalone, CDN, and wllama variants (dist/hermit-ui-standalone.html, dist/hermit-ui-cdn.html, dist/hermit-ui-wllama.html) are committed so they are browsable and downloadable straight from GitHub; the local variant dist/hermit-ui-local.html and the downloaded libs/ are generated-only and stay gitignored.

🤝 Contributing

Bug reports, questions and ideas are welcome in Issues; pull requests are welcome too. A bug report travels much further with your browser and version, whether WebGPU was on, and — for in-browser inference — the model and the output of the debug console at verbosity Debug.

Before opening a PR, two things will save you a rewrite:

  • Read AGENTS.md. It is the single source of truth for this project's rules, and they are unusually strict on purpose: single-file output, vanilla JS only, no build tools or frameworks or CSS libraries, no localStorage/IndexedDB/cookies for any reason, and every AI-rendered string sanitized through DOMPurify. A change that breaks one of these can't be merged no matter how good it is — the constraints are the product.
  • Edit src/, never the generated files. dist/*.html and the root index.html are build artifacts and get overwritten; run python build.py and commit the regenerated outputs along with your source change.

There is no test suite or linter. Verify a change by opening the rebuilt dist/hermit-ui-standalone.html (or dist/hermit-ui-wllama.html) in a browser and exercising the affected path by hand. If your change touches inference performance, the benchmark harness produces numbers that can go straight into a PR description.

🛠️ Built with

  • Vanilla HTML5 / CSS3 / ES6 JavaScript
  • wllama — llama.cpp in WebAssembly, for in-browser inference
  • Marked.js — Markdown parsing
  • DOMPurify — HTML sanitization / XSS prevention
  • Highlight.js — code syntax highlighting
  • KaTeX — LaTeX math rendering (MathML output)
  • Mermaid — diagram rendering from ```mermaid fences
  • Google Fonts (Inter) — typography

🗺️ Roadmap

  • Split-GGUF support: load sharded models (-00001-of-000NN.gguf) through the in-browser URL loader.
  • Save the in-flight download: write the model to disk as it downloads (File System Access API), so keeping a copy costs one transfer instead of two. Chrome/Edge only, hence the simpler two-download link today.
  • Companion text-analysis app: a separate ultra-light single-file build (Transformers.js + WebGPU, sub-200 MB models) dedicated to summarization and text analysis.

The full list of ideas and tasks lives in docs/backlog.md.

📄 License

This project is open-source and available under the terms of the GNU AGPL v3. See the included LICENSE file for the full text.

air-gapped
air-gapped-ai
chat-ui
local-llm
openai-api
privacy-first
single-file
zero-telemetry

moooff/HermitUI

Keep your AI conversations private. A zero-dependency, single-file local chat UI that connects to any AI endpoint.

HTML

7

145 commits

updated Sep 28, 2026

See the code

README

hermitui-logo

A lightweight, modern, and ephemeral single-page web interface for local AI models.

License: AGPL v3 Vanilla JS Zero Install

🌐 Online Demo • Try it • In-Browser AI • Benchmarks • Connect a server • Features

HermitUI demo: a GGUF model is downloaded from Hugging Face and answers in real time — WebGPU-accelerated, fully in-browser

In-browser inference with WebGPU acceleration: the model downloads straight into memory (timelapsed) and the answer streams in real time.

Try the Live Online Demo Try the In-Browser AI Demo

Left: connect it to your local AI server. Right: one click downloads a small but capable model (Qwen3-0.6B, ~380 MB) and chats fully inside your browser — no server at all.

HermitUI is a chat interface that is one .html file. No install, no server, no build step, no npm — double-click it and it opens.

Two things set it apart, and the combination is the point:

  • 🧠 It runs models itself. GGUF models execute entirely in your browser via llama.cpp compiled to WebAssembly, with WebGPU acceleration. A 12.1 GB model loads in a tab and decodes at 43 tok/s — and we measured it properly.
  • 🔒 It stores absolutely nothing. No localStorage, no IndexedDB, no cookies, no model cache, no telemetry. Close the tab and the conversation and the model are gone — and you can check that in a minute.

Or ignore all of that and point it at LM Studio, Ollama, llama.cpp, or vLLM as a normal client.

Built for the machines where nothing else fits: air-gapped boxes, locked-down corporate and government networks, shared kiosks and hot desks.

⚡ Try it in 60 seconds

One click: open the 🧠 In-Browser AI Demo — it pre-fills Qwen3-0.6B (~380 MB) via the #gguf= hash parameter; confirm the banner and chat.

Or do it by hand:

  1. Save dist/hermit-ui-wllama.html to disk (right-click → Save link as…, since GitHub serves raw .html as plain text), then open it in your browser.
  2. Settings → Backend Mode → True Offline (Wllama GGUF), then paste into the URL field: hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
  3. Hit ⬇️ Load (~380 MB download) and chat — no server, no install, and nothing persisted.

🧠 In-browser inference

HermitUI can run true offline inference entirely in the browser — no local server or OpenAI-compatible endpoint required. It's powered by wllama (llama.cpp compiled to WebAssembly, with optional WebGPU acceleration): you load a .gguf model file and chat with it directly on the page.

This ships as a dedicated build output, dist/hermit-ui-wllama.html — the regular standalone app plus a Backend Mode switch in Settings (Remote / Local API ↔ True Offline (Wllama GGUF)). The main builds stay lean: the feature is stripped out of them at build time.

  • 🔌 The app needs no network: The wllama engine (JS + WASM) is embedded directly into the file at build time (gzipped, decompressed in-browser via the native DecompressionStream API), so the ~6 MB file is complete on its own — perfect for USB-stick distribution to air-gapped machines. Pair it with a .gguf from disk and the whole stack is offline; only the optional download-by-URL path touches the network.
  • 📂 Local GGUF loading: Pick a .gguf file from disk and run it fully client-side, with an optional WebGPU toggle for hardware acceleration. Download a model once, keep it, and re-pick it every session — no network involved, so this is also the fastest way to use HermitUI repeatedly (and the only way on an air-gapped machine).
  • 🔗 Load by URL / Hugging Face: Paste a direct .gguf link, a Hugging Face /blob/ page URL (auto-rewritten to /resolve/), or the hf:user/repo/file.gguf shorthand, then hit Load. The model streams straight into memory with a live progress bar — true to the ephemerality promise, nothing is written to browser storage. A URL-loaded model is therefore fetched again next session — unless you hit 💾 Save a copy to disk under the URL field, which hands the same file to your browser's download flow so the file picker can load it from then on. A model can also be baked into a shareable link: hermit-ui-wllama.html#gguf=hf:user/repo/file.gguf (see Configuration via URL).
  • 🎚️ Configurable inference: Adjustable context window (n_ctx, default 32k — automatically halved until it fits in memory, with the effective size shown in the status line) and max output tokens per reply (default 4096); temperature, top-p, and seed from the regular settings apply too.
  • 🧩 Layered chat-template handling: Uses the model's own embedded tokenizer.chat_template when present, otherwise auto-detects a sane format from the model architecture (ChatML, Llama 3, Mistral, Gemma, Phi-3, Zephyr, Alpaca, …), with a manual override.
  • 🐛 Quake-style debug console: A drop-down console with graduated verbosity levels (Off → Errors → Warnings → Info → Debug) that surfaces engine init, download/load progress, model metadata, the exact prompt sent, and native llama.cpp logs.
  • ⏱️ Live tokens/s: A real-time generation-speed readout while the model streams.

Every link below opens the wllama build with that model pre-filled via #gguf= — confirm the banner and it streams straight into memory. Nothing is written to browser storage, so a link-loaded model is fetched again next session unless you keep a copy of the .gguf. Start small; the bigger rungs need a modern Chrome/Edge (see Browser support).

ModelDownloadTry it
Qwen3-0.6B0.4 GB▶ Run in browser
Qwen3-1.7B1.1 GB▶ Run in browser
Qwen3-4B ⭐2.5 GB▶ Run in browser
Qwen3-8B5.0 GB▶ Run in browser
Gemma-4-E2B3.1 GB▶ Run in browser
Gemma-4-E4B5.0 GB▶ Run in browser
Gemma-4-12B7.1 GB▶ Run in browser
gpt-oss-20b ⚠️12.1 GB▶ Run in browser

⭐ = best speed/quality trade-off on a WebGPU machine. ⚠️ = the fastest model above 7 GB measured here (~43 t/s), but a 12 GB download that needs Chrome/Edge and a mostly-free GPU — see the note under the results table. Quants are Q4_K_M from unsloth except gpt-oss-20b, which is MXFP4 from ggml-org. On a CPU-only machine, use Qwen3-1.7B or smaller.

Download once, reuse offline

You only pay for the transfer once. The links above re-fetch the model each session, which is fine for a first try but wasteful afterwards. Keep the .gguf on disk instead:

  1. Put the model URL in Settings → Backend Mode → True Offline, then click 💾 Save a copy to disk under the URL field. That downloads the same file through your browser's normal download flow, to a location you pick — the app itself still writes nothing.
  2. Next session, load that file with the file picker in the same panel. No transfer, no network at all — and it's the only way to work on an air-gapped machine.

(Downloading the .gguf straight from Hugging Face works identically; the save link just points at the same file.)

Browsers don't allow a file path to be pre-filled from a link, so it stays a manual pick each session, and you still pay the model's load time — just not the download. The table below puts that at ≈6 s for Qwen3-0.6B up to ≈80 s for gpt-oss-20b, measured over a local HTTP server, so a disk pick is if anything quicker.

📊 Benchmarks

A Playwright harness drives the unmodified dist/hermit-ui-wllama.html — via #gguf= and the app's own buttons, exactly as a user would — and has every model answer the same 10 questions, scored for correctness by hand. 16 threads, RTX 5070 Ti, Edge (WebGPU), 3+ runs per model on a verified-idle GPU (the harness refuses to start otherwise).

ModelSizeLoadavg TTFTdecode t/send-to-end t/s
Qwen3-0.6B0.4 GB5.7s0.66s7956
Qwen3-1.7B1.1 GB11.3s0.91s7354
Qwen3-4B2.5 GB21.4s1.03s6443
Qwen3-8B5.0 GB41.5s1.17s5535
Gemma-4-E2B3.1 GB26.5s9.3s406.1
Gemma-4-E4B5.0 GB36.6s11.7s325.6
Gemma-4-12B7.1 GB58.6s15.2s364.0
gpt-oss-20b †12.1 GB~80s3.4s4315

Load = engine init + model transfer from a local HTTP server (not your Hugging Face download time). Decode = pure generation speed with TTFT excluded; end-to-end = what the app's own stats readout shows, prompt processing included.

† gpt-oss-20b is MXFP4, not Q4_K_M, and it is a reasoning model. Its hidden thinking trace inflates the end-to-end denominator — one terse-but-correct logic answer read 2.9 t/s end-to-end while decoding at 40.5 t/s — so that column is not comparable to the rest. Judge it on the decode column, which is excellent — and the answers are genuinely good. It is the strongest large rung here if you have the VRAM to spare.

CPU-only (WebGPU off, same machine): Qwen3-0.6B ≈ 16 t/s, Qwen3-1.7B ≈ 10 t/s. Everything larger is unusable — Qwen3-4B ≈ 3 t/s, Gemma-4-E2B ≈ 1.6 t/s.

What the numbers say:

  • Qwen3 is the sweet spot in a browser. Even 8B stays interactive on WebGPU, and TTFT is ~1 s across the whole family.
  • Gemma-4 decodes fine but prompt-processes slowly under wllama — 9–15 s before the first token drags the end-to-end figure into single digits even though tokens then arrive at 30–40 t/s. All its answers were correct: an engine-side prompt-eval gap, not a model-quality one.
  • 12.1 GB runs, and runs well — and 13.0 GB does too. Qwen3.8-27B at 4 bpw loads and holds ~8 t/s, so the practical ceiling is your VRAM rather than any engine limit. gpt-oss-20b out-decodes the 7.1 GB Gemma-4-12B on a 16 GB card — a sparse MoE activating only ~3.6B params per token beats a dense model despite being 70% larger on disk. Architecture predicts throughput, not file size. WASM Memory64 (Chrome/Edge) is required: without it, anything above ~4 GB fails outright.
  • The GPU matters more than the model. 0.6B → 8B costs ~30% of decode throughput; dropping to CPU costs ~80%. Free VRAM matters most of all — a contended GPU understated these same runs by 23×, so if your numbers look nothing like these, check what else is on your card first.

Reproduce it yourself — one pip install, no Node. Each run writes review.md with the timings and every answer in full, so quality is reviewable and not just asserted, plus a machine-readable run.json.

Qwen3.8-27B — verified working in the browser

Qwen3.8-27B at 4 bpw runs in the web build. At 13.0 GB it is the largest model verified here, ahead of gpt-oss-20b.

Modelbyteshape/Qwen3.8-27B-GGUF → Qwen3.8-27B-IQ4_XS-4.00bpw.gguf
Size / quant13.0 GB, IQ4_XS (4.00 bpw)
Architectureqwen35 — hybrid attention + SSM, 65 layers, 248k vocab
Load time18–23 s from a local file (engine init + weights)
VRAM14.2–14.9 GB of a 16 GB card, at n_ctx 8192
Speed (app readout)≈8 tok/s — 7.7–8.5 across seven replies of 550–920 tokens

The speed figure is exactly what the app's own live tokens/s readout shows, i.e. end-to-end with prompt processing included — the same column as end-to-end t/s in the table above, not a decode-only number. Short replies read lower (≈4–6 t/s at 50–170 tokens, and ≈2.5 t/s on the very first reply after a load) purely because a fixed TTFT dominates a small denominator. Judge it on the longer generations.

Two caveats worth stating plainly:

  • Verified via the file picker, not #gguf=. A 13 GB one-click link was not tested, and at this size the download-once-and-re-pick flow is the sensible route anyway. Grab the .gguf from the repo above, then load it with the file picker.
  • Chrome/Edge only, and it needs the VRAM. WASM Memory64 is required (as for every rung above ~4 GB), and 13 GB of weights leaves little headroom on a 16 GB card — keep n_ctx modest (8192 was used here; this model's KV runs ~64 KB per token).

It is a reasoning model that thinks at xhigh by default, which is slow enough to matter on a 27B: use the 🧠 Think control in the chatbox to drop to low or Off for routine questions. Measured on three trivial prompts, Off finished in 24.9 s with zero reasoning versus 57.7 s and ~935 characters of reasoning at the default — see Flexible thinking control.

This rung was verified by hand rather than through the benchmark harness, so it has no avg TTFT / decode t/s entry in the table above.

For comparison, the same file under native llama-server on the same GPU decodes at 77–101 t/s — about 10× faster. The browser build's appeal is that it needs no install and nothing leaves the tab, not throughput.

Browser support & model size limits

How large a model you can load — and how fast it runs — depends on two WebAssembly/GPU features of your browser, which wllama detects at load time:

CapabilityWhat it enablesChrome / EdgeFirefoxSafari
JSPI (WebAssembly.Suspending)Streams the GGUF straight into the engine instead of copying it whole into the WASM heap → model size limited only by your RAM/VRAM✅ Chrome 137+⚠️ 153+ only❌ none → ~3 GB cap
WebGPU (in workers)Hardware-accelerated inference✅ mature⚠️ new / may fail to initialize → CPU fallback⚠️ present in recent versions, untested here

HermitUI probes both at load time (by capability, not user-agent) and warns you in the panel before you start a download that can't succeed.

In practice:

  • Chrome / Edge: Multi-GB models (7B+ quants) load and run fine, with WebGPU acceleration. The limit is your actual RAM/VRAM.
  • Firefox before 153: Without JSPI, wllama falls back to copying the entire model file into the 4 GiB WASM heap. Models larger than roughly 3 GB fail with the cryptic error source array is too long (an unchecked allocation failure inside wllama). Fix: update to Firefox 153+, which enables JSPI by default. You can verify support by typing !!WebAssembly.Suspending into the DevTools console — it must print true.
  • Firefox speed: Even with JSPI, Firefox's WebGPU support is much newer than Chrome's and may not initialize inside the wllama worker, dropping inference to single-threaded CPU WASM — noticeably slower than Chrome on the same machine. Check the debug console (verbosity Debug, then reload the model) to see whether a WebGPU device or the CPU backend was picked. If WebGPU misbehaves, try unchecking the WebGPU toggle — a clean CPU run can beat a broken GPU path.
  • Safari: no JSPI at all, so the same ~3 GB ceiling applies with no version to upgrade to — Qwen3-0.6B and 1.7B are fine, Qwen3-4B is the realistic top rung, and anything larger fails. The benchmarks above were not run on Safari. The rest of HermitUI (the normal client mode) works everywhere.

🔌 Connect to your own endpoint

Prefer to run the model outside the browser? The standalone build (index.html / dist/hermit-ui-standalone.html) talks to anything that speaks the OpenAI chat completions API.

  1. Start your local AI server — LM Studio, Ollama, llama.cpp, or vLLM.
  2. Open HermitUI — double-click index.html; it opens in any modern browser.
  3. Configure — click ⚙️ Settings in the top right to set the API URL, model name, API key, or system prompt.
Configuration examples for popular servers

LM Studio (the default)

  1. Launch LM Studio and start the Local Server.
  2. API URL: http://localhost:1234/v1/chat/completions
  3. Model Name: leave blank, or set to the specific model identifier you loaded.
  4. Tip: ensure CORS is enabled in the LM Studio settings.

Ollama

  1. Start your Ollama server from the terminal, making sure to enable CORS:
    OLLAMA_ORIGINS="*" ollama serve
    
  2. API URL: http://localhost:11434/v1/chat/completions
  3. Model Name: the name of the model you pulled (e.g., llama3, mistral, deepseek-coder).

vLLM

  1. Start your vLLM server — it serves an OpenAI-compatible endpoint and already allows every origin by default:
    vllm serve meta-llama/Llama-3.1-8B-Instruct
    
    To pin the allowlist instead, pass --allowed-origins a JSON array: --allowed-origins '["https://example.com"]'.
  2. API URL: http://localhost:8000/v1/chat/completions
  3. Model Name: the model you served (e.g., meta-llama/Llama-3.1-8B-Instruct).

llama.cpp (llama-server)

  1. Start llama-server with a GGUF. --jinja is the important flag — it applies the model's own chat template, which is what makes the 🧠 Think control work:
    llama-server --model model.gguf --jinja --host 0.0.0.0 --port 8080
    
  2. API URL: http://localhost:8080/v1/chat/completions
  3. Model Name: anything — llama-server serves whichever model it loaded.
  4. Tip: add --reasoning-format deepseek for a reasoning model. It returns the trace in reasoning_content, which HermitUI renders in the collapsible think block. Settings → Check Reasoning Support reads this server's /props and will report reasoning support exactly.

Reference config: a 27B on a 16 GB card, measured

If you have a 16 GB GPU, this is a config that fits and what each flag costs. Qwen3.8-27B at 4 bpw, one RTX 5070 Ti, using the model's own MTP layer for speculative decoding. Nothing here is card-specific except the numbers — the same reasoning applies to any 16 GB card.

llama-server --model Qwen3.8-27B-IQ4_XS-4.00bpw.gguf \
  --ctx-size 40000 --jinja --reasoning-format deepseek --reasoning-preserve \
  --flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
  --spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
  --cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \
  --n-gpu-layers all --parallel 1 --threads 32 --batch-size 1024 --ubatch-size 1024 \
  --temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0 \
  --host 0.0.0.0 --port 8080

Speed (llama-server's own timings, so decode-only — TTFT excluded):

WorkloadWith --spec-typeWithoutGain
Prose (200-word explainer, ×3)77.1 t/s50.4 t/s1.53×
Reasoning prompt91–96 t/s—~1.85×
Code generation100.7 t/s50.2 t/s2.01×

Prompt processing reached 1,489 t/s on a 6,742-token prompt. The no-speculation baseline was flat to within 0.3 t/s across four runs, so those gains are signal rather than noise.

Draft acceptance explains the spread — 98.7% on code, 82–88% on the reasoning prompt, 50–65% on prose. MTP predicts structured text far better than free prose, so the speedup is largest exactly where you generate code.

⚠️ --parallel 1 is not optional — it is worth 3.6×. Omit it and the server defaults to 4 slots, which flips llama.cpp to a unified KV cache (kv_unified = 'true') sized across all of them, so every decoded token does far more attention work. Measured on this machine, same GPU, same everything else: 20.4 t/s prose / 29.2 t/s code with 4 slots, versus 73.8 / 97.1 with one. VRAM barely moves (~15.5 GB either way), so nothing looks wrong — you just silently lose three quarters of your throughput. Check total_slots at http://localhost:8080/props; it must read 1.

Where the VRAM goes, at --ctx-size 40000 — as nvidia-smi and the server's own load log report it:

VRAM
Model weights on GPU (4 bpw)12,376 MiB
Compute buffers (--ubatch-size 1024)396 MiB
KV cache @ 40k q8_0, plus CUDA context overhead≈2,260 MiB
MTP draft context (what --spec-type costs)≈860 MiB
Total resident15,893 of 16,303 MiB

That leaves roughly 410 MiB spare — it runs, but with no room for anything else on the card. Don't run a WebGPU browser (including HermitUI's own in-browser mode) against the same GPU at this context size.

If it doesn't fit on your card, cut in this order:

  1. Lower --ctx-size. The KV cache scales linearly with it, so halving 40k → 20k returns roughly 1,000 MiB. This is the cheapest VRAM you will find.
  2. Keep --cache-type-k q8_0 --cache-type-v q8_0 (and --flash-attn on, which the quantized-KV path wants). q8_0 stores one byte per element where f16 stores two, so reverting to f16 would roughly double the KV figure above — at 40k context that is the difference between fitting and not.
  3. Drop --spec-type last. It returns ~860 MiB but costs the 1.5–2.0× decode speedup above, so trade it away only once the context is as small as you can live with.

For scale: this is roughly 10× the ~8 t/s the same model reaches in the in-browser build. Native is far faster — the browser build buys zero-install and privacy, not speed.

Cloud models (OpenRouter, OpenAI, Groq, …)

  1. API URL: the provider's chat completions endpoint (e.g., https://openrouter.ai/api/v1/chat/completions).
  2. Model Name: the model you want (e.g., anthropic/claude-opus-4.8) — check your provider's model list for the exact slug.
  3. API Key: enter your provider's key in the settings menu.

[!WARNING] Privacy note: using cloud models is generally not advised if you require strict privacy. Your data leaves your machine, and it is unclear how these providers handle, store, or train on it. For true ephemerality, stick to local models.

Troubleshooting (CORS)

If HermitUI fails to connect to your local AI server (e.g., a "Network Error"), it is most likely CORS. Because HermitUI runs as a local file (file://), browsers block its requests to http://localhost unless the server explicitly allows it.

  • LM Studio: "Local Server" tab → find the CORS toggle → turn it ON.
  • Ollama: set OLLAMA_ORIGINS before starting, e.g. OLLAMA_ORIGINS="*" ollama serve.
  • vLLM: allows all origins out of the box — if you narrowed it, widen --allowed-origins (it takes a JSON array, e.g. '["*"]').

✨ Features

  • 📦 Zero-dependency setup: all external libraries (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid) and the Inter font are bundled directly into the file. No installation, no build step. (A CDN-linked developer version lives in dist/hermit-ui-cdn.html.)
  • 🔒 Privacy first & ephemeral: no localStorage, IndexedDB, or cookies — nothing survives the tab.
  • 🧠 Thinking-model support: built-in parser formats <think>, <thought>, and <reasoning> tags as they stream from reasoning models.
  • 🎚️ Flexible thinking control: a 🧠 Think selector beside the persona picker sets reasoning depth — Off / Low / Medium / High — for models that support it. Off genuinely stops the reasoning trace rather than shortening it, which is the difference between a 25 s and a 58 s answer on a 27B. It only appears where it can actually do something: for a local GGUF the model's own chat template is scanned for the reasoning variables, and for a remote endpoint Settings → Check Reasoning Support reads llama.cpp's /props capability flags (or Ollama's template). Levels are validated against the set the template accepts, so a value it would reject is never sent.
  • ⚡ Real-time streaming with 📊 live performance stats — prompt tokens, completion tokens, tokens/second, and total duration.
  • 🖼️ Image & vision support: upload, drag-and-drop, or paste (Ctrl+V) images for vision-capable models, sent as image_url content per the OpenAI schema with automatic vision-model detection.
  • 📝 Rich rendering: Markdown with per-block copy buttons, syntax highlighting, 🧮 LaTeX math ($…$, $$…$$, \(…\), \[…\]) rendered via KaTeX to native MathML — no webfonts, works mid-stream and offline — and 📈 Mermaid diagrams from ```mermaid fences.
  • 🎭 Personas: switch between preset system prompts (technical, general, writing, tutor) on the fly.
  • ✏️ Edit & regenerate any previous message without restarting the conversation.
More features
  • 📎 Context attachments: drag-and-drop or upload text files to inject their contents into your prompt.
  • 🎛️ Advanced sampling controls: temperature, max_tokens, top_p, presence_penalty, frequency_penalty, and seed from a collapsible Settings panel. Params are only sent when set, keeping payloads compatible with minimal backends.
  • 🎨 Modern UI/UX: clean, responsive design with smooth micro-animations, comprehensive CSS variables for theming, and a glassmorphism feel. Light and dark themes.
  • 💾 Chat export & import: download the entire conversation as a formatted Markdown file (📤, Ctrl+E), and load one back in later with 📂 (Ctrl+I) — messages, system prompt and attached context are restored, and the model picks the conversation up where you left it. The exported file is the only copy: nothing is stored in the browser. Images can't come back, since the export records them as [N images attached].
  • ⚙️ Customizable settings: API URL, model name, API key, and system prompt via the on-page settings overlay.

🔗 Configuration via URL

You can pre-configure HermitUI through the URL fragment (the part after #), so a single link or bookmark carries the whole connection setup:

hermit-ui-standalone.html#api=http://localhost:8080/v1&model=qwen3-8b
hermit-ui-standalone.html#api=https://api.groq.com/openai/v1&key=gsk_...&model=llama-3.3-70b
hermit-ui-wllama.html#gguf=hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
ParameterEffect
apiAPI base URL (same as the Settings field)
modelModel name
keyAPI key
personaPreset persona: technical, general, writing, or tutor
gguf(wllama build only) GGUF model to load in-browser — direct URL, Hugging Face link, or hf:user/repo/file.gguf shorthand. Shows a one-click confirmation banner before downloading.

Why the fragment and not ?query: the part after # never leaves your browser — it is not sent in any HTTP request — and nothing is stored, so this stays true to the ephemerality promise (the URL is the config; refresh keeps your setup). Applied settings are always announced in a toast, so a shared link can't reconfigure the app invisibly. Free-text system prompts are deliberately not supported as a parameter, since a link could smuggle a malicious prompt.

[!NOTE] A key in the URL is never transmitted, but it does end up in your browser history (and any bookmark you save). Prefer entering keys in Settings on shared machines.

🎯 Ideal use cases

  • Heavily regulated environments: enterprise or government networks where software installation is restricted, but a secure local or remote inference endpoint is accessible.
  • Air-gapped systems: distribute on a USB stick and run on disconnected machines — either against a local network LLM server, or with the wllama build and a .gguf, against nothing at all.
  • Ephemeral kiosks & shared terminals: no chat history is ever written, making it safe for public workstations and desk-sharing environments.

🏗️ Architecture & philosophy

HermitUI enforces strict architectural constraints to remain lightweight and accessible:

  • Single file constraint: the final product is always a single, standalone .html file. The src/ directory is a blueprint only — its split into index.html, style.css, and script.js exists for maintainability, and build.py assembles them back into one file.
  • Vanilla only: no React, Vue, Angular, or other frontend frameworks.
  • No build tools: no package.json, npm, Webpack, or Vite.
  • No CSS frameworks: pure vanilla CSS, no Tailwind or Bootstrap.
  • Security: all rendered AI responses are sanitized with DOMPurify to prevent XSS.

The live build runs on GitHub Pages at moooff.github.io/HermitUI.

Verify the privacy claim

"Stores nothing" is the whole pitch, so don't take it on trust — checking takes about a minute.

  • At runtime: open DevTools → Application → Storage and use the app normally. Local Storage, Session Storage, IndexedDB, Cookies and Cache Storage stay empty for the entire session, including after a model has loaded. This is the authoritative check: it covers the bundled libraries and the embedded wllama engine, not just HermitUI's own code.
  • In the source: grep the file you downloaded.
    grep -o "localStorage\|sessionStorage\|indexedDB\|document\.cookie" dist/hermit-ui-standalone.html | wc -l
    
    The answer is 2, and both are false positives — Highlight.js's list of JavaScript keywords contains the strings localStorage and sessionStorage. There is not one call site. (In the wllama build the engine ships gzipped, so grep sees the app but not the engine; use the runtime check above to cover it.)
  • On the network: the Network tab shows requests only to endpoints you configured yourself — your API server, or the model URL if you chose to load a GGUF that way. No analytics, no phone-home, and no CDN at runtime in the standalone builds.

Model weights get the same treatment. A downloaded GGUF is streamed into an in-memory Blob instead of going through wllama's own URL loader, specifically because that one would persist the model to OPFS — see downloadGgufToBlob in src/script.js.

📦 Building & development

The root index.html (a copy of dist/hermit-ui-standalone.html) is a completely offline, standalone build: web fonts and images are base64-encoded and external JS/CSS libraries are injected directly into the file, which is what makes it work in air-gapped environments.

To modify it, edit the modular sources in src/ — index.html, style.css, and script.js, which reference libraries via CDN for convenient local development — then run:

python build.py        # or python3 build.py

Prerequisites: Python 3 and, on the first run, an internet connection. That is the whole list — build.py uses only the standard library, so there is no pip install, no package.json, and no Node.

The network requirement is worth spelling out, since it cuts against the rest of the project: the output runs offline, but producing it does not. On the first build, build.py downloads the pinned library versions (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid, the Inter font, and the wllama engine — ~14 MB total) into libs/, verifying each one against the SRI hash pinned in src/index.html. Everything after that is cached and offline; pass --refresh to force a re-download. So if you need to build on an air-gapped machine, copy a populated libs/ directory across with the repo.

This generates the standalone build at dist/hermit-ui-standalone.html, copies it to the root index.html for GitHub Pages, and creates the alternative builds in dist/. The standalone, CDN, and wllama variants (dist/hermit-ui-standalone.html, dist/hermit-ui-cdn.html, dist/hermit-ui-wllama.html) are committed so they are browsable and downloadable straight from GitHub; the local variant dist/hermit-ui-local.html and the downloaded libs/ are generated-only and stay gitignored.

🤝 Contributing

Bug reports, questions and ideas are welcome in Issues; pull requests are welcome too. A bug report travels much further with your browser and version, whether WebGPU was on, and — for in-browser inference — the model and the output of the debug console at verbosity Debug.

Before opening a PR, two things will save you a rewrite:

  • Read AGENTS.md. It is the single source of truth for this project's rules, and they are unusually strict on purpose: single-file output, vanilla JS only, no build tools or frameworks or CSS libraries, no localStorage/IndexedDB/cookies for any reason, and every AI-rendered string sanitized through DOMPurify. A change that breaks one of these can't be merged no matter how good it is — the constraints are the product.
  • Edit src/, never the generated files. dist/*.html and the root index.html are build artifacts and get overwritten; run python build.py and commit the regenerated outputs along with your source change.

There is no test suite or linter. Verify a change by opening the rebuilt dist/hermit-ui-standalone.html (or dist/hermit-ui-wllama.html) in a browser and exercising the affected path by hand. If your change touches inference performance, the benchmark harness produces numbers that can go straight into a PR description.

🛠️ Built with

  • Vanilla HTML5 / CSS3 / ES6 JavaScript
  • wllama — llama.cpp in WebAssembly, for in-browser inference
  • Marked.js — Markdown parsing
  • DOMPurify — HTML sanitization / XSS prevention
  • Highlight.js — code syntax highlighting
  • KaTeX — LaTeX math rendering (MathML output)
  • Mermaid — diagram rendering from ```mermaid fences
  • Google Fonts (Inter) — typography

🗺️ Roadmap

  • Split-GGUF support: load sharded models (-00001-of-000NN.gguf) through the in-browser URL loader.
  • Save the in-flight download: write the model to disk as it downloads (File System Access API), so keeping a copy costs one transfer instead of two. Chrome/Edge only, hence the simpler two-download link today.
  • Companion text-analysis app: a separate ultra-light single-file build (Transformers.js + WebGPU, sub-200 MB models) dedicated to summarization and text analysis.

The full list of ideas and tasks lives in docs/backlog.md.

📄 License

This project is open-source and available under the terms of the GNU AGPL v3. See the included LICENSE file for the full text.

air-gapped
air-gapped-ai
chat-ui
local-llm
openai-api
privacy-first
single-file
zero-telemetry

Languages

HTML

90.1%

JavaScript

6.6%

Python

2.1%

CSS

1.3%