Keep your AI conversations private. A zero-dependency, single-file local chat UI that connects to any AI endpoint.
HTML
7
145 commits
updated Sep 28, 2026
A lightweight, modern, and ephemeral single-page web interface for local AI models.
🌐 Online Demo • Try it • In-Browser AI • Benchmarks • Connect a server • Features
In-browser inference with WebGPU acceleration: the model downloads straight into memory (timelapsed) and the answer streams in real time.
Left: connect it to your local AI server. Right: one click downloads a small but capable model (Qwen3-0.6B, ~380 MB) and chats fully inside your browser — no server at all.
HermitUI is a chat interface that is one .html file. No install, no server, no build step, no npm — double-click it and it opens.
Two things set it apart, and the combination is the point:
localStorage, no IndexedDB, no cookies, no model cache, no telemetry. Close the tab and the conversation and the model are gone — and you can check that in a minute.Or ignore all of that and point it at LM Studio, Ollama, llama.cpp, or vLLM as a normal client.
Built for the machines where nothing else fits: air-gapped boxes, locked-down corporate and government networks, shared kiosks and hot desks.
One click: open the 🧠 In-Browser AI Demo — it pre-fills Qwen3-0.6B (~380 MB) via the #gguf= hash parameter; confirm the banner and chat.
Or do it by hand:
dist/hermit-ui-wllama.html to disk (right-click → Save link as…, since GitHub serves raw .html as plain text), then open it in your browser.hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.ggufHermitUI can run true offline inference entirely in the browser — no local server or OpenAI-compatible endpoint required. It's powered by wllama (llama.cpp compiled to WebAssembly, with optional WebGPU acceleration): you load a .gguf model file and chat with it directly on the page.
This ships as a dedicated build output, dist/hermit-ui-wllama.html — the regular standalone app plus a Backend Mode switch in Settings (Remote / Local API ↔ True Offline (Wllama GGUF)). The main builds stay lean: the feature is stripped out of them at build time.
DecompressionStream API), so the ~6 MB file is complete on its own — perfect for USB-stick distribution to air-gapped machines. Pair it with a .gguf from disk and the whole stack is offline; only the optional download-by-URL path touches the network..gguf file from disk and run it fully client-side, with an optional WebGPU toggle for hardware acceleration. Download a model once, keep it, and re-pick it every session — no network involved, so this is also the fastest way to use HermitUI repeatedly (and the only way on an air-gapped machine)..gguf link, a Hugging Face /blob/ page URL (auto-rewritten to /resolve/), or the hf:user/repo/file.gguf shorthand, then hit Load. The model streams straight into memory with a live progress bar — true to the ephemerality promise, nothing is written to browser storage. A URL-loaded model is therefore fetched again next session — unless you hit 💾 Save a copy to disk under the URL field, which hands the same file to your browser's download flow so the file picker can load it from then on. A model can also be baked into a shareable link: hermit-ui-wllama.html#gguf=hf:user/repo/file.gguf (see Configuration via URL).n_ctx, default 32k — automatically halved until it fits in memory, with the effective size shown in the status line) and max output tokens per reply (default 4096); temperature, top-p, and seed from the regular settings apply too.tokenizer.chat_template when present, otherwise auto-detects a sane format from the model architecture (ChatML, Llama 3, Mistral, Gemma, Phi-3, Zephyr, Alpaca, …), with a manual override.Every link below opens the wllama build with that model pre-filled via #gguf= — confirm the banner and it streams straight into memory. Nothing is written to browser storage, so a link-loaded model is fetched again next session unless you keep a copy of the .gguf. Start small; the bigger rungs need a modern Chrome/Edge (see Browser support).
| Model | Download | Try it |
|---|---|---|
| Qwen3-0.6B | 0.4 GB | ▶ Run in browser |
| Qwen3-1.7B | 1.1 GB | ▶ Run in browser |
| Qwen3-4B ⭐ | 2.5 GB | ▶ Run in browser |
| Qwen3-8B | 5.0 GB | ▶ Run in browser |
| Gemma-4-E2B | 3.1 GB | ▶ Run in browser |
| Gemma-4-E4B | 5.0 GB | ▶ Run in browser |
| Gemma-4-12B | 7.1 GB | ▶ Run in browser |
| gpt-oss-20b ⚠️ | 12.1 GB | ▶ Run in browser |
⭐ = best speed/quality trade-off on a WebGPU machine. ⚠️ = the fastest model above 7 GB measured here (~43 t/s), but a 12 GB download that needs Chrome/Edge and a mostly-free GPU — see the note under the results table. Quants are Q4_K_M from unsloth except gpt-oss-20b, which is MXFP4 from ggml-org. On a CPU-only machine, use Qwen3-1.7B or smaller.
You only pay for the transfer once. The links above re-fetch the model each session, which is fine for a first try but wasteful afterwards. Keep the .gguf on disk instead:
(Downloading the .gguf straight from Hugging Face works identically; the save link just points at the same file.)
Browsers don't allow a file path to be pre-filled from a link, so it stays a manual pick each session, and you still pay the model's load time — just not the download. The table below puts that at ≈6 s for Qwen3-0.6B up to ≈80 s for gpt-oss-20b, measured over a local HTTP server, so a disk pick is if anything quicker.
A Playwright harness drives the unmodified dist/hermit-ui-wllama.html — via #gguf= and the app's own buttons, exactly as a user would — and has every model answer the same 10 questions, scored for correctness by hand. 16 threads, RTX 5070 Ti, Edge (WebGPU), 3+ runs per model on a verified-idle GPU (the harness refuses to start otherwise).
| Model | Size | Load | avg TTFT | decode t/s | end-to-end t/s |
|---|---|---|---|---|---|
| Qwen3-0.6B | 0.4 GB | 5.7s | 0.66s | 79 | 56 |
| Qwen3-1.7B | 1.1 GB | 11.3s | 0.91s | 73 | 54 |
| Qwen3-4B | 2.5 GB | 21.4s | 1.03s | 64 | 43 |
| Qwen3-8B | 5.0 GB | 41.5s | 1.17s | 55 | 35 |
| Gemma-4-E2B | 3.1 GB | 26.5s | 9.3s | 40 | 6.1 |
| Gemma-4-E4B | 5.0 GB | 36.6s | 11.7s | 32 | 5.6 |
| Gemma-4-12B | 7.1 GB | 58.6s | 15.2s | 36 | 4.0 |
| gpt-oss-20b † | 12.1 GB | ~80s | 3.4s | 43 | 15 |
Load = engine init + model transfer from a local HTTP server (not your Hugging Face download time). Decode = pure generation speed with TTFT excluded; end-to-end = what the app's own stats readout shows, prompt processing included.
† gpt-oss-20b is MXFP4, not Q4_K_M, and it is a reasoning model. Its hidden thinking trace inflates the end-to-end denominator — one terse-but-correct logic answer read 2.9 t/s end-to-end while decoding at 40.5 t/s — so that column is not comparable to the rest. Judge it on the decode column, which is excellent — and the answers are genuinely good. It is the strongest large rung here if you have the VRAM to spare.
CPU-only (WebGPU off, same machine): Qwen3-0.6B ≈ 16 t/s, Qwen3-1.7B ≈ 10 t/s. Everything larger is unusable — Qwen3-4B ≈ 3 t/s, Gemma-4-E2B ≈ 1.6 t/s.
What the numbers say:
Reproduce it yourself — one pip install, no Node. Each run writes review.md with the timings and every answer in full, so quality is reviewable and not just asserted, plus a machine-readable run.json.
Qwen3.8-27B at 4 bpw runs in the web build. At 13.0 GB it is the largest model verified here, ahead of gpt-oss-20b.
| Model | byteshape/Qwen3.8-27B-GGUF → Qwen3.8-27B-IQ4_XS-4.00bpw.gguf |
| Size / quant | 13.0 GB, IQ4_XS (4.00 bpw) |
| Architecture | qwen35 — hybrid attention + SSM, 65 layers, 248k vocab |
| Load time | 18–23 s from a local file (engine init + weights) |
| VRAM | 14.2–14.9 GB of a 16 GB card, at n_ctx 8192 |
| Speed (app readout) | ≈8 tok/s — 7.7–8.5 across seven replies of 550–920 tokens |
The speed figure is exactly what the app's own live tokens/s readout shows, i.e. end-to-end with prompt processing included — the same column as end-to-end t/s in the table above, not a decode-only number. Short replies read lower (≈4–6 t/s at 50–170 tokens, and ≈2.5 t/s on the very first reply after a load) purely because a fixed TTFT dominates a small denominator. Judge it on the longer generations.
Two caveats worth stating plainly:
#gguf=. A 13 GB one-click link was not tested, and at this size the download-once-and-re-pick flow is the sensible route anyway. Grab the .gguf from the repo above, then load it with the file picker.n_ctx modest (8192 was used here; this model's KV runs ~64 KB per token).It is a reasoning model that thinks at xhigh by default, which is slow enough to matter on a 27B: use the 🧠 Think control in the chatbox to drop to low or Off for routine questions. Measured on three trivial prompts, Off finished in 24.9 s with zero reasoning versus 57.7 s and ~935 characters of reasoning at the default — see Flexible thinking control.
This rung was verified by hand rather than through the benchmark harness, so it has no avg TTFT / decode t/s entry in the table above.
For comparison, the same file under native llama-server on the same GPU decodes at 77–101 t/s — about 10× faster. The browser build's appeal is that it needs no install and nothing leaves the tab, not throughput.
How large a model you can load — and how fast it runs — depends on two WebAssembly/GPU features of your browser, which wllama detects at load time:
| Capability | What it enables | Chrome / Edge | Firefox | Safari |
|---|---|---|---|---|
JSPI (WebAssembly.Suspending) | Streams the GGUF straight into the engine instead of copying it whole into the WASM heap → model size limited only by your RAM/VRAM | ✅ Chrome 137+ | ⚠️ 153+ only | ❌ none → ~3 GB cap |
| WebGPU (in workers) | Hardware-accelerated inference | ✅ mature | ⚠️ new / may fail to initialize → CPU fallback | ⚠️ present in recent versions, untested here |
HermitUI probes both at load time (by capability, not user-agent) and warns you in the panel before you start a download that can't succeed.
In practice:
source array is too long (an unchecked allocation failure inside wllama). Fix: update to Firefox 153+, which enables JSPI by default. You can verify support by typing !!WebAssembly.Suspending into the DevTools console — it must print true.Prefer to run the model outside the browser? The standalone build (index.html / dist/hermit-ui-standalone.html) talks to anything that speaks the OpenAI chat completions API.
index.html; it opens in any modern browser.http://localhost:1234/v1/chat/completionsOLLAMA_ORIGINS="*" ollama serve
http://localhost:11434/v1/chat/completionsllama3, mistral, deepseek-coder).vllm serve meta-llama/Llama-3.1-8B-Instruct
To pin the allowlist instead, pass --allowed-origins a JSON array: --allowed-origins '["https://example.com"]'.http://localhost:8000/v1/chat/completionsmeta-llama/Llama-3.1-8B-Instruct).llama-server)llama-server with a GGUF. --jinja is the important flag — it applies the model's own chat template, which is what makes the 🧠 Think control work:
llama-server --model model.gguf --jinja --host 0.0.0.0 --port 8080
http://localhost:8080/v1/chat/completionsllama-server serves whichever model it loaded.--reasoning-format deepseek for a reasoning model. It returns the trace in reasoning_content, which HermitUI renders in the collapsible think block. Settings → Check Reasoning Support reads this server's /props and will report reasoning support exactly.If you have a 16 GB GPU, this is a config that fits and what each flag costs. Qwen3.8-27B at 4 bpw, one RTX 5070 Ti, using the model's own MTP layer for speculative decoding. Nothing here is card-specific except the numbers — the same reasoning applies to any 16 GB card.
llama-server --model Qwen3.8-27B-IQ4_XS-4.00bpw.gguf \
--ctx-size 40000 --jinja --reasoning-format deepseek --reasoning-preserve \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \
--n-gpu-layers all --parallel 1 --threads 32 --batch-size 1024 --ubatch-size 1024 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0 \
--host 0.0.0.0 --port 8080
Speed (llama-server's own timings, so decode-only — TTFT excluded):
| Workload | With --spec-type | Without | Gain |
|---|---|---|---|
| Prose (200-word explainer, ×3) | 77.1 t/s | 50.4 t/s | 1.53× |
| Reasoning prompt | 91–96 t/s | — | ~1.85× |
| Code generation | 100.7 t/s | 50.2 t/s | 2.01× |
Prompt processing reached 1,489 t/s on a 6,742-token prompt. The no-speculation baseline was flat to within 0.3 t/s across four runs, so those gains are signal rather than noise.
Draft acceptance explains the spread — 98.7% on code, 82–88% on the reasoning prompt, 50–65% on prose. MTP predicts structured text far better than free prose, so the speedup is largest exactly where you generate code.
⚠️ --parallel 1 is not optional — it is worth 3.6×. Omit it and the server defaults to 4 slots, which flips llama.cpp to a unified KV cache (kv_unified = 'true') sized across all of them, so every decoded token does far more attention work. Measured on this machine, same GPU, same everything else: 20.4 t/s prose / 29.2 t/s code with 4 slots, versus 73.8 / 97.1 with one. VRAM barely moves (~15.5 GB either way), so nothing looks wrong — you just silently lose three quarters of your throughput. Check total_slots at http://localhost:8080/props; it must read 1.
Where the VRAM goes, at --ctx-size 40000 — as nvidia-smi and the server's own load log report it:
| VRAM | |
|---|---|
| Model weights on GPU (4 bpw) | 12,376 MiB |
Compute buffers (--ubatch-size 1024) | 396 MiB |
KV cache @ 40k q8_0, plus CUDA context overhead | ≈2,260 MiB |
MTP draft context (what --spec-type costs) | ≈860 MiB |
| Total resident | 15,893 of 16,303 MiB |
That leaves roughly 410 MiB spare — it runs, but with no room for anything else on the card. Don't run a WebGPU browser (including HermitUI's own in-browser mode) against the same GPU at this context size.
If it doesn't fit on your card, cut in this order:
--ctx-size. The KV cache scales linearly with it, so halving 40k → 20k returns roughly 1,000 MiB. This is the cheapest VRAM you will find.--cache-type-k q8_0 --cache-type-v q8_0 (and --flash-attn on, which the quantized-KV path wants). q8_0 stores one byte per element where f16 stores two, so reverting to f16 would roughly double the KV figure above — at 40k context that is the difference between fitting and not.--spec-type last. It returns ~860 MiB but costs the 1.5–2.0× decode speedup above, so trade it away only once the context is as small as you can live with.For scale: this is roughly 10× the ~8 t/s the same model reaches in the in-browser build. Native is far faster — the browser build buys zero-install and privacy, not speed.
https://openrouter.ai/api/v1/chat/completions).anthropic/claude-opus-4.8) — check your provider's model list for the exact slug.[!WARNING] Privacy note: using cloud models is generally not advised if you require strict privacy. Your data leaves your machine, and it is unclear how these providers handle, store, or train on it. For true ephemerality, stick to local models.
If HermitUI fails to connect to your local AI server (e.g., a "Network Error"), it is most likely CORS. Because HermitUI runs as a local file (file://), browsers block its requests to http://localhost unless the server explicitly allows it.
OLLAMA_ORIGINS before starting, e.g. OLLAMA_ORIGINS="*" ollama serve.--allowed-origins (it takes a JSON array, e.g. '["*"]').dist/hermit-ui-cdn.html.)localStorage, IndexedDB, or cookies — nothing survives the tab.<think>, <thought>, and <reasoning> tags as they stream from reasoning models./props capability flags (or Ollama's template). Levels are validated against the set the template accepts, so a value it would reject is never sent.image_url content per the OpenAI schema with automatic vision-model detection.$…$, $$…$$, \(…\), \[…\]) rendered via KaTeX to native MathML — no webfonts, works mid-stream and offline — and 📈 Mermaid diagrams from ```mermaid fences.temperature, max_tokens, top_p, presence_penalty, frequency_penalty, and seed from a collapsible Settings panel. Params are only sent when set, keeping payloads compatible with minimal backends.Ctrl+E), and load one back in later with 📂 (Ctrl+I) — messages, system prompt and attached context are restored, and the model picks the conversation up where you left it. The exported file is the only copy: nothing is stored in the browser. Images can't come back, since the export records them as [N images attached].You can pre-configure HermitUI through the URL fragment (the part after #), so a single link or bookmark carries the whole connection setup:
hermit-ui-standalone.html#api=http://localhost:8080/v1&model=qwen3-8b
hermit-ui-standalone.html#api=https://api.groq.com/openai/v1&key=gsk_...&model=llama-3.3-70b
hermit-ui-wllama.html#gguf=hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
| Parameter | Effect |
|---|---|
api | API base URL (same as the Settings field) |
model | Model name |
key | API key |
persona | Preset persona: technical, general, writing, or tutor |
gguf | (wllama build only) GGUF model to load in-browser — direct URL, Hugging Face link, or hf:user/repo/file.gguf shorthand. Shows a one-click confirmation banner before downloading. |
Why the fragment and not ?query: the part after # never leaves your browser — it is not sent in any HTTP request — and nothing is stored, so this stays true to the ephemerality promise (the URL is the config; refresh keeps your setup). Applied settings are always announced in a toast, so a shared link can't reconfigure the app invisibly. Free-text system prompts are deliberately not supported as a parameter, since a link could smuggle a malicious prompt.
[!NOTE] A
keyin the URL is never transmitted, but it does end up in your browser history (and any bookmark you save). Prefer entering keys in Settings on shared machines.
.gguf, against nothing at all.HermitUI enforces strict architectural constraints to remain lightweight and accessible:
.html file. The src/ directory is a blueprint only — its split into index.html, style.css, and script.js exists for maintainability, and build.py assembles them back into one file.package.json, npm, Webpack, or Vite.DOMPurify to prevent XSS.The live build runs on GitHub Pages at moooff.github.io/HermitUI.
"Stores nothing" is the whole pitch, so don't take it on trust — checking takes about a minute.
grep -o "localStorage\|sessionStorage\|indexedDB\|document\.cookie" dist/hermit-ui-standalone.html | wc -l
The answer is 2, and both are false positives — Highlight.js's list of JavaScript keywords contains the strings localStorage and sessionStorage. There is not one call site. (In the wllama build the engine ships gzipped, so grep sees the app but not the engine; use the runtime check above to cover it.)Model weights get the same treatment. A downloaded GGUF is streamed into an in-memory Blob instead of going through wllama's own URL loader, specifically because that one would persist the model to OPFS — see downloadGgufToBlob in src/script.js.
The root index.html (a copy of dist/hermit-ui-standalone.html) is a completely offline, standalone build: web fonts and images are base64-encoded and external JS/CSS libraries are injected directly into the file, which is what makes it work in air-gapped environments.
To modify it, edit the modular sources in src/ — index.html, style.css, and script.js, which reference libraries via CDN for convenient local development — then run:
python build.py # or python3 build.py
Prerequisites: Python 3 and, on the first run, an internet connection. That is the whole list — build.py uses only the standard library, so there is no pip install, no package.json, and no Node.
The network requirement is worth spelling out, since it cuts against the rest of the project: the output runs offline, but producing it does not. On the first build, build.py downloads the pinned library versions (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid, the Inter font, and the wllama engine — ~14 MB total) into libs/, verifying each one against the SRI hash pinned in src/index.html. Everything after that is cached and offline; pass --refresh to force a re-download. So if you need to build on an air-gapped machine, copy a populated libs/ directory across with the repo.
This generates the standalone build at dist/hermit-ui-standalone.html, copies it to the root index.html for GitHub Pages, and creates the alternative builds in dist/. The standalone, CDN, and wllama variants (dist/hermit-ui-standalone.html, dist/hermit-ui-cdn.html, dist/hermit-ui-wllama.html) are committed so they are browsable and downloadable straight from GitHub; the local variant dist/hermit-ui-local.html and the downloaded libs/ are generated-only and stay gitignored.
Bug reports, questions and ideas are welcome in Issues; pull requests are welcome too. A bug report travels much further with your browser and version, whether WebGPU was on, and — for in-browser inference — the model and the output of the debug console at verbosity Debug.
Before opening a PR, two things will save you a rewrite:
AGENTS.md. It is the single source of truth for this project's rules, and they are unusually strict on purpose: single-file output, vanilla JS only, no build tools or frameworks or CSS libraries, no localStorage/IndexedDB/cookies for any reason, and every AI-rendered string sanitized through DOMPurify. A change that breaks one of these can't be merged no matter how good it is — the constraints are the product.src/, never the generated files. dist/*.html and the root index.html are build artifacts and get overwritten; run python build.py and commit the regenerated outputs along with your source change.There is no test suite or linter. Verify a change by opening the rebuilt dist/hermit-ui-standalone.html (or dist/hermit-ui-wllama.html) in a browser and exercising the affected path by hand. If your change touches inference performance, the benchmark harness produces numbers that can go straight into a PR description.
```mermaid fences-00001-of-000NN.gguf) through the in-browser URL loader.The full list of ideas and tasks lives in docs/backlog.md.
This project is open-source and available under the terms of the GNU AGPL v3. See the included LICENSE file for the full text.
HTML
90.1%
JavaScript
6.6%
Python
2.1%
CSS
1.3%
Keep your AI conversations private. A zero-dependency, single-file local chat UI that connects to any AI endpoint.
HTML
7
145 commits
updated Sep 28, 2026
A lightweight, modern, and ephemeral single-page web interface for local AI models.
🌐 Online Demo • Try it • In-Browser AI • Benchmarks • Connect a server • Features
In-browser inference with WebGPU acceleration: the model downloads straight into memory (timelapsed) and the answer streams in real time.
Left: connect it to your local AI server. Right: one click downloads a small but capable model (Qwen3-0.6B, ~380 MB) and chats fully inside your browser — no server at all.
HermitUI is a chat interface that is one .html file. No install, no server, no build step, no npm — double-click it and it opens.
Two things set it apart, and the combination is the point:
localStorage, no IndexedDB, no cookies, no model cache, no telemetry. Close the tab and the conversation and the model are gone — and you can check that in a minute.Or ignore all of that and point it at LM Studio, Ollama, llama.cpp, or vLLM as a normal client.
Built for the machines where nothing else fits: air-gapped boxes, locked-down corporate and government networks, shared kiosks and hot desks.
One click: open the 🧠 In-Browser AI Demo — it pre-fills Qwen3-0.6B (~380 MB) via the #gguf= hash parameter; confirm the banner and chat.
Or do it by hand:
dist/hermit-ui-wllama.html to disk (right-click → Save link as…, since GitHub serves raw .html as plain text), then open it in your browser.hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.ggufHermitUI can run true offline inference entirely in the browser — no local server or OpenAI-compatible endpoint required. It's powered by wllama (llama.cpp compiled to WebAssembly, with optional WebGPU acceleration): you load a .gguf model file and chat with it directly on the page.
This ships as a dedicated build output, dist/hermit-ui-wllama.html — the regular standalone app plus a Backend Mode switch in Settings (Remote / Local API ↔ True Offline (Wllama GGUF)). The main builds stay lean: the feature is stripped out of them at build time.
DecompressionStream API), so the ~6 MB file is complete on its own — perfect for USB-stick distribution to air-gapped machines. Pair it with a .gguf from disk and the whole stack is offline; only the optional download-by-URL path touches the network..gguf file from disk and run it fully client-side, with an optional WebGPU toggle for hardware acceleration. Download a model once, keep it, and re-pick it every session — no network involved, so this is also the fastest way to use HermitUI repeatedly (and the only way on an air-gapped machine)..gguf link, a Hugging Face /blob/ page URL (auto-rewritten to /resolve/), or the hf:user/repo/file.gguf shorthand, then hit Load. The model streams straight into memory with a live progress bar — true to the ephemerality promise, nothing is written to browser storage. A URL-loaded model is therefore fetched again next session — unless you hit 💾 Save a copy to disk under the URL field, which hands the same file to your browser's download flow so the file picker can load it from then on. A model can also be baked into a shareable link: hermit-ui-wllama.html#gguf=hf:user/repo/file.gguf (see Configuration via URL).n_ctx, default 32k — automatically halved until it fits in memory, with the effective size shown in the status line) and max output tokens per reply (default 4096); temperature, top-p, and seed from the regular settings apply too.tokenizer.chat_template when present, otherwise auto-detects a sane format from the model architecture (ChatML, Llama 3, Mistral, Gemma, Phi-3, Zephyr, Alpaca, …), with a manual override.Every link below opens the wllama build with that model pre-filled via #gguf= — confirm the banner and it streams straight into memory. Nothing is written to browser storage, so a link-loaded model is fetched again next session unless you keep a copy of the .gguf. Start small; the bigger rungs need a modern Chrome/Edge (see Browser support).
| Model | Download | Try it |
|---|---|---|
| Qwen3-0.6B | 0.4 GB | ▶ Run in browser |
| Qwen3-1.7B | 1.1 GB | ▶ Run in browser |
| Qwen3-4B ⭐ | 2.5 GB | ▶ Run in browser |
| Qwen3-8B | 5.0 GB | ▶ Run in browser |
| Gemma-4-E2B | 3.1 GB | ▶ Run in browser |
| Gemma-4-E4B | 5.0 GB | ▶ Run in browser |
| Gemma-4-12B | 7.1 GB | ▶ Run in browser |
| gpt-oss-20b ⚠️ | 12.1 GB | ▶ Run in browser |
⭐ = best speed/quality trade-off on a WebGPU machine. ⚠️ = the fastest model above 7 GB measured here (~43 t/s), but a 12 GB download that needs Chrome/Edge and a mostly-free GPU — see the note under the results table. Quants are Q4_K_M from unsloth except gpt-oss-20b, which is MXFP4 from ggml-org. On a CPU-only machine, use Qwen3-1.7B or smaller.
You only pay for the transfer once. The links above re-fetch the model each session, which is fine for a first try but wasteful afterwards. Keep the .gguf on disk instead:
(Downloading the .gguf straight from Hugging Face works identically; the save link just points at the same file.)
Browsers don't allow a file path to be pre-filled from a link, so it stays a manual pick each session, and you still pay the model's load time — just not the download. The table below puts that at ≈6 s for Qwen3-0.6B up to ≈80 s for gpt-oss-20b, measured over a local HTTP server, so a disk pick is if anything quicker.
A Playwright harness drives the unmodified dist/hermit-ui-wllama.html — via #gguf= and the app's own buttons, exactly as a user would — and has every model answer the same 10 questions, scored for correctness by hand. 16 threads, RTX 5070 Ti, Edge (WebGPU), 3+ runs per model on a verified-idle GPU (the harness refuses to start otherwise).
| Model | Size | Load | avg TTFT | decode t/s | end-to-end t/s |
|---|---|---|---|---|---|
| Qwen3-0.6B | 0.4 GB | 5.7s | 0.66s | 79 | 56 |
| Qwen3-1.7B | 1.1 GB | 11.3s | 0.91s | 73 | 54 |
| Qwen3-4B | 2.5 GB | 21.4s | 1.03s | 64 | 43 |
| Qwen3-8B | 5.0 GB | 41.5s | 1.17s | 55 | 35 |
| Gemma-4-E2B | 3.1 GB | 26.5s | 9.3s | 40 | 6.1 |
| Gemma-4-E4B | 5.0 GB | 36.6s | 11.7s | 32 | 5.6 |
| Gemma-4-12B | 7.1 GB | 58.6s | 15.2s | 36 | 4.0 |
| gpt-oss-20b † | 12.1 GB | ~80s | 3.4s | 43 | 15 |
Load = engine init + model transfer from a local HTTP server (not your Hugging Face download time). Decode = pure generation speed with TTFT excluded; end-to-end = what the app's own stats readout shows, prompt processing included.
† gpt-oss-20b is MXFP4, not Q4_K_M, and it is a reasoning model. Its hidden thinking trace inflates the end-to-end denominator — one terse-but-correct logic answer read 2.9 t/s end-to-end while decoding at 40.5 t/s — so that column is not comparable to the rest. Judge it on the decode column, which is excellent — and the answers are genuinely good. It is the strongest large rung here if you have the VRAM to spare.
CPU-only (WebGPU off, same machine): Qwen3-0.6B ≈ 16 t/s, Qwen3-1.7B ≈ 10 t/s. Everything larger is unusable — Qwen3-4B ≈ 3 t/s, Gemma-4-E2B ≈ 1.6 t/s.
What the numbers say:
Reproduce it yourself — one pip install, no Node. Each run writes review.md with the timings and every answer in full, so quality is reviewable and not just asserted, plus a machine-readable run.json.
Qwen3.8-27B at 4 bpw runs in the web build. At 13.0 GB it is the largest model verified here, ahead of gpt-oss-20b.
| Model | byteshape/Qwen3.8-27B-GGUF → Qwen3.8-27B-IQ4_XS-4.00bpw.gguf |
| Size / quant | 13.0 GB, IQ4_XS (4.00 bpw) |
| Architecture | qwen35 — hybrid attention + SSM, 65 layers, 248k vocab |
| Load time | 18–23 s from a local file (engine init + weights) |
| VRAM | 14.2–14.9 GB of a 16 GB card, at n_ctx 8192 |
| Speed (app readout) | ≈8 tok/s — 7.7–8.5 across seven replies of 550–920 tokens |
The speed figure is exactly what the app's own live tokens/s readout shows, i.e. end-to-end with prompt processing included — the same column as end-to-end t/s in the table above, not a decode-only number. Short replies read lower (≈4–6 t/s at 50–170 tokens, and ≈2.5 t/s on the very first reply after a load) purely because a fixed TTFT dominates a small denominator. Judge it on the longer generations.
Two caveats worth stating plainly:
#gguf=. A 13 GB one-click link was not tested, and at this size the download-once-and-re-pick flow is the sensible route anyway. Grab the .gguf from the repo above, then load it with the file picker.n_ctx modest (8192 was used here; this model's KV runs ~64 KB per token).It is a reasoning model that thinks at xhigh by default, which is slow enough to matter on a 27B: use the 🧠 Think control in the chatbox to drop to low or Off for routine questions. Measured on three trivial prompts, Off finished in 24.9 s with zero reasoning versus 57.7 s and ~935 characters of reasoning at the default — see Flexible thinking control.
This rung was verified by hand rather than through the benchmark harness, so it has no avg TTFT / decode t/s entry in the table above.
For comparison, the same file under native llama-server on the same GPU decodes at 77–101 t/s — about 10× faster. The browser build's appeal is that it needs no install and nothing leaves the tab, not throughput.
How large a model you can load — and how fast it runs — depends on two WebAssembly/GPU features of your browser, which wllama detects at load time:
| Capability | What it enables | Chrome / Edge | Firefox | Safari |
|---|---|---|---|---|
JSPI (WebAssembly.Suspending) | Streams the GGUF straight into the engine instead of copying it whole into the WASM heap → model size limited only by your RAM/VRAM | ✅ Chrome 137+ | ⚠️ 153+ only | ❌ none → ~3 GB cap |
| WebGPU (in workers) | Hardware-accelerated inference | ✅ mature | ⚠️ new / may fail to initialize → CPU fallback | ⚠️ present in recent versions, untested here |
HermitUI probes both at load time (by capability, not user-agent) and warns you in the panel before you start a download that can't succeed.
In practice:
source array is too long (an unchecked allocation failure inside wllama). Fix: update to Firefox 153+, which enables JSPI by default. You can verify support by typing !!WebAssembly.Suspending into the DevTools console — it must print true.Prefer to run the model outside the browser? The standalone build (index.html / dist/hermit-ui-standalone.html) talks to anything that speaks the OpenAI chat completions API.
index.html; it opens in any modern browser.http://localhost:1234/v1/chat/completionsOLLAMA_ORIGINS="*" ollama serve
http://localhost:11434/v1/chat/completionsllama3, mistral, deepseek-coder).vllm serve meta-llama/Llama-3.1-8B-Instruct
To pin the allowlist instead, pass --allowed-origins a JSON array: --allowed-origins '["https://example.com"]'.http://localhost:8000/v1/chat/completionsmeta-llama/Llama-3.1-8B-Instruct).llama-server)llama-server with a GGUF. --jinja is the important flag — it applies the model's own chat template, which is what makes the 🧠 Think control work:
llama-server --model model.gguf --jinja --host 0.0.0.0 --port 8080
http://localhost:8080/v1/chat/completionsllama-server serves whichever model it loaded.--reasoning-format deepseek for a reasoning model. It returns the trace in reasoning_content, which HermitUI renders in the collapsible think block. Settings → Check Reasoning Support reads this server's /props and will report reasoning support exactly.If you have a 16 GB GPU, this is a config that fits and what each flag costs. Qwen3.8-27B at 4 bpw, one RTX 5070 Ti, using the model's own MTP layer for speculative decoding. Nothing here is card-specific except the numbers — the same reasoning applies to any 16 GB card.
llama-server --model Qwen3.8-27B-IQ4_XS-4.00bpw.gguf \
--ctx-size 40000 --jinja --reasoning-format deepseek --reasoning-preserve \
--flash-attn on --cache-type-k q8_0 --cache-type-v q8_0 \
--spec-type draft-mtp,ngram-mod --spec-draft-n-max 2 \
--cache-type-k-draft q8_0 --cache-type-v-draft q8_0 \
--n-gpu-layers all --parallel 1 --threads 32 --batch-size 1024 --ubatch-size 1024 \
--temp 1.0 --top-p 0.95 --top-k 20 --min-p 0.0 --repeat-penalty 1.0 \
--host 0.0.0.0 --port 8080
Speed (llama-server's own timings, so decode-only — TTFT excluded):
| Workload | With --spec-type | Without | Gain |
|---|---|---|---|
| Prose (200-word explainer, ×3) | 77.1 t/s | 50.4 t/s | 1.53× |
| Reasoning prompt | 91–96 t/s | — | ~1.85× |
| Code generation | 100.7 t/s | 50.2 t/s | 2.01× |
Prompt processing reached 1,489 t/s on a 6,742-token prompt. The no-speculation baseline was flat to within 0.3 t/s across four runs, so those gains are signal rather than noise.
Draft acceptance explains the spread — 98.7% on code, 82–88% on the reasoning prompt, 50–65% on prose. MTP predicts structured text far better than free prose, so the speedup is largest exactly where you generate code.
⚠️ --parallel 1 is not optional — it is worth 3.6×. Omit it and the server defaults to 4 slots, which flips llama.cpp to a unified KV cache (kv_unified = 'true') sized across all of them, so every decoded token does far more attention work. Measured on this machine, same GPU, same everything else: 20.4 t/s prose / 29.2 t/s code with 4 slots, versus 73.8 / 97.1 with one. VRAM barely moves (~15.5 GB either way), so nothing looks wrong — you just silently lose three quarters of your throughput. Check total_slots at http://localhost:8080/props; it must read 1.
Where the VRAM goes, at --ctx-size 40000 — as nvidia-smi and the server's own load log report it:
| VRAM | |
|---|---|
| Model weights on GPU (4 bpw) | 12,376 MiB |
Compute buffers (--ubatch-size 1024) | 396 MiB |
KV cache @ 40k q8_0, plus CUDA context overhead | ≈2,260 MiB |
MTP draft context (what --spec-type costs) | ≈860 MiB |
| Total resident | 15,893 of 16,303 MiB |
That leaves roughly 410 MiB spare — it runs, but with no room for anything else on the card. Don't run a WebGPU browser (including HermitUI's own in-browser mode) against the same GPU at this context size.
If it doesn't fit on your card, cut in this order:
--ctx-size. The KV cache scales linearly with it, so halving 40k → 20k returns roughly 1,000 MiB. This is the cheapest VRAM you will find.--cache-type-k q8_0 --cache-type-v q8_0 (and --flash-attn on, which the quantized-KV path wants). q8_0 stores one byte per element where f16 stores two, so reverting to f16 would roughly double the KV figure above — at 40k context that is the difference between fitting and not.--spec-type last. It returns ~860 MiB but costs the 1.5–2.0× decode speedup above, so trade it away only once the context is as small as you can live with.For scale: this is roughly 10× the ~8 t/s the same model reaches in the in-browser build. Native is far faster — the browser build buys zero-install and privacy, not speed.
https://openrouter.ai/api/v1/chat/completions).anthropic/claude-opus-4.8) — check your provider's model list for the exact slug.[!WARNING] Privacy note: using cloud models is generally not advised if you require strict privacy. Your data leaves your machine, and it is unclear how these providers handle, store, or train on it. For true ephemerality, stick to local models.
If HermitUI fails to connect to your local AI server (e.g., a "Network Error"), it is most likely CORS. Because HermitUI runs as a local file (file://), browsers block its requests to http://localhost unless the server explicitly allows it.
OLLAMA_ORIGINS before starting, e.g. OLLAMA_ORIGINS="*" ollama serve.--allowed-origins (it takes a JSON array, e.g. '["*"]').dist/hermit-ui-cdn.html.)localStorage, IndexedDB, or cookies — nothing survives the tab.<think>, <thought>, and <reasoning> tags as they stream from reasoning models./props capability flags (or Ollama's template). Levels are validated against the set the template accepts, so a value it would reject is never sent.image_url content per the OpenAI schema with automatic vision-model detection.$…$, $$…$$, \(…\), \[…\]) rendered via KaTeX to native MathML — no webfonts, works mid-stream and offline — and 📈 Mermaid diagrams from ```mermaid fences.temperature, max_tokens, top_p, presence_penalty, frequency_penalty, and seed from a collapsible Settings panel. Params are only sent when set, keeping payloads compatible with minimal backends.Ctrl+E), and load one back in later with 📂 (Ctrl+I) — messages, system prompt and attached context are restored, and the model picks the conversation up where you left it. The exported file is the only copy: nothing is stored in the browser. Images can't come back, since the export records them as [N images attached].You can pre-configure HermitUI through the URL fragment (the part after #), so a single link or bookmark carries the whole connection setup:
hermit-ui-standalone.html#api=http://localhost:8080/v1&model=qwen3-8b
hermit-ui-standalone.html#api=https://api.groq.com/openai/v1&key=gsk_...&model=llama-3.3-70b
hermit-ui-wllama.html#gguf=hf:unsloth/Qwen3-0.6B-GGUF/Qwen3-0.6B-Q4_K_M.gguf
| Parameter | Effect |
|---|---|
api | API base URL (same as the Settings field) |
model | Model name |
key | API key |
persona | Preset persona: technical, general, writing, or tutor |
gguf | (wllama build only) GGUF model to load in-browser — direct URL, Hugging Face link, or hf:user/repo/file.gguf shorthand. Shows a one-click confirmation banner before downloading. |
Why the fragment and not ?query: the part after # never leaves your browser — it is not sent in any HTTP request — and nothing is stored, so this stays true to the ephemerality promise (the URL is the config; refresh keeps your setup). Applied settings are always announced in a toast, so a shared link can't reconfigure the app invisibly. Free-text system prompts are deliberately not supported as a parameter, since a link could smuggle a malicious prompt.
[!NOTE] A
keyin the URL is never transmitted, but it does end up in your browser history (and any bookmark you save). Prefer entering keys in Settings on shared machines.
.gguf, against nothing at all.HermitUI enforces strict architectural constraints to remain lightweight and accessible:
.html file. The src/ directory is a blueprint only — its split into index.html, style.css, and script.js exists for maintainability, and build.py assembles them back into one file.package.json, npm, Webpack, or Vite.DOMPurify to prevent XSS.The live build runs on GitHub Pages at moooff.github.io/HermitUI.
"Stores nothing" is the whole pitch, so don't take it on trust — checking takes about a minute.
grep -o "localStorage\|sessionStorage\|indexedDB\|document\.cookie" dist/hermit-ui-standalone.html | wc -l
The answer is 2, and both are false positives — Highlight.js's list of JavaScript keywords contains the strings localStorage and sessionStorage. There is not one call site. (In the wllama build the engine ships gzipped, so grep sees the app but not the engine; use the runtime check above to cover it.)Model weights get the same treatment. A downloaded GGUF is streamed into an in-memory Blob instead of going through wllama's own URL loader, specifically because that one would persist the model to OPFS — see downloadGgufToBlob in src/script.js.
The root index.html (a copy of dist/hermit-ui-standalone.html) is a completely offline, standalone build: web fonts and images are base64-encoded and external JS/CSS libraries are injected directly into the file, which is what makes it work in air-gapped environments.
To modify it, edit the modular sources in src/ — index.html, style.css, and script.js, which reference libraries via CDN for convenient local development — then run:
python build.py # or python3 build.py
Prerequisites: Python 3 and, on the first run, an internet connection. That is the whole list — build.py uses only the standard library, so there is no pip install, no package.json, and no Node.
The network requirement is worth spelling out, since it cuts against the rest of the project: the output runs offline, but producing it does not. On the first build, build.py downloads the pinned library versions (Marked.js, DOMPurify, Highlight.js, KaTeX, Mermaid, the Inter font, and the wllama engine — ~14 MB total) into libs/, verifying each one against the SRI hash pinned in src/index.html. Everything after that is cached and offline; pass --refresh to force a re-download. So if you need to build on an air-gapped machine, copy a populated libs/ directory across with the repo.
This generates the standalone build at dist/hermit-ui-standalone.html, copies it to the root index.html for GitHub Pages, and creates the alternative builds in dist/. The standalone, CDN, and wllama variants (dist/hermit-ui-standalone.html, dist/hermit-ui-cdn.html, dist/hermit-ui-wllama.html) are committed so they are browsable and downloadable straight from GitHub; the local variant dist/hermit-ui-local.html and the downloaded libs/ are generated-only and stay gitignored.
Bug reports, questions and ideas are welcome in Issues; pull requests are welcome too. A bug report travels much further with your browser and version, whether WebGPU was on, and — for in-browser inference — the model and the output of the debug console at verbosity Debug.
Before opening a PR, two things will save you a rewrite:
AGENTS.md. It is the single source of truth for this project's rules, and they are unusually strict on purpose: single-file output, vanilla JS only, no build tools or frameworks or CSS libraries, no localStorage/IndexedDB/cookies for any reason, and every AI-rendered string sanitized through DOMPurify. A change that breaks one of these can't be merged no matter how good it is — the constraints are the product.src/, never the generated files. dist/*.html and the root index.html are build artifacts and get overwritten; run python build.py and commit the regenerated outputs along with your source change.There is no test suite or linter. Verify a change by opening the rebuilt dist/hermit-ui-standalone.html (or dist/hermit-ui-wllama.html) in a browser and exercising the affected path by hand. If your change touches inference performance, the benchmark harness produces numbers that can go straight into a PR description.
```mermaid fences-00001-of-000NN.gguf) through the in-browser URL loader.The full list of ideas and tasks lives in docs/backlog.md.
This project is open-source and available under the terms of the GNU AGPL v3. See the included LICENSE file for the full text.
HTML
90.1%
JavaScript
6.6%
Python
2.1%
CSS
1.3%