(Probably) The fastest browser based Gemma 4 E2B implementation with vision support. Comes with a workstation to interact with the model. Use AI fully air-gapped and offline, you just need a browser.
See the codeRun Gemma 4 E2B (QAT Mobile) entirely in your browser — text and vision — 100% on-device with WebGPU. No server, no API calls, no transformers.js. Weights download once from Hugging Face, cache in IndexedDB, and every token is produced on your own GPU.
This repository is a fork of the
webml-community/gemma-4-webgpu-kernels
Hugging Face Space by Xenova, with three big additions on top of the custom
kernel:
| Model | google/gemma-4-E2B-it-qat-mobile-transformers |
| Effective size | ~2.3B params (QAT mobile, w8a8o8) |
| Context | Supports up to 128K architectural context, subject to device and runtime memory limits. The top-bar selector defaults to a 32K cap. |
| Runtime | WebGPU compute — custom WGSL kernels |
| Multimodal | Text + images, fully on-device |
| Privacy | Prompts never leave your machine |
subgroupAdd reduce
into a portable butterfly (or falls back to disabling subgroups) so Windows +
NVIDIA output stays correct.One model load powers four apps (hash-routed views in a single page):
| App | What it does |
|---|---|
| Chat | Streaming chat + vision, with conversation history (IndexedDB), export to .md, and the kernels viewer. Context cap follows the global top-bar selector. |
| Research | Upload PDF / DOCX / TXT / MD / CSV / images and reason over them. Supports up to 128K architectural context, subject to device and runtime memory limits. Research automatically switches to BM25 retrieval when the effective prompt budget is exceeded — every selected document contributes at least one retrieved block, and chunks are grouped under their document in both the prompt and the context inspector, so documents can be compared. A context inspector shows exactly what the model sees. Scanned PDFs get OCR'd by the on-device vision tower with per-page progress and ETA, then export as structured Markdown with page headings and OCR markers. Global context cap is controlled from the top bar (defaults to 32K). |
| Code | File-aware chat workstation: explorer + editor + chat over project files. See Code — Chat with files below. |
| Reports | Staged report generation tuned for greedy decoding: JSON outline → one bounded section at a time → charts emitted as JSON specs rendered by Chart.js (never model-written JS). Export produces a self-contained .html with charts baked in as PNGs — it renders offline with zero JavaScript. Global context cap from the top bar applies to each stage. |
Everything is saved locally in IndexedDB (gemma4-workstation-v1): conversations,
reports, parsed documents and code projects (code-project-v3). Nothing syncs anywhere.
The Context selector in the workstation top bar (Auto / 8K / 16K / 32K / 64K / 128K) caps the effective context window shared across Chat, Research, Code and Reports. It defaults to 32K; a one-time migration moves existing profiles to that default, after which your own choice is remembered. Auto lets the device use its full runtime capacity (reported in the status pill as runtime X K / 128K). Selecting e.g. 64K triggers an on-device KV-cache re-allocation and the pill then shows runtime 64K / 128K once the allocation succeeds. Lower caps are useful to simulate smaller GPUs or to keep prompts bounded. Research shows a live inspector of what fits; Code's chat panel shows a token estimate and truncates oversize attachments with a visible notice.
A browser-native workspace for iterating over code with the model:
Layout
280px): collapsible tree with folders/files (create / rename / move by drag-and-drop or ⋯ menu, delete). Toolbar: + File, + Folder, ⬆ Zip (upload a .zip — fflate-unzipped, text files merged into the virtual FS), ⬇ Zip (download entire project), Reset (restores README.md + main.py + index.html/style.css/script.js).Console vs Preview.
Console · Python for Pyodide stdout/stderr/plots, Console · Web for the preview's console.* output — so the two never interleave. Opening or running a file points it at that runner, and output arriving while its console is off screen is counted on the tab. clear only empties the console you are looking at.allow-scripts only, no same-origin) and its tab appears only when relevant: for an HTML file, or for an asset the previewed page references (so editing a stylesheet shows its live effect immediately). A Python-only session shows no Preview tab.▶ Run always acts on the active file only: .py → Pyodide with full FS sync so import utils.helpers works; .html → renders that HTML file in Preview. CSS/JS/Markdown have no runner (the button is disabled rather than guessing another file). Editing an HTML file re-renders it, and editing a stylesheet/script that the previewed page references re-renders that same page in place..py file and a ▶ Run selection button appears next to ▶ Run Python. It runs just the highlighted code, exactly as written — the file itself is not executed. Output appends to the Python console; plots render inline.clear resets Python: clearing the Python console also drops all variables, functions and runtime imports, so the next run starts clean. Installed packages stay cached; the web console is untouched.400px): read-only Q&A over your project files. Attach files with @-mentions, the attachment bar, editor selections, or right-click → Explain / Review / Add to Chat. Suggested code ships with a Copy button per snippet — you paste it into the editor yourself. Chat never edits files.Package management (offline-ready)
📦 Packages dialog lists 34 bundled pure-Python packages that install from disk with no network:
.py file scans its imports and installs any bundled package it needs — so import seaborn works offline. Data-science packages that build on Pyodide's binary wheels (numpy, pandas, matplotlib, scipy, scikit-learn, statsmodels) get those loaded automatically, and each entry shows what it needs.vendor/python-packages/ and are generated by npm run vendor:python (node tools/vendor-python-packages.mjs, --force to re-download, --list to show the curated set). Pyodide's own binary packages are not duplicated; curated packages with no pure-Python wheel (e.g. PyYAML) are reported as PyPI-only.sw.js caches every .whl (vendored and downloaded from PyPI) cache-first, so installed packages keep working offline; manifest.json is deliberately revalidated so a re-vendor is picked up.How file chat works
@-mention a file to attach it (autocomplete popup), right-click → Explain / Review / Add to Chat, or highlight code → Add selection. Up to 8 attachments persist across turns so follow-ups ("now fix the loop") keep working; oversize files are head-truncated with a visible notice.→ Add selection. Clicking attaches that snippet to the next chat turn, so you can scope questions without stuffing the whole repo.app.py and web/ in one project; runners dispatch by the active file's own extension (.py → Pyodide, .html → sandboxed preview; everything else is edit-only). import works across folders (a/b.py → from a.b import x via Pyodide FS), and the preview inlines whatever relative CSS/JS an HTML entry point references — nested folders and any file names, resolved from the project itself.⚠️ Keep code runs bounded. Both runtimes share the browser's main thread, so a
while True:loop with nobreak(Python) or an endlesswhile (true)loop in previewed JavaScript freezes the whole tab — including the⏹ Stopbutton — and the tab has to be reloaded. The app warns before running a Pythonwhile True:loop without abreak, and tells you on the next load if a previous run never finished. Move the runtime into a Web Worker if you need pre-emptible execution.
Quick start
main.py or index.html.@main.py (or right-click → Explain this file) and ask, e.g. “what does this do?” → Send.▶ Run to verify.Add selection → “fix the off-by-one here”.⬆ Zip to import an existing codebase, ⬇ Zip to export, Reset to start fresh. All files persist in IndexedDB.src/services/model-service.js) — a global
generation lock (src/services/generation.js) guarantees only one stream at a
time across all apps.sw.js); weight downloads are explicitly excluded so HTTP
Range streaming keeps working.localhost or HTTPS (browsers require this for WebGPU)The app is fully static and Pages-friendly: serve the repo root and open
index.html. All module imports are relative, and weights stream from the
Hugging Face Hub (CORS + Range enabled) so no server and no local models are
needed.
⚠️ First load downloads ~2.4 GB of weights from Hugging Face. After that they live in IndexedDB and load from cache.
To deploy: repo Settings → Pages → Deploy from a branch → main / root.
No build step required.
⚠️ You MUST use a server that supports HTTP
Rangerequests. The kernel and the vision loader stream the 2.4 GBmodel.safetensorswith byte-range fetches.python3 -m http.serverignoresRangeand returns the whole file, which crashes withRangeError: Array buffer allocation failed. Use the bundled server:
# from this repo's root
node tools/serve.mjs 4173
# or: npm run serve
# throttle defaults to 1000 Mbit/s (~125 MB/s global) to avoid the
# localhost stall at ~91% (see note below).
# THROTTLE_MBPS=0 node tools/serve.mjs 4173 # disable
# node tools/serve.mjs 4173 . --throttle-mbps 500 # 500 Mbit/s
Then open http://localhost:4173/ → Load model → chat.
💡 Fully offline? Place
model.safetensorsundermodels/google/gemma-4-E2B-it-qat-mobile-transformers/and open the page with?localweights=1(bothindex.htmlandtest-vision.htmlrespect this).
⚠️ Local testing stall at 91%? At loopback the kernel streams
4×128 MiBRange requests concurrently (hd=4, md=128 MiB). Without a cap that bursts>>1 Gbit, saturating Chrome'sReadableStream+IndexedDB pipeline (streamAll→writeTensorper chunk) and freezing progress at e.g.Loading cached weights: 1.79 GB / 1.97 GB (91%).tools/serve.mjsnow caps output to ~1 Gbit/s global (token-bucket,THROTTLE_MBPSenv /--throttle-mbpsflag) so IDB commits can keep up. The same file also now handlesHEADcorrectly (engine probes size) and aborts streams onclose. Disable with--no-throttleif you need full loopback speed.
Ship the whole workstation as one HTML file plus an assets folder:
npm install # release-only build deps
npm run vendor:portable # Pyodide runtime + document parsers (~15 MB)
npm run build:portable # → dist/portable/
Copy dist/portable/ anywhere, open gemma4-workstation.html in Chrome or Edge, and either
pick the assets folder (fully offline) or click Stream from Hugging Face and let the
model download itself. No server and no installation either way.
A file:// page cannot read its sibling files and has no HTTP Range support, so
in offline mode the app reads the weights by slicing the File you picked (Blob.slice)
through a shimmed fetch — served under a reserved virtual origin and supporting Range.
The engine's own options.fetch / knownSize / knownAcceptsRanges seam takes it from
there, and the same shim serves the vendored Pyodide runtime and wheels from disk.
To package a downloadable release:
npm run build:release # → dist/release/gemma4-workstation-portable-<version>.zip (67 MB)
The weights are deliberately not in it: GitHub release assets must be under 2 GiB per file, and the checkpoint is 2.29 GiB raw / 2.02 GiB compressed (the int8 tensors are ~90% incompressible). Mode B supplies them instead. See PORTABLE.md for the full picture, the measurements, limits and fallbacks.
Releases are automated: pushing a v* tag runs .github/workflows/release.yml, which
checks the tag against package.json, vendors the runtime and the model sidecars, runs the
suite, then attaches the zip and its SHA256SUMS.txt. The release page opens with the
how-to-use steps from docs/RELEASE-HOWTO.md followed by a condensed changelog section —
the full detail stays in CHANGELOG.md. Run the workflow manually to get a
draft first. npm run changelog <ver> [--brief] [--usage <file>] prints the notes locally.
Browser page (index.html)
│
├─ landing.js Three.js hero (WebGL); pauses when chat is active
│
└─ load model
│
├─ gemma4-sg-guard.js self-test + patch bare subgroupAdd (or disable subgroups)
│
└─ gemma-4-e2b.js Gemma4Mobile runtime (patched with vision hooks)
│
├─ request WebGPU device
├─ fetch tokenizer + chat template (HF Hub or local)
├─ fetch / cache safetensors weights (IndexedDB)
├─ compile device-selected WGSL kernels
└─ generate() streams tokens on-GPU
└─ vision (optional)
└─ gemma4-vision.js custom 16-layer WGSL vision tower
└─ gemma4-vision-inject.js injects image features into the LLM kernel
index.html — vanilla HTML/CSS/JS, no bundler. Loads the model, runs the
subgroup guard, wraps the model with vision, then streams
model.generate(messages, { maxNewTokens: 4096 }). Has a “View Kernels”
overlay that shows the actually compiled WGSL for your GPU.gemma-4-e2b.js — self-contained ES module engine: tokenizer + Jinja chat
template, IndexedDB weight cache, and fused WebGPU/WGSL op templates for Gemma 4
decode/prefill. Patched with 8 surgical vision hooks (already applied; see
VISION.md).gemma4-vision.js — faithful port of Gemma4VisionModel from
huggingface/transformers (mobile QAT w8a8o8): preprocessing, patch embedder,
2-D RoPE, bidirectional attention, 3×3 pooling, and the embedding projection →
[num_soft_tokens, 1536] image features.gemma4-sg-guard.js — MIT
drop-in guard that patches the NVIDIA/Windows subgroup bug at load time.Attach an image in the composer and the model answers with vision. The image encoder is a custom WebGPU/WGSL port of the Gemma 4 vision tower that loads the vision weights from the same mobile QAT safetensors the text engine uses (HTTP Range requests, ~190 MB, no transformers.js / onnxruntime). See VISION.md for the full architecture, the chunked-prefill injection design, and the implementation notes.
Highlights:
matmulQat/matmulQkv/matmulGateUp use packed int8
activations + WGSL dot4I8Packed (i32 accumulate), bit-exact vs the f32 path.| Metric | f32 baseline | int8 + attention | Gain |
|---|---|---|---|
| Vision encode (warm) | ~5250 ms | ~2050 ms | 2.6× |
| App time-to-first-token (vision prompt) | 6601 ms | 3093 ms | 2.1× |
| Decode throughput | ~129 tok/s | ~129 tok/s | unchanged |
On some Windows + NVIDIA (D3D12) stacks, a bare WGSL subgroupAdd after
lane-divergent stores returns wrong sums inside QatMatMul. The model still
runs fast and reports no errors, but tokens become gibberish / repetition loops.
Before load, this fork:
force: true), and never takes the Apple-style
“exact 32/32 → do nothing” shortcut on Windows.createShaderModule and rewrites
bare reduces to subgroupShuffleXor butterflies (patched-sg).subgroups for the engine
(nosubgroups-fallback, slightly slower but correct).Details, console expectations, and upstream links: NVIDIA-WINDOWS-GIBBERISH-FIX.md
.
├── index.html Workstation shell: landing hero + hash-routed app frame
├── landing.js Three.js landing scene
├── sw.js Service worker: app shell + CDN/library + wheel caches
├── manifest.webmanifest, icon.svg PWA install metadata
├── gemma-4-e2b.js WebGPU inference engine + embedded WGSL (vision hooks applied)
├── gemma4-vision.js Vision tower: preprocessing + 16-layer WGSL encoder + pooling + projection
├── gemma4-vision-inject.js Wraps Gemma4Mobile so generate() accepts image content + injects features
├── gemma4-sg-guard.js NVIDIA/Windows subgroup correctness guard (MIT, from Ar5en1c)
├── test-vision.html WebGPU vision test harness (QAT matmul vs CPU + encode sanity)
├── apps/
│ ├── chat/app.js Chat: vision input, conversation history, .md export
│ ├── research/app.js Documents: PDF/DOCX/TXT parsing, retrieval Q&A, OCR
│ ├── reports/app.js Staged report generation, charts, self-contained HTML export
│ └── code/ Code app
│ ├── app.js Explorer + editor + preview/console + file chat
│ ├── components/ explorer.js, editor-cm.js (CodeMirror 6 wrapper)
│ └── runners/ web-runner.js, pyodide-runner.js, python-packages.js
├── src/
│ ├── model-config.js Weight URL config (HF Hub default, ?localweights=1 override)
│ ├── lib/ markdown.js, chat-thread.js, zip-utils.js, document-markdown.js
│ ├── services/ model-service, generation, db, context, code-project, settings, …
│ └── shell/router.js Hash router (lazy app mounting)
├── vendor/python-packages/ Offline pure-Python wheels + manifest.json + LICENSES.md
├── models/README.md Optional local weights drop-in (see below)
├── tools/
│ ├── serve.mjs Range-capable static server (required for weight streaming)
│ ├── vendor-python-packages.mjs Downloads/refreshes the offline Python bundle
│ └── check-release.mjs Pre-release checks (module graph, import map, bundle)
├── VISION.md Vision architecture + implementation notes
├── NVIDIA-WINDOWS-GIBBERISH-FIX.md Deep dive on the Windows gibberish bug
├── CHANGELOG.md, THIRD_PARTY_NOTICES.md
└── README.md This file
No build step or bundler — the page is plain static files.
After Load model, you should see something like:
[gemma4-sg-guard] patched-sg { …adapter… } { bare: "FAIL(…)", butterfly: "PASS" }
or nosubgroups-fallback. On Windows you should not see exact32-stock {}
(that path leaves the broken kernels unpatched).
158f16ae). The text kernels were written and optimized by
Fable 5 for the upstream Space.gemma4-sg-guard.js
runtime guard (MIT).gemma4-vision.js),
the multimodal extension (gemma4-vision-inject.js), the 8 kernel patches,
and the test harness, as custom WebGPU kernels.vendor/python-packages/ — 49 pure-Python wheels
redistributed unmodified (MIT / BSD / Apache-2.0 / 0BSD / PSF-2.0, plus
pingouin GPL-3.0 and tqdm MPL-2.0 AND MIT). Per-package licenses:
vendor/python-packages/LICENSES.md.Upstream Space: https://huggingface.co/spaces/webml-community/gemma-4-webgpu-kernels
This repo: https://github.com//gemma-4-webgpu-kernels
The code written for this fork — the vision tower, multimodal extension, kernel patches, tools, and documentation — is released under the MIT License.
Third-party components keep their own terms; full attribution is in
THIRD_PARTY_NOTICES.md. Please note the upstream
Xenova engine (gemma-4-e2b.js, upstream index.html/landing.js) carries
no explicit license grant at the time of writing, and the model weights
are governed by Google's
Gemma Terms of Use.
JavaScript
84.4%
HTML
15.5%
(Probably) The fastest browser based Gemma 4 E2B implementation with vision support. Comes with a workstation to interact with the model. Use AI fully air-gapped and offline, you just need a browser.
See the codeRun Gemma 4 E2B (QAT Mobile) entirely in your browser — text and vision — 100% on-device with WebGPU. No server, no API calls, no transformers.js. Weights download once from Hugging Face, cache in IndexedDB, and every token is produced on your own GPU.
This repository is a fork of the
webml-community/gemma-4-webgpu-kernels
Hugging Face Space by Xenova, with three big additions on top of the custom
kernel:
| Model | google/gemma-4-E2B-it-qat-mobile-transformers |
| Effective size | ~2.3B params (QAT mobile, w8a8o8) |
| Context | Supports up to 128K architectural context, subject to device and runtime memory limits. The top-bar selector defaults to a 32K cap. |
| Runtime | WebGPU compute — custom WGSL kernels |
| Multimodal | Text + images, fully on-device |
| Privacy | Prompts never leave your machine |
subgroupAdd reduce
into a portable butterfly (or falls back to disabling subgroups) so Windows +
NVIDIA output stays correct.One model load powers four apps (hash-routed views in a single page):
| App | What it does |
|---|---|
| Chat | Streaming chat + vision, with conversation history (IndexedDB), export to .md, and the kernels viewer. Context cap follows the global top-bar selector. |
| Research | Upload PDF / DOCX / TXT / MD / CSV / images and reason over them. Supports up to 128K architectural context, subject to device and runtime memory limits. Research automatically switches to BM25 retrieval when the effective prompt budget is exceeded — every selected document contributes at least one retrieved block, and chunks are grouped under their document in both the prompt and the context inspector, so documents can be compared. A context inspector shows exactly what the model sees. Scanned PDFs get OCR'd by the on-device vision tower with per-page progress and ETA, then export as structured Markdown with page headings and OCR markers. Global context cap is controlled from the top bar (defaults to 32K). |
| Code | File-aware chat workstation: explorer + editor + chat over project files. See Code — Chat with files below. |
| Reports | Staged report generation tuned for greedy decoding: JSON outline → one bounded section at a time → charts emitted as JSON specs rendered by Chart.js (never model-written JS). Export produces a self-contained .html with charts baked in as PNGs — it renders offline with zero JavaScript. Global context cap from the top bar applies to each stage. |
Everything is saved locally in IndexedDB (gemma4-workstation-v1): conversations,
reports, parsed documents and code projects (code-project-v3). Nothing syncs anywhere.
The Context selector in the workstation top bar (Auto / 8K / 16K / 32K / 64K / 128K) caps the effective context window shared across Chat, Research, Code and Reports. It defaults to 32K; a one-time migration moves existing profiles to that default, after which your own choice is remembered. Auto lets the device use its full runtime capacity (reported in the status pill as runtime X K / 128K). Selecting e.g. 64K triggers an on-device KV-cache re-allocation and the pill then shows runtime 64K / 128K once the allocation succeeds. Lower caps are useful to simulate smaller GPUs or to keep prompts bounded. Research shows a live inspector of what fits; Code's chat panel shows a token estimate and truncates oversize attachments with a visible notice.
A browser-native workspace for iterating over code with the model:
Layout
280px): collapsible tree with folders/files (create / rename / move by drag-and-drop or ⋯ menu, delete). Toolbar: + File, + Folder, ⬆ Zip (upload a .zip — fflate-unzipped, text files merged into the virtual FS), ⬇ Zip (download entire project), Reset (restores README.md + main.py + index.html/style.css/script.js).Console vs Preview.
Console · Python for Pyodide stdout/stderr/plots, Console · Web for the preview's console.* output — so the two never interleave. Opening or running a file points it at that runner, and output arriving while its console is off screen is counted on the tab. clear only empties the console you are looking at.allow-scripts only, no same-origin) and its tab appears only when relevant: for an HTML file, or for an asset the previewed page references (so editing a stylesheet shows its live effect immediately). A Python-only session shows no Preview tab.▶ Run always acts on the active file only: .py → Pyodide with full FS sync so import utils.helpers works; .html → renders that HTML file in Preview. CSS/JS/Markdown have no runner (the button is disabled rather than guessing another file). Editing an HTML file re-renders it, and editing a stylesheet/script that the previewed page references re-renders that same page in place..py file and a ▶ Run selection button appears next to ▶ Run Python. It runs just the highlighted code, exactly as written — the file itself is not executed. Output appends to the Python console; plots render inline.clear resets Python: clearing the Python console also drops all variables, functions and runtime imports, so the next run starts clean. Installed packages stay cached; the web console is untouched.400px): read-only Q&A over your project files. Attach files with @-mentions, the attachment bar, editor selections, or right-click → Explain / Review / Add to Chat. Suggested code ships with a Copy button per snippet — you paste it into the editor yourself. Chat never edits files.Package management (offline-ready)
📦 Packages dialog lists 34 bundled pure-Python packages that install from disk with no network:
.py file scans its imports and installs any bundled package it needs — so import seaborn works offline. Data-science packages that build on Pyodide's binary wheels (numpy, pandas, matplotlib, scipy, scikit-learn, statsmodels) get those loaded automatically, and each entry shows what it needs.vendor/python-packages/ and are generated by npm run vendor:python (node tools/vendor-python-packages.mjs, --force to re-download, --list to show the curated set). Pyodide's own binary packages are not duplicated; curated packages with no pure-Python wheel (e.g. PyYAML) are reported as PyPI-only.sw.js caches every .whl (vendored and downloaded from PyPI) cache-first, so installed packages keep working offline; manifest.json is deliberately revalidated so a re-vendor is picked up.How file chat works
@-mention a file to attach it (autocomplete popup), right-click → Explain / Review / Add to Chat, or highlight code → Add selection. Up to 8 attachments persist across turns so follow-ups ("now fix the loop") keep working; oversize files are head-truncated with a visible notice.→ Add selection. Clicking attaches that snippet to the next chat turn, so you can scope questions without stuffing the whole repo.app.py and web/ in one project; runners dispatch by the active file's own extension (.py → Pyodide, .html → sandboxed preview; everything else is edit-only). import works across folders (a/b.py → from a.b import x via Pyodide FS), and the preview inlines whatever relative CSS/JS an HTML entry point references — nested folders and any file names, resolved from the project itself.⚠️ Keep code runs bounded. Both runtimes share the browser's main thread, so a
while True:loop with nobreak(Python) or an endlesswhile (true)loop in previewed JavaScript freezes the whole tab — including the⏹ Stopbutton — and the tab has to be reloaded. The app warns before running a Pythonwhile True:loop without abreak, and tells you on the next load if a previous run never finished. Move the runtime into a Web Worker if you need pre-emptible execution.
Quick start
main.py or index.html.@main.py (or right-click → Explain this file) and ask, e.g. “what does this do?” → Send.▶ Run to verify.Add selection → “fix the off-by-one here”.⬆ Zip to import an existing codebase, ⬇ Zip to export, Reset to start fresh. All files persist in IndexedDB.src/services/model-service.js) — a global
generation lock (src/services/generation.js) guarantees only one stream at a
time across all apps.sw.js); weight downloads are explicitly excluded so HTTP
Range streaming keeps working.localhost or HTTPS (browsers require this for WebGPU)The app is fully static and Pages-friendly: serve the repo root and open
index.html. All module imports are relative, and weights stream from the
Hugging Face Hub (CORS + Range enabled) so no server and no local models are
needed.
⚠️ First load downloads ~2.4 GB of weights from Hugging Face. After that they live in IndexedDB and load from cache.
To deploy: repo Settings → Pages → Deploy from a branch → main / root.
No build step required.
⚠️ You MUST use a server that supports HTTP
Rangerequests. The kernel and the vision loader stream the 2.4 GBmodel.safetensorswith byte-range fetches.python3 -m http.serverignoresRangeand returns the whole file, which crashes withRangeError: Array buffer allocation failed. Use the bundled server:
# from this repo's root
node tools/serve.mjs 4173
# or: npm run serve
# throttle defaults to 1000 Mbit/s (~125 MB/s global) to avoid the
# localhost stall at ~91% (see note below).
# THROTTLE_MBPS=0 node tools/serve.mjs 4173 # disable
# node tools/serve.mjs 4173 . --throttle-mbps 500 # 500 Mbit/s
Then open http://localhost:4173/ → Load model → chat.
💡 Fully offline? Place
model.safetensorsundermodels/google/gemma-4-E2B-it-qat-mobile-transformers/and open the page with?localweights=1(bothindex.htmlandtest-vision.htmlrespect this).
⚠️ Local testing stall at 91%? At loopback the kernel streams
4×128 MiBRange requests concurrently (hd=4, md=128 MiB). Without a cap that bursts>>1 Gbit, saturating Chrome'sReadableStream+IndexedDB pipeline (streamAll→writeTensorper chunk) and freezing progress at e.g.Loading cached weights: 1.79 GB / 1.97 GB (91%).tools/serve.mjsnow caps output to ~1 Gbit/s global (token-bucket,THROTTLE_MBPSenv /--throttle-mbpsflag) so IDB commits can keep up. The same file also now handlesHEADcorrectly (engine probes size) and aborts streams onclose. Disable with--no-throttleif you need full loopback speed.
Ship the whole workstation as one HTML file plus an assets folder:
npm install # release-only build deps
npm run vendor:portable # Pyodide runtime + document parsers (~15 MB)
npm run build:portable # → dist/portable/
Copy dist/portable/ anywhere, open gemma4-workstation.html in Chrome or Edge, and either
pick the assets folder (fully offline) or click Stream from Hugging Face and let the
model download itself. No server and no installation either way.
A file:// page cannot read its sibling files and has no HTTP Range support, so
in offline mode the app reads the weights by slicing the File you picked (Blob.slice)
through a shimmed fetch — served under a reserved virtual origin and supporting Range.
The engine's own options.fetch / knownSize / knownAcceptsRanges seam takes it from
there, and the same shim serves the vendored Pyodide runtime and wheels from disk.
To package a downloadable release:
npm run build:release # → dist/release/gemma4-workstation-portable-<version>.zip (67 MB)
The weights are deliberately not in it: GitHub release assets must be under 2 GiB per file, and the checkpoint is 2.29 GiB raw / 2.02 GiB compressed (the int8 tensors are ~90% incompressible). Mode B supplies them instead. See PORTABLE.md for the full picture, the measurements, limits and fallbacks.
Releases are automated: pushing a v* tag runs .github/workflows/release.yml, which
checks the tag against package.json, vendors the runtime and the model sidecars, runs the
suite, then attaches the zip and its SHA256SUMS.txt. The release page opens with the
how-to-use steps from docs/RELEASE-HOWTO.md followed by a condensed changelog section —
the full detail stays in CHANGELOG.md. Run the workflow manually to get a
draft first. npm run changelog <ver> [--brief] [--usage <file>] prints the notes locally.
Browser page (index.html)
│
├─ landing.js Three.js hero (WebGL); pauses when chat is active
│
└─ load model
│
├─ gemma4-sg-guard.js self-test + patch bare subgroupAdd (or disable subgroups)
│
└─ gemma-4-e2b.js Gemma4Mobile runtime (patched with vision hooks)
│
├─ request WebGPU device
├─ fetch tokenizer + chat template (HF Hub or local)
├─ fetch / cache safetensors weights (IndexedDB)
├─ compile device-selected WGSL kernels
└─ generate() streams tokens on-GPU
└─ vision (optional)
└─ gemma4-vision.js custom 16-layer WGSL vision tower
└─ gemma4-vision-inject.js injects image features into the LLM kernel
index.html — vanilla HTML/CSS/JS, no bundler. Loads the model, runs the
subgroup guard, wraps the model with vision, then streams
model.generate(messages, { maxNewTokens: 4096 }). Has a “View Kernels”
overlay that shows the actually compiled WGSL for your GPU.gemma-4-e2b.js — self-contained ES module engine: tokenizer + Jinja chat
template, IndexedDB weight cache, and fused WebGPU/WGSL op templates for Gemma 4
decode/prefill. Patched with 8 surgical vision hooks (already applied; see
VISION.md).gemma4-vision.js — faithful port of Gemma4VisionModel from
huggingface/transformers (mobile QAT w8a8o8): preprocessing, patch embedder,
2-D RoPE, bidirectional attention, 3×3 pooling, and the embedding projection →
[num_soft_tokens, 1536] image features.gemma4-sg-guard.js — MIT
drop-in guard that patches the NVIDIA/Windows subgroup bug at load time.Attach an image in the composer and the model answers with vision. The image encoder is a custom WebGPU/WGSL port of the Gemma 4 vision tower that loads the vision weights from the same mobile QAT safetensors the text engine uses (HTTP Range requests, ~190 MB, no transformers.js / onnxruntime). See VISION.md for the full architecture, the chunked-prefill injection design, and the implementation notes.
Highlights:
matmulQat/matmulQkv/matmulGateUp use packed int8
activations + WGSL dot4I8Packed (i32 accumulate), bit-exact vs the f32 path.| Metric | f32 baseline | int8 + attention | Gain |
|---|---|---|---|
| Vision encode (warm) | ~5250 ms | ~2050 ms | 2.6× |
| App time-to-first-token (vision prompt) | 6601 ms | 3093 ms | 2.1× |
| Decode throughput | ~129 tok/s | ~129 tok/s | unchanged |
On some Windows + NVIDIA (D3D12) stacks, a bare WGSL subgroupAdd after
lane-divergent stores returns wrong sums inside QatMatMul. The model still
runs fast and reports no errors, but tokens become gibberish / repetition loops.
Before load, this fork:
force: true), and never takes the Apple-style
“exact 32/32 → do nothing” shortcut on Windows.createShaderModule and rewrites
bare reduces to subgroupShuffleXor butterflies (patched-sg).subgroups for the engine
(nosubgroups-fallback, slightly slower but correct).Details, console expectations, and upstream links: NVIDIA-WINDOWS-GIBBERISH-FIX.md
.
├── index.html Workstation shell: landing hero + hash-routed app frame
├── landing.js Three.js landing scene
├── sw.js Service worker: app shell + CDN/library + wheel caches
├── manifest.webmanifest, icon.svg PWA install metadata
├── gemma-4-e2b.js WebGPU inference engine + embedded WGSL (vision hooks applied)
├── gemma4-vision.js Vision tower: preprocessing + 16-layer WGSL encoder + pooling + projection
├── gemma4-vision-inject.js Wraps Gemma4Mobile so generate() accepts image content + injects features
├── gemma4-sg-guard.js NVIDIA/Windows subgroup correctness guard (MIT, from Ar5en1c)
├── test-vision.html WebGPU vision test harness (QAT matmul vs CPU + encode sanity)
├── apps/
│ ├── chat/app.js Chat: vision input, conversation history, .md export
│ ├── research/app.js Documents: PDF/DOCX/TXT parsing, retrieval Q&A, OCR
│ ├── reports/app.js Staged report generation, charts, self-contained HTML export
│ └── code/ Code app
│ ├── app.js Explorer + editor + preview/console + file chat
│ ├── components/ explorer.js, editor-cm.js (CodeMirror 6 wrapper)
│ └── runners/ web-runner.js, pyodide-runner.js, python-packages.js
├── src/
│ ├── model-config.js Weight URL config (HF Hub default, ?localweights=1 override)
│ ├── lib/ markdown.js, chat-thread.js, zip-utils.js, document-markdown.js
│ ├── services/ model-service, generation, db, context, code-project, settings, …
│ └── shell/router.js Hash router (lazy app mounting)
├── vendor/python-packages/ Offline pure-Python wheels + manifest.json + LICENSES.md
├── models/README.md Optional local weights drop-in (see below)
├── tools/
│ ├── serve.mjs Range-capable static server (required for weight streaming)
│ ├── vendor-python-packages.mjs Downloads/refreshes the offline Python bundle
│ └── check-release.mjs Pre-release checks (module graph, import map, bundle)
├── VISION.md Vision architecture + implementation notes
├── NVIDIA-WINDOWS-GIBBERISH-FIX.md Deep dive on the Windows gibberish bug
├── CHANGELOG.md, THIRD_PARTY_NOTICES.md
└── README.md This file
No build step or bundler — the page is plain static files.
After Load model, you should see something like:
[gemma4-sg-guard] patched-sg { …adapter… } { bare: "FAIL(…)", butterfly: "PASS" }
or nosubgroups-fallback. On Windows you should not see exact32-stock {}
(that path leaves the broken kernels unpatched).
158f16ae). The text kernels were written and optimized by
Fable 5 for the upstream Space.gemma4-sg-guard.js
runtime guard (MIT).gemma4-vision.js),
the multimodal extension (gemma4-vision-inject.js), the 8 kernel patches,
and the test harness, as custom WebGPU kernels.vendor/python-packages/ — 49 pure-Python wheels
redistributed unmodified (MIT / BSD / Apache-2.0 / 0BSD / PSF-2.0, plus
pingouin GPL-3.0 and tqdm MPL-2.0 AND MIT). Per-package licenses:
vendor/python-packages/LICENSES.md.Upstream Space: https://huggingface.co/spaces/webml-community/gemma-4-webgpu-kernels
This repo: https://github.com//gemma-4-webgpu-kernels
The code written for this fork — the vision tower, multimodal extension, kernel patches, tools, and documentation — is released under the MIT License.
Third-party components keep their own terms; full attribution is in
THIRD_PARTY_NOTICES.md. Please note the upstream
Xenova engine (gemma-4-e2b.js, upstream index.html/landing.js) carries
no explicit license grant at the time of writing, and the model weights
are governed by Google's
Gemma Terms of Use.
JavaScript
84.4%
HTML
15.5%