A static Vite demo inspired by Shridhar Rathi's animated image search, using the stack and search pipeline from Hugging Face's semantic-image-search-web.
Use Node.js 20.19+ or 22.12+ (Node 24 recommended).
npm install
npm run dev
npm run build
npm run preview
The downloadable ZIP includes both source and a prebuilt dist/ folder.
Deploy the contents of dist/ to any static host. There is no server, API key, account, database service, or upload endpoint. Vite uses relative asset paths, so the build can also live under a subdirectory. Use HTTP/HTTPS rather than opening index.html through file://.
Xenova/clip-vit-base-patch16, AutoTokenizer, and CLIPTextModelWithProjection in a module Web Worker.q8 weights and WASM execution, matching the original example's quantized CPU path. WebGPU is not required.Xenova/semantic-image-search-assets.@eslint/js stay on 9.39.5 for compatibility with eslint-plugin-react 7.37.5, whose peer range does not yet include ESLint 10. All other direct dependencies were current at the upgrade check.This reproduces the reference's interaction and visual composition with the original Transformers.js example's photographic dataset. It does not include the reference video's private/design image collection or claim identical results.
grass, blue, red flowers, misty forest, or interior./ to focus search, and Escape to close the full-screen viewer or clear search.Inference and ranking happen on the user's device. Search text is not sent to an inference API. The browser downloads model files and embedding data from Hugging Face, ONNX Runtime from its version-matched jsDelivr URL, and photos from Unsplash. These ordinary asset requests still contact those providers. First use requires an internet connection; model/database assets are cached where browser storage is available. This is not a fully offline photo archive.
The Vite config selects ONNX Runtime's supported onnxruntime-web-use-extern-wasm export condition. Transformers.js manages the version-matched runtime download/cache; this avoids bundling an unused WASM binary and keeps static-host uploads small.
| File | Purpose |
|---|---|
src/App.jsx | Query state, worker lifecycle, loading/errors, photo selection |
src/PhotoViewer.jsx | Full-screen photo viewer, zoom, and focus restoration |
src/relevance.js | Sigmoid score-to-opacity mapping and blur, without stale query scores |
src/worker.js | CLIP initialization, latest-query queue, inference |
src/search.js | Normalization and cosine ranking |
src/ImageField.jsx | Image lifecycle and shared scene population |
src/layout.js | Deterministic scatter and unified ranked layouts |
src/motion.js | Frame-rate-independent curves with velocity-preserving retargeting |
src/assets.js | Image URLs and resilient database caching |
src/image-loader.js | Prioritized thumbnail loading, decoded-image cache, placeholder worker client |
src/placeholder.worker.js | Off-thread BlurHash decoding and PNG encoding |
src/collection.json | Initial sample from the original metadata |
src/index.css | Appearance, responsive rules, reduced motion |
An optional, feature-detected search_images WebMCP action calls the same search flow. It has no effect in browsers without document.modelContext.
To use a different collection, replace the metadata, initial sample, and image embeddings together. Encode the images using the same CLIP checkpoint as the text encoder, preserve metadata/embedding row alignment, and update DIMENSIONS if necessary. Replacing only the photos or assigning textual tags is not equivalent to semantic image search.
Photo rights remain with their respective Unsplash contributors. The reference is credited for its visual inspiration; the inference pipeline is based on the Hugging Face example.
A static Vite demo inspired by Shridhar Rathi's animated image search, using the stack and search pipeline from Hugging Face's semantic-image-search-web.
Use Node.js 20.19+ or 22.12+ (Node 24 recommended).
npm install
npm run dev
npm run build
npm run preview
The downloadable ZIP includes both source and a prebuilt dist/ folder.
Deploy the contents of dist/ to any static host. There is no server, API key, account, database service, or upload endpoint. Vite uses relative asset paths, so the build can also live under a subdirectory. Use HTTP/HTTPS rather than opening index.html through file://.
Xenova/clip-vit-base-patch16, AutoTokenizer, and CLIPTextModelWithProjection in a module Web Worker.q8 weights and WASM execution, matching the original example's quantized CPU path. WebGPU is not required.Xenova/semantic-image-search-assets.@eslint/js stay on 9.39.5 for compatibility with eslint-plugin-react 7.37.5, whose peer range does not yet include ESLint 10. All other direct dependencies were current at the upgrade check.This reproduces the reference's interaction and visual composition with the original Transformers.js example's photographic dataset. It does not include the reference video's private/design image collection or claim identical results.
grass, blue, red flowers, misty forest, or interior./ to focus search, and Escape to close the full-screen viewer or clear search.Inference and ranking happen on the user's device. Search text is not sent to an inference API. The browser downloads model files and embedding data from Hugging Face, ONNX Runtime from its version-matched jsDelivr URL, and photos from Unsplash. These ordinary asset requests still contact those providers. First use requires an internet connection; model/database assets are cached where browser storage is available. This is not a fully offline photo archive.
The Vite config selects ONNX Runtime's supported onnxruntime-web-use-extern-wasm export condition. Transformers.js manages the version-matched runtime download/cache; this avoids bundling an unused WASM binary and keeps static-host uploads small.
| File | Purpose |
|---|---|
src/App.jsx | Query state, worker lifecycle, loading/errors, photo selection |
src/PhotoViewer.jsx | Full-screen photo viewer, zoom, and focus restoration |
src/relevance.js | Sigmoid score-to-opacity mapping and blur, without stale query scores |
src/worker.js | CLIP initialization, latest-query queue, inference |
src/search.js | Normalization and cosine ranking |
src/ImageField.jsx | Image lifecycle and shared scene population |
src/layout.js | Deterministic scatter and unified ranked layouts |
src/motion.js | Frame-rate-independent curves with velocity-preserving retargeting |
src/assets.js | Image URLs and resilient database caching |
src/image-loader.js | Prioritized thumbnail loading, decoded-image cache, placeholder worker client |
src/placeholder.worker.js | Off-thread BlurHash decoding and PNG encoding |
src/collection.json | Initial sample from the original metadata |
src/index.css | Appearance, responsive rules, reduced motion |
An optional, feature-detected search_images WebMCP action calls the same search flow. It has no effect in browsers without document.modelContext.
To use a different collection, replace the metadata, initial sample, and image embeddings together. Encode the images using the same CLIP checkpoint as the text encoder, preserve metadata/embedding row alignment, and update DIMENSIONS if necessary. Replacing only the photos or assigning textual tags is not equivalent to semantic image search.
Photo rights remain with their respective Unsplash contributors. The reference is credited for its visual inspiration; the inference pipeline is based on the Hugging Face example.