webml-community/semantic-image-field

Space

Semantic Image Field

15

6 commits

1 linked in READMEs

updated Sep 22, 2026

See the code

README

Semantic Image Field

A static Vite demo inspired by Shridhar Rathi's animated image search, using the stack and search pipeline from Hugging Face's semantic-image-search-web.

Run

Use Node.js 20.19+ or 22.12+ (Node 24 recommended).

npm install
npm run dev
npm run build
npm run preview

The downloadable ZIP includes both source and a prebuilt dist/ folder.

Deploy the contents of dist/ to any static host. There is no server, API key, account, database service, or upload endpoint. Vite uses relative asset paths, so the build can also live under a subdirectory. Use HTTP/HTTPS rather than opening index.html through file://.

Stack

  • Transformers.js 4.3.0, pinned exactly.
  • React 19.3, Vite 8.3, Tailwind CSS 4.3, BlurHash: the same libraries as the original example, upgraded to stable releases.
  • Xenova/clip-vit-base-patch16, AutoTokenizer, and CLIPTextModelWithProjection in a module Web Worker.
  • Explicit q8 weights and WASM execution, matching the original example's quantized CPU path. WebGPU is not required.
  • The original 24,989-image metadata and 512-dimensional float32 embeddings from Xenova/semantic-image-search-assets.
  • ESLint and @eslint/js stay on 9.39.5 for compatibility with eslint-plugin-react 7.37.5, whose peer range does not yet include ESLint 10. All other direct dependencies were current at the upgrade check.
  • No animation library. One requestAnimationFrame loop drives continuous curved paths and subtle shared pointer parallax.

How search works

  1. The initial field displays a deterministic sample of 96 larger photographs (56 on narrow screens) from the original dataset.
  2. After the first render, a worker loads the CLIP text encoder, tokenizer, metadata, and approximately 49 MiB of image embeddings. It warms up the model before marking search ready.
  3. Input is sent immediately; the worker coalesces queued requests so the newest query wins. CLIP embeds the text locally; cosine similarity ranks the full collection. The database vectors are normalized once, avoiding allocations for each comparison.
  4. The closest 60 photographs enter a single shared population of 96 images (56 on narrow screens, showing the first 56 matches). Existing photos are retained where possible. Every image uses the same coordinate plane and identity-based stacking order. A normalized sigmoid maps current-query cosine scores to opacity: a broad bright shoulder keeps good matches clear, with a steeper fade and slight blur for the weak tail; retained images outside the result set receive the softest treatment. The strongest matches remain crisp, and hovering or keyboard focusing a photo reveals it. Relevance changes position, with gently larger images toward the center; there is no separate background/results overlay.
  5. Each image retains its DOM node and motion state while it remains in the scene. Quintic curves preserve position and velocity if a query interrupts the movement. Movement starts as soon as search returns, without waiting for photo downloads. BlurHash placeholders cover incoming assets, which fade in as they load. Images glide in from the edges while departing photographs flow out; transitions finish in 600–950 ms. Clearing search restores the original arrangement.
  6. A separate worker generates BlurHash placeholders, keeping decoding and PNG encoding off the animation thread. Thumbnails use 256–512 px widths based on display size and pixel density. A priority queue loads and decodes up to six images at once, prioritizing the strongest matches, and attaches at most two decoded photos per frame. A bounded cache retains 64 detached, decoded images for repeat searches. The full-screen viewer reuses the existing thumbnail while its high-resolution image loads. First-time photo downloads still depend on network speed.

This reproduces the reference's interaction and visual composition with the original Transformers.js example's photographic dataset. It does not include the reference video's private/design image collection or claim identical results.

Interaction

  • Type a word, color, feeling, or phrase. Try grass, blue, red flowers, misty forest, or interior.
  • Click a photograph to open the full-screen viewer. The photo zooms from its position in the field, with a high-resolution image loading over the existing thumbnail. Click the image to zoom further, and scroll to explore it. Close using Escape, the close button, or the backdrop.
  • Press / to focus search, and Escape to close the full-screen viewer or clear search.
  • The first 12 ranked matches are keyboard-focusable. The native dialog traps focus and returns it on close. The field pauses while viewing a photo.
  • Reduced-motion preferences disable transitions, drift, parallax, and fades.

Network and privacy

Inference and ranking happen on the user's device. Search text is not sent to an inference API. The browser downloads model files and embedding data from Hugging Face, ONNX Runtime from its version-matched jsDelivr URL, and photos from Unsplash. These ordinary asset requests still contact those providers. First use requires an internet connection; model/database assets are cached where browser storage is available. This is not a fully offline photo archive.

The Vite config selects ONNX Runtime's supported onnxruntime-web-use-extern-wasm export condition. Transformers.js manages the version-matched runtime download/cache; this avoids bundling an unused WASM binary and keeps static-host uploads small.

Source map

FilePurpose
src/App.jsxQuery state, worker lifecycle, loading/errors, photo selection
src/PhotoViewer.jsxFull-screen photo viewer, zoom, and focus restoration
src/relevance.jsSigmoid score-to-opacity mapping and blur, without stale query scores
src/worker.jsCLIP initialization, latest-query queue, inference
src/search.jsNormalization and cosine ranking
src/ImageField.jsxImage lifecycle and shared scene population
src/layout.jsDeterministic scatter and unified ranked layouts
src/motion.jsFrame-rate-independent curves with velocity-preserving retargeting
src/assets.jsImage URLs and resilient database caching
src/image-loader.jsPrioritized thumbnail loading, decoded-image cache, placeholder worker client
src/placeholder.worker.jsOff-thread BlurHash decoding and PNG encoding
src/collection.jsonInitial sample from the original metadata
src/index.cssAppearance, responsive rules, reduced motion

An optional, feature-detected search_images WebMCP action calls the same search flow. It has no effect in browsers without document.modelContext.

To use a different collection, replace the metadata, initial sample, and image embeddings together. Encode the images using the same CLIP checkpoint as the text encoder, preserve metadata/embedding row alignment, and update DIMENSIONS if necessary. Replacing only the photos or assigning textual tags is not equivalent to semantic image search.

Photo rights remain with their respective Unsplash contributors. The reference is credited for its visual inspiration; the inference pipeline is based on the Hugging Face example.

static

webml-community/semantic-image-field

Space

Semantic Image Field

15

6 commits

1 linked in READMEs

updated Sep 22, 2026

See the code

README

Semantic Image Field

A static Vite demo inspired by Shridhar Rathi's animated image search, using the stack and search pipeline from Hugging Face's semantic-image-search-web.

Run

Use Node.js 20.19+ or 22.12+ (Node 24 recommended).

npm install
npm run dev
npm run build
npm run preview

The downloadable ZIP includes both source and a prebuilt dist/ folder.

Deploy the contents of dist/ to any static host. There is no server, API key, account, database service, or upload endpoint. Vite uses relative asset paths, so the build can also live under a subdirectory. Use HTTP/HTTPS rather than opening index.html through file://.

Stack

  • Transformers.js 4.3.0, pinned exactly.
  • React 19.3, Vite 8.3, Tailwind CSS 4.3, BlurHash: the same libraries as the original example, upgraded to stable releases.
  • Xenova/clip-vit-base-patch16, AutoTokenizer, and CLIPTextModelWithProjection in a module Web Worker.
  • Explicit q8 weights and WASM execution, matching the original example's quantized CPU path. WebGPU is not required.
  • The original 24,989-image metadata and 512-dimensional float32 embeddings from Xenova/semantic-image-search-assets.
  • ESLint and @eslint/js stay on 9.39.5 for compatibility with eslint-plugin-react 7.37.5, whose peer range does not yet include ESLint 10. All other direct dependencies were current at the upgrade check.
  • No animation library. One requestAnimationFrame loop drives continuous curved paths and subtle shared pointer parallax.

How search works

  1. The initial field displays a deterministic sample of 96 larger photographs (56 on narrow screens) from the original dataset.
  2. After the first render, a worker loads the CLIP text encoder, tokenizer, metadata, and approximately 49 MiB of image embeddings. It warms up the model before marking search ready.
  3. Input is sent immediately; the worker coalesces queued requests so the newest query wins. CLIP embeds the text locally; cosine similarity ranks the full collection. The database vectors are normalized once, avoiding allocations for each comparison.
  4. The closest 60 photographs enter a single shared population of 96 images (56 on narrow screens, showing the first 56 matches). Existing photos are retained where possible. Every image uses the same coordinate plane and identity-based stacking order. A normalized sigmoid maps current-query cosine scores to opacity: a broad bright shoulder keeps good matches clear, with a steeper fade and slight blur for the weak tail; retained images outside the result set receive the softest treatment. The strongest matches remain crisp, and hovering or keyboard focusing a photo reveals it. Relevance changes position, with gently larger images toward the center; there is no separate background/results overlay.
  5. Each image retains its DOM node and motion state while it remains in the scene. Quintic curves preserve position and velocity if a query interrupts the movement. Movement starts as soon as search returns, without waiting for photo downloads. BlurHash placeholders cover incoming assets, which fade in as they load. Images glide in from the edges while departing photographs flow out; transitions finish in 600–950 ms. Clearing search restores the original arrangement.
  6. A separate worker generates BlurHash placeholders, keeping decoding and PNG encoding off the animation thread. Thumbnails use 256–512 px widths based on display size and pixel density. A priority queue loads and decodes up to six images at once, prioritizing the strongest matches, and attaches at most two decoded photos per frame. A bounded cache retains 64 detached, decoded images for repeat searches. The full-screen viewer reuses the existing thumbnail while its high-resolution image loads. First-time photo downloads still depend on network speed.

This reproduces the reference's interaction and visual composition with the original Transformers.js example's photographic dataset. It does not include the reference video's private/design image collection or claim identical results.

Interaction

  • Type a word, color, feeling, or phrase. Try grass, blue, red flowers, misty forest, or interior.
  • Click a photograph to open the full-screen viewer. The photo zooms from its position in the field, with a high-resolution image loading over the existing thumbnail. Click the image to zoom further, and scroll to explore it. Close using Escape, the close button, or the backdrop.
  • Press / to focus search, and Escape to close the full-screen viewer or clear search.
  • The first 12 ranked matches are keyboard-focusable. The native dialog traps focus and returns it on close. The field pauses while viewing a photo.
  • Reduced-motion preferences disable transitions, drift, parallax, and fades.

Network and privacy

Inference and ranking happen on the user's device. Search text is not sent to an inference API. The browser downloads model files and embedding data from Hugging Face, ONNX Runtime from its version-matched jsDelivr URL, and photos from Unsplash. These ordinary asset requests still contact those providers. First use requires an internet connection; model/database assets are cached where browser storage is available. This is not a fully offline photo archive.

The Vite config selects ONNX Runtime's supported onnxruntime-web-use-extern-wasm export condition. Transformers.js manages the version-matched runtime download/cache; this avoids bundling an unused WASM binary and keeps static-host uploads small.

Source map

FilePurpose
src/App.jsxQuery state, worker lifecycle, loading/errors, photo selection
src/PhotoViewer.jsxFull-screen photo viewer, zoom, and focus restoration
src/relevance.jsSigmoid score-to-opacity mapping and blur, without stale query scores
src/worker.jsCLIP initialization, latest-query queue, inference
src/search.jsNormalization and cosine ranking
src/ImageField.jsxImage lifecycle and shared scene population
src/layout.jsDeterministic scatter and unified ranked layouts
src/motion.jsFrame-rate-independent curves with velocity-preserving retargeting
src/assets.jsImage URLs and resilient database caching
src/image-loader.jsPrioritized thumbnail loading, decoded-image cache, placeholder worker client
src/placeholder.worker.jsOff-thread BlurHash decoding and PNG encoding
src/collection.jsonInitial sample from the original metadata
src/index.cssAppearance, responsive rules, reduced motion

An optional, feature-detected search_images WebMCP action calls the same search flow. It has no effect in browsers without document.modelContext.

To use a different collection, replace the metadata, initial sample, and image embeddings together. Encode the images using the same CLIP checkpoint as the text encoder, preserve metadata/embedding row alignment, and update DIMENSIONS if necessary. Replacing only the photos or assigning textual tags is not equivalent to semantic image search.

Photo rights remain with their respective Unsplash contributors. The reference is credited for its visual inspiration; the inference pipeline is based on the Hugging Face example.

static