Deep AI search for every photo and every frame of video in any folder on macOS
JavaScript
7
8 commits
updated Oct 4, 2026
Deep AI search for every photo and every frame of video in any folder on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac.
What makes it different
| Mode | Finds |
|---|---|
| Files | Whole photos/videos by meaning — vision rank with filename and phrase boosts |
| Scenes | Moments inside video — search a shot, jump to its timecode |
| OCR | Text visible in images and frames, matched literally (Tesseract; eng + 35 language toggles) |
| Dialogue | Exact spoken words in videos (Whisper), tiered exactness |
| LLMs (opt-in) | Local chat over the dialogue, OCR, and filenames your Mac already extracted — cited answers |
bun as package manager and
runnerbun installThe easiest install is via Homebrew (Apple Silicon, macOS 12+). The tap's cask clears the macOS quarantine flag automatically on every install and upgrade, so the app launches with no manual Gatekeeper steps:
brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm
Upgrades keep the same behavior:
brew upgrade --cask allenv0/scm/scm
Prefer least privilege? Trust just the cask instead of the whole tap:
brew tap allenv0/scm
brew trust --cask allenv0/scm/scm
brew install --cask scm
The tap lives at allenv0/homebrew-scm.
bun run dev # build the renderer bundle, then launch the Electron app
bun start # launch the Electron app without rebuilding
bun run build # just rebuild the renderer bundle into dist/
bun run dist # signed if an identity is in the keychain
bun run dist:unsigned # skip code-sign discovery
This runs two steps in sequence:
vite build — compiles the React renderer into dist/ (picks up all
changes under src/).electron-builder --mac — packages the app. It bundles the fresh
dist/ bundle together with main.js, preload.js, main-lib/, and
indexer/ (the file list is configured under build.files in
package.json), then produces the installers.Output: the installers land in dist-app/ (see build.directories.output
in package.json) — look for SCM-0.2.4.dmg and SCM-0.2.4.zip.
Typing starts an instant filename-keyword pre-pass, then the vision model takes over: results are scored by cosine similarity against image embeddings, with gated phrase and filename boosts, an honesty floor calibrated per model, and a near-duplicate diversity filter. Every tile carries a "why it matched" badge (Visual match / Filename match / …) and a hover tooltip with the per-component score breakdown. CJK queries search as overlapping bigrams ("台北車站" also matches 台北, 車站).
Every scene segment across all videos is scored, so a hit lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps straight to that moment. A noise gate returns "no scene match" instead of flooding the grid with gibberish, and each video contributes at most 3 scenes.
Matches the fraction of query tokens literally visible in each image's OCR text — the filename is ignored and no vision model is involved, so it works even while the AI engine is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox.
Exact literal retrieval over Whisper transcripts — no embeddings, no thresholds, works with the AI engine down. Results come in three tiers: Exact line (contiguous phrase in one utterance), Exact words (all words in one utterance or an ≤8s window), and Words spoken (all words in the same video). Matching words are highlighted in a speech snippet; opening a result seeks straight to the line.
Opt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A
llama.cpp sidecar bound to loopback answers your question from evidence the
app already extracted — dialogue lines, OCR text, and filename keyword hits
— with numbered citations you can click, streamed token-by-token with a live
tok/s readout. Leading /screenshots, /videos, /email narrow the
corpus; Stop keeps the partial answer; empty evidence short-circuits before
the model ever runs.
| Chat model | Size | Notes |
|---|---|---|
| Qwen3 1.7B (default) | ~1.1GB | Fast everyday chat; fits 8GB Macs |
| Llama 3.2 3B | ~2GB | Stronger long answers; needs headroom |
/screenshots, /videos,
/email narrow the corpus before the model ever runs.Surfaces photos whose visible OCR text contains an email address — an overlapping view (a photo keeps its category too). Detection is OCR-tolerant: it reassembles addresses Tesseract fractures across word boxes, and handles comma-for-dot noise ("gmail,com"), split TLDs ("gmail. com"), bracketed obfuscation ("allen [at] gmail [dot] com"), and dictated addresses ("allen at gmail dot com"). Tiles show a contact strip; expand it to copy or compose.
Screenshot classification is rename-proof. Four signals, in priority order: a manual override (right-click any tile) → filename vocabulary (30+ localized OS screenshot names in 20+ languages) → a PNG/JPEG metadata probe (reads "screenshot" from PNG text chunks / EXIF UserComment, so a renamed Bildschirmfoto still classifies) → source-folder hint. Everything else lands in Projects.
Four switchable models via ONNX Runtime; the active one is chosen per library:
| Model | Role | Speed (CPU) | Download |
|---|---|---|---|
| CLIP ViT-L/14@336 (default) | Best real-world video scene-search | ~480–570ms/img | ~435MB |
| SigLIP-2-B/16 | Fastest bulk import | ~50–100ms/img | ~412MB |
| SigLIP-2-L/16@256 | High-detail (1024-dim) — small objects, signs, on-screen text | ~200ms/img | ~850MB |
| SigLIP-B/16@384 | Maximum detail | ~480ms/img | ~214MB |
Switching models re-embeds the whole library: the flip lands instantly with the tail filled in the background, and search falls back to filename keywords until it completes. Per-model text-mean centering de-biases text embeddings so similarity scores stay honest across models.
ffmpeg scans each video for shot boundaries and builds a segment plan sampled to the density you pick in Settings → Video Search — each preset shows its measured time and disk cost before you commit:
| Preset | Seconds per point | Segment budget |
|---|---|---|
| Eco | 60 | 4–32 |
| Balanced (default) | 30 | 8–128 |
| Detailed | 15 | 12–256 |
| Ultra | 5 | 16–1024 |
| Ultra Pro | 2.5 | 24–2048 (confirm required) |
Each segment embeds its midpoint frame and keeps a poster; shot plans are cached per file (path + size + mtime + config fingerprint), so re-imports skip detection entirely.
Dialogue transcription: Whisper tiny.en (~150MB, default) or base.en
(~300MB) — switching re-transcribes every video. Whole videos embed three
frames (20/50/80%) averaged; GIFs embed an average of middle frames.
Tesseract runs in its own worker, separate from the vision model. English is always on; 35 more languages are toggleable in Settings → Photo Search (default: Simplified + Traditional Chinese, Japanese, Korean). Each language pack downloads once (~2.4–5MB; ~17MB for the default set), then everything is offline. Word boxes are stored with the text so matches highlight in place; CJK text is joined without spaces and email fragments fractured across word boxes are reassembled.
fs.watch plus a re-sync at every launch).
Problem files retry up to 3 times, then sit out watch-syncs until they
change.~/Library/Application Support/scm
(MEMORIES_DATA_DIR overrides it): the index JSON, Float32 embedding bins
per model, scene and transcript sidecars, thumbnails, and posters.app:// bundle — contextIsolation, OS
sandbox, and a CSP pinned to 'self' (+ Google Fonts CDN for display
type, with a monospace fallback when offline).A macOS-style settings sheet with ten panels (Library, Appearance, Grid, Smart Tabs, Photo Search, Video Search, LLMs Chat, Global Shortcut, Keyboard, Menu Bar). Around the core: light/dark/system theme, a CRT screen effect for the lightbox, 6 alternate app icons, menu-bar-only mode, a recordable global shortcut, a first-run onboarding tour, background-work trays (scenes / transcripts / OCR) with a global pause, and a status bar with version and indexed-video counts.
main.js, preload.js,
main-lib/, and indexer/ are packaged from source (not cached), a
bun run dist after editing any of them produces a fresh package.bun run test:all # the full verification battery (unit + smoke suites)
bun run smoke:indexer # headless smoke test of the CLIP indexer
bun test test/ # unit tests (pure modules, no Electron needed)
bun run lint # eslint
bun run format:check # prettier
test:* scripts in package.json).ELECTRON_SMOKE_*
environment variables — a dozen-plus scenarios from boot/protocol checks
to search matrices, model migration, and Ask mode (drivers in
scripts/e2e/, dispatched in main.js; e.g. bun run smoke:ask,
smoke:deep, smoke:grid).bun run bench:inference, bench:enrich, and
bench:detect write JSON reports into MDs/bench-*.main.js Electron main process: library, IPC, indexer worker pool, app:// protocol
preload.js contextBridge — exposes window.memories to the renderer
main-lib/ main-process modules split out of main.js (settings, library store,
rank search, Ask retrieval, LLM sidecar config, embedding versions, …)
indexer/ vision/OCR/ASR workers (utility processes) + video utils + model registry
src/ React renderer (grid, search modes, lightbox, tabs, settings, onboarding)
scripts/ bench scripts, E2E drivers, the smoke battery (smoke-all.sh)
test/ unit + integration tests
MDs/ design docs, bench reports, plans
dist/ vite build output (renderer bundle)
dist-app/ electron-builder output (DMG / ZIP)
JavaScript
64.0%
TypeScript
28.0%
CSS
5.6%
Deep AI search for every photo and every frame of video in any folder on macOS
JavaScript
7
8 commits
updated Oct 4, 2026
Deep AI search for every photo and every frame of video in any folder on macOS. Local-first — no accounts, no cloud, no uploads. Inference runs on your Mac.
What makes it different
| Mode | Finds |
|---|---|
| Files | Whole photos/videos by meaning — vision rank with filename and phrase boosts |
| Scenes | Moments inside video — search a shot, jump to its timecode |
| OCR | Text visible in images and frames, matched literally (Tesseract; eng + 35 language toggles) |
| Dialogue | Exact spoken words in videos (Whisper), tiered exactness |
| LLMs (opt-in) | Local chat over the dialogue, OCR, and filenames your Mac already extracted — cited answers |
bun as package manager and
runnerbun installThe easiest install is via Homebrew (Apple Silicon, macOS 12+). The tap's cask clears the macOS quarantine flag automatically on every install and upgrade, so the app launches with no manual Gatekeeper steps:
brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm
Upgrades keep the same behavior:
brew upgrade --cask allenv0/scm/scm
Prefer least privilege? Trust just the cask instead of the whole tap:
brew tap allenv0/scm
brew trust --cask allenv0/scm/scm
brew install --cask scm
The tap lives at allenv0/homebrew-scm.
bun run dev # build the renderer bundle, then launch the Electron app
bun start # launch the Electron app without rebuilding
bun run build # just rebuild the renderer bundle into dist/
bun run dist # signed if an identity is in the keychain
bun run dist:unsigned # skip code-sign discovery
This runs two steps in sequence:
vite build — compiles the React renderer into dist/ (picks up all
changes under src/).electron-builder --mac — packages the app. It bundles the fresh
dist/ bundle together with main.js, preload.js, main-lib/, and
indexer/ (the file list is configured under build.files in
package.json), then produces the installers.Output: the installers land in dist-app/ (see build.directories.output
in package.json) — look for SCM-0.2.4.dmg and SCM-0.2.4.zip.
Typing starts an instant filename-keyword pre-pass, then the vision model takes over: results are scored by cosine similarity against image embeddings, with gated phrase and filename boosts, an honesty floor calibrated per model, and a near-duplicate diversity filter. Every tile carries a "why it matched" badge (Visual match / Filename match / …) and a hover tooltip with the per-component score breakdown. CJK queries search as overlapping bigrams ("台北車站" also matches 台北, 車站).
Every scene segment across all videos is scored, so a hit lands on the exact shot: tiles show the scene poster with a timecode badge, and opening the video jumps straight to that moment. A noise gate returns "no scene match" instead of flooding the grid with gibberish, and each video contributes at most 3 scenes.
Matches the fraction of query tokens literally visible in each image's OCR text — the filename is ignored and no vision model is involved, so it works even while the AI engine is warming up or offline. Matched words are boxed in amber on tiles and in the lightbox.
Exact literal retrieval over Whisper transcripts — no embeddings, no thresholds, works with the AI engine down. Results come in three tiers: Exact line (contiguous phrase in one utterance), Exact words (all words in one utterance or an ≤8s window), and Words spoken (all words in the same video). Matching words are highlighted in a speech snippet; opening a result seeks straight to the line.
Opt-in — nothing downloads or runs until enabled in Settings → LLMs Chat. A
llama.cpp sidecar bound to loopback answers your question from evidence the
app already extracted — dialogue lines, OCR text, and filename keyword hits
— with numbered citations you can click, streamed token-by-token with a live
tok/s readout. Leading /screenshots, /videos, /email narrow the
corpus; Stop keeps the partial answer; empty evidence short-circuits before
the model ever runs.
| Chat model | Size | Notes |
|---|---|---|
| Qwen3 1.7B (default) | ~1.1GB | Fast everyday chat; fits 8GB Macs |
| Llama 3.2 3B | ~2GB | Stronger long answers; needs headroom |
/screenshots, /videos,
/email narrow the corpus before the model ever runs.Surfaces photos whose visible OCR text contains an email address — an overlapping view (a photo keeps its category too). Detection is OCR-tolerant: it reassembles addresses Tesseract fractures across word boxes, and handles comma-for-dot noise ("gmail,com"), split TLDs ("gmail. com"), bracketed obfuscation ("allen [at] gmail [dot] com"), and dictated addresses ("allen at gmail dot com"). Tiles show a contact strip; expand it to copy or compose.
Screenshot classification is rename-proof. Four signals, in priority order: a manual override (right-click any tile) → filename vocabulary (30+ localized OS screenshot names in 20+ languages) → a PNG/JPEG metadata probe (reads "screenshot" from PNG text chunks / EXIF UserComment, so a renamed Bildschirmfoto still classifies) → source-folder hint. Everything else lands in Projects.
Four switchable models via ONNX Runtime; the active one is chosen per library:
| Model | Role | Speed (CPU) | Download |
|---|---|---|---|
| CLIP ViT-L/14@336 (default) | Best real-world video scene-search | ~480–570ms/img | ~435MB |
| SigLIP-2-B/16 | Fastest bulk import | ~50–100ms/img | ~412MB |
| SigLIP-2-L/16@256 | High-detail (1024-dim) — small objects, signs, on-screen text | ~200ms/img | ~850MB |
| SigLIP-B/16@384 | Maximum detail | ~480ms/img | ~214MB |
Switching models re-embeds the whole library: the flip lands instantly with the tail filled in the background, and search falls back to filename keywords until it completes. Per-model text-mean centering de-biases text embeddings so similarity scores stay honest across models.
ffmpeg scans each video for shot boundaries and builds a segment plan sampled to the density you pick in Settings → Video Search — each preset shows its measured time and disk cost before you commit:
| Preset | Seconds per point | Segment budget |
|---|---|---|
| Eco | 60 | 4–32 |
| Balanced (default) | 30 | 8–128 |
| Detailed | 15 | 12–256 |
| Ultra | 5 | 16–1024 |
| Ultra Pro | 2.5 | 24–2048 (confirm required) |
Each segment embeds its midpoint frame and keeps a poster; shot plans are cached per file (path + size + mtime + config fingerprint), so re-imports skip detection entirely.
Dialogue transcription: Whisper tiny.en (~150MB, default) or base.en
(~300MB) — switching re-transcribes every video. Whole videos embed three
frames (20/50/80%) averaged; GIFs embed an average of middle frames.
Tesseract runs in its own worker, separate from the vision model. English is always on; 35 more languages are toggleable in Settings → Photo Search (default: Simplified + Traditional Chinese, Japanese, Korean). Each language pack downloads once (~2.4–5MB; ~17MB for the default set), then everything is offline. Word boxes are stored with the text so matches highlight in place; CJK text is joined without spaces and email fragments fractured across word boxes are reassembled.
fs.watch plus a re-sync at every launch).
Problem files retry up to 3 times, then sit out watch-syncs until they
change.~/Library/Application Support/scm
(MEMORIES_DATA_DIR overrides it): the index JSON, Float32 embedding bins
per model, scene and transcript sidecars, thumbnails, and posters.app:// bundle — contextIsolation, OS
sandbox, and a CSP pinned to 'self' (+ Google Fonts CDN for display
type, with a monospace fallback when offline).A macOS-style settings sheet with ten panels (Library, Appearance, Grid, Smart Tabs, Photo Search, Video Search, LLMs Chat, Global Shortcut, Keyboard, Menu Bar). Around the core: light/dark/system theme, a CRT screen effect for the lightbox, 6 alternate app icons, menu-bar-only mode, a recordable global shortcut, a first-run onboarding tour, background-work trays (scenes / transcripts / OCR) with a global pause, and a status bar with version and indexed-video counts.
main.js, preload.js,
main-lib/, and indexer/ are packaged from source (not cached), a
bun run dist after editing any of them produces a fresh package.bun run test:all # the full verification battery (unit + smoke suites)
bun run smoke:indexer # headless smoke test of the CLIP indexer
bun test test/ # unit tests (pure modules, no Electron needed)
bun run lint # eslint
bun run format:check # prettier
test:* scripts in package.json).ELECTRON_SMOKE_*
environment variables — a dozen-plus scenarios from boot/protocol checks
to search matrices, model migration, and Ask mode (drivers in
scripts/e2e/, dispatched in main.js; e.g. bun run smoke:ask,
smoke:deep, smoke:grid).bun run bench:inference, bench:enrich, and
bench:detect write JSON reports into MDs/bench-*.main.js Electron main process: library, IPC, indexer worker pool, app:// protocol
preload.js contextBridge — exposes window.memories to the renderer
main-lib/ main-process modules split out of main.js (settings, library store,
rank search, Ask retrieval, LLM sidecar config, embedding versions, …)
indexer/ vision/OCR/ASR workers (utility processes) + video utils + model registry
src/ React renderer (grid, search modes, lightbox, tabs, settings, onboarding)
scripts/ bench scripts, E2E drivers, the smoke battery (smoke-all.sh)
test/ unit + integration tests
MDs/ design docs, bench reports, plans
dist/ vite build output (renderer bundle)
dist-app/ electron-builder output (DMG / ZIP)
JavaScript
64.0%
TypeScript
28.0%
CSS
5.6%