SCM for macOS in 2026: Free Offline AI Search for Every Photo and Video Frame — and Which Mac You Actually Need
TL;DR: SCM is a free, MIT-licensed macOS app that indexes every photo and every frame of video in any folder, then searches them by plain-English description — fully offline, on your own Mac. Inference runs on the CPU via ONNX Runtime, so any Apple Silicon Mac handles it; there is no reason to buy hardware for this tool. The only real costs are first-index time and disk space.
| SCM on the Mac you own | Mac Mini M6 16GB as a dedicated appliance | A bigger Mac “for AI search” | |
|---|---|---|---|
| Best for | Everyone — it’s free and CPU-bound | Always-on indexer + home server duty | Nobody, for this workload |
| Price / Cost | $0 (MIT license) | $899 (Apple list, Oct 2026) | $1,699+ |
| The catch | First index of 50K photos takes 42 min–8 hrs depending on model choice | Indexing speed, not memory, is the constraint — 16GB is plenty | SCM doesn’t use the GPU or Neural Engine, so extra bandwidth buys nothing here |
Honest take: install SCM on whatever Apple Silicon Mac you already have — it was designed for exactly that. If a purchase is happening anyway, buy the Mac for your other local AI workloads and treat SCM as a free bonus; this app alone justifies zero dollars of new hardware.
Private photo libraries are exactly the data people least want in a cloud index, and until recently the local options were thin: Apple Photos searches only its own library, Spotlight doesn’t understand image content, and the credible open-source tools indexed photos but not video. SCM — an open-source project by GitHub user allenv0 that hit the Hacker News front page in the first week of October 2026 — closes that gap: it points at any folder, splits videos into shots, and lets you type “kid blowing out birthday candles” or “whiteboard with the Q3 roadmap” and land on the exact frame.
We dug through the repo and its documentation to answer the question this site exists for: what hardware does it actually need, and does it change what Mac you should buy? The short answers: surprisingly little, and no. The longer answers involve some genuinely interesting engineering choices — including one that will annoy Neural Engine optimists.
What SCM actually is
SCM (“Screen Memories,” github.com/allenv0/SCM) is an Electron app — main process, React renderer, electron-builder packaging — released under the MIT license. As of October 9, 2026 the repo sits at 448 stars, with the build docs referencing v0.2.4 artifacts, so this is an early-stage project, not a polished commercial product. The Homebrew install targets Apple Silicon Macs on macOS 12 or later; Intel Macs aren’t mentioned anywhere in the install docs, so treat them as unsupported.
It ships five search modes, each backed by a different local model:
- Files — whole photos and videos matched by meaning, via CLIP/SigLIP embeddings
- Scenes — specific moments inside videos; ffmpeg splits each file at shot boundaries so results land on the matching shot, not just the filename
- OCR — text visible in images and frames, via Tesseract (English always on, 35 more languages toggleable; the default set adds Simplified and Traditional Chinese, Japanese, and Korean for ~17MB)
- Dialogue — spoken words, via Whisper transcription (
tiny.en, ~150MB, default;base.en, ~300MB, optional) - LLMs — an opt-in “Ask” chat over the extracted evidence with citations, served by a llama.cpp sidecar running Qwen3 1.7B (~1.1GB) or Llama 3.2 3B (~2GB)
Nothing downloads until the mode that needs it is first used, every model download is sha256-verified, and the chat sidecar stays completely off until you enable it in Settings.
The indexing engine, with real numbers
The vision model is the knob that matters. SCM offers four, one active per library, and the repo publishes per-image CPU embedding times for each:
| Vision model | Download | Per-image embed (CPU) | Notes |
|---|---|---|---|
| CLIP ViT-L/14@336 | ~435MB | ~480–570 ms | Default; strongest retrieval quality |
| SigLIP-2-B/16 | ~412MB | ~50–100 ms | ”Fastest bulk import” per the repo |
| SigLIP-2-L/16@256 | ~850MB | ~200 ms | 1024-dim embeddings |
| SigLIP-B/16@384 | ~214MB | ~480 ms | Legacy option |
Those timings are the repo’s own CPU figures (it doesn’t state which chip produced them), but the ratios are what matter: the default CLIP model is roughly 5–10× slower per image than SigLIP-2-B/16.
Scale that to the library sizes real people have, and the model choice decides whether first indexing is a coffee break or an overnight job. For a 50,000-photo library — our math from the repo’s per-image figures, not a measured benchmark:
- SigLIP-2-B/16: 50,000 × 50–100 ms ≈ 42–83 minutes
- CLIP ViT-L/14@336: 50,000 × 480–570 ms ≈ 6.7–7.9 hours
Video is where the design gets clever. Instead of embedding every frame (a 2-hour movie at 30 fps is 216,000 frames — nobody has that CPU budget), ffmpeg detects shot boundaries, then a sampling preset budgets how many segments each file gets: Eco (one sample per 60 seconds, 4–32 segments per file), Balanced (30s, 8–128, the default), Detailed (15s, 12–256), Ultra (5s, 16–1,024), and Ultra Pro (2.5s, 24–2,048, gated behind a confirmation dialog). Each segment embeds its midpoint frame and keeps a poster image. At Balanced, a 2-hour video caps out at 128 segments — the embedding cost of 128 photos, a few seconds of CPU on the fast model. Shot plans are cached per file (keyed on path, size, mtime, and config), so re-imports skip detection entirely.
The index itself is refreshingly boring: no database, just a JSON index, raw Float32 embedding binaries per model, transcript and scene sidecars, and thumbnails under ~/Library/Application Support/scm (relocatable via MEMORIES_DATA_DIR). Embeddings are small — at CLIP’s 768 dimensions, 50,000 photos work out to roughly 150MB of Float32 vectors (our arithmetic: 50,000 × 768 × 4 bytes). Thumbnails and video posters will dominate the actual disk cost, and the app shows each preset’s disk estimate in Settings before you commit.
Install and first run
brew tap allenv0/scm
brew trust allenv0/scm
brew install --cask allenv0/scm/scm
The cask clears macOS’s quarantine flag, so the app opens normally. A real problem you’ll hit if you skip Homebrew and build the DMG yourself (bun run dist): without an Apple Developer ID in your keychain the build is unsigned, and macOS will refuse to open it with a “cannot be opened because the developer cannot be verified” dialog. The fix is the standard one — right-click the app → Open → Open — and it’s documented on the project site, but it catches people every time. Stick with the cask and you never see it.
After the onboarding tour, you add media with ⌘I, drag-and-drop, or a watched folder, and the first use of each search mode pulls its model from Hugging Face or the Tesseract CDN. Plan for that: “offline” means offline after roughly 600MB–1GB of one-time downloads (vision model + Whisper + OCR packs), which is worth knowing before a flight.
The hardware question: it’s a CPU app, and that’s the story
Here’s the part that matters for this site’s readers. The queue of coverage around SCM’s launch assumed what everyone assumes about Mac AI apps: that it rides the Neural Engine, and that a bigger Mac means proportionally faster indexing. The repo says otherwise. SCM runs its vision models through ONNX Runtime with published CPU timings — nowhere in the documentation does it claim CoreML, Neural Engine, or GPU execution. The llama.cpp chat sidecar can use Metal, but the indexing pipeline that does 99% of the work is CPU inference.
That has three practical consequences:
1. Memory bandwidth is irrelevant here. The spec this site tells you to buy Macs for — 153 GB/s on an M6 Mac Mini base versus 614 GB/s on an M5 Max — governs LLM decode speed, not small-batch CPU vision inference. A Mac Studio will not index your library 4× faster than a Mac Mini. Performance-core count and clocks move the needle; the $2,000 you’d spend on bandwidth does not.
2. RAM needs are modest. The repo’s only memory guidance concerns the optional chat model: Qwen3 1.7B “fits 8GB Macs,” Llama 3.2 3B “needs headroom.” Vision models in the 214–850MB range plus an Electron shell run comfortably inside 16GB alongside your normal workload. This is not a “buy 64GB” application.
3. The Neural Engine stays idle — again. This is the same pattern we documented in our Apple Neural Engine deep dive: the ANE is genuinely efficient silicon that almost no third-party software uses, because targeting it (CoreML conversion, per-chip quirks) is far more work than shipping portable ONNX CPU inference. eiln’s reverse-engineering work showed even a fixed ANE pipeline losing to the GPU on the same chip. SCM shipping CPU-only is a rational engineering choice — and one more data point that “40 TOPS of NPU” on a spec sheet tells you nothing about what your apps will actually use.
If you were waiting for a reason to upgrade, this isn’t it. An M1 MacBook Air from 2020 runs SCM; it just indexes the back catalog more slowly — and indexing is a one-time cost, with incremental adds nearly free thanks to the shot-plan cache.
Privacy: the architecture backs the claim
“Private AI photo search” is a claim worth auditing, because your photo library is the most sensitive dataset you own. SCM’s architecture holds up: no accounts, no telemetry, no uploads; inference on-device; network traffic limited to one-time, sha256-verified model downloads (plus a CSP allowance for Google Fonts, with a monospace fallback offline). Because it’s MIT-licensed, all of this is inspectable rather than promised — the same bar our sister site aifoss.dev applies to every self-hosted AI tool it covers.
One honest caveat the marketing copy won’t give you: the index is plain files on disk — JSON, Float32 embedding bins, thumbnails, transcripts. Nothing is encrypted at the application layer. On a personal Mac with FileVault on (the default for years), that’s fine. But those thumbnails and Whisper transcripts are a readable distillation of your private media, so if you point SCM at sensitive folders on a shared or unencrypted volume, you’ve created a second, smaller copy of exactly the content you were protecting. Relocate the data dir accordingly.
It’s also worth being clear about what SCM is not: it won’t replace Apple Photos’ face recognition or its tight iPhone sync, it has no mobile client, and the “Ask” chat is a 1.7B–3B model doing retrieval over captions and transcripts — useful for “find the video where someone mentions the lease,” not a reasoning engine. For that class of local AI you still want the hardware tiers in our buying guides, or a browser-based zero-install path like the ones we covered in the WebGPU browser-LLM piece.
If you’re buying a Mac anyway
To repeat the verdict: SCM justifies no new hardware. But it stacks nicely onto a purchase you’re already making for other local AI reasons, and it’s a natural fit for an always-on machine that can chew through a watched folder overnight:
- The Mac Mini M6 16GB at $899 (Apple list price, October 2026; 32GB is a $400 build-to-order bump to $1,299) is the cheapest always-on Mac appliance — SCM, Home Assistant, and a small Ollama model all fit, within the 153–170 GB/s bandwidth limits we covered in the M6 Mini guide.
- The Mac Mini M5 Pro 64GB at $1,699 (307 GB/s) is the floor where a Mac starts making sense as an actual LLM box rather than an appliance — see our MacBook vs Studio comparison for where the tiers above it land.
If your media library lives on a NAS, note our NAS-for-local-AI warning still applies: run SCM on the Mac against a mounted share, don’t try to make the NAS do the thinking.
FAQ
Does SCM run on Intel Macs? The Homebrew install is documented for Apple Silicon on macOS 12+, and the project site describes the app as designed for Apple Silicon. Intel support isn’t claimed anywhere; assume no.
Does it use the Neural Engine or GPU? No. The vision pipeline runs through ONNX Runtime with CPU-labeled timings in the repo’s own docs. Only the optional llama.cpp chat sidecar can touch the GPU.
How long will my library take to index? From the repo’s per-image CPU numbers: a 50,000-photo library is roughly 42–83 minutes on SigLIP-2-B/16, or 6.7–7.9 hours on the default CLIP ViT-L/14@336. Videos cost far less than their length suggests — the default preset caps a 2-hour file at 128 embedded frames.
Is anything uploaded? No — no accounts, no telemetry. Network use is limited to one-time, sha256-verified model downloads (Hugging Face, Tesseract CDN, and GitHub if you opt into chat). After that it runs offline.
Does it replace Apple Photos search? It complements it. Photos only searches its own library and keeps its index opaque; SCM indexes any folder — external drives, project directories, video archives — and adds OCR, dialogue, and in-video scene search that Photos doesn’t offer.
Sources
- allenv0/SCM — GitHub repository (license, models, timings, install)
- SCM — Deep AI Search for Every Photo and Frame of Video on Mac (official site)
- SCM Ships MIT-Licensed, Offline CLIP Search for Mac Media — AI Weekly
- SCM launches offline AI search tool for photos and videos on macOS — Complete AI Training
- SCM: AI search across macOS photos and video frames — PromptZone
- New Mac mini M6 / M5 Pro: specs, price, preorder — travis.media
- Mac mini (2026) — Michael Tsai
Last updated October 9, 2026. Prices and specs change; verify current rates before purchasing. Indexing-time figures marked as computed are derived from the repo’s published per-image timings, not independent benchmarks.
Recommended Gear
- Mac Mini M6 — $899 (16GB) / $1,299 (32GB BTO), the always-on appliance tier
- Mac Mini M5 Pro — $1,699 (64GB, 307 GB/s), the cheapest Mac that’s also a real LLM box
Was this article helpful?
Thanks for the feedback — it helps improve future articles.
Need hands-on help?
I offer 1-on-1 technical consulting for local AI setup, GPU selection, and AI coding tool configuration — same topics covered on this site.
Book a session — $49 / hour →Get the numbers before you buy
New GPU and mini-PC benchmarks, VRAM thresholds, and price checks — sent when there's something worth acting on, not on a schedule. No spam, unsubscribe anytime.