Seen those expensive ai usb flash drives? With this you make your own.
C#
2
559 commits
updated Jun 6, 2026
Plug in a drive. Ask your AI anything. No internet required.
Prepare the drive once on a Windows or Mac machine with internet access — download the models, load your documents, and finalize the SSD. PrepApp ships for both hosts: a WPF app on Windows and a native SwiftUI app on Mac, so a Mac-only user can prep without owning a Windows machine. On Windows, the Runner provides the full offline assistant: document-grounded chat, voice I/O, HOTAS PTT, DCS binding import, and the LAN API. The macOS beta Runner provides a subset: RAG-backed chat against an already-indexed library, encrypted config unlock, and the API sidecar. See Platform Availability below for a full feature comparison.
New to local AI? This app runs an AI model on your own hardware — nothing is sent to the cloud, and you don't need an account or a subscription. A few terms you'll see below:
- Local LLM — the AI "brain" (a large language model) that runs on your machine instead of a remote server.
- Ollama — the small engine that loads and runs those models locally. Free-AI-SSD bundles it on the drive for you.
- Model / GGUF — a downloadable AI model file (GGUF is the format Ollama uses); you choose which model to stage on the drive.
- RAG — Retrieval-Augmented Generation: the AI reads your documents and answers from them (citing sources), instead of relying on training data alone.
- Embeddings — how the app turns your documents into searchable math so it can find the right passage to answer a question.
In short: prep the drive once online, then carry it anywhere and ask questions about your own documents — fully offline.
Quick start: download from Releases, run FreeAiSsd.PrepApp.exe (Windows) or PrepApp.app (macOS), then follow Setup & Installation below. A condensed version is at docs/QUICKSTART.txt.
Windows PrepApp — Model Manager (top) and Drive Setup (bottom).
macOS screenshots coming soon.
Open bugs, fixes, and feature status are tracked on GitHub rather than hand-maintained here (so this list can't drift out of date):
This started as a way to take AI into the field with no cell signal — ham radio manuals, band plans, and reference docs loaded onto a pocket SSD so an LLM could answer questions about them miles from civilization. Then it turned out the same setup works really well as a voice-activated copilot in DCS: load aircraft manuals, import your HOTAS bindings, and hit a button on the throttle to ask questions mid-sortie without taking the VR headset off. Same drive. Same AI. Same offline-first idea.
A clear picture of what works where today. Mac support is actively expanding but not yet at feature parity with Windows.
| Feature | Windows PrepApp | macOS PrepApp |
|---|---|---|
| Format drive | ✅ NTFS or exFAT | ✅ exFAT only |
| Stage Ollama + pull models | ✅ Full | ✅ Full |
| Pull models from Hugging Face | ✅ Full (with token auth) | ✅ Full (with token auth) |
| Manage models on encrypted drives | ✅ Full | ✅ Full |
| Detect pre-configured drive | ✅ Full | ✅ Full |
| Target: Windows-only (NTFS) | ✅ Yes | ❌ macOS cannot format NTFS |
| Target: exFAT (cross-platform or Mac-only) | ✅ Yes | ✅ Yes |
| Encrypted config roundtrip | ✅ Full | ✅ Full (CryptoKit port) |
| Feature | Windows Runner | macOS Runner (beta) |
|---|---|---|
| Chat (non-RAG) | ✅ Full | ✅ Full |
| RAG chat (sources panel; inline citations opt-in) | ✅ Full | ⚠️ Query-only — reads a library indexed on Windows |
| Add / sweep / rebuild document library | ✅ Full | ❌ Not supported yet |
| Voice input (speech-to-text) | ✅ Whisper.cpp (fully local) | ✅ On-device dictation (SFSpeechRecognizer) |
| Voice output (text-to-speech) | ✅ SAPI + Piper | ✅ Native (AVSpeechSynthesizer) |
| HOTAS Push-to-Talk | ✅ DirectInput | ❌ Not ported yet |
| DCS Bindings Import | ✅ Full | ❌ Not supported yet |
| Network Mode (LAN API) | ✅ Full (v2) | ✅ Sidecar-hosted |
Web chat UI (/chat/, browser) | ✅ Full | ✅ Sidecar-hosted |
| Companion tray app (second PC) | ✅ Full | — |
| Headless CLI (RunnerCli) | ✅ Full | — |
| Unlock Windows-prepped encrypted SSD | ✅ Full | ✅ Full (CryptoKit) |
RAG on macOS: The Mac Runner can answer questions using a document library that was ingested and indexed on Windows. It cannot add new documents, watch folders, sweep, or rebuild the index — those operations require the Windows Runner. If you only have a Mac, you can still get RAG-backed answers by preparing the drive on Windows first (or on a second Windows machine).
[guide.pdf §Engine Start p.12], plus opt-in OCR to recover text from scanned/diagram pageshttp://HOST:41555/chat/); full assistant minus voice, with model/library pickers, RAG sources, and per-device history. Works from any LAN device including an iPadFreeAiSsd.RunnerCli) — terminal REPL for SSH/Tailscale access; streams chat, shows RAG sources, zero GUI deps<SSD>/Runner.app (drive root, double-click to launch; no zip to expand)SFSpeechRecognizer) and spoken responses (AVSpeechSynthesizer); audio never leaves the MacKnown limitations across all platforms:
pcm16le only; other codecs not implementedFor headless access from a terminal (including an iPad over Tailscale), FreeAiSsd.RunnerCli ships alongside the Windows Runner. It's a thin HTTP client against Runner's LAN API — same RAG pipeline, same source citations, no GUI.
$ FreeAiSsd.RunnerCli --help
$ FREEAI_URL=http://my-desk:41555 FREEAI_API_KEY=... FreeAiSsd.RunnerCli --model phi3
Target: http://my-desk:41555
Host reachable (ollamaRunning=True). Type /help for commands. Ctrl-C or 'exit' to quit.
phi3> what aircraft can I fly in DCS Open Beta?
...streamed response...
— sources: dcs-aircraft-list.pdf
phi3> /quit
Precedence for configuration: --url / --api-key flag > FREEAI_URL / FREEAI_API_KEY env var > default (http://127.0.0.1:41555, no key). Use --no-stream on very flaky links to fall back to a single-response round-trip.
You're in VR, mid-sortie, and can't remember the sequence to uncage an AIM-9. You reach for your HOTAS, key the mic, and ask. The AI answers with the buttons on your stick — sourced from the aircraft manual sitting on the drive. No internet. No cloud. No subscription.
What the Windows Runner does for flight sim:
Saved Games\DCS folder, scans your aircraft, and writes a per-aircraft reference file with your real button assignmentsSupported now in the Windows Runner: DCS World (stable and Open Beta), any aircraft with binding files in Config/Input, multi-device merging (stick + throttle + rudder pedals)
Planned: IL-2 Sturmovik and War Thunder binding parsers (see Roadmap)
Camping, deployed for emergency comms, or away from a desk — you need to reference your radio manual or band plan and there's no cell signal.
Load your manuals and reference documents onto the drive before you go. The AI indexes everything and answers from your own library, completely offline, from a drive that fits in your pocket.
Maybe you don't trust cloud AI with your data. Maybe your workplace restricts internet access. Maybe you want the same staged SSD available across your machines, with the full feature set on Windows and the current direct-chat beta on macOS.
Prepare the drive once. Your models, your documents, your config — nothing leaves the drive, no account needed, no telemetry.
Load first aid guides, plant identification references, equipment specs, survival manuals — whatever you need when there's no connectivity. The AI indexes it all and answers from your library when you're completely off-grid.
Which prep host (source OS) can produce which target drive:
| Source OS | Target | Filesystem | Supported |
|---|---|---|---|
| Windows | Windows-only | NTFS | Yes |
| Windows | Cross-platform (Windows + Mac) | exFAT | Yes |
| Windows | Mac-only | exFAT | Yes (APFS not available from Windows) |
| Mac | Mac-only | exFAT | Yes (APFS deferred from supported targets) |
| Mac | Cross-platform (Windows + Mac) | exFAT | Yes |
| Mac | Windows-only | NTFS | Not supported — use Windows PrepApp (macOS cannot natively format NTFS) |
Encrypted-config roundtrip is bidirectional. A drive prepped on Windows unlocks cleanly on Mac, and a drive prepped on Mac unlocks cleanly on Windows. The on-disk encrypted format (AES-256-GCM + PBKDF2-SHA256) is identical on both platforms and is pinned by cross-language tests.
Notes on the unsupported cells:
Stable (recommended): Download Free-AI-SSD-win.zip from Releases. Extract anywhere on Windows. Run FreeAiSsd.PrepApp.exe.
Beta cross-platform bundle: Free-AI-SSD-crossplatform.zip includes the Mac PrepApp (PrepApp.app) and Mac Runner beta (Runner.app) alongside the Windows artifacts. The macOS builds are currently unsigned/not notarized — see macOS first launch below before opening either app.
The download root is intentionally minimal — the prep tool(s), LICENSE, QUICKSTART.txt, and one dependencies/ folder for everything the prep tool consumes (no duplicated/nested copies):
Free-AI-SSD-win.zip Free-AI-SSD-crossplatform.zip
├── FreeAiSsd.PrepApp.exe ├── FreeAiSsd.PrepApp.exe (Windows prep)
├── LICENSE ├── PrepApp.app (macOS prep)
├── QUICKSTART.txt ├── LICENSE
└── dependencies/ ├── QUICKSTART.txt
├── runner/ └── dependencies/
├── companion/ ├── runner/ companion/ prereqs/
└── prereqs/ └── mac/ (Runner.app, ollama, manifest)
Cross-platform note: because macOS
.appbundles carry symlinks Windows archivers strip, prep a cross-platform drive's macOS side from a Mac (PrepApp.app). The Windows side works from either host. The Mac Runner is staged to the SSD root as<SSD>/Runner.app— unzipped, launchable directly, no zip to expand each run.
Until the signed/notarized release ships, the unsigned ad-hoc Mac apps trip Gatekeeper as soon as Safari stamps the downloaded ZIP with a quarantine xattr. The dialog reads "FreeAiSsd is damaged and can't be opened. You should move it to the Trash." even though nothing is corrupted — and right-click → Open / "Allow apps from anywhere" do not clear this state.
Strip the quarantine xattr once in Terminal, replacing the path with wherever you extracted the bundle:
xattr -dr com.apple.quarantine /path/to/PrepApp.app
xattr -dr com.apple.quarantine /path/to/Runner.app
Both apps then launch normally on double-click. This workaround goes away with the next signed release.
CI artifacts: Available from GitHub Actions for validation and testing. Prefer Releases for normal use.
Phase 1 — Prepare (online, once):
On Windows:
FreeAiSsd.PrepApp.exeOn Mac:
PrepApp.app from the cross-platform bundle. First launch only: run the Gatekeeper unblock xattr command above — the build is unsigned/not notarized, so Safari quarantine makes Gatekeeper claim the app is "damaged" until that bit is cleareddiskutil directly to format the drive as exFAT and lay out the canonical SSD directory structureThe resulting drive is byte-for-byte interchangeable with a Windows-prepped drive of the same target compatibility.
Phase 2 — Run (offline, anywhere):
<SSD>\windows\runner\FreeAiSsd.Runner.exe<SSD>/Runner.app (at the drive root — double-click)| Operation | Internet Required? |
|---|---|
| PrepApp — download, pull, staging | Yes |
| Windows Runner start / chat | No |
| macOS beta Runner start / sidecar-backed chat | No |
| macOS beta Runner — unlock Windows-prepped encrypted SSD | No |
| Reference Documents indexing and retrieval (Windows Runner) | No |
| Pull embedding model (if missing from SSD, Windows Runner) | Once |
| DCS Bindings Import (Windows Runner) | No |
| Voice input (Whisper transcription, Windows Runner) | No (model download is once) |
| Text-to-speech (Windows Runner) | No |
Runner won't start / dependency warnings
Missing embedding model while offline
PDF citations seem wrong or sparse
.NET / runtime prerequisites on target machine
macOS beta limitations
Every prerequisite and bundled third-party tool is fetched, verified, and recorded at runtime via a single shared resolver (shared/Prereqs/PrereqResolver.cs) used by both PrepApp and the CI offline-bundle builder (tools/FreeAiSsd.PrereqFetch). There are no hardcoded per-version SHA pins in the workflow or the catalog — stale pins were the failure mode we were hitting most often. Instead:
| Upstream | Version discovery | Integrity check | Trust basis |
|---|---|---|---|
| .NET 8 Desktop Runtime (x64) | https://builds.dotnet.microsoft.com/dotnet/release-metadata/8.0/releases.json → latest-release (rejects preview/rc builds) | SHA-512 from the same releases.json entry | Vendor-published hash over HTTPS to Microsoft's CDN |
| VC++ Redistributable (x64) | https://aka.ms/vs/17/release/vc_redist.x64.exe evergreen permalink | Observed SHA-256 recorded in manifest only | HTTPS-only trust to Microsoft aka.ms (no vendor per-version hash is published at a predictable URL) |
| Ollama (macOS, universal) | GitHub API releases/latest → picks Ollama-darwin.zip / ollama-darwin.zip asset | SHA-256 from the release's sha256sum.txt asset | Vendor-published hash over HTTPS to github.com |
Fail-closed invariants (CI and PrepApp both enforce):
latest-releaseThe prereqs-manifest.json that ships on the SSD records the resolved upstream URL, the vendor hash (when one was available), the observed SHA-256, and a short trustNote describing which trust basis was used — so offline installs can be audited without calling back to the upstream.
The Windows Runner includes a Reference Documents panel. Add your own files and the AI references them when answering instead of relying on training data alone. Retrieved chunks are cited inline so you can see exactly where an answer came from.
Supported formats: .pdf, .txt, .md, .json, .csv
Workflow:
How retrieval works: A hybrid retriever combines semantic (vector) search with keyword (BM25) search, fuses the results, and pulls neighboring chunks for surrounding context. The Sources list always shows what was used. Inline citations like [guide.pdf §Engine Start p.12] are opt-in (ragInlineCitations, off by default — answers stay concise) and are stripped before text-to-speech when enabled. If nothing relevant is found, the model is told so and won't invent context.
macOS: RAG queries work against a library indexed on Windows. Adding documents, folder sweeps, and index rebuilds require the Windows Runner.
Limitations:
ocrEnabled, off by default) to recover text from embedded images.The Windows Runner reads your DCS World controller bindings and writes them into the document library as a per-aircraft reference file. After import, when you ask "how do I uncage my AIM-9?" the AI answers with the button on your stick — not a generic keybind table.
How to import:
Saved Games\DCS folder — browse manually if detection failsdiff.lua (stick, throttle, rudder pedals), merges them into one file per aircraft, and writes it to your librarySupported:
Config/InputNot yet supported: IL-2 Sturmovik and War Thunder (see Roadmap)
In the Windows Runner, speak your questions and hear the answers. The entire pipeline runs locally — no cloud STT, no cloud TTS.
Speaking to the AI:
autoSendVoiceInput)AI voice response: Enable TTS in settings. Two engines available:
piper.exe (~22 MB) plus the default en_US-amy-medium voice (~60 MB) into windows/tools/piper/ or mac/tools/piper/, both SHA-256 verified.You can route TTS to a specific audio output device — useful for sending AI voice to your VR headset while system audio goes elsewhere.
Whisper model sizes (stored at models/whisper/ on the SSD):
| Size | File | Approx. disk | Notes |
|---|---|---|---|
| Tiny | ggml-tiny.bin | ~75 MB | Fastest; lower accuracy |
| Base | ggml-base.bin | ~142 MB | Default; good for most use |
| Small | ggml-small.bin | ~466 MB | Better accuracy |
| Medium | ggml-medium.bin | ~1.5 GB | Best accuracy; more RAM required |
The first time voice is used, Runner downloads the selected Whisper model (internet required for that one step). After that, fully offline.
In the Windows Runner, bind a button on your HOTAS to start and stop voice recording — no keyboard, no mouse. Built for VR where hands-free activation matters.
Setup:
"X-56 Rhino Throttle")push_to_talk — hold the button to record, release to sendtoggle — press once to start recording, press again to stop and sendOptional overlay: A small always-on-top window shows recording status. Disable it for VR where it would be distracting (pttOverlayEnabled).
Optional sound: A short beep plays on PTT activation/deactivation. Toggle with pttActivationSoundEnabled.
Full VR voice loop: HOTAS button → mic opens → speak → button release → Whisper transcribes → prompt sent → AI responds → TTS speaks into headset. Hands never leave the controls.
Companion tray app: Remote HOTAS/PTT is supported — the Companion app can run on a second PC and drive the full voice loop against Runner over LAN, including its own PTT activation beep and overlay.
Network Mode lets one Windows machine run Runner + Ollama locally, while other devices on your LAN call Runner's HTTP API.
Architecture:
127.0.0.1) on the host machineSecurity model (home LAN baseline):
127.0.0.1 (loopback) by default. Binding to 0.0.0.0 (all interfaces) is an explicit opt-in set in portable-config.json; Runner logs a WARNING on startup whenever the effective bind address is not loopback.Authorization: Bearer <key> or X-API-Key)Endpoints:
GET /api/healthGET /api/modelsPOST /api/chatPOST /api/chat/stream (newline-delimited JSON stream)POST /api/stt/transcribe (multipart upload: audio)POST /api/voice/query (multipart upload: audio, optional model, autoSendToChat, speakResponse, returnAudio). When returnAudio=true the response includes AudioBase64 + AudioMime so the client can play TTS locally instead of on the host.POST /api/tts/speakPOST /api/tts/stopExample cURL requests:
# health (no API key required)
curl http://RUNNER_HOST:41555/api/health
# list models
curl -H "Authorization: Bearer YOUR_KEY" \
http://RUNNER_HOST:41555/api/models
# non-stream chat
curl -X POST http://RUNNER_HOST:41555/api/chat \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi3","prompt":"Summarize startup checklist"}'
# stream chat (NDJSON)
curl -N -X POST http://RUNNER_HOST:41555/api/chat/stream \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi3","prompt":"Step-by-step A-10C startup"}'
# trigger host-side TTS
curl -X POST http://RUNNER_HOST:41555/api/tts/speak \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Radio check complete."}'
# STT transcription (WAV upload)
curl -X POST http://RUNNER_HOST:41555/api/stt/transcribe \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav"
# Voice query (upload -> transcribe -> chat -> optional host-side TTS)
curl -X POST http://RUNNER_HOST:41555/api/voice/query \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav" \
-F "model=phi3" \
-F "autoSendToChat=true" \
-F "speakResponse=true"
# Voice query with client-side TTS playback (returnAudio)
# Response JSON contains AudioBase64 + AudioMime ("audio/wav") for local playback.
curl -X POST http://RUNNER_HOST:41555/api/voice/query \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav" \
-F "model=phi3" \
-F "autoSendToChat=true" \
-F "speakResponse=true" \
-F "returnAudio=true"
Remote voice upload formats and limits:
format=pcm16le)networkMaxAudioUploadMBreturnAudio=true returns synthesized TTS as WAV PCM 16-bit mono 16kHz (AudioMime: "audio/wav", AudioBase64), bounded by networkMaxAudioUploadMBWeb chat / LAN access (no install):
Runner serves a standalone browser chat client from the same LAN API host — no app to install on the other device. It's the full assistant minus voice: model picker, document-library selection, RAG-grounded chat with sources, a collapsible thinking view, and per-request temperature/thinking controls. Chat history is saved per-device in the browser (localStorage).
On the host PC the web UI runs on loopback automatically whenever Ollama is up — just click "Open Chat UI" in Runner (or open http://127.0.0.1:41555/chat/). No Network Mode, no API key for local use.
To reach it from other devices:
http://RUNNER_HOST:41555/chat/ (use the host's LAN IP or HOSTNAME.local).Per-request model parameters set in the web UI (temperature, thinking) apply only to that request and never overwrite the host's saved configuration. Plain HTTP over a trusted home LAN, API-key gated — do not expose port 41555 to the public internet.
All settings live in config/portable-config.json on the SSD.
| Property | Default | Description |
|---|---|---|
ollamaPort | 11434 | TCP port for the local Ollama server |
preferredCompute | "auto" | Compute mode: "auto" (detected GPU — AMD→Vulkan, NVIDIA→CUDA, Intel→Vulkan) or "cpu" (force CPU). Legacy "cuda"/"rocm" are treated as "auto" |
useStreamingChat | true | Stream tokens as they generate; falls back to non-streaming if streaming fails |
Per-chat overrides for the active model. Sentinel values mean "use the model's built-in default," so the app stays compatible with any GGUF model. The web/desktop UI can set these per request without changing the saved config.
| Property | Default | Description |
|---|---|---|
modelContextWindow | 0 | Override Ollama num_ctx; 0 = model default |
modelTemperature | -1 | Override temperature (0.0–2.0); -1 = model default |
modelTopP | -1 | Override top_p (0.0–1.0); -1 = model default |
modelMaxOutputTokens | -1 | Override num_predict (max tokens per response); -1 = unbounded. Also caps the thinking budget. |
modelThinkMode | "" | Reasoning control: "" = model default, "off", "low", "medium", "high" (only for models that support thinking) |
| Property | Default | Description |
|---|---|---|
activeDocumentLibraryId | null | Active library ID; null disables RAG |
retrievalTopK | 8 | Number of chunks retrieved per query |
hybridRetrievalEnabled | true | Fuse semantic (vector) + keyword (BM25) search; false = vector-only. No reindex needed. |
retrievalNeighborRadius | 1 | Chunks pulled on each side of a hit for context; 0 disables |
ragInlineCitations | false | true appends inline labels like [guide.pdf §Engine Start p.12] (stripped before TTS); false = concise answers, sources still shown in the Sources panel |
chunkSize | 1200 | Characters per chunk during indexing |
chunkOverlap | 200 | Characters of overlap between adjacent chunks |
embeddingModelName | "nomic-embed-text" | Embedding model served by local Ollama |
minimumSimilarityThreshold | 0.3 | Minimum cosine similarity (0.0–1.0) for a chunk to be included; lower = more permissive |
maxEmbeddingConcurrency | 4 | Concurrent embedding requests during ingestion |
maxDocumentSizeMB | 512 | Max file size (MB) accepted for ingestion |
Off by default. When enabled (and a Tesseract bundle is staged on the SSD), OCR recovers text baked into images inside PDFs (e.g. cockpit MFD labels, diagrams) and adds it as supplementary searchable chunks — the clean text layer is never replaced.
| Property | Default | Description |
|---|---|---|
ocrEnabled | false | Run OCR over embedded PDF images during ingestion |
ocrLanguage | "eng" | Tesseract language code(s) passed to -l |
ocrMinImagePixels | 10000 | Skip images smaller than this (width × height) — filters out icons, rules, logos |
ocrMinWordConfidence | 55 | Drop OCR words below this confidence (0–100) to suppress garble |
ocrMaxImagesPerFile | 4000 | Hard cap on images OCR'd per file |
ocrPerImageTimeoutSeconds | 30 | Per-image OCR timeout (a stuck image is skipped) |
| Property | Default | Description |
|---|---|---|
whisperModelSize | "Base" | Whisper model: "Tiny", "Base", "Small", or "Medium" |
selectedMicrophoneDevice | null | Microphone device name; null = system default |
autoSendVoiceInput | true | true sends transcribed text immediately; false puts it in the prompt field for review |
| Property | Default | Description |
|---|---|---|
ttsEnabled | false | Enable TTS for AI responses |
ttsEngine | "system" | "system" (Windows SAPI) or "piper" (neural TTS) |
ttsVoiceName | null | Voice name for the selected engine; null = engine default |
ttsRate | 0 | Speech rate: -10 (slowest) to 10 (fastest) |
ttsVolume | 100 | Volume: 0 (silent) to 100 (max) |
ttsOutputDevice | null | Audio output device for TTS; null = system default |
| Property | Default | Description |
|---|---|---|
pttEnabled | false | Enable HOTAS push-to-talk |
pttDeviceName | null | DirectInput device name (e.g., "X-56 Rhino Throttle") |
pttButtonIndex | 0 | Zero-based button index on the joystick device |
pttMode | "push_to_talk" | "push_to_talk" (hold to record) or "toggle" (press to start/stop) |
pttActivationSoundEnabled | true | Play a beep on PTT activation/deactivation |
pttOverlayEnabled | true | Show the always-on-top PTT status overlay |
pttOverlayX / pttOverlayY | 20 / 20 | Overlay window position in pixels from top-left |
The Companion tray app exposes the same two cues under identical key names (pttActivationSoundEnabled, pttOverlayEnabled) in companion-config.json, toggleable from Companion's Settings window.
| Property | Default | Description |
|---|---|---|
networkModeEnabled | false | Expose the Runner API on the LAN. The API always runs on loopback when Ollama is up (for the on-PC web UI); this flag rebinds it to networkBindAddress and enforces the API key. |
networkBindAddress | "127.0.0.1" | Bind address used when exposed on the LAN (typically 0.0.0.0). Ignored — forced to loopback — when networkModeEnabled is off. |
networkPort | 41555 | TCP port for Runner API |
networkApiKey | "" | Shared secret for API auth |
networkRequireApiKey | true | Require API key on all non-health endpoints |
networkAllowTts | false | Allow remote callers to trigger host-side TTS |
networkAllowRemoteStt | false | Allow remote audio upload transcription via /api/stt/transcribe |
networkAllowRemoteVoiceQuery | false | Allow remote voice-query orchestration via /api/voice/query |
networkVoiceAutoSendToChat | true | Default for voice query: auto-send transcription to chat when request omits override |
networkMaxAudioUploadMB | 10 | Maximum upload size in MB for remote STT/voice endpoints |
Core infrastructure:
Windows Runner (full feature set):
returnAudio local TTSModel management (Windows + Mac):
num_ctx / temperature / top_p / max-output (num_predict), plus a Thinking control (Ollama think: Off / Low / Medium / High) to disable or cap reasoning models that loop; live on both the Windows and Mac runnersmacOS:
Free-AI-SSD ships several components backed by a shared cross-platform library:
prep-app/) — runs on an online Windows machine to configure the SSD: picks drive, downloads and stages Ollama, pulls models (Ollama or Hugging Face), bundles prerequisites, finalizes layoutmac-prep-app/, SwiftUI) — native macOS PrepApp for the cross-platform bundle. Drives diskutil directly to format target SSDs as exFAT, stages the runner / Ollama / prereq payloads via the mac-prep-host net8.0 sidecar (which consumes prep-core/), and writes encrypted config in a format byte-identical to the Windows PrepApp. Apple Silicon (arm64), macOS 11+runner/) — runs from the SSD on the target machine; starts Ollama, provides the chat interface, manages document libraries, voice pipeline, HOTAS PTT, and the LAN API hostmac-runner/, Swift) — thin macOS app shipped at <SSD>/Runner.app (drive root). It selects/infers the SSD, unlocks encrypted config, reads installed models, starts macOS Ollama, spawns the local Runner API sidecar, and sends chat through the shared RAG pipeline when an active indexed library existsAudioCaptureService → WhisperSpeechToTextService → ChatService → SystemTextToSpeechService / PiperTextToSpeechService, orchestrated by PttVoicePipelineService when HOTAS PTT is enabledshared/Documents/) — DcsSavedGamesLocator finds DCS installs, DcsAircraftScanner enumerates aircraft, DcsBindingParser parses diff.lua, DcsBatchProcessor merges devices and writes RAG documentscompanion/, WPF tray app) — optional lightweight client for a second LAN machine; no SSD required; talks to the Runner LAN API for chat / STT upload / voice-query / host-side TTS. Supports its own HOTAS PTT loop, activation beep, status overlay, and mic-preflight checkFreeAiSsd.Shared, net8.0) — common core logic for encryption, trust policy, path guards, config, dependency checking, download management, MVVM infrastructure, DCS binding models, document library, and RAG pipelineRunner's business logic lives in injectable services with no UI dependencies, enabling unit testing without a WPF host:
| Service | Purpose |
|---|---|
OllamaLifecycleService | Process start/stop, port resolution, trust validation |
ModelManagementService | Installed model listing, sizing warnings, embedding model pull |
DocumentOperationsService | Library CRUD, file ingestion, folder sweep, index rebuild |
ChatService | RAG-augmented prompt construction and Ollama /api/generate calls |
DcsBindingsImportService | DCS installation detection, aircraft scanning, batch binding import |
WhisperSpeechToTextService | Whisper.cpp transcription via Whisper.net; model download management |
SystemTextToSpeechService | Windows SAPI TTS with optional NAudio device targeting |
PiperTextToSpeechService | Piper neural TTS; spawns piper.exe, streams raw PCM through NAudio |
AudioCaptureService | Microphone capture at 16 kHz/16-bit mono (Whisper's required format) |
HotasInputService | DirectInput polling for HOTAS PTT button state |
PttVoicePipelineService | Orchestrates the full PTT → record → transcribe → send → TTS loop |
minimumSimilarityThreshold (default 0.3) are discarded; the model is told explicitly when nothing relevant was foundmaxEmbeddingConcurrency)System.Numerics.Vector<float>; top-K uses an O(N log K) priority queue| Control | Detail |
|---|---|
| Encrypted config | AES-256-GCM with PBKDF2-SHA256 (210,000 iterations) |
| Config write guard | ConfigStore is the only path for config writes; a plaintext write to an encrypted drive throws InvalidOperationException (fail-closed) |
| Package trust | Ollama downloads validated against URL allowlist + SHA-256 digest before execution; macOS payloads additionally verified as arm64 Mach-O |
| Fail-closed write guard | PrepDriveWriteGuard blocks all writes to encrypted drives if encryption state is ambiguous |
| Path traversal prevention | PathGuards enforces sibling boundary checks with platform-aware case sensitivity |
| Shell injection prevention | ProcessRunner uses ArgumentList, not string concatenation |
A security review on 2026-02-19 found no critical vulnerabilities in the audited surface. The invariants above are enforced in code and tests; ongoing issues are tracked on GitHub Issues.
config/ — portable-config.json (plaintext) or portable-config.encrypted.json (opt-in)
models/ — Ollama model store
models/whisper/ — Whisper STT model files (ggml-*.bin)
logs/ — app logs
docs/libraries/ — Reference Documents library files, manifests, index DB
windows/runner/ — Runner app
windows/tools/ollama/ — staged Ollama runtime + trust attestation
windows/tools/piper/ — optional Piper TTS binary and voice models (user-installed)
windows/tools/tesseract/ — optional Tesseract OCR engine + tessdata
windows/tools/prereqs/ — offline prerequisite installers + manifest
Runner.app/ — macOS Runner bundle (root-level, launchable directly)
mac/tools/ollama/ — staged macOS Ollama runtime + trust attestation
mac/tools/piper/ — optional Piper TTS binary and voice models
mac/tools/tesseract/ — optional Tesseract OCR engine + tessdata
cache/ — prep-time download cache
SsdLayout in the shared library is the single source of truth for these paths — always use it rather than constructing paths manually.
| Directory | Target | Purpose |
|---|---|---|
shared/ | net8.0 | Cross-platform shared library (FreeAiSsd.Shared) |
runner-core/ | net8.0 | Platform-neutral Runner business logic (chat, RAG, library, local API) shared by Windows Runner and the Mac runner-host sidecar |
prep-core/ | net8.0 | Platform-neutral PrepApp business logic (manifest, staging, prereq, encrypted config, HF/Ollama model pulls) shared by Windows PrepApp and the Mac prep-host sidecar |
prep-app/ | net8.0-windows | WPF PrepApp (Windows) |
mac-prep-app/ | macOS (Swift) | Native SwiftUI PrepApp (Mac); produces drives byte-identical to Windows PrepApp |
mac-prep-host/ | net8.0 | osx-arm64 sidecar that runs prep-core/ business logic for the Mac PrepApp over a stdin command protocol |
runner/ | net8.0-windows | WPF Runner (Windows) |
mac-runner/ | macOS (Swift) | Swift macOS beta Runner over the local Runner API sidecar |
mac-runner-host/ | net8.0 | osx-arm64 sidecar that hosts RunnerLocalApiService for the Mac Runner |
runner-cli/ | net8.0 | Headless CLI client (FreeAiSsd.RunnerCli) — SSH/Tailscale terminal access to Runner API |
companion/ | net8.0-windows | WPF Companion tray client (LAN second-PC use) |
tools/FreeAiSsd.PrereqFetch/ | net8.0 | CI helper that pre-builds the offline prereq bundle via the shared PrereqResolver |
tests/ | net10.0 | xUnit test project (FreeAiSsd.Tests) |
tests-ocr/ | net10.0 | xUnit OCR test project (FreeAiSsd.Tests.Ocr) — Tesseract OCR coverage |
docs/ | — | Documentation (includes QUICKSTART.txt) |
| File | Purpose |
|---|---|
DependencyChecker.cs | Detects missing VC++ / .NET runtimes via registry + process checks |
DownloadManager.cs | Resumable HTTP downloads with progress callbacks |
DriveInspector.cs | Enumerates candidate drives |
ModelSizing.cs | Maps model tags to RAM/VRAM/disk requirements for sizing warnings |
NetUtils.cs | Port availability checking |
OllamaPackageTrustPolicy.cs | URL allowlisting + SHA-256 digest verification |
PathGuards.cs | Path traversal prevention |
PortableConfig.cs | JSON config serialization with atomic writes |
PrepDriveWriteGuard.cs | Blocks writes to encrypted drives (fail-closed) |
PrereqInstallValidator.cs | Validates installer integrity (SHA-256) before execution |
Prereqs/PrereqResolver.cs | Runtime discovery of the latest stable upstream prereq versions + vendor-hash verification. Shared by PrepApp and CI. |
ProcessRunner.cs | Safe process spawning via ArgumentList, not string concatenation |
SsdEncryption.cs | AES-256-GCM config encryption |
SsdLayout.cs | Canonical directory structure constants and creation |
SsdLogger.cs | File-based logger writing to the SSD's logs directory |
SystemCompatibility.cs | GPU/CPU/OS detection for compatibility display |
Documents/DcsBindingParser.cs | Parses DCS diff.lua files into structured data for RAG |
Documents/DcsAircraftScanner.cs | Scans Config/Input for aircraft folders and device files |
Documents/DcsBatchProcessor.cs | Batch import: merges devices, formats output, writes to library |
Documents/DcsSavedGamesLocator.cs | Auto-detects Saved Games\DCS and .openbeta; supports manual override |
PrepViewModel lives in shared/ (net8.0) so it can be unit tested on Linux without WPFIDialogService abstracts all MessageBox/dialog interactionsshared/, implementations in prep-app/Services/ (net8.0-windows)MainWindow.xaml.cs reduced from ~1,800 lines to ~95 lines; all logic in PrepViewModel and servicesShared + tests (all platforms):
dotnet build shared/FreeAiSsd.Shared.csproj
dotnet build tests/FreeAiSsd.Tests.csproj
dotnet test tests/FreeAiSsd.Tests.csproj --verbosity normal
1 test (
IsPathUnderRoot_WindowsBoundaryIsRespected) is expected to fail on Linux — it tests Windows-specific path behavior.
Full build (Windows only):
dotnet restore FreeAiSsd.sln
dotnet build FreeAiSsd.sln -c Release
dotnet test FreeAiSsd.sln -c Release
Stage Runner payload into PrepApp output:
./build.ps1 -Configuration Release -Runtime win-x64
Key dependencies: xUnit, System.Management, Moq, PdfPig, SharpDX (DirectInput), ASP.NET Core (in-process LAN host), Microsoft.Extensions.DependencyInjection, SQLite
~1,000+ test cases ([Fact]/[Theory]; Theories expand to more at runtime) across ~100 test files in tests/ and tests-ocr/. One Windows-specific path test is expected to fail on Linux. Coverage spans:
CI (.github/workflows/build.yml) is the source of truth for the current count and pass/fail status.
Signing is disabled by default in CI (MAC_SIGNING_ENABLED=false). Supported via repository secrets: MACOS_CERT_P12_BASE64, MACOS_CERT_PASSWORD, APPLE_TEAM_ID, APPLE_ID, APPLE_APP_SPECIFIC_PASSWORD, MACOS_SIGN_IDENTITY.
See GitHub Releases for the full version history and release notes.
C#
64.1%
HTML
19.9%
Swift
14.2%
Seen those expensive ai usb flash drives? With this you make your own.
C#
2
559 commits
updated Jun 6, 2026
Plug in a drive. Ask your AI anything. No internet required.
Prepare the drive once on a Windows or Mac machine with internet access — download the models, load your documents, and finalize the SSD. PrepApp ships for both hosts: a WPF app on Windows and a native SwiftUI app on Mac, so a Mac-only user can prep without owning a Windows machine. On Windows, the Runner provides the full offline assistant: document-grounded chat, voice I/O, HOTAS PTT, DCS binding import, and the LAN API. The macOS beta Runner provides a subset: RAG-backed chat against an already-indexed library, encrypted config unlock, and the API sidecar. See Platform Availability below for a full feature comparison.
New to local AI? This app runs an AI model on your own hardware — nothing is sent to the cloud, and you don't need an account or a subscription. A few terms you'll see below:
- Local LLM — the AI "brain" (a large language model) that runs on your machine instead of a remote server.
- Ollama — the small engine that loads and runs those models locally. Free-AI-SSD bundles it on the drive for you.
- Model / GGUF — a downloadable AI model file (GGUF is the format Ollama uses); you choose which model to stage on the drive.
- RAG — Retrieval-Augmented Generation: the AI reads your documents and answers from them (citing sources), instead of relying on training data alone.
- Embeddings — how the app turns your documents into searchable math so it can find the right passage to answer a question.
In short: prep the drive once online, then carry it anywhere and ask questions about your own documents — fully offline.
Quick start: download from Releases, run FreeAiSsd.PrepApp.exe (Windows) or PrepApp.app (macOS), then follow Setup & Installation below. A condensed version is at docs/QUICKSTART.txt.
Windows PrepApp — Model Manager (top) and Drive Setup (bottom).
macOS screenshots coming soon.
Open bugs, fixes, and feature status are tracked on GitHub rather than hand-maintained here (so this list can't drift out of date):
This started as a way to take AI into the field with no cell signal — ham radio manuals, band plans, and reference docs loaded onto a pocket SSD so an LLM could answer questions about them miles from civilization. Then it turned out the same setup works really well as a voice-activated copilot in DCS: load aircraft manuals, import your HOTAS bindings, and hit a button on the throttle to ask questions mid-sortie without taking the VR headset off. Same drive. Same AI. Same offline-first idea.
A clear picture of what works where today. Mac support is actively expanding but not yet at feature parity with Windows.
| Feature | Windows PrepApp | macOS PrepApp |
|---|---|---|
| Format drive | ✅ NTFS or exFAT | ✅ exFAT only |
| Stage Ollama + pull models | ✅ Full | ✅ Full |
| Pull models from Hugging Face | ✅ Full (with token auth) | ✅ Full (with token auth) |
| Manage models on encrypted drives | ✅ Full | ✅ Full |
| Detect pre-configured drive | ✅ Full | ✅ Full |
| Target: Windows-only (NTFS) | ✅ Yes | ❌ macOS cannot format NTFS |
| Target: exFAT (cross-platform or Mac-only) | ✅ Yes | ✅ Yes |
| Encrypted config roundtrip | ✅ Full | ✅ Full (CryptoKit port) |
| Feature | Windows Runner | macOS Runner (beta) |
|---|---|---|
| Chat (non-RAG) | ✅ Full | ✅ Full |
| RAG chat (sources panel; inline citations opt-in) | ✅ Full | ⚠️ Query-only — reads a library indexed on Windows |
| Add / sweep / rebuild document library | ✅ Full | ❌ Not supported yet |
| Voice input (speech-to-text) | ✅ Whisper.cpp (fully local) | ✅ On-device dictation (SFSpeechRecognizer) |
| Voice output (text-to-speech) | ✅ SAPI + Piper | ✅ Native (AVSpeechSynthesizer) |
| HOTAS Push-to-Talk | ✅ DirectInput | ❌ Not ported yet |
| DCS Bindings Import | ✅ Full | ❌ Not supported yet |
| Network Mode (LAN API) | ✅ Full (v2) | ✅ Sidecar-hosted |
Web chat UI (/chat/, browser) | ✅ Full | ✅ Sidecar-hosted |
| Companion tray app (second PC) | ✅ Full | — |
| Headless CLI (RunnerCli) | ✅ Full | — |
| Unlock Windows-prepped encrypted SSD | ✅ Full | ✅ Full (CryptoKit) |
RAG on macOS: The Mac Runner can answer questions using a document library that was ingested and indexed on Windows. It cannot add new documents, watch folders, sweep, or rebuild the index — those operations require the Windows Runner. If you only have a Mac, you can still get RAG-backed answers by preparing the drive on Windows first (or on a second Windows machine).
[guide.pdf §Engine Start p.12], plus opt-in OCR to recover text from scanned/diagram pageshttp://HOST:41555/chat/); full assistant minus voice, with model/library pickers, RAG sources, and per-device history. Works from any LAN device including an iPadFreeAiSsd.RunnerCli) — terminal REPL for SSH/Tailscale access; streams chat, shows RAG sources, zero GUI deps<SSD>/Runner.app (drive root, double-click to launch; no zip to expand)SFSpeechRecognizer) and spoken responses (AVSpeechSynthesizer); audio never leaves the MacKnown limitations across all platforms:
pcm16le only; other codecs not implementedFor headless access from a terminal (including an iPad over Tailscale), FreeAiSsd.RunnerCli ships alongside the Windows Runner. It's a thin HTTP client against Runner's LAN API — same RAG pipeline, same source citations, no GUI.
$ FreeAiSsd.RunnerCli --help
$ FREEAI_URL=http://my-desk:41555 FREEAI_API_KEY=... FreeAiSsd.RunnerCli --model phi3
Target: http://my-desk:41555
Host reachable (ollamaRunning=True). Type /help for commands. Ctrl-C or 'exit' to quit.
phi3> what aircraft can I fly in DCS Open Beta?
...streamed response...
— sources: dcs-aircraft-list.pdf
phi3> /quit
Precedence for configuration: --url / --api-key flag > FREEAI_URL / FREEAI_API_KEY env var > default (http://127.0.0.1:41555, no key). Use --no-stream on very flaky links to fall back to a single-response round-trip.
You're in VR, mid-sortie, and can't remember the sequence to uncage an AIM-9. You reach for your HOTAS, key the mic, and ask. The AI answers with the buttons on your stick — sourced from the aircraft manual sitting on the drive. No internet. No cloud. No subscription.
What the Windows Runner does for flight sim:
Saved Games\DCS folder, scans your aircraft, and writes a per-aircraft reference file with your real button assignmentsSupported now in the Windows Runner: DCS World (stable and Open Beta), any aircraft with binding files in Config/Input, multi-device merging (stick + throttle + rudder pedals)
Planned: IL-2 Sturmovik and War Thunder binding parsers (see Roadmap)
Camping, deployed for emergency comms, or away from a desk — you need to reference your radio manual or band plan and there's no cell signal.
Load your manuals and reference documents onto the drive before you go. The AI indexes everything and answers from your own library, completely offline, from a drive that fits in your pocket.
Maybe you don't trust cloud AI with your data. Maybe your workplace restricts internet access. Maybe you want the same staged SSD available across your machines, with the full feature set on Windows and the current direct-chat beta on macOS.
Prepare the drive once. Your models, your documents, your config — nothing leaves the drive, no account needed, no telemetry.
Load first aid guides, plant identification references, equipment specs, survival manuals — whatever you need when there's no connectivity. The AI indexes it all and answers from your library when you're completely off-grid.
Which prep host (source OS) can produce which target drive:
| Source OS | Target | Filesystem | Supported |
|---|---|---|---|
| Windows | Windows-only | NTFS | Yes |
| Windows | Cross-platform (Windows + Mac) | exFAT | Yes |
| Windows | Mac-only | exFAT | Yes (APFS not available from Windows) |
| Mac | Mac-only | exFAT | Yes (APFS deferred from supported targets) |
| Mac | Cross-platform (Windows + Mac) | exFAT | Yes |
| Mac | Windows-only | NTFS | Not supported — use Windows PrepApp (macOS cannot natively format NTFS) |
Encrypted-config roundtrip is bidirectional. A drive prepped on Windows unlocks cleanly on Mac, and a drive prepped on Mac unlocks cleanly on Windows. The on-disk encrypted format (AES-256-GCM + PBKDF2-SHA256) is identical on both platforms and is pinned by cross-language tests.
Notes on the unsupported cells:
Stable (recommended): Download Free-AI-SSD-win.zip from Releases. Extract anywhere on Windows. Run FreeAiSsd.PrepApp.exe.
Beta cross-platform bundle: Free-AI-SSD-crossplatform.zip includes the Mac PrepApp (PrepApp.app) and Mac Runner beta (Runner.app) alongside the Windows artifacts. The macOS builds are currently unsigned/not notarized — see macOS first launch below before opening either app.
The download root is intentionally minimal — the prep tool(s), LICENSE, QUICKSTART.txt, and one dependencies/ folder for everything the prep tool consumes (no duplicated/nested copies):
Free-AI-SSD-win.zip Free-AI-SSD-crossplatform.zip
├── FreeAiSsd.PrepApp.exe ├── FreeAiSsd.PrepApp.exe (Windows prep)
├── LICENSE ├── PrepApp.app (macOS prep)
├── QUICKSTART.txt ├── LICENSE
└── dependencies/ ├── QUICKSTART.txt
├── runner/ └── dependencies/
├── companion/ ├── runner/ companion/ prereqs/
└── prereqs/ └── mac/ (Runner.app, ollama, manifest)
Cross-platform note: because macOS
.appbundles carry symlinks Windows archivers strip, prep a cross-platform drive's macOS side from a Mac (PrepApp.app). The Windows side works from either host. The Mac Runner is staged to the SSD root as<SSD>/Runner.app— unzipped, launchable directly, no zip to expand each run.
Until the signed/notarized release ships, the unsigned ad-hoc Mac apps trip Gatekeeper as soon as Safari stamps the downloaded ZIP with a quarantine xattr. The dialog reads "FreeAiSsd is damaged and can't be opened. You should move it to the Trash." even though nothing is corrupted — and right-click → Open / "Allow apps from anywhere" do not clear this state.
Strip the quarantine xattr once in Terminal, replacing the path with wherever you extracted the bundle:
xattr -dr com.apple.quarantine /path/to/PrepApp.app
xattr -dr com.apple.quarantine /path/to/Runner.app
Both apps then launch normally on double-click. This workaround goes away with the next signed release.
CI artifacts: Available from GitHub Actions for validation and testing. Prefer Releases for normal use.
Phase 1 — Prepare (online, once):
On Windows:
FreeAiSsd.PrepApp.exeOn Mac:
PrepApp.app from the cross-platform bundle. First launch only: run the Gatekeeper unblock xattr command above — the build is unsigned/not notarized, so Safari quarantine makes Gatekeeper claim the app is "damaged" until that bit is cleareddiskutil directly to format the drive as exFAT and lay out the canonical SSD directory structureThe resulting drive is byte-for-byte interchangeable with a Windows-prepped drive of the same target compatibility.
Phase 2 — Run (offline, anywhere):
<SSD>\windows\runner\FreeAiSsd.Runner.exe<SSD>/Runner.app (at the drive root — double-click)| Operation | Internet Required? |
|---|---|
| PrepApp — download, pull, staging | Yes |
| Windows Runner start / chat | No |
| macOS beta Runner start / sidecar-backed chat | No |
| macOS beta Runner — unlock Windows-prepped encrypted SSD | No |
| Reference Documents indexing and retrieval (Windows Runner) | No |
| Pull embedding model (if missing from SSD, Windows Runner) | Once |
| DCS Bindings Import (Windows Runner) | No |
| Voice input (Whisper transcription, Windows Runner) | No (model download is once) |
| Text-to-speech (Windows Runner) | No |
Runner won't start / dependency warnings
Missing embedding model while offline
PDF citations seem wrong or sparse
.NET / runtime prerequisites on target machine
macOS beta limitations
Every prerequisite and bundled third-party tool is fetched, verified, and recorded at runtime via a single shared resolver (shared/Prereqs/PrereqResolver.cs) used by both PrepApp and the CI offline-bundle builder (tools/FreeAiSsd.PrereqFetch). There are no hardcoded per-version SHA pins in the workflow or the catalog — stale pins were the failure mode we were hitting most often. Instead:
| Upstream | Version discovery | Integrity check | Trust basis |
|---|---|---|---|
| .NET 8 Desktop Runtime (x64) | https://builds.dotnet.microsoft.com/dotnet/release-metadata/8.0/releases.json → latest-release (rejects preview/rc builds) | SHA-512 from the same releases.json entry | Vendor-published hash over HTTPS to Microsoft's CDN |
| VC++ Redistributable (x64) | https://aka.ms/vs/17/release/vc_redist.x64.exe evergreen permalink | Observed SHA-256 recorded in manifest only | HTTPS-only trust to Microsoft aka.ms (no vendor per-version hash is published at a predictable URL) |
| Ollama (macOS, universal) | GitHub API releases/latest → picks Ollama-darwin.zip / ollama-darwin.zip asset | SHA-256 from the release's sha256sum.txt asset | Vendor-published hash over HTTPS to github.com |
Fail-closed invariants (CI and PrepApp both enforce):
latest-releaseThe prereqs-manifest.json that ships on the SSD records the resolved upstream URL, the vendor hash (when one was available), the observed SHA-256, and a short trustNote describing which trust basis was used — so offline installs can be audited without calling back to the upstream.
The Windows Runner includes a Reference Documents panel. Add your own files and the AI references them when answering instead of relying on training data alone. Retrieved chunks are cited inline so you can see exactly where an answer came from.
Supported formats: .pdf, .txt, .md, .json, .csv
Workflow:
How retrieval works: A hybrid retriever combines semantic (vector) search with keyword (BM25) search, fuses the results, and pulls neighboring chunks for surrounding context. The Sources list always shows what was used. Inline citations like [guide.pdf §Engine Start p.12] are opt-in (ragInlineCitations, off by default — answers stay concise) and are stripped before text-to-speech when enabled. If nothing relevant is found, the model is told so and won't invent context.
macOS: RAG queries work against a library indexed on Windows. Adding documents, folder sweeps, and index rebuilds require the Windows Runner.
Limitations:
ocrEnabled, off by default) to recover text from embedded images.The Windows Runner reads your DCS World controller bindings and writes them into the document library as a per-aircraft reference file. After import, when you ask "how do I uncage my AIM-9?" the AI answers with the button on your stick — not a generic keybind table.
How to import:
Saved Games\DCS folder — browse manually if detection failsdiff.lua (stick, throttle, rudder pedals), merges them into one file per aircraft, and writes it to your librarySupported:
Config/InputNot yet supported: IL-2 Sturmovik and War Thunder (see Roadmap)
In the Windows Runner, speak your questions and hear the answers. The entire pipeline runs locally — no cloud STT, no cloud TTS.
Speaking to the AI:
autoSendVoiceInput)AI voice response: Enable TTS in settings. Two engines available:
piper.exe (~22 MB) plus the default en_US-amy-medium voice (~60 MB) into windows/tools/piper/ or mac/tools/piper/, both SHA-256 verified.You can route TTS to a specific audio output device — useful for sending AI voice to your VR headset while system audio goes elsewhere.
Whisper model sizes (stored at models/whisper/ on the SSD):
| Size | File | Approx. disk | Notes |
|---|---|---|---|
| Tiny | ggml-tiny.bin | ~75 MB | Fastest; lower accuracy |
| Base | ggml-base.bin | ~142 MB | Default; good for most use |
| Small | ggml-small.bin | ~466 MB | Better accuracy |
| Medium | ggml-medium.bin | ~1.5 GB | Best accuracy; more RAM required |
The first time voice is used, Runner downloads the selected Whisper model (internet required for that one step). After that, fully offline.
In the Windows Runner, bind a button on your HOTAS to start and stop voice recording — no keyboard, no mouse. Built for VR where hands-free activation matters.
Setup:
"X-56 Rhino Throttle")push_to_talk — hold the button to record, release to sendtoggle — press once to start recording, press again to stop and sendOptional overlay: A small always-on-top window shows recording status. Disable it for VR where it would be distracting (pttOverlayEnabled).
Optional sound: A short beep plays on PTT activation/deactivation. Toggle with pttActivationSoundEnabled.
Full VR voice loop: HOTAS button → mic opens → speak → button release → Whisper transcribes → prompt sent → AI responds → TTS speaks into headset. Hands never leave the controls.
Companion tray app: Remote HOTAS/PTT is supported — the Companion app can run on a second PC and drive the full voice loop against Runner over LAN, including its own PTT activation beep and overlay.
Network Mode lets one Windows machine run Runner + Ollama locally, while other devices on your LAN call Runner's HTTP API.
Architecture:
127.0.0.1) on the host machineSecurity model (home LAN baseline):
127.0.0.1 (loopback) by default. Binding to 0.0.0.0 (all interfaces) is an explicit opt-in set in portable-config.json; Runner logs a WARNING on startup whenever the effective bind address is not loopback.Authorization: Bearer <key> or X-API-Key)Endpoints:
GET /api/healthGET /api/modelsPOST /api/chatPOST /api/chat/stream (newline-delimited JSON stream)POST /api/stt/transcribe (multipart upload: audio)POST /api/voice/query (multipart upload: audio, optional model, autoSendToChat, speakResponse, returnAudio). When returnAudio=true the response includes AudioBase64 + AudioMime so the client can play TTS locally instead of on the host.POST /api/tts/speakPOST /api/tts/stopExample cURL requests:
# health (no API key required)
curl http://RUNNER_HOST:41555/api/health
# list models
curl -H "Authorization: Bearer YOUR_KEY" \
http://RUNNER_HOST:41555/api/models
# non-stream chat
curl -X POST http://RUNNER_HOST:41555/api/chat \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi3","prompt":"Summarize startup checklist"}'
# stream chat (NDJSON)
curl -N -X POST http://RUNNER_HOST:41555/api/chat/stream \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"model":"phi3","prompt":"Step-by-step A-10C startup"}'
# trigger host-side TTS
curl -X POST http://RUNNER_HOST:41555/api/tts/speak \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{"text":"Radio check complete."}'
# STT transcription (WAV upload)
curl -X POST http://RUNNER_HOST:41555/api/stt/transcribe \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav"
# Voice query (upload -> transcribe -> chat -> optional host-side TTS)
curl -X POST http://RUNNER_HOST:41555/api/voice/query \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav" \
-F "model=phi3" \
-F "autoSendToChat=true" \
-F "speakResponse=true"
# Voice query with client-side TTS playback (returnAudio)
# Response JSON contains AudioBase64 + AudioMime ("audio/wav") for local playback.
curl -X POST http://RUNNER_HOST:41555/api/voice/query \
-H "Authorization: Bearer YOUR_KEY" \
-F "audio=@question.wav;type=audio/wav" \
-F "model=phi3" \
-F "autoSendToChat=true" \
-F "speakResponse=true" \
-F "returnAudio=true"
Remote voice upload formats and limits:
format=pcm16le)networkMaxAudioUploadMBreturnAudio=true returns synthesized TTS as WAV PCM 16-bit mono 16kHz (AudioMime: "audio/wav", AudioBase64), bounded by networkMaxAudioUploadMBWeb chat / LAN access (no install):
Runner serves a standalone browser chat client from the same LAN API host — no app to install on the other device. It's the full assistant minus voice: model picker, document-library selection, RAG-grounded chat with sources, a collapsible thinking view, and per-request temperature/thinking controls. Chat history is saved per-device in the browser (localStorage).
On the host PC the web UI runs on loopback automatically whenever Ollama is up — just click "Open Chat UI" in Runner (or open http://127.0.0.1:41555/chat/). No Network Mode, no API key for local use.
To reach it from other devices:
http://RUNNER_HOST:41555/chat/ (use the host's LAN IP or HOSTNAME.local).Per-request model parameters set in the web UI (temperature, thinking) apply only to that request and never overwrite the host's saved configuration. Plain HTTP over a trusted home LAN, API-key gated — do not expose port 41555 to the public internet.
All settings live in config/portable-config.json on the SSD.
| Property | Default | Description |
|---|---|---|
ollamaPort | 11434 | TCP port for the local Ollama server |
preferredCompute | "auto" | Compute mode: "auto" (detected GPU — AMD→Vulkan, NVIDIA→CUDA, Intel→Vulkan) or "cpu" (force CPU). Legacy "cuda"/"rocm" are treated as "auto" |
useStreamingChat | true | Stream tokens as they generate; falls back to non-streaming if streaming fails |
Per-chat overrides for the active model. Sentinel values mean "use the model's built-in default," so the app stays compatible with any GGUF model. The web/desktop UI can set these per request without changing the saved config.
| Property | Default | Description |
|---|---|---|
modelContextWindow | 0 | Override Ollama num_ctx; 0 = model default |
modelTemperature | -1 | Override temperature (0.0–2.0); -1 = model default |
modelTopP | -1 | Override top_p (0.0–1.0); -1 = model default |
modelMaxOutputTokens | -1 | Override num_predict (max tokens per response); -1 = unbounded. Also caps the thinking budget. |
modelThinkMode | "" | Reasoning control: "" = model default, "off", "low", "medium", "high" (only for models that support thinking) |
| Property | Default | Description |
|---|---|---|
activeDocumentLibraryId | null | Active library ID; null disables RAG |
retrievalTopK | 8 | Number of chunks retrieved per query |
hybridRetrievalEnabled | true | Fuse semantic (vector) + keyword (BM25) search; false = vector-only. No reindex needed. |
retrievalNeighborRadius | 1 | Chunks pulled on each side of a hit for context; 0 disables |
ragInlineCitations | false | true appends inline labels like [guide.pdf §Engine Start p.12] (stripped before TTS); false = concise answers, sources still shown in the Sources panel |
chunkSize | 1200 | Characters per chunk during indexing |
chunkOverlap | 200 | Characters of overlap between adjacent chunks |
embeddingModelName | "nomic-embed-text" | Embedding model served by local Ollama |
minimumSimilarityThreshold | 0.3 | Minimum cosine similarity (0.0–1.0) for a chunk to be included; lower = more permissive |
maxEmbeddingConcurrency | 4 | Concurrent embedding requests during ingestion |
maxDocumentSizeMB | 512 | Max file size (MB) accepted for ingestion |
Off by default. When enabled (and a Tesseract bundle is staged on the SSD), OCR recovers text baked into images inside PDFs (e.g. cockpit MFD labels, diagrams) and adds it as supplementary searchable chunks — the clean text layer is never replaced.
| Property | Default | Description |
|---|---|---|
ocrEnabled | false | Run OCR over embedded PDF images during ingestion |
ocrLanguage | "eng" | Tesseract language code(s) passed to -l |
ocrMinImagePixels | 10000 | Skip images smaller than this (width × height) — filters out icons, rules, logos |
ocrMinWordConfidence | 55 | Drop OCR words below this confidence (0–100) to suppress garble |
ocrMaxImagesPerFile | 4000 | Hard cap on images OCR'd per file |
ocrPerImageTimeoutSeconds | 30 | Per-image OCR timeout (a stuck image is skipped) |
| Property | Default | Description |
|---|---|---|
whisperModelSize | "Base" | Whisper model: "Tiny", "Base", "Small", or "Medium" |
selectedMicrophoneDevice | null | Microphone device name; null = system default |
autoSendVoiceInput | true | true sends transcribed text immediately; false puts it in the prompt field for review |
| Property | Default | Description |
|---|---|---|
ttsEnabled | false | Enable TTS for AI responses |
ttsEngine | "system" | "system" (Windows SAPI) or "piper" (neural TTS) |
ttsVoiceName | null | Voice name for the selected engine; null = engine default |
ttsRate | 0 | Speech rate: -10 (slowest) to 10 (fastest) |
ttsVolume | 100 | Volume: 0 (silent) to 100 (max) |
ttsOutputDevice | null | Audio output device for TTS; null = system default |
| Property | Default | Description |
|---|---|---|
pttEnabled | false | Enable HOTAS push-to-talk |
pttDeviceName | null | DirectInput device name (e.g., "X-56 Rhino Throttle") |
pttButtonIndex | 0 | Zero-based button index on the joystick device |
pttMode | "push_to_talk" | "push_to_talk" (hold to record) or "toggle" (press to start/stop) |
pttActivationSoundEnabled | true | Play a beep on PTT activation/deactivation |
pttOverlayEnabled | true | Show the always-on-top PTT status overlay |
pttOverlayX / pttOverlayY | 20 / 20 | Overlay window position in pixels from top-left |
The Companion tray app exposes the same two cues under identical key names (pttActivationSoundEnabled, pttOverlayEnabled) in companion-config.json, toggleable from Companion's Settings window.
| Property | Default | Description |
|---|---|---|
networkModeEnabled | false | Expose the Runner API on the LAN. The API always runs on loopback when Ollama is up (for the on-PC web UI); this flag rebinds it to networkBindAddress and enforces the API key. |
networkBindAddress | "127.0.0.1" | Bind address used when exposed on the LAN (typically 0.0.0.0). Ignored — forced to loopback — when networkModeEnabled is off. |
networkPort | 41555 | TCP port for Runner API |
networkApiKey | "" | Shared secret for API auth |
networkRequireApiKey | true | Require API key on all non-health endpoints |
networkAllowTts | false | Allow remote callers to trigger host-side TTS |
networkAllowRemoteStt | false | Allow remote audio upload transcription via /api/stt/transcribe |
networkAllowRemoteVoiceQuery | false | Allow remote voice-query orchestration via /api/voice/query |
networkVoiceAutoSendToChat | true | Default for voice query: auto-send transcription to chat when request omits override |
networkMaxAudioUploadMB | 10 | Maximum upload size in MB for remote STT/voice endpoints |
Core infrastructure:
Windows Runner (full feature set):
returnAudio local TTSModel management (Windows + Mac):
num_ctx / temperature / top_p / max-output (num_predict), plus a Thinking control (Ollama think: Off / Low / Medium / High) to disable or cap reasoning models that loop; live on both the Windows and Mac runnersmacOS:
Free-AI-SSD ships several components backed by a shared cross-platform library:
prep-app/) — runs on an online Windows machine to configure the SSD: picks drive, downloads and stages Ollama, pulls models (Ollama or Hugging Face), bundles prerequisites, finalizes layoutmac-prep-app/, SwiftUI) — native macOS PrepApp for the cross-platform bundle. Drives diskutil directly to format target SSDs as exFAT, stages the runner / Ollama / prereq payloads via the mac-prep-host net8.0 sidecar (which consumes prep-core/), and writes encrypted config in a format byte-identical to the Windows PrepApp. Apple Silicon (arm64), macOS 11+runner/) — runs from the SSD on the target machine; starts Ollama, provides the chat interface, manages document libraries, voice pipeline, HOTAS PTT, and the LAN API hostmac-runner/, Swift) — thin macOS app shipped at <SSD>/Runner.app (drive root). It selects/infers the SSD, unlocks encrypted config, reads installed models, starts macOS Ollama, spawns the local Runner API sidecar, and sends chat through the shared RAG pipeline when an active indexed library existsAudioCaptureService → WhisperSpeechToTextService → ChatService → SystemTextToSpeechService / PiperTextToSpeechService, orchestrated by PttVoicePipelineService when HOTAS PTT is enabledshared/Documents/) — DcsSavedGamesLocator finds DCS installs, DcsAircraftScanner enumerates aircraft, DcsBindingParser parses diff.lua, DcsBatchProcessor merges devices and writes RAG documentscompanion/, WPF tray app) — optional lightweight client for a second LAN machine; no SSD required; talks to the Runner LAN API for chat / STT upload / voice-query / host-side TTS. Supports its own HOTAS PTT loop, activation beep, status overlay, and mic-preflight checkFreeAiSsd.Shared, net8.0) — common core logic for encryption, trust policy, path guards, config, dependency checking, download management, MVVM infrastructure, DCS binding models, document library, and RAG pipelineRunner's business logic lives in injectable services with no UI dependencies, enabling unit testing without a WPF host:
| Service | Purpose |
|---|---|
OllamaLifecycleService | Process start/stop, port resolution, trust validation |
ModelManagementService | Installed model listing, sizing warnings, embedding model pull |
DocumentOperationsService | Library CRUD, file ingestion, folder sweep, index rebuild |
ChatService | RAG-augmented prompt construction and Ollama /api/generate calls |
DcsBindingsImportService | DCS installation detection, aircraft scanning, batch binding import |
WhisperSpeechToTextService | Whisper.cpp transcription via Whisper.net; model download management |
SystemTextToSpeechService | Windows SAPI TTS with optional NAudio device targeting |
PiperTextToSpeechService | Piper neural TTS; spawns piper.exe, streams raw PCM through NAudio |
AudioCaptureService | Microphone capture at 16 kHz/16-bit mono (Whisper's required format) |
HotasInputService | DirectInput polling for HOTAS PTT button state |
PttVoicePipelineService | Orchestrates the full PTT → record → transcribe → send → TTS loop |
minimumSimilarityThreshold (default 0.3) are discarded; the model is told explicitly when nothing relevant was foundmaxEmbeddingConcurrency)System.Numerics.Vector<float>; top-K uses an O(N log K) priority queue| Control | Detail |
|---|---|
| Encrypted config | AES-256-GCM with PBKDF2-SHA256 (210,000 iterations) |
| Config write guard | ConfigStore is the only path for config writes; a plaintext write to an encrypted drive throws InvalidOperationException (fail-closed) |
| Package trust | Ollama downloads validated against URL allowlist + SHA-256 digest before execution; macOS payloads additionally verified as arm64 Mach-O |
| Fail-closed write guard | PrepDriveWriteGuard blocks all writes to encrypted drives if encryption state is ambiguous |
| Path traversal prevention | PathGuards enforces sibling boundary checks with platform-aware case sensitivity |
| Shell injection prevention | ProcessRunner uses ArgumentList, not string concatenation |
A security review on 2026-02-19 found no critical vulnerabilities in the audited surface. The invariants above are enforced in code and tests; ongoing issues are tracked on GitHub Issues.
config/ — portable-config.json (plaintext) or portable-config.encrypted.json (opt-in)
models/ — Ollama model store
models/whisper/ — Whisper STT model files (ggml-*.bin)
logs/ — app logs
docs/libraries/ — Reference Documents library files, manifests, index DB
windows/runner/ — Runner app
windows/tools/ollama/ — staged Ollama runtime + trust attestation
windows/tools/piper/ — optional Piper TTS binary and voice models (user-installed)
windows/tools/tesseract/ — optional Tesseract OCR engine + tessdata
windows/tools/prereqs/ — offline prerequisite installers + manifest
Runner.app/ — macOS Runner bundle (root-level, launchable directly)
mac/tools/ollama/ — staged macOS Ollama runtime + trust attestation
mac/tools/piper/ — optional Piper TTS binary and voice models
mac/tools/tesseract/ — optional Tesseract OCR engine + tessdata
cache/ — prep-time download cache
SsdLayout in the shared library is the single source of truth for these paths — always use it rather than constructing paths manually.
| Directory | Target | Purpose |
|---|---|---|
shared/ | net8.0 | Cross-platform shared library (FreeAiSsd.Shared) |
runner-core/ | net8.0 | Platform-neutral Runner business logic (chat, RAG, library, local API) shared by Windows Runner and the Mac runner-host sidecar |
prep-core/ | net8.0 | Platform-neutral PrepApp business logic (manifest, staging, prereq, encrypted config, HF/Ollama model pulls) shared by Windows PrepApp and the Mac prep-host sidecar |
prep-app/ | net8.0-windows | WPF PrepApp (Windows) |
mac-prep-app/ | macOS (Swift) | Native SwiftUI PrepApp (Mac); produces drives byte-identical to Windows PrepApp |
mac-prep-host/ | net8.0 | osx-arm64 sidecar that runs prep-core/ business logic for the Mac PrepApp over a stdin command protocol |
runner/ | net8.0-windows | WPF Runner (Windows) |
mac-runner/ | macOS (Swift) | Swift macOS beta Runner over the local Runner API sidecar |
mac-runner-host/ | net8.0 | osx-arm64 sidecar that hosts RunnerLocalApiService for the Mac Runner |
runner-cli/ | net8.0 | Headless CLI client (FreeAiSsd.RunnerCli) — SSH/Tailscale terminal access to Runner API |
companion/ | net8.0-windows | WPF Companion tray client (LAN second-PC use) |
tools/FreeAiSsd.PrereqFetch/ | net8.0 | CI helper that pre-builds the offline prereq bundle via the shared PrereqResolver |
tests/ | net10.0 | xUnit test project (FreeAiSsd.Tests) |
tests-ocr/ | net10.0 | xUnit OCR test project (FreeAiSsd.Tests.Ocr) — Tesseract OCR coverage |
docs/ | — | Documentation (includes QUICKSTART.txt) |
| File | Purpose |
|---|---|
DependencyChecker.cs | Detects missing VC++ / .NET runtimes via registry + process checks |
DownloadManager.cs | Resumable HTTP downloads with progress callbacks |
DriveInspector.cs | Enumerates candidate drives |
ModelSizing.cs | Maps model tags to RAM/VRAM/disk requirements for sizing warnings |
NetUtils.cs | Port availability checking |
OllamaPackageTrustPolicy.cs | URL allowlisting + SHA-256 digest verification |
PathGuards.cs | Path traversal prevention |
PortableConfig.cs | JSON config serialization with atomic writes |
PrepDriveWriteGuard.cs | Blocks writes to encrypted drives (fail-closed) |
PrereqInstallValidator.cs | Validates installer integrity (SHA-256) before execution |
Prereqs/PrereqResolver.cs | Runtime discovery of the latest stable upstream prereq versions + vendor-hash verification. Shared by PrepApp and CI. |
ProcessRunner.cs | Safe process spawning via ArgumentList, not string concatenation |
SsdEncryption.cs | AES-256-GCM config encryption |
SsdLayout.cs | Canonical directory structure constants and creation |
SsdLogger.cs | File-based logger writing to the SSD's logs directory |
SystemCompatibility.cs | GPU/CPU/OS detection for compatibility display |
Documents/DcsBindingParser.cs | Parses DCS diff.lua files into structured data for RAG |
Documents/DcsAircraftScanner.cs | Scans Config/Input for aircraft folders and device files |
Documents/DcsBatchProcessor.cs | Batch import: merges devices, formats output, writes to library |
Documents/DcsSavedGamesLocator.cs | Auto-detects Saved Games\DCS and .openbeta; supports manual override |
PrepViewModel lives in shared/ (net8.0) so it can be unit tested on Linux without WPFIDialogService abstracts all MessageBox/dialog interactionsshared/, implementations in prep-app/Services/ (net8.0-windows)MainWindow.xaml.cs reduced from ~1,800 lines to ~95 lines; all logic in PrepViewModel and servicesShared + tests (all platforms):
dotnet build shared/FreeAiSsd.Shared.csproj
dotnet build tests/FreeAiSsd.Tests.csproj
dotnet test tests/FreeAiSsd.Tests.csproj --verbosity normal
1 test (
IsPathUnderRoot_WindowsBoundaryIsRespected) is expected to fail on Linux — it tests Windows-specific path behavior.
Full build (Windows only):
dotnet restore FreeAiSsd.sln
dotnet build FreeAiSsd.sln -c Release
dotnet test FreeAiSsd.sln -c Release
Stage Runner payload into PrepApp output:
./build.ps1 -Configuration Release -Runtime win-x64
Key dependencies: xUnit, System.Management, Moq, PdfPig, SharpDX (DirectInput), ASP.NET Core (in-process LAN host), Microsoft.Extensions.DependencyInjection, SQLite
~1,000+ test cases ([Fact]/[Theory]; Theories expand to more at runtime) across ~100 test files in tests/ and tests-ocr/. One Windows-specific path test is expected to fail on Linux. Coverage spans:
CI (.github/workflows/build.yml) is the source of truth for the current count and pass/fail status.
Signing is disabled by default in CI (MAC_SIGNING_ENABLED=false). Supported via repository secrets: MACOS_CERT_P12_BASE64, MACOS_CERT_PASSWORD, APPLE_TEAM_ID, APPLE_ID, APPLE_APP_SPECIFIC_PASSWORD, MACOS_SIGN_IDENTITY.
See GitHub Releases for the full version history and release notes.
C#
64.1%
HTML
19.9%
Swift
14.2%