Native local-first macOS speech-to-text, live captions, subtitle editing, offline translation, and on-device Gemma 4 transcript enhancement.
See the codeEnglish | Chinese (Simplified) | Japanese
Turn microphones, media files, and the sound playing on your Mac into editable text and subtitles—locally.
LocalScribe is a free, open-source native speech-to-text and audio/video transcription app for Apple silicon Macs, with no account required. It combines local recognition, floating live captions, subtitle editing and export, offline translation, long-task recovery, and on-device Gemma 4 transcript enhancement behind one SwiftUI interface. Audio, imported transcripts, and AI processing content are not uploaded by the app.
Current version: 1.7.0 (40) · Download DMG · First-launch installation guide · User documentation
[!TIP] New in 1.7.0 — original microphone audio: Optionally save the original recording for a new microphone transcription, play it back, and export an M4A alongside the transcript. Choose one of three quality profiles. Language changes now update menu titles, the AI prompt editor, and transcript export dialogs without restarting. Speech synthesis is disabled in this release.
An eight-second loop of the home screen and editable transcript view, captured from the real 1.6.6 app with a non-private English sample. Images and GIFs include light and dark versions that follow your viewing theme.
[!IMPORTANT] Apple SpeechAnalyzer recognition and floating live captions require macOS 26. The app itself supports macOS 15.5 or later. On macOS 15.5–25, manually select Whisper, SenseVoice, or Parakeet on the home screen. SenseVoice and Parakeet currently support file transcription only.
Installing for the first time on macOS 26 or 27? Follow the step-by-step first-launch guide, including what to do if Open Anyway is missing.
For illustrated steps, troubleshooting, and SHA-256 verification, see the Download and Installation Guide.
[!WARNING] This build has an ad-hoc integrity signature only. It is neither Developer ID signed nor notarized by Apple. Override the warning only if you trust this repository and its Release. SHA-256 files are included with every Release.
Transcription is only the starting point. After transcription or transcript import, open AI Enhancement in the right inspector to turn the result into a cleaner, deliverable document without copying it to another AI service or uploading it.
Gemma removes filler, repetition, and obvious transcription errors while preserving the intended meaning. Original and suggested text update in a live diff alongside batch progress, model status, and the on-device privacy indicator. Segments that fail validation keep their original text.
Summaries replace the text preview while keeping an undo snapshot of the original. Add one-time instructions for names, terminology, writing style, summary length, priorities, or format, and optionally save reusable instructions in Settings. AI output should still be reviewed.
The default Gemma 4 E2B IT Q4 model is about 2.8 GB; an optional E4B model of about 4.6 GB can be enabled in Settings. Models are downloaded on demand, verified, and run on the Mac through llama.cpp and Metal. On Macs with 8 GB of physical memory or less, AI Enhancement is off by default and can be enabled manually in Settings after confirming the memory warning. Gemma starts loading only after you select Proofread or Summarize. Completing a transcription or opening its result never preloads Gemma. Before loading Gemma, LocalScribe releases its active recognition and NLLB translation runtimes; the Gemma helper exits when the task finishes.
The screenshots are from version 1.6.6 (37) of the real macOS app and use non-private English sample text. Both appearances are captured from the app, without the Computer Use pointer. LocalScribe includes complete English, Simplified Chinese, and Japanese interfaces.
For a new microphone transcription, enable Save Original Audio in the settings sidebar before starting. This option is off for each new task and becomes locked once recording starts. Storage Priority, Quality Priority, and Highest Quality produce M4A files using AAC or ALAC. The app adapts to the microphone format automatically.
The recording keeps silence and other captured sounds independently of recognition. Pause gaps are omitted. Once saving finishes, use play/pause, seeking, and 0.5×, 1×, or 2× playback. Export the original recording alongside TXT, Markdown, JSON, PDF, SRT, or WebVTT with the same base filename. Editing or translating the transcript does not change the original audio.
Original-audio saving is unavailable for media-file transcription, Mac system audio, floating captions, and microphone transcription appended to an imported transcript. If audio saving fails, recognition can continue; any recoverable partial audio is labeled. The latest recoverable task retains its audio locally. Speech synthesis and its settings are disabled in this release.
LocalScribe follows the preferred language order in macOS by default. You can also switch the app immediately among English, Simplified Chinese, and Japanese in Settings without changing the system language; the menu bar and live-caption source controls update at the same time. Recognition-language menus place English and the Mac's language in a deduplicated Recommended Languages section. Settings also provide System, Light, and Dark appearance choices, plus current microphone, speech-recognition, and system-audio permission status.
Prefer Apple's built-in engines whenever possible. If your macOS version and languages are supported, use Apple Speech for transcription and Apple Translation for translation. Their native optimization and output quality are noticeably better than those of the third-party transcription and translation models available in LocalScribe.
| Engine | Use | Runtime |
|---|---|---|
| Apple Speech | Microphone, files, live captions | SpeechAnalyzer / SpeechTranscriber |
| Whisper | Microphone and files | whisper.cpp GGML, Metal with CPU fallback |
| SenseVoice | Files | sherpa-onnx, Core ML eligible path with CPU fallback |
| NVIDIA Parakeet | Files | sherpa-onnx, Core ML eligible path with CPU fallback |
| Apple Translation | Default post-transcription translation | macOS Translation framework |
| NLLB | Optional post-transcription translation | CTranslate2 CPU/int8 |
| Gemma 4 | Transcript proofreading, refinement, and summarization | On-device llama.cpp inference with Metal |
Whisper file transcription uses the model's internal sliding windows rather than fixed, non-overlapping application chunks. For longer media, the bundled Silero VAD can skip silence while retaining speech padding and overlap. Output filtering considers silence, confidence, repetition, and known hallucination patterns.
For supported third-party engines, the right inspector includes an Advanced section that is collapsed when a task opens. Whisper exposes a model prompt, temperature and fallback controls, beam or greedy search settings, context length, silence and confidence filters, VAD, and a guarded automatic thread policy. SenseVoice and Parakeet expose the options their local runtimes actually support. File-only and microphone-only settings appear only for the matching source.
Numeric options support both direct keyboard entry and sliders or steppers, with range validation and short hover explanations. The app remembers the last engine, model, and advanced configuration for the next task; Restore Model Defaults resets only the selected engine.
arm64); Intel Macs are not supported.Microphone input requires microphone permission. Capturing Mac audio requires Screen & System Audio Recording permission. Apple Speech and Apple Translation may download language assets managed by macOS.
Install Xcode and its Command Line Tools. The repository includes the native runtimes required by the app; large recognition, NLLB, and Gemma 4 models are downloaded only when selected.
ruby generate_project.rb
xcodebuild \
-project LocalScribe.xcodeproj \
-scheme LocalScribe \
-destination 'platform=macOS,arch=arm64' \
build
Run the test suite:
xcodebuild \
-project LocalScribe.xcodeproj \
-scheme LocalScribe \
-destination 'platform=macOS,arch=arm64' \
test
CI runs the tests and Release static analysis on GitHub's macOS 26 Apple silicon runner.
The local packaging script creates ZIP and DMG artifacts and validates nested signatures, architectures, deployment targets, embedded Mach-O files, archive extraction, DMG mounting, and helper startup:
./tools/package_local_release.sh
By default it uses the configured local certificate. Set CODESIGN_IDENTITY=- to make the same ad-hoc package produced by GitHub Actions. Pushing a tag that matches the version in Info.plist (for example, v1.7.0) or includes the build number (for example, v1.7.0-build40) runs release-unsigned.yml, verifies the package, and creates the GitHub Release without storing a certificate or password in GitHub Secrets. Developer ID signing, timestamping, notarization, and stapling remain the preferred public distribution path.
LocalScribe.app/Contents/MacOS/LocalScribe --cli help
LocalScribe.app/Contents/MacOS/LocalScribe --cli models --json
LocalScribe.app/Contents/MacOS/LocalScribe --cli transcribe input.mp4 \
--engine whisper --language en_US --format srt --output output.srt
LocalScribe does not upload recognition audio, imported transcripts, or content processed by Gemma. Recognition, translation, and AI transcript enhancement run on the Mac. Network access is used for on-demand recognition, translation, and Gemma model downloads, automatic or manual update checks and downloads, and opening external documentation. Automatic updates can be disabled in Settings.
The built-in updater checks at launch and every six hours by default, reads update.json from the latest GitHub Release, downloads its ZIP, verifies SHA-256, checks the bundle identifier and version, and asks before replacing and relaunching the app. This update path is not a substitute for a notarized public release.
LocalScribe source code is available under the MIT License. Third-party components and models retain their original licenses; see THIRD_PARTY_NOTICES.md and the license files under Vendor.
The optional NLLB model is distributed upstream under CC-BY-NC-4.0.
Native local-first macOS speech-to-text, live captions, subtitle editing, offline translation, and on-device Gemma 4 transcript enhancement.
See the codeEnglish | Chinese (Simplified) | Japanese
Turn microphones, media files, and the sound playing on your Mac into editable text and subtitles—locally.
LocalScribe is a free, open-source native speech-to-text and audio/video transcription app for Apple silicon Macs, with no account required. It combines local recognition, floating live captions, subtitle editing and export, offline translation, long-task recovery, and on-device Gemma 4 transcript enhancement behind one SwiftUI interface. Audio, imported transcripts, and AI processing content are not uploaded by the app.
Current version: 1.7.0 (40) · Download DMG · First-launch installation guide · User documentation
[!TIP] New in 1.7.0 — original microphone audio: Optionally save the original recording for a new microphone transcription, play it back, and export an M4A alongside the transcript. Choose one of three quality profiles. Language changes now update menu titles, the AI prompt editor, and transcript export dialogs without restarting. Speech synthesis is disabled in this release.
An eight-second loop of the home screen and editable transcript view, captured from the real 1.6.6 app with a non-private English sample. Images and GIFs include light and dark versions that follow your viewing theme.
[!IMPORTANT] Apple SpeechAnalyzer recognition and floating live captions require macOS 26. The app itself supports macOS 15.5 or later. On macOS 15.5–25, manually select Whisper, SenseVoice, or Parakeet on the home screen. SenseVoice and Parakeet currently support file transcription only.
Installing for the first time on macOS 26 or 27? Follow the step-by-step first-launch guide, including what to do if Open Anyway is missing.
For illustrated steps, troubleshooting, and SHA-256 verification, see the Download and Installation Guide.
[!WARNING] This build has an ad-hoc integrity signature only. It is neither Developer ID signed nor notarized by Apple. Override the warning only if you trust this repository and its Release. SHA-256 files are included with every Release.
Transcription is only the starting point. After transcription or transcript import, open AI Enhancement in the right inspector to turn the result into a cleaner, deliverable document without copying it to another AI service or uploading it.
Gemma removes filler, repetition, and obvious transcription errors while preserving the intended meaning. Original and suggested text update in a live diff alongside batch progress, model status, and the on-device privacy indicator. Segments that fail validation keep their original text.
Summaries replace the text preview while keeping an undo snapshot of the original. Add one-time instructions for names, terminology, writing style, summary length, priorities, or format, and optionally save reusable instructions in Settings. AI output should still be reviewed.
The default Gemma 4 E2B IT Q4 model is about 2.8 GB; an optional E4B model of about 4.6 GB can be enabled in Settings. Models are downloaded on demand, verified, and run on the Mac through llama.cpp and Metal. On Macs with 8 GB of physical memory or less, AI Enhancement is off by default and can be enabled manually in Settings after confirming the memory warning. Gemma starts loading only after you select Proofread or Summarize. Completing a transcription or opening its result never preloads Gemma. Before loading Gemma, LocalScribe releases its active recognition and NLLB translation runtimes; the Gemma helper exits when the task finishes.
The screenshots are from version 1.6.6 (37) of the real macOS app and use non-private English sample text. Both appearances are captured from the app, without the Computer Use pointer. LocalScribe includes complete English, Simplified Chinese, and Japanese interfaces.
For a new microphone transcription, enable Save Original Audio in the settings sidebar before starting. This option is off for each new task and becomes locked once recording starts. Storage Priority, Quality Priority, and Highest Quality produce M4A files using AAC or ALAC. The app adapts to the microphone format automatically.
The recording keeps silence and other captured sounds independently of recognition. Pause gaps are omitted. Once saving finishes, use play/pause, seeking, and 0.5×, 1×, or 2× playback. Export the original recording alongside TXT, Markdown, JSON, PDF, SRT, or WebVTT with the same base filename. Editing or translating the transcript does not change the original audio.
Original-audio saving is unavailable for media-file transcription, Mac system audio, floating captions, and microphone transcription appended to an imported transcript. If audio saving fails, recognition can continue; any recoverable partial audio is labeled. The latest recoverable task retains its audio locally. Speech synthesis and its settings are disabled in this release.
LocalScribe follows the preferred language order in macOS by default. You can also switch the app immediately among English, Simplified Chinese, and Japanese in Settings without changing the system language; the menu bar and live-caption source controls update at the same time. Recognition-language menus place English and the Mac's language in a deduplicated Recommended Languages section. Settings also provide System, Light, and Dark appearance choices, plus current microphone, speech-recognition, and system-audio permission status.
Prefer Apple's built-in engines whenever possible. If your macOS version and languages are supported, use Apple Speech for transcription and Apple Translation for translation. Their native optimization and output quality are noticeably better than those of the third-party transcription and translation models available in LocalScribe.
| Engine | Use | Runtime |
|---|---|---|
| Apple Speech | Microphone, files, live captions | SpeechAnalyzer / SpeechTranscriber |
| Whisper | Microphone and files | whisper.cpp GGML, Metal with CPU fallback |
| SenseVoice | Files | sherpa-onnx, Core ML eligible path with CPU fallback |
| NVIDIA Parakeet | Files | sherpa-onnx, Core ML eligible path with CPU fallback |
| Apple Translation | Default post-transcription translation | macOS Translation framework |
| NLLB | Optional post-transcription translation | CTranslate2 CPU/int8 |
| Gemma 4 | Transcript proofreading, refinement, and summarization | On-device llama.cpp inference with Metal |
Whisper file transcription uses the model's internal sliding windows rather than fixed, non-overlapping application chunks. For longer media, the bundled Silero VAD can skip silence while retaining speech padding and overlap. Output filtering considers silence, confidence, repetition, and known hallucination patterns.
For supported third-party engines, the right inspector includes an Advanced section that is collapsed when a task opens. Whisper exposes a model prompt, temperature and fallback controls, beam or greedy search settings, context length, silence and confidence filters, VAD, and a guarded automatic thread policy. SenseVoice and Parakeet expose the options their local runtimes actually support. File-only and microphone-only settings appear only for the matching source.
Numeric options support both direct keyboard entry and sliders or steppers, with range validation and short hover explanations. The app remembers the last engine, model, and advanced configuration for the next task; Restore Model Defaults resets only the selected engine.
arm64); Intel Macs are not supported.Microphone input requires microphone permission. Capturing Mac audio requires Screen & System Audio Recording permission. Apple Speech and Apple Translation may download language assets managed by macOS.
Install Xcode and its Command Line Tools. The repository includes the native runtimes required by the app; large recognition, NLLB, and Gemma 4 models are downloaded only when selected.
ruby generate_project.rb
xcodebuild \
-project LocalScribe.xcodeproj \
-scheme LocalScribe \
-destination 'platform=macOS,arch=arm64' \
build
Run the test suite:
xcodebuild \
-project LocalScribe.xcodeproj \
-scheme LocalScribe \
-destination 'platform=macOS,arch=arm64' \
test
CI runs the tests and Release static analysis on GitHub's macOS 26 Apple silicon runner.
The local packaging script creates ZIP and DMG artifacts and validates nested signatures, architectures, deployment targets, embedded Mach-O files, archive extraction, DMG mounting, and helper startup:
./tools/package_local_release.sh
By default it uses the configured local certificate. Set CODESIGN_IDENTITY=- to make the same ad-hoc package produced by GitHub Actions. Pushing a tag that matches the version in Info.plist (for example, v1.7.0) or includes the build number (for example, v1.7.0-build40) runs release-unsigned.yml, verifies the package, and creates the GitHub Release without storing a certificate or password in GitHub Secrets. Developer ID signing, timestamping, notarization, and stapling remain the preferred public distribution path.
LocalScribe.app/Contents/MacOS/LocalScribe --cli help
LocalScribe.app/Contents/MacOS/LocalScribe --cli models --json
LocalScribe.app/Contents/MacOS/LocalScribe --cli transcribe input.mp4 \
--engine whisper --language en_US --format srt --output output.srt
LocalScribe does not upload recognition audio, imported transcripts, or content processed by Gemma. Recognition, translation, and AI transcript enhancement run on the Mac. Network access is used for on-demand recognition, translation, and Gemma model downloads, automatic or manual update checks and downloads, and opening external documentation. Automatic updates can be disabled in Settings.
The built-in updater checks at launch and every six hours by default, reads update.json from the latest GitHub Release, downloads its ZIP, verifies SHA-256, checks the bundle identifier and version, and asks before replacing and relaunching the app. This update path is not a substitute for a notarized public release.
LocalScribe source code is available under the MIT License. Third-party components and models retain their original licenses; see THIRD_PARTY_NOTICES.md and the license files under Vendor.
The optional NLLB model is distributed upstream under CC-BY-NC-4.0.