Local-first, privacy-preserving voice-to-text for Windows, powered by on-device speech recognition.
See the codeSpeak freely. Type privately.
Parlotype is a local-by-default voice-to-text desktop application: on-device speech recognition is the default, and your voice never leaves your machine in local mode. Three local engines ship today — NVIDIA Parakeet TDT v3 via sherpa-onnx (the default: CPU-only and fastest), Whisper (widest language coverage, GPU-accelerated), and Gemma 4 (Google's multimodal model, run via a local llama.cpp sidecar). Two opt-in cloud engines ship alongside them for users whose hardware can't deliver the latency they need — bring-your-own-key, never selected automatically. See Provider Modes.
org.k2fsa.sherpa.onnx) — Parakeet TDT v3 speech recognition, in-process on CPUllama-server) — Gemma 4 speech recognition sidecarSystem.Security.Cryptography.ProtectedData) — encrypted storage for cloud API keysParlotype currently runs on Windows only. macOS and Linux support are planned for the future. Two features are explicitly Windows-only today: DPAPI encryption of cloud API keys (other platforms fall back to base64 with a logged warning) and keyboard-layout source detection (other platforms fall back to auto-detect).
GPU acceleration applies to the Whisper engine and runs on any Vulkan-capable GPU (AMD, Intel, NVIDIA), with automatic CPU fallback if no compatible GPU is detected. CUDA was dropped in ADR-049 — NVIDIA cards are still accelerated, via Vulkan. The active runtime can be changed under Settings → Speech engine → Whisper runtime. Parakeet is CPU-only by design (see ADR-041) and needs no GPU to be fast; Gemma 4 acceleration depends on the llama-server build you install.
Pre-built Windows binaries are published on the Releases page. Each
release ships one self-contained win-x64 build — no .NET runtime install required.
What changed in each one is in CHANGELOG.md.
Recommended: Parlotype-win-Setup.exe. It installs per-user into
%LOCALAPPDATA%\Parlotype with no administrator prompt, adds Start Menu and desktop
shortcuts, and can update itself in place (ADR-053).
Parlotype-win-Portable.zip is also published for anyone who wants to unzip and run
without installing. The portable build cannot update itself — you re-download it by
hand each time.
Being self-contained it is sizeable (~280 MB installed). Updates after the first install are incremental, so a typical update downloads a fraction of that. Parlotype falls back to CPU automatically if no Vulkan-capable GPU is found, and the default Parakeet engine needs no GPU at all. Builds are currently unsigned, so Windows SmartScreen may warn on first launch — choose More info → Run anyway.
Speech models are not bundled — they download on demand into
%LOCALAPPDATA%\parlotype-data\models\ the first time you need them, and survive both
updates and uninstalls.
Upgrading from a version before the installer? Older builds stored data in
%LOCALAPPDATA%\parlotype, which is now where the installer itself lives. Rename that folder to%LOCALAPPDATA%\parlotype-databefore installing to keep your settings and downloaded models; otherwise Parlotype starts fresh and re-downloads models on demand. See docs/RELEASING.md.
Parlotype checks for new releases automatically, shortly after launch and every six hours after that. New versions download in the background and install when you quit Parlotype — or immediately, via Install and restart now on that page.
The check is an anonymous HTTPS GET of
https://api.github.com/repos/mdemin729/parlotype/releases — the public list of releases
for this repository — followed by a download of the release file itself if a newer version
exists. No account, no machine identifier, no install id, no usage data, and no custom
headers are sent. Nothing about you or your machine is transmitted.
This is the only network request Parlotype makes while your speech engines are local. Turn it off at Settings → Application → Updates, which also shows when the feed was last reached and offers a manual Check now. With the toggle off, the updater makes no outbound requests at all.
Parlotype ships five interchangeable speech-to-text engines — three local, two cloud. Switch between them under Settings → Speech engine → Engine.
| Engine | Where it runs | Download | Languages | Translation |
|---|---|---|---|---|
| Parakeet TDT v3 (default) | Local, CPU only | ~670 MB (INT8) | 25 European, always auto-detected | — |
| Whisper | Local, Vulkan / CPU | 75 MiB – 2.9 GiB by model | ~99, selectable | To English |
| Gemma 4 | Local, llama-server sidecar | ~5.5–15 GiB by variant | Full catalog | To any language |
| OpenAI-compatible (cloud) | Your provider (OpenAI, Groq, …) | — | Auto-detected | — |
| xAI Grok (cloud) | xAI | — | Auto-detected | — |
Parlotype is local by default. Cloud by choice.
See ADR-032 for the positioning principles and brand commitments, and ADR-043 for the implementation.
NVIDIA's Parakeet TDT 0.6B v3 (FastConformer encoder + Token-and-Duration Transducer decoder, CC-BY-4.0), run in-process through the official sherpa-onnx .NET bindings — no sidecar process, no port management.
parakeet-tdt-0.6b-v3-int8 (~670 MB, the default) and a full-precision parakeet-tdt-0.6b-v3-fp32 entry (~2.6 GB, better on accented speech at ~2× decode time and ~3× RAM). Selectable under Settings → Speech engine → Parakeet model; downloaded from HuggingFace (csukuangfj/...) with a progress dialog on first use.OpenAI Whisper via Whisper.net. Widest language coverage and the only engine with GPU acceleration. Pick any GGML model from Tiny through Large v3 Turbo under Settings → Speech engine → Whisper model; models download on demand.
translate task. Only multilingual models can do this — the English-only models (*.en) and Large v3 Turbo cannot, so the toggle is disabled with an explanation and the pipeline gates the flag regardless of UI state (ADR-033). Your preference is preserved and resumes when you pick a capable model.Google's multimodal Gemma 4 model, run through a local llama-server process. Parlotype manages the sidecar for you: audio is sent over a loopback HTTP connection, so nothing leaves your machine.
mmproj), downloaded from HuggingFace (ggml-org/gemma-4-E2B-it-GGUF / gemma-4-E4B-it-GGUF) into %LOCALAPPDATA%\parlotype\models\. Sizes range from ~5.5 GiB to ~15 GiB depending on variant and quantization; the default is E4B (Q4_K_M) (~5.9 GiB).llama-server binary is installed and managed from the Settings UI (or you can point Parlotype at your own copy).{speech_lang} (the spoken language) and {text_lang} (the output language — the same as the spoken language whenever translation is off); the built-in default carries purpose-specific bodies for transcription, translation, and auto-detect. A collapsible How prompts work panel on the page explains both placeholders, the conditions that trigger translation, and how the built-in's three bodies differ from a custom single-body prompt. The legacy {language} token is retired.Both cloud engines are bring-your-own-key and configured under Settings → Speech engine → Cloud providers (visible only while a cloud engine is selected).
POST {base URL}/audio/transcriptions, the OpenAI transcription protocol. Defaults to https://api.openai.com/v1 with gpt-4o-mini-transcribe. Pointing the base URL at Groq, a self-hosted Whisper server, or any other compatible host is the supported way to use other providers.POST {base URL}/stt. Defaults to https://api.x.ai/v1 with grok-stt.Both always auto-detect the language (Parlotype sends no language-forcing parameter) and are transcribe-only, so the Language UI hides while they are active. Base URLs must be HTTPS unless the host is loopback, which keeps LM Studio / llama.cpp self-hosting working. Failures surface as dialogs rather than silent no-ops: a missing or rejected key offers an Open settings deep link; quota, rate-limit, and provider-outage errors show an informational message and recording keeps running.
The language surface is capability-driven — each engine declares what it supports and the UI renders only choices that actually take effect (ADR-036, ADR-042).
[source] → [target]. The connector is the master translation toggle; each side opens a searchable floating picker with recently-used languages pinned on top.Dictation has two genuinely different modes — holding a key for one sentence, and going hands-free — so Parlotype binds both out of the box rather than making you choose (ADR-047):
| Gesture | Mode | Why this one |
|---|---|---|
| Hold Right Ctrl | Push-to-talk | The same physical key on every platform, and comfortable to hold. Left and right Ctrl are distinct keys to the hook, so all your normal Ctrl shortcuts keep working. |
| Double-tap Ctrl | Toggle | macOS Dictation's own default on external keyboards, and it collides with nothing on Windows or Linux. |
| Ctrl+Alt+Space | Toggle | An explicit chord for anywhere bare-modifier detection isn't available. |
Esc cancels. While recording, Escape stops and discards the take — nothing is transcribed and nothing is typed. It passes straight through to whatever app you're in the rest of the time.
Manage all of this under Settings → Input → Hotkeys: add gestures from a preset menu or record a chord, remove them, and flip a chord between push-to-talk and toggle. Hold and double-tap gestures don't offer that choice — releasing a held key has to mean "stop", and a double-tap has no release to hang "stop" on.
New bindings are validated before they're accepted. Shortcuts the OS has already claimed (Win+L, Win+H, Win+Ctrl+S, …) and duplicates of your own bindings are refused; combinations that merely tend to collide are accepted with a note — Ctrl+Shift+Space shows parameter hints in Visual Studio and VS Code, and any Ctrl+Alt+<letter> can fire while typing accented characters, because AltGr is Ctrl+Alt on European layouts. (Parlotype ignores chord matches while right Alt is held for exactly that reason.)
Two behaviours worth knowing:
Upgrading? If you'd picked your own hotkey, it carries over as your only binding and the new defaults are not added on top — no surprise global shortcuts. Anyone still on the old Ctrl+Shift+Space default gets the new set, since that combination fights the IDEs most of this audience lives in.
The Transcribe window is a frameless always-on-top mini card (172 px wide, 88 px tall — 118 px when the language strip is shown), modelled on the Windows Voice Typing widget (ADR-040).
window-state.json, and self-heals to centre-screen if the saved spot is off-screen (monitor unplugged, resolution change).Findings from the 2026-07 security audit are addressed in ADR-046:
Information; the console keeps Debug for development.settings.json — in secrets.json, encrypted with DPAPI (CurrentUser scope) on Windows, so settings backups, sync, or diagnostics never carry credentials. Undecryptable values are treated as absent and re-prompted. On non-Windows platforms they are base64-encoded with a logged warning (OS keychain support is a known gap).settings.json and secrets.json are written atomically (temp file + move), and the llama-server sidecar is spawned with an argument list rather than a quoted command line.User data lives under %LOCALAPPDATA%\parlotype-data\, deliberately outside the
install directory. The installer owns %LOCALAPPDATA%\Parlotype and erases it on
uninstall, so nothing of yours is kept there — your models and settings survive updates
and uninstalls alike (ADR-053).
| Path | Contents |
|---|---|
settings.json | User-configured settings (microphone, engine, model, hotkeys, languages, theme) |
window-state.json | Transient window chrome state (Transcribe widget position) |
secrets.json | Cloud API keys, DPAPI-encrypted on Windows |
models\ | Downloaded Whisper / Parakeet / Gemma 4 model files |
logs\ | Rolling log files (Information and above; no transcript text), plus velopack.log for install/update events |
Uninstalling keeps this folder by default — it can hold several GB of models that are expensive to re-download, and plenty of uninstalls are really reinstalls. If you'd rather have a clean removal, turn on Delete everything when I uninstall Parlotype at Settings → Application → Data before uninstalling; uninstall then removes the whole folder. You can also delete it by hand at any time.
That page also shows the folder's exact location — with Copy path and Open folder buttons — how much space your downloaded models occupy, and a Delete downloaded models… button if you want the disk space back without giving up your settings and API keys.
Vulkan is not required unless you want GPU-accelerated Whisper — the default Parakeet engine runs on CPU.
Vulkan is the only GPU runtime for Whisper, on every vendor. Most modern GPU drivers (Radeon, Intel Arc, GeForce) already ship the Vulkan loader (vulkan-1.dll) — no extra install is needed for end users.
If Parlotype reports that the Vulkan loader is missing, install the Vulkan SDK from vulkan.lunarg.com/sdk/home (the SDK bundles a system-wide Vulkan loader and is also useful for development).
Parlotype will automatically detect and use Vulkan when available, in priority order Vulkan → CPU (Auto mode). You can pin a specific runtime under Settings → Speech engine → Whisper runtime; changes take effect after an app restart.
dotnet build Parlotype.slnx
dotnet run --project src\Parlotype.Desktop
The app starts minimized to the system tray. Click the tray icon for an Open / Settings / Exit menu, or use a dictation hotkey — hold Right Ctrl, or double-tap Ctrl — to open the Transcribe window and start recording.
The desktop app supports the official Avalonia 12 Developer Tools (Essentials edition — free under AvaloniaUI's community licence for organisations under €1M revenue). Setup is per-developer:
Install the standalone tool once:
dotnet tool install --global AvaloniaUI.DeveloperTools
Launch it in a separate window:
avdt
Run the app in Debug configuration, give a Parlotype window focus, and press F12. The inspector will connect and show the visual tree, properties, layout, and styles.
First-time activation requires a free AvaloniaUI Portal
account. The AvaloniaUI.DiagnosticsSupport package is referenced with a
Configuration == Debug condition, so Release builds carry no extra binaries.
See ADR 016.
dotnet test
Evaluate speech recognition quality with the built-in benchmark tool:
# Run a benchmark (Whisper smoke test)
dotnet run --project src\Parlotype.Benchmark -- run `
--config datasets\smoke-test-config.json `
--datasets datasets `
--output results
# Run the Parakeet smoke test (model auto-downloads headlessly)
dotnet run --project src\Parlotype.Benchmark -- run `
--config datasets\parakeet-smoke-config.json `
--datasets datasets `
--output results
# List historical benchmark runs
dotnet run --project src\Parlotype.Benchmark -- list --output results
# Compare two runs side by side
dotnet run --project src\Parlotype.Benchmark -- compare `
--run-a <run-id-a> --run-b <run-id-b> --output results
# Export a run as CSV, Markdown, or JSON
dotnet run --project src\Parlotype.Benchmark -- export `
--run-id <run-id> --format markdown --output results
# Rebuild SQLite index from existing JSON result files
dotnet run --project src\Parlotype.Benchmark -- import --output results
# Run a parameter sweep across configurations
dotnet run --project src\Parlotype.Benchmark -- sweep `
--config datasets\sweep-config.json `
--datasets datasets `
--output results
# Check for regressions against a baseline (for CI)
dotnet run --project src\Parlotype.Benchmark -- check `
--baseline <run-id> --current latest `
--output results --max-wer-delta 2.0
The benchmark computes WER (Word Error Rate), CER (Character Error Rate), and RTF (Real-Time Factor) against WAV/FLAC datasets with ground-truth transcriptions. Results are saved as JSON and auto-indexed into SQLite for historical queries. Supports tag/sample filtering (--tags, --samples), side-by-side comparison with delta metrics, and export to CSV, Markdown, or JSON.
Parlotype.Benchmark measures transcription quality. Allocation and hot-path latency
questions ("how many bytes does one WAV encode allocate?") are answered by a separate
BenchmarkDotNet project (ADR-044). It is
not a test project, so dotnet test is unaffected — run it manually in Release:
dotnet run -c Release --project src\Parlotype.MicroBenchmarks -- --filter *
src/
├── Parlotype.Core/ # Domain interfaces and models (zero external deps)
├── Parlotype.Platform/ # Platform implementations (Whisper, sherpa-onnx, NAudio, SharpHook, DPAPI)
├── Parlotype.Desktop/ # Avalonia 12 desktop app (tray-based, entry point)
├── Parlotype.Desktop.Tests/ # Avalonia headless UI tests (xUnit v3)
├── Parlotype.Benchmark/ # CLI quality benchmark (WER/CER/RTF, sweep, compare, CI check)
├── Parlotype.Benchmark.Tests/ # Benchmark unit tests
├── Parlotype.MicroBenchmarks/ # BenchmarkDotNet allocation/latency suites (not a test project)
└── Parlotype.Tests/ # Core + Platform unit tests (xUnit)
datasets/ # Benchmark configs + sample datasets with ground truth
docs/
├── decisions/ # Architecture Decision Records
├── architecture/ # Architecture notes
├── research/ # Engine and provider research
└── security/ # Security audits
ISpeechRecognizer, IAudioCaptureService, ISettingsService, IWindowStateService, ISecretStore, etc.)SpeechRecognizerFactory resolves one of five ISpeechRecognizer implementations (Parakeet, Whisper, Gemma 4, OpenAI-compatible, xAI Grok) from the persisted SpeechEngine setting. Cloud engines needed zero changes to the audio pipeline or the ISpeechRecognizer contract.LanguageCapabilities, and both the Settings page and the Transcribe widget render from one shared LanguageRelationshipViewModel — a new engine gets correct language UI by declaring capabilities, not by touching viewsArrayPool<float> (ADR-045)WhisperOptions for benchmarkingdocs/decisions/ — Gemma 4 spans ADRs 024–030 and 037, the language UX 034–036 and 042, and the cloud providers 032 and 043This project is licensed under the MIT License.
227 followers · starred Aug 2026
C#
96.1%
PowerShell
2.0%
HTML
1.3%
Local-first, privacy-preserving voice-to-text for Windows, powered by on-device speech recognition.
See the codeSpeak freely. Type privately.
Parlotype is a local-by-default voice-to-text desktop application: on-device speech recognition is the default, and your voice never leaves your machine in local mode. Three local engines ship today — NVIDIA Parakeet TDT v3 via sherpa-onnx (the default: CPU-only and fastest), Whisper (widest language coverage, GPU-accelerated), and Gemma 4 (Google's multimodal model, run via a local llama.cpp sidecar). Two opt-in cloud engines ship alongside them for users whose hardware can't deliver the latency they need — bring-your-own-key, never selected automatically. See Provider Modes.
org.k2fsa.sherpa.onnx) — Parakeet TDT v3 speech recognition, in-process on CPUllama-server) — Gemma 4 speech recognition sidecarSystem.Security.Cryptography.ProtectedData) — encrypted storage for cloud API keysParlotype currently runs on Windows only. macOS and Linux support are planned for the future. Two features are explicitly Windows-only today: DPAPI encryption of cloud API keys (other platforms fall back to base64 with a logged warning) and keyboard-layout source detection (other platforms fall back to auto-detect).
GPU acceleration applies to the Whisper engine and runs on any Vulkan-capable GPU (AMD, Intel, NVIDIA), with automatic CPU fallback if no compatible GPU is detected. CUDA was dropped in ADR-049 — NVIDIA cards are still accelerated, via Vulkan. The active runtime can be changed under Settings → Speech engine → Whisper runtime. Parakeet is CPU-only by design (see ADR-041) and needs no GPU to be fast; Gemma 4 acceleration depends on the llama-server build you install.
Pre-built Windows binaries are published on the Releases page. Each
release ships one self-contained win-x64 build — no .NET runtime install required.
What changed in each one is in CHANGELOG.md.
Recommended: Parlotype-win-Setup.exe. It installs per-user into
%LOCALAPPDATA%\Parlotype with no administrator prompt, adds Start Menu and desktop
shortcuts, and can update itself in place (ADR-053).
Parlotype-win-Portable.zip is also published for anyone who wants to unzip and run
without installing. The portable build cannot update itself — you re-download it by
hand each time.
Being self-contained it is sizeable (~280 MB installed). Updates after the first install are incremental, so a typical update downloads a fraction of that. Parlotype falls back to CPU automatically if no Vulkan-capable GPU is found, and the default Parakeet engine needs no GPU at all. Builds are currently unsigned, so Windows SmartScreen may warn on first launch — choose More info → Run anyway.
Speech models are not bundled — they download on demand into
%LOCALAPPDATA%\parlotype-data\models\ the first time you need them, and survive both
updates and uninstalls.
Upgrading from a version before the installer? Older builds stored data in
%LOCALAPPDATA%\parlotype, which is now where the installer itself lives. Rename that folder to%LOCALAPPDATA%\parlotype-databefore installing to keep your settings and downloaded models; otherwise Parlotype starts fresh and re-downloads models on demand. See docs/RELEASING.md.
Parlotype checks for new releases automatically, shortly after launch and every six hours after that. New versions download in the background and install when you quit Parlotype — or immediately, via Install and restart now on that page.
The check is an anonymous HTTPS GET of
https://api.github.com/repos/mdemin729/parlotype/releases — the public list of releases
for this repository — followed by a download of the release file itself if a newer version
exists. No account, no machine identifier, no install id, no usage data, and no custom
headers are sent. Nothing about you or your machine is transmitted.
This is the only network request Parlotype makes while your speech engines are local. Turn it off at Settings → Application → Updates, which also shows when the feed was last reached and offers a manual Check now. With the toggle off, the updater makes no outbound requests at all.
Parlotype ships five interchangeable speech-to-text engines — three local, two cloud. Switch between them under Settings → Speech engine → Engine.
| Engine | Where it runs | Download | Languages | Translation |
|---|---|---|---|---|
| Parakeet TDT v3 (default) | Local, CPU only | ~670 MB (INT8) | 25 European, always auto-detected | — |
| Whisper | Local, Vulkan / CPU | 75 MiB – 2.9 GiB by model | ~99, selectable | To English |
| Gemma 4 | Local, llama-server sidecar | ~5.5–15 GiB by variant | Full catalog | To any language |
| OpenAI-compatible (cloud) | Your provider (OpenAI, Groq, …) | — | Auto-detected | — |
| xAI Grok (cloud) | xAI | — | Auto-detected | — |
Parlotype is local by default. Cloud by choice.
See ADR-032 for the positioning principles and brand commitments, and ADR-043 for the implementation.
NVIDIA's Parakeet TDT 0.6B v3 (FastConformer encoder + Token-and-Duration Transducer decoder, CC-BY-4.0), run in-process through the official sherpa-onnx .NET bindings — no sidecar process, no port management.
parakeet-tdt-0.6b-v3-int8 (~670 MB, the default) and a full-precision parakeet-tdt-0.6b-v3-fp32 entry (~2.6 GB, better on accented speech at ~2× decode time and ~3× RAM). Selectable under Settings → Speech engine → Parakeet model; downloaded from HuggingFace (csukuangfj/...) with a progress dialog on first use.OpenAI Whisper via Whisper.net. Widest language coverage and the only engine with GPU acceleration. Pick any GGML model from Tiny through Large v3 Turbo under Settings → Speech engine → Whisper model; models download on demand.
translate task. Only multilingual models can do this — the English-only models (*.en) and Large v3 Turbo cannot, so the toggle is disabled with an explanation and the pipeline gates the flag regardless of UI state (ADR-033). Your preference is preserved and resumes when you pick a capable model.Google's multimodal Gemma 4 model, run through a local llama-server process. Parlotype manages the sidecar for you: audio is sent over a loopback HTTP connection, so nothing leaves your machine.
mmproj), downloaded from HuggingFace (ggml-org/gemma-4-E2B-it-GGUF / gemma-4-E4B-it-GGUF) into %LOCALAPPDATA%\parlotype\models\. Sizes range from ~5.5 GiB to ~15 GiB depending on variant and quantization; the default is E4B (Q4_K_M) (~5.9 GiB).llama-server binary is installed and managed from the Settings UI (or you can point Parlotype at your own copy).{speech_lang} (the spoken language) and {text_lang} (the output language — the same as the spoken language whenever translation is off); the built-in default carries purpose-specific bodies for transcription, translation, and auto-detect. A collapsible How prompts work panel on the page explains both placeholders, the conditions that trigger translation, and how the built-in's three bodies differ from a custom single-body prompt. The legacy {language} token is retired.Both cloud engines are bring-your-own-key and configured under Settings → Speech engine → Cloud providers (visible only while a cloud engine is selected).
POST {base URL}/audio/transcriptions, the OpenAI transcription protocol. Defaults to https://api.openai.com/v1 with gpt-4o-mini-transcribe. Pointing the base URL at Groq, a self-hosted Whisper server, or any other compatible host is the supported way to use other providers.POST {base URL}/stt. Defaults to https://api.x.ai/v1 with grok-stt.Both always auto-detect the language (Parlotype sends no language-forcing parameter) and are transcribe-only, so the Language UI hides while they are active. Base URLs must be HTTPS unless the host is loopback, which keeps LM Studio / llama.cpp self-hosting working. Failures surface as dialogs rather than silent no-ops: a missing or rejected key offers an Open settings deep link; quota, rate-limit, and provider-outage errors show an informational message and recording keeps running.
The language surface is capability-driven — each engine declares what it supports and the UI renders only choices that actually take effect (ADR-036, ADR-042).
[source] → [target]. The connector is the master translation toggle; each side opens a searchable floating picker with recently-used languages pinned on top.Dictation has two genuinely different modes — holding a key for one sentence, and going hands-free — so Parlotype binds both out of the box rather than making you choose (ADR-047):
| Gesture | Mode | Why this one |
|---|---|---|
| Hold Right Ctrl | Push-to-talk | The same physical key on every platform, and comfortable to hold. Left and right Ctrl are distinct keys to the hook, so all your normal Ctrl shortcuts keep working. |
| Double-tap Ctrl | Toggle | macOS Dictation's own default on external keyboards, and it collides with nothing on Windows or Linux. |
| Ctrl+Alt+Space | Toggle | An explicit chord for anywhere bare-modifier detection isn't available. |
Esc cancels. While recording, Escape stops and discards the take — nothing is transcribed and nothing is typed. It passes straight through to whatever app you're in the rest of the time.
Manage all of this under Settings → Input → Hotkeys: add gestures from a preset menu or record a chord, remove them, and flip a chord between push-to-talk and toggle. Hold and double-tap gestures don't offer that choice — releasing a held key has to mean "stop", and a double-tap has no release to hang "stop" on.
New bindings are validated before they're accepted. Shortcuts the OS has already claimed (Win+L, Win+H, Win+Ctrl+S, …) and duplicates of your own bindings are refused; combinations that merely tend to collide are accepted with a note — Ctrl+Shift+Space shows parameter hints in Visual Studio and VS Code, and any Ctrl+Alt+<letter> can fire while typing accented characters, because AltGr is Ctrl+Alt on European layouts. (Parlotype ignores chord matches while right Alt is held for exactly that reason.)
Two behaviours worth knowing:
Upgrading? If you'd picked your own hotkey, it carries over as your only binding and the new defaults are not added on top — no surprise global shortcuts. Anyone still on the old Ctrl+Shift+Space default gets the new set, since that combination fights the IDEs most of this audience lives in.
The Transcribe window is a frameless always-on-top mini card (172 px wide, 88 px tall — 118 px when the language strip is shown), modelled on the Windows Voice Typing widget (ADR-040).
window-state.json, and self-heals to centre-screen if the saved spot is off-screen (monitor unplugged, resolution change).Findings from the 2026-07 security audit are addressed in ADR-046:
Information; the console keeps Debug for development.settings.json — in secrets.json, encrypted with DPAPI (CurrentUser scope) on Windows, so settings backups, sync, or diagnostics never carry credentials. Undecryptable values are treated as absent and re-prompted. On non-Windows platforms they are base64-encoded with a logged warning (OS keychain support is a known gap).settings.json and secrets.json are written atomically (temp file + move), and the llama-server sidecar is spawned with an argument list rather than a quoted command line.User data lives under %LOCALAPPDATA%\parlotype-data\, deliberately outside the
install directory. The installer owns %LOCALAPPDATA%\Parlotype and erases it on
uninstall, so nothing of yours is kept there — your models and settings survive updates
and uninstalls alike (ADR-053).
| Path | Contents |
|---|---|
settings.json | User-configured settings (microphone, engine, model, hotkeys, languages, theme) |
window-state.json | Transient window chrome state (Transcribe widget position) |
secrets.json | Cloud API keys, DPAPI-encrypted on Windows |
models\ | Downloaded Whisper / Parakeet / Gemma 4 model files |
logs\ | Rolling log files (Information and above; no transcript text), plus velopack.log for install/update events |
Uninstalling keeps this folder by default — it can hold several GB of models that are expensive to re-download, and plenty of uninstalls are really reinstalls. If you'd rather have a clean removal, turn on Delete everything when I uninstall Parlotype at Settings → Application → Data before uninstalling; uninstall then removes the whole folder. You can also delete it by hand at any time.
That page also shows the folder's exact location — with Copy path and Open folder buttons — how much space your downloaded models occupy, and a Delete downloaded models… button if you want the disk space back without giving up your settings and API keys.
Vulkan is not required unless you want GPU-accelerated Whisper — the default Parakeet engine runs on CPU.
Vulkan is the only GPU runtime for Whisper, on every vendor. Most modern GPU drivers (Radeon, Intel Arc, GeForce) already ship the Vulkan loader (vulkan-1.dll) — no extra install is needed for end users.
If Parlotype reports that the Vulkan loader is missing, install the Vulkan SDK from vulkan.lunarg.com/sdk/home (the SDK bundles a system-wide Vulkan loader and is also useful for development).
Parlotype will automatically detect and use Vulkan when available, in priority order Vulkan → CPU (Auto mode). You can pin a specific runtime under Settings → Speech engine → Whisper runtime; changes take effect after an app restart.
dotnet build Parlotype.slnx
dotnet run --project src\Parlotype.Desktop
The app starts minimized to the system tray. Click the tray icon for an Open / Settings / Exit menu, or use a dictation hotkey — hold Right Ctrl, or double-tap Ctrl — to open the Transcribe window and start recording.
The desktop app supports the official Avalonia 12 Developer Tools (Essentials edition — free under AvaloniaUI's community licence for organisations under €1M revenue). Setup is per-developer:
Install the standalone tool once:
dotnet tool install --global AvaloniaUI.DeveloperTools
Launch it in a separate window:
avdt
Run the app in Debug configuration, give a Parlotype window focus, and press F12. The inspector will connect and show the visual tree, properties, layout, and styles.
First-time activation requires a free AvaloniaUI Portal
account. The AvaloniaUI.DiagnosticsSupport package is referenced with a
Configuration == Debug condition, so Release builds carry no extra binaries.
See ADR 016.
dotnet test
Evaluate speech recognition quality with the built-in benchmark tool:
# Run a benchmark (Whisper smoke test)
dotnet run --project src\Parlotype.Benchmark -- run `
--config datasets\smoke-test-config.json `
--datasets datasets `
--output results
# Run the Parakeet smoke test (model auto-downloads headlessly)
dotnet run --project src\Parlotype.Benchmark -- run `
--config datasets\parakeet-smoke-config.json `
--datasets datasets `
--output results
# List historical benchmark runs
dotnet run --project src\Parlotype.Benchmark -- list --output results
# Compare two runs side by side
dotnet run --project src\Parlotype.Benchmark -- compare `
--run-a <run-id-a> --run-b <run-id-b> --output results
# Export a run as CSV, Markdown, or JSON
dotnet run --project src\Parlotype.Benchmark -- export `
--run-id <run-id> --format markdown --output results
# Rebuild SQLite index from existing JSON result files
dotnet run --project src\Parlotype.Benchmark -- import --output results
# Run a parameter sweep across configurations
dotnet run --project src\Parlotype.Benchmark -- sweep `
--config datasets\sweep-config.json `
--datasets datasets `
--output results
# Check for regressions against a baseline (for CI)
dotnet run --project src\Parlotype.Benchmark -- check `
--baseline <run-id> --current latest `
--output results --max-wer-delta 2.0
The benchmark computes WER (Word Error Rate), CER (Character Error Rate), and RTF (Real-Time Factor) against WAV/FLAC datasets with ground-truth transcriptions. Results are saved as JSON and auto-indexed into SQLite for historical queries. Supports tag/sample filtering (--tags, --samples), side-by-side comparison with delta metrics, and export to CSV, Markdown, or JSON.
Parlotype.Benchmark measures transcription quality. Allocation and hot-path latency
questions ("how many bytes does one WAV encode allocate?") are answered by a separate
BenchmarkDotNet project (ADR-044). It is
not a test project, so dotnet test is unaffected — run it manually in Release:
dotnet run -c Release --project src\Parlotype.MicroBenchmarks -- --filter *
src/
├── Parlotype.Core/ # Domain interfaces and models (zero external deps)
├── Parlotype.Platform/ # Platform implementations (Whisper, sherpa-onnx, NAudio, SharpHook, DPAPI)
├── Parlotype.Desktop/ # Avalonia 12 desktop app (tray-based, entry point)
├── Parlotype.Desktop.Tests/ # Avalonia headless UI tests (xUnit v3)
├── Parlotype.Benchmark/ # CLI quality benchmark (WER/CER/RTF, sweep, compare, CI check)
├── Parlotype.Benchmark.Tests/ # Benchmark unit tests
├── Parlotype.MicroBenchmarks/ # BenchmarkDotNet allocation/latency suites (not a test project)
└── Parlotype.Tests/ # Core + Platform unit tests (xUnit)
datasets/ # Benchmark configs + sample datasets with ground truth
docs/
├── decisions/ # Architecture Decision Records
├── architecture/ # Architecture notes
├── research/ # Engine and provider research
└── security/ # Security audits
ISpeechRecognizer, IAudioCaptureService, ISettingsService, IWindowStateService, ISecretStore, etc.)SpeechRecognizerFactory resolves one of five ISpeechRecognizer implementations (Parakeet, Whisper, Gemma 4, OpenAI-compatible, xAI Grok) from the persisted SpeechEngine setting. Cloud engines needed zero changes to the audio pipeline or the ISpeechRecognizer contract.LanguageCapabilities, and both the Settings page and the Transcribe widget render from one shared LanguageRelationshipViewModel — a new engine gets correct language UI by declaring capabilities, not by touching viewsArrayPool<float> (ADR-045)WhisperOptions for benchmarkingdocs/decisions/ — Gemma 4 spans ADRs 024–030 and 037, the language UX 034–036 and 042, and the cloud providers 032 and 043This project is licensed under the MIT License.
227 followers · starred Aug 2026
C#
96.1%
PowerShell
2.0%
HTML
1.3%