FlowLocal is a Windows 11 x64 WPF dictation app. Hold the global shortcut, speak, and release to insert only your recognized speech into the active target. Recognition, speaker diarization, and voice matching run locally on CPU. There is no transcript cleanup or language-model rewrite; the inserted text is the raw Multitalker transcript.
%LOCALAPPDATA%\FlowLocal\Models (set FLOWLOCAL_MODELS_DIR to use another directory).Open System status → Your voice → Enroll my voice. Speak alone near the microphone for the 11-second recording; at least eight seconds of sufficiently audible speech are required. The 192-dimensional voice embedding is stored at %LOCALAPPDATA%\FlowLocal\user-embedding.json; it can be replaced by enrolling again. Audio for enrollment is not uploaded. Dictation refuses to start without enrollment.
During dictation, Sortformer v2.1 identifies up to four simultaneous speaker tracks, and Multitalker Parakeet Streaming INT8 transcribes tracks continuously via parakeet-rs. A local ECAPA-TDNN speaker embedding compares clean non-overlapping speech against the enrollment. Only a verified speaker channel can contribute text; ambiguous/overlapping speech is not used to establish identity. If no confident match occurs, nothing is inserted. This is speaker filtering, not a security-grade biometric authentication system; noisy, short, or changing voices can be missed or falsely matched.
From the repository root in PowerShell:
dotnet restore .\FlowLocal.slnx
dotnet build .\FlowLocal.slnx -c Release
dotnet test .\FlowLocal.slnx -c Release
dotnet run --project .\src\FlowLocal.App\FlowLocal.App.csproj -c Release
Builds require Rust/Cargo on PATH and compile the native worker. To package a self-contained release, install Inno Setup 6 and run ./pack.ps1 -Configuration Release -Version X.Y.Z. Use -PortableOnly to publish without Inno Setup. The app and native worker reside in the output together; first-run model acquisition happens in the worker, not in the installer.
Recordings and raw transcripts remain in %LOCALAPPDATA%\FlowLocal; History supports copying or pasting the raw text, retrying recognition from saved audio, and deleting entries. Older rows can still display prior cleaned text, but new dictations do not create it. The app never automatically presses Enter in a target. On uninstall, the local history, models, and voice embedding are removed.
The app checks GitHub Releases for signed update metadata and verifies the downloaded setup SHA-256 before launching it. First-run model downloads and update checks require network access; speech and transcripts do not leave the local worker.
For manual coverage and measured performance, see compatibility, performance, and known limitations. These are procedures, not completed benchmark claims.
FlowLocal is a Windows 11 x64 WPF dictation app. Hold the global shortcut, speak, and release to insert only your recognized speech into the active target. Recognition, speaker diarization, and voice matching run locally on CPU. There is no transcript cleanup or language-model rewrite; the inserted text is the raw Multitalker transcript.
%LOCALAPPDATA%\FlowLocal\Models (set FLOWLOCAL_MODELS_DIR to use another directory).Open System status → Your voice → Enroll my voice. Speak alone near the microphone for the 11-second recording; at least eight seconds of sufficiently audible speech are required. The 192-dimensional voice embedding is stored at %LOCALAPPDATA%\FlowLocal\user-embedding.json; it can be replaced by enrolling again. Audio for enrollment is not uploaded. Dictation refuses to start without enrollment.
During dictation, Sortformer v2.1 identifies up to four simultaneous speaker tracks, and Multitalker Parakeet Streaming INT8 transcribes tracks continuously via parakeet-rs. A local ECAPA-TDNN speaker embedding compares clean non-overlapping speech against the enrollment. Only a verified speaker channel can contribute text; ambiguous/overlapping speech is not used to establish identity. If no confident match occurs, nothing is inserted. This is speaker filtering, not a security-grade biometric authentication system; noisy, short, or changing voices can be missed or falsely matched.
From the repository root in PowerShell:
dotnet restore .\FlowLocal.slnx
dotnet build .\FlowLocal.slnx -c Release
dotnet test .\FlowLocal.slnx -c Release
dotnet run --project .\src\FlowLocal.App\FlowLocal.App.csproj -c Release
Builds require Rust/Cargo on PATH and compile the native worker. To package a self-contained release, install Inno Setup 6 and run ./pack.ps1 -Configuration Release -Version X.Y.Z. Use -PortableOnly to publish without Inno Setup. The app and native worker reside in the output together; first-run model acquisition happens in the worker, not in the installer.
Recordings and raw transcripts remain in %LOCALAPPDATA%\FlowLocal; History supports copying or pasting the raw text, retrying recognition from saved audio, and deleting entries. Older rows can still display prior cleaned text, but new dictations do not create it. The app never automatically presses Enter in a target. On uninstall, the local history, models, and voice embedding are removed.
The app checks GitHub Releases for signed update metadata and verifies the downloaded setup SHA-256 before launching it. First-run model downloads and update checks require network access; speech and transcripts do not leave the local worker.
For manual coverage and measured performance, see compatibility, performance, and known limitations. These are procedures, not completed benchmark claims.