Install: Give Copilot this one-line request:
Install gujiguji on this Windows PC from https://github.com/fafa-npu/VoiceInput. Preserve my existing settings and downloaded models, use the repository's supported installation flow, resolve PowerShell execution-policy and startup-shortcut permission issues, verify any downloaded release against its GitHub SHA-256 digest (or build from source), launch the app, and ask me only when Windows requires a security confirmation.
A native tray/menu-bar voice input method. Hold a key while speaking, or press once to start and again to stop. gujiguji transcribes the recording and inserts it only if the original window and input control still have focus.
The existing Windows client remains C# / .NET 10 + WPF. The macOS client is a separate native Swift + AppKit target with the same profiles, settings, onboarding, overlay, engines, recovery, and privacy behavior.
The recommended installation is to give Copilot the first-line request above. Copilot can inspect the repository, choose the supported release or source-build path, handle PowerShell and shortcut permissions, preserve an existing installation, and stop only for Windows security confirmations.
Manual one-line install in PowerShell:
$s="$env:TEMP\vi.ps1"; iwr https://github.com/fafa-npu/VoiceInput/raw/main/scripts/install.ps1 -OutFile $s; powershell -ExecutionPolicy Bypass -File $s
The installer verifies the official GitHub Release asset against the SHA-256 digest returned by
GitHub before installing it to %LOCALAPPDATA%\VoiceInput, creating Start Menu and auto-start
shortcuts, and launching it. Reinstalling preserves %APPDATA%\VoiceInput settings and downloaded
models under %LOCALAPPDATA%\VoiceInput\FunASR. Current release executables are not Authenticode
signed, so Windows may still show a security confirmation.
You can also download VoiceInput.exe from
Releases and run it directly. It is self-contained;
the .NET runtime is not required.
On first launch, the Windows guide offers to download Qwen3-ASR 0.6B int8 (about 987 MB). The download is resumable and does not require the separate FunASR runtime or VAD. Windows dictation remains available as a lower-accuracy fallback.
Uninstall:
powershell -File "$env:LOCALAPPDATA\VoiceInput\uninstall.ps1" -Uninstall
This also removes %APPDATA%\VoiceInput (settings, logs, and encrypted correction samples). Add
-KeepUserData to retain it.
The first macOS build targets macOS 15+ on Apple Silicon. It uses the official universal Azure Speech SDK; the pinned local FunASR runtime is currently arm64-only.
swift test --package-path src/VoiceInputMac
scripts/build-macos.sh
open dist/gujiguji.app
On first launch, allow Microphone, Accessibility, and Input Monitoring. The two-page guide then
downloads the resumable, SHA-256-verified Qwen3-ASR 0.6B model (about 850 MB) or lets you
explicitly choose macOS Speech. Settings and models live under
~/Library/Application Support/gujiguji; secrets are stored in Keychain and logs under
~/Library/Logs/gujiguji.
For a signed release, configure a Developer ID certificate and notarytool profile, then run:
SIGN_IDENTITY="Developer ID Application: …" NOTARY_PROFILE=gujiguji scripts/release-macos.sh
gujiguji defaults new Windows and macOS users to Qwen3-ASR 0.6B. An optional speech-aware LLM refinement layer can be configured for:
瑞法克特 → query-inspector-refactor).Alt+Shift+G switches profiles
while idle, and the tray menu provides direct selection. The last active profile is restored.| Action | How |
|---|---|
| Talk | Use the activation key and behavior configured for the active Profile |
| Switch profile | Alt+Shift+G (Windows) / Option+Shift+G (macOS), or use the tray/menu bar |
| Start | Start Menu → gujiguji, or it auto-starts at login |
| Quit | Tray icon → Quit |
| Pause / resume | Tray → Pause / Resume listening |
| Context-aware refine | Settings → Language intelligence (off by default) |
| Setup | Tray → Settings… |
On first launch, Windows and macOS select Qwen3-ASR 0.6B. After setup, open Settings → Model Selection to install, switch, or remove local models on demand. Qwen3-ASR 1.7B is optional on both platforms and is never downloaded automatically.
| Model | Platform and download | Languages | Intended use |
|---|---|---|---|
| SenseVoiceSmall q8 | Windows/macOS · 254 MB | English, Chinese, Japanese, Korean | Balanced CPU model |
| Paraformer q8 | Windows/macOS · 237 MB | English, Chinese | Faster Chinese/English dictation |
| Fun-ASR Nano q4 | Windows/macOS · 954 MB | English, Chinese, Japanese | Difficult vocabulary and accents |
| Qwen3-ASR 0.6B | Windows int8 ONNX · 987 MB; macOS Q8 GGUF · 850 MB | English, Chinese, Japanese, Korean, Vietnamese | Default on both platforms; higher-quality multilingual recognition and automatic language detection |
| Qwen3-ASR 1.7B | Windows int8 ONNX · 2.40 GB; macOS Q5_K_M GGUF · 1.52 GB | English, Chinese, Japanese, Korean, Vietnamese | Optional accuracy-first model; larger memory footprint and slower CPU inference |
The three FunASR models share the official FunASR llama.cpp runtime (Windows x64 or macOS arm64) and an FSMN-VAD model. Qwen3-ASR uses a separate backend and does not download that runtime or VAD. All model downloads are resumable and SHA-256 verified before activation, and models can be removed from Setup without touching other app settings.
FunASR runs as a hidden native child process; its temporary local WAV is deleted after each batch. Qwen3-ASR stays loaded in-process between dictations: Windows uses sherpa-onnx 1.13.4 on CPU and macOS uses transcribe.cpp 0.1.3 with Metal/CPU. Both Qwen backends use automatic language detection. Windows accepts a bounded vocabulary prompt (first 10 terms, up to 96 characters) and limits each Qwen dictation to 25 seconds so the fixed ONNX context cannot silently truncate it. The macOS transcribe.cpp backend does not currently expose native hotword prompting. The 1.7B model is best suited to Apple silicon or a Windows machine with ample RAM; the 0.6B default is substantially more practical in CPU-only virtual machines.
No local backend requires Python, PyTorch, Docker, a local HTTP server, or a listening port. The three FunASR GGUF models do not support Vietnamese; Qwen3-ASR does. Model weights and runtimes use the licenses linked from their source cards; the pinned Qwen and sherpa artifacts are Apache 2.0.
Tray → Settings… opens the Setup Hub. Model Selection contains speech engines, cloud authentication, and local-model downloads; Profiles configures the two activation and overlay presets; Language intelligence combines recognition terms, the OpenAI-compatible model connection, live refinement, and reviewed learning from locally encrypted corrections; App contains language, privacy, startup, update, and logging controls. Saving corrections makes no network request. Review learning sends them only to the configured language-model endpoint and never clears the history automatically. Secret fields are DPAPI-encrypted per Windows user or stored in macOS Keychain.
On Windows, Microsoft Entra authentication uses Windows Authentication Manager (WAM), the signed-in
work/school account, and an OS-protected token cache. A tenant Conditional Access policy can still
require MFA or another explicit challenge, but gujiguji serializes that interaction. Foundry batch
transcription begins the check as soon as PCM capture is live and retains that recording while the
single WAM dialog completes. Use Switch Azure account… in Model Selection only when a different
work/school account should access the configured resource. The WAM integration is pinned to
Azure.Identity.Broker 1.3.1 (NuGet package: 73,665 bytes; SHA-256
2348dc1830c3966904e89fd0c952661cac13d5eb7970e284743ab42112469cd4).
On macOS, Azure Speech with Microsoft Entra ID also requires the Speech resource's Region and
full Azure Resource ID (for example /subscriptions/.../resourceGroups/.../providers/Microsoft.CognitiveServices/accounts/...).
The native Speech SDK uses these to construct Microsoft's required aad#resource-id#token
authorization token; the custom-domain endpoint and tenant remain configured as on Windows.
Windows needs the .NET 10 SDK; macOS needs Xcode/Swift 6.
make run # run from source
make install # build + install to %LOCALAPPDATA% + auto-start + launch
make release VERSION=vX.Y.Z SIGN_PFX=publisher.pfx # prompts securely for the PFX password
dotnet test tests/VoiceInput.Tests/VoiceInput.Tests.csproj
dotnet build src/VoiceInput/VoiceInput.csproj -p:EnableWindowsTargeting=true
swift test --package-path src/VoiceInputMac
scripts/build-macos.sh
The app shows its version in Settings and offers Update to vX.Y.Z… when a newer release exists. Updates are user-initiated, verified against a pinned Authenticode publisher when signed or the GitHub Release SHA-256 digest otherwise, atomically replaced, and rolled back if the new process does not stay running.
gpt-4o-transcribe deployment (e.g. in eastus2 / swedencentral).C#
57.2%
Swift
41.0%
PowerShell
1.1%
Install: Give Copilot this one-line request:
Install gujiguji on this Windows PC from https://github.com/fafa-npu/VoiceInput. Preserve my existing settings and downloaded models, use the repository's supported installation flow, resolve PowerShell execution-policy and startup-shortcut permission issues, verify any downloaded release against its GitHub SHA-256 digest (or build from source), launch the app, and ask me only when Windows requires a security confirmation.
A native tray/menu-bar voice input method. Hold a key while speaking, or press once to start and again to stop. gujiguji transcribes the recording and inserts it only if the original window and input control still have focus.
The existing Windows client remains C# / .NET 10 + WPF. The macOS client is a separate native Swift + AppKit target with the same profiles, settings, onboarding, overlay, engines, recovery, and privacy behavior.
The recommended installation is to give Copilot the first-line request above. Copilot can inspect the repository, choose the supported release or source-build path, handle PowerShell and shortcut permissions, preserve an existing installation, and stop only for Windows security confirmations.
Manual one-line install in PowerShell:
$s="$env:TEMP\vi.ps1"; iwr https://github.com/fafa-npu/VoiceInput/raw/main/scripts/install.ps1 -OutFile $s; powershell -ExecutionPolicy Bypass -File $s
The installer verifies the official GitHub Release asset against the SHA-256 digest returned by
GitHub before installing it to %LOCALAPPDATA%\VoiceInput, creating Start Menu and auto-start
shortcuts, and launching it. Reinstalling preserves %APPDATA%\VoiceInput settings and downloaded
models under %LOCALAPPDATA%\VoiceInput\FunASR. Current release executables are not Authenticode
signed, so Windows may still show a security confirmation.
You can also download VoiceInput.exe from
Releases and run it directly. It is self-contained;
the .NET runtime is not required.
On first launch, the Windows guide offers to download Qwen3-ASR 0.6B int8 (about 987 MB). The download is resumable and does not require the separate FunASR runtime or VAD. Windows dictation remains available as a lower-accuracy fallback.
Uninstall:
powershell -File "$env:LOCALAPPDATA\VoiceInput\uninstall.ps1" -Uninstall
This also removes %APPDATA%\VoiceInput (settings, logs, and encrypted correction samples). Add
-KeepUserData to retain it.
The first macOS build targets macOS 15+ on Apple Silicon. It uses the official universal Azure Speech SDK; the pinned local FunASR runtime is currently arm64-only.
swift test --package-path src/VoiceInputMac
scripts/build-macos.sh
open dist/gujiguji.app
On first launch, allow Microphone, Accessibility, and Input Monitoring. The two-page guide then
downloads the resumable, SHA-256-verified Qwen3-ASR 0.6B model (about 850 MB) or lets you
explicitly choose macOS Speech. Settings and models live under
~/Library/Application Support/gujiguji; secrets are stored in Keychain and logs under
~/Library/Logs/gujiguji.
For a signed release, configure a Developer ID certificate and notarytool profile, then run:
SIGN_IDENTITY="Developer ID Application: …" NOTARY_PROFILE=gujiguji scripts/release-macos.sh
gujiguji defaults new Windows and macOS users to Qwen3-ASR 0.6B. An optional speech-aware LLM refinement layer can be configured for:
瑞法克特 → query-inspector-refactor).Alt+Shift+G switches profiles
while idle, and the tray menu provides direct selection. The last active profile is restored.| Action | How |
|---|---|
| Talk | Use the activation key and behavior configured for the active Profile |
| Switch profile | Alt+Shift+G (Windows) / Option+Shift+G (macOS), or use the tray/menu bar |
| Start | Start Menu → gujiguji, or it auto-starts at login |
| Quit | Tray icon → Quit |
| Pause / resume | Tray → Pause / Resume listening |
| Context-aware refine | Settings → Language intelligence (off by default) |
| Setup | Tray → Settings… |
On first launch, Windows and macOS select Qwen3-ASR 0.6B. After setup, open Settings → Model Selection to install, switch, or remove local models on demand. Qwen3-ASR 1.7B is optional on both platforms and is never downloaded automatically.
| Model | Platform and download | Languages | Intended use |
|---|---|---|---|
| SenseVoiceSmall q8 | Windows/macOS · 254 MB | English, Chinese, Japanese, Korean | Balanced CPU model |
| Paraformer q8 | Windows/macOS · 237 MB | English, Chinese | Faster Chinese/English dictation |
| Fun-ASR Nano q4 | Windows/macOS · 954 MB | English, Chinese, Japanese | Difficult vocabulary and accents |
| Qwen3-ASR 0.6B | Windows int8 ONNX · 987 MB; macOS Q8 GGUF · 850 MB | English, Chinese, Japanese, Korean, Vietnamese | Default on both platforms; higher-quality multilingual recognition and automatic language detection |
| Qwen3-ASR 1.7B | Windows int8 ONNX · 2.40 GB; macOS Q5_K_M GGUF · 1.52 GB | English, Chinese, Japanese, Korean, Vietnamese | Optional accuracy-first model; larger memory footprint and slower CPU inference |
The three FunASR models share the official FunASR llama.cpp runtime (Windows x64 or macOS arm64) and an FSMN-VAD model. Qwen3-ASR uses a separate backend and does not download that runtime or VAD. All model downloads are resumable and SHA-256 verified before activation, and models can be removed from Setup without touching other app settings.
FunASR runs as a hidden native child process; its temporary local WAV is deleted after each batch. Qwen3-ASR stays loaded in-process between dictations: Windows uses sherpa-onnx 1.13.4 on CPU and macOS uses transcribe.cpp 0.1.3 with Metal/CPU. Both Qwen backends use automatic language detection. Windows accepts a bounded vocabulary prompt (first 10 terms, up to 96 characters) and limits each Qwen dictation to 25 seconds so the fixed ONNX context cannot silently truncate it. The macOS transcribe.cpp backend does not currently expose native hotword prompting. The 1.7B model is best suited to Apple silicon or a Windows machine with ample RAM; the 0.6B default is substantially more practical in CPU-only virtual machines.
No local backend requires Python, PyTorch, Docker, a local HTTP server, or a listening port. The three FunASR GGUF models do not support Vietnamese; Qwen3-ASR does. Model weights and runtimes use the licenses linked from their source cards; the pinned Qwen and sherpa artifacts are Apache 2.0.
Tray → Settings… opens the Setup Hub. Model Selection contains speech engines, cloud authentication, and local-model downloads; Profiles configures the two activation and overlay presets; Language intelligence combines recognition terms, the OpenAI-compatible model connection, live refinement, and reviewed learning from locally encrypted corrections; App contains language, privacy, startup, update, and logging controls. Saving corrections makes no network request. Review learning sends them only to the configured language-model endpoint and never clears the history automatically. Secret fields are DPAPI-encrypted per Windows user or stored in macOS Keychain.
On Windows, Microsoft Entra authentication uses Windows Authentication Manager (WAM), the signed-in
work/school account, and an OS-protected token cache. A tenant Conditional Access policy can still
require MFA or another explicit challenge, but gujiguji serializes that interaction. Foundry batch
transcription begins the check as soon as PCM capture is live and retains that recording while the
single WAM dialog completes. Use Switch Azure account… in Model Selection only when a different
work/school account should access the configured resource. The WAM integration is pinned to
Azure.Identity.Broker 1.3.1 (NuGet package: 73,665 bytes; SHA-256
2348dc1830c3966904e89fd0c952661cac13d5eb7970e284743ab42112469cd4).
On macOS, Azure Speech with Microsoft Entra ID also requires the Speech resource's Region and
full Azure Resource ID (for example /subscriptions/.../resourceGroups/.../providers/Microsoft.CognitiveServices/accounts/...).
The native Speech SDK uses these to construct Microsoft's required aad#resource-id#token
authorization token; the custom-domain endpoint and tenant remain configured as on Windows.
Windows needs the .NET 10 SDK; macOS needs Xcode/Swift 6.
make run # run from source
make install # build + install to %LOCALAPPDATA% + auto-start + launch
make release VERSION=vX.Y.Z SIGN_PFX=publisher.pfx # prompts securely for the PFX password
dotnet test tests/VoiceInput.Tests/VoiceInput.Tests.csproj
dotnet build src/VoiceInput/VoiceInput.csproj -p:EnableWindowsTargeting=true
swift test --package-path src/VoiceInputMac
scripts/build-macos.sh
The app shows its version in Settings and offers Update to vX.Y.Z… when a newer release exists. Updates are user-initiated, verified against a pinned Authenticode publisher when signed or the GitHub Release SHA-256 digest otherwise, atomically replaced, and rolled back if the new process does not stay running.
gpt-4o-transcribe deployment (e.g. in eastus2 / swedencentral).C#
57.2%
Swift
41.0%
PowerShell
1.1%