WinUI 3 (Windows App SDK) desktop app for offline speech transcription. All ASR inference runs locally after model download.
This repo also supports an optional Translate mode (offline MT + optional TTS) in the same app.
| Main Screen | File Transcription | Settings |
|---|---|---|
![]() | ![]() | ![]() |

Select an audio file (.wav / .mp3) and get the transcript in seconds — all processing runs locally.
Transcribe: transcription only.Translate: transcription + offline translation + optional TTS playback.Microphone, System Audio (WASAPI loopback)..wav, .mp3).CPU, RAM, tok/s, elapsed audio). Note: tok/s is a rough word-per-second estimate.translation.txt and tts.wav in the exported ZIP.Evidence Mode (events.jsonl + model manifests + screenshots; export one ZIP).Defined in src/OfflineTranscription/Models/ModelInfo.cs.
Models are downloaded from Hugging Face at runtime and stored under %LOCALAPPDATA%\\OfflineTranscription\\Models\\<model-id>\\.
Benchmark: JFK inauguration excerpt (11s, 22 words). Device: i5-1035G1 (4C/8T), 8 GB RAM, CPU-only.
| Model ID | Engine | Params (weights) | Disk (download) | Languages | Inference | RTF | Words/s |
|---|---|---|---|---|---|---|---|
moonshine-tiny | sherpa-onnx offline | 27 M | ~125 MB | English | 435 ms | 0.040x | 50.6 |
sensevoice-small | sherpa-onnx offline | 234 M | ~240 MB | zh/en/ja/ko/yue | 462 ms | 0.042x | 47.6 |
moonshine-base | sherpa-onnx offline | 61 M | ~290 MB | English | 534 ms | 0.049x | 41.2 |
parakeet-tdt-v2 | sherpa-onnx offline | 600 M | ~660 MB | English | 1,239 ms | 0.113x | 17.8 |
zipformer-20m | sherpa-onnx streaming | 20 M | ~73 MB | English | 1,775 ms | 0.161x | 12.4 |
whisper-tiny | whisper.cpp | 39 M | ~80 MB | 99 languages | 2,325 ms | 0.211x | 9.5 |
omnilingual-300m | sherpa-onnx offline | 300 M | ~365 MB | 1,600+ languages | 2,360 ms | 0.215x | — |
whisper-base | whisper.cpp | 74 M | ~150 MB | 99 languages | 6,501 ms | 0.591x | 3.4 |
qwen3-asr-0.6b | qwen-asr (C) | 600 M | ~1.9 GB | 52 languages | 13,359 ms | 1.214x | 1.6 |
whisper-small | whisper.cpp | 244 M | ~500 MB | 99 languages | 21,260 ms | 1.933x | 1.0 |
whisper-large-v3-turbo | whisper.cpp | 809 M | ~834 MB | 99 languages | 92,845 ms | 8.440x | 0.2 |
windows-speech | Windows Speech API | N/A | 0 MB | Installed packs | — | — | — |


Model weights are not distributed with this repo; model licensing varies. See NOTICE.
Want to see a new model or device benchmark? If there is an offline ASR model you would like added, or you have benchmark results to share from your hardware, please open an issue. Community contributions of inference benchmarks on different Windows devices are welcome.
This app has multiple inference paths:
sherpa-onnx streaming (Zipformer):
whisper.cpp and sherpa-onnx offline:
Provider selection:
sherpa-onnx offline probes DirectML first and falls back to CPU.sherpa-onnx streaming uses CPU for stable real-time behavior.qwen-asr (antirez/qwen-asr): native C engine for Qwen3-ASR models.qwen_asr.dll from source (see Native Dependencies).WindowsSpeech: built-in Windows recognizer via WinRT SpeechRecognizer.When you select Translate in the left navigation:
Translation settings:
Source Language and Target Language in Settings (Translate mode only).%LOCALAPPDATA%\\OfflineTranscription\\TranslationModels\\.On first run, the app performs a best-effort migration from the legacy Windows app data folder
%LOCALAPPDATA%\\OfflineSpeechTranslation\\ into %LOCALAPPDATA%\\OfflineTranscription\\ when the new location does not already exist.
This app uses native engines via P/Invoke:
whisper.dll (see src/OfflineTranscription/Interop/WhisperNative.cs)sherpa-onnx-c-api.dll (see src/OfflineTranscription/Interop/SherpaOnnxNative.cs)qwen_asr.dll + libopenblas.dll (see src/OfflineTranscription/Interop/QwenAsrNative.cs)OfflineTranscription.NativeTranslation.dll (see src/OfflineTranscription/Interop/NativeTranslation.cs)Place required .dll files in:
libs/runtimes/win-x64/They are copied to the build output directory by src/OfflineTranscription/OfflineTranscription.csproj.
If you use the sherpa-onnx release bundle, make sure you also include its dependent DLLs
(for example onnxruntime.dll and DirectML-related DLLs if applicable).
qwen_asr.dll (Windows)For qwen-asr, compile qwen_asr.dll from antirez/qwen-asr source
and place alongside libopenblas.dll (from OpenMathLib/OpenBLAS releases).
OfflineTranscription.NativeTranslation.dll (Windows)The native translation wrapper lives in src/OfflineTranscription.NativeTranslation/ and is built with CMake.
src/OfflineTranscription.NativeTranslation/README.md.OfflineTranscription.NativeTranslation.dll (and any dependency DLLs) into:
libs/runtimes/win-x64/Open OfflineTranscription.sln and build x64 (Debug or Release).
Save Screenshot writes PNGs under %LOCALAPPDATA%\\OfflineTranscription\\Diagnostics\\Evidence Mode writes under %LOCALAPPDATA%\\OfflineTranscription\\Evidence\\ and can export a single evidence ZIP.See DEVICE_TESTING.md for a real-device test checklist.
dotnet test tests/OfflineTranscription.Tests/OfflineTranscription.Tests.csproj -c Release
Apache License 2.0. See LICENSE and NOTICE.
C#
90.9%
PowerShell
4.0%
Python
3.6%
WinUI 3 (Windows App SDK) desktop app for offline speech transcription. All ASR inference runs locally after model download.
This repo also supports an optional Translate mode (offline MT + optional TTS) in the same app.
| Main Screen | File Transcription | Settings |
|---|---|---|
![]() | ![]() | ![]() |

Select an audio file (.wav / .mp3) and get the transcript in seconds — all processing runs locally.
Transcribe: transcription only.Translate: transcription + offline translation + optional TTS playback.Microphone, System Audio (WASAPI loopback)..wav, .mp3).CPU, RAM, tok/s, elapsed audio). Note: tok/s is a rough word-per-second estimate.translation.txt and tts.wav in the exported ZIP.Evidence Mode (events.jsonl + model manifests + screenshots; export one ZIP).Defined in src/OfflineTranscription/Models/ModelInfo.cs.
Models are downloaded from Hugging Face at runtime and stored under %LOCALAPPDATA%\\OfflineTranscription\\Models\\<model-id>\\.
Benchmark: JFK inauguration excerpt (11s, 22 words). Device: i5-1035G1 (4C/8T), 8 GB RAM, CPU-only.
| Model ID | Engine | Params (weights) | Disk (download) | Languages | Inference | RTF | Words/s |
|---|---|---|---|---|---|---|---|
moonshine-tiny | sherpa-onnx offline | 27 M | ~125 MB | English | 435 ms | 0.040x | 50.6 |
sensevoice-small | sherpa-onnx offline | 234 M | ~240 MB | zh/en/ja/ko/yue | 462 ms | 0.042x | 47.6 |
moonshine-base | sherpa-onnx offline | 61 M | ~290 MB | English | 534 ms | 0.049x | 41.2 |
parakeet-tdt-v2 | sherpa-onnx offline | 600 M | ~660 MB | English | 1,239 ms | 0.113x | 17.8 |
zipformer-20m | sherpa-onnx streaming | 20 M | ~73 MB | English | 1,775 ms | 0.161x | 12.4 |
whisper-tiny | whisper.cpp | 39 M | ~80 MB | 99 languages | 2,325 ms | 0.211x | 9.5 |
omnilingual-300m | sherpa-onnx offline | 300 M | ~365 MB | 1,600+ languages | 2,360 ms | 0.215x | — |
whisper-base | whisper.cpp | 74 M | ~150 MB | 99 languages | 6,501 ms | 0.591x | 3.4 |
qwen3-asr-0.6b | qwen-asr (C) | 600 M | ~1.9 GB | 52 languages | 13,359 ms | 1.214x | 1.6 |
whisper-small | whisper.cpp | 244 M | ~500 MB | 99 languages | 21,260 ms | 1.933x | 1.0 |
whisper-large-v3-turbo | whisper.cpp | 809 M | ~834 MB | 99 languages | 92,845 ms | 8.440x | 0.2 |
windows-speech | Windows Speech API | N/A | 0 MB | Installed packs | — | — | — |


Model weights are not distributed with this repo; model licensing varies. See NOTICE.
Want to see a new model or device benchmark? If there is an offline ASR model you would like added, or you have benchmark results to share from your hardware, please open an issue. Community contributions of inference benchmarks on different Windows devices are welcome.
This app has multiple inference paths:
sherpa-onnx streaming (Zipformer):
whisper.cpp and sherpa-onnx offline:
Provider selection:
sherpa-onnx offline probes DirectML first and falls back to CPU.sherpa-onnx streaming uses CPU for stable real-time behavior.qwen-asr (antirez/qwen-asr): native C engine for Qwen3-ASR models.qwen_asr.dll from source (see Native Dependencies).WindowsSpeech: built-in Windows recognizer via WinRT SpeechRecognizer.When you select Translate in the left navigation:
Translation settings:
Source Language and Target Language in Settings (Translate mode only).%LOCALAPPDATA%\\OfflineTranscription\\TranslationModels\\.On first run, the app performs a best-effort migration from the legacy Windows app data folder
%LOCALAPPDATA%\\OfflineSpeechTranslation\\ into %LOCALAPPDATA%\\OfflineTranscription\\ when the new location does not already exist.
This app uses native engines via P/Invoke:
whisper.dll (see src/OfflineTranscription/Interop/WhisperNative.cs)sherpa-onnx-c-api.dll (see src/OfflineTranscription/Interop/SherpaOnnxNative.cs)qwen_asr.dll + libopenblas.dll (see src/OfflineTranscription/Interop/QwenAsrNative.cs)OfflineTranscription.NativeTranslation.dll (see src/OfflineTranscription/Interop/NativeTranslation.cs)Place required .dll files in:
libs/runtimes/win-x64/They are copied to the build output directory by src/OfflineTranscription/OfflineTranscription.csproj.
If you use the sherpa-onnx release bundle, make sure you also include its dependent DLLs
(for example onnxruntime.dll and DirectML-related DLLs if applicable).
qwen_asr.dll (Windows)For qwen-asr, compile qwen_asr.dll from antirez/qwen-asr source
and place alongside libopenblas.dll (from OpenMathLib/OpenBLAS releases).
OfflineTranscription.NativeTranslation.dll (Windows)The native translation wrapper lives in src/OfflineTranscription.NativeTranslation/ and is built with CMake.
src/OfflineTranscription.NativeTranslation/README.md.OfflineTranscription.NativeTranslation.dll (and any dependency DLLs) into:
libs/runtimes/win-x64/Open OfflineTranscription.sln and build x64 (Debug or Release).
Save Screenshot writes PNGs under %LOCALAPPDATA%\\OfflineTranscription\\Diagnostics\\Evidence Mode writes under %LOCALAPPDATA%\\OfflineTranscription\\Evidence\\ and can export a single evidence ZIP.See DEVICE_TESTING.md for a real-device test checklist.
dotnet test tests/OfflineTranscription.Tests/OfflineTranscription.Tests.csproj -c Release
Apache License 2.0. See LICENSE and NOTICE.
C#
90.9%
PowerShell
4.0%
Python
3.6%