Me-in-U/LiveDialogue-Translator

A local-first Windows desktop app for live speech recognition, speaker diarization, captions, and translation.

C#

3

6 commits

updated Aug 26, 2026

See the code

README

LiveDialogue-Translator

LiveDialogue Translator is a local-first Windows desktop captioning app. It captures system audio and microphone audio directly, runs local speech-to-text, assigns speaker labels with local diarization, and can show translated captions in the main window or a transparent overlay.

The project is inspired by LiveCaptions-Translator, but it does not depend on Windows Live Captions. Audio is processed through the app's own capture pipeline and Python worker.

Screenshots

Screenshot assets are managed under docs/assets/screenshots. Additional notes and the overlay placeholder are kept in docs/screenshots.md.

Caption Workspace

LiveDialogue Translator caption workspace

The caption workspace is the default screen. It keeps live speaker captions, translation output, model status, and capture state visible in a compact window.

Settings

LiveDialogue Translator settings

The settings screen separates audio input, ASR model, diarization, translation, overlay, model management, and debug controls into compact groups.

Info

LiveDialogue Translator info

The info screen lists the project links, reference project, supported ASR/STT and diarization backends, license, runtime path, and local data directory.

Python Console

LiveDialogue Translator Python console

The console screen shows Python worker logs separately from captions, with quick controls for clearing logs and keeping the view pinned to the latest output.

Overlay

Overlay screenshot placeholder. Add the final overlay capture to docs/assets/screenshots/overlay.png and replace this block when the overlay image is ready.

Current Scope

  • Capture system audio, microphone audio, or a mixed device.
  • Normalize captured audio to 16 kHz mono PCM before worker processing.
  • Run a newline-delimited JSON protocol between the WPF app and Python worker.
  • Run local ASR/STT with faster-whisper, Qwen3-ASR, WhisperLiveKit, or WhisperX. faster-whisper is the default engine.
  • Install optional ASR engines into isolated package folders so engine-specific dependencies do not overwrite the base runtime.
  • Run local speaker diarization with pyannote.audio==4.0.4 and pyannote/speaker-diarization-community-1, Diart, or Sortformer.
  • Translate captions with the no-key Google provider. Other providers are currently shown as placeholders.
  • Show captions in a compact WPF shell and in a configurable transparent overlay.
  • Store app settings and downloaded runtime/model files under LocalAppData.
  • Follow the Windows UI language at startup. Korean and English strings are included.

Requirements

  • Windows desktop environment.
  • .NET 8 SDK for development builds.
  • Internet access on first setup so the app can download the managed Python runtime, Python packages, and selected model files.
  • Hugging Face access is required only when local pyannote diarization is used. Accept the selected model terms and provide a fine-grained token with User permissions > Repositories > Read access to contents of all public gated repos you can access.
  • NVIDIA CUDA is optional. The app can install CUDA-enabled PyTorch when an NVIDIA GPU is detected.

End users do not need to install Python manually. On first Start or Prepare, the app downloads the official Python 3.11.9 x64 embeddable runtime from python.org, extracts it under %LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9, bootstraps pip, and installs the worker dependencies there.

Model Combination Hardware Guide

The numbers below are conservative minimums for the current selectable model families in this app. They are intended for short live captures and model loading without out-of-memory failures. Real-time stability improves significantly with the recommended CUDA path and lower preset/model choices. They are app-level planning baselines, not official vendor guarantees.

Important assumptions:

  • CPU mode means Compute = CPU. It is practical for light Whisper models, but large Whisper, Qwen3-ASR, WhisperX, and Sortformer can fall behind real time.
  • CUDA mode means an NVIDIA GPU with current drivers and enough free VRAM. Avoid running another heavy CUDA workload at the same time.
  • Qwen3-ASR uses the Qwen3 forced aligner by default, so the 0.6B aligner must also fit in memory.
  • Diart uses CPU by default when Compute = Auto; choose Compute = CUDA if you want Diart to run on the GPU.
  • Sortformer uses the WhisperLiveKit Sortformer package even when the selected ASR engine is not WhisperLiveKit.
  • Keep at least 30 GB free disk for one heavy setup and 50-80 GB if you install every optional ASR engine and model cache.

ASR-only minimums

Use this table when speaker diarization is disabled.

ASR engine / modelCPU minimumCUDA minimumNotes
Faster-Whisper tiny, base, small4 CPU cores, 8 GB RAM4 GB VRAM, 8 GB RAMBest CPU-compatible path.
Faster-Whisper medium, large-v3, large-v3-turbo8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMCPU works, but large models may not keep up in live use.
Qwen3-ASR 0.6B + forced aligner8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCPU is mainly for testing; CUDA is strongly preferred.
Qwen3-ASR 1.7B + forced aligner12 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMDefault Qwen model; use CUDA for realistic latency.
WhisperLiveKit default (large-v3-turbo)8 CPU cores, 24 GB RAM8 GB VRAM, 16 GB RAMStreaming stack; CUDA recommended.
WhisperX tiny, base, small8 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMAlignment adds memory and startup cost over faster-whisper.
WhisperX medium, large-v3, large-v3-turbo12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMLower batch/model size if CUDA memory is tight.

ASR + speaker diarization minimums

Use this table when speaker diarization is enabled. Community-1 and Diart require accepted Hugging Face model terms and a valid token. Sortformer does not require pyannote model access, but it is the most CUDA-oriented diarization path.

ASR engine / modelDiarization modelCPU minimumCUDA minimumNotes
Faster-Whisper tiny/base/smallCommunity-14 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMLowest balanced setup with speaker labels.
Faster-Whisper tiny/base/smallDiart4 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMGood low-latency choice; Auto keeps Diart on CPU.
Faster-Whisper tiny/base/smallSortformer8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMCPU is usable only for light testing.
Faster-Whisper medium/large-v3/large-v3-turboCommunity-18 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCUDA recommended for live captions.
Faster-Whisper medium/large-v3/large-v3-turboDiart8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMStable if ASR model fits comfortably.
Faster-Whisper medium/large-v3/large-v3-turboSortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMPrefer 12 GB+ VRAM for large-v3.
Qwen3-ASR 0.6B + forced alignerCommunity-18 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCPU latency is high; use smaller chunks/presets if needed.
Qwen3-ASR 0.6B + forced alignerDiart8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCUDA leaves more CPU headroom for capture and translation.
Qwen3-ASR 0.6B + forced alignerSortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMRuns two heavy neural stacks; CUDA strongly preferred.
Qwen3-ASR 1.7B + forced alignerCommunity-112 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMPractical minimum for the default Qwen setup.
Qwen3-ASR 1.7B + forced alignerDiart12 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMUse CUDA and close other GPU workloads.
Qwen3-ASR 1.7B + forced alignerSortformer16 CPU cores, 64 GB RAM12 GB VRAM, 32 GB RAM16 GB VRAM is more comfortable for long sessions.
WhisperLiveKit default (large-v3-turbo)Community-112 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMWhisperLiveKit handles ASR; pyannote handles diarization.
WhisperLiveKit default (large-v3-turbo)Diart12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMUse CUDA if Diart should not consume CPU headroom.
WhisperLiveKit default (large-v3-turbo)Sortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMNative WhisperLiveKit + Sortformer streaming path.
WhisperX tiny/base/smallCommunity-18 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMWord alignment and diarization both add overhead.
WhisperX tiny/base/smallDiart8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMGood if word timestamps matter more than lowest latency.
WhisperX tiny/base/smallSortformer8 CPU cores, 24 GB RAM10 GB VRAM, 24 GB RAMSortformer adds the WhisperLiveKit runtime package.
WhisperX medium/large-v3/large-v3-turboCommunity-112 CPU cores, 32 GB RAM12 GB VRAM, 32 GB RAMReduce WhisperX batch size if memory is tight.
WhisperX medium/large-v3/large-v3-turboDiart12 CPU cores, 32 GB RAM12 GB VRAM, 32 GB RAMHeavy but reasonable on 12 GB+ NVIDIA GPUs.
WhisperX medium/large-v3/large-v3-turboSortformer16 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAM16 GB VRAM is preferred for long sessions.

For a general-purpose Windows desktop setup using CUDA, the practical baseline is a modern 8-core CPU, 32 GB system RAM, and an NVIDIA GPU with 12 GB VRAM. That class of machine can cover every selectable combination, although Qwen3-ASR 1.7B plus Sortformer or WhisperX large plus Sortformer benefits from 16 GB VRAM.

Quick Start

  1. Build or package the app with the commands below.
  2. Launch LiveDialogueTranslator.exe.
  3. Open Model Manager if the app asks for Hugging Face access.
  4. Choose the ASR model, diarization mode, input source, and translation target.
  5. Press Start capture.

The first run can take several minutes because the app prepares Python packages and model files. Later runs reuse the LocalAppData runtime and cache.

Build and Package

This repository can use a repo-local SDK at .dotnet-sdk\dotnet.exe or a system dotnet SDK.

.\.dotnet-sdk\dotnet.exe run --project tests\LiveDialogueTranslator.Tests\LiveDialogueTranslator.Tests.csproj
.\.dotnet-sdk\dotnet.exe build src\LiveDialogueTranslator.App\LiveDialogueTranslator.App.csproj -c Release
.\scripts\package.ps1

scripts\package.ps1 publishes to publish\win-x64. If Inno Setup is installed, it also creates artifacts\installer\LiveDialogueTranslatorSetup-x64.exe.

Manual Worker Testing

For developer-only worker testing, use the app-managed runtime after it has been prepared:

%LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9\python.exe -m pip install --no-warn-script-location -r worker\requirements.txt
%LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9\python.exe worker\speaker_worker.py

The worker reads JSON commands from stdin and writes JSON events to stdout. Use the WPF app for normal operation because it owns audio capture, model setup, translation, and worker lifecycle management.

Repository Layout

  • src\LiveDialogueTranslator.App - WPF desktop app, overlay window, settings, and worker orchestration.
  • src\LiveDialogueTranslator.Core - shared protocol, startup planning, transcript, speaker, runtime, and history logic.
  • worker - Python speech worker, package requirements, and engine environment presets.
  • tests\LiveDialogueTranslator.Tests - lightweight executable tests.
  • scripts - packaging helpers.
  • installer - Inno Setup installer definition.

Privacy

Audio is processed locally by the Python worker. Hugging Face is contacted only to download gated model files after you provide a token. Translation requests are sent to the selected translation provider.

Credits

License

LiveDialogue Translator is licensed under the Apache License 2.0. See LICENSE.

asr
captions
desktop-app
diarization
dotnet
huggingface
python
speaker-diarization
speech-recognition
translation
whisper
windows
wpf

Me-in-U/LiveDialogue-Translator

A local-first Windows desktop app for live speech recognition, speaker diarization, captions, and translation.

C#

3

6 commits

updated Aug 26, 2026

See the code

README

LiveDialogue-Translator

LiveDialogue Translator is a local-first Windows desktop captioning app. It captures system audio and microphone audio directly, runs local speech-to-text, assigns speaker labels with local diarization, and can show translated captions in the main window or a transparent overlay.

The project is inspired by LiveCaptions-Translator, but it does not depend on Windows Live Captions. Audio is processed through the app's own capture pipeline and Python worker.

Screenshots

Screenshot assets are managed under docs/assets/screenshots. Additional notes and the overlay placeholder are kept in docs/screenshots.md.

Caption Workspace

LiveDialogue Translator caption workspace

The caption workspace is the default screen. It keeps live speaker captions, translation output, model status, and capture state visible in a compact window.

Settings

LiveDialogue Translator settings

The settings screen separates audio input, ASR model, diarization, translation, overlay, model management, and debug controls into compact groups.

Info

LiveDialogue Translator info

The info screen lists the project links, reference project, supported ASR/STT and diarization backends, license, runtime path, and local data directory.

Python Console

LiveDialogue Translator Python console

The console screen shows Python worker logs separately from captions, with quick controls for clearing logs and keeping the view pinned to the latest output.

Overlay

Overlay screenshot placeholder. Add the final overlay capture to docs/assets/screenshots/overlay.png and replace this block when the overlay image is ready.

Current Scope

  • Capture system audio, microphone audio, or a mixed device.
  • Normalize captured audio to 16 kHz mono PCM before worker processing.
  • Run a newline-delimited JSON protocol between the WPF app and Python worker.
  • Run local ASR/STT with faster-whisper, Qwen3-ASR, WhisperLiveKit, or WhisperX. faster-whisper is the default engine.
  • Install optional ASR engines into isolated package folders so engine-specific dependencies do not overwrite the base runtime.
  • Run local speaker diarization with pyannote.audio==4.0.4 and pyannote/speaker-diarization-community-1, Diart, or Sortformer.
  • Translate captions with the no-key Google provider. Other providers are currently shown as placeholders.
  • Show captions in a compact WPF shell and in a configurable transparent overlay.
  • Store app settings and downloaded runtime/model files under LocalAppData.
  • Follow the Windows UI language at startup. Korean and English strings are included.

Requirements

  • Windows desktop environment.
  • .NET 8 SDK for development builds.
  • Internet access on first setup so the app can download the managed Python runtime, Python packages, and selected model files.
  • Hugging Face access is required only when local pyannote diarization is used. Accept the selected model terms and provide a fine-grained token with User permissions > Repositories > Read access to contents of all public gated repos you can access.
  • NVIDIA CUDA is optional. The app can install CUDA-enabled PyTorch when an NVIDIA GPU is detected.

End users do not need to install Python manually. On first Start or Prepare, the app downloads the official Python 3.11.9 x64 embeddable runtime from python.org, extracts it under %LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9, bootstraps pip, and installs the worker dependencies there.

Model Combination Hardware Guide

The numbers below are conservative minimums for the current selectable model families in this app. They are intended for short live captures and model loading without out-of-memory failures. Real-time stability improves significantly with the recommended CUDA path and lower preset/model choices. They are app-level planning baselines, not official vendor guarantees.

Important assumptions:

  • CPU mode means Compute = CPU. It is practical for light Whisper models, but large Whisper, Qwen3-ASR, WhisperX, and Sortformer can fall behind real time.
  • CUDA mode means an NVIDIA GPU with current drivers and enough free VRAM. Avoid running another heavy CUDA workload at the same time.
  • Qwen3-ASR uses the Qwen3 forced aligner by default, so the 0.6B aligner must also fit in memory.
  • Diart uses CPU by default when Compute = Auto; choose Compute = CUDA if you want Diart to run on the GPU.
  • Sortformer uses the WhisperLiveKit Sortformer package even when the selected ASR engine is not WhisperLiveKit.
  • Keep at least 30 GB free disk for one heavy setup and 50-80 GB if you install every optional ASR engine and model cache.

ASR-only minimums

Use this table when speaker diarization is disabled.

ASR engine / modelCPU minimumCUDA minimumNotes
Faster-Whisper tiny, base, small4 CPU cores, 8 GB RAM4 GB VRAM, 8 GB RAMBest CPU-compatible path.
Faster-Whisper medium, large-v3, large-v3-turbo8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMCPU works, but large models may not keep up in live use.
Qwen3-ASR 0.6B + forced aligner8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCPU is mainly for testing; CUDA is strongly preferred.
Qwen3-ASR 1.7B + forced aligner12 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMDefault Qwen model; use CUDA for realistic latency.
WhisperLiveKit default (large-v3-turbo)8 CPU cores, 24 GB RAM8 GB VRAM, 16 GB RAMStreaming stack; CUDA recommended.
WhisperX tiny, base, small8 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMAlignment adds memory and startup cost over faster-whisper.
WhisperX medium, large-v3, large-v3-turbo12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMLower batch/model size if CUDA memory is tight.

ASR + speaker diarization minimums

Use this table when speaker diarization is enabled. Community-1 and Diart require accepted Hugging Face model terms and a valid token. Sortformer does not require pyannote model access, but it is the most CUDA-oriented diarization path.

ASR engine / modelDiarization modelCPU minimumCUDA minimumNotes
Faster-Whisper tiny/base/smallCommunity-14 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMLowest balanced setup with speaker labels.
Faster-Whisper tiny/base/smallDiart4 CPU cores, 16 GB RAM6 GB VRAM, 16 GB RAMGood low-latency choice; Auto keeps Diart on CPU.
Faster-Whisper tiny/base/smallSortformer8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMCPU is usable only for light testing.
Faster-Whisper medium/large-v3/large-v3-turboCommunity-18 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCUDA recommended for live captions.
Faster-Whisper medium/large-v3/large-v3-turboDiart8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMStable if ASR model fits comfortably.
Faster-Whisper medium/large-v3/large-v3-turboSortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMPrefer 12 GB+ VRAM for large-v3.
Qwen3-ASR 0.6B + forced alignerCommunity-18 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCPU latency is high; use smaller chunks/presets if needed.
Qwen3-ASR 0.6B + forced alignerDiart8 CPU cores, 32 GB RAM8 GB VRAM, 24 GB RAMCUDA leaves more CPU headroom for capture and translation.
Qwen3-ASR 0.6B + forced alignerSortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMRuns two heavy neural stacks; CUDA strongly preferred.
Qwen3-ASR 1.7B + forced alignerCommunity-112 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMPractical minimum for the default Qwen setup.
Qwen3-ASR 1.7B + forced alignerDiart12 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAMUse CUDA and close other GPU workloads.
Qwen3-ASR 1.7B + forced alignerSortformer16 CPU cores, 64 GB RAM12 GB VRAM, 32 GB RAM16 GB VRAM is more comfortable for long sessions.
WhisperLiveKit default (large-v3-turbo)Community-112 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMWhisperLiveKit handles ASR; pyannote handles diarization.
WhisperLiveKit default (large-v3-turbo)Diart12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMUse CUDA if Diart should not consume CPU headroom.
WhisperLiveKit default (large-v3-turbo)Sortformer12 CPU cores, 32 GB RAM10 GB VRAM, 24 GB RAMNative WhisperLiveKit + Sortformer streaming path.
WhisperX tiny/base/smallCommunity-18 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMWord alignment and diarization both add overhead.
WhisperX tiny/base/smallDiart8 CPU cores, 16 GB RAM8 GB VRAM, 16 GB RAMGood if word timestamps matter more than lowest latency.
WhisperX tiny/base/smallSortformer8 CPU cores, 24 GB RAM10 GB VRAM, 24 GB RAMSortformer adds the WhisperLiveKit runtime package.
WhisperX medium/large-v3/large-v3-turboCommunity-112 CPU cores, 32 GB RAM12 GB VRAM, 32 GB RAMReduce WhisperX batch size if memory is tight.
WhisperX medium/large-v3/large-v3-turboDiart12 CPU cores, 32 GB RAM12 GB VRAM, 32 GB RAMHeavy but reasonable on 12 GB+ NVIDIA GPUs.
WhisperX medium/large-v3/large-v3-turboSortformer16 CPU cores, 48 GB RAM12 GB VRAM, 32 GB RAM16 GB VRAM is preferred for long sessions.

For a general-purpose Windows desktop setup using CUDA, the practical baseline is a modern 8-core CPU, 32 GB system RAM, and an NVIDIA GPU with 12 GB VRAM. That class of machine can cover every selectable combination, although Qwen3-ASR 1.7B plus Sortformer or WhisperX large plus Sortformer benefits from 16 GB VRAM.

Quick Start

  1. Build or package the app with the commands below.
  2. Launch LiveDialogueTranslator.exe.
  3. Open Model Manager if the app asks for Hugging Face access.
  4. Choose the ASR model, diarization mode, input source, and translation target.
  5. Press Start capture.

The first run can take several minutes because the app prepares Python packages and model files. Later runs reuse the LocalAppData runtime and cache.

Build and Package

This repository can use a repo-local SDK at .dotnet-sdk\dotnet.exe or a system dotnet SDK.

.\.dotnet-sdk\dotnet.exe run --project tests\LiveDialogueTranslator.Tests\LiveDialogueTranslator.Tests.csproj
.\.dotnet-sdk\dotnet.exe build src\LiveDialogueTranslator.App\LiveDialogueTranslator.App.csproj -c Release
.\scripts\package.ps1

scripts\package.ps1 publishes to publish\win-x64. If Inno Setup is installed, it also creates artifacts\installer\LiveDialogueTranslatorSetup-x64.exe.

Manual Worker Testing

For developer-only worker testing, use the app-managed runtime after it has been prepared:

%LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9\python.exe -m pip install --no-warn-script-location -r worker\requirements.txt
%LOCALAPPDATA%\LiveDialogue Translator\runtime\python-3.11.9\python.exe worker\speaker_worker.py

The worker reads JSON commands from stdin and writes JSON events to stdout. Use the WPF app for normal operation because it owns audio capture, model setup, translation, and worker lifecycle management.

Repository Layout

  • src\LiveDialogueTranslator.App - WPF desktop app, overlay window, settings, and worker orchestration.
  • src\LiveDialogueTranslator.Core - shared protocol, startup planning, transcript, speaker, runtime, and history logic.
  • worker - Python speech worker, package requirements, and engine environment presets.
  • tests\LiveDialogueTranslator.Tests - lightweight executable tests.
  • scripts - packaging helpers.
  • installer - Inno Setup installer definition.

Privacy

Audio is processed locally by the Python worker. Hugging Face is contacted only to download gated model files after you provide a token. Translation requests are sent to the selected translation provider.

Credits

License

LiveDialogue Translator is licensed under the Apache License 2.0. See LICENSE.

asr
captions
desktop-app
diarization
dotnet
huggingface
python
speaker-diarization
speech-recognition
translation
whisper
windows
wpf

Languages

C#

55.0%

Python

44.6%