ONNX speech pipeline library for ASR (diarization, VAD), and TTS
C#
23
834 commits
updated Sep 30, 2026
A .NET 10 speech pipeline library and toolset for local, offline inference using ONNX models.
No cloud. No telemetry. Runs entirely on your hardware.
Vernacula converts audio into accurate, multi-speaker transcripts on your own computer. It ships as a reusable library (Vernacula.Base), a command-line tool (Vernacula.CLI), and a cross-platform desktop app (Vernacula-Desktop, built on Avalonia UI).
Powered by NVIDIA's Parakeet TDT v3 and Sortformer by default, with optional pluggable backends (Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1). Parakeet v3 posts a Word Error Rate of 4.85 on Google's FLEURS benchmark. Most modern computers will transcribe one hour of audio in about five minutes; GPU-accelerated systems are significantly faster.
https://github.com/user-attachments/assets/42015635-03b9-4c6b-868c-248e8c29c352

More screenshots and a feature tour live in docs/desktop-app.md.
Vernacula's models are converted in-house from upstream PyTorch / NeMo / HuggingFace checkpoints into the ONNX contract its C# inference code expects. The export tooling lives in scripts/ and is usable independently of the rest of the project — the export scripts are dev-time only and never ship as a runtime dependency.
Most of these graphs (split KV-cache decoders, transducer/TDT decoder state, streaming GRU hidden-state I/O, six-input Sortformer chunked diarization) require non-trivial graph surgery beyond torch.onnx.export defaults. Each export folder has its own README with the contract, parity checks, and tuning notes.
A KenLM build pipeline for Parakeet shallow fusion lives in scripts/kenlm_build; an in-progress IndicConformer export spike is in scripts/indicconformer_export.
Install prerequisites — .NET 10 SDK. FFmpeg is optional: WAV, MP3, AIFF, Ogg Vorbis and Ogg Opus are decoded in-process, and FFmpeg is only needed for FLAC, M4A/AAC, WMA and video containers — on Windows the desktop app can fetch it for you. Full setup (including GPU) is in docs/installation.md.
Run the desktop app:
cd src/Vernacula.Avalonia
# Windows — native WinMM audio output, no external player needed
dotnet run -f net10.0-windows
# Linux / macOS — playback goes through ffplay
dotnet run -f net10.0
-f is required because the desktop app targets two frameworks. NAudio 3 hands the
Windows audio backend (WaveOut) only to a Windows target framework, so net10.0-windows
is what carries native playback; net10.0 is the portable build and uses ffplay.
On Linux, ./install.sh from the repo root builds a self-contained package and registers a .desktop entry.
On macOS, ./package-macos.sh builds dist/Vernacula.app, ready to copy into /Applications.
Run the CLI:
dotnet run --project src/Vernacula.CLI -p:EP=Cuda -- \
--audio meeting.wav
Full argument reference and more examples in docs/cli-reference.md. Build configurations (CUDA / CPU / DirectML) in docs/building.md.
Full documentation lives in docs/.
Getting started
Reference
Project
Vernacula.Base and Vernacula.CLI — MITVernacula.Avalonia — PolyForm Shield 1.0.0 (free to use and build; may not be used to create a competing commercial product)See docs/licensing.md for the full breakdown.
C#
60.6%
Python
30.9%
JavaScript
4.6%
TypeScript
3.4%
ONNX speech pipeline library for ASR (diarization, VAD), and TTS
C#
23
834 commits
updated Sep 30, 2026
A .NET 10 speech pipeline library and toolset for local, offline inference using ONNX models.
No cloud. No telemetry. Runs entirely on your hardware.
Vernacula converts audio into accurate, multi-speaker transcripts on your own computer. It ships as a reusable library (Vernacula.Base), a command-line tool (Vernacula.CLI), and a cross-platform desktop app (Vernacula-Desktop, built on Avalonia UI).
Powered by NVIDIA's Parakeet TDT v3 and Sortformer by default, with optional pluggable backends (Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1). Parakeet v3 posts a Word Error Rate of 4.85 on Google's FLEURS benchmark. Most modern computers will transcribe one hour of audio in about five minutes; GPU-accelerated systems are significantly faster.
https://github.com/user-attachments/assets/42015635-03b9-4c6b-868c-248e8c29c352

More screenshots and a feature tour live in docs/desktop-app.md.
Vernacula's models are converted in-house from upstream PyTorch / NeMo / HuggingFace checkpoints into the ONNX contract its C# inference code expects. The export tooling lives in scripts/ and is usable independently of the rest of the project — the export scripts are dev-time only and never ship as a runtime dependency.
Most of these graphs (split KV-cache decoders, transducer/TDT decoder state, streaming GRU hidden-state I/O, six-input Sortformer chunked diarization) require non-trivial graph surgery beyond torch.onnx.export defaults. Each export folder has its own README with the contract, parity checks, and tuning notes.
A KenLM build pipeline for Parakeet shallow fusion lives in scripts/kenlm_build; an in-progress IndicConformer export spike is in scripts/indicconformer_export.
Install prerequisites — .NET 10 SDK. FFmpeg is optional: WAV, MP3, AIFF, Ogg Vorbis and Ogg Opus are decoded in-process, and FFmpeg is only needed for FLAC, M4A/AAC, WMA and video containers — on Windows the desktop app can fetch it for you. Full setup (including GPU) is in docs/installation.md.
Run the desktop app:
cd src/Vernacula.Avalonia
# Windows — native WinMM audio output, no external player needed
dotnet run -f net10.0-windows
# Linux / macOS — playback goes through ffplay
dotnet run -f net10.0
-f is required because the desktop app targets two frameworks. NAudio 3 hands the
Windows audio backend (WaveOut) only to a Windows target framework, so net10.0-windows
is what carries native playback; net10.0 is the portable build and uses ffplay.
On Linux, ./install.sh from the repo root builds a self-contained package and registers a .desktop entry.
On macOS, ./package-macos.sh builds dist/Vernacula.app, ready to copy into /Applications.
Run the CLI:
dotnet run --project src/Vernacula.CLI -p:EP=Cuda -- \
--audio meeting.wav
Full argument reference and more examples in docs/cli-reference.md. Build configurations (CUDA / CPU / DirectML) in docs/building.md.
Full documentation lives in docs/.
Getting started
Reference
Project
Vernacula.Base and Vernacula.CLI — MITVernacula.Avalonia — PolyForm Shield 1.0.0 (free to use and build; may not be used to create a competing commercial product)See docs/licensing.md for the full breakdown.
C#
60.6%
Python
30.9%
JavaScript
4.6%
TypeScript
3.4%