christopherthompson81/vernacula

ONNX speech pipeline library for ASR (diarization, VAD), and TTS

C#

23

834 commits

updated Sep 30, 2026

See the code

README

Vernacula

A .NET 10 speech pipeline library and toolset for local, offline inference using ONNX models.
No cloud. No telemetry. Runs entirely on your hardware.

Vernacula-Desktop

Vernacula converts audio into accurate, multi-speaker transcripts on your own computer. It ships as a reusable library (Vernacula.Base), a command-line tool (Vernacula.CLI), and a cross-platform desktop app (Vernacula-Desktop, built on Avalonia UI).

Powered by NVIDIA's Parakeet TDT v3 and Sortformer by default, with optional pluggable backends (Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1). Parakeet v3 posts a Word Error Rate of 4.85 on Google's FLEURS benchmark. Most modern computers will transcribe one hour of audio in about five minutes; GPU-accelerated systems are significantly faster.

Demo

https://github.com/user-attachments/assets/42015635-03b9-4c6b-868c-248e8c29c352

Results view

More screenshots and a feature tour live in docs/desktop-app.md.

Highlights

  • Local, private transcription — audio never leaves your computer
  • Multi-speaker detection — identifies and labels up to four concurrent speakers
  • No audio length limits — the default pipeline streams and segments, so file length is unbounded. The VibeVoice backends hold the whole recording in context to attribute speakers without a diarizer, so anything past their context window (about 68 minutes for VibeVoice-ASR Streaming 1.5B, 2 hours for the 7B, less on a card that cannot hold that much context alongside the model) is decoded in several passes. Each pass gets its own set of speaker labels, since nothing carries identity across the reset — merge them in the editor if the same person appears in more than one.
  • Transcript editor with confidence colouring, audio playback, and word-level timestamps
  • Pluggable ASR backends — Parakeet TDT v3, Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1
  • Shallow KenLM fusion for domain-specific English (general, medical)
  • Export to XLSX, CSV, JSON, SRT, Markdown, DOCX, and SQLite
  • GPU acceleration via CUDA (DirectML on Windows), with automatic CPU fallback
  • 52 languages covered across the four backends — see the support matrix

Model conversion pipelines

Vernacula's models are converted in-house from upstream PyTorch / NeMo / HuggingFace checkpoints into the ONNX contract its C# inference code expects. The export tooling lives in scripts/ and is usable independently of the rest of the project — the export scripts are dev-time only and never ship as a runtime dependency.

Most of these graphs (split KV-cache decoders, transducer/TDT decoder state, streaming GRU hidden-state I/O, six-input Sortformer chunked diarization) require non-trivial graph surgery beyond torch.onnx.export defaults. Each export folder has its own README with the contract, parity checks, and tuning notes.

A KenLM build pipeline for Parakeet shallow fusion lives in scripts/kenlm_build; an in-progress IndicConformer export spike is in scripts/indicconformer_export.

Quick start

Install prerequisites — .NET 10 SDK. FFmpeg is optional: WAV, MP3, AIFF, Ogg Vorbis and Ogg Opus are decoded in-process, and FFmpeg is only needed for FLAC, M4A/AAC, WMA and video containers — on Windows the desktop app can fetch it for you. Full setup (including GPU) is in docs/installation.md.

Run the desktop app:

cd src/Vernacula.Avalonia

# Windows — native WinMM audio output, no external player needed
dotnet run -f net10.0-windows

# Linux / macOS — playback goes through ffplay
dotnet run -f net10.0

-f is required because the desktop app targets two frameworks. NAudio 3 hands the Windows audio backend (WaveOut) only to a Windows target framework, so net10.0-windows is what carries native playback; net10.0 is the portable build and uses ffplay.

On Linux, ./install.sh from the repo root builds a self-contained package and registers a .desktop entry. On macOS, ./package-macos.sh builds dist/Vernacula.app, ready to copy into /Applications.

Run the CLI:

dotnet run --project src/Vernacula.CLI -p:EP=Cuda -- \
  --audio meeting.wav

Full argument reference and more examples in docs/cli-reference.md. Build configurations (CUDA / CPU / DirectML) in docs/building.md.

Documentation

Full documentation lives in docs/.

Getting started

Reference

Project

License

  • Vernacula.Base and Vernacula.CLI — MIT
  • Vernacula.Avalonia — PolyForm Shield 1.0.0 (free to use and build; may not be used to create a competing commercial product)
  • Model weights — see respective HuggingFace repository licenses

See docs/licensing.md for the full breakdown.

asr
avalonia
cohere-transcribe
cross-platform
csharp
deepfilternet
diarization
dotnet
kenlm
local-first
offline
onnx
onnxruntime
parakeet
qwen3-asr
sortformer
speaker-diarization
speech-to-text
transcription
vibevoice

christopherthompson81/vernacula

ONNX speech pipeline library for ASR (diarization, VAD), and TTS

C#

23

834 commits

updated Sep 30, 2026

See the code

README

Vernacula

A .NET 10 speech pipeline library and toolset for local, offline inference using ONNX models.
No cloud. No telemetry. Runs entirely on your hardware.

Vernacula-Desktop

Vernacula converts audio into accurate, multi-speaker transcripts on your own computer. It ships as a reusable library (Vernacula.Base), a command-line tool (Vernacula.CLI), and a cross-platform desktop app (Vernacula-Desktop, built on Avalonia UI).

Powered by NVIDIA's Parakeet TDT v3 and Sortformer by default, with optional pluggable backends (Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1). Parakeet v3 posts a Word Error Rate of 4.85 on Google's FLEURS benchmark. Most modern computers will transcribe one hour of audio in about five minutes; GPU-accelerated systems are significantly faster.

Demo

https://github.com/user-attachments/assets/42015635-03b9-4c6b-868c-248e8c29c352

Results view

More screenshots and a feature tour live in docs/desktop-app.md.

Highlights

  • Local, private transcription — audio never leaves your computer
  • Multi-speaker detection — identifies and labels up to four concurrent speakers
  • No audio length limits — the default pipeline streams and segments, so file length is unbounded. The VibeVoice backends hold the whole recording in context to attribute speakers without a diarizer, so anything past their context window (about 68 minutes for VibeVoice-ASR Streaming 1.5B, 2 hours for the 7B, less on a card that cannot hold that much context alongside the model) is decoded in several passes. Each pass gets its own set of speaker labels, since nothing carries identity across the reset — merge them in the editor if the same person appears in more than one.
  • Transcript editor with confidence colouring, audio playback, and word-level timestamps
  • Pluggable ASR backends — Parakeet TDT v3, Cohere Transcribe, Qwen3-ASR, VibeVoice-ASR, VibeVoice-ASR Streaming, Whisper large-v3-turbo, IndicConformer, Granite Speech 4.1
  • Shallow KenLM fusion for domain-specific English (general, medical)
  • Export to XLSX, CSV, JSON, SRT, Markdown, DOCX, and SQLite
  • GPU acceleration via CUDA (DirectML on Windows), with automatic CPU fallback
  • 52 languages covered across the four backends — see the support matrix

Model conversion pipelines

Vernacula's models are converted in-house from upstream PyTorch / NeMo / HuggingFace checkpoints into the ONNX contract its C# inference code expects. The export tooling lives in scripts/ and is usable independently of the rest of the project — the export scripts are dev-time only and never ship as a runtime dependency.

Most of these graphs (split KV-cache decoders, transducer/TDT decoder state, streaming GRU hidden-state I/O, six-input Sortformer chunked diarization) require non-trivial graph surgery beyond torch.onnx.export defaults. Each export folder has its own README with the contract, parity checks, and tuning notes.

A KenLM build pipeline for Parakeet shallow fusion lives in scripts/kenlm_build; an in-progress IndicConformer export spike is in scripts/indicconformer_export.

Quick start

Install prerequisites — .NET 10 SDK. FFmpeg is optional: WAV, MP3, AIFF, Ogg Vorbis and Ogg Opus are decoded in-process, and FFmpeg is only needed for FLAC, M4A/AAC, WMA and video containers — on Windows the desktop app can fetch it for you. Full setup (including GPU) is in docs/installation.md.

Run the desktop app:

cd src/Vernacula.Avalonia

# Windows — native WinMM audio output, no external player needed
dotnet run -f net10.0-windows

# Linux / macOS — playback goes through ffplay
dotnet run -f net10.0

-f is required because the desktop app targets two frameworks. NAudio 3 hands the Windows audio backend (WaveOut) only to a Windows target framework, so net10.0-windows is what carries native playback; net10.0 is the portable build and uses ffplay.

On Linux, ./install.sh from the repo root builds a self-contained package and registers a .desktop entry. On macOS, ./package-macos.sh builds dist/Vernacula.app, ready to copy into /Applications.

Run the CLI:

dotnet run --project src/Vernacula.CLI -p:EP=Cuda -- \
  --audio meeting.wav

Full argument reference and more examples in docs/cli-reference.md. Build configurations (CUDA / CPU / DirectML) in docs/building.md.

Documentation

Full documentation lives in docs/.

Getting started

Reference

Project

License

  • Vernacula.Base and Vernacula.CLI — MIT
  • Vernacula.Avalonia — PolyForm Shield 1.0.0 (free to use and build; may not be used to create a competing commercial product)
  • Model weights — see respective HuggingFace repository licenses

See docs/licensing.md for the full breakdown.

asr
avalonia
cohere-transcribe
cross-platform
csharp
deepfilternet
diarization
dotnet
kenlm
local-first
offline
onnx
onnxruntime
parakeet
qwen3-asr
sortformer
speaker-diarization
speech-to-text
transcription
vibevoice

Languages

C#

60.6%

Python

30.9%

JavaScript

4.6%

TypeScript

3.4%