A 100% local audio transcription, editing, and review tool
See the codeA small local voice-to-text starter that can:
Whisper.net (GPU via CUDA runtime if available)ollama pull llama3.2HF_TOKEN$env:HF_TOKEN = "hf_xxx"From the repo root:
cd src\LocalTranscriber.Cli
dotnet restore
dotnet run -- devices
dotnet run -- record --device 0 --out ..\..\output\note.wav
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --model SmallEn --format-provider auto
# Optional: use Semantic Kernel + Hugging Face instead of Ollama
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --format-provider huggingface --hf-endpoint https://router.huggingface.co --hf-model openai/gpt-oss-20b:groq
# Optional: enable speaker labels quickly
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --format-provider local --speakers 2
dotnet run -- record-and-transcribe --device 0 --wav ..\..\output\meeting.wav --out ..\..\output\meeting.md --model SmallEn
cd src\LocalTranscriber.Web
dotnet run
Then open the app URL (for example http://localhost:5078) and use the recording panel.
Run In Browser (Experimental) when the capability badge reports support.WebGPU, AudioContext, and MediaRecorder.%LOCALAPPDATA%\LocalTranscriber\modelsSmallEn is a very good default.MediumEn / LargeV3 are more accurate but heavier.--max-seg-length <int> to cap segment length (legacy --max-seg-seconds is still accepted).--format-provider auto|local|ollama|huggingface instead of juggling multiple boolean switches.--speakers <int> as a shortcut for speaker labeling.--speaker-sensitivity 0-100 (low = conservative, high = aggressive)--speaker-min-score-gain--speaker-max-switch-rate--speaker-min-separation--speaker-min-cluster-size--speaker-max-auto--speaker-global-variance-gate--speaker-short-run-merge-seconds--format-provider auto tries Ollama first, then Hugging Face (if HF_TOKEN / --hf-api-key is configured), then local formatting.--format-provider ollama or --format-provider huggingface force a provider, with local fallback if unavailable.--format-sensitivity 0-100--format-strict-transcript true|false--format-overlap-threshold--format-summary-min / --format-summary-max--format-include-action-items true|false--format-temperature / --format-max-tokens--format-local-big-gap / --format-local-small-gapaudio checksum + normalized transcription/format/speaker settings.src/LocalTranscriber.Web/output/cache.Run LocalTranscriber anywhere with Docker:
# Start web UI with Ollama
docker compose up web
# Or use GPU acceleration (requires NVIDIA Docker runtime)
docker compose --profile cuda up web-cuda
# Transcribe a file
docker compose run --rm cli transcribe --in /app/input/meeting.wav --out /app/output/meeting.md
# With GPU
docker compose --profile cli-cuda run --rm cli-cuda transcribe --in /app/input/meeting.wav --out /app/output/meeting.md
| Image | Description |
|---|---|
local-transcriber:web | Web UI, CPU-only |
local-transcriber:web-cuda | Web UI with CUDA GPU support |
local-transcriber:cli | CLI, CPU-only |
local-transcriber:cli-cuda | CLI with CUDA GPU support |
./models → /app/models — Whisper model cache./output → /app/output — Transcription output./input → /app/input — Input files (CLI only)# CPU version
docker build --target web -t local-transcriber:web .
# CUDA version
docker build --target web --build-arg VARIANT=cuda -t local-transcriber:web-cuda .
Whisper.net is a .NET wrapper around whisper.cpp and supports multiple runtimes (CPU, CUDA, etc.) and a built-in GGML model downloader.
This repo includes baseline GitHub project automation:
.github/workflows/ci.yml.github/workflows/dependency-review.yml.github/workflows/pr-labeler.yml + .github/labeler.yml.github/workflows/label-sync.yml + .github/labels.yml.github/workflows/release.yml.github/dependabot.yml.github/Project governance/docs:
CONTRIBUTING.mdCODE_OF_CONDUCT.mdSECURITY.mdSUPPORT.mdCHANGELOG.mdLICENSE.github/ISSUE_TEMPLATE/config.yml discussion URL (<owner>/<repo> placeholder).CODEOWNERS with real usernames/teams.main, then run Label Sync workflow once to seed labels.v0.1.0) to validate the release workflow.JavaScript
35.0%
HTML
30.1%
C#
20.4%
CSS
13.5%
A 100% local audio transcription, editing, and review tool
See the codeA small local voice-to-text starter that can:
Whisper.net (GPU via CUDA runtime if available)ollama pull llama3.2HF_TOKEN$env:HF_TOKEN = "hf_xxx"From the repo root:
cd src\LocalTranscriber.Cli
dotnet restore
dotnet run -- devices
dotnet run -- record --device 0 --out ..\..\output\note.wav
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --model SmallEn --format-provider auto
# Optional: use Semantic Kernel + Hugging Face instead of Ollama
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --format-provider huggingface --hf-endpoint https://router.huggingface.co --hf-model openai/gpt-oss-20b:groq
# Optional: enable speaker labels quickly
dotnet run -- transcribe --in ..\..\output\note.wav --out ..\..\output\note.md --format-provider local --speakers 2
dotnet run -- record-and-transcribe --device 0 --wav ..\..\output\meeting.wav --out ..\..\output\meeting.md --model SmallEn
cd src\LocalTranscriber.Web
dotnet run
Then open the app URL (for example http://localhost:5078) and use the recording panel.
Run In Browser (Experimental) when the capability badge reports support.WebGPU, AudioContext, and MediaRecorder.%LOCALAPPDATA%\LocalTranscriber\modelsSmallEn is a very good default.MediumEn / LargeV3 are more accurate but heavier.--max-seg-length <int> to cap segment length (legacy --max-seg-seconds is still accepted).--format-provider auto|local|ollama|huggingface instead of juggling multiple boolean switches.--speakers <int> as a shortcut for speaker labeling.--speaker-sensitivity 0-100 (low = conservative, high = aggressive)--speaker-min-score-gain--speaker-max-switch-rate--speaker-min-separation--speaker-min-cluster-size--speaker-max-auto--speaker-global-variance-gate--speaker-short-run-merge-seconds--format-provider auto tries Ollama first, then Hugging Face (if HF_TOKEN / --hf-api-key is configured), then local formatting.--format-provider ollama or --format-provider huggingface force a provider, with local fallback if unavailable.--format-sensitivity 0-100--format-strict-transcript true|false--format-overlap-threshold--format-summary-min / --format-summary-max--format-include-action-items true|false--format-temperature / --format-max-tokens--format-local-big-gap / --format-local-small-gapaudio checksum + normalized transcription/format/speaker settings.src/LocalTranscriber.Web/output/cache.Run LocalTranscriber anywhere with Docker:
# Start web UI with Ollama
docker compose up web
# Or use GPU acceleration (requires NVIDIA Docker runtime)
docker compose --profile cuda up web-cuda
# Transcribe a file
docker compose run --rm cli transcribe --in /app/input/meeting.wav --out /app/output/meeting.md
# With GPU
docker compose --profile cli-cuda run --rm cli-cuda transcribe --in /app/input/meeting.wav --out /app/output/meeting.md
| Image | Description |
|---|---|
local-transcriber:web | Web UI, CPU-only |
local-transcriber:web-cuda | Web UI with CUDA GPU support |
local-transcriber:cli | CLI, CPU-only |
local-transcriber:cli-cuda | CLI with CUDA GPU support |
./models → /app/models — Whisper model cache./output → /app/output — Transcription output./input → /app/input — Input files (CLI only)# CPU version
docker build --target web -t local-transcriber:web .
# CUDA version
docker build --target web --build-arg VARIANT=cuda -t local-transcriber:web-cuda .
Whisper.net is a .NET wrapper around whisper.cpp and supports multiple runtimes (CPU, CUDA, etc.) and a built-in GGML model downloader.
This repo includes baseline GitHub project automation:
.github/workflows/ci.yml.github/workflows/dependency-review.yml.github/workflows/pr-labeler.yml + .github/labeler.yml.github/workflows/label-sync.yml + .github/labels.yml.github/workflows/release.yml.github/dependabot.yml.github/Project governance/docs:
CONTRIBUTING.mdCODE_OF_CONDUCT.mdSECURITY.mdSUPPORT.mdCHANGELOG.mdLICENSE.github/ISSUE_TEMPLATE/config.yml discussion URL (<owner>/<repo> placeholder).CODEOWNERS with real usernames/teams.main, then run Label Sync workflow once to seed labels.v0.1.0) to validate the release workflow.JavaScript
35.0%
HTML
30.1%
C#
20.4%
CSS
13.5%