Transcribe Audio files to text with AI, and make searchable via AI.
C#
0
53 commits
updated Jun 7, 2026
The Enterprise-Grade Audio Transcription Solution You Didn't Know You Needed™
A fully local, privacy-first desktop application that transcribes audio files with OpenAI Whisper, stores searchable vector embeddings, and answers questions about your recordings using an on-device LLM — no cloud, no subscriptions, no "we updated our privacy policy" emails.
all-MiniLM-L6-v2 embedding model and cosine similarity. Ask questions. Get answers. Revolutionary.| Component | Minimum | Recommended |
|---|---|---|
| OS | Windows 10, Ubuntu 20.04, macOS 12 | Latest stable release |
| CPU | x64 with AVX2 | Modern multi-core (AVX2 required for Whisper) |
| RAM | 8 GB | 16 GB (LLM uses ~2 GB) |
| Storage | 500 MB (app) | 4 GB (all models: Whisper small + ONNX + Phi-3 Q4) |
| Runtime | .NET 10 | .NET 10 |
| Linux extra | ffmpeg on $PATH | ffmpeg 6+ |
Download the installer from the Releases page. Double-click. Follow prompts. We've made it simple enough that even management can do it.
wget https://github.com/davidpizon/Phonematic/releases/latest/download/Phonematic-linux-x64.tar.gz
tar -xzf Phonematic-linux-x64.tar.gz
./Phonematic
Install ffmpeg if you haven't already: sudo apt install ffmpeg
You're using Linux. We trust you can figure out the rest.
For those who trust no one (we respect that):
git clone https://github.com/davidpizon/Phonematic.git
cd Phonematic
dotnet build src/Phonematic.slnx -c Release
dotnet run --project src/Phonematic.Gui/Phonematic.Gui.csproj
Requires .NET 10 SDK. Yes, we're living in the future.
Prefer the terminal? There's a headless CLI too — point it at an audio file or a folder and it
writes PhoScript (.phos) files:
dotnet run --project src/Phonematic/Phonematic.csproj -- ./recordings --recursive
See docs/CLI.md for all flags and exit codes.
On first launch Phonematic will automatically download the three AI models it needs:
| Model | Size | Purpose |
|---|---|---|
Whisper tiny.en (default) | ~75 MB | Speech-to-text |
all-MiniLM-L6-v2 (ONNX) | ~23 MB | Sentence embeddings |
| Phi-3 Mini 4k Q4 (GGUF) | ~2.2 GB | On-device LLM for Q&A |
Downloads go to %LOCALAPPDATA%\Phonematic\models\. You can switch to a larger Whisper model (e.g. small.en, medium.en) in Settings at any time.
.txt files to your configured Output Directory (default: ~/Documents/Phonematic/).Browse and read all previously transcribed files. Select any entry to view its full timestamped transcript.
http://localhost:27839/plaud-token.Phonematic also ships a headless audio → PhoScript (.phos) converter for scripts and automation.
Examples below use the Phonematic command; from source, substitute
dotnet run --project src/Phonematic/Phonematic.csproj -- <args>. Result paths print to stdout;
progress and logs go to stderr. Full flag/exit-code reference: docs/CLI.md.
# Default: writes song.phos next to the source
Phonematic song.mp3
# Explicit output path
Phonematic song.mp3 -o transcripts/song.phos
# Every supported audio file in the folder (top level only), .phos next to each source
Phonematic ./recordings
# Recurse into subfolders, mirror the tree into ./out, overwrite existing targets
Phonematic ./recordings --recursive --output-dir ./out --overwrite
# Quiet mode — capture just the written paths
Phonematic ./recordings -q > written.txt
Supply the exact words spoken; the phones are aligned to them so <word orth="…"> matches your text.
# Single file + its transcript
Phonematic interview.mp3 --transcript interview.txt -o interview.phos
# Folder: each audio file is paired with its sibling <name>.txt automatically
Phonematic ./recordings --recursive
Whisper supplies the words; wav2vec2 forced-aligns the phones.
# Use the default (config) Whisper model size
Phonematic lecture.mp3 --whisper
# Pick a larger Whisper model for the whole folder
Phonematic ./recordings --whisper --whisper-model small
# Improve recognition of a known speaker's new, transcript-less audio
Phonematic new-recording.mp3 --voice-model models/speaker-A.phonematic -o out.phos
# Batch a folder through the same speaker model
Phonematic ./recordings -r --voice-model models/speaker-A.phonematic --output-dir ./out
train)Learns a portable .phonematic model from (audio, sibling-<name>.txt) pairs.
# Train from a folder of recordings + matching transcripts
Phonematic train ./speaker-A --output models/speaker-A.phonematic --recursive
# Different speaker, more epochs, an explicitly chosen base model
Phonematic train ./speaker-B -o models/speaker-B.phonematic --epochs 80 --base-model wav2vec2-phoneme
models)The only commands that download. Conversion/training otherwise exit 3 if a model is missing.
# Download the default base model
Phonematic models download
# Download a base model AND a Whisper model for hybrid mode
Phonematic models download --whisper --whisper-model small
# See which models are present on disk
Phonematic models status
┌─────────────────────────────────────────────────────────────┐
│ Avalonia UI (Views) │
├─────────────────────────────────────────────────────────────┤
│ ViewModels (MVVM / CommunityToolkit) │
├──────────────┬──────────────┬──────────────┬────────────────┤
│ Transcription│ Embedding & │ PLAUD API │ Config & │
│ Service │ Vector Search│ Service │ Model Manager │
│ (Whisper.net)│ (ONNX + LLM)│ (HttpClient)│ Service │
├──────────────┴──────────────┴──────────────┴────────────────┤
│ EF Core + SQLite (PhonematicDbContext) │
├─────────────────────────────────────────────────────────────┤
│ %LOCALAPPDATA%\Phonematic\ (models, DB, logs) │
└─────────────────────────────────────────────────────────────┘
See docs/ARCHITECTURE.md for a detailed description of all layers, data flows, and the database schema.
| Document | Description |
|---|---|
| docs/ARCHITECTURE.md | System architecture, data flows, DB schema, file layout |
| docs/API.md | Full reference for all classes, interfaces, and records |
| docs/CLI.md | Command-line interface: synopsis, flags, examples, and exit codes |
| docs/CONTRIBUTING.md | Development workflow, coding standards, PR checklist |
| docs/TESTING.md | Test suite structure, patterns, and how to run tests |
| docs/AGENTS.md | Guidelines for AI coding agents working on this repo |
| docs/PHOSCRIPT.md | PhoScript 1.0 specification — prosodic markup format for .phos files |
| docs/IPA_REFERENCE.md | International Phonetic Alphabet reference — how sounds are represented in IPA |
Q: Why not just use [Cloud Service X]?
A: Because you read the terms of service, didn't you? ...You didn't? We recommend reading them. With a lawyer present.
Q: How accurate is the transcription?
A: Depends on your audio quality and chosen model. The default tiny.en is fast; small.en or medium.en gives noticeably better results for unclear audio. Garbage in, garbage out — this is a universal constant.
Q: Can I transcribe in languages other than English?
A: Yes. Switch to a non-.en model (e.g. small) in Settings. Whisper supports ~100 languages. Your mileage may vary based on your accent's deviation from the training data.
Q: Where is my data stored?
A: Everything — database, models, config, logs — lives under %LOCALAPPDATA%\Phonematic\ on your own machine. Transcripts go to ~/Documents/Phonematic/ by default (configurable).
Q: Can I use a different LLM?
A: The LLM path is managed by ModelManagerService.GetLlmModelPath(). Swap in any GGUF compatible with LLamaSharp and point llm/ to it.
See LICENSE for details.
Q: Why is the first transcription slow?
A: The AI models need to be loaded into memory. Subsequent transcriptions are faster. Patience is a virtue.
Q: My transcription has errors.
A: See "Garbage in, garbage out" above. Also, AI is not perfect. Neither are humans. We're all doing our best here.
Pull requests are welcome. Please ensure your code:
MIT License. Use it, modify it, sell it, tattoo it on your forearm. We don't care. See LICENSE for the legal boilerplate.
Phonematic: Because your audio files aren't going to transcribe themselves.
Built with caffeine and mass quantities of reasonable expectations
C#
99.1%
Transcribe Audio files to text with AI, and make searchable via AI.
C#
0
53 commits
updated Jun 7, 2026
The Enterprise-Grade Audio Transcription Solution You Didn't Know You Needed™
A fully local, privacy-first desktop application that transcribes audio files with OpenAI Whisper, stores searchable vector embeddings, and answers questions about your recordings using an on-device LLM — no cloud, no subscriptions, no "we updated our privacy policy" emails.
all-MiniLM-L6-v2 embedding model and cosine similarity. Ask questions. Get answers. Revolutionary.| Component | Minimum | Recommended |
|---|---|---|
| OS | Windows 10, Ubuntu 20.04, macOS 12 | Latest stable release |
| CPU | x64 with AVX2 | Modern multi-core (AVX2 required for Whisper) |
| RAM | 8 GB | 16 GB (LLM uses ~2 GB) |
| Storage | 500 MB (app) | 4 GB (all models: Whisper small + ONNX + Phi-3 Q4) |
| Runtime | .NET 10 | .NET 10 |
| Linux extra | ffmpeg on $PATH | ffmpeg 6+ |
Download the installer from the Releases page. Double-click. Follow prompts. We've made it simple enough that even management can do it.
wget https://github.com/davidpizon/Phonematic/releases/latest/download/Phonematic-linux-x64.tar.gz
tar -xzf Phonematic-linux-x64.tar.gz
./Phonematic
Install ffmpeg if you haven't already: sudo apt install ffmpeg
You're using Linux. We trust you can figure out the rest.
For those who trust no one (we respect that):
git clone https://github.com/davidpizon/Phonematic.git
cd Phonematic
dotnet build src/Phonematic.slnx -c Release
dotnet run --project src/Phonematic.Gui/Phonematic.Gui.csproj
Requires .NET 10 SDK. Yes, we're living in the future.
Prefer the terminal? There's a headless CLI too — point it at an audio file or a folder and it
writes PhoScript (.phos) files:
dotnet run --project src/Phonematic/Phonematic.csproj -- ./recordings --recursive
See docs/CLI.md for all flags and exit codes.
On first launch Phonematic will automatically download the three AI models it needs:
| Model | Size | Purpose |
|---|---|---|
Whisper tiny.en (default) | ~75 MB | Speech-to-text |
all-MiniLM-L6-v2 (ONNX) | ~23 MB | Sentence embeddings |
| Phi-3 Mini 4k Q4 (GGUF) | ~2.2 GB | On-device LLM for Q&A |
Downloads go to %LOCALAPPDATA%\Phonematic\models\. You can switch to a larger Whisper model (e.g. small.en, medium.en) in Settings at any time.
.txt files to your configured Output Directory (default: ~/Documents/Phonematic/).Browse and read all previously transcribed files. Select any entry to view its full timestamped transcript.
http://localhost:27839/plaud-token.Phonematic also ships a headless audio → PhoScript (.phos) converter for scripts and automation.
Examples below use the Phonematic command; from source, substitute
dotnet run --project src/Phonematic/Phonematic.csproj -- <args>. Result paths print to stdout;
progress and logs go to stderr. Full flag/exit-code reference: docs/CLI.md.
# Default: writes song.phos next to the source
Phonematic song.mp3
# Explicit output path
Phonematic song.mp3 -o transcripts/song.phos
# Every supported audio file in the folder (top level only), .phos next to each source
Phonematic ./recordings
# Recurse into subfolders, mirror the tree into ./out, overwrite existing targets
Phonematic ./recordings --recursive --output-dir ./out --overwrite
# Quiet mode — capture just the written paths
Phonematic ./recordings -q > written.txt
Supply the exact words spoken; the phones are aligned to them so <word orth="…"> matches your text.
# Single file + its transcript
Phonematic interview.mp3 --transcript interview.txt -o interview.phos
# Folder: each audio file is paired with its sibling <name>.txt automatically
Phonematic ./recordings --recursive
Whisper supplies the words; wav2vec2 forced-aligns the phones.
# Use the default (config) Whisper model size
Phonematic lecture.mp3 --whisper
# Pick a larger Whisper model for the whole folder
Phonematic ./recordings --whisper --whisper-model small
# Improve recognition of a known speaker's new, transcript-less audio
Phonematic new-recording.mp3 --voice-model models/speaker-A.phonematic -o out.phos
# Batch a folder through the same speaker model
Phonematic ./recordings -r --voice-model models/speaker-A.phonematic --output-dir ./out
train)Learns a portable .phonematic model from (audio, sibling-<name>.txt) pairs.
# Train from a folder of recordings + matching transcripts
Phonematic train ./speaker-A --output models/speaker-A.phonematic --recursive
# Different speaker, more epochs, an explicitly chosen base model
Phonematic train ./speaker-B -o models/speaker-B.phonematic --epochs 80 --base-model wav2vec2-phoneme
models)The only commands that download. Conversion/training otherwise exit 3 if a model is missing.
# Download the default base model
Phonematic models download
# Download a base model AND a Whisper model for hybrid mode
Phonematic models download --whisper --whisper-model small
# See which models are present on disk
Phonematic models status
┌─────────────────────────────────────────────────────────────┐
│ Avalonia UI (Views) │
├─────────────────────────────────────────────────────────────┤
│ ViewModels (MVVM / CommunityToolkit) │
├──────────────┬──────────────┬──────────────┬────────────────┤
│ Transcription│ Embedding & │ PLAUD API │ Config & │
│ Service │ Vector Search│ Service │ Model Manager │
│ (Whisper.net)│ (ONNX + LLM)│ (HttpClient)│ Service │
├──────────────┴──────────────┴──────────────┴────────────────┤
│ EF Core + SQLite (PhonematicDbContext) │
├─────────────────────────────────────────────────────────────┤
│ %LOCALAPPDATA%\Phonematic\ (models, DB, logs) │
└─────────────────────────────────────────────────────────────┘
See docs/ARCHITECTURE.md for a detailed description of all layers, data flows, and the database schema.
| Document | Description |
|---|---|
| docs/ARCHITECTURE.md | System architecture, data flows, DB schema, file layout |
| docs/API.md | Full reference for all classes, interfaces, and records |
| docs/CLI.md | Command-line interface: synopsis, flags, examples, and exit codes |
| docs/CONTRIBUTING.md | Development workflow, coding standards, PR checklist |
| docs/TESTING.md | Test suite structure, patterns, and how to run tests |
| docs/AGENTS.md | Guidelines for AI coding agents working on this repo |
| docs/PHOSCRIPT.md | PhoScript 1.0 specification — prosodic markup format for .phos files |
| docs/IPA_REFERENCE.md | International Phonetic Alphabet reference — how sounds are represented in IPA |
Q: Why not just use [Cloud Service X]?
A: Because you read the terms of service, didn't you? ...You didn't? We recommend reading them. With a lawyer present.
Q: How accurate is the transcription?
A: Depends on your audio quality and chosen model. The default tiny.en is fast; small.en or medium.en gives noticeably better results for unclear audio. Garbage in, garbage out — this is a universal constant.
Q: Can I transcribe in languages other than English?
A: Yes. Switch to a non-.en model (e.g. small) in Settings. Whisper supports ~100 languages. Your mileage may vary based on your accent's deviation from the training data.
Q: Where is my data stored?
A: Everything — database, models, config, logs — lives under %LOCALAPPDATA%\Phonematic\ on your own machine. Transcripts go to ~/Documents/Phonematic/ by default (configurable).
Q: Can I use a different LLM?
A: The LLM path is managed by ModelManagerService.GetLlmModelPath(). Swap in any GGUF compatible with LLamaSharp and point llm/ to it.
See LICENSE for details.
Q: Why is the first transcription slow?
A: The AI models need to be loaded into memory. Subsequent transcriptions are faster. Patience is a virtue.
Q: My transcription has errors.
A: See "Garbage in, garbage out" above. Also, AI is not perfect. Neither are humans. We're all doing our best here.
Pull requests are welcome. Please ensure your code:
MIT License. Use it, modify it, sell it, tattoo it on your forearm. We don't care. See LICENSE for the legal boilerplate.
Phonematic: Because your audio files aren't going to transcribe themselves.
Built with caffeine and mass quantities of reasonable expectations
C#
99.1%