Local Whisper speech-to-text in .NET with ONNX Runtime. Auto-downloads models from HuggingFace.
C#
10
35 commits
updated Aug 3, 2026
Transcribe audio to text in .NET using OpenAI's Whisper model. Powered by ONNX Runtime with automatic model download from HuggingFace.
| Package | NuGet | Downloads | Description |
|---|---|---|---|
ElBruno.Whisper | Core speech-to-text library with Whisper ONNX models | ||
ElBruno.Whisper.BlazorComponents | Reusable Blazor components for transcription workflows |
AddWhisper() in ASP.NET CoreISpeechToTextClient for standard speech-to-text integrationGetStreamingTextAsync().en variants for best accuracy on English audiodotnet add package ElBruno.Whisper
using ElBruno.Whisper;
// Create client (downloads tiny.en model on first run)
using var client = await WhisperClient.CreateAsync();
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine(result.Text);
The first time you create a WhisperClient, the model is downloaded from HuggingFace to your local cache directory (~75 MB - 3 GB depending on model size). This typically takes 10-60 seconds depending on your internet connection and chosen model.
Track download progress:
using var client = await WhisperClient.CreateAsync(
progress: new Progress<ElBruno.HuggingFace.DownloadProgress>(p =>
{
if (p.Stage == ElBruno.HuggingFace.DownloadStage.Downloading)
Console.WriteLine($"{p.CurrentFile}: {p.PercentComplete:F0}%");
else
Console.WriteLine($"{p.Stage}: {p.Message}");
})
);
Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/Whisper/models).
Whisper offers various model sizes. English-optimized models (.en suffix) are smaller and faster for English audio:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
Model = KnownWhisperModels.WhisperSmallEn
});
var result = await client.TranscribeAsync("english-audio.wav");
Console.WriteLine(result.Text);
| Size | English | Multilingual | Parameters | Approx Size | Speed |
|---|---|---|---|---|---|
| tiny | tiny.en | tiny | 39M | 75 MB | β‘β‘β‘β‘β‘ |
| base | base.en | base | 74M | 140 MB | β‘β‘β‘β‘ |
| small | small.en | small | 244M | 460 MB | β‘β‘β‘ |
| medium | medium.en | medium | 769M | 1.5 GB | β‘β‘ |
| large | β | large | 1550M | 3.0 GB | β‘ |
Use English-optimized (.en) models for:
Use Multilingual models for:
Monitor both file downloads and transcription progress:
var downloadProgress = new Progress<ElBruno.HuggingFace.DownloadProgress>(p =>
{
if (p.Stage == ElBruno.HuggingFace.DownloadStage.Downloading)
Console.Write($"\rβ¬οΈ {p.PercentComplete:F0}%");
else
Console.WriteLine($"\nβ {p.Message}");
});
using var client = await WhisperClient.CreateAsync(progress: downloadProgress);
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine($"β Transcribed: {result.Text}");
Register Whisper options in ASP.NET Core or other DI-enabled applications:
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperBaseEn;
options.Concurrency.MaximumConcurrentRequests = 2;
});
AddWhisper() registers WhisperOptions, WhisperSpeechToTextClient, and ISpeechToTextClient. Create and share a WhisperClient yourself when you want direct control over startup and disposal.
Use the adapter when you want the standard ISpeechToTextClient contract:
using ElBruno.Whisper;
using Microsoft.Extensions.AI;
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperTinyEn;
options.Language = "en";
});
var speechToText = builder.Services
.BuildServiceProvider()
.GetRequiredService<ISpeechToTextClient>();
await using var audioStream = File.OpenRead("audio.wav");
var response = await speechToText.GetTextAsync(
audioStream,
new SpeechToTextOptions
{
SpeechLanguage = "en",
AdditionalProperties = new()
{
["elbruno.whisper.enable_timestamps"] = true
}
});
Console.WriteLine(response.Text);
Console.WriteLine(response.AdditionalProperties?["elbruno.whisper.detected_language"]);
The adapter keeps caller-owned streams open, supports cancellation, exposes SpeechToTextClientMetadata, and returns these response metadata keys through AdditionalProperties:
elbruno.whisper.detected_languageelbruno.whisper.audio_duration_mselbruno.whisper.segmentselbruno.whisper.wordselbruno.whisper.model_idelbruno.whisper.execution_providerWhisperClient can now be shared across concurrent callers. By default it still processes one transcription at a time. Raise the concurrency limit to allow parallel work and reuse pooled ONNX sessions:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
Model = KnownWhisperModels.WhisperTinyEn,
Concurrency = new WhisperConcurrencyOptions
{
MaximumConcurrentRequests = 2,
QueueTimeout = TimeSpan.FromSeconds(15),
EnableSessionPooling = true
}
});
var results = await Task.WhenAll(
client.TranscribeAsync("audio-1.wav"),
client.TranscribeAsync("audio-2.wav"));
If all inference slots are busy longer than QueueTimeout, TranscribeAsync throws TimeoutException. Cancelling the request also aborts queue waiting and the next safe decode checkpoint.
WhisperClient.GetStreamingTextAsync() runs Whisper over rolling windows and emits ordered updates with both stable and provisional text:
using var client = await WhisperClient.CreateAsync();
await foreach (var update in client.GetStreamingTextAsync(
"audio.wav",
new WhisperStreamingOptions
{
WindowSize = TimeSpan.FromSeconds(8),
StepSize = TimeSpan.FromSeconds(1),
ContextOverlap = TimeSpan.FromSeconds(2),
UseLocalAgreement = true,
AgreementIterations = 2
}))
{
Console.WriteLine($"Committed: {update.CommittedText}");
Console.WriteLine($"Provisional: {update.ProvisionalText}");
if (update.IsFinal)
Console.WriteLine("Final update received.");
}
Each update exposes:
CommittedText β text that has stabilized across rolling windowsProvisionalText β the newest hypothesis that may still changeText β the combined transcript for the current updateIsFinal β true exactly once, after the final flushLimitations: Whisper is not a native streaming model. This API reads completed file or stream content, re-runs inference over overlapping windows, and uses local agreement to reduce duplicate committed text. Provisional text can still change between updates.
Use the explicit-audio overloads when your pipeline already has PCM in memory and you want to avoid temporary WAV files:
using var client = await WhisperClient.CreateAsync();
var pcm16Format = new WhisperAudioFormat(
sampleRate: 48000,
channels: 2,
sampleFormat: WhisperAudioSampleFormat.Pcm16);
await using var rawAudioStream = File.OpenRead("call.raw");
var streamResult = await client.TranscribeAsync(rawAudioStream, pcm16Format);
ReadOnlyMemory<byte> pcmBytes = await File.ReadAllBytesAsync("call.raw");
var byteResult = await client.TranscribeAsync(pcmBytes, pcm16Format);
ReadOnlyMemory<float> monoFloatSamples = GetNormalizedSamples();
var floatResult = await client.TranscribeAsync(monoFloatSamples, sampleRate: 16000);
Notes:
TranscribeAsync(Stream) auto-detects WAV headers and keeps the caller-owned stream open.WhisperAudioFormat so the client can downmix and resample to Whisper's 16 kHz mono input.ReadOnlyMemory<float> overloads expect normalized PCM samples in the [-1, 1] range.Build speech-to-text interfaces quickly with ElBruno.Whisper.BlazorComponents:
dotnet add package ElBruno.Whisper.BlazorComponents
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperBaseEn;
});
builder.Services.AddWhisperBlazorComponents();
Available components:
WhisperModelSelectorFileTranscriptionPanelLiveTranscriptViewerTranscriptionHistoryListThe TranscriptionResult includes:
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine(result.Text); // Transcribed text
Console.WriteLine(result.DetectedLanguage); // Detected language (for multilingual models)
Console.WriteLine(result.Duration); // Audio duration
Enable timestamps to access both segment-level and word-level timing metadata:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
EnableTimestamps = true
});
var result = await client.TranscribeAsync("audio.wav");
foreach (var segment in result.Segments ?? [])
{
Console.WriteLine($"[{segment.Start:mm\\:ss\\.ff} - {segment.End:mm\\:ss\\.ff}] {segment.Text}");
foreach (var word in segment.Words)
{
Console.WriteLine($" {word.Start:mm\\:ss\\.ff} - {word.End:mm\\:ss\\.ff}: {word.Text}");
}
}
Word timings are derived from the timestamped transcript spans produced by Whisper. If a model returns text without explicit spans, the library falls back to a single full-duration segment and derives word timings within that range.
Model download fails?
HF_TOKEN environment variableOut of memory?
For detailed troubleshooting, see docs.
| Sample | Description |
|---|---|
| HelloWhisper | Minimal console transcription |
| BlazorWhisper | Blazor app with audio recording and real-time transcription |
| Blazor Components demo page | Demonstrates the reusable Blazor components package |
ElBruno.Whisper.BlazorComponents with model selector, file transcription panel, live transcript viewer, and transcription history list.BlazorWhisper sample.TranscriptionResult.NuGet/login.git clone https://github.com/elbruno/ElBruno.Whisper
cd ElBruno.Whisper
dotnet build ElBruno.Whisper.slnx
dotnet test ElBruno.Whisper.slnx --filter "Category!=Integration"
The repository includes comprehensive unit and integration tests:
Quick test run (unit tests, no model download):
dotnet test ElBruno.Whisper.slnx --filter "Category!=Integration"
Full test run (includes integration with real models):
dotnet test ElBruno.Whisper.slnx
Test audio files are provided in testdata/audio/ for validation and transcription testing. For details, see the Testing Guide.
Contributions are welcome! Please:
git checkout -b feature/amazing-feature)git commit -m 'Add amazing feature')git push origin feature/amazing-feature)This project is licensed under the MIT License β see the LICENSE file for details.
Hi! I'm ElBruno π§‘, a passionate developer and content creator exploring AI, .NET, and modern development practices.
Made with β€οΈ by ElBruno
If you like this project, consider following my work across platforms:
Local Whisper speech-to-text in .NET with ONNX Runtime. Auto-downloads models from HuggingFace.
C#
10
35 commits
updated Aug 3, 2026
Transcribe audio to text in .NET using OpenAI's Whisper model. Powered by ONNX Runtime with automatic model download from HuggingFace.
| Package | NuGet | Downloads | Description |
|---|---|---|---|
ElBruno.Whisper | Core speech-to-text library with Whisper ONNX models | ||
ElBruno.Whisper.BlazorComponents | Reusable Blazor components for transcription workflows |
AddWhisper() in ASP.NET CoreISpeechToTextClient for standard speech-to-text integrationGetStreamingTextAsync().en variants for best accuracy on English audiodotnet add package ElBruno.Whisper
using ElBruno.Whisper;
// Create client (downloads tiny.en model on first run)
using var client = await WhisperClient.CreateAsync();
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine(result.Text);
The first time you create a WhisperClient, the model is downloaded from HuggingFace to your local cache directory (~75 MB - 3 GB depending on model size). This typically takes 10-60 seconds depending on your internet connection and chosen model.
Track download progress:
using var client = await WhisperClient.CreateAsync(
progress: new Progress<ElBruno.HuggingFace.DownloadProgress>(p =>
{
if (p.Stage == ElBruno.HuggingFace.DownloadStage.Downloading)
Console.WriteLine($"{p.CurrentFile}: {p.PercentComplete:F0}%");
else
Console.WriteLine($"{p.Stage}: {p.Message}");
})
);
Subsequent runs load instantly from cache (%LOCALAPPDATA%/ElBruno/Whisper/models).
Whisper offers various model sizes. English-optimized models (.en suffix) are smaller and faster for English audio:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
Model = KnownWhisperModels.WhisperSmallEn
});
var result = await client.TranscribeAsync("english-audio.wav");
Console.WriteLine(result.Text);
| Size | English | Multilingual | Parameters | Approx Size | Speed |
|---|---|---|---|---|---|
| tiny | tiny.en | tiny | 39M | 75 MB | β‘β‘β‘β‘β‘ |
| base | base.en | base | 74M | 140 MB | β‘β‘β‘β‘ |
| small | small.en | small | 244M | 460 MB | β‘β‘β‘ |
| medium | medium.en | medium | 769M | 1.5 GB | β‘β‘ |
| large | β | large | 1550M | 3.0 GB | β‘ |
Use English-optimized (.en) models for:
Use Multilingual models for:
Monitor both file downloads and transcription progress:
var downloadProgress = new Progress<ElBruno.HuggingFace.DownloadProgress>(p =>
{
if (p.Stage == ElBruno.HuggingFace.DownloadStage.Downloading)
Console.Write($"\rβ¬οΈ {p.PercentComplete:F0}%");
else
Console.WriteLine($"\nβ {p.Message}");
});
using var client = await WhisperClient.CreateAsync(progress: downloadProgress);
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine($"β Transcribed: {result.Text}");
Register Whisper options in ASP.NET Core or other DI-enabled applications:
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperBaseEn;
options.Concurrency.MaximumConcurrentRequests = 2;
});
AddWhisper() registers WhisperOptions, WhisperSpeechToTextClient, and ISpeechToTextClient. Create and share a WhisperClient yourself when you want direct control over startup and disposal.
Use the adapter when you want the standard ISpeechToTextClient contract:
using ElBruno.Whisper;
using Microsoft.Extensions.AI;
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperTinyEn;
options.Language = "en";
});
var speechToText = builder.Services
.BuildServiceProvider()
.GetRequiredService<ISpeechToTextClient>();
await using var audioStream = File.OpenRead("audio.wav");
var response = await speechToText.GetTextAsync(
audioStream,
new SpeechToTextOptions
{
SpeechLanguage = "en",
AdditionalProperties = new()
{
["elbruno.whisper.enable_timestamps"] = true
}
});
Console.WriteLine(response.Text);
Console.WriteLine(response.AdditionalProperties?["elbruno.whisper.detected_language"]);
The adapter keeps caller-owned streams open, supports cancellation, exposes SpeechToTextClientMetadata, and returns these response metadata keys through AdditionalProperties:
elbruno.whisper.detected_languageelbruno.whisper.audio_duration_mselbruno.whisper.segmentselbruno.whisper.wordselbruno.whisper.model_idelbruno.whisper.execution_providerWhisperClient can now be shared across concurrent callers. By default it still processes one transcription at a time. Raise the concurrency limit to allow parallel work and reuse pooled ONNX sessions:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
Model = KnownWhisperModels.WhisperTinyEn,
Concurrency = new WhisperConcurrencyOptions
{
MaximumConcurrentRequests = 2,
QueueTimeout = TimeSpan.FromSeconds(15),
EnableSessionPooling = true
}
});
var results = await Task.WhenAll(
client.TranscribeAsync("audio-1.wav"),
client.TranscribeAsync("audio-2.wav"));
If all inference slots are busy longer than QueueTimeout, TranscribeAsync throws TimeoutException. Cancelling the request also aborts queue waiting and the next safe decode checkpoint.
WhisperClient.GetStreamingTextAsync() runs Whisper over rolling windows and emits ordered updates with both stable and provisional text:
using var client = await WhisperClient.CreateAsync();
await foreach (var update in client.GetStreamingTextAsync(
"audio.wav",
new WhisperStreamingOptions
{
WindowSize = TimeSpan.FromSeconds(8),
StepSize = TimeSpan.FromSeconds(1),
ContextOverlap = TimeSpan.FromSeconds(2),
UseLocalAgreement = true,
AgreementIterations = 2
}))
{
Console.WriteLine($"Committed: {update.CommittedText}");
Console.WriteLine($"Provisional: {update.ProvisionalText}");
if (update.IsFinal)
Console.WriteLine("Final update received.");
}
Each update exposes:
CommittedText β text that has stabilized across rolling windowsProvisionalText β the newest hypothesis that may still changeText β the combined transcript for the current updateIsFinal β true exactly once, after the final flushLimitations: Whisper is not a native streaming model. This API reads completed file or stream content, re-runs inference over overlapping windows, and uses local agreement to reduce duplicate committed text. Provisional text can still change between updates.
Use the explicit-audio overloads when your pipeline already has PCM in memory and you want to avoid temporary WAV files:
using var client = await WhisperClient.CreateAsync();
var pcm16Format = new WhisperAudioFormat(
sampleRate: 48000,
channels: 2,
sampleFormat: WhisperAudioSampleFormat.Pcm16);
await using var rawAudioStream = File.OpenRead("call.raw");
var streamResult = await client.TranscribeAsync(rawAudioStream, pcm16Format);
ReadOnlyMemory<byte> pcmBytes = await File.ReadAllBytesAsync("call.raw");
var byteResult = await client.TranscribeAsync(pcmBytes, pcm16Format);
ReadOnlyMemory<float> monoFloatSamples = GetNormalizedSamples();
var floatResult = await client.TranscribeAsync(monoFloatSamples, sampleRate: 16000);
Notes:
TranscribeAsync(Stream) auto-detects WAV headers and keeps the caller-owned stream open.WhisperAudioFormat so the client can downmix and resample to Whisper's 16 kHz mono input.ReadOnlyMemory<float> overloads expect normalized PCM samples in the [-1, 1] range.Build speech-to-text interfaces quickly with ElBruno.Whisper.BlazorComponents:
dotnet add package ElBruno.Whisper.BlazorComponents
builder.Services.AddWhisper(options =>
{
options.Model = KnownWhisperModels.WhisperBaseEn;
});
builder.Services.AddWhisperBlazorComponents();
Available components:
WhisperModelSelectorFileTranscriptionPanelLiveTranscriptViewerTranscriptionHistoryListThe TranscriptionResult includes:
var result = await client.TranscribeAsync("audio.wav");
Console.WriteLine(result.Text); // Transcribed text
Console.WriteLine(result.DetectedLanguage); // Detected language (for multilingual models)
Console.WriteLine(result.Duration); // Audio duration
Enable timestamps to access both segment-level and word-level timing metadata:
using var client = await WhisperClient.CreateAsync(new WhisperOptions
{
EnableTimestamps = true
});
var result = await client.TranscribeAsync("audio.wav");
foreach (var segment in result.Segments ?? [])
{
Console.WriteLine($"[{segment.Start:mm\\:ss\\.ff} - {segment.End:mm\\:ss\\.ff}] {segment.Text}");
foreach (var word in segment.Words)
{
Console.WriteLine($" {word.Start:mm\\:ss\\.ff} - {word.End:mm\\:ss\\.ff}: {word.Text}");
}
}
Word timings are derived from the timestamped transcript spans produced by Whisper. If a model returns text without explicit spans, the library falls back to a single full-duration segment and derives word timings within that range.
Model download fails?
HF_TOKEN environment variableOut of memory?
For detailed troubleshooting, see docs.
| Sample | Description |
|---|---|
| HelloWhisper | Minimal console transcription |
| BlazorWhisper | Blazor app with audio recording and real-time transcription |
| Blazor Components demo page | Demonstrates the reusable Blazor components package |
ElBruno.Whisper.BlazorComponents with model selector, file transcription panel, live transcript viewer, and transcription history list.BlazorWhisper sample.TranscriptionResult.NuGet/login.git clone https://github.com/elbruno/ElBruno.Whisper
cd ElBruno.Whisper
dotnet build ElBruno.Whisper.slnx
dotnet test ElBruno.Whisper.slnx --filter "Category!=Integration"
The repository includes comprehensive unit and integration tests:
Quick test run (unit tests, no model download):
dotnet test ElBruno.Whisper.slnx --filter "Category!=Integration"
Full test run (includes integration with real models):
dotnet test ElBruno.Whisper.slnx
Test audio files are provided in testdata/audio/ for validation and transcription testing. For details, see the Testing Guide.
Contributions are welcome! Please:
git checkout -b feature/amazing-feature)git commit -m 'Add amazing feature')git push origin feature/amazing-feature)This project is licensed under the MIT License β see the LICENSE file for details.
Hi! I'm ElBruno π§‘, a passionate developer and content creator exploring AI, .NET, and modern development practices.
Made with β€οΈ by ElBruno
If you like this project, consider following my work across platforms: