qourex/FasterWhisper.NET

FasterWhisper.NET is a high-performance, cross-platform .NET SDK for Whisper speech recognition, powered by CTranslate2. Built for production workloads with streaming, Silero VAD, batching, multi-replica inference, audio diagnostics, subtitle generation, and local CPU/CUDA execution without Python dependencies.

C#

6

65 commits

updated Aug 31, 2026

See the code

README

FasterWhisper.NET Banner

FasterWhisper.NET

by Qourex — High-Performance Speech Recognition for .NET

Build & Test NuGet Downloads Documentation License: MIT .NET

Documentation Portal — Guides, API references, .NET 10.0 samples, and mobile deployment walkthroughs.


FasterWhisper.NET is a production-ready .NET SDK for OpenAI Whisper built on top of the high-performance CTranslate2 inference engine.

The library delivers high-throughput offline transcription, real-time streaming, batch inference pipelines, audio quality analytics, hallucination diagnostics, subtitle formatting, and cross-platform native execution for modern .NET workloads.


Why FasterWhisper.NET?

FasterWhisper.NET delivers an optimized, native .NET developer experience for Whisper speech recognition:

  • Idiomatic .NET API Surface — Clean, type-safe builder patterns and asynchronous APIs (async/await and IAsyncEnumerable<T>).
  • CTranslate2 Inference Engine — High-throughput execution with INT8, FP16, and INT16 quantization support.
  • Shared-Weight Replica Pools — Concurrent multi-threaded inference where replicas share loaded weight tensors in memory to minimize RAM and VRAM overhead.
  • Real-Time Streaming — Asynchronous push pipelines for live audio capture and incremental segment generation.
  • Integrated Voice Activity Detection (VAD) — Embedded Silero VAD v5 ONNX model for silence filtering and chunk segmentation.
  • Audio Quality Assessment — Built-in non-intrusive analyzer evaluating SNR, clipping, and signal clarity prior to inference.
  • Hallucination Diagnostics — Automated detection and mitigation of repetition loops and silent-region hallucinations.
  • Batched Inference Pipelines — High-throughput processing for long audio files using concurrent chunk batches.
  • Telemetry and Performance Profiling — Integrated measurement for Real-Time Factor (RTF), memory consumption, and execution duration.
  • Export Pipelines — Standardized output formatters for SubRip (SRT), WebVTT, TSV, JSON, and Markdown transcripts.
  • Automated Test Coverage — Comprehensive suite of 152 automated unit and integration tests validating interop boundaries and reliability.

Ecosystem Positioning

FasterWhisper.NET and Python's faster-whisper both leverage CTranslate2 for model inference. FasterWhisper.NET provides a dedicated .NET experience:

  • Type-Safe Native Interop — Built with source-generated P/Invoke ([LibraryImport]) for Native AOT compatibility and zero garbage collection overhead on hot paths.
  • Thread-Safe by Design — Internal replica coordination and resource semaphores prevent native re-entrancy conflicts in multi-threaded web applications.
  • Enterprise Framework Integration — First-class patterns for ASP.NET Core Dependency Injection, Blazor Server, Windows Forms, and .NET MAUI.
  • Self-Contained Cross-Platform Packaging — Pre-compiled native binaries packaged directly into NuGet packages for Windows, Linux, macOS, Android, and iOS.

Project Health

MetricDetails
Test Suite152 Automated Tests (Unit, Integration, Concurrency, Interop)
Native EngineCTranslate2 v4.7.0 (C++ / CUDA)
LicenseMIT License
Target Frameworks.NET 8.0, .NET 9.0, .NET 10.0
Supported Operating SystemsWindows (x64), Linux (x64), macOS (x64, ARM64), Android (ARM64), iOS (ARM64)
Core LanguagesC#, C++, CUDA

Feature Matrix

CapabilityStatusNotes
Standard Audio TranscriptionSupportedFile paths, streams, or raw PCM float[] arrays
Real-Time StreamingSupportedAsynchronous push pipeline with IAsyncEnumerable<T>
Voice Activity Detection (VAD)SupportedEmbedded Silero VAD v5 ONNX integration
Batched Inference PipelinesSupportedConcurrent chunk batching for high-throughput GPU workloads
Shared Replica PoolsSupportedWeight-shared multi-replica concurrency
Word-Level TimestampsSupportedCross-attention matrix alignment with median filtering
Audio Signal Quality AnalysisSupportedSNR calculation, clipping detection, quality grading
Hallucination MitigationSupportedCompression ratio validation and temperature fallback sequences
Subtitle and Transcript ExportSupportedSRT, WebVTT, TSV, JSON, and Markdown formatters
Memory-Mapped Weight LoadingSupportedRapid initialization via virtual memory mapping
Native AOT CompatibilitySupportedSource-generated P/Invoke declarations

Architecture Overview

graph TD
    App[Application Layer] --> SDK[FasterWhisper.NET Managed Layer]
    subgraph SDK Features
        SDK --> Audio[Audio Preprocessing & Resampling]
        SDK --> Stream[Streaming & IAsyncEnumerable]
        SDK --> VAD[Silero VAD v5 ONNX]
        SDK --> Diag[Diagnostics & Quality Analysis]
        SDK --> Export[Subtitle & Transcript Exporters]
        SDK --> Pool[Replica Pool & Semaphores]
    end
    SDK --> Native[Native Interop Bridge]
    Native --> CT2[CTranslate2 Engine C++]
    CT2 --> Models[Whisper Model Weights]

Table of Contents


Installation

Install the package via the .NET CLI:

dotnet add package FasterWhisper.NET

Or via the Package Manager Console:

Install-Package FasterWhisper.NET

[!NOTE] The base package includes pre-compiled native binaries for Windows (win-x64), Linux (linux-x64), macOS (osx-x64, osx-arm64), Android (arm64), and iOS (arm64). For GPU acceleration on Windows and Linux, install FasterWhisper.NET.Gpu and refer to CUDA Prerequisites.


Quick Start

using System;
using System.Threading.Tasks;
using Qourex.FasterWhisper.NET;

// 1. Download and initialize the model (cached to ~/.cache/qourex-fasterwhisper)
using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "base",       // "tiny", "base", "small", "medium", "large-v3", etc.
    device:          "cpu",        // "cpu" or "cuda"
    computeType:     "default"     // "float32", "float16", "int8", "int8_float16", etc.
);

// 2. Configure transcription options
var options = new WhisperOptions
{
    BeamSize = 5,
    WordTimestamps = false
};

// 3. Configure Voice Activity Detection (optional)
var vadOptions = new VadOptions
{
    Enabled   = true,
    Threshold = 0.5f
};

// 4. Transcribe audio (WAV, MP3, MP4, Opus)
var segments = model.Transcribe(
    mediaPath:  "audio.wav",
    language:   "en",           // Pass null for automatic language detection
    options:    options,
    vadOptions: vadOptions
);

// 5. Output results
foreach (var segment in segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");
}

Sample Applications

A suite of 10 sample applications targeting .NET 10.0 is provided under the samples/ directory:

  • Console Application (Cpu / Gpu) — Minimal CLI showcasing model downloading progress, parameter configuration, and transcription output.
  • ASP.NET Core Minimal API (Cpu / Gpu) — Production-grade REST API (POST /api/transcribe) demonstrating thread pool offloading and singleton model registration.
  • Blazor Web App (Cpu / Gpu) — Interactive server dashboard with SignalR progress indicators and interactive timeline segment inspection.
  • Windows Forms (Cpu / Gpu) — Desktop interface utilizing native .NET 10.0 Dark Mode and background worker threads for UI responsiveness.
  • .NET MAUI (Cpu / Gpu) — Cross-platform application demonstrating mobile asset extraction and native file picker integration.

For setup details and execution commands, see the Samples Documentation.


Available Models

Models are resolved automatically from Hugging Face on first use and cached locally:

ModelParametersDisk SizeApprox. VRAM (FP16)Relative SpeedTarget Use Case
tiny39 M~75 MB~1 GBFastestMobile devices, unit testing, quick prototyping
base74 M~142 MB~1 GBVery FastLightweight applications, desktop utilities
small244 M~466 MB~2 GBFastGeneral production balance
medium769 M~1.5 GB~5 GBModerateHigh-accuracy transcription
large-v11550 M~3.1 GB~10 GBStandardLegacy large model checkpoint
large-v21550 M~3.1 GB~10 GBStandardImproved large model checkpoint
large-v31550 M~3.1 GB~10 GBStandardHighest overall accuracy and multilingual quality
large-v3-turbo809 M~1.6 GB~3 GBFastHigh accuracy with reduced decoder depth
faster-distil-whisper-large-v3756 M~1.5 GB~3 GBFastOptimized English-only speed and accuracy

Devices and Compute Types

DeviceDescription
"cpu"CPU execution with Intel oneMKL / OpenBLAS acceleration
"cuda"NVIDIA GPU execution via CUDA and cuDNN runtimes
Compute TypeDescription
"default"Automatically selects the optimal precision for the host hardware
"float32"Full 32-bit floating point precision
"float16"16-bit half precision (recommended for CUDA devices)
"int8"8-bit integer quantization (fastest, lowest memory footprint)
"int8_float16"INT8 quantized compute with FP16 activation storage
"int16"16-bit integer quantization

Performance Benchmarks

The following benchmarks demonstrate throughput and memory characteristics across CPU and GPU configurations.

Benchmark Hardware Environment

  • CPU: Intel Core i7-4790 (4 Cores / 8 Threads)
  • RAM: 32 GB DDR3
  • GPU: NVIDIA GeForce GTX 1070 Ti (8 GB VRAM)
  • CUDA: 12.4 / cuDNN 8.9.7
  • Operating System: Windows 11 Pro
  • Model: faster-distil-whisper-large-v3 (756 million parameters)
  • Audio Duration: 972.29 seconds (16.2 minutes)

Initialization and Memory Overhead: Standard vs. Memory-Mapped

Loading StrategyLoad TimeCPU RAM DeltaGPU VRAM DeltaStartup Speedup
Standard Path-Based Load3,172.8 ms110.1 MB3,814.0 MBBaseline
Memory-Mapped Load2,506.2 ms963.1 MB3,713.0 MB1.27x

[!NOTE] Standard loading delegates allocations to the native C++ heap. Memory-mapped loading allocates virtual memory buffers within C# before passing pinned pointers to CTranslate2, reflecting in managed process working set telemetry.

Multi-Replica Resource Scaling (Shared Model Weights)

When scaling concurrent execution threads via NumReplicas, CTranslate2 shares model weights across replicas:

ConfigurationLoad TimeCPU RAMGPU VRAM
NumReplicas = 12,516.7 ms963.1 MB3,712.0 MB
NumReplicas = 22,406.2 ms1,454.2 MB3,700.0 MB
NumReplicas = 42,496.4 ms1,447.3 MB3,712.0 MB

Throughput and Optimization Results

Phase / ConfigurationDurationThroughput Metric
Quantization 'default' (float32)44,403.2 msReal-Time Factor (RTF): 0.0457
Quantization 'int8'31,708.0 msReal-Time Factor (RTF): 0.0326
Quantization Improvement (int8 vs. float32)—1.40x Speedup
4 Sequential Requests (Replica = 1)179,217.2 msBaseline
4 Concurrent Requests (Replica = 2)155,992.5 ms1.15x Speedup
Sequential File Transcription (2x Audio)89,092.7 msBaseline
Batched Pipeline Transcription (2x Audio)57,802.7 ms1.54x Speedup

Advanced Features

Fluent Model Builder

using Qourex.FasterWhisper.NET;

using var model = await WhisperModelBuilder.Create("base")
    .WithDevice("cuda")
    .WithComputeType("float16")
    .WithNumReplicas(2)
    .WithVad(threshold: 0.5f)
    .WithWordTimestamps()
    .WithDenoising()
    .BuildAsync();

var segments = model.Transcribe("meeting.wav");

Concurrency and Multi-Replica Execution

// Load model with 2 replicas sharing weights in memory
using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "base",
    device: "cpu",
    numReplicas: 2
);

// Transcribe multiple files concurrently
var tasks = new[] { "audio1.wav", "audio2.wav" }.Select(file => Task.Run(() =>
{
    var segments = model.Transcribe(file);
    Console.WriteLine($"Finished transcribing: {file}");
}));

await Task.WhenAll(tasks);

Batched Inference Pipeline

using Qourex.FasterWhisper.NET;

using var model = await WhisperModel.LoadAsync("base", device: "cuda");
using var pipeline = new BatchedInferencePipeline(model, batchSize: 8);

var result = pipeline.Transcribe("podcast.mp3");
foreach (var segment in result.Segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");
}

Word-Level Timestamps

var options = new WhisperOptions
{
    WordTimestamps    = true,
    MedianFilterWidth = 7 // Smoothing kernel width for cross-attention matrix
};

var segments = model.Transcribe("interview.wav", language: "en", options: options);

foreach (var segment in segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");

    foreach (var word in segment.Words)
    {
        Console.WriteLine($"  '{word.Word}' [{word.Start:F2}s -> {word.End:F2}s] (p={word.Probability:F3})");
    }
}

Real-Time Streaming Transcription

async IAsyncEnumerable<float[]> GetAudioStream()
{
    // Capture 16 kHz mono float32 PCM buffers
    while (isCapturing)
    {
        yield return await microphone.ReadChunkAsync();
    }
}

var options = new WhisperOptions { BeamSize = 1 }; // Greedy decoding for lowest latency
var vadOptions = new VadOptions
{
    Enabled              = true,
    Threshold            = 0.5f,
    MinSpeechDurationMs  = 250,
    MinSilenceDurationMs = 100
};

await foreach (var segment in model.TranscribeStreamAsync(
    GetAudioStream(),
    language: "en",
    options: options,
    vadOptions: vadOptions))
{
    Console.WriteLine($"[Live] {segment.Text}");
}

Voice Activity Detection (VAD)

var vadOptions = new VadOptions
{
    Enabled              = true,   // Enable Silero VAD segmentation
    Threshold            = 0.5f,   // Speech probability threshold (0.0 to 1.0)
    MinSpeechDurationMs  = 250,    // Minimum duration of speech intervals (ms)
    MinSilenceDurationMs = 1000    // Minimum silence required to split chunks (ms)
};

var segments = model.Transcribe("meeting.mp3", language: null, vadOptions: vadOptions);

In-Memory Model Loading

[!WARNING] The dictionary must include either vocabulary.txt or vocabulary.json alongside model.bin and config.json. Omitting vocabulary data results in a KeyNotFoundException during tokenizer initialization.

var modelFiles = new Dictionary<string, byte[]>
{
    ["model.bin"]       = File.ReadAllBytes("path/to/model.bin"),
    ["config.json"]     = File.ReadAllBytes("path/to/config.json"),
    ["vocabulary.txt"]  = File.ReadAllBytes("path/to/vocabulary.txt")
};

using var model = new WhisperModel(
    modelFiles,
    device: "cpu",
    computeType: "int8",
    cpuThreads: 4
);

var segments = model.Transcribe("audio.wav", language: "en");

Language Detection

float[] pcm = audioProcessor.LoadWav("speech.wav");
var detectedLanguages = model.DetectLanguage(pcm);

foreach (var (language, probability) in detectedLanguages.Take(5))
{
    Console.WriteLine($"  {language}: {probability:P1}");
}

Audio Preprocessing Options

var options = new WhisperOptions
{
    NormalizeAudio    = true,   // RMS amplitude normalization (target -20 dBFS)
    CutLowFrequencies = true,   // 80 Hz high-pass filter (removes DC offset and hum)
    PreEmphasis       = false,  // High-frequency emphasis filter
    DenoiseAudio      = false   // Spectral subtraction noise gate
};

Text Post-Processing Filters

var options = new WhisperOptions
{
    FilterFillerWords       = true,   // Removes vocal hesitations ("uh", "um", "ah", "eh", "mhm")
    PruneStutters           = true,   // Removes consecutive duplicate words
    ConditionOnPreviousText = true    // Retains preceding context for window continuity
};

Audio Quality Assessment

float[] samples = WhisperModel.LoadAudio("input.wav");
var report = AudioQualityReport.Assess(samples);

Console.WriteLine($"Quality Grade: {report.OverallGrade}");
Console.WriteLine($"Signal-to-Noise Ratio: {report.SignalToNoiseRatio:F1} dB");

foreach (var suggestion in report.Suggestions)
{
    Console.WriteLine($"Recommendation: {suggestion}");
}

Subtitle and Export Formats

var segments = model.Transcribe("presentation.wav");

// 1. Export as SRT string
string srtContent = SubtitleExporter.ToSrt(segments);
File.WriteAllText("presentation.srt", srtContent);

// 2. Export directly to WebVTT file
SubtitleExporter.WriteVtt(segments, "presentation.vtt");

// 3. Export as TSV or JSON data
string tsvContent = SubtitleExporter.ToTsv(segments);
string jsonContent = SubtitleExporter.ToJson(segments);

API Reference

WhisperOptions

PropertyTypeDefaultDescription
BeamSizeint5Beam size for beam search decoding. Set to 1 for greedy decoding
Patiencefloat1.0Beam search patience factor
LengthPenaltyfloat1.0Exponential penalty applied to sequence length
RepetitionPenaltyfloat1.0Penalty applied to previously generated tokens
NoRepeatNgramSizeint0Prevent repetition of n-grams of this size (0 disables)
MaxLengthint448Maximum tokens generated per 30-second window
SamplingTopKint1Top-K sampling pool size (1 = deterministic greedy)
SamplingTemperaturefloat1.0Softmax temperature for non-greedy sampling
NumHypothesesint1Number of hypothesis candidates returned
ReturnScoresbooltrueInclude token log-probability scores in output
ReturnNoSpeechProbbooltrueInclude silence probability scores in output
MaxInitialTimestampIndexint50Maximum index of the initial predicted timestamp token
SuppressBlankbooltrueSuppress blank outputs at start of sampling
SuppressTokensint[]?[-1]Explicit token IDs to suppress during decoding
WordTimestampsboolfalseExtract per-word timestamp boundaries via cross-attention
MedianFilterWidthint7Smoothing kernel width for cross-attention matrix
Temperaturesfloat[][0.0, 0.2, 0.4, 0.6, 0.8, 1.0]Temperature sequence used for validation fallbacks
LogProbThresholdfloat-1.0Minimum average log-probability threshold
NoSpeechThresholdfloat0.6Maximum no-speech confidence before classifying as silence
CompressionRatioThresholdfloat2.4Maximum gzip compression ratio before flagging repetitive loops
Prefixstring?nullText prefix used to constrain initial chunk generation
WithoutTimestampsboolfalseSuppress timestamp token generation
NormalizeAudiobooltrueStandardize signal levels via RMS normalization
CutLowFrequenciesbooltrueApply 80 Hz high-pass filter
ConditionOnPreviousTextbooltruePass preceding transcript into subsequent window prompt
FilterFillerWordsboolfalseRemove vocal filler words from output text
PruneStuttersboolfalseRemove consecutive duplicate words
PreEmphasisboolfalseApply high-frequency pre-emphasis filter
DenoiseAudioboolfalseApply spectral subtraction noise gate
InitialPromptstring?nullContextual text prompt guiding vocabulary and style
Hotwordsstring?nullComma-separated list of prioritized domain words
HallucinationSilenceThresholdfloat0Skip generation across silent regions exceeding duration (s)
PrependPunctuationsstring"\"'“¿([{-"Punctuation prepended to following word
AppendPunctuationsstring"\".。,,!!??::)”)]}、"Punctuation appended to preceding word
MaxNewTokensint0Maximum new tokens generated per chunk (0 = use MaxLength)
BestOfint5Number of candidate sequences evaluated when temperature > 0
PromptResetOnTemperaturefloat0.5Discard previous context when fallback temperature reaches threshold
ClipTimestampsList<(float, float)>?nullTemporal boundaries restricting transcription
MultilingualboolfalsePerform language detection per 30-second window
AdaptiveBeamSizebooltrueUse greedy decoding at temp=0 and expand during fallback
RestoreTextFormattingboolfalseApply grammar-based capitalization and punctuation rules
VocabularyBiasDictionary<string, float>?nullDirect logit probability biases for specific token strings
MultiPassEnabledboolfalseEnable second-pass decoding for low-confidence segments
MultiPassConfidenceThresholdfloat0.6Confidence threshold triggering second-pass decoding
MultiPassBeamSizeint10Beam size used during second-pass decoding

VadOptions

PropertyTypeDefaultDescription
EnabledboolfalseEnable or disable Silero VAD segmentation
Thresholdfloat0.5Speech probability threshold (0.0 to 1.0)
MinSpeechDurationMsint250Minimum speech duration in milliseconds
MinSilenceDurationMsint2000Minimum silence duration in milliseconds to trigger chunk split

WhisperSegment

PropertyTypeDescription
TextstringTranscribed text content
Tokensint[]Raw token IDs generated by the model tokenizer
ScorefloatAverage log-probability score
NoSpeechProbfloatProbability that the segment contains non-speech
StartfloatStart timestamp in seconds
EndfloatEnd timestamp in seconds
WordsList<WhisperWord>Word-level alignments (populated when WordTimestamps = true)

WhisperWord

PropertyTypeDescription
WordstringWord text content
StartfloatStart timestamp in seconds
EndfloatEnd timestamp in seconds
ProbabilityfloatAlignment confidence score (0.0 to 1.0)

Building from Source

Prerequisites

ComponentTarget RequirementDownload Link
CMake 3.18+Native C++ build systemcmake.org
Visual Studio 2022 (MSVC)C++ compilervisualstudio.com
CUDA Toolkit 12.xGPU builds onlyNVIDIA Developer
cuDNN 8.9.xGPU builds onlyNVIDIA Developer Archive
.NET SDK 8.0+Managed library builddotnet.microsoft.com

Automated Build Script

The repository includes a PowerShell automation script (build.ps1):

# Build with CUDA GPU acceleration (default)
.\build.ps1

# Build CPU-only (no CUDA or NVCC required)
.\build.ps1 -CpuOnly

The script automatically executes the following steps:

  1. Configures and compiles the native C++ wrapper (qourex_fasterwhisper_native.dll).
  2. Copies native dynamic libraries into runtimes/win-x64/native/.
  3. Builds the .NET solution and packs NuGet packages into ./artifacts/.

CUDA Prerequisites

[!IMPORTANT] CUDA and cuDNN runtimes are required only when initializing models with device: "cuda". CPU execution has no external GPU dependencies.

  1. NVIDIA CUDA Toolkit 12.x — Download
  2. NVIDIA cuDNN 8.9.x — Download

    Note on cuDNN versions: FasterWhisper.NET wraps CTranslate2, which natively links against cuDNN 8.x (cudnn64_8.dll on Windows / libcudnn.so.8 on Linux). If you have cuDNN 9 installed on your system, ensure cudnn64_8.dll is present in your system PATH or in your application's output directory.

Ensure the following dynamic libraries are present in your system PATH:

  • cudart64_12.dll
  • cublas64_12.dll
  • cublasLt64_12.dll
  • cudnn64_8.dll

Flash Attention Support

Flash Attention accelerates inference on supported NVIDIA GPUs (Ampere architecture or newer, compute capability ≥ 8.0):

using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "large-v3",
    device:          "cuda",
    computeType:     "float16",
    flashAttention:  true
);

Project Structure

Qourex.FasterWhisper/
├── src/
│   ├── Qourex.FasterWhisper.Native/   # C++ CMake wrapper for CTranslate2
│   │   ├── CMakeLists.txt
│   │   ├── qourex_fasterwhisper_native.cpp
│   │   └── qourex_fasterwhisper_native.h
│   ├── Qourex.FasterWhisper.NET/      # Core managed library (CPU Package)
│   │   ├── AudioProcessor.cs          # WAV decoding, resampling, Mel extraction
│   │   ├── AudioQualityReport.cs      # Signal quality assessment
│   │   ├── BatchedInferencePipeline.cs# High-throughput batch inference pipeline
│   │   ├── HallucinationDetector.cs   # Repetition and hallucination detection
│   │   ├── ModelDownloader.cs         # Hugging Face model downloader
│   │   ├── NativeMethods.cs           # P/Invoke declarations
│   │   ├── SileroVad.cs               # Silero VAD v5 ONNX integration
│   │   ├── StreamingMelExtractor.cs   # Real-time streaming Mel extractor
│   │   ├── SubtitleExporter.cs        # SRT, VTT, TSV, JSON export pipeline
│   │   ├── WhisperModel.cs            # Primary model API
│   │   └── WhisperModelBuilder.cs     # Fluent builder API
│   └── Qourex.FasterWhisper.NET.Gpu/  # GPU package source and assets
├── samples/
│   ├── Qourex.FasterWhisper.NET.Samples.Console.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Console.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.AspNetCore.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.AspNetCore.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Blazor.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Blazor.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.WinForms.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.WinForms.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Maui.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Maui.Gpu/
│   └── README.md
├── tests/
│   └── Qourex.FasterWhisper.NET.Tests/ # xUnit test suite (152 tests)
├── docs/                              # VitePress documentation portal
├── build.ps1                          # PowerShell build script
├── build.sh                           # Bash build script
├── Qourex.FasterWhisper.slnx          # Solution file
└── LICENSE                            # MIT License

Deployment and Production Guidelines

ASP.NET Core Integration

WhisperModel allocates native weights and initializes CTranslate2 execution engines. Model instances should be registered as a Singleton in your ASP.NET Core dependency injection container:

// Program.cs
builder.Services.AddSingleton<WhisperModel>(sp =>
{
    return WhisperModel.LoadAsync(
        modelNameOrPath: "base",
        device:          "cpu",
        computeType:     "default",
        numReplicas:     2 // Enable 2 concurrent worker replicas
    ).GetAwaiter().GetResult();
});

Because WhisperModel coordinates concurrent calls internally via replica pools and semaphores, it can be safely injected into scoped controllers or minimal APIs:

[ApiController]
[Route("api/transcribe")]
public class TranscriptionController : ControllerBase
{
    private readonly WhisperModel _model;

    public TranscriptionController(WhisperModel model)
    {
        _model = model;
    }

    [HttpPost]
    public async Task<IActionResult> TranscribeAudio(IFormFile file)
    {
        using var stream = file.OpenReadStream();
        var segments = await Task.Run(() => _model.Transcribe(stream));
        return Ok(segments);
    }
}

Docker GPU Deployment (Linux Container)

FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04

RUN apt-get update && apt-get install -y \
    dotnet-sdk-8.0 \
    libcublas-12-4 \
    libcudnn8 \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app
COPY . .
RUN dotnet publish -c Release -o out

ENTRYPOINT ["dotnet", "out/YourApplication.dll"]

Execute with NVIDIA GPU access:

docker run --gpus all -it your-whisper-app

Troubleshooting and Diagnostics

Issue: DllNotFoundException (Native Library Missing)

  • Cause: On Windows, the native wrapper depends on the Visual C++ Redistributable and Intel MKL runtime libraries. On GPU builds, CUDA 12.x and cuDNN runtime DLLs must be located in the system PATH.
  • Resolution:
    1. Install the Visual C++ Redistributable.
    2. For GPU execution, verify that cudart64_12.dll, cublas64_12.dll, and cudnn64_8.dll reside in your PATH. Verify environment setup by testing device: "cpu" first.

Issue: StackOverflowException in StreamingMelExtractor

  • Cause: Supplying an excessively large FFT window size exceeds stack allocation thresholds.
  • Resolution: Maintain fftSize below 8192 (default is 400 or 512).

Downstream Licensing Obligations

When deploying applications utilizing FasterWhisper.NET, downstream developers must comply with the licenses of bundled and external dependencies:

  1. Intel oneMKL (ISSL License): Windows packages bundle Intel MKL runtime binaries under the Intel Simplified Software License (ISSL). Downstream commercial users should note that the ISSL contains reverse-engineering restrictions. For environments where ISSL terms cannot be met, run on Linux (which links dynamically to OpenBLAS) or compile the native wrapper with alternative BLAS engines.
  2. FFmpeg (LGPL / GPL): The library invokes the external ffmpeg CLI as a fallback subprocess for non-WAV media. Ensure compliance with FFmpeg redistribution terms if bundling FFmpeg binaries with your application. WAV files are decoded using a 100% managed C# decoder and do not require FFmpeg.
  3. ONNX Runtime and Silero VAD (MIT License): Both dependencies are distributed under the MIT license. Downloaded Silero VAD model assets are verified via SHA-256 integrity checks.

License

This project is licensed under the MIT License — see the LICENSE file for details.

MIT License · Copyright (c) 2026 Qourex

Maintained by Qourex
Report Issue · Discussions

android
asr
audio-processing
c-sharp
csharp
ctranslate2
cuda
cudnn
dotnet
faster-whisper
gpu-acceleration
ios
offline-speech-to-text
speech-recognition
speech-to-text
streaming-transcription
transcription
vad
voice-activity-detection
whisper

qourex/FasterWhisper.NET

FasterWhisper.NET is a high-performance, cross-platform .NET SDK for Whisper speech recognition, powered by CTranslate2. Built for production workloads with streaming, Silero VAD, batching, multi-replica inference, audio diagnostics, subtitle generation, and local CPU/CUDA execution without Python dependencies.

C#

6

65 commits

updated Aug 31, 2026

See the code

README

FasterWhisper.NET Banner

FasterWhisper.NET

by Qourex — High-Performance Speech Recognition for .NET

Build & Test NuGet Downloads Documentation License: MIT .NET

Documentation Portal — Guides, API references, .NET 10.0 samples, and mobile deployment walkthroughs.


FasterWhisper.NET is a production-ready .NET SDK for OpenAI Whisper built on top of the high-performance CTranslate2 inference engine.

The library delivers high-throughput offline transcription, real-time streaming, batch inference pipelines, audio quality analytics, hallucination diagnostics, subtitle formatting, and cross-platform native execution for modern .NET workloads.


Why FasterWhisper.NET?

FasterWhisper.NET delivers an optimized, native .NET developer experience for Whisper speech recognition:

  • Idiomatic .NET API Surface — Clean, type-safe builder patterns and asynchronous APIs (async/await and IAsyncEnumerable<T>).
  • CTranslate2 Inference Engine — High-throughput execution with INT8, FP16, and INT16 quantization support.
  • Shared-Weight Replica Pools — Concurrent multi-threaded inference where replicas share loaded weight tensors in memory to minimize RAM and VRAM overhead.
  • Real-Time Streaming — Asynchronous push pipelines for live audio capture and incremental segment generation.
  • Integrated Voice Activity Detection (VAD) — Embedded Silero VAD v5 ONNX model for silence filtering and chunk segmentation.
  • Audio Quality Assessment — Built-in non-intrusive analyzer evaluating SNR, clipping, and signal clarity prior to inference.
  • Hallucination Diagnostics — Automated detection and mitigation of repetition loops and silent-region hallucinations.
  • Batched Inference Pipelines — High-throughput processing for long audio files using concurrent chunk batches.
  • Telemetry and Performance Profiling — Integrated measurement for Real-Time Factor (RTF), memory consumption, and execution duration.
  • Export Pipelines — Standardized output formatters for SubRip (SRT), WebVTT, TSV, JSON, and Markdown transcripts.
  • Automated Test Coverage — Comprehensive suite of 152 automated unit and integration tests validating interop boundaries and reliability.

Ecosystem Positioning

FasterWhisper.NET and Python's faster-whisper both leverage CTranslate2 for model inference. FasterWhisper.NET provides a dedicated .NET experience:

  • Type-Safe Native Interop — Built with source-generated P/Invoke ([LibraryImport]) for Native AOT compatibility and zero garbage collection overhead on hot paths.
  • Thread-Safe by Design — Internal replica coordination and resource semaphores prevent native re-entrancy conflicts in multi-threaded web applications.
  • Enterprise Framework Integration — First-class patterns for ASP.NET Core Dependency Injection, Blazor Server, Windows Forms, and .NET MAUI.
  • Self-Contained Cross-Platform Packaging — Pre-compiled native binaries packaged directly into NuGet packages for Windows, Linux, macOS, Android, and iOS.

Project Health

MetricDetails
Test Suite152 Automated Tests (Unit, Integration, Concurrency, Interop)
Native EngineCTranslate2 v4.7.0 (C++ / CUDA)
LicenseMIT License
Target Frameworks.NET 8.0, .NET 9.0, .NET 10.0
Supported Operating SystemsWindows (x64), Linux (x64), macOS (x64, ARM64), Android (ARM64), iOS (ARM64)
Core LanguagesC#, C++, CUDA

Feature Matrix

CapabilityStatusNotes
Standard Audio TranscriptionSupportedFile paths, streams, or raw PCM float[] arrays
Real-Time StreamingSupportedAsynchronous push pipeline with IAsyncEnumerable<T>
Voice Activity Detection (VAD)SupportedEmbedded Silero VAD v5 ONNX integration
Batched Inference PipelinesSupportedConcurrent chunk batching for high-throughput GPU workloads
Shared Replica PoolsSupportedWeight-shared multi-replica concurrency
Word-Level TimestampsSupportedCross-attention matrix alignment with median filtering
Audio Signal Quality AnalysisSupportedSNR calculation, clipping detection, quality grading
Hallucination MitigationSupportedCompression ratio validation and temperature fallback sequences
Subtitle and Transcript ExportSupportedSRT, WebVTT, TSV, JSON, and Markdown formatters
Memory-Mapped Weight LoadingSupportedRapid initialization via virtual memory mapping
Native AOT CompatibilitySupportedSource-generated P/Invoke declarations

Architecture Overview

graph TD
    App[Application Layer] --> SDK[FasterWhisper.NET Managed Layer]
    subgraph SDK Features
        SDK --> Audio[Audio Preprocessing & Resampling]
        SDK --> Stream[Streaming & IAsyncEnumerable]
        SDK --> VAD[Silero VAD v5 ONNX]
        SDK --> Diag[Diagnostics & Quality Analysis]
        SDK --> Export[Subtitle & Transcript Exporters]
        SDK --> Pool[Replica Pool & Semaphores]
    end
    SDK --> Native[Native Interop Bridge]
    Native --> CT2[CTranslate2 Engine C++]
    CT2 --> Models[Whisper Model Weights]

Table of Contents


Installation

Install the package via the .NET CLI:

dotnet add package FasterWhisper.NET

Or via the Package Manager Console:

Install-Package FasterWhisper.NET

[!NOTE] The base package includes pre-compiled native binaries for Windows (win-x64), Linux (linux-x64), macOS (osx-x64, osx-arm64), Android (arm64), and iOS (arm64). For GPU acceleration on Windows and Linux, install FasterWhisper.NET.Gpu and refer to CUDA Prerequisites.


Quick Start

using System;
using System.Threading.Tasks;
using Qourex.FasterWhisper.NET;

// 1. Download and initialize the model (cached to ~/.cache/qourex-fasterwhisper)
using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "base",       // "tiny", "base", "small", "medium", "large-v3", etc.
    device:          "cpu",        // "cpu" or "cuda"
    computeType:     "default"     // "float32", "float16", "int8", "int8_float16", etc.
);

// 2. Configure transcription options
var options = new WhisperOptions
{
    BeamSize = 5,
    WordTimestamps = false
};

// 3. Configure Voice Activity Detection (optional)
var vadOptions = new VadOptions
{
    Enabled   = true,
    Threshold = 0.5f
};

// 4. Transcribe audio (WAV, MP3, MP4, Opus)
var segments = model.Transcribe(
    mediaPath:  "audio.wav",
    language:   "en",           // Pass null for automatic language detection
    options:    options,
    vadOptions: vadOptions
);

// 5. Output results
foreach (var segment in segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");
}

Sample Applications

A suite of 10 sample applications targeting .NET 10.0 is provided under the samples/ directory:

  • Console Application (Cpu / Gpu) — Minimal CLI showcasing model downloading progress, parameter configuration, and transcription output.
  • ASP.NET Core Minimal API (Cpu / Gpu) — Production-grade REST API (POST /api/transcribe) demonstrating thread pool offloading and singleton model registration.
  • Blazor Web App (Cpu / Gpu) — Interactive server dashboard with SignalR progress indicators and interactive timeline segment inspection.
  • Windows Forms (Cpu / Gpu) — Desktop interface utilizing native .NET 10.0 Dark Mode and background worker threads for UI responsiveness.
  • .NET MAUI (Cpu / Gpu) — Cross-platform application demonstrating mobile asset extraction and native file picker integration.

For setup details and execution commands, see the Samples Documentation.


Available Models

Models are resolved automatically from Hugging Face on first use and cached locally:

ModelParametersDisk SizeApprox. VRAM (FP16)Relative SpeedTarget Use Case
tiny39 M~75 MB~1 GBFastestMobile devices, unit testing, quick prototyping
base74 M~142 MB~1 GBVery FastLightweight applications, desktop utilities
small244 M~466 MB~2 GBFastGeneral production balance
medium769 M~1.5 GB~5 GBModerateHigh-accuracy transcription
large-v11550 M~3.1 GB~10 GBStandardLegacy large model checkpoint
large-v21550 M~3.1 GB~10 GBStandardImproved large model checkpoint
large-v31550 M~3.1 GB~10 GBStandardHighest overall accuracy and multilingual quality
large-v3-turbo809 M~1.6 GB~3 GBFastHigh accuracy with reduced decoder depth
faster-distil-whisper-large-v3756 M~1.5 GB~3 GBFastOptimized English-only speed and accuracy

Devices and Compute Types

DeviceDescription
"cpu"CPU execution with Intel oneMKL / OpenBLAS acceleration
"cuda"NVIDIA GPU execution via CUDA and cuDNN runtimes
Compute TypeDescription
"default"Automatically selects the optimal precision for the host hardware
"float32"Full 32-bit floating point precision
"float16"16-bit half precision (recommended for CUDA devices)
"int8"8-bit integer quantization (fastest, lowest memory footprint)
"int8_float16"INT8 quantized compute with FP16 activation storage
"int16"16-bit integer quantization

Performance Benchmarks

The following benchmarks demonstrate throughput and memory characteristics across CPU and GPU configurations.

Benchmark Hardware Environment

  • CPU: Intel Core i7-4790 (4 Cores / 8 Threads)
  • RAM: 32 GB DDR3
  • GPU: NVIDIA GeForce GTX 1070 Ti (8 GB VRAM)
  • CUDA: 12.4 / cuDNN 8.9.7
  • Operating System: Windows 11 Pro
  • Model: faster-distil-whisper-large-v3 (756 million parameters)
  • Audio Duration: 972.29 seconds (16.2 minutes)

Initialization and Memory Overhead: Standard vs. Memory-Mapped

Loading StrategyLoad TimeCPU RAM DeltaGPU VRAM DeltaStartup Speedup
Standard Path-Based Load3,172.8 ms110.1 MB3,814.0 MBBaseline
Memory-Mapped Load2,506.2 ms963.1 MB3,713.0 MB1.27x

[!NOTE] Standard loading delegates allocations to the native C++ heap. Memory-mapped loading allocates virtual memory buffers within C# before passing pinned pointers to CTranslate2, reflecting in managed process working set telemetry.

Multi-Replica Resource Scaling (Shared Model Weights)

When scaling concurrent execution threads via NumReplicas, CTranslate2 shares model weights across replicas:

ConfigurationLoad TimeCPU RAMGPU VRAM
NumReplicas = 12,516.7 ms963.1 MB3,712.0 MB
NumReplicas = 22,406.2 ms1,454.2 MB3,700.0 MB
NumReplicas = 42,496.4 ms1,447.3 MB3,712.0 MB

Throughput and Optimization Results

Phase / ConfigurationDurationThroughput Metric
Quantization 'default' (float32)44,403.2 msReal-Time Factor (RTF): 0.0457
Quantization 'int8'31,708.0 msReal-Time Factor (RTF): 0.0326
Quantization Improvement (int8 vs. float32)—1.40x Speedup
4 Sequential Requests (Replica = 1)179,217.2 msBaseline
4 Concurrent Requests (Replica = 2)155,992.5 ms1.15x Speedup
Sequential File Transcription (2x Audio)89,092.7 msBaseline
Batched Pipeline Transcription (2x Audio)57,802.7 ms1.54x Speedup

Advanced Features

Fluent Model Builder

using Qourex.FasterWhisper.NET;

using var model = await WhisperModelBuilder.Create("base")
    .WithDevice("cuda")
    .WithComputeType("float16")
    .WithNumReplicas(2)
    .WithVad(threshold: 0.5f)
    .WithWordTimestamps()
    .WithDenoising()
    .BuildAsync();

var segments = model.Transcribe("meeting.wav");

Concurrency and Multi-Replica Execution

// Load model with 2 replicas sharing weights in memory
using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "base",
    device: "cpu",
    numReplicas: 2
);

// Transcribe multiple files concurrently
var tasks = new[] { "audio1.wav", "audio2.wav" }.Select(file => Task.Run(() =>
{
    var segments = model.Transcribe(file);
    Console.WriteLine($"Finished transcribing: {file}");
}));

await Task.WhenAll(tasks);

Batched Inference Pipeline

using Qourex.FasterWhisper.NET;

using var model = await WhisperModel.LoadAsync("base", device: "cuda");
using var pipeline = new BatchedInferencePipeline(model, batchSize: 8);

var result = pipeline.Transcribe("podcast.mp3");
foreach (var segment in result.Segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");
}

Word-Level Timestamps

var options = new WhisperOptions
{
    WordTimestamps    = true,
    MedianFilterWidth = 7 // Smoothing kernel width for cross-attention matrix
};

var segments = model.Transcribe("interview.wav", language: "en", options: options);

foreach (var segment in segments)
{
    Console.WriteLine($"[{segment.Start:F2}s -> {segment.End:F2}s] {segment.Text}");

    foreach (var word in segment.Words)
    {
        Console.WriteLine($"  '{word.Word}' [{word.Start:F2}s -> {word.End:F2}s] (p={word.Probability:F3})");
    }
}

Real-Time Streaming Transcription

async IAsyncEnumerable<float[]> GetAudioStream()
{
    // Capture 16 kHz mono float32 PCM buffers
    while (isCapturing)
    {
        yield return await microphone.ReadChunkAsync();
    }
}

var options = new WhisperOptions { BeamSize = 1 }; // Greedy decoding for lowest latency
var vadOptions = new VadOptions
{
    Enabled              = true,
    Threshold            = 0.5f,
    MinSpeechDurationMs  = 250,
    MinSilenceDurationMs = 100
};

await foreach (var segment in model.TranscribeStreamAsync(
    GetAudioStream(),
    language: "en",
    options: options,
    vadOptions: vadOptions))
{
    Console.WriteLine($"[Live] {segment.Text}");
}

Voice Activity Detection (VAD)

var vadOptions = new VadOptions
{
    Enabled              = true,   // Enable Silero VAD segmentation
    Threshold            = 0.5f,   // Speech probability threshold (0.0 to 1.0)
    MinSpeechDurationMs  = 250,    // Minimum duration of speech intervals (ms)
    MinSilenceDurationMs = 1000    // Minimum silence required to split chunks (ms)
};

var segments = model.Transcribe("meeting.mp3", language: null, vadOptions: vadOptions);

In-Memory Model Loading

[!WARNING] The dictionary must include either vocabulary.txt or vocabulary.json alongside model.bin and config.json. Omitting vocabulary data results in a KeyNotFoundException during tokenizer initialization.

var modelFiles = new Dictionary<string, byte[]>
{
    ["model.bin"]       = File.ReadAllBytes("path/to/model.bin"),
    ["config.json"]     = File.ReadAllBytes("path/to/config.json"),
    ["vocabulary.txt"]  = File.ReadAllBytes("path/to/vocabulary.txt")
};

using var model = new WhisperModel(
    modelFiles,
    device: "cpu",
    computeType: "int8",
    cpuThreads: 4
);

var segments = model.Transcribe("audio.wav", language: "en");

Language Detection

float[] pcm = audioProcessor.LoadWav("speech.wav");
var detectedLanguages = model.DetectLanguage(pcm);

foreach (var (language, probability) in detectedLanguages.Take(5))
{
    Console.WriteLine($"  {language}: {probability:P1}");
}

Audio Preprocessing Options

var options = new WhisperOptions
{
    NormalizeAudio    = true,   // RMS amplitude normalization (target -20 dBFS)
    CutLowFrequencies = true,   // 80 Hz high-pass filter (removes DC offset and hum)
    PreEmphasis       = false,  // High-frequency emphasis filter
    DenoiseAudio      = false   // Spectral subtraction noise gate
};

Text Post-Processing Filters

var options = new WhisperOptions
{
    FilterFillerWords       = true,   // Removes vocal hesitations ("uh", "um", "ah", "eh", "mhm")
    PruneStutters           = true,   // Removes consecutive duplicate words
    ConditionOnPreviousText = true    // Retains preceding context for window continuity
};

Audio Quality Assessment

float[] samples = WhisperModel.LoadAudio("input.wav");
var report = AudioQualityReport.Assess(samples);

Console.WriteLine($"Quality Grade: {report.OverallGrade}");
Console.WriteLine($"Signal-to-Noise Ratio: {report.SignalToNoiseRatio:F1} dB");

foreach (var suggestion in report.Suggestions)
{
    Console.WriteLine($"Recommendation: {suggestion}");
}

Subtitle and Export Formats

var segments = model.Transcribe("presentation.wav");

// 1. Export as SRT string
string srtContent = SubtitleExporter.ToSrt(segments);
File.WriteAllText("presentation.srt", srtContent);

// 2. Export directly to WebVTT file
SubtitleExporter.WriteVtt(segments, "presentation.vtt");

// 3. Export as TSV or JSON data
string tsvContent = SubtitleExporter.ToTsv(segments);
string jsonContent = SubtitleExporter.ToJson(segments);

API Reference

WhisperOptions

PropertyTypeDefaultDescription
BeamSizeint5Beam size for beam search decoding. Set to 1 for greedy decoding
Patiencefloat1.0Beam search patience factor
LengthPenaltyfloat1.0Exponential penalty applied to sequence length
RepetitionPenaltyfloat1.0Penalty applied to previously generated tokens
NoRepeatNgramSizeint0Prevent repetition of n-grams of this size (0 disables)
MaxLengthint448Maximum tokens generated per 30-second window
SamplingTopKint1Top-K sampling pool size (1 = deterministic greedy)
SamplingTemperaturefloat1.0Softmax temperature for non-greedy sampling
NumHypothesesint1Number of hypothesis candidates returned
ReturnScoresbooltrueInclude token log-probability scores in output
ReturnNoSpeechProbbooltrueInclude silence probability scores in output
MaxInitialTimestampIndexint50Maximum index of the initial predicted timestamp token
SuppressBlankbooltrueSuppress blank outputs at start of sampling
SuppressTokensint[]?[-1]Explicit token IDs to suppress during decoding
WordTimestampsboolfalseExtract per-word timestamp boundaries via cross-attention
MedianFilterWidthint7Smoothing kernel width for cross-attention matrix
Temperaturesfloat[][0.0, 0.2, 0.4, 0.6, 0.8, 1.0]Temperature sequence used for validation fallbacks
LogProbThresholdfloat-1.0Minimum average log-probability threshold
NoSpeechThresholdfloat0.6Maximum no-speech confidence before classifying as silence
CompressionRatioThresholdfloat2.4Maximum gzip compression ratio before flagging repetitive loops
Prefixstring?nullText prefix used to constrain initial chunk generation
WithoutTimestampsboolfalseSuppress timestamp token generation
NormalizeAudiobooltrueStandardize signal levels via RMS normalization
CutLowFrequenciesbooltrueApply 80 Hz high-pass filter
ConditionOnPreviousTextbooltruePass preceding transcript into subsequent window prompt
FilterFillerWordsboolfalseRemove vocal filler words from output text
PruneStuttersboolfalseRemove consecutive duplicate words
PreEmphasisboolfalseApply high-frequency pre-emphasis filter
DenoiseAudioboolfalseApply spectral subtraction noise gate
InitialPromptstring?nullContextual text prompt guiding vocabulary and style
Hotwordsstring?nullComma-separated list of prioritized domain words
HallucinationSilenceThresholdfloat0Skip generation across silent regions exceeding duration (s)
PrependPunctuationsstring"\"'“¿([{-"Punctuation prepended to following word
AppendPunctuationsstring"\".。,,!!??::)”)]}、"Punctuation appended to preceding word
MaxNewTokensint0Maximum new tokens generated per chunk (0 = use MaxLength)
BestOfint5Number of candidate sequences evaluated when temperature > 0
PromptResetOnTemperaturefloat0.5Discard previous context when fallback temperature reaches threshold
ClipTimestampsList<(float, float)>?nullTemporal boundaries restricting transcription
MultilingualboolfalsePerform language detection per 30-second window
AdaptiveBeamSizebooltrueUse greedy decoding at temp=0 and expand during fallback
RestoreTextFormattingboolfalseApply grammar-based capitalization and punctuation rules
VocabularyBiasDictionary<string, float>?nullDirect logit probability biases for specific token strings
MultiPassEnabledboolfalseEnable second-pass decoding for low-confidence segments
MultiPassConfidenceThresholdfloat0.6Confidence threshold triggering second-pass decoding
MultiPassBeamSizeint10Beam size used during second-pass decoding

VadOptions

PropertyTypeDefaultDescription
EnabledboolfalseEnable or disable Silero VAD segmentation
Thresholdfloat0.5Speech probability threshold (0.0 to 1.0)
MinSpeechDurationMsint250Minimum speech duration in milliseconds
MinSilenceDurationMsint2000Minimum silence duration in milliseconds to trigger chunk split

WhisperSegment

PropertyTypeDescription
TextstringTranscribed text content
Tokensint[]Raw token IDs generated by the model tokenizer
ScorefloatAverage log-probability score
NoSpeechProbfloatProbability that the segment contains non-speech
StartfloatStart timestamp in seconds
EndfloatEnd timestamp in seconds
WordsList<WhisperWord>Word-level alignments (populated when WordTimestamps = true)

WhisperWord

PropertyTypeDescription
WordstringWord text content
StartfloatStart timestamp in seconds
EndfloatEnd timestamp in seconds
ProbabilityfloatAlignment confidence score (0.0 to 1.0)

Building from Source

Prerequisites

ComponentTarget RequirementDownload Link
CMake 3.18+Native C++ build systemcmake.org
Visual Studio 2022 (MSVC)C++ compilervisualstudio.com
CUDA Toolkit 12.xGPU builds onlyNVIDIA Developer
cuDNN 8.9.xGPU builds onlyNVIDIA Developer Archive
.NET SDK 8.0+Managed library builddotnet.microsoft.com

Automated Build Script

The repository includes a PowerShell automation script (build.ps1):

# Build with CUDA GPU acceleration (default)
.\build.ps1

# Build CPU-only (no CUDA or NVCC required)
.\build.ps1 -CpuOnly

The script automatically executes the following steps:

  1. Configures and compiles the native C++ wrapper (qourex_fasterwhisper_native.dll).
  2. Copies native dynamic libraries into runtimes/win-x64/native/.
  3. Builds the .NET solution and packs NuGet packages into ./artifacts/.

CUDA Prerequisites

[!IMPORTANT] CUDA and cuDNN runtimes are required only when initializing models with device: "cuda". CPU execution has no external GPU dependencies.

  1. NVIDIA CUDA Toolkit 12.x — Download
  2. NVIDIA cuDNN 8.9.x — Download

    Note on cuDNN versions: FasterWhisper.NET wraps CTranslate2, which natively links against cuDNN 8.x (cudnn64_8.dll on Windows / libcudnn.so.8 on Linux). If you have cuDNN 9 installed on your system, ensure cudnn64_8.dll is present in your system PATH or in your application's output directory.

Ensure the following dynamic libraries are present in your system PATH:

  • cudart64_12.dll
  • cublas64_12.dll
  • cublasLt64_12.dll
  • cudnn64_8.dll

Flash Attention Support

Flash Attention accelerates inference on supported NVIDIA GPUs (Ampere architecture or newer, compute capability ≥ 8.0):

using var model = await WhisperModel.LoadAsync(
    modelNameOrPath: "large-v3",
    device:          "cuda",
    computeType:     "float16",
    flashAttention:  true
);

Project Structure

Qourex.FasterWhisper/
├── src/
│   ├── Qourex.FasterWhisper.Native/   # C++ CMake wrapper for CTranslate2
│   │   ├── CMakeLists.txt
│   │   ├── qourex_fasterwhisper_native.cpp
│   │   └── qourex_fasterwhisper_native.h
│   ├── Qourex.FasterWhisper.NET/      # Core managed library (CPU Package)
│   │   ├── AudioProcessor.cs          # WAV decoding, resampling, Mel extraction
│   │   ├── AudioQualityReport.cs      # Signal quality assessment
│   │   ├── BatchedInferencePipeline.cs# High-throughput batch inference pipeline
│   │   ├── HallucinationDetector.cs   # Repetition and hallucination detection
│   │   ├── ModelDownloader.cs         # Hugging Face model downloader
│   │   ├── NativeMethods.cs           # P/Invoke declarations
│   │   ├── SileroVad.cs               # Silero VAD v5 ONNX integration
│   │   ├── StreamingMelExtractor.cs   # Real-time streaming Mel extractor
│   │   ├── SubtitleExporter.cs        # SRT, VTT, TSV, JSON export pipeline
│   │   ├── WhisperModel.cs            # Primary model API
│   │   └── WhisperModelBuilder.cs     # Fluent builder API
│   └── Qourex.FasterWhisper.NET.Gpu/  # GPU package source and assets
├── samples/
│   ├── Qourex.FasterWhisper.NET.Samples.Console.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Console.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.AspNetCore.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.AspNetCore.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Blazor.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Blazor.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.WinForms.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.WinForms.Gpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Maui.Cpu/
│   ├── Qourex.FasterWhisper.NET.Samples.Maui.Gpu/
│   └── README.md
├── tests/
│   └── Qourex.FasterWhisper.NET.Tests/ # xUnit test suite (152 tests)
├── docs/                              # VitePress documentation portal
├── build.ps1                          # PowerShell build script
├── build.sh                           # Bash build script
├── Qourex.FasterWhisper.slnx          # Solution file
└── LICENSE                            # MIT License

Deployment and Production Guidelines

ASP.NET Core Integration

WhisperModel allocates native weights and initializes CTranslate2 execution engines. Model instances should be registered as a Singleton in your ASP.NET Core dependency injection container:

// Program.cs
builder.Services.AddSingleton<WhisperModel>(sp =>
{
    return WhisperModel.LoadAsync(
        modelNameOrPath: "base",
        device:          "cpu",
        computeType:     "default",
        numReplicas:     2 // Enable 2 concurrent worker replicas
    ).GetAwaiter().GetResult();
});

Because WhisperModel coordinates concurrent calls internally via replica pools and semaphores, it can be safely injected into scoped controllers or minimal APIs:

[ApiController]
[Route("api/transcribe")]
public class TranscriptionController : ControllerBase
{
    private readonly WhisperModel _model;

    public TranscriptionController(WhisperModel model)
    {
        _model = model;
    }

    [HttpPost]
    public async Task<IActionResult> TranscribeAudio(IFormFile file)
    {
        using var stream = file.OpenReadStream();
        var segments = await Task.Run(() => _model.Transcribe(stream));
        return Ok(segments);
    }
}

Docker GPU Deployment (Linux Container)

FROM nvidia/cuda:12.4.1-runtime-ubuntu22.04

RUN apt-get update && apt-get install -y \
    dotnet-sdk-8.0 \
    libcublas-12-4 \
    libcudnn8 \
    && rm -rf /var/lib/apt/lists/*

WORKDIR /app
COPY . .
RUN dotnet publish -c Release -o out

ENTRYPOINT ["dotnet", "out/YourApplication.dll"]

Execute with NVIDIA GPU access:

docker run --gpus all -it your-whisper-app

Troubleshooting and Diagnostics

Issue: DllNotFoundException (Native Library Missing)

  • Cause: On Windows, the native wrapper depends on the Visual C++ Redistributable and Intel MKL runtime libraries. On GPU builds, CUDA 12.x and cuDNN runtime DLLs must be located in the system PATH.
  • Resolution:
    1. Install the Visual C++ Redistributable.
    2. For GPU execution, verify that cudart64_12.dll, cublas64_12.dll, and cudnn64_8.dll reside in your PATH. Verify environment setup by testing device: "cpu" first.

Issue: StackOverflowException in StreamingMelExtractor

  • Cause: Supplying an excessively large FFT window size exceeds stack allocation thresholds.
  • Resolution: Maintain fftSize below 8192 (default is 400 or 512).

Downstream Licensing Obligations

When deploying applications utilizing FasterWhisper.NET, downstream developers must comply with the licenses of bundled and external dependencies:

  1. Intel oneMKL (ISSL License): Windows packages bundle Intel MKL runtime binaries under the Intel Simplified Software License (ISSL). Downstream commercial users should note that the ISSL contains reverse-engineering restrictions. For environments where ISSL terms cannot be met, run on Linux (which links dynamically to OpenBLAS) or compile the native wrapper with alternative BLAS engines.
  2. FFmpeg (LGPL / GPL): The library invokes the external ffmpeg CLI as a fallback subprocess for non-WAV media. Ensure compliance with FFmpeg redistribution terms if bundling FFmpeg binaries with your application. WAV files are decoded using a 100% managed C# decoder and do not require FFmpeg.
  3. ONNX Runtime and Silero VAD (MIT License): Both dependencies are distributed under the MIT license. Downloaded Silero VAD model assets are verified via SHA-256 integrity checks.

License

This project is licensed under the MIT License — see the LICENSE file for details.

MIT License · Copyright (c) 2026 Qourex

Maintained by Qourex
Report Issue · Discussions

android
asr
audio-processing
c-sharp
csharp
ctranslate2
cuda
cudnn
dotnet
faster-whisper
gpu-acceleration
ios
offline-speech-to-text
speech-recognition
speech-to-text
streaming-transcription
transcription
vad
voice-activity-detection
whisper

Languages

C#

86.7%

C++

5.3%

PowerShell

3.9%

C

2.1%

Shell

1.4%