iamdroppy/WhisperSharp

Whisper implementation in C#

C#

0

3 commits

updated Sep 21, 2026

See the code

README

WhisperSharp

A from-scratch implementation of OpenAI's Whisper speech-recognition model (Robust Speech Recognition via Large-Scale Weak Supervision, Radford et al. 2022) in C# 14 / .NET 10, running on TorchSharp / libtorch with CUDA GPU acceleration (FP16) and automatic CPU fallback.

The architecture, audio frontend, tokenizer, weight loader, and decoding loop are all implemented in this repository — the only runtime dependency is TorchSharp, which supplies the tensor kernels (cuBLAS/cuDNN on GPU). The only inputs are a standard Hugging Face Whisper checkpoint and a WAV file.

using var whisper = Whisper.LoadFromDirectory("models/whisper-base"); // CUDA + FP16 if available
var result = whisper.Transcribe("speech.wav");
Console.WriteLine(result.Text);

Features

  • GPU inference — all model math (convs, attention, MLPs, mel STFT) runs on the torch device. Defaults to CUDA with FP16 weights/compute when an NVIDIA GPU is present, FP32 on CPU otherwise; override via WhisperOptions. LayerNorm and attention softmax are computed in FP32 (as in the reference implementation) so FP16 output stays stable.
  • Faithful architecture — 2× GELU conv frontend, sinusoidal/learned position embeddings, pre-norm residual attention blocks, causal decoder self-attention, cross-attention to audio, tied output projection. Matches whisper/model.py and Hugging Face WhisperModel.
  • Exact audio frontend — centered STFT (reflect padding, periodic Hann, n_fft=400, hop=160) via torch.fft.rfft over windowed frames (chunked so hours-long audio stays memory-bounded), a librosa-compatible Slaney mel filterbank (80 or 128 bins), and Whisper's log10/clamp/rescale normalization.
  • Real tokenizer — GPT-2 byte-level BPE with the GPT-2 pre-tokenization regex, plus Whisper's structured special tokens (start-of-transcript, 99/100 language tokens, task tokens, 1501 timestamp tokens) derived exactly for English-only, multilingual, and large-v3 vocabularies.
  • Greedy decoding with timestamps — language detection, blank/non-speech suppression, the full timestamp-rule logit grammar, an incremental on-device KV cache, and the 30-second seek loop that stitches timestamped segments (text / segments / SRT output).
  • Weight loading — reads Hugging Face safetensors (F32 / F16 / BF16), streamed tensor-by-tensor onto the device so peak host memory stays low.

Requirements

  • .NET 10 SDK (tested with 10.0.203). C# 14 / LangVersion latest.
  • Windows: the build references TorchSharp-cuda-windows (libtorch 2.10 + CUDA 12.8 bundled — a ~3 GB NuGet download on first restore; no separate CUDA toolkit install needed). Works without an NVIDIA GPU too (runs on CPU).
  • Linux/macOS: the build references TorchSharp-cpu. (Swap in TorchSharp-cuda-linux for Linux GPU support.)

Build, test, run

dotnet build WhisperSharp.slnx -c Release

# Self-test: verifies the torch backend (rfft vs direct DFT, layernorm/GELU/linear/conv1d vs
# naive references, mel frontend invariants, tokenizer special-token layout, safetensors parsing)
# and runs the full encoder→decoder pipeline on a tiny synthetic model, including a
# KV-cache == full-forward equivalence check. Runs on CPU/FP32 for determinism.
dotnet run -c Release --project tests/WhisperSharp.SelfTest

Getting a model

WhisperSharp consumes a directory containing config.json, a single *.safetensors weights file, vocab.json, and merges.txt — i.e. a standard Hugging Face Whisper repo.

# Option A: built-in downloader
dotnet run -c Release --project src/WhisperSharp.Cli -- download openai/whisper-base models/whisper-base

# Option B: huggingface-cli
huggingface-cli download openai/whisper-base --local-dir models/whisper-base

Supported checkpoints include openai/whisper-tiny[.en], whisper-base[.en], whisper-small[.en], whisper-medium[.en], whisper-large-v2, and whisper-large-v3 (128 mel bins, 100 languages).

Library API

using WhisperSharp;
using WhisperSharp.Decoding;

// Device/precision are auto-selected; override explicitly if needed:
using var whisper = Whisper.LoadFromDirectory("models/whisper-small", new WhisperOptions
{
    Device = "cuda",   // "cuda", "cpu", or null/"auto"
    Fp16 = true,       // null = FP16 on CUDA, FP32 on CPU
});

Console.WriteLine($"Running on {whisper.Device} (fp16: {whisper.Fp16})");

var result = whisper.Transcribe("interview.wav", new DecodingOptions
{
    Language = null,        // auto-detect
    Translate = false,
    Timestamps = true,
});

Console.WriteLine($"Detected language: {result.Language}");
foreach (var seg in result.Segments)
    Console.WriteLine($"[{seg.Start:F2}-{seg.End:F2}] {seg.Text}");

File.WriteAllText("interview.srt", result.ToSrt());

Whisper owns native GPU/CPU tensors — dispose it (e.g. using var) to release them deterministically.

Project layout

src/WhisperSharp/
  Numerics/      torch device/dtype helpers, CPU softmax/argmax for logit rows
  Audio/         WAV reader, mel filterbank, torch STFT log-mel spectrogram
  Tokenization/  byte-level BPE, Whisper special-token layout, language list
  Serialization/ safetensors reader, config, HF downloader
  Models/        Linear/LayerNorm/Conv1d, attention, blocks, encoder, decoder, loader (TorchSharp)
  Decoding/      options, logit filters (timestamp rules), greedy decoder
  Transcription/ 30s seek loop, segmentation, SRT
  Whisper.cs     high-level facade (device/precision selection)
src/WhisperSharp.Cli/    command-line interface
tests/WhisperSharp.SelfTest/  verification harness (CPU/FP32)

Implementation notes

  • Attention scaling. Hugging Face scales queries by head_dim^-0.5; OpenAI splits the scale as head_dim^-0.25 on both queries and keys. These are mathematically identical, and the loader targets HF weights, so the single-scale form is used.
  • FP16 numerics. Weights and matmuls run in FP16 on CUDA; LayerNorm (with FP32 parameters) and the attention softmax are evaluated in FP32 and cast back, mirroring whisper/model.py. The mel spectrogram is computed in FP32 and cast to the model dtype per window.
  • Decode loop. The transformer forward pass stays on the device; only the single vocab-sized logit row per step is copied to the CPU, where the timestamp-grammar logit filters and argmax run unchanged.
  • Memory. Intermediate tensors are managed with TorchSharp dispose scopes; the KV cache is pre-allocated on the device per 30-second window and freed deterministically.
  • Position embeddings. Loaded directly from the checkpoint (embed_positions.weight); the encoder's are the precomputed sinusoids, the decoder's are learned.
  • Output projection is tied to the token embedding unless a separate proj_out.weight exists.

Limitations / not (yet) implemented

  • Greedy decoding only (temperature 0). No beam search or the temperature-fallback ladder, so there is no automatic re-decode on low-confidence/high-compression-ratio windows.
  • No condition_on_previous_text prompting across windows.
  • No word-level timestamps (cross-attention DTW alignment).
  • Single-file safetensors checkpoints only (no sharded *.safetensors.index.json).
  • WAV input only; transcode other formats to WAV first (any rate/channel count is accepted and resampled to 16 kHz mono).
  • Batch size 1 (one audio stream at a time).

License & attribution

This is an independent reimplementation for educational and practical use. Whisper, the model weights, and the reference implementation are © OpenAI (MIT-licensed).

iamdroppy/WhisperSharp

Whisper implementation in C#

C#

0

3 commits

updated Sep 21, 2026

See the code

README

WhisperSharp

A from-scratch implementation of OpenAI's Whisper speech-recognition model (Robust Speech Recognition via Large-Scale Weak Supervision, Radford et al. 2022) in C# 14 / .NET 10, running on TorchSharp / libtorch with CUDA GPU acceleration (FP16) and automatic CPU fallback.

The architecture, audio frontend, tokenizer, weight loader, and decoding loop are all implemented in this repository — the only runtime dependency is TorchSharp, which supplies the tensor kernels (cuBLAS/cuDNN on GPU). The only inputs are a standard Hugging Face Whisper checkpoint and a WAV file.

using var whisper = Whisper.LoadFromDirectory("models/whisper-base"); // CUDA + FP16 if available
var result = whisper.Transcribe("speech.wav");
Console.WriteLine(result.Text);

Features

  • GPU inference — all model math (convs, attention, MLPs, mel STFT) runs on the torch device. Defaults to CUDA with FP16 weights/compute when an NVIDIA GPU is present, FP32 on CPU otherwise; override via WhisperOptions. LayerNorm and attention softmax are computed in FP32 (as in the reference implementation) so FP16 output stays stable.
  • Faithful architecture — 2× GELU conv frontend, sinusoidal/learned position embeddings, pre-norm residual attention blocks, causal decoder self-attention, cross-attention to audio, tied output projection. Matches whisper/model.py and Hugging Face WhisperModel.
  • Exact audio frontend — centered STFT (reflect padding, periodic Hann, n_fft=400, hop=160) via torch.fft.rfft over windowed frames (chunked so hours-long audio stays memory-bounded), a librosa-compatible Slaney mel filterbank (80 or 128 bins), and Whisper's log10/clamp/rescale normalization.
  • Real tokenizer — GPT-2 byte-level BPE with the GPT-2 pre-tokenization regex, plus Whisper's structured special tokens (start-of-transcript, 99/100 language tokens, task tokens, 1501 timestamp tokens) derived exactly for English-only, multilingual, and large-v3 vocabularies.
  • Greedy decoding with timestamps — language detection, blank/non-speech suppression, the full timestamp-rule logit grammar, an incremental on-device KV cache, and the 30-second seek loop that stitches timestamped segments (text / segments / SRT output).
  • Weight loading — reads Hugging Face safetensors (F32 / F16 / BF16), streamed tensor-by-tensor onto the device so peak host memory stays low.

Requirements

  • .NET 10 SDK (tested with 10.0.203). C# 14 / LangVersion latest.
  • Windows: the build references TorchSharp-cuda-windows (libtorch 2.10 + CUDA 12.8 bundled — a ~3 GB NuGet download on first restore; no separate CUDA toolkit install needed). Works without an NVIDIA GPU too (runs on CPU).
  • Linux/macOS: the build references TorchSharp-cpu. (Swap in TorchSharp-cuda-linux for Linux GPU support.)

Build, test, run

dotnet build WhisperSharp.slnx -c Release

# Self-test: verifies the torch backend (rfft vs direct DFT, layernorm/GELU/linear/conv1d vs
# naive references, mel frontend invariants, tokenizer special-token layout, safetensors parsing)
# and runs the full encoder→decoder pipeline on a tiny synthetic model, including a
# KV-cache == full-forward equivalence check. Runs on CPU/FP32 for determinism.
dotnet run -c Release --project tests/WhisperSharp.SelfTest

Getting a model

WhisperSharp consumes a directory containing config.json, a single *.safetensors weights file, vocab.json, and merges.txt — i.e. a standard Hugging Face Whisper repo.

# Option A: built-in downloader
dotnet run -c Release --project src/WhisperSharp.Cli -- download openai/whisper-base models/whisper-base

# Option B: huggingface-cli
huggingface-cli download openai/whisper-base --local-dir models/whisper-base

Supported checkpoints include openai/whisper-tiny[.en], whisper-base[.en], whisper-small[.en], whisper-medium[.en], whisper-large-v2, and whisper-large-v3 (128 mel bins, 100 languages).

Library API

using WhisperSharp;
using WhisperSharp.Decoding;

// Device/precision are auto-selected; override explicitly if needed:
using var whisper = Whisper.LoadFromDirectory("models/whisper-small", new WhisperOptions
{
    Device = "cuda",   // "cuda", "cpu", or null/"auto"
    Fp16 = true,       // null = FP16 on CUDA, FP32 on CPU
});

Console.WriteLine($"Running on {whisper.Device} (fp16: {whisper.Fp16})");

var result = whisper.Transcribe("interview.wav", new DecodingOptions
{
    Language = null,        // auto-detect
    Translate = false,
    Timestamps = true,
});

Console.WriteLine($"Detected language: {result.Language}");
foreach (var seg in result.Segments)
    Console.WriteLine($"[{seg.Start:F2}-{seg.End:F2}] {seg.Text}");

File.WriteAllText("interview.srt", result.ToSrt());

Whisper owns native GPU/CPU tensors — dispose it (e.g. using var) to release them deterministically.

Project layout

src/WhisperSharp/
  Numerics/      torch device/dtype helpers, CPU softmax/argmax for logit rows
  Audio/         WAV reader, mel filterbank, torch STFT log-mel spectrogram
  Tokenization/  byte-level BPE, Whisper special-token layout, language list
  Serialization/ safetensors reader, config, HF downloader
  Models/        Linear/LayerNorm/Conv1d, attention, blocks, encoder, decoder, loader (TorchSharp)
  Decoding/      options, logit filters (timestamp rules), greedy decoder
  Transcription/ 30s seek loop, segmentation, SRT
  Whisper.cs     high-level facade (device/precision selection)
src/WhisperSharp.Cli/    command-line interface
tests/WhisperSharp.SelfTest/  verification harness (CPU/FP32)

Implementation notes

  • Attention scaling. Hugging Face scales queries by head_dim^-0.5; OpenAI splits the scale as head_dim^-0.25 on both queries and keys. These are mathematically identical, and the loader targets HF weights, so the single-scale form is used.
  • FP16 numerics. Weights and matmuls run in FP16 on CUDA; LayerNorm (with FP32 parameters) and the attention softmax are evaluated in FP32 and cast back, mirroring whisper/model.py. The mel spectrogram is computed in FP32 and cast to the model dtype per window.
  • Decode loop. The transformer forward pass stays on the device; only the single vocab-sized logit row per step is copied to the CPU, where the timestamp-grammar logit filters and argmax run unchanged.
  • Memory. Intermediate tensors are managed with TorchSharp dispose scopes; the KV cache is pre-allocated on the device per 30-second window and freed deterministically.
  • Position embeddings. Loaded directly from the checkpoint (embed_positions.weight); the encoder's are the precomputed sinusoids, the decoder's are learned.
  • Output projection is tied to the token embedding unless a separate proj_out.weight exists.

Limitations / not (yet) implemented

  • Greedy decoding only (temperature 0). No beam search or the temperature-fallback ladder, so there is no automatic re-decode on low-confidence/high-compression-ratio windows.
  • No condition_on_previous_text prompting across windows.
  • No word-level timestamps (cross-attention DTW alignment).
  • Single-file safetensors checkpoints only (no sharded *.safetensors.index.json).
  • WAV input only; transcode other formats to WAV first (any rate/channel count is accepted and resampled to 16 kHz mono).
  • Batch size 1 (one audio stream at a time).

License & attribution

This is an independent reimplementation for educational and practical use. Whisper, the model weights, and the reference implementation are © OpenAI (MIT-licensed).