A from-scratch implementation of OpenAI's Whisper speech-recognition model (Robust Speech Recognition via Large-Scale Weak Supervision, Radford et al. 2022) in C# 14 / .NET 10, running on TorchSharp / libtorch with CUDA GPU acceleration (FP16) and automatic CPU fallback.
The architecture, audio frontend, tokenizer, weight loader, and decoding loop are all implemented in this repository — the only runtime dependency is TorchSharp, which supplies the tensor kernels (cuBLAS/cuDNN on GPU). The only inputs are a standard Hugging Face Whisper checkpoint and a WAV file.
using var whisper = Whisper.LoadFromDirectory("models/whisper-base"); // CUDA + FP16 if available
var result = whisper.Transcribe("speech.wav");
Console.WriteLine(result.Text);
WhisperOptions. LayerNorm and attention softmax are computed in FP32 (as in the
reference implementation) so FP16 output stays stable.whisper/model.py and Hugging Face WhisperModel.n_fft=400,
hop=160) via torch.fft.rfft over windowed frames (chunked so hours-long audio stays
memory-bounded), a librosa-compatible Slaney mel filterbank (80 or 128 bins), and Whisper's
log10/clamp/rescale normalization.segments / SRT output).safetensors (F32 / F16 / BF16), streamed
tensor-by-tensor onto the device so peak host memory stays low.10.0.203). C# 14 / LangVersion latest.TorchSharp-cuda-windows (libtorch 2.10 + CUDA 12.8 bundled — a
~3 GB NuGet download on first restore; no separate CUDA toolkit install needed). Works
without an NVIDIA GPU too (runs on CPU).TorchSharp-cpu. (Swap in TorchSharp-cuda-linux for Linux
GPU support.)dotnet build WhisperSharp.slnx -c Release
# Self-test: verifies the torch backend (rfft vs direct DFT, layernorm/GELU/linear/conv1d vs
# naive references, mel frontend invariants, tokenizer special-token layout, safetensors parsing)
# and runs the full encoder→decoder pipeline on a tiny synthetic model, including a
# KV-cache == full-forward equivalence check. Runs on CPU/FP32 for determinism.
dotnet run -c Release --project tests/WhisperSharp.SelfTest
WhisperSharp consumes a directory containing config.json, a single *.safetensors weights file,
vocab.json, and merges.txt — i.e. a standard Hugging Face Whisper repo.
# Option A: built-in downloader
dotnet run -c Release --project src/WhisperSharp.Cli -- download openai/whisper-base models/whisper-base
# Option B: huggingface-cli
huggingface-cli download openai/whisper-base --local-dir models/whisper-base
Supported checkpoints include openai/whisper-tiny[.en], whisper-base[.en], whisper-small[.en],
whisper-medium[.en], whisper-large-v2, and whisper-large-v3 (128 mel bins, 100 languages).
using WhisperSharp;
using WhisperSharp.Decoding;
// Device/precision are auto-selected; override explicitly if needed:
using var whisper = Whisper.LoadFromDirectory("models/whisper-small", new WhisperOptions
{
Device = "cuda", // "cuda", "cpu", or null/"auto"
Fp16 = true, // null = FP16 on CUDA, FP32 on CPU
});
Console.WriteLine($"Running on {whisper.Device} (fp16: {whisper.Fp16})");
var result = whisper.Transcribe("interview.wav", new DecodingOptions
{
Language = null, // auto-detect
Translate = false,
Timestamps = true,
});
Console.WriteLine($"Detected language: {result.Language}");
foreach (var seg in result.Segments)
Console.WriteLine($"[{seg.Start:F2}-{seg.End:F2}] {seg.Text}");
File.WriteAllText("interview.srt", result.ToSrt());
Whisper owns native GPU/CPU tensors — dispose it (e.g. using var) to release them
deterministically.
src/WhisperSharp/
Numerics/ torch device/dtype helpers, CPU softmax/argmax for logit rows
Audio/ WAV reader, mel filterbank, torch STFT log-mel spectrogram
Tokenization/ byte-level BPE, Whisper special-token layout, language list
Serialization/ safetensors reader, config, HF downloader
Models/ Linear/LayerNorm/Conv1d, attention, blocks, encoder, decoder, loader (TorchSharp)
Decoding/ options, logit filters (timestamp rules), greedy decoder
Transcription/ 30s seek loop, segmentation, SRT
Whisper.cs high-level facade (device/precision selection)
src/WhisperSharp.Cli/ command-line interface
tests/WhisperSharp.SelfTest/ verification harness (CPU/FP32)
head_dim^-0.5; OpenAI splits the scale as
head_dim^-0.25 on both queries and keys. These are mathematically identical, and the loader
targets HF weights, so the single-scale form is used.whisper/model.py. The mel
spectrogram is computed in FP32 and cast to the model dtype per window.embed_positions.weight); the
encoder's are the precomputed sinusoids, the decoder's are learned.proj_out.weight exists.condition_on_previous_text prompting across windows.*.safetensors.index.json).This is an independent reimplementation for educational and practical use. Whisper, the model weights, and the reference implementation are © OpenAI (MIT-licensed).
A from-scratch implementation of OpenAI's Whisper speech-recognition model (Robust Speech Recognition via Large-Scale Weak Supervision, Radford et al. 2022) in C# 14 / .NET 10, running on TorchSharp / libtorch with CUDA GPU acceleration (FP16) and automatic CPU fallback.
The architecture, audio frontend, tokenizer, weight loader, and decoding loop are all implemented in this repository — the only runtime dependency is TorchSharp, which supplies the tensor kernels (cuBLAS/cuDNN on GPU). The only inputs are a standard Hugging Face Whisper checkpoint and a WAV file.
using var whisper = Whisper.LoadFromDirectory("models/whisper-base"); // CUDA + FP16 if available
var result = whisper.Transcribe("speech.wav");
Console.WriteLine(result.Text);
WhisperOptions. LayerNorm and attention softmax are computed in FP32 (as in the
reference implementation) so FP16 output stays stable.whisper/model.py and Hugging Face WhisperModel.n_fft=400,
hop=160) via torch.fft.rfft over windowed frames (chunked so hours-long audio stays
memory-bounded), a librosa-compatible Slaney mel filterbank (80 or 128 bins), and Whisper's
log10/clamp/rescale normalization.segments / SRT output).safetensors (F32 / F16 / BF16), streamed
tensor-by-tensor onto the device so peak host memory stays low.10.0.203). C# 14 / LangVersion latest.TorchSharp-cuda-windows (libtorch 2.10 + CUDA 12.8 bundled — a
~3 GB NuGet download on first restore; no separate CUDA toolkit install needed). Works
without an NVIDIA GPU too (runs on CPU).TorchSharp-cpu. (Swap in TorchSharp-cuda-linux for Linux
GPU support.)dotnet build WhisperSharp.slnx -c Release
# Self-test: verifies the torch backend (rfft vs direct DFT, layernorm/GELU/linear/conv1d vs
# naive references, mel frontend invariants, tokenizer special-token layout, safetensors parsing)
# and runs the full encoder→decoder pipeline on a tiny synthetic model, including a
# KV-cache == full-forward equivalence check. Runs on CPU/FP32 for determinism.
dotnet run -c Release --project tests/WhisperSharp.SelfTest
WhisperSharp consumes a directory containing config.json, a single *.safetensors weights file,
vocab.json, and merges.txt — i.e. a standard Hugging Face Whisper repo.
# Option A: built-in downloader
dotnet run -c Release --project src/WhisperSharp.Cli -- download openai/whisper-base models/whisper-base
# Option B: huggingface-cli
huggingface-cli download openai/whisper-base --local-dir models/whisper-base
Supported checkpoints include openai/whisper-tiny[.en], whisper-base[.en], whisper-small[.en],
whisper-medium[.en], whisper-large-v2, and whisper-large-v3 (128 mel bins, 100 languages).
using WhisperSharp;
using WhisperSharp.Decoding;
// Device/precision are auto-selected; override explicitly if needed:
using var whisper = Whisper.LoadFromDirectory("models/whisper-small", new WhisperOptions
{
Device = "cuda", // "cuda", "cpu", or null/"auto"
Fp16 = true, // null = FP16 on CUDA, FP32 on CPU
});
Console.WriteLine($"Running on {whisper.Device} (fp16: {whisper.Fp16})");
var result = whisper.Transcribe("interview.wav", new DecodingOptions
{
Language = null, // auto-detect
Translate = false,
Timestamps = true,
});
Console.WriteLine($"Detected language: {result.Language}");
foreach (var seg in result.Segments)
Console.WriteLine($"[{seg.Start:F2}-{seg.End:F2}] {seg.Text}");
File.WriteAllText("interview.srt", result.ToSrt());
Whisper owns native GPU/CPU tensors — dispose it (e.g. using var) to release them
deterministically.
src/WhisperSharp/
Numerics/ torch device/dtype helpers, CPU softmax/argmax for logit rows
Audio/ WAV reader, mel filterbank, torch STFT log-mel spectrogram
Tokenization/ byte-level BPE, Whisper special-token layout, language list
Serialization/ safetensors reader, config, HF downloader
Models/ Linear/LayerNorm/Conv1d, attention, blocks, encoder, decoder, loader (TorchSharp)
Decoding/ options, logit filters (timestamp rules), greedy decoder
Transcription/ 30s seek loop, segmentation, SRT
Whisper.cs high-level facade (device/precision selection)
src/WhisperSharp.Cli/ command-line interface
tests/WhisperSharp.SelfTest/ verification harness (CPU/FP32)
head_dim^-0.5; OpenAI splits the scale as
head_dim^-0.25 on both queries and keys. These are mathematically identical, and the loader
targets HF weights, so the single-scale form is used.whisper/model.py. The mel
spectrogram is computed in FP32 and cast to the model dtype per window.embed_positions.weight); the
encoder's are the precomputed sinusoids, the decoder's are learned.proj_out.weight exists.condition_on_previous_text prompting across windows.*.safetensors.index.json).This is an independent reimplementation for educational and practical use. Whisper, the model weights, and the reference implementation are © OpenAI (MIT-licensed).