luisquintanilla/mlnet-audio-custom-transforms

ML.NET custom transforms for audio AI: classification, embeddings, VAD, speech-to-text, text-to-speech using local ONNX models

C#

0

18 commits

updated Mar 24, 2026

See the code

README

ML.NET Audio Custom Transforms

Multi-task audio inference transforms for ML.NET using local ONNX models. Brings the same patterns from text-based ML.NET transforms to the audio domain — classification, embeddings, speech-to-text, text-to-speech, and voice activity detection.

New to audio ML? Start with the Audio Processing Primer — no prior audio knowledge required.

Packages

PackageDescriptionKey Dependencies
MLNet.Audio.CoreAudio primitives: AudioData, WAV I/O, mel spectrogram, WhisperTokenizerNWaves, System.Numerics.Tensors
MLNet.Audio.TokenizersText tokenizer extensions for audio models: SentencePieceCharTokenizerMicrosoft.ML.Tokenizers
MLNet.AudioInference.OnnxML.NET transforms: classification, embeddings, VAD, raw ONNX ASR/TTSMicrosoft.ML, OnnxRuntime, ML.Tokenizers, MEAI
MLNet.ASR.OnnxGenAILocal Whisper speech-to-text via ORT GenAIMicrosoft.ML, OnnxRuntimeGenAI, MEAI
MLNet.Audio.DataIngestionDataIngestion components: audio document reader, chunker, embedding processorDataIngestion.Abstractions, MEAI, Audio.Core

Supported Audio Tasks

TaskStatusKey TypesMEAI Interface
Audio Classification✅OnnxAudioClassificationTransformer—
Audio Embeddings✅OnnxAudioEmbeddingTransformer, OnnxAudioEmbeddingGeneratorIEmbeddingGenerator<AudioData, Embedding<float>>
Voice Activity Detection✅OnnxVadTransformerIVoiceActivityDetector (custom)
Speech-to-Text (Provider)✅SpeechToTextClientTransformerISpeechToTextClient (any provider)
Speech-to-Text (ORT GenAI)✅OnnxSpeechToTextTransformer, OnnxSpeechToTextClientISpeechToTextClient
Speech-to-Text (Raw ONNX)✅OnnxWhisperTransformer, OnnxWhisperSpeechToTextClientISpeechToTextClient
Text-to-Speech (SpeechT5)✅OnnxSpeechT5TtsTransformer, OnnxTextToSpeechClientITextToSpeechClient
Text-to-Speech (KittenTTS)✅OnnxTextToSpeechClient via OnnxKittenTtsOptionsITextToSpeechClient

Quick Start

using Microsoft.ML;
using MLNet.Audio.Core;
using MLNet.AudioInference.Onnx;

var mlContext = new MLContext();
var audio = AudioIO.LoadWav("audio.wav");

// --- Audio Classification (AST) — full Fit/Transform pattern ---
var options = new OnnxAudioClassificationOptions
{
    ModelPath = "models/ast/onnx/model.onnx",
    FeatureExtractor = new MelSpectrogramExtractor(16000) { NumMelBins = 128 },
    Labels = new[] { "Speech", "Music", "Silence" }
};

// Load data
var data = mlContext.Data.LoadFromEnumerable(new[] { new AudioInput { Audio = audio.Samples } });

// Create pipeline, fit, and transform
var pipeline = mlContext.Transforms.OnnxAudioClassification(options);
var model = pipeline.Fit(data);
var output = model.Transform(data);

// Enumerate results
var results = mlContext.Data.CreateEnumerable<ClassificationOutput>(output, reuseRowObject: false);

The other tasks follow the same Fit/Transform pattern. Pipeline creation for each:

// --- Audio Embeddings (CLAP) ---
var pipeline = mlContext.Transforms.OnnxAudioEmbedding(new OnnxAudioEmbeddingOptions
{
    ModelPath = "models/clap/onnx/model.onnx",
    FeatureExtractor = new MelSpectrogramExtractor(16000),
    Pooling = AudioPoolingStrategy.MeanPooling,
    Normalize = true
});

// --- Voice Activity Detection (Silero) ---
var pipeline = mlContext.Transforms.OnnxVad(new OnnxVadOptions
{
    ModelPath = "models/silero-vad/silero_vad.onnx",
    Threshold = 0.5f
});

// --- Speech-to-Text (provider-agnostic) ---
ISpeechToTextClient sttClient = /* Azure, OpenAI, local, etc. */;
var pipeline = mlContext.Transforms.SpeechToText(sttClient);

// --- Speech-to-Text (Raw ONNX Whisper via ISpeechToTextClient) ---
ISpeechToTextClient whisperClient = new OnnxWhisperSpeechToTextClient(whisperOptions);
var pipeline = mlContext.Transforms.SpeechToText(whisperClient);
// or: mlContext.Transforms.OnnxWhisperSpeechToText(whisperOptions);

// --- Speech-to-Text (Raw ONNX Whisper — direct transformer) ---
var pipeline = mlContext.Transforms.OnnxWhisper(new OnnxWhisperOptions
{
    EncoderModelPath = "models/whisper-base/encoder_model.onnx",
    DecoderModelPath = "models/whisper-base/decoder_model_merged.onnx",
    Language = "en"
});

// --- Text-to-Speech (SpeechT5) ---
var pipeline = mlContext.Transforms.SpeechT5Tts(new OnnxSpeechT5Options
{
    EncoderModelPath = "models/speecht5/encoder_model.onnx",
    DecoderModelPath = "models/speecht5/decoder_model_merged.onnx",
    VocoderModelPath = "models/speecht5/decoder_postnet_and_vocoder.onnx",
});

// --- Text-to-Speech (KittenTTS — lightweight single-model) ---
var pipeline = mlContext.Transforms.KittenTts(new OnnxKittenTtsOptions
{
    ModelPath = "models/kittentts/model.onnx"
});

MEAI Integration

Integrates with Microsoft.Extensions.AI:

// Audio Embeddings via MEAI
var estimator = mlContext.Transforms.OnnxAudioEmbedding(embeddingOptions);
var transformer = estimator.Fit(mlContext.Data.LoadFromEnumerable(Array.Empty<AudioInput>()));
IEmbeddingGenerator<AudioData, Embedding<float>> generator =
    new OnnxAudioEmbeddingGenerator(transformer);
var embeddings = await generator.GenerateAsync([audio]);

// Speech-to-Text via MEAI (ORT GenAI)
ISpeechToTextClient sttClient = new OnnxSpeechToTextClient(sttOptions);
var response = await sttClient.GetTextAsync(audioStream);

// Speech-to-Text via MEAI (Raw ONNX — no ORT GenAI dep)
ISpeechToTextClient rawClient = new OnnxWhisperSpeechToTextClient(whisperOptions);
var response = await rawClient.GetTextAsync(audioStream);
Console.WriteLine($"{response.Text} [{response.StartTime} → {response.EndTime}]");

// Speech-to-Text with middleware pipeline
var client = new OnnxSpeechToTextClient(sttOptions)
    .AsBuilder()
    .UseLogging()
    .UseOpenTelemetry()
    .Build();

// Text-to-Speech via official MEAI ITextToSpeechClient
ITextToSpeechClient ttsClient = new OnnxTextToSpeechClient(ttsOptions);
var response = await ttsClient.GetAudioAsync("Hello, world!");
var audioContent = response.Contents.OfType<DataContent>().First();
File.WriteAllBytes("output.wav", audioContent.Data.ToArray());

See MEAI Integration Guide for DI patterns and middleware.

DataIngestion Integration

Integrates with Microsoft.Extensions.DataIngestion to prove that DataIngestion is modality-agnostic — not just for text/PDF:

using MLNet.Audio.DataIngestion;

// Layer 3: DataIngestion — Read audio files into documents
var reader = new AudioDocumentReader(targetSampleRate: 16000);
var doc = await reader.ReadAsync(stream, "audio.wav", "audio/wav");

// Layer 3: DataIngestion — Chunk into fixed time-windows
var chunker = new AudioSegmentChunker(segmentDuration: TimeSpan.FromSeconds(2));
var chunks = chunker.ProcessAsync(doc);

// Layer 3: DataIngestion — Enrich with embeddings via MEAI → ML.NET
var processor = new AudioEmbeddingChunkProcessor(generator);
await foreach (var chunk in processor.ProcessAsync(chunks))
{
    var embedding = (float[])chunk.Metadata["embedding"];
    // Use for similarity search, clustering, RAG, etc.
}

See Architecture Guide for the full layered design.

Samples

SampleTaskDescription
AudioClassificationClassificationClassify audio using AST (Audio Spectrogram Transformer)
AudioEmbeddingsEmbeddingsGenerate vector embeddings + cosine similarity
VoiceActivityDetectionVADDetect speech segments using Silero VAD
SpeechToTextSTTProvider-agnostic ASR patterns + multi-modal pipeline
WhisperTranscriptionSTTLocal Whisper via ORT GenAI
WhisperRawOnnxSTTFull-control Whisper with manual KV cache
TextToSpeechTTSSpeechT5 encoder-decoder-vocoder synthesis
KittenTTSTTSLightweight KittenTTS with espeak-ng phonemization
AudioDataIngestionDataIngestionEnd-to-end Read → Chunk → Embed → Similarity Search

All samples run without models — they show API patterns and download instructions as graceful fallback.

Architecture

Layer 1 (ML.NET):         Audio (PCM) → Feature Extraction → ONNX Scoring → Post-processing → Result
Layer 2 (MEAI):           IEmbeddingGenerator<AudioData, Embedding<float>> / ISpeechToTextClient / ITextToSpeechClient
Layer 3 (DataIngestion):  AudioDocumentReader → AudioSegmentChunker → AudioEmbeddingChunkProcessor

Three-stage pipeline pattern mirroring the text transform architecture. See Architecture Guide.

  • Encoder-only (classification, embeddings, VAD): single-pass, mel → ONNX → result. Uses a composed 3-stage pattern with lazy IDataView wrappers (feature extraction → scoring → post-processing as separate ITransformer stages).
  • Encoder-decoder (Whisper ASR): mel → encoder → decoder loop with KV cache → text. Stays monolithic — the autoregressive decode loop cannot be split across lazy stages.
  • Encoder-decoder-vocoder (SpeechT5 TTS): tokens → encoder → decoder loop with KV cache → mel → vocoder → audio. Also monolithic due to the sequential decode loop.

.NET Primitives Used

PrimitiveWhere Used
System.Numerics.Tensors / TensorPrimitivesSoftmax, normalization, mel features, argmax, temperature sampling
Microsoft.ML.Tokenizers (SentencePiece)SpeechT5 text tokenization (with SentencePieceCharTokenizer fallback from Audio.Tokenizers for Char models)
Microsoft.Extensions.AIIEmbeddingGenerator, ISpeechToTextClient, ITextToSpeechClient
Microsoft.Extensions.DataIngestionIngestionDocumentReader, IngestionChunker<AudioData>, IngestionChunkProcessor<AudioData>
Custom WhisperTokenizerWhisper BPE + timestamps + language codes
AudioFeatureExtractor (abstract)Audio's equivalent of Tokenizer
AudioDataCore audio type: float[] samples + sample rate + channels

Documentation

📖 Full Documentation — Architecture, transforms guide, audio primer, MEAI integration, models guide, extending the framework, and samples walkthrough.

Prerequisites

  • .NET 10 SDK
  • ONNX models from HuggingFace (see Models Guide)

Getting Started

Codespaces / DevContainer

Open in GitHub Codespaces for a pre-configured environment with .NET 10, Python, huggingface-cli (for model downloads), espeak-ng (for KittenTTS phonemization), and C# Dev Kit.

NuGet Packages

Published to GitHub Packages:

  • MLNet.Audio.Core
  • MLNet.Audio.Tokenizers
  • MLNet.AudioInference.Onnx
  • MLNet.ASR.OnnxGenAI
  • MLNet.Audio.DataIngestion

Add the GitHub Packages source to your nuget.config:

<add key="github" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
ProjectDescription
mlnet-text-inference-custom-transformsText-based ML.NET transforms (same architectural patterns)
model-packages-prototypeModelPackages SDK for NuGet-wrapped AI models
dotnet-model-garden-prototypeModel garden with pre-packaged AI models
dotnet-tokenizers-guideMicrosoft.ML.Tokenizers guide
dotnet-tensors-guideSystem.Numerics.Tensors guide

luisquintanilla/mlnet-audio-custom-transforms

ML.NET custom transforms for audio AI: classification, embeddings, VAD, speech-to-text, text-to-speech using local ONNX models

C#

0

18 commits

updated Mar 24, 2026

See the code

README

ML.NET Audio Custom Transforms

Multi-task audio inference transforms for ML.NET using local ONNX models. Brings the same patterns from text-based ML.NET transforms to the audio domain — classification, embeddings, speech-to-text, text-to-speech, and voice activity detection.

New to audio ML? Start with the Audio Processing Primer — no prior audio knowledge required.

Packages

PackageDescriptionKey Dependencies
MLNet.Audio.CoreAudio primitives: AudioData, WAV I/O, mel spectrogram, WhisperTokenizerNWaves, System.Numerics.Tensors
MLNet.Audio.TokenizersText tokenizer extensions for audio models: SentencePieceCharTokenizerMicrosoft.ML.Tokenizers
MLNet.AudioInference.OnnxML.NET transforms: classification, embeddings, VAD, raw ONNX ASR/TTSMicrosoft.ML, OnnxRuntime, ML.Tokenizers, MEAI
MLNet.ASR.OnnxGenAILocal Whisper speech-to-text via ORT GenAIMicrosoft.ML, OnnxRuntimeGenAI, MEAI
MLNet.Audio.DataIngestionDataIngestion components: audio document reader, chunker, embedding processorDataIngestion.Abstractions, MEAI, Audio.Core

Supported Audio Tasks

TaskStatusKey TypesMEAI Interface
Audio Classification✅OnnxAudioClassificationTransformer—
Audio Embeddings✅OnnxAudioEmbeddingTransformer, OnnxAudioEmbeddingGeneratorIEmbeddingGenerator<AudioData, Embedding<float>>
Voice Activity Detection✅OnnxVadTransformerIVoiceActivityDetector (custom)
Speech-to-Text (Provider)✅SpeechToTextClientTransformerISpeechToTextClient (any provider)
Speech-to-Text (ORT GenAI)✅OnnxSpeechToTextTransformer, OnnxSpeechToTextClientISpeechToTextClient
Speech-to-Text (Raw ONNX)✅OnnxWhisperTransformer, OnnxWhisperSpeechToTextClientISpeechToTextClient
Text-to-Speech (SpeechT5)✅OnnxSpeechT5TtsTransformer, OnnxTextToSpeechClientITextToSpeechClient
Text-to-Speech (KittenTTS)✅OnnxTextToSpeechClient via OnnxKittenTtsOptionsITextToSpeechClient

Quick Start

using Microsoft.ML;
using MLNet.Audio.Core;
using MLNet.AudioInference.Onnx;

var mlContext = new MLContext();
var audio = AudioIO.LoadWav("audio.wav");

// --- Audio Classification (AST) — full Fit/Transform pattern ---
var options = new OnnxAudioClassificationOptions
{
    ModelPath = "models/ast/onnx/model.onnx",
    FeatureExtractor = new MelSpectrogramExtractor(16000) { NumMelBins = 128 },
    Labels = new[] { "Speech", "Music", "Silence" }
};

// Load data
var data = mlContext.Data.LoadFromEnumerable(new[] { new AudioInput { Audio = audio.Samples } });

// Create pipeline, fit, and transform
var pipeline = mlContext.Transforms.OnnxAudioClassification(options);
var model = pipeline.Fit(data);
var output = model.Transform(data);

// Enumerate results
var results = mlContext.Data.CreateEnumerable<ClassificationOutput>(output, reuseRowObject: false);

The other tasks follow the same Fit/Transform pattern. Pipeline creation for each:

// --- Audio Embeddings (CLAP) ---
var pipeline = mlContext.Transforms.OnnxAudioEmbedding(new OnnxAudioEmbeddingOptions
{
    ModelPath = "models/clap/onnx/model.onnx",
    FeatureExtractor = new MelSpectrogramExtractor(16000),
    Pooling = AudioPoolingStrategy.MeanPooling,
    Normalize = true
});

// --- Voice Activity Detection (Silero) ---
var pipeline = mlContext.Transforms.OnnxVad(new OnnxVadOptions
{
    ModelPath = "models/silero-vad/silero_vad.onnx",
    Threshold = 0.5f
});

// --- Speech-to-Text (provider-agnostic) ---
ISpeechToTextClient sttClient = /* Azure, OpenAI, local, etc. */;
var pipeline = mlContext.Transforms.SpeechToText(sttClient);

// --- Speech-to-Text (Raw ONNX Whisper via ISpeechToTextClient) ---
ISpeechToTextClient whisperClient = new OnnxWhisperSpeechToTextClient(whisperOptions);
var pipeline = mlContext.Transforms.SpeechToText(whisperClient);
// or: mlContext.Transforms.OnnxWhisperSpeechToText(whisperOptions);

// --- Speech-to-Text (Raw ONNX Whisper — direct transformer) ---
var pipeline = mlContext.Transforms.OnnxWhisper(new OnnxWhisperOptions
{
    EncoderModelPath = "models/whisper-base/encoder_model.onnx",
    DecoderModelPath = "models/whisper-base/decoder_model_merged.onnx",
    Language = "en"
});

// --- Text-to-Speech (SpeechT5) ---
var pipeline = mlContext.Transforms.SpeechT5Tts(new OnnxSpeechT5Options
{
    EncoderModelPath = "models/speecht5/encoder_model.onnx",
    DecoderModelPath = "models/speecht5/decoder_model_merged.onnx",
    VocoderModelPath = "models/speecht5/decoder_postnet_and_vocoder.onnx",
});

// --- Text-to-Speech (KittenTTS — lightweight single-model) ---
var pipeline = mlContext.Transforms.KittenTts(new OnnxKittenTtsOptions
{
    ModelPath = "models/kittentts/model.onnx"
});

MEAI Integration

Integrates with Microsoft.Extensions.AI:

// Audio Embeddings via MEAI
var estimator = mlContext.Transforms.OnnxAudioEmbedding(embeddingOptions);
var transformer = estimator.Fit(mlContext.Data.LoadFromEnumerable(Array.Empty<AudioInput>()));
IEmbeddingGenerator<AudioData, Embedding<float>> generator =
    new OnnxAudioEmbeddingGenerator(transformer);
var embeddings = await generator.GenerateAsync([audio]);

// Speech-to-Text via MEAI (ORT GenAI)
ISpeechToTextClient sttClient = new OnnxSpeechToTextClient(sttOptions);
var response = await sttClient.GetTextAsync(audioStream);

// Speech-to-Text via MEAI (Raw ONNX — no ORT GenAI dep)
ISpeechToTextClient rawClient = new OnnxWhisperSpeechToTextClient(whisperOptions);
var response = await rawClient.GetTextAsync(audioStream);
Console.WriteLine($"{response.Text} [{response.StartTime} → {response.EndTime}]");

// Speech-to-Text with middleware pipeline
var client = new OnnxSpeechToTextClient(sttOptions)
    .AsBuilder()
    .UseLogging()
    .UseOpenTelemetry()
    .Build();

// Text-to-Speech via official MEAI ITextToSpeechClient
ITextToSpeechClient ttsClient = new OnnxTextToSpeechClient(ttsOptions);
var response = await ttsClient.GetAudioAsync("Hello, world!");
var audioContent = response.Contents.OfType<DataContent>().First();
File.WriteAllBytes("output.wav", audioContent.Data.ToArray());

See MEAI Integration Guide for DI patterns and middleware.

DataIngestion Integration

Integrates with Microsoft.Extensions.DataIngestion to prove that DataIngestion is modality-agnostic — not just for text/PDF:

using MLNet.Audio.DataIngestion;

// Layer 3: DataIngestion — Read audio files into documents
var reader = new AudioDocumentReader(targetSampleRate: 16000);
var doc = await reader.ReadAsync(stream, "audio.wav", "audio/wav");

// Layer 3: DataIngestion — Chunk into fixed time-windows
var chunker = new AudioSegmentChunker(segmentDuration: TimeSpan.FromSeconds(2));
var chunks = chunker.ProcessAsync(doc);

// Layer 3: DataIngestion — Enrich with embeddings via MEAI → ML.NET
var processor = new AudioEmbeddingChunkProcessor(generator);
await foreach (var chunk in processor.ProcessAsync(chunks))
{
    var embedding = (float[])chunk.Metadata["embedding"];
    // Use for similarity search, clustering, RAG, etc.
}

See Architecture Guide for the full layered design.

Samples

SampleTaskDescription
AudioClassificationClassificationClassify audio using AST (Audio Spectrogram Transformer)
AudioEmbeddingsEmbeddingsGenerate vector embeddings + cosine similarity
VoiceActivityDetectionVADDetect speech segments using Silero VAD
SpeechToTextSTTProvider-agnostic ASR patterns + multi-modal pipeline
WhisperTranscriptionSTTLocal Whisper via ORT GenAI
WhisperRawOnnxSTTFull-control Whisper with manual KV cache
TextToSpeechTTSSpeechT5 encoder-decoder-vocoder synthesis
KittenTTSTTSLightweight KittenTTS with espeak-ng phonemization
AudioDataIngestionDataIngestionEnd-to-end Read → Chunk → Embed → Similarity Search

All samples run without models — they show API patterns and download instructions as graceful fallback.

Architecture

Layer 1 (ML.NET):         Audio (PCM) → Feature Extraction → ONNX Scoring → Post-processing → Result
Layer 2 (MEAI):           IEmbeddingGenerator<AudioData, Embedding<float>> / ISpeechToTextClient / ITextToSpeechClient
Layer 3 (DataIngestion):  AudioDocumentReader → AudioSegmentChunker → AudioEmbeddingChunkProcessor

Three-stage pipeline pattern mirroring the text transform architecture. See Architecture Guide.

  • Encoder-only (classification, embeddings, VAD): single-pass, mel → ONNX → result. Uses a composed 3-stage pattern with lazy IDataView wrappers (feature extraction → scoring → post-processing as separate ITransformer stages).
  • Encoder-decoder (Whisper ASR): mel → encoder → decoder loop with KV cache → text. Stays monolithic — the autoregressive decode loop cannot be split across lazy stages.
  • Encoder-decoder-vocoder (SpeechT5 TTS): tokens → encoder → decoder loop with KV cache → mel → vocoder → audio. Also monolithic due to the sequential decode loop.

.NET Primitives Used

PrimitiveWhere Used
System.Numerics.Tensors / TensorPrimitivesSoftmax, normalization, mel features, argmax, temperature sampling
Microsoft.ML.Tokenizers (SentencePiece)SpeechT5 text tokenization (with SentencePieceCharTokenizer fallback from Audio.Tokenizers for Char models)
Microsoft.Extensions.AIIEmbeddingGenerator, ISpeechToTextClient, ITextToSpeechClient
Microsoft.Extensions.DataIngestionIngestionDocumentReader, IngestionChunker<AudioData>, IngestionChunkProcessor<AudioData>
Custom WhisperTokenizerWhisper BPE + timestamps + language codes
AudioFeatureExtractor (abstract)Audio's equivalent of Tokenizer
AudioDataCore audio type: float[] samples + sample rate + channels

Documentation

📖 Full Documentation — Architecture, transforms guide, audio primer, MEAI integration, models guide, extending the framework, and samples walkthrough.

Prerequisites

  • .NET 10 SDK
  • ONNX models from HuggingFace (see Models Guide)

Getting Started

Codespaces / DevContainer

Open in GitHub Codespaces for a pre-configured environment with .NET 10, Python, huggingface-cli (for model downloads), espeak-ng (for KittenTTS phonemization), and C# Dev Kit.

NuGet Packages

Published to GitHub Packages:

  • MLNet.Audio.Core
  • MLNet.Audio.Tokenizers
  • MLNet.AudioInference.Onnx
  • MLNet.ASR.OnnxGenAI
  • MLNet.Audio.DataIngestion

Add the GitHub Packages source to your nuget.config:

<add key="github" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
ProjectDescription
mlnet-text-inference-custom-transformsText-based ML.NET transforms (same architectural patterns)
model-packages-prototypeModelPackages SDK for NuGet-wrapped AI models
dotnet-model-garden-prototypeModel garden with pre-packaged AI models
dotnet-tokenizers-guideMicrosoft.ML.Tokenizers guide
dotnet-tensors-guideSystem.Numerics.Tensors guide

Languages

C#

99.6%