luisquintanilla/mlnet-text-inference-custom-transforms

ML.NET custom transforms for text inference using ONNX encoder transformer models (embeddings, classification, NER, reranking, QA)

C#

3

54 commits

updated Sep 30, 2026

See the code

README

ML.NET Text Inference Custom Transforms

Open in GitHub Codespaces

A multi-task text inference platform for ML.NET that runs local HuggingFace ONNX encoder models. Provides a shared foundation of tokenization + ONNX scoring, with task-specific post-processing transforms for embeddings, classification, NER, reranking, question answering, and more.

                      TextTokenizerTransformer
                               │
                               ▼
                    OnnxTextModelScorerTransformer          (task-agnostic)
                               │
        ┌──────────┬───────────┼───────────┬──────────┐
        │          │           │           │          │
  EmbeddingPool SoftmaxClass SigmoidScor NerDecoding QaSpanExtract
  Transformer   Transformer  Transformer Transformer Transformer

  ChatClientTransformer (text generation — provider-agnostic, separate pipeline)

Task Status

TaskStatusPost-processorFacade
Embeddings✅ ImplementedEmbeddingPoolingTransformerOnnxTextEmbeddingEstimator
Classification✅ ImplementedSoftmaxClassificationTransformerOnnxTextClassificationEstimator
Reranking✅ ImplementedSigmoidScorerTransformerOnnxRerankerEstimator
NER✅ ImplementedNerDecodingTransformerOnnxNerEstimator
QA✅ ImplementedQaSpanExtractionTransformerOnnxQaEstimator
Text Generation✅ ImplementedChatClientTransformerN/A (provider-agnostic)
Text Generation (local)✅ ImplementedOnnxTextGenerationTransformerOnnxTextGenerationEstimator
Typed decisions (Laya English FP32)✅ ImplementedML.NET stages: PrepareDecisionInputs, ScoreOnnxDecisionModel, DecodeDecisionsOnnxTypedDecisionsEstimator

Why This Exists

This project is forked from mlnet-embedding-custom-transforms, which provides embedding generation only. This fork extends the platform to support all encoder transformer tasks — embeddings, classification, NER, reranking, and question answering — by sharing a task-agnostic tokenization and ONNX scoring foundation and adding task-specific post-processing transforms.

ML.NET has no built-in transform for modern HuggingFace encoder models (all-MiniLM-L6-v2, BGE, E5, DeBERTa, etc.). Building one is hard because ML.NET's convenient internal base classes (RowToRowTransformerBase, OneToOneTransformerBase) have private protected constructors — they can't be subclassed from external projects.

This project implements custom transforms using direct IEstimator<T> / ITransformer interfaces (Approach C from the ML.NET Custom Transformer Guide), enhanced with custom zip-based save/load for model persistence.

Features

  • Composable modular pipeline — task-agnostic transforms (TokenizeText → ScoreOnnxTextModel) plus task-specific post-processing that can be inspected, swapped, and reused
  • Convenience facades — OnnxTextEmbeddingEstimator, OnnxTextClassificationEstimator, OnnxRerankerEstimator, OnnxNerEstimator, OnnxQaEstimator each wrap all transforms for their task in a single call
  • Provider-agnostic MEAI integration — EmbeddingGeneratorEstimator wraps any IEmbeddingGenerator<string, Embedding<float>> as an ML.NET transform; ChatClientEstimator wraps any IChatClient for text generation
  • Text-pair tokenization — cross-encoder reranking uses [CLS] A [SEP] B [SEP] with token type IDs for query-document pairs
  • Token offset tracking — NER tokenization preserves character offsets via EncodeToTokens() for mapping entities back to source text
  • Multi-output ONNX scoring — QA models produce separate start/end logit tensors via AdditionalOutputTensorNames
  • Smart tokenizer resolution — point to a directory; auto-detects from tokenizer_config.json, known vocab files (BPE, SentencePiece, WordPiece), or HuggingFace tokenizer.json (fast tokenizer)
  • ONNX auto-discovery — automatically detects input/output tensor names, shapes, and dimensions from model metadata
  • Self-contained save/load — serializes to a portable .mlnet zip file containing the ONNX model, tokenizer, and config
  • SIMD-accelerated post-processing — pooling and normalization use TensorPrimitives for hardware-vectorized math
  • Configurable batching — process rows in configurable batch sizes to bound memory usage
  • Multiple pooling strategies — Mean, CLS token, and Max pooling (for embeddings)
  • Typed decisions — an ML.NET-first facade and composable stages backed by shared task-specific Laya kernels for choice, score, and Boolean questions

Typed-decision naming follows the same conventions as the other transforms. The ML.NET surface keeps the user-facing TransformsCatalog verbs and the approved OnnxTypedDecisions, PrepareDecisionInputs, ScoreOnnxDecisionModel, and DecodeDecisions names, while role types use the repository's *Options, *Estimator, and *Transformer conventions: DecisionInputPreparation*, OnnxDecisionModelScorer*, and DecisionDecoding*. The compiled ML.NET facade can be appended with AppendOnnxTypedDecisions.

Typed decisions

Typed decisions use a versioned local bundle for the English FP32 Laya graph from receptron/laya-onnx revision 68f27dfe5a27a54fb2b1fefc432f43f972e90868. The ML.NET facade is the primary entry point; the direct transformer API and inspectable native stages use the same implementation. The feature scores caller-supplied alternatives rather than generating prose: Choice selects a label, Score returns an expected zero-based option index, and Noul returns a Boolean plus the probability of the true option. For this pretrained typed-decision estimator, Fit validates an ML.NET schema and initializes resources; it does not train the ONNX model.

In this repository, the typed-decision implementation is part of the MLNet.TextInference.Onnx assembly/package. It uses Microsoft.ML.Tokenizers for the selected byte-level BPE contract, the managed/native ONNX Runtime packages for the five-input graph, and C# decoding with stable tensor primitives. State is text (including caller-serialized JSON); there is no implicit Python-compatible object serializer. Inference never downloads model assets.

The normal model-assets directory contains the model, external-data sidecar, Laya configuration, and tokenizer directory. An optional manifest-backed archive is also supported. The explicit acceptance launcher requires model assets prepared locally:

.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
  -ModelAssetsPath .\models\laya-english-fp32.bundle -Mode facade
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
  -ModelAssetsPath .\models\laya-english-fp32.bundle -Mode stages

The typed-decision samples have deliberately different jobs: the short orientation page points to the canonical guided tutorial, which owns the exact questions, file-based run path, captured output, tensor shapes, decoder semantics, native stage columns, and portable deployment boundary. The tracked PortableProcessHarness is a developer/acceptance utility for fresh-process offline persistence, not the beginner walkthrough.

The internal Laya preparation kernel loads and configures the existing Microsoft.ML.Tokenizers BPE engine, while separate profile metadata carries the special-token IDs and mask token. This keeps Microsoft.ML.Tokenizers as the encoding boundary without introducing a second tokenizer interface or reimplementing BPE.

Quick Start

Option A: GitHub Codespaces (fastest)

Click "Open in GitHub Codespaces" above. The dev container automatically:

  1. Installs .NET 10, Python 3.12, and CUDA toolkit
  2. Restores packages and builds the solution
  3. Downloads the starter model (all-MiniLM-L6-v2, ~86MB)

Once ready, run the first sample:

cd samples/BasicUsage && dotnet run

Download models for other tasks:

bash scripts/download-models.sh classification  # sentiment, emotion, zero-shot
bash scripts/download-models.sh reranking       # cross-encoder reranking
bash scripts/download-models.sh ner             # named entity recognition
bash scripts/download-models.sh qa              # question answering
bash scripts/download-models.sh all             # everything (~3.5GB)
bash scripts/download-models.sh --help          # see all options

Option B: Local setup

Prerequisites: .NET 10 SDK. Python 3.x needed only for NER/QA model export.

dotnet restore && dotnet build
bash scripts/download-models.sh embeddings-core   # downloads ~86MB starter model
cd samples/BasicUsage && dotnet run

Or download manually:

mkdir samples/BasicUsage/models
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" -OutFile "samples/BasicUsage/models/model.onnx"
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt" -OutFile "samples/BasicUsage/models/vocab.txt"

Code Examples (Embeddings)

Note: Embeddings are the simplest task. Browse the samples/ directory for classification, reranking, NER, QA, and text generation examples.

using Microsoft.ML;
using MLNet.TextInference.Onnx;

var mlContext = new MLContext();
var data = mlContext.Data.LoadFromEnumerable(new[]
{
    new { Text = "What is machine learning?" },
    new { Text = "How to cook pasta" }
});

// Step-by-step: tokenize → score → pool
var tokenizer = mlContext.Transforms.TokenizeText(new TextTokenizerOptions
{
    TokenizerPath = "models/",  // directory — auto-detects tokenizer
    InputColumnName = "Text"
}).Fit(data);

var scorer = mlContext.Transforms.ScoreOnnxTextModel(new OnnxTextModelScorerOptions
{
    ModelPath = "models/model.onnx"
}).Fit(tokenizer.Transform(data));

var pooler = mlContext.Transforms.PoolEmbedding(new EmbeddingPoolingOptions
{
    Pooling = PoolingStrategy.MeanPooling,
    Normalize = true,
    HiddenDim = scorer.HiddenDim,
    IsPrePooled = scorer.HasPooledOutput,
    SequenceLength = 128
}).Fit(scorer.Transform(tokenizer.Transform(data)));

Or chain with .Append() for the idiomatic ML.NET pattern:

var pipeline = mlContext.Transforms.TokenizeText(tokenizerOpts)
    .Append(mlContext.Transforms.ScoreOnnxTextModel(scorerOpts))
    .Append(mlContext.Transforms.PoolEmbedding(poolingOpts));
var model = pipeline.Fit(data);

Convenience facade (single-shot)

var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
});
var transformer = estimator.Fit(data);
var embeddings = transformer.Transform(data);

Provider-agnostic MEAI embedding

using Microsoft.Extensions.AI;

// Works with ANY IEmbeddingGenerator — ONNX, OpenAI, Azure, Ollama...
IEmbeddingGenerator<string, Embedding<float>> generator =
    new OnnxEmbeddingGenerator(mlContext, transformer);

var estimator = mlContext.Transforms.TextEmbedding(generator);

Save and load

// Save — bundles ONNX model + tokenizer + config into a portable zip
transformer.Save("my-embedding-model.mlnet");

// Load — fully self-contained, no external file dependencies
var loaded = OnnxTextEmbeddingTransformer.Load(mlContext, "my-embedding-model.mlnet");

Model Compatibility

The shared foundation supports any encoder transformer ONNX model. Compatible model architectures include:

  • BERT and derivatives (all-MiniLM, all-mpnet)
  • RoBERTa (including XLM-RoBERTa for multilingual)
  • DistilBERT
  • DeBERTa / DeBERTa-v2
  • MiniLM (Microsoft)
  • MPNet (Microsoft)
  • E5 (intfloat)
  • BGE (BAAI)
  • GTE (Alibaba)

Tested Embedding Models

Compatible Classification Models

Any encoder model fine-tuned for sequence classification (outputs logits over label classes):

  • Sentiment: DistilBERT-SST2, BERT-SST2
  • Emotion: RoBERTa-emotion, GoEmotions
  • NLI/Zero-shot: DeBERTa-v3-NLI, BART-large-MNLI

Compatible Reranking Models

Cross-encoder models that take text pairs and output a relevance score:

  • MS MARCO: cross-encoder/ms-marco-MiniLM-L-6-v2
  • BGE: BAAI/bge-reranker-base, bge-reranker-v2-m3

Compatible NER Models

Token classification models with BIO tagging:

  • BERT-NER: dslim/bert-base-NER
  • Multilingual: Davlan/xlm-roberta-base-ner-hrl

Compatible QA Models

Extractive QA models outputting start/end logits:

  • RoBERTa: deepset/roberta-base-squad2
  • MiniLM: deepset/minilm-uncased-squad2

Models with sentence_embedding output (pre-pooled) are auto-detected and pooling is skipped.

Project Structure

mlnet-text-inference-custom-transforms/
├── src/
│   ├── MLNet.TextInference.Onnx/
│   │   ├── TextTokenizerEstimator.cs         — Transform 1: tokenization (BPE/WordPiece/SentencePiece)
│   │   ├── OnnxTextModelScorerEstimator.cs   — Transform 2: ONNX inference with lookahead batching
│   │   ├── ChatClientEstimator.cs            — Provider-agnostic text generation (IChatClient)
│   │   ├── MLContextExtensions.cs            — Extension methods for fluent API
│   │   ├── Embeddings/
│   │   │   ├── EmbeddingPoolingEstimator.cs  — Pooling + L2 normalization
│   │   │   ├── OnnxTextEmbeddingEstimator.cs — Convenience facade
│   │   │   ├── EmbeddingGeneratorEstimator.cs— MEAI IEmbeddingGenerator wrapper
│   │   │   └── ...                           — OnnxEmbeddingGenerator, ModelPackager, PoolingStrategy
│   │   ├── Classification/
│   │   │   ├── SoftmaxClassificationEstimator.cs — Softmax post-processing
│   │   │   ├── OnnxTextClassificationEstimator.cs— Full pipeline facade
│   │   │   └── ...                              — Options, Transformer, ClassificationResult
│   │   ├── Reranking/
│   │   │   ├── SigmoidScorerEstimator.cs     — Sigmoid scoring transform
│   │   │   ├── OnnxRerankerEstimator.cs      — Cross-encoder facade
│   │   │   └── ...                           — Options, Transformer
│   │   ├── NER/
│   │   │   ├── NerDecodingEstimator.cs       — BIO tag decoding to entity spans
│   │   │   ├── OnnxNerEstimator.cs           — End-to-end NER facade
│   │   │   └── ...                           — Options, Transformer, NerEntity
│   │   └── QA/
│   │       ├── QaSpanExtractionEstimator.cs  — Span extraction from start/end logits
│   │       ├── OnnxQaEstimator.cs            — End-to-end QA facade
│   │       └── ...                           — Options, Transformer, QaResult
│   └── MLNet.TextGeneration.OnnxGenAI/
│       ├── OnnxTextGenerationEstimator.cs    — ORT GenAI local text generation
│       ├── MLContextExtensions.cs            — OnnxTextGeneration() extension method
│       └── ...                               — Options, Transformer
├── samples/
│   ├── BasicUsage/                           — all-MiniLM-L6-v2: all embedding API surfaces
│   ├── BgeSmallEmbedding/                    — BGE-small: query prefix pattern
│   ├── E5SmallEmbedding/                     — E5-small: query/passage prefixes
│   ├── GteSmallEmbedding/                    — GTE-small: semantic search (no prefix)
│   ├── ComposablePoolingComparison/          — 3 pooling strategies, shared inference
│   ├── IntermediateInspection/               — Inspect tokens, masks, raw output
│   ├── MeaiProviderAgnostic/                 — Provider-agnostic MEAI transform
│   ├── Classification/
│   │   ├── SentimentDistilBERT/              — Sentiment analysis with DistilBERT
│   │   ├── EmotionRoBERTa/                   — Multi-class emotion classification
│   │   └── ZeroShotDeBERTa/                  — Zero-shot NLI classification with DeBERTa
│   ├── Reranking/
│   │   ├── MsMarcoMiniLM/                    — MS MARCO cross-encoder reranking
│   │   └── BgeReranker/                      — BGE reranker for retrieval
│   ├── NER/
│   │   ├── BertBaseNER/                      — Named entity recognition with BERT-base
│   │   └── MultilingualNER/                  — Multilingual NER with XLM-RoBERTa
│   ├── QA/
│   │   ├── RobertaSquad2/                    — Extractive QA with RoBERTa on SQuAD 2.0
│   │   └── MiniLMSquad2/                     — Extractive QA with MiniLM on SQuAD 2.0
│   ├── TextGenerationMeai/                   — Provider-agnostic text generation (IChatClient)
│   └── TextGenerationLocal/                  — Local ORT GenAI text generation
├── docs/                                     — Detailed documentation
│   ├── architecture-decision-record.md       — ADR for multi-task platform decisions
│   ├── architecture.md                       — Component walkthrough + pipeline stages
│   ├── design-decisions.md                   — Why every choice was made
│   ├── extending.md                          — How to modify, extend, and add new tasks
│   ├── tensor-deep-dive.md                   — System.Numerics.Tensors for AI workloads
│   └── references.md                         — All sources and further reading
├── proposals/                                — Design proposals for the modular architecture
└── nuget.config                              — NuGet source (nuget.org only)

Samples

SampleModelPattern Demonstrated
BasicUsageall-MiniLM-L6-v2All API surfaces: facade, composable pipeline, .Append(), save/load, MEAI
BgeSmallEmbeddingBGE-small-en-v1.5Composable pipeline + BGE query prefix for asymmetric retrieval
E5SmallEmbeddingE5-small-v2Composable pipeline + E5 dual query:/passage: prefix pattern
GteSmallEmbeddingGTE-smallComposable pipeline + semantic search (no prefix needed)
ComposablePoolingComparisonall-MiniLM-L6-v23 pooling strategies, shared tokenizer+scorer (key modularization demo)
IntermediateInspectionall-MiniLM-L6-v2Inspect token IDs, attention masks, raw output at each pipeline stage
MeaiProviderAgnosticall-MiniLM-L6-v2EmbeddingGeneratorEstimator wrapping IEmbeddingGenerator
SentimentDistilBERTDistilBERT-SST2Sentiment classification with softmax post-processing
EmotionRoBERTaRoBERTa-emotionMulti-class emotion classification
ZeroShotDeBERTaDeBERTa-v3-NLIZero-shot classification via NLI entailment
MsMarcoMiniLMMS MARCO MiniLMCross-encoder reranking with text-pair tokenization
BgeRerankerBAAI/bge-rerankerBGE cross-encoder reranking for retrieval
BertBaseNERBERT-base-NERNamed entity recognition with BIO tag decoding
MultilingualNERXLM-RoBERTa-NERMultilingual named entity recognition
RobertaSquad2RoBERTa-SQuAD2Extractive question answering with span extraction
MiniLMSquad2MiniLM-SQuAD2Extractive question answering (lighter model)
TextGenerationMeaiAny IChatClientProvider-agnostic text generation via MEAI
TextGenerationLocalORT GenAI modelLocal text generation with ONNX Runtime GenAI

API at a Glance

Shared Foundation

ClassRoleKey Methods
TextTokenizerEstimatorTransform 1: TokenizationFit(IDataView)
OnnxTextModelScorerEstimatorTransform 2: ONNX ScoringFit(IDataView) → .HiddenDim, .HasPooledOutput

Embeddings

ClassRoleKey Methods
EmbeddingPoolingEstimatorTransform 3: PoolingFit(IDataView)
OnnxTextEmbeddingEstimatorFacade (chains 1→2→3)Fit(IDataView), GetOutputSchema()
OnnxTextEmbeddingTransformerFacade transformerTransform(IDataView), Save(path), Load(ctx, path)
EmbeddingGeneratorEstimatorMEAI wrapper transformFit(IDataView)
OnnxEmbeddingGeneratorMEAI IEmbeddingGeneratorGenerateAsync(texts)

Classification

ClassRoleKey Methods
SoftmaxClassificationEstimatorSoftmax post-processingFit(IDataView)
OnnxTextClassificationEstimatorFacade (tokenize→score→softmax)Fit(IDataView)
OnnxTextClassificationTransformerFacade transformerTransform(IDataView), Classify(texts)

Reranking

ClassRoleKey Methods
SigmoidScorerEstimatorSigmoid scoring transformFit(IDataView)
OnnxRerankerEstimatorFacade (text-pair tokenize→score→sigmoid)Fit(IDataView)
OnnxRerankerTransformerFacade transformerTransform(IDataView), Rerank(query, docs)

Named Entity Recognition

ClassRoleKey Methods
NerDecodingEstimatorBIO tag decodingFit(IDataView)
OnnxNerEstimatorFacade (tokenize→score→decode)Fit(IDataView)
OnnxNerTransformerFacade transformerTransform(IDataView), RecognizeEntities(texts)

Question Answering

ClassRoleKey Methods
QaSpanExtractionEstimatorSpan extraction from logitsFit(IDataView)
OnnxQaEstimatorFacade (tokenize→multi-score→extract)Fit(IDataView)
OnnxQaTransformerFacade transformerTransform(IDataView), Answer(questions)

Typed decisions

SurfaceTypes and extensionsRole
ML.NET facadeOnnxTypedDecisionsOptions, OnnxTypedDecisionsEstimator, OnnxTypedDecisionsTransformerSchema-aware lazy end-to-end transform
ML.NET stagesDecisionInputPreparation*, OnnxDecisionModelScorer*, DecisionDecoding*Options / Estimator / Transformer stage types
ML.NET compositionTransformsCatalog extensions and AppendOnnxTypedDecisionsStandard ML.NET pipeline composition

The ML.NET facade adds question-specific typed columns. Choice exposes PredictedLabel and a calibrated Probabilities vector; Score exposes the expected zero-based Score and its probability vector; Noul exposes Boolean PredictedLabel and Probability (P(true)). Each question also gets its own entropy Confidence and ActionProbability. DecisionResults remains an optional full diagnostic JSON column. Preparation and scoring stages use native ML.NET numeric/vector/Boolean columns, and the facade is the normal cursor-batched path. Native single-row PredictionEngine mapping is supported through the transformer and native stage row mappers. Native ML.NET Save/Load remains unsupported for typed decisions in this release; use the explicit portable API instead. External assets are a packaging consideration, not an inherent limitation of row mapping. Typed-decision transformers provide an explicit portable ZIP API (Save/Load and the supported facade/stage pipeline helpers); this is separate from native MLContext.Model.Save/Load, which remains unsupported for these custom components. The supported pipeline boundary is the demonstrated flat preparation -> scoring -> decoding chain and naturally inferred appended typed-decision facades; arbitrary/native chains are rejected. All sources in a saved pipeline must reference the same complete asset payload, so separately loaded selective stage archives cannot currently be recombined.

Text Generation

ClassRoleKey Methods
ChatClientEstimatorProvider-agnostic (wraps IChatClient)Fit(IDataView)
ChatClientTransformerProvider-agnostic transformerTransform(IDataView)
OnnxTextGenerationEstimatorORT GenAI local generationFit(IDataView)
OnnxTextGenerationTransformerORT GenAI transformerTransform(IDataView)

GPU Support (CUDA)

The library supports GPU-accelerated ONNX inference via CUDA. The library itself ships with no native binaries — you control the execution provider by choosing your OnnxRuntime package.

GPU Prerequisites

GPU inference requires the CUDA Toolkit and cuDNN installed on the host machine, plus an NVIDIA GPU with a compatible driver.

Windows (winget + direct download):

# 1. Install CUDA Toolkit 12.6
winget install Nvidia.CUDA --version 12.6 --source winget

# 2. Download and install cuDNN 9.x for CUDA 12
#    Download the zip from NVIDIA's redistributable endpoint:
$url = "https://developer.download.nvidia.com/compute/cudnn/redist/cudnn/windows-x86_64/cudnn-windows-x86_64-9.8.0.87_cuda12-archive.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:TEMP\cudnn.zip"
Expand-Archive "$env:TEMP\cudnn.zip" -DestinationPath "$env:TEMP\cudnn"

# 3. Copy cuDNN DLLs to CUDA bin (requires admin)
Copy-Item "$env:TEMP\cudnn\cudnn-*\bin\*.dll" "$env:CUDA_PATH\bin" -Force

Linux (apt):

# CUDA Toolkit
sudo apt-get install -y nvidia-cuda-toolkit

# cuDNN (via NVIDIA's apt repository)
# See: https://docs.nvidia.com/deeplearning/cudnn/installation/linux.html

Verify installation:

nvidia-smi                    # Should show GPU + driver version
nvcc --version                # Should show CUDA 12.x

Package Setup

Replace Microsoft.ML.OnnxRuntime with Microsoft.ML.OnnxRuntime.Gpu in your application:

<PackageReference Include="Microsoft.ML.OnnxRuntime.Gpu" Version="1.24.2" />

Samples auto-detect GPU: The sample projects use a Directory.Build.props that checks for CUDA_PATH (set by the CUDA Toolkit installer) and automatically switches to Microsoft.ML.OnnxRuntime.Gpu. Override with dotnet build -p:UseGpuRuntime=true or dotnet build -p:UseGpuRuntime=false.

Usage

// Pattern 1: Context-level (applies to all ONNX estimators)
var mlContext = new MLContext();
mlContext.GpuDeviceId = 0; // Use first CUDA device

var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
});

// Pattern 2: Per-estimator override
var scorerOptions = new OnnxTextModelScorerOptions
{
    ModelPath = "models/model.onnx",
    GpuDeviceId = 0,       // Override for this estimator only
    FallbackToCpu = true,  // Graceful degradation if CUDA unavailable
};

Resolution order: Per-estimator GpuDeviceId → MLContext.GpuDeviceId → null (CPU).

When FallbackToCpu = true, if CUDA initialization fails the estimator silently falls back to CPU execution.

NuGet Dependencies

MLNet.TextInference.Onnx (encoder tasks)

PackageVersionPurpose
Microsoft.ML5.0.0IEstimator/ITransformer, IDataView, MLContext
Microsoft.ML.OnnxRuntime.Managed1.24.2InferenceSession, OrtValue (managed API; bring your own native runtime)
Microsoft.ML.Tokenizers2.0.0BertTokenizer (WordPiece), BPE, SentencePiece
Microsoft.Extensions.AI.Abstractions10.3.0IEmbeddingGenerator, IChatClient
System.Numerics.Tensors10.0.3Tensor<T>, TensorPrimitives

MLNet.TextGeneration.OnnxGenAI (local text generation)

PackageVersionPurpose
Microsoft.ML5.0.0IEstimator/ITransformer, IDataView, MLContext
Microsoft.ML.OnnxRuntimeGenAI0.7.1ORT GenAI for local autoregressive generation
Microsoft.Extensions.AI.Abstractions10.3.0IChatClient

Using as a NuGet Package

Pre-built packages are published to GitHub Packages. No need to copy source code.

1. Add the GitHub Packages NuGet source (one-time setup):

<!-- nuget.config in your project root -->
<configuration>
  <packageSources>
    <add key="nuget.org" value="https://api.nuget.org/v3/index.json" />
    <add key="github-packages" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
  </packageSources>
  <packageSourceCredentials>
    <github-packages>
      <add key="Username" value="YOUR_GITHUB_USERNAME" />
      <add key="ClearTextPassword" value="YOUR_GITHUB_PAT" />
    </github-packages>
  </packageSourceCredentials>
</configuration>

Note: GitHub Packages requires authentication even for public packages. Create a Personal Access Token with read:packages scope.

2. Add the package reference:

<!-- Encoder tasks: embeddings, classification, reranking, NER, QA, text generation -->
<PackageReference Include="MLNet.TextInference.Onnx" Version="0.1.0-preview.1" />

<!-- Local text generation with ORT GenAI (separate package, heavier dependency) -->
<PackageReference Include="MLNet.TextGeneration.OnnxGenAI" Version="0.1.0-preview.1" />

3. Use it:

using MLNet.TextInference.Onnx;

var mlContext = new MLContext();
var estimator = mlContext.Transforms.OnnxTextClassification(new OnnxTextClassificationOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
    Labels = ["negative", "positive"]
});

Releasing a new version

Packages are versioned via git tags using MinVer:

git tag v0.1.0-preview.1
git push origin v0.1.0-preview.1
# → GitHub Actions builds, packs, and publishes to GitHub Packages

Documentation

Target Framework

.NET 10 (LTS). Requires the .NET 10 SDK.

License

MIT

luisquintanilla/mlnet-text-inference-custom-transforms

ML.NET custom transforms for text inference using ONNX encoder transformer models (embeddings, classification, NER, reranking, QA)

C#

3

54 commits

updated Sep 30, 2026

See the code

README

ML.NET Text Inference Custom Transforms

Open in GitHub Codespaces

A multi-task text inference platform for ML.NET that runs local HuggingFace ONNX encoder models. Provides a shared foundation of tokenization + ONNX scoring, with task-specific post-processing transforms for embeddings, classification, NER, reranking, question answering, and more.

                      TextTokenizerTransformer
                               │
                               ▼
                    OnnxTextModelScorerTransformer          (task-agnostic)
                               │
        ┌──────────┬───────────┼───────────┬──────────┐
        │          │           │           │          │
  EmbeddingPool SoftmaxClass SigmoidScor NerDecoding QaSpanExtract
  Transformer   Transformer  Transformer Transformer Transformer

  ChatClientTransformer (text generation — provider-agnostic, separate pipeline)

Task Status

TaskStatusPost-processorFacade
Embeddings✅ ImplementedEmbeddingPoolingTransformerOnnxTextEmbeddingEstimator
Classification✅ ImplementedSoftmaxClassificationTransformerOnnxTextClassificationEstimator
Reranking✅ ImplementedSigmoidScorerTransformerOnnxRerankerEstimator
NER✅ ImplementedNerDecodingTransformerOnnxNerEstimator
QA✅ ImplementedQaSpanExtractionTransformerOnnxQaEstimator
Text Generation✅ ImplementedChatClientTransformerN/A (provider-agnostic)
Text Generation (local)✅ ImplementedOnnxTextGenerationTransformerOnnxTextGenerationEstimator
Typed decisions (Laya English FP32)✅ ImplementedML.NET stages: PrepareDecisionInputs, ScoreOnnxDecisionModel, DecodeDecisionsOnnxTypedDecisionsEstimator

Why This Exists

This project is forked from mlnet-embedding-custom-transforms, which provides embedding generation only. This fork extends the platform to support all encoder transformer tasks — embeddings, classification, NER, reranking, and question answering — by sharing a task-agnostic tokenization and ONNX scoring foundation and adding task-specific post-processing transforms.

ML.NET has no built-in transform for modern HuggingFace encoder models (all-MiniLM-L6-v2, BGE, E5, DeBERTa, etc.). Building one is hard because ML.NET's convenient internal base classes (RowToRowTransformerBase, OneToOneTransformerBase) have private protected constructors — they can't be subclassed from external projects.

This project implements custom transforms using direct IEstimator<T> / ITransformer interfaces (Approach C from the ML.NET Custom Transformer Guide), enhanced with custom zip-based save/load for model persistence.

Features

  • Composable modular pipeline — task-agnostic transforms (TokenizeText → ScoreOnnxTextModel) plus task-specific post-processing that can be inspected, swapped, and reused
  • Convenience facades — OnnxTextEmbeddingEstimator, OnnxTextClassificationEstimator, OnnxRerankerEstimator, OnnxNerEstimator, OnnxQaEstimator each wrap all transforms for their task in a single call
  • Provider-agnostic MEAI integration — EmbeddingGeneratorEstimator wraps any IEmbeddingGenerator<string, Embedding<float>> as an ML.NET transform; ChatClientEstimator wraps any IChatClient for text generation
  • Text-pair tokenization — cross-encoder reranking uses [CLS] A [SEP] B [SEP] with token type IDs for query-document pairs
  • Token offset tracking — NER tokenization preserves character offsets via EncodeToTokens() for mapping entities back to source text
  • Multi-output ONNX scoring — QA models produce separate start/end logit tensors via AdditionalOutputTensorNames
  • Smart tokenizer resolution — point to a directory; auto-detects from tokenizer_config.json, known vocab files (BPE, SentencePiece, WordPiece), or HuggingFace tokenizer.json (fast tokenizer)
  • ONNX auto-discovery — automatically detects input/output tensor names, shapes, and dimensions from model metadata
  • Self-contained save/load — serializes to a portable .mlnet zip file containing the ONNX model, tokenizer, and config
  • SIMD-accelerated post-processing — pooling and normalization use TensorPrimitives for hardware-vectorized math
  • Configurable batching — process rows in configurable batch sizes to bound memory usage
  • Multiple pooling strategies — Mean, CLS token, and Max pooling (for embeddings)
  • Typed decisions — an ML.NET-first facade and composable stages backed by shared task-specific Laya kernels for choice, score, and Boolean questions

Typed-decision naming follows the same conventions as the other transforms. The ML.NET surface keeps the user-facing TransformsCatalog verbs and the approved OnnxTypedDecisions, PrepareDecisionInputs, ScoreOnnxDecisionModel, and DecodeDecisions names, while role types use the repository's *Options, *Estimator, and *Transformer conventions: DecisionInputPreparation*, OnnxDecisionModelScorer*, and DecisionDecoding*. The compiled ML.NET facade can be appended with AppendOnnxTypedDecisions.

Typed decisions

Typed decisions use a versioned local bundle for the English FP32 Laya graph from receptron/laya-onnx revision 68f27dfe5a27a54fb2b1fefc432f43f972e90868. The ML.NET facade is the primary entry point; the direct transformer API and inspectable native stages use the same implementation. The feature scores caller-supplied alternatives rather than generating prose: Choice selects a label, Score returns an expected zero-based option index, and Noul returns a Boolean plus the probability of the true option. For this pretrained typed-decision estimator, Fit validates an ML.NET schema and initializes resources; it does not train the ONNX model.

In this repository, the typed-decision implementation is part of the MLNet.TextInference.Onnx assembly/package. It uses Microsoft.ML.Tokenizers for the selected byte-level BPE contract, the managed/native ONNX Runtime packages for the five-input graph, and C# decoding with stable tensor primitives. State is text (including caller-serialized JSON); there is no implicit Python-compatible object serializer. Inference never downloads model assets.

The normal model-assets directory contains the model, external-data sidecar, Laya configuration, and tokenizer directory. An optional manifest-backed archive is also supported. The explicit acceptance launcher requires model assets prepared locally:

.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
  -ModelAssetsPath .\models\laya-english-fp32.bundle -Mode facade
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
  -ModelAssetsPath .\models\laya-english-fp32.bundle -Mode stages

The typed-decision samples have deliberately different jobs: the short orientation page points to the canonical guided tutorial, which owns the exact questions, file-based run path, captured output, tensor shapes, decoder semantics, native stage columns, and portable deployment boundary. The tracked PortableProcessHarness is a developer/acceptance utility for fresh-process offline persistence, not the beginner walkthrough.

The internal Laya preparation kernel loads and configures the existing Microsoft.ML.Tokenizers BPE engine, while separate profile metadata carries the special-token IDs and mask token. This keeps Microsoft.ML.Tokenizers as the encoding boundary without introducing a second tokenizer interface or reimplementing BPE.

Quick Start

Option A: GitHub Codespaces (fastest)

Click "Open in GitHub Codespaces" above. The dev container automatically:

  1. Installs .NET 10, Python 3.12, and CUDA toolkit
  2. Restores packages and builds the solution
  3. Downloads the starter model (all-MiniLM-L6-v2, ~86MB)

Once ready, run the first sample:

cd samples/BasicUsage && dotnet run

Download models for other tasks:

bash scripts/download-models.sh classification  # sentiment, emotion, zero-shot
bash scripts/download-models.sh reranking       # cross-encoder reranking
bash scripts/download-models.sh ner             # named entity recognition
bash scripts/download-models.sh qa              # question answering
bash scripts/download-models.sh all             # everything (~3.5GB)
bash scripts/download-models.sh --help          # see all options

Option B: Local setup

Prerequisites: .NET 10 SDK. Python 3.x needed only for NER/QA model export.

dotnet restore && dotnet build
bash scripts/download-models.sh embeddings-core   # downloads ~86MB starter model
cd samples/BasicUsage && dotnet run

Or download manually:

mkdir samples/BasicUsage/models
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" -OutFile "samples/BasicUsage/models/model.onnx"
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt" -OutFile "samples/BasicUsage/models/vocab.txt"

Code Examples (Embeddings)

Note: Embeddings are the simplest task. Browse the samples/ directory for classification, reranking, NER, QA, and text generation examples.

using Microsoft.ML;
using MLNet.TextInference.Onnx;

var mlContext = new MLContext();
var data = mlContext.Data.LoadFromEnumerable(new[]
{
    new { Text = "What is machine learning?" },
    new { Text = "How to cook pasta" }
});

// Step-by-step: tokenize → score → pool
var tokenizer = mlContext.Transforms.TokenizeText(new TextTokenizerOptions
{
    TokenizerPath = "models/",  // directory — auto-detects tokenizer
    InputColumnName = "Text"
}).Fit(data);

var scorer = mlContext.Transforms.ScoreOnnxTextModel(new OnnxTextModelScorerOptions
{
    ModelPath = "models/model.onnx"
}).Fit(tokenizer.Transform(data));

var pooler = mlContext.Transforms.PoolEmbedding(new EmbeddingPoolingOptions
{
    Pooling = PoolingStrategy.MeanPooling,
    Normalize = true,
    HiddenDim = scorer.HiddenDim,
    IsPrePooled = scorer.HasPooledOutput,
    SequenceLength = 128
}).Fit(scorer.Transform(tokenizer.Transform(data)));

Or chain with .Append() for the idiomatic ML.NET pattern:

var pipeline = mlContext.Transforms.TokenizeText(tokenizerOpts)
    .Append(mlContext.Transforms.ScoreOnnxTextModel(scorerOpts))
    .Append(mlContext.Transforms.PoolEmbedding(poolingOpts));
var model = pipeline.Fit(data);

Convenience facade (single-shot)

var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
});
var transformer = estimator.Fit(data);
var embeddings = transformer.Transform(data);

Provider-agnostic MEAI embedding

using Microsoft.Extensions.AI;

// Works with ANY IEmbeddingGenerator — ONNX, OpenAI, Azure, Ollama...
IEmbeddingGenerator<string, Embedding<float>> generator =
    new OnnxEmbeddingGenerator(mlContext, transformer);

var estimator = mlContext.Transforms.TextEmbedding(generator);

Save and load

// Save — bundles ONNX model + tokenizer + config into a portable zip
transformer.Save("my-embedding-model.mlnet");

// Load — fully self-contained, no external file dependencies
var loaded = OnnxTextEmbeddingTransformer.Load(mlContext, "my-embedding-model.mlnet");

Model Compatibility

The shared foundation supports any encoder transformer ONNX model. Compatible model architectures include:

  • BERT and derivatives (all-MiniLM, all-mpnet)
  • RoBERTa (including XLM-RoBERTa for multilingual)
  • DistilBERT
  • DeBERTa / DeBERTa-v2
  • MiniLM (Microsoft)
  • MPNet (Microsoft)
  • E5 (intfloat)
  • BGE (BAAI)
  • GTE (Alibaba)

Tested Embedding Models

Compatible Classification Models

Any encoder model fine-tuned for sequence classification (outputs logits over label classes):

  • Sentiment: DistilBERT-SST2, BERT-SST2
  • Emotion: RoBERTa-emotion, GoEmotions
  • NLI/Zero-shot: DeBERTa-v3-NLI, BART-large-MNLI

Compatible Reranking Models

Cross-encoder models that take text pairs and output a relevance score:

  • MS MARCO: cross-encoder/ms-marco-MiniLM-L-6-v2
  • BGE: BAAI/bge-reranker-base, bge-reranker-v2-m3

Compatible NER Models

Token classification models with BIO tagging:

  • BERT-NER: dslim/bert-base-NER
  • Multilingual: Davlan/xlm-roberta-base-ner-hrl

Compatible QA Models

Extractive QA models outputting start/end logits:

  • RoBERTa: deepset/roberta-base-squad2
  • MiniLM: deepset/minilm-uncased-squad2

Models with sentence_embedding output (pre-pooled) are auto-detected and pooling is skipped.

Project Structure

mlnet-text-inference-custom-transforms/
├── src/
│   ├── MLNet.TextInference.Onnx/
│   │   ├── TextTokenizerEstimator.cs         — Transform 1: tokenization (BPE/WordPiece/SentencePiece)
│   │   ├── OnnxTextModelScorerEstimator.cs   — Transform 2: ONNX inference with lookahead batching
│   │   ├── ChatClientEstimator.cs            — Provider-agnostic text generation (IChatClient)
│   │   ├── MLContextExtensions.cs            — Extension methods for fluent API
│   │   ├── Embeddings/
│   │   │   ├── EmbeddingPoolingEstimator.cs  — Pooling + L2 normalization
│   │   │   ├── OnnxTextEmbeddingEstimator.cs — Convenience facade
│   │   │   ├── EmbeddingGeneratorEstimator.cs— MEAI IEmbeddingGenerator wrapper
│   │   │   └── ...                           — OnnxEmbeddingGenerator, ModelPackager, PoolingStrategy
│   │   ├── Classification/
│   │   │   ├── SoftmaxClassificationEstimator.cs — Softmax post-processing
│   │   │   ├── OnnxTextClassificationEstimator.cs— Full pipeline facade
│   │   │   └── ...                              — Options, Transformer, ClassificationResult
│   │   ├── Reranking/
│   │   │   ├── SigmoidScorerEstimator.cs     — Sigmoid scoring transform
│   │   │   ├── OnnxRerankerEstimator.cs      — Cross-encoder facade
│   │   │   └── ...                           — Options, Transformer
│   │   ├── NER/
│   │   │   ├── NerDecodingEstimator.cs       — BIO tag decoding to entity spans
│   │   │   ├── OnnxNerEstimator.cs           — End-to-end NER facade
│   │   │   └── ...                           — Options, Transformer, NerEntity
│   │   └── QA/
│   │       ├── QaSpanExtractionEstimator.cs  — Span extraction from start/end logits
│   │       ├── OnnxQaEstimator.cs            — End-to-end QA facade
│   │       └── ...                           — Options, Transformer, QaResult
│   └── MLNet.TextGeneration.OnnxGenAI/
│       ├── OnnxTextGenerationEstimator.cs    — ORT GenAI local text generation
│       ├── MLContextExtensions.cs            — OnnxTextGeneration() extension method
│       └── ...                               — Options, Transformer
├── samples/
│   ├── BasicUsage/                           — all-MiniLM-L6-v2: all embedding API surfaces
│   ├── BgeSmallEmbedding/                    — BGE-small: query prefix pattern
│   ├── E5SmallEmbedding/                     — E5-small: query/passage prefixes
│   ├── GteSmallEmbedding/                    — GTE-small: semantic search (no prefix)
│   ├── ComposablePoolingComparison/          — 3 pooling strategies, shared inference
│   ├── IntermediateInspection/               — Inspect tokens, masks, raw output
│   ├── MeaiProviderAgnostic/                 — Provider-agnostic MEAI transform
│   ├── Classification/
│   │   ├── SentimentDistilBERT/              — Sentiment analysis with DistilBERT
│   │   ├── EmotionRoBERTa/                   — Multi-class emotion classification
│   │   └── ZeroShotDeBERTa/                  — Zero-shot NLI classification with DeBERTa
│   ├── Reranking/
│   │   ├── MsMarcoMiniLM/                    — MS MARCO cross-encoder reranking
│   │   └── BgeReranker/                      — BGE reranker for retrieval
│   ├── NER/
│   │   ├── BertBaseNER/                      — Named entity recognition with BERT-base
│   │   └── MultilingualNER/                  — Multilingual NER with XLM-RoBERTa
│   ├── QA/
│   │   ├── RobertaSquad2/                    — Extractive QA with RoBERTa on SQuAD 2.0
│   │   └── MiniLMSquad2/                     — Extractive QA with MiniLM on SQuAD 2.0
│   ├── TextGenerationMeai/                   — Provider-agnostic text generation (IChatClient)
│   └── TextGenerationLocal/                  — Local ORT GenAI text generation
├── docs/                                     — Detailed documentation
│   ├── architecture-decision-record.md       — ADR for multi-task platform decisions
│   ├── architecture.md                       — Component walkthrough + pipeline stages
│   ├── design-decisions.md                   — Why every choice was made
│   ├── extending.md                          — How to modify, extend, and add new tasks
│   ├── tensor-deep-dive.md                   — System.Numerics.Tensors for AI workloads
│   └── references.md                         — All sources and further reading
├── proposals/                                — Design proposals for the modular architecture
└── nuget.config                              — NuGet source (nuget.org only)

Samples

SampleModelPattern Demonstrated
BasicUsageall-MiniLM-L6-v2All API surfaces: facade, composable pipeline, .Append(), save/load, MEAI
BgeSmallEmbeddingBGE-small-en-v1.5Composable pipeline + BGE query prefix for asymmetric retrieval
E5SmallEmbeddingE5-small-v2Composable pipeline + E5 dual query:/passage: prefix pattern
GteSmallEmbeddingGTE-smallComposable pipeline + semantic search (no prefix needed)
ComposablePoolingComparisonall-MiniLM-L6-v23 pooling strategies, shared tokenizer+scorer (key modularization demo)
IntermediateInspectionall-MiniLM-L6-v2Inspect token IDs, attention masks, raw output at each pipeline stage
MeaiProviderAgnosticall-MiniLM-L6-v2EmbeddingGeneratorEstimator wrapping IEmbeddingGenerator
SentimentDistilBERTDistilBERT-SST2Sentiment classification with softmax post-processing
EmotionRoBERTaRoBERTa-emotionMulti-class emotion classification
ZeroShotDeBERTaDeBERTa-v3-NLIZero-shot classification via NLI entailment
MsMarcoMiniLMMS MARCO MiniLMCross-encoder reranking with text-pair tokenization
BgeRerankerBAAI/bge-rerankerBGE cross-encoder reranking for retrieval
BertBaseNERBERT-base-NERNamed entity recognition with BIO tag decoding
MultilingualNERXLM-RoBERTa-NERMultilingual named entity recognition
RobertaSquad2RoBERTa-SQuAD2Extractive question answering with span extraction
MiniLMSquad2MiniLM-SQuAD2Extractive question answering (lighter model)
TextGenerationMeaiAny IChatClientProvider-agnostic text generation via MEAI
TextGenerationLocalORT GenAI modelLocal text generation with ONNX Runtime GenAI

API at a Glance

Shared Foundation

ClassRoleKey Methods
TextTokenizerEstimatorTransform 1: TokenizationFit(IDataView)
OnnxTextModelScorerEstimatorTransform 2: ONNX ScoringFit(IDataView) → .HiddenDim, .HasPooledOutput

Embeddings

ClassRoleKey Methods
EmbeddingPoolingEstimatorTransform 3: PoolingFit(IDataView)
OnnxTextEmbeddingEstimatorFacade (chains 1→2→3)Fit(IDataView), GetOutputSchema()
OnnxTextEmbeddingTransformerFacade transformerTransform(IDataView), Save(path), Load(ctx, path)
EmbeddingGeneratorEstimatorMEAI wrapper transformFit(IDataView)
OnnxEmbeddingGeneratorMEAI IEmbeddingGeneratorGenerateAsync(texts)

Classification

ClassRoleKey Methods
SoftmaxClassificationEstimatorSoftmax post-processingFit(IDataView)
OnnxTextClassificationEstimatorFacade (tokenize→score→softmax)Fit(IDataView)
OnnxTextClassificationTransformerFacade transformerTransform(IDataView), Classify(texts)

Reranking

ClassRoleKey Methods
SigmoidScorerEstimatorSigmoid scoring transformFit(IDataView)
OnnxRerankerEstimatorFacade (text-pair tokenize→score→sigmoid)Fit(IDataView)
OnnxRerankerTransformerFacade transformerTransform(IDataView), Rerank(query, docs)

Named Entity Recognition

ClassRoleKey Methods
NerDecodingEstimatorBIO tag decodingFit(IDataView)
OnnxNerEstimatorFacade (tokenize→score→decode)Fit(IDataView)
OnnxNerTransformerFacade transformerTransform(IDataView), RecognizeEntities(texts)

Question Answering

ClassRoleKey Methods
QaSpanExtractionEstimatorSpan extraction from logitsFit(IDataView)
OnnxQaEstimatorFacade (tokenize→multi-score→extract)Fit(IDataView)
OnnxQaTransformerFacade transformerTransform(IDataView), Answer(questions)

Typed decisions

SurfaceTypes and extensionsRole
ML.NET facadeOnnxTypedDecisionsOptions, OnnxTypedDecisionsEstimator, OnnxTypedDecisionsTransformerSchema-aware lazy end-to-end transform
ML.NET stagesDecisionInputPreparation*, OnnxDecisionModelScorer*, DecisionDecoding*Options / Estimator / Transformer stage types
ML.NET compositionTransformsCatalog extensions and AppendOnnxTypedDecisionsStandard ML.NET pipeline composition

The ML.NET facade adds question-specific typed columns. Choice exposes PredictedLabel and a calibrated Probabilities vector; Score exposes the expected zero-based Score and its probability vector; Noul exposes Boolean PredictedLabel and Probability (P(true)). Each question also gets its own entropy Confidence and ActionProbability. DecisionResults remains an optional full diagnostic JSON column. Preparation and scoring stages use native ML.NET numeric/vector/Boolean columns, and the facade is the normal cursor-batched path. Native single-row PredictionEngine mapping is supported through the transformer and native stage row mappers. Native ML.NET Save/Load remains unsupported for typed decisions in this release; use the explicit portable API instead. External assets are a packaging consideration, not an inherent limitation of row mapping. Typed-decision transformers provide an explicit portable ZIP API (Save/Load and the supported facade/stage pipeline helpers); this is separate from native MLContext.Model.Save/Load, which remains unsupported for these custom components. The supported pipeline boundary is the demonstrated flat preparation -> scoring -> decoding chain and naturally inferred appended typed-decision facades; arbitrary/native chains are rejected. All sources in a saved pipeline must reference the same complete asset payload, so separately loaded selective stage archives cannot currently be recombined.

Text Generation

ClassRoleKey Methods
ChatClientEstimatorProvider-agnostic (wraps IChatClient)Fit(IDataView)
ChatClientTransformerProvider-agnostic transformerTransform(IDataView)
OnnxTextGenerationEstimatorORT GenAI local generationFit(IDataView)
OnnxTextGenerationTransformerORT GenAI transformerTransform(IDataView)

GPU Support (CUDA)

The library supports GPU-accelerated ONNX inference via CUDA. The library itself ships with no native binaries — you control the execution provider by choosing your OnnxRuntime package.

GPU Prerequisites

GPU inference requires the CUDA Toolkit and cuDNN installed on the host machine, plus an NVIDIA GPU with a compatible driver.

Windows (winget + direct download):

# 1. Install CUDA Toolkit 12.6
winget install Nvidia.CUDA --version 12.6 --source winget

# 2. Download and install cuDNN 9.x for CUDA 12
#    Download the zip from NVIDIA's redistributable endpoint:
$url = "https://developer.download.nvidia.com/compute/cudnn/redist/cudnn/windows-x86_64/cudnn-windows-x86_64-9.8.0.87_cuda12-archive.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:TEMP\cudnn.zip"
Expand-Archive "$env:TEMP\cudnn.zip" -DestinationPath "$env:TEMP\cudnn"

# 3. Copy cuDNN DLLs to CUDA bin (requires admin)
Copy-Item "$env:TEMP\cudnn\cudnn-*\bin\*.dll" "$env:CUDA_PATH\bin" -Force

Linux (apt):

# CUDA Toolkit
sudo apt-get install -y nvidia-cuda-toolkit

# cuDNN (via NVIDIA's apt repository)
# See: https://docs.nvidia.com/deeplearning/cudnn/installation/linux.html

Verify installation:

nvidia-smi                    # Should show GPU + driver version
nvcc --version                # Should show CUDA 12.x

Package Setup

Replace Microsoft.ML.OnnxRuntime with Microsoft.ML.OnnxRuntime.Gpu in your application:

<PackageReference Include="Microsoft.ML.OnnxRuntime.Gpu" Version="1.24.2" />

Samples auto-detect GPU: The sample projects use a Directory.Build.props that checks for CUDA_PATH (set by the CUDA Toolkit installer) and automatically switches to Microsoft.ML.OnnxRuntime.Gpu. Override with dotnet build -p:UseGpuRuntime=true or dotnet build -p:UseGpuRuntime=false.

Usage

// Pattern 1: Context-level (applies to all ONNX estimators)
var mlContext = new MLContext();
mlContext.GpuDeviceId = 0; // Use first CUDA device

var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
});

// Pattern 2: Per-estimator override
var scorerOptions = new OnnxTextModelScorerOptions
{
    ModelPath = "models/model.onnx",
    GpuDeviceId = 0,       // Override for this estimator only
    FallbackToCpu = true,  // Graceful degradation if CUDA unavailable
};

Resolution order: Per-estimator GpuDeviceId → MLContext.GpuDeviceId → null (CPU).

When FallbackToCpu = true, if CUDA initialization fails the estimator silently falls back to CPU execution.

NuGet Dependencies

MLNet.TextInference.Onnx (encoder tasks)

PackageVersionPurpose
Microsoft.ML5.0.0IEstimator/ITransformer, IDataView, MLContext
Microsoft.ML.OnnxRuntime.Managed1.24.2InferenceSession, OrtValue (managed API; bring your own native runtime)
Microsoft.ML.Tokenizers2.0.0BertTokenizer (WordPiece), BPE, SentencePiece
Microsoft.Extensions.AI.Abstractions10.3.0IEmbeddingGenerator, IChatClient
System.Numerics.Tensors10.0.3Tensor<T>, TensorPrimitives

MLNet.TextGeneration.OnnxGenAI (local text generation)

PackageVersionPurpose
Microsoft.ML5.0.0IEstimator/ITransformer, IDataView, MLContext
Microsoft.ML.OnnxRuntimeGenAI0.7.1ORT GenAI for local autoregressive generation
Microsoft.Extensions.AI.Abstractions10.3.0IChatClient

Using as a NuGet Package

Pre-built packages are published to GitHub Packages. No need to copy source code.

1. Add the GitHub Packages NuGet source (one-time setup):

<!-- nuget.config in your project root -->
<configuration>
  <packageSources>
    <add key="nuget.org" value="https://api.nuget.org/v3/index.json" />
    <add key="github-packages" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
  </packageSources>
  <packageSourceCredentials>
    <github-packages>
      <add key="Username" value="YOUR_GITHUB_USERNAME" />
      <add key="ClearTextPassword" value="YOUR_GITHUB_PAT" />
    </github-packages>
  </packageSourceCredentials>
</configuration>

Note: GitHub Packages requires authentication even for public packages. Create a Personal Access Token with read:packages scope.

2. Add the package reference:

<!-- Encoder tasks: embeddings, classification, reranking, NER, QA, text generation -->
<PackageReference Include="MLNet.TextInference.Onnx" Version="0.1.0-preview.1" />

<!-- Local text generation with ORT GenAI (separate package, heavier dependency) -->
<PackageReference Include="MLNet.TextGeneration.OnnxGenAI" Version="0.1.0-preview.1" />

3. Use it:

using MLNet.TextInference.Onnx;

var mlContext = new MLContext();
var estimator = mlContext.Transforms.OnnxTextClassification(new OnnxTextClassificationOptions
{
    ModelPath = "models/model.onnx",
    TokenizerPath = "models/",
    Labels = ["negative", "positive"]
});

Releasing a new version

Packages are versioned via git tags using MinVer:

git tag v0.1.0-preview.1
git push origin v0.1.0-preview.1
# → GitHub Actions builds, packs, and publishes to GitHub Packages

Documentation

Target Framework

.NET 10 (LTS). Requires the .NET 10 SDK.

License

MIT

Languages

C#

97.9%

Shell

1.9%