ML.NET custom transforms for text inference using ONNX encoder transformer models (embeddings, classification, NER, reranking, QA)
C#
3
54 commits
updated Sep 30, 2026
A multi-task text inference platform for ML.NET that runs local HuggingFace ONNX encoder models. Provides a shared foundation of tokenization + ONNX scoring, with task-specific post-processing transforms for embeddings, classification, NER, reranking, question answering, and more.
TextTokenizerTransformer
│
▼
OnnxTextModelScorerTransformer (task-agnostic)
│
┌──────────┬───────────┼───────────┬──────────┐
│ │ │ │ │
EmbeddingPool SoftmaxClass SigmoidScor NerDecoding QaSpanExtract
Transformer Transformer Transformer Transformer Transformer
ChatClientTransformer (text generation — provider-agnostic, separate pipeline)
| Task | Status | Post-processor | Facade |
|---|---|---|---|
| Embeddings | ✅ Implemented | EmbeddingPoolingTransformer | OnnxTextEmbeddingEstimator |
| Classification | ✅ Implemented | SoftmaxClassificationTransformer | OnnxTextClassificationEstimator |
| Reranking | ✅ Implemented | SigmoidScorerTransformer | OnnxRerankerEstimator |
| NER | ✅ Implemented | NerDecodingTransformer | OnnxNerEstimator |
| QA | ✅ Implemented | QaSpanExtractionTransformer | OnnxQaEstimator |
| Text Generation | ✅ Implemented | ChatClientTransformer | N/A (provider-agnostic) |
| Text Generation (local) | ✅ Implemented | OnnxTextGenerationTransformer | OnnxTextGenerationEstimator |
| Typed decisions (Laya English FP32) | ✅ Implemented | ML.NET stages: PrepareDecisionInputs, ScoreOnnxDecisionModel, DecodeDecisions | OnnxTypedDecisionsEstimator |
This project is forked from mlnet-embedding-custom-transforms, which provides embedding generation only. This fork extends the platform to support all encoder transformer tasks — embeddings, classification, NER, reranking, and question answering — by sharing a task-agnostic tokenization and ONNX scoring foundation and adding task-specific post-processing transforms.
ML.NET has no built-in transform for modern HuggingFace encoder models (all-MiniLM-L6-v2, BGE, E5, DeBERTa, etc.). Building one is hard because ML.NET's convenient internal base classes (RowToRowTransformerBase, OneToOneTransformerBase) have private protected constructors — they can't be subclassed from external projects.
This project implements custom transforms using direct IEstimator<T> / ITransformer interfaces (Approach C from the ML.NET Custom Transformer Guide), enhanced with custom zip-based save/load for model persistence.
TokenizeText → ScoreOnnxTextModel) plus task-specific post-processing that can be inspected, swapped, and reusedOnnxTextEmbeddingEstimator, OnnxTextClassificationEstimator, OnnxRerankerEstimator, OnnxNerEstimator, OnnxQaEstimator each wrap all transforms for their task in a single callEmbeddingGeneratorEstimator wraps any IEmbeddingGenerator<string, Embedding<float>> as an ML.NET transform; ChatClientEstimator wraps any IChatClient for text generation[CLS] A [SEP] B [SEP] with token type IDs for query-document pairsEncodeToTokens() for mapping entities back to source textAdditionalOutputTensorNamestokenizer_config.json, known vocab files (BPE, SentencePiece, WordPiece), or HuggingFace tokenizer.json (fast tokenizer).mlnet zip file containing the ONNX model, tokenizer, and configTensorPrimitives for hardware-vectorized mathTyped-decision naming follows the same conventions as the other transforms. The
ML.NET surface keeps the user-facing TransformsCatalog verbs and the approved
OnnxTypedDecisions, PrepareDecisionInputs, ScoreOnnxDecisionModel, and
DecodeDecisions names, while role types use the repository's
*Options, *Estimator, and *Transformer conventions:
DecisionInputPreparation*, OnnxDecisionModelScorer*, and
DecisionDecoding*. The compiled ML.NET facade can be appended with
AppendOnnxTypedDecisions.
Typed decisions use a versioned local bundle for the English FP32 Laya graph
from receptron/laya-onnx revision
68f27dfe5a27a54fb2b1fefc432f43f972e90868. The ML.NET facade is the primary
entry point; the direct transformer API and inspectable native stages use the
same implementation. The feature scores caller-supplied alternatives rather
than generating prose: Choice selects a label, Score returns an expected
zero-based option index, and Noul returns a Boolean plus the probability of
the true option. For this pretrained typed-decision estimator, Fit
validates an ML.NET schema and initializes resources; it does not train the
ONNX model.
In this repository, the typed-decision implementation is part of the
MLNet.TextInference.Onnx assembly/package. It uses
Microsoft.ML.Tokenizers for the selected byte-level BPE contract,
the managed/native ONNX Runtime packages for the five-input graph, and C#
decoding with stable tensor primitives. State is text (including
caller-serialized JSON); there is no implicit Python-compatible object
serializer. Inference never downloads model assets.
The normal model-assets directory contains the model, external-data sidecar, Laya configuration, and tokenizer directory. An optional manifest-backed archive is also supported. The explicit acceptance launcher requires model assets prepared locally:
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
-ModelAssetsPath .\models\laya-english-fp32.bundle -Mode facade
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
-ModelAssetsPath .\models\laya-english-fp32.bundle -Mode stages
The typed-decision samples have deliberately different jobs: the short orientation page points to the canonical guided tutorial, which owns the exact questions, file-based run path, captured output, tensor shapes, decoder semantics, native stage columns, and portable deployment boundary. The tracked PortableProcessHarness is a developer/acceptance utility for fresh-process offline persistence, not the beginner walkthrough.
The internal Laya preparation kernel loads and configures the existing Microsoft.ML.Tokenizers BPE engine, while separate profile metadata carries the special-token IDs and mask token. This keeps Microsoft.ML.Tokenizers as the encoding boundary without introducing a second tokenizer interface or reimplementing BPE.
Click "Open in GitHub Codespaces" above. The dev container automatically:
Once ready, run the first sample:
cd samples/BasicUsage && dotnet run
Download models for other tasks:
bash scripts/download-models.sh classification # sentiment, emotion, zero-shot
bash scripts/download-models.sh reranking # cross-encoder reranking
bash scripts/download-models.sh ner # named entity recognition
bash scripts/download-models.sh qa # question answering
bash scripts/download-models.sh all # everything (~3.5GB)
bash scripts/download-models.sh --help # see all options
Prerequisites: .NET 10 SDK. Python 3.x needed only for NER/QA model export.
dotnet restore && dotnet build
bash scripts/download-models.sh embeddings-core # downloads ~86MB starter model
cd samples/BasicUsage && dotnet run
Or download manually:
mkdir samples/BasicUsage/models
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" -OutFile "samples/BasicUsage/models/model.onnx"
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt" -OutFile "samples/BasicUsage/models/vocab.txt"
Note: Embeddings are the simplest task. Browse the samples/ directory for classification, reranking, NER, QA, and text generation examples.
using Microsoft.ML;
using MLNet.TextInference.Onnx;
var mlContext = new MLContext();
var data = mlContext.Data.LoadFromEnumerable(new[]
{
new { Text = "What is machine learning?" },
new { Text = "How to cook pasta" }
});
// Step-by-step: tokenize → score → pool
var tokenizer = mlContext.Transforms.TokenizeText(new TextTokenizerOptions
{
TokenizerPath = "models/", // directory — auto-detects tokenizer
InputColumnName = "Text"
}).Fit(data);
var scorer = mlContext.Transforms.ScoreOnnxTextModel(new OnnxTextModelScorerOptions
{
ModelPath = "models/model.onnx"
}).Fit(tokenizer.Transform(data));
var pooler = mlContext.Transforms.PoolEmbedding(new EmbeddingPoolingOptions
{
Pooling = PoolingStrategy.MeanPooling,
Normalize = true,
HiddenDim = scorer.HiddenDim,
IsPrePooled = scorer.HasPooledOutput,
SequenceLength = 128
}).Fit(scorer.Transform(tokenizer.Transform(data)));
Or chain with .Append() for the idiomatic ML.NET pattern:
var pipeline = mlContext.Transforms.TokenizeText(tokenizerOpts)
.Append(mlContext.Transforms.ScoreOnnxTextModel(scorerOpts))
.Append(mlContext.Transforms.PoolEmbedding(poolingOpts));
var model = pipeline.Fit(data);
var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
});
var transformer = estimator.Fit(data);
var embeddings = transformer.Transform(data);
using Microsoft.Extensions.AI;
// Works with ANY IEmbeddingGenerator — ONNX, OpenAI, Azure, Ollama...
IEmbeddingGenerator<string, Embedding<float>> generator =
new OnnxEmbeddingGenerator(mlContext, transformer);
var estimator = mlContext.Transforms.TextEmbedding(generator);
// Save — bundles ONNX model + tokenizer + config into a portable zip
transformer.Save("my-embedding-model.mlnet");
// Load — fully self-contained, no external file dependencies
var loaded = OnnxTextEmbeddingTransformer.Load(mlContext, "my-embedding-model.mlnet");
The shared foundation supports any encoder transformer ONNX model. Compatible model architectures include:
| Model | Dimensions | Size | Tested | Sample |
|---|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | ~86 MB | ✅ | BasicUsage |
| BAAI/bge-small-en-v1.5 | 384 | ~127 MB | ✅ | BgeSmallEmbedding |
| intfloat/e5-small-v2 | 384 | ~127 MB | ✅ | E5SmallEmbedding |
| thenlper/gte-small | 384 | ~127 MB | ✅ | GteSmallEmbedding |
| all-MiniLM-L12-v2 | 384 | ~120 MB | — | |
| all-mpnet-base-v2 | 768 | ~420 MB | — |
Any encoder model fine-tuned for sequence classification (outputs logits over label classes):
Cross-encoder models that take text pairs and output a relevance score:
Token classification models with BIO tagging:
Extractive QA models outputting start/end logits:
Models with sentence_embedding output (pre-pooled) are auto-detected and pooling is skipped.
mlnet-text-inference-custom-transforms/
├── src/
│ ├── MLNet.TextInference.Onnx/
│ │ ├── TextTokenizerEstimator.cs — Transform 1: tokenization (BPE/WordPiece/SentencePiece)
│ │ ├── OnnxTextModelScorerEstimator.cs — Transform 2: ONNX inference with lookahead batching
│ │ ├── ChatClientEstimator.cs — Provider-agnostic text generation (IChatClient)
│ │ ├── MLContextExtensions.cs — Extension methods for fluent API
│ │ ├── Embeddings/
│ │ │ ├── EmbeddingPoolingEstimator.cs — Pooling + L2 normalization
│ │ │ ├── OnnxTextEmbeddingEstimator.cs — Convenience facade
│ │ │ ├── EmbeddingGeneratorEstimator.cs— MEAI IEmbeddingGenerator wrapper
│ │ │ └── ... — OnnxEmbeddingGenerator, ModelPackager, PoolingStrategy
│ │ ├── Classification/
│ │ │ ├── SoftmaxClassificationEstimator.cs — Softmax post-processing
│ │ │ ├── OnnxTextClassificationEstimator.cs— Full pipeline facade
│ │ │ └── ... — Options, Transformer, ClassificationResult
│ │ ├── Reranking/
│ │ │ ├── SigmoidScorerEstimator.cs — Sigmoid scoring transform
│ │ │ ├── OnnxRerankerEstimator.cs — Cross-encoder facade
│ │ │ └── ... — Options, Transformer
│ │ ├── NER/
│ │ │ ├── NerDecodingEstimator.cs — BIO tag decoding to entity spans
│ │ │ ├── OnnxNerEstimator.cs — End-to-end NER facade
│ │ │ └── ... — Options, Transformer, NerEntity
│ │ └── QA/
│ │ ├── QaSpanExtractionEstimator.cs — Span extraction from start/end logits
│ │ ├── OnnxQaEstimator.cs — End-to-end QA facade
│ │ └── ... — Options, Transformer, QaResult
│ └── MLNet.TextGeneration.OnnxGenAI/
│ ├── OnnxTextGenerationEstimator.cs — ORT GenAI local text generation
│ ├── MLContextExtensions.cs — OnnxTextGeneration() extension method
│ └── ... — Options, Transformer
├── samples/
│ ├── BasicUsage/ — all-MiniLM-L6-v2: all embedding API surfaces
│ ├── BgeSmallEmbedding/ — BGE-small: query prefix pattern
│ ├── E5SmallEmbedding/ — E5-small: query/passage prefixes
│ ├── GteSmallEmbedding/ — GTE-small: semantic search (no prefix)
│ ├── ComposablePoolingComparison/ — 3 pooling strategies, shared inference
│ ├── IntermediateInspection/ — Inspect tokens, masks, raw output
│ ├── MeaiProviderAgnostic/ — Provider-agnostic MEAI transform
│ ├── Classification/
│ │ ├── SentimentDistilBERT/ — Sentiment analysis with DistilBERT
│ │ ├── EmotionRoBERTa/ — Multi-class emotion classification
│ │ └── ZeroShotDeBERTa/ — Zero-shot NLI classification with DeBERTa
│ ├── Reranking/
│ │ ├── MsMarcoMiniLM/ — MS MARCO cross-encoder reranking
│ │ └── BgeReranker/ — BGE reranker for retrieval
│ ├── NER/
│ │ ├── BertBaseNER/ — Named entity recognition with BERT-base
│ │ └── MultilingualNER/ — Multilingual NER with XLM-RoBERTa
│ ├── QA/
│ │ ├── RobertaSquad2/ — Extractive QA with RoBERTa on SQuAD 2.0
│ │ └── MiniLMSquad2/ — Extractive QA with MiniLM on SQuAD 2.0
│ ├── TextGenerationMeai/ — Provider-agnostic text generation (IChatClient)
│ └── TextGenerationLocal/ — Local ORT GenAI text generation
├── docs/ — Detailed documentation
│ ├── architecture-decision-record.md — ADR for multi-task platform decisions
│ ├── architecture.md — Component walkthrough + pipeline stages
│ ├── design-decisions.md — Why every choice was made
│ ├── extending.md — How to modify, extend, and add new tasks
│ ├── tensor-deep-dive.md — System.Numerics.Tensors for AI workloads
│ └── references.md — All sources and further reading
├── proposals/ — Design proposals for the modular architecture
└── nuget.config — NuGet source (nuget.org only)
| Sample | Model | Pattern Demonstrated |
|---|---|---|
| BasicUsage | all-MiniLM-L6-v2 | All API surfaces: facade, composable pipeline, .Append(), save/load, MEAI |
| BgeSmallEmbedding | BGE-small-en-v1.5 | Composable pipeline + BGE query prefix for asymmetric retrieval |
| E5SmallEmbedding | E5-small-v2 | Composable pipeline + E5 dual query:/passage: prefix pattern |
| GteSmallEmbedding | GTE-small | Composable pipeline + semantic search (no prefix needed) |
| ComposablePoolingComparison | all-MiniLM-L6-v2 | 3 pooling strategies, shared tokenizer+scorer (key modularization demo) |
| IntermediateInspection | all-MiniLM-L6-v2 | Inspect token IDs, attention masks, raw output at each pipeline stage |
| MeaiProviderAgnostic | all-MiniLM-L6-v2 | EmbeddingGeneratorEstimator wrapping IEmbeddingGenerator |
| SentimentDistilBERT | DistilBERT-SST2 | Sentiment classification with softmax post-processing |
| EmotionRoBERTa | RoBERTa-emotion | Multi-class emotion classification |
| ZeroShotDeBERTa | DeBERTa-v3-NLI | Zero-shot classification via NLI entailment |
| MsMarcoMiniLM | MS MARCO MiniLM | Cross-encoder reranking with text-pair tokenization |
| BgeReranker | BAAI/bge-reranker | BGE cross-encoder reranking for retrieval |
| BertBaseNER | BERT-base-NER | Named entity recognition with BIO tag decoding |
| MultilingualNER | XLM-RoBERTa-NER | Multilingual named entity recognition |
| RobertaSquad2 | RoBERTa-SQuAD2 | Extractive question answering with span extraction |
| MiniLMSquad2 | MiniLM-SQuAD2 | Extractive question answering (lighter model) |
| TextGenerationMeai | Any IChatClient | Provider-agnostic text generation via MEAI |
| TextGenerationLocal | ORT GenAI model | Local text generation with ONNX Runtime GenAI |
| Class | Role | Key Methods |
|---|---|---|
TextTokenizerEstimator | Transform 1: Tokenization | Fit(IDataView) |
OnnxTextModelScorerEstimator | Transform 2: ONNX Scoring | Fit(IDataView) → .HiddenDim, .HasPooledOutput |
| Class | Role | Key Methods |
|---|---|---|
EmbeddingPoolingEstimator | Transform 3: Pooling | Fit(IDataView) |
OnnxTextEmbeddingEstimator | Facade (chains 1→2→3) | Fit(IDataView), GetOutputSchema() |
OnnxTextEmbeddingTransformer | Facade transformer | Transform(IDataView), Save(path), Load(ctx, path) |
EmbeddingGeneratorEstimator | MEAI wrapper transform | Fit(IDataView) |
OnnxEmbeddingGenerator | MEAI IEmbeddingGenerator | GenerateAsync(texts) |
| Class | Role | Key Methods |
|---|---|---|
SoftmaxClassificationEstimator | Softmax post-processing | Fit(IDataView) |
OnnxTextClassificationEstimator | Facade (tokenize→score→softmax) | Fit(IDataView) |
OnnxTextClassificationTransformer | Facade transformer | Transform(IDataView), Classify(texts) |
| Class | Role | Key Methods |
|---|---|---|
SigmoidScorerEstimator | Sigmoid scoring transform | Fit(IDataView) |
OnnxRerankerEstimator | Facade (text-pair tokenize→score→sigmoid) | Fit(IDataView) |
OnnxRerankerTransformer | Facade transformer | Transform(IDataView), Rerank(query, docs) |
| Class | Role | Key Methods |
|---|---|---|
NerDecodingEstimator | BIO tag decoding | Fit(IDataView) |
OnnxNerEstimator | Facade (tokenize→score→decode) | Fit(IDataView) |
OnnxNerTransformer | Facade transformer | Transform(IDataView), RecognizeEntities(texts) |
| Class | Role | Key Methods |
|---|---|---|
QaSpanExtractionEstimator | Span extraction from logits | Fit(IDataView) |
OnnxQaEstimator | Facade (tokenize→multi-score→extract) | Fit(IDataView) |
OnnxQaTransformer | Facade transformer | Transform(IDataView), Answer(questions) |
| Surface | Types and extensions | Role |
|---|---|---|
| ML.NET facade | OnnxTypedDecisionsOptions, OnnxTypedDecisionsEstimator, OnnxTypedDecisionsTransformer | Schema-aware lazy end-to-end transform |
| ML.NET stages | DecisionInputPreparation*, OnnxDecisionModelScorer*, DecisionDecoding* | Options / Estimator / Transformer stage types |
| ML.NET composition | TransformsCatalog extensions and AppendOnnxTypedDecisions | Standard ML.NET pipeline composition |
The ML.NET facade adds question-specific typed columns. Choice exposes
PredictedLabel and a calibrated Probabilities vector; Score exposes the
expected zero-based Score and its probability vector; Noul exposes Boolean
PredictedLabel and Probability (P(true)). Each question also gets its
own entropy Confidence and ActionProbability. DecisionResults remains
an optional full diagnostic JSON column. Preparation and scoring stages use
native ML.NET numeric/vector/Boolean columns, and the facade is the normal
cursor-batched path. Native single-row PredictionEngine mapping is supported
through the transformer and native stage row mappers. Native ML.NET Save/Load
remains unsupported for typed decisions in this release; use the explicit
portable API instead. External assets are a packaging consideration, not an
inherent limitation of row mapping.
Typed-decision transformers provide an explicit portable ZIP API
(Save/Load and the supported facade/stage pipeline helpers); this is separate from
native MLContext.Model.Save/Load, which remains unsupported for these custom
components. The supported pipeline boundary is the demonstrated flat
preparation -> scoring -> decoding chain and naturally inferred appended
typed-decision facades; arbitrary/native chains are rejected. All sources in a
saved pipeline must reference the same complete asset payload, so separately
loaded selective stage archives cannot currently be recombined.
| Class | Role | Key Methods |
|---|---|---|
ChatClientEstimator | Provider-agnostic (wraps IChatClient) | Fit(IDataView) |
ChatClientTransformer | Provider-agnostic transformer | Transform(IDataView) |
OnnxTextGenerationEstimator | ORT GenAI local generation | Fit(IDataView) |
OnnxTextGenerationTransformer | ORT GenAI transformer | Transform(IDataView) |
The library supports GPU-accelerated ONNX inference via CUDA. The library itself ships with no native binaries — you control the execution provider by choosing your OnnxRuntime package.
GPU inference requires the CUDA Toolkit and cuDNN installed on the host machine, plus an NVIDIA GPU with a compatible driver.
Windows (winget + direct download):
# 1. Install CUDA Toolkit 12.6
winget install Nvidia.CUDA --version 12.6 --source winget
# 2. Download and install cuDNN 9.x for CUDA 12
# Download the zip from NVIDIA's redistributable endpoint:
$url = "https://developer.download.nvidia.com/compute/cudnn/redist/cudnn/windows-x86_64/cudnn-windows-x86_64-9.8.0.87_cuda12-archive.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:TEMP\cudnn.zip"
Expand-Archive "$env:TEMP\cudnn.zip" -DestinationPath "$env:TEMP\cudnn"
# 3. Copy cuDNN DLLs to CUDA bin (requires admin)
Copy-Item "$env:TEMP\cudnn\cudnn-*\bin\*.dll" "$env:CUDA_PATH\bin" -Force
Linux (apt):
# CUDA Toolkit
sudo apt-get install -y nvidia-cuda-toolkit
# cuDNN (via NVIDIA's apt repository)
# See: https://docs.nvidia.com/deeplearning/cudnn/installation/linux.html
Verify installation:
nvidia-smi # Should show GPU + driver version
nvcc --version # Should show CUDA 12.x
Replace Microsoft.ML.OnnxRuntime with Microsoft.ML.OnnxRuntime.Gpu in your application:
<PackageReference Include="Microsoft.ML.OnnxRuntime.Gpu" Version="1.24.2" />
Samples auto-detect GPU: The sample projects use a
Directory.Build.propsthat checks forCUDA_PATH(set by the CUDA Toolkit installer) and automatically switches toMicrosoft.ML.OnnxRuntime.Gpu. Override withdotnet build -p:UseGpuRuntime=trueordotnet build -p:UseGpuRuntime=false.
// Pattern 1: Context-level (applies to all ONNX estimators)
var mlContext = new MLContext();
mlContext.GpuDeviceId = 0; // Use first CUDA device
var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
});
// Pattern 2: Per-estimator override
var scorerOptions = new OnnxTextModelScorerOptions
{
ModelPath = "models/model.onnx",
GpuDeviceId = 0, // Override for this estimator only
FallbackToCpu = true, // Graceful degradation if CUDA unavailable
};
Resolution order: Per-estimator GpuDeviceId → MLContext.GpuDeviceId → null (CPU).
When FallbackToCpu = true, if CUDA initialization fails the estimator silently falls back to CPU execution.
| Package | Version | Purpose |
|---|---|---|
Microsoft.ML | 5.0.0 | IEstimator/ITransformer, IDataView, MLContext |
Microsoft.ML.OnnxRuntime.Managed | 1.24.2 | InferenceSession, OrtValue (managed API; bring your own native runtime) |
Microsoft.ML.Tokenizers | 2.0.0 | BertTokenizer (WordPiece), BPE, SentencePiece |
Microsoft.Extensions.AI.Abstractions | 10.3.0 | IEmbeddingGenerator, IChatClient |
System.Numerics.Tensors | 10.0.3 | Tensor<T>, TensorPrimitives |
| Package | Version | Purpose |
|---|---|---|
Microsoft.ML | 5.0.0 | IEstimator/ITransformer, IDataView, MLContext |
Microsoft.ML.OnnxRuntimeGenAI | 0.7.1 | ORT GenAI for local autoregressive generation |
Microsoft.Extensions.AI.Abstractions | 10.3.0 | IChatClient |
Pre-built packages are published to GitHub Packages. No need to copy source code.
1. Add the GitHub Packages NuGet source (one-time setup):
<!-- nuget.config in your project root -->
<configuration>
<packageSources>
<add key="nuget.org" value="https://api.nuget.org/v3/index.json" />
<add key="github-packages" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
</packageSources>
<packageSourceCredentials>
<github-packages>
<add key="Username" value="YOUR_GITHUB_USERNAME" />
<add key="ClearTextPassword" value="YOUR_GITHUB_PAT" />
</github-packages>
</packageSourceCredentials>
</configuration>
Note: GitHub Packages requires authentication even for public packages. Create a Personal Access Token with
read:packagesscope.
2. Add the package reference:
<!-- Encoder tasks: embeddings, classification, reranking, NER, QA, text generation -->
<PackageReference Include="MLNet.TextInference.Onnx" Version="0.1.0-preview.1" />
<!-- Local text generation with ORT GenAI (separate package, heavier dependency) -->
<PackageReference Include="MLNet.TextGeneration.OnnxGenAI" Version="0.1.0-preview.1" />
3. Use it:
using MLNet.TextInference.Onnx;
var mlContext = new MLContext();
var estimator = mlContext.Transforms.OnnxTextClassification(new OnnxTextClassificationOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
Labels = ["negative", "positive"]
});
Packages are versioned via git tags using MinVer:
git tag v0.1.0-preview.1
git push origin v0.1.0-preview.1
# → GitHub Actions builds, packs, and publishes to GitHub Packages
.NET 10 (LTS). Requires the .NET 10 SDK.
C#
97.9%
Shell
1.9%
ML.NET custom transforms for text inference using ONNX encoder transformer models (embeddings, classification, NER, reranking, QA)
C#
3
54 commits
updated Sep 30, 2026
A multi-task text inference platform for ML.NET that runs local HuggingFace ONNX encoder models. Provides a shared foundation of tokenization + ONNX scoring, with task-specific post-processing transforms for embeddings, classification, NER, reranking, question answering, and more.
TextTokenizerTransformer
│
▼
OnnxTextModelScorerTransformer (task-agnostic)
│
┌──────────┬───────────┼───────────┬──────────┐
│ │ │ │ │
EmbeddingPool SoftmaxClass SigmoidScor NerDecoding QaSpanExtract
Transformer Transformer Transformer Transformer Transformer
ChatClientTransformer (text generation — provider-agnostic, separate pipeline)
| Task | Status | Post-processor | Facade |
|---|---|---|---|
| Embeddings | ✅ Implemented | EmbeddingPoolingTransformer | OnnxTextEmbeddingEstimator |
| Classification | ✅ Implemented | SoftmaxClassificationTransformer | OnnxTextClassificationEstimator |
| Reranking | ✅ Implemented | SigmoidScorerTransformer | OnnxRerankerEstimator |
| NER | ✅ Implemented | NerDecodingTransformer | OnnxNerEstimator |
| QA | ✅ Implemented | QaSpanExtractionTransformer | OnnxQaEstimator |
| Text Generation | ✅ Implemented | ChatClientTransformer | N/A (provider-agnostic) |
| Text Generation (local) | ✅ Implemented | OnnxTextGenerationTransformer | OnnxTextGenerationEstimator |
| Typed decisions (Laya English FP32) | ✅ Implemented | ML.NET stages: PrepareDecisionInputs, ScoreOnnxDecisionModel, DecodeDecisions | OnnxTypedDecisionsEstimator |
This project is forked from mlnet-embedding-custom-transforms, which provides embedding generation only. This fork extends the platform to support all encoder transformer tasks — embeddings, classification, NER, reranking, and question answering — by sharing a task-agnostic tokenization and ONNX scoring foundation and adding task-specific post-processing transforms.
ML.NET has no built-in transform for modern HuggingFace encoder models (all-MiniLM-L6-v2, BGE, E5, DeBERTa, etc.). Building one is hard because ML.NET's convenient internal base classes (RowToRowTransformerBase, OneToOneTransformerBase) have private protected constructors — they can't be subclassed from external projects.
This project implements custom transforms using direct IEstimator<T> / ITransformer interfaces (Approach C from the ML.NET Custom Transformer Guide), enhanced with custom zip-based save/load for model persistence.
TokenizeText → ScoreOnnxTextModel) plus task-specific post-processing that can be inspected, swapped, and reusedOnnxTextEmbeddingEstimator, OnnxTextClassificationEstimator, OnnxRerankerEstimator, OnnxNerEstimator, OnnxQaEstimator each wrap all transforms for their task in a single callEmbeddingGeneratorEstimator wraps any IEmbeddingGenerator<string, Embedding<float>> as an ML.NET transform; ChatClientEstimator wraps any IChatClient for text generation[CLS] A [SEP] B [SEP] with token type IDs for query-document pairsEncodeToTokens() for mapping entities back to source textAdditionalOutputTensorNamestokenizer_config.json, known vocab files (BPE, SentencePiece, WordPiece), or HuggingFace tokenizer.json (fast tokenizer).mlnet zip file containing the ONNX model, tokenizer, and configTensorPrimitives for hardware-vectorized mathTyped-decision naming follows the same conventions as the other transforms. The
ML.NET surface keeps the user-facing TransformsCatalog verbs and the approved
OnnxTypedDecisions, PrepareDecisionInputs, ScoreOnnxDecisionModel, and
DecodeDecisions names, while role types use the repository's
*Options, *Estimator, and *Transformer conventions:
DecisionInputPreparation*, OnnxDecisionModelScorer*, and
DecisionDecoding*. The compiled ML.NET facade can be appended with
AppendOnnxTypedDecisions.
Typed decisions use a versioned local bundle for the English FP32 Laya graph
from receptron/laya-onnx revision
68f27dfe5a27a54fb2b1fefc432f43f972e90868. The ML.NET facade is the primary
entry point; the direct transformer API and inspectable native stages use the
same implementation. The feature scores caller-supplied alternatives rather
than generating prose: Choice selects a label, Score returns an expected
zero-based option index, and Noul returns a Boolean plus the probability of
the true option. For this pretrained typed-decision estimator, Fit
validates an ML.NET schema and initializes resources; it does not train the
ONNX model.
In this repository, the typed-decision implementation is part of the
MLNet.TextInference.Onnx assembly/package. It uses
Microsoft.ML.Tokenizers for the selected byte-level BPE contract,
the managed/native ONNX Runtime packages for the five-input graph, and C#
decoding with stable tensor primitives. State is text (including
caller-serialized JSON); there is no implicit Python-compatible object
serializer. Inference never downloads model assets.
The normal model-assets directory contains the model, external-data sidecar, Laya configuration, and tokenizer directory. An optional manifest-backed archive is also supported. The explicit acceptance launcher requires model assets prepared locally:
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
-ModelAssetsPath .\models\laya-english-fp32.bundle -Mode facade
.\scripts\typed-decisions\Invoke-LayaAcceptance.ps1 `
-ModelAssetsPath .\models\laya-english-fp32.bundle -Mode stages
The typed-decision samples have deliberately different jobs: the short orientation page points to the canonical guided tutorial, which owns the exact questions, file-based run path, captured output, tensor shapes, decoder semantics, native stage columns, and portable deployment boundary. The tracked PortableProcessHarness is a developer/acceptance utility for fresh-process offline persistence, not the beginner walkthrough.
The internal Laya preparation kernel loads and configures the existing Microsoft.ML.Tokenizers BPE engine, while separate profile metadata carries the special-token IDs and mask token. This keeps Microsoft.ML.Tokenizers as the encoding boundary without introducing a second tokenizer interface or reimplementing BPE.
Click "Open in GitHub Codespaces" above. The dev container automatically:
Once ready, run the first sample:
cd samples/BasicUsage && dotnet run
Download models for other tasks:
bash scripts/download-models.sh classification # sentiment, emotion, zero-shot
bash scripts/download-models.sh reranking # cross-encoder reranking
bash scripts/download-models.sh ner # named entity recognition
bash scripts/download-models.sh qa # question answering
bash scripts/download-models.sh all # everything (~3.5GB)
bash scripts/download-models.sh --help # see all options
Prerequisites: .NET 10 SDK. Python 3.x needed only for NER/QA model export.
dotnet restore && dotnet build
bash scripts/download-models.sh embeddings-core # downloads ~86MB starter model
cd samples/BasicUsage && dotnet run
Or download manually:
mkdir samples/BasicUsage/models
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/onnx/model.onnx" -OutFile "samples/BasicUsage/models/model.onnx"
Invoke-WebRequest -Uri "https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2/resolve/main/vocab.txt" -OutFile "samples/BasicUsage/models/vocab.txt"
Note: Embeddings are the simplest task. Browse the samples/ directory for classification, reranking, NER, QA, and text generation examples.
using Microsoft.ML;
using MLNet.TextInference.Onnx;
var mlContext = new MLContext();
var data = mlContext.Data.LoadFromEnumerable(new[]
{
new { Text = "What is machine learning?" },
new { Text = "How to cook pasta" }
});
// Step-by-step: tokenize → score → pool
var tokenizer = mlContext.Transforms.TokenizeText(new TextTokenizerOptions
{
TokenizerPath = "models/", // directory — auto-detects tokenizer
InputColumnName = "Text"
}).Fit(data);
var scorer = mlContext.Transforms.ScoreOnnxTextModel(new OnnxTextModelScorerOptions
{
ModelPath = "models/model.onnx"
}).Fit(tokenizer.Transform(data));
var pooler = mlContext.Transforms.PoolEmbedding(new EmbeddingPoolingOptions
{
Pooling = PoolingStrategy.MeanPooling,
Normalize = true,
HiddenDim = scorer.HiddenDim,
IsPrePooled = scorer.HasPooledOutput,
SequenceLength = 128
}).Fit(scorer.Transform(tokenizer.Transform(data)));
Or chain with .Append() for the idiomatic ML.NET pattern:
var pipeline = mlContext.Transforms.TokenizeText(tokenizerOpts)
.Append(mlContext.Transforms.ScoreOnnxTextModel(scorerOpts))
.Append(mlContext.Transforms.PoolEmbedding(poolingOpts));
var model = pipeline.Fit(data);
var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
});
var transformer = estimator.Fit(data);
var embeddings = transformer.Transform(data);
using Microsoft.Extensions.AI;
// Works with ANY IEmbeddingGenerator — ONNX, OpenAI, Azure, Ollama...
IEmbeddingGenerator<string, Embedding<float>> generator =
new OnnxEmbeddingGenerator(mlContext, transformer);
var estimator = mlContext.Transforms.TextEmbedding(generator);
// Save — bundles ONNX model + tokenizer + config into a portable zip
transformer.Save("my-embedding-model.mlnet");
// Load — fully self-contained, no external file dependencies
var loaded = OnnxTextEmbeddingTransformer.Load(mlContext, "my-embedding-model.mlnet");
The shared foundation supports any encoder transformer ONNX model. Compatible model architectures include:
| Model | Dimensions | Size | Tested | Sample |
|---|---|---|---|---|
| all-MiniLM-L6-v2 | 384 | ~86 MB | ✅ | BasicUsage |
| BAAI/bge-small-en-v1.5 | 384 | ~127 MB | ✅ | BgeSmallEmbedding |
| intfloat/e5-small-v2 | 384 | ~127 MB | ✅ | E5SmallEmbedding |
| thenlper/gte-small | 384 | ~127 MB | ✅ | GteSmallEmbedding |
| all-MiniLM-L12-v2 | 384 | ~120 MB | — | |
| all-mpnet-base-v2 | 768 | ~420 MB | — |
Any encoder model fine-tuned for sequence classification (outputs logits over label classes):
Cross-encoder models that take text pairs and output a relevance score:
Token classification models with BIO tagging:
Extractive QA models outputting start/end logits:
Models with sentence_embedding output (pre-pooled) are auto-detected and pooling is skipped.
mlnet-text-inference-custom-transforms/
├── src/
│ ├── MLNet.TextInference.Onnx/
│ │ ├── TextTokenizerEstimator.cs — Transform 1: tokenization (BPE/WordPiece/SentencePiece)
│ │ ├── OnnxTextModelScorerEstimator.cs — Transform 2: ONNX inference with lookahead batching
│ │ ├── ChatClientEstimator.cs — Provider-agnostic text generation (IChatClient)
│ │ ├── MLContextExtensions.cs — Extension methods for fluent API
│ │ ├── Embeddings/
│ │ │ ├── EmbeddingPoolingEstimator.cs — Pooling + L2 normalization
│ │ │ ├── OnnxTextEmbeddingEstimator.cs — Convenience facade
│ │ │ ├── EmbeddingGeneratorEstimator.cs— MEAI IEmbeddingGenerator wrapper
│ │ │ └── ... — OnnxEmbeddingGenerator, ModelPackager, PoolingStrategy
│ │ ├── Classification/
│ │ │ ├── SoftmaxClassificationEstimator.cs — Softmax post-processing
│ │ │ ├── OnnxTextClassificationEstimator.cs— Full pipeline facade
│ │ │ └── ... — Options, Transformer, ClassificationResult
│ │ ├── Reranking/
│ │ │ ├── SigmoidScorerEstimator.cs — Sigmoid scoring transform
│ │ │ ├── OnnxRerankerEstimator.cs — Cross-encoder facade
│ │ │ └── ... — Options, Transformer
│ │ ├── NER/
│ │ │ ├── NerDecodingEstimator.cs — BIO tag decoding to entity spans
│ │ │ ├── OnnxNerEstimator.cs — End-to-end NER facade
│ │ │ └── ... — Options, Transformer, NerEntity
│ │ └── QA/
│ │ ├── QaSpanExtractionEstimator.cs — Span extraction from start/end logits
│ │ ├── OnnxQaEstimator.cs — End-to-end QA facade
│ │ └── ... — Options, Transformer, QaResult
│ └── MLNet.TextGeneration.OnnxGenAI/
│ ├── OnnxTextGenerationEstimator.cs — ORT GenAI local text generation
│ ├── MLContextExtensions.cs — OnnxTextGeneration() extension method
│ └── ... — Options, Transformer
├── samples/
│ ├── BasicUsage/ — all-MiniLM-L6-v2: all embedding API surfaces
│ ├── BgeSmallEmbedding/ — BGE-small: query prefix pattern
│ ├── E5SmallEmbedding/ — E5-small: query/passage prefixes
│ ├── GteSmallEmbedding/ — GTE-small: semantic search (no prefix)
│ ├── ComposablePoolingComparison/ — 3 pooling strategies, shared inference
│ ├── IntermediateInspection/ — Inspect tokens, masks, raw output
│ ├── MeaiProviderAgnostic/ — Provider-agnostic MEAI transform
│ ├── Classification/
│ │ ├── SentimentDistilBERT/ — Sentiment analysis with DistilBERT
│ │ ├── EmotionRoBERTa/ — Multi-class emotion classification
│ │ └── ZeroShotDeBERTa/ — Zero-shot NLI classification with DeBERTa
│ ├── Reranking/
│ │ ├── MsMarcoMiniLM/ — MS MARCO cross-encoder reranking
│ │ └── BgeReranker/ — BGE reranker for retrieval
│ ├── NER/
│ │ ├── BertBaseNER/ — Named entity recognition with BERT-base
│ │ └── MultilingualNER/ — Multilingual NER with XLM-RoBERTa
│ ├── QA/
│ │ ├── RobertaSquad2/ — Extractive QA with RoBERTa on SQuAD 2.0
│ │ └── MiniLMSquad2/ — Extractive QA with MiniLM on SQuAD 2.0
│ ├── TextGenerationMeai/ — Provider-agnostic text generation (IChatClient)
│ └── TextGenerationLocal/ — Local ORT GenAI text generation
├── docs/ — Detailed documentation
│ ├── architecture-decision-record.md — ADR for multi-task platform decisions
│ ├── architecture.md — Component walkthrough + pipeline stages
│ ├── design-decisions.md — Why every choice was made
│ ├── extending.md — How to modify, extend, and add new tasks
│ ├── tensor-deep-dive.md — System.Numerics.Tensors for AI workloads
│ └── references.md — All sources and further reading
├── proposals/ — Design proposals for the modular architecture
└── nuget.config — NuGet source (nuget.org only)
| Sample | Model | Pattern Demonstrated |
|---|---|---|
| BasicUsage | all-MiniLM-L6-v2 | All API surfaces: facade, composable pipeline, .Append(), save/load, MEAI |
| BgeSmallEmbedding | BGE-small-en-v1.5 | Composable pipeline + BGE query prefix for asymmetric retrieval |
| E5SmallEmbedding | E5-small-v2 | Composable pipeline + E5 dual query:/passage: prefix pattern |
| GteSmallEmbedding | GTE-small | Composable pipeline + semantic search (no prefix needed) |
| ComposablePoolingComparison | all-MiniLM-L6-v2 | 3 pooling strategies, shared tokenizer+scorer (key modularization demo) |
| IntermediateInspection | all-MiniLM-L6-v2 | Inspect token IDs, attention masks, raw output at each pipeline stage |
| MeaiProviderAgnostic | all-MiniLM-L6-v2 | EmbeddingGeneratorEstimator wrapping IEmbeddingGenerator |
| SentimentDistilBERT | DistilBERT-SST2 | Sentiment classification with softmax post-processing |
| EmotionRoBERTa | RoBERTa-emotion | Multi-class emotion classification |
| ZeroShotDeBERTa | DeBERTa-v3-NLI | Zero-shot classification via NLI entailment |
| MsMarcoMiniLM | MS MARCO MiniLM | Cross-encoder reranking with text-pair tokenization |
| BgeReranker | BAAI/bge-reranker | BGE cross-encoder reranking for retrieval |
| BertBaseNER | BERT-base-NER | Named entity recognition with BIO tag decoding |
| MultilingualNER | XLM-RoBERTa-NER | Multilingual named entity recognition |
| RobertaSquad2 | RoBERTa-SQuAD2 | Extractive question answering with span extraction |
| MiniLMSquad2 | MiniLM-SQuAD2 | Extractive question answering (lighter model) |
| TextGenerationMeai | Any IChatClient | Provider-agnostic text generation via MEAI |
| TextGenerationLocal | ORT GenAI model | Local text generation with ONNX Runtime GenAI |
| Class | Role | Key Methods |
|---|---|---|
TextTokenizerEstimator | Transform 1: Tokenization | Fit(IDataView) |
OnnxTextModelScorerEstimator | Transform 2: ONNX Scoring | Fit(IDataView) → .HiddenDim, .HasPooledOutput |
| Class | Role | Key Methods |
|---|---|---|
EmbeddingPoolingEstimator | Transform 3: Pooling | Fit(IDataView) |
OnnxTextEmbeddingEstimator | Facade (chains 1→2→3) | Fit(IDataView), GetOutputSchema() |
OnnxTextEmbeddingTransformer | Facade transformer | Transform(IDataView), Save(path), Load(ctx, path) |
EmbeddingGeneratorEstimator | MEAI wrapper transform | Fit(IDataView) |
OnnxEmbeddingGenerator | MEAI IEmbeddingGenerator | GenerateAsync(texts) |
| Class | Role | Key Methods |
|---|---|---|
SoftmaxClassificationEstimator | Softmax post-processing | Fit(IDataView) |
OnnxTextClassificationEstimator | Facade (tokenize→score→softmax) | Fit(IDataView) |
OnnxTextClassificationTransformer | Facade transformer | Transform(IDataView), Classify(texts) |
| Class | Role | Key Methods |
|---|---|---|
SigmoidScorerEstimator | Sigmoid scoring transform | Fit(IDataView) |
OnnxRerankerEstimator | Facade (text-pair tokenize→score→sigmoid) | Fit(IDataView) |
OnnxRerankerTransformer | Facade transformer | Transform(IDataView), Rerank(query, docs) |
| Class | Role | Key Methods |
|---|---|---|
NerDecodingEstimator | BIO tag decoding | Fit(IDataView) |
OnnxNerEstimator | Facade (tokenize→score→decode) | Fit(IDataView) |
OnnxNerTransformer | Facade transformer | Transform(IDataView), RecognizeEntities(texts) |
| Class | Role | Key Methods |
|---|---|---|
QaSpanExtractionEstimator | Span extraction from logits | Fit(IDataView) |
OnnxQaEstimator | Facade (tokenize→multi-score→extract) | Fit(IDataView) |
OnnxQaTransformer | Facade transformer | Transform(IDataView), Answer(questions) |
| Surface | Types and extensions | Role |
|---|---|---|
| ML.NET facade | OnnxTypedDecisionsOptions, OnnxTypedDecisionsEstimator, OnnxTypedDecisionsTransformer | Schema-aware lazy end-to-end transform |
| ML.NET stages | DecisionInputPreparation*, OnnxDecisionModelScorer*, DecisionDecoding* | Options / Estimator / Transformer stage types |
| ML.NET composition | TransformsCatalog extensions and AppendOnnxTypedDecisions | Standard ML.NET pipeline composition |
The ML.NET facade adds question-specific typed columns. Choice exposes
PredictedLabel and a calibrated Probabilities vector; Score exposes the
expected zero-based Score and its probability vector; Noul exposes Boolean
PredictedLabel and Probability (P(true)). Each question also gets its
own entropy Confidence and ActionProbability. DecisionResults remains
an optional full diagnostic JSON column. Preparation and scoring stages use
native ML.NET numeric/vector/Boolean columns, and the facade is the normal
cursor-batched path. Native single-row PredictionEngine mapping is supported
through the transformer and native stage row mappers. Native ML.NET Save/Load
remains unsupported for typed decisions in this release; use the explicit
portable API instead. External assets are a packaging consideration, not an
inherent limitation of row mapping.
Typed-decision transformers provide an explicit portable ZIP API
(Save/Load and the supported facade/stage pipeline helpers); this is separate from
native MLContext.Model.Save/Load, which remains unsupported for these custom
components. The supported pipeline boundary is the demonstrated flat
preparation -> scoring -> decoding chain and naturally inferred appended
typed-decision facades; arbitrary/native chains are rejected. All sources in a
saved pipeline must reference the same complete asset payload, so separately
loaded selective stage archives cannot currently be recombined.
| Class | Role | Key Methods |
|---|---|---|
ChatClientEstimator | Provider-agnostic (wraps IChatClient) | Fit(IDataView) |
ChatClientTransformer | Provider-agnostic transformer | Transform(IDataView) |
OnnxTextGenerationEstimator | ORT GenAI local generation | Fit(IDataView) |
OnnxTextGenerationTransformer | ORT GenAI transformer | Transform(IDataView) |
The library supports GPU-accelerated ONNX inference via CUDA. The library itself ships with no native binaries — you control the execution provider by choosing your OnnxRuntime package.
GPU inference requires the CUDA Toolkit and cuDNN installed on the host machine, plus an NVIDIA GPU with a compatible driver.
Windows (winget + direct download):
# 1. Install CUDA Toolkit 12.6
winget install Nvidia.CUDA --version 12.6 --source winget
# 2. Download and install cuDNN 9.x for CUDA 12
# Download the zip from NVIDIA's redistributable endpoint:
$url = "https://developer.download.nvidia.com/compute/cudnn/redist/cudnn/windows-x86_64/cudnn-windows-x86_64-9.8.0.87_cuda12-archive.zip"
Invoke-WebRequest -Uri $url -OutFile "$env:TEMP\cudnn.zip"
Expand-Archive "$env:TEMP\cudnn.zip" -DestinationPath "$env:TEMP\cudnn"
# 3. Copy cuDNN DLLs to CUDA bin (requires admin)
Copy-Item "$env:TEMP\cudnn\cudnn-*\bin\*.dll" "$env:CUDA_PATH\bin" -Force
Linux (apt):
# CUDA Toolkit
sudo apt-get install -y nvidia-cuda-toolkit
# cuDNN (via NVIDIA's apt repository)
# See: https://docs.nvidia.com/deeplearning/cudnn/installation/linux.html
Verify installation:
nvidia-smi # Should show GPU + driver version
nvcc --version # Should show CUDA 12.x
Replace Microsoft.ML.OnnxRuntime with Microsoft.ML.OnnxRuntime.Gpu in your application:
<PackageReference Include="Microsoft.ML.OnnxRuntime.Gpu" Version="1.24.2" />
Samples auto-detect GPU: The sample projects use a
Directory.Build.propsthat checks forCUDA_PATH(set by the CUDA Toolkit installer) and automatically switches toMicrosoft.ML.OnnxRuntime.Gpu. Override withdotnet build -p:UseGpuRuntime=trueordotnet build -p:UseGpuRuntime=false.
// Pattern 1: Context-level (applies to all ONNX estimators)
var mlContext = new MLContext();
mlContext.GpuDeviceId = 0; // Use first CUDA device
var estimator = new OnnxTextEmbeddingEstimator(mlContext, new OnnxTextEmbeddingOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
});
// Pattern 2: Per-estimator override
var scorerOptions = new OnnxTextModelScorerOptions
{
ModelPath = "models/model.onnx",
GpuDeviceId = 0, // Override for this estimator only
FallbackToCpu = true, // Graceful degradation if CUDA unavailable
};
Resolution order: Per-estimator GpuDeviceId → MLContext.GpuDeviceId → null (CPU).
When FallbackToCpu = true, if CUDA initialization fails the estimator silently falls back to CPU execution.
| Package | Version | Purpose |
|---|---|---|
Microsoft.ML | 5.0.0 | IEstimator/ITransformer, IDataView, MLContext |
Microsoft.ML.OnnxRuntime.Managed | 1.24.2 | InferenceSession, OrtValue (managed API; bring your own native runtime) |
Microsoft.ML.Tokenizers | 2.0.0 | BertTokenizer (WordPiece), BPE, SentencePiece |
Microsoft.Extensions.AI.Abstractions | 10.3.0 | IEmbeddingGenerator, IChatClient |
System.Numerics.Tensors | 10.0.3 | Tensor<T>, TensorPrimitives |
| Package | Version | Purpose |
|---|---|---|
Microsoft.ML | 5.0.0 | IEstimator/ITransformer, IDataView, MLContext |
Microsoft.ML.OnnxRuntimeGenAI | 0.7.1 | ORT GenAI for local autoregressive generation |
Microsoft.Extensions.AI.Abstractions | 10.3.0 | IChatClient |
Pre-built packages are published to GitHub Packages. No need to copy source code.
1. Add the GitHub Packages NuGet source (one-time setup):
<!-- nuget.config in your project root -->
<configuration>
<packageSources>
<add key="nuget.org" value="https://api.nuget.org/v3/index.json" />
<add key="github-packages" value="https://nuget.pkg.github.com/luisquintanilla/index.json" />
</packageSources>
<packageSourceCredentials>
<github-packages>
<add key="Username" value="YOUR_GITHUB_USERNAME" />
<add key="ClearTextPassword" value="YOUR_GITHUB_PAT" />
</github-packages>
</packageSourceCredentials>
</configuration>
Note: GitHub Packages requires authentication even for public packages. Create a Personal Access Token with
read:packagesscope.
2. Add the package reference:
<!-- Encoder tasks: embeddings, classification, reranking, NER, QA, text generation -->
<PackageReference Include="MLNet.TextInference.Onnx" Version="0.1.0-preview.1" />
<!-- Local text generation with ORT GenAI (separate package, heavier dependency) -->
<PackageReference Include="MLNet.TextGeneration.OnnxGenAI" Version="0.1.0-preview.1" />
3. Use it:
using MLNet.TextInference.Onnx;
var mlContext = new MLContext();
var estimator = mlContext.Transforms.OnnxTextClassification(new OnnxTextClassificationOptions
{
ModelPath = "models/model.onnx",
TokenizerPath = "models/",
Labels = ["negative", "positive"]
});
Packages are versioned via git tags using MinVer:
git tag v0.1.0-preview.1
git push origin v0.1.0-preview.1
# → GitHub Actions builds, packs, and publishes to GitHub Packages
.NET 10 (LTS). Requires the .NET 10 SDK.
C#
97.9%
Shell
1.9%