Fast, dependency-light sentence embeddings for .NET. This library wraps a set of
ONNX embedding models behind a single, simple ISentenceEncoder interface so you
can turn text into vectors — for semantic search, clustering, retrieval-augmented generation (RAG),
deduplication, recommendations and similarity scoring — entirely in-process, with no Python runtime
and no external API calls.
It is built and maintained by Curiosity and powers the AI search / vector indexing features of Curiosity Workspace.
using SentenceTransformers.MiniLM;
using var encoder = new SentenceEncoder();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"The quick brown fox jumps over the lazy dog",
"A fast auburn fox leaps above a sleepy hound",
});
// vectors[0] and vectors[1] are L2-normalized float[384] embeddings.
using — every encoder implements
ISentenceEncoder.| Package | Model | Dimensions | Max tokens | Languages | Weights |
|---|---|---|---|---|---|
| Core interfaces & chunking helpers | — | — | — | — | |
| all-MiniLM-L6-v2 | 384 | 256 | English | Embedded | |
| snowflake-arctic-embed-xs | 384 | 512 | English | Embedded | |
| Qwen3-Embedding-0.6B | 1024 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-0.6b | 1024 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-270m | 640 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-270m (pure C#, no ONNX) | 640 | 32768 | Multilingual | Downloaded on first use |
Pick MiniLM for the smallest/fastest footprint, Arctic XS for a strong English default, Harrier Small when you need multilingual coverage without paying for the larger 0.6b weights, and Qwen3 or Harrier Medium when you want the highest-quality embeddings (1024 dim) and the full 32k-token context window.
Install only the model package(s) you need (the core SentenceTransformers package is pulled in as a
dependency):
dotnet add package SentenceTransformers.MiniLM
dotnet add package SentenceTransformers.ArcticXs
dotnet add package SentenceTransformers.Qwen3
dotnet add package SentenceTransformers.Harrier.Medium
dotnet add package SentenceTransformers.Harrier.Small
Targets .NET 10.
Embedded models are ready to use as soon as you construct them:
using SentenceTransformers.ArcticXs;
using var encoder = new SentenceEncoder();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"How do I reset my password?",
"I forgot my login credentials.",
});
Larger models download their ONNX weights on first use. Create them with the async CreateAsync
factory — the download is cached, so subsequent runs are instant:
using SentenceTransformers.Qwen3;
// Downloads the model to a temp folder on first use, then loads it.
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[] { "Hello world" });
// vectors[0] is a float[1024]
Harrier Medium is multilingual:
using SentenceTransformers.Harrier.Medium;
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"Good morning", // English
"Buenos días", // Spanish
"おはよう", // Japanese
});
Harrier Small is the same multilingual family at ~270M parameters (640-dim embeddings), suitable when you want multilingual coverage without paying for the 0.6b weights:
using SentenceTransformers.Harrier.Small;
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"Good morning",
"Buenos días",
"おはよう",
});
// vectors[0] is a float[640]
SentenceTransformers.Harrier.Small.Pure is a 100% managed reimplementation of Harrier Small. It runs
the Gemma3 forward pass and the Gemma BPE tokenizer entirely in C# (on top of
System.Numerics.Tensors), with no ONNX Runtime and no native tokenizer — so there is not a single
.so/.dll/.dylib to ship. That makes it trim/AOT-friendly and portable to anywhere .NET runs,
including Blazor WebAssembly and mobile. The API mirrors the ONNX package:
using SentenceTransformers.Harrier.Small.Pure;
// Downloads the original bfloat16 safetensors weights (~540 MB) on first use, then loads them.
using var encoder = await SentenceEncoder.CreateAsync();
// Queries take a task instruction prefix; documents are encoded as-is.
float[][] queryVectors = await encoder.EncodeQueriesAsync(
new[] { "how much protein should a female eat" },
SentenceEncoder.Prompts.WebSearchQuery);
float[][] docVectors = await encoder.EncodeAsync(new[] { "…a passage about dietary protein…" });
// vectors are L2-normalized float[640]
It produces the same embeddings as the reference: pure fp32 reproduces the query/document
similarity matrix published on the model card
to within 0.01 — actually closer to the reference than the shipped ONNX Q4F16 build.
Choosing a quantization. The transformer weights can be loaded at reduced precision to cut both
memory and inference time. Pass a Quantization to CreateAsync (or the constructor):
using SentenceTransformers.Harrier.Small.Pure;
using SentenceTransformers.Harrier.Small.Pure.Model;
// fp32 (default, most faithful), Int8 (recommended — fastest & ~40% less memory), or Int4 (smallest).
using var encoder = await SentenceEncoder.CreateAsync(quantization: Quantization.Int8);
The Int8/Int4 paths run as true int8 GEMMs and pick the best instruction set available at runtime:
vpdpbsud on 512-bit registers (AvxVnniInt8.V512, 64 int8 MACs/instruction), vpdpbusd/vpdpbsud
on 256-bit (AvxVnni/AvxVnniInt8), or a widen + vpmaddwd sequence on AVX-512 / AVX2 CPUs. On an
AVX-512 host, also set DOTNET_PreferredVectorBitWidth=512 to let the JIT emit 512-bit vectors.
Benchmark — pure C# vs ONNX (harrier-oss-v1-270m, .NET 10, 4-core Xeon, single-text encode):
| Variant | Native deps | Max err vs model card¹ | Short text | ~512-token text | Resident weights² |
|---|---|---|---|---|---|
ONNX Q4F16 (SentenceTransformers.Harrier.Small) | ONNX Runtime | 2.30 | ~40 ms | ~0.7 s | ~172 MB (native) |
| Pure fp32 | none | 0.01 | 226 ms | 4.4 s | ~740 MB |
| Pure int8 | none | 0.97 | 86 ms | 1.6 s | ~440 MB |
| Pure int4 | none | 1.42 | 260 ms | 3.8 s | ~390 MB |
¹ Largest absolute deviation (on a 0–100 cosine×100 scale) from the published query/document score
matrix; lower is more faithful — every pure variant tracks the reference more closely than the ONNX
Q4F16 weights. ² Approximate model-weight memory; all pure variants share the same bfloat16
token-embedding table (~335 MB), which is the floor. Numbers are hardware-dependent — the run above
is an AVX-512 server CPU where the int8 dot uses widen + vpmaddwd; CPUs with the int8-VNNI
instructions are substantially faster (see below).
How to read this. Int8 is the recommended pure setting — ~2.5× faster than fp32, ~40 % smaller, and
still more faithful than ONNX Q4F16.
How close it gets to ONNX depends on the CPU's int8 instruction set. ONNX Runtime's MLAS uses hand-tuned
assembly with 4-bit weights and int8-VNNI (vpdpbusd/vpdpbsud). The pure build emits int8-VNNI
too when the runtime exposes it: AvxVnni (256-bit, Alder Lake and newer client CPUs) or
AvxVnniInt8.V512 (512-bit, AVX10.2 / Granite Rapids-class CPUs) — on those it is within ~1.5–2× of ONNX
and can approach parity with the 512-bit path. The one gap is "classic" AVX-512 servers that report the
avx512_vnni CPUID flag but not AvxVnni/AvxVnniInt8/AVX10: .NET 10 has no standalone Avx512Vnni
intrinsic, so there the pure build must fall back to widen + vpmaddwd (~6× the instructions) and lands
~2.5× off ONNX. Either way the pure build's case is zero native dependencies (trim/AOT/WASM/mobile,
one managed package) and higher fidelity, at a CPU-inference cost within a small multiple of ONNX.
Because every model returns L2-normalized vectors, cosine similarity is simply the dot product:
static float CosineSimilarity(float[] a, float[] b)
{
float dot = 0f;
for (int i = 0; i < a.Length; i++)
{
dot += a[i] * b[i];
}
return dot; // vectors are unit-length, so dot product == cosine similarity
}
using var encoder = new SentenceTransformers.MiniLM.SentenceEncoder();
var v = await encoder.EncodeAsync(new[] { "cat", "kitten", "spaceship" });
Console.WriteLine(CosineSimilarity(v[0], v[1])); // cat vs kitten -> high
Console.WriteLine(CosineSimilarity(v[0], v[2])); // cat vs spaceship -> low
EncodeAsync expects each input to fit within the model's context window
(encoder.MaxChunkLength tokens). For longer text, use the built-in chunking helpers, which split on
token boundaries and embed each chunk:
using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync();
EncodedChunk[] chunks = await encoder.ChunkAndEncodeAsync(
longDocument,
chunkLength: 512, // tokens per chunk (clamped to MaxChunkLength)
chunkOverlap: 64, // tokens of overlap between consecutive chunks
reportProgress: p => Console.WriteLine($"{p:P0}"));
foreach (var chunk in chunks)
{
// chunk.Text -> the chunk's source text
// chunk.Vector -> its embedding
Index(chunk.Text, chunk.Vector);
}
Need to map results back to their position in the original text (e.g. to highlight a passage)? Use
ChunkAndEncodeAlignedAsync, which additionally returns character offsets (Start, LastStart,
ApproximateEnd) into the source. There are also tagged variants
(ChunkAndEncodeTaggedAsync / …AlignedAsync) for carrying per-chunk metadata such as page numbers
through the chunking pipeline.
For the downloaded models you can control the cache location and the source URL:
using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync(
downloadToPath: "/var/models/qwen3.onnx");
// Or point at your own mirror:
using var harrier = await SentenceTransformers.Harrier.Medium.SentenceEncoder.CreateAsync(
modelUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx",
modelDataUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx_data");
You can also pass a custom Microsoft.ML.OnnxRuntime.SessionOptions to any constructor / CreateAsync
to tune threading or enable hardware execution providers.
Both Harrier packages ship multiple quantization formats — pick the one that fits your CPU / GPU
memory budget. URLs for every variant are exposed as constants on SentenceEncoder.Quantizations:
| Variant | Constant | Harrier Medium (0.6b) weights | Harrier Small (270m) weights |
|---|---|---|---|
| Full (fp32) | Quantizations.FullModelUrl | 2.09 GB (+306 MB) | 1.11 GB |
| FP16 | Quantizations.Fp16ModelUrl | 1.20 GB | 553 MB |
| Q4 | Quantizations.Q4ModelUrl | 399 MB | 205 MB |
| Q4 + FP16 | Quantizations.Q4Fp16ModelUrl (default) | 353 MB | 172 MB |
| Quantized | Quantizations.QuantizedModelUrl | 706 MB | 344 MB |
Q4F16 is the default — it's the smallest variant on disk and keeps multilingual retrieval quality
close to the unquantized reference. Pick Quantized when you need a pure float32 output and
broader ONNX-Runtime execution-provider compatibility; FP16 for the most precision per byte;
Full for the unquantized reference. Each entry has a matching …ModelDataUrl constant for the
external weights file (and the Harrier Medium Full variant additionally has FullModelDataUrl2
because its fp32 weights are split into two files).
using SentenceTransformers.Harrier.Small;
// Use the unquantized reference instead of the default Q4F16:
using var encoder = await SentenceEncoder.CreateAsync(
modelUrl: SentenceEncoder.Quantizations.FullModelUrl,
modelDataUrl: SentenceEncoder.Quantizations.FullModelDataUrl);
📖 See LORA.md for the full guide: how it works internally, every option, negative-sample handling, stop conditions, and library + CLI training walkthroughs.
You can specialize the pure-C# encoders — MiniLM, Arctic XS and Harrier Small — for a specific
domain (support tickets, legal clauses, product descriptions, a particular language pair) by training a
real weight-space LoRA adapter from a set of related pairs (a query and a relevant passage, two
paraphrases, a question and its duplicate…). This is proper LoRA: small low-rank factors are injected
inside the transformer's attention/MLP projections (W + (α/r)·B·A), and the loss is backpropagated
through the whole (frozen) network. The forward and backward passes run entirely in pure C# — no ONNX
Runtime, and no PyTorch/autodiff dependency. A shared tensor autograd engine (SentenceTransformers.Training.Autograd)
powers both model families.
MiniLM / Arctic XS (
SentenceTransformers.Bert.Pure, theBertModelarchitecture) read their full-precision weights directly from the fp32 ONNX graphs already embedded in the MiniLM / Arctic packages, so training is fully self-contained — no model download.Harrier Small (
SentenceTransformers.Harrier.Small.Pure, a Gemma3 decoder) trains through the same autograd engine and LoRA objectives viaGemma3LoraTrainer/Gemma3LoraEncoder; its bf16 weights are downloaded on first use, and because it is a ~270M-param decoder each step is much heavier than the small BERTs (keep the batch / sequence length modest). LoRA there targets q/k/v/o and the gate/up/down MLP projections.The remaining ONNX-only models (Qwen3, Harrier Medium) are inference-only and are not trainable here.
using SentenceTransformers.Bert.Pure;
using SentenceTransformers.Bert.Pure.Model;
using SentenceTransformers.Bert.Pure.Training;
using SentenceTransformers.Training;
// Pure-C# MiniLM (weights fetched once from HuggingFace, or use SentenceEncoder.LoadFromOnnx to reuse
// the weights embedded in the SentenceTransformers.MiniLM package with no download at all).
using var baseEncoder = await SentenceEncoder.CreateMiniLMAsync();
var dataset = new SentencePairDataset(new[]
{
new SentencePair("how do I reset my password", "Use the ‘Forgot password’ link on the sign-in page.", 0.95f),
new SentencePair("cancel my subscription", "Go to Billing → Manage plan → Cancel.", 0.95f),
// … a few hundred to a few thousand related (optionally scored) pairs …
});
var report = await BertLoraTrainer.TrainAsync(baseEncoder, dataset, new BertLoraTrainingOptions
{
Objective = BertTrainingObjective.CoSent, // pairwise ranking loss — directly targets STS Spearman
Rank = 8, // LoRA rank
Targets = LoraTargets.Attention, // inject into Q, K, V and the attention output
Epochs = 10,
});
report.Adapter.Save("support-faq.lora");
// The tuned model is a drop-in ISentenceEncoder (the adapter is folded into the transformer at runtime):
using var tuned = baseEncoder.WithAdapter(report.Adapter);
float[][] vectors = await tuned.EncodeAsync(new[] { "I can't log in" });
How it works. A tiny tensor-level autograd engine records the BERT forward pass; the exact gradients for the injected LoRA factors are obtained by replaying the tape backward (verified against finite differences in the test-suite). Only the low-rank factors — and optionally a learned output-centering bias — are trainable; every base weight stays frozen, so the gradients for the frozen network are skipped (the LoRA efficiency win). The held-out validation set is scored with STS Spearman and top-1 retrieval accuracy and the best adapter is kept.
Objectives (BertLoraTrainingOptions.Objective):
Contrastive (default) — symmetric InfoNCE (MultipleNegativesRanking) with in-batch and hard
negatives (explicit triplets via SentencePair.Negative, plus optional mined negatives), with
false-negative masking. Best for retrieval / nearest-neighbour separation.CoSent — the CoSENT pairwise-ranking loss over graded pairs. It optimizes the ordering that Spearman
measures and consistently beats plain cosine-MSE on STS. Best when you care about graded similarity.CosineRegression — mean-squared error between adapted cosine and the gold [0,1] score.Extras (all optional): warmup + cosine LR schedule, a learned temperature (LearnableTemperature),
multi-seed model selection (NumSeeds), a learned output-centering bias (UseOutputBias) and post-hoc
ZCA whitening (ApplyWhitening) to counter embedding anisotropy, Matryoshka sub-dimension losses
(MatryoshkaDims), asymmetric instruction prefixes (QueryPrefix / DocumentPrefix), and a choice of
which projections carry adapters (Targets: Attention, Mlp or All).
The SentenceTransformers.LoraTraining project is a ready-to-run console app that fine-tunes MiniLM,
Arctic XS or Harrier Small against one of two bundled example datasets (--dataset):
stsb — the English STS Benchmark, a broad
general-English similarity set, downloaded on demand.patent — the Google Patent Phrase Similarity
dataset (CC BY 4.0), embedded directly in the app (no download). Terse, domain-specific technical
phrases where a general encoder has real headroom.cd SentenceTransformers.LoraTraining
# General-English STS Benchmark (needs a one-time dataset download; model weights are embedded):
dotnet run -c Release -- download
dotnet run -c Release -- train --model minilm --objective cosent --rank 8 --epochs 10 --output-bias
dotnet run -c Release -- eval --model minilm --adapter ./adapters/minilm-stsb.lora --split test
# Domain-specific patent phrases (embedded, no download) — graded scores:
dotnet run -c Release -- train --model arctic --dataset patent --objective cosent --rank 8 --whitening
dotnet run -c Release -- eval --model arctic --dataset patent --adapter ./adapters/arctic-patent.lora --split test
# Harrier Small (Gemma3, pure C#) — bf16 weights downloaded on first use; heavier, so keep batch small:
dotnet run -c Release -- train --model harrier-small --dataset patent --objective cosent --rank 8 --batch 8 --max-tokens 64
--model accepts minilm, arctic or harrier-small. Run dotnet run -- help for the full option list (objective,
targets, rank, α, learning rate, warmup, temperature, mined negatives, learnable temperature, output bias,
whitening, Matryoshka, query/doc prefixes, multi-seed, …). Training prints per-epoch validation retrieval
accuracy and STS Spearman and a base-vs-tuned summary at the end.
Each model package contains:
SentenceEncoder that tokenizes, runs ONNX Runtime inference, pools the token outputs and
L2-normalizes the result.The shared SentenceTransformers package provides the ISentenceEncoder
contract and the default-implemented chunking helpers, so the model packages only implement
EncodeAsync and expose their MaxChunkLength / Tokenizer.
The SentenceTransformers.Harrier.Small.Pure package is the exception: instead of ONNX Runtime it
ships its own pure-managed implementation of the model. It reads the original
safetensors weights directly, runs the Gemma3
decoder forward pass (RMSNorm, grouped-query attention with Q/K-norm and RoPE, GeGLU MLP, last-token
pooling) on TensorPrimitives,
and tokenizes with a from-scratch Gemma byte-level BPE tokenizer — so it depends only on the .NET base
class library and System.Numerics.Tensors.
dotnet build SentenceTransformers.sln -c Release
dotnet test SentenceTransformers.sln
NuGet packages are produced and published by the Azure DevOps pipeline in
.devops/azure-pipelines.yml on pushes to main.
MIT. The BERT tokenizers are derived from
BERTTokenizers (MIT, © 2021 Othneil Drew). Each wrapped
model is distributed under its own upstream license — see the linked Hugging Face model pages. The
Google Patent Phrase Similarity
dataset bundled with the SentenceTransformers.LoraTraining example is © Google, licensed
CC BY 4.0.
Fast, dependency-light sentence embeddings for .NET. This library wraps a set of
ONNX embedding models behind a single, simple ISentenceEncoder interface so you
can turn text into vectors — for semantic search, clustering, retrieval-augmented generation (RAG),
deduplication, recommendations and similarity scoring — entirely in-process, with no Python runtime
and no external API calls.
It is built and maintained by Curiosity and powers the AI search / vector indexing features of Curiosity Workspace.
using SentenceTransformers.MiniLM;
using var encoder = new SentenceEncoder();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"The quick brown fox jumps over the lazy dog",
"A fast auburn fox leaps above a sleepy hound",
});
// vectors[0] and vectors[1] are L2-normalized float[384] embeddings.
using — every encoder implements
ISentenceEncoder.| Package | Model | Dimensions | Max tokens | Languages | Weights |
|---|---|---|---|---|---|
| Core interfaces & chunking helpers | — | — | — | — | |
| all-MiniLM-L6-v2 | 384 | 256 | English | Embedded | |
| snowflake-arctic-embed-xs | 384 | 512 | English | Embedded | |
| Qwen3-Embedding-0.6B | 1024 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-0.6b | 1024 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-270m | 640 | 32768 | Multilingual | Downloaded on first use | |
| harrier-oss-v1-270m (pure C#, no ONNX) | 640 | 32768 | Multilingual | Downloaded on first use |
Pick MiniLM for the smallest/fastest footprint, Arctic XS for a strong English default, Harrier Small when you need multilingual coverage without paying for the larger 0.6b weights, and Qwen3 or Harrier Medium when you want the highest-quality embeddings (1024 dim) and the full 32k-token context window.
Install only the model package(s) you need (the core SentenceTransformers package is pulled in as a
dependency):
dotnet add package SentenceTransformers.MiniLM
dotnet add package SentenceTransformers.ArcticXs
dotnet add package SentenceTransformers.Qwen3
dotnet add package SentenceTransformers.Harrier.Medium
dotnet add package SentenceTransformers.Harrier.Small
Targets .NET 10.
Embedded models are ready to use as soon as you construct them:
using SentenceTransformers.ArcticXs;
using var encoder = new SentenceEncoder();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"How do I reset my password?",
"I forgot my login credentials.",
});
Larger models download their ONNX weights on first use. Create them with the async CreateAsync
factory — the download is cached, so subsequent runs are instant:
using SentenceTransformers.Qwen3;
// Downloads the model to a temp folder on first use, then loads it.
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[] { "Hello world" });
// vectors[0] is a float[1024]
Harrier Medium is multilingual:
using SentenceTransformers.Harrier.Medium;
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"Good morning", // English
"Buenos días", // Spanish
"おはよう", // Japanese
});
Harrier Small is the same multilingual family at ~270M parameters (640-dim embeddings), suitable when you want multilingual coverage without paying for the 0.6b weights:
using SentenceTransformers.Harrier.Small;
using var encoder = await SentenceEncoder.CreateAsync();
float[][] vectors = await encoder.EncodeAsync(new[]
{
"Good morning",
"Buenos días",
"おはよう",
});
// vectors[0] is a float[640]
SentenceTransformers.Harrier.Small.Pure is a 100% managed reimplementation of Harrier Small. It runs
the Gemma3 forward pass and the Gemma BPE tokenizer entirely in C# (on top of
System.Numerics.Tensors), with no ONNX Runtime and no native tokenizer — so there is not a single
.so/.dll/.dylib to ship. That makes it trim/AOT-friendly and portable to anywhere .NET runs,
including Blazor WebAssembly and mobile. The API mirrors the ONNX package:
using SentenceTransformers.Harrier.Small.Pure;
// Downloads the original bfloat16 safetensors weights (~540 MB) on first use, then loads them.
using var encoder = await SentenceEncoder.CreateAsync();
// Queries take a task instruction prefix; documents are encoded as-is.
float[][] queryVectors = await encoder.EncodeQueriesAsync(
new[] { "how much protein should a female eat" },
SentenceEncoder.Prompts.WebSearchQuery);
float[][] docVectors = await encoder.EncodeAsync(new[] { "…a passage about dietary protein…" });
// vectors are L2-normalized float[640]
It produces the same embeddings as the reference: pure fp32 reproduces the query/document
similarity matrix published on the model card
to within 0.01 — actually closer to the reference than the shipped ONNX Q4F16 build.
Choosing a quantization. The transformer weights can be loaded at reduced precision to cut both
memory and inference time. Pass a Quantization to CreateAsync (or the constructor):
using SentenceTransformers.Harrier.Small.Pure;
using SentenceTransformers.Harrier.Small.Pure.Model;
// fp32 (default, most faithful), Int8 (recommended — fastest & ~40% less memory), or Int4 (smallest).
using var encoder = await SentenceEncoder.CreateAsync(quantization: Quantization.Int8);
The Int8/Int4 paths run as true int8 GEMMs and pick the best instruction set available at runtime:
vpdpbsud on 512-bit registers (AvxVnniInt8.V512, 64 int8 MACs/instruction), vpdpbusd/vpdpbsud
on 256-bit (AvxVnni/AvxVnniInt8), or a widen + vpmaddwd sequence on AVX-512 / AVX2 CPUs. On an
AVX-512 host, also set DOTNET_PreferredVectorBitWidth=512 to let the JIT emit 512-bit vectors.
Benchmark — pure C# vs ONNX (harrier-oss-v1-270m, .NET 10, 4-core Xeon, single-text encode):
| Variant | Native deps | Max err vs model card¹ | Short text | ~512-token text | Resident weights² |
|---|---|---|---|---|---|
ONNX Q4F16 (SentenceTransformers.Harrier.Small) | ONNX Runtime | 2.30 | ~40 ms | ~0.7 s | ~172 MB (native) |
| Pure fp32 | none | 0.01 | 226 ms | 4.4 s | ~740 MB |
| Pure int8 | none | 0.97 | 86 ms | 1.6 s | ~440 MB |
| Pure int4 | none | 1.42 | 260 ms | 3.8 s | ~390 MB |
¹ Largest absolute deviation (on a 0–100 cosine×100 scale) from the published query/document score
matrix; lower is more faithful — every pure variant tracks the reference more closely than the ONNX
Q4F16 weights. ² Approximate model-weight memory; all pure variants share the same bfloat16
token-embedding table (~335 MB), which is the floor. Numbers are hardware-dependent — the run above
is an AVX-512 server CPU where the int8 dot uses widen + vpmaddwd; CPUs with the int8-VNNI
instructions are substantially faster (see below).
How to read this. Int8 is the recommended pure setting — ~2.5× faster than fp32, ~40 % smaller, and
still more faithful than ONNX Q4F16.
How close it gets to ONNX depends on the CPU's int8 instruction set. ONNX Runtime's MLAS uses hand-tuned
assembly with 4-bit weights and int8-VNNI (vpdpbusd/vpdpbsud). The pure build emits int8-VNNI
too when the runtime exposes it: AvxVnni (256-bit, Alder Lake and newer client CPUs) or
AvxVnniInt8.V512 (512-bit, AVX10.2 / Granite Rapids-class CPUs) — on those it is within ~1.5–2× of ONNX
and can approach parity with the 512-bit path. The one gap is "classic" AVX-512 servers that report the
avx512_vnni CPUID flag but not AvxVnni/AvxVnniInt8/AVX10: .NET 10 has no standalone Avx512Vnni
intrinsic, so there the pure build must fall back to widen + vpmaddwd (~6× the instructions) and lands
~2.5× off ONNX. Either way the pure build's case is zero native dependencies (trim/AOT/WASM/mobile,
one managed package) and higher fidelity, at a CPU-inference cost within a small multiple of ONNX.
Because every model returns L2-normalized vectors, cosine similarity is simply the dot product:
static float CosineSimilarity(float[] a, float[] b)
{
float dot = 0f;
for (int i = 0; i < a.Length; i++)
{
dot += a[i] * b[i];
}
return dot; // vectors are unit-length, so dot product == cosine similarity
}
using var encoder = new SentenceTransformers.MiniLM.SentenceEncoder();
var v = await encoder.EncodeAsync(new[] { "cat", "kitten", "spaceship" });
Console.WriteLine(CosineSimilarity(v[0], v[1])); // cat vs kitten -> high
Console.WriteLine(CosineSimilarity(v[0], v[2])); // cat vs spaceship -> low
EncodeAsync expects each input to fit within the model's context window
(encoder.MaxChunkLength tokens). For longer text, use the built-in chunking helpers, which split on
token boundaries and embed each chunk:
using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync();
EncodedChunk[] chunks = await encoder.ChunkAndEncodeAsync(
longDocument,
chunkLength: 512, // tokens per chunk (clamped to MaxChunkLength)
chunkOverlap: 64, // tokens of overlap between consecutive chunks
reportProgress: p => Console.WriteLine($"{p:P0}"));
foreach (var chunk in chunks)
{
// chunk.Text -> the chunk's source text
// chunk.Vector -> its embedding
Index(chunk.Text, chunk.Vector);
}
Need to map results back to their position in the original text (e.g. to highlight a passage)? Use
ChunkAndEncodeAlignedAsync, which additionally returns character offsets (Start, LastStart,
ApproximateEnd) into the source. There are also tagged variants
(ChunkAndEncodeTaggedAsync / …AlignedAsync) for carrying per-chunk metadata such as page numbers
through the chunking pipeline.
For the downloaded models you can control the cache location and the source URL:
using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync(
downloadToPath: "/var/models/qwen3.onnx");
// Or point at your own mirror:
using var harrier = await SentenceTransformers.Harrier.Medium.SentenceEncoder.CreateAsync(
modelUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx",
modelDataUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx_data");
You can also pass a custom Microsoft.ML.OnnxRuntime.SessionOptions to any constructor / CreateAsync
to tune threading or enable hardware execution providers.
Both Harrier packages ship multiple quantization formats — pick the one that fits your CPU / GPU
memory budget. URLs for every variant are exposed as constants on SentenceEncoder.Quantizations:
| Variant | Constant | Harrier Medium (0.6b) weights | Harrier Small (270m) weights |
|---|---|---|---|
| Full (fp32) | Quantizations.FullModelUrl | 2.09 GB (+306 MB) | 1.11 GB |
| FP16 | Quantizations.Fp16ModelUrl | 1.20 GB | 553 MB |
| Q4 | Quantizations.Q4ModelUrl | 399 MB | 205 MB |
| Q4 + FP16 | Quantizations.Q4Fp16ModelUrl (default) | 353 MB | 172 MB |
| Quantized | Quantizations.QuantizedModelUrl | 706 MB | 344 MB |
Q4F16 is the default — it's the smallest variant on disk and keeps multilingual retrieval quality
close to the unquantized reference. Pick Quantized when you need a pure float32 output and
broader ONNX-Runtime execution-provider compatibility; FP16 for the most precision per byte;
Full for the unquantized reference. Each entry has a matching …ModelDataUrl constant for the
external weights file (and the Harrier Medium Full variant additionally has FullModelDataUrl2
because its fp32 weights are split into two files).
using SentenceTransformers.Harrier.Small;
// Use the unquantized reference instead of the default Q4F16:
using var encoder = await SentenceEncoder.CreateAsync(
modelUrl: SentenceEncoder.Quantizations.FullModelUrl,
modelDataUrl: SentenceEncoder.Quantizations.FullModelDataUrl);
📖 See LORA.md for the full guide: how it works internally, every option, negative-sample handling, stop conditions, and library + CLI training walkthroughs.
You can specialize the pure-C# encoders — MiniLM, Arctic XS and Harrier Small — for a specific
domain (support tickets, legal clauses, product descriptions, a particular language pair) by training a
real weight-space LoRA adapter from a set of related pairs (a query and a relevant passage, two
paraphrases, a question and its duplicate…). This is proper LoRA: small low-rank factors are injected
inside the transformer's attention/MLP projections (W + (α/r)·B·A), and the loss is backpropagated
through the whole (frozen) network. The forward and backward passes run entirely in pure C# — no ONNX
Runtime, and no PyTorch/autodiff dependency. A shared tensor autograd engine (SentenceTransformers.Training.Autograd)
powers both model families.
MiniLM / Arctic XS (
SentenceTransformers.Bert.Pure, theBertModelarchitecture) read their full-precision weights directly from the fp32 ONNX graphs already embedded in the MiniLM / Arctic packages, so training is fully self-contained — no model download.Harrier Small (
SentenceTransformers.Harrier.Small.Pure, a Gemma3 decoder) trains through the same autograd engine and LoRA objectives viaGemma3LoraTrainer/Gemma3LoraEncoder; its bf16 weights are downloaded on first use, and because it is a ~270M-param decoder each step is much heavier than the small BERTs (keep the batch / sequence length modest). LoRA there targets q/k/v/o and the gate/up/down MLP projections.The remaining ONNX-only models (Qwen3, Harrier Medium) are inference-only and are not trainable here.
using SentenceTransformers.Bert.Pure;
using SentenceTransformers.Bert.Pure.Model;
using SentenceTransformers.Bert.Pure.Training;
using SentenceTransformers.Training;
// Pure-C# MiniLM (weights fetched once from HuggingFace, or use SentenceEncoder.LoadFromOnnx to reuse
// the weights embedded in the SentenceTransformers.MiniLM package with no download at all).
using var baseEncoder = await SentenceEncoder.CreateMiniLMAsync();
var dataset = new SentencePairDataset(new[]
{
new SentencePair("how do I reset my password", "Use the ‘Forgot password’ link on the sign-in page.", 0.95f),
new SentencePair("cancel my subscription", "Go to Billing → Manage plan → Cancel.", 0.95f),
// … a few hundred to a few thousand related (optionally scored) pairs …
});
var report = await BertLoraTrainer.TrainAsync(baseEncoder, dataset, new BertLoraTrainingOptions
{
Objective = BertTrainingObjective.CoSent, // pairwise ranking loss — directly targets STS Spearman
Rank = 8, // LoRA rank
Targets = LoraTargets.Attention, // inject into Q, K, V and the attention output
Epochs = 10,
});
report.Adapter.Save("support-faq.lora");
// The tuned model is a drop-in ISentenceEncoder (the adapter is folded into the transformer at runtime):
using var tuned = baseEncoder.WithAdapter(report.Adapter);
float[][] vectors = await tuned.EncodeAsync(new[] { "I can't log in" });
How it works. A tiny tensor-level autograd engine records the BERT forward pass; the exact gradients for the injected LoRA factors are obtained by replaying the tape backward (verified against finite differences in the test-suite). Only the low-rank factors — and optionally a learned output-centering bias — are trainable; every base weight stays frozen, so the gradients for the frozen network are skipped (the LoRA efficiency win). The held-out validation set is scored with STS Spearman and top-1 retrieval accuracy and the best adapter is kept.
Objectives (BertLoraTrainingOptions.Objective):
Contrastive (default) — symmetric InfoNCE (MultipleNegativesRanking) with in-batch and hard
negatives (explicit triplets via SentencePair.Negative, plus optional mined negatives), with
false-negative masking. Best for retrieval / nearest-neighbour separation.CoSent — the CoSENT pairwise-ranking loss over graded pairs. It optimizes the ordering that Spearman
measures and consistently beats plain cosine-MSE on STS. Best when you care about graded similarity.CosineRegression — mean-squared error between adapted cosine and the gold [0,1] score.Extras (all optional): warmup + cosine LR schedule, a learned temperature (LearnableTemperature),
multi-seed model selection (NumSeeds), a learned output-centering bias (UseOutputBias) and post-hoc
ZCA whitening (ApplyWhitening) to counter embedding anisotropy, Matryoshka sub-dimension losses
(MatryoshkaDims), asymmetric instruction prefixes (QueryPrefix / DocumentPrefix), and a choice of
which projections carry adapters (Targets: Attention, Mlp or All).
The SentenceTransformers.LoraTraining project is a ready-to-run console app that fine-tunes MiniLM,
Arctic XS or Harrier Small against one of two bundled example datasets (--dataset):
stsb — the English STS Benchmark, a broad
general-English similarity set, downloaded on demand.patent — the Google Patent Phrase Similarity
dataset (CC BY 4.0), embedded directly in the app (no download). Terse, domain-specific technical
phrases where a general encoder has real headroom.cd SentenceTransformers.LoraTraining
# General-English STS Benchmark (needs a one-time dataset download; model weights are embedded):
dotnet run -c Release -- download
dotnet run -c Release -- train --model minilm --objective cosent --rank 8 --epochs 10 --output-bias
dotnet run -c Release -- eval --model minilm --adapter ./adapters/minilm-stsb.lora --split test
# Domain-specific patent phrases (embedded, no download) — graded scores:
dotnet run -c Release -- train --model arctic --dataset patent --objective cosent --rank 8 --whitening
dotnet run -c Release -- eval --model arctic --dataset patent --adapter ./adapters/arctic-patent.lora --split test
# Harrier Small (Gemma3, pure C#) — bf16 weights downloaded on first use; heavier, so keep batch small:
dotnet run -c Release -- train --model harrier-small --dataset patent --objective cosent --rank 8 --batch 8 --max-tokens 64
--model accepts minilm, arctic or harrier-small. Run dotnet run -- help for the full option list (objective,
targets, rank, α, learning rate, warmup, temperature, mined negatives, learnable temperature, output bias,
whitening, Matryoshka, query/doc prefixes, multi-seed, …). Training prints per-epoch validation retrieval
accuracy and STS Spearman and a base-vs-tuned summary at the end.
Each model package contains:
SentenceEncoder that tokenizes, runs ONNX Runtime inference, pools the token outputs and
L2-normalizes the result.The shared SentenceTransformers package provides the ISentenceEncoder
contract and the default-implemented chunking helpers, so the model packages only implement
EncodeAsync and expose their MaxChunkLength / Tokenizer.
The SentenceTransformers.Harrier.Small.Pure package is the exception: instead of ONNX Runtime it
ships its own pure-managed implementation of the model. It reads the original
safetensors weights directly, runs the Gemma3
decoder forward pass (RMSNorm, grouped-query attention with Q/K-norm and RoPE, GeGLU MLP, last-token
pooling) on TensorPrimitives,
and tokenizes with a from-scratch Gemma byte-level BPE tokenizer — so it depends only on the .NET base
class library and System.Numerics.Tensors.
dotnet build SentenceTransformers.sln -c Release
dotnet test SentenceTransformers.sln
NuGet packages are produced and published by the Azure DevOps pipeline in
.devops/azure-pipelines.yml on pushes to main.
MIT. The BERT tokenizers are derived from
BERTTokenizers (MIT, © 2021 Othneil Drew). Each wrapped
model is distributed under its own upstream license — see the linked Hugging Face model pages. The
Google Patent Phrase Similarity
dataset bundled with the SentenceTransformers.LoraTraining example is © Google, licensed
CC BY 4.0.