curiosity-ai/sentence-transformers-sharp

C#

25

161 commits

updated Sep 18, 2026

See the code

README

sentence-transformers-sharp

Fast, dependency-light sentence embeddings for .NET. This library wraps a set of ONNX embedding models behind a single, simple ISentenceEncoder interface so you can turn text into vectors — for semantic search, clustering, retrieval-augmented generation (RAG), deduplication, recommendations and similarity scoring — entirely in-process, with no Python runtime and no external API calls.

It is built and maintained by Curiosity and powers the AI search / vector indexing features of Curiosity Workspace.

using SentenceTransformers.MiniLM;

using var encoder = new SentenceEncoder();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "The quick brown fox jumps over the lazy dog",
    "A fast auburn fox leaps above a sleepy hound",
});

// vectors[0] and vectors[1] are L2-normalized float[384] embeddings.

Why this library

  • No Python, no servers. Inference runs locally via the ONNX Runtime; the embedded models ship inside the NuGet package.
  • One interface, many models. Swap models by changing a single using — every encoder implements ISentenceEncoder.
  • Tokenizer-aware chunking built in. Long documents are split on token boundaries (never exceeding the model's context window) and encoded in one call, with optional overlap, progress reporting and offset alignment back into the source text.
  • Normalized vectors. All models return L2-normalized embeddings, so cosine similarity is just a dot product.

Models

PackageModelDimensionsMax tokensLanguagesWeights
NuGet DownloadsCore interfaces & chunking helpers————
NuGet Downloadsall-MiniLM-L6-v2384256EnglishEmbedded
NuGet Downloadssnowflake-arctic-embed-xs384512EnglishEmbedded
NuGet DownloadsQwen3-Embedding-0.6B102432768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-0.6b102432768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-270m64032768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-270m (pure C#, no ONNX)64032768MultilingualDownloaded on first use
  • Embedded models bundle the ONNX weights inside the NuGet package, so the encoder is ready immediately after construction.
  • Downloaded models are larger; their weights are fetched once on first use and cached on disk (under the system temp folder by default — see Choosing where weights are stored).

Pick MiniLM for the smallest/fastest footprint, Arctic XS for a strong English default, Harrier Small when you need multilingual coverage without paying for the larger 0.6b weights, and Qwen3 or Harrier Medium when you want the highest-quality embeddings (1024 dim) and the full 32k-token context window.

Installation

Install only the model package(s) you need (the core SentenceTransformers package is pulled in as a dependency):

dotnet add package SentenceTransformers.MiniLM
dotnet add package SentenceTransformers.ArcticXs
dotnet add package SentenceTransformers.Qwen3
dotnet add package SentenceTransformers.Harrier.Medium
dotnet add package SentenceTransformers.Harrier.Small

Targets .NET 10.

Usage

Embedded models (MiniLM, Arctic XS)

Embedded models are ready to use as soon as you construct them:

using SentenceTransformers.ArcticXs;

using var encoder = new SentenceEncoder();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "How do I reset my password?",
    "I forgot my login credentials.",
});

Downloaded models (Qwen3, Harrier Medium, Harrier Small)

Larger models download their ONNX weights on first use. Create them with the async CreateAsync factory — the download is cached, so subsequent runs are instant:

using SentenceTransformers.Qwen3;

// Downloads the model to a temp folder on first use, then loads it.
using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[] { "Hello world" });
// vectors[0] is a float[1024]

Harrier Medium is multilingual:

using SentenceTransformers.Harrier.Medium;

using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "Good morning",   // English
    "Buenos días",    // Spanish
    "おはよう",          // Japanese
});

Harrier Small is the same multilingual family at ~270M parameters (640-dim embeddings), suitable when you want multilingual coverage without paying for the 0.6b weights:

using SentenceTransformers.Harrier.Small;

using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "Good morning",
    "Buenos días",
    "おはよう",
});
// vectors[0] is a float[640]

Harrier Small, pure C# — no ONNX, no native dependencies

SentenceTransformers.Harrier.Small.Pure is a 100% managed reimplementation of Harrier Small. It runs the Gemma3 forward pass and the Gemma BPE tokenizer entirely in C# (on top of System.Numerics.Tensors), with no ONNX Runtime and no native tokenizer — so there is not a single .so/.dll/.dylib to ship. That makes it trim/AOT-friendly and portable to anywhere .NET runs, including Blazor WebAssembly and mobile. The API mirrors the ONNX package:

using SentenceTransformers.Harrier.Small.Pure;

// Downloads the original bfloat16 safetensors weights (~540 MB) on first use, then loads them.
using var encoder = await SentenceEncoder.CreateAsync();

// Queries take a task instruction prefix; documents are encoded as-is.
float[][] queryVectors = await encoder.EncodeQueriesAsync(
    new[] { "how much protein should a female eat" },
    SentenceEncoder.Prompts.WebSearchQuery);

float[][] docVectors = await encoder.EncodeAsync(new[] { "…a passage about dietary protein…" });
// vectors are L2-normalized float[640]

It produces the same embeddings as the reference: pure fp32 reproduces the query/document similarity matrix published on the model card to within 0.01 — actually closer to the reference than the shipped ONNX Q4F16 build.

Choosing a quantization. The transformer weights can be loaded at reduced precision to cut both memory and inference time. Pass a Quantization to CreateAsync (or the constructor):

using SentenceTransformers.Harrier.Small.Pure;
using SentenceTransformers.Harrier.Small.Pure.Model;

// fp32 (default, most faithful), Int8 (recommended — fastest & ~40% less memory), or Int4 (smallest).
using var encoder = await SentenceEncoder.CreateAsync(quantization: Quantization.Int8);

The Int8/Int4 paths run as true int8 GEMMs and pick the best instruction set available at runtime: vpdpbsud on 512-bit registers (AvxVnniInt8.V512, 64 int8 MACs/instruction), vpdpbusd/vpdpbsud on 256-bit (AvxVnni/AvxVnniInt8), or a widen + vpmaddwd sequence on AVX-512 / AVX2 CPUs. On an AVX-512 host, also set DOTNET_PreferredVectorBitWidth=512 to let the JIT emit 512-bit vectors.

Benchmark — pure C# vs ONNX (harrier-oss-v1-270m, .NET 10, 4-core Xeon, single-text encode):

VariantNative depsMax err vs model card¹Short text~512-token textResident weights²
ONNX Q4F16 (SentenceTransformers.Harrier.Small)ONNX Runtime2.30~40 ms~0.7 s~172 MB (native)
Pure fp32none0.01226 ms4.4 s~740 MB
Pure int8none0.9786 ms1.6 s~440 MB
Pure int4none1.42260 ms3.8 s~390 MB

¹ Largest absolute deviation (on a 0–100 cosine×100 scale) from the published query/document score matrix; lower is more faithful — every pure variant tracks the reference more closely than the ONNX Q4F16 weights. ² Approximate model-weight memory; all pure variants share the same bfloat16 token-embedding table (~335 MB), which is the floor. Numbers are hardware-dependent — the run above is an AVX-512 server CPU where the int8 dot uses widen + vpmaddwd; CPUs with the int8-VNNI instructions are substantially faster (see below).

How to read this. Int8 is the recommended pure setting — ~2.5× faster than fp32, ~40 % smaller, and still more faithful than ONNX Q4F16.

How close it gets to ONNX depends on the CPU's int8 instruction set. ONNX Runtime's MLAS uses hand-tuned assembly with 4-bit weights and int8-VNNI (vpdpbusd/vpdpbsud). The pure build emits int8-VNNI too when the runtime exposes it: AvxVnni (256-bit, Alder Lake and newer client CPUs) or AvxVnniInt8.V512 (512-bit, AVX10.2 / Granite Rapids-class CPUs) — on those it is within ~1.5–2× of ONNX and can approach parity with the 512-bit path. The one gap is "classic" AVX-512 servers that report the avx512_vnni CPUID flag but not AvxVnni/AvxVnniInt8/AVX10: .NET 10 has no standalone Avx512Vnni intrinsic, so there the pure build must fall back to widen + vpmaddwd (~6× the instructions) and lands ~2.5× off ONNX. Either way the pure build's case is zero native dependencies (trim/AOT/WASM/mobile, one managed package) and higher fidelity, at a CPU-inference cost within a small multiple of ONNX.

Comparing two texts (cosine similarity)

Because every model returns L2-normalized vectors, cosine similarity is simply the dot product:

static float CosineSimilarity(float[] a, float[] b)
{
    float dot = 0f;
    for (int i = 0; i < a.Length; i++)
    {
        dot += a[i] * b[i];
    }
    return dot; // vectors are unit-length, so dot product == cosine similarity
}

using var encoder = new SentenceTransformers.MiniLM.SentenceEncoder();
var v = await encoder.EncodeAsync(new[] { "cat", "kitten", "spaceship" });

Console.WriteLine(CosineSimilarity(v[0], v[1])); // cat vs kitten   -> high
Console.WriteLine(CosineSimilarity(v[0], v[2])); // cat vs spaceship -> low

Embedding long documents (chunking)

EncodeAsync expects each input to fit within the model's context window (encoder.MaxChunkLength tokens). For longer text, use the built-in chunking helpers, which split on token boundaries and embed each chunk:

using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync();

EncodedChunk[] chunks = await encoder.ChunkAndEncodeAsync(
    longDocument,
    chunkLength:  512,   // tokens per chunk (clamped to MaxChunkLength)
    chunkOverlap: 64,    // tokens of overlap between consecutive chunks
    reportProgress: p => Console.WriteLine($"{p:P0}"));

foreach (var chunk in chunks)
{
    // chunk.Text   -> the chunk's source text
    // chunk.Vector -> its embedding
    Index(chunk.Text, chunk.Vector);
}

Need to map results back to their position in the original text (e.g. to highlight a passage)? Use ChunkAndEncodeAlignedAsync, which additionally returns character offsets (Start, LastStart, ApproximateEnd) into the source. There are also tagged variants (ChunkAndEncodeTaggedAsync / …AlignedAsync) for carrying per-chunk metadata such as page numbers through the chunking pipeline.

Choosing where weights are stored

For the downloaded models you can control the cache location and the source URL:

using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync(
    downloadToPath: "/var/models/qwen3.onnx");

// Or point at your own mirror:
using var harrier = await SentenceTransformers.Harrier.Medium.SentenceEncoder.CreateAsync(
    modelUrl:     "https://my-mirror.example.com/harrier/model_quantized.onnx",
    modelDataUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx_data");

You can also pass a custom Microsoft.ML.OnnxRuntime.SessionOptions to any constructor / CreateAsync to tune threading or enable hardware execution providers.

Choosing a Harrier quantization

Both Harrier packages ship multiple quantization formats — pick the one that fits your CPU / GPU memory budget. URLs for every variant are exposed as constants on SentenceEncoder.Quantizations:

VariantConstantHarrier Medium (0.6b) weightsHarrier Small (270m) weights
Full (fp32)Quantizations.FullModelUrl2.09 GB (+306 MB)1.11 GB
FP16Quantizations.Fp16ModelUrl1.20 GB553 MB
Q4Quantizations.Q4ModelUrl399 MB205 MB
Q4 + FP16Quantizations.Q4Fp16ModelUrl (default)353 MB172 MB
QuantizedQuantizations.QuantizedModelUrl706 MB344 MB

Q4F16 is the default — it's the smallest variant on disk and keeps multilingual retrieval quality close to the unquantized reference. Pick Quantized when you need a pure float32 output and broader ONNX-Runtime execution-provider compatibility; FP16 for the most precision per byte; Full for the unquantized reference. Each entry has a matching …ModelDataUrl constant for the external weights file (and the Harrier Medium Full variant additionally has FullModelDataUrl2 because its fp32 weights are split into two files).

using SentenceTransformers.Harrier.Small;

// Use the unquantized reference instead of the default Q4F16:
using var encoder = await SentenceEncoder.CreateAsync(
    modelUrl:     SentenceEncoder.Quantizations.FullModelUrl,
    modelDataUrl: SentenceEncoder.Quantizations.FullModelDataUrl);

Fine-tuning for your use case (real weight-space LoRA)

📖 See LORA.md for the full guide: how it works internally, every option, negative-sample handling, stop conditions, and library + CLI training walkthroughs.

You can specialize the pure-C# encoders — MiniLM, Arctic XS and Harrier Small — for a specific domain (support tickets, legal clauses, product descriptions, a particular language pair) by training a real weight-space LoRA adapter from a set of related pairs (a query and a relevant passage, two paraphrases, a question and its duplicate…). This is proper LoRA: small low-rank factors are injected inside the transformer's attention/MLP projections (W + (α/r)·B·A), and the loss is backpropagated through the whole (frozen) network. The forward and backward passes run entirely in pure C# — no ONNX Runtime, and no PyTorch/autodiff dependency. A shared tensor autograd engine (SentenceTransformers.Training.Autograd) powers both model families.

MiniLM / Arctic XS (SentenceTransformers.Bert.Pure, the BertModel architecture) read their full-precision weights directly from the fp32 ONNX graphs already embedded in the MiniLM / Arctic packages, so training is fully self-contained — no model download.

Harrier Small (SentenceTransformers.Harrier.Small.Pure, a Gemma3 decoder) trains through the same autograd engine and LoRA objectives via Gemma3LoraTrainer / Gemma3LoraEncoder; its bf16 weights are downloaded on first use, and because it is a ~270M-param decoder each step is much heavier than the small BERTs (keep the batch / sequence length modest). LoRA there targets q/k/v/o and the gate/up/down MLP projections.

The remaining ONNX-only models (Qwen3, Harrier Medium) are inference-only and are not trainable here.

using SentenceTransformers.Bert.Pure;
using SentenceTransformers.Bert.Pure.Model;
using SentenceTransformers.Bert.Pure.Training;
using SentenceTransformers.Training;

// Pure-C# MiniLM (weights fetched once from HuggingFace, or use SentenceEncoder.LoadFromOnnx to reuse
// the weights embedded in the SentenceTransformers.MiniLM package with no download at all).
using var baseEncoder = await SentenceEncoder.CreateMiniLMAsync();

var dataset = new SentencePairDataset(new[]
{
    new SentencePair("how do I reset my password", "Use the ‘Forgot password’ link on the sign-in page.", 0.95f),
    new SentencePair("cancel my subscription",      "Go to Billing → Manage plan → Cancel.",              0.95f),
    // … a few hundred to a few thousand related (optionally scored) pairs …
});

var report = await BertLoraTrainer.TrainAsync(baseEncoder, dataset, new BertLoraTrainingOptions
{
    Objective = BertTrainingObjective.CoSent,     // pairwise ranking loss — directly targets STS Spearman
    Rank      = 8,                                // LoRA rank
    Targets   = LoraTargets.Attention,            // inject into Q, K, V and the attention output
    Epochs    = 10,
});

report.Adapter.Save("support-faq.lora");

// The tuned model is a drop-in ISentenceEncoder (the adapter is folded into the transformer at runtime):
using var tuned = baseEncoder.WithAdapter(report.Adapter);
float[][] vectors = await tuned.EncodeAsync(new[] { "I can't log in" });

How it works. A tiny tensor-level autograd engine records the BERT forward pass; the exact gradients for the injected LoRA factors are obtained by replaying the tape backward (verified against finite differences in the test-suite). Only the low-rank factors — and optionally a learned output-centering bias — are trainable; every base weight stays frozen, so the gradients for the frozen network are skipped (the LoRA efficiency win). The held-out validation set is scored with STS Spearman and top-1 retrieval accuracy and the best adapter is kept.

Objectives (BertLoraTrainingOptions.Objective):

  • Contrastive (default) — symmetric InfoNCE (MultipleNegativesRanking) with in-batch and hard negatives (explicit triplets via SentencePair.Negative, plus optional mined negatives), with false-negative masking. Best for retrieval / nearest-neighbour separation.
  • CoSent — the CoSENT pairwise-ranking loss over graded pairs. It optimizes the ordering that Spearman measures and consistently beats plain cosine-MSE on STS. Best when you care about graded similarity.
  • CosineRegression — mean-squared error between adapted cosine and the gold [0,1] score.

Extras (all optional): warmup + cosine LR schedule, a learned temperature (LearnableTemperature), multi-seed model selection (NumSeeds), a learned output-centering bias (UseOutputBias) and post-hoc ZCA whitening (ApplyWhitening) to counter embedding anisotropy, Matryoshka sub-dimension losses (MatryoshkaDims), asymmetric instruction prefixes (QueryPrefix / DocumentPrefix), and a choice of which projections carry adapters (Targets: Attention, Mlp or All).

Training CLI + example dataset

The SentenceTransformers.LoraTraining project is a ready-to-run console app that fine-tunes MiniLM, Arctic XS or Harrier Small against one of two bundled example datasets (--dataset):

  • stsb — the English STS Benchmark, a broad general-English similarity set, downloaded on demand.
  • patent — the Google Patent Phrase Similarity dataset (CC BY 4.0), embedded directly in the app (no download). Terse, domain-specific technical phrases where a general encoder has real headroom.
cd SentenceTransformers.LoraTraining

# General-English STS Benchmark (needs a one-time dataset download; model weights are embedded):
dotnet run -c Release -- download
dotnet run -c Release -- train --model minilm --objective cosent --rank 8 --epochs 10 --output-bias
dotnet run -c Release -- eval  --model minilm --adapter ./adapters/minilm-stsb.lora --split test

# Domain-specific patent phrases (embedded, no download) — graded scores:
dotnet run -c Release -- train --model arctic --dataset patent --objective cosent --rank 8 --whitening
dotnet run -c Release -- eval  --model arctic --dataset patent --adapter ./adapters/arctic-patent.lora --split test

# Harrier Small (Gemma3, pure C#) — bf16 weights downloaded on first use; heavier, so keep batch small:
dotnet run -c Release -- train --model harrier-small --dataset patent --objective cosent --rank 8 --batch 8 --max-tokens 64

--model accepts minilm, arctic or harrier-small. Run dotnet run -- help for the full option list (objective, targets, rank, α, learning rate, warmup, temperature, mined negatives, learnable temperature, output bias, whitening, Matryoshka, query/doc prefixes, multi-seed, …). Training prints per-epoch validation retrieval accuracy and STS Spearman and a base-vs-tuned summary at the end.

How it works

Each model package contains:

  • a tokenizer (WordPiece for the BERT-family embedded models, BPE via Hugging Face tokenizers for the Qwen3 / Harrier Medium / Harrier Small models),
  • the ONNX graph (embedded, or downloaded on first use), and
  • a thin SentenceEncoder that tokenizes, runs ONNX Runtime inference, pools the token outputs and L2-normalizes the result.

The shared SentenceTransformers package provides the ISentenceEncoder contract and the default-implemented chunking helpers, so the model packages only implement EncodeAsync and expose their MaxChunkLength / Tokenizer.

The SentenceTransformers.Harrier.Small.Pure package is the exception: instead of ONNX Runtime it ships its own pure-managed implementation of the model. It reads the original safetensors weights directly, runs the Gemma3 decoder forward pass (RMSNorm, grouped-query attention with Q/K-norm and RoPE, GeGLU MLP, last-token pooling) on TensorPrimitives, and tokenizes with a from-scratch Gemma byte-level BPE tokenizer — so it depends only on the .NET base class library and System.Numerics.Tensors.

Contributing & building

dotnet build SentenceTransformers.sln -c Release
dotnet test  SentenceTransformers.sln

NuGet packages are produced and published by the Azure DevOps pipeline in .devops/azure-pipelines.yml on pushes to main.

License

MIT. The BERT tokenizers are derived from BERTTokenizers (MIT, © 2021 Othneil Drew). Each wrapped model is distributed under its own upstream license — see the linked Hugging Face model pages. The Google Patent Phrase Similarity dataset bundled with the SentenceTransformers.LoraTraining example is © Google, licensed CC BY 4.0.

curiosity-ai/sentence-transformers-sharp

C#

25

161 commits

updated Sep 18, 2026

See the code

README

sentence-transformers-sharp

Fast, dependency-light sentence embeddings for .NET. This library wraps a set of ONNX embedding models behind a single, simple ISentenceEncoder interface so you can turn text into vectors — for semantic search, clustering, retrieval-augmented generation (RAG), deduplication, recommendations and similarity scoring — entirely in-process, with no Python runtime and no external API calls.

It is built and maintained by Curiosity and powers the AI search / vector indexing features of Curiosity Workspace.

using SentenceTransformers.MiniLM;

using var encoder = new SentenceEncoder();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "The quick brown fox jumps over the lazy dog",
    "A fast auburn fox leaps above a sleepy hound",
});

// vectors[0] and vectors[1] are L2-normalized float[384] embeddings.

Why this library

  • No Python, no servers. Inference runs locally via the ONNX Runtime; the embedded models ship inside the NuGet package.
  • One interface, many models. Swap models by changing a single using — every encoder implements ISentenceEncoder.
  • Tokenizer-aware chunking built in. Long documents are split on token boundaries (never exceeding the model's context window) and encoded in one call, with optional overlap, progress reporting and offset alignment back into the source text.
  • Normalized vectors. All models return L2-normalized embeddings, so cosine similarity is just a dot product.

Models

PackageModelDimensionsMax tokensLanguagesWeights
NuGet DownloadsCore interfaces & chunking helpers————
NuGet Downloadsall-MiniLM-L6-v2384256EnglishEmbedded
NuGet Downloadssnowflake-arctic-embed-xs384512EnglishEmbedded
NuGet DownloadsQwen3-Embedding-0.6B102432768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-0.6b102432768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-270m64032768MultilingualDownloaded on first use
NuGet Downloadsharrier-oss-v1-270m (pure C#, no ONNX)64032768MultilingualDownloaded on first use
  • Embedded models bundle the ONNX weights inside the NuGet package, so the encoder is ready immediately after construction.
  • Downloaded models are larger; their weights are fetched once on first use and cached on disk (under the system temp folder by default — see Choosing where weights are stored).

Pick MiniLM for the smallest/fastest footprint, Arctic XS for a strong English default, Harrier Small when you need multilingual coverage without paying for the larger 0.6b weights, and Qwen3 or Harrier Medium when you want the highest-quality embeddings (1024 dim) and the full 32k-token context window.

Installation

Install only the model package(s) you need (the core SentenceTransformers package is pulled in as a dependency):

dotnet add package SentenceTransformers.MiniLM
dotnet add package SentenceTransformers.ArcticXs
dotnet add package SentenceTransformers.Qwen3
dotnet add package SentenceTransformers.Harrier.Medium
dotnet add package SentenceTransformers.Harrier.Small

Targets .NET 10.

Usage

Embedded models (MiniLM, Arctic XS)

Embedded models are ready to use as soon as you construct them:

using SentenceTransformers.ArcticXs;

using var encoder = new SentenceEncoder();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "How do I reset my password?",
    "I forgot my login credentials.",
});

Downloaded models (Qwen3, Harrier Medium, Harrier Small)

Larger models download their ONNX weights on first use. Create them with the async CreateAsync factory — the download is cached, so subsequent runs are instant:

using SentenceTransformers.Qwen3;

// Downloads the model to a temp folder on first use, then loads it.
using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[] { "Hello world" });
// vectors[0] is a float[1024]

Harrier Medium is multilingual:

using SentenceTransformers.Harrier.Medium;

using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "Good morning",   // English
    "Buenos días",    // Spanish
    "おはよう",          // Japanese
});

Harrier Small is the same multilingual family at ~270M parameters (640-dim embeddings), suitable when you want multilingual coverage without paying for the 0.6b weights:

using SentenceTransformers.Harrier.Small;

using var encoder = await SentenceEncoder.CreateAsync();

float[][] vectors = await encoder.EncodeAsync(new[]
{
    "Good morning",
    "Buenos días",
    "おはよう",
});
// vectors[0] is a float[640]

Harrier Small, pure C# — no ONNX, no native dependencies

SentenceTransformers.Harrier.Small.Pure is a 100% managed reimplementation of Harrier Small. It runs the Gemma3 forward pass and the Gemma BPE tokenizer entirely in C# (on top of System.Numerics.Tensors), with no ONNX Runtime and no native tokenizer — so there is not a single .so/.dll/.dylib to ship. That makes it trim/AOT-friendly and portable to anywhere .NET runs, including Blazor WebAssembly and mobile. The API mirrors the ONNX package:

using SentenceTransformers.Harrier.Small.Pure;

// Downloads the original bfloat16 safetensors weights (~540 MB) on first use, then loads them.
using var encoder = await SentenceEncoder.CreateAsync();

// Queries take a task instruction prefix; documents are encoded as-is.
float[][] queryVectors = await encoder.EncodeQueriesAsync(
    new[] { "how much protein should a female eat" },
    SentenceEncoder.Prompts.WebSearchQuery);

float[][] docVectors = await encoder.EncodeAsync(new[] { "…a passage about dietary protein…" });
// vectors are L2-normalized float[640]

It produces the same embeddings as the reference: pure fp32 reproduces the query/document similarity matrix published on the model card to within 0.01 — actually closer to the reference than the shipped ONNX Q4F16 build.

Choosing a quantization. The transformer weights can be loaded at reduced precision to cut both memory and inference time. Pass a Quantization to CreateAsync (or the constructor):

using SentenceTransformers.Harrier.Small.Pure;
using SentenceTransformers.Harrier.Small.Pure.Model;

// fp32 (default, most faithful), Int8 (recommended — fastest & ~40% less memory), or Int4 (smallest).
using var encoder = await SentenceEncoder.CreateAsync(quantization: Quantization.Int8);

The Int8/Int4 paths run as true int8 GEMMs and pick the best instruction set available at runtime: vpdpbsud on 512-bit registers (AvxVnniInt8.V512, 64 int8 MACs/instruction), vpdpbusd/vpdpbsud on 256-bit (AvxVnni/AvxVnniInt8), or a widen + vpmaddwd sequence on AVX-512 / AVX2 CPUs. On an AVX-512 host, also set DOTNET_PreferredVectorBitWidth=512 to let the JIT emit 512-bit vectors.

Benchmark — pure C# vs ONNX (harrier-oss-v1-270m, .NET 10, 4-core Xeon, single-text encode):

VariantNative depsMax err vs model card¹Short text~512-token textResident weights²
ONNX Q4F16 (SentenceTransformers.Harrier.Small)ONNX Runtime2.30~40 ms~0.7 s~172 MB (native)
Pure fp32none0.01226 ms4.4 s~740 MB
Pure int8none0.9786 ms1.6 s~440 MB
Pure int4none1.42260 ms3.8 s~390 MB

¹ Largest absolute deviation (on a 0–100 cosine×100 scale) from the published query/document score matrix; lower is more faithful — every pure variant tracks the reference more closely than the ONNX Q4F16 weights. ² Approximate model-weight memory; all pure variants share the same bfloat16 token-embedding table (~335 MB), which is the floor. Numbers are hardware-dependent — the run above is an AVX-512 server CPU where the int8 dot uses widen + vpmaddwd; CPUs with the int8-VNNI instructions are substantially faster (see below).

How to read this. Int8 is the recommended pure setting — ~2.5× faster than fp32, ~40 % smaller, and still more faithful than ONNX Q4F16.

How close it gets to ONNX depends on the CPU's int8 instruction set. ONNX Runtime's MLAS uses hand-tuned assembly with 4-bit weights and int8-VNNI (vpdpbusd/vpdpbsud). The pure build emits int8-VNNI too when the runtime exposes it: AvxVnni (256-bit, Alder Lake and newer client CPUs) or AvxVnniInt8.V512 (512-bit, AVX10.2 / Granite Rapids-class CPUs) — on those it is within ~1.5–2× of ONNX and can approach parity with the 512-bit path. The one gap is "classic" AVX-512 servers that report the avx512_vnni CPUID flag but not AvxVnni/AvxVnniInt8/AVX10: .NET 10 has no standalone Avx512Vnni intrinsic, so there the pure build must fall back to widen + vpmaddwd (~6× the instructions) and lands ~2.5× off ONNX. Either way the pure build's case is zero native dependencies (trim/AOT/WASM/mobile, one managed package) and higher fidelity, at a CPU-inference cost within a small multiple of ONNX.

Comparing two texts (cosine similarity)

Because every model returns L2-normalized vectors, cosine similarity is simply the dot product:

static float CosineSimilarity(float[] a, float[] b)
{
    float dot = 0f;
    for (int i = 0; i < a.Length; i++)
    {
        dot += a[i] * b[i];
    }
    return dot; // vectors are unit-length, so dot product == cosine similarity
}

using var encoder = new SentenceTransformers.MiniLM.SentenceEncoder();
var v = await encoder.EncodeAsync(new[] { "cat", "kitten", "spaceship" });

Console.WriteLine(CosineSimilarity(v[0], v[1])); // cat vs kitten   -> high
Console.WriteLine(CosineSimilarity(v[0], v[2])); // cat vs spaceship -> low

Embedding long documents (chunking)

EncodeAsync expects each input to fit within the model's context window (encoder.MaxChunkLength tokens). For longer text, use the built-in chunking helpers, which split on token boundaries and embed each chunk:

using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync();

EncodedChunk[] chunks = await encoder.ChunkAndEncodeAsync(
    longDocument,
    chunkLength:  512,   // tokens per chunk (clamped to MaxChunkLength)
    chunkOverlap: 64,    // tokens of overlap between consecutive chunks
    reportProgress: p => Console.WriteLine($"{p:P0}"));

foreach (var chunk in chunks)
{
    // chunk.Text   -> the chunk's source text
    // chunk.Vector -> its embedding
    Index(chunk.Text, chunk.Vector);
}

Need to map results back to their position in the original text (e.g. to highlight a passage)? Use ChunkAndEncodeAlignedAsync, which additionally returns character offsets (Start, LastStart, ApproximateEnd) into the source. There are also tagged variants (ChunkAndEncodeTaggedAsync / …AlignedAsync) for carrying per-chunk metadata such as page numbers through the chunking pipeline.

Choosing where weights are stored

For the downloaded models you can control the cache location and the source URL:

using var encoder = await SentenceTransformers.Qwen3.SentenceEncoder.CreateAsync(
    downloadToPath: "/var/models/qwen3.onnx");

// Or point at your own mirror:
using var harrier = await SentenceTransformers.Harrier.Medium.SentenceEncoder.CreateAsync(
    modelUrl:     "https://my-mirror.example.com/harrier/model_quantized.onnx",
    modelDataUrl: "https://my-mirror.example.com/harrier/model_quantized.onnx_data");

You can also pass a custom Microsoft.ML.OnnxRuntime.SessionOptions to any constructor / CreateAsync to tune threading or enable hardware execution providers.

Choosing a Harrier quantization

Both Harrier packages ship multiple quantization formats — pick the one that fits your CPU / GPU memory budget. URLs for every variant are exposed as constants on SentenceEncoder.Quantizations:

VariantConstantHarrier Medium (0.6b) weightsHarrier Small (270m) weights
Full (fp32)Quantizations.FullModelUrl2.09 GB (+306 MB)1.11 GB
FP16Quantizations.Fp16ModelUrl1.20 GB553 MB
Q4Quantizations.Q4ModelUrl399 MB205 MB
Q4 + FP16Quantizations.Q4Fp16ModelUrl (default)353 MB172 MB
QuantizedQuantizations.QuantizedModelUrl706 MB344 MB

Q4F16 is the default — it's the smallest variant on disk and keeps multilingual retrieval quality close to the unquantized reference. Pick Quantized when you need a pure float32 output and broader ONNX-Runtime execution-provider compatibility; FP16 for the most precision per byte; Full for the unquantized reference. Each entry has a matching …ModelDataUrl constant for the external weights file (and the Harrier Medium Full variant additionally has FullModelDataUrl2 because its fp32 weights are split into two files).

using SentenceTransformers.Harrier.Small;

// Use the unquantized reference instead of the default Q4F16:
using var encoder = await SentenceEncoder.CreateAsync(
    modelUrl:     SentenceEncoder.Quantizations.FullModelUrl,
    modelDataUrl: SentenceEncoder.Quantizations.FullModelDataUrl);

Fine-tuning for your use case (real weight-space LoRA)

📖 See LORA.md for the full guide: how it works internally, every option, negative-sample handling, stop conditions, and library + CLI training walkthroughs.

You can specialize the pure-C# encoders — MiniLM, Arctic XS and Harrier Small — for a specific domain (support tickets, legal clauses, product descriptions, a particular language pair) by training a real weight-space LoRA adapter from a set of related pairs (a query and a relevant passage, two paraphrases, a question and its duplicate…). This is proper LoRA: small low-rank factors are injected inside the transformer's attention/MLP projections (W + (α/r)·B·A), and the loss is backpropagated through the whole (frozen) network. The forward and backward passes run entirely in pure C# — no ONNX Runtime, and no PyTorch/autodiff dependency. A shared tensor autograd engine (SentenceTransformers.Training.Autograd) powers both model families.

MiniLM / Arctic XS (SentenceTransformers.Bert.Pure, the BertModel architecture) read their full-precision weights directly from the fp32 ONNX graphs already embedded in the MiniLM / Arctic packages, so training is fully self-contained — no model download.

Harrier Small (SentenceTransformers.Harrier.Small.Pure, a Gemma3 decoder) trains through the same autograd engine and LoRA objectives via Gemma3LoraTrainer / Gemma3LoraEncoder; its bf16 weights are downloaded on first use, and because it is a ~270M-param decoder each step is much heavier than the small BERTs (keep the batch / sequence length modest). LoRA there targets q/k/v/o and the gate/up/down MLP projections.

The remaining ONNX-only models (Qwen3, Harrier Medium) are inference-only and are not trainable here.

using SentenceTransformers.Bert.Pure;
using SentenceTransformers.Bert.Pure.Model;
using SentenceTransformers.Bert.Pure.Training;
using SentenceTransformers.Training;

// Pure-C# MiniLM (weights fetched once from HuggingFace, or use SentenceEncoder.LoadFromOnnx to reuse
// the weights embedded in the SentenceTransformers.MiniLM package with no download at all).
using var baseEncoder = await SentenceEncoder.CreateMiniLMAsync();

var dataset = new SentencePairDataset(new[]
{
    new SentencePair("how do I reset my password", "Use the ‘Forgot password’ link on the sign-in page.", 0.95f),
    new SentencePair("cancel my subscription",      "Go to Billing → Manage plan → Cancel.",              0.95f),
    // … a few hundred to a few thousand related (optionally scored) pairs …
});

var report = await BertLoraTrainer.TrainAsync(baseEncoder, dataset, new BertLoraTrainingOptions
{
    Objective = BertTrainingObjective.CoSent,     // pairwise ranking loss — directly targets STS Spearman
    Rank      = 8,                                // LoRA rank
    Targets   = LoraTargets.Attention,            // inject into Q, K, V and the attention output
    Epochs    = 10,
});

report.Adapter.Save("support-faq.lora");

// The tuned model is a drop-in ISentenceEncoder (the adapter is folded into the transformer at runtime):
using var tuned = baseEncoder.WithAdapter(report.Adapter);
float[][] vectors = await tuned.EncodeAsync(new[] { "I can't log in" });

How it works. A tiny tensor-level autograd engine records the BERT forward pass; the exact gradients for the injected LoRA factors are obtained by replaying the tape backward (verified against finite differences in the test-suite). Only the low-rank factors — and optionally a learned output-centering bias — are trainable; every base weight stays frozen, so the gradients for the frozen network are skipped (the LoRA efficiency win). The held-out validation set is scored with STS Spearman and top-1 retrieval accuracy and the best adapter is kept.

Objectives (BertLoraTrainingOptions.Objective):

  • Contrastive (default) — symmetric InfoNCE (MultipleNegativesRanking) with in-batch and hard negatives (explicit triplets via SentencePair.Negative, plus optional mined negatives), with false-negative masking. Best for retrieval / nearest-neighbour separation.
  • CoSent — the CoSENT pairwise-ranking loss over graded pairs. It optimizes the ordering that Spearman measures and consistently beats plain cosine-MSE on STS. Best when you care about graded similarity.
  • CosineRegression — mean-squared error between adapted cosine and the gold [0,1] score.

Extras (all optional): warmup + cosine LR schedule, a learned temperature (LearnableTemperature), multi-seed model selection (NumSeeds), a learned output-centering bias (UseOutputBias) and post-hoc ZCA whitening (ApplyWhitening) to counter embedding anisotropy, Matryoshka sub-dimension losses (MatryoshkaDims), asymmetric instruction prefixes (QueryPrefix / DocumentPrefix), and a choice of which projections carry adapters (Targets: Attention, Mlp or All).

Training CLI + example dataset

The SentenceTransformers.LoraTraining project is a ready-to-run console app that fine-tunes MiniLM, Arctic XS or Harrier Small against one of two bundled example datasets (--dataset):

  • stsb — the English STS Benchmark, a broad general-English similarity set, downloaded on demand.
  • patent — the Google Patent Phrase Similarity dataset (CC BY 4.0), embedded directly in the app (no download). Terse, domain-specific technical phrases where a general encoder has real headroom.
cd SentenceTransformers.LoraTraining

# General-English STS Benchmark (needs a one-time dataset download; model weights are embedded):
dotnet run -c Release -- download
dotnet run -c Release -- train --model minilm --objective cosent --rank 8 --epochs 10 --output-bias
dotnet run -c Release -- eval  --model minilm --adapter ./adapters/minilm-stsb.lora --split test

# Domain-specific patent phrases (embedded, no download) — graded scores:
dotnet run -c Release -- train --model arctic --dataset patent --objective cosent --rank 8 --whitening
dotnet run -c Release -- eval  --model arctic --dataset patent --adapter ./adapters/arctic-patent.lora --split test

# Harrier Small (Gemma3, pure C#) — bf16 weights downloaded on first use; heavier, so keep batch small:
dotnet run -c Release -- train --model harrier-small --dataset patent --objective cosent --rank 8 --batch 8 --max-tokens 64

--model accepts minilm, arctic or harrier-small. Run dotnet run -- help for the full option list (objective, targets, rank, α, learning rate, warmup, temperature, mined negatives, learnable temperature, output bias, whitening, Matryoshka, query/doc prefixes, multi-seed, …). Training prints per-epoch validation retrieval accuracy and STS Spearman and a base-vs-tuned summary at the end.

How it works

Each model package contains:

  • a tokenizer (WordPiece for the BERT-family embedded models, BPE via Hugging Face tokenizers for the Qwen3 / Harrier Medium / Harrier Small models),
  • the ONNX graph (embedded, or downloaded on first use), and
  • a thin SentenceEncoder that tokenizes, runs ONNX Runtime inference, pools the token outputs and L2-normalizes the result.

The shared SentenceTransformers package provides the ISentenceEncoder contract and the default-implemented chunking helpers, so the model packages only implement EncodeAsync and expose their MaxChunkLength / Tokenizer.

The SentenceTransformers.Harrier.Small.Pure package is the exception: instead of ONNX Runtime it ships its own pure-managed implementation of the model. It reads the original safetensors weights directly, runs the Gemma3 decoder forward pass (RMSNorm, grouped-query attention with Q/K-norm and RoPE, GeGLU MLP, last-token pooling) on TensorPrimitives, and tokenizes with a from-scratch Gemma byte-level BPE tokenizer — so it depends only on the .NET base class library and System.Numerics.Tensors.

Contributing & building

dotnet build SentenceTransformers.sln -c Release
dotnet test  SentenceTransformers.sln

NuGet packages are produced and published by the Azure DevOps pipeline in .devops/azure-pipelines.yml on pushes to main.

License

MIT. The BERT tokenizers are derived from BERTTokenizers (MIT, © 2021 Othneil Drew). Each wrapped model is distributed under its own upstream license — see the linked Hugging Face model pages. The Google Patent Phrase Similarity dataset bundled with the SentenceTransformers.LoraTraining example is © Google, licensed CC BY 4.0.