elbruno/elbruno.localembeddings

.NET library for local embedding generation using ONNX and Microsoft.Extensions.AI

C#

40

248 commits

updated Sep 28, 2026

See the code

README

LocalEmbeddings

NuGet NuGet Downloads Build Status License: MIT GitHub stars Twitter Follow

A .NET library for generating text embeddings locally using ONNX Runtime and Microsoft.Extensions.AI abstractions β€” no external API calls required.

πŸŽ₯ Quick Overview

New to local embeddings? Watch this 5-minute video explaining the main goal of the library.

Want to build RAG applications? Read this blog post about 3 RAG approaches in .NET with local embeddings and zero cloud calls.

Interested in image embeddings? Check out the YouTube video and blog post about local image embeddings with CLIP and ONNX.

Features

  • Local Embedding Generation β€” Run inference entirely on your machine using ONNX Runtime
  • Multi-Framework β€” Targets .NET 8.0 (LTS) and .NET 10.0
  • Microsoft.Extensions.AI Integration β€” Implements IEmbeddingGenerator<string, Embedding<float>>
  • πŸ¦… Harrier model support β€” Microsoft Harrier-OSS-v1 (270M, 640-dim, 94+ languages, instruction-tuned) via ElBruno.LocalEmbeddings.Harrier
  • Kernel Memory Integration β€” Companion package ElBruno.LocalEmbeddings.KernelMemory provides a native ITextEmbeddingGenerator adapter for Microsoft Kernel Memory
  • VectorData Integration β€” Companion package ElBruno.LocalEmbeddings.VectorData adds DI helpers for Microsoft.Extensions.VectorData (VectorStore and typed collections)
  • Blazor Components β€” Companion package ElBruno.LocalEmbeddings.BlazorComponents provides 9 ready-to-use Razor components: model gallery, similarity meter, semantic search box, embedding explorer, dimension viewer, health badge, and metrics panel
  • Built-in In-Memory Vector Store β€” ElBruno.LocalEmbeddings.VectorData includes InMemoryVectorStore (no Semantic Kernel connector dependency required)
  • HuggingFace Model Support β€” Use popular sentence transformer models from HuggingFace Hub
  • Automatic Model Caching β€” Models are downloaded once and cached locally
  • Dependency Injection Support β€” First-class IServiceCollection integration
  • Single-String Convenience API β€” GenerateAsync("text") and GenerateEmbeddingAsync("text") β€” no array wrapping needed
  • Similarity Helpers β€” Cosine similarity, all-pairs Similarity(...) matrix, and one-line FindClosestAsync(...) semantic search
  • Thread-Safe & Batched β€” Concurrent generation and efficient multi-text processing

πŸ†• What's New

The last 5 notable additions to the library. Updated with each NuGet release.

VersionDateHighlight
v1.6.32026-09-28Uncased BERT/MiniLM tokenization β€” normalizes accents so accented words map to their expected tokens
v1.5.92026-08-02ElBruno.LocalEmbeddings.BlazorComponents β€” full test coverage (77 xUnit + bUnit tests for all 9 components), EmbeddingDimensionViewer PCA demo added to BlazorDemo, sample README, fixed auto-release β†’ NuGet publish pipeline
v1.5.72026-08-02New package: ElBruno.LocalEmbeddings.BlazorComponents β€” 9 ready-to-use Razor components: EmbeddingModelGallery, SimilarityMeter, SemanticSearchBox, EmbeddingExplorer, EmbeddingDimensionViewer (PCA 2-D), EmbeddingModelSelector, EmbeddingHealthBadge, EmbeddingMetricsPanel, EmbeddingModelStatusCard
v1.5.32026-08-02ElBruno.LocalEmbeddings.OpenTelemetry β€” OpenTelemetry instrumentation with Activity tracing, metrics (tokens/sec, latency), and configurable exporters for embedding pipelines
v1.4.x2026NPU hardware acceleration β€” ElBruno.LocalEmbeddings.Npu.Intel (Intel Core Ultra via OpenVINO) and ElBruno.LocalEmbeddings.Npu.Qualcomm (Snapdragon X via QNN) for on-device NPU inference

πŸ“¦ NuGet Packages

PackageVersionDownloadsDescription
ElBruno.LocalEmbeddingsNuGetDownloadsCore library β€” ONNX Runtime + Microsoft.Extensions.AI
ElBruno.LocalEmbeddings.HarrierNuGetDownloadsMicrosoft Harrier-OSS-v1 (270M, 640-dim, 94+ languages)
ElBruno.LocalEmbeddings.ImageEmbeddingsNuGetDownloadsCLIP-based image embeddings for multimodal search
ElBruno.LocalEmbeddings.ImageEmbeddings.DownloaderNuGetDownloadsModel downloader for image embeddings
ElBruno.LocalEmbeddings.KernelMemoryNuGetDownloadsMicrosoft Kernel Memory adapter
ElBruno.LocalEmbeddings.VectorDataNuGetDownloadsMicrosoft.Extensions.VectorData + InMemoryVectorStore
ElBruno.LocalEmbeddings.NpuNuGetDownloadsNPU-accelerated embeddings via DirectML
ElBruno.LocalEmbeddings.Npu.IntelNuGetDownloadsIntel Core Ultra NPU via OpenVINO
ElBruno.LocalEmbeddings.Npu.QualcommNuGetDownloadsQualcomm Snapdragon X NPU via QNN
ElBruno.LocalEmbeddings.BlazorComponentsNuGetDownloadsBlazor UI components β€” model gallery, similarity meter, semantic search, embedding explorer

Installation

dotnet add package ElBruno.LocalEmbeddings

For Harrier model support, install the companion package:

dotnet add package ElBruno.LocalEmbeddings.Harrier

For Kernel Memory integration, also install:

dotnet add package ElBruno.LocalEmbeddings.KernelMemory

For VectorData integration, install:

dotnet add package ElBruno.LocalEmbeddings.VectorData

Quick Start

1) Generate one embedding

using ElBruno.LocalEmbeddings;

await using var generator = await LocalEmbeddingGenerator.CreateAsync();
var embedding = await generator.GenerateEmbeddingAsync("Hello, world!");
Console.WriteLine(embedding.Vector.Length); // 384

2) Generate embeddings for multiple texts

var inputs = new[] { "first text", "second text", "third text" };
var embeddings = await generator.GenerateAsync(inputs);
Console.WriteLine(embeddings.Count); // 3

3) Compare two texts with cosine similarity

using ElBruno.LocalEmbeddings.Extensions;

var pair = await generator.GenerateAsync(["I love coding", "I enjoy programming"]);
var score = pair[0].CosineSimilarity(pair[1]);
Console.WriteLine(score);

4) Semantic search in one line

var corpus = new[]
{
    "Python for data science",
    "JavaScript for web apps",
    "Swift for iOS development"
};

var corpusEmbeddings = await generator.GenerateAsync(corpus);
var results = await generator.FindClosestAsync(
    "best language for websites",
    corpus,
    corpusEmbeddings,
    topK: 2,
    minScore: 0.2f);

foreach (var result in results)
    Console.WriteLine($"{result.Score:F3} - {result.Text}");

5) Using Harrier model (higher quality, multilingual)

using ElBruno.LocalEmbeddings.Harrier;

var generator = await HarrierEmbeddingGenerator.CreateAsync();
var embedding = await generator.GenerateEmbeddingAsync("Hello, world!");
Console.WriteLine($"Dimensions: {embedding.Vector.Length}"); // 640

For custom models and runtime behavior, use the options-based constructor: new LocalEmbeddingGenerator(new LocalEmbeddingsOptions { ... }).

Note: The synchronous constructor remains available for backward compatibility, but performs blocking initialization when downloads are needed.

Want to go further? Read the Getting Started guide and the other docs in this repo for DI, configuration, VectorData, Kernel Memory, and full RAG examples.

Prefer a containerized dev environment? See the Dev Container section in the Contributing guide.

Samples

All sample applications are located in src/Samples/. See the samples README for prerequisites and run instructions.

SampleWhat It Shows
HelloWorldAltModelMinimal hello world with sentence-transformers/all-MiniLM-L12-v2
ConsoleAppAll the basics: single/batch embeddings, similarity, semantic search, DI
HarrierConsoleAppHarrier embedding model usage and similarity search
RagChatEmbedding-only semantic search Q&A using shared VectorData InMemoryVectorStore (no LLM needed)
RagOllamaFull RAG with Ollama + phi4-mini + Kernel Memory
RagFoundryLocalFull RAG with Foundry Local + phi4-mini
ImageRagSimpleMinimal image RAG: index images β†’ search by text
ImageRagChatInteractive image RAG chat with text and image-to-image search

Configuration

var options = new LocalEmbeddingsOptions
{
    ModelName = "sentence-transformers/all-MiniLM-L6-v2",  // HuggingFace model
    MaxSequenceLength = 512,                                // Max tokens
    CacheDirectory = null,                                  // Auto-detect per platform
    EnsureModelDownloaded = true,                           // Download if missing
    NormalizeEmbeddings = false                              // L2 normalize vectors
};

See Configuration docs for supported models, local model paths, and cache locations.

Common model options (with model cards)

Estimated download sizes below are approximate and can vary by ONNX variant (fp32/int8) and tokenizer assets.

Documentation

TopicDescription
Getting StartedStep-by-step guide from hello world to RAG
API ReferenceClasses, methods, and extension methods
ConfigurationOptions, supported models, cache locations
Alternative ModelsNon-default free models, local download workflow, and license notes
Dependency InjectionAll DI overloads and IConfiguration binding
Harrier IntegrationMicrosoft Harrier-OSS-v1 local embedding model
Kernel Memory IntegrationUsing local embeddings with Microsoft Kernel Memory
VectorData IntegrationUsing local embeddings with Microsoft.Extensions.VectorData abstractions
Blazor Components9 ready-to-use Blazor components for embedding-powered web apps
ContributingBuild from source, repo structure, guidelines
RoadmapPlanned and completed features/samples with priorities
PublishingNuGet publishing with GitHub Actions + Trusted Publishing
ChangelogVersioned summary of notable changes

Have an idea for a new feature or sample? Please open an issue and share your suggestion.

Building from Source

git clone https://github.com/elbruno/elbruno.localembeddings.git
cd elbruno.localembeddings
dotnet build
dotnet test

Requirements

  • .NET 8.0 SDK (LTS) or .NET 10.0 SDK
  • ONNX Runtime compatible platform (Windows, Linux, macOS)

πŸ‘‹ About the Author

Hi! I'm ElBruno 🧑, a passionate developer and content creator exploring AI, .NET, and modern development practices.

Made with ❀️ by ElBruno

If you like this project, consider following my work across platforms:

  • πŸ“» Podcast: No Tienen Nombre β€” Spanish-language episodes on AI, development, and tech culture
  • πŸ’» Blog: ElBruno.com β€” Deep dives on embeddings, RAG, .NET, and local AI
  • πŸ“Ί YouTube: youtube.com/elbruno β€” Demos, tutorials, and live coding
  • πŸ”— LinkedIn: @elbruno β€” Professional updates and insights
  • 𝕏 Twitter: @elbruno β€” Quick tips, releases, and tech news

License

This project is licensed under the MIT License β€” see the LICENSE file for details.

cosine-similarity
csharp
dependency-injection
dotnet
embeddings
huggingface
kernel-memory
local-ai
machine-learning
microsoft-extensions-ai
nuget
onnx
onnx-runtime
rag
semantic-search
text-embeddings
vector-embeddings

elbruno/elbruno.localembeddings

.NET library for local embedding generation using ONNX and Microsoft.Extensions.AI

C#

40

248 commits

updated Sep 28, 2026

See the code

README

LocalEmbeddings

NuGet NuGet Downloads Build Status License: MIT GitHub stars Twitter Follow

A .NET library for generating text embeddings locally using ONNX Runtime and Microsoft.Extensions.AI abstractions β€” no external API calls required.

πŸŽ₯ Quick Overview

New to local embeddings? Watch this 5-minute video explaining the main goal of the library.

Want to build RAG applications? Read this blog post about 3 RAG approaches in .NET with local embeddings and zero cloud calls.

Interested in image embeddings? Check out the YouTube video and blog post about local image embeddings with CLIP and ONNX.

Features

  • Local Embedding Generation β€” Run inference entirely on your machine using ONNX Runtime
  • Multi-Framework β€” Targets .NET 8.0 (LTS) and .NET 10.0
  • Microsoft.Extensions.AI Integration β€” Implements IEmbeddingGenerator<string, Embedding<float>>
  • πŸ¦… Harrier model support β€” Microsoft Harrier-OSS-v1 (270M, 640-dim, 94+ languages, instruction-tuned) via ElBruno.LocalEmbeddings.Harrier
  • Kernel Memory Integration β€” Companion package ElBruno.LocalEmbeddings.KernelMemory provides a native ITextEmbeddingGenerator adapter for Microsoft Kernel Memory
  • VectorData Integration β€” Companion package ElBruno.LocalEmbeddings.VectorData adds DI helpers for Microsoft.Extensions.VectorData (VectorStore and typed collections)
  • Blazor Components β€” Companion package ElBruno.LocalEmbeddings.BlazorComponents provides 9 ready-to-use Razor components: model gallery, similarity meter, semantic search box, embedding explorer, dimension viewer, health badge, and metrics panel
  • Built-in In-Memory Vector Store β€” ElBruno.LocalEmbeddings.VectorData includes InMemoryVectorStore (no Semantic Kernel connector dependency required)
  • HuggingFace Model Support β€” Use popular sentence transformer models from HuggingFace Hub
  • Automatic Model Caching β€” Models are downloaded once and cached locally
  • Dependency Injection Support β€” First-class IServiceCollection integration
  • Single-String Convenience API β€” GenerateAsync("text") and GenerateEmbeddingAsync("text") β€” no array wrapping needed
  • Similarity Helpers β€” Cosine similarity, all-pairs Similarity(...) matrix, and one-line FindClosestAsync(...) semantic search
  • Thread-Safe & Batched β€” Concurrent generation and efficient multi-text processing

πŸ†• What's New

The last 5 notable additions to the library. Updated with each NuGet release.

VersionDateHighlight
v1.6.32026-09-28Uncased BERT/MiniLM tokenization β€” normalizes accents so accented words map to their expected tokens
v1.5.92026-08-02ElBruno.LocalEmbeddings.BlazorComponents β€” full test coverage (77 xUnit + bUnit tests for all 9 components), EmbeddingDimensionViewer PCA demo added to BlazorDemo, sample README, fixed auto-release β†’ NuGet publish pipeline
v1.5.72026-08-02New package: ElBruno.LocalEmbeddings.BlazorComponents β€” 9 ready-to-use Razor components: EmbeddingModelGallery, SimilarityMeter, SemanticSearchBox, EmbeddingExplorer, EmbeddingDimensionViewer (PCA 2-D), EmbeddingModelSelector, EmbeddingHealthBadge, EmbeddingMetricsPanel, EmbeddingModelStatusCard
v1.5.32026-08-02ElBruno.LocalEmbeddings.OpenTelemetry β€” OpenTelemetry instrumentation with Activity tracing, metrics (tokens/sec, latency), and configurable exporters for embedding pipelines
v1.4.x2026NPU hardware acceleration β€” ElBruno.LocalEmbeddings.Npu.Intel (Intel Core Ultra via OpenVINO) and ElBruno.LocalEmbeddings.Npu.Qualcomm (Snapdragon X via QNN) for on-device NPU inference

πŸ“¦ NuGet Packages

PackageVersionDownloadsDescription
ElBruno.LocalEmbeddingsNuGetDownloadsCore library β€” ONNX Runtime + Microsoft.Extensions.AI
ElBruno.LocalEmbeddings.HarrierNuGetDownloadsMicrosoft Harrier-OSS-v1 (270M, 640-dim, 94+ languages)
ElBruno.LocalEmbeddings.ImageEmbeddingsNuGetDownloadsCLIP-based image embeddings for multimodal search
ElBruno.LocalEmbeddings.ImageEmbeddings.DownloaderNuGetDownloadsModel downloader for image embeddings
ElBruno.LocalEmbeddings.KernelMemoryNuGetDownloadsMicrosoft Kernel Memory adapter
ElBruno.LocalEmbeddings.VectorDataNuGetDownloadsMicrosoft.Extensions.VectorData + InMemoryVectorStore
ElBruno.LocalEmbeddings.NpuNuGetDownloadsNPU-accelerated embeddings via DirectML
ElBruno.LocalEmbeddings.Npu.IntelNuGetDownloadsIntel Core Ultra NPU via OpenVINO
ElBruno.LocalEmbeddings.Npu.QualcommNuGetDownloadsQualcomm Snapdragon X NPU via QNN
ElBruno.LocalEmbeddings.BlazorComponentsNuGetDownloadsBlazor UI components β€” model gallery, similarity meter, semantic search, embedding explorer

Installation

dotnet add package ElBruno.LocalEmbeddings

For Harrier model support, install the companion package:

dotnet add package ElBruno.LocalEmbeddings.Harrier

For Kernel Memory integration, also install:

dotnet add package ElBruno.LocalEmbeddings.KernelMemory

For VectorData integration, install:

dotnet add package ElBruno.LocalEmbeddings.VectorData

Quick Start

1) Generate one embedding

using ElBruno.LocalEmbeddings;

await using var generator = await LocalEmbeddingGenerator.CreateAsync();
var embedding = await generator.GenerateEmbeddingAsync("Hello, world!");
Console.WriteLine(embedding.Vector.Length); // 384

2) Generate embeddings for multiple texts

var inputs = new[] { "first text", "second text", "third text" };
var embeddings = await generator.GenerateAsync(inputs);
Console.WriteLine(embeddings.Count); // 3

3) Compare two texts with cosine similarity

using ElBruno.LocalEmbeddings.Extensions;

var pair = await generator.GenerateAsync(["I love coding", "I enjoy programming"]);
var score = pair[0].CosineSimilarity(pair[1]);
Console.WriteLine(score);

4) Semantic search in one line

var corpus = new[]
{
    "Python for data science",
    "JavaScript for web apps",
    "Swift for iOS development"
};

var corpusEmbeddings = await generator.GenerateAsync(corpus);
var results = await generator.FindClosestAsync(
    "best language for websites",
    corpus,
    corpusEmbeddings,
    topK: 2,
    minScore: 0.2f);

foreach (var result in results)
    Console.WriteLine($"{result.Score:F3} - {result.Text}");

5) Using Harrier model (higher quality, multilingual)

using ElBruno.LocalEmbeddings.Harrier;

var generator = await HarrierEmbeddingGenerator.CreateAsync();
var embedding = await generator.GenerateEmbeddingAsync("Hello, world!");
Console.WriteLine($"Dimensions: {embedding.Vector.Length}"); // 640

For custom models and runtime behavior, use the options-based constructor: new LocalEmbeddingGenerator(new LocalEmbeddingsOptions { ... }).

Note: The synchronous constructor remains available for backward compatibility, but performs blocking initialization when downloads are needed.

Want to go further? Read the Getting Started guide and the other docs in this repo for DI, configuration, VectorData, Kernel Memory, and full RAG examples.

Prefer a containerized dev environment? See the Dev Container section in the Contributing guide.

Samples

All sample applications are located in src/Samples/. See the samples README for prerequisites and run instructions.

SampleWhat It Shows
HelloWorldAltModelMinimal hello world with sentence-transformers/all-MiniLM-L12-v2
ConsoleAppAll the basics: single/batch embeddings, similarity, semantic search, DI
HarrierConsoleAppHarrier embedding model usage and similarity search
RagChatEmbedding-only semantic search Q&A using shared VectorData InMemoryVectorStore (no LLM needed)
RagOllamaFull RAG with Ollama + phi4-mini + Kernel Memory
RagFoundryLocalFull RAG with Foundry Local + phi4-mini
ImageRagSimpleMinimal image RAG: index images β†’ search by text
ImageRagChatInteractive image RAG chat with text and image-to-image search

Configuration

var options = new LocalEmbeddingsOptions
{
    ModelName = "sentence-transformers/all-MiniLM-L6-v2",  // HuggingFace model
    MaxSequenceLength = 512,                                // Max tokens
    CacheDirectory = null,                                  // Auto-detect per platform
    EnsureModelDownloaded = true,                           // Download if missing
    NormalizeEmbeddings = false                              // L2 normalize vectors
};

See Configuration docs for supported models, local model paths, and cache locations.

Common model options (with model cards)

Estimated download sizes below are approximate and can vary by ONNX variant (fp32/int8) and tokenizer assets.

Documentation

TopicDescription
Getting StartedStep-by-step guide from hello world to RAG
API ReferenceClasses, methods, and extension methods
ConfigurationOptions, supported models, cache locations
Alternative ModelsNon-default free models, local download workflow, and license notes
Dependency InjectionAll DI overloads and IConfiguration binding
Harrier IntegrationMicrosoft Harrier-OSS-v1 local embedding model
Kernel Memory IntegrationUsing local embeddings with Microsoft Kernel Memory
VectorData IntegrationUsing local embeddings with Microsoft.Extensions.VectorData abstractions
Blazor Components9 ready-to-use Blazor components for embedding-powered web apps
ContributingBuild from source, repo structure, guidelines
RoadmapPlanned and completed features/samples with priorities
PublishingNuGet publishing with GitHub Actions + Trusted Publishing
ChangelogVersioned summary of notable changes

Have an idea for a new feature or sample? Please open an issue and share your suggestion.

Building from Source

git clone https://github.com/elbruno/elbruno.localembeddings.git
cd elbruno.localembeddings
dotnet build
dotnet test

Requirements

  • .NET 8.0 SDK (LTS) or .NET 10.0 SDK
  • ONNX Runtime compatible platform (Windows, Linux, macOS)

πŸ‘‹ About the Author

Hi! I'm ElBruno 🧑, a passionate developer and content creator exploring AI, .NET, and modern development practices.

Made with ❀️ by ElBruno

If you like this project, consider following my work across platforms:

  • πŸ“» Podcast: No Tienen Nombre β€” Spanish-language episodes on AI, development, and tech culture
  • πŸ’» Blog: ElBruno.com β€” Deep dives on embeddings, RAG, .NET, and local AI
  • πŸ“Ί YouTube: youtube.com/elbruno β€” Demos, tutorials, and live coding
  • πŸ”— LinkedIn: @elbruno β€” Professional updates and insights
  • 𝕏 Twitter: @elbruno β€” Quick tips, releases, and tech news

License

This project is licensed under the MIT License β€” see the LICENSE file for details.

cosine-similarity
csharp
dependency-injection
dotnet
embeddings
huggingface
kernel-memory
local-ai
machine-learning
microsoft-extensions-ai
nuget
onnx
onnx-runtime
rag
semantic-search
text-embeddings
vector-embeddings

Languages

C#

94.0%

HTML

2.7%

JavaScript

1.2%