A lightweight lexical text search and classification library for .NET — BM25/TF-IDF tuning, explainable scoring, hybrid federation, PostgreSQL backends
See the codeA composable information retrieval toolkit for .NET — build, measure and inspect search pipelines, from lexical BM25 to hybrid and reranked retrieval.
Index, retrieve, rank and judge a search pipeline: an in-memory inverted index, four ranking strategies and their BM25 variants, rank fusion, reranking, optional PostgreSQL backends, and model-agnostic seams for dense, learned-sparse and neural scoring — the models stay in your application. The core package references no NuGet package at all.
📖 Full documentation → — the guide, the
reference, and every measurement with the command that reproduces it. Published from
docs/; this README is the short version and the details are delegated to those
pages.
dotnet add package LexiSharp # the core: index, scorers, engines, decorators — no dependencies
net10.0, MIT. Optional: LexiSharp.MessagePack (binary index persistence),
LexiSharp.AspNetCore (a GET /search minimal-API endpoint), LexiSharp.Postgres
(lexical, vector, sparse, fuzzy and true BM25 backends) — packages.
using LexiSharp.Core;
using LexiSharp.Indexing;
using LexiSharp.Ranking;
// Three pieces, one contract: the index owns the corpus statistics, the scorer is a pure
// ranking strategy reading from it, the engine orchestrates. Each of the three is an
// interface, so you replace one without touching the others.
ITextSearchEngine engine = new RankedTextSearchEngine(
new InMemoryTextIndex(),
new Bm25Scorer());
engine.Index(new[]
{
new SearchDocument("1", "The search engine uses BM25 to rank the results"),
new SearchDocument("2", "TF-IDF is a classic method of textual search"),
new SearchDocument("3", "Italian cuisine is renowned in Rome"),
});
foreach (var result in engine.Search("textual search"))
Console.WriteLine($"{result.DocumentId} - {result.Score:0.###}: {result.Document.Text}");
That is the smallest thing the library does. The rest is composition: every engine above
implements ITextSearchEngine, and a pipeline is stages wrapping each other. Given two
engines over one index, HashingEmbeddingProvider standing in for your
IEmbeddingProvider (no model, no service):
var hybrid = new HybridTextSearchEngine(
new[] { lexical, dense },
new ReciprocalRankFusionMerger()); // a BM25 score and a cosine, fused by rank, uncalibrated
ITextSearchEngine pipeline = new RerankedTextSearchEngine(
hybrid, new ProximityReranker(index)); // ...or MMR, a cascade, MaxSim, a cross-encoder
Replacing a piece is the whole extension model — a PostgreSQL, vector, sparse or fuzzy
backend takes the same slot, and IEmbeddingProvider, ISparseEmbeddingProvider and
ICrossEncoderScorer are yours to implement: pipelines, and
backends.
LexiSharpIndex<T> is the typed facade over the same engine if you would rather hand it your
own objects — Getting started.
dotnet run --project samples/LexiSharp.Demo # → http://localhost:5000
Five retrieval strategies over one corpus, compared live — BM25, corpus-derived semantic expansion, dense hashing embeddings, RRF fusion and a term-overlap rerank — with per-lane latency, highlighting and a click-through "why did this rank here?" panel. No model, no external service.

Stated plainly, so nothing is implied. The full list, with the measurement behind each claim, is Scope and limits.
trec_eval — the standard evaluator, not this repository's code — reads both figures off
a run this harness writes, and the 15,466 returned scores for the 1,406 queries match the
reference's own searcher on the raw bits, so the ranking is that ranking rather than a lookalike.
The same scores read under the library defaults are 0.308 / 0.662 / 0.320; that difference is
the analyzer, the parameters and one task convention, not the ranking. Corpora are md5-verified on
download, and the numbers are pinned and re-checked by the Pinned reference workflow, which
replays every pinned configuration and exits non-zero on drift. It runs on a dispatch, on a push
that touches the library or the harness, and weekly — see evaluation.δ, tie a tuned BM25 on the reference corpus and NFCorpus and edge it by 0.002–0.004 on
SciFact — an in-sample margin, so an upper bound rather than a result. On ArguAna, untuned,
they lose, and no δ-tuned ArguAna row exists, so whether tuning closes that gap is
unmeasured (ranking).dotnet build LexiSharp.slnx
dotnet test tests/LexiSharp.Tests # xUnit suite; the Postgres suites need POSTGRES_TEST_CONNECTION
The retrieval quality gate replays every pinned configuration on the three BEIR corpora and exits non-zero on any drift. It downloads the corpora on first run, so it is not part of the xUnit suite:
dotnet run --project bench/LexiSharp.Eval -c Release -- --verify-reference
Benchmarks, the evaluation harness and the behavioural gate: Reference and Benchmarks.
MIT — see LICENSE. The ParadeDB pg_search extension used by the BM25 backend is
licensed separately, under AGPL-3.
C#
99.1%
A lightweight lexical text search and classification library for .NET — BM25/TF-IDF tuning, explainable scoring, hybrid federation, PostgreSQL backends
See the codeA composable information retrieval toolkit for .NET — build, measure and inspect search pipelines, from lexical BM25 to hybrid and reranked retrieval.
Index, retrieve, rank and judge a search pipeline: an in-memory inverted index, four ranking strategies and their BM25 variants, rank fusion, reranking, optional PostgreSQL backends, and model-agnostic seams for dense, learned-sparse and neural scoring — the models stay in your application. The core package references no NuGet package at all.
📖 Full documentation → — the guide, the
reference, and every measurement with the command that reproduces it. Published from
docs/; this README is the short version and the details are delegated to those
pages.
dotnet add package LexiSharp # the core: index, scorers, engines, decorators — no dependencies
net10.0, MIT. Optional: LexiSharp.MessagePack (binary index persistence),
LexiSharp.AspNetCore (a GET /search minimal-API endpoint), LexiSharp.Postgres
(lexical, vector, sparse, fuzzy and true BM25 backends) — packages.
using LexiSharp.Core;
using LexiSharp.Indexing;
using LexiSharp.Ranking;
// Three pieces, one contract: the index owns the corpus statistics, the scorer is a pure
// ranking strategy reading from it, the engine orchestrates. Each of the three is an
// interface, so you replace one without touching the others.
ITextSearchEngine engine = new RankedTextSearchEngine(
new InMemoryTextIndex(),
new Bm25Scorer());
engine.Index(new[]
{
new SearchDocument("1", "The search engine uses BM25 to rank the results"),
new SearchDocument("2", "TF-IDF is a classic method of textual search"),
new SearchDocument("3", "Italian cuisine is renowned in Rome"),
});
foreach (var result in engine.Search("textual search"))
Console.WriteLine($"{result.DocumentId} - {result.Score:0.###}: {result.Document.Text}");
That is the smallest thing the library does. The rest is composition: every engine above
implements ITextSearchEngine, and a pipeline is stages wrapping each other. Given two
engines over one index, HashingEmbeddingProvider standing in for your
IEmbeddingProvider (no model, no service):
var hybrid = new HybridTextSearchEngine(
new[] { lexical, dense },
new ReciprocalRankFusionMerger()); // a BM25 score and a cosine, fused by rank, uncalibrated
ITextSearchEngine pipeline = new RerankedTextSearchEngine(
hybrid, new ProximityReranker(index)); // ...or MMR, a cascade, MaxSim, a cross-encoder
Replacing a piece is the whole extension model — a PostgreSQL, vector, sparse or fuzzy
backend takes the same slot, and IEmbeddingProvider, ISparseEmbeddingProvider and
ICrossEncoderScorer are yours to implement: pipelines, and
backends.
LexiSharpIndex<T> is the typed facade over the same engine if you would rather hand it your
own objects — Getting started.
dotnet run --project samples/LexiSharp.Demo # → http://localhost:5000
Five retrieval strategies over one corpus, compared live — BM25, corpus-derived semantic expansion, dense hashing embeddings, RRF fusion and a term-overlap rerank — with per-lane latency, highlighting and a click-through "why did this rank here?" panel. No model, no external service.

Stated plainly, so nothing is implied. The full list, with the measurement behind each claim, is Scope and limits.
trec_eval — the standard evaluator, not this repository's code — reads both figures off
a run this harness writes, and the 15,466 returned scores for the 1,406 queries match the
reference's own searcher on the raw bits, so the ranking is that ranking rather than a lookalike.
The same scores read under the library defaults are 0.308 / 0.662 / 0.320; that difference is
the analyzer, the parameters and one task convention, not the ranking. Corpora are md5-verified on
download, and the numbers are pinned and re-checked by the Pinned reference workflow, which
replays every pinned configuration and exits non-zero on drift. It runs on a dispatch, on a push
that touches the library or the harness, and weekly — see evaluation.δ, tie a tuned BM25 on the reference corpus and NFCorpus and edge it by 0.002–0.004 on
SciFact — an in-sample margin, so an upper bound rather than a result. On ArguAna, untuned,
they lose, and no δ-tuned ArguAna row exists, so whether tuning closes that gap is
unmeasured (ranking).dotnet build LexiSharp.slnx
dotnet test tests/LexiSharp.Tests # xUnit suite; the Postgres suites need POSTGRES_TEST_CONNECTION
The retrieval quality gate replays every pinned configuration on the three BEIR corpora and exits non-zero on any drift. It downloads the corpora on first run, so it is not part of the xUnit suite:
dotnet run --project bench/LexiSharp.Eval -c Release -- --verify-reference
Benchmarks, the evaluation harness and the behavioural gate: Reference and Benchmarks.
MIT — see LICENSE. The ParadeDB pg_search extension used by the BM25 backend is
licensed separately, under AGPL-3.
C#
99.1%