harshil-sh/WeaveLLM

WeaveLLM — Production-grade AI orchestration for .NET. Strongly typed chains, RAG pipelines, multi-agent graphs, and tool use. The LangChain alternative .NET developers actually deserve.

C#

2

41 commits

updated May 8, 2026

See the code

README

WeaveLLM

NuGet Version Build Status License: MIT

A composable AI orchestration framework for .NET — build LLM chains, RAG pipelines, and autonomous agents with idiomatic C#.


Why WeaveLLM?

Pain points with LangChain / Semantic Kernel:

  • LangChain is Python-first; .NET ports lag behind and feel bolted on
  • Semantic Kernel's kernel-centric model adds ceremony around simple chat calls
  • Both frameworks hide errors inside exceptions, making retry logic hard to compose
  • Neither integrates naturally with ASP.NET Core DI, health checks, or OpenTelemetry

WeaveLLM advantages:

  • Native .NET 8 — built on IServiceCollection, IAsyncEnumerable, and the BCL; no Python bridges
  • Railway-oriented results — every call returns ChainResult<T> instead of throwing; errors are first-class values
  • ASP.NET-style middleware pipeline — add retry, caching, rate limiting, PII scrubbing, and cost tracking as composable layers
  • Single fluent registration — one AddWeaveLLM() call wires providers, memory, agents, health checks, and telemetry

Installation

dotnet add package WeaveLLM.Core
dotnet add package WeaveLLM.Providers      # OpenAI, Anthropic, Ollama, HuggingFace
dotnet add package WeaveLLM.Memory         # RAG, vector stores, chunking
dotnet add package WeaveLLM.Observability  # OpenTelemetry, cost tracking, PII scrubbing
dotnet add package WeaveLLM.Extensions.DependencyInjection  # ASP.NET Core integration

Quick Start

using WeaveLLM.Core.Models;
using WeaveLLM.Extensions.DependencyInjection;

var builder = WebApplication.CreateBuilder(args);

builder.Services
    .AddWeaveLLM()
    .AddOpenAI(apiKey: builder.Configuration["OpenAI:ApiKey"]!, modelId: "gpt-4o")
    .AddInMemoryMemory()
    .AddToolRegistry()
    .AddReActAgent();

var app = builder.Build();

app.MapPost("/chat", async (ChatRequest req, IChatModel model) =>
{
    var result = await model.ChatAsync([Message.User(req.Message)]);
    return result.IsSuccess
        ? Results.Ok(result.Value!.Content)
        : Results.Problem(result.Error!.Message);
});

app.Run();

record ChatRequest(string Message);

Features

Composable Chains

Build type-safe pipelines from reusable IChain<TInput, TOutput> blocks. Wrap any chain with middleware layers using PipelineBuilder:

var pipeline = new PipelineBuilder<IReadOnlyList<Message>, ChatResponse>(chatChain)
    .UseMiddleware(new RetryMiddleware<IReadOnlyList<Message>, ChatResponse>(maxRetries: 3))
    .UseMiddleware(new CacheMiddleware<IReadOnlyList<Message>, ChatResponse>(ttl: TimeSpan.FromMinutes(5)));

var chain = await pipeline.BuildAsync();
var result = await chain.ExecuteAsync([Message.User("Hello")]);

Streaming

Every provider supports IAsyncEnumerable<string> streaming out of the box:

app.MapGet("/chat/stream", async (string message, IChatModel model, HttpContext http) =>
{
    http.Response.ContentType = "text/event-stream";
    await foreach (var chunk in model.StreamChatAsync([Message.User(message)]))
        await http.Response.WriteAsync($"data: {chunk}\n\n");
});

Middleware

Cross-cutting concerns slot into the chain pipeline without modifying chain logic. Implement one interface:

public sealed class LoggingMiddleware<TIn, TOut> : IChainMiddleware<TIn, TOut>
{
    public async Task<ChainResult<TOut>> InvokeAsync(
        TIn input, ChainContext ctx, ChainDelegate<TIn, TOut> next, CancellationToken ct)
    {
        Console.WriteLine($"[{ctx.ChainName}] invoking");
        var result = await next(input, ctx, ct);
        Console.WriteLine($"[{ctx.ChainName}] success={result.IsSuccess}");
        return result;
    }
}

Built-in: RetryMiddleware (exponential backoff), CacheMiddleware, RateLimitingMiddleware, TracingMiddleware, CostMiddleware, PiiScrubbingMiddleware.

Memory

Persist conversation history and enable semantic search with a swappable IMemoryStore / IVectorStore:

builder.Services.AddWeaveLLM()
    .AddOpenAI(apiKey)
    .AddInMemoryMemory();  // swap for QdrantVectorStore or PostgresVectorStore

// In a handler:
var history = new List<Message>();
await foreach (var entry in memory.GetAsync(sessionId, limit: 20, ct))
    history.Add(entry.Message);

history.Add(Message.User(userInput));
var response = await model.ChatAsync(history, cancellationToken: ct);
await memory.AddAsync(
    new MemoryEntry(sessionId, Message.Assistant(response.Value!.Content), DateTimeOffset.UtcNow), ct);

Supported vector backends: In-Memory, Qdrant, PostgreSQL + pgvector.

Agents

Two strategies ship out of the box, both composable with [LLMTool]-annotated methods:

public sealed class CalculatorTool
{
    [LLMTool("calculate", "Evaluate a mathematical expression")]
    public string Calculate([Description("e.g. '2 + 2 * 10'")] string expression)
    {
        var table = new System.Data.DataTable();
        return table.Compute(expression, string.Empty)?.ToString() ?? "0";
    }
}

builder.Services.AddWeaveLLM()
    .AddOpenAI(apiKey)
    .AddToolRegistry(r => r.RegisterFromObject(new CalculatorTool()))
    .AddReActAgent(maxSteps: 10);

// Inject IAgent and run:
var result = await agent.RunAsync("What is (123 * 456) + 789?", ct);
Console.WriteLine(result.Value!.FinalAnswer);

ReActAgent — Thought → Action → Observation loop until the model emits a Final Answer
PlanAndExecuteAgent — separate planning and execution phases for complex multi-step tasks


Supported Providers

ProviderChatStreamingEmbeddingsNotes
OpenAI✅✅✅Default model: gpt-4o
Anthropic✅✅—Default model: claude-sonnet-4-5
Azure OpenAIPlannedPlannedPlannedv0.2.0-alpha
Ollama✅✅✅Local inference; default model: llama3
HuggingFace✅—✅Inference API; bring your own model ID

Examples

SampleDescription
samples/WeaveLLM.Sample.BasicChainMulti-provider ASP.NET Core API with session memory, SSE streaming, agent endpoint, and RAG index/query routes
samples/ChatWithDocsConsole RAG chatbot — loads a local docs/ folder, indexes into a vector store, answers questions with source citations

Roadmap

VersionStatusHighlights
v0.1.0-alpha✅ CurrentCore chains · OpenAI / Anthropic / Ollama / HuggingFace providers · ReAct + PlanAndExecute agents · RAG pipeline · In-Memory / Qdrant / Postgres vector stores · OpenTelemetry · PII scrubbing · cost tracking
v0.2.0-alphaPlannedAzure OpenAI provider · multi-modal (image) input · streaming agents · evaluation suite GA
v1.0.0Target: Q4 2026Stable public API · NuGet stable release · documentation site

Testing with WeaveLLM

Add the testing helper package to your test project — no real API keys or network calls required:

dotnet add package WeaveLLM.Testing

FakeStreamingChatModel is a fully configurable in-process test double for IChatModel. Set Tokens for streaming assertions, BlockingResponse for ChatAsync, and ErrorAfterTokens to simulate mid-stream failures:

using WeaveLLM.Testing;

var fake = new FakeStreamingChatModel
{
    Tokens = ["Hello", ", ", "world", "!"],
    BlockingResponse = "Hello, world!"
};

// Blocking path
var result = await fake.ChatAsync([Message.User("Hi")]);
result.IsSuccess.Should().BeTrue();
result.Value!.Content.Should().Be("Hello, world!");

// Streaming path
var tokens = new List<string>();
await foreach (var chunk in fake.StreamChatAsync([Message.User("Hi")]))
    tokens.Add(chunk);
tokens.Should().Equal("Hello", ", ", "world", "!");

// Error injection — yields 2 tokens then a failure
fake.ErrorAfterTokens = WeaveLLMError.ProviderError("fake", "Quota exceeded");
fake.TokensBeforeError = 2;
var results = new List<ChainResult<string>>();
await foreach (var item in fake.StreamChatSafeAsync([Message.User("Hi")]))
    results.Add(item);
results[2].IsFailure.Should().BeTrue();
results[2].Error!.Code.Should().Be("PROVIDER_ERROR");

Option B — NSubstitute manual pattern

When you need a mock rather than a fake, use NSubstitute with a static async-iterator helper. The [EnumeratorCancellation] attribute on the CancellationToken parameter is required — without it, the token passed via .WithCancellation(ct) is silently ignored and cancellation tests never fire:

using NSubstitute;
using System.Runtime.CompilerServices;

var model = Substitute.For<IChatModel>();

model.StreamChatAsync(Arg.Any<IReadOnlyList<Message>>(), Arg.Any<LLMOptions?>(), Arg.Any<CancellationToken>())
     .Returns(FakeStream("Hello", ", ", "world"));

static async IAsyncEnumerable<string> FakeStream(
    params string[] tokens,
    [EnumeratorCancellation] CancellationToken ct = default)   // ← required
{
    foreach (var token in tokens)
    {
        ct.ThrowIfCancellationRequested();
        yield return token;
        await Task.Yield();
    }
}

Contributing

Contributions are welcome. Please read CONTRIBUTING.md for branch conventions, coding guidelines, and the pull-request checklist before opening a PR.


License

WeaveLLM is released under the MIT License.


Star History

Star History Chart

agents
ai
anthropic
aspnetcore
csharp
donet8
dotnet
embeddings
generative-ai
langchain
langchain-alternative
llm
multi-agent
nuget
ollama
openai
opentelemetry
orchestration
rag
vector-search

harshil-sh/WeaveLLM

WeaveLLM — Production-grade AI orchestration for .NET. Strongly typed chains, RAG pipelines, multi-agent graphs, and tool use. The LangChain alternative .NET developers actually deserve.

C#

2

41 commits

updated May 8, 2026

See the code

README

WeaveLLM

NuGet Version Build Status License: MIT

A composable AI orchestration framework for .NET — build LLM chains, RAG pipelines, and autonomous agents with idiomatic C#.


Why WeaveLLM?

Pain points with LangChain / Semantic Kernel:

  • LangChain is Python-first; .NET ports lag behind and feel bolted on
  • Semantic Kernel's kernel-centric model adds ceremony around simple chat calls
  • Both frameworks hide errors inside exceptions, making retry logic hard to compose
  • Neither integrates naturally with ASP.NET Core DI, health checks, or OpenTelemetry

WeaveLLM advantages:

  • Native .NET 8 — built on IServiceCollection, IAsyncEnumerable, and the BCL; no Python bridges
  • Railway-oriented results — every call returns ChainResult<T> instead of throwing; errors are first-class values
  • ASP.NET-style middleware pipeline — add retry, caching, rate limiting, PII scrubbing, and cost tracking as composable layers
  • Single fluent registration — one AddWeaveLLM() call wires providers, memory, agents, health checks, and telemetry

Installation

dotnet add package WeaveLLM.Core
dotnet add package WeaveLLM.Providers      # OpenAI, Anthropic, Ollama, HuggingFace
dotnet add package WeaveLLM.Memory         # RAG, vector stores, chunking
dotnet add package WeaveLLM.Observability  # OpenTelemetry, cost tracking, PII scrubbing
dotnet add package WeaveLLM.Extensions.DependencyInjection  # ASP.NET Core integration

Quick Start

using WeaveLLM.Core.Models;
using WeaveLLM.Extensions.DependencyInjection;

var builder = WebApplication.CreateBuilder(args);

builder.Services
    .AddWeaveLLM()
    .AddOpenAI(apiKey: builder.Configuration["OpenAI:ApiKey"]!, modelId: "gpt-4o")
    .AddInMemoryMemory()
    .AddToolRegistry()
    .AddReActAgent();

var app = builder.Build();

app.MapPost("/chat", async (ChatRequest req, IChatModel model) =>
{
    var result = await model.ChatAsync([Message.User(req.Message)]);
    return result.IsSuccess
        ? Results.Ok(result.Value!.Content)
        : Results.Problem(result.Error!.Message);
});

app.Run();

record ChatRequest(string Message);

Features

Composable Chains

Build type-safe pipelines from reusable IChain<TInput, TOutput> blocks. Wrap any chain with middleware layers using PipelineBuilder:

var pipeline = new PipelineBuilder<IReadOnlyList<Message>, ChatResponse>(chatChain)
    .UseMiddleware(new RetryMiddleware<IReadOnlyList<Message>, ChatResponse>(maxRetries: 3))
    .UseMiddleware(new CacheMiddleware<IReadOnlyList<Message>, ChatResponse>(ttl: TimeSpan.FromMinutes(5)));

var chain = await pipeline.BuildAsync();
var result = await chain.ExecuteAsync([Message.User("Hello")]);

Streaming

Every provider supports IAsyncEnumerable<string> streaming out of the box:

app.MapGet("/chat/stream", async (string message, IChatModel model, HttpContext http) =>
{
    http.Response.ContentType = "text/event-stream";
    await foreach (var chunk in model.StreamChatAsync([Message.User(message)]))
        await http.Response.WriteAsync($"data: {chunk}\n\n");
});

Middleware

Cross-cutting concerns slot into the chain pipeline without modifying chain logic. Implement one interface:

public sealed class LoggingMiddleware<TIn, TOut> : IChainMiddleware<TIn, TOut>
{
    public async Task<ChainResult<TOut>> InvokeAsync(
        TIn input, ChainContext ctx, ChainDelegate<TIn, TOut> next, CancellationToken ct)
    {
        Console.WriteLine($"[{ctx.ChainName}] invoking");
        var result = await next(input, ctx, ct);
        Console.WriteLine($"[{ctx.ChainName}] success={result.IsSuccess}");
        return result;
    }
}

Built-in: RetryMiddleware (exponential backoff), CacheMiddleware, RateLimitingMiddleware, TracingMiddleware, CostMiddleware, PiiScrubbingMiddleware.

Memory

Persist conversation history and enable semantic search with a swappable IMemoryStore / IVectorStore:

builder.Services.AddWeaveLLM()
    .AddOpenAI(apiKey)
    .AddInMemoryMemory();  // swap for QdrantVectorStore or PostgresVectorStore

// In a handler:
var history = new List<Message>();
await foreach (var entry in memory.GetAsync(sessionId, limit: 20, ct))
    history.Add(entry.Message);

history.Add(Message.User(userInput));
var response = await model.ChatAsync(history, cancellationToken: ct);
await memory.AddAsync(
    new MemoryEntry(sessionId, Message.Assistant(response.Value!.Content), DateTimeOffset.UtcNow), ct);

Supported vector backends: In-Memory, Qdrant, PostgreSQL + pgvector.

Agents

Two strategies ship out of the box, both composable with [LLMTool]-annotated methods:

public sealed class CalculatorTool
{
    [LLMTool("calculate", "Evaluate a mathematical expression")]
    public string Calculate([Description("e.g. '2 + 2 * 10'")] string expression)
    {
        var table = new System.Data.DataTable();
        return table.Compute(expression, string.Empty)?.ToString() ?? "0";
    }
}

builder.Services.AddWeaveLLM()
    .AddOpenAI(apiKey)
    .AddToolRegistry(r => r.RegisterFromObject(new CalculatorTool()))
    .AddReActAgent(maxSteps: 10);

// Inject IAgent and run:
var result = await agent.RunAsync("What is (123 * 456) + 789?", ct);
Console.WriteLine(result.Value!.FinalAnswer);

ReActAgent — Thought → Action → Observation loop until the model emits a Final Answer
PlanAndExecuteAgent — separate planning and execution phases for complex multi-step tasks


Supported Providers

ProviderChatStreamingEmbeddingsNotes
OpenAI✅✅✅Default model: gpt-4o
Anthropic✅✅—Default model: claude-sonnet-4-5
Azure OpenAIPlannedPlannedPlannedv0.2.0-alpha
Ollama✅✅✅Local inference; default model: llama3
HuggingFace✅—✅Inference API; bring your own model ID

Examples

SampleDescription
samples/WeaveLLM.Sample.BasicChainMulti-provider ASP.NET Core API with session memory, SSE streaming, agent endpoint, and RAG index/query routes
samples/ChatWithDocsConsole RAG chatbot — loads a local docs/ folder, indexes into a vector store, answers questions with source citations

Roadmap

VersionStatusHighlights
v0.1.0-alpha✅ CurrentCore chains · OpenAI / Anthropic / Ollama / HuggingFace providers · ReAct + PlanAndExecute agents · RAG pipeline · In-Memory / Qdrant / Postgres vector stores · OpenTelemetry · PII scrubbing · cost tracking
v0.2.0-alphaPlannedAzure OpenAI provider · multi-modal (image) input · streaming agents · evaluation suite GA
v1.0.0Target: Q4 2026Stable public API · NuGet stable release · documentation site

Testing with WeaveLLM

Add the testing helper package to your test project — no real API keys or network calls required:

dotnet add package WeaveLLM.Testing

FakeStreamingChatModel is a fully configurable in-process test double for IChatModel. Set Tokens for streaming assertions, BlockingResponse for ChatAsync, and ErrorAfterTokens to simulate mid-stream failures:

using WeaveLLM.Testing;

var fake = new FakeStreamingChatModel
{
    Tokens = ["Hello", ", ", "world", "!"],
    BlockingResponse = "Hello, world!"
};

// Blocking path
var result = await fake.ChatAsync([Message.User("Hi")]);
result.IsSuccess.Should().BeTrue();
result.Value!.Content.Should().Be("Hello, world!");

// Streaming path
var tokens = new List<string>();
await foreach (var chunk in fake.StreamChatAsync([Message.User("Hi")]))
    tokens.Add(chunk);
tokens.Should().Equal("Hello", ", ", "world", "!");

// Error injection — yields 2 tokens then a failure
fake.ErrorAfterTokens = WeaveLLMError.ProviderError("fake", "Quota exceeded");
fake.TokensBeforeError = 2;
var results = new List<ChainResult<string>>();
await foreach (var item in fake.StreamChatSafeAsync([Message.User("Hi")]))
    results.Add(item);
results[2].IsFailure.Should().BeTrue();
results[2].Error!.Code.Should().Be("PROVIDER_ERROR");

Option B — NSubstitute manual pattern

When you need a mock rather than a fake, use NSubstitute with a static async-iterator helper. The [EnumeratorCancellation] attribute on the CancellationToken parameter is required — without it, the token passed via .WithCancellation(ct) is silently ignored and cancellation tests never fire:

using NSubstitute;
using System.Runtime.CompilerServices;

var model = Substitute.For<IChatModel>();

model.StreamChatAsync(Arg.Any<IReadOnlyList<Message>>(), Arg.Any<LLMOptions?>(), Arg.Any<CancellationToken>())
     .Returns(FakeStream("Hello", ", ", "world"));

static async IAsyncEnumerable<string> FakeStream(
    params string[] tokens,
    [EnumeratorCancellation] CancellationToken ct = default)   // ← required
{
    foreach (var token in tokens)
    {
        ct.ThrowIfCancellationRequested();
        yield return token;
        await Task.Yield();
    }
}

Contributing

Contributions are welcome. Please read CONTRIBUTING.md for branch conventions, coding guidelines, and the pull-request checklist before opening a PR.


License

WeaveLLM is released under the MIT License.


Star History

Star History Chart

agents
ai
anthropic
aspnetcore
csharp
donet8
dotnet
embeddings
generative-ai
langchain
langchain-alternative
llm
multi-agent
nuget
ollama
openai
opentelemetry
orchestration
rag
vector-search