WeaveLLM — Production-grade AI orchestration for .NET. Strongly typed chains, RAG pipelines, multi-agent graphs, and tool use. The LangChain alternative .NET developers actually deserve.
C#
2
41 commits
updated May 8, 2026
A composable AI orchestration framework for .NET — build LLM chains, RAG pipelines, and autonomous agents with idiomatic C#.
Pain points with LangChain / Semantic Kernel:
WeaveLLM advantages:
IServiceCollection, IAsyncEnumerable, and the BCL; no Python bridgesChainResult<T> instead of throwing; errors are first-class valuesAddWeaveLLM() call wires providers, memory, agents, health checks, and telemetrydotnet add package WeaveLLM.Core
dotnet add package WeaveLLM.Providers # OpenAI, Anthropic, Ollama, HuggingFace
dotnet add package WeaveLLM.Memory # RAG, vector stores, chunking
dotnet add package WeaveLLM.Observability # OpenTelemetry, cost tracking, PII scrubbing
dotnet add package WeaveLLM.Extensions.DependencyInjection # ASP.NET Core integration
using WeaveLLM.Core.Models;
using WeaveLLM.Extensions.DependencyInjection;
var builder = WebApplication.CreateBuilder(args);
builder.Services
.AddWeaveLLM()
.AddOpenAI(apiKey: builder.Configuration["OpenAI:ApiKey"]!, modelId: "gpt-4o")
.AddInMemoryMemory()
.AddToolRegistry()
.AddReActAgent();
var app = builder.Build();
app.MapPost("/chat", async (ChatRequest req, IChatModel model) =>
{
var result = await model.ChatAsync([Message.User(req.Message)]);
return result.IsSuccess
? Results.Ok(result.Value!.Content)
: Results.Problem(result.Error!.Message);
});
app.Run();
record ChatRequest(string Message);
Build type-safe pipelines from reusable IChain<TInput, TOutput> blocks. Wrap any chain with middleware layers using PipelineBuilder:
var pipeline = new PipelineBuilder<IReadOnlyList<Message>, ChatResponse>(chatChain)
.UseMiddleware(new RetryMiddleware<IReadOnlyList<Message>, ChatResponse>(maxRetries: 3))
.UseMiddleware(new CacheMiddleware<IReadOnlyList<Message>, ChatResponse>(ttl: TimeSpan.FromMinutes(5)));
var chain = await pipeline.BuildAsync();
var result = await chain.ExecuteAsync([Message.User("Hello")]);
Every provider supports IAsyncEnumerable<string> streaming out of the box:
app.MapGet("/chat/stream", async (string message, IChatModel model, HttpContext http) =>
{
http.Response.ContentType = "text/event-stream";
await foreach (var chunk in model.StreamChatAsync([Message.User(message)]))
await http.Response.WriteAsync($"data: {chunk}\n\n");
});
Cross-cutting concerns slot into the chain pipeline without modifying chain logic. Implement one interface:
public sealed class LoggingMiddleware<TIn, TOut> : IChainMiddleware<TIn, TOut>
{
public async Task<ChainResult<TOut>> InvokeAsync(
TIn input, ChainContext ctx, ChainDelegate<TIn, TOut> next, CancellationToken ct)
{
Console.WriteLine($"[{ctx.ChainName}] invoking");
var result = await next(input, ctx, ct);
Console.WriteLine($"[{ctx.ChainName}] success={result.IsSuccess}");
return result;
}
}
Built-in: RetryMiddleware (exponential backoff), CacheMiddleware, RateLimitingMiddleware, TracingMiddleware, CostMiddleware, PiiScrubbingMiddleware.
Persist conversation history and enable semantic search with a swappable IMemoryStore / IVectorStore:
builder.Services.AddWeaveLLM()
.AddOpenAI(apiKey)
.AddInMemoryMemory(); // swap for QdrantVectorStore or PostgresVectorStore
// In a handler:
var history = new List<Message>();
await foreach (var entry in memory.GetAsync(sessionId, limit: 20, ct))
history.Add(entry.Message);
history.Add(Message.User(userInput));
var response = await model.ChatAsync(history, cancellationToken: ct);
await memory.AddAsync(
new MemoryEntry(sessionId, Message.Assistant(response.Value!.Content), DateTimeOffset.UtcNow), ct);
Supported vector backends: In-Memory, Qdrant, PostgreSQL + pgvector.
Two strategies ship out of the box, both composable with [LLMTool]-annotated methods:
public sealed class CalculatorTool
{
[LLMTool("calculate", "Evaluate a mathematical expression")]
public string Calculate([Description("e.g. '2 + 2 * 10'")] string expression)
{
var table = new System.Data.DataTable();
return table.Compute(expression, string.Empty)?.ToString() ?? "0";
}
}
builder.Services.AddWeaveLLM()
.AddOpenAI(apiKey)
.AddToolRegistry(r => r.RegisterFromObject(new CalculatorTool()))
.AddReActAgent(maxSteps: 10);
// Inject IAgent and run:
var result = await agent.RunAsync("What is (123 * 456) + 789?", ct);
Console.WriteLine(result.Value!.FinalAnswer);
ReActAgent — Thought → Action → Observation loop until the model emits a Final Answer
PlanAndExecuteAgent — separate planning and execution phases for complex multi-step tasks
| Provider | Chat | Streaming | Embeddings | Notes |
|---|---|---|---|---|
| OpenAI | ✅ | ✅ | ✅ | Default model: gpt-4o |
| Anthropic | ✅ | ✅ | — | Default model: claude-sonnet-4-5 |
| Azure OpenAI | Planned | Planned | Planned | v0.2.0-alpha |
| Ollama | ✅ | ✅ | ✅ | Local inference; default model: llama3 |
| HuggingFace | ✅ | — | ✅ | Inference API; bring your own model ID |
| Sample | Description |
|---|---|
samples/WeaveLLM.Sample.BasicChain | Multi-provider ASP.NET Core API with session memory, SSE streaming, agent endpoint, and RAG index/query routes |
samples/ChatWithDocs | Console RAG chatbot — loads a local docs/ folder, indexes into a vector store, answers questions with source citations |
| Version | Status | Highlights |
|---|---|---|
| v0.1.0-alpha | ✅ Current | Core chains · OpenAI / Anthropic / Ollama / HuggingFace providers · ReAct + PlanAndExecute agents · RAG pipeline · In-Memory / Qdrant / Postgres vector stores · OpenTelemetry · PII scrubbing · cost tracking |
| v0.2.0-alpha | Planned | Azure OpenAI provider · multi-modal (image) input · streaming agents · evaluation suite GA |
| v1.0.0 | Target: Q4 2026 | Stable public API · NuGet stable release · documentation site |
Add the testing helper package to your test project — no real API keys or network calls required:
dotnet add package WeaveLLM.Testing
FakeStreamingChatModel is a fully configurable in-process test double for IChatModel.
Set Tokens for streaming assertions, BlockingResponse for ChatAsync, and
ErrorAfterTokens to simulate mid-stream failures:
using WeaveLLM.Testing;
var fake = new FakeStreamingChatModel
{
Tokens = ["Hello", ", ", "world", "!"],
BlockingResponse = "Hello, world!"
};
// Blocking path
var result = await fake.ChatAsync([Message.User("Hi")]);
result.IsSuccess.Should().BeTrue();
result.Value!.Content.Should().Be("Hello, world!");
// Streaming path
var tokens = new List<string>();
await foreach (var chunk in fake.StreamChatAsync([Message.User("Hi")]))
tokens.Add(chunk);
tokens.Should().Equal("Hello", ", ", "world", "!");
// Error injection — yields 2 tokens then a failure
fake.ErrorAfterTokens = WeaveLLMError.ProviderError("fake", "Quota exceeded");
fake.TokensBeforeError = 2;
var results = new List<ChainResult<string>>();
await foreach (var item in fake.StreamChatSafeAsync([Message.User("Hi")]))
results.Add(item);
results[2].IsFailure.Should().BeTrue();
results[2].Error!.Code.Should().Be("PROVIDER_ERROR");
When you need a mock rather than a fake, use NSubstitute with a static async-iterator
helper. The [EnumeratorCancellation] attribute on the CancellationToken parameter
is required — without it, the token passed via .WithCancellation(ct) is silently
ignored and cancellation tests never fire:
using NSubstitute;
using System.Runtime.CompilerServices;
var model = Substitute.For<IChatModel>();
model.StreamChatAsync(Arg.Any<IReadOnlyList<Message>>(), Arg.Any<LLMOptions?>(), Arg.Any<CancellationToken>())
.Returns(FakeStream("Hello", ", ", "world"));
static async IAsyncEnumerable<string> FakeStream(
params string[] tokens,
[EnumeratorCancellation] CancellationToken ct = default) // ← required
{
foreach (var token in tokens)
{
ct.ThrowIfCancellationRequested();
yield return token;
await Task.Yield();
}
}
Contributions are welcome. Please read CONTRIBUTING.md for branch conventions, coding guidelines, and the pull-request checklist before opening a PR.
WeaveLLM is released under the MIT License.
WeaveLLM — Production-grade AI orchestration for .NET. Strongly typed chains, RAG pipelines, multi-agent graphs, and tool use. The LangChain alternative .NET developers actually deserve.
C#
2
41 commits
updated May 8, 2026
A composable AI orchestration framework for .NET — build LLM chains, RAG pipelines, and autonomous agents with idiomatic C#.
Pain points with LangChain / Semantic Kernel:
WeaveLLM advantages:
IServiceCollection, IAsyncEnumerable, and the BCL; no Python bridgesChainResult<T> instead of throwing; errors are first-class valuesAddWeaveLLM() call wires providers, memory, agents, health checks, and telemetrydotnet add package WeaveLLM.Core
dotnet add package WeaveLLM.Providers # OpenAI, Anthropic, Ollama, HuggingFace
dotnet add package WeaveLLM.Memory # RAG, vector stores, chunking
dotnet add package WeaveLLM.Observability # OpenTelemetry, cost tracking, PII scrubbing
dotnet add package WeaveLLM.Extensions.DependencyInjection # ASP.NET Core integration
using WeaveLLM.Core.Models;
using WeaveLLM.Extensions.DependencyInjection;
var builder = WebApplication.CreateBuilder(args);
builder.Services
.AddWeaveLLM()
.AddOpenAI(apiKey: builder.Configuration["OpenAI:ApiKey"]!, modelId: "gpt-4o")
.AddInMemoryMemory()
.AddToolRegistry()
.AddReActAgent();
var app = builder.Build();
app.MapPost("/chat", async (ChatRequest req, IChatModel model) =>
{
var result = await model.ChatAsync([Message.User(req.Message)]);
return result.IsSuccess
? Results.Ok(result.Value!.Content)
: Results.Problem(result.Error!.Message);
});
app.Run();
record ChatRequest(string Message);
Build type-safe pipelines from reusable IChain<TInput, TOutput> blocks. Wrap any chain with middleware layers using PipelineBuilder:
var pipeline = new PipelineBuilder<IReadOnlyList<Message>, ChatResponse>(chatChain)
.UseMiddleware(new RetryMiddleware<IReadOnlyList<Message>, ChatResponse>(maxRetries: 3))
.UseMiddleware(new CacheMiddleware<IReadOnlyList<Message>, ChatResponse>(ttl: TimeSpan.FromMinutes(5)));
var chain = await pipeline.BuildAsync();
var result = await chain.ExecuteAsync([Message.User("Hello")]);
Every provider supports IAsyncEnumerable<string> streaming out of the box:
app.MapGet("/chat/stream", async (string message, IChatModel model, HttpContext http) =>
{
http.Response.ContentType = "text/event-stream";
await foreach (var chunk in model.StreamChatAsync([Message.User(message)]))
await http.Response.WriteAsync($"data: {chunk}\n\n");
});
Cross-cutting concerns slot into the chain pipeline without modifying chain logic. Implement one interface:
public sealed class LoggingMiddleware<TIn, TOut> : IChainMiddleware<TIn, TOut>
{
public async Task<ChainResult<TOut>> InvokeAsync(
TIn input, ChainContext ctx, ChainDelegate<TIn, TOut> next, CancellationToken ct)
{
Console.WriteLine($"[{ctx.ChainName}] invoking");
var result = await next(input, ctx, ct);
Console.WriteLine($"[{ctx.ChainName}] success={result.IsSuccess}");
return result;
}
}
Built-in: RetryMiddleware (exponential backoff), CacheMiddleware, RateLimitingMiddleware, TracingMiddleware, CostMiddleware, PiiScrubbingMiddleware.
Persist conversation history and enable semantic search with a swappable IMemoryStore / IVectorStore:
builder.Services.AddWeaveLLM()
.AddOpenAI(apiKey)
.AddInMemoryMemory(); // swap for QdrantVectorStore or PostgresVectorStore
// In a handler:
var history = new List<Message>();
await foreach (var entry in memory.GetAsync(sessionId, limit: 20, ct))
history.Add(entry.Message);
history.Add(Message.User(userInput));
var response = await model.ChatAsync(history, cancellationToken: ct);
await memory.AddAsync(
new MemoryEntry(sessionId, Message.Assistant(response.Value!.Content), DateTimeOffset.UtcNow), ct);
Supported vector backends: In-Memory, Qdrant, PostgreSQL + pgvector.
Two strategies ship out of the box, both composable with [LLMTool]-annotated methods:
public sealed class CalculatorTool
{
[LLMTool("calculate", "Evaluate a mathematical expression")]
public string Calculate([Description("e.g. '2 + 2 * 10'")] string expression)
{
var table = new System.Data.DataTable();
return table.Compute(expression, string.Empty)?.ToString() ?? "0";
}
}
builder.Services.AddWeaveLLM()
.AddOpenAI(apiKey)
.AddToolRegistry(r => r.RegisterFromObject(new CalculatorTool()))
.AddReActAgent(maxSteps: 10);
// Inject IAgent and run:
var result = await agent.RunAsync("What is (123 * 456) + 789?", ct);
Console.WriteLine(result.Value!.FinalAnswer);
ReActAgent — Thought → Action → Observation loop until the model emits a Final Answer
PlanAndExecuteAgent — separate planning and execution phases for complex multi-step tasks
| Provider | Chat | Streaming | Embeddings | Notes |
|---|---|---|---|---|
| OpenAI | ✅ | ✅ | ✅ | Default model: gpt-4o |
| Anthropic | ✅ | ✅ | — | Default model: claude-sonnet-4-5 |
| Azure OpenAI | Planned | Planned | Planned | v0.2.0-alpha |
| Ollama | ✅ | ✅ | ✅ | Local inference; default model: llama3 |
| HuggingFace | ✅ | — | ✅ | Inference API; bring your own model ID |
| Sample | Description |
|---|---|
samples/WeaveLLM.Sample.BasicChain | Multi-provider ASP.NET Core API with session memory, SSE streaming, agent endpoint, and RAG index/query routes |
samples/ChatWithDocs | Console RAG chatbot — loads a local docs/ folder, indexes into a vector store, answers questions with source citations |
| Version | Status | Highlights |
|---|---|---|
| v0.1.0-alpha | ✅ Current | Core chains · OpenAI / Anthropic / Ollama / HuggingFace providers · ReAct + PlanAndExecute agents · RAG pipeline · In-Memory / Qdrant / Postgres vector stores · OpenTelemetry · PII scrubbing · cost tracking |
| v0.2.0-alpha | Planned | Azure OpenAI provider · multi-modal (image) input · streaming agents · evaluation suite GA |
| v1.0.0 | Target: Q4 2026 | Stable public API · NuGet stable release · documentation site |
Add the testing helper package to your test project — no real API keys or network calls required:
dotnet add package WeaveLLM.Testing
FakeStreamingChatModel is a fully configurable in-process test double for IChatModel.
Set Tokens for streaming assertions, BlockingResponse for ChatAsync, and
ErrorAfterTokens to simulate mid-stream failures:
using WeaveLLM.Testing;
var fake = new FakeStreamingChatModel
{
Tokens = ["Hello", ", ", "world", "!"],
BlockingResponse = "Hello, world!"
};
// Blocking path
var result = await fake.ChatAsync([Message.User("Hi")]);
result.IsSuccess.Should().BeTrue();
result.Value!.Content.Should().Be("Hello, world!");
// Streaming path
var tokens = new List<string>();
await foreach (var chunk in fake.StreamChatAsync([Message.User("Hi")]))
tokens.Add(chunk);
tokens.Should().Equal("Hello", ", ", "world", "!");
// Error injection — yields 2 tokens then a failure
fake.ErrorAfterTokens = WeaveLLMError.ProviderError("fake", "Quota exceeded");
fake.TokensBeforeError = 2;
var results = new List<ChainResult<string>>();
await foreach (var item in fake.StreamChatSafeAsync([Message.User("Hi")]))
results.Add(item);
results[2].IsFailure.Should().BeTrue();
results[2].Error!.Code.Should().Be("PROVIDER_ERROR");
When you need a mock rather than a fake, use NSubstitute with a static async-iterator
helper. The [EnumeratorCancellation] attribute on the CancellationToken parameter
is required — without it, the token passed via .WithCancellation(ct) is silently
ignored and cancellation tests never fire:
using NSubstitute;
using System.Runtime.CompilerServices;
var model = Substitute.For<IChatModel>();
model.StreamChatAsync(Arg.Any<IReadOnlyList<Message>>(), Arg.Any<LLMOptions?>(), Arg.Any<CancellationToken>())
.Returns(FakeStream("Hello", ", ", "world"));
static async IAsyncEnumerable<string> FakeStream(
params string[] tokens,
[EnumeratorCancellation] CancellationToken ct = default) // ← required
{
foreach (var token in tokens)
{
ct.ThrowIfCancellationRequested();
yield return token;
await Task.Yield();
}
}
Contributions are welcome. Please read CONTRIBUTING.md for branch conventions, coding guidelines, and the pull-request checklist before opening a PR.
WeaveLLM is released under the MIT License.