OrihuelaConde/LiteRtLmSharp

Unofficial .NET bindings for Google's LiteRT-LM — on-device LLM inference (Gemma) for Windows, Linux, Android, macOS & MAUI

C#

11

203 commits

updated Oct 7, 2026

See the code

README

LiteRtLmSharp

Run LLMs on-device from any .NET app — Windows, Linux, Android, macOS, MAUI. No server, no cloud.

CI NuGet Downloads Docs License

Documentation · API Reference · Samples · Changelog · Roadmap

Streaming chat fully on-device in the MAUI sample (Android, GPU backend)

.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) via P/Invoke over its C API, with native binaries distributed per-RID as NuGet packages (LLamaSharp-style). Status: stable (1.x).

PlatformNativeNuGetCPUGPUValidated on
win-x64✅✅✅✅real hardware
linux-x64✅✅✅✅real hardware (GPU last checked with the v0.13.1 libraries)
linux-arm64✅⏳✅—CI
android-arm64✅✅✅✅real device
android-x64 (emulator)✅⏳✅—emulator
osx-arm64✅✅✅✅CI
ios-arm64✅⏳——pending

CPU / GPU = inference validated on that backend. macOS GPU runs in CI on the WebGPU (Dawn→Metal) delegate; the native Metal delegate ships as a real-hardware fallback. linux-arm64 and android-x64 publish with the next release. The iOS runtime package ships once on-device validation lands.

Quick start

<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
<!-- or LiteRtLmSharp.runtime.linux-x64 / linux-arm64 / android-arm64 / android-x64 (emulator) / osx-arm64, per target -->

Install the managed package plus the runtime package for your platform, always with the same version number. Which LiteRT-LM native build each release wraps:

LiteRtLmSharpLiteRT-LM native
1.2.0v0.16.0 (Google's official C API prebuilts)
1.1.1v0.14.0
1.1.0v0.14.0
1.0.0v0.13.1
0.1.0-preview.3v0.13.1
0.1.0-preview.2v0.13.1
0.1.0-preview.1v0.13.1
using LiteRtLmSharp;

using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm",   // from huggingface.co/litert-community
    Backend = LiteRtBackend.Cpu,              // or .Gpu (WebGPU -> D3D12/Vulkan/Metal)
    MaxNumTokens = 4096,                       // total context window
});

using var chat = engine.CreateConversation();

// Blocking:
Console.WriteLine(chat.Send("Hello!").Text);

// Awaitable (pass a CancellationToken to cancel mid-generation):
Console.WriteLine((await chat.SendAsync("Hello!")).Text);

// Streaming: each piece is tagged (answer / thinking / tool call):
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
    Console.Write(chunk.Text);

Features

Every feature has a full guide on the documentation site:

ChatBlocking, awaitable + cancellable, and streaming sends — guide
Function callingReal tool calls with constrained decoding for reliable JSON arguments — guide
Reasoning modeGemma "thinking", surfaced separately from the answer (blocking & streaming) — guide
MultimodalImage and audio attachments, in-memory or memory-mapped from disk — guide
Conversation statePersist/restore chats across restarts; clone a live conversation to branch it — guide
EmbeddingsOn-device text embeddings (EmbeddingGemma 2) for semantic search and RAG, next to a chat model — guide
Model metadataRead a model file's type, context size, inputs and backends without loading it — guide
Token countingTokenize/detokenize with the model's own tokenizer; budget the context window — guide
Speculative decodingMTP drafter support plus a built-in benchmark API (tok/s, TTFT) — guide
Engine tuningActivation precision, prefill chunking, thread counts, cache control — guide
.NET AI ecosystemIChatClient, IEmbeddingGenerator + Semantic Kernel connectors (below)
AOT & trimmingSource-generated P/Invoke, no runtime marshalling — Native AOT compatible

.NET AI integrations

Two optional companion packages plug the on-device model into the .NET AI ecosystem:

PackageExposes the model asWorks with
LiteRtLmSharp.Extensions.AIMicrosoft.Extensions.AI.IChatClient, IEmbeddingGeneratorMicrosoft Agent Framework, Semantic Kernel, plain MEAI
LiteRtLmSharp.SemanticKernelIChatCompletionService, embedding generatorSemantic Kernel
using IChatClient client = new LiteRtChatClient(engine);
Console.WriteLine((await client.GetResponseAsync("One upbeat sentence about on-device AI.")).Text);
// var agent = new ChatClientAgent(client, "You are helpful.");   // Microsoft Agent Framework

Function calling (auto-invocation), reasoning content, multimodal and opt-in stateful conversations (MEAI ConversationId) are supported — see the Extensions.AI guide and the Semantic Kernel guide.

Samples

Models tab — download and manage Gemma models Chat tab — multimodal: image attachment answered on-device Tools tab — on-device function calling against real device APIs

  • samples/Maui — full Android/Windows chat app: model download with resume, streaming, multimodal attachments, function calling against real device APIs, speculative-decoding and reasoning toggles.
  • samples/Console — chat loop with --tools, --spec and --thinking demos.
  • samples/SemanticKernel — kernel registration, streaming, and a [KernelFunction] plugin.

Important notes

  • One chat engine alive at a time. Loading a second LiteRtEngine while one is alive throws (it would hang in the native layer). To switch model or backend, dispose the conversations and the engine, then LiteRtEngine.Load again — same pattern as Google's Edge Gallery. An embedding engine (LiteRtEmbeddingEngine) does not count, so a chat model and an embedding model can stay loaded together.
  • MaxNumTokens is the total context window (prompt + response, across turns). Use >= 1024; too small can make blocking generation return nothing.
  • Conversations are not thread-safe — serialize sends per engine (the Microsoft.Extensions.AI client does this for you).
  • Linux needs no extra system package for the CPU backend since 1.3.0 (1.2.0's library needed libvulkan1). The official library has no hard dependency on the Vulkan loader. The GPU backend runs on Vulkan: install your GPU's Vulkan driver and the Vulkan loader (libvulkan1 on Debian/Ubuntu).
  • win-x64 needs no Visual C++ Redistributable since 1.2.0 (the official library links the CRT statically). The GPU backend's shader compiler (dxcompiler.dll, dxil.dll) ships in the runtime package.
  • Android GPU needs manifest declarations. Android 12+ only grants access to vendor native libraries declared via <uses-native-library>; without libOpenCL.so the engine silently picks a Vulkan path that produces garbage on older Adreno drivers. Copy the <uses-native-library> block from the MAUI sample's AndroidManifest. Full diagnosis in the Android guide.

Why .NET 10 only?

.NET 10 is the current LTS. Targeting it exclusively lets the binding use the modern interop stack as designed — source-generated P/Invoke ([LibraryImport]) and [UnmanagedCallersOnly] callbacks with no runtime marshalling, which is what makes it AOT- and trim-compatible — and a single net10.0 TFM is directly consumable from net10.0-android/-ios/-windows MAUI apps without multi-targeting.

Building from source

Native binaries are not committed — restore them into runtimes/<rid>/native/ from the native-v* GitHub release, then build:

pwsh scripts/restore-natives.ps1    # -All for every RID, -Rid android-arm64 for one
dotnet build LiteRtLmSharp.slnx && dotnet test

The samples have their own solution (samples/LiteRtLmSharp.Samples.slnx; the MAUI sample needs dotnet workload install maui). To run the model-backed tests, point LITERTLM_TEST_MODEL at a .litertlm file. CI: native-release.yml repackages Google's official LiteRT-LM C API prebuilts for a pinned release (verified against the sha256 digests GitHub or PyPI publish, each library inspected) into the native-v* release, pack-nuget.yml packs and publishes. Internals docs: native ABI, native build, packaging.

Contributing

Issues and PRs are welcome — see CONTRIBUTING.md for the dev setup and guidelines. Please open an issue first for anything beyond a small fix, and use Discussions for questions.

License and trademarks

Apache-2.0 (see LICENSE.txt, NOTICE and THIRD-PARTY-NOTICES.md).

This is an unofficial, community-maintained project. It is not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are Google's official LiteRT-LM prebuilts (Apache-2.0) at pinned release tags.

ai
csharp
dotnet
edge-ai
gemma
inference
litert
litert-lm
llm
local-llm
maui
nuget
on-device
on-device-ai
semantic-kernel

OrihuelaConde/LiteRtLmSharp

Unofficial .NET bindings for Google's LiteRT-LM — on-device LLM inference (Gemma) for Windows, Linux, Android, macOS & MAUI

C#

11

203 commits

updated Oct 7, 2026

See the code

README

LiteRtLmSharp

Run LLMs on-device from any .NET app — Windows, Linux, Android, macOS, MAUI. No server, no cloud.

CI NuGet Downloads Docs License

Documentation · API Reference · Samples · Changelog · Roadmap

Streaming chat fully on-device in the MAUI sample (Android, GPU backend)

.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) via P/Invoke over its C API, with native binaries distributed per-RID as NuGet packages (LLamaSharp-style). Status: stable (1.x).

PlatformNativeNuGetCPUGPUValidated on
win-x64✅✅✅✅real hardware
linux-x64✅✅✅✅real hardware (GPU last checked with the v0.13.1 libraries)
linux-arm64✅⏳✅—CI
android-arm64✅✅✅✅real device
android-x64 (emulator)✅⏳✅—emulator
osx-arm64✅✅✅✅CI
ios-arm64✅⏳——pending

CPU / GPU = inference validated on that backend. macOS GPU runs in CI on the WebGPU (Dawn→Metal) delegate; the native Metal delegate ships as a real-hardware fallback. linux-arm64 and android-x64 publish with the next release. The iOS runtime package ships once on-device validation lands.

Quick start

<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
<!-- or LiteRtLmSharp.runtime.linux-x64 / linux-arm64 / android-arm64 / android-x64 (emulator) / osx-arm64, per target -->

Install the managed package plus the runtime package for your platform, always with the same version number. Which LiteRT-LM native build each release wraps:

LiteRtLmSharpLiteRT-LM native
1.2.0v0.16.0 (Google's official C API prebuilts)
1.1.1v0.14.0
1.1.0v0.14.0
1.0.0v0.13.1
0.1.0-preview.3v0.13.1
0.1.0-preview.2v0.13.1
0.1.0-preview.1v0.13.1
using LiteRtLmSharp;

using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
    ModelPath = "gemma-4-E2B-it.litertlm",   // from huggingface.co/litert-community
    Backend = LiteRtBackend.Cpu,              // or .Gpu (WebGPU -> D3D12/Vulkan/Metal)
    MaxNumTokens = 4096,                       // total context window
});

using var chat = engine.CreateConversation();

// Blocking:
Console.WriteLine(chat.Send("Hello!").Text);

// Awaitable (pass a CancellationToken to cancel mid-generation):
Console.WriteLine((await chat.SendAsync("Hello!")).Text);

// Streaming: each piece is tagged (answer / thinking / tool call):
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
    Console.Write(chunk.Text);

Features

Every feature has a full guide on the documentation site:

ChatBlocking, awaitable + cancellable, and streaming sends — guide
Function callingReal tool calls with constrained decoding for reliable JSON arguments — guide
Reasoning modeGemma "thinking", surfaced separately from the answer (blocking & streaming) — guide
MultimodalImage and audio attachments, in-memory or memory-mapped from disk — guide
Conversation statePersist/restore chats across restarts; clone a live conversation to branch it — guide
EmbeddingsOn-device text embeddings (EmbeddingGemma 2) for semantic search and RAG, next to a chat model — guide
Model metadataRead a model file's type, context size, inputs and backends without loading it — guide
Token countingTokenize/detokenize with the model's own tokenizer; budget the context window — guide
Speculative decodingMTP drafter support plus a built-in benchmark API (tok/s, TTFT) — guide
Engine tuningActivation precision, prefill chunking, thread counts, cache control — guide
.NET AI ecosystemIChatClient, IEmbeddingGenerator + Semantic Kernel connectors (below)
AOT & trimmingSource-generated P/Invoke, no runtime marshalling — Native AOT compatible

.NET AI integrations

Two optional companion packages plug the on-device model into the .NET AI ecosystem:

PackageExposes the model asWorks with
LiteRtLmSharp.Extensions.AIMicrosoft.Extensions.AI.IChatClient, IEmbeddingGeneratorMicrosoft Agent Framework, Semantic Kernel, plain MEAI
LiteRtLmSharp.SemanticKernelIChatCompletionService, embedding generatorSemantic Kernel
using IChatClient client = new LiteRtChatClient(engine);
Console.WriteLine((await client.GetResponseAsync("One upbeat sentence about on-device AI.")).Text);
// var agent = new ChatClientAgent(client, "You are helpful.");   // Microsoft Agent Framework

Function calling (auto-invocation), reasoning content, multimodal and opt-in stateful conversations (MEAI ConversationId) are supported — see the Extensions.AI guide and the Semantic Kernel guide.

Samples

Models tab — download and manage Gemma models Chat tab — multimodal: image attachment answered on-device Tools tab — on-device function calling against real device APIs

  • samples/Maui — full Android/Windows chat app: model download with resume, streaming, multimodal attachments, function calling against real device APIs, speculative-decoding and reasoning toggles.
  • samples/Console — chat loop with --tools, --spec and --thinking demos.
  • samples/SemanticKernel — kernel registration, streaming, and a [KernelFunction] plugin.

Important notes

  • One chat engine alive at a time. Loading a second LiteRtEngine while one is alive throws (it would hang in the native layer). To switch model or backend, dispose the conversations and the engine, then LiteRtEngine.Load again — same pattern as Google's Edge Gallery. An embedding engine (LiteRtEmbeddingEngine) does not count, so a chat model and an embedding model can stay loaded together.
  • MaxNumTokens is the total context window (prompt + response, across turns). Use >= 1024; too small can make blocking generation return nothing.
  • Conversations are not thread-safe — serialize sends per engine (the Microsoft.Extensions.AI client does this for you).
  • Linux needs no extra system package for the CPU backend since 1.3.0 (1.2.0's library needed libvulkan1). The official library has no hard dependency on the Vulkan loader. The GPU backend runs on Vulkan: install your GPU's Vulkan driver and the Vulkan loader (libvulkan1 on Debian/Ubuntu).
  • win-x64 needs no Visual C++ Redistributable since 1.2.0 (the official library links the CRT statically). The GPU backend's shader compiler (dxcompiler.dll, dxil.dll) ships in the runtime package.
  • Android GPU needs manifest declarations. Android 12+ only grants access to vendor native libraries declared via <uses-native-library>; without libOpenCL.so the engine silently picks a Vulkan path that produces garbage on older Adreno drivers. Copy the <uses-native-library> block from the MAUI sample's AndroidManifest. Full diagnosis in the Android guide.

Why .NET 10 only?

.NET 10 is the current LTS. Targeting it exclusively lets the binding use the modern interop stack as designed — source-generated P/Invoke ([LibraryImport]) and [UnmanagedCallersOnly] callbacks with no runtime marshalling, which is what makes it AOT- and trim-compatible — and a single net10.0 TFM is directly consumable from net10.0-android/-ios/-windows MAUI apps without multi-targeting.

Building from source

Native binaries are not committed — restore them into runtimes/<rid>/native/ from the native-v* GitHub release, then build:

pwsh scripts/restore-natives.ps1    # -All for every RID, -Rid android-arm64 for one
dotnet build LiteRtLmSharp.slnx && dotnet test

The samples have their own solution (samples/LiteRtLmSharp.Samples.slnx; the MAUI sample needs dotnet workload install maui). To run the model-backed tests, point LITERTLM_TEST_MODEL at a .litertlm file. CI: native-release.yml repackages Google's official LiteRT-LM C API prebuilts for a pinned release (verified against the sha256 digests GitHub or PyPI publish, each library inspected) into the native-v* release, pack-nuget.yml packs and publishes. Internals docs: native ABI, native build, packaging.

Contributing

Issues and PRs are welcome — see CONTRIBUTING.md for the dev setup and guidelines. Please open an issue first for anything beyond a small fix, and use Discussions for questions.

License and trademarks

Apache-2.0 (see LICENSE.txt, NOTICE and THIRD-PARTY-NOTICES.md).

This is an unofficial, community-maintained project. It is not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are Google's official LiteRT-LM prebuilts (Apache-2.0) at pinned release tags.

ai
csharp
dotnet
edge-ai
gemma
inference
litert
litert-lm
llm
local-llm
maui
nuget
on-device
on-device-ai
semantic-kernel