Unofficial .NET bindings for Google's LiteRT-LM — on-device LLM inference (Gemma) for Windows, Linux, Android, macOS & MAUI
See the codeRun LLMs on-device from any .NET app — Windows, Linux, Android, macOS, MAUI. No server, no cloud.
Documentation · API Reference · Samples · Changelog · Roadmap
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) via P/Invoke over its C API, with native binaries distributed per-RID as NuGet packages (LLamaSharp-style). Status: stable (1.x).
| Platform | Native | NuGet | CPU | GPU | Validated on |
|---|---|---|---|---|---|
| win-x64 | ✅ | ✅ | ✅ | ✅ | real hardware |
| linux-x64 | ✅ | ✅ | ✅ | ✅ | real hardware (GPU last checked with the v0.13.1 libraries) |
| linux-arm64 | ✅ | ⏳ | ✅ | — | CI |
| android-arm64 | ✅ | ✅ | ✅ | ✅ | real device |
| android-x64 (emulator) | ✅ | ⏳ | ✅ | — | emulator |
| osx-arm64 | ✅ | ✅ | ✅ | ✅ | CI |
| ios-arm64 | ✅ | ⏳ | — | — | pending |
CPU / GPU = inference validated on that backend. macOS GPU runs in CI on the WebGPU (Dawn→Metal) delegate; the native Metal delegate ships as a real-hardware fallback. linux-arm64 and android-x64 publish with the next release. The iOS runtime package ships once on-device validation lands.
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
<!-- or LiteRtLmSharp.runtime.linux-x64 / linux-arm64 / android-arm64 / android-x64 (emulator) / osx-arm64, per target -->
Install the managed package plus the runtime package for your platform, always with the same version number. Which LiteRT-LM native build each release wraps:
| LiteRtLmSharp | LiteRT-LM native |
|---|---|
| 1.2.0 | v0.16.0 (Google's official C API prebuilts) |
| 1.1.1 | v0.14.0 |
| 1.1.0 | v0.14.0 |
| 1.0.0 | v0.13.1 |
| 0.1.0-preview.3 | v0.13.1 |
| 0.1.0-preview.2 | v0.13.1 |
| 0.1.0-preview.1 | v0.13.1 |
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu (WebGPU -> D3D12/Vulkan/Metal)
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
// Blocking:
Console.WriteLine(chat.Send("Hello!").Text);
// Awaitable (pass a CancellationToken to cancel mid-generation):
Console.WriteLine((await chat.SendAsync("Hello!")).Text);
// Streaming: each piece is tagged (answer / thinking / tool call):
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
Every feature has a full guide on the documentation site:
| Chat | Blocking, awaitable + cancellable, and streaming sends — guide |
| Function calling | Real tool calls with constrained decoding for reliable JSON arguments — guide |
| Reasoning mode | Gemma "thinking", surfaced separately from the answer (blocking & streaming) — guide |
| Multimodal | Image and audio attachments, in-memory or memory-mapped from disk — guide |
| Conversation state | Persist/restore chats across restarts; clone a live conversation to branch it — guide |
| Embeddings | On-device text embeddings (EmbeddingGemma 2) for semantic search and RAG, next to a chat model — guide |
| Model metadata | Read a model file's type, context size, inputs and backends without loading it — guide |
| Token counting | Tokenize/detokenize with the model's own tokenizer; budget the context window — guide |
| Speculative decoding | MTP drafter support plus a built-in benchmark API (tok/s, TTFT) — guide |
| Engine tuning | Activation precision, prefill chunking, thread counts, cache control — guide |
| .NET AI ecosystem | IChatClient, IEmbeddingGenerator + Semantic Kernel connectors (below) |
| AOT & trimming | Source-generated P/Invoke, no runtime marshalling — Native AOT compatible |
Two optional companion packages plug the on-device model into the .NET AI ecosystem:
| Package | Exposes the model as | Works with |
|---|---|---|
LiteRtLmSharp.Extensions.AI | Microsoft.Extensions.AI.IChatClient, IEmbeddingGenerator | Microsoft Agent Framework, Semantic Kernel, plain MEAI |
LiteRtLmSharp.SemanticKernel | IChatCompletionService, embedding generator | Semantic Kernel |
using IChatClient client = new LiteRtChatClient(engine);
Console.WriteLine((await client.GetResponseAsync("One upbeat sentence about on-device AI.")).Text);
// var agent = new ChatClientAgent(client, "You are helpful."); // Microsoft Agent Framework
Function calling (auto-invocation), reasoning content, multimodal and opt-in stateful
conversations (MEAI ConversationId) are supported — see the
Extensions.AI guide and the
Semantic Kernel guide.
samples/Maui — full
Android/Windows chat app: model download with resume, streaming, multimodal attachments, function
calling against real device APIs, speculative-decoding and reasoning toggles.samples/Console —
chat loop with --tools, --spec and --thinking demos.samples/SemanticKernel —
kernel registration, streaming, and a [KernelFunction] plugin.LiteRtEngine while one is alive throws (it
would hang in the native layer). To switch model or backend, dispose the conversations and the
engine, then LiteRtEngine.Load again — same pattern as Google's Edge Gallery. An embedding engine
(LiteRtEmbeddingEngine) does not count, so a chat model and an embedding model can stay loaded together.MaxNumTokens is the total context window (prompt + response, across turns). Use >= 1024;
too small can make blocking generation return nothing.libvulkan1). The official library has no hard dependency on the Vulkan loader. The GPU backend runs on Vulkan: install your GPU's Vulkan driver and
the Vulkan loader (libvulkan1 on Debian/Ubuntu).dxcompiler.dll, dxil.dll) ships in the runtime package.<uses-native-library>; without libOpenCL.so the engine silently picks a
Vulkan path that produces garbage on older Adreno drivers. Copy the <uses-native-library> block
from the MAUI sample's AndroidManifest.
Full diagnosis in the Android guide..NET 10 is the current LTS. Targeting it exclusively lets the binding use the modern interop stack
as designed — source-generated P/Invoke ([LibraryImport]) and [UnmanagedCallersOnly] callbacks
with no runtime marshalling, which is what makes it AOT- and trim-compatible — and a single
net10.0 TFM is directly consumable from net10.0-android/-ios/-windows MAUI apps without
multi-targeting.
Native binaries are not committed — restore them into runtimes/<rid>/native/ from the
native-v* GitHub release, then build:
pwsh scripts/restore-natives.ps1 # -All for every RID, -Rid android-arm64 for one
dotnet build LiteRtLmSharp.slnx && dotnet test
The samples have their own solution (samples/LiteRtLmSharp.Samples.slnx; the MAUI sample needs
dotnet workload install maui). To run the model-backed tests, point LITERTLM_TEST_MODEL at a
.litertlm file. CI: native-release.yml
repackages Google's official LiteRT-LM C API prebuilts for a pinned release (verified against the
sha256 digests GitHub or PyPI publish, each library inspected) into the native-v* release,
pack-nuget.yml
packs and publishes. Internals docs:
native ABI,
native build,
packaging.
Issues and PRs are welcome — see CONTRIBUTING.md for the dev setup and guidelines. Please open an issue first for anything beyond a small fix, and use Discussions for questions.
Apache-2.0 (see LICENSE.txt, NOTICE and THIRD-PARTY-NOTICES.md).
This is an unofficial, community-maintained project. It is not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are Google's official LiteRT-LM prebuilts (Apache-2.0) at pinned release tags.
Unofficial .NET bindings for Google's LiteRT-LM — on-device LLM inference (Gemma) for Windows, Linux, Android, macOS & MAUI
See the codeRun LLMs on-device from any .NET app — Windows, Linux, Android, macOS, MAUI. No server, no cloud.
Documentation · API Reference · Samples · Changelog · Roadmap
.NET 10 bindings for Google's LiteRT-LM — on-device LLM inference (e.g. Gemma) via P/Invoke over its C API, with native binaries distributed per-RID as NuGet packages (LLamaSharp-style). Status: stable (1.x).
| Platform | Native | NuGet | CPU | GPU | Validated on |
|---|---|---|---|---|---|
| win-x64 | ✅ | ✅ | ✅ | ✅ | real hardware |
| linux-x64 | ✅ | ✅ | ✅ | ✅ | real hardware (GPU last checked with the v0.13.1 libraries) |
| linux-arm64 | ✅ | ⏳ | ✅ | — | CI |
| android-arm64 | ✅ | ✅ | ✅ | ✅ | real device |
| android-x64 (emulator) | ✅ | ⏳ | ✅ | — | emulator |
| osx-arm64 | ✅ | ✅ | ✅ | ✅ | CI |
| ios-arm64 | ✅ | ⏳ | — | — | pending |
CPU / GPU = inference validated on that backend. macOS GPU runs in CI on the WebGPU (Dawn→Metal) delegate; the native Metal delegate ships as a real-hardware fallback. linux-arm64 and android-x64 publish with the next release. The iOS runtime package ships once on-device validation lands.
<PackageReference Include="LiteRtLmSharp" Version="1.2.0" />
<PackageReference Include="LiteRtLmSharp.runtime.win-x64" Version="1.2.0" />
<!-- or LiteRtLmSharp.runtime.linux-x64 / linux-arm64 / android-arm64 / android-x64 (emulator) / osx-arm64, per target -->
Install the managed package plus the runtime package for your platform, always with the same version number. Which LiteRT-LM native build each release wraps:
| LiteRtLmSharp | LiteRT-LM native |
|---|---|
| 1.2.0 | v0.16.0 (Google's official C API prebuilts) |
| 1.1.1 | v0.14.0 |
| 1.1.0 | v0.14.0 |
| 1.0.0 | v0.13.1 |
| 0.1.0-preview.3 | v0.13.1 |
| 0.1.0-preview.2 | v0.13.1 |
| 0.1.0-preview.1 | v0.13.1 |
using LiteRtLmSharp;
using var engine = LiteRtEngine.Load(new LiteRtEngineOptions
{
ModelPath = "gemma-4-E2B-it.litertlm", // from huggingface.co/litert-community
Backend = LiteRtBackend.Cpu, // or .Gpu (WebGPU -> D3D12/Vulkan/Metal)
MaxNumTokens = 4096, // total context window
});
using var chat = engine.CreateConversation();
// Blocking:
Console.WriteLine(chat.Send("Hello!").Text);
// Awaitable (pass a CancellationToken to cancel mid-generation):
Console.WriteLine((await chat.SendAsync("Hello!")).Text);
// Streaming: each piece is tagged (answer / thinking / tool call):
await foreach (var chunk in chat.SendStreamingAsync("Tell me a joke"))
Console.Write(chunk.Text);
Every feature has a full guide on the documentation site:
| Chat | Blocking, awaitable + cancellable, and streaming sends — guide |
| Function calling | Real tool calls with constrained decoding for reliable JSON arguments — guide |
| Reasoning mode | Gemma "thinking", surfaced separately from the answer (blocking & streaming) — guide |
| Multimodal | Image and audio attachments, in-memory or memory-mapped from disk — guide |
| Conversation state | Persist/restore chats across restarts; clone a live conversation to branch it — guide |
| Embeddings | On-device text embeddings (EmbeddingGemma 2) for semantic search and RAG, next to a chat model — guide |
| Model metadata | Read a model file's type, context size, inputs and backends without loading it — guide |
| Token counting | Tokenize/detokenize with the model's own tokenizer; budget the context window — guide |
| Speculative decoding | MTP drafter support plus a built-in benchmark API (tok/s, TTFT) — guide |
| Engine tuning | Activation precision, prefill chunking, thread counts, cache control — guide |
| .NET AI ecosystem | IChatClient, IEmbeddingGenerator + Semantic Kernel connectors (below) |
| AOT & trimming | Source-generated P/Invoke, no runtime marshalling — Native AOT compatible |
Two optional companion packages plug the on-device model into the .NET AI ecosystem:
| Package | Exposes the model as | Works with |
|---|---|---|
LiteRtLmSharp.Extensions.AI | Microsoft.Extensions.AI.IChatClient, IEmbeddingGenerator | Microsoft Agent Framework, Semantic Kernel, plain MEAI |
LiteRtLmSharp.SemanticKernel | IChatCompletionService, embedding generator | Semantic Kernel |
using IChatClient client = new LiteRtChatClient(engine);
Console.WriteLine((await client.GetResponseAsync("One upbeat sentence about on-device AI.")).Text);
// var agent = new ChatClientAgent(client, "You are helpful."); // Microsoft Agent Framework
Function calling (auto-invocation), reasoning content, multimodal and opt-in stateful
conversations (MEAI ConversationId) are supported — see the
Extensions.AI guide and the
Semantic Kernel guide.
samples/Maui — full
Android/Windows chat app: model download with resume, streaming, multimodal attachments, function
calling against real device APIs, speculative-decoding and reasoning toggles.samples/Console —
chat loop with --tools, --spec and --thinking demos.samples/SemanticKernel —
kernel registration, streaming, and a [KernelFunction] plugin.LiteRtEngine while one is alive throws (it
would hang in the native layer). To switch model or backend, dispose the conversations and the
engine, then LiteRtEngine.Load again — same pattern as Google's Edge Gallery. An embedding engine
(LiteRtEmbeddingEngine) does not count, so a chat model and an embedding model can stay loaded together.MaxNumTokens is the total context window (prompt + response, across turns). Use >= 1024;
too small can make blocking generation return nothing.libvulkan1). The official library has no hard dependency on the Vulkan loader. The GPU backend runs on Vulkan: install your GPU's Vulkan driver and
the Vulkan loader (libvulkan1 on Debian/Ubuntu).dxcompiler.dll, dxil.dll) ships in the runtime package.<uses-native-library>; without libOpenCL.so the engine silently picks a
Vulkan path that produces garbage on older Adreno drivers. Copy the <uses-native-library> block
from the MAUI sample's AndroidManifest.
Full diagnosis in the Android guide..NET 10 is the current LTS. Targeting it exclusively lets the binding use the modern interop stack
as designed — source-generated P/Invoke ([LibraryImport]) and [UnmanagedCallersOnly] callbacks
with no runtime marshalling, which is what makes it AOT- and trim-compatible — and a single
net10.0 TFM is directly consumable from net10.0-android/-ios/-windows MAUI apps without
multi-targeting.
Native binaries are not committed — restore them into runtimes/<rid>/native/ from the
native-v* GitHub release, then build:
pwsh scripts/restore-natives.ps1 # -All for every RID, -Rid android-arm64 for one
dotnet build LiteRtLmSharp.slnx && dotnet test
The samples have their own solution (samples/LiteRtLmSharp.Samples.slnx; the MAUI sample needs
dotnet workload install maui). To run the model-backed tests, point LITERTLM_TEST_MODEL at a
.litertlm file. CI: native-release.yml
repackages Google's official LiteRT-LM C API prebuilts for a pinned release (verified against the
sha256 digests GitHub or PyPI publish, each library inspected) into the native-v* release,
pack-nuget.yml
packs and publishes. Internals docs:
native ABI,
native build,
packaging.
Issues and PRs are welcome — see CONTRIBUTING.md for the dev setup and guidelines. Please open an issue first for anything beyond a small fix, and use Discussions for questions.
Apache-2.0 (see LICENSE.txt, NOTICE and THIRD-PARTY-NOTICES.md).
This is an unofficial, community-maintained project. It is not affiliated with, sponsored, or endorsed by Google. LiteRT, LiteRT-LM and Gemma are trademarks of Google LLC. The native binaries are Google's official LiteRT-LM prebuilts (Apache-2.0) at pinned release tags.