zhongkaifu/TensorSharp

A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability

C#

513

975 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop (r/SideProject)

I’ve been building an open-source side project called **TensorSharp**, a local LLM inference engine and agent runtime written in C#/.NET: [TensorSharp on GitHub](https://github.com/zhongkaifu/TensorSharp?utm_source=chatgpt.com) One of the problems I’ve been experimenting with is: **How far can we…

2

Oct 3, 2026

Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD (r/LocalLLM)

How much hardware do you actually need to run a **176B-parameter Qwen3.8 Flash Next** locally? Turns out, this can be enough: **RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD.** Yes — a **176B MoE model on a laptop**. I’ve been working on this in **TensorSharp**, my open-source local LLM…

3

Oct 3, 2026

Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD (r/LocalLLaMA)

I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model. Turns out, **Qwen3.8 Flash Next 176B** can run on: **RTX 3080 Laptop — 16GB VRAM** **32GB system RAM** **SSD** No 128GB/256GB RAM workstation and no multi-GPU setup. I’m running it with **TensorSharp**, my…

9

Oct 3, 2026

README

TensorSharp

TensorSharp logo

English | 中文

TensorSharp is a .NET 10 inference engine for local GGUF models. Run it on Windows, macOS or Linux through the CLI, browser chat, or Ollama/OpenAI-compatible APIs, or embed it in your own .NET application. Use managed C# CPU kernels or native CUDA, Metal and Vulkan backends, with support that varies by model.

Current source covers text and reasoning, multimodal input, embeddings, image generation/editing, video with audio, and Agent Skills with code tools. TensorSharp also powers TensorAgent, the local app for iPhone, iPad, Mac and Windows. Start below, or check the project status for capabilities and validation limits; source changes may be ahead of published packages.

Highlights

  • Local models, several interfaces. One engine for CLI use, browser chat and Ollama/OpenAI-compatible APIs, across managed CPU and native accelerator backends.
  • Text and multimodal models. Dense and MoE GGUF models, reasoning, image/audio input and document questions. See Supported models for each family's capabilities.
  • Embeddings and media generation. Text/code embeddings, Qwen-Image-2.1 generation and masked editing, and video models including MiniMax-H3 with audio.
  • Efficient serving. Continuous batching and a paged, Radix prefix-shared KV cache are on by default. Speculative decoding and multi-GPU placement are available for supported models. See Features.
  • Agent Skills and code tools. TensorSharp.AgentHost adds file, shell and document workflows. Server and TensorAgent chats also support bounded sub-agent delegation with private workspaces and read-only defaults.
  • Recorded comparisons. Benchmarks against llama.cpp use identical GGUF files and hardware; results apply to the measured model, backend and workload. See Benchmarks.
  • TensorAgent apps. Local chat, attachments, skills and saved work on phones and desktops. See the Mac/Windows installation guide and app coverage notes.

Quick Start

TensorSharp CLI and server

The Releases page provides self-contained CLI and Server archives for Windows x64 (CPU/CUDA), Linux x64 (CPU/CUDA), and macOS arm64.

To build from source you need the full .NET 10 SDK (how to install it), git, curl, CMake 3.20+, and the toolchain for your GPU. Then run the verified Gemma 4 E4B model (7.48 GiB). On Windows with an NVIDIA GPU (PowerShell):

git clone https://github.com/zhongkaifu/TensorSharp.git; Set-Location TensorSharp
New-Item -ItemType Directory -Force models | Out-Null
curl.exe -L --fail "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models\gemma-4-E4B-it-Q8_0.gguf
'Answer in one short sentence: what is TensorSharp?' | Set-Content prompt.txt
$env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA = 'ON'
dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models\gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cuda

On other machines, change the backend (see Pick a backend):

  • macOS (Apple Silicon): drop the CUDA environment variable and use --backend ggml_metal.
  • Linux + NVIDIA: prefix the dotnet run with TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON and use --backend ggml_cuda.
  • AMD / Intel / NVIDIA Vulkan: set TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON and use --backend ggml_vulkan.

Host the same model as a server: a browser chat at http://localhost:5000 plus Ollama- and OpenAI-compatible APIs.

dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 512

The server listens on 0.0.0.0:5000 with no built-in authentication or TLS; keep it behind a firewall or an authenticated HTTPS reverse proxy.

Pick a backend

Your hardwareBackend
Apple Silicon (Mac)--backend ggml_metal
Windows / Linux + NVIDIA GPU--backend ggml_cuda
Windows / Linux + AMD / Intel / NVIDIA GPU--backend ggml_vulkan
No GPU--backend ggml_cpu (native kernels), or --backend cpu (pure C#, no native dependencies)

The Getting started guide has the rest: installing the SDK on each platform, multi-GPU and multi-node runs, NVIDIA DGX Spark, multimodal input, embeddings, and making it fast. Every option is in the CLI and Server references, and both programs print them with --help.

dotnet build TensorSharp.slnx also builds TensorAgent's available desktop heads and the iOS simulator head on Apple Silicon when the selected SDK has the required MAUI workloads and staged native/Python files. Missing prerequisites skip the affected app head with a warning; see TensorAgent build instructions.

TensorAgent Desktop: download, install, chat

Download the latest release and expand Assets. Choose a file beginning with tensoragent-desktop-:

PlatformDownload and install
macOS 14+ on Apple Silicontensoragent-desktop-<version>-osx-arm64.dmg: open and drag TensorAgent to Applications. A PKG installer and ZIP are also available.
Windows x64tensoragent-desktop-<version>-win-x64-cpu.msi, or win-x64-cuda.msi for a compatible NVIDIA GPU/driver: install and open TensorAgent from Start. ZIPs are also available.

The app includes its .NET runtime and native engine. Windows needs WebView2 Evergreen Runtime if missing; Python/Node are optional tools for skills. The current built-in catalog needs at least 12 GB system RAM, and model weights download separately. Open ☰ → Models → Download → Use, then type a message. The Desktop user guide covers package verification, unsigned-app prompts, first-run setup, attachments, skills, updates and troubleshooting. Historical releases may have no Desktop assets until a release runs the updated workflow; iPhone/iPad remain source builds.

Supported model families at a glance

Backend, modality, feature support, and validation coverage vary by model. See the supported models tables, the model cards, and the embedding guide for details.

Recent source additions include Qwen-Image-2.1 masked edits with exact protected pixels and optional processing of the selected region, twelve TensorAgent LoRA plug-ins for speed, style and editing, and Qwen3.8 Flash Next on a 48 GB Mac using SSD-backed weights. Multi-GPU --layer-split and --tp are separate controls; support and performance depend on the architecture and quantization. These source features may be ahead of published packages; Desktop installers appear in releases built with the updated Release Binaries workflow, while older releases may lack them.

See it in action

One engine, four ways to use it, each an unedited capture of a real run.

TensorSharp.Cli in a terminal: an interactive chat with Gemma 4 E4B that reads this README and answers questions about it
TensorSharp.Cli
Models in your terminal
The TensorSharp Web UI: Qwen3.8 27B compared two mortgages by writing and running a Python script
TensorSharp.Server.Host
Web UI chat and Ollama/OpenAI-compatible APIs
TensorAgent in the iPhone 17 Pro simulator: Gemma 4 E2B uses ggml_cpu and runs an in-app Python script to scale a recipe
TensorAgent · iPhone simulator
CPU inference and in-app Python; this is a simulator capture
TensorAgent on a Mac: a saved Qwen-Image 2.1 edit changes the TensorSharp banner background to a starry blue night sky
TensorAgent on Mac
A saved Qwen-Image 2.1 edit in the Mac app

What each run shows, step by step: Screenshots.

Benchmarks

TensorSharp and llama.cpp run identical GGUF files on the same NVIDIA RTX 3080 Laptop GPU (16 GB), each on its GGML CUDA and Vulkan builds. Each number is TensorSharp's speedup over llama.cpp on the same backend (geomean, single-stream, greedy, MTP off); above 1.0× means TensorSharp is faster.

ModelBackenddecodeprefillTTFT
Gemma 4 E4B it (Q8_0, dense multimodal)CUDA1.02×1.28×1.27×
Gemma 4 E4B it (Q8_0, dense multimodal)Vulkan1.00×1.05×1.03×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense)CUDA1.04×1.17×1.16×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense)Vulkan1.21×1.04×1.03×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE)CUDA0.98×1.28×1.27×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE)Vulkan0.87×1.04×1.03×
Qwen 3.6 27B (UD-IQ2_XXS, dense)CUDA1.07×0.96×0.95×
Qwen 3.6 27B (UD-IQ2_XXS, dense)Vulkan1.02×0.85×0.84×

What these numbers mean, how to rerun them, and the head-to-heads of models too large for this GPU: Benchmarks.

Documentation

New here? The sections above are all you need to get running. Everything else is detailed reference:

DocWhat's inside
TensorAgent Desktop user guideMac/Windows downloads, DMG/PKG/MSI/ZIP installation, model setup, first chat, attachments, skills, updates and troubleshooting
TensorSharp and TensorAgent book guideBuilding LLM Inference Engines and Agentic Runtimes from Scratch, plus From Tensors to Tokens: introductions, Amazon links, and repository reading paths
Getting startedThe full first-run guide: the .NET SDK on each platform, every backend, multi-GPU and multi-node runs, NVIDIA DGX Spark, embeddings, choosing a backend, and making it fast
Supported modelsImplemented model families and their validation scope: example GGUFs, modalities, thinking, tools, and speculative decoding
BenchmarksTensorSharp against llama.cpp on the same GPU and files, and the head-to-heads of larger models
ScreenshotsThe CLI, the Web UI, and TensorAgent on iPhone and Mac at work, with what each run did
Model DownloadsPer-model huggingface-cli download + run quick reference (quant tiers, projectors, companions)
UsageFull CLI reference (options, interactive REPL, JSONL batch), server hosting, logging, HTTP API examples, backends, and the env-var matrix
FeaturesDeep dives on continuous batching, speculative decoding, tool calling, thinking mode, multimodal, MoE, KV codecs, and more
Configuration filesPut options in a reusable JSON file with ${variables} and auto-downloading models
DevelopmentPrerequisites, building the native GGML/MLX libraries, repository layout, package boundaries, internal architecture, and the test harness
Per-model architecture cardsEnd-to-end docs of each architecture (forward graph, components, parameters, prefill/decode optimizations)
Paged attention & continuous batchingThe vLLM-style paged KV cache, prefix sharing, and iteration-level scheduler
Agent Skills & agentic workThe SKILL.md format, progressive disclosure and its budget, the in-process tool loop, sandboxed code execution, workspaces and artifacts, the path/ZIP/exec security model, and the HTTP + C# surfaces
Multiple agentsAutomatic task delegation, private child workspaces, dependency scheduling, permission limits, server controls, and reproducible evaluation
Browser automation skill (Playwright)Running the bundled playwright skill, which drives a browser through @playwright/cli via skills_run: the flags it needs, the macOS Chromium-sandbox config, account handoff, and TensorAgent desktop hosting (not iOS)
Speculative decodingThe three-layer design (model adapter / algorithm / speculator weights), the shipped auto / draft-head / block / ngram algorithms, and what to write to add a new one
Environment variable feature matrixWhich high-impact runtime flags affect which models, backends, and prompt types
Engine comparison reportFull per-scenario TensorSharp vs llama.cpp tables
ggml_metal vs llama.cppHead-to-head prefill/decode on Apple Silicon, the four graph-construction gaps it found, and what each was worth
Test/benchmark matrix runnerSweep model × backend × feature × env-var cells and generate regression reports
Server API examplesComplete curl and Python examples for the server surface

Current Status

Actively developed, and the source tree runs ahead of the published packages.

AreaWhere it stands
ModelsA dozen autoregressive families plus text diffusion, image generation and editing, and video with audio. See Supported models.
Inference hostsCLI, Web UI, Ollama- and OpenAI-compatible APIs, and TensorAgent for iPhone, iPad, Mac and Windows. The release workflow packages Mac/Windows Desktop; historical releases may lack those assets. iPhone/iPad use source builds.
BackendsPure C# CPU, direct CUDA/cuBLAS, MLX Metal, and GGML CPU/Metal/CUDA/Vulkan, with per-architecture exceptions.
Serving featuresContinuous batching with a shared prefix cache, speculative decoding, tensor parallelism, structured output, and tool calling.
Agentic workAgent Skills, sandboxed file and shell tools, and bounded sub-agents. See Agent Skills and Multiple agents.
TensorAgentTwelve catalog entries, saved chats and artifacts, masked image edits and LoRA choices, eight interface languages, and persisted text-turn statistics. Media generation has been measured on a Mac; iOS media generation and Windows image/audio/video generation remain unverified.

Per-area detail (which architecture runs on which backend, which features each family supports, and the known limits) is in the status matrix.

Author

Zhongkai Fu

License

See LICENSE for details.

Learn with the books

Qwen inference and agentic runtimesGemma 4 and multimodal inference
Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgentFrom Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B
Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgentFrom Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B
Build Qwen dense/MoE inference and controlled agent workflows in C#. Follow tensors, tokenization, attention, expert routing, quantization, and caching through GPU acceleration, multimodal execution, tools, skills, sandboxed code, and desktop/mobile deployment with TensorSharp and TensorAgent.Build a multimodal inference engine in C#/.NET with Gemma 4 E4B, from tensors, GGUF model loading, quantization, and tokenization to text, image, video, and audio execution. Connect correctness checks and serving optimizations to the running TensorSharp code.
Buy on AmazonBuy on Amazon

Explore both books and their repository reading paths

Buy Me A Coffee
TensorSharp/TensorAgent is free. If you like it, a coffee keeps the work on it going.

ai
csharp
cuda
cuda-programming
deepseek
dotnet
dspark
gemma
gemma4
gguf
glm
inference
inference-engine
llm
metal
minimax-h3
multimodal
qwen
qwen35
transformers

Significant stargazers

Oisin Grehan

117 followers · starred May 2026

Steely Wing

43 followers · starred Sep 2026

Islam Nofl

155 followers · starred Jun 2026

InCerryGit

169 followers · starred Jul 2026

zhongkaifu/TensorSharp

A native .NET LLM inference engine and agent runtime for GGUF models. TensorSharp provides a console application, a web-based chatbot interface, iPhone App, and Ollama/OpenAI-compatible HTTP APIs for programmatic access. It supports Windows/MacOS/iOS/Linux with full GPU capability

C#

513

975 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built an open-source inference engine that runs a 176B MoE model on my RTX 3080 laptop (r/SideProject)

I’ve been building an open-source side project called **TensorSharp**, a local LLM inference engine and agent runtime written in C#/.NET: [TensorSharp on GitHub](https://github.com/zhongkaifu/TensorSharp?utm_source=chatgpt.com) One of the problems I’ve been experimenting with is: **How far can we…

2

Oct 3, 2026

Running a 176B Qwen3.8 Flash Next on a 16GB RTX 3080 Laptop + 32GB RAM + SSD (r/LocalLLM)

How much hardware do you actually need to run a **176B-parameter Qwen3.8 Flash Next** locally? Turns out, this can be enough: **RTX 3080 Laptop — 16GB VRAM + 32GB system RAM + SSD.** Yes — a **176B MoE model on a laptop**. I’ve been working on this in **TensorSharp**, my open-source local LLM…

3

Oct 3, 2026

Running Qwen3.8 Flash Next 176B on a 16GB RTX 3080 Laptop + 32GB RAM + SSD (r/LocalLLaMA)

I wanted to see how far I could push a fairly ordinary laptop with a huge MoE model. Turns out, **Qwen3.8 Flash Next 176B** can run on: **RTX 3080 Laptop — 16GB VRAM** **32GB system RAM** **SSD** No 128GB/256GB RAM workstation and no multi-GPU setup. I’m running it with **TensorSharp**, my…

9

Oct 3, 2026

README

TensorSharp

TensorSharp logo

English | 中文

TensorSharp is a .NET 10 inference engine for local GGUF models. Run it on Windows, macOS or Linux through the CLI, browser chat, or Ollama/OpenAI-compatible APIs, or embed it in your own .NET application. Use managed C# CPU kernels or native CUDA, Metal and Vulkan backends, with support that varies by model.

Current source covers text and reasoning, multimodal input, embeddings, image generation/editing, video with audio, and Agent Skills with code tools. TensorSharp also powers TensorAgent, the local app for iPhone, iPad, Mac and Windows. Start below, or check the project status for capabilities and validation limits; source changes may be ahead of published packages.

Highlights

  • Local models, several interfaces. One engine for CLI use, browser chat and Ollama/OpenAI-compatible APIs, across managed CPU and native accelerator backends.
  • Text and multimodal models. Dense and MoE GGUF models, reasoning, image/audio input and document questions. See Supported models for each family's capabilities.
  • Embeddings and media generation. Text/code embeddings, Qwen-Image-2.1 generation and masked editing, and video models including MiniMax-H3 with audio.
  • Efficient serving. Continuous batching and a paged, Radix prefix-shared KV cache are on by default. Speculative decoding and multi-GPU placement are available for supported models. See Features.
  • Agent Skills and code tools. TensorSharp.AgentHost adds file, shell and document workflows. Server and TensorAgent chats also support bounded sub-agent delegation with private workspaces and read-only defaults.
  • Recorded comparisons. Benchmarks against llama.cpp use identical GGUF files and hardware; results apply to the measured model, backend and workload. See Benchmarks.
  • TensorAgent apps. Local chat, attachments, skills and saved work on phones and desktops. See the Mac/Windows installation guide and app coverage notes.

Quick Start

TensorSharp CLI and server

The Releases page provides self-contained CLI and Server archives for Windows x64 (CPU/CUDA), Linux x64 (CPU/CUDA), and macOS arm64.

To build from source you need the full .NET 10 SDK (how to install it), git, curl, CMake 3.20+, and the toolchain for your GPU. Then run the verified Gemma 4 E4B model (7.48 GiB). On Windows with an NVIDIA GPU (PowerShell):

git clone https://github.com/zhongkaifu/TensorSharp.git; Set-Location TensorSharp
New-Item -ItemType Directory -Force models | Out-Null
curl.exe -L --fail "https://huggingface.co/ggml-org/gemma-4-E4B-it-GGUF/resolve/main/gemma-4-E4B-it-Q8_0.gguf?download=true" -o models\gemma-4-E4B-it-Q8_0.gguf
'Answer in one short sentence: what is TensorSharp?' | Set-Content prompt.txt
$env:TENSORSHARP_GGML_NATIVE_ENABLE_CUDA = 'ON'
dotnet run --project TensorSharp.Cli -c Release -p:TensorSharpSkipMlxNative=true -- --model models\gemma-4-E4B-it-Q8_0.gguf --input prompt.txt --max-tokens 128 --backend ggml_cuda

On other machines, change the backend (see Pick a backend):

  • macOS (Apple Silicon): drop the CUDA environment variable and use --backend ggml_metal.
  • Linux + NVIDIA: prefix the dotnet run with TENSORSHARP_GGML_NATIVE_ENABLE_CUDA=ON and use --backend ggml_cuda.
  • AMD / Intel / NVIDIA Vulkan: set TENSORSHARP_GGML_NATIVE_ENABLE_VULKAN=ON and use --backend ggml_vulkan.

Host the same model as a server: a browser chat at http://localhost:5000 plus Ollama- and OpenAI-compatible APIs.

dotnet run --project TensorSharp.Server.Host -c Release -p:TensorSharpSkipMlxNative=true -- --model models/gemma-4-E4B-it-Q8_0.gguf --backend ggml_cuda --max-tokens 512

The server listens on 0.0.0.0:5000 with no built-in authentication or TLS; keep it behind a firewall or an authenticated HTTPS reverse proxy.

Pick a backend

Your hardwareBackend
Apple Silicon (Mac)--backend ggml_metal
Windows / Linux + NVIDIA GPU--backend ggml_cuda
Windows / Linux + AMD / Intel / NVIDIA GPU--backend ggml_vulkan
No GPU--backend ggml_cpu (native kernels), or --backend cpu (pure C#, no native dependencies)

The Getting started guide has the rest: installing the SDK on each platform, multi-GPU and multi-node runs, NVIDIA DGX Spark, multimodal input, embeddings, and making it fast. Every option is in the CLI and Server references, and both programs print them with --help.

dotnet build TensorSharp.slnx also builds TensorAgent's available desktop heads and the iOS simulator head on Apple Silicon when the selected SDK has the required MAUI workloads and staged native/Python files. Missing prerequisites skip the affected app head with a warning; see TensorAgent build instructions.

TensorAgent Desktop: download, install, chat

Download the latest release and expand Assets. Choose a file beginning with tensoragent-desktop-:

PlatformDownload and install
macOS 14+ on Apple Silicontensoragent-desktop-<version>-osx-arm64.dmg: open and drag TensorAgent to Applications. A PKG installer and ZIP are also available.
Windows x64tensoragent-desktop-<version>-win-x64-cpu.msi, or win-x64-cuda.msi for a compatible NVIDIA GPU/driver: install and open TensorAgent from Start. ZIPs are also available.

The app includes its .NET runtime and native engine. Windows needs WebView2 Evergreen Runtime if missing; Python/Node are optional tools for skills. The current built-in catalog needs at least 12 GB system RAM, and model weights download separately. Open ☰ → Models → Download → Use, then type a message. The Desktop user guide covers package verification, unsigned-app prompts, first-run setup, attachments, skills, updates and troubleshooting. Historical releases may have no Desktop assets until a release runs the updated workflow; iPhone/iPad remain source builds.

Supported model families at a glance

Backend, modality, feature support, and validation coverage vary by model. See the supported models tables, the model cards, and the embedding guide for details.

Recent source additions include Qwen-Image-2.1 masked edits with exact protected pixels and optional processing of the selected region, twelve TensorAgent LoRA plug-ins for speed, style and editing, and Qwen3.8 Flash Next on a 48 GB Mac using SSD-backed weights. Multi-GPU --layer-split and --tp are separate controls; support and performance depend on the architecture and quantization. These source features may be ahead of published packages; Desktop installers appear in releases built with the updated Release Binaries workflow, while older releases may lack them.

See it in action

One engine, four ways to use it, each an unedited capture of a real run.

TensorSharp.Cli in a terminal: an interactive chat with Gemma 4 E4B that reads this README and answers questions about it
TensorSharp.Cli
Models in your terminal
The TensorSharp Web UI: Qwen3.8 27B compared two mortgages by writing and running a Python script
TensorSharp.Server.Host
Web UI chat and Ollama/OpenAI-compatible APIs
TensorAgent in the iPhone 17 Pro simulator: Gemma 4 E2B uses ggml_cpu and runs an in-app Python script to scale a recipe
TensorAgent · iPhone simulator
CPU inference and in-app Python; this is a simulator capture
TensorAgent on a Mac: a saved Qwen-Image 2.1 edit changes the TensorSharp banner background to a starry blue night sky
TensorAgent on Mac
A saved Qwen-Image 2.1 edit in the Mac app

What each run shows, step by step: Screenshots.

Benchmarks

TensorSharp and llama.cpp run identical GGUF files on the same NVIDIA RTX 3080 Laptop GPU (16 GB), each on its GGML CUDA and Vulkan builds. Each number is TensorSharp's speedup over llama.cpp on the same backend (geomean, single-stream, greedy, MTP off); above 1.0× means TensorSharp is faster.

ModelBackenddecodeprefillTTFT
Gemma 4 E4B it (Q8_0, dense multimodal)CUDA1.02×1.28×1.27×
Gemma 4 E4B it (Q8_0, dense multimodal)Vulkan1.00×1.05×1.03×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense)CUDA1.04×1.17×1.16×
Gemma 4 12B it (QAT UD-Q4_K_XL, dense)Vulkan1.21×1.04×1.03×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE)CUDA0.98×1.28×1.27×
Qwen 3.6 35B-A3B (UD-IQ2_XXS, MoE)Vulkan0.87×1.04×1.03×
Qwen 3.6 27B (UD-IQ2_XXS, dense)CUDA1.07×0.96×0.95×
Qwen 3.6 27B (UD-IQ2_XXS, dense)Vulkan1.02×0.85×0.84×

What these numbers mean, how to rerun them, and the head-to-heads of models too large for this GPU: Benchmarks.

Documentation

New here? The sections above are all you need to get running. Everything else is detailed reference:

DocWhat's inside
TensorAgent Desktop user guideMac/Windows downloads, DMG/PKG/MSI/ZIP installation, model setup, first chat, attachments, skills, updates and troubleshooting
TensorSharp and TensorAgent book guideBuilding LLM Inference Engines and Agentic Runtimes from Scratch, plus From Tensors to Tokens: introductions, Amazon links, and repository reading paths
Getting startedThe full first-run guide: the .NET SDK on each platform, every backend, multi-GPU and multi-node runs, NVIDIA DGX Spark, embeddings, choosing a backend, and making it fast
Supported modelsImplemented model families and their validation scope: example GGUFs, modalities, thinking, tools, and speculative decoding
BenchmarksTensorSharp against llama.cpp on the same GPU and files, and the head-to-heads of larger models
ScreenshotsThe CLI, the Web UI, and TensorAgent on iPhone and Mac at work, with what each run did
Model DownloadsPer-model huggingface-cli download + run quick reference (quant tiers, projectors, companions)
UsageFull CLI reference (options, interactive REPL, JSONL batch), server hosting, logging, HTTP API examples, backends, and the env-var matrix
FeaturesDeep dives on continuous batching, speculative decoding, tool calling, thinking mode, multimodal, MoE, KV codecs, and more
Configuration filesPut options in a reusable JSON file with ${variables} and auto-downloading models
DevelopmentPrerequisites, building the native GGML/MLX libraries, repository layout, package boundaries, internal architecture, and the test harness
Per-model architecture cardsEnd-to-end docs of each architecture (forward graph, components, parameters, prefill/decode optimizations)
Paged attention & continuous batchingThe vLLM-style paged KV cache, prefix sharing, and iteration-level scheduler
Agent Skills & agentic workThe SKILL.md format, progressive disclosure and its budget, the in-process tool loop, sandboxed code execution, workspaces and artifacts, the path/ZIP/exec security model, and the HTTP + C# surfaces
Multiple agentsAutomatic task delegation, private child workspaces, dependency scheduling, permission limits, server controls, and reproducible evaluation
Browser automation skill (Playwright)Running the bundled playwright skill, which drives a browser through @playwright/cli via skills_run: the flags it needs, the macOS Chromium-sandbox config, account handoff, and TensorAgent desktop hosting (not iOS)
Speculative decodingThe three-layer design (model adapter / algorithm / speculator weights), the shipped auto / draft-head / block / ngram algorithms, and what to write to add a new one
Environment variable feature matrixWhich high-impact runtime flags affect which models, backends, and prompt types
Engine comparison reportFull per-scenario TensorSharp vs llama.cpp tables
ggml_metal vs llama.cppHead-to-head prefill/decode on Apple Silicon, the four graph-construction gaps it found, and what each was worth
Test/benchmark matrix runnerSweep model × backend × feature × env-var cells and generate regression reports
Server API examplesComplete curl and Python examples for the server surface

Current Status

Actively developed, and the source tree runs ahead of the published packages.

AreaWhere it stands
ModelsA dozen autoregressive families plus text diffusion, image generation and editing, and video with audio. See Supported models.
Inference hostsCLI, Web UI, Ollama- and OpenAI-compatible APIs, and TensorAgent for iPhone, iPad, Mac and Windows. The release workflow packages Mac/Windows Desktop; historical releases may lack those assets. iPhone/iPad use source builds.
BackendsPure C# CPU, direct CUDA/cuBLAS, MLX Metal, and GGML CPU/Metal/CUDA/Vulkan, with per-architecture exceptions.
Serving featuresContinuous batching with a shared prefix cache, speculative decoding, tensor parallelism, structured output, and tool calling.
Agentic workAgent Skills, sandboxed file and shell tools, and bounded sub-agents. See Agent Skills and Multiple agents.
TensorAgentTwelve catalog entries, saved chats and artifacts, masked image edits and LoRA choices, eight interface languages, and persisted text-turn statistics. Media generation has been measured on a Mac; iOS media generation and Windows image/audio/video generation remain unverified.

Per-area detail (which architecture runs on which backend, which features each family supports, and the known limits) is in the status matrix.

Author

Zhongkai Fu

License

See LICENSE for details.

Learn with the books

Qwen inference and agentic runtimesGemma 4 and multimodal inference
Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgentFrom Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B
Building LLM Inference Engines and Agentic Runtimes from Scratch: Qwen Dense and MoE Models with TensorSharp and TensorAgentFrom Tensors to Tokens: Building a Multimodal LLM Inference Engine from Scratch with TensorSharp and Gemma 4 E4B
Build Qwen dense/MoE inference and controlled agent workflows in C#. Follow tensors, tokenization, attention, expert routing, quantization, and caching through GPU acceleration, multimodal execution, tools, skills, sandboxed code, and desktop/mobile deployment with TensorSharp and TensorAgent.Build a multimodal inference engine in C#/.NET with Gemma 4 E4B, from tensors, GGUF model loading, quantization, and tokenization to text, image, video, and audio execution. Connect correctness checks and serving optimizations to the running TensorSharp code.
Buy on AmazonBuy on Amazon

Explore both books and their repository reading paths

Buy Me A Coffee
TensorSharp/TensorAgent is free. If you like it, a coffee keeps the work on it going.

ai
csharp
cuda
cuda-programming
deepseek
dotnet
dspark
gemma
gemma4
gguf
glm
inference
inference-engine
llm
metal
minimax-h3
multimodal
qwen
qwen35
transformers

Significant stargazers

Oisin Grehan

117 followers · starred May 2026

Steely Wing

43 followers · starred Sep 2026

Islam Nofl

155 followers · starred Jun 2026

InCerryGit

169 followers · starred Jul 2026

Languages

C#

70.7%

C++

12.5%

Python

8.2%

HTML

4.5%

Cuda

2.0%