Local AI in C#: LLM, image/video diffusion, vision, and speech (TTS/ASR) — GGUF on CPU (AVX2/AVX-512), Vulkan, or CUDA. NativeAOT, no Python.
C#
0
1,814 commits
updated Oct 4, 2026
Local AI for .NET. Chat with language models, turn text into speech, transcribe audio, read images and generate pictures, inside your own .NET process. There is no Python, no sidecar server and no native binaries to ship: the engine is managed C# that runs on your CPU, or your GPU through Vulkan or CUDA.
> dotnet run -- chat qwen2.5-0.5b-instruct-q4_k_m.gguf "In one sentence, why do developers write unit tests?"
Developers write unit tests to ensure individual components of a program work as intended,
enhancing code quality, maintaining stability, and facilitating testing and debugging.
That answer came from a 469 MB model running on an ordinary desktop CPU, in about three seconds.
dotnet add package OpenTail.Stingray and your app can run
models. Nothing to install on the user's machine, nothing to keep in sync, and it publishes with
NativeAOT into a single executable.New here? What can I do with Stingray? is a one-page tour of every task it handles, what you need and where the limits are, and the task guides show how to do each one. If something goes wrong, see troubleshooting.
You need the .NET 10 SDK and a 64-bit x86 CPU with AVX2 (most PCs from 2015 on). A GPU is optional.
1. Create a project and add the package
dotnet new console -n HelloStingray && cd HelloStingray
dotnet add package OpenTail.Stingray
2. Download a small model (Qwen2.5 0.5B Instruct, 469 MB). The stingray command-line tool
fetches GGUF files from Hugging Face:
dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
This saves models/qwen2.5-0.5b-instruct-q4_k_m.gguf. You can also download it from the
model page by hand.
3. Replace Program.cs
using OpenTail.Stingray.Core;
using OpenTail.Stingray.Cpu;
using OpenTail.Stingray.Engine;
// Load the model and its tokenizer, and run it on the CPU.
using var model = GgufModel.Open("models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var hp = ModelHyperparams.FromGgufMetadata(model.Metadata, model);
var tokenizer = GgufTokenizer.FromGgufModel(model);
using var cpu = new CpuBackend();
var forward = new ForwardPass(model, cpu, hp, maxContextLength: 4096);
await using var engine = new InferenceEngine(forward, tokenizer, "qwen", forward);
// Format the question with the model's own chat template, then stream the answer.
string prompt = tokenizer.ChatTemplate!.Render(new Dictionary<string, object?>
{
["messages"] = JinjaChatTemplate.BuildMessages("In one sentence, why do developers write unit tests?"),
["add_generation_prompt"] = true,
});
await foreach (string piece in engine.GenerateAsync(prompt, new SamplingParams { Temperature = 0.7f, MaxNewTokens = 200 }))
Console.Write(piece);
4. Run it
dotnet run
Yes, that is more wiring than it should be. A one-line Open("model") API is being designed
(docs/103). The code above is what works today, and it is
compiled and run as samples/QuickStart.
Text to speech with a Piper voice. Download
both files of a voice, for example en_US-lessac-medium.onnx (63 MB) and
en_US-lessac-medium.onnx.json
(here):
using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Piper;
using var tts = PiperPipeline.FromConfigFile("en_US-lessac-medium.onnx.json");
tts.Generate(new AudioGenerationRequest { Text = "Hello from a voice that never left this computer.", OutputPath = "hello.wav" });
Speech to text with Whisper. Download
ggml-base.bin (148 MB):
using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Whisper;
using var stt = WhisperPipeline.Load("ggml-base.bin");
var (samples, sampleRate, _) = WavReader.ReadWav("hello.wav");
Console.WriteLine(stt.Transcribe(new SpeechToTextRequest { AudioSamples = samples, SampleRate = sampleRate }).Text);
// Hello from a voice that never left this computer.
On the test machine below, the voice generates 2.7 seconds of audio in 1.2 seconds, and Whisper transcribes it back word for word.
Add OpenTail.Stingray.Server to an ASP.NET project, and existing OpenAI clients can talk to a
local model:
using OpenTail.Stingray.Server;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddOpenTailStingray(builder.Configuration, o => o.ModelPath = "models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var app = builder.Build();
app.MapOpenTailStingray(); // /v1/chat/completions, /v1/models, Anthropic and Responses APIs
app.Run("http://localhost:5080");
curl localhost:5080/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model":"local","messages":[{"role":"user","content":"Name three planets, comma-separated."}]}'
# ... "content":"Mercury, Venus, and Earth." ...
See samples/ChatServer.
The same engine as a command-line tool, which is handy for trying models before writing code:
dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf # interactive chat
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf -p "What is a unit test?"
stingray tts -e piper -m en_US-lessac-medium.onnx.json -t "Hello!" -o hello.wav
stingray stt -m base --model-file ggml-base.bin -i hello.wav
Add -g -1 to run a language model on the GPU. Flag names follow llama.cpp's llama-cli where
they mean the same thing.
Built from source (not yet in the published package), stingray setup fetches the recommended
model for a task into a per-user folder (%LOCALAPPDATA%\stingray\models, or
~/.cache/stingray/models; set STINGRAY_MODEL_HOME to move it), checks each file's SHA-256 and
prints the command to run it. stingray models shows which tasks are ready:
stingray setup chat # Qwen2.5 0.5B Instruct, 469 MB
stingray setup speak # Piper en_US-lessac-medium, 60 MB (asks you to accept the voice data licence)
stingray setup transcribe # Whisper base, 141 MB
stingray models
docs/MODELS.md is a short, curated list of models to start with, one table per task (chat, images, speech, transcription, image generation, search). Each entry links straight to the file on Hugging Face and shows its size, licence and whether we have tested that exact file. docs/RUNNING.md then says how to run each one well: the command, the RAM it really needs and its measured speed.
The recipes above are the verified starting points. The engine covers much more, at varying levels of polish:
What is verified, what is partial and what is experimental is tracked per model, with dated evidence, in docs/STATUS.md. Speed comparisons between the speech engines are in the same file.
The timings in this README come from an AMD Ryzen 7 5700G (8 cores, integrated graphics only) with 64 GB of RAM, CPU only, on Windows 11. Rough guide:
| You want to | Download | RAM | On that CPU |
|---|---|---|---|
| Chat with a small model | 0.5 GB | 2 GB | ~60-70 tokens/s (Qwen2.5 0.5B) |
| Chat with a 7–8B model | 4–5 GB | 8 GB | a few tokens/s; a GPU helps a lot |
| Text to speech (Piper) | 63 MB | < 1 GB | faster than real time |
| Speech to text (Whisper base) | 148 MB | < 1 GB | ~2x real time |
| Image generation | 5–12 GB | 16 GB+ | minutes per image; a GPU is recommended |
A GPU is optional: any Vulkan-capable card, or NVIDIA with CUDA 12.
A guided "front door" is under way in docs/103. Done: a small
catalog of verified models per task and stingray setup <task> / stingray models. Next: task
commands that need no file paths, and a one-line C# API.
git clone https://github.com/opentail-net/OpenTail.Stingray && cd OpenTail.Stingray
dotnet build -c Release
dotnet run --project samples/QuickStart -c Release -- chat models/qwen2.5-0.5b-instruct-q4_k_m.gguf "Hello"
The documentation index sorts docs/ into user guides, design notes and active
engineering work; contributors should read
CLAUDE.md for the build and test conventions.
MIT — Copyright (c) 2026 OpenTail. Model files have their own licenses; check each model's page.
C#
99.0%
Local AI in C#: LLM, image/video diffusion, vision, and speech (TTS/ASR) — GGUF on CPU (AVX2/AVX-512), Vulkan, or CUDA. NativeAOT, no Python.
C#
0
1,814 commits
updated Oct 4, 2026
Local AI for .NET. Chat with language models, turn text into speech, transcribe audio, read images and generate pictures, inside your own .NET process. There is no Python, no sidecar server and no native binaries to ship: the engine is managed C# that runs on your CPU, or your GPU through Vulkan or CUDA.
> dotnet run -- chat qwen2.5-0.5b-instruct-q4_k_m.gguf "In one sentence, why do developers write unit tests?"
Developers write unit tests to ensure individual components of a program work as intended,
enhancing code quality, maintaining stability, and facilitating testing and debugging.
That answer came from a 469 MB model running on an ordinary desktop CPU, in about three seconds.
dotnet add package OpenTail.Stingray and your app can run
models. Nothing to install on the user's machine, nothing to keep in sync, and it publishes with
NativeAOT into a single executable.New here? What can I do with Stingray? is a one-page tour of every task it handles, what you need and where the limits are, and the task guides show how to do each one. If something goes wrong, see troubleshooting.
You need the .NET 10 SDK and a 64-bit x86 CPU with AVX2 (most PCs from 2015 on). A GPU is optional.
1. Create a project and add the package
dotnet new console -n HelloStingray && cd HelloStingray
dotnet add package OpenTail.Stingray
2. Download a small model (Qwen2.5 0.5B Instruct, 469 MB). The stingray command-line tool
fetches GGUF files from Hugging Face:
dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
This saves models/qwen2.5-0.5b-instruct-q4_k_m.gguf. You can also download it from the
model page by hand.
3. Replace Program.cs
using OpenTail.Stingray.Core;
using OpenTail.Stingray.Cpu;
using OpenTail.Stingray.Engine;
// Load the model and its tokenizer, and run it on the CPU.
using var model = GgufModel.Open("models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var hp = ModelHyperparams.FromGgufMetadata(model.Metadata, model);
var tokenizer = GgufTokenizer.FromGgufModel(model);
using var cpu = new CpuBackend();
var forward = new ForwardPass(model, cpu, hp, maxContextLength: 4096);
await using var engine = new InferenceEngine(forward, tokenizer, "qwen", forward);
// Format the question with the model's own chat template, then stream the answer.
string prompt = tokenizer.ChatTemplate!.Render(new Dictionary<string, object?>
{
["messages"] = JinjaChatTemplate.BuildMessages("In one sentence, why do developers write unit tests?"),
["add_generation_prompt"] = true,
});
await foreach (string piece in engine.GenerateAsync(prompt, new SamplingParams { Temperature = 0.7f, MaxNewTokens = 200 }))
Console.Write(piece);
4. Run it
dotnet run
Yes, that is more wiring than it should be. A one-line Open("model") API is being designed
(docs/103). The code above is what works today, and it is
compiled and run as samples/QuickStart.
Text to speech with a Piper voice. Download
both files of a voice, for example en_US-lessac-medium.onnx (63 MB) and
en_US-lessac-medium.onnx.json
(here):
using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Piper;
using var tts = PiperPipeline.FromConfigFile("en_US-lessac-medium.onnx.json");
tts.Generate(new AudioGenerationRequest { Text = "Hello from a voice that never left this computer.", OutputPath = "hello.wav" });
Speech to text with Whisper. Download
ggml-base.bin (148 MB):
using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Whisper;
using var stt = WhisperPipeline.Load("ggml-base.bin");
var (samples, sampleRate, _) = WavReader.ReadWav("hello.wav");
Console.WriteLine(stt.Transcribe(new SpeechToTextRequest { AudioSamples = samples, SampleRate = sampleRate }).Text);
// Hello from a voice that never left this computer.
On the test machine below, the voice generates 2.7 seconds of audio in 1.2 seconds, and Whisper transcribes it back word for word.
Add OpenTail.Stingray.Server to an ASP.NET project, and existing OpenAI clients can talk to a
local model:
using OpenTail.Stingray.Server;
var builder = WebApplication.CreateBuilder(args);
builder.Services.AddOpenTailStingray(builder.Configuration, o => o.ModelPath = "models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var app = builder.Build();
app.MapOpenTailStingray(); // /v1/chat/completions, /v1/models, Anthropic and Responses APIs
app.Run("http://localhost:5080");
curl localhost:5080/v1/chat/completions -H "Content-Type: application/json" \
-d '{"model":"local","messages":[{"role":"user","content":"Name three planets, comma-separated."}]}'
# ... "content":"Mercury, Venus, and Earth." ...
See samples/ChatServer.
The same engine as a command-line tool, which is handy for trying models before writing code:
dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf # interactive chat
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf -p "What is a unit test?"
stingray tts -e piper -m en_US-lessac-medium.onnx.json -t "Hello!" -o hello.wav
stingray stt -m base --model-file ggml-base.bin -i hello.wav
Add -g -1 to run a language model on the GPU. Flag names follow llama.cpp's llama-cli where
they mean the same thing.
Built from source (not yet in the published package), stingray setup fetches the recommended
model for a task into a per-user folder (%LOCALAPPDATA%\stingray\models, or
~/.cache/stingray/models; set STINGRAY_MODEL_HOME to move it), checks each file's SHA-256 and
prints the command to run it. stingray models shows which tasks are ready:
stingray setup chat # Qwen2.5 0.5B Instruct, 469 MB
stingray setup speak # Piper en_US-lessac-medium, 60 MB (asks you to accept the voice data licence)
stingray setup transcribe # Whisper base, 141 MB
stingray models
docs/MODELS.md is a short, curated list of models to start with, one table per task (chat, images, speech, transcription, image generation, search). Each entry links straight to the file on Hugging Face and shows its size, licence and whether we have tested that exact file. docs/RUNNING.md then says how to run each one well: the command, the RAM it really needs and its measured speed.
The recipes above are the verified starting points. The engine covers much more, at varying levels of polish:
What is verified, what is partial and what is experimental is tracked per model, with dated evidence, in docs/STATUS.md. Speed comparisons between the speech engines are in the same file.
The timings in this README come from an AMD Ryzen 7 5700G (8 cores, integrated graphics only) with 64 GB of RAM, CPU only, on Windows 11. Rough guide:
| You want to | Download | RAM | On that CPU |
|---|---|---|---|
| Chat with a small model | 0.5 GB | 2 GB | ~60-70 tokens/s (Qwen2.5 0.5B) |
| Chat with a 7–8B model | 4–5 GB | 8 GB | a few tokens/s; a GPU helps a lot |
| Text to speech (Piper) | 63 MB | < 1 GB | faster than real time |
| Speech to text (Whisper base) | 148 MB | < 1 GB | ~2x real time |
| Image generation | 5–12 GB | 16 GB+ | minutes per image; a GPU is recommended |
A GPU is optional: any Vulkan-capable card, or NVIDIA with CUDA 12.
A guided "front door" is under way in docs/103. Done: a small
catalog of verified models per task and stingray setup <task> / stingray models. Next: task
commands that need no file paths, and a one-line C# API.
git clone https://github.com/opentail-net/OpenTail.Stingray && cd OpenTail.Stingray
dotnet build -c Release
dotnet run --project samples/QuickStart -c Release -- chat models/qwen2.5-0.5b-instruct-q4_k_m.gguf "Hello"
The documentation index sorts docs/ into user guides, design notes and active
engineering work; contributors should read
CLAUDE.md for the build and test conventions.
MIT — Copyright (c) 2026 OpenTail. Model files have their own licenses; check each model's page.
C#
99.0%