opentail-net/OpenTail.Stingray

Local AI in C#: LLM, image/video diffusion, vision, and speech (TTS/ASR) — GGUF on CPU (AVX2/AVX-512), Vulkan, or CUDA. NativeAOT, no Python.

C#

0

1,814 commits

updated Oct 4, 2026

See the code

README

OpenTail.Stingray

Local AI for .NET. Chat with language models, turn text into speech, transcribe audio, read images and generate pictures, inside your own .NET process. There is no Python, no sidecar server and no native binaries to ship: the engine is managed C# that runs on your CPU, or your GPU through Vulkan or CUDA.

NuGet .NET 10 License: MIT

> dotnet run -- chat qwen2.5-0.5b-instruct-q4_k_m.gguf "In one sentence, why do developers write unit tests?"
Developers write unit tests to ensure individual components of a program work as intended,
enhancing code quality, maintaining stability, and facilitating testing and debugging.

That answer came from a 469 MB model running on an ordinary desktop CPU, in about three seconds.

Why Stingray

  • It's just a NuGet package. dotnet add package OpenTail.Stingray and your app can run models. Nothing to install on the user's machine, nothing to keep in sync, and it publishes with NativeAOT into a single executable.
  • One library for text, speech and images. The same package does chat, text-to-speech, speech-to-text, image understanding and image generation, so you don't glue five tools together.
  • It reads the models people already use. GGUF files from Hugging Face (the llama.cpp format), plus the common ONNX/safetensors releases for speech and diffusion.
  • Checked against the reference implementations. Language models are compared token by token with llama.cpp. What is verified, and how, is recorded model by model in docs/STATUS.md.

New here? What can I do with Stingray? is a one-page tour of every task it handles, what you need and where the limits are, and the task guides show how to do each one. If something goes wrong, see troubleshooting.

Quick start: chat from C#

You need the .NET 10 SDK and a 64-bit x86 CPU with AVX2 (most PCs from 2015 on). A GPU is optional.

1. Create a project and add the package

dotnet new console -n HelloStingray && cd HelloStingray
dotnet add package OpenTail.Stingray

2. Download a small model (Qwen2.5 0.5B Instruct, 469 MB). The stingray command-line tool fetches GGUF files from Hugging Face:

dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF

This saves models/qwen2.5-0.5b-instruct-q4_k_m.gguf. You can also download it from the model page by hand.

3. Replace Program.cs

using OpenTail.Stingray.Core;
using OpenTail.Stingray.Cpu;
using OpenTail.Stingray.Engine;

// Load the model and its tokenizer, and run it on the CPU.
using var model = GgufModel.Open("models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var hp = ModelHyperparams.FromGgufMetadata(model.Metadata, model);
var tokenizer = GgufTokenizer.FromGgufModel(model);
using var cpu = new CpuBackend();
var forward = new ForwardPass(model, cpu, hp, maxContextLength: 4096);
await using var engine = new InferenceEngine(forward, tokenizer, "qwen", forward);

// Format the question with the model's own chat template, then stream the answer.
string prompt = tokenizer.ChatTemplate!.Render(new Dictionary<string, object?>
{
    ["messages"] = JinjaChatTemplate.BuildMessages("In one sentence, why do developers write unit tests?"),
    ["add_generation_prompt"] = true,
});
await foreach (string piece in engine.GenerateAsync(prompt, new SamplingParams { Temperature = 0.7f, MaxNewTokens = 200 }))
    Console.Write(piece);

4. Run it

dotnet run

Yes, that is more wiring than it should be. A one-line Open("model") API is being designed (docs/103). The code above is what works today, and it is compiled and run as samples/QuickStart.

Speak and listen

Text to speech with a Piper voice. Download both files of a voice, for example en_US-lessac-medium.onnx (63 MB) and en_US-lessac-medium.onnx.json (here):

using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Piper;

using var tts = PiperPipeline.FromConfigFile("en_US-lessac-medium.onnx.json");
tts.Generate(new AudioGenerationRequest { Text = "Hello from a voice that never left this computer.", OutputPath = "hello.wav" });

Speech to text with Whisper. Download ggml-base.bin (148 MB):

using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Whisper;

using var stt = WhisperPipeline.Load("ggml-base.bin");
var (samples, sampleRate, _) = WavReader.ReadWav("hello.wav");
Console.WriteLine(stt.Transcribe(new SpeechToTextRequest { AudioSamples = samples, SampleRate = sampleRate }).Text);
// Hello from a voice that never left this computer.

On the test machine below, the voice generates 2.7 seconds of audio in 1.2 seconds, and Whisper transcribes it back word for word.

Serve an OpenAI-compatible API

Add OpenTail.Stingray.Server to an ASP.NET project, and existing OpenAI clients can talk to a local model:

using OpenTail.Stingray.Server;

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddOpenTailStingray(builder.Configuration, o => o.ModelPath = "models/qwen2.5-0.5b-instruct-q4_k_m.gguf");

var app = builder.Build();
app.MapOpenTailStingray();   // /v1/chat/completions, /v1/models, Anthropic and Responses APIs
app.Run("http://localhost:5080");
curl localhost:5080/v1/chat/completions -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Name three planets, comma-separated."}]}'
# ... "content":"Mercury, Venus, and Earth." ...

See samples/ChatServer.

From the command line

The same engine as a command-line tool, which is handy for trying models before writing code:

dotnet tool install -g OpenTail.Stingray.Cli

stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf                        # interactive chat
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf -p "What is a unit test?"
stingray tts -e piper -m en_US-lessac-medium.onnx.json -t "Hello!" -o hello.wav
stingray stt -m base --model-file ggml-base.bin -i hello.wav

Add -g -1 to run a language model on the GPU. Flag names follow llama.cpp's llama-cli where they mean the same thing.

Built from source (not yet in the published package), stingray setup fetches the recommended model for a task into a per-user folder (%LOCALAPPDATA%\stingray\models, or ~/.cache/stingray/models; set STINGRAY_MODEL_HOME to move it), checks each file's SHA-256 and prints the command to run it. stingray models shows which tasks are ready:

stingray setup chat          # Qwen2.5 0.5B Instruct, 469 MB
stingray setup speak         # Piper en_US-lessac-medium, 60 MB (asks you to accept the voice data licence)
stingray setup transcribe    # Whisper base, 141 MB
stingray models

Finding models

docs/MODELS.md is a short, curated list of models to start with, one table per task (chat, images, speech, transcription, image generation, search). Each entry links straight to the file on Hugging Face and shows its size, licence and whether we have tested that exact file. docs/RUNNING.md then says how to run each one well: the command, the RAM it really needs and its measured speed.

What else it can do

The recipes above are the verified starting points. The engine covers much more, at varying levels of polish:

  • Language models: Llama, Qwen, Gemma, Mistral, Phi, DeepSeek, gpt-oss, GLM-4, Granite 4.0-H, Nemotron-H, LFM2 and many more GGUF architectures, on CPU, Vulkan or CUDA. Tool calling, JSON-schema constrained output, speculative decoding.
  • Images in, text out: Gemma, Qwen-VL, LLaVA, Pixtral, InternVL, OCR models such as dots.ocr.
  • Speech: text-to-speech engines including Kokoro, XTTS-v2 (voice cloning), Qwen3-TTS, Fish Speech and Chatterbox; speech recognition with Whisper, Parakeet and Qwen3-ASR.
  • Image and video generation: FLUX.1 / FLUX.2, Stable Diffusion 1.5 / XL / 3.5, Z-Image-Turbo, Qwen Image, Wan, LTX-Video, HunyuanVideo. These need several model files each; a guided setup is on the way.
  • Music and sound: Stable Audio 3, ACE-Step 1.5, MiniMax-Music3, MusicGen and AudioGen.

What is verified, what is partial and what is experimental is tracked per model, with dated evidence, in docs/STATUS.md. Speed comparisons between the speech engines are in the same file.

Hardware and speed

The timings in this README come from an AMD Ryzen 7 5700G (8 cores, integrated graphics only) with 64 GB of RAM, CPU only, on Windows 11. Rough guide:

You want toDownloadRAMOn that CPU
Chat with a small model0.5 GB2 GB~60-70 tokens/s (Qwen2.5 0.5B)
Chat with a 7–8B model4–5 GB8 GBa few tokens/s; a GPU helps a lot
Text to speech (Piper)63 MB< 1 GBfaster than real time
Speech to text (Whisper base)148 MB< 1 GB~2x real time
Image generation5–12 GB16 GB+minutes per image; a GPU is recommended

A GPU is optional: any Vulkan-capable card, or NVIDIA with CUDA 12.

What's next

A guided "front door" is under way in docs/103. Done: a small catalog of verified models per task and stingray setup <task> / stingray models. Next: task commands that need no file paths, and a one-line C# API.

Building from source

git clone https://github.com/opentail-net/OpenTail.Stingray && cd OpenTail.Stingray
dotnet build -c Release
dotnet run --project samples/QuickStart -c Release -- chat models/qwen2.5-0.5b-instruct-q4_k_m.gguf "Hello"

The documentation index sorts docs/ into user guides, design notes and active engineering work; contributors should read CLAUDE.md for the build and test conventions.

License

MIT — Copyright (c) 2026 OpenTail. Model files have their own licenses; check each model's page.

opentail-net/OpenTail.Stingray

Local AI in C#: LLM, image/video diffusion, vision, and speech (TTS/ASR) — GGUF on CPU (AVX2/AVX-512), Vulkan, or CUDA. NativeAOT, no Python.

C#

0

1,814 commits

updated Oct 4, 2026

See the code

README

OpenTail.Stingray

Local AI for .NET. Chat with language models, turn text into speech, transcribe audio, read images and generate pictures, inside your own .NET process. There is no Python, no sidecar server and no native binaries to ship: the engine is managed C# that runs on your CPU, or your GPU through Vulkan or CUDA.

NuGet .NET 10 License: MIT

> dotnet run -- chat qwen2.5-0.5b-instruct-q4_k_m.gguf "In one sentence, why do developers write unit tests?"
Developers write unit tests to ensure individual components of a program work as intended,
enhancing code quality, maintaining stability, and facilitating testing and debugging.

That answer came from a 469 MB model running on an ordinary desktop CPU, in about three seconds.

Why Stingray

  • It's just a NuGet package. dotnet add package OpenTail.Stingray and your app can run models. Nothing to install on the user's machine, nothing to keep in sync, and it publishes with NativeAOT into a single executable.
  • One library for text, speech and images. The same package does chat, text-to-speech, speech-to-text, image understanding and image generation, so you don't glue five tools together.
  • It reads the models people already use. GGUF files from Hugging Face (the llama.cpp format), plus the common ONNX/safetensors releases for speech and diffusion.
  • Checked against the reference implementations. Language models are compared token by token with llama.cpp. What is verified, and how, is recorded model by model in docs/STATUS.md.

New here? What can I do with Stingray? is a one-page tour of every task it handles, what you need and where the limits are, and the task guides show how to do each one. If something goes wrong, see troubleshooting.

Quick start: chat from C#

You need the .NET 10 SDK and a 64-bit x86 CPU with AVX2 (most PCs from 2015 on). A GPU is optional.

1. Create a project and add the package

dotnet new console -n HelloStingray && cd HelloStingray
dotnet add package OpenTail.Stingray

2. Download a small model (Qwen2.5 0.5B Instruct, 469 MB). The stingray command-line tool fetches GGUF files from Hugging Face:

dotnet tool install -g OpenTail.Stingray.Cli
stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF

This saves models/qwen2.5-0.5b-instruct-q4_k_m.gguf. You can also download it from the model page by hand.

3. Replace Program.cs

using OpenTail.Stingray.Core;
using OpenTail.Stingray.Cpu;
using OpenTail.Stingray.Engine;

// Load the model and its tokenizer, and run it on the CPU.
using var model = GgufModel.Open("models/qwen2.5-0.5b-instruct-q4_k_m.gguf");
var hp = ModelHyperparams.FromGgufMetadata(model.Metadata, model);
var tokenizer = GgufTokenizer.FromGgufModel(model);
using var cpu = new CpuBackend();
var forward = new ForwardPass(model, cpu, hp, maxContextLength: 4096);
await using var engine = new InferenceEngine(forward, tokenizer, "qwen", forward);

// Format the question with the model's own chat template, then stream the answer.
string prompt = tokenizer.ChatTemplate!.Render(new Dictionary<string, object?>
{
    ["messages"] = JinjaChatTemplate.BuildMessages("In one sentence, why do developers write unit tests?"),
    ["add_generation_prompt"] = true,
});
await foreach (string piece in engine.GenerateAsync(prompt, new SamplingParams { Temperature = 0.7f, MaxNewTokens = 200 }))
    Console.Write(piece);

4. Run it

dotnet run

Yes, that is more wiring than it should be. A one-line Open("model") API is being designed (docs/103). The code above is what works today, and it is compiled and run as samples/QuickStart.

Speak and listen

Text to speech with a Piper voice. Download both files of a voice, for example en_US-lessac-medium.onnx (63 MB) and en_US-lessac-medium.onnx.json (here):

using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Piper;

using var tts = PiperPipeline.FromConfigFile("en_US-lessac-medium.onnx.json");
tts.Generate(new AudioGenerationRequest { Text = "Hello from a voice that never left this computer.", OutputPath = "hello.wav" });

Speech to text with Whisper. Download ggml-base.bin (148 MB):

using OpenTail.Stingray.Audio;
using OpenTail.Stingray.Audio.Whisper;

using var stt = WhisperPipeline.Load("ggml-base.bin");
var (samples, sampleRate, _) = WavReader.ReadWav("hello.wav");
Console.WriteLine(stt.Transcribe(new SpeechToTextRequest { AudioSamples = samples, SampleRate = sampleRate }).Text);
// Hello from a voice that never left this computer.

On the test machine below, the voice generates 2.7 seconds of audio in 1.2 seconds, and Whisper transcribes it back word for word.

Serve an OpenAI-compatible API

Add OpenTail.Stingray.Server to an ASP.NET project, and existing OpenAI clients can talk to a local model:

using OpenTail.Stingray.Server;

var builder = WebApplication.CreateBuilder(args);
builder.Services.AddOpenTailStingray(builder.Configuration, o => o.ModelPath = "models/qwen2.5-0.5b-instruct-q4_k_m.gguf");

var app = builder.Build();
app.MapOpenTailStingray();   // /v1/chat/completions, /v1/models, Anthropic and Responses APIs
app.Run("http://localhost:5080");
curl localhost:5080/v1/chat/completions -H "Content-Type: application/json" \
  -d '{"model":"local","messages":[{"role":"user","content":"Name three planets, comma-separated."}]}'
# ... "content":"Mercury, Venus, and Earth." ...

See samples/ChatServer.

From the command line

The same engine as a command-line tool, which is handy for trying models before writing code:

dotnet tool install -g OpenTail.Stingray.Cli

stingray pull -r Qwen/Qwen2.5-0.5B-Instruct-GGUF
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf                        # interactive chat
stingray -m models/qwen2.5-0.5b-instruct-q4_k_m.gguf -p "What is a unit test?"
stingray tts -e piper -m en_US-lessac-medium.onnx.json -t "Hello!" -o hello.wav
stingray stt -m base --model-file ggml-base.bin -i hello.wav

Add -g -1 to run a language model on the GPU. Flag names follow llama.cpp's llama-cli where they mean the same thing.

Built from source (not yet in the published package), stingray setup fetches the recommended model for a task into a per-user folder (%LOCALAPPDATA%\stingray\models, or ~/.cache/stingray/models; set STINGRAY_MODEL_HOME to move it), checks each file's SHA-256 and prints the command to run it. stingray models shows which tasks are ready:

stingray setup chat          # Qwen2.5 0.5B Instruct, 469 MB
stingray setup speak         # Piper en_US-lessac-medium, 60 MB (asks you to accept the voice data licence)
stingray setup transcribe    # Whisper base, 141 MB
stingray models

Finding models

docs/MODELS.md is a short, curated list of models to start with, one table per task (chat, images, speech, transcription, image generation, search). Each entry links straight to the file on Hugging Face and shows its size, licence and whether we have tested that exact file. docs/RUNNING.md then says how to run each one well: the command, the RAM it really needs and its measured speed.

What else it can do

The recipes above are the verified starting points. The engine covers much more, at varying levels of polish:

  • Language models: Llama, Qwen, Gemma, Mistral, Phi, DeepSeek, gpt-oss, GLM-4, Granite 4.0-H, Nemotron-H, LFM2 and many more GGUF architectures, on CPU, Vulkan or CUDA. Tool calling, JSON-schema constrained output, speculative decoding.
  • Images in, text out: Gemma, Qwen-VL, LLaVA, Pixtral, InternVL, OCR models such as dots.ocr.
  • Speech: text-to-speech engines including Kokoro, XTTS-v2 (voice cloning), Qwen3-TTS, Fish Speech and Chatterbox; speech recognition with Whisper, Parakeet and Qwen3-ASR.
  • Image and video generation: FLUX.1 / FLUX.2, Stable Diffusion 1.5 / XL / 3.5, Z-Image-Turbo, Qwen Image, Wan, LTX-Video, HunyuanVideo. These need several model files each; a guided setup is on the way.
  • Music and sound: Stable Audio 3, ACE-Step 1.5, MiniMax-Music3, MusicGen and AudioGen.

What is verified, what is partial and what is experimental is tracked per model, with dated evidence, in docs/STATUS.md. Speed comparisons between the speech engines are in the same file.

Hardware and speed

The timings in this README come from an AMD Ryzen 7 5700G (8 cores, integrated graphics only) with 64 GB of RAM, CPU only, on Windows 11. Rough guide:

You want toDownloadRAMOn that CPU
Chat with a small model0.5 GB2 GB~60-70 tokens/s (Qwen2.5 0.5B)
Chat with a 7–8B model4–5 GB8 GBa few tokens/s; a GPU helps a lot
Text to speech (Piper)63 MB< 1 GBfaster than real time
Speech to text (Whisper base)148 MB< 1 GB~2x real time
Image generation5–12 GB16 GB+minutes per image; a GPU is recommended

A GPU is optional: any Vulkan-capable card, or NVIDIA with CUDA 12.

What's next

A guided "front door" is under way in docs/103. Done: a small catalog of verified models per task and stingray setup <task> / stingray models. Next: task commands that need no file paths, and a one-line C# API.

Building from source

git clone https://github.com/opentail-net/OpenTail.Stingray && cd OpenTail.Stingray
dotnet build -c Release
dotnet run --project samples/QuickStart -c Release -- chat models/qwen2.5-0.5b-instruct-q4_k_m.gguf "Hello"

The documentation index sorts docs/ into user guides, design notes and active engineering work; contributors should read CLAUDE.md for the build and test conventions.

License

MIT — Copyright (c) 2026 OpenTail. Model files have their own licenses; check each model's page.

Languages

C#

99.0%