SmallMind is a production-ready, local language model inference engine built entirely in C#
C#
2
1,422 commits
updated May 19, 2026
Pure C# LLM inference runtime for CPU execution with zero native dependencies.
dotnet add package SmallMind
git clone https://github.com/justinamiller/SmallMind.git
cd SmallMind
dotnet build SmallMind.sln -c Release
using SmallMind;
// Create engine with model
var options = new SmallMindOptions
{
ModelPath = "model.smq", // or .gguf format
MaxContextTokens = 2048,
EnableKvCache = true
};
using var engine = SmallMindFactory.Create(options);
// Generate text
var sessionOpts = new TextGenerationOptions
{
Temperature = 0.7f,
MaxOutputTokens = 100
};
using var session = engine.CreateTextGenerationSession(sessionOpts);
var request = new TextGenerationRequest { Prompt = "The future of AI is".AsMemory() };
var result = session.Generate(request);
Console.WriteLine(result.Text);
Console.WriteLine($"Speed: {result.Timings.TokensPerSecond:F2} tok/s");
Streaming:
await foreach (var token in session.GenerateStreaming(request))
{
Console.Write(token.TokenText);
}
See examples/GoldenPath for complete working code.
| OS | Arch | SIMD Paths | Notes |
|---|---|---|---|
| Windows | x64 | AVX-512, AVX2, fallback | Tested in CI (windows-latest) |
| Linux | x64 | AVX-512, AVX2, fallback | Tested in CI (ubuntu-latest) |
| macOS | x64 | AVX2, fallback | Tested in CI (macos-latest) |
| macOS | arm64 | NEON, fallback | M1/M2/M3 Macs |
Runtime: .NET 10.0 or later
| Format | Purpose | Quantization Support |
|---|---|---|
.smq | SmallMind native format | F32, F16, Q8_0, Q4_0 |
.gguf | GGUF import (read-only) | 14 formats (see below) |
Floating point: F32, F16
Basic quant: Q4_0, Q4_1, Q5_0, Q5_1, Q8_0
K-quant: Q4_K, Q5_K, Q6_K, Q8_K
Note: Importance-weighted quantization (IQ2_XXS, IQ3_S, etc.) and Q2_K/Q3_K are not supported.
| Model | Params | FP32 | Q8_0 | Q4_0 |
|---|---|---|---|---|
| GPT-2 Small | 124M | 500 MB | 125 MB | 62 MB |
| GPT-2 Medium | 350M | 1.4 GB | 350 MB | 175 MB |
| GPT-2 Large | 774M | 3.1 GB | 775 MB | 388 MB |
| SmolLM-135M | 135M | 540 MB | 135 MB | 68 MB |
| LLaMA-1B | 1B | 4 GB | 1 GB | 500 MB |
Size limits: Models up to ~2B parameters (constrained by .NET array limits). See docs/LARGE_MODEL_SUPPORT.md.
Benchmark infrastructure: Automated CI on every PR (bench-ci.yml) and nightly extended runs (bench-nightly.yml).
Metrics measured:
Run benchmarks locally:
# Matrix multiplication benchmark
dotnet run --project benchmarks/specialized/MatMulBenchmark -c Release
# Full profiler suite
dotnet run --project benchmarks/specialized/ProfilerBenchmarks -c Release
# Inference allocation profiling
dotnet run --project benchmarks/specialized/InferenceAllocationBenchmark -c Release
Performance reports: See docs/PERFORMANCE_ANALYSIS_REPORT_2026-02-11.md and docs/benchmarks/.
Note: CI benchmarks use structural test fixtures, not full production models. Throughput estimates (15-25 tok/s for 7B Q4 models) are documented but not empirically validated in CI.
Entry point: SmallMindFactory.Create(SmallMindOptions) → returns ISmallMindEngine
Main abstractions:
ISmallMindEngine — Thread-safe factory for creating sessionsITextGenerationSession — Stateful generation context (not thread-safe)TextGenerationOptions — Sampling params (Temperature, TopP, TopK, StopSequences)TextGenerationRequest — Input promptGenerationResult — Output with text, usage stats, timingsDisposable pattern: Both engine and sessions implement IDisposable for resource cleanup.
See docs/PublicApi.md for complete API reference.
tools/ and examples/)See docs/roadmap/ for detailed plans.
Contributions welcome! See docs/CONTRIBUTING.md for guidelines.
Development setup:
git clone https://github.com/justinamiller/SmallMind.git
cd SmallMind
dotnet build SmallMind.sln -c Release
dotnet test
Report vulnerabilities via SECURITY.md.
Security tooling:
MIT License — see LICENSE
Copyright (c) 2024 Justin Miller
Project status: Production-ready for CPU inference with models up to ~2B parameters. Actively maintained.
C#
99.3%
SmallMind is a production-ready, local language model inference engine built entirely in C#
C#
2
1,422 commits
updated May 19, 2026
Pure C# LLM inference runtime for CPU execution with zero native dependencies.
dotnet add package SmallMind
git clone https://github.com/justinamiller/SmallMind.git
cd SmallMind
dotnet build SmallMind.sln -c Release
using SmallMind;
// Create engine with model
var options = new SmallMindOptions
{
ModelPath = "model.smq", // or .gguf format
MaxContextTokens = 2048,
EnableKvCache = true
};
using var engine = SmallMindFactory.Create(options);
// Generate text
var sessionOpts = new TextGenerationOptions
{
Temperature = 0.7f,
MaxOutputTokens = 100
};
using var session = engine.CreateTextGenerationSession(sessionOpts);
var request = new TextGenerationRequest { Prompt = "The future of AI is".AsMemory() };
var result = session.Generate(request);
Console.WriteLine(result.Text);
Console.WriteLine($"Speed: {result.Timings.TokensPerSecond:F2} tok/s");
Streaming:
await foreach (var token in session.GenerateStreaming(request))
{
Console.Write(token.TokenText);
}
See examples/GoldenPath for complete working code.
| OS | Arch | SIMD Paths | Notes |
|---|---|---|---|
| Windows | x64 | AVX-512, AVX2, fallback | Tested in CI (windows-latest) |
| Linux | x64 | AVX-512, AVX2, fallback | Tested in CI (ubuntu-latest) |
| macOS | x64 | AVX2, fallback | Tested in CI (macos-latest) |
| macOS | arm64 | NEON, fallback | M1/M2/M3 Macs |
Runtime: .NET 10.0 or later
| Format | Purpose | Quantization Support |
|---|---|---|
.smq | SmallMind native format | F32, F16, Q8_0, Q4_0 |
.gguf | GGUF import (read-only) | 14 formats (see below) |
Floating point: F32, F16
Basic quant: Q4_0, Q4_1, Q5_0, Q5_1, Q8_0
K-quant: Q4_K, Q5_K, Q6_K, Q8_K
Note: Importance-weighted quantization (IQ2_XXS, IQ3_S, etc.) and Q2_K/Q3_K are not supported.
| Model | Params | FP32 | Q8_0 | Q4_0 |
|---|---|---|---|---|
| GPT-2 Small | 124M | 500 MB | 125 MB | 62 MB |
| GPT-2 Medium | 350M | 1.4 GB | 350 MB | 175 MB |
| GPT-2 Large | 774M | 3.1 GB | 775 MB | 388 MB |
| SmolLM-135M | 135M | 540 MB | 135 MB | 68 MB |
| LLaMA-1B | 1B | 4 GB | 1 GB | 500 MB |
Size limits: Models up to ~2B parameters (constrained by .NET array limits). See docs/LARGE_MODEL_SUPPORT.md.
Benchmark infrastructure: Automated CI on every PR (bench-ci.yml) and nightly extended runs (bench-nightly.yml).
Metrics measured:
Run benchmarks locally:
# Matrix multiplication benchmark
dotnet run --project benchmarks/specialized/MatMulBenchmark -c Release
# Full profiler suite
dotnet run --project benchmarks/specialized/ProfilerBenchmarks -c Release
# Inference allocation profiling
dotnet run --project benchmarks/specialized/InferenceAllocationBenchmark -c Release
Performance reports: See docs/PERFORMANCE_ANALYSIS_REPORT_2026-02-11.md and docs/benchmarks/.
Note: CI benchmarks use structural test fixtures, not full production models. Throughput estimates (15-25 tok/s for 7B Q4 models) are documented but not empirically validated in CI.
Entry point: SmallMindFactory.Create(SmallMindOptions) → returns ISmallMindEngine
Main abstractions:
ISmallMindEngine — Thread-safe factory for creating sessionsITextGenerationSession — Stateful generation context (not thread-safe)TextGenerationOptions — Sampling params (Temperature, TopP, TopK, StopSequences)TextGenerationRequest — Input promptGenerationResult — Output with text, usage stats, timingsDisposable pattern: Both engine and sessions implement IDisposable for resource cleanup.
See docs/PublicApi.md for complete API reference.
tools/ and examples/)See docs/roadmap/ for detailed plans.
Contributions welcome! See docs/CONTRIBUTING.md for guidelines.
Development setup:
git clone https://github.com/justinamiller/SmallMind.git
cd SmallMind
dotnet build SmallMind.sln -c Release
dotnet test
Report vulnerabilities via SECURITY.md.
Security tooling:
MIT License — see LICENSE
Copyright (c) 2024 Justin Miller
Project status: Production-ready for CPU inference with models up to ~2B parameters. Actively maintained.
C#
99.3%