ooples/AiDotNet.Tensors

The fastest .NET tensor library. Beats MathNet (6x), NumSharp (3200x), matches TorchSharp CPU - pure managed C# with hand-tuned AVX2/FMA SIMD kernels. Optional CUDA/OpenCL GPU acceleration.

C#

12

751 commits

updated Oct 4, 2026

See the code

README

AiDotNet.Tensors

NuGet Build License

A high-performance .NET tensor library with hand-written AVX2/AVX-512 SIMD kernels in SimdKernels.cs / SimdGemm.cs / SimdConvHelper.cs. Every hot path runs through our own managed-C# kernels — we do NOT call into System.Numerics.Tensors, MKL.NET, or oneDNN through the standard wrappers. Beats ML.NET, TensorFlow.NET, MathNet, and NumSharp outright on every measured op. Against libtorch (TorchSharp's hand-tuned C++ kernels), wins on Mish 2.3×, Mish (double) 2.2×, GELU (double) 1.6× ahead, Tanh (double) within noise, Tanh (float) 1.4×, TensorMean/Min/Max, MaxPool2D, TensorAdd 100K, and TensorAdd 1M (vs single-thread torch) — all using pure managed C# with hand-tuned AVX2/FMA SIMD kernels and JIT-compiled machine code.

Note on dependencies. The .nupkg ships with the following PackageReferences: Microsoft.Extensions.Logging.Abstractions, System.Text.Json, System.Threading.Channels, K4os.Compression.LZ4 (LZ4 compression for serialized tensor blobs), AiDotNet.Native.OpenBLAS (transitive native OpenBLAS for fallback paths only — our SimdGemm beats it for d=128 transformer hot paths), and MKL via Microsoft.ML.Mkl.Redist (~66 MB on win-x64) + intelmkl.redist.win-x64 (~500 MB on win-x64) for the FP64 kernels that haven't yet been ported to pure-managed AVX2 (Phase 0 remediation work tracks the port). For air-gapped / federal deployments we ship a custom build with MKL/OpenBLAS removed and the entire telemetry namespace compiled out — see aidotnet.dev/enterprise for the Enterprise tier including air-gapped builds.

Performance numbers above assume net8.0+. On net471 the SIMD/intrinsics helpers are excluded (System.Runtime.Intrinsics is unavailable pre-net6); a custom net471 SIMD path that beats System.Numerics.Vector<T> is on the roadmap as Phase 5.

Features

  • Zero Allocations: In-place operations with ArrayPool<T> and Span<T> for hot paths
  • Hand-Tuned SIMD: Custom AVX2/FMA kernels with 4x loop unrolling, not just Vector<T> wrappers
  • JIT-Compiled Kernels: Runtime x86-64 machine code generation for size-specialized operations
  • BLIS-Style GEMM: Tiled matrix multiply with FMA micro-kernel, cache-aware panel packing
  • GPU Acceleration: Optional CUDA, HIP/ROCm, OpenCL, Metal, Vulkan, and WebGPU support via separate packages, with CPU-vs-GPU op-parity validated across every backend (#775)
  • Native ANN Index: Dependency-free approximate-nearest-neighbour search (Flat / IVF / PQ / IVFPQ) via AnnIndex, with fused IAnnBackend GPU kernels across all seven backends — no FAISS / MKL dependency (#824)
  • Multi-Target: Supports .NET 10.0 and .NET Framework 4.7.1
  • Generic Math: Works with any numeric type via INumericOperations<T> interface

Installation

# Core package (CPU SIMD acceleration)
dotnet add package AiDotNet.Tensors

# Optional: OpenBLAS for optimized CPU BLAS operations
dotnet add package AiDotNet.Native.OpenBLAS

# Optional: CLBlast for OpenCL GPU acceleration (AMD/Intel/NVIDIA)
dotnet add package AiDotNet.Native.CLBlast

# Optional: CUDA for NVIDIA GPU acceleration (requires NVIDIA GPU)
dotnet add package AiDotNet.Native.CUDA

Quick Start

using AiDotNet.Tensors.LinearAlgebra;

// Create vectors
var v1 = new Vector<double>(new[] { 1.0, 2.0, 3.0, 4.0 });
var v2 = new Vector<double>(new[] { 5.0, 6.0, 7.0, 8.0 });

// SIMD-accelerated operations
var sum = v1 + v2;
var dot = v1.Dot(v2);

// Create matrices
var m1 = new Matrix<double>(3, 3);
var m2 = Matrix<double>.Identity(3);

// Matrix operations
var product = m1 * m2;
var transpose = m1.Transpose();

CPU Benchmarks

All numbers from the latest BenchmarkDotNet run on AMD Ryzen 9 3950X (16 cores, AVX2/FMA, no AVX-512), .NET 10.0. Reproduce with:

dotnet run -c Release --project tests/AiDotNet.Tensors.Benchmarks --framework net10.0 -- --vs-all

The full per-op result set with error bars lives in tests/AiDotNet.Tensors.Benchmarks/BENCHMARK_RESULTS.md. The summary below is a hand-curated subset.

vs TorchSharp CPU (libtorch C++ backend)

Latest BDN run, post-#209 perf fixes — captured after removing System.Numerics.Tensors entirely and routing every hot path through our in-house SimdKernels. All comparisons are eager-vs-eager — neither side uses torch.compile or AiDotNet compiled plans, so this is libtorch's hand-rolled C++ kernels against AiDotNet's pure managed C# + AVX2 SIMD. See tests/AiDotNet.Tensors.Benchmarks/BENCHMARK_RESULTS.md for the full per-op table with error bars.

Big wins — AiDotNet beats TorchSharp by 2× or more:

OperationSizeAiDotNetTorchSharpSpeedup
Mish1M377 µs884 µs2.3× faster
Mish (double)1M1,038 µs2,313 µs2.2× faster

Wins — AiDotNet beats TorchSharp:

OperationSizeAiDotNetTorchSharpSpeedup
GELU (double)1M481 µs753 µs1.6× faster (was 3.6× behind!)
Tanh (double)1M586 µs627 µs1.07× faster (was 3.3× behind!)
Tanh (float)1M282 µs406 µs1.4× faster
TensorAdd100K33 µs42 µs1.3× faster
TensorMean1M189 µs243 µs1.3× faster
TensorAdd1M (vs 1-thread torch)350 µs468 µs1.3× vs 1-thread torch
MaxPool2D—250 µs285 µs1.1× faster
TensorMin1M205 µs215 µswithin noise (slight win)
TensorMultiply100K37 µs39 µswithin noise (slight win)

Closer-to-parity — AiDotNet within ~1.5× of libtorch:

OperationSizeAiDotNetTorchSharpRatio
ReLU1M261 µs191 µs1.4×
Sigmoid1M326 µs223 µs1.5×
TensorMaxValue1M195 µs189 µs1.03×
TensorExp1M296 µs306 µswithin noise
GELU (float)1M354 µs332 µs1.07×
TensorSum1M229 µs212 µs1.08×
TensorAbs1M362 µs221 µs1.6×
LeakyReLU1M409 µs273 µs1.5×
Exp (double)1M753 µs284 µs2.6× (was 4.3×)
Log (double)1M612 µs355 µs1.7× (was 16×!)

This PR's #209 close-parity wins — validated against the pre-fix baseline by fresh BDN re-runs and same-process micro-benchmarks:

OperationPre-fixPost-fixImprovement
Softmax_Double 512×10243,766 µs185 µs (slightly AHEAD of torch's 206!)20× faster
GELU_Double 1M2,782 µs481 µs (now 1.6× ahead of torch!)5.8× faster
Tanh_Double 1M2,067 µs586 µs (within noise of torch)3.5× faster
Log_Double 1M5,785 µs612 µs9.4× faster
Exp_Double 1M1,634 µs753 µs2.2× faster
LayerNorm 32k×641,347 µs890 µs1.5× faster
TensorAdd 1M480 µs350 µs1.4× faster
AttentionQKT 512×64599 µs419 µs (parallel-M pre-transpose)1.4× faster
AttentionQKT 512×128(not measured)451 µs(149 GFLOPS, parallel-M kernel)
MatMul 256³510 µs196 µs (parallel-M SgemmDirect)2.6× faster
MatMul 512³1,074 µs930 µs1.15× faster
Conv2D 1×16×64×64→32458 µs (regressed to 764 with naive 4-oc)397 µs (Auto policy picks PerChannel)back to baseline + 13%

Residual tracked gaps — areas where libtorch's Intel MKL-DNN (with AVX-512 inner kernels on Intel hardware) still wins. These need multi-day kernel rewrites (single-pass register-resident LayerNorm, fused QKᵀ attention kernel, BLIS-style 6×16 micro-kernel prefetch tuning) and are left as follow-up work:

OperationSizeAiDotNetTorchSharpRatio
TensorMatMul (float)256196 µs (parallel-M SgemmDirect)109 µs1.8× — was 4.7×
TensorMatMul (float)512930 µs534 µs1.7× — was 2.0×
LayerNorm32k×64890 µs303 µs2.9×
BatchNorm32×64×32×322,201 µs745 µs3.0×
Conv2D (float)1×16×64×64→32~397 µs (Auto picks PerChannel)310 µs1.3× — was 2.3× before A/B fix
Conv2D (double)4×3×32×32438 µs115 µs3.8× — unchanged this PR
AttentionQKT512×64419 µs (parallel-M pre-transpose)135 µs3.1× — was 4.3×
AttentionQKT512×128451 µs (parallel-M)—149 GFLOPS, was 1,102 µs
Softmax_Double 512×1024—185 µs206 µsslight win ✓ closed

Zero-external-dependency policy. Every hot path runs through our hand-tuned SimdKernels AVX2/AVX-512 implementations. We deliberately do NOT reference System.Numerics.Tensors, MKL, MKL.NET, or oneDNN — both for supply-chain hygiene and because we measured several TensorPrimitives entry points to regress 4–20× vs our in-house kernels on Ryzen 9 3950X (notably Tanh(float) 20× slower, Sigmoid(double) 12× slower, Log(double) 4× slower). All double-precision and single-precision paths now go through the same hand-tuned SIMD kernels — no fallback to any external library.

vs ML.NET (Microsoft.ML, eager-vs-eager)

Latest BDN run, validated post-#209-perf. Microsoft's general-purpose ML framework — same Ryzen 9 3950X, same .NET 10.0.7.

OperationSizeAiDotNetML.NETSpeedup
TensorMean1M80 µs180 µs2.2× faster
TensorSum1M92 µs104 µs1.1× faster
TensorAdd100K106 µs55 µs0.5× (memory-bound — ML.NET stayed allocator-warm)
TensorMultiply100K106 µs60 µs0.6× (memory-bound)
TensorAdd1M800 µs601 µs0.75× (memory-bound)
TensorMultiply1M782 µs595 µs0.76× (memory-bound)

The 1M-element bulk ops are memory-bandwidth-bound: at ~50 GB/s sustained DRAM bandwidth on Zen 2, a 4 MB read + 4 MB read + 4 MB write = 12 MB of traffic per call → 240 µs theoretical floor before any allocator overhead. Both libraries are within 2× of that floor.

vs TensorFlow.NET CPU (eager-vs-eager)

Latest BDN run, validated post-#209-perf. SciSharp's TensorFlow .NET binding (eager mode, no graph compile). Same hardware. AiDotNet wins outright on every measured op except small-Conv2D and 256×256 MatMul.

OperationSizeAiDotNetTensorFlow.NETSpeedup
TensorSum1M77 µs259 µs3.4× faster
TensorMean1M76 µs189 µs2.5× faster
TensorMultiply100K119 µs202 µs1.7× faster
Sigmoid1M1,264 µs1,941 µs1.5× faster
TensorAdd100K141 µs211 µs1.5× faster
TensorMatMul5121,286 µs1,554 µs1.2× faster
TensorAdd1M1,340 µs1,478 µs1.1× faster
ReLU1M1,680 µs1,606 µswithin noise (high stddev 713 µs)
TensorMultiply1M1,655 µs1,347 µs0.81× (memory-bound)
TensorMatMul256432 µs398 µs0.92×
Conv2D4×3×32×32719 µs428 µs0.6×

The fresh validation run captured full data on bulk Add/Multiply + 256/512 MatMul (the original fcb7fea baseline showed NA because SciSharp's TensorFlow.NET was crashing at those shapes; later runtime versions stabilized).

vs MathNet.Numerics (Linear Algebra, double, N=1000)

OperationAiDotNetMathNetSpeedup
Matrix Multiply 1000×10008.3 ms49.2 ms6× faster
Matrix Add1.87 ms2.50 ms1.3× faster
Matrix Subtract2.08 ms2.47 ms1.2× faster
Matrix Scalar Multiply1.66 ms2.14 ms1.3× faster
Transpose2.85 ms3.68 ms1.3× faster
Dot Product97 ns817 ns8.4× faster
L2 Norm92 ns11,552 ns125× faster

vs NumSharp (N=1000)

OperationAiDotNetNumSharpSpeedup
Matrix Multiply 1000×10008.3 ms26.5 s3,200× faster
Matrix Add1.87 ms1.98 ms1.1× faster
Transpose2.85 ms13.7 ms4.8× faster
Vector Add1.47 us54.5 us37× faster

vs System.Numerics.Tensors.TensorPrimitives (historical — REMOVED)

We previously referenced System.Numerics.Tensors and benchmarked our kernels against TensorPrimitives.* directly. As of #209 the dependency is removed entirely — every elementwise op now runs through our in-house SimdKernels, both for supply-chain hygiene and because we measured several TensorPrimitives entry points to regress 4–20× vs our in-house kernels on Ryzen 9 3950X (notably Tanh(float) ~20× slower, Sigmoid(double) ~12× slower, Log(double) ~4× slower).

OperationAiDotNetTensorPrimitives (raw)Speedup
Sigmoid (1M, float)284 µs7,295 µs25× faster
TensorAdd (100K, float)24 µs138 µs5.7× faster
TensorAdd (1M, float)379 µs614 µs1.6× faster
TensorSum (1M, float)196 µs298 µs1.5× faster
Dot Product (1K, double, in-place)97 ns185 ns1.9× faster
L2 Norm (1K, double, in-place)92 ns187 ns2.0× faster

Small Matrix Multiply (double)

SizeAiDotNetMathNetNumSharp
4×4172 ns165 ns2,198 ns
16×162.1 us2.9 us107.5 us
32×3210.5 us36.2 us774.8 us

AiDotNet is 1.4× faster at 16×16 and 3.4× faster at 32×32 than MathNet.

SIMD Instruction Support

The library automatically detects and uses the best available SIMD instructions:

Instruction SetVector WidthSupported
AVX-512512-bit (16 floats).NET 8+
AVX2 + FMA256-bit (8 floats).NET 6+
AVX256-bit (8 floats).NET 6+
SSE4.2128-bit (4 floats).NET 6+
ARM NEON128-bit (4 floats).NET 6+

Check Available Acceleration

using AiDotNet.Tensors.Engines;

var caps = PlatformDetector.Capabilities;

// SIMD capabilities
Console.WriteLine($"AVX2: {caps.HasAVX2}");
Console.WriteLine($"AVX-512: {caps.HasAVX512F}");

// GPU support
Console.WriteLine($"CUDA: {caps.HasCudaSupport}");
Console.WriteLine($"OpenCL: {caps.HasOpenCLSupport}");

// Native library availability
Console.WriteLine($"OpenBLAS: {caps.HasOpenBlas}");
Console.WriteLine($"CLBlast: {caps.HasClBlast}");

// Or get a full status summary
Console.WriteLine(NativeLibraryDetector.GetStatusSummary());

Optional Acceleration Packages

AiDotNet.Native.OpenBLAS

Provides optimized CPU BLAS operations using OpenBLAS:

dotnet add package AiDotNet.Native.OpenBLAS

Performance: Accelerated BLAS operations for matrix multiply and decompositions.

AiDotNet.Native.CLBlast

Provides GPU acceleration via OpenCL (works on AMD, Intel, and NVIDIA GPUs):

dotnet add package AiDotNet.Native.CLBlast

Performance: 10x+ faster for large matrix operations on GPU.

AiDotNet.Native.CUDA

Provides GPU acceleration via NVIDIA CUDA (NVIDIA GPUs only):

dotnet add package AiDotNet.Native.CUDA

Performance: 30,000+ GFLOPS for matrix operations on modern NVIDIA GPUs.

Requirements:

  • NVIDIA GPU (GeForce, Quadro, or Tesla)
  • NVIDIA display driver 525.60+ (includes CUDA driver)

Usage with helpful error messages:

using AiDotNet.Tensors.Engines.DirectGpu.CUDA;

// Recommended: throws beginner-friendly exception if CUDA unavailable
using var cuda = CudaBackend.CreateOrThrow();

// Or check availability first
if (CudaBackend.IsCudaAvailable)
{
    using var backend = new CudaBackend();
    // Use CUDA acceleration
}

If CUDA is not available, you'll get detailed troubleshooting steps explaining exactly what's missing and how to fix it.

Requirements

  • .NET 10.0 or .NET Framework 4.7.1+
  • Windows x64, Linux x64, or macOS x64/arm64

License

Apache 2.0 - See LICENSE for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

ooples/AiDotNet.Tensors

The fastest .NET tensor library. Beats MathNet (6x), NumSharp (3200x), matches TorchSharp CPU - pure managed C# with hand-tuned AVX2/FMA SIMD kernels. Optional CUDA/OpenCL GPU acceleration.

C#

12

751 commits

updated Oct 4, 2026

See the code

README

AiDotNet.Tensors

NuGet Build License

A high-performance .NET tensor library with hand-written AVX2/AVX-512 SIMD kernels in SimdKernels.cs / SimdGemm.cs / SimdConvHelper.cs. Every hot path runs through our own managed-C# kernels — we do NOT call into System.Numerics.Tensors, MKL.NET, or oneDNN through the standard wrappers. Beats ML.NET, TensorFlow.NET, MathNet, and NumSharp outright on every measured op. Against libtorch (TorchSharp's hand-tuned C++ kernels), wins on Mish 2.3×, Mish (double) 2.2×, GELU (double) 1.6× ahead, Tanh (double) within noise, Tanh (float) 1.4×, TensorMean/Min/Max, MaxPool2D, TensorAdd 100K, and TensorAdd 1M (vs single-thread torch) — all using pure managed C# with hand-tuned AVX2/FMA SIMD kernels and JIT-compiled machine code.

Note on dependencies. The .nupkg ships with the following PackageReferences: Microsoft.Extensions.Logging.Abstractions, System.Text.Json, System.Threading.Channels, K4os.Compression.LZ4 (LZ4 compression for serialized tensor blobs), AiDotNet.Native.OpenBLAS (transitive native OpenBLAS for fallback paths only — our SimdGemm beats it for d=128 transformer hot paths), and MKL via Microsoft.ML.Mkl.Redist (~66 MB on win-x64) + intelmkl.redist.win-x64 (~500 MB on win-x64) for the FP64 kernels that haven't yet been ported to pure-managed AVX2 (Phase 0 remediation work tracks the port). For air-gapped / federal deployments we ship a custom build with MKL/OpenBLAS removed and the entire telemetry namespace compiled out — see aidotnet.dev/enterprise for the Enterprise tier including air-gapped builds.

Performance numbers above assume net8.0+. On net471 the SIMD/intrinsics helpers are excluded (System.Runtime.Intrinsics is unavailable pre-net6); a custom net471 SIMD path that beats System.Numerics.Vector<T> is on the roadmap as Phase 5.

Features

  • Zero Allocations: In-place operations with ArrayPool<T> and Span<T> for hot paths
  • Hand-Tuned SIMD: Custom AVX2/FMA kernels with 4x loop unrolling, not just Vector<T> wrappers
  • JIT-Compiled Kernels: Runtime x86-64 machine code generation for size-specialized operations
  • BLIS-Style GEMM: Tiled matrix multiply with FMA micro-kernel, cache-aware panel packing
  • GPU Acceleration: Optional CUDA, HIP/ROCm, OpenCL, Metal, Vulkan, and WebGPU support via separate packages, with CPU-vs-GPU op-parity validated across every backend (#775)
  • Native ANN Index: Dependency-free approximate-nearest-neighbour search (Flat / IVF / PQ / IVFPQ) via AnnIndex, with fused IAnnBackend GPU kernels across all seven backends — no FAISS / MKL dependency (#824)
  • Multi-Target: Supports .NET 10.0 and .NET Framework 4.7.1
  • Generic Math: Works with any numeric type via INumericOperations<T> interface

Installation

# Core package (CPU SIMD acceleration)
dotnet add package AiDotNet.Tensors

# Optional: OpenBLAS for optimized CPU BLAS operations
dotnet add package AiDotNet.Native.OpenBLAS

# Optional: CLBlast for OpenCL GPU acceleration (AMD/Intel/NVIDIA)
dotnet add package AiDotNet.Native.CLBlast

# Optional: CUDA for NVIDIA GPU acceleration (requires NVIDIA GPU)
dotnet add package AiDotNet.Native.CUDA

Quick Start

using AiDotNet.Tensors.LinearAlgebra;

// Create vectors
var v1 = new Vector<double>(new[] { 1.0, 2.0, 3.0, 4.0 });
var v2 = new Vector<double>(new[] { 5.0, 6.0, 7.0, 8.0 });

// SIMD-accelerated operations
var sum = v1 + v2;
var dot = v1.Dot(v2);

// Create matrices
var m1 = new Matrix<double>(3, 3);
var m2 = Matrix<double>.Identity(3);

// Matrix operations
var product = m1 * m2;
var transpose = m1.Transpose();

CPU Benchmarks

All numbers from the latest BenchmarkDotNet run on AMD Ryzen 9 3950X (16 cores, AVX2/FMA, no AVX-512), .NET 10.0. Reproduce with:

dotnet run -c Release --project tests/AiDotNet.Tensors.Benchmarks --framework net10.0 -- --vs-all

The full per-op result set with error bars lives in tests/AiDotNet.Tensors.Benchmarks/BENCHMARK_RESULTS.md. The summary below is a hand-curated subset.

vs TorchSharp CPU (libtorch C++ backend)

Latest BDN run, post-#209 perf fixes — captured after removing System.Numerics.Tensors entirely and routing every hot path through our in-house SimdKernels. All comparisons are eager-vs-eager — neither side uses torch.compile or AiDotNet compiled plans, so this is libtorch's hand-rolled C++ kernels against AiDotNet's pure managed C# + AVX2 SIMD. See tests/AiDotNet.Tensors.Benchmarks/BENCHMARK_RESULTS.md for the full per-op table with error bars.

Big wins — AiDotNet beats TorchSharp by 2× or more:

OperationSizeAiDotNetTorchSharpSpeedup
Mish1M377 µs884 µs2.3× faster
Mish (double)1M1,038 µs2,313 µs2.2× faster

Wins — AiDotNet beats TorchSharp:

OperationSizeAiDotNetTorchSharpSpeedup
GELU (double)1M481 µs753 µs1.6× faster (was 3.6× behind!)
Tanh (double)1M586 µs627 µs1.07× faster (was 3.3× behind!)
Tanh (float)1M282 µs406 µs1.4× faster
TensorAdd100K33 µs42 µs1.3× faster
TensorMean1M189 µs243 µs1.3× faster
TensorAdd1M (vs 1-thread torch)350 µs468 µs1.3× vs 1-thread torch
MaxPool2D—250 µs285 µs1.1× faster
TensorMin1M205 µs215 µswithin noise (slight win)
TensorMultiply100K37 µs39 µswithin noise (slight win)

Closer-to-parity — AiDotNet within ~1.5× of libtorch:

OperationSizeAiDotNetTorchSharpRatio
ReLU1M261 µs191 µs1.4×
Sigmoid1M326 µs223 µs1.5×
TensorMaxValue1M195 µs189 µs1.03×
TensorExp1M296 µs306 µswithin noise
GELU (float)1M354 µs332 µs1.07×
TensorSum1M229 µs212 µs1.08×
TensorAbs1M362 µs221 µs1.6×
LeakyReLU1M409 µs273 µs1.5×
Exp (double)1M753 µs284 µs2.6× (was 4.3×)
Log (double)1M612 µs355 µs1.7× (was 16×!)

This PR's #209 close-parity wins — validated against the pre-fix baseline by fresh BDN re-runs and same-process micro-benchmarks:

OperationPre-fixPost-fixImprovement
Softmax_Double 512×10243,766 µs185 µs (slightly AHEAD of torch's 206!)20× faster
GELU_Double 1M2,782 µs481 µs (now 1.6× ahead of torch!)5.8× faster
Tanh_Double 1M2,067 µs586 µs (within noise of torch)3.5× faster
Log_Double 1M5,785 µs612 µs9.4× faster
Exp_Double 1M1,634 µs753 µs2.2× faster
LayerNorm 32k×641,347 µs890 µs1.5× faster
TensorAdd 1M480 µs350 µs1.4× faster
AttentionQKT 512×64599 µs419 µs (parallel-M pre-transpose)1.4× faster
AttentionQKT 512×128(not measured)451 µs(149 GFLOPS, parallel-M kernel)
MatMul 256³510 µs196 µs (parallel-M SgemmDirect)2.6× faster
MatMul 512³1,074 µs930 µs1.15× faster
Conv2D 1×16×64×64→32458 µs (regressed to 764 with naive 4-oc)397 µs (Auto policy picks PerChannel)back to baseline + 13%

Residual tracked gaps — areas where libtorch's Intel MKL-DNN (with AVX-512 inner kernels on Intel hardware) still wins. These need multi-day kernel rewrites (single-pass register-resident LayerNorm, fused QKᵀ attention kernel, BLIS-style 6×16 micro-kernel prefetch tuning) and are left as follow-up work:

OperationSizeAiDotNetTorchSharpRatio
TensorMatMul (float)256196 µs (parallel-M SgemmDirect)109 µs1.8× — was 4.7×
TensorMatMul (float)512930 µs534 µs1.7× — was 2.0×
LayerNorm32k×64890 µs303 µs2.9×
BatchNorm32×64×32×322,201 µs745 µs3.0×
Conv2D (float)1×16×64×64→32~397 µs (Auto picks PerChannel)310 µs1.3× — was 2.3× before A/B fix
Conv2D (double)4×3×32×32438 µs115 µs3.8× — unchanged this PR
AttentionQKT512×64419 µs (parallel-M pre-transpose)135 µs3.1× — was 4.3×
AttentionQKT512×128451 µs (parallel-M)—149 GFLOPS, was 1,102 µs
Softmax_Double 512×1024—185 µs206 µsslight win ✓ closed

Zero-external-dependency policy. Every hot path runs through our hand-tuned SimdKernels AVX2/AVX-512 implementations. We deliberately do NOT reference System.Numerics.Tensors, MKL, MKL.NET, or oneDNN — both for supply-chain hygiene and because we measured several TensorPrimitives entry points to regress 4–20× vs our in-house kernels on Ryzen 9 3950X (notably Tanh(float) 20× slower, Sigmoid(double) 12× slower, Log(double) 4× slower). All double-precision and single-precision paths now go through the same hand-tuned SIMD kernels — no fallback to any external library.

vs ML.NET (Microsoft.ML, eager-vs-eager)

Latest BDN run, validated post-#209-perf. Microsoft's general-purpose ML framework — same Ryzen 9 3950X, same .NET 10.0.7.

OperationSizeAiDotNetML.NETSpeedup
TensorMean1M80 µs180 µs2.2× faster
TensorSum1M92 µs104 µs1.1× faster
TensorAdd100K106 µs55 µs0.5× (memory-bound — ML.NET stayed allocator-warm)
TensorMultiply100K106 µs60 µs0.6× (memory-bound)
TensorAdd1M800 µs601 µs0.75× (memory-bound)
TensorMultiply1M782 µs595 µs0.76× (memory-bound)

The 1M-element bulk ops are memory-bandwidth-bound: at ~50 GB/s sustained DRAM bandwidth on Zen 2, a 4 MB read + 4 MB read + 4 MB write = 12 MB of traffic per call → 240 µs theoretical floor before any allocator overhead. Both libraries are within 2× of that floor.

vs TensorFlow.NET CPU (eager-vs-eager)

Latest BDN run, validated post-#209-perf. SciSharp's TensorFlow .NET binding (eager mode, no graph compile). Same hardware. AiDotNet wins outright on every measured op except small-Conv2D and 256×256 MatMul.

OperationSizeAiDotNetTensorFlow.NETSpeedup
TensorSum1M77 µs259 µs3.4× faster
TensorMean1M76 µs189 µs2.5× faster
TensorMultiply100K119 µs202 µs1.7× faster
Sigmoid1M1,264 µs1,941 µs1.5× faster
TensorAdd100K141 µs211 µs1.5× faster
TensorMatMul5121,286 µs1,554 µs1.2× faster
TensorAdd1M1,340 µs1,478 µs1.1× faster
ReLU1M1,680 µs1,606 µswithin noise (high stddev 713 µs)
TensorMultiply1M1,655 µs1,347 µs0.81× (memory-bound)
TensorMatMul256432 µs398 µs0.92×
Conv2D4×3×32×32719 µs428 µs0.6×

The fresh validation run captured full data on bulk Add/Multiply + 256/512 MatMul (the original fcb7fea baseline showed NA because SciSharp's TensorFlow.NET was crashing at those shapes; later runtime versions stabilized).

vs MathNet.Numerics (Linear Algebra, double, N=1000)

OperationAiDotNetMathNetSpeedup
Matrix Multiply 1000×10008.3 ms49.2 ms6× faster
Matrix Add1.87 ms2.50 ms1.3× faster
Matrix Subtract2.08 ms2.47 ms1.2× faster
Matrix Scalar Multiply1.66 ms2.14 ms1.3× faster
Transpose2.85 ms3.68 ms1.3× faster
Dot Product97 ns817 ns8.4× faster
L2 Norm92 ns11,552 ns125× faster

vs NumSharp (N=1000)

OperationAiDotNetNumSharpSpeedup
Matrix Multiply 1000×10008.3 ms26.5 s3,200× faster
Matrix Add1.87 ms1.98 ms1.1× faster
Transpose2.85 ms13.7 ms4.8× faster
Vector Add1.47 us54.5 us37× faster

vs System.Numerics.Tensors.TensorPrimitives (historical — REMOVED)

We previously referenced System.Numerics.Tensors and benchmarked our kernels against TensorPrimitives.* directly. As of #209 the dependency is removed entirely — every elementwise op now runs through our in-house SimdKernels, both for supply-chain hygiene and because we measured several TensorPrimitives entry points to regress 4–20× vs our in-house kernels on Ryzen 9 3950X (notably Tanh(float) ~20× slower, Sigmoid(double) ~12× slower, Log(double) ~4× slower).

OperationAiDotNetTensorPrimitives (raw)Speedup
Sigmoid (1M, float)284 µs7,295 µs25× faster
TensorAdd (100K, float)24 µs138 µs5.7× faster
TensorAdd (1M, float)379 µs614 µs1.6× faster
TensorSum (1M, float)196 µs298 µs1.5× faster
Dot Product (1K, double, in-place)97 ns185 ns1.9× faster
L2 Norm (1K, double, in-place)92 ns187 ns2.0× faster

Small Matrix Multiply (double)

SizeAiDotNetMathNetNumSharp
4×4172 ns165 ns2,198 ns
16×162.1 us2.9 us107.5 us
32×3210.5 us36.2 us774.8 us

AiDotNet is 1.4× faster at 16×16 and 3.4× faster at 32×32 than MathNet.

SIMD Instruction Support

The library automatically detects and uses the best available SIMD instructions:

Instruction SetVector WidthSupported
AVX-512512-bit (16 floats).NET 8+
AVX2 + FMA256-bit (8 floats).NET 6+
AVX256-bit (8 floats).NET 6+
SSE4.2128-bit (4 floats).NET 6+
ARM NEON128-bit (4 floats).NET 6+

Check Available Acceleration

using AiDotNet.Tensors.Engines;

var caps = PlatformDetector.Capabilities;

// SIMD capabilities
Console.WriteLine($"AVX2: {caps.HasAVX2}");
Console.WriteLine($"AVX-512: {caps.HasAVX512F}");

// GPU support
Console.WriteLine($"CUDA: {caps.HasCudaSupport}");
Console.WriteLine($"OpenCL: {caps.HasOpenCLSupport}");

// Native library availability
Console.WriteLine($"OpenBLAS: {caps.HasOpenBlas}");
Console.WriteLine($"CLBlast: {caps.HasClBlast}");

// Or get a full status summary
Console.WriteLine(NativeLibraryDetector.GetStatusSummary());

Optional Acceleration Packages

AiDotNet.Native.OpenBLAS

Provides optimized CPU BLAS operations using OpenBLAS:

dotnet add package AiDotNet.Native.OpenBLAS

Performance: Accelerated BLAS operations for matrix multiply and decompositions.

AiDotNet.Native.CLBlast

Provides GPU acceleration via OpenCL (works on AMD, Intel, and NVIDIA GPUs):

dotnet add package AiDotNet.Native.CLBlast

Performance: 10x+ faster for large matrix operations on GPU.

AiDotNet.Native.CUDA

Provides GPU acceleration via NVIDIA CUDA (NVIDIA GPUs only):

dotnet add package AiDotNet.Native.CUDA

Performance: 30,000+ GFLOPS for matrix operations on modern NVIDIA GPUs.

Requirements:

  • NVIDIA GPU (GeForce, Quadro, or Tesla)
  • NVIDIA display driver 525.60+ (includes CUDA driver)

Usage with helpful error messages:

using AiDotNet.Tensors.Engines.DirectGpu.CUDA;

// Recommended: throws beginner-friendly exception if CUDA unavailable
using var cuda = CudaBackend.CreateOrThrow();

// Or check availability first
if (CudaBackend.IsCudaAvailable)
{
    using var backend = new CudaBackend();
    // Use CUDA acceleration
}

If CUDA is not available, you'll get detailed troubleshooting steps explaining exactly what's missing and how to fix it.

Requirements

  • .NET 10.0 or .NET Framework 4.7.1+
  • Windows x64, Linux x64, or macOS x64/arm64

License

Apache 2.0 - See LICENSE for details.

Contributing

Contributions are welcome! Please feel free to submit a Pull Request.

Languages

C#

98.9%