A high-performance, columnar DataFrame library for .NET
See the codeA high-performance, columnar DataFrame library for .NET, focused on type safety, explicit null semantics, query planning, and clean interop with platform tensor and data APIs.
Nivara is designed for developers who want predictable behavior, strong typing, and performance-oriented data processing without relying on dynamic or NaN-based conventions.
Most DataFrame-style libraries trade correctness and type safety for convenience. Nivara takes a different approach:
System.Numerics.Tensors for tensor mathIf you care about correctness, debuggability, and performance in .NET data processing, Nivara is built for you.
Core library:
dotnet add package Nivara
Optional extensions and I/O integrations (install when you need file formats, Arrow interoperability, or ML integration):
dotnet add package Nivara.Extensions
using Nivara;
using Nivara.Linq;
// Create typed columns
NivaraColumn<int> ages = [25, 30, 35];
var names = NivaraColumn<string>.CreateForReferenceType(new[] { "Alice", "Bob", "Charlie" });
// Combine into a DataFrame
var frame = NivaraFrame.Create(
("Name", names),
("Age", ages)
);
// Query with lazy evaluation — strongly typed lambdas over a POCO
public sealed class Person { public string Name { get; set; } public int Age { get; set; } }
var typed = frame.Query<Person>()
.Where(p => p.Age > 30)
.Select(p => new { p.Name })
.ToObjects(); // IReadOnlyList<anonymous> — { Name = "Charlie" }
// Or materialize to a NivaraFrame
var adults = frame.Query<Person>()
.Where(p => p.Age > 30)
.Collect(); // NivaraFrame — 1 row (Charlie)
frame.Query<T>() maps a POCO to the frame schema and compiles typed lambdas into query plans (predicates, projections, conditional expressions, OrderBy/ThenBy with per-key SortDirection/NullOrdering, Distinct/DistinctBy, SelectRows, Skip/Take, GroupBy with g.Key + Average/Sum/Count/Min/Max/Quantile/Median/StdDev/Variance aggregates), materializing to a NivaraFrame or IReadOnlyList<TResult>Over() / WindowSpec builder for SQL-style partitioned windows: rolling (Sum/Mean/Min/Max), cumulative (Sum/Max/Min/Product/Count), Shift/Lead, and rank family (RowNumber/Rank/DenseRank/PercentRank), on both eager NivaraFrame and lazy QueryFrameQueryFrame.AsStream(chunkSize) and NivaraQuery<T>.AsStream yield one NivaraFrame per source chunk for async processing; ScanAsQueryFrame factories open streaming directly from CSV/JSON/Parquet filesJson.ScanQuery<T>() (core) and Csv.ScanQuery<T>() (Extensions) defer I/O until execution; ReadFrame/ScanFrame cover eager/lazy frame loadingCollectAsync/ToListAsync and integrated performance diagnosticsTensor<T> for platform math APIsNullableTensor<T> when crossing tensor boundariesSystem.Numerics.Tensors, not custom DataFrame APIsGradientUtils.Grad(), and module state can be copied via StateDict() / LoadStateDict()IFloatingPointIeee754<T> constraint — Half/F16 and BFloat16 now pass runtime validation alongside float and doubleEmbedding<T>, SparseEmbedding<T>, Conv1d<T> (im2col-rewritten, PyTorch-compatible layout), Conv2d<T> (grouped conv, 1×1 fast path, PatchLocation lookup, InputGrad specializations), ConvTranspose2d<T>, BatchNorm1d<T> (now accepts 3D [B,C,L] input), BatchNorm2d<T>, LayerNorm<T> (SIMD via TensorPrimitives.Dot), DepthwiseSeparableConv2d<T>, TransformerBlock<T> (RMSNorm/LayerNorm + GELU), MultiheadAttention<T> (self/cross/causal), ConvVAE<T>, VAE<T> (optional conditioning), MaxPool2d<T>, AdaptiveAvgPool2d<T>, GELU, TextTokenizer, and Sampler<T> — all differentiable and composable with the existing module system (ready-to-use TextClassifierModel<T> / TokenClassifierModel<T> ship as sample code in samples/Nivara.Samples/)ScanAsQueryFrame for lazy streaming entry pointsNivara.Extensions)Nivara.Extensions)CollectAsync/ToListAsync run genuinely asynchronously with cancellation supportQueryPlan, QueryPlanAnalyzer, QueryDiagnostics, ExecutionEngine, ExecutionProgress — all public)For detailed examples and tutorials, see GETTING-STARTED.md.
For comprehensive API documentation and advanced usage patterns, explore the samples/ directory — including a character-level GPT trained on Nivara AutoDiff, a neural chess evaluator, a hybrid Nivara+LLM agent workflow, a variational autoencoder for synthetic pattern generation, a PyTorch parity benchmark suite showing <0.04% loss-curve divergence, a MiniLM inference pipeline, a DistilBERT fine-tuning pipeline for SST-2 (samples/NivaraFineTuning), a MobileNetV2/ResNet-18 inference pipeline (samples/NivaraInference), and a time-series anomaly detection sample (samples/NivaraTimeSeries).
Nivara aims to bring predictable, high-performance data processing to the .NET ecosystem — without sacrificing correctness or clarity.
Nivara currently supports:
Tensor<T> and nullable tensor conversion helpers, plus matrix/labeled-row ingestionOperationType constants, diagnostics and plan inspectionframe.Query<T>() with eager POCO→column mapping, typed predicates/projections, GroupBy aggregates, and row-factory materialization (Collect/ToList → NivaraFrame, ToObjects/ToRows → IReadOnlyList<TResult>); unsupported expressions fail fast with UnsupportedQueryExpressionExceptionNivara.Extensions)Nivara.Extensions)Nivara.Extensions)ExplainPlan(), per-operation timings)GradientUtils.Grad(). Type constraint broadened to IFloatingPointIeee754<T> — Half/F16 and BFloat16 supported alongside float/double. Full training stack: module system (Linear, Sequential, Embedding, SparseEmbedding, Conv1d (im2col + Dot, PyTorch-compatible layout), Conv2d (grouped conv, 1×1 fast path, PatchLocation, InputGrad specializations), ConvTranspose2d, BatchNorm1d/2d (fused span-kernel, 3D input support), LayerNorm (SIMD TensorPrimitives.Dot), DepthwiseSeparableConv2d, TransformerBlock (RMSNorm/LayerNorm + GELU), MultiheadAttention, ConvVAE, VAE (optional conditioning), MaxPool2d, AdaptiveAvgPool2d), NLP utilities (TextTokenizer, Sampler), activations (GELU), operations (MeanPool, TransposeAxes, SparseEmbeddingBag, Gather with zero-copy forward, Softmax, LogSoftmax, Dropout), optimizers (SGD, Adam, AdamW) with SIMD-accelerated kernels, training loops, data-parallel training, model serialization, and 55 PyTorch-validated functional testsBFloat16 is a first-class numeric column type (mirroring Half) — vectorized arithmetic with null-mask preservation, window functions, sorting, aggregation (Sum/Mean/Quantile → double), and fused query expressions (BFloat16 + int promotes to double). It is also a first-class AutoDiff type; see BFLOAT16.mdC#
99.8%
A high-performance, columnar DataFrame library for .NET
See the codeA high-performance, columnar DataFrame library for .NET, focused on type safety, explicit null semantics, query planning, and clean interop with platform tensor and data APIs.
Nivara is designed for developers who want predictable behavior, strong typing, and performance-oriented data processing without relying on dynamic or NaN-based conventions.
Most DataFrame-style libraries trade correctness and type safety for convenience. Nivara takes a different approach:
System.Numerics.Tensors for tensor mathIf you care about correctness, debuggability, and performance in .NET data processing, Nivara is built for you.
Core library:
dotnet add package Nivara
Optional extensions and I/O integrations (install when you need file formats, Arrow interoperability, or ML integration):
dotnet add package Nivara.Extensions
using Nivara;
using Nivara.Linq;
// Create typed columns
NivaraColumn<int> ages = [25, 30, 35];
var names = NivaraColumn<string>.CreateForReferenceType(new[] { "Alice", "Bob", "Charlie" });
// Combine into a DataFrame
var frame = NivaraFrame.Create(
("Name", names),
("Age", ages)
);
// Query with lazy evaluation — strongly typed lambdas over a POCO
public sealed class Person { public string Name { get; set; } public int Age { get; set; } }
var typed = frame.Query<Person>()
.Where(p => p.Age > 30)
.Select(p => new { p.Name })
.ToObjects(); // IReadOnlyList<anonymous> — { Name = "Charlie" }
// Or materialize to a NivaraFrame
var adults = frame.Query<Person>()
.Where(p => p.Age > 30)
.Collect(); // NivaraFrame — 1 row (Charlie)
frame.Query<T>() maps a POCO to the frame schema and compiles typed lambdas into query plans (predicates, projections, conditional expressions, OrderBy/ThenBy with per-key SortDirection/NullOrdering, Distinct/DistinctBy, SelectRows, Skip/Take, GroupBy with g.Key + Average/Sum/Count/Min/Max/Quantile/Median/StdDev/Variance aggregates), materializing to a NivaraFrame or IReadOnlyList<TResult>Over() / WindowSpec builder for SQL-style partitioned windows: rolling (Sum/Mean/Min/Max), cumulative (Sum/Max/Min/Product/Count), Shift/Lead, and rank family (RowNumber/Rank/DenseRank/PercentRank), on both eager NivaraFrame and lazy QueryFrameQueryFrame.AsStream(chunkSize) and NivaraQuery<T>.AsStream yield one NivaraFrame per source chunk for async processing; ScanAsQueryFrame factories open streaming directly from CSV/JSON/Parquet filesJson.ScanQuery<T>() (core) and Csv.ScanQuery<T>() (Extensions) defer I/O until execution; ReadFrame/ScanFrame cover eager/lazy frame loadingCollectAsync/ToListAsync and integrated performance diagnosticsTensor<T> for platform math APIsNullableTensor<T> when crossing tensor boundariesSystem.Numerics.Tensors, not custom DataFrame APIsGradientUtils.Grad(), and module state can be copied via StateDict() / LoadStateDict()IFloatingPointIeee754<T> constraint — Half/F16 and BFloat16 now pass runtime validation alongside float and doubleEmbedding<T>, SparseEmbedding<T>, Conv1d<T> (im2col-rewritten, PyTorch-compatible layout), Conv2d<T> (grouped conv, 1×1 fast path, PatchLocation lookup, InputGrad specializations), ConvTranspose2d<T>, BatchNorm1d<T> (now accepts 3D [B,C,L] input), BatchNorm2d<T>, LayerNorm<T> (SIMD via TensorPrimitives.Dot), DepthwiseSeparableConv2d<T>, TransformerBlock<T> (RMSNorm/LayerNorm + GELU), MultiheadAttention<T> (self/cross/causal), ConvVAE<T>, VAE<T> (optional conditioning), MaxPool2d<T>, AdaptiveAvgPool2d<T>, GELU, TextTokenizer, and Sampler<T> — all differentiable and composable with the existing module system (ready-to-use TextClassifierModel<T> / TokenClassifierModel<T> ship as sample code in samples/Nivara.Samples/)ScanAsQueryFrame for lazy streaming entry pointsNivara.Extensions)Nivara.Extensions)CollectAsync/ToListAsync run genuinely asynchronously with cancellation supportQueryPlan, QueryPlanAnalyzer, QueryDiagnostics, ExecutionEngine, ExecutionProgress — all public)For detailed examples and tutorials, see GETTING-STARTED.md.
For comprehensive API documentation and advanced usage patterns, explore the samples/ directory — including a character-level GPT trained on Nivara AutoDiff, a neural chess evaluator, a hybrid Nivara+LLM agent workflow, a variational autoencoder for synthetic pattern generation, a PyTorch parity benchmark suite showing <0.04% loss-curve divergence, a MiniLM inference pipeline, a DistilBERT fine-tuning pipeline for SST-2 (samples/NivaraFineTuning), a MobileNetV2/ResNet-18 inference pipeline (samples/NivaraInference), and a time-series anomaly detection sample (samples/NivaraTimeSeries).
Nivara aims to bring predictable, high-performance data processing to the .NET ecosystem — without sacrificing correctness or clarity.
Nivara currently supports:
Tensor<T> and nullable tensor conversion helpers, plus matrix/labeled-row ingestionOperationType constants, diagnostics and plan inspectionframe.Query<T>() with eager POCO→column mapping, typed predicates/projections, GroupBy aggregates, and row-factory materialization (Collect/ToList → NivaraFrame, ToObjects/ToRows → IReadOnlyList<TResult>); unsupported expressions fail fast with UnsupportedQueryExpressionExceptionNivara.Extensions)Nivara.Extensions)Nivara.Extensions)ExplainPlan(), per-operation timings)GradientUtils.Grad(). Type constraint broadened to IFloatingPointIeee754<T> — Half/F16 and BFloat16 supported alongside float/double. Full training stack: module system (Linear, Sequential, Embedding, SparseEmbedding, Conv1d (im2col + Dot, PyTorch-compatible layout), Conv2d (grouped conv, 1×1 fast path, PatchLocation, InputGrad specializations), ConvTranspose2d, BatchNorm1d/2d (fused span-kernel, 3D input support), LayerNorm (SIMD TensorPrimitives.Dot), DepthwiseSeparableConv2d, TransformerBlock (RMSNorm/LayerNorm + GELU), MultiheadAttention, ConvVAE, VAE (optional conditioning), MaxPool2d, AdaptiveAvgPool2d), NLP utilities (TextTokenizer, Sampler), activations (GELU), operations (MeanPool, TransposeAxes, SparseEmbeddingBag, Gather with zero-copy forward, Softmax, LogSoftmax, Dropout), optimizers (SGD, Adam, AdamW) with SIMD-accelerated kernels, training loops, data-parallel training, model serialization, and 55 PyTorch-validated functional testsBFloat16 is a first-class numeric column type (mirroring Half) — vectorized arithmetic with null-mask preservation, window functions, sorting, aggregation (Sum/Mean/Quantile → double), and fused query expressions (BFloat16 + int promotes to double). It is also a first-class AutoDiff type; see BFLOAT16.mdC#
99.8%