eminsk/nanogemm

Minimalist, bare-metal SIMD & Assembly GEMM engine for Python. Sub-microsecond CPU matrix multiplication for AI & scientific computing.

2

stars

4

commits

Python

primary language

Sep 11, 2026

updated

assembly
avx2
cpu-inference
deep-learning
fasm
gemm
high-performance-computing
matrix-multiplication
python
simd
Browse cluster: SIMD and vectorization libraries

README

NanoGEMM ⚡

PyPI CI License: MIT Python SIMD Habr Open In Colab Footprint

NanoGEMM is a minimalist, bare-metal General Matrix Multiplication (GEMM) engine designed for sub-microsecond CPU inference and high-performance computing in Python.

Built with direct AVX2 / FMA (256-bit SIMD) assembly-level register tiling and cache blocking, NanoGEMM eliminates the heavy function-call dispatch, thread-pool barriers, and memory-packing overhead of heavyweight BLAS libraries (OpenBLAS, MKL) for small-to-medium tensors.

📖 Deep Dive: Read the architecture breakdown and CPU profiling post-mortem on Habr (Хабр).


🚀 Performance Benchmarks

Measured on Intel/AMD x86-64 CPU (AVX2 + FMA) against NumPy 2.2.3 (single-precision float32):

Matrix DimensionNumPy 2.2.3 LatencyNanoGEMM LatencySpeedup FactorNanoGEMM Throughput
16 x 163.21 µs1.23 µs (C: 0.65 µs)🚀 2.83x FASTER2.83 GFLOPS
32 x 325.75 µs2.74 µs (C: 2.18 µs)🚀 2.26x FASTER15.13 GFLOPS
64 x 6418.70 µs17.76 µs (C: 16.39 µs)🚀 1.10x FASTER27.08 GFLOPS
128 x 128114.07 µs182.46 µs0.60x23.42 GFLOPS
256 x 256426.24 µs1501.24 µs0.28x22.35 GFLOPS

💡 Why is NanoGEMM faster on small/medium matrices?
Traditional BLAS engines incur 3–10 µs of fixed overhead per invocation due to dynamic runtime dispatch, argument sanitization, thread synchronization, and packing buffers. NanoGEMM utilizes a zero-allocation, direct register-tiled microkernel that executes in sub-microsecond time immediately upon invocation.


🛠 Architectural Design

1. Register Tiling ($6 \times 16$ Microkernel)

  • Register allocation: Utilizes 12 ymm registers (ymm0ymm11) as 256-bit floating-point accumulators storing a $6 \times 16$ tile of matrix $C$.
  • Vector broadcast & FMA: Two ymm registers load vectors from $B$, while individual scalar elements of $A$ are broadcast across ymm using _mm256_set1_ps and accumulated via fused multiply-add (_mm256_fmadd_ps).
  • Zero Spilling: Fits completely inside the 16 available x86-64 YMM registers without stack eviction.

2. Multi-Level Cache Blocking

  • $L_1$ / $L_2$ Cache Tiling: Matrices are processed in cache blocks ($M_c = 64, N_c = 128, K_c = 128$) to maintain maximum L1/L2 data cache hit ratios and eliminate memory bus thrashing.
  • Vectorized Edge Handling: Arbitrary matrix dimensions (non-multiples of 6 or 16) are processed using boundary SIMD edge loops without padding or buffer allocations.
       Matrix A (M x K)              Matrix B (K x N)
     [ . . . . . . . . ]           [ . . . ymm0 . . . ]
     [ . . . . . . . . ]           [ . . . ymm1 . . . ]
     [ a0 a1 a2 a3 . . ]     x     [ . . . . .  . . . ]
     [ . . . . . . . . ]           [ . . . . .  . . . ]
     [ . . . . . . . . ]
             │                             │
             └──────────────┬──────────────┘
                            ▼
                Matrix C (6 x 16 Tile)
             [ ymm0  ymm1  ] -> Row 0
             [ ymm2  ymm3  ] -> Row 1
             [ ymm4  ymm5  ] -> Row 2
             [ ymm6  ymm7  ] -> Row 3
             [ ymm8  ymm9  ] -> Row 4
             [ ymm10 ymm11 ] -> Row 5

📦 Installation & Quickstart

pip install nanogemm

Build from Source

git clone https://github.com/eminsk/nanogemm.git
cd nanogemm
pip install -e .

Python Usage

import nanogemm as ng
import numpy as np

# Verify SIMD hardware acceleration
print("Active ISA:", ng.get_simd_isa())
# Output: Active ISA: AVX2+FMA (256-bit SIMD, 6x16 register tiling)

# Allocate input matrices
A = np.random.randn(32, 64).astype(np.float32)
B = np.random.randn(64, 128).astype(np.float32)

# Direct hardware-accelerated MatMul: C = A @ B
C = ng.matmul(A, B)

# Or with pre-allocated zero-copy output buffer for maximum performance:
out = np.empty((32, 128), dtype=np.float32)
ng.matmul(A, B, out=out)

# Standard BLAS SGEMM interface: C = alpha * (A @ B) + beta * C
res = ng.sgemm(A, B, alpha=2.0, beta=0.5, c=out)

🧪 Testing & Verification

Run the comprehensive correctness test suite comparing NanoGEMM with NumPy reference outputs across random uniforms, normals, non-square dimensions, and prime shapes:

python tests/test_correctness.py

Run the official benchmark against your installed NumPy BLAS:

python benchmarks/bench_vs_numpy.py

🌐 High-Performance Systems Ecosystem

NanoGEMM is developed by @eminsk as part of an engineering ecosystem focused on low-level hardware performance, assembly programming, and native desktop computing:

  • 🎥 screenvideo — Lightweight desktop screen recorder featuring WASAPI loopback audio and a standalone pure x64 Flat Assembler (FASM) native edition.
  • 📊 xlsx_vievers — Desktop spreadsheet processor with 80+ formula functions, Chart Wizard, and hardware-accelerated SIMD SSE2 math engine.
  • 📈 yfinance-ta-patterns — Candlestick pattern scanner and AI ranking suite powered by TA-Lib and quantitative backtesting.
  • 🔍 StackOverflowAPI — Desktop client for Stack Overflow built with CustomTkinter and native FASM x64 search client.

📄 License

MIT License — Copyright (c) 2026 eminsk.

Contributors

MNNik91

4 commits

eminsk/nanogemm

Minimalist, bare-metal SIMD & Assembly GEMM engine for Python. Sub-microsecond CPU matrix multiplication for AI & scientific computing.

2

stars

4

commits

Python

primary language

Sep 11, 2026

updated

assembly
avx2
cpu-inference
deep-learning
fasm
gemm
high-performance-computing
matrix-multiplication
python
simd
Browse cluster: SIMD and vectorization libraries

README

NanoGEMM ⚡

PyPI CI License: MIT Python SIMD Habr Open In Colab Footprint

NanoGEMM is a minimalist, bare-metal General Matrix Multiplication (GEMM) engine designed for sub-microsecond CPU inference and high-performance computing in Python.

Built with direct AVX2 / FMA (256-bit SIMD) assembly-level register tiling and cache blocking, NanoGEMM eliminates the heavy function-call dispatch, thread-pool barriers, and memory-packing overhead of heavyweight BLAS libraries (OpenBLAS, MKL) for small-to-medium tensors.

📖 Deep Dive: Read the architecture breakdown and CPU profiling post-mortem on Habr (Хабр).


🚀 Performance Benchmarks

Measured on Intel/AMD x86-64 CPU (AVX2 + FMA) against NumPy 2.2.3 (single-precision float32):

Matrix DimensionNumPy 2.2.3 LatencyNanoGEMM LatencySpeedup FactorNanoGEMM Throughput
16 x 163.21 µs1.23 µs (C: 0.65 µs)🚀 2.83x FASTER2.83 GFLOPS
32 x 325.75 µs2.74 µs (C: 2.18 µs)🚀 2.26x FASTER15.13 GFLOPS
64 x 6418.70 µs17.76 µs (C: 16.39 µs)🚀 1.10x FASTER27.08 GFLOPS
128 x 128114.07 µs182.46 µs0.60x23.42 GFLOPS
256 x 256426.24 µs1501.24 µs0.28x22.35 GFLOPS

💡 Why is NanoGEMM faster on small/medium matrices?
Traditional BLAS engines incur 3–10 µs of fixed overhead per invocation due to dynamic runtime dispatch, argument sanitization, thread synchronization, and packing buffers. NanoGEMM utilizes a zero-allocation, direct register-tiled microkernel that executes in sub-microsecond time immediately upon invocation.


🛠 Architectural Design

1. Register Tiling ($6 \times 16$ Microkernel)

  • Register allocation: Utilizes 12 ymm registers (ymm0ymm11) as 256-bit floating-point accumulators storing a $6 \times 16$ tile of matrix $C$.
  • Vector broadcast & FMA: Two ymm registers load vectors from $B$, while individual scalar elements of $A$ are broadcast across ymm using _mm256_set1_ps and accumulated via fused multiply-add (_mm256_fmadd_ps).
  • Zero Spilling: Fits completely inside the 16 available x86-64 YMM registers without stack eviction.

2. Multi-Level Cache Blocking

  • $L_1$ / $L_2$ Cache Tiling: Matrices are processed in cache blocks ($M_c = 64, N_c = 128, K_c = 128$) to maintain maximum L1/L2 data cache hit ratios and eliminate memory bus thrashing.
  • Vectorized Edge Handling: Arbitrary matrix dimensions (non-multiples of 6 or 16) are processed using boundary SIMD edge loops without padding or buffer allocations.
       Matrix A (M x K)              Matrix B (K x N)
     [ . . . . . . . . ]           [ . . . ymm0 . . . ]
     [ . . . . . . . . ]           [ . . . ymm1 . . . ]
     [ a0 a1 a2 a3 . . ]     x     [ . . . . .  . . . ]
     [ . . . . . . . . ]           [ . . . . .  . . . ]
     [ . . . . . . . . ]
             │                             │
             └──────────────┬──────────────┘
                            ▼
                Matrix C (6 x 16 Tile)
             [ ymm0  ymm1  ] -> Row 0
             [ ymm2  ymm3  ] -> Row 1
             [ ymm4  ymm5  ] -> Row 2
             [ ymm6  ymm7  ] -> Row 3
             [ ymm8  ymm9  ] -> Row 4
             [ ymm10 ymm11 ] -> Row 5

📦 Installation & Quickstart

pip install nanogemm

Build from Source

git clone https://github.com/eminsk/nanogemm.git
cd nanogemm
pip install -e .

Python Usage

import nanogemm as ng
import numpy as np

# Verify SIMD hardware acceleration
print("Active ISA:", ng.get_simd_isa())
# Output: Active ISA: AVX2+FMA (256-bit SIMD, 6x16 register tiling)

# Allocate input matrices
A = np.random.randn(32, 64).astype(np.float32)
B = np.random.randn(64, 128).astype(np.float32)

# Direct hardware-accelerated MatMul: C = A @ B
C = ng.matmul(A, B)

# Or with pre-allocated zero-copy output buffer for maximum performance:
out = np.empty((32, 128), dtype=np.float32)
ng.matmul(A, B, out=out)

# Standard BLAS SGEMM interface: C = alpha * (A @ B) + beta * C
res = ng.sgemm(A, B, alpha=2.0, beta=0.5, c=out)

🧪 Testing & Verification

Run the comprehensive correctness test suite comparing NanoGEMM with NumPy reference outputs across random uniforms, normals, non-square dimensions, and prime shapes:

python tests/test_correctness.py

Run the official benchmark against your installed NumPy BLAS:

python benchmarks/bench_vs_numpy.py

🌐 High-Performance Systems Ecosystem

NanoGEMM is developed by @eminsk as part of an engineering ecosystem focused on low-level hardware performance, assembly programming, and native desktop computing:

  • 🎥 screenvideo — Lightweight desktop screen recorder featuring WASAPI loopback audio and a standalone pure x64 Flat Assembler (FASM) native edition.
  • 📊 xlsx_vievers — Desktop spreadsheet processor with 80+ formula functions, Chart Wizard, and hardware-accelerated SIMD SSE2 math engine.
  • 📈 yfinance-ta-patterns — Candlestick pattern scanner and AI ranking suite powered by TA-Lib and quantitative backtesting.
  • 🔍 StackOverflowAPI — Desktop client for Stack Overflow built with CustomTkinter and native FASM x64 search client.

📄 License

MIT License — Copyright (c) 2026 eminsk.

Contributors

MNNik91

4 commits

Languages

Python

46.7%

C

42.5%

Jupyter Notebook

10.7%