Skip to main content

NanoGEMM ⚡

CI License: MIT Python SIMD Footprint

NanoGEMM is a minimalist, bare-metal General Matrix Multiplication (GEMM) engine designed for sub-microsecond CPU inference and high-performance computing in Python.

Built with direct AVX2 / FMA (256-bit SIMD) assembly-level register tiling and cache blocking, NanoGEMM eliminates the heavy function-call dispatch, thread-pool barriers, and memory-packing overhead of heavyweight BLAS libraries (OpenBLAS, MKL) for small-to-medium tensors.


🚀 Performance Benchmarks

Measured on Intel/AMD x86-64 CPU (AVX2 + FMA) against NumPy 2.2.3 (single-precision float32):

Matrix Dimension NumPy 2.2.3 Latency NanoGEMM Latency Speedup Factor NanoGEMM Throughput
16 x 16 3.21 µs 1.23 µs (C: 0.65 µs) 🚀 2.83x FASTER 2.83 GFLOPS
32 x 32 5.75 µs 2.74 µs (C: 2.18 µs) 🚀 2.26x FASTER 15.13 GFLOPS
64 x 64 18.70 µs 17.76 µs (C: 16.39 µs) 🚀 1.10x FASTER 27.08 GFLOPS
128 x 128 114.07 µs 182.46 µs 0.60x 23.42 GFLOPS
256 x 256 426.24 µs 1501.24 µs 0.28x 22.35 GFLOPS

💡 Why is NanoGEMM faster on small/medium matrices?
Traditional BLAS engines incur 3–10 µs of fixed overhead per invocation due to dynamic runtime dispatch, argument sanitization, thread synchronization, and packing buffers. NanoGEMM utilizes a zero-allocation, direct register-tiled microkernel that executes in sub-microsecond time immediately upon invocation.


🛠 Architectural Design

1. Register Tiling ($6 \times 16$ Microkernel)

  • Register allocation: Utilizes 12 ymm registers (ymm0ymm11) as 256-bit floating-point accumulators storing a $6 \times 16$ tile of matrix $C$.
  • Vector broadcast & FMA: Two ymm registers load vectors from $B$, while individual scalar elements of $A$ are broadcast across ymm using _mm256_set1_ps and accumulated via fused multiply-add (_mm256_fmadd_ps).
  • Zero Spilling: Fits completely inside the 16 available x86-64 YMM registers without stack eviction.

2. Multi-Level Cache Blocking

  • $L_1$ / $L_2$ Cache Tiling: Matrices are processed in cache blocks ($M_c = 64, N_c = 128, K_c = 128$) to maintain maximum L1/L2 data cache hit ratios and eliminate memory bus thrashing.
  • Vectorized Edge Handling: Arbitrary matrix dimensions (non-multiples of 6 or 16) are processed using boundary SIMD edge loops without padding or buffer allocations.
       Matrix A (M x K)              Matrix B (K x N)
     [ . . . . . . . . ]           [ . . . ymm0 . . . ]
     [ . . . . . . . . ]           [ . . . ymm1 . . . ]
     [ a0 a1 a2 a3 . . ]     x     [ . . . . .  . . . ]
     [ . . . . . . . . ]           [ . . . . .  . . . ]
     [ . . . . . . . . ]
             │                             │
             └──────────────┬──────────────┘
                            ▼
                Matrix C (6 x 16 Tile)
             [ ymm0  ymm1  ] -> Row 0
             [ ymm2  ymm3  ] -> Row 1
             [ ymm4  ymm5  ] -> Row 2
             [ ymm6  ymm7  ] -> Row 3
             [ ymm8  ymm9  ] -> Row 4
             [ ymm10 ymm11 ] -> Row 5

📦 Installation & Quickstart

Installation

git clone https://github.com/eminsk/nanogemm.git
cd nanogemm
python setup.py build_ext --inplace

Python Usage

import nanogemm as ng
import numpy as np

# Verify SIMD hardware acceleration
print("Active ISA:", ng.get_simd_isa())
# Output: Active ISA: AVX2+FMA (256-bit SIMD, 6x16 register tiling)

# Allocate input matrices
A = np.random.randn(32, 64).astype(np.float32)
B = np.random.randn(64, 128).astype(np.float32)

# Direct hardware-accelerated MatMul: C = A @ B
C = ng.matmul(A, B)

# Or with pre-allocated zero-copy output buffer for maximum performance:
out = np.empty((32, 128), dtype=np.float32)
ng.matmul(A, B, out=out)

# Standard BLAS SGEMM interface: C = alpha * (A @ B) + beta * C
res = ng.sgemm(A, B, alpha=2.0, beta=0.5, c=out)

🧪 Testing & Verification

Run the comprehensive correctness test suite comparing NanoGEMM with NumPy reference outputs across random uniforms, normals, non-square dimensions, and prime shapes:

python tests/test_correctness.py

Run the official benchmark against your installed NumPy BLAS:

python benchmarks/bench_vs_numpy.py

🌐 High-Performance Systems Ecosystem

NanoGEMM is developed by @eminsk as part of an engineering ecosystem focused on low-level hardware performance, assembly programming, and native desktop computing:

  • 🎥 screenvideo — Lightweight desktop screen recorder featuring WASAPI loopback audio and a standalone pure x64 Flat Assembler (FASM) native edition.
  • 📊 xlsx_vievers — Desktop spreadsheet processor with 80+ formula functions, Chart Wizard, and hardware-accelerated SIMD SSE2 math engine.
  • 📈 yfinance-ta-patterns — Candlestick pattern scanner and AI ranking suite powered by TA-Lib and quantitative backtesting.
  • 🔍 StackOverflowAPI — Desktop client for Stack Overflow built with CustomTkinter and native FASM x64 search client.

📄 License

MIT License — Copyright (c) 2026 eminsk.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nanogemm-0.1.0.tar.gz (56.6 kB view details)

Uploaded Source

File details

Details for the file nanogemm-0.1.0.tar.gz.

File metadata

  • Download URL: nanogemm-0.1.0.tar.gz
  • Upload date:
  • Size: 56.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.12.9

File hashes

Hashes for nanogemm-0.1.0.tar.gz
Algorithm Hash digest
SHA256 9b7203f964b32c0ad3b3d4ee8c79322440979b7d88728237dd211da88fec9066
MD5 16932b3638463d1c595028441cca4aab
BLAKE2b-256 46051be44b0a535051c7ef51f439a4891debd7060cd0280c64cbd835db90ead2

See more details on using hashes here.

Release history Release notifications | RSS feed

0.2.0

5 files

This release

0.1.0 This release

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page