Skip to main content

A fast, memory-lean neural-network core in C, with a Python binding.

Project description

mantissa

ci License Language Python Dependencies

A fast, memory-lean neural-network core in C, with a Python binding.

mantissa is an ultra-lightweight, high-performance neural network engine designed to execute forward and backward passes with minimal time and memory overhead per parameter. By pairing a bare-metal, highly optimized C compute core with a thin Python interface, mantissa enables the training and deployment of large-scale models directly within tightly constrained RAM/VRAM environments.

Started by Tekin Ertekin (2024); later refactored with Claude Code — see AUTHORS.md.

Version history and per-release benchmark scores: RELEASES.md.


How it works

mantissa is a shared library with a plain C ABI, so any language that can call C — Python (via ctypes), C/C++, Rust, Go — drives the same engine. The caller passes ordinary float32 arrays in and gets float32 back; inside, weights are kept in a narrow type to save memory.

mantissa architecture: callers to the C ABI to the engine core

A call is one of three kinds:

  • Build — quantize weights from float32 into the compact storage type (once).
  • Forward pass (tk_linear_forward) — compute y = activation(W·x + b). Called forward because data flows forward through the network, input → output (also feed-forward). There is no "fast" in it — the opposite direction is the backward pass.
  • Backward pass / training (tk_linear_backwardtk_sgd_step) — gradients flow backward, output → input, and the weights are updated.

Whichever it is, the compute core does the same underneath: read the narrow weights, accumulate in float32, apply an activation, return float32.

Performance at a glance

A 2048×2048 dense layer (4.2M parameters) on an Apple M-series laptop, default bfloat16 (make bench / make benchbp):

value
forward pass (10 threads) 0.10 ms (~83 GFLOP/s)
forward pass (1 thread) 0.35 ms (~29 GFLOP/s)
backward pass (10 threads) 0.33 ms (~39 GFLOP/s)
backward pass (1 thread) 1.01 ms (~12 GFLOP/s)
weight memory 8 MB — half of float32
relu over 4M values 0.41 ms

Forward and backward use explicit NEON kernels (bfloat16 leads — half the bytes, same FMAs) and a persistent thread pool that splits rows across cores for large layers. Numbers are for the default bfloat16 on a 10-core Apple laptop.

Memory is the headline at scale — it decides whether a model loads at all:

model float32 mantissa (bf16) mantissa (1-byte)
1B params 4.0 GB 2.0 GB 1.0 GB
7B params 28 GB 14 GB 7 GB

How it stays small and fast

  1. Narrow weight storage — weights are the bulk of a model, so storing each in 2 bytes (or 1) instead of 4 is where the RAM/VRAM savings come from, and since a dense layer is memory-bandwidth bound, moving fewer bytes is also faster. The precision is a build-time dial (below).
  2. float32 accumulation — compute always sums in float32 so accuracy holds across a layer's millions of terms (mixed precision; Micikevicius 2017).
  3. Tuned kernels — register-blocked GEMV with explicit SIMD FMA kernels (NEON on arm64, AVX2 on x86-64) and a persistent thread pool across cores, branchless activations (sign-bit / fmax / copysign, chosen by benchmark), and stochastic rounding so training survives narrow weights. Details and measurements in docs/DESIGN.md.

The precision dial

Storage precision is one build-time knob (DTYPE); the default needs no config. It is how the core trades accuracy for memory — a means to the speed/RAM goal, not the goal itself.

DTYPE Name Bytes When
0 float32 4 reference / exact
2 bfloat16 2 default — half the RAM, training-safe
1 fp16 2 half RAM, more precision, less range
3 tekin32 4 custom high-fidelity accumulator
4 tekin8 1 FP8 E4M3 — 4× smaller
5 fp8_e5m2 1 FP8, wider range
6 fp4_e2m1 ½* FP4 — the extreme

*fp4 is stored one value per byte today (1 B); the "½" is the logical E2M1 width. Two-per-byte packing is a roadmap item gated on block-scaling — see docs/DESIGN.md.

Bit layouts, the tekin formats' design rationale, the hot/cold conversion split, and the current-research context (MX / NVFP4 / posit / IEEE P3109) all live in docs/DESIGN.md. Every value is stored in the selected type but computed in float32.

Accuracy cost, seen

The same value stored in each format and read back (make DTYPE=<n> test). This is the accuracy you trade away for the memory you save — float32 is the reference column:

input float32 tekin32 fp16 bfloat16 e5m2 tekin8
3.14159 3.14159 3.14159 3.14062 3.14062 3.0 3.25
100.0 100 100 100 100 96 104
0.01 0.01 0.01 0.0100021 0.0100098 0.0098 0.00977

fp4_e2m1 is coarser still: its 8 magnitudes are {0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6} — everything rounds onto that grid.

Full benchmarks

The numbers behind Performance at a glance, per dtype. make DTYPE=<n> bench — a 2048×2048 dense layer (4.2M params), plus the activation-dispatch micro-test. Apple M-series laptop, clang -O3; indicative, not absolute.

dtype weight memory GEMV ms/pass GEMV GFLOP/s vs baseline
bfloat16 8.0 MB 0.39 ~21 2.5× (NEON)
float32 16.0 MB 0.55 ~15 1.9× (NEON)
tekin8 4.0 MB 2.6 ~3.2 1.25× (blocking)

The forward pass is register-blocked (4 output rows share each x load and run 4 independent FMA chains), with explicit SIMD FMA kernels for float32 and bfloat16 — NEON on arm64, AVX2 on x86-64 (runtime-dispatched via __builtin_cpu_supports, scalar fallback on older CPUs). bfloat16 leads on both axes: it moves half the bytes of float32 for the same FMAs. tekin8 stays conversion- bound (E4M3→float unpack, no SIMD; native FP8 hardware removes that).

On top, a thread pool splits the output rows across cores for large layers, both forward and backward (the per-kernel numbers above are single-thread):

bfloat16 (2048×2048) 1 thread 10 threads
forward GFLOP/s ~29 ~83 (2.9×)
backward GFLOP/s ~12 ~39 (3.1×)

Scaling is sub-linear because GEMV does only ~2 FLOPs per byte, so it saturates memory bandwidth before compute — float32 (twice the bytes) barely gains, which is exactly why the narrow default matters. Small layers run serially (below a work threshold) so the pool never hurts the millions-of-small-calls case; set MANTISSA_THREADS to tune. Numbers are laptop-noisy.

Backward pass (make DTYPE=<n> benchbp, same 2048×2048 layer):

dtype backward ms/pass backward GFLOP/s SGD update (M weights/s)
float32 1.03 12.21 9379
bfloat16 0.98 12.83 994
tekin8 3.04 4.14 321

Activation dispatch (4M elements): a per-element switch beats a resolved function pointer ~7× for relu and ~1.5× for sigmoid on Apple Silicon — on the x86 CI runner the gap shrinks to ≤1.2× (parity on some cases), so the win is real but platform-dependent. The inline switch vectorizes (relu compiles to a single fmax per lane); an indirect call per element does not. So tk_activate keeps the switch, and the function-pointer API (tk_act_resolve) is reserved for genuinely pluggable dispatch. The step/sign/relu kernels themselves are branchless (sign-bit read, fmax, copysign), each picked by benchmark over the obvious comparison. Measure, don't assume.

Mixed architectures

Every layer configures itself independently — bias is a per-call NULL-able pointer, activation is a per-call argument:

tk_linear_forward(W1, x,  b1,   h1, 6, 4, TK_ACT_TANH);     /* bias + tanh    */
tk_linear_forward(W2, h1, NULL, h2, 5, 6, TK_ACT_RELU);     /* no bias + relu */
tk_linear_forward(W3, h2, b3,   y,  2, 5, TK_ACT_SIGMOID);  /* bias + sigmoid */

(Under the default narrow storage each float output is requantized to tk_scalar_t before it feeds the next layer — see examples/mlp_example.c; the snippet elides that for readability. It compiles as-is only under DTYPE=0.)

That is a full 3-layer MLP with three different bias/activation setups — exactly the heterogeneity a Transformer needs (bias-free attention projections, bias-using feed-forward). Run it: make mlp.

Training (back-propagation)

The reverse pass mirrors the forward core: tk_linear_backward computes the weight, bias, and input gradients for a dense layer; tk_sgd_step updates narrow-stored weights; tk_loss (MSE / BCE) seeds the gradient. No autograd graph — the caller drives the layers, keeping everything explicit and inspectable.

One training step is a loop: run forward, measure the loss against the target, propagate gradients backward, update the weights, repeat.

training loop: forward, loss, backward, update

Correctness is proven by a gradient check (make testbp): analytic gradients vs central finite differences, matching to <1e-2 relative error for tanh, sigmoid, relu, and gelu.

make train learns XOR — the problem a single perceptron cannot solve — end to end, in the default bfloat16:

Training XOR  (dtype=bfloat16, stochastic_rounding=1)
  epoch    0  loss 0.31523
  epoch 1000  loss 0.00035
  epoch 4000  loss 0.00007
predictions:
  (0,0) -> 0.004   (0,1) -> 0.990   (1,0) -> 0.991   (1,1) -> 0.009

That it converges in bf16 is the point of stochastic rounding (config TK_USE_STOCHASTIC_ROUNDING): under plain round-to-nearest, a weight update smaller than the storage type's precision rounds to zero and training stalls; SR rounds up/down with probability proportional to distance, so tiny updates accumulate in expectation (Gupta et al., 2015; the technique behind FP8 training on Hopper/Blackwell). It needs no fp32 master copy of the weights.

Measured — same XOR run, 4000 epochs, round-to-nearest vs SR (make DTYPE=<n> benchbp):

dtype round-to-nearest stochastic rounding
float32 0.00008 0.00008 (SR is a no-op)
bfloat16 0.01090 0.00009
tekin8 0.24862 — stalled 0.00008 — converged

In the 1-byte type, plain rounding never learns XOR; stochastic rounding does. That is the whole reason the technique exists.

Training config (all OFF by default)

Flag Effect
TK_USE_DROPOUT / TK_DROPOUT_RATE inverted dropout on activations
TK_USE_L1 / TK_L1_LAMBDA L1 weight penalty in the update
TK_USE_L2 / TK_L2_LAMBDA L2 weight penalty in the update
TK_USE_STOCHASTIC_ROUNDING SR on the weight write-back

Each sets a default; the runtime tk_optim / dropout calls override per layer. Back-propagation covers the dense layer (the shared primitive) and, since v0.2.1, the conv/pool family (conv.h, float32 — see the roadmap section); recurrent backward reuses the same pieces later.

Zero-config

Never touch config.h and you get Google's bfloat16 with bias enabled — the safe general-purpose default. Override only when you want to:

make DTYPE=4 test     # switch storage to tekin8

Install

pip install mantissa-core       # import name stays `mantissa`

The wheel ships a prebuilt libmantissa (default bfloat16) — no compiler or make needed. numpy is optional; install mantissa-core[numpy] for the zero-copy training path. Then:

from mantissa import Mantissa, STEP
tk = Mantissa()
print(tk.dtype)                # 'bfloat16'

The distribution is named mantissa-core on PyPI (the bare mantissa name is taken by an unrelated project); the Python import name is still mantissa.

Installing from an sdist (or pip install . from a checkout) compiles the C core on your machine, so a C toolchain (cc, make-grade compiler) is required for that path — wheel users need nothing. Override the storage dtype at build time with MANTISSA_DTYPE=<id> (same ids as DTYPE; default 2 = bfloat16):

pip install .                       # from a source checkout
MANTISSA_DTYPE=4 pip install .      # build the tekin8 (FP8) library instead

Download

Prebuilt shared libraries are attached to every release — built by CI for Linux, macOS, and Windows. Grab the one for your OS and load it straight from Python, no compiler or make needed:

OS file
Linux libmantissa-linux-x86_64.so
macOS libmantissa-macos-arm64.dylib
Windows libmantissa-windows-x86_64.dll
import ctypes
lib = ctypes.CDLL("./libmantissa-linux-x86_64.so")   # the file you downloaded
lib.tk_dtype_name.restype = ctypes.c_char_p
print(lib.tk_dtype_name().decode())                  # -> e.g. "bfloat16"

Or use the ctypes wrapper in python/mantissa/ (point it at the downloaded file) — pass float32 numpy arrays (or array('f')) for a zero-copy, in-place training path, ~200× faster per step than plain lists (numpy optional). For an epoch loop, Mantissa.trainer() binds the buffers' pointers once and roughly halves the per-epoch call cost again (measured 9.8 → 4.8 µs at 1030×4 — the conversion, not the C, was the cost); see docs/USAGE.md. Runnable forward + back-prop examples for Python, C++, C#, Java, JavaScript, and Rust live in clients/ — all calling the same C ABI. To build from source instead, see below.

For a full, installable client with honest benchmarks, see the sister project mantissa-perceptron — a Python perceptron classifier built on this engine.

Quick start

make test                     # tests, default bfloat16
make example                  # C perceptron
make mlp                      # mixed 3-layer MLP
make DTYPE=0 bench            # benchmark, float32
make dist && python3 python/perceptron_example.py

The Python binding is dtype-agnostic: it calls the float32 entry point, so the same script runs against any storage type the library was built for. Full API and examples in docs/USAGE.md.

Numeric landscape & roadmap

The low-precision frontier moves fast; mantissa tracks it deliberately:

  • Microscaling (MX) — OCP MX v1.0 (2023): blocks of 32 elements share one E8M0 scale, mitigating the tiny dynamic range of 4/6-bit elements. MXFP8, MXFP6, MXFP4.
  • NVFP4 — NVIDIA Blackwell (2024 hardware): 16-element blocks with an FP8 E4M3 block scale; used to pretrain LLMs at 4 bits (arXiv:2509.25149, 2025).
  • Posit / takum — tapered-precision alternatives to IEEE floats, strong for ≤8-bit inference on zero-centered weights.
  • IEEE P3109 — an emerging standard for ML arithmetic formats.

mantissa implements the element formats (E4M3, E5M2, E2M1); block-level microscaling (a shared per-block scale) is the next planned step. The conv/pool primitives landed in v0.2.1 (conv.h): im2col + register-blocked GEMM convolution, max pooling with argmax scatter, a batched dense head, fused softmax-cross-entropy and plain-f32 SGD — the full CNN training family, gradient-checked (make testconv) and benched at LeNet-5/VGG shapes (make benchconv; VGG 64→64@3×3 forward 199 GFLOP/s threaded on M4). This family is deliberately pure float32: narrow-storage conv stays on the roadmap, gated on the same block-scaling work as the 4-bit packing. Recurrent backward will reuse the same gradient and optimizer pieces.

Project layout

include/   config, dtypes + conversions, activations, ops, loss, backprop, conv, pool
src/       implementations (conv.c: the float32 CNN family — conv2d, maxpool, softmax-xent)
tests/     forward round-trip checks (7 formats) + backprop & conv gradient checks
examples/  perceptron, mixed-MLP, XOR training
bench/     GEMV + activation-dispatch benchmark, conv shapes (benchconv)
python/    ctypes binding + Python perceptron & training examples
clients/   forward + back-prop demos in C++, C#, Java, JavaScript, Rust
docs/      DESIGN.md (numerics, optimization), USAGE.md (API + examples)

License

MIT — see LICENSE. © 2024 Tekin Ertekin.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mantissa_core-0.2.5.tar.gz (105.8 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

mantissa_core-0.2.5-py3-none-win_amd64.whl (162.1 kB view details)

Uploaded Python 3Windows x86-64

mantissa_core-0.2.5-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl (123.3 kB view details)

Uploaded Python 3manylinux: glibc 2.27+ x86-64manylinux: glibc 2.28+ x86-64

mantissa_core-0.2.5-py3-none-macosx_11_0_arm64.whl (68.5 kB view details)

Uploaded Python 3macOS 11.0+ ARM64

File details

Details for the file mantissa_core-0.2.5.tar.gz.

File metadata

  • Download URL: mantissa_core-0.2.5.tar.gz
  • Upload date:
  • Size: 105.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mantissa_core-0.2.5.tar.gz
Algorithm Hash digest
SHA256 abce64d0de7b043a77367a073b2ce60fab00bf814e894613c44cab943a52cadb
MD5 4bc4f4f0c4258abc6b65a68eb1c76d6c
BLAKE2b-256 c2e52172adf941aa4ae32700fa2ba2e8340850ac57899a4890d61f097a87497c

See more details on using hashes here.

Provenance

The following attestation bundles were made for mantissa_core-0.2.5.tar.gz:

Publisher: release.yml on tekinertekin/mantissa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mantissa_core-0.2.5-py3-none-win_amd64.whl.

File metadata

  • Download URL: mantissa_core-0.2.5-py3-none-win_amd64.whl
  • Upload date:
  • Size: 162.1 kB
  • Tags: Python 3, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for mantissa_core-0.2.5-py3-none-win_amd64.whl
Algorithm Hash digest
SHA256 df2c2ce4c4aae069869620d78c1889f282acc0457944b9ea737dbab1d81867de
MD5 cfcdfa2bef833799448895300e961243
BLAKE2b-256 399c73b7bfb492c125b63d32d49c65fe8f6320273ac6c5da6bbcd091edf9ee3b

See more details on using hashes here.

Provenance

The following attestation bundles were made for mantissa_core-0.2.5-py3-none-win_amd64.whl:

Publisher: release.yml on tekinertekin/mantissa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mantissa_core-0.2.5-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for mantissa_core-0.2.5-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 f9143ded744fa10f7a4f0f147c35ddb54c073e256c4b97c4be711429514581e9
MD5 1f5a6ad6cc153bd486fc662f56cd0155
BLAKE2b-256 d967be1a238316d4a782b230513e37592a3b91c843ccb2d2a10dce0022776dfa

See more details on using hashes here.

Provenance

The following attestation bundles were made for mantissa_core-0.2.5-py3-none-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl:

Publisher: release.yml on tekinertekin/mantissa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mantissa_core-0.2.5-py3-none-macosx_11_0_arm64.whl.

File metadata

File hashes

Hashes for mantissa_core-0.2.5-py3-none-macosx_11_0_arm64.whl
Algorithm Hash digest
SHA256 b0dce37180f150fb7106c4de5296a92763fcf20c29acccb3e1dac59ca59a7969
MD5 ed613bf4eef1191f017bd491477e1629
BLAKE2b-256 acd9bb0de7aa380892b955b455928a8521d01ca0005e02da73762c51d81abc75

See more details on using hashes here.

Provenance

The following attestation bundles were made for mantissa_core-0.2.5-py3-none-macosx_11_0_arm64.whl:

Publisher: release.yml on tekinertekin/mantissa

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page