Skip to main content

Exact BF16 GEMM on FP16 tensor cores (V100 / sm_70)

CUDA port + extensions of "Exact BF16 Product Packing into FP16 Matrix Multiplication". Volta has FP16 tensor cores but no BF16 support; this suite provides BF16 GEMM as a menu of speed/accuracy modes, all built on the same CUTLASS-sm70 mainloop with an empirically probed accumulator mapping.

Mode (impl header) passes speed @9216x4096x4096 accuracy
scaled_impl (scaled FP16) 1 5.6 ms, 6x FP32 route == FP32 route on real ML data; survives extreme exponents
sigcut_impl (exact 1:1) 1 5.4 ms bit-exact (block-separable exponents, full BF16 value set)
gexp_impl (general exp.) 1-9 adaptive 5.4-55 ms bit-exact to 16-bit/row exponent spread; 50x better than FP32 route
  • src/ implementation: common.cuh (helpers), exact_mma.cuh (CUTLASS config, rebase mainloop, probe), rowscan.cuh, one *_impl.cuh per mode.
  • tests/ correctness: make test runs bit-exact/accuracy validation of all three modes (zeros, subnormals, Inf/NaN, odd shapes, permutation path).
  • experiments/ benchmarks (bench_*.cu) and historical prototypes: bf16_pack_gemm.cu (paper's 9:8 K<=32 repro), deep_k_pack.cu, full_k_pack.cu (12:11 limbs), sig_gemm.cu (WMMA 1:1), cublas_roofline.cu, cutlass_hgemm_ref.cu (attribution baselines).
  • analysis/ real-ML BF16 studies (SmolLM2 weights/activations).
  • python/ the voltabf16 PyTorch bindings (see below), docs/ their Sphinx site.

Build on the GPU host: make -j4 && make test (needs CUDA 12.6 for sm_70 and a CUTLASS checkout; see CUTLASS/NVCC vars).

PyTorch bindings: voltabf16

pip install voltabf16 exposes the scaled_impl mode to PyTorch, so bfloat16 training on a V100 uses the FP16 tensor cores instead of torch's non-tensor-core fallback:

import voltabf16

with voltabf16.autobf16():
    loss = model(x.bfloat16()).sum()
    loss.backward()

Nothing in the model changes -- a TorchDispatchMode intercepts bfloat16 aten::mm/aten::addmm below autograd, so forward and both backward GEMMs are routed. MNIST MLP (784 -> 4096x3 -> 10), batch 2048, 5 epochs:

mode ms/step vs FP32 test acc
FP32 39.2 1.00x 97.93%
FP16 8.4 4.68x 97.90%
BF16 (torch native) 51.6 0.76x 97.93%
BF16 (voltabf16) 11.9 3.30x 97.91%

4.3x faster than torch's bfloat16, and FP32-class accurate (~1e-7 relative error against float64, versus ~1.7e-3 native) because the products are exact. It does not reach FP16 parity and cannot: the mainloop is ~78-83% of cuBLAS FP16, and the exponent scan and pack passes are inherent. docs/performance.md has the full accounting.

The binding is built against this repository's own src/*.cuh -- there is one copy of those headers, shared with make test. setup.py copies them into wheels at build time so an installed wheel is self-contained; an editable install or a run from the checkout uses src/ directly.

pip install -e '.[test,docs,bench]'
pytest                                    # needs a V100; skips otherwise
make -C docs html
python python/examples/bench_mnist.py --epochs 5 --batch 2048 --lr 0.15
make bin/tile_sweep && ./bin/tile_sweep 2048     # CUTLASS tile-shape sweep

.github/workflows/docs.yml builds and publishes the docs; .github/workflows/pypi.yml publishes voltabf16 to PyPI via Trusted Publishing when a GitHub Release is cut.

Metadata

Release files for voltabf16 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for voltabf16 0.1.0
File Size Uploaded
voltabf16-0.1.0.tar.gz 47.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for voltabf16 0.1.0
File Interpreter ABI Platform
voltabf16-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 83.1 kB

Release files / voltabf16-0.1.0.tar.gz

Download URL voltabf16-0.1.0.tar.gz
Size 47.6 kB
Tags Source
SHA-256 checksum
How to use checksums
b6ce05053ff8ac58ab0c2edddc27f0afc29ea726ac4ee36ab5857b2453154285
BLAKE2b-256 checksum
How to use checksums
57714dbe73bb81927270d2b51d65d49617ee51f5aa9427f807e26dc621d18219
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.

Transparency log

Release files / voltabf16-0.1.0-py3-none-any.whl

Download URL voltabf16-0.1.0-py3-none-any.whl
Size 35.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
046afd35cecd22f6d8e0a9a62d5f0089eed143a46f625461c759f41e63dcb6d8
BLAKE2b-256 checksum
How to use checksums
7b81959879feae6513305a8cb8679aaf68df7f8a9ff4de271ef5306266b37ace
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page