Exact BF16 GEMM on FP16 tensor cores (V100 / sm_70)
CUDA port + extensions of "Exact BF16 Product Packing into FP16 Matrix Multiplication". Volta has FP16 tensor cores but no BF16 support; this suite provides BF16 GEMM as a menu of speed/accuracy modes, all built on the same CUTLASS-sm70 mainloop with an empirically probed accumulator mapping.
| Mode (impl header) | passes | speed @9216x4096x4096 | accuracy |
|---|---|---|---|
| scaled_impl (scaled FP16) | 1 | 5.6 ms, 6x FP32 route | == FP32 route on real ML data; survives extreme exponents |
| sigcut_impl (exact 1:1) | 1 | 5.4 ms | bit-exact (block-separable exponents, full BF16 value set) |
| gexp_impl (general exp.) | 1-9 adaptive | 5.4-55 ms | bit-exact to 16-bit/row exponent spread; 50x better than FP32 route |
src/implementation:common.cuh(helpers),exact_mma.cuh(CUTLASS config, rebase mainloop, probe),rowscan.cuh, one*_impl.cuhper mode.tests/correctness:make testruns bit-exact/accuracy validation of all three modes (zeros, subnormals, Inf/NaN, odd shapes, permutation path).experiments/benchmarks (bench_*.cu) and historical prototypes:bf16_pack_gemm.cu(paper's 9:8 K<=32 repro),deep_k_pack.cu,full_k_pack.cu(12:11 limbs),sig_gemm.cu(WMMA 1:1),cublas_roofline.cu,cutlass_hgemm_ref.cu(attribution baselines).analysis/real-ML BF16 studies (SmolLM2 weights/activations).python/thevoltabf16PyTorch bindings (see below),docs/their Sphinx site.
Build on the GPU host: make -j4 && make test (needs CUDA 12.6 for sm_70 and
a CUTLASS checkout; see CUTLASS/NVCC vars).
PyTorch bindings: voltabf16
pip install voltabf16 exposes the scaled_impl mode to PyTorch, so bfloat16
training on a V100 uses the FP16 tensor cores instead of torch's non-tensor-core
fallback:
import voltabf16
with voltabf16.autobf16():
loss = model(x.bfloat16()).sum()
loss.backward()
Nothing in the model changes -- a TorchDispatchMode intercepts bfloat16
aten::mm/aten::addmm below autograd, so forward and both backward GEMMs are
routed. MNIST MLP (784 -> 4096x3 -> 10), batch 2048, 5 epochs:
| mode | ms/step | vs FP32 | test acc |
|---|---|---|---|
| FP32 | 39.2 | 1.00x | 97.93% |
| FP16 | 8.4 | 4.68x | 97.90% |
| BF16 (torch native) | 51.6 | 0.76x | 97.93% |
| BF16 (voltabf16) | 11.9 | 3.30x | 97.91% |
4.3x faster than torch's bfloat16, and FP32-class accurate (~1e-7 relative
error against float64, versus ~1.7e-3 native) because the products are exact.
It does not reach FP16 parity and cannot: the mainloop is ~78-83% of cuBLAS
FP16, and the exponent scan and pack passes are inherent. docs/performance.md
has the full accounting.
The binding is built against this repository's own src/*.cuh -- there is one
copy of those headers, shared with make test. setup.py copies them into
wheels at build time so an installed wheel is self-contained; an editable
install or a run from the checkout uses src/ directly.
pip install -e '.[test,docs,bench]'
pytest # needs a V100; skips otherwise
make -C docs html
python python/examples/bench_mnist.py --epochs 5 --batch 2048 --lr 0.15
make bin/tile_sweep && ./bin/tile_sweep 2048 # CUTLASS tile-shape sweep
.github/workflows/docs.yml builds and publishes the docs;
.github/workflows/pypi.yml publishes voltabf16 to PyPI via Trusted
Publishing when a GitHub Release is cut.
Metadata
Release files for voltabf16 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| voltabf16-0.1.0.tar.gz | 47.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| voltabf16-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 83.1 kB
Release files / voltabf16-0.1.0.tar.gz
| Download URL | voltabf16-0.1.0.tar.gz |
|---|---|
| Size | 47.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
b6ce05053ff8ac58ab0c2edddc27f0afc29ea726ac4ee36ab5857b2453154285
|
|
BLAKE2b-256 checksum How to use checksums |
57714dbe73bb81927270d2b51d65d49617ee51f5aa9427f807e26dc621d18219
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.
Transparency logRelease files / voltabf16-0.1.0-py3-none-any.whl
| Download URL | voltabf16-0.1.0-py3-none-any.whl |
|---|---|
| Size | 35.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
046afd35cecd22f6d8e0a9a62d5f0089eed143a46f625461c759f41e63dcb6d8
|
|
BLAKE2b-256 checksum How to use checksums |
7b81959879feae6513305a8cb8679aaf68df7f8a9ff4de271ef5306266b37ace
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 29, 2026.
Transparency log