This release is a pre-release and may not be stable for production use.
BioNeMo Inference Runtime
Easy, fast, and memory-efficient structure prediction inference
GPU-accelerated inference for protein, nucleic-acid, and ligand structure prediction models — from FASTA/MSA to PDB/mmCIF.
About
BioNeMo Inference Runtime (BioIR) is NVIDIA's library for structure-prediction
inference. A five-stage GPU pipeline turns AlphaFold-lineage and all-atom models
into PDB/mmCIF with confidence scores. Models stay ordinary nn.Modules — no
TensorRT engine build.
Getting Started
Prerequisites
- Linux, x86_64 or aarch64, with an NVIDIA GPU. The wheels are
manylinux_2_34, so the host needs glibc 2.34 or newer — Ubuntu 22.04, RHEL 9 or later. - Driver 580 or newer. The dev image carries a CUDA 13.2 build of PyTorch. An older driver runs it only through the forward-compatibility shim, which we have measured hanging and crashing part-way through a run rather than merely running slowly — results taken on one are discarded, not corrected.
- Python 3.12. The released wheels are tagged
cp312, so pip finds no matching build on a newer interpreter. - Docker and the NVIDIA Container Toolkit, to build from source in the dev container. Installing the wheel needs neither.
PyTorch and the CUDA math libraries arrive as wheel dependencies, or in
nvcr.io/nvidia/pytorch:26.05-py3 when you use the container. Building the
extension from source outside a container needs a C++17 compiler and CUDA
headers as well —
docs/dev.md.
Release-qualified GPUs
H200, H100, A100, L40S, GB200 and GB300. Measured speedup, memory and accuracy
for each: docs/ref/benchmark.md.
BioIR runs on more than these. The support matrix lists every architecture the backend covers and which fused kernels apply to each; those devices work but are not part of this release's qualification.
Install
BioIR is published on PyPI, one wheel per CPU architecture:
pip install bionemo-ir
The wheel ships the kernels precompiled, so nothing in the install builds CUDA
and running it needs only the driver's libcuda.so.1. That is the whole
install if you are calling BioIR from your own code — the container below is
for working on BioIR itself. docs/install.md covers the
environment setup and the requirements in full.
Build from source
Configure SSH authentication with GitHub, then clone the repository and fetch its submodules and LFS objects.
git lfs install &&
GIT_LFS_SKIP_SMUDGE=0 \
git clone --recurse-submodules \
git@github.com:NVIDIA-BioNeMo/BioNeMo-Inference-Runtime.git &&
cd BioNeMo-Inference-Runtime
Then, build the dev image and open a shell in it:
docker/dev.sh
The image carries the dependencies; your checkout is bind-mounted, so install the package once inside and fold something:
pip install -e '.[dev]'
scripts/fetch_weights.sh --model boltz-2
python examples/folding/run_demo.py --output-dir output
Checkpoints come from their upstream publishers and need no NVIDIA credentials;
anything that cannot be fetched is skipped, and the tests needing it skip too.
Running scripts/run_tests.sh stages weights and runs the suite the way CI
does.
Building without a container needs more than a Python environment — see
docs/dev.md for
the prerequisites and the wheel build. The rest of that page covers daily
development; docs/ref/docker-images.md covers the
images and what docker/dev.sh mounts.
Documentation
BioIR documentation lives under docs/ and is published with Fern:
- Overview — what BioIR is and how to start
- Installation — requirements and release-wheel installation
- Quickstart — run a serial Boltz-2 prediction
- Ray multi-GPU inference — scale independent requests across visible GPUs
- Developer guide — build, test, stage weights, contribute
- API reference —
build_processor, model constructors, inputs/outputs - Architecture — five-stage pipeline and runtime design
- Config architecture — model
BaseConfigtree and pipeline stage configs - Support matrix — models, GPUs, and fused kernels
- Benchmarks — measured speedup and memory against OSS PyTorch
- Model weights — checkpoint resolution and staging
- Docker images — development and runtime images
- Coding guidelines — style, naming, and tooling
- Folding example — runnable
build_processordemo
Benchmarks
Methodology
Folding benchmarks over a bench set the shipped
rebuild_dataset.py builds from RCSB
and NVIDIA's MSA Search NIM — there is no dataset release to download.
Template-bearing samples included: both sides load every bundled MSA and attach
every listed template.
- One GPU, serial, one structure per forward call.
- Time only GPU-synchronized
model.forward(). Featurization, transfers, postprocessing, writing, and scoring stay outside the window. - Discard one warmup forward, then report one measured forward per sample.
- BioIR runs its default optimized config, with a CUDA graph on the diffusion module where supported.
- OSS runs its own inference script: eager always, plus
torch.compilewhen it passes a dynamic-shape probe. - Runtime knobs match on both sides — 200 sampling steps, 3 or 5 diffusion samples, and per-model recycling.
- Score written structures with OpenStructure lDDT and DockQ. Speedup is
OSS forward / BioIR forward; above 1 favors BioIR. - Future work will add additional Blackwell-optimized kernels.
Results
| Model | H100 | H200 |
|---|---|---|
| Boltz-2 | 1.78x / 2.65x | 1.74x / 2.54x |
| OpenFold3 | 1.55x / 2.02x | 1.54x / 2.03x |
| OpenFold2 / AlphaFold2 monomer | 2.55x / 2.60x | 2.61x / 2.66x |
| OpenFold2 / AlphaFold2 multimer | 2.66x / 2.77x | 2.61x / 2.75x |
| Protenix | — / 1.87x | — / 1.84x |
Geomean speedup, vs OSS torch.compile / vs OSS PyTorch eager; above 1 favours
BioIR. Protenix has no torch.compile path. Fourteen GPUs, per-model accuracy
and peak memory, and how to reproduce any of it:
docs/ref/benchmark.md.
The bench-perf-oss agent skill has
the full gates, environment isolation, result schema, and charting protocol.
Contributing
We welcome contributions. See contributing.md for
policy and docs/dev.md for the development workflow.
Citation
If you use BioIR in your research, please cite it via
CITATION.cff.
Contact / Support
- Bugs and feature requests: GitHub Issues
- Usage questions: GitHub Discussions
- Security vulnerabilities: see
SECURITY.md— do not file a public issue
License
NVIDIA-authored BioIR code is licensed under the Apache License 2.0. Distribution compliance material is available here:
- Third-party notices and attributions
- Full third-party license texts
- Gemmi 0.6.5 corresponding source, licensed under MPL-2.0 or LGPL-3.0-or-later; BioIR distributes it under the MPL-2.0 option
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distributions
Built Distributions
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file bionemo_ir-0.1.0rc1-cp312-cp312-manylinux_2_34_x86_64.whl.
File metadata
- Download URL: bionemo_ir-0.1.0rc1-cp312-cp312-manylinux_2_34_x86_64.whl
- Upload date:
- Size: 11.7 MB
- Tags: CPython 3.12, manylinux: glibc 2.34+ x86-64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
678afea10145e58678d09fefce955960584ed81262326c4b6ec9e6267db395e9
|
|
| MD5 |
f891fa1148aecbe97368fd5667b118ed
|
|
| BLAKE2b-256 |
61e35c6c96819918029df6ff3140801816ebaa38f3be8ee43892778d1826da9c
|
File details
Details for the file bionemo_ir-0.1.0rc1-cp312-cp312-manylinux_2_34_aarch64.whl.
File metadata
- Download URL: bionemo_ir-0.1.0rc1-cp312-cp312-manylinux_2_34_aarch64.whl
- Upload date:
- Size: 11.7 MB
- Tags: CPython 3.12, manylinux: glibc 2.34+ ARM64
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.14.7
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6b7abaf25a9f87fa94a51d7f4ff644eba3cff27f326eace5da3c751955861721
|
|
| MD5 |
19bce2f8170e9f3666b44472f56cfcd6
|
|
| BLAKE2b-256 |
4c09bdd2befe9ac7c67903586615ba640011f08dd1aa49ee97911e4eaee0c61e
|