Skip to main content

NAB Python

Neighbourhood Algorithm Bayes (NAB) — a tool for Bayesian analysis using Markov Chain Monte Carlo (MCMC) and Gibbs Sampling based on Voronoi tessellations.

At a glance

What Bayesian appraisal of parameter ensembles: MCMC with Gibbs sampling over Voronoi cells of sampled models (Neighbourhood Algorithm, Bayesian stage).
Install pip install nab-bayes (PyPI); Docker image published by CI.
Quality 470+ automated tests; CI on every push.
Performance Benchmarks and their limits: performance summary.
Maintainer Denis Samatov (Heriot-Watt TPU Center); see the commit history for contributors.

🚀 Quick Start

1. Install Dependencies

Install the published package from PyPI:

pip install nab-bayes

For notebook usage with inline Plotly figures, install the visualization extras:

pip install "nab-bayes[viz]"

For a reproducible developer setup from a repository checkout:

To automatically detect your OS/GPU configuration and install the correct accelerated PyTorch build:

git clone https://github.com/denis-samatov/neighbourhood-algorithm-bayes.git
cd neighbourhood-algorithm-bayes
python install_gpu.py

This script handles the heavy lifting of resolving CUDA on Windows/Linux or MPS on Apple Silicon, then installs all optional components.

📦 Custom/Pip Installation (with CUDA 12.1)

pip install -r requirements-gpu.txt

💻 Standard CPU Installation

pip install -e ".[dev,viz,nf,hmc]"

After installation, run tests to verify:

pytest -q

2. Minimum Example

Using the test data provided in data/:

nab run data/punq20/Distributions.txt data/punq20/misfit.tsv \
    output/solutionProb.tsv output/sortedProb.tsv \
    --chains 5 --burn-in 10000 --length 50000 --refresh 1000 \
    --plot --name quickstart_run

This will run the MCMC sampling using 5 chains and generate an interactive HTML report. See User Guide for full CLI options.

3. Python / Notebook API

import nab

result = nab.fit(
    distribution_file="data/punq20/Distributions.txt",
    misfit_file="data/punq20/misfit.tsv",
    run_dir="output/notebook_run",
    chains=5,
    burn_in=10_000,
    chain_length=50_000,
    sampler="gibbs",
    diagnostics=True,
)

result.posterior.head()
result.plot_posterior()

For notebook-oriented usage, see Python API.

📚 Documentation

Detailed documentation is available in the docs/ directory:

  1. User Guide — Installation, data preparation, CLI usage, and output formats.
  2. End-to-End Pipeline — Post-run workflow for adaptive prior narrowing, Top-N posterior analysis, and related diagnostics.
  3. Mathematical Description — Theory behind NAB, Voronoi approximation, and MCMC sampling.
  4. Concepts — High-level explanation and analogies.
  5. Python API — nab.fit(...), nab.load_run(...), nab.analyze(...), and notebook workflows.
  6. Analytical Validation — 6 synthetic validation experiments against analytical ground truth.
  7. Architecture — Code structure, data flow, and module descriptions.
  8. Style Guide — Rules for contributing text and docstrings.

🏗️ Architecture Overview & Data Flow

NAB is designed as a pipeline with clear separation between data loading, evaluation, simulation, and reporting.

Data Flow Pipeline:

  1. Input: The user provides a Distributions file defining parameter bounds/types, and a Misfit (solutions) file containing previously evaluated models and their errors.
  2. Initialization: The BayesianEvaluator constructs the Voronoi tessellation space across the bounded parameters.
  3. Sampling: NAB supports five sampler modes: Gibbs (default), adaptive Gibbs, parallel tempering, surrogate-assisted HMC, and Normalizing Flows. Depending on the mode, the algorithm either walks the Voronoi cells directly or fits a smooth approximation to accelerate posterior exploration.
  4. Validation/Diagnostics: The chains are evaluated for convergence (R-hat, Effective Sample Size).
  5. Output: Posterior probabilities are calculated, normalized, and saved to solutionProb.tsv and sortedProb.tsv. A chain_history trace is optional.
  6. Visualization: If --plot is enabled, trace plots and posterior distributions are rendered to an interactive HTML report.

⚙️ Configuration & CLI

The primary entry point is the nab CLI. You can also invoke programmatically via nab.core.orchestrator.run_analysis.

For notebook-first usage, prefer the higher-level public API:

import nab

result = nab.fit(...)
analysis = nab.analyze(result, top_n=20)

Supported sampler modes are gibbs, adaptive, tempered, hmc, and nf. The User Guide includes a comparison table describing what each sampler does and when to use it.

gibbs, adaptive, and tempered are the reference discrete NAB samplers. hmc and nf are intentionally marked as approximate / experimental in manifests, reports, and diagnostics.

Evidence Tiers

NAB distinguishes baseline_scientific (gibbs, adaptive, tempered) and approximate_exploratory (hmc, nf) evidence tiers. For full details on trusted benchmark regeneration, verification, and cross-language equivalence criteria, see Evidence Tiers.

For CLI arguments, supported distributions, and input formats, see the User Guide. For the post-processing workflow, see the Pipeline Guide.

🎯 Practical Interpretation of the Bayesian Result

In this repository, the Bayesian NAB output is positioned as a starting approximation rather than a final answer by itself.

Use it as:

  1. A warm start (MAP / top-probability models) for building an approximation of the target function.
  2. A probabilistic region-of-interest selector (where to spend expensive forward evaluations next).
  3. An initialization stage before local refinement, surrogate fitting, or physics-constrained optimization.

If you want a narrower follow-up analysis instead of only the raw NAB outputs, the recommended next step is:

  1. Run nab adaptive-prior to propose a tighter prior with ESS/PSIS checks.
  2. Run nab dashboard --top-n N to export and visualize the highest-probability Top-N subset under the adaptive prior.

🧪 Testing & CI

To ensure reliability, NAB maintains an extensive pytest suite and uses Ruff for code quality.

Run tests:

pytest -q

Run linter and formatter:

ruff check .
ruff format .

Benchmark data/nab_data:

A reproducible performance benchmark for the prepared data/nab_data ensembles measures data-loading (cold/warm + peak heap) and discrete Gibbs throughput:

python scripts/benchmark_nab_data.py --dataset static_maas --burn-in 100 --length 1000
python scripts/benchmark_nab_data.py --dataset dyn --no-sampling          # loading only

The script is read-only with respect to data/ and warms up Numba JIT before timing. Pass --json out.json to capture machine-readable results.

Note on Java Parity: This is a Python port of the original Java implementation. Posterior-level equivalence is validated rather than trajectory identity. See Evidence Tiers for acceptance criteria.

🚑 Troubleshooting

Top issues:

  1. DistributionParseError: The Distributions.txt format is extremely strict. Ensure no trailing whitespaces on lines and use tab/space delimiters as documented.
  2. NaNs in output: Check if the misfit values in your input TSV are too large, leading to numerical underflow. Use the --normalize flag (enabled by default) to handle large scales.
  3. MCMC Chains taking too long: Reduce --length for a quick test, or increase OMP_NUM_THREADS / NUMBA_NUM_THREADS env vars to leverage multi-core CPU capabilities in the Voronoi kernels.
  4. No interactive plot: Ensure you added --plot and check your console for the file:// link generated.

🛡️ Security

Do not commit sensitive data or proprietary configurations into data/ or output/. Outputs are added to .gitignore, but double-check before pushing any custom project data. No secret keys are required to use this software.

🔬 Origins

This repository is a Python port of the Java implementation of NAB, which itself is a re-implementation of the original NA-Bayes algorithm by Malcolm Sambridge (Australian National University).

Founding papers:

  1. Sambridge, M. Geophysical Inversion with a Neighbourhood Algorithm I: searching a parameter space. Geophys. J. Int., 138, 479–494, 1999.
  2. Sambridge, M. Geophysical Inversion with a Neighbourhood Algorithm II: appraising the ensemble. Geophys. J. Int., 138, 727–746, 1999.

Metadata

Release files for nab-bayes 1.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nab-bayes 1.2.0
File Size Uploaded
nab_bayes-1.2.0.tar.gz 1.2 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for nab-bayes 1.2.0
File Interpreter ABI Platform
nab_bayes-1.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.5 MB

Release files / nab_bayes-1.2.0.tar.gz

Download URL nab_bayes-1.2.0.tar.gz
Size 1.2 MB
Tags Source
SHA-256 checksum
How to use checksums
1c1597fbe10e07880e684a75258c39c36e96f12880cd50e38b4b12d6f79f2f9d
BLAKE2b-256 checksum
How to use checksums
5a1fed43f89d0441638d9afb07c8f94597bab3bef1159d80ac1152163e4fcff1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / nab_bayes-1.2.0-py3-none-any.whl

Download URL nab_bayes-1.2.0-py3-none-any.whl
Size 238.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
306959f2418cd9a9b70ed5dd24dd966a209984860f5eb9c574a08a72cef63255
BLAKE2b-256 checksum
How to use checksums
30803c6018af69a3637e3479c7e7447807b4a1384825663fb99a5f6f5c4d59e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

1.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page