Skip to main content

nvSubquadratic

A unified PyTorch-native library for subquadratic alternatives to quadratic attention methods.

Overview

nvSubquadratic consolidates efforts from across NVIDIA Research teams (nvResearch, NeMo, BioNeMo) into a single, consistent API. Currently supporting multi-dimensional (1D, 2D, 3D) Hyena operators with optimized CUDA kernels. Hyena operators provide subquadratic alternatives to attention mechanisms, achieving O(N) or O(N log N) complexity compared to the O(N²) scaling of traditional attention.

Dependencies

subquadratic-ops: This library depends on subquadratic-ops for high-performance CUDA kernels. The subquadratic-ops library provides optimized implementations of:

  • B2B CausalConv1d: Back-to-back causal convolutions for striped Hyena architectures
  • CausalConv1d: Standard causal convolutions with various kernel sizes (2-256)
  • FFT CausalConv1d: FFT-based causal convolutions for large kernel sizes (up to 8K-16M)
  • Fused FFT Conv2d: single-launch 2D FFT convolution running natively in fp32/fp16/bf16 (spatial dims up to 64 per axis); requires subquadratic-ops-torch >= 0.3.0, and compute capability 9.0+ (Hopper/Blackwell) for extents above 32 per axis

Requirements:

  • CUDA-compatible NVIDIA GPU (Ampere or newer)
  • CUDA Toolkit 13.0 or higher
  • NVIDIA driver >= 580 (the CUDA 13.x minimum), or the CUDA forward-compatibility package
  • Python 3.10 or higher

quack-kernels (optional):

  • Used by RMSNorm for a fused CUDA kernel when available.
  • Supported GPUs: Hopper and Blackwell only (H100, B200, B300). There is no quack-kernels build for Ampere (e.g. RTX A6000, A100) or older architectures.
  • On unsupported GPUs or when quack is not installed, RMSNorm uses a pure-PyTorch fallback automatically. No separate “version” fixes Ampere; use the fallback or run on H100/B200 for the kernel.
  • Optional install: pip install quack-kernels (on H100/B200); or pip install -e ".[quack]" if using the optional dependency group.

Architecture

nvSubquadratic provides a high-level PyTorch interface that depends on the subquadratic-ops library for optimized CUDA kernels. This separation provides clear boundaries between API design and performance optimization:

  • nvSubquadratic: Focuses on API design, user experience, and PyTorch integration
  • subquadratic-ops: Focuses on kernel optimization and CUDA performance
  • megatron-core: Provides distributed training and model parallelism capabilities

Installation

PyPI

pip install nvsubquadratic

This installs the full training/experiment stack — nvSubquadratic targets GPU workflows. Requires Python 3.10+.

Optional extras:

pip install "nvsubquadratic[cuda]"         # accelerated fused FFT-conv / causal-conv CUDA kernels
pip install "nvsubquadratic[quack]"        # fused RMSNorm kernel (Hopper/Blackwell only)
pip install "nvsubquadratic[dali]"         # NVIDIA DALI data pipelines for the ImageNet/Well examples (~400 MB)
pip install "nvsubquadratic[distributed]"  # megatron-core, for context-parallel / distributed training
pip install "nvsubquadratic[baselines]"    # timm, for the ConvNeXt UNet baseline models
pip install "nvsubquadratic[all]"          # all of the above

The accelerated CUDA kernels ([cuda]) are a source build that requires nvcc, so they are kept out of the core install — this is what lets pip install nvsubquadratic succeed in environments without the CUDA toolkit (e.g. a downstream project's CPU CI). The operators default to the portable torch.fft backend; selecting fft_backend="subq_ops" (or "subq_ops_fused") without [cuda] installed raises a clear ImportError pointing you to the extra.

On 2D problems with spatial dims of at most 64 per axis, fft_backend="subq_ops_fused" is the fastest option: it fuses the whole FFT-conv pipeline into one launch and runs it natively in bf16/fp16 instead of upcasting to fp32 — measured at 1.9-4.9x over torch_fft and 1.3-2.5x over subq_ops on an H100 across spatial extents 16-64 (batch 8, hidden 768, forward+backward, bf16). The margin over torch_fft grows with spatial extent; reproduce with benchmarks/ops/bench_fused_fftconv2d.py. Models already written against torch_fft can pick up the same kernel under torch.compile without a config change — see the torch.compile lowering.

Package Manager

This project uses pip with pyproject.toml for dependency management. A Pipfile.lock is maintained for nSpect security scanning compliance.

Open in VS Code and select "Reopen in Container". The devcontainer extension will automatically build the Docker image and set up the development environment with all dependencies pre-installed.

Docker

# Build and run
docker build -t nvsubquadratic:dev .
docker run --gpus all -p 8888:8888 -v $(pwd):/workspaces/nvSubquadratic nvsubquadratic:dev

The Dockerfile builds NVIDIA Apex from source for a broad set of NVIDIA archs by default (7.5;8.0;8.6;8.9;9.0;10.0;12.0 — Turing through Blackwell). Build-args let you tune the compile:

  • TORCH_CUDA_ARCH_LIST — narrow to your GPU(s) to speed up the build (e.g. 9.0 for H100, 8.6 for A6000, 8.9 for L4). Also applied to the patched mamba-ssm / causal-conv1d source builds (upstream otherwise hardcodes sm_75..sm_120).
  • MAX_JOBS — number of parallel nvcc/ninja jobs. Defaults to unconstrained. Set to 1 if the build OOMs or gcc ICEs (typical under qemu emulation for arm64).
  • NVCC_THREADS — nvcc --threads for the mamba stack (default 4; use 1 under QEMU).
docker build \
    --build-arg TORCH_CUDA_ARCH_LIST="9.0" \
    -t nvsubquadratic:dev .

Enroot (SLURM clusters)

For SLURM deployments that use enroot/pyxis, scripts/slurm/enroot/build_sqsh.sh builds the Docker image and converts it to an enroot .sqsh in one step. It selects the right TORCH_CUDA_ARCH_LIST, MAX_JOBS, and NVCC_THREADS per platform:

# H100 (x86-64, default)
scripts/slurm/enroot/build_sqsh.sh

# GB200 (ARM64) — QEMU on x86: keep ≥64GB free RAM+swap; defaults MAX_JOBS=1 NVCC_THREADS=1
PLATFORM=arm64 scripts/slurm/enroot/build_sqsh.sh

Apptainer

# Build SIF (add --fakeroot if required on your system)
apptainer build nvsubquadratic.sif nvsubquadratic.def

# Interactive shell with GPUs and live code from your checkout
apptainer shell --nv --bind $(pwd):/workspaces/nvSubquadratic nvsubquadratic.sif

# Run a command inside the image (example: tests)
apptainer exec --nv --bind $(pwd):/workspaces/nvSubquadratic nvsubquadratic.sif python -m pytest nvsubquadratic/ tests/

# Use the default runscript (starts Jupyter Lab as defined in the .def)
apptainer run --nv --bind $(pwd):/workspaces/nvSubquadratic nvsubquadratic.sif --no-browser
bash setup_conda_env.sh
conda activate nvsubquadratic

This script creates the nvsubquadratic conda environment with Python 3.12 and PyTorch 2.14 (CUDA 13.0), installs all dev dependencies, builds NVIDIA Apex from source, and installs quack-kernels.

Local Installation (venv)

# Create and activate a virtual environment
python3 -m venv venv
source venv/bin/activate

# Install PyTorch with CUDA support first (before package dependencies)
pip install torch==2.14.0 torchvision==0.29.0 --index-url https://download.pytorch.org/whl/cu130

# Install development dependencies
pip install -r requirements-dev.txt

# Install the package in editable mode (installs remaining dependencies from pyproject.toml)
pip install --no-build-isolation -e .

Development

Pre-commit hooks are automatically installed in the dev container. For other installation methods:

pre-commit install
pre-commit install --hook-type pre-push

Updating Dependencies for Security Scanning

This project maintains a Pipfile.lock for nSpect security scanning compliance. When you update dependencies in pyproject.toml, regenerate the lock file:

# Install pipenv (if not already installed)
pip install pipenv

# Regenerate Pipfile.lock
pipenv lock

# Note: Continue using pip for actual installations (pip install -e .)

Testing

See tests/README.md for full details on test suites, markers, and SLURM usage.

# Unit tests (CPU-safe, no external data needed)
PYTHONPATH=. python -m pytest tests/ -m "not nightly" -v -o addopts=""

# Nightly validation (requires GPU, DALI, ImageNet, wandb)
source .env && PYTHONPATH=. python -m pytest tests/ -m nightly -v -o addopts=""

Documentation

All public classes and functions carry Google-style docstrings with math context, shape annotations, and paper references. See CONVENTIONS.md for the style guide and PR checklist.

Viewing the docs

The API reference is built with Sphinx. Sources live under docs/ and the rendered site is published to the gh-pages branch on every push to main via .github/workflows/docs.yml.

Build and preview locally:

pip install -r docs/requirements.txt
pip install -e . --no-deps
make -C docs html SPHINXBUILD="python -m sphinx"
python -m http.server 8000 --directory docs/_build/html

Open http://localhost:8000 to browse. The autosummary stubs under docs/api/generated/ are regenerated on every build (gitignored).

While editing, you can also hover over any symbol in VS Code / PyCharm to see the rendered docstring, or run help(SomeClass) in a REPL.

CI

GPU tests run automatically on pull requests via a self-hosted runner. Runner provisioning is maintained out-of-tree; contact the maintainers for access.

Pre-commit Hooks

On commit:

  • Code formatting (Ruff)
  • Import sorting (Ruff)
  • YAML validation
  • Markdown formatting
  • Secret detection

Contributing

Contributions are welcome — see CONTRIBUTING.md for the DCO sign-off requirement and the PR/issue flow. Pull requests from external forks run through the same CI pipeline; the GPU stage requires a codeowner (.github/CODEOWNERS) to approve workflow runs from outside collaborators before the self-hosted runner picks them up — this is the standard GitHub "Require approval for outside collaborators" gate.

For security-sensitive findings, please follow SECURITY.md instead of opening a public issue or PR.

Metadata

Release files for nvsubquadratic 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for nvsubquadratic 0.2.0
File Size Uploaded
nvsubquadratic-0.2.0.tar.gz 280.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for nvsubquadratic 0.2.0
File Interpreter ABI Platform
nvsubquadratic-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 566.2 kB

Release files / nvsubquadratic-0.2.0.tar.gz

Download URL nvsubquadratic-0.2.0.tar.gz
Size 280.1 kB
Tags Source
SHA-256 checksum
How to use checksums
ff93a2a25e143491f8094e9a70d988c8582c01a04dcd8c00804288363c1c113a
BLAKE2b-256 checksum
How to use checksums
b001cc3bcf336694228d75e54b1f4d5ee8872a5530bface8399491d592273d9f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release files / nvsubquadratic-0.2.0-py3-none-any.whl

Download URL nvsubquadratic-0.2.0-py3-none-any.whl
Size 286.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bc482a75fc57c96a01c8518c4f5c859a2d85a9f9d3f95e9b601085929c0bdf74
BLAKE2b-256 checksum
How to use checksums
8a0d51e3829534e88f94b79d6bea95a3f54fa8a02cb2445750998db195d16913
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page