Skip to main content

hawk

Write per-sample kernels in Python, with derivatives. Part of the RAPTOR family.

hawk: one kernel, two machines, and its derivatives

Write a numerical kernel for one sample as a plain Python function. hawk compiles it for the GPU or the CPU and derives its gradient (reverse mode) and tangent (forward mode).

CI Docs License

The RAPTOR family: hawk (write it), eagle (run it), aether (the C++/CUDA numerics underneath) and raptor (the shared contract); you are looking at hawk.

aether · hawk · eagle · raptor · the family

hawk is the kernel-authoring layer of the RAPTOR family: write a kernel once at the level of your application, and hawk derives its forward- and reverse-mode derivatives and compiles it for host or device through eagle. See where RAPTOR fits for the full four-part story.

Read the documentation — installation, a five-minute quickstart, tutorials, examples and the full Python API reference.

Write a kernel, get its gradient

import pathlib, tempfile

import hawk
from hawk import Kernel, Mutable, Scalar, Vector
from hawk.diff import vjp
from hawk.math import dot

@hawk.kernel
def energy(v: Vector[3], out: Mutable[Scalar]):
    out = 0.5 * dot(v, v)

energy_vjp = Kernel("energy_vjp", vjp(energy, wrt=("v",)))
work = pathlib.Path(tempfile.mkdtemp())
bundle = hawk.build([energy, energy_vjp], work, targets=("host",))
import numpy as np

v = np.array([[1.0, 2.0, 3.0, 4.0], [0.0, 1.0, 0.0, 1.0], [0.0, 0.0, 1.0, 1.0]])
e = np.zeros(4)
hawk.run(hawk.load(work, "energy"), v=v, out=e)
print("energy:", e)
# energy: [0.5 2.5 5.  9. ]
bar_out = np.ones(4)             # seed: d(loss)/d(energy) = 1 for every sample
bar_v = np.zeros((3, 4))
hawk.run(hawk.load(work, "energy_vjp"), v=v, bar_out=bar_out, bar_v=bar_v)
print("gradient d(energy)/dv:")
print(bar_v)
# [[1. 2. 3. 4.]
#  [0. 1. 0. 1.]
#  [0. 0. 1. 1.]]

energy = 0.5 |v|^2, so d(energy)/dv = v exactly — the printed gradient above is the input v. Output above is the notebook's own: docs/content/tutorials/04_gradients_backward.ipynb. hawk.diff.jvp derives the forward-mode (tangent) twin the same way.

Five-minute quickstart · install: pip install raptor-hawk (CPU route; it brings aether-dsc along), or pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]" for the GPU route

Writing kernels

A kernel is a plain Python function decorated @hawk.kernel, one sample at a time; its parameters are its declared planes (Scalar, Vector[3], Mutable, ...) and its body writes ordinary arithmetic:

import hawk
from hawk import Mutable, Scalar, Terminated, Vector
from hawk.math import maximum, norm

@hawk.kernel
def running_apogee(position: Vector[3], terminated: Terminated,
                    r_max: Mutable[Scalar]):
    r_max = maximum(r_max, norm(position))
Running statistics: read it before you write it

A Mutable parameter such as r_max is an ordinary local variable: read it before you assign it and you get the value it held when THIS launch began, which is how a running statistic is written — a launch reads what an earlier launch left behind, and writes a new value on top of it. Your buffer remembers; the kernel reads what it remembered. The host owns that memory: it must initialise r_max before the first launch and preserve it between launches — hawk never zeroes or scratch-allocates a plane a kernel reads from. That memory is not differentiable: a launch-start read is a recurrence across launches, and hawk keeps no tape across them, so hawk.diff.vjp/jvp refuse a kernel that reads one — differentiate the per-launch term on its own instead.

Kernel kinds and returned outputs

A kernel may commit through several sinks — the places a kernel's result goes: declared Mutable/Accum parameters (Output.named(...)), a return <expr> committed into a synthesised slot (a result sink hawk creates for you rather than one you declared as a parameter; Output.returned(ttype, slot=...) names it), or both together — a returned acceleration alongside an ordinary diagnostic Mutable the body also writes:

from hawk import Param
from hawk.ext import Kind, Output

ACCEL = Kind("drag", output=Output.returned(Vector[3], slot="acc"))

@ACCEL
def drag(velocity: Vector[3], k: Param, speed: Mutable[Scalar]):
    speed = norm(velocity)          # an ordinary named sink
    return (-k * norm(velocity)) * velocity   # the returned slot, "acc"

Kind(...) also has a class-form spelling — sugar over the same object, one kernel-authoring style closer to a plugin author's own vocabulary: an annotated class attribute is a vocabulary entry, a plain-assigned one sets one of Kind's own fields (output, sink, guard, ...), and anything else refuses.

from hawk.ext import KernelKind, Output

class Accel(KernelKind, slug="drag"):
    output = Output.returned(Vector[3], slot="acc")

@Accel
def drag(velocity: Vector[3], k: Param):
    return (-k * norm(velocity)) * velocity

Install

pip install raptor-hawk                                     # CPU only: no GPU, driver or CUDA packages needed
pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"    # GPU: hawk compiles, eagle runs; [cuda13] on both for CUDA 13

Linux x86_64, CPython 3.10-3.13, and a host g++ 11 or newer. raptor-hawk pulls aether-dsc (the sealed C++ headers hawk compiles against) automatically. No nvcc or CUDA toolkit is needed. Tested with CUDA 12.6 and CUDA 13.0 (CUDA 12.6 or newer).

NVIDIA packages come only through the extras. raptor-hawk[cuda12] pulls cuda-bindings 12, nvidia-cuda-nvrtc-cu12 and nvidia-cuda-cccl-cu12; raptor-hawk[cuda13] pulls cuda-bindings 13, nvidia-cuda-nvrtc 13 and nvidia-cuda-cccl 13 (CUDA 13's wheels have no -cu13 suffix); raptor-eagle[cuda12] / [cuda13] pull CuPy (cupy-cuda12x / cupy-cuda13x, with the CUDA headers CuPy compiles against). Without an extra pip installs no NVIDIA package: you get the CPU route, or the GPU route through a CUDA setup you already have. Pick the extra matching the CUDA version your driver reports (nvidia-smi, top right).

Platforms: built and tested on Linux x86_64 only so far (CPython 3.10–3.13), on NVIDIA GPUs from Pascal (Quadro P2000) and Turing (Tesla T4). There are no wheels for macOS, Windows or ARM yet, and WSL2 is untested. raptor-core and aether-dsc are pure Python and install anywhere.

See installation for the CPU-only route (no GPU, driver or NVRTC needed at all) and every other detail.

Performance

Every number in this section is read from eagle's committed cards (benchmarks/perf_card/card_quadro-p2000.md, benchmarks/rk78_card/card_quadro-p2000.md). Kernels authored with hawk are what eagle's performance card times: hawk + eagle (eagle.simulate) is 3.9×–47× faster than NVIDIA Warp (3.9×), JAX (8.7×), CuPy and PyTorch (47×) at 1,000,000 RK4 oscillators finishing at different times — 156 ms on a Quadro P2000, using 3.8×–5.2× less GPU memory than CuPy and PyTorch (a million samples in 50 MiB, 1.31× the bare minimum state; CuPy 5.1×, PyTorch 6.9×; NVIDIA Warp 64 MiB, 1.7×) — and the same hawk kernel, unmodified, integrates 1,000,000 adaptive RK7(8) orbits in 13.0 s on the GPU or 7.15 s on 8 CPU threads alone. This is the regime hawk + eagle's compaction targets (many samples stopping at different times); in other regimes the other tools hold their own — on a dense workload where nothing finishes early Warp ties it (693 ms vs 699 ms at N = 1,000,000) and is level at N = 10,000 (7.22 ms vs 7.22 ms), and on this FP64-weak development card the CPU arm is the faster one there. See eagle's performance page for every number, every arm, including the ones where it doesn't win.

Going deeper

The pip-only device compile, directly

The device compile hawk.artifact.build/build_bundle use when no nvcc is on $PATH goes through hawk.compile.device/hawk.compile.cubin; the same entry points are usable directly, for an ad hoc CUDA C++ source string that was never authored as a traced hawk kernel at all:

from hawk.compile import cubin, cubin_available

if cubin_available():
    image = cubin(source)          # DeviceImage(image=<SASS bytes>, target="cubin", arch="sm_XX")

cubin_available() never raises; device(source, target="ptx") is the alternative when PTX (the portable intermediate form NVRTC can also emit) is what a caller needs, guarded against a driver too old to JIT a newer NVRTC's PTX (target="cubin" needs no such guard — it is already SASS, the GPU's own machine code). Headers come from aether_dsc's sealed payload — a copy of the headers hawk's device compiler reads from directly, bundled into the wheel — never from disk.

Build and test from a source checkout (development)

Clone aether and eagle next to this checkout (so ../aether and ../eagle sit beside hawk/), then install aether's sealed header payload before hawk itself:

pip install ../aether/dsc
pip install -e .[test]
pytest -m "not gpu"

The compiled half (hawk._core) is a nanobind extension built through scikit-build-core; pip install -e . builds it via CMakeLists.txt, resolving the aether/eagle C++ header roots from those sibling checkouts (or $HAWK_AETHER_INCLUDE/$HAWK_EAGLE_INCLUDE). -m "not gpu" deselects the tests that need a CUDA GPU and cupy; most of the rest additionally need eagle installed as a Python package to execute traced kernels on host or device (see the "Python package" section of eagle's README).


hawk write it (this repo) · eagle run it · aether the numerics core · raptor the shared contracts — the RAPTOR family

Apache-2.0 (see LICENSE and NOTICE) · cite via "Cite this repository" (CITATION.cff; each tagged release is archived on Zenodo with its own DOI once the first one exists) · built to make GPU computing accessible on modest hardware, for research and education. Collaboration is the point, and a citation is the currency — get in touch.

Metadata

Release files for raptor-hawk 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for raptor-hawk 0.3.0
File Size Uploaded
raptor_hawk-0.3.0.tar.gz 1.4 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for raptor-hawk 0.3.0
File
raptor_hawk-0.3.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.3.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.26+ x86-64, Linux glibc 2.28+ x86-64 Details
raptor_hawk-0.3.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.26+ x86-64, Linux glibc 2.28+ x86-64 Details
raptor_hawk-0.3.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.10 CPython 3.10 Linux glibc 2.26+ x86-64, Linux glibc 2.28+ x86-64 Details

Total release size: 3.3 MB

Release files / raptor_hawk-0.3.0.tar.gz

Download URL raptor_hawk-0.3.0.tar.gz
Size 1.4 MB
Tags Source
SHA-256 checksum
How to use checksums
920554ff50b35a64b310dce8fa3e32816baea2514aef34349524ae40f9bdca3e
BLAKE2b-256 checksum
How to use checksums
ccd7d5c12a19d4996587c035eb3625ed259705e83a32920ec363ee71b8b05b57
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / raptor_hawk-0.3.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.3.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 484.1 kB
Tags CPython 3.13 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
a301b217b209820ba05624e30c2d63e9aab7a75fb2218cd4ea05cf173b965906
BLAKE2b-256 checksum
How to use checksums
74ed4883be33ec0b2b7ca2634e3b9cd49a206750198036b45670346d09e065ca
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / raptor_hawk-0.3.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.3.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 484.2 kB
Tags CPython 3.12 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
4e7995777e6fc3c8c6184caf178c7d6ecf528045907571b5f6db2e6b2421aa89
BLAKE2b-256 checksum
How to use checksums
4998cb7988a8ed505d20864b186a297adfd8f337935daf5169d7c16bc6e65e69
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / raptor_hawk-0.3.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.3.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 484.6 kB
Tags CPython 3.11 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
233d4e41fa822a36e24d8a6b15caaf1d92b784dfc801d784ebedc7f5829ad1c0
BLAKE2b-256 checksum
How to use checksums
c197af301b4f42983c1ff5a01e598c17ae83a4f6fb44acad1e47b7f283d4e9bf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / raptor_hawk-0.3.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.3.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 485.0 kB
Tags CPython 3.10 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
338a37a08a05f98dd4321f1a2efacaa4555e1fbd3187294e998ef666ba6d094f
BLAKE2b-256 checksum
How to use checksums
8e39f39f4c5e090c10c054cecbbd13499d627672737e0bf331ade393fa9ecbd8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.4.0

9 release files

0.3.1

9 release files

This release

0.3.0 This release

5 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page