Skip to main content

hawk

Write per-sample kernels in Python, with derivatives. Part of the RAPTOR family.

hawk: one kernel, two machines, and its derivatives

Write a numerical kernel for one sample as a plain Python function. hawk compiles it for the GPU or the CPU and derives its gradient (reverse mode) and tangent (forward mode).

CI Docs License PyPI DOI

The RAPTOR family
hawk eagle
aether raptor

hawk is the kernel-authoring layer of the RAPTOR family: write a kernel once at the level of your application, and hawk derives its forward- and reverse-mode derivatives and compiles it for host or device through eagle. See where RAPTOR fits for the full four-part story.

Read the documentation — installation, a five-minute quickstart, tutorials, examples and the full Python API reference.

Write a kernel, get its gradient

import pathlib, tempfile

import hawk
from hawk import Kernel, Mutable, Scalar, Vector
from hawk.diff import vjp
from hawk.math import dot

@hawk.kernel
def energy(v: Vector[3], out: Mutable[Scalar]):
    out = 0.5 * dot(v, v)

energy_vjp = Kernel("energy_vjp", vjp(energy, wrt=("v",)))
work = pathlib.Path(tempfile.mkdtemp())
bundle = hawk.build([energy, energy_vjp], work, targets=("host",))
import numpy as np

v = np.array([[1.0, 2.0, 3.0, 4.0], [0.0, 1.0, 0.0, 1.0], [0.0, 0.0, 1.0, 1.0]])
e = np.zeros(4)
hawk.run(hawk.load(work, "energy"), v=v, out=e)
print("energy:", e)
# energy: [0.5 2.5 5.  9. ]
bar_out = np.ones(4)             # seed: d(loss)/d(energy) = 1 for every sample
bar_v = np.zeros((3, 4))
hawk.run(hawk.load(work, "energy_vjp"), v=v, bar_out=bar_out, bar_v=bar_v)
print("gradient d(energy)/dv:")
print(bar_v)
# [[1. 2. 3. 4.]
#  [0. 1. 0. 1.]
#  [0. 0. 1. 1.]]

energy = 0.5 |v|^2, so d(energy)/dv = v exactly — the printed gradient above is the input v. Output above is the notebook's own: docs/content/tutorials/04_gradients_backward.ipynb. hawk.diff.jvp derives the forward-mode (tangent) twin the same way.

Five-minute quickstart · install: pip install raptor-hawk (CPU route; it brings aether-dsc along), or pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]" for the GPU route

Writing kernels

A kernel is a plain Python function decorated @hawk.kernel, one sample at a time; its parameters are its declared planes (Scalar, Vector[3], Mutable, ...) and its body writes ordinary arithmetic:

import hawk
from hawk import Mutable, Scalar, Terminated, Vector
from hawk.math import maximum, norm

@hawk.kernel
def running_apogee(position: Vector[3], terminated: Terminated,
                    r_max: Mutable[Scalar]):
    r_max = maximum(r_max, norm(position))
Running statistics: read it before you write it

A Mutable parameter such as r_max is an ordinary local variable: read it before you assign it and you get the value it held when THIS launch began, which is how a running statistic is written — a launch reads what an earlier launch left behind, and writes a new value on top of it. Your buffer remembers; the kernel reads what it remembered. The host owns that memory: it must initialise r_max before the first launch and preserve it between launches — hawk never zeroes or scratch-allocates a plane a kernel reads from. That memory is not differentiable: a launch-start read is a recurrence across launches, and hawk keeps no tape across them, so hawk.diff.vjp/jvp refuse a kernel that reads one — differentiate the per-launch term on its own instead.

Kernel kinds and returned outputs

A kernel may commit through several sinks — the places a kernel's result goes: declared Mutable/Accum parameters (Output.named(...)), a return <expr> committed into a synthesised slot (a result sink hawk creates for you rather than one you declared as a parameter; Output.returned(ttype, slot=...) names it), or both together — a returned acceleration alongside an ordinary diagnostic Mutable the body also writes:

from hawk import Param
from hawk.ext import Kind, Output

ACCEL = Kind("drag", output=Output.returned(Vector[3], slot="acc"))

@ACCEL
def drag(velocity: Vector[3], k: Param, speed: Mutable[Scalar]):
    speed = norm(velocity)          # an ordinary named sink
    return (-k * norm(velocity)) * velocity   # the returned slot, "acc"

Kind(...) also has a class-form spelling — sugar over the same object, one kernel-authoring style closer to a plugin author's own vocabulary: an annotated class attribute is a vocabulary entry, a plain-assigned one sets one of Kind's own fields (output, sink, guard, ...), and anything else refuses.

from hawk.ext import KernelKind, Output

class Accel(KernelKind, slug="drag"):
    output = Output.returned(Vector[3], slot="acc")

@Accel
def drag(velocity: Vector[3], k: Param):
    return (-k * norm(velocity)) * velocity

Threading

Free-threaded CPython (3.13t/3.14t) is supported: the compiled modules declare GIL-free operation and the GIL stays disabled after import. Any number of threads may call module-level functions, build, compile, plan and cache concurrently. Distinct objects may be used from distinct threads without synchronisation. One stateful object (a stream, capture, launcher, graph, composer, plan, pipeline, active set, host kernel or arg block) shared by several threads is memory-safe — its calls serialise and a consumed object raises — but the ORDER of those calls is the caller's responsibility, exactly as for a NumPy array or a CuPy stream. CUDA adds two rules of its own: a stream capture is begun, filled and ended by one thread, and while any capture is open no thread may synchronise the whole device (stream-level synchronisation is fine). GPU routes are supported on free-threaded 3.14; on 3.13t the CPU route runs. hawk's compile cache is shared safely between threads, and a sealed-path kernel's cache key follows the user headers it includes.

Install

pip install raptor-hawk                                     # CPU only: no GPU, driver or CUDA packages needed
pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"    # GPU: hawk compiles, eagle runs; [cuda13] on both for CUDA 13

Linux x86_64, CPython 3.9-3.14 (including free-threaded 3.13t and 3.14t), and a host g++ 11 or newer. raptor-hawk pulls aether-dsc (the sealed C++ headers hawk compiles against) automatically. No nvcc or CUDA toolkit is needed. Tested with CUDA 12.6 and CUDA 13.0 (CUDA 12.6 or newer).

NVIDIA packages come only through the extras. raptor-hawk[cuda12] pulls cuda-bindings 12, nvidia-cuda-nvrtc-cu12 and nvidia-cuda-cccl-cu12; raptor-hawk[cuda13] pulls cuda-bindings 13, nvidia-cuda-nvrtc 13 and nvidia-cuda-cccl 13 (CUDA 13's wheels have no -cu13 suffix); raptor-eagle[cuda12] / [cuda13] pull CuPy (cupy-cuda12x / cupy-cuda13x, with the CUDA headers CuPy compiles against). Without an extra pip installs no NVIDIA package: you get the CPU route, or the GPU route through a CUDA setup you already have. Pick the extra matching the CUDA version your driver reports (nvidia-smi, top right).

Platforms: built and tested on Linux x86_64 only so far (CPython 3.9–3.14, including free-threaded 3.13t and 3.14t), on NVIDIA GPUs from Pascal (Quadro P2000) and Turing (Tesla T4). There are no wheels for macOS, Windows or ARM yet, and WSL2 is untested. raptor-core and aether-dsc are pure Python and install anywhere. On free-threaded 3.13t and 3.14t, hawk runs GIL-free.

See installation for the CPU-only route (no GPU, driver or NVRTC needed at all) and every other detail.

Performance

Every number in this section is read from eagle's committed cards (benchmarks/perf_card/card_quadro-p2000.md, benchmarks/rk78_card/card_quadro-p2000.md). Kernels authored with hawk are what eagle's performance card times: hawk + eagle (eagle.simulate) is 3.9×–47× faster than NVIDIA Warp (3.9×), JAX (8.7×), CuPy and PyTorch (47×) at 1,000,000 RK4 oscillators finishing at different times — 156 ms on a Quadro P2000, using 3.8×–5.2× less GPU memory than CuPy and PyTorch (a million samples in 50 MiB, 1.31× the bare minimum state; CuPy 5.1×, PyTorch 6.9×; NVIDIA Warp 64 MiB, 1.7×) — and the same hawk kernel, unmodified, integrates 1,000,000 adaptive RK7(8) orbits in 13.0 s on the GPU or 7.15 s on 8 CPU threads alone. This is the regime hawk + eagle's compaction targets (many samples stopping at different times); in other regimes the other tools hold their own — on a dense workload where nothing finishes early Warp ties it (693 ms vs 699 ms at N = 1,000,000) and is level at N = 10,000 (7.22 ms vs 7.22 ms), and on this FP64-weak development card the CPU arm is the faster one there. See eagle's performance page for every number, every arm, including the ones where it doesn't win.

Going deeper

The pip-only device compile, directly

The device compile hawk.artifact.build/build_bundle use when no nvcc is on $PATH goes through hawk.compile.device/hawk.compile.cubin; the same entry points are usable directly, for an ad hoc CUDA C++ source string that was never authored as a traced hawk kernel at all:

from hawk.compile import cubin, cubin_available

if cubin_available():
    image = cubin(source)          # DeviceImage(image=<SASS bytes>, target="cubin", arch="sm_XX")

cubin_available() never raises; device(source, target="ptx") is the alternative when PTX (the portable intermediate form NVRTC can also emit) is what a caller needs, guarded against a driver too old to JIT a newer NVRTC's PTX (target="cubin" needs no such guard — it is already SASS, the GPU's own machine code). Headers come from aether_dsc's sealed payload — a copy of the headers hawk's device compiler reads from directly, bundled into the wheel — never from disk.

Build and test from a source checkout (development)

Clone aether and eagle next to this checkout (so ../aether and ../eagle sit beside hawk/), then install aether's sealed header payload before hawk itself:

pip install ../aether/dsc
pip install -e .[test]
pytest -m "not gpu"

The compiled half (hawk._core) is a nanobind extension built through scikit-build-core; pip install -e . builds it via CMakeLists.txt, resolving the aether/eagle C++ header roots from those sibling checkouts (or $HAWK_AETHER_INCLUDE/$HAWK_EAGLE_INCLUDE). -m "not gpu" deselects the tests that need a CUDA GPU and cupy; most of the rest additionally need eagle installed as a Python package to execute traced kernels on host or device (see the "Python package" section of eagle's README).


hawk write it (this repo) · eagle run it · aether the numerics core · raptor the shared contracts — the RAPTOR family

Apache-2.0 (see LICENSE and NOTICE) · cite via "Cite this repository" (CITATION.cff; every tagged release is archived on Zenodo: doi:10.5281/zenodo.23250242) · built to make GPU computing accessible on modest hardware, for research and education. Collaboration is the point, and a citation is the currency — get in touch.

Metadata

Release files for raptor-hawk 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for raptor-hawk 0.4.0
File Size Uploaded
raptor_hawk-0.4.0.tar.gz 1.4 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for raptor-hawk 0.4.0
File
raptor_hawk-0.4.0-cp314-cp314t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.14 CPython 3.14 free-threading Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp314-cp314-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.14 CPython 3.14 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp313-cp313t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.13 CPython 3.13 free-threading Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.10 CPython 3.10 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details
raptor_hawk-0.4.0-cp39-cp39-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl CPython 3.9 CPython 3.9 Linux glibc 2.28+ x86-64, Linux glibc 2.26+ x86-64 Details

Total release size: 5.3 MB

Release files / raptor_hawk-0.4.0.tar.gz

Download URL raptor_hawk-0.4.0.tar.gz
Size 1.4 MB
Tags Source
SHA-256 checksum
How to use checksums
8f50d6b2824c3fa1bf58f2b1f240c06c83612d4b1f4ee7ea5f5940f307abb566
BLAKE2b-256 checksum
How to use checksums
10927e0dc481ec7c04a6c8c77a0f5ec19248dcf0fc93a883d4000622341b126b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp314-cp314t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp314-cp314t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 492.6 kB
Tags CPython 3.14 CPython 3.14 free-threading Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
e14f42ecc19e5249201ba87a476a56d9c7cbda66212cc305d3ca68bd386db965
BLAKE2b-256 checksum
How to use checksums
49456c44406252b87e727fff53ad25e1f241f99af1a445a42157e7ea28f67bed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp314-cp314-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp314-cp314-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 489.8 kB
Tags CPython 3.14 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
3b69031e0ba512569686e22fc41c86dbd4dbc048a580a162cbb4d1651fdd0654
BLAKE2b-256 checksum
How to use checksums
017ae239ec6f0ad89ca8d8041b727da54817ceb1f3a60594ba88cc776f3bc6a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp313-cp313t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp313-cp313t-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 492.4 kB
Tags CPython 3.13 CPython 3.13 free-threading Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
8b8c99129a875805a1bdf2318330c179ab469c365d32a07eee52a3e8c387789b
BLAKE2b-256 checksum
How to use checksums
adfd6cc7f0fef150e3ae078749dbeba954877d8414e42e1dbe186d2bc2383887
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp313-cp313-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 489.5 kB
Tags CPython 3.13 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
d594a459d9d0a228177dc71b41954e14d212677b82d90fd2e0fd733d5f49976c
BLAKE2b-256 checksum
How to use checksums
5ba7db708c6eeb0aa0db6b7d40543804544687004a008f1ac7b534345b355fa2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp312-cp312-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 489.6 kB
Tags CPython 3.12 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
99a70f8d3de13df23bcd5bbf60f842d9d47b48de9a6dd592f7dde04dadf084ec
BLAKE2b-256 checksum
How to use checksums
084221a9b88ea093e918ad5c6f9ad5c1402681d9a8e4b37b78d9db8e53933603
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp311-cp311-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 490.0 kB
Tags CPython 3.11 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
4643e5d993fd344d90bdbe060bb51f543d53e2eda7d65a191bc5d8eadeb1ea19
BLAKE2b-256 checksum
How to use checksums
f26ee9d3a4797855e4e7ca41cc14ce8654623ff93bb781a37addd417352c50bb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp310-cp310-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 490.4 kB
Tags CPython 3.10 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
7da93dbbd89584331ae0e92c0b5a62b52e1b88cf7ed69884073a7fdf56260ab8
BLAKE2b-256 checksum
How to use checksums
0ff1af32cb98e05248101d384efe60deb611a0a73bfc7156c3f91566e169d99d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release files / raptor_hawk-0.4.0-cp39-cp39-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl

Download URL raptor_hawk-0.4.0-cp39-cp39-manylinux_2_26_x86_64.manylinux_2_28_x86_64.whl
Size 483.9 kB
Tags CPython 3.9 Linux glibc 2.26+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
0bae2f0d5a46a0bcda759a61f6ac78c311d48878e75641d6b25cf1cd5ee17a9f
BLAKE2b-256 checksum
How to use checksums
db4de8fcbcaf0937afccf081fda034ca53e642a79eb12946c92ab8a2e3fab81c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.4.0 This release

9 release files

0.3.1

9 release files

0.3.0

5 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page