Skip to main content

raptor

raptor: one manifest, any engine

A compiled GPU kernel ships with a small JSON file that says what it takes, what it returns and where it can run. raptor defines that file, so one package can build a kernel and another can run it.

Runs on: CPU, no GPU needed · You need: Python 3.9+, nothing else to install raptor itself

CI Docs License

The RAPTOR family: hawk (write it), eagle (run it), aether (the C++/CUDA numerics underneath) and raptor (the shared contract); you are looking at raptor.

aether · hawk · eagle · raptor · the family

raptor holds the shared contracts of the RAPTOR family: the manifest format and certification contracts that let a kernel built once work with the code you already have. See where RAPTOR fits for how the four pieces combine.

0 4 1
hard dependencies (pure Python) array libraries certified: NumPy, CuPy, PyTorch, Warp JSON manifest per kernel bundle (a kernel's compiled artifact(s) plus its manifest)

See it working: a batch of orbits, one kernel, run on CPU threads — Demo below has the commands and what each printed line means.

Docs · the family site

For package authors

from raptor.schema import validate_manifest

manifest = {
    "schema_version": 2,
    "pattern": "pure",
    "aether_abi": "aether-abi/2",
    "exec_targets": ["host", "device"],
    "exec_access": "sample_local",
    "plugins": [{"id": "two_body_step", "order": 0, "enabled": True, "artifact": "two_body_step.ptx",
                "sidecar": "two_body_step.json", "format": "ptx"}],
}
validate_manifest(manifest)  # raises on a malformed or out-of-order document
key meaning
schema_version which rules this document follows
pattern which kind of kernel this is; pure = no side effects
aether_abi the binary-format tag paired with schema_version (e.g. "aether-abi/2" for v2)
exec_targets host / device: where this kernel may run
exec_access how it touches per-sample vs. cross-sample data
plugins the compiled artifact(s) this manifest describes (ptx = NVIDIA's GPU assembly format)

Any class with a matching build(directory) method satisfies ManifestProvider — "duck typing": judged by which methods exist, checked with isinstance, never by which class it inherits from:

from raptor.protocols import ManifestProvider

class Bundle:
    def build(self, directory):
        return f"{directory}/manifest.json"

isinstance(Bundle(), ManifestProvider)  # True — duck-typed, no inheritance needed

Install

pip install raptor-core      # Python 3.9+, no dependencies

The import name is raptor. The rest of the family installs the same way: pip install raptor-hawk for the CPU route (Linux x86_64, CPython 3.10-3.13, a host g++ 11 or newer; it brings aether-dsc, the sealed C++ headers it compiles against, automatically), and for the GPU route pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]" (use [cuda13] on both for CUDA 13; no nvcc needed).

NVIDIA packages come only through the extras. raptor-hawk[cuda12] pulls cuda-bindings 12, nvidia-cuda-nvrtc-cu12 and nvidia-cuda-cccl-cu12; raptor-hawk[cuda13] pulls cuda-bindings 13, nvidia-cuda-nvrtc 13 and nvidia-cuda-cccl 13; raptor-eagle[cuda12] / [cuda13] pull CuPy (cupy-cuda12x / cupy-cuda13x, with the CUDA headers CuPy compiles against). Without an extra pip installs no NVIDIA package: you get the CPU route, or the GPU route through a CUDA setup you already have. Pick the extra matching the CUDA version your driver reports (nvidia-smi, top right).

Platforms: built and tested on Linux x86_64 only so far (CPython 3.10–3.13), on NVIDIA GPUs from Pascal (Quadro P2000) and Turing (Tesla T4). There are no wheels for macOS, Windows or ARM yet, and WSL2 is untested. raptor-core and aether-dsc are pure Python and install anywhere.

To work on the code, from a clone: git clone https://github.com/amasat01/raptor.git, then pip install -e .. Test with pip install pytest && pytest tests -m "not cross_repo" (-m "not cross_repo" leaves out the tests that check eagle against raptor's contracts — they need an eagle checkout next to this one, and fail rather than skip without it). To set up a development machine for the whole family, see tools/devbox/.

Demo

examples/two_body_hybrid.py is a small end-to-end example: a batch of orbits propagated by a hand-written physics formula plus a small learned correction, traced into one compiled kernel that this package's manifest describes. See Going deeper below for exactly what it computes and prints.

# hawk and eagle come from PyPI (no nvcc needed); [demo] adds numpy and torch
pip install "raptor-hawk[cuda12]" "raptor-eagle[cuda12]"   # or [cuda13] on both
pip install -e .[demo]                                     # from a clone of this repository
python examples/two_body_hybrid.py --host
python examples/two_body_hybrid.py --device --steps 400 --batch 65536

--host prints the following (output of python examples/two_body_hybrid.py --host; the wall-per-step line varies by machine, everything above it is exact, the same seed every run):

target            : host (manifest exec_targets ['host'])
batch x steps     : 1024 x 200 at dt = 0.01
training loss     : 1.230e-06
(a) max |compiled - torch eager| : 1.788e-06 (tolerance 1e-04)
(b) final-state error vs truth   : corrected 6.491e-03 vs two-body only 1.930e-01
(c) wall per step (demo measurement, not a benchmark): compiled 1.306 ms, torch eager 0.612 ms

Going deeper (optional)

What the demo computes, package internals, and three scripts' own machine-local measurements — none of it needed to install or use raptor.

What's inside
  • raptor.schema — the kernel-manifest schema, two layers: schema.manifest (the core layer every producer writes against) and schema.blocks (the neural_block extension; hawk never imports it). schema.dtypes carries the scalar-type tag vocabulary. raptor.schema.validate_manifest(doc) is a minimal shape validator (key order, top keys, schema_version, pattern vocabulary) — not a port of eagle's deep validate_sidecar. Schema versions 1 and 2 both load (manifest.MAX_SCHEMA_VERSION): v2 adds three flat, versioned top-level keys — exec_targets (device/host list), exec_access (sample_local / cross_sample_read / cross_sample_write / mapreduce), exec_op (required iff exec_access="mapreduce") — and requires aether_abi: "aether-abi/2"; v1 rejects any exec_* key outright (manifest.check_execution_axis).
  • raptor.protocols — the typing.Protocol definitions: Backend / Executable (program-level execution), KernelLauncher / Marshal (kernel-level launch/marshal, mirroring eagle's public surface 1:1), ManifestProvider (the kernel-provider protocol's one facet).
  • raptor.conformance — the interop certification declaration (UNIVERSE / CERTIFIED / ROADMAP / ROWS) and its reusable matrix-hygiene scanners; the full row-by-row certification matrix is on the interop protocols page.
  • goldens/ — the golden conformance corpus (harvested from eagle's test fixtures) that raptor.schema.validate_manifest is exercised against.
  • tests/*cross_repo*: the eagle conformance battery (test_verbatim_against_eagle.py, test_goldens.py), marked cross_repo and failing rather than skipping when a sibling checkout is absent — deselected here only via -m "not cross_repo", run for real by the local gate and a downstream package's own CI. raptor.schema.manifest.SCHEMA_VERSION is the version of record; eagle.roles.SCHEMA_VERSION derives from it directly, so no standalone text-scan cross-check is kept as a redundant leg.
What the demo computes

A batch of planar orbits is propagated with RK4 (a standard 4th-order numerical-integration formula). The textbook two-body formula is written by hand; the piece it misses — drag — is learned by a small (4 → 32 → 2) neural network trained in torch on CPU. hawk traces ONE kernel that performs a whole RK4 step with the network evaluated inline, compiles it, and publishes it with a manifest this package's schema validates; eagle loads that manifest and runs the kernel, either on CPU threads or as a replayed CUDA graph (a GPU's sequence of launches recorded once and replayed without re-issuing each one). The script prints how closely the compiled rollout tracks a plain torch rollout of the same arithmetic, how much closer to the truth the corrected orbits land than uncorrected ones, and a wall-per-step line.

early termination: samples that finish at different times (measured when the script runs — numbers vary by machine)

examples/early_termination.py shows the other half of this one-kernel idea: irregular work whose samples finish at different times, with the host out of the loop. A batch of planar orbits ends when it crosses a collision or an escape radius, and a device-side live count guards each step with a conditional graph node (a graph step that can skip itself at replay time), so the whole horizon is queued as graph replays with zero host round-trips until every sample is done — later replays then do no work at all. Measured against the identical captured graph with the guard removed (so the guard is the only difference), the skip is a real wall-clock saving that scales with how much of the horizon runs past the last termination — this script's own measurement, on the machine that built this page: over 13x at a million samples over 4000 steps (13.8x in the committed card). The script also runs the same kernel eagerly, host-side guard check and all: at a million samples the guarded graph and this host-guarded eager loop tie (1.01x in the card; the kernel itself dominates there); at the smallest batches (1, 32 samples) the skip does not pay for itself — the conditional node's own fixed cost exceeds a tiny kernel's, so the unguarded graph is the fastest arm. The eager loop also carries one host round-trip per two-step replay against the graph's one for the whole captured horizon, a gap the numbers above already price in. CPU mode (--host) runs the same kernel with the guard evaluated eagerly, so it needs no GPU. --torch (device only) times the same RK4 two-body physics written in idiomatic torch — eager, eager with a host check, and both as a replayed CUDA graph — against the guarded graph and prints each arm's ratio. Run it yourself for your own machine's numbers.

python examples/early_termination.py --host
python examples/early_termination.py --device
python examples/early_termination.py --device --torch
zero-copy interop (measured when the script runs — numbers vary by machine)

examples/zero_copy_interop.py backs the third leg of the same idea: exchanging GPU arrays with NumPy, CuPy and PyTorch without copies, and only where the family's own certification matrix says there is none. A CUDA-resident torch tensor is handed straight into one hawk kernel through eagle's zero-copy to_cupy view, and the kernel writes its result into a second torch CUDA buffer through the same view — no allocation, no copy, matching the aliasing rows of raptor's interop matrix exactly. The same tensor read from CPU memory instead takes the honest upload the matrix declares as a copy. The script prints each crossing's measured pointer identity next to the row that covers it, and a behavioural proof for the input alias: mutating the torch tensor in place and replaying the same kernel with no rebind shows the output tracking the change. --device also times that alias path against the copying round-trip a caller without this protocol would naturally write — this script's own measurement, on the machine that built this page: at 256 MB, the copies cost over 100x the alias path (113.6x in the committed card). CPU mode (--host) needs neither a GPU nor torch: it runs the same kernel on CPU threads with plain numpy in and out.

python examples/zero_copy_interop.py --host
python examples/zero_copy_interop.py --device
autodiff vs. torch (measured when the script runs — numbers vary by machine)

examples/autodiff_vs_torch.py closes the loop on the same one-kernel idea: hawk differentiates the traced kernel itself rather than executing a separate tape. One RK4 two-body step is traced once, and hawk.diff.vjp/ jvp derive a reverse-mode gradient kernel (vjp: given a cotangent — a weighting of the output — produce the matching weighting of the input) and a forward-mode directional-derivative kernel (jvp: given a tangent — a direction in the input — produce the matching direction in the output) from it, published in the same manifest as the primal (the original, undifferentiated kernel). --device checks both against torch.autograd.grad and torch.func.jvp on the identical physics written in torch, and against a float64 finite-difference Jacobian (the matrix of every output's derivative with respect to every input, approximated numerically) so the check does not rest on torch alone; every plane (one array per sample) in and out — the state, the cotangent, the tangent, both derivative outputs — is a torch CUDA tensor crossing through eagle's zero-copy to_cupy view, read back from torch's own storage. It also times hawk's vjp/jvp kernels against torch's, and each side's primal alone so the derivative's own overhead is visible — this script's own sweep, on the machine that built this page (benchmarks/autodiff_vs_torch_card.md): across a batch sweep from 1 to a million samples, hawk's vjp measured roughly 90-550x faster per call than torch.autograd.grad, and hawk's jvp roughly 60-780x faster than torch.func.jvp. At small batch those ratios are mostly torch's own per-call overhead; at a million samples (about 90x and 60x) they are memory traffic — one fused kernel against every intermediate written out. hawk's derivatives cost about 1.0-1.5x their own primal, where torch's forward-plus-backward pass costs several times its primal alone. CPU mode (--host) needs no GPU or cupy; its reference is the float64 finite difference alone, with torch (if installed) added as an optional CPU cross-check.

python examples/autodiff_vs_torch.py --host
python examples/autodiff_vs_torch.py --device
python examples/autodiff_vs_torch.py --sweep

hawk write it · eagle run it · aether the numerics core · raptor the shared contracts (this repo) — the RAPTOR family

Apache-2.0 (see LICENSE and NOTICE) · cite via "Cite this repository" (CITATION.cff; each tagged release is archived on Zenodo with its own DOI once the first one exists) · built to make GPU computing accessible on modest hardware, for research and education. Collaboration is the point, and a citation is the currency — get in touch.

Metadata

Release files for raptor-core 0.2.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for raptor-core 0.2.0
File Size Uploaded
raptor_core-0.2.0.tar.gz 62.7 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for raptor-core 0.2.0
File Interpreter ABI Platform
raptor_core-0.2.0-py3-none-any.whl Python 3 none any Details

Total release size: 104.3 kB

Release files / raptor_core-0.2.0.tar.gz

Download URL raptor_core-0.2.0.tar.gz
Size 62.7 kB
Tags Source
SHA-256 checksum
How to use checksums
9cff4e9a54715394fedd73dde2ab6b971ddf4ab3c9bae213466138252882ef9e
BLAKE2b-256 checksum
How to use checksums
a20c1feea49516a0f54e98ccfa689ea59e5f9671cfd3a9ca0274b53de1a08e61
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release files / raptor_core-0.2.0-py3-none-any.whl

Download URL raptor_core-0.2.0-py3-none-any.whl
Size 41.6 kB
Tags Python 3
SHA-256 checksum
How to use checksums
9468509b2847fa75adefe23a921a701d11cf478a3eba789aa1c536a695dfcebe
BLAKE2b-256 checksum
How to use checksums
7e9b5cf61b689d558066a5b748252b5619f6f387b6942e972965258905a8aff4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 8, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

This release

0.2.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page