Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Foundry

Instantaneous CUDA graph restoration via execution context materialization.

Python arXiv License

Foundry is a system that persists CUDA graph states through template-based context materialization. It materializes both the structure and execution context of captured CUDA graphs, making graph restoration kernel-agnostic and eliminating the need for hand-crafted patching rules. By intercepting CUDA driver calls, Foundry enforces a deterministic memory layout and automatically detects and serializes the binaries of kernels used in the CUDA graphs.

With Foundry, LLM serving engines can directly reload CUDA states from disk and skip the warmup process to start in a few seconds.

Foundry in Action

Demonstration of serving Qwen/Qwen3-30B-A3B-FP8 with expert parallel size = 2.

Foundry vs. baseline vLLM cold start

Baseline vLLM (top) vs. vLLM + Foundry (bottom). Foundry reconstructs 256 graphs in 1 sec (plus ~2 sec sampler warmup + API server init), while original vLLM spends around 30 sec to warmup and capture graphs.

How it works

Foundry intercepts three classes of CUDA driver calls and routes each through its own piece of in-process state, so SAVE can serialize that state to disk and LOAD can rebuild an identical driver-side state from the archive.

flowchart LR
    subgraph APP["Application calls"]
        direction TB
        A1["cudaMemXXX"]
        A2["cudaModuleLoadXXX<br/>cudaLibraryLoadXXX"]
        A3["cudaGraph capture<br/>cudaGraphLaunch"]
    end

    subgraph FDR["Foundry: indirection &amp; interception"]
        direction TB
        F1["<b>VMM region</b><br/>monotonic cursor →<br/>byte-deterministic offset"]
        F2["<b>Module registry</b><br/>fatbin bytes,<br/>entry_name → CUfunction"]
        F3["<b>Captured graphs</b><br/>each topology group:<br/>1 template + N on-demand"]
    end

    subgraph DRV["CUDA driver / context"]
        direction TB
        D1["cuMemCreate / cuMemMap<br/>cuMemAddressReserve"]
        D2["cuLibraryLoadData<br/>cuLibraryGetKernel"]
        D3["CUgraph<br/>CUgraphExec"]
    end

    A1 --> F1 --> D1
    A2 --> F2 --> D2
    A3 --> F3 --> D3

    classDef app fill:#dfe8ff,stroke:#4060c0,color:#000
    classDef fdr fill:#ffe8d6,stroke:#c08040,color:#000
    classDef drv fill:#d8f0d8,stroke:#40a040,color:#000
    class APP app
    class FDR fdr
    class DRV drv
  • Memory. Every device allocation is funneled into a single VMM region [base_addr, base_addr + region_size). A monotonic cursor gives every tensor a byte-deterministic offset; the underlying physical mapping is done with cuMemCreate + cuMemMap.
  • Modules / libraries. As device code is loaded, the fatbin bytes and the entry_name → CUfunction table are recorded.
  • Captured graphs. Each captured CUgraph is serialized and then grouped with other graphs that share the same topology (same kernels, dependency DAG and cluster dimensions; they differ only in kernel parameters). One graph per group is kept as the template; the rest are stored as on-demand members that carry only their per-node kernel parameters.

SAVE writes all three pieces to an archive. LOAD pre-maps the same VMM range, re-loads the same modules from the packed fatbins, builds each template's CUgraph node by node once, and then produces every member by rewriting the template's node parameters, either into a dedicated CUgraphExec per member (default, eager or lazy) or into one shared exec per template switched with cuGraphExecUpdate. Either way restored graphs replay at native speed; kernel handles embedded in the captured graphs resolve to the same device addresses they had at SAVE time. See docs/graph-templates.md for the design and measurements.

Inference-Engine Integrations

Foundry ships engine integrations under foundry/python/foundry/integration/. Per-engine setup instructions live under recipe/.

Engine Integration code Documentation Setup Instructions
SGLang integration/sglang/ docs/sglang/overview.md recipe/sglang/README.md
vLLM integration/vllm/ docs/vllm/overview.md recipe/vllm/README.md
TensorRT-LLM integration/trtllm/ docs/trtllm/overview.md recipe/trtllm/README.md

Status

Engine Single GPU DP TP EP
SGLang ✅ ✅ ✅ ✅
vLLM ✅ ✅ 🚧 ✅
TensorRT-LLM 🚧 🚧 🚧 🚧

✅ validated end-to-end (SAVE → LOAD → query)  ·  🚧 not yet

SGLang TP uses torch symmetric-memory allreduce inside the decode graphs (TP=2/4); SGLang EP covers DeepEP low-latency (NVSHMEM) and DeepEP v2 (NCCL symmetric windows). Every SGLang configuration is validated with the full decode-graph set (batch sizes 1..256) for restore time, per-token latency and greedy-output equality against unmodified SGLang; see recipe/sglang/README.md.

The adapted SGLang fork is published at foundry-org/sglang, branch foundry (current head 6272eb04c5 = upstream main 03ea13a545 + one integration commit; the v0.0.3 pairing f1d688e52 is kept as foundry-0.0.3, the 0.0.2-era integration on foundry-0.0.2). The vLLM and TensorRT-LLM forks will follow at foundry-org/vllm and foundry-org/TensorRT-LLM.

Performance

SGLang on 8×H100 (2 GPUs per run), all 256 decode graphs captured or restored, prefill graphs off on both sides, median TPOT of restored vs unmodified SGLang:

Config Model Capture Restore Init to /health: SGLang → foundry LOAD TPOT delta (bs 1 / 8 / 32 / 128)
TP=2 (symm-mem) Qwen3-32B 26.7 s 3.5 s 63 s → 41 s -0.4 / -0.3 / +0.0 / -0.7 %
EP=2 (DeepEP LL) Qwen3-30B-A3B 38.0 s 2.5 s 75 s → 45 s -0.0 / -0.0 / +0.1 / +0.6 %
EP=2 (DeepEP v2, NCCL) Qwen3-30B-A3B-FP8 56 s 2.1 s 93 s → 45 s +0.0 / +0.0 / +0.2 / +0.2 %

Restored graphs keep their programmatic-dependent-launch edges, so per-token latency matches the captured graphs within run-to-run noise (docs/pdl-edge-batching.md). Greedy completions are identical to unmodified SGLang for dense TP and within SGLang's own run-to-run nondeterminism for MoE.

Roadmap

See ROADMAP.md for the full development plan and progress.

Requirements

  • Linux x86_64, NVIDIA driver with CUDA 12.0+ (the driver's libcuda.so.1 is loaded at runtime)
  • PyTorch: foundry.ops is a torch C++ extension bound to the torch it was built with (major.minor and CUDA major), like sglang-kernel or flashinfer. Importing it under another torch raises a readable error.

Installation

Prebuilt wheels are published to PyPI as foundry-core (the import name is foundry). One torch/CUDA pairing per release line:

foundry-core torch CUDA CPython Platform
0.1.x 2.13 (torch==2.13.*, cu130 build) 13.0 3.10-3.13 manylinux_2_28 x86_64
pip install "torch==2.13.0" --index-url https://download.pytorch.org/whl/cu130
pip install "foundry-core>=0.1.0rc1,<0.2"
python -c "import foundry; print(foundry.__version__)"

The wheel ships foundry/ops.*.so and foundry/libcuda_hook.so (the hook SGLang/vLLM preload) and registers the SGLang plugin entry point. No Boost or other C++ runtime dependency is needed. Wheels for other torch/CUDA pairs, when built, are attached to the GitHub Release with a local version (0.1.0+cu128.torch2.12) and install by URL.

From source

Needed for any other torch, or for development. Requirements:

  • CMake 4.0+ and ninja (pip install "cmake>=4.0" ninja if the system ones are older)
  • the torch you will run with, already installed (Foundry compiles against it; rebuild after changing torch)
  • CUDA Toolkit with nvcc (CUDA 12+; CUDA 13 with torch cu130)
  • Boost headers, header-only (nothing is linked): the vendored copy in third_party/boost (populated by tools/release/vendor_boost.sh), or a system Boost >= 1.83 (Ubuntu 24.04: apt-get install libboost-dev; conda: conda install -c conda-forge boost-cpp; or point FOUNDRY_BOOST_INCLUDE_DIR at the directory that contains boost/version.hpp)
pip install "cmake>=4.0" ninja
# Torch 2.13 with CUDA 13.0
pip install torch==2.13.0 --index-url https://download.pytorch.org/whl/cu130
pip install -e . --no-build-isolation

--no-build-isolation matters: an isolated build compiles against whatever torch pip resolves, not the one in your environment. From PyPI's sdist into an existing environment: pip install --no-build-isolation --no-binary foundry-core foundry-core.

Release process (vendoring Boost, tagging, trusted publishing): docs/release.md.

Debugging

Verbose C++ logging (allocator hook and graph replay) is gated behind a single compile-time flag, FOUNDRY_DEBUG. Uncomment the // #define FOUNDRY_DEBUG line at the top of csrc/hook.cpp or csrc/CUDAGraph.cpp (or build with -DFOUNDRY_DEBUG) and reinstall. With the flag on, every LOAD also logs a per-graph check that the restored edge records (PDL programmatic edges) survived insertion; see docs/pdl-edge-batching.md.

Quick Start

Graph Capture and Save

Foundry requires LD_PRELOAD to intercept CUDA driver calls. The graph capture and save must run in a subprocess with the hook library preloaded.

import foundry as fdry
import torch

torch.cuda.init()
device = torch.device('cuda:0')
torch.set_default_device(device)

# Set up VMM allocation region for deterministic memory addresses
BASE_ADDR = 0x400000000000
region_size = fdry.parse_size('1GB')
fdry.set_allocation_region(BASE_ADDR, region_size)

# Allocate input tensors
input_a = torch.full((100, 100), 2.0, device=device)
input_b = torch.full((100, 100), 3.0, device=device)

# Warm up the model
model = MyModel()
model(input_a, input_b)
torch.cuda.synchronize()

# Capture CUDA graph
graph = fdry.CUDAGraph()
with fdry.graph(graph):
    result = model(input_a, input_b)

# Replay and verify
graph.replay()
torch.cuda.synchronize()

# Save graph with output tensors
# Produces BOTH graph.json and graph.cugraph (optimized binary format)
graph.save('graph.json', output_tensors=result)

fdry.stop_allocation_region()

Graph Load and Replay

Loading a saved graph also requires LD_PRELOAD and must use the same allocation region base address.

import foundry as fdry
import torch

torch.cuda.init()
device = torch.device('cuda:0')
torch.set_default_device(device)

# Load CUDA modules and libraries from archive
fdry.load_cuda_modules_and_libraries('hook_archive')

# Set up the same allocation region as capture
BASE_ADDR = 0x400000000000
region_size = fdry.parse_size('1GB')
fdry.set_allocation_region(BASE_ADDR, region_size)

# Allocate input tensors (can have different values)
input_a = torch.full((100, 100), 5.0, device=device)
input_b = torch.full((100, 100), 3.0, device=device)

# Load and replay the graph (NOTE: auto-loads .cugraph binary when available)
graph, output_tensor = fdry.CUDAGraph.load('graph.json')
graph.replay()
torch.cuda.synchronize()

# output_tensor now contains the result
fdry.stop_allocation_region()

Async Graph Loading

Load graphs asynchronously with background template building. Graphs with the same topology share a single CUgraphExec template — only node parameters are updated before each launch (on-demand replay). Two finish APIs are available:

  • finish_graph_loads(pending) — bulk, waits for all templates then returns the full list. Simplest call shape.
  • finish_one_graph_load(pending, index) — per-graph; first call waits on background build completion, later calls just finalize. Use this when you want to interleave finalization with other work (the foundry vLLM integration uses it to walk the same VMM cursor trajectory on LOAD that SAVE recorded).
import foundry as fdry

# Phase 1: parse .cugraph binaries + build topology groups + templates in background
pending = fdry.CUDAGraph.start_graph_builds(
    ["graph_0.json", "graph_1.json", ...], num_threads=24
)

# Background threads now race against whatever the caller does next.
# In practice we found that overlapping with model weight loading is
# net-negative (driver contention), so the vLLM integration kicks off
# start_graph_builds *after* weight load and lets it overlap with the
# cheaper post-load init phases instead.

# Bulk finish:
results = fdry.CUDAGraph.finish_graph_loads(pending)
for graph, output in results:
    graph.replay()

# OR per-graph finish (interleavable):
# for i in range(num_graphs):
#     graph, output = fdry.CUDAGraph.finish_one_graph_load(pending, i)
#     graph.replay()

Graph Manifest and Topology Groups

After capturing all graphs, call save_graph_manifest() to group graphs by topology and assign templates. On-demand (non-template) graphs strip dependencies to reduce file size.

import foundry as fdry

# After all graphs are captured and saved
fdry.save_graph_manifest('hook_archive')

Memory Preallocation for Fast Graph Reload

The preallocation API physically allocates memory upfront, enabling subsequent allocations to use a fast path (pointer bump only, no VMM driver calls).

import foundry as fdry

# With preallocation - allocations within 8GB use fast path
with fdry.allocation_region(0x500000000000, '16GB', prealloc_size='8GB'):
    graph, outputs = fdry.CUDAGraph.load('model.json')
    graph.replay()
Function Description
set_allocation_region(base, size) Set VMM allocation region for deterministic memory addresses
stop_allocation_region() Stop the allocation region
resume_allocation_region() Re-enable a previously stopped allocation region
allocation_region(base, size, prealloc_size=None) Context manager to set up VMM allocation region with optional preallocation
preallocate_region(size) Manually preallocate memory inside an allocation region
free_preallocated_region() Free manually preallocated memory
get_current_alloc_offset() / set_current_alloc_offset(offset) Read or fast-forward the in-region cursor
parse_size(size) Parse a size string ("1GB", "16MB", …) to bytes
load_cuda_modules_and_libraries(archive_dir) Load CUDA modules and libraries for graph loading
save_graph_manifest(archive_dir) Write graph_manifest.json with topology groups and template assignments
CUDAGraph.start_graph_builds(paths, num_threads) Kick off background template build for a list of saved graphs
CUDAGraph.finish_graph_loads(pending) Wait for all builds and return [(graph, output), ...]
CUDAGraph.finish_one_graph_load(pending, i) Finalize one graph by index; interleavable with other work
init_nvshmem_for_loaded_modules() After prepare_communication_buffer_for_model has bootstrapped NVSHMEM, finalize nvshmemx_cumodule_init for each module queued by load_cuda_modules_and_libraries

Testing

Run the test suite:

pytest tests/

Setting up clangd

conda install -c conda-forge libstdcxx-ng libgcc-ng
conda install -c conda-forge bear
bear -- python setup.py build_ext --inplace

Contributors

  • Xueshen Liu
  • Yongji Wu

Metadata

Release files for foundry-core 0.1.0rc1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for foundry-core 0.1.0rc1
File Size Uploaded
foundry_core-0.1.0rc1.tar.gz 1.3 MB Details

Built distributions (wheels)

Table of built distributions (wheels) for foundry-core 0.1.0rc1
File
foundry_core-0.1.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl CPython 3.13 CPython 3.13 Linux glibc 2.27+ x86-64, Linux glibc 2.28+ x86-64 Details
foundry_core-0.1.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl CPython 3.12 CPython 3.12 Linux glibc 2.27+ x86-64, Linux glibc 2.28+ x86-64 Details
foundry_core-0.1.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl CPython 3.11 CPython 3.11 Linux glibc 2.27+ x86-64, Linux glibc 2.28+ x86-64 Details
foundry_core-0.1.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl CPython 3.10 CPython 3.10 Linux glibc 2.27+ x86-64, Linux glibc 2.28+ x86-64 Details

Total release size: 41.3 MB

Release files / foundry_core-0.1.0rc1.tar.gz

Download URL foundry_core-0.1.0rc1.tar.gz
Size 1.3 MB
Tags Source
SHA-256 checksum
How to use checksums
ec9a044dbfc74cfb98940ffe27afb7c1264527bbdc84bc14e001c055c05311b7
BLAKE2b-256 checksum
How to use checksums
39fe05abf181fa2f1b49e8d832c49c2d34ca871f9d494c3a5e9483d21a6f85b2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / foundry_core-0.1.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Download URL foundry_core-0.1.0rc1-cp313-cp313-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Size 10.0 MB
Tags CPython 3.13 Linux glibc 2.27+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
f4b3c3dcabddfe4f3046e10e609c9a25375e4d799f646f921dd8fba1dee2c007
BLAKE2b-256 checksum
How to use checksums
050eb4041e67453a5b6ca5e177c1003958ff114e7183ab4ba4b3cc198ffb0699
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / foundry_core-0.1.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Download URL foundry_core-0.1.0rc1-cp312-cp312-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Size 10.0 MB
Tags CPython 3.12 Linux glibc 2.27+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
e7ad8927a2295d94ca9ed5fb41cc5f1a47ac899993de0d87cf3b74f08d00db45
BLAKE2b-256 checksum
How to use checksums
8fbad0190d245c5a1d2bfb50e5a4a84234ea285acb3f3e875a962a584b7e07c4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / foundry_core-0.1.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Download URL foundry_core-0.1.0rc1-cp311-cp311-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Size 10.0 MB
Tags CPython 3.11 Linux glibc 2.27+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
d5f81cd49293e2517fe3f2dc19744e18ba5a0b9961e1f509b20299fb3ee10576
BLAKE2b-256 checksum
How to use checksums
3f351975680c4d49e174c3bad3334a7cb98ceee22936e2787b251b82d914ea39
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release files / foundry_core-0.1.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl

Download URL foundry_core-0.1.0rc1-cp310-cp310-manylinux_2_27_x86_64.manylinux_2_28_x86_64.whl
Size 10.0 MB
Tags CPython 3.10 Linux glibc 2.27+ x86-64 Linux glibc 2.28+ x86-64
SHA-256 checksum
How to use checksums
983ce330340e222b5c9608c93c5b8f561ddd7bd8748f82e05f60a3abb3710a57
BLAKE2b-256 checksum
How to use checksums
2cbf5ebf64d16a55974cf79ac38668fb21e6dfc62fc94b897ac17f3a794b720f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0rc1 This release

5 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page