Skip to main content

TensorTorrent logo

TensorTorrent

A heterogeneous PyTorch compiler and runtime for one machine with many CPUs, GPUs, and memory tiers.

CI status Latest version tag Python 3.10 to 3.13 Rust 1.85 or newer Apache-2.0 license

TensorTorrent exports a PyTorch model, partitions its graph, places regions across available compute, and runs the resulting schedule through a Rust data plane. Parameters can stream from slower storage and activations can spill when the model exceeds device or host memory.

Python compiles. Rust schedules. One immutable ExecutableArtifact describes the program.

[!IMPORTANT] Supported target: Linux with Python 3.10–3.13 and PyTorch 2.4 or newer. Validate every deployment machine with tensortorrent validate-hardware before serving production traffic.

Installation

pip install torch --index-url https://download.pytorch.org/whl/cpu   # or CUDA/ROCm from pytorch.org
pip install tensortorrent

Linux, Python 3.10–3.13, PyTorch ≥2.4. Install torch first if you need a specific build. Wheels: PyPI, Releases. Dev from source: uv + Rust 1.85+ (Quick start).

Quick start

git clone https://github.com/alhussein-jamil/TensorTorrent.git
cd TensorTorrent
make sync
make doctor

Compile a module and compare it with eager PyTorch:

import torch
import torch.nn as nn
import tensortorrent as tt  # import alias: tt

model = nn.Sequential(
    nn.Linear(256, 256),
    nn.ReLU(),
    nn.Linear(256, 10),
).eval()
x = torch.randn(32, 256)

compiled = tt.compile(model, example_inputs=(x,))
torch.testing.assert_close(compiled(x), model(x), check_device=False)

compiled.save("artifact/")
reloaded = tt.load_compiled("artifact/")

Run uv run python examples/public_api_demo.py for hardware discovery, compile, and schedule output in one executable example.

What it handles

Area Implementation
PyTorch export and graph partitioning python/tensortorrent/compile
CPU, CUDA, ROCm, Intel XPU, and plugin discovery python/tensortorrent/backends
Resource budget resolver (host memory, VRAM, CPU, disk) python/tensortorrent/hardware/budget.py
NUMA-aware host allocation and CPU budget enforcement crates/tt-backend-cpu
Scheduling, residency, transfer, stall watchdog, and cancellation crates/tt-runtime
Parameter streaming and activation spill crates/tt-storage
Atomic, checksummed artifact bundles python/tensortorrent/artifact_io.py
Concurrent request serving (HTTP, auth, metrics) python/tensortorrent/serve
Virtual accelerators for deterministic tests crates/tt-backend-virtual

The runtime supports NCCL, RCCL, oneCCL, Gloo, and explicit host-staged collective fallbacks where the installed hardware and libraries allow them.

Architecture

flowchart LR
    M[PyTorch module] --> E[Export and normalize]
    E --> P[Partition and place]
    P --> A[ExecutableArtifact]
    A --> R[Rust dispatcher]
    R --> C[CPU / GPU regions]
    R --> S[Memory / storage tiers]

The Python control plane owns export, normalization, partitioning, region compilation, public APIs, and diagnostics. The Rust data plane owns the artifact, schedule, workers, residency, transfers, storage, cancellation, and telemetry. Torch compute regions may call back into Python; scheduling and data movement remain in Rust.

See the architecture guide for ownership boundaries and backend contracts for extension points.

Module composition

Compile a sequence as one graph to avoid opaque transfers between separately compiled artifacts:

compiled = tt.compile_modules(
    [encoder, projector, decoder],
    example_inputs=(x,),
    names=["encoder", "projector", "decoder"],
)

For branches, joins, structured arguments, or nested outputs, build a ModuleGraph from ModuleNode, GraphInput, and NodeOutput. Invalid names, forward references, and output paths are rejected before export.

Opt-in training

Compilation is inference-only by default. Set allow_training=True to use the same heterogeneous schedule with autograd:

config = tt.CompileConfig(allow_training=True)
compiled = tt.compile(model, example_inputs=(x,), config=config)

optimizer = torch.optim.Adam(compiled.parameters())
compiled.train()
optimizer.zero_grad()
loss = compiled(x).sum()
loss.backward()
optimizer.step()
compiled.eval()

Training cannot currently be combined with NVMe parameter streaming, activation spill budgets, or process workers. See the full product scope for intentional limits.

Does it actually work?

On a single device TensorTorrent reaches eager parity at scale — matching or beating PyTorch on large MLPs and transformers. Eligible resident single-region graphs use the direct path by default. Measured resident CPU+accelerator branch plans can use the same low-overhead path after synchronized timing beats both schedule execution and full fusion (prefer_direct_path; override with TT_DIRECT_PATH=0/1). The product focus beyond that is multi-device placement, parameter streaming, and activation spill.

Measured tables and the same-device harness pin live in Benchmarks.

Resource budgets and guardrails

Every memory limit, CPU count, and disk quota flows through a single resolver that reads cgroup v2/v1 limits, live OS availability, and explicit config values — in that precedence order. Containers automatically see their cgroup limits, not host totals. The resolver provenance is shown by tensortorrent doctor.

See Resource budgets and guardrails for the full precedence chain, spill lifecycle, stall watchdog, and worked examples.

Development

make sync                 # create the environment and build the native extension
make check                # lint, types, Rust tests, Python tests, doctor
make audit                # cargo-audit (Rust) + pip-audit (Python)
make coverage             # run tests with coverage gate (Python 3.12)
make native-gate          # native extension smoke and execution checks
make hardware-test        # explicit: may consume most available VRAM or spill space

On a machine with a GPU, run everything that needs real hardware in one go:

bash tools/run_everything.sh     # tests + hardware suite + all benchmarks

It writes logs, JSON, and a SUMMARY.md to bench-results/<timestamp>/. Install the benchmark baselines first with uv sync --extra bench so the ONNX Runtime and Accelerate comparisons run instead of reporting as missing.

CI: PRs and pushes to main (Python 3.10 + 3.13, x86-64/ARM64). Hardware tests are opt-in. See CONTRIBUTING.md.

Repository map

python/tensortorrent/   Python control plane, public API, and serving
crates/tt-*/            Rust IR, runtime, memory, storage, backends, and FFI
tests/                  Unit, integration, end-to-end, property, and hardware tests
docs/                   Product, architecture, deployment, and reference guides
examples/               Small public API programs
bench/                  Runtime and planner comparisons
tools/                  Local quality and native-extension gates
deploy/                 Docker Compose and Kubernetes examples
Dockerfile              CPU-only production container
Dockerfile.cuda         CUDA GPU production container (validate on GPU host before use)

Documentation

Versions and releases

Versions follow Semantic Versioning; tags are vMAJOR.MINOR.PATCH. A tag builds wheels, a GitHub Release, and a PyPI publish — see docs/RELEASING.md.

License

Apache-2.0. See LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

tensortorrent-0.2.6.tar.gz (368.5 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.4 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.17+ x86-64

tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (1.3 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.17+ ARM64

tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.4 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.17+ x86-64

tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl (1.3 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.17+ ARM64

tensortorrent-0.2.6-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.4 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.17+ x86-64

tensortorrent-0.2.6-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl (1.4 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.17+ x86-64

File details

Details for the file tensortorrent-0.2.6.tar.gz.

File metadata

  • Download URL: tensortorrent-0.2.6.tar.gz
  • Upload date:
  • Size: 368.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for tensortorrent-0.2.6.tar.gz
Algorithm Hash digest
SHA256 d272df70e18f66916c6e3bb52449db0e1a8960235c9bae9834cb7edeeea66906
MD5 2944ab7013e310c2154aadf4e69b8aa2
BLAKE2b-256 5c951126771dd400b7d04d5588df8018dc64e03644f4fd8cbdc250f1e9449ab6

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6.tar.gz:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 bd929212e8b4b564e423017f20a248c040c92425ca6979b93d5735257e7723c4
MD5 775e0c9c0db4a04d57acbce20c5eece8
BLAKE2b-256 217b80c73811c689d1cfa428ce98ddbcf6e5cbcbd0a2b3813ba0f732cd24d2ab

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 ea4b189a76e2efb6a4ceb960c95317b5143905ba4a9f8eba69d4ce836eba8223
MD5 41b88701106819d66f377d3d9e815395
BLAKE2b-256 c66dd2cadae264910ef82f28c33eb01e2f6812a9c52504fee9f01023a22596b3

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp313-cp313-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 482a4444c33463052cc2a81fd6ea97e298cbb33498a246695324da8f9aad2a6b
MD5 7d94af08f315dd3a08cc176140b14263
BLAKE2b-256 1f2f534626320eff8bf6941058855a284c8cb7c8f7c6d69627e491ddcfbce11e

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl
Algorithm Hash digest
SHA256 37e04d7af879e054dcb95bd9332e5fe943ccea16847741c75696389c2707b972
MD5 61b832b28925302f69c813703e2040e0
BLAKE2b-256 21514bf01fdf8783693788f73d084ebb33857d781173ee1cddebb05ce18f0eb8

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp312-cp312-manylinux_2_17_aarch64.manylinux2014_aarch64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 6a419dc298b72822227e7f35e0adead4e1ce70ca0505897ce338d330f9fa2951
MD5 bb3e583940ac598b9aa6ce06ac3dd050
BLAKE2b-256 4ba6706d0f5ada2c71563006d435c57453cbe011b81237184339d7b45c841690

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp311-cp311-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file tensortorrent-0.2.6-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl.

File metadata

File hashes

Hashes for tensortorrent-0.2.6-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Algorithm Hash digest
SHA256 195643d77f12d3e6c5f52a9c79562ee682246f3b24e20ad04097703b65501bd4
MD5 9164cb44b00701bbae6f84f63f520950
BLAKE2b-256 571f8b2e1438b19c44ad09b5a4664fe2e18ad0ff06f8647c94f647f4f1a5f965

See more details on using hashes here.

Provenance

The following attestation bundles were made for tensortorrent-0.2.6-cp310-cp310-manylinux_2_17_x86_64.manylinux2014_x86_64.whl:

Publisher: release.yml on alhussein-jamil/TensorTorrent

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page