Skip to main content

PyTriton

PyTriton - a Flask/FastAPI-like framework designed to streamline the use of NVIDIA’s Triton Inference Server.

For comprehensive guidance on how to deploy your models, optimize performance, and explore the API, delve into the extensive resources found in our documentation.

Features at a Glance

The distinct capabilities of PyTriton are summarized in the feature matrix:

Feature

Description

Native Python support

You can create any Python function and expose it as an HTTP/gRPC API.

Framework-agnostic

You can run any Python code with any framework of your choice, such as: PyTorch, TensorFlow, or JAX.

Performance optimization

You can benefit from dynamic batching, response cache, model pipelining, clusters, performance tracing, and GPU/CPU inference.

Decorators

You can use batching decorators to handle batching and other pre-processing tasks for your inference function.

Easy installation and setup

You can use a simple and familiar interface based on Flask/FastAPI for easy installation and setup.

Model clients

You can access high-level model clients for HTTP/gRPC requests with configurable options and both synchronous and asynchronous API.

Streaming (alpha)

You can stream partial responses from a model by serving it in a decoupled mode.

Learn more about PyTriton’s architecture.

Prerequisites

Before proceeding with the installation of PyTriton, ensure your system meets the following criteria:

  • Operating System: Compatible with glibc version 2.35 or higher. - Primarily tested on Ubuntu 22.04. - Other supported OS include Debian 11+, Rocky Linux 9+, and Red Hat UBI 9+. - Use ldd --version to verify your glibc version.

  • Python: Version 3.8 or newer.

  • pip: Version 20.3 or newer.

  • libpython: Ensure libpython3.*.so is installed, corresponding to your Python version.

Install

The PyTriton can be installed from pypi.org by running the following command:

pip install nvidia-pytriton

Important: The Triton Inference Server binary is installed as part of the PyTriton package.

Discover more about PyTriton’s installation procedures, including Docker usage, prerequisites, and insights into building binaries from source to match your specific Triton server versions.

Quick Start

The quick start presents how to run Python model in Triton Inference Server without need to change the current working environment. In the example we are using a simple Linear model.

The infer_fn is a function that takes an data tensor and returns a list with single output tensor. The @batch from batching decorators is used to handle batching for the model.

import numpy as np
from pytriton.decorators import batch

@batch
def infer_fn(data):
    result = data * np.array([[-1]], dtype=np.float32)  # Process inputs and produce result
    return [result]

In the next step, you can create the binding between the inference callable and Triton Inference Server using the bind method from pyTriton. This method takes the model name, the inference callable, the inputs and outputs tensors, and an optional model configuration object.

from pytriton.model_config import Tensor
from pytriton.triton import Triton
triton = Triton()
triton.bind(
    model_name="Linear",
    infer_func=infer_fn,
    inputs=[Tensor(name="data", dtype=np.float32, shape=(-1,)),],
    outputs=[Tensor(name="result", dtype=np.float32, shape=(-1,)),],
)
triton.run()

Finally, you can send an inference query to the model using the ModelClient class. The infer_sample method takes the input data as a numpy array and returns the output data as a numpy array. You can learn more about the ModelClient class in the clients section.

from pytriton.client import ModelClient

client = ModelClient("localhost", "Linear")
data = np.array([1, 2, ], dtype=np.float32)
print(client.infer_sample(data=data))

After the inference is done, you can stop the Triton Inference Server and close the client:

client.close()
triton.stop()

The output of the inference should be:

{'result': array([-1., -2.], dtype=float32)}

For the full example, including defining the model and binding it to the Triton server, check out our detailed Quick Start instructions. Get your model up and running, explore how to serve it, and learn how to invoke it from client applications.

The full example code can be found in examples/linear_random_pytorch.

Examples

The examples page showcases various use cases of serving models using PyTriton. This includes simple examples of running models in PyTorch, TensorFlow2, JAX, and plain Python. In addition, more advanced scenarios are covered, such as online learning, multi-node models, and deployment on Kubernetes using PyTriton. Each example is accompanied by instructions on how to build and run it. Discover more about utilizing PyTriton by exploring our examples.

Metadata

Release files for nvidia-pytriton 0.7.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Built distributions (wheels)

Table of built distributions (wheels) for nvidia-pytriton 0.7.0
File Interpreter ABI Platform
nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_x86_64.whl Python 3 none Linux glibc 2.35+ x86-64 Details
nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_aarch64.whl Python 3 none Linux glibc 2.35+ ARM64 Details

Total release size: 109.2 MB

Release files / nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_x86_64.whl

Download URL nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_x86_64.whl
Size 55.4 MB
Tags Linux glibc 2.35+ x86-64 Python 3
SHA-256 checksum
How to use checksums
21f063ff93bb426d0c06b7a55e7eb9efcfe62c0dddd9a908e9877e18710128df
BLAKE2b-256 checksum
How to use checksums
65159192f7ffe7fab32f545d2a51bdde01e8a50363b8694f0d441efe0cc5310d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.0.0 CPython/3.10.18

Release files / nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_aarch64.whl

Download URL nvidia_pytriton-0.7.0-py3-none-manylinux_2_35_aarch64.whl
Size 53.8 MB
Tags Linux glibc 2.35+ ARM64 Python 3
SHA-256 checksum
How to use checksums
f99e07ee56fe065337453956afba17d2673e02ec6a632c3de26ce6e7b32a0898
BLAKE2b-256 checksum
How to use checksums
a5ea278744f3a731b1106fc69038b5173469d6d16be31a8f580faf3893bfffcb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/5.0.0 CPython/3.10.18

Release history Release notifications | RSS feed

This release

0.7.0 This release

2 release files

0.6.0

2 release files

0.5.14

2 release files

0.5.13

2 release files

0.5.12

2 release files

0.5.11

2 release files

0.5.9

2 release files

0.5.8

2 release files

0.5.7

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.2

2 release files

0.4.1

1 release file

0.4.0

1 release file

0.3.1

1 release file

0.3.0

1 release file

0.2.5

1 release file

0.2.4

1 release file

0.2.3

1 release file

0.2.2

1 release file

0.2.1

1 release file

0.2.0

1 release file

0.1.5

1 release file

0.1.4

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page