Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Truss

The simplest way to serve AI/ML models in production

PyPI version Python versions ci_status

Truss is the CLI for deploying and serving ML models on Baseten. Package your model's serving logic in Python, launch training jobs, and deploy to production—Truss handles containerization, dependency management, and GPU configuration.

Truss lets you serve models with the Baseten Inference Stack as well as deploy models from any open-source framework: vLLM, SGLang, TensorRT-LLM, transformers, diffusers, PyTorch, TensorFlow, and more.

Get started | 100+ examples | Documentation

Why Truss?

  • Write once, run anywhere: Package model code, weights, and dependencies with a model server that behaves the same in development and production.
  • Fast developer loop: Iterate with live reload, skip Docker and Kubernetes configuration, and use a batteries-included serving environment.
  • Support for all Python frameworks: From transformers and diffusers to PyTorch and TensorFlow to vLLM, SGLang, and TensorRT-LLM, Truss supports models created and served with any framework.
  • Production-ready: Built-in support for GPUs, secrets, caching, and autoscaling when deployed to Baseten or your own infrastructure.

Installation

Install Truss with:

pip install --upgrade truss

Quickstart

Deploying a model to Baseten via Truss turns a Hugging Face model into a production-ready API endpoint. You write a config.yaml that specifies the model, the hardware, and the engine, then uvx truss push builds a TensorRT-optimized container and deploys it. No Python code, no Dockerfile, no container management.

This guide walks through deploying Qwen 2.5 3B Instruct, a small but capable LLM, from a config file to a production API. You'll set up Truss, write a config, deploy to Baseten, and call the model's OpenAI-compatible endpoint.

Set up your environment

Before you begin:

  • Sign up or sign in to Baseten.
  • Install uv, a fast Python package manager. This guide uses uvx to run Truss commands without a separate install step.

Authenticate with Baseten

Log in:

uvx truss login

The CLI asks how you want to authenticate. Generate an API key from Settings > API keys if you pick the paste path:

💻 Let's add a Baseten remote!
How would you like to authenticate?
  Paste an API key
  Log in via browser (OAuth)
🤫 Quietly paste your API_KEY:

Skip the picker with --browser or --api-key:

uvx truss login --browser
uvx truss login --api-key "paste-your-api-key-here"

BASETEN_API_KEY is what the OpenAI client uses later to call the model. It does not skip truss login.

Create a Truss project

Scaffold a new project:

uvx truss init qwen-2.5-3b && cd qwen-2.5-3b

When prompted, name the model Qwen 2.5 3B.

? 📦 Name this model: Qwen 2.5 3B
Truss Qwen 2.5 3B was created in ~/qwen-2.5-3b

This creates a directory with a config.yaml, a model/ directory, and supporting files. For engine-based deployments like this one, you only need config.yaml. The model/ directory is for custom Python code when you need custom preprocessing, postprocessing, or unsupported model architectures.

Write the config

Replace the contents of config.yaml with:

model_metadata:
  tags:
    - openai-compatible
model_name: Qwen-2.5-3B
resources:
  accelerator: L4
  use_gpu: true
trt_llm:
  build:
    base_model: decoder
    checkpoint_repository:
      source: HF
      repo: "Qwen/Qwen2.5-3B-Instruct"
    max_seq_len: 8192
    quantization_type: fp8
    tensor_parallel_count: 1
    num_builder_gpus: 2

That's the entire deployment specification.

  • model_name identifies the model in your Baseten dashboard.
  • resources selects an L4 GPU (24 GB VRAM), which is plenty for a 3B parameter model.
  • trt_llm tells Baseten to use Engine-Builder-LLM, which compiles the model with TensorRT-LLM for optimized inference.
  • checkpoint_repository points to the model weights on Hugging Face. Qwen 2.5 3B Instruct is ungated, so no access token is needed.
  • quantization_type: fp8 compresses weights to 8-bit floating point, cutting memory usage roughly in half with negligible quality loss.
  • max_seq_len: 8192 sets the maximum context length for requests.
  • num_builder_gpus: 2 uses two GPUs during the build phase. FP8 quantization needs more GPU memory at compile time than at inference. The first-model guide says a single L4 can run out of memory during compilation without this.

Deploy

Push the model to Baseten. A default push is a published deployment. --watch (development mode with live reload) is not supported for TRT-LLM. For custom Python models, see Customize a model.

uvx truss push

You should see:

Deploying as a published deployment. Use --watch for a development deployment.

✨ Model Qwen 2.5 3B was successfully pushed ✨

   Model ID:      abc1d2ef
   Deployment ID: xyz123
   Endpoint:      https://model-abc1d2ef.api.baseten.co
   Logs:          https://app.baseten.co/models/abc1d2ef/logs/xyz123

You'll need the model ID to call the model's API. You can also find it in your Baseten dashboard.

Baseten now downloads the model weights from Hugging Face, compiles them with TensorRT-LLM, and deploys the resulting container to an L4 GPU. You can watch progress in the logs linked above.

Call the model

Engine-based deployments serve an OpenAI-compatible API. Once the deployment shows "Active" in the dashboard, call it using the OpenAI SDK or cURL. Replace {model_id} with your model ID from the deployment output.

Install the OpenAI SDK if you don't have it, and export the API key the client will send:

uv pip install openai
export BASETEN_API_KEY="paste-your-api-key-here"

Create a chat completion:

import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["BASETEN_API_KEY"],
    base_url="https://model-{model_id}.api.baseten.co/environments/production/sync/v1",
)

response = client.chat.completions.create(
    model="Qwen-2.5-3B",
    messages=[
        {"role": "user", "content": "What is machine learning?"}
    ],
)

print(response.choices[0].message.content)

You should see a response like:

Machine learning is a branch of artificial intelligence where systems learn
patterns from data to make predictions or decisions without being explicitly
programmed for each task...

Any code that works with the OpenAI SDK works with your deployment. Just point the base_url at your model's endpoint.

Your model ID is the string after /models/ in the logs URL from uvx truss push. You can also find it in your Baseten dashboard.

IDE support

Truss ships a JSON schema for config.yaml. Projects created with truss init include a schema reference automatically, giving you autocompletion, hover docs, and validation in any editor that supports the YAML language server (VS Code, JetBrains, Neovim, and others).

To add schema support to an existing config.yaml, add this comment as the first line:

# yaml-language-server: $schema=https://raw.githubusercontent.com/basetenlabs/truss/main/truss/config.schema.json

Release files for truss 0.18.32rc2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for truss 0.18.32rc2
File Size Uploaded
truss-0.18.32rc2.tar.gz 637.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for truss 0.18.32rc2
File Interpreter ABI Platform
truss-0.18.32rc2-py3-none-any.whl Python 3 none any Details

Total release size: 1.4 MB

Release files / truss-0.18.32rc2.tar.gz

Download URL truss-0.18.32rc2.tar.gz
Size 637.2 kB
Tags Source
SHA-256 checksum
How to use checksums
2ce0443d9f2cab78ae6b6c6d1e397e1d997011a72b6bac1706c66b9b674bdda3
BLAKE2b-256 checksum
How to use checksums
bf16f81748c1efe95476e3792b5d67dee0c419305429ed7b824f76deb44d8b3c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release files / truss-0.18.32rc2-py3-none-any.whl

Download URL truss-0.18.32rc2-py3-none-any.whl
Size 760.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
40007c46a2c0342b3d9cd4c0c463983905ce3ca3963bb770fcb833d75bb43d4f
BLAKE2b-256 checksum
How to use checksums
69e3358a60c7cc5a93e1bcd06c6a6ffe6c5a36ad0fcd72a997a6b341a53c05eb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 18, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.18.32rc2 This release

2 release files

0.18.9

2 release files

0.18.8

2 release files

0.18.3

2 release files

0.18.2

2 release files

0.18.1

2 release files

0.18.0

2 release files

0.17.0

2 release files

0.16.2

2 release files

0.16.1

2 release files

0.16.0

2 release files

0.15.9

2 release files

0.15.8

2 release files

0.15.7

2 release files

0.15.6

2 release files

0.15.5

2 release files

0.15.4

2 release files

0.14.2

2 release files

0.14.1

2 release files

0.14.0

2 release files

0.13.5

2 release files

0.13.4

2 release files

0.12.9

2 release files

0.12.8

2 release files

0.12.7

2 release files

0.12.6

2 release files

0.12.5

2 release files

0.12.4

2 release files

0.12.0

2 release files

0.11.8

2 release files

0.11.7

2 release files

0.11.6

2 release files

0.11.5

2 release files

0.11.4

2 release files

0.11.3

2 release files

0.11.2

2 release files

0.11.1

2 release files

0.10.9

2 release files

0.10.8

2 release files

0.10.7

2 release files

0.10.6

2 release files

0.10.5

2 release files

0.10.4

2 release files

0.10.3

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.97

2 release files

0.9.96

2 release files

0.9.95

2 release files

0.9.94

2 release files

0.9.93

2 release files

0.9.92

2 release files

0.9.91

2 release files

0.9.90

2 release files

0.9.87

2 release files

0.9.86

2 release files

0.9.85

2 release files

0.9.84

2 release files

0.9.83

2 release files

0.9.82

2 release files

0.9.81

2 release files

0.9.80

2 release files

0.9.79

2 release files

0.9.74

2 release files

0.9.73

2 release files

0.9.72

2 release files

0.9.71

2 release files

0.9.70

2 release files

0.9.69

2 release files

0.9.68

2 release files

0.9.64

2 release files

0.9.63

2 release files

0.9.62

2 release files

0.9.60

2 release files

0.9.59

2 release files

0.9.57

2 release files

0.9.53

2 release files

0.9.51

2 release files

0.9.50

2 release files

0.9.46

2 release files

0.9.45

2 release files

0.9.44

2 release files

0.9.40

2 release files

0.9.39

2 release files

0.9.38

2 release files

0.9.37

2 release files

0.9.36

2 release files

0.9.35

2 release files

0.9.34

2 release files

0.9.29

2 release files

0.9.28

2 release files

0.9.25

2 release files

0.9.24

2 release files

0.9.23

2 release files

0.9.22

2 release files

0.9.21

2 release files

0.9.18

2 release files

0.9.17

2 release files

0.9.16

2 release files

0.9.15

2 release files

0.9.14

2 release files

0.9.13

2 release files

0.9.12

2 release files

0.9.10

2 release files

0.9.9

2 release files

0.9.8

2 release files

0.9.7

2 release files

0.9.6

2 release files

0.9.5

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.7.23

2 release files

0.7.22

2 release files

0.7.19

2 release files

0.7.18

2 release files

0.7.17

2 release files

0.7.14

2 release files

0.7.13

2 release files

0.7.12

2 release files

0.7.11

2 release files

0.7.9

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.2

2 release files

0.7.1

2 release files

0.7.0

2 release files

0.6.4

2 release files

0.6.3

2 release files

0.6.2

2 release files

0.6.1

2 release files

0.6.0

2 release files

0.5.8

2 release files

0.5.6

2 release files

0.5.5

2 release files

0.5.4

2 release files

0.5.3

2 release files

0.5.2

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.0

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.9

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.33

2 release files

0.0.30

2 release files

0.0.29

2 release files

0.0.28

2 release files

0.0.27

2 release files

0.0.26

2 release files

0.0.25

2 release files

0.0.24

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page