Skip to main content

High performance client for Baseten.co

This library provides a high-performance Python client for Baseten.co endpoints including embeddings, reranking, and classification. It was built for massive concurrent post requests to any URL, also outside of baseten.co. PerformanceClient releases the GIL while performing requests in the Rust, and supports simultaneous sync and async usage. It was benchmarked with >1200 rps per client in our blog. PerformanceClient is built on top of pyo3, reqwest and tokio and is MIT licensed.

benchmarks

Installation

pip install baseten_performance_client

Usage

import os
import asyncio
from baseten_performance_client import PerformanceClient, OpenAIEmbeddingsResponse, RerankResponse, ClassificationResponse

api_key = os.environ.get("BASETEN_API_KEY")
base_url_embed = "https://model-yqv4yjjq.api.baseten.co/environments/production/sync"
# Also works with OpenAI or Mixedbread.
# base_url_embed = "https://api.openai.com" or "https://api.mixedbread.com"

# Basic client setup
client = PerformanceClient(base_url=base_url_embed, api_key=api_key)

# Advanced setup with HTTP version selection and connection pooling
from baseten_performance_client import HttpClientWrapper
http_wrapper = HttpClientWrapper(http_version=1)  # HTTP/1.1 (default)
advanced_client = PerformanceClient(
    base_url=base_url_embed,
    api_key=api_key,
    http_version=1,  # HTTP/1.1
    client_wrapper=http_wrapper  # Share connection pool
)

Embeddings

Synchronous Embedding

from baseten_performance_client import RequestProcessingPreference

texts = ["Hello world", "Example text", "Another sample"]
preference = RequestProcessingPreference(
    batch_size=16,
    max_concurrent_requests=32,
    timeout_s=360,
    max_chars_per_request=256000,  # Character limit per request
    hedge_delay=0.5,  # Enable hedging with 0.5s delay
    total_timeout_s=360  # Total operation timeout
)
response = client.embed(
    input=texts,
    model="my_model",
    preference=preference
)

# Accessing embedding data
print(f"Model used: {response.model}")
print(f"Total tokens used: {response.usage.total_tokens}")
print(f"Total time: {response.total_time:.4f}s")
if response.individual_batch_request_times:
    for i, batch_time in enumerate(response.individual_batch_request_times):
        print(f"  Time for batch {i}: {batch_time:.4f}s")

for i, embedding_data in enumerate(response.data):
    print(f"Embedding for text {i} (original input index {embedding_data.index}):")
    # embedding_data.embedding can be List[float] or str (base64)
    if isinstance(embedding_data.embedding, list):
        print(f"  First 3 dimensions: {embedding_data.embedding[:3]}")
        print(f"  Length: {len(embedding_data.embedding)}")

# Using the numpy() method (requires numpy to be installed)
import numpy as np
numpy_array = response.numpy()
print("\nEmbeddings as NumPy array:")
print(f"  Shape: {numpy_array.shape}")
print(f"  Data type: {numpy_array.dtype}")
if numpy_array.shape[0] > 0:
    print(f"  First 3 dimensions of the first embedding: {numpy_array[0][:3]}")

Note: The embed method is versatile and can be used with any embeddings service, e.g. OpenAI API embeddings, not just for Baseten deployments.

Asynchronous Embedding

async def async_embed():
    from baseten_performance_client import RequestProcessingPreference

    texts = ["Async hello", "Async example"]
    preference = RequestProcessingPreference(
        batch_size=16,
        max_concurrent_requests=32,
        timeout_s=360,
        max_chars_per_request=256000,  # Character limit per request
        hedge_delay=0.5,  # Enable hedging with 0.5s delay
        total_timeout_s=360  # Total operation timeout
    )
    response = await client.async_embed(
        input=texts,
        model="my_model",
        preference=preference
    )
    print("Async embedding response:", response.data)

# To run:
# asyncio.run(async_embed())

Embedding Benchmarks

Comparison against pip install openai for /v1/embeddings. Tested with the ./scripts/compare_latency_openai.py with mini_batch_size of 128, and 4 server-side replicas. Results with OpenAI similar, OpenAI allows a max mini_batch_size of 2048.

Number of inputs / embeddings Number of Tasks PerformanceClient (s) AsyncOpenAI (s) Speedup
128 1 0.12 0.13 1.08×
512 4 0.14 0.21 1.50×
8 192 64 0.83 1.95 2.35×
131 072 1 024 4.63 39.07 8.44×
2 097 152 16 384 70.92 903.68 12.74×

General Batch POST

The batch_post method is generic. It can be used to send POST requests to any URL, not limited to Baseten endpoints. The input and output can be any JSON item.

Synchronous Batch POST

from baseten_performance_client import RequestProcessingPreference

payload1 = {"model": "my_model", "input": ["Batch request sample 1"]}
payload2 = {"model": "my_model", "input": ["Batch request sample 2"]}
preference = RequestProcessingPreference(
    max_concurrent_requests=32,
    timeout_s=360,
    hedge_delay=0.5,  # Enable hedging with 0.5s delay
    total_timeout_s=360,  # Total operation timeout
    extra_headers={"x-custom-header": "value"}  # Custom headers
)
response_obj = client.batch_post(
    url_path="/v1/embeddings", # Example path, adjust to your needs
    payloads=[payload1, payload2],
    preference=preference
)
print(f"Total time for batch POST: {response_obj.total_time:.4f}s")
for i, (resp_data, headers, time_taken) in enumerate(zip(response_obj.data, response_obj.response_headers, response_obj.individual_request_times)):
    print(f"Response {i+1}:")
    print(f"  Data: {resp_data}")
    print(f"  Headers: {headers}")
    print(f"  Time taken: {time_taken:.4f}s")

Asynchronous Batch POST

async def async_batch_post_example():
    from baseten_performance_client import RequestProcessingPreference

    payload1 = {"model": "my_model", "input": ["Async batch sample 1"]}
    payload2 = {"model": "my_model", "input": ["Async batch sample 2"]}
preference = RequestProcessingPreference(
    max_concurrent_requests=32,
    timeout_s=360,
    hedge_delay=0.5,  # Enable hedging with 0.5s delay
    total_timeout_s=360,  # Total operation timeout
    extra_headers={"x-custom-header": "value"}  # Custom headers
)
    response_obj = await client.async_batch_post(
        url_path="/v1/embeddings",
        payloads=[payload1, payload2],
        preference=preference
    )
    print(f"Async total time for batch POST: {response_obj.total_time:.4f}s")
    for i, (resp_data, headers, time_taken) in enumerate(zip(response_obj.data, response_obj.response_headers, response_obj.individual_request_times)):
        print(f"Async Response {i+1}:")
        print(f"  Data: {resp_data}")
        print(f"  Headers: {headers}")
        print(f"  Time taken: {time_taken:.4f}s")

# To run:
# asyncio.run(async_batch_post_example())

Reranking

Reranking compatible with BEI or text-embeddings-inference.

Synchronous Reranking

from baseten_performance_client import RequestProcessingPreference

query = "What is the best framework?"
documents = ["Doc 1 text", "Doc 2 text", "Doc 3 text"]
preference = RequestProcessingPreference(
    batch_size=16,
    max_concurrent_requests=32,
    timeout_s=360,
    max_chars_per_request=256000,  # Character limit per request
    hedge_delay=0.5,  # Enable hedging with 0.5s delay
    total_timeout_s=360  # Total operation timeout
)
rerank_response = client.rerank(
    query=query,
    texts=documents,
    model="rerank-model",  # Optional model specification
    return_text=True,
    preference=preference
)
for res in rerank_response.data:
    print(f"Index: {res.index} Score: {res.score}")

Asynchronous Reranking

async def async_rerank():
    from baseten_performance_client import RequestProcessingPreference

    query = "Async query sample"
    docs = ["Async doc1", "Async doc2"]
    preference = RequestProcessingPreference(
        batch_size=16,
        max_concurrent_requests=32,
        timeout_s=360,
        max_chars_per_request=256000,  # Character limit per request
        hedge_delay=0.5,  # Enable hedging with 0.5s delay
        total_timeout_s=360  # Total operation timeout
    )
    response = await client.async_rerank(
        query=query,
        texts=docs,
        model="rerank-model",  # Optional model specification
        return_text=True,
        preference=preference
    )
    for res in response.data:
        print(f"Async Index: {res.index} Score: {res.score}")

# To run:
# asyncio.run(async_rerank())

Classification

Predict (classification endpoint) compatible with BEI or text-embeddings-inference.

Synchronous Classification

from baseten_performance_client import RequestProcessingPreference

texts_to_classify = [
    "This is great!",
    "I did not like it.",
    "Neutral experience."
]
preference = RequestProcessingPreference(
    batch_size=16,
    max_concurrent_requests=32,
    timeout_s=360,
    max_chars_per_request=256000,  # Character limit per request
    hedge_delay=0.5,  # Enable hedging with 0.5s delay
    total_timeout_s=360  # Total operation timeout
)
classify_response = client.classify(
    inputs=texts_to_classify,
    model="classification-model",  # Optional model specification
    preference=preference
)
for group in classify_response.data:
    for result in group:
        print(f"Label: {result.label}, Score: {result.score}")

Asynchronous Classification

async def async_classify():
    from baseten_performance_client import RequestProcessingPreference

    texts = ["Async positive", "Async negative"]
    preference = RequestProcessingPreference(
        batch_size=16,
        max_concurrent_requests=32,
        timeout_s=360,
        max_chars_per_request=256000,  # Character limit per request
        hedge_delay=0.5,  # Enable hedging with 0.5s delay
        total_timeout_s=360  # Total operation timeout
    )
    response = await client.async_classify(
        inputs=texts,
        model="classification-model",  # Optional model specification
        preference=preference
    )
    for group in response.data:
        for res in group:
            print(f"Async Label: {res.label}, Score: {res.score}")

# To run:
# asyncio.run(async_classify())

Advanced Features

RequestProcessingPreference

The RequestProcessingPreference class provides a unified way to configure all request processing parameters. This is the recommended approach for advanced configuration as it provides better type safety and clearer intent.

The framework defaults are 256 concurrent requests, a batch size of 8, and 8,000 characters per request. max_chars_per_request accepts values from 50 through 1,048,576.

from baseten_performance_client import RequestProcessingPreference

# Create a preference with custom settings
preference = RequestProcessingPreference(
    max_concurrent_requests=64,        # Parallel requests (default: 256)
    batch_size=32,                     # Items per batch (default: 8)
    timeout_s=30.0,                   # Per-request timeout (default: 3600.0)
    hedge_delay=0.5,                  # Hedging delay (default: None)
    hedge_budget_pct=0.15,            # Hedge budget percentage (default: 0.10)
    retry_budget_pct=0.08,            # Retry budget percentage (default: 0.05)
    max_retries=5,                    # Maximum HTTP retries (default: 5)
    initial_backoff_ms=250,           # Initial backoff in milliseconds (default: 125)
    total_timeout_s=300.0              # Total operation timeout (default: None)
)

# Use with any method
response = client.embed(
    input=["text1", "text2"],
    model="my_model",
    preference=preference
)

# Also works with async methods
response = await client.async_embed(
    input=["text1", "text2"],
    model="my_model",
    preference=preference
)

Property-based Configuration: You can also modify preferences after creation using property setters:

# Create preference and modify properties
preference = RequestProcessingPreference()
preference.max_concurrent_requests = 64        # Set parallel requests
preference.batch_size = 32                     # Set batch size
preference.timeout_s = 30.0                    # Set timeout
preference.hedge_delay = 0.5                   # Enable hedging
preference.hedge_budget_pct = 0.15            # Set hedge budget
preference.retry_budget_pct = 0.08            # Set retry budget
preference.max_retries = 3                     # Set max retries
preference.initial_backoff_ms = 250            # Set backoff

# Use with any method
response = client.embed(
    input=["text1", "text2"],
    model="my_model",
    preference=preference
)

Budget Percentages:

  • hedge_budget_pct: Percentage of total requests allocated for hedging (default: 10%)
  • retry_budget_pct: Percentage of total requests allocated for retries (default: 5%)
  • Maximum allowed: 300% for both budgets

Retry Configuration:

  • HTTP status-code retries are controlled by max_retries, not by retry_budget_pct.
  • Retryable status codes by default: 408, 409, 429, and 500 through 599.
  • Use non_retryable_status_codes={529} to opt specific statuses out of the default retry policy.
  • max_retries: Maximum HTTP status-code retries per request (default: 5, max: 6). Set to 0 to disable these retries.
  • retry_budget_pct: Budget for timeout and network-error retry paths (default: 5%, max: 300%).
  • initial_backoff_ms: Initial backoff duration in milliseconds (default: 125, range: 50-45000).
  • Backoff multiplies by 4 after each retry, caps at 45000ms, and adds 0-99ms jitter. With defaults, the retry sleeps are about 125ms, 500ms, 2000ms, 8000ms, and 32000ms; a sixth retry sleeps about 45000ms.

Request Hedging

The client supports request hedging for improved latency by sending duplicate requests after a specified delay:

# Enable hedging with 0.5 second delay
preference = RequestProcessingPreference(
    hedge_delay=0.5,  # Send hedge request after 0.5s
    max_chars_per_request=256000,
    total_timeout_s=360
)
response = client.embed(
    input=texts,
    model="my_model",
    preference=preference
)

Custom Headers

Use custom headers with batch_post:

preference = RequestProcessingPreference(
    extra_headers={
        "x-custom-header": "value",
        "authorization": "Bearer token"
    }
)
response = client.batch_post(
    url_path="/v1/embeddings",
    payloads=payloads,
    preference=preference
)

HTTP Version Selection

Choose between HTTP/1.1 and HTTP/2:

# HTTP/1.1 (default, better for high concurrency)
client_http1 = PerformanceClient(base_url, api_key, http_version=1)

# HTTP/2 (better for single requests)
client_http2 = PerformanceClient(base_url, api_key, http_version=2)

Connection Pooling

Share connection pools across multiple clients:

from baseten_performance_client import HttpClientWrapper

# Create shared wrapper
wrapper = HttpClientWrapper(http_version=1)

# Reuse across multiple clients
client1 = PerformanceClient(base_url="https://api1.example.com", client_wrapper=wrapper)
client2 = PerformanceClient(base_url="https://api2.example.com", client_wrapper=wrapper)

HTTP Proxy Support

Route all HTTP requests through a proxy (e.g., for connection pooling with Envoy):

from baseten_performance_client import HttpClientWrapper

# Create wrapper with HTTP proxy
wrapper = HttpClientWrapper(
    http_version=1,
    proxy="http://envoy-proxy.local:8080"
)

# Share the wrapper across multiple clients
client1 = PerformanceClient(
    base_url="https://api1.example.com",
    api_key="your_key",
    client_wrapper=wrapper
)
client2 = PerformanceClient(
    base_url="https://api2.example.com",
    api_key="your_key",
    client_wrapper=wrapper
)
# Both clients will use the same connection pool and proxy

You can also specify the proxy directly when creating a client:

client = PerformanceClient(
    base_url="https://api.example.com",
    api_key="your_key",
    proxy="http://envoy-proxy.local:8080"
)

Endpoint Pool and Health Checks

Route traffic across reusable endpoints with deterministic weighted routing. Each Endpoint owns its own health worker, so the same endpoint object can be shared across many pools without duplicate probes:

from baseten_performance_client import Endpoint, EndpointPool, HttpClientWrapper, PerformanceClient

health_wrapper = HttpClientWrapper(http_version=1)
endpoint_a = Endpoint(
    base_url="https://model-AAAA.api.baseten.co/environments/production/sync",
    api_key="your_key",
    client_wrapper=health_wrapper,
    deployment_health_path="/health",
    deployment_timeout_is_no_vote=False,
)
endpoint_b = Endpoint(
    base_url="https://model-BBBB.api.baseten.co/environments/production/sync",
    api_key="your_key",
    client_wrapper=health_wrapper,
    deployment_health_path="/health",
    deployment_timeout_is_no_vote=False,
)

endpoint_pool = EndpointPool(
    endpoints=[endpoint_a, endpoint_b],
    endpoint_weights=[0.8, 0.2],  # deterministic weighted routing
)

client = PerformanceClient(
    base_url="https://model-AAAA.api.baseten.co/environments/production/sync",
    api_key="your_key",
    endpoint_pool=endpoint_pool,
)

Health semantics:

  • Weights are deterministic weighted routing, not weighted round robin.
  • Each configured health check is retried up to health_check_retries, and one successful retry is enough for that check.
  • If an endpoint has deep_health_url configured, both the shallow deployment health path and the deep health URL are evaluated.
  • health_fail_on_first=True short-circuits on the first hard failing check within an endpoint refresh cycle.

Error Handling

The client can raise several types of errors. Here's how to handle common ones:

  • requests.exceptions.HTTPError: This error is raised for HTTP issues, such as authentication failures (e.g., 403 Forbidden if the API key is wrong), server errors (e.g., 5xx), or if the endpoint is not found (404). You can inspect e.response.status_code and e.response.text (or e.response.json() if the body is JSON) for more details.
  • ValueError: This error can occur due to invalid input parameters (e.g., an empty input list for embed, invalid batch_size or max_concurrent_requests values). It can also be raised by response.numpy() if embeddings are not float vectors or have inconsistent dimensions.

Here's an example demonstrating how to catch these errors for the embed method:

import requests
from baseten_performance_client import RequestProcessingPreference

# client = PerformanceClient(base_url="your_baseten_url", api_key="your_baseten_api_key")

texts_to_embed = ["Hello world", "Another text example"]
try:
    preference = RequestProcessingPreference(
        batch_size=2,
        max_concurrent_requests=4,
        timeout_s=60 # Timeout in seconds
    )
    response = client.embed(
        input=texts_to_embed,
        model="your_embedding_model", # Replace with your actual model name
        preference=preference
    )
    # Process successful response
    print(f"Model used: {response.model}")
    print(f"Total tokens: {response.usage.total_tokens}")
    for item in response.data:
        embedding_preview = item.embedding[:3] if isinstance(item.embedding, list) else "Base64 Data"
        print(f"Index {item.index}, Embedding (first 3 dims or type): {embedding_preview}")

except requests.exceptions.HTTPError as e:
    print(f"An HTTP error occurred: {e}, code {e.args[0]}")

For asynchronous methods (async_embed, async_rerank, async_classify, async_batch_post), the same exceptions will be raised by the await call and can be caught using a try...except block within an async def function.

Development

# Install prerequisites
sudo apt-get install patchelf
# Install cargo if not already installed.

# Set up a Python virtual environment
python -m venv .venv
source .venv/bin/activate

# Install development dependencies
pip install maturin[patchelf] pytest requests numpy

# Build and install the Rust extension in development mode
maturin develop
cargo fmt
# Run tests
pytest tests

Contributions

Feel free to contribute to this repo, tag @michaelfeil for review.

License

MIT License

Release files for baseten-performance-client 0.1.15

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for baseten-performance-client 0.1.15
File Size Uploaded
baseten_performance_client-0.1.15.tar.gz 119.4 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for baseten-performance-client 0.1.15
File
baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_x86_64.whl CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ x86-64 Details
baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_i686.whl CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ x86-32 Details
baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_armv7l.whl CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ ARMv7l Details
baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_aarch64.whl CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ ARM64 Details
baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ x86-64 Details
baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ PowerPC 64-le Details
baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_i686.manylinux2014_i686.whl CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ x86-32 Details
baseten_performance_client-0.1.15-cp313-cp313t-macosx_11_0_arm64.whl CPython 3.13 CPython 3.13 free-threading macOS 11.0+ ARM64 Details
baseten_performance_client-0.1.15-cp313-cp313t-macosx_10_12_x86_64.whl CPython 3.13 CPython 3.13 free-threading macOS 10.12+ x86-64 Details
baseten_performance_client-0.1.15-cp38-abi3-win_amd64.whl CPython 3.8 abi3 Windows x86-64 Details
baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_x86_64.whl CPython 3.8 abi3 Linux musl 1.2+ x86-64 Details
baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_i686.whl CPython 3.8 abi3 Linux musl 1.2+ x86-32 Details
baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_armv7l.whl CPython 3.8 abi3 Linux musl 1.2+ ARMv7l Details
baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_aarch64.whl CPython 3.8 abi3 Linux musl 1.2+ ARM64 Details
baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_armv7l.whl CPython 3.8 abi3 Linux glibc 2.28+ ARMv7l Details
baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_aarch64.whl CPython 3.8 abi3 Linux glibc 2.28+ ARM64 Details
baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.8 abi3 Linux glibc 2.17+ x86-64 Details
baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl CPython 3.8 abi3 Linux glibc 2.17+ PowerPC 64-le Details
baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_i686.manylinux2014_i686.whl CPython 3.8 abi3 Linux glibc 2.17+ x86-32 Details
baseten_performance_client-0.1.15-cp38-abi3-macosx_11_0_arm64.whl CPython 3.8 abi3 macOS 11.0+ ARM64 Details
baseten_performance_client-0.1.15-cp38-abi3-macosx_10_12_x86_64.whl CPython 3.8 abi3 macOS 10.12+ x86-64 Details

Total release size: 117.9 MB

Release files / baseten_performance_client-0.1.15.tar.gz

Download URL baseten_performance_client-0.1.15.tar.gz
Size 119.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a98bf440d3ddd6c3229decfe9b714595e41c39a454d0bfa7009d84c7a2b396dd
BLAKE2b-256 checksum
How to use checksums
e0ff4606667e34b4282420f15d32c91d7361e595a112b2d97adc10802619d0a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_x86_64.whl
Size 6.3 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ x86-64
SHA-256 checksum
How to use checksums
ff14316b51a18ef6a5e4f4b8f0646cdd891759b86036c1d1e68004d698e113b5
BLAKE2b-256 checksum
How to use checksums
8c8e02a2bfe0a4128ed2634cdcc895850c72a8166f357d4758ffce9a1cebbf9e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_i686.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_i686.whl
Size 6.1 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ x86-32
SHA-256 checksum
How to use checksums
a2b07fa4b7fcbb5574b836f934ce4d4510ac23ca528b27a085557fb0ea627759
BLAKE2b-256 checksum
How to use checksums
88d8bf4e485f2e5a5dcd26dbf7e8cbb04a0e249e376dec4aaa847ce62679eeac
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_armv7l.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_armv7l.whl
Size 5.6 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ ARMv7l
SHA-256 checksum
How to use checksums
09f4d23c1bf26b204c34a772e10211b6e086b47227927a07df0144cff54e8328
BLAKE2b-256 checksum
How to use checksums
2982fc66d16633c8efc11c1647b1b21faaf3b1fbd5ee9362a753db6739b1ba66
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_aarch64.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-musllinux_1_2_aarch64.whl
Size 6.6 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux musl 1.2+ ARM64
SHA-256 checksum
How to use checksums
b3a937984e4377fdd681cff4970d154e7465654122fb0b2ff6648e6f0f042ecb
BLAKE2b-256 checksum
How to use checksums
400e1e7bfc3c87046af565b1307d51afa35a2929f9effbee638810387ca849ae
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 5.9 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ x86-64
SHA-256 checksum
How to use checksums
0aaa1de5f62fc35417c7ee256ac84fe31de6077532f6de8bf147f6b93e6e1115
BLAKE2b-256 checksum
How to use checksums
9b5756f46b193180b83729a2f1aee408cb0982c2240c4c855f42a7b9a255c0a1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl
Size 6.6 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ PowerPC 64-le
SHA-256 checksum
How to use checksums
f61f28d3125a431b8825098a580e1a589517a10d611627134bf5f42e8ac6eeda
BLAKE2b-256 checksum
How to use checksums
67009af7811e39ef4623f0db2858d54312ba4db6a36af4f4013e58f5ae98bc14
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_i686.manylinux2014_i686.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-manylinux_2_17_i686.manylinux2014_i686.whl
Size 6.1 MB
Tags CPython 3.13 CPython 3.13 free-threading Linux glibc 2.17+ x86-32
SHA-256 checksum
How to use checksums
5c40d2ec8f90cf48d362ba470aac7aca213cec07c020e68325da9067c0370eba
BLAKE2b-256 checksum
How to use checksums
9b0a93920b76a5bc8d83b9319b962ddcbe307c9a7771bef6f6d31b3847c2c5e9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-macosx_11_0_arm64.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-macosx_11_0_arm64.whl
Size 3.5 MB
Tags CPython 3.13 CPython 3.13 free-threading macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
5ebef8dbd31d3c42e0637abcc47951a3c17a9b0ea8981ce90c8b0abd1d777b03
BLAKE2b-256 checksum
How to use checksums
77e84ffa102ea42f4841eba6b684b600e7219dc6b5ea6c7092134e8490efd461
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp313-cp313t-macosx_10_12_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp313-cp313t-macosx_10_12_x86_64.whl
Size 3.6 MB
Tags CPython 3.13 CPython 3.13 free-threading macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
72a335485dcb6f92f7a7cb0c5ff6675a9eb4876f90d83743af7d8c61fde92081
BLAKE2b-256 checksum
How to use checksums
301b80ec909ae30290f248add4713881d3b173ee705c5d26e753258e22bddd61
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-win_amd64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-win_amd64.whl
Size 3.5 MB
Tags CPython 3.8 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
60d0def91ba6179c83da173dac25d447b6aa2bf313142adab3155e52322c01f2
BLAKE2b-256 checksum
How to use checksums
7ef4d0ebd656c700f35b733c6a08db1904e0a93b43ecd6e0015b7f8359283cb0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_x86_64.whl
Size 6.5 MB
Tags CPython 3.8 Linux musl 1.2+ x86-64 abi3
SHA-256 checksum
How to use checksums
c062592af56d23415813b3fa2a87afdab2fdaccfb34204fad1abb254880dfcfc
BLAKE2b-256 checksum
How to use checksums
787c468fa8c9ecb1198490ca2eb297d9e65426416486c892bdcd3bede9cb8edd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_i686.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_i686.whl
Size 6.3 MB
Tags CPython 3.8 Linux musl 1.2+ x86-32 abi3
SHA-256 checksum
How to use checksums
f72b15329bec2fc7655cd5500d69c8b85c7ac56587e7337b577a2784612c93bf
BLAKE2b-256 checksum
How to use checksums
833fc103b5352e8264a2d92b9a22ffcdb32e2ab87a57aeb6c92fe21643304fc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_armv7l.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_armv7l.whl
Size 5.8 MB
Tags CPython 3.8 Linux musl 1.2+ ARMv7l abi3
SHA-256 checksum
How to use checksums
3dce074288c6a0f8a6d8b93a1bd5b1da5e573848b67b338759ad523a3718c59f
BLAKE2b-256 checksum
How to use checksums
747b140ee30d18c5c8e237913008cf2442508493f18cd3837facb1ff5d930500
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_aarch64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-musllinux_1_2_aarch64.whl
Size 6.8 MB
Tags CPython 3.8 Linux musl 1.2+ ARM64 abi3
SHA-256 checksum
How to use checksums
b9a66069a8447fc04bc36ae87890ce0836e3c071239d3c052a3759b634d533a3
BLAKE2b-256 checksum
How to use checksums
dacd86810aafdb9a8606b96df7989895460dc0e8eb5e7bf365031e8831c71988
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_armv7l.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_armv7l.whl
Size 5.6 MB
Tags CPython 3.8 Linux glibc 2.28+ ARMv7l abi3
SHA-256 checksum
How to use checksums
beeac0bac70d285b47c6d7de73c31f8ce33794274891ab431ed5fa8aaa4bae4a
BLAKE2b-256 checksum
How to use checksums
0d1fbecba92d538184982f0ee85f2b0b8612fc341c1ca0a0b2a8a7f68761bee1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_aarch64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_28_aarch64.whl
Size 6.5 MB
Tags CPython 3.8 Linux glibc 2.28+ ARM64 abi3
SHA-256 checksum
How to use checksums
95be0581a0a6790e66d6789b5d3e4f8b6455430c952620e0913f775e8fa813ea
BLAKE2b-256 checksum
How to use checksums
50c550531fe0a6b8026609daee7580d451b948ced9403069c69cf126ed04b2d5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 6.1 MB
Tags CPython 3.8 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
a793df6834f143d3a8af99a3479b3776b9a6df3ef3d3588282fb435d2fc8dc19
BLAKE2b-256 checksum
How to use checksums
18f2f187079f4ec6f463a631b56f9a49611c78081df5098e7e4cd2b6d8066517
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_ppc64le.manylinux2014_ppc64le.whl
Size 6.7 MB
Tags CPython 3.8 Linux glibc 2.17+ PowerPC 64-le abi3
SHA-256 checksum
How to use checksums
c5aabb6ad6da06a717a2a0438c2ac97cc69da9b20f6ce3526cd1102c0f3349bb
BLAKE2b-256 checksum
How to use checksums
1a141c9ff667c749e36549883db46134fd09c4ec87afc8bb4a04d5346d8afde1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_i686.manylinux2014_i686.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-manylinux_2_17_i686.manylinux2014_i686.whl
Size 6.3 MB
Tags CPython 3.8 Linux glibc 2.17+ x86-32 abi3
SHA-256 checksum
How to use checksums
d4a3017804523ec67e657411dc862ceb4f07588a181aeaa3c1a5d59a589bca3b
BLAKE2b-256 checksum
How to use checksums
2f9b6f118b46166c1e2067d6c70a3d8ae400e18071733dcedd5d874c3b1f0fa9
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-macosx_11_0_arm64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-macosx_11_0_arm64.whl
Size 3.6 MB
Tags CPython 3.8 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
80e0a9a42d6a9e4cfff39b3aadd0316fc9270e5ab7343151cd1fdab2fc73856d
BLAKE2b-256 checksum
How to use checksums
69f891a40eea2a088623ef5ec9799958088979ba7111dbf7c0b76f9396f82033
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release files / baseten_performance_client-0.1.15-cp38-abi3-macosx_10_12_x86_64.whl

Download URL baseten_performance_client-0.1.15-cp38-abi3-macosx_10_12_x86_64.whl
Size 3.7 MB
Tags CPython 3.8 abi3 macOS 10.12+ x86-64
SHA-256 checksum
How to use checksums
58e351b278a1001f3c2e7384c47f311f5180357b79d84cb51ad93c33382ff555
BLAKE2b-256 checksum
How to use checksums
ac6c8b0125cdf08ddd8609c17ab94365ff81979dd9988e9cdea5c4a4866b048c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via maturin/1.15.0

Release history Release notifications | RSS feed

This release

0.1.15 This release

22 release files

0.1.8

22 release files

0.1.7

22 release files

0.1.6

22 release files

0.1.5

30 release files

0.1.4

30 release files

0.1.3

30 release files

0.1.2

30 release files

0.1.0

30 release files

0.0.9

23 release files

0.0.8

23 release files

0.0.7

23 release files

0.0.6

25 release files

0.0.2

25 release files

0.0.1

25 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page