Skip to main content

tokn-requests Python SDK

The Python package embeds the same Rust routing engine as tokn-sdk. It uses the existing config.toml, config.d, auth.yaml, and auth.d sources and does not require a gateway process.

Installation

Install the PyPI distribution tokn-requests, then import tokn_requests:

python -m pip install tokn-requests==0.2.3

Release wheels use CPython's Python 3.10 stable ABI (abi3-py310): one cp310-abi3 wheel per platform supports regular CPython 3.10 and newer. CI tests the same wheels on CPython 3.10 through 3.14 on Linux x86-64 (glibc 2.17 or newer), macOS 11 or newer on Apple Silicon, and Windows x86-64. Free-threaded CPython builds require a different ABI and are not covered by these wheels. Other platforms can build the source distribution with a current stable Rust toolchain and a C/C++ compiler. Type annotations and native extension stubs are included in the installed package.

Friendly generation API

For a one-off request, start with the client-bound builder:

from tokn_requests import Client

client = Client()

response = await (
  client.generate("smart")
  .system("You are a Python expert.")
  .prompt("Explain this function.")
  .temperature(0.2)
  .send()
)

print(response.text)

Responses normalize text, reasoning, tool calls, token usage, and finish reasons across supported providers.

Common generation controls are available directly on both the client-bound and detached builders:

from tokn_requests import (
  ReasoningEffort,
  ReasoningMode,
  ReasoningSummary,
)

openai_call = (
  client.generate("gpt-5")
  .prompt("Solve this step by step.")
  .max_tokens(2048)
  .reasoning_effort(ReasoningEffort.HIGH)
  .reasoning_summary(ReasoningSummary.AUTO)
)

llama_call = (
  client.generate("local-llama")
  .prompt("Compare these implementations.")
  .top_p(0.9)
  .top_k(40)
  .max_tokens(2048)
)

claude_call = (
  client.generate("claude-sonnet-4.6")
  .prompt("Plan this migration.")
  .max_tokens(2048)
  .reasoning_mode(ReasoningMode.ADAPTIVE)
  .reasoning_effort(ReasoningEffort.HIGH)
)

max_tokens() is a convenience alias for the provider-neutral max_output_tokens() control. Managed routes serialize that limit as max_output_tokens for Responses, max_completion_tokens for OpenAI Chat Completions, and max_tokens for other Chat Completions or Messages routes. Codex account routes reject an explicit limit because that backend does not preserve it.

These examples assume the model selectors route to OpenAI Responses, llama.cpp Chat Completions, and Copilot's Claude Chat Completions fallback, respectively; use selectors from your own configuration. top_p() is portable across compatible routes and accepts values from 0 through 1. Responses supports reasoning effort and summary but not top_k or an enabled/adaptive mode. Typed top_k is currently supported on llama.cpp Chat Completions; llama.cpp has no portable reasoning control. Known non-reasoning models reject typed reasoning locally.

Effort levels are validated against cached upstream model metadata, falling back to the provider-specific models.dev catalogue. Discovery exposes x_tokn_router.capabilities.reasoning_efforts: null means unknown and [] means no effort control. Unknown support does not reject an effort value. DeepSeek V4 Flash advertises low, high, and max; V4 Pro advertises high and max. DeepSeek thinking also rejects temperature and top_p, which that backend would ignore. Claude supports adaptive reasoning on 4.6 and newer models but not a reasoning summary. Manual Claude reasoning requires ReasoningMode.ENABLED, an explicit max_tokens() limit, a budget of at least 1024 tokens, and budget_tokens < max_tokens; manual mode is rejected on 4.7 and newer models, while adaptive mode is rejected on 4.5 and older models. Claude effort levels use discovery metadata; sampling compatibility uses the selected model generation. Explicit controls unsupported by the selected route fail clearly after routing instead of being silently dropped or reinterpreted.

passthrough and switch profiles preserve the generated Responses payload verbatim, so they reject typed top_k and reasoning controls that would require post-route lowering. Use an exact, route, or fuzzy profile for the provider-neutral control API.

The snippets below continue using the client created above. Snippets after the detached-request example also reuse its request.

Build an owned request when it needs to be serialized, transformed, queued, or reused independently of a client:

from tokn_requests import GenerateRequest

request = (
  GenerateRequest.builder("smart")
  .prompt("Explain this function.")
  .temperature(0.2)
  .build()
)

serialized = request.to_json()
request = GenerateRequest.from_json(serialized)
request = request.with_changes(max_output_tokens=128)

response = await client.send(request)

As an alternative to client.send(request), use await request.bind(client).send() when fluent binding is more convenient.

Semantic streaming returns typed events:

from tokn_requests import Completed, TextDelta

stream = await client.generate("smart").prompt("Write a haiku.").stream()
async with stream:
  async for event in stream:
    if isinstance(event, TextDelta):
      print(event.text, end="")
    elif isinstance(event, Completed):
      print(f"\nfinish reason: {event.finish_reason}")

Use stream_text() when only generated text is needed:

stream = await client.stream_text(request)
async with stream:
  async for text in stream:
    print(text, end="")

Local request-model validation raises ValueError before native execution. Execution failures derive from ToknError (and remain compatible with RuntimeError). Catch a specific subtype when recovery depends on the cause:

from tokn_requests import APIStatusError, ToknError

try:
  response = await client.send(request)
except APIStatusError as error:
  print(error.status, error.body)
except ToknError as error:
  print(f"request failed: {error}")

Raw endpoint escape hatches

Raw endpoint clients are the exact-wire escape hatch for endpoint- or provider-specific fields. Unlike the friendly generation API, the mapping is sent in the selected endpoint's native shape:

raw = await client.responses.create({
  "model": "gpt-5",
  "input": "Explain this function.",
})

stream = await client.chat.completions.stream({
  "model": "claude-sonnet-4",
  "messages": [{"role": "user", "content": "Hello"}],
})
async with stream:
  payload = b"".join([chunk async for chunk in stream])

Raw stream chunks are transport bytes and do not necessarily align with SSE event or UTF-8 boundaries.

Pass config_path, auth_path, or profile to Client to override the same defaults used by the gateway.

Preparing a PyPI release

Pushing v0.2.3-sdk runs the Python release workflow in .github/workflows/release-python.yml. It builds three stable-ABI wheels, audits their Python symbols, installs the same artifacts on CPython 3.10–3.14, rebuilds the source distribution with locked Cargo dependencies, and uploads the distributions as workflow artifacts. Branch runs build only. Use the successful CI artifacts for the manual publication steps in the SDK release guide. VERSION, the Cargo workspace version, and pyproject.toml must agree.

For optional automated publication in future releases, create the GitHub environment pypi and register a PyPI trusted publisher for:

  • Project: tokn-requests
  • Owner: tokn-ai
  • Repository: tokn
  • Workflow: release-python.yml
  • Environment: pypi

Use a pending publisher in PyPI's account publishing settings if the project does not exist yet. Once workflow dispatch is available, publish=false rehearses a release and publish=true enables the publishing job. This job uses GitHub OIDC. Configure any desired release approval rules on the pypi environment before enabling publishing. See the PyPI trusted publishing setup and publishing documentation.

Use Python 3.12 or newer to run the release packaging helper. To build and test a local wheel from the repository root:

python -m venv tmp/release-python/venv
tmp/release-python/venv/bin/python -m pip install 'maturin==1.14.1' twine
tmp/release-python/venv/bin/maturin build --release --locked \
  --manifest-path bindings/python/Cargo.toml --out tmp/release-python/dist
cargo fetch --locked --manifest-path bindings/python/Cargo.toml
tmp/release-python/venv/bin/python bindings/python/scripts/build_sdist.py \
  --out tmp/release-python/dist
tmp/release-python/venv/bin/python -m twine check --strict tmp/release-python/dist/*
tmp/release-python/venv/bin/python -m pip install tmp/release-python/dist/*.whl
tmp/release-python/venv/bin/python -m unittest discover -s bindings/python/tests

The local wheel targets the current operating system and CPython's Python 3.10 stable ABI. The release workflow builds the portable Linux wheel inside a manylinux2014 container. Check a wheel's package metadata and stable ABI with:

tmp/release-python/venv/bin/python bindings/python/scripts/check_wheel.py \
  tmp/release-python/dist/*.whl
tmp/release-python/venv/bin/python -m pip install abi3audit==0.0.26
tmp/release-python/venv/bin/python -m abi3audit --strict --summary \
  tmp/release-python/dist/*.whl

Use the source-distribution helper above when preparing releases. Maturin removes unrelated Cargo workspace members from an sdist but currently leaves their lockfile entries behind (upstream issue). The helper reconciles the archive's lockfile offline, verifies that every remaining dependency retains its original version and checksum, and checks that Cargo accepts it with --locked. The repository's lockfile is preserved.

Metadata

Release files for tokn-requests 0.2.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for tokn-requests 0.2.3
File Size Uploaded
tokn_requests-0.2.3.tar.gz 765.5 kB Details

Built distributions (wheels)

Table of built distributions (wheels) for tokn-requests 0.2.3
File Interpreter ABI Platform
tokn_requests-0.2.3-cp310-abi3-win_amd64.whl CPython 3.10 abi3 Windows x86-64 Details
tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl CPython 3.10 abi3 Linux glibc 2.17+ x86-64 Details
tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl CPython 3.10 abi3 macOS 11.0+ ARM64 Details

Total release size: 17.7 MB

Release files / tokn_requests-0.2.3.tar.gz

Download URL tokn_requests-0.2.3.tar.gz
Size 765.5 kB
Tags Source
SHA-256 checksum
How to use checksums
f76d46035516c618f2bfe2d940d4edd22a567e2b015eaff400365562d989084f
BLAKE2b-256 checksum
How to use checksums
8cd72f3cfc869165adccabefc908f3c95fadb7ae0b11e6a320d8d0ca3992afd1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / tokn_requests-0.2.3-cp310-abi3-win_amd64.whl

Download URL tokn_requests-0.2.3-cp310-abi3-win_amd64.whl
Size 5.6 MB
Tags CPython 3.10 Windows x86-64 abi3
SHA-256 checksum
How to use checksums
34f934e4c09d26cb0576d2d6818ac00817d5b1a154a3499c9aeef0363ff17158
BLAKE2b-256 checksum
How to use checksums
dfa3dd7afa4b2cbde463dd18304abc73a68d89518d11a680a5b54ed7baa3b082
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl

Download URL tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
Size 6.0 MB
Tags CPython 3.10 Linux glibc 2.17+ x86-64 abi3
SHA-256 checksum
How to use checksums
e95284b38b50ef70b78644eaf311446ddce1018bc4b2082e257f2ed03ead6c37
BLAKE2b-256 checksum
How to use checksums
12737076837ab185c56753596492f570b50af322b125fe5988bdbd829d638deb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release files / tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl

Download URL tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl
Size 5.4 MB
Tags CPython 3.10 abi3 macOS 11.0+ ARM64
SHA-256 checksum
How to use checksums
c1d5429faa4697467c82c5e5ef6330a12506e0c72aa6710482ea2b5b1f0c2932
BLAKE2b-256 checksum
How to use checksums
38a0d6f5bf4381bd2eac883906af756931b613ee8498cb71ec8a67c05fec1add
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.3

Release history Release notifications | RSS feed

This release

0.2.3 This release

4 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page