tokn-requests Python SDK
The Python package embeds the same Rust routing engine as tokn-sdk. It uses
the existing config.toml, config.d, auth.yaml, and auth.d sources and
does not require a gateway process.
Installation
Install the PyPI distribution tokn-requests, then import tokn_requests:
python -m pip install tokn-requests==0.2.3
Release wheels use CPython's Python 3.10 stable ABI (abi3-py310): one
cp310-abi3 wheel per platform supports regular CPython 3.10 and newer.
CI tests the same wheels on CPython 3.10 through 3.14 on Linux x86-64
(glibc 2.17 or newer), macOS 11 or newer on Apple Silicon, and Windows x86-64.
Free-threaded CPython builds require a different ABI and are not covered by
these wheels.
Other platforms can build the source distribution with a current stable Rust
toolchain and a C/C++ compiler. Type annotations and native extension stubs
are included in the installed package.
Friendly generation API
For a one-off request, start with the client-bound builder:
from tokn_requests import Client
client = Client()
response = await (
client.generate("smart")
.system("You are a Python expert.")
.prompt("Explain this function.")
.temperature(0.2)
.send()
)
print(response.text)
Responses normalize text, reasoning, tool calls, token usage, and finish reasons across supported providers.
Common generation controls are available directly on both the client-bound and detached builders:
from tokn_requests import (
ReasoningEffort,
ReasoningMode,
ReasoningSummary,
)
openai_call = (
client.generate("gpt-5")
.prompt("Solve this step by step.")
.max_tokens(2048)
.reasoning_effort(ReasoningEffort.HIGH)
.reasoning_summary(ReasoningSummary.AUTO)
)
llama_call = (
client.generate("local-llama")
.prompt("Compare these implementations.")
.top_p(0.9)
.top_k(40)
.max_tokens(2048)
)
claude_call = (
client.generate("claude-sonnet-4.6")
.prompt("Plan this migration.")
.max_tokens(2048)
.reasoning_mode(ReasoningMode.ADAPTIVE)
.reasoning_effort(ReasoningEffort.HIGH)
)
max_tokens() is a convenience alias for the provider-neutral
max_output_tokens() control. Managed routes serialize that limit as
max_output_tokens for Responses, max_completion_tokens for OpenAI Chat
Completions, and max_tokens for other Chat Completions or Messages routes.
Codex account routes reject an explicit limit because that backend does not
preserve it.
These examples assume the model selectors route to OpenAI Responses,
llama.cpp Chat Completions, and Copilot's Claude Chat Completions fallback,
respectively; use selectors from your own configuration. top_p() is
portable across compatible routes and accepts values from 0 through 1.
Responses supports reasoning effort and summary but not top_k or an
enabled/adaptive mode. Typed top_k is currently supported on llama.cpp Chat
Completions; llama.cpp has no portable reasoning control. Known non-reasoning
models reject typed reasoning locally.
Effort levels are validated against cached upstream model metadata, falling
back to the provider-specific models.dev catalogue. Discovery exposes
x_tokn_router.capabilities.reasoning_efforts: null means unknown and []
means no effort control. Unknown support does not reject an effort value.
DeepSeek V4 Flash advertises low, high, and max; V4 Pro advertises high
and max. DeepSeek thinking also rejects
temperature and top_p, which that backend would ignore. Claude supports
adaptive reasoning on 4.6 and newer models but not a reasoning summary. Manual
Claude reasoning requires ReasoningMode.ENABLED, an explicit max_tokens()
limit, a budget of at least 1024 tokens, and
budget_tokens < max_tokens; manual mode is rejected on 4.7 and newer models,
while adaptive mode is rejected on 4.5 and older models. Claude effort levels
use discovery metadata; sampling compatibility uses the selected model generation.
Explicit controls unsupported by the selected route fail clearly after routing
instead of being silently dropped or reinterpreted.
passthrough and switch profiles preserve the generated Responses payload
verbatim, so they reject typed top_k and reasoning controls that would
require post-route lowering. Use an exact, route, or fuzzy profile for
the provider-neutral control API.
The snippets below continue using the client created above. Snippets after
the detached-request example also reuse its request.
Build an owned request when it needs to be serialized, transformed, queued, or reused independently of a client:
from tokn_requests import GenerateRequest
request = (
GenerateRequest.builder("smart")
.prompt("Explain this function.")
.temperature(0.2)
.build()
)
serialized = request.to_json()
request = GenerateRequest.from_json(serialized)
request = request.with_changes(max_output_tokens=128)
response = await client.send(request)
As an alternative to client.send(request), use
await request.bind(client).send() when fluent binding is more convenient.
Semantic streaming returns typed events:
from tokn_requests import Completed, TextDelta
stream = await client.generate("smart").prompt("Write a haiku.").stream()
async with stream:
async for event in stream:
if isinstance(event, TextDelta):
print(event.text, end="")
elif isinstance(event, Completed):
print(f"\nfinish reason: {event.finish_reason}")
Use stream_text() when only generated text is needed:
stream = await client.stream_text(request)
async with stream:
async for text in stream:
print(text, end="")
Local request-model validation raises ValueError before native execution.
Execution failures derive from ToknError (and remain compatible with
RuntimeError). Catch a specific subtype when recovery depends on the cause:
from tokn_requests import APIStatusError, ToknError
try:
response = await client.send(request)
except APIStatusError as error:
print(error.status, error.body)
except ToknError as error:
print(f"request failed: {error}")
Raw endpoint escape hatches
Raw endpoint clients are the exact-wire escape hatch for endpoint- or provider-specific fields. Unlike the friendly generation API, the mapping is sent in the selected endpoint's native shape:
raw = await client.responses.create({
"model": "gpt-5",
"input": "Explain this function.",
})
stream = await client.chat.completions.stream({
"model": "claude-sonnet-4",
"messages": [{"role": "user", "content": "Hello"}],
})
async with stream:
payload = b"".join([chunk async for chunk in stream])
Raw stream chunks are transport bytes and do not necessarily align with SSE event or UTF-8 boundaries.
Pass config_path, auth_path, or profile to Client to override the same
defaults used by the gateway.
Preparing a PyPI release
Pushing v0.2.3-sdk runs the Python release workflow in
.github/workflows/release-python.yml. It builds three stable-ABI wheels,
audits their Python symbols, installs the same artifacts on CPython 3.10–3.14,
rebuilds the source distribution with locked Cargo dependencies, and uploads
the distributions as workflow artifacts. Branch runs build only. Use the
successful CI artifacts for the manual publication steps in the
SDK release guide. VERSION, the Cargo workspace
version, and pyproject.toml must agree.
For optional automated publication in future releases, create the GitHub
environment pypi and register a PyPI trusted publisher for:
- Project:
tokn-requests - Owner:
tokn-ai - Repository:
tokn - Workflow:
release-python.yml - Environment:
pypi
Use a pending publisher in PyPI's account publishing settings if the project
does not exist yet. Once workflow dispatch is available, publish=false
rehearses a release and publish=true enables the publishing job. This job
uses GitHub OIDC. Configure any desired release approval rules on the pypi
environment before enabling publishing. See the
PyPI trusted publishing setup
and publishing documentation.
Use Python 3.12 or newer to run the release packaging helper. To build and test a local wheel from the repository root:
python -m venv tmp/release-python/venv
tmp/release-python/venv/bin/python -m pip install 'maturin==1.14.1' twine
tmp/release-python/venv/bin/maturin build --release --locked \
--manifest-path bindings/python/Cargo.toml --out tmp/release-python/dist
cargo fetch --locked --manifest-path bindings/python/Cargo.toml
tmp/release-python/venv/bin/python bindings/python/scripts/build_sdist.py \
--out tmp/release-python/dist
tmp/release-python/venv/bin/python -m twine check --strict tmp/release-python/dist/*
tmp/release-python/venv/bin/python -m pip install tmp/release-python/dist/*.whl
tmp/release-python/venv/bin/python -m unittest discover -s bindings/python/tests
The local wheel targets the current operating system and CPython's Python 3.10 stable ABI. The release workflow builds the portable Linux wheel inside a manylinux2014 container. Check a wheel's package metadata and stable ABI with:
tmp/release-python/venv/bin/python bindings/python/scripts/check_wheel.py \
tmp/release-python/dist/*.whl
tmp/release-python/venv/bin/python -m pip install abi3audit==0.0.26
tmp/release-python/venv/bin/python -m abi3audit --strict --summary \
tmp/release-python/dist/*.whl
Use the source-distribution helper above when preparing releases. Maturin
removes unrelated Cargo workspace members from an sdist but currently leaves
their lockfile entries behind (upstream issue).
The helper reconciles the archive's lockfile offline, verifies that every
remaining dependency retains its original version and checksum, and checks
that Cargo accepts it with --locked. The repository's lockfile is preserved.
Metadata
Release files for tokn-requests 0.2.3
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tokn_requests-0.2.3.tar.gz | 765.5 kB | Details |
Built distributions (wheels)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tokn_requests-0.2.3-cp310-abi3-win_amd64.whl | CPython 3.10 | abi3 | Windows x86-64 | Details |
| tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl | CPython 3.10 | abi3 | Linux glibc 2.17+ x86-64 | Details |
| tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl | CPython 3.10 | abi3 | macOS 11.0+ ARM64 | Details |
Total release size: 17.7 MB
Release files / tokn_requests-0.2.3.tar.gz
| Download URL | tokn_requests-0.2.3.tar.gz |
|---|---|
| Size | 765.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f76d46035516c618f2bfe2d940d4edd22a567e2b015eaff400365562d989084f
|
|
BLAKE2b-256 checksum How to use checksums |
8cd72f3cfc869165adccabefc908f3c95fadb7ae0b11e6a320d8d0ca3992afd1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / tokn_requests-0.2.3-cp310-abi3-win_amd64.whl
| Download URL | tokn_requests-0.2.3-cp310-abi3-win_amd64.whl |
|---|---|
| Size | 5.6 MB |
| Tags | CPython 3.10 Windows x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
34f934e4c09d26cb0576d2d6818ac00817d5b1a154a3499c9aeef0363ff17158
|
|
BLAKE2b-256 checksum How to use checksums |
dfa3dd7afa4b2cbde463dd18304abc73a68d89518d11a680a5b54ed7baa3b082
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl
| Download URL | tokn_requests-0.2.3-cp310-abi3-manylinux_2_17_x86_64.manylinux2014_x86_64.whl |
|---|---|
| Size | 6.0 MB |
| Tags | CPython 3.10 Linux glibc 2.17+ x86-64 abi3 |
|
SHA-256 checksum How to use checksums |
e95284b38b50ef70b78644eaf311446ddce1018bc4b2082e257f2ed03ead6c37
|
|
BLAKE2b-256 checksum How to use checksums |
12737076837ab185c56753596492f570b50af322b125fe5988bdbd829d638deb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|
Release files / tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl
| Download URL | tokn_requests-0.2.3-cp310-abi3-macosx_11_0_arm64.whl |
|---|---|
| Size | 5.4 MB |
| Tags | CPython 3.10 abi3 macOS 11.0+ ARM64 |
|
SHA-256 checksum How to use checksums |
c1d5429faa4697467c82c5e5ef6330a12506e0c72aa6710482ea2b5b1f0c2932
|
|
BLAKE2b-256 checksum How to use checksums |
38a0d6f5bf4381bd2eac883906af756931b613ee8498cb71ec8a67c05fec1add
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.12.3
|