tai-sdk
Official Python SDK for the TAI Assistant API (contract v1).
- Sync (
TAI) and async (AsyncTAI) clients — both are context managers. - Assistants, threads, messages, stateless chat, models and usage.
- Incremental SSE streaming with typed events.
- Zero-config endpoints: the API base URL is discovered from a GitHub Pages discovery document and cached for 12 hours, so a rotating ngrok tunnel does not need a code change.
- One runtime dependency:
httpx. Python 3.9+.
Install
pip install tai-sdk
From a checkout:
pip install -e ./sdk # or: pip install ./sdk
pip install -e "./sdk[dev]" # plus pytest + pytest-asyncio
Quickstart
export TAI_API_KEY="sk-tai-..." # developer dashboard → API keys
from tai_sdk import TAI
with TAI() as client:
print(client.base_url, client.base_url_source) # what it resolved, and why
reply = client.chat.create(
model="tfmf",
messages=[{"role": "user", "content": "Explain tail recursion."}],
)
print(reply.content) # attribute access
print(reply.usage.output_tokens) # typed nested object
print(reply.raw) # the untouched JSON payload
Async is the same shape, awaited:
import asyncio
from tai_sdk import AsyncTAI
async def main():
async with AsyncTAI() as client:
page = await client.models.list()
for model in page.data:
print(model.id, model.context_window)
async for event in client.chat.stream(
model=page.data[0].id,
messages=[{"role": "user", "content": "Count to five."}],
):
if event.event == "message.delta":
print(event.delta, end="", flush=True)
asyncio.run(main())
Runnable examples live in examples/:
quickstart.py, streaming.py,
async_chat.py,
assistant_thread.py.
Configuration
Base URL resolution
Checked in order; the first usable value wins. client.base_url_source tells
you which one was used ("argument", "env", "config", "discovery" or
"default").
| # | Source | Notes |
|---|---|---|
| 1 | TAI(base_url="https://…") |
always wins |
| 2 | TAI_BASE_URL environment variable |
|
| 3 | ~/.tai/config.json |
{"base_url": "https://…"} |
| 4 | Discovery document | see below; 3 s timeout, cached 12 h |
| 5 | https://api.tai-research.dev |
last-resort default |
A value that is not a URL (no scheme://) is skipped, not used. Any trailing
slashes are stripped.
Endpoint discovery (ngrok tunnels)
GitHub Pages cannot proxy requests, so it hosts a discovery document that simply points at the current tunnel:
GET https://ltyleo.github.io/platform/api-endpoint.json
{
"object": "endpoint",
"base_url": "https://xxxx.ngrok-free.app",
"updated_at": "2026-09-30T15:04:05+00:00",
"note": "TAI Developer Platform API. Update this file when the tunnel changes."
}
- Short timeout (3 s). Any failure (offline, 404, bad JSON) is non-fatal — the
SDK falls through to the next source, ending at
https://api.tai-research.dev. - The result is cached in
~/.tai/endpoint-cache.jsonfor 12 hours, so the SDK works offline and does not hit Pages on every client construction. TAI_DISCOVERY_URLpoints the SDK at a different document.TAI(discover=False)disables discovery entirely.TAI_DISCOVERY=0disables it for the whole process.- Free ngrok tunnels inject an interstitial HTML page unless they see
ngrok-skip-browser-warning: true— the SDK sends that header on every request, so you never have to think about it.
client = TAI(discover=False, base_url="http://127.0.0.1:8000") # local server
API key resolution
TAI(api_key="sk-tai-…")TAI_API_KEYenvironment variable
If neither is set, the first request raises TAIError before touching the
network, with a message telling you to set TAI_API_KEY. (The server also
enforces this, with 401 missing_api_key / 401 invalid_api_key.)
Timeouts and retries
client = TAI(timeout=60.0, max_retries=4)
client = TAI(timeout=httpx.Timeout(connect=5.0, read=60.0, write=5.0, pool=5.0))
timeoutdefaults to 30 s connect/read for normal requests. Streams keep the connect timeout but disable the read timeout, so a slow generation is never cut off mid-reply.max_retriesdefaults to 2.
| Failure | Retried? |
|---|---|
Connection error, GET/DELETE/HEAD/OPTIONS/PUT |
yes |
Connection error, POST |
yes — it provably never left the client |
| Timeout (connect or read) | no — the server may have already processed it |
HTTP 429, 5xx, idempotent method |
yes, honouring Retry-After |
HTTP 429, 5xx, POST |
no — a generation may already have been recorded |
HTTP 4xx (other than 429) |
no |
Backoff is exponential (0.5 s, 1 s, 2 s … capped at 60 s) with full jitter, and
Retry-After wins when the server sends it.
Other options
TAI(
api_key=None, # or TAI_API_KEY
base_url=None, # see the table above
timeout=30.0,
max_retries=2,
default_headers={"X-Trace-Id": "abc"}, # merged into every request
discover=True,
discovery_url=None, # or TAI_DISCOVERY_URL
discovery_timeout=3.0,
http_client=None, # bring your own httpx.Client
transport=None, # bring your own httpx transport
)
http_client= / transport= are how the test suite injects
httpx.MockTransport; they are also useful for custom proxies, TLS settings
and connection pooling. When you pass your own client, the SDK does not
close it for you.
API tour
All paths are relative to the resolved base URL and prefixed with /api/v1.
Timestamps are UTC ISO-8601 (2026-09-30T15:04:05+00:00). Object ids are
prefixed: asst_, thrd_, msg_, chat_.
Assistants
assistant = client.assistants.create(
model="tfmf", # required, must be live
name="Study Buddy", # ≤ 80 chars
instructions="You are a patient tutor.", # ≤ 8000 chars
metadata={"course": "algorithms"}, # ≤ 16 flat string values
)
page = client.assistants.list(limit=20, after=None) # newest first
page.data, page.has_more, page.first_id, page.last_id
client.assistants.get(assistant.id)
client.assistants.update(assistant.id, name="Tutor", instructions="Be brief.")
client.assistants.delete(assistant.id) # -> None (HTTP 204)
update() sends only the fields you pass, so you can PATCH one field.
Passing None explicitly sends a JSON null.
Threads
thread = client.threads.create(assistant_id=assistant.id, title="Factorial help")
# `model=` instead of `assistant_id=` is the other option; if you set neither,
# you must pass `model` when you send a message.
client.threads.get(thread.id)
client.threads.list(limit=20)
client.threads.update(thread.id, title="Renamed") # title / assistant_id / model / metadata
client.threads.update(thread.id, assistant_id=None) # clear the assistant
client.threads.delete(thread.id) # messages cascade
Messages (client.threads.messages)
One POST appends your message and generates the reply, returning both plus usage — no second round trip:
exchange = client.threads.messages.create(
thread.id,
content="How do I reverse a list?",
temperature=0.7, # 0.0–2.0
max_output_tokens=512, # 1–8192
)
exchange.user_message.id # msg_…
exchange.assistant_message.content
exchange.assistant_message.model
exchange.assistant_message.usage.output_tokens
exchange.usage.cost_cny
exchange.content # convenience: assistant_message.content
Reading them back (order defaults to desc):
page = client.threads.messages.list(thread.id, limit=20, after=None, order="asc")
for message in page.data:
print(message.role, message.content)
client.messages is a shortcut for client.threads.messages.
Stateless chat
Nothing is persisted; thread_id is null in stream events.
completion = client.chat.create(
model="tfmf", # required unless you pass assistant_id
messages=[ # 1–100 items
{"role": "system", "content": "You are terse."},
{"role": "user", "content": "Hi"},
],
)
completion.id, completion.model, completion.output, completion.usage, completion.created_at
completion.content # convenience: output["content"]
# An assistant supplies both model and instructions:
client.chat.create(assistant_id="asst_…", messages=[{"role": "user", "content": "Hi"}])
Messages are accepted as plain dicts, (role, content) tuples, or
ChatMessage objects.
Models and usage
for model in client.models.list().data:
model.id, model.name, model.live, model.context_window
model.input_cny_per_1m, model.output_cny_per_1m
client.models.retrieve("tfmf") # client-side filter of list(); v1 has no single-model route
summary = client.usage.retrieve() # the dashboard usage_summary() payload
summary.raw["total_input_tokens"]
Streaming
When stream is true the server sends text/event-stream:
event: message.start
data: {"id":"msg_…","model":"tfmf","thread_id":"thrd_…"}
event: message.delta
data: {"delta":"Hello"}
event: message.done
data: {"id":"msg_…","content":"Hello there","usage":{…}}
event: error
data: {"code":"backend_unavailable","message":"…"}
The SDK parses this incrementally (splits inside a JSON payload or even inside a multi-byte UTF-8 character are fine) and yields typed events:
| Event | Object | Useful attributes |
|---|---|---|
message.start |
MessageStart |
.id, .model, .thread_id (None for /chat) |
message.delta |
MessageDelta |
.delta |
message.done |
MessageDone |
.id, .content, .usage |
error |
raises | APIError subclass with .code, .message |
stream = client.chat.stream(
model="tfmf",
messages=[{"role": "user", "content": "Tell me a story."}],
)
with stream: # optional, but explicit
for event in stream:
if event.event == "message.delta":
print(event.delta, end="", flush=True)
elif event.event == "message.done":
print("\n", event.usage)
Async:
async for event in client.chat.stream(model="tfmf", messages=[...]):
...
Thread streams persist the assistant message and record usage exactly once:
for event in client.threads.messages.create(thread.id, content="Hi", stream=True):
...
Cleanup is guaranteed. The HTTP response is closed when the iterator is
exhausted, when it is closed (.close() / aclose() / with), when the caller
breaks out of the loop, when an exception propagates, and when the iterator is
garbage-collected. The request is opened eagerly, so HTTP errors (401, 404,
422, …) raise from the stream(...) call itself rather than at first
iteration — with the async client, that means when you start iterating.
Errors
Every failure is a TAIError. Anything from a server response is also an
APIError carrying code, message, param, status_code, request_id
(when the server sent one) and the decoded body.
from tai_sdk import errors
try:
client.assistants.create(model="does-not-exist")
except errors.ModelNotFoundError as exc: # 404 model_not_found
print(exc.code, exc.message, exc.status_code)
except errors.InvalidRequestError as exc: # 422 invalid_request
print("bad field:", exc.param)
except errors.RateLimitError as exc: # 429 rate_limited
print("retry after", exc.retry_after, exc.request_id)
except errors.APIConnectionError as exc: # never reached the server
print("tried:", exc.base_url)
except errors.APIError as exc: # any other non-2xx
print(exc.status_code, exc.code, exc.message)
except errors.TAIError as exc: # anything the SDK raises
print(exc)
| Code | HTTP | Exception |
|---|---|---|
missing_api_key, invalid_api_key |
401 | AuthenticationError |
account_disabled |
403 | PermissionDeniedError |
insufficient_balance |
402 | InsufficientBalanceError |
not_found |
404 | NotFoundError |
model_not_found |
404 | ModelNotFoundError (a NotFoundError) |
model_not_available |
409 | ModelNotAvailableError (a ConflictError) |
invalid_request |
422 | InvalidRequestError |
rate_limited |
429 | RateLimitError |
backend_unavailable |
502 | BackendUnavailableError |
stream_aborted |
499 | StreamAbortedError |
| anything unmapped | any | UnknownAPIError (keeps the server's .code) |
APIConnectionError and APITimeoutError are TransportErrors whose message
includes the resolved base URL, so you can see exactly which tunnel was
tried.
Response objects
Small hand-written dataclasses (no pydantic). Every one supports attribute
access, .raw (the untouched dict), repr() and .to_dict():
assistant.id # "asst_…"
assistant.raw # the full JSON payload from the server
repr(assistant) # Assistant(id='asst_…', object='assistant', …)
page = client.assistants.list()
len(page), list(page), page[0] # list envelopes are sequences
Nested objects are typed too (message.usage.output_tokens,
exchange.assistant_message.model). Unknown fields are never dropped — they
stay in .raw, so a server-side addition cannot break your code.
ChatMessage is the one object that is input, not output:
from tai_sdk import ChatMessage
ChatMessage(role="user", content="Hi")
Low-level escape hatch
If the contract grows a route before the SDK does:
payload = client.request("GET", "/api/v1/models") # -> raw dict
stream = client.stream("POST", "/api/v1/chat", json_body={...}) # -> events
await client.request("GET", "/api/v1/models") # AsyncTAI
async for event in await client.astream("POST", "/api/v1/chat", json_body={...}):
...
Resource lifecycle
with TAI() as client: # closes the connection pool on exit
...
client = TAI()
try:
...
finally:
client.close() # idempotent
async with AsyncTAI() as client: # aclose() on exit
...
await client.aclose()
Using a closed client raises RuntimeError.
Development
python -m pytest tests -q # no server needed: httpx.MockTransport
python -m build # -> dist/tai_sdk-0.1.0-py3-none-any.whl
The suite covers URL and header construction, base-URL/API-key resolution
order, discovery caching, every resource happy path, SSE parsing (including a
split chunk boundary and a mid-stream error event), error mapping for
401/404/409/422/429, retry-then-succeed, and stream cleanup on early exit.
License
MIT © TAI Research — see LICENSE.
Metadata
Release files for tai-sdk 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| tai_sdk-0.1.0.tar.gz | 52.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| tai_sdk-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 89.9 kB
Release files / tai_sdk-0.1.0.tar.gz
| Download URL | tai_sdk-0.1.0.tar.gz |
|---|---|
| Size | 52.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
c483085d25ddf8f211821ef0670d8f01aded1af225041a53654e1a0d5ec3caa8
|
|
BLAKE2b-256 checksum How to use checksums |
f5a57d1898429da394a9697e720e58eb0deed219478ff6474eb39f4914cd32a4
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|
Release files / tai_sdk-0.1.0-py3-none-any.whl
| Download URL | tai_sdk-0.1.0-py3-none-any.whl |
|---|---|
| Size | 37.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
e35cb94e29a69414f2fbea582645644e37b702694755b0d097c155e5a701b371
|
|
BLAKE2b-256 checksum How to use checksums |
b78908b662969c7d654be5612d06e9eb89957a4e3eba56d57f12c46aaee45684
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|