Skip to main content

nvoken Python SDK

An Invocation is one durable turn by a deliberately created, tenant-scoped Agent. An Agent binds your agent_key to one App-owned, versioned Agent Definition; a Session is one conversation with that Agent.

The package has three deliberate levels:

  • Agent is the ordinary workflow facade: text, run, invoke, stream, and locally serialized bound Sessions.
  • Client and InvocationHandle expose durable operations, transcript drains, provider-key lifecycle, iterators, configurable waits, and resumable streams.
  • nvoken_generated is the complete generated Runtime transport and raw escape hatch.
python -m pip install nvoken
NVOKEN_BASE_URL=http://localhost:8080 NVOKEN_API_KEY=... \
  python examples/quickstart.py

The async facade provides durable handles, replay-safe retries, typed errors, cursor iterators, resumable SSE, composed result reads (result, list_messages, output_text), and callback verification. Session-scoped messages use Client.list_session_messages.

List or read the Agent record without admitting work:

agents = await client.list_agents(agent_key="support")
record = await client.get_agent(agents.items[0].id)

AgentResource is that record: tenant, key, display name, Definition binding, optional revision pin, lifecycle timestamps, and archive state. Agent is the object that runs its turns, and it corresponds to the same row — declare one from the keys you already own and it creates its record on first use:

agent = client.agent(AgentOptions(
    tenant_key=user_id,
    agent_key="support",
    definition_key="support",   # the Definition this instance follows
    tools=(lookup_order,),      # this process's handlers; never on the server
))

An Agent's identity and configuration live on the server; its tool handlers are supplied by whichever process runs the turn. await agent.ensure() creates the record at a moment you choose instead of on first use, and never mutates: the same keys and Definition resolve onto what exists, a different Definition is agent_key_conflict, an archived record is agent_archived, and a declared pinned_revision the record does not follow is refused. agent.resource and agent.id report the record once it is known, and agent.with_tools() attaches handlers to an Agent read back from the server.

Opt into the fixed guarded public-web reader with fetch_tool():

from nvoken import AgentDefinition, AgentOptions, Model, fetch_tool

definition = await client.create_agent_definition(
    "research",
    "Research",
    AgentDefinition(
        instructions="Use nvoken_fetch for public URLs, then summarize the source.",
        model=Model(provider="anthropic", id="claude-sonnet-5"),
        tools=(fetch_tool(),),
    ),
)
options = AgentOptions(agent_key="research", definition_key=definition.definition_key)

The Runtime accepts only {"name":"nvoken_fetch","mode":"builtin"}. It owns public-address checks, up to five guarded redirects, one transient retry, HTML-to-Markdown conversion, and the ten-second and 64 KiB limits. Run python examples/fetch.py to summarize NVOKEN_FETCH_URL.

Use an Agent for the common path:

agent = client.agent(AgentOptions(
    agent_key="support",
    definition_key="support",
))

print(await agent.text("Why was I charged twice?"))
continued = agent.session(session_key="customer-123")
print(await continued.text("What should I do next?"))

A bound Session serializes admission only within that local binding. The Runtime remains authoritative across processes and rejects a second nonterminal turn. Agent operations dispatch configured host-tool handlers. If a waiting call has no handler, the Agent cancels before raising MissingToolHandlerError by default; set InvocationOptions(leave_waiting_on_missing_handler=True) only when another worker deliberately owns it. NoOutputTextError.result_kind distinguishes structured, tool-only, and empty completions.

For an intentional replace/regenerate action, use a new idempotency key and the typed option:

handle = await agent.invoke(
    "Try that answer again.",
    options=InvocationOptions(
        idempotency_key="customer-123:regenerate-2",
        if_active="supersede",
        session_key="customer-123",
    ),
)

Omission or "reject" preserves the default conflict response. Low-level callers set the same policy on InvokeRequest.if_active.

if_active="interrupt" is the keep-the-work variant: the active Invocation stops at its next execution seam and settles completed with stop_reason "interrupted", so the replacement turn builds on what it already produced. await handle.interrupt() asks for the same graceful stop without admitting a replacement, and Invocation.stop_reason names why any turn ended.

A turn can also stop without ending: "incomplete" means the Runtime enforced a budget at a seam, with stop_reason naming which one. It is terminal — the wait helpers stop there — and its work is kept, so treat it as an unfinished answer rather than an error. SessionMessage.phase says which assistant message was the reply: "final_answer" on the one that ended a completed turn, "commentary" on everything else, so an incomplete turn has none.

InvocationOptions(timeout=...) is one overall local deadline. Cancelling the calling task still raises native asyncio.CancelledError; it does not imply a durable Runtime cancellation. Call handle.cancel() when that is intended.

Recovery reads accept a status union, and a known Invocation can stream only durable frames:

page = await client.list_invocations(
    status=["queued", "running", "waiting"],
)
async for event in handle.events(deltas=False):
    ...

Equivalent status sets share cursor identity regardless of input order. Session get/list models expose typed nullable usage, computed from durable Invocation usage as a convenience estimate rather than a billing ledger.

Install restart-stable compaction on a new or existing Session:

from nvoken import ContextCompaction, SessionOptions

request = InvokeRequest(
    agent_key="support",
    session_key="support:123",
    session_options=SessionOptions(
        compaction=ContextCompaction(trigger_tokens="auto"),
    ),
    input="hello",
)

Use an integer trigger and optional same-provider model for explicit policy. A Session without a policy accepts late opt-in; once installed, the policy is immutable. Supplied options on an existing Session must equal stored values or admission returns session_options_conflict.

Summary usage appears in Session usage rather than Invocation usage. Read applied and fell-through diagnostics with client.list_session_compactions(session_id).

InvocationOptions.metadata correlates a turn with your own records from the Agent binding. It is part of the admitted input, so it is immutable and material to idempotency: a replay carrying different metadata conflicts rather than updating it. That is why it is per-call rather than an AgentOptions default.

Pass a stored or one-turn provider key directly through InvokeRequest:

request = InvokeRequest(
    agent_key="support",
    input="hello",
    provider_keys=(
        ProviderKeySelection(
            provider="openai",
            source="caller_ephemeral",
            api_key=provider_key,
        ),
    ),
)

Stored sources are app_byok, tenant_byok, and platform and do not accept an api_key. Client.stream_session(session_id, reducer, consume) follows the Session until its task is cancelled; a terminal turn does not end the Session stream. For catch-up reads, use get_transcript_page when checkpointing each page or drain_transcript to consume one fixed cut.

Discover models through the same async facade:

catalog = await client.list_models(provider="openai")
selected = await client.get_model(
    Model(provider="openai", id=catalog.items[0].id)
)
print(selected.cataloged, selected.pricing.status)

The list is curated discovery metadata, not proof of provider-account access. Exact inspection also accepts uncataloged IDs.

Set an explicit portable temperature on the Agent Definition or a safe per-turn override:

from nvoken import AgentDefinitionOverrides, InvokeRequest, Model, Sampling

request = InvokeRequest(
    agent_key="support",
    input="hello",
    overrides=AgentDefinitionOverrides(
        model=Model(provider="anthropic", id="claude-haiku-4-5"),
        sampling=Sampling(temperature=0),
    ),
)

Omit sampling to preserve the provider default. Check selected.controls.sampling.temperature first; missing controls are unknown, and unsupported or unknown selections fail before durable admission. The portable range is [0,1]. top_p and stop sequences are intentionally absent; limits.max_output_tokens is the output guardrail.

Reasoning is typed and fail closed:

from nvoken import AgentDefinitionOverrides, InvokeRequest, Model, Reasoning

request = InvokeRequest(
    agent_key="support",
    input="hello",
    overrides=AgentDefinitionOverrides(
        model=Model(provider="anthropic", id="claude-opus-5"),
        reasoning=Reasoning(effort="high"),
    ),
)

Check selected.controls.reasoning first. budget_tokens requires a larger explicit limits.max_output_tokens. Omission preserves provider defaults; unsupported values and combinations are rejected without aliasing. OpenAI reasoning remains unavailable until its complete continuation representation is durable.

Structured-output schema preflight

Client.invoke and Agent operations call preflight_output_schema(schema) before transport when InvokeRequest.overrides.output_schema is present. Rejection is an NvokenError with code schema_preflight_failed; its safe details contain the portable issue code, RFC 6901 path, and optional keyword. A successful local check means eligible for admission. Generated APIs reached through client.raw() still rely on the authoritative Runtime check.

Define and instantiate an Agent

Every turn runs against an App-owned, versioned Agent Definition. Create a tenant-scoped Agent instance that binds to the template before admitting work:

resource = await client.create_agent_definition(
    "support",
    "Support",
    AgentDefinition(
        instructions="Help with billing questions.",
        model=Model(provider="anthropic", id="claude-sonnet-5"),
    ),
    idempotency_key="support-definition-v1",
)
agent = client.agent(AgentOptions(
    agent_key="support",
    definition_key=resource.definition_key,
))

handle = await agent.invoke("Why was I charged twice?")

Creating a Definition starts no turn. It has an immutable definition_key, a stable ID, and an increasing revision; get_agent_definition_revision() reads historical revisions. Updating a Definition does not rewrite an Agent's binding. An Agent or Session may pin a revision, while an Invocation may select one revision for that turn. Safe overrides cover model, sampling, reasoning, tool choice, limits, and output schema; they cannot expand tools, data access, memory authority, or instructions. Host tool handlers remain local to the SDK facade.

Record changing application state

Keep instructions static. Product state that changes between turns — a board snapshot, customer facts, the current policy — belongs in context:

answer = await agent.text(
    "Can I refund the duplicate charge?",
    InvocationOptions(
        session_key="ticket-483",
        context=(
            ContextItem(name="customer", tier="contextual", content="plan: pro"),
            ContextItem(
                name="refund-policy",
                tier="operator",
                content="Self-serve refunds cap at 50 USD",
            ),
        ),
    ),
)

A name is a stable identity. Send it once and nvoken records it as a leading message the model reads as app-customer; omit that reserved prefix here. Send the same name again only when its value changes — a byte-identical resend is accepted but adds no message, so a stateless host may resend its whole snapshot every turn and get the same transcript as a host that tracks changes.

Use contextual for conversation-adjacent facts and operator for policy or other application-authoritative state. Context is Session history, not an Agent Definition field: it never changes definition_id, and later turns keep sending it to the model even when you omit it. That is what keeps the prompt prefix stable enough for provider caching, which rewriting the same state into instructions would break on every turn.

The list is order-sensitive and part of idempotency, so a replay that reorders or edits an item conflicts rather than updating it. A request accepts at most eight items, 8 KiB per item, and 16 KiB in total; the SDK checks all three before the request leaves the process. A Session may accumulate at most 16 distinct names, which only the service can check. Retire a name by superseding it with a short current value such as "ticket: closed".

Remote MCP tools

Use the handwritten declaration for discovery. Store the same declaration on the Agent Definition, then pass only its one-turn secret headers at admission:

server = MCPServer(
    name="support",
    url="https://mcp.example.com/rpc",
    allowed_tools=("lookup_order",),
    timeouts=MCPTimeouts(discovery_seconds=10, call_seconds=30),
)
headers = {"Authorization": f"Bearer {mcp_token}"}

catalog = await client.list_mcp_tools(server, headers)
request = InvokeRequest(
    agent_key="support",
    input="hello",
    mcp_server_headers=(MCPServerHeaders(name="support", headers=headers),),
)

The declaration carries no secrets. An Agent Definition may be reused across turns, so authentication headers travel per Invocation in mcp_server_headers, keyed to the server name. They are hidden from dataclass representation, are one-Invocation secret material, and never appear in durable Agent Definitions or public recovery surfaces.

Callback tools

A callback tool runs on an HTTPS endpoint nvoken posts to. Verify the signed delivery with verify_callback, then answer with one of two replies. callback_result(content, is_error=False) settles the ToolCall inline and the turn resumes as soon as nvoken records the reply. acknowledge_callback() returns 202 with no body instead: it accepts the delivery without settling the call, for work that will outlive this tool's reply deadline — its declared timeout_seconds, or the App's default when it declares none. Settle it later with client.submit_tool_results, reusing the delivery's ToolCall ID.

Every delivery names its tool inside the signed body, so one endpoint can serve several tools: dispatch on verified.tool_name rather than on a path suffix nothing signs.

Acknowledging trades away the fail-loud guarantee. nvoken marks an unacknowledged delivery failed once its retries are exhausted, so the turn always moves on. An acknowledged call instead waits under your responsibility, bounded only by the Invocation's limits.waiting_timeout_seconds. Acknowledge only when something durable will settle the call. Such a call appears in a waiting Invocation's pending calls the same way a host call does, so answerable_tool_calls includes it. host_tool_calls is the narrower set you must run yourself — answerable and mode host — and an Agent dispatches exactly that, whatever its own definition declares.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

nvoken-0.21.0.tar.gz (341.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

nvoken-0.21.0-py3-none-any.whl (1.5 MB view details)

Uploaded Python 3

File details

Details for the file nvoken-0.21.0.tar.gz.

File metadata

  • Download URL: nvoken-0.21.0.tar.gz
  • Upload date:
  • Size: 341.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for nvoken-0.21.0.tar.gz
Algorithm Hash digest
SHA256 397e69d1a65f917848ae405a6c9bb6df3d3bac9412aec2602354c9f7522b9290
MD5 ea28e28758ec6739f0546c37f8ba1049
BLAKE2b-256 17a42b34446b409b6749d819b091df389c216a1aeaf58af54e59c9e70cdfad49

See more details on using hashes here.

Provenance

The following attestation bundles were made for nvoken-0.21.0.tar.gz:

Publisher: release-pypi.yml on deepnoodle-ai/nvoken

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file nvoken-0.21.0-py3-none-any.whl.

File metadata

  • Download URL: nvoken-0.21.0-py3-none-any.whl
  • Upload date:
  • Size: 1.5 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for nvoken-0.21.0-py3-none-any.whl
Algorithm Hash digest
SHA256 e0031a8c3cf202c41faeee628927b6112987decc17ffdcb4bdad803c4155b419
MD5 120b04b2172de132f2a2cd521a626af5
BLAKE2b-256 c9e8bc1732f1f5a2d8e2254906a2c53a33bcccb205b89ee5864d0bbfc925201e

See more details on using hashes here.

Provenance

The following attestation bundles were made for nvoken-0.21.0-py3-none-any.whl:

Publisher: release-pypi.yml on deepnoodle-ai/nvoken

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.35.0

2 files

0.34.0

2 files

0.33.0

2 files

0.32.0

2 files

0.31.0

2 files

0.30.0

2 files

0.29.0

2 files

0.28.0

2 files

0.27.0

2 files

0.26.0

2 files

0.25.0

2 files

0.24.0

2 files

0.23.0

2 files

0.22.0

2 files

This release

0.21.0 This release

2 files

0.20.0

2 files

0.19.0

2 files

0.18.0

2 files

0.17.0

2 files

0.16.0

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page