Skip to main content

A Python runtime for production AI agents with durable sessions, controlled tool execution, context management, reviewed knowledge, recovery, evals, and replay.

Project description

Cayu

Cayu is a production agent runtime for building and operating AI agents in Python.

A harness turns a model into an agent by supplying its context, tools, permissions, and execution logic. Cayu gives applications control of the full agent execution lifecycle: how context is assembled, models and tools are invoked, where agent code runs, how state is persisted, authority is governed, failures are recovered, and behavior is observed and evaluated.

Cayu provides durable agent-runtime primitives including sessions, task dispatch, leased workers, resumable workflow steps, approvals, and recovery. Applications can use them directly without a separate workflow engine.

Applications retain control of their UI, authentication, domain logic, and business workflows.

Cayu is designed for agents that do consequential or long-running work. You compose its runtime primitives directly in your application.

Why we built Cayu

Cayu was extracted from the production runtime behind an agent-operated software factory that built and deployed thousands of business applications. Specialized agents worked together as the AI SRE, AI product manager, AI coder, and FDE assistant behind that delivery process.

We began by building agents with SDKs and frameworks including the Claude Agent SDK, Mastra, and LangGraph. They helped us implement the model-and-tool loop quickly. Production quality required deeper control of the loop itself: context assembly, model and tool invocation, output validation, and failure handling.

Production quality also depended on everything around that loop: scheduling, durable state, credentials, execution environments, human intervention, recovery, cost attribution, and evaluation. Cayu gives applications control of that entire agent execution lifecycle.

Why Cayu

Agent prototypes are easy to start. Production failures happen at the boundaries:

  • a process dies after a side effect but before state is recorded;
  • a model requests a valid tool with the wrong authority;
  • a run needs human input or approval halfway through;
  • context grows until a provider rejects the next request;
  • retries, forks, or subagents lose cost and causal attribution;
  • operators cannot reconstruct what happened from prompt text alone; or
  • evals test final prose while missing the runtime trajectory.

Cayu treats these as runtime contracts. Important actions become structured events; tool authority and recovery are explicit; configured durable stores let transcripts and checkpoints survive process boundaries; and the same public seams support local development, tests, control-plane inspection, and hosted deployments.

What Cayu provides

Need Cayu primitive
Long-running work Durable sessions, transcripts, events, resume, fork, interruption
Safe effects Typed tools, effect declarations, policies, approvals, idempotency keys
Human interaction User-input checkpoints, approval resolution, manual recovery
Context pressure Token counting, projection, compaction, overflow recovery
Cost control Usage events, run limits, budgets, pricing, causal-budget summaries
Execution boundaries Environments, workspaces, runners, artifacts, vaults, egress
Reviewed knowledge Durable entries, approval state, keyword/vector retrieval, recall tools
Provider flexibility OpenAI API, experimental OpenAI subscription login, Anthropic, Bedrock, Vertex, OpenAI-compatible APIs
Agent operations Tasks, dispatchers, event watchers, subagents, runtime hooks
Behavioral proof Runtime tests, trajectory assertions, replay, eval reports
Operations FastAPI control plane and a packaged inspection dashboard

Quickstart

Start a project

The generated project is the recommended path for both humans and coding agents. Cayu requires Python 3.11 or newer.

You can give a coding agent one request: “Run pip install cayu and create a code review agent.”

pip install cayu pytest
cayu new myagent
cd myagent

cayu inspect --json
cayu check --json
pytest
cayu eval run

# After configuring the provider in app.py:
python run.py --message "Review this change."

The scaffold is credential-free and includes:

  • a process-scoped build_app() factory with bounded local-store initialization;
  • one model-only agent with no required tools;
  • a hermetic runtime test and output eval; and
  • a project-local AGENTS.md with the exact build and verification contract.

Open the generated project, describe the requested job in the existing agent, and keep its public test/eval seam intact.

Run an agent

This compact example shows the core API. Real projects should put the same registrations in the generated build_app() factory instead of constructing a module-global app.

import asyncio

from cayu import (
    AgentSpec,
    CayuApp,
    Message,
    OpenAIProvider,
    RunRequest,
    run_to_completion,
)


async def main() -> None:
    app = CayuApp()
    app.register_provider(OpenAIProvider(), default=True)  # reads OPENAI_API_KEY
    app.register_agent(AgentSpec(name="assistant", model="gpt-5.6"))

    outcome = await run_to_completion(
        app,
        RunRequest(
            agent_name="assistant",
            messages=[Message.text("user", "Explain durable agent sessions.")],
        ),
    )

    if outcome.ok:
        print(outcome.final_text)
    else:
        print(f"{outcome.status}: {outcome.error}")


asyncio.run(main())

CayuApp() uses in-memory stores by default, which is appropriate for this one-shot example and for tests. The generated project configures all local Cayu stores in data/cayu.db so sessions survive process restarts. Multi-process production deployments should select a conforming shared store such as PostgreSQL.

CayuApp.run(...) is the lower-level event-stream API. Runtime failures arrive as terminal session.failed events instead of exceptions raised from iteration. run_to_completion(...) consumes that same stream and returns a typed outcome when an application only needs the result. It retains the complete event stream in RunOutcome.events; use it for bounded runs. Consume CayuApp.run(...) incrementally for long-lived or high-volume runs.

For a credential-free domain-tool tracer bullet, run cayu guide references#domain-tool, then use cayu generate tool. To add workspace tools and command execution, see examples/local_environment_runtime.py.

Build with a coding agent

The generated AGENTS.md is the project-local source of truth. Ask the coding agent to read it first, then use Cayu's package-shipped guides and structured inspection:

cayu guide anatomy
cayu guide authoring
cayu inspect --json
cayu check --fail-on warning --json

The package-shipped cayu guide authoring#cayu-map routes each optional capability to the smallest version-matched local reference. Its online source mirror is secondary. The examples index provides runnable references without making them required project structure.

The supported authoring loop is:

understand -> inspect -> plan -> change -> test -> eval -> exercise -> report evidence

Start by editing the existing model-only agent, test, and eval. Add a generated tool-backed slice only when the requested job needs a capability outside the model; generated slices remain unfinished until their placeholder behavior, test, and eval have been replaced.

Mental model

Cayu separates the agent's identity from the resources and durable state used for one execution:

AgentSpec
  identity, model, system prompt, defaults, runtime policies

Environment
  workspace, runner, artifacts, vault, proxy, knowledge, MCP

Session
  durable identity, transcript, events, status, checkpoints

ToolContext
  the active environment services and call identity for one tool execution
  • Agent describes who is acting and how model work is configured.
  • Environment describes what that agent can touch.
  • Session records one durable execution and its lineage.
  • Tool is an explicitly registered, application-owned capability that the model may request. A native Python Tool runs inside the trusted Cayu application process; ToolPolicy gates its invocation but does not sandbox its implementation.
  • Task is an optional durable unit of background or orchestrated work.
  • Workflow is deterministic application orchestration around agent steps.

An environment is optional for a conversational agent. It becomes important when tools need files, commands, artifacts, secrets, network policy, or a sandbox. Static environments are useful for trusted local work; EnvironmentFactory creates or reattaches session-specific environments in production.

Choose the execution surface according to where code should run and which boundary should contain it:

Surface Execution location Boundary
Native Python Tool Cayu application process Trusted application code; policy controls invocation, not host-process access
Runner-backed operation (ctx.runner) Selected runner for that operation; the enclosing Tool.run() stays in the Cayu application process Isolation, environment, network, and filesystem guarantees for the operation come from the admitted runner and environment
MCP tool Configured MCP process or server Separate integration boundary whose process, transport, credentials, and isolation remain deployment choices
Virtual egress Selected runner plus a trusted broker outside it A conforming adapter can keep the real credential out of the workload; code isolation still depends on the runner

Use the smallest runtime shape

Do not add every Cayu primitive to every application.

Desired behavior Start with
One model-driven interaction CayuApp, AgentSpec, provider, RunRequest
Deterministic model-callable action Tool, ToolSpec, explicit ToolEffect
Authority over an effect ToolPolicy; approval only where a human gate is required
Mutable files or commands Explicit Environment, Workspace, and Runner
Durable uploaded or generated files ArtifactStore
Long-lived conversation or recovery Durable SessionStore and checkpoint APIs
Background durable work TaskStore plus an explicitly started worker
Delegated model work Subagent tools with bounded child-session policy
Behavioral regression proof EvalSuite and trajectory assertions

Start a conversation agent with the model and state it needs. Add workflows, task queues, environments, memory stores, servers, or multi-agent topology when the behavior requires them. Give coding agents narrow domain tools before granting broader shell access.

Application UI and control plane

Your application should own:

  • end-user prompts and domain forms;
  • product authentication and authorization;
  • business-specific workflow and state;
  • user-facing streaming, notifications, and presentation; and
  • decisions about when a run, task, approval, or interruption is allowed.

Cayu owns runtime execution and the operational state recorded by the application's configured stores. Its optional dashboard is a control plane for developers and operators: inspect sessions, events, transcripts, tasks, usage, artifacts, pending actions, and recovery state. Your application remains responsible for the product experience.

Start work through the API that matches the trigger:

  • run for an immediate new session;
  • resume for a deliberate continuation;
  • dispatch for placement through a dispatcher;
  • a task worker for durable queued work;
  • a subagent for model-selected bounded delegation; or
  • an event watcher for durable reactions to already-persisted events.

See Triggering runs for the decision guide and lifecycle responsibilities.

Providers and environments

The base package includes the provider contracts and built-in OpenAI, Anthropic, OpenAI-compatible HTTP, and experimental OpenAI-subscription adapters. Optional extras add integrations without forcing their dependencies into every deployment:

Extra Adds
cayu[server] FastAPI control plane and packaged dashboard
cayu[server-settings] Server extra plus typed environment and .env loading
cayu[postgres] PostgreSQL session, task, knowledge, and related stores
cayu[aws] Amazon Bedrock and Lambda MicroVM support
cayu[vertex] Anthropic models through Google Cloud Vertex AI
cayu[e2b] E2B runner and workspace
cayu[microsandbox] Local microVM-backed untrusted-code runner
cayu[egress] Virtual egress and credential-broker primitives
cayu[files] Image and PDF inspection
cayu[console] Interactive application console

Providers normalize text, thinking, tool calls, usage, completion reasons, and typed failures behind one runtime contract. Applications register providers explicitly and may add deterministic model-pattern routing; an arbitrary model name never selects a provider.

For local development without separate OpenAI API billing, users can sign in with their own ChatGPT subscription:

cayu auth openai login
# For SSH or a remote machine:
cayu auth openai login --headless
from cayu import OpenAISubscriptionProvider

app.register_provider(OpenAISubscriptionProvider(), default=True)

This experimental integration uses the Codex backend. It does not use the documented OpenAI Platform API. Cayu identifies itself with originator: cayu and preserves upstream rejections. OpenAI has not documented this raw backend as a general third-party provider API, so support may change or stop.

Intended-use boundary: Use this path only for a subscription holder's own local development and evaluation. For production, customer-facing or multi-user services, use the OpenAI Platform API or another officially supported provider. Do not share or resell credentials or bypass plan limits.

See OpenAI subscription authentication for the support boundary, credential storage, and fallback options.

The same agent can run in a local workspace, trusted Docker container, E2B, Microsandbox, Lambda MicroVM, or an application-owned runner without changing its identity or transcript contract.

Production boundaries

Cayu makes safety boundaries explicit, but configuration still matters:

  • Native Python Tool implementations are trusted host-process code and can access authority available to the Cayu application. ToolPolicy controls whether the model may call a tool and with which arguments; it is not an OS isolation boundary. Run model-authored or otherwise untrusted code through an admitted runner or separately governed external tool boundary. Native tools that need credentials should use explicit SecretRef values through ctx.proxy or ctx.vault and return only safe results; ambient host environment values are not automatically mediated as workload credentials.
  • LocalRunner executes directly on a trusted local machine and provides no sandbox isolation.
  • DockerRunner is useful for development and CI; ordinary Docker isolation is not presented as a secure untrusted-code boundary.
  • Environment registration does not imply selection: mark a default explicitly or name the environment on the request. Provider defaults and model-pattern routing should likewise be configured deliberately and kept unambiguous.
  • Tool effects do not authorize themselves. Use policies, approvals, scoped credentials, and destination controls where consequences require them.
  • SQLite is appropriate for local and single-writer deployments. Use PostgreSQL or another conforming shared store for sustained multi-process concurrency.
  • The FastAPI control plane requires an explicit ServerConfig access policy. Use AuthenticatedAccess for deployed operator surfaces; OpenAccess and ServerConfig.local_development() are deliberate local-only choices. Deployment names are descriptive metadata and never relax security policy. See server configuration. AuthContext.tenant records authenticated operator provenance but does not filter or isolate Cayu data. See Server authentication and tenant isolation. Generated API documentation is a separate exposure decision.
  • When embedding with mount_cayu(..., path="/your/path") or the lower-level mount_dashboard(...), use /your/path/ as the canonical dashboard URL. Cayu redirects an exact GET or HEAD of the slashless non-root mount after a successful dashboard mount. That public 307 may be returned without credentials; dashboard HTML, assets, deep links, and other protected content at the canonical target still require configured authentication. mount_cayu(...) places its control-plane API under /your/path/api; mount_dashboard(...) configures apiBaseUrl independently and defaults it to /api.
  • Usage is derived from recorded events and survives restarts when those events use a durable store; cost remains an estimate against the price book your application selects.
  • Recovery never invents the outcome of an ambiguous external side effect. Reconcile it through the typed recovery APIs.

Read Runtime contracts before changing persistence, replay, approval, interruption, budget, provider, runner, or recovery behavior.

Documentation

Start with the document that matches the job:

Goal Guide
Choose Cayu concepts and build an application, by hand or with an AI coding agent cayu guide authoring#cayu-map
Classify and verify tool mutation and replay behavior cayu guide tool-effects
Understand factories, process roles, and lifecycle cayu guide anatomy (source)
Choose how work starts Triggering runs
Create per-session workspaces and runners Environment factories
Implement a runner for your platform Build a runner
Configure network and credential boundaries Virtual egress
Run GitHub CLI without giving the runner a real token GitHub CLI through virtual egress
Design assertions and trajectory evals Evals
Estimate and govern cost Cost optimization
Use the application console Console
Start a configured server process Project server
Start a named worker process Project workers
Configure CLI session-store discovery Session-store targets
Inspect durable sessions safely Session inspection
Configure a control-plane server deployment Server configuration
Embed Cayu behind tenant-aware product APIs Server authentication and tenant isolation
Inspect supported model metadata Model catalog
Look up exact runtime behavior Runtime contracts
Track prerelease behavior and migrations Release notes

Maintainer-facing architecture is documented in Architecture, Project layout, and the Glossary.

Examples

Advanced examples are executable runtime specifications. Each example states its evidence boundary instead of presenting one strategy as suitable for every workload. Their measured results are described in Advanced runtime strategies.

Contributing and security

Cayu contributors should read CONTRIBUTING.md for placement policy, setup, validation commands, and pull-request requirements. New third-party integrations normally live in their own packages against Cayu's public extension contracts.

Report suspected vulnerabilities privately as described in SECURITY.md. Do not open a public issue or pull request for a suspected security vulnerability.

For questions and project discussion, join Discord. Use GitHub issues for actionable bugs and concrete feature proposals.

License

Cayu is licensed under the Apache License 2.0.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

cayu-0.1.0rc5.tar.gz (1.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

cayu-0.1.0rc5-py3-none-any.whl (2.0 MB view details)

Uploaded Python 3

File details

Details for the file cayu-0.1.0rc5.tar.gz.

File metadata

  • Download URL: cayu-0.1.0rc5.tar.gz
  • Upload date:
  • Size: 1.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for cayu-0.1.0rc5.tar.gz
Algorithm Hash digest
SHA256 94fb4103fa247b40412daee5add6e609fee59a419e041d6b284217995ebae78b
MD5 c56ebb31192fce5845719b64a32ebbf4
BLAKE2b-256 50b3fd028760447712cc3728193aacbbbadda79620291af1e26c7d9091a85722

See more details on using hashes here.

Provenance

The following attestation bundles were made for cayu-0.1.0rc5.tar.gz:

Publisher: ci.yml on cayu-dev/cayu

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file cayu-0.1.0rc5-py3-none-any.whl.

File metadata

  • Download URL: cayu-0.1.0rc5-py3-none-any.whl
  • Upload date:
  • Size: 2.0 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for cayu-0.1.0rc5-py3-none-any.whl
Algorithm Hash digest
SHA256 5c0798c38b4c7f980e48011b1b2cb8bdead36fd89915c4f3cd1bde76dd45c9ae
MD5 927e22fbc91bbba75a66b5af8a6942e7
BLAKE2b-256 37473e9380476825673b93bb958b1337e2d9d0728abf2f6f007154dc3075c95a

See more details on using hashes here.

Provenance

The following attestation bundles were made for cayu-0.1.0rc5-py3-none-any.whl:

Publisher: ci.yml on cayu-dev/cayu

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page