Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

Kitaru

Traces you can run, not just read.

Kitaru (来る, "to arrive") records every agent run as a full trace — every model call, tool call, and decision — and replays it against your real code. Reproduce the trace exactly. Fork it with one thing changed. Trust the diff. It works underneath whatever framework you already use, self-hosted on your own infrastructure, and it can deploy and run your agents too.

PyPI Python License

Docs · Quick Start · Examples · Getting Started Guide · Roadmap · Community


Kitaru Dashboard

🎯 Why Kitaru?

Most traces are transcripts — you read them. A Kitaru trace re-executes: your actual code runs again, with the trace answering for everything the original run saw. Kitaru is a debugger with a memory, sitting beside your observability stack — it tells you what happened; Kitaru re-runs it. That turns production traffic into the eval suite you never had to write: every incident is a reproducible test case, and "would the cheaper model have held?" is an experiment over real traces instead of a guess.

  • Every trace is a recording. Each checkpoint output — model call, tool call, decision — is written to your object store as a typed, versioned artifact. Step through it, diff it against other runs, trace a bad output back to the step that produced it.
  • Replay is re-execution, not re-scoring. An unchanged replay reproduces the original exactly — and that faithful baseline is what lets you fork from any checkpoint with one thing changed and trust that the diff is your change, not replay noise.
  • Decide with evidence. Every trace includes the model traffic — prompt, response, tokens, latency, estimated cost — recorded automatically by the framework adapters, or by kitaru.llm() in raw Python.

🔁 The loop

uv add "kitaru[pydantic-ai]"   # plain `kitaru` for the raw @flow/@checkpoint path
kitaru init

No decorators, no graph, no rewrite. Wrap the agent you already have and run it — Kitaru opens a flow around the call and records every model request and tool call as a checkpoint:

# agent.py
from pydantic_ai import Agent
from kitaru.adapters.pydantic_ai import KitaruAgent

agent = Agent("openai:gpt-5.4", name="support-agent",
              system_prompt="You resolve support tickets.")

@agent.tool_plain
def refund_payment(order_id: str) -> str:
    return payments.refund(order_id)  # your real API

support = KitaruAgent(agent)
support.run_sync("Refund order #4821 — the card reader was double-charged.")

Traces recorded elsewhere land the same way — import them, and they become executions like any other:

from kitaru import KitaruClient

client = KitaruClient()
client.executions.import_traces("support-traces.jsonl", format="otel")
client.imports.langfuse(
    "langfuse-observations.jsonl",
    source_project_id="prod",
    agent_name="support-agent",
)

Every run is now a trace you can replay:

trace = client.executions.latest()

# Replay — start from the agent's first model call, and your real code
# runs again against the recorded world. Unchanged, it reproduces the
# original exactly. That's your baseline.
client.executions.replay(trace.exec_id, at="support-agent_model_request")

# Fork — same trace, one thing changed: patch the recorded tool output.
# What would the agent have done if the refund had succeeded?
client.executions.replay(
    trace.exec_id,
    at="refund_payment_tool",
    checkpoint_overrides={
        "refund_payment_tool": {"output": "refund issued: $129.00"},
    },
)

# Widen — the same call takes a list. Replay last week's traces against
# the code in your working tree, and the cohort is a regression test.
traces = client.executions.list(limit=20)
client.executions.replay(
    [t.exec_id for t in traces],
    at="support-agent_model_request",
    tag="pr-1234-check",
)

Overrides can also swap the model on an LLM call, edit tool arguments, or swap a checkpoint's code — see Replay and overrides. Explicit @flow/@checkpoint decorators are there when you want named replay boundaries or multi-turn workflows, and flow.deploy() ships a winner as a versioned deployment invoked by name — optional; stopping at the regression test is a fine place to stop.

Durable execution (the plumbing)

Recording a run means surviving one. Checkpoints double as crash recovery — a crash or pod eviction resumes from cached outputs instead of re-burning tokens. kitaru.wait() pauses a flow for hours or days until a human or webhook responds. flow.deploy() freezes versioned snapshots that consumers invoke by name, and @checkpoint(runtime="isolated") runs heavy steps in their own pod on Kubernetes, AWS, GCP, or Azure. This is how a faithful recording gets minted — not the reason you reach for Kitaru.

Works with your agent SDK

Adapters for six agent frameworks — wrap your existing agent, no rewrite:

Framework Adapter
PydanticAI kitaru.adapters.pydantic_ai.KitaruAgent
OpenAI Agents SDK kitaru_openai_agents.KitaruRunner (install kitaru-openai-agents)
Claude Agent SDK kitaru.adapters.claude_agent_sdk.KitaruClaudeRunner
LangGraph, LangChain, and Deep Agents kitaru_langgraph.KitaruGraphRunner (install kitaru-langgraph)
Gemini kitaru.adapters.gemini.KitaruGeminiInteractionsRunner
Google ADK kitaru.adapters.google_adk.KitaruADKRunner

For raw-Python agents, @flow and @checkpoint around your calls give you the same recording without an adapter. Your model, your tools, your framework — Kitaru wraps them, not the other way around.

The OpenAI Agents SDK adapter currently lives only in this repository's plugin workspace. It is not exported by the installed kitaru package.

Inspect it from your coding agent

Kitaru's optional v2 MCP server gives Claude Code, Codex, Cursor, and other MCP clients a compact typed interface for agents, sessions, cohorts, experiments, evaluators, replays, and asynchronous jobs. It starts in read-only mode with two tools; write and destructive tools are absent unless you explicitly select a broader mode.

uv add kitaru --extra mcp
claude mcp add --scope project kitaru -- kitaru-mcp

See the MCP server guide before enabling standard or destructive mode.

Claude Code users can also install the kitaru-skills plugin — quickstart, workflow authoring, and adapter-migration skills:

/plugin marketplace add zenml-io/kitaru-skills
/plugin install kitaru@kitaru

Self-hosted, batteries included

A single server on your own infra. Flows run on whichever stack you pick — local, Kubernetes, GCP, AWS, or Azure — with artifacts in your own S3/GCS/Azure Blob bucket, and a built-in UI to step through executions, diff replays, and approve human-in-the-loop wait steps. No mandatory SaaS control plane.

📚 Learn more

Resource Description
Getting Started Guide Full setup walkthrough with all examples
Documentation Complete reference and guides
Agents guide Run, replay, and improve production agents end to end
Examples Runnable workflows for every feature
Stacks Deploy to Kubernetes, AWS, GCP, or Azure
MCP server Inspect and operate the Kitaru v2 API from a compact, read-only-by-default MCP server

🌱 Origins

Kitaru is built by the team behind ZenML, drawing on five years of production orchestration experience (JetBrains, Adeo, Brevo). The orchestration primitives (stacks, artifacts, lineage) are purpose-rebuilt here for autonomous agents.

🤝 Contributing

We welcome contributions! See CONTRIBUTING.md for development setup, code style, and how to submit changes. The default branch is develop — all PRs should target it.

💬 Community and support

  • Discussions — ask questions, share ideas
  • Issues — report bugs, request features
  • Roadmap — see what's coming next
  • Docs — guides and reference

📄 License

Apache 2.0

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

kitaru-0.22.0rc0.tar.gz (3.8 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

kitaru-0.22.0rc0-py3-none-any.whl (2.6 MB view details)

Uploaded Python 3

File details

Details for the file kitaru-0.22.0rc0.tar.gz.

File metadata

  • Download URL: kitaru-0.22.0rc0.tar.gz
  • Upload date:
  • Size: 3.8 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for kitaru-0.22.0rc0.tar.gz
Algorithm Hash digest
SHA256 9dd9ab074a02aa27c06106796bcad919eb6fb7ec6fcc47727ec3f8b9e9c20fd1
MD5 9390bf27deb3b99cee44fdf7b38f3775
BLAKE2b-256 a93f351c5f25ab5fc88494921f1b56149088a984c3ae3ad45e7b941a04cc41a2

See more details on using hashes here.

Provenance

The following attestation bundles were made for kitaru-0.22.0rc0.tar.gz:

Publisher: release-plugins.yml on zenml-io/kitaru

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file kitaru-0.22.0rc0-py3-none-any.whl.

File metadata

  • Download URL: kitaru-0.22.0rc0-py3-none-any.whl
  • Upload date:
  • Size: 2.6 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for kitaru-0.22.0rc0-py3-none-any.whl
Algorithm Hash digest
SHA256 a23237e6bc92e54a4cd90740b52a72b9f63587d596c5440a75e1b7de6e101d03
MD5 0482951aeeec77c08de625460d6ca838
BLAKE2b-256 88a98ce89fb9dc88731423753710a388b03f15567ec893225626d9d7bdfd2a34

See more details on using hashes here.

Provenance

The following attestation bundles were made for kitaru-0.22.0rc0-py3-none-any.whl:

Publisher: release-plugins.yml on zenml-io/kitaru

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page