Skip to main content

Record-and-replay for agent decision graphs: reproduce a prod agent failure as a committed regression test — and re-run your fix without live LLM calls.

Project description

Chronicle

Record-and-replay for agent decision graphs.
Turn a production agent failure into a committed regression test, and re-run your fix without live LLM calls.

CI PyPI Downloads Python License: MIT Ruff Stars

Chronicle: record an incident, then verify the fix with a cut-point test

Watch the full walkthrough

Chronicle records what your agent did at each decision boundary (its LLM calls, tool calls, and routing choices) so you can reproduce a production failure as a committed regression test and re-run your fix without live LLM calls. The target is one specific, real problem: control-flow and tool-safety regressions in multi-agent systems, caught deterministically from recorded incidents.

Why · Architecture · Install · Quick start · Verification · Demos · Comparison · CLI · Roadmap

Why Chronicle

  • Record every LLM call, tool call, and routing decision as an immutable Envelope.
  • Cut-point replay: change one boundary, freeze the rest of the incident, and assert deterministically with no LLM calls.
  • Two-layer verification: structural replay for control flow and tool safety, plus LLM-as-judge for meaning.
  • Commit incidents as regression tests so a fixed failure never silently returns.
  • Batteries included: secret redaction, real model-version capture, and LangGraph / OpenInference integration.

Architecture

Chronicle is two systems that share one artifact, the Envelope: a recorder that captures boundary crossings during a live run, and a test bench that replays them without touching the model.

flowchart LR
    subgraph REC["Record (LIVE run)"]
        A["Your agent<br/>llm · tool · router calls"] -->|"@boundary"| B["Envelope Recorder"]
        B --> C["Append-only<br/>envelope store (.jsonl)"]
        C --> D["Execution graph<br/>(side graph, no topology change)"]
    end

    D -->|"export / extract"| E["fixtures/ committed to git<br/>envelopes/ · traces/"]

    subgraph REP["Replay & Verify (no live LLM)"]
        E --> F["ReplayPlan<br/>stub upstream · run one boundary live"]
        F --> G["Layer 1: structural replay<br/>control flow &amp; tool safety"]
        E --> H["Layer 2: LLM-as-judge<br/>grounding · safety · refusal"]
    end

    G --> I["pytest / CI<br/>regression tests"]
    H --> I
    D -.->|"on_crossing hook"| J["Cost / governance<br/>(external, e.g. TokenOps)"]

The Envelope is an immutable, append-only record of one boundary crossing:

Field Contents
Contextual metadata Model version, sampling parameters, runtime build ID
Input state Assembled prompt, graph state, retrieved context chunks
Action / result Structured tool calls and model completion
Graph linkage parent_envelope_id, sequence, invocation_index for retries

Optional OpenInference and Arize Phoenix integrations feed framework-agnostic tracing into this same envelope format.

Install

# From source (development):
pip install -e ".[dev]"

# From PyPI:
pip install agent-chronicle

Quick start

Annotate a decision boundary once with @boundary; it records in live mode and stubs from a fixture in replay mode. Record a run, and freeze it as a committed fixture, in a single block:

import chronicle
from chronicle import boundary

@boundary("agent", kind="llm")
def agent_plan(state: dict) -> dict:
    ...

@boundary("delete_file", kind="tool")
def delete_file(path: str, environment: str) -> dict:
    ...

with chronicle.record(
    "incident-001",
    store=".chronicle/runs/incident.jsonl",
    export="fixtures/traces/incident-001/",
):
    run_agent(...)

record() wraps reset_session(); drop to the session API when you need finer control.

Cut-point replay

Test a fix in one boundary while the rest of the incident stays frozen. Upstream boundaries are stubbed from the fixture, your changed boundary runs live, and you assert on its captured result.

import chronicle
from chronicle import ReplayPlan

with chronicle.replay_trace(
    "fixtures/traces/deletion-incident-001/",
    ReplayPlan()
    .stub("agent", 1)          # upstream: frozen from fixture
    .live("delete_file", 1)    # cut-point: run new code
    .live("agent", 2)          # downstream: observe the effect
) as session:
    run_agent(...)
    assert session.captured_result("delete_file", 1)["blocked"] is True

One decorator, two behaviors: in live mode your function runs and its input/output are recorded into an Envelope; in replay + stub mode it does not run and Chronicle returns the recorded output. A cut-point is the one boundary you flip back to live to test new code against real upstream inputs.

Verification layers

Layer Goal Mechanism
Layer 1: replay Validate control flow and tool safety Structural assertions over recorded fixtures; never calls the LLM
Layer 2: evaluation Validate generation quality LLM-as-a-judge on meaning (grounding, safety, refusal), not bitwise equality
Cut-point replay Test a change in one boundary Stub upstream from fixtures, run the target boundary live

Layer 1 (single-envelope injector)

from chronicle.replay import ReplayInjector
from chronicle import Envelope

envelope = Envelope.from_file("fixtures/envelopes/incident-2026-06-17-001.json")
injector = ReplayInjector(envelope)

def agent(state, inj):
    inj.stub_llm()
    inj.stub_tool("search_docs", {"query": "reset API key"})
    return {"finish_reason": "tool_calls"}

_, _, assertions = injector.replay(agent)
assert all(a.passed for a in assertions)

Layer 2 (LLM-as-judge)

from chronicle.judge import JudgeRunner, OpenAIJudgeClient

runner = JudgeRunner(OpenAIJudgeClient(model="gpt-4o-mini"))
result = runner.evaluate(envelope)
assert result.overall_passed

How Chronicle compares

Chronicle is not a tracing dashboard or an eval framework. It is the piece that makes a recorded agent run replayable and testable, and it sits alongside the tools you already use.

Chronicle LangSmith / Langfuse / Phoenix promptfoo VCR.py
Trace agent runs Partial HTTP only
Deterministic replay, no live LLM No No ✅ (HTTP)
Cut-point: change one boundary, freeze the rest No No No
Commit incidents as regression tests Via datasets
Structural + LLM-judge verification Judge only Judge only No
Agent-graph aware (boundaries, retries) No No

Demos

Each demo records an incident from an ungated tool, then a cut-point test verifies the gated fix. All share the agent@1 -> tool@1 -> agent@2 shape.

Demo What goes wrong Run the cut-point test
Refund $9.8M refund on a $47 order (amount read from the order ID) python examples/financial_incidents/run.py refund test
Invoice EUR 2M invoice sent as USD python examples/financial_incidents/run.py invoice test
Trade ~$190k sell instead of ~$1k (notional read as share count) python examples/financial_incidents/run.py trade test
Deletion Ungated delete_file wipes prod python examples/deletion_agent/run_cutpoint_demo.py
Record the incident, visualize the trace, run the full suite
# Financial incidents: record the bad run, then cut-point test the fix
python examples/financial_incidents/run.py refund record
python examples/financial_incidents/run.py all test
pytest tests/test_financial_incidents.py -v

# Deletion agent: record, visualize the trace, cut-point test
python examples/deletion_agent/record_incident.py
python examples/deletion_agent/show_trace.py --ui   # interactive timeline + graph
pytest tests/test_deletion_cutpoint.py -v

The gated fix refuses when an amount exceeds a flat cap (MAX_REFUND_CENTS, MAX_INVOICE_CENTS, MAX_ORDER_NOTIONAL_CENTS). Source lives under examples/financial_incidents/ and examples/deletion_agent/.

CLI

Command reference
chronicle record                                    # bootstrap tracing + instrumentation
chronicle extract --trace-id ID                     # export envelopes to fixtures/
chronicle replay FIXTURE.json                       # Layer 1 deterministic replay
chronicle verify FIXTURE.json --layer2 --mock-judge # Layer 1 + Layer 2
chronicle show-graph fixtures/traces/TRACE --ui     # interactive trace visualization
chronicle show-graph TRACE --html out.html          # static HTML export
chronicle schema                                    # print Envelope JSON Schema
chronicle list-fixtures                             # list committed envelope fixtures

LangGraph integration (optional)

Wrap LangGraph nodes as an alternative to @boundary
from chronicle.envelope.capture import EnvelopeRecorder
from chronicle.envelope.store import EnvelopeStore
from chronicle.instrumentation import instrument_graph_nodes

recorder = EnvelopeRecorder(
    store=EnvelopeStore(".chronicle/runs/envelopes.jsonl"),
    model_version="gpt-4o-2024-08-06",
    build_id="deploy-abc123",
)
wrapped_nodes = instrument_graph_nodes(recorder, {"agent": agent_node})

See examples/langgraph_demo/agent.py.

Cost and governance observers (on_crossing)

Chronicle does not embed cost management. External systems (for example TokenOps) attach an observer that fires after each live crossing:

session = reset_session()
session.on_crossing = my_observer  # (boundary_id, kind, input_state, result) -> None

It runs after a live envelope record and a live cut-point capture, and does not run on stub replay. See tests/test_cost_management_e2e.py for an end-to-end ledger and budget pattern.

Environment variables

Variable Purpose
CHRONICLE_BUILD_ID Pin runtime build ID in envelope metadata
CHRONICLE_STORE Default envelope store path
PHOENIX_COLLECTOR_ENDPOINT Phoenix OTLP endpoint (default http://localhost:4317)

Project structure

Only chronicle/ is the installable library. Demos and interactive benches stay under examples/; committed regression traces live in fixtures/.

chronicle/                 # installable package
├── boundary.py            # @boundary decorator (record + replay + cut-point)
├── session.py             # runtime session, on_crossing hook, stub/live modes
├── execution_graph.py     # side graph builder (load/save/render)
├── visualizer.py          # HTML trace UI (library + CLI)
├── envelope/              # schema, capture, append-only store
├── replay/                # ReplayPlan, ReplayInjector, structural assertions
├── judge/                 # Layer 2 rubric + LLM-as-judge runner
├── instrumentation/       # OpenInference + LangGraph hooks
└── cli.py
fixtures/                  # committed regression data (envelopes/ · traces/)
examples/                  # demos and test benches (not imported by the package)
scripts/                   # demo and test runners
tests/                     # unit + e2e

Talks and writing

This work was presented at the AI Engineer World's Fair 2026.

Roadmap

Chronicle is early (0.x), and the Envelope schema may still change. Near-term:

  • Drop-in provider capture: chronicle.wrap(client) for OpenAI and Anthropic.
  • Async @boundary for async agents and graph nodes.
  • A pytest plugin so a committed incident becomes a one-decorator regression test.
  • One-call LangGraph instrumentation and a documentation site.

Ideas and priorities are welcome in Discussions.

Contributing

Contributions are welcome. See CONTRIBUTING.md for dev setup, the DCO sign-off, and the record-and-replay reviewer checklist, and please read our Code of Conduct.

Security

Chronicle captures prompts, agent state, and retrieved context, so a recording can contain secrets. Turn on redaction before recording production traffic, so nothing sensitive reaches a committed fixture:

import chronicle

session = chronicle.reset_session()
session.redactors = chronicle.default_redactors()   # mask API keys, tokens, JWTs

Redaction runs at record time and keeps the structure your tests assert on (message roles, tool names, argument keys) while masking the values. Add your own (str) -> str redactors for PII. Read the data-handling guidance in SECURITY.md and report vulnerabilities privately per that policy.

Contributors

Thanks to everyone who has contributed.

Contributors

Questions or ideas? Open a Discussion. For security issues, follow SECURITY.md.

License

MIT (c) 2026 Susheem Koul and Tisha Chawla.


If Chronicle saves you a debugging session, please ⭐ star the repo so more people can find it.

Built by Susheem Koul and Tisha Chawla

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agent_chronicle-0.1.1.tar.gz (406.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agent_chronicle-0.1.1-py3-none-any.whl (41.4 kB view details)

Uploaded Python 3

File details

Details for the file agent_chronicle-0.1.1.tar.gz.

File metadata

  • Download URL: agent_chronicle-0.1.1.tar.gz
  • Upload date:
  • Size: 406.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for agent_chronicle-0.1.1.tar.gz
Algorithm Hash digest
SHA256 def9e5bc71f2652f894185f8ca62b987197d6da06e2f557a86856202353af80a
MD5 fa4129c44b6f3590cac143beea6744f7
BLAKE2b-256 e047ec2776466a3a47fb1b40f33da536cc1c9daa3b3efc6610f837dd2064b05d

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_chronicle-0.1.1.tar.gz:

Publisher: release.yml on theagentplane/chronicle

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agent_chronicle-0.1.1-py3-none-any.whl.

File metadata

  • Download URL: agent_chronicle-0.1.1-py3-none-any.whl
  • Upload date:
  • Size: 41.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for agent_chronicle-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 4886ebec21da2ec62fb9ac5a1ffd61645833fb715420d55e1d0cf3ba316232dd
MD5 764a7974ab43eae3005b67606ea6ca56
BLAKE2b-256 8a22b8320b67a98faea289d34722ec7e02669e43a064803e83e74401c4afc2f6

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_chronicle-0.1.1-py3-none-any.whl:

Publisher: release.yml on theagentplane/chronicle

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page