agentreplay
Framework-agnostic regression testing for AI agents
Record real model and tool interactions once, replay them offline in pytest with zero API calls, and detect behavioural regressions with structured trajectory diffs.
Overview • Demo • Features • Installation • Quick Start • How It Works • Cassette Format • CLI & Fixture Reference • Development
Demo
Overview
Testing AI agents in continuous integration is often painful:
- Live LLM calls in CI are slow, expensive, and flaky.
- Traditional mocks are brittle and easily miss subtle agent drifts (such as skipping a verification tool or altering argument payloads).
- Raw snapshot tests generate massive, noisy JSON diffs cluttered with timestamps, request IDs, and non-deterministic tokens.
agentreplay brings deterministic VCR-style testing to AI agents:
- Record once against live models and tools during local test development.
- Replay offline in CI with zero network and zero model API calls.
- Catch behavioural drift with step-by-step trajectory diffs whenever tools, arguments, or execution sequences change.
[!NOTE]
agentreplayis a testing tool, not an agent runtime or orchestration system. You don't need to rewrite your agent or replace your framework runtime.
Features
- ⚡ Zero-API Replay: Substitutes model responses and tool executions offline — test suites execute in milliseconds.
- 🔍 Structural Trajectory Diffs: Highlights the exact step where an agent diverged instead of dumping raw JSON walls.
- 🎯 Non-Invasive Adapter: Integrates with PydanticAI via public capability hooks without monkey-patching HTTP clients.
- 🛡️ No Silent Fallback: Replay never secretly falls back to live network calls; cassette exhaustion and unexpected tool calls fail loudly.
- 📦 Git-Friendly Cassettes: Canonical JSONL serialization (sorted keys, compact separators) produces byte-identical files across operating systems.
- 🧪 Pytest-Native: Integrated
--agentreplayCLI option and test fixture for seamless workflow switching.
Installation
Install pytest-agentreplay using uv or pip:
# Using uv (recommended)
uv add pytest-agentreplay
# Using pip
pip install pytest-agentreplay
To include development dependencies:
uv add --dev pytest-agentreplay[pydantic-ai,pytest]
Quick Start
1. Using the pytest Fixture (Recommended)
Add the agentreplay fixture to your existing test function. The cassette path is automatically derived from the test module and function name:
from your_app import support_agent
def test_refund_flow(agentreplay):
caps = [c for c in [agentreplay.capability()] if c is not None]
result = support_agent.run_sync("Refund order 123", capabilities=caps)
assert "refund" in result.output.lower()
Step 1: Record Real Interactions
Run pytest with --agentreplay=record to capture live model and tool events into a cassette:
pytest --agentreplay=record tests/test_refund.py
RECORD tests/cassettes/test_refund/test_refund_flow.jsonl
✓ model interactions recorded
✓ tool interactions recorded
Step 2: Replay Offline in CI
Run with --agentreplay=replay to run tests entirely offline:
pytest --agentreplay=replay tests/test_refund.py
REPLAY tests/cassettes/test_refund/test_refund_flow.jsonl
✓ zero model API calls
✓ zero network
✓ deterministic
[!TIP] Running
pytestwithout the--agentreplayflag executes tests normally without recording or intercepting interactions.
2. Programmatic Usage
You can also control recording and replaying explicitly in code without pytest CLI flags:
import agentreplay
from your_app import support_agent
def test_custom_refund():
result = support_agent.run_sync(
"Refund order 123",
capabilities=[
agentreplay.pydantic_ai(
mode="replay", # or "record"
cassette_path="tests/cassettes/custom_refund.jsonl",
)
],
)
assert "refund" in result.output.lower()
Interactive Live Example (Click to expand)
Here is a complete, self-contained example you can run immediately without API keys using PydanticAI's built-in TestModel:
from pathlib import Path
from pydantic_ai import Agent
from pydantic_ai.models.test import TestModel
from pydantic_ai.tools import RunContext
import agentreplay
# 1. Define your agent & tools
def lookup_customer(ctx: RunContext[None], customer_id: str) -> str:
return f"Customer {customer_id}: Alice (tier=gold)"
support_agent = Agent(
TestModel(call_tools=["lookup_customer"]),
tools=[lookup_customer],
)
cassette = Path("test_refund.jsonl")
# 2. Record: live interaction captured to JSONL
cap_record = agentreplay.pydantic_ai(mode="record", cassette_path=cassette)
record_result = support_agent.run_sync("Find customer 123", capabilities=[cap_record])
print("Recorded output:", record_result.output)
# 3. Replay: 100% offline, zero model/tool execution
cap_replay = agentreplay.pydantic_ai(mode="replay", cassette_path=cassette)
replay_result = support_agent.run_sync("Find customer 123", capabilities=[cap_replay])
print("Replayed output:", replay_result.output)
assert record_result.output == replay_result.output
Behavioural Trajectory Diffs
When an agent changes its decision path (e.g. prompt changes, tool parameter updates, or model upgrade drifts), agentreplay provides a clear, numbered trajectory diff:
FAILED tests/test_refund.py::test_refund_flow - DivergenceError:
Agent trajectory changed
Expected:
1. model_request
2. tool_call: lookup_customer(id='123')
3. tool_call: check_refund_policy(tier='gold')
4. tool_call: refund_customer(amount=39)
5. model_response → "Refund processed."
Actual:
1. model_request
2. tool_call: lookup_customer(id='123')
3. tool_call: refund_customer(amount=39)
Divergence at step 3:
- tool_call: check_refund_policy(tier='gold')
+ tool_call: refund_customer(amount=39)
How It Works
┌────────────────────────────────────────────────────────┐
│ Agent Test │
└───────────────────────────┬────────────────────────────┘
│
┌─────────────┴─────────────┐
▼ ▼
[Record Mode] [Replay Mode]
│ │
Live Model & Tools Cassette JSONL File
│ │
Intercepts via Capability Intercepts via Hooks:
• Model Requests/Responses • SkipModelRequest
• Tool Calls/Results • SkipToolExecution
│ │
▼ ▼
Saves Canonical JSONL Zero Network / API Calls
- Record Mode: Intercepts model requests, model responses, and tool executions via PydanticAI's
AbstractCapabilityhooks. Deep-copies all arguments and results to prevent mutation side-effects. - Replay Mode: Uses PydanticAI's
SkipModelRequestto return recorded responses without model invocations, andSkipToolExecutionto substitute recorded tool outputs. - Divergence Engine: Tracks execution position with a stateful cursor. Rejects mismatches in event kinds, unexpected tool names, altered arguments, cassette exhaustion, and unconsumed leftover events.
Cassette Format
Cassettes are stored as streamable, Git-diffable JSON Lines (JSONL) files.
Line 1 contains the cassette header with format versioning and metadata:
{"created_at":"2026-08-27T08:30:00+00:00","format_version":1,"framework":"pydantic-ai"}
Subsequent lines contain canonical TraceEvent records:
{"event_id":"a1b2c3d4e5f6","kind":"run_start","timestamp":1724747400.0}
{"event_id":"b2c3d4e5f6a1","kind":"model_request","name":"gpt-4o","timestamp":1724747401.0}
{"arguments":{"customer_id":"123"},"event_id":"c3d4e5f6a1b2","kind":"tool_result","name":"lookup_customer","result":{"tier":"gold"},"timestamp":1724747402.0}
{"event_id":"d4e5f6a1b2c3","kind":"model_response","result":{"parts":[{"content":"Processed.","part_kind":"text"}]},"timestamp":1724747403.0}
{"event_id":"e5f6a1b2c3d4","kind":"run_end","timestamp":1724747404.0}
[!IMPORTANT] Serialization is deterministic: dictionary keys are sorted, compact separators are enforced, and encoding is UTF-8. Logically identical traces produce byte-identical files across platforms.
CLI & Fixture Reference
Pytest CLI Flags
| Flag | Description |
|---|---|
--agentreplay=record |
Run live tests and record interactions into cassette files. |
--agentreplay=replay |
Run tests offline using recorded cassettes; fail on behavioural divergence. |
| (no flag) | Standard pytest execution without recording or replay interception. |
agentreplay Fixture Methods
agentreplay.mode— Returns"record","replay", orNone.agentreplay.capability(cassette_path=None)— Creates anAgentReplayCapabilityinstance configured with the active mode and target cassette path.agentreplay.default_cassette_path— Returns the auto-derived path:tests/cassettes/{module_name}/{test_name}.jsonl.
Development
Set up a local development environment with uv:
# Clone the repository
git clone https://github.com/aafre/agentreplay.git
cd agentreplay
# Install dependencies
uv sync --all-extras
# Run quality gates
uv run ruff check .
uv run ruff format --check .
uv run mypy .
uv run pytest -v
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pytest_agentreplay-0.1.1.tar.gz.
File metadata
- Download URL: pytest_agentreplay-0.1.1.tar.gz
- Upload date:
- Size: 169.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
1b9ab7ea59c3db1051afcfcda3b83f12d79cccce7843923242cd8bcda65f7d1b
|
|
| MD5 |
ec5f632ebf1c8af21eacbff35d5e4bd0
|
|
| BLAKE2b-256 |
462d4638dc9b37aaa908d45c95ff636c3c7f1c489a5168997ef45eea050074df
|
Provenance
The following attestation bundles were made for pytest_agentreplay-0.1.1.tar.gz:
Publisher:
publish.yml on aafre/agentreplay
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pytest_agentreplay-0.1.1.tar.gz -
Subject digest:
1b9ab7ea59c3db1051afcfcda3b83f12d79cccce7843923242cd8bcda65f7d1b - Sigstore transparency entry: 2617999286
- Sigstore integration time:
-
Permalink:
aafre/agentreplay@5b6a79b805d58146cbcb0102df17518d63fa78d2 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/aafre
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b6a79b805d58146cbcb0102df17518d63fa78d2 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file pytest_agentreplay-0.1.1-py3-none-any.whl.
File metadata
- Download URL: pytest_agentreplay-0.1.1-py3-none-any.whl
- Upload date:
- Size: 20.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9bd29139fc9bb82d4f4b9f91386bfbfd3a426cb38df6f7e2ac38e241ebc5776b
|
|
| MD5 |
054f546b5dc6b78ef31aa0644b1eaa78
|
|
| BLAKE2b-256 |
532a44c58685badf7e283a5539b0353cd57f86f2f89a78f229883dbbb398c7bd
|
Provenance
The following attestation bundles were made for pytest_agentreplay-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on aafre/agentreplay
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pytest_agentreplay-0.1.1-py3-none-any.whl -
Subject digest:
9bd29139fc9bb82d4f4b9f91386bfbfd3a426cb38df6f7e2ac38e241ebc5776b - Sigstore transparency entry: 2617999299
- Sigstore integration time:
-
Permalink:
aafre/agentreplay@5b6a79b805d58146cbcb0102df17518d63fa78d2 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/aafre
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@5b6a79b805d58146cbcb0102df17518d63fa78d2 -
Trigger Event:
workflow_dispatch
-
Statement type: