Capture your agent's outputs, store them as baselines, and catch regressions in CI — with one decorator.
English · 中文 · Quick Start · How It Works · How It Compares
The Problem
You ship an AI agent. It works great. Two weeks later, you update a prompt, swap a model, or bump a dependency — and something breaks. But you don't notice until a user complains, because there's no test that catches agent behavior regressions.
Traditional unit tests don't work for agents. The outputs are non-deterministic natural language, so you can't just assertEqual — and writing fixtures by hand costs more than writing the agent.
AgentProbe fixes this. One decorator captures your agent's output and saves it as a baseline snapshot. On the next run, it compares the new output against the baseline — exact match or semantic similarity. If something changed, the test fails. Run it in CI, and you catch regressions before they hit production.
How It Works
Quick Start
pip install agentpoke
Heads up: the PyPI distribution is
agentpoke(the nameagentprobewas taken), but you import it asagentprobein code —from agentprobe import ....
1. Snapshot Testing
from agentprobe import snapshot
@snapshot("summarize_article")
def test_summarize():
result = my_agent.summarize("The quick brown fox jumps over the lazy dog.")
return result
First run: creates a baseline in .agentprobe/snapshots/summarize_article.json. Next runs: compares the output against the baseline and fails if they differ. async def tests work the same way. Snapshots are meant to be committed, so CI can compare against them.
Non-deterministic fields and credentials are handled before comparison:
# mask volatile fields at any depth so they don't cause spurious mismatches
@snapshot("summarize_article", redact=["timestamp", "request_id"])
# scrub API keys, tokens, JWTs and emails even inside free text,
# plus your own regex shapes (also globally via --agentprobe-redact-secrets)
@snapshot("summarize_article", redact_secrets=True, redact_patterns=[r"internal-\d{4}"])
2. Mock LLM
MockLLM is a drop-in replacement for openai.Client that returns scripted responses, so agent logic tests hit no API:
from agentprobe import MockLLM
mock = MockLLM(responses=[
"The document discusses three main topics.",
{"tool_calls": [{"id": "1", "function": {"name": "search", "arguments": '{"q": "test"}'}}]},
])
result = mock.chat.completions.create(messages=[{"role": "user", "content": "Summarize this doc"}])
assert "three main topics" in result.choices[0].message.content
assert mock.call_count == 1 # mock.calls records everything; mock.reset() for reuse
Scripted responses are consumed in order; default_response= covers the tail once they run out.
3. Tool Call Assertions
Verify your agent calls the right tools, in the right shape:
from agentprobe import assert_no_tool_called, assert_tool_called, assert_tool_sequence
assert_tool_called(tool_calls, "web_search", times=1)
assert_tool_called(tool_calls, "web_search", with_args={"query": "latest news"})
assert_tool_sequence(tool_calls, ["web_search", "summarize"])
assert_no_tool_called(tool_calls, "delete_file")
The variants cover the messy real cases:
assert_tool_sequence(..., contiguous=True)catches planner reorderings where a tool must immediately follow another.min_times/max_timesreplacetimeswhen the exact count is non-deterministic:assert_tool_called(tool_calls, "api_call", max_times=3)bounds a flaky retry.assert_max_tool_calls(tool_calls, 10)budgets the whole run, not just one tool (zero calls still passes).with_argsis a nested subset match and also accepts OpenAI-style JSON string arguments.assert_tool_not_called_with(tool_calls, "run", {"sudo": True})allows the tool but fails on a dangerous argument subset.
4. Schema Validation
from pydantic import BaseModel
from agentprobe import assert_schema
class AgentResponse(BaseModel):
answer: str
confidence: float
sources: list[str]
def test_output_structure():
output = my_agent.run("What is the capital of France?")
result = assert_schema(output, AgentResponse)
assert result.confidence > 0.8
5. Multi-Step Tracing
Record what an agent did step by step, then assert over the trace or snapshot it:
from agentprobe import Trace, assert_tool_sequence
trace = Trace()
trace.record_llm("planning the search")
trace.record_tool_call("search", {"query": "rainfall 2023"})
trace.record_event("retry", attempt=2)
trace.record_tool_call("fetch", {"url": "https://example.com"})
assert_tool_sequence(trace.tool_calls, ["search", "fetch"])
assert trace.names == ["llm", "search", "retry", "fetch"]
# trace.to_dict() is snapshot-friendly for full-run regression tests
6. Cost Tracking
Record token usage and assert the run stayed under a USD budget — catching regressions that quietly burn more money. Pricing comes from a dict, a callable, or TokenTracker's price table (pip install toktally):
from agentprobe import assert_cost_under
assert_cost_under(trace, 0.05, pricing={"gpt-4o": (0.005, 0.015)}) # (input, output) per 1k tokens
Pytest Integration
AgentProbe registers as a pytest plugin automatically:
def test_with_fixture(agentprobe):
output = my_agent.run("Hello")
result = agentprobe.capture("greeting_test", output)
assert result.passed
pytest tests/ # run tests
pytest tests/ --agentprobe-update # regenerate baselines after intentional changes
pytest tests/ --agentprobe-mode=semantic --agentprobe-threshold=0.85
When a snapshot changes, AgentProbe prints a unified diff between the stored JSON and the current output, so CI logs show the exact field or sentence that drifted. The standalone CLI mirrors the flags: agentprobe run, agentprobe run --mode semantic --threshold 0.9, agentprobe update.
Reviewing failures locally
A failed comparison also saves the actual output to .agentprobe/last_run/, so you can review and accept drift without re-running the test suite:
agentprobe diff # baseline vs last failing run, with similarity scores
agentprobe diff summarize # just one snapshot
agentprobe accept # promote all last-run outputs to baselines
agentprobe accept summarize # promote just one
That is the everyday loop: CI goes red, agentprobe diff shows exactly which sentence moved, agentprobe accept blesses the new normal. No hand-editing JSON, no blind update of everything.
Comparison Modes
| Mode | How it works | When to use |
|---|---|---|
exact (default) |
String equality after serialization | Deterministic agents, structured outputs |
semantic |
Cosine similarity via sentence-transformers (pip install agentpoke[semantic]) |
Non-deterministic LLM outputs |
How It Compares
| Feature | AgentProbe | DeepEval | Promptfoo |
|---|---|---|---|
| pytest native | Yes (plugin) | Separate runner | CLI only |
| Snapshot baselines | Yes | No | No |
| Semantic comparison | Yes | Yes | Yes |
| Mock LLM | Yes (built-in) | No | Partial |
| Tool call assertions | Yes | No | No |
| Schema validation | Yes (Pydantic) | Partial | No |
| Cloud required | No | Optional | No |
| Config format | Python code | Python code | YAML |
GitHub Actions
- name: Run agent tests
run: |
pip install agentpoke
pytest tests/ -v
Commit .agentprobe/snapshots/ so CI can compare against them.
FAQ
Do I need an API key?
No. MockLLM gives you deterministic tests without any API calls. Testing against a real LLM needs that provider's key, which is your agent's dependency, not AgentProbe's.
What about flaky tests from non-deterministic outputs?
Use semantic mode with an appropriate threshold, or use MockLLM to make the underlying LLM deterministic.
Does it work with LangChain / CrewAI / AutoGen? Yes. AgentProbe tests your agent's output, not its internals. Call your agent inside the test function and return the result.
Roadmap
Shipped: async agent tests, tool-call assertions (presence, count bounds, ordering, forbidden-argument checks), multi-step tracing, cost tracking via TokenTracker, in-terminal visual diffs for snapshot mismatches, pytest-xdist parallel runs with atomic snapshot writes, and pattern-based secret scrubbing for snapshots.
Planned:
- Interactive snapshot review — an
--agentprobe-reviewmode that walks each changed snapshot and lets you accept or reject it one at a time. - Framework adapters — first-class step capture for LangChain, LlamaIndex, and the OpenAI Assistants API.
- Offline semantic mode — a local embedding backend, so threshold checks need no API call per assertion.
Contributing
Contributions welcome. If you're testing AI agents in production and have ideas for what's missing, open an issue.
Related Projects
- CoreCoder — understand how a coding agent really works by reading the whole ~1k-line engine end to end.
- RepoWiki — a guided wiki and where-to-start reading path for unfamiliar codebases, a self-hostable DeepWiki alternative.
- LiteBench — benchmark any LLM in one command: HumanEval, GSM8K and MMLU built in, plus your own tasks.
- agentcikit — the CI safety layer for LLM agents: replay runs, fence tool calls, and triage failures before they ship.
License
Stop shipping untested agents.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agentpoke-0.3.0.tar.gz.
File metadata
- Download URL: agentpoke-0.3.0.tar.gz
- Upload date:
- Size: 223.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
29dc7e7a1f69eb9822a6ece6ae83fecb86941365c33a9ff06e12ed778f9273e5
|
|
| MD5 |
152bafdd7ecd602732b054686bfd73eb
|
|
| BLAKE2b-256 |
0cc5d2733d91e1ae26af02266fef58d244313730a1fc7c205f3af7f8dd32a6b1
|
File details
Details for the file agentpoke-0.3.0-py3-none-any.whl.
File metadata
- Download URL: agentpoke-0.3.0-py3-none-any.whl
- Upload date:
- Size: 23.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/6.2.0 CPython/3.14.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
05a7271290047e2ee8ad9ddf98899c3bbe5fdac76a392e62e695c1321a6dcf73
|
|
| MD5 |
09886e30bdc2fb69f7b768c745dfe5bd
|
|
| BLAKE2b-256 |
c682d265dbad05d74d6b471281ad2153f2f54d6f82baf0f0aada9742494e90b2
|