Skip to main content

AgentRunProof

CI PyPI Python License: MIT

AgentRunProof: runtime bugs deserve proofs, not screenshots

Catch OpenAI Agents SDK Runner regressions without an API key.

AgentRunProof runs deterministic scenarios against the real Runner, compares observable behavior across run and run_streamed, and writes content-addressed JSON records for stream, session, tool-linkage, and RunState resume invariants. A failing record carries the normalized counterexample observations.

The 0.3.0 source contract declares openai-agents>=0.20,<0.23 on Python 3.10–3.14. Its packaged-wheel CI matrix requires the exact 0.20.0, 0.21.0, and 0.22.0 release baselines. Published availability is shown by the PyPI badge and immutable GitHub Releases. Built-in scenarios make no model API call and require no API key.

AgentRunProof-backed reports are referenced by two merged maintainer fixes, #4413 and #4414. This is upstream diagnostic impact—not OpenAI adoption, dependency, or endorsement.

Read the five-minute RunState case study for the released failure, the recursive follow-up, and the exact before/after evidence chain. For complete usage details, start with the documentation index, Python API reference, or CLI reference.

30-second local check

python -m pip install agentrunproof
agentrunproof probe basic-tool-session-parity --certificate proof.json
agentrunproof check-certificate proof.json

Expected output:

PASS basic-tool-session-parity
  PASS    execution_outcome: OK
  PASS    stream_parity: OK
  PASS    tool_linkage: OK
  PASS    exactly_once: OK
  PASS    model_script_consumed: OK
certificate_id: sha256:...
written: proof.json
VALID sha256:... PASS

Exit 0 means PASS, 1 means an observed invariant violation, and 2 means invalid or unverifiable evidence. See the provider-free real Runner example and the OpenAI Agents integration guide.

Maintaining a downstream library? The isolated, test-only CI guide provides a copyable real-Runner contract test and an ephemeral uv matrix. Its published v0.2.0 example covers exact SDK 0.20.0 and 0.21.0; the 0.3.0 source contract adds SDK 0.22.0 and may be pinned downstream only after its immutable release and PyPI pages exist. The isolated pattern keeps AgentRunProof out of runtime metadata and the project lockfile.

For artifact review, pin an exact published version and use its matching immutable GitHub Release. Each SHA256SUMS binds the wheel and sdist. The release workflow rebuilds and byte-compares those artifacts, smoke-tests the wheel, and publishes the same verified files to PyPI through OIDC trusted publishing. The CI guide explains the remaining third-party-code trust boundary.

Where it fits

Use the SDK's public agents.testing.ScriptedModel with pytest for a focused deterministic application or SDK test. AgentRunProof delegates to ScriptedModel on supported SDKs 0.21 and 0.22 and adds reusable scenario orchestration, automatic run/run_streamed comparison, multi-phase RunState checks, cross-version evidence, and content-addressed records.

You need to… Start with
Script model responses and assert one application behavior agents.testing.ScriptedModel + pytest
Compare the same contract across runner modes or SDK versions AgentRunProof
Check approval/rejection and JSON-restored RunState flows AgentRunProof
Share a normalized record that can be checked without a provider call AgentRunProof
Evaluate model-output quality An eval framework, not AgentRunProof

What AgentRunProof checks

  • declared completion, interruption, or Runner-exception outcomes for every scenario phase;
  • post-run parity between non-streaming execution and scripted terminal-event streaming (response.output_item.done plus response.completed);
  • ordered function-call/output linkage in generated items, Session snapshots, and every model input;
  • declared counts for scenario-owned local tool invocations;
  • consumption of each deterministic model script;
  • selected public RunState transitions: JSON transport, from_json() reconstruction, restored-state equality, interruption identities, and exact approve/reject decisions;
  • direct sibling-RunState approval isolation from repeated RunResult.to_state() calls;
  • recursive approval routing through two Agent.as_tool checkpoints while preserving an untouched direct sibling state;
  • recursive approval routing after a public RunState.to_json() / RunState.from_json() boundary, with one exact approval applied to the restored interruption;
  • per-phase tool-count deltas, scenario probes, and replay of persisted tool history.
  • output-guardrail tool-pair durability, including SDK 0.22's removal of a rejected raw tool result from durable replay without depending on its replacement wording.

The terminal-event profile does not claim token/delta, timing, backpressure, or cancellation-stream equivalence. Generic handoff, retry, cancellation, max-turn, generalized snapshot-isolation, and task-cleanup contracts remain future scenarios unless a certificate explicitly names and observes them.

AgentRunProof checks SDK runtime semantics. It is not a model-quality evaluator, tracing backend, HTTP recorder, hosted service, or general agent framework.

For observability integration tests, DeterministicModel(..., emit_traces=True) emits the SDK's ordinary generation span while the real Runner emits its agent and tool spans. The default is False, and run_scenario() still disables tracing. Built-ins remain provider-free, but arbitrary scenario tools and hooks are not network-sandboxed. Use the opt-in by passing the model directly to Runner.run() or Runner.run_streamed(); any installed trace processor may export data or make network requests.

Development quickstart

python -m pip install -e ".[test,dev]"
agentrunproof --version
agentrunproof list-scenarios
agentrunproof probe basic-tool-session-parity --certificate build/basic.json
agentrunproof check-certificate build/basic.json

A successful probe exits 0; an observed invariant violation exits 1; invalid input or unverifiable evidence exits 2.

The sibling-isolation probe intentionally exposes a released SDK counterexample:

agentrunproof probe runstate-sibling-approval-isolation \
  --certificate build/runstate-sibling-isolation.json
agentrunproof check-certificate build/runstate-sibling-isolation.json

On openai-agents==0.20.0, approving one sibling state also mutates an untouched sibling; resuming that untouched state executes the protected tool. The certificate records state_fork_isolation: SIBLING_STATE_MUTATED and the associated unexpected outcome and side effect. This adjacent gap was reported on upstream PR #4409; the report is not a claim that #4409 introduced the bug.

The recursive routing probe exercises the remaining boundary after upstream #4413:

agentrunproof probe runstate-recursive-agent-tool-approval-routing \
  --certificate build/runstate-recursive-approval.json
agentrunproof check-certificate build/runstate-recursive-approval.json

It pauses a protected effect behind two Agent.as_tool edges, creates two direct sibling states, approves only one flattened interruption, and resumes both branches. On upstream commit 0b93ce8, the untouched sibling correctly remains pending but the approved sibling also remains interrupted; the focused result is recursive_approval_routing: APPROVED_NESTED_STATE_REMAINED_INTERRUPTED. A corrected runtime must finish the approved branch with exactly one effect in both runner modes while leaving the untouched branch at zero effects.

The serialized-routing probe checks the durable form of the same contract:

agentrunproof probe runstate-recursive-agent-tool-approval-serialization \
  --certificate build/runstate-recursive-approval-serialization.json
agentrunproof check-certificate build/runstate-recursive-approval-serialization.json

The initial head of upstream PR #4414 (9dc7da9) fixed the live path but remained interrupted after JSON restoration. The revised head 1725a898 passes the built-in restored-approval scenario in both runner modes and was squash-merged as 50d65f65; the upstream 24-case regression also covers approval and rejection before and after restoration across two and three nested edges. The immutable v0.1.2 comparison bundle pins the released failure, the intermediate merged behavior, and the final recursive and serialized PASS results by wheel hash and Git provenance. The v0.2.0 release carries a freshly reproduced bundle for the same causal ladder, bound to the v0.2.0 harness wheel.

For library scenarios, the top-level package exposes Scenario/ScenarioCase for one run and ScenarioPlan/ScenarioPhase/ResumeInput/StateProbe for ordered multi-run contracts, together with DeterministicModel, RecordingSession, run_scenario(), and certificate helpers. The built-in scenario and the two multi-phase historical scenarios are executable examples.

Historical falsification matrix

The development matrix uses only released SDK wheels and public runtime interfaces:

Upstream case Buggy boundary Fixed boundary Required fingerprint
#4322 0.19.4 FAIL 0.20.0 PASS session limiting must not send an orphan function output to the model
#4244 0.19.4 FAIL 0.20.0 PASS serialized approval must survive a context-overridden resume and execute once
#4125 0.19.2 streamed FAIL 0.19.3 PASS a committed tool call/output pair must survive a resumed output-guardrail tripwire

Run a local, non-canonical rehearsal with:

python scripts/run_history_matrix.py --output-directory build/history-rehearsal
agentrunproof check-history-matrix build/history-rehearsal/matrix.json

Canonical evidence is stricter: Linux x86_64 CPython 3.12, fresh environments, hash-locked wheel closures, isolated worker processes, a Python socket-deny guard during scenario execution, an exact clean Git commit, and a bundle marker written last. Artifact acquisition occurs before the network guard and is explicitly recorded as a limitation. The immutable v0.1.0 Gate 2 bundle is published under evidence/history/v1.

The 0.19.x rows are historical-only compatibility probes, not supported installations: the harness wheel is installed with --no-deps over each locked legacy SDK closure, and that dependency-metadata bypass is explicit in the canonical bundle.

Evidence and trust boundary

Certificate and history identifiers are SHA-256 addresses over canonical JSON. The independent checker rejects schema drift, non-finite or duplicate-key JSON, semantic inconsistencies, forged phase transitions, altered historical fingerprints, a missing or tampered required matrix/marker, and internally inconsistent source-state metadata. The referenced wheel is optional beside a local marker and is separately bound by CI or release artifacts.

Checking a record re-evaluates its normalized observations; it does not rerun the SDK, authenticate an untrusted publisher, or prove that the stated command executed. Public claims therefore require the clean source commit plus a visible CI or release anchor.

The private profile prevents raw observed payloads from being serialized, but values still exist in the scenario process. Its deterministic unsalted hashes are correlatable and may be dictionary-guessed for low-entropy values. Arbitrary user-defined tools, hooks, and probes are not sandboxed. Treat private records as local diagnostics and publish only reviewed synthetic evidence.

Project contract

The exact release gates and exclusions are in the project charter. The execution plan tracks releases, evidence, and external-adoption work.

AgentRunProof received its first maintainer-level citation when OpenAI Agents follow-up #4413 cited the reported checkpoint isolation defect. The next adoption target is reuse of the recursive regression fixture, an optional CI check, or a documentation reference—not a default SDK dependency. A community-tool entry was proposed on the official v0.21 testing-guide PR. The maintainer kept that guide limited to SDK-maintained APIs while explicitly welcoming future reproducible findings backed by the tool. AgentRunProof therefore remains an external project rather than an official SDK listing or dependency.

Contribute a runtime contract

Found a public-API Runner inconsistency? Open a scenario request with the exact SDK version and a minimal reproducer. Want to make it permanent? See the contribution guide and add the smallest failing scenario. For usage questions and early contract ideas, use Discussions.

If AgentRunProof belongs in your regression toolbox, star the repository so other SDK maintainers can find it.

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agentrunproof-0.3.0.tar.gz (158.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agentrunproof-0.3.0-py3-none-any.whl (99.0 kB view details)

Uploaded Python 3

File details

Details for the file agentrunproof-0.3.0.tar.gz.

File metadata

  • Download URL: agentrunproof-0.3.0.tar.gz
  • Upload date:
  • Size: 158.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentrunproof-0.3.0.tar.gz
Algorithm Hash digest
SHA256 f04624893d54d6d6708a5d24f9079b5c4e4dd3d1d2364068e73cd86ae9d5c313
MD5 9682a445135c2fb337192463c9bfe285
BLAKE2b-256 23bd72a535626417a3a4e760b4c0c8e14b2f04df3b1b547c14fbb195f1f82eda

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentrunproof-0.3.0.tar.gz:

Publisher: publish.yml on FU-max-boop/agentrunproof

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agentrunproof-0.3.0-py3-none-any.whl.

File metadata

  • Download URL: agentrunproof-0.3.0-py3-none-any.whl
  • Upload date:
  • Size: 99.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for agentrunproof-0.3.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7375d7eab4ad42377ceb9f7387e1151cc30dad87538d4f518da21afe7f41901f
MD5 bb3e171faa3f430ba7a4298b34cf9831
BLAKE2b-256 40fe2c2f25380047365358b8d6a15133305bbd3a05279794c49db97ed40b8df7

See more details on using hashes here.

Provenance

The following attestation bundles were made for agentrunproof-0.3.0-py3-none-any.whl:

Publisher: publish.yml on FU-max-boop/agentrunproof

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page