Skip to main content

Local-first process assurance for agentic AI pipelines.

Project description

agent-assure

PyPI Python CI License: MIT Offline demo

Catch agent process regressions that final-answer evals miss.

agent-assure is a local-first process assurance toolkit for agentic AI pipelines. It produces deterministic review artifacts and CI-gate signals when a candidate agent preserves the visible decision but changes the governed process around it.

Core thesis: Output equivalence is not process equivalence.

Agent Assure evidence diff report showing same approval, broken evidence trail, and CI blocked.

Flagship evidence diff: the visible approval stayed stable, but the material evidence trail regressed and the CI gate blocked the candidate.

The 30-second story

The flagship demo compares a passing baseline with an evidence-normalization candidate under the same deterministic fixtures:

baseline:  recommendation=approve; outcome=approve
candidate: recommendation=approve; outcome=approve

decision fields: preserved
missing evidence link: claim-duration
classification: new_failure
CI gate: blocked as expected

The point is deliberately narrow and reviewable: the business decision did not change, but the governed evidence path did. agent-assure catches that process regression before release.

Flagship regression at a glance

The diagram makes the gate logic explicit: fixture equivalence gates the comparison, the visible answer stays stable, and the candidate still fails the material evidence invariant.

flowchart LR
    subgraph OutputCheck["Ordinary visible-output check"]
        BOut["Baseline output<br/>recommendation=approve<br/>outcome=approve"]
        COut["Candidate output<br/>recommendation=approve<br/>outcome=approve"]
        Same["Visible answer unchanged"]
        BOut --> Same
        COut --> Same
    end

    subgraph InvariantCheck["agent-assure invariant check"]
        BEv["Baseline evidence<br/>claim-duration linked"]
        CEv["Candidate evidence<br/>claim-duration missing link"]
        Pass["Baseline evaluation: pass"]
        Fail["Candidate evaluation: fail<br/>MATERIAL_CLAIM_MISSING_EVIDENCE"]
        BEv --> Pass
        CEv --> Fail
    end

    Same --> Tension["Output unchanged<br/>but governance invariant regressed"]
    Equiv["Fixture equivalence: pass"] --> Compare["Baseline-to-candidate comparison"]
    Pass --> Compare
    Fail --> Compare
    Tension --> Compare

    Compare --> NewFailure["Classification: new_failure"]

Start in one command

pip install agent-assure
agent-assure demo flagship

The demo runs offline with bundled deterministic fixtures. It writes local review artifacts under .tmp/demo/flagship by default, including the generated evidence-diff.html report previewed above as a PNG.

How agent-assure is different

agent-assure is a local-first assurance layer designed for agentic AI release review. It runs where engineering changes are already reviewed, writes evidence into the caller's workspace, and can block a PR or release when a declared process invariant fails.

Common governance-tool pattern agent-assure approach
Central hosted dashboard Local-first CLI and workspace files
Separate sign-in and platform workflow Runs directly from pip install and CI
Manual or dashboard-driven approval Deterministic exit codes that block PRs or releases
Final-answer or trace-inspection focus Process invariants for evidence, boundaries, routing, redaction, and provenance
Full-lifecycle governance platform Focused release-review evidence and CI gates
Code/text diffs or output-similarity scoring Deterministic process invariants, plus protocol-bound stochastic review for live behavior
Centralized dashboard artifacts Portable, versionable artifacts: JSON, Markdown, static HTML, digests, and evidence packets
Broad trust claims Explicit boundary: measured evidence for human review

Why it matters:

  • Local-first review: teams can inspect the evidence packet, Markdown reports, and static HTML diff from the same workspace where the change was built.
  • CI-native release gates: the gate logic ships in the package and returns ordinary exit codes, so nothing needs an external service to decide pass or fail.
  • Process-aware regression checks: a candidate can keep the same visible decision while losing evidence, changing review routing, or crossing a provider/tool boundary.
  • Statistical live review: when live provider behavior is involved, declared protocols preserve repeated observations, clustering, interval bounds, paired tests, drift summaries, and trajectory signals instead of relying only on deterministic diffs or output similarity.

For AI leaders

Use agent-assure when a team needs release-review evidence that an agent change preserved declared expectations around the process, not just the answer.

  • It makes hidden process regressions visible: evidence links, review routing, provider/tool boundaries, redaction behavior, retries, and provenance.
  • It creates evidence packets, Markdown reports, and a static HTML evidence diff that reviewers can inspect, archive, or attach to release review.
  • It complements answer-quality evals. Those evals ask whether the response is good; agent-assure asks whether the governed path to that response still matches the controls reviewers expected.

For AI and security engineers

agent-assure is built around declared, observable controls:

  • YAML suites and live protocols compile to strict JSON artifacts.
  • Fixture mode runs baseline and candidate variants offline with no provider API key, network call, or token spend.
  • Evaluators check expectations, policy controls, privacy filters, tool/provider boundaries, review routing, and material claim-evidence links.
  • Baseline-to-candidate comparison runs after fixture equivalence passes, so verdicts are tied to controlled input material rather than incidental drift.
  • Evidence packets bundle summaries, limitations, artifact digests, dependency inventory, environment context, and CI-gate state.
  • CI commands exit nonzero on blocking findings, while the demo wrapper treats the known blocked candidate as a successful demonstration.

What it produces

The flagship run writes reproducible local review artifacts:

  • .tmp/demo/flagship/demo-summary.json
  • .tmp/demo/flagship/baseline-report/evaluation-report.md
  • .tmp/demo/flagship/evidence-report/evaluation-report.md
  • .tmp/demo/flagship/comparison-report/comparison-report.md
  • .tmp/demo/flagship/ci-report/evidence-packet.json
  • .tmp/demo/flagship/evidence-diff.html

The evidence diff is a single local HTML file with inline CSS and escaped dynamic content. It does not load external JavaScript, CSS, fonts, or network resources.

Architecture

At the highest level, agent-assure turns declared expectations into local release-review evidence:

flowchart LR
  A[Declared expectations] --> B[Baseline and candidate runs]
  B --> C[Evaluate controls]
  C --> D[Compare process evidence]
  D --> E[Evidence packet]
  E --> F[CI gate]
  F -->|blocking regression| G[Release blocked]

The broader toolkit includes YAML authoring, strict schemas, canonical digests, fixture and live execution paths, privacy-filtered reporting, release replay, and optional OpenTelemetry-aligned span-plan export. See docs/architecture.md for the implementation map.

Five-minute fixture walkthrough

Run these commands one at a time from the repository root. The candidate evaluation, comparison, and CI commands are expected to exit 1 after writing their reports. See docs/showcase.md for a GitHub Actions snippet that asserts those expected failures in set -e contexts.

pip install -e ".[dev]"
mkdir -p .tmp/showcase
agent-assure suite compile examples/prior_auth_synthetic/suite.yaml --out .tmp/showcase/prior-auth.compiled.json --manifest .tmp/showcase/prior-auth.fixtures.json
agent-assure suite run .tmp/showcase/prior-auth.compiled.json --variant examples/prior_auth_synthetic/variants/baseline.yaml --manifest .tmp/showcase/prior-auth.fixtures.json --out .tmp/showcase/prior-auth.baseline.json
agent-assure suite run .tmp/showcase/prior-auth.compiled.json --variant examples/prior_auth_synthetic/variants/candidate_evidence_normalization.yaml --manifest .tmp/showcase/prior-auth.fixtures.json --out .tmp/showcase/prior-auth.evidence-candidate.json
agent-assure evaluate .tmp/showcase/prior-auth.baseline.json --suite .tmp/showcase/prior-auth.compiled.json --out-dir .tmp/showcase/baseline-report
agent-assure evaluate .tmp/showcase/prior-auth.evidence-candidate.json --suite .tmp/showcase/prior-auth.compiled.json --out-dir .tmp/showcase/evidence-report
agent-assure compare .tmp/showcase/prior-auth.baseline.json .tmp/showcase/prior-auth.evidence-candidate.json --suite .tmp/showcase/prior-auth.compiled.json --out-dir .tmp/showcase/comparison-report
agent-assure ci .tmp/showcase/prior-auth.evidence-candidate.json --suite .tmp/showcase/prior-auth.compiled.json --baseline .tmp/showcase/prior-auth.baseline.json --out-dir .tmp/showcase/ci-report --report-mode full

The baseline evaluation exits 0 and writes a pass summary with ten evaluated cases and zero blocking findings. The candidate evaluation writes one blocking finding for shared-source-multi-claim with reason code MATERIAL_CLAIM_MISSING_EVIDENCE.

The comparison report classifies the change as new_failure under passing fixture equivalence. For the failing case, the baseline and candidate both keep recommendation=approve; outcome=approve; the material regression is the missing claim-duration evidence link.

After reports exist, an evidence packet can also be built and gated from summaries:

agent-assure packet build .tmp/showcase/evidence-report/evaluation-summary.json --comparison .tmp/showcase/comparison-report/comparison-summary.json --out .tmp/showcase/evidence-packet.json
agent-assure ci gate .tmp/showcase/evidence-packet.json

For this known failing candidate, both the CI command and packet gate are expected to exit 1. The CI command writes JSON/Markdown reports, evidence-packet.json, evidence-packet.md, dependency-inventory.json, release-artifact-manifest.json, and ci-diagnostics.json.

Small generic example

The expense-approval example is a compact non-healthcare suite that uses the same offline fixture and expectation method. It is a generic demonstration, not a benchmark.

agent-assure suite compile examples/expense_approval_minimal/suite.yaml --out .tmp/expense.compiled.json --manifest .tmp/expense.fixtures.json
agent-assure suite run .tmp/expense.compiled.json --variant examples/expense_approval_minimal/variants/baseline.yaml --manifest .tmp/expense.fixtures.json --out .tmp/expense.baseline.json
agent-assure suite run .tmp/expense.compiled.json --variant examples/expense_approval_minimal/variants/candidate_provider_policy.yaml --manifest .tmp/expense.fixtures.json --out .tmp/expense.candidate.json
agent-assure evaluate .tmp/expense.baseline.json --suite .tmp/expense.compiled.json --out-dir .tmp/expense.baseline-report
agent-assure evaluate .tmp/expense.candidate.json --suite .tmp/expense.compiled.json --out-dir .tmp/expense.candidate-report

The baseline evaluation exits 0. The provider-policy candidate is expected to exit 1 with deterministic provider, outcome, and human-review control findings.

Schemas and development

Schema changes are versioned. Development work uses schemas/unreleased/. Stable releases freeze a copy into schemas/vX.Y.Z/. The release gate verifies the latest frozen schema directory, while schema staging exports the current development schema surface to schemas/unreleased/.

From a repository checkout:

pip install -e ".[dev]"
git config core.hooksPath .githooks
python scripts/check_docs_alignment.py
ruff check .
mypy src scripts
pytest
python -m build

Dependency locking for release builds is documented in docs/dependency_locking.md. Release bundle reproduction, SBOM generation, and cosign verification are documented in docs/release_evidence.md.

The installed package includes bundled deterministic examples for reproducible local demos. The top-level examples/ tree mirrors those packaged resources for repository-oriented docs and tests; scripts/check_packaged_examples.py keeps the copies aligned. They are not a stable extension API; see docs/api_surface.md.

Claim and trust boundary

agent-assure produces local review evidence, traceability, evidence mapping, artifact digests, and CI-gate signals. It does not replace legal, regulatory, clinical, provider-quality, model-quality, or business-impact review.

This project is not a compliance attestation.

It is not a safety claim.

Live results remain bounded by the declared protocol, data boundary, provider/model configuration, and execution window. They are not general model-quality, safety, or clinical-validation claims.

Learn more

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

agent_assure-0.3.1.tar.gz (506.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

agent_assure-0.3.1-py3-none-any.whl (388.3 kB view details)

Uploaded Python 3

File details

Details for the file agent_assure-0.3.1.tar.gz.

File metadata

  • Download URL: agent_assure-0.3.1.tar.gz
  • Upload date:
  • Size: 506.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for agent_assure-0.3.1.tar.gz
Algorithm Hash digest
SHA256 06b05649b77bbb0df52f155174449617d381783297393d12c1f84ee19aa5c9fa
MD5 c75d7ba2fa6a1f463d09df31099f1364
BLAKE2b-256 b475542f94a6852530353e1451fa60c8a8ecdd63e5b7db95108b74f28733102c

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_assure-0.3.1.tar.gz:

Publisher: release.yml on acblabs/agent-assure

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file agent_assure-0.3.1-py3-none-any.whl.

File metadata

  • Download URL: agent_assure-0.3.1-py3-none-any.whl
  • Upload date:
  • Size: 388.3 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.13

File hashes

Hashes for agent_assure-0.3.1-py3-none-any.whl
Algorithm Hash digest
SHA256 00b38ad8a4846883951f715ecb341206a26b1d82563fef5646fb62cd263dd439
MD5 a80a099b0e88f02b257fc97edb458388
BLAKE2b-256 a83bdf8b71f95f5cffb79859d851e09babd508d9369d49269048eef5bdff1479

See more details on using hashes here.

Provenance

The following attestation bundles were made for agent_assure-0.3.1-py3-none-any.whl:

Publisher: release.yml on acblabs/agent-assure

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page