Local-first process assurance for agentic AI pipelines.
Project description
agent-assure
Quickstart · Integrations · LangGraph · Google ADK · RAG provenance · Trust boundary
Block agent process regressions that answer evals miss.
agent-assure is a local-first process assurance toolkit for agentic AI
pipelines. It produces deterministic review artifacts and CI-gate signals when
a candidate agent preserves the visible decision but changes the governed
process, evidence path, RAG retrieval provenance, or framework workflow metadata
around it.
Core thesis: Output equivalence is not process equivalence.
Teams shipping agent changes often discover too late that the answer stayed the same while the evidence trail, retrieval sources, provider boundary, or review route quietly changed.
Flagship evidence diff: the visible approval stayed stable, but the material evidence trail regressed and the CI gate blocked the candidate.
What it catches
agent-assure is built for the release-review gap between "the answer still
looks right" and "the governed process still matched the controls reviewers
expected."
| Surface | Regression it can make visible | Review signal |
|---|---|---|
| Final answer checks | Candidate keeps recommendation=approve; outcome=approve while losing the evidence path |
new_failure after fixture equivalence passes |
| LangGraph workflows | A graph update keeps the visible decision but drops required policy evidence or decision-node metadata | Privacy-filtered agent_assure metadata becomes evaluable run evidence |
| Google ADK workflows | A multi-agent ADK route keeps the visible decision and evidence but bypasses required performed review | The same framework observation contract feeds ordinary review-boundary controls |
| RAG retrieval | Same answer and same retrieval_corpus_digest, but the retrieved source backing a material claim disappears |
MATERIAL_CLAIM_MISSING_EVIDENCE |
| RAG provenance drift | Evidence links stay intact, but the retrieval corpus digest changes | provenance_only_change for review, not a blocking finding |
| Boundaries and routing | Provider, tool, review route, or redaction state changes unexpectedly | Deterministic invariant findings |
| Measured usage | Candidate uses more or less measured tokens, tool calls, retries, or declared estimated cost | Usage delta evidence beside, not instead of, governance findings |
| Streaming agents | Replayed, duplicated, or out-of-order events hide mid-run evidence removal, review bypass, or retry bursts | Idempotent ingestion, jitter-tolerant ordering, and ordinary process findings |
| CI release gates | A blocking process invariant fails before merge or release | Nonzero exit code plus local evidence packet |
The 30-second story
The flagship demo compares a passing baseline with an evidence-normalization candidate under the same deterministic fixtures:
baseline: recommendation=approve; outcome=approve
candidate: recommendation=approve; outcome=approve
decision fields: preserved
missing evidence link: claim-duration
classification: new_failure
CI gate: blocked as expected
The point is deliberately narrow and reviewable: the business decision did not
change, but the governed evidence path did. agent-assure catches that process
regression before release.
The same evaluator model is used by the RAG provenance demo and LangGraph adapter, so retrieval and graph-process regressions become reviewable in the same packet-and-gate flow.
Flagship regression at a glance
The diagram makes the gate logic explicit: fixture equivalence gates the comparison, the visible answer stays stable, and the candidate still fails the material evidence invariant.
flowchart LR
subgraph OutputCheck["Ordinary visible-output check"]
BOut["Baseline output<br/>recommendation=approve<br/>outcome=approve"]
COut["Candidate output<br/>recommendation=approve<br/>outcome=approve"]
Same["Visible answer unchanged"]
BOut --> Same
COut --> Same
end
subgraph InvariantCheck["agent-assure invariant check"]
BEv["Baseline evidence<br/>claim-duration linked"]
CEv["Candidate evidence<br/>claim-duration missing link"]
Pass["Baseline evaluation: pass"]
Fail["Candidate evaluation: fail<br/>MATERIAL_CLAIM_MISSING_EVIDENCE"]
BEv --> Pass
CEv --> Fail
end
Same --> Tension["Output unchanged<br/>but governance invariant regressed"]
Equiv["Fixture equivalence: pass"] --> Compare["Baseline-to-candidate comparison"]
Pass --> Compare
Fail --> Compare
Tension --> Compare
Compare --> NewFailure["Classification: new_failure"]
Quickstart
Published alpha package: agent-assure on PyPI.
The current published release is 0.5.0.
Run the flagship demo offline:
pip install agent-assure
agent-assure demo flagship
The demo runs with bundled deterministic fixtures. It writes local review
artifacts under .tmp/demo/flagship by default, including the generated
evidence-diff.html report previewed above as a PNG.
Try the RAG provenance demo:
agent-assure demo rag --out .tmp/demo/rag --clean
Explore the broader process-assurance fixture cases:
agent-assure demo measurement-cases --out .tmp/measurement-cases --clean
From a repository checkout, try the LangGraph expense-assurance example:
pip install "agent-assure[langgraph]"
python examples/langgraph_expense_assurance/run_example.py
If LangGraph is not installed, the example uses the same deterministic fallback stream shape so adapter and evaluator behavior remain testable without network calls or token spend.
Try the Google ADK process-assurance example:
python -m pip install -e ".[adk]"
python examples/adk_process_assurance/run_example.py
The ADK example also runs offline. It uses a synthetic ADK event transcript to show a same-decision candidate that preserves evidence while bypassing the observed human-review requirement.
Use it in GitHub Actions:
name: agent-assure
on: [pull_request]
jobs:
assure:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
- run: python -m pip install agent-assure==0.5.0
- uses: acblabs/agent-assure/.github/actions/agent-assure@v0.5.0
with:
suite: examples/prior_auth_synthetic/suite.yaml
baseline-variant: examples/prior_auth_synthetic/variants/baseline.yaml
candidate-variant: examples/prior_auth_synthetic/variants/candidate_evidence_normalization.yaml
report-mode: full
The composite action lives at
.github/actions/agent-assure/action.yml. Use report-mode: full when you
want complete review artifacts and report-mode: fail-fast when you want
shorter blocking feedback. Point suite, baseline-variant, and
candidate-variant at files in your repository; the paths above are this
project's bundled fixtures, shown as a runnable illustration. The bundled
candidate is expected to fail because it keeps the same visible decision while
dropping a material evidence link.
Integrations
Streaming event ingestion
Status: experimental.
agent-assure can ingest privacy-filtered JSONL event streams from asynchronous
or multi-agent runtimes, then evaluate the projected run with the same
expectation engine used by fixture, RAG, and framework-adapter flows.
agent-assure stream ingest examples/streaming_process_regression/events/candidate_evidence_removed.jsonl \
--sequence-scope global \
--out .tmp/streaming/stream-run.json
agent-assure stream evaluate .tmp/streaming/stream-run.json \
--suite examples/streaming_process_regression/suite.yaml \
--out-dir .tmp/streaming/report
Stream ingestion is intentionally explicit about idempotency and arrival jitter:
- Idempotency: the sequencing contract declares the composite key. For
global streams it is
run_id + sequence_number; for producer-local streams it isrun_id + producer_id|node_id|span_id + sequence_number. Duplicate events with the same composite key,event_id, and canonical payload digest are deduplicated and counted. Conflicting duplicateevent_idor digest values fail closed. At-least-once producers should redeliver the same logical event with stable timestamp and privacy-filtered payload fields. - Jitter handling: accepted events are persisted in deterministic order by run ID, sequence number, timestamp, and digest for globally sequenced streams. Producer-local streams use timestamp to merge independent producer counters, so timestamped events are required for that mode. Out-of-order arrival does not change the projected process trajectory.
The stream projection derives evidence-link state, review-route state, retry/rate-limit counters, usage summaries, and ordered span-plan events. It does not persist raw prompts, raw tool arguments, raw token chunks, or unredacted model output. The bundled retry-burst stream is review evidence by default, not a blocking gate failure unless the suite declares a retry-related expectation or policy.
LangGraph
Status: experimental.
The LangGraph adapter reads only privacy-filtered agent_assure metadata from
LangGraph events or streamed node updates. It ignores raw event data.input,
data.output, messages, completions, and tool arguments. Application nodes
emit compact labels or digests, then agent-assure converts those observations
into ordinary AgentRunRecord artifacts for the same evaluator used by fixture
and CI flows.
{
"agent_assure": {
"case_id": "lg-exp-001",
"event_type": "tool_call",
"sequence_number": 2,
"tool_name": "expense_policy_lookup",
"evidence_refs": ["ref-expense-policy-v3"],
"redaction_state": "redacted",
"privacy_filtered_attributes": {
"policy_version": "expense-policy-v3"
},
}
}
The included examples/langgraph_expense_assurance case keeps the same final
recommendation and review route across baseline and candidate. The candidate
omits the required policy evidence reference, so deterministic evaluation
blocks the process regression.
See docs/integrations/langgraph.md for the
adapter contract and current version notes.
Google ADK
Status: experimental.
The Google ADK adapter reads privacy-filtered agent_assure metadata from
ADK-style event mappings or event objects. Real ADK apps should prefer
custom_metadata for event labels and actions.state_delta for state-update
events. It ignores raw event content, message parts, function-call arguments,
completions, and unredacted summaries. Like the LangGraph adapter, it produces
ordinary FrameworkObservation and AgentRunRecord artifacts; evaluation
remains framework-neutral.
{
"custom_metadata": {
"agent_assure": {
"case_id": "adk-benefit-001",
"event_type": "review_route",
"sequence_number": 3,
"node_name": "review_agent",
"review_route": "clinical_review",
"redaction_state": "redacted",
"privacy_filtered_attributes": {
"human_review_required": "true",
"human_review_performed": "true"
},
}
}
}
The included examples/adk_process_assurance case keeps the same final
recommendation and policy evidence across baseline and candidate. The candidate
changes to an automatic path and reports human_review_required=false, so
deterministic evaluation blocks the process regression. The route token remains
observable process evidence; this minimal suite gates on human-review routing
and performed-review observation, not route-string equality.
See docs/integrations/google_adk.md for
the adapter contract and current version notes.
RAG provenance
The RAG demo stages a synthetic prior-authorization case with committed policy chunks, a corpus manifest, scaled-integer cached vectors, retrieval outputs, and counterfactual query-family fixtures.
agent-assure demo rag --out .tmp/demo/rag --clean
Expected punchline:
output equivalence: preserved
retrieval corpus digest: unchanged
missing evidence link: claim-duration
classification: new_failure
CI gate: blocked as expected
The hero candidate keeps recommendation=approve; outcome=approve and the same
retrieval_corpus_digest, but drops the retrieved source supporting
claim-duration; the existing material-claim evidence invariant catches the
regression.
| Candidate shape | Decision | Corpus digest | Evidence links | Classification |
|---|---|---|---|---|
| Reranker regression | Preserved | Unchanged | Missing claim-duration source |
new_failure |
| Corpus-version skew | Preserved | Changed | Preserved | provenance_only_change |
The counterfactual query-family fixtures use committed query-vector keys and report query digests rather than raw query text. They recompute retrieval evidence support for authored variants; they do not prove semantic equivalence between natural-language queries.
See docs/demo_rag.md for the full demo boundary.
Governance crosswalks
agent-assure controls are tagged against four external governance frameworks
so reviewers can line up local process evidence with the language their
programs already use. The source mapping lives in the machine-readable
docs/threat_coverage_matrix.yaml, with CI
checks for taxonomy tag shape, pinned ATLAS IDs, and selected Markdown/YAML
consistency.
| Framework | What is mapped | Crosswalk |
|---|---|---|
| NIST AI RMF | Controls tagged by Govern / Map / Measure / Manage function |
NIST AI RMF crosswalk |
| OWASP Top 10 for LLM Applications 2025 | Controls tagged by related LLM01–LLM10 risk IDs |
OWASP LLM Top 10 crosswalk |
| ISO/IEC 42001 | Controls tagged by reviewer-facing concept areas | ISO/IEC 42001 crosswalk |
| MITRE ATLAS 2026.06 | Controls mapped to adversary tactics and techniques with a stated mapping strength | MITRE ATLAS crosswalk |
These are planning crosswalks and review aids, not framework conformance, complete coverage, or third-party assurance claims. Scope limits and declared gaps are kept visible where the current mapping records them.
How agent-assure is different
agent-assure is a local-first assurance layer designed for agentic AI release
review. It runs where engineering changes are already reviewed, writes evidence
into the caller's workspace, and can block a PR or release when a declared
process invariant fails.
| Common pattern | agent-assure approach |
|---|---|
| Dashboard waits for a human to notice drift | Local CLI writes review artifacts where the change is built |
| A separate platform owns the release workflow | pip install and a composite GitHub Action fit ordinary CI |
| Approval depends on manual dashboard review | Deterministic exit codes can block PRs or releases |
| Final-answer checks miss changed evidence paths | Process invariants check evidence, boundaries, routing, redaction, and provenance |
| Full-lifecycle governance platform required | Focused release-review evidence and CI gates |
| Code diffs or output similarity stand in for process review | Deterministic process invariants, plus protocol-bound stochastic review for live behavior |
| Artifacts stay in a hosted dashboard | Portable JSON, Markdown, static HTML, digests, and evidence packets stay in the workspace |
| Broad trust claims blur the boundary | Explicit boundary: measured evidence for human review |
Why it matters:
- Local-first review: teams can inspect evidence packets, Markdown reports, and static HTML diffs from the same workspace where the change was built.
- CI-native release gates: gate logic ships in the package and returns ordinary exit codes, so no external service decides pass or fail.
- Process-aware regression checks: a candidate can keep the same visible decision while losing evidence, changing review routing, or crossing a provider/tool boundary.
- Protocol-bound live review: when live provider behavior is involved, declared protocols preserve repeated observations, clustering, interval bounds, paired tests, drift summaries, and trajectory signals instead of relying only on deterministic diffs or output similarity.
Who it is for
AI leaders
An agent change can pass every answer-quality eval and still quietly stop
citing the policy that justified its decision. That regression may not show up
in a final-answer test; agent-assure turns silent process drift into a
blocking CI signal before release.
- Hidden process drift: surface lost evidence links, changed review routes, provider/tool boundary changes, redaction changes, retries, and provenance drift.
- Release evidence: produce evidence packets, Markdown reports, and static HTML evidence diffs reviewers can inspect, archive, or attach to release review.
- Eval complement: answer-quality evals ask whether the response is good;
agent-assureasks whether the governed path to that response still matches declared controls.
AI engineers
Use agent-assure when you need reproducible checks around agent pipelines,
framework integrations, and retrieval systems.
- Strict artifacts: compile YAML suites and live protocols into strict JSON artifacts.
- Offline fixtures: run baseline and candidate variants with no provider API key, network call, or token spend.
- Controlled comparisons: compare only after fixture equivalence passes, so verdicts are tied to controlled input material rather than incidental drift.
- Shared run model: project LangGraph metadata, fixture runs, and RAG provenance into the same run-record and evaluator model.
AI security engineers
Use agent-assure when review needs to inspect observable controls without
persisting raw sensitive material in reports.
- Observable controls: check material claim-evidence links, provider/tool boundaries, review routing, redaction state, and policy controls.
- Reviewable provenance: inspect fixture manifest digests, retrieval corpus digests, source IDs, query digests, artifact digests, and dependency inventory.
- Explicit boundary: keep findings as release-review evidence, not safety, compliance, clinical, or provider-quality decisions.
What it produces
The flagship run writes artifacts that reviewers can inspect, archive, attach to release review, or upload from CI:
.tmp/demo/flagship/demo-summary.json.tmp/demo/flagship/baseline-report/evaluation-report.md.tmp/demo/flagship/evidence-report/evaluation-report.md.tmp/demo/flagship/comparison-report/comparison-report.md.tmp/demo/flagship/ci-report/evidence-packet.json.tmp/demo/flagship/evidence-diff.html
The evidence diff is a single local HTML file with inline CSS and escaped dynamic content. It does not load external JavaScript, CSS, fonts, or network resources.
Evidence packets can also include summaries, limitations, artifact digests, dependency inventory, environment context, release manifests, measured usage summaries, declared estimated cost deltas, and CI-gate state.
Architecture
At the highest level, agent-assure turns declared expectations and observed
agent behavior into local release-review evidence:
flowchart LR
A[Declared expectations] --> B[Compiled suite]
LG[LangGraph metadata] --> R[Run records]
ADK[Google ADK metadata] --> R
RAG[RAG fixtures and provenance] --> R
F[Fixture or live runs] --> R
B --> E[Evaluate controls]
R --> E
E --> C[Compare process evidence]
C --> P[Evidence packet]
P --> G[CI gate]
G -->|blocking regression| X[Release blocked]
The broader toolkit includes YAML authoring, strict schemas, canonical digests,
fixture and live execution paths, privacy-filtered reporting, release replay,
and optional OpenTelemetry-aligned span-plan export. See
docs/architecture.md for the implementation map.
Detailed walkthroughs
Five-minute fixture walkthrough
The full command-by-command showcase lives in
docs/showcase.md. It compiles the synthetic
prior-authorization suite, runs baseline and candidate variants under identical
fixtures, evaluates both run sets, compares them, builds an evidence packet,
and shows the expected new_failure gate for the missing claim-duration
evidence link.
Small generic example
The expense-approval example is a compact non-healthcare suite using the same
offline fixture and expectation method. See
docs/demo_expense.md and
examples/expense_approval_minimal/.
Schemas and development
Schema changes are versioned. Development work uses schemas/unreleased/.
Stable releases freeze a copy into schemas/vX.Y.Z/. The release gate verifies
the latest frozen schema directory, while schema staging exports the current
development schema surface to schemas/unreleased/.
From a repository checkout:
pip install -e ".[dev]"
git config core.hooksPath .githooks
python scripts/check_docs_alignment.py
ruff check .
mypy src scripts
pytest
python -m build
Dependency locking for release builds is documented in
docs/dependency_locking.md. Release bundle
reproduction, SBOM generation, and cosign verification are documented in
docs/release_evidence.md.
The installed package includes bundled deterministic examples for reproducible
local demos. The top-level examples/ tree mirrors those packaged resources
for repository-oriented docs and tests; scripts/check_packaged_examples.py
keeps the copies aligned. They are not a stable extension API; see
docs/api_surface.md.
Claim boundary and what this is not
agent-assure produces local review evidence, traceability, evidence mapping,
artifact digests, and CI-gate signals. It does not replace legal, regulatory,
clinical, provider-quality, model-quality, or business-impact review.
This project is not a compliance attestation.
It is not a safety claim.
The governance crosswalks above map local controls to NIST AI RMF, the OWASP LLM Top 10, ISO/IEC 42001, and MITRE ATLAS as planning and review aids. They do not establish conformance with, complete coverage of, or third-party acceptance by any of those frameworks.
agent-assure is |
agent-assure is not |
|---|---|
| Release-review evidence for declared process expectations | A substitute for legal, regulatory, clinical, provider-quality, or business-impact review |
| A deterministic and protocol-bound measurement toolkit | A safety determination |
| A way to surface evidence, routing, redaction, boundary, and provenance regressions | A general model-quality benchmark |
| A local artifact and CI-gate workflow | A hosted governance platform or legal approval workflow |
Live results remain bounded by the declared protocol, data boundary, provider/model configuration, and execution window. They are not general model-quality, safety, or clinical-validation claims.
Learn more
- Audience: AI leaders · Engineers
- Demos: Flagship demo · RAG provenance demo
- Integrations: LangGraph integration · Google ADK integration
- Assurance model: What this measures · Evidence diff
- Security and boundaries: Threat model · Current claim boundary
- Governance crosswalks: NIST AI RMF · OWASP LLM Top 10 · ISO/IEC 42001 · MITRE ATLAS
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agent_assure-0.5.0.tar.gz.
File metadata
- Download URL: agent_assure-0.5.0.tar.gz
- Upload date:
- Size: 755.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b24ccdc6c1ddf5352794b64a3eeec9c4003c31d106d38a59414aff1fd422cff5
|
|
| MD5 |
86f834b743ad7ec250de3029fbbca768
|
|
| BLAKE2b-256 |
f0ee4116c86b667f7c88f33c6431c96bfc81d3fd6009dcf9975d51c9f43f8020
|
Provenance
The following attestation bundles were made for agent_assure-0.5.0.tar.gz:
Publisher:
release.yml on acblabs/agent-assure
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
agent_assure-0.5.0.tar.gz -
Subject digest:
b24ccdc6c1ddf5352794b64a3eeec9c4003c31d106d38a59414aff1fd422cff5 - Sigstore transparency entry: 2188123873
- Sigstore integration time:
-
Permalink:
acblabs/agent-assure@c59934c254fd01a0437f90070f3beae45eac3e42 -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/acblabs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c59934c254fd01a0437f90070f3beae45eac3e42 -
Trigger Event:
push
-
Statement type:
File details
Details for the file agent_assure-0.5.0-py3-none-any.whl.
File metadata
- Download URL: agent_assure-0.5.0-py3-none-any.whl
- Upload date:
- Size: 641.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.13
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dcceb7366a67f8b616e467b3e169b759ddf1adfc306f6350c58e565f54b1b62f
|
|
| MD5 |
68096717c67baf143925aa4e03ed5265
|
|
| BLAKE2b-256 |
95c1266aa290b18ceda944f5476474eccac6c79352d4b27c48a9133cd554fa38
|
Provenance
The following attestation bundles were made for agent_assure-0.5.0-py3-none-any.whl:
Publisher:
release.yml on acblabs/agent-assure
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
agent_assure-0.5.0-py3-none-any.whl -
Subject digest:
dcceb7366a67f8b616e467b3e169b759ddf1adfc306f6350c58e565f54b1b62f - Sigstore transparency entry: 2188123894
- Sigstore integration time:
-
Permalink:
acblabs/agent-assure@c59934c254fd01a0437f90070f3beae45eac3e42 -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/acblabs
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@c59934c254fd01a0437f90070f3beae45eac3e42 -
Trigger Event:
push
-
Statement type: