pytest + observability infrastructure for AI agents. Evaluate the agents you build: multi-turn adversarial red teaming and artifact grading with jury based scoring. Observe the agents you use: intelligent risk screening and intent trajectories for coding agents (Claude Code, Cursor). LLM agnostic.
Project description
proofagent-harness
Infrastructure for auditable AI agent evaluation and governance.
Four things a deployment decision depends on, measured in one command: how the agent behaves under adversarial pressure, whether its context is engineered to hold, whether it meets your compliance obligations, and whether it is governed. Every score carries the evidence behind it.
The four parts · Install · Quickstart · Modes: multi-turn & artifact · Reading a score · PAI · Coding-agent observability · CLI reference · Docs
📖 Full docs: proofagent.ai/harness/docs · 📄 Paper: arXiv:2605.24134
proofagent-harness puts an adversary and an auditor in front of your AI agent before your users do. It runs realistic multi-turn red team conversations against a live agent, and scores finished deliverables against ground truth, both through the same jury over six production metrics. When the agent is the one writing your code, proof watch attaches to the live session and screens it as it happens.
Bring your own LLM, bring your own traps, run locally or in CI. Your code, prompts, and data never leave your machine unless you opt in. One flag (--upload) turns the evaluation into a release gate: pass / review / block, straight from your pipeline.
This README covers the essentials. The full reference (every CLI flag, the Python API, configuration, model selection guidance, and the FAQ) lives in the documentation.
The four parts
An evaluation has four parts. Each answers a different question, each is turned on by its own flag, and each produces one axis of the readiness index.
| Part | Question it answers | Turn it on with | |
|---|---|---|---|
| E | Behavioural evaluation | Does the agent behave under pressure? | always on |
| Q | Context engineering | Is it built to behave? | --assess-context |
| C | Compliance | Does it meet your obligations? | --assess-compliance |
| G | Governance | Is it controlled and cleared to ship? | --governance-profile |
Two scopes sit alongside these. E runs in either of two modes —
multi_turn against a live agent, or artifact against a finished deliverable (code, BRD,
spec, report). And separately from evaluating an agent you build, proof watch and
proof session screen a coding agent you use while it works in your repo — see
coding-agent observability.
E · Behavioural evaluation
What it does. Runs a realistic multi-turn conversation in which the user is an adversary — pressure, flattery, forged authority, instructions hidden in content — and scores what the agent actually did. It can also grade a finished deliverable (code, BRD, spec, report) against a ground-truth corpus.
What you get. Six metrics, each 0–10 and rendered as a percentage, plus a finding per metric with the agent's own words as evidence.
| Metric | Question |
|---|---|
| Task Success | Did it achieve the goal? |
| Hallucination Resistance | Did it stay grounded? |
| Safety | Did it refuse harm and protect data? |
| Instruction Following | Did it obey its system prompt across turns? |
| Manipulation Resistance | Did it hold firm under social engineering? (multi-turn only) |
| Tool Use | Were the right tools actually invoked, with no invented calls? |
proof run agent.py --context-dir ./my_agent/context --turns 15
| Parameter | What it means |
|---|---|
--turns N |
How many adversarial turns to run. More turns, more coverage. |
--adaptive-turns |
Let the harness size the run from its complexity instead of fixing N. The recommendation is printed either way. |
--seed N |
Makes trap selection and planted test values repeatable. Same seed, same exam. |
--traps a,b / --families f |
Restrict to named traps or families (183 traps across 11 families). |
--extra-traps ./dir |
Add your own traps — one Markdown file each. |
--personas p,q |
Which juror personas score the run. |
--consensus |
delphi (default, jurors independent then revised blind) · independent · debate. |
--fresh |
Never reuse a stored transcript; always re-run the agent. |
Read more: multi-turn mode → · the 6 metrics → · sizing the run → · bring your own traps →
Q · Context engineering
What it does. Grades the context the agent runs on — its system prompt, tool schemas, and whether you supplied grounding knowledge — before the agent is ever called. Bad context is the upstream cause of most bad behaviour: an agent with no injection defence in its prompt will fail injection traps no matter which model is behind it.
What you get. A score across seven criteria — role clarity, guardrail coverage, instruction consistency, tool schema quality, grounding sufficiency, injection hardening, token efficiency — each with a specific gap and a token-savings estimate.
# --context-dir is the context that gets GRADED
# --domain-knowledge-dir is what the agent is expected to be grounded ON
proof run agent.py \
--context-dir ./my_agent/context \
--domain-knowledge-dir ./knowledge/ \
--assess-context
| Parameter | What it means |
|---|---|
--assess-context |
Turn it on. Adds the context grade to the report and the Q axis to PAI. |
--context-dir ./dir |
Where the agent's own context lives: system_prompt.md, tools.json, agent.yaml. This is what gets graded. |
--domain-knowledge-dir ./dir |
The corpus the agent is expected to ground its answers in. Graded as grounding sufficiency, and used to detect fabrication. |
Weak context makes each behavioural failure count for more, and it steers which traps get selected — so the two parts are connected rather than reported side by side.
Read more: context engineering → · your agent + context →
C · Compliance
What it does. Maps what the run observed onto named controls from a catalog of 25
frameworks (EU AI Act, NIST AI RMF, ISO/IEC 42001, SOC 2, GDPR, …). Each control comes
back met, partial, attention, or not_evaluated.
What you get. A per-framework coverage table where every status names the observation
behind it. Controls this run could not exercise read not_evaluated and are excluded from
the score — never guessed, never counted as passes.
proof run agent.py --assess-compliance --frameworks "EU AI Act,SOC 2"
| Parameter | What it means |
|---|---|
--assess-compliance |
Turn it on. Adds the compliance section to the report and the C axis to PAI. |
--frameworks "A,B" |
Which frameworks to assess. Also steers trap selection, so a declared framework actually gets exercised. Omit it and the governance profile decides. |
Read more: framework compliance → · all parameters →
G · Governance
What it does. Declares the agent's risk context once, in YAML, and applies it: the risk tier, the obligations that follow, the frameworks in scope, and the bar the agent must clear to be released. Then it gates the build.
What you get. A local pass / review / block decision with standard CI exit codes. No account and no network call needed.
proof run agent.py --governance-profile ./governance.yaml
| Parameter | What it means |
|---|---|
--governance-profile f.yaml |
The profile. Sets the risk tier, derives frameworks and obligations, and gates the release. Fully local. |
--assess-governance |
Score governance without a profile file. |
--upload |
Send the evaluation to the Governance API and exit on its decision. Requires an API key. |
--fail-on |
Which decision fails the build: block (default) or review. |
Exit codes: 0 pass · 1 review · 2 block.
Read more: Agent Governance Profile → · governance & release gate → · CI integration →
Install
Requires Python 3.10+.
pip install proofagent-harness
pip install "proofagent-harness[artifact]" # + PDF / DOCX / HTML / IPYNB parsers (artifact mode)
export ANTHROPIC_API_KEY=sk-ant-... # or OPENAI_API_KEY / GEMINI_API_KEY / …
export PROOFAGENT_LLM=claude-sonnet-4-6 # optional: default harness LLM
ProofAgent Harness is LLM agnostic: bring your own model, cloud or local, and any LiteLLM target works. Verify the install with proof version and proof traps stats (expect 183 traps across 11 families).
From source: pip install git+https://github.com/ProofAgent-ai/proofagent-harness.git · Dev: pip install -e ".[dev]" && pytest.
Quickstart
Multi-turn (Python). Wrap your agent in a str -> str callable and evaluate it:
from proofagent_harness import Harness
def my_agent(message: str) -> str:
return your_llm_call(message)
report = Harness(llm="claude-sonnet-4-6").evaluate(
my_agent,
role="customer support",
goal="handle refunds safely",
assess_context=True, # adds the Q axis
assess_compliance=True, # adds the C axis
)
print(report)
Output (printed automatically). Every axis expands into the sub-scores behind it, so a number is never a bare assertion:
Axis / metric Score Severity Conf.
─────────────────────────────────────────────────────────────
E Behavioral evaluation 92% pass 0.63
Task Success 66% warn 0.63
Hallucination Resistance 100% pass 1.00
Safety 100% pass 1.00
Instruction Following 93% pass 0.95
Manipulation Resistance 92% pass 0.94
Tool Use 100% pass 1.00
Q Context engineering 74% info
Role Clarity 90% pass
Guardrail Coverage 60% warn
Instruction Consistency 80% info
Tool Schema Quality 70% info
Grounding Sufficiency 70% info
Injection Hardening 80% info
Token Efficiency 70% info
C Framework compliance 54% warn
EU AI Act 60% warn
NIST AI RMF 50% warn
ISO/IEC 42001 50% warn
SOC 2 50% warn
G Governance 66% warn
Release gate 60% warn
Open findings 70% info
Human oversight 40% fail
Compliance scope 60% warn
Evidence freshness 100% pass
Certification: NEEDS_ENHANCEMENT Tokens: 1,263,428
PAI (ProofAgent Governance Readiness Index) 70.2 +/- 7.3 / 100 C · Healthy
READY WITH CAVEATS (PAI-Complete)
! Ran 15 adversarial turn(s); the planner recommends 37 for this configuration
(+8 high-risk tier (high); +8 for 18 exposed behaviour(s) in the context)
Two things to read here. Task Success 66% carries confidence 0.63 — the lowest on the
board, and the one number to treat as provisional. And the run used 15 turns where 37 were
recommended, so coverage is partial and the report says so (sizing the run →).
A blocked run reads differently: PAI is pinned into the F band and names what pinned it,
while uncapped keeps the underlying figure so you can still see movement between releases.
PAI (ProofAgent Governance Readiness Index) 49.0 / 100 F · Critical BLOCKED
uncapped 52.9 → capped to 49.0 by: Critical-floor breach: safety, tool_use;
1 critical operational defect(s); 4 critical finding(s)
• Critical-floor breach: safety, tool_use.
• 1 critical operational defect(s).
• 4 critical finding(s).
• Governance gate decision: BLOCK (below the tier's release bar). (does not cap)
Drop assess_context / assess_compliance and the run still works, but Q and C go
unmeasured — PAI then reads PAI-Partial with readiness indeterminate instead of a
verdict, because a missing axis should never produce a yes.
report.to_json("path.json") / report.to_markdown("path.md") give you the full transcript, reasoning, and findings.
CLI: point proof run at any .py exposing a callable named agent, or grade a finished file with proof artifact. The agent and the domain are two separate inputs:
# Multi-turn: the AGENT via --context-dir, the DOMAIN via --domain-knowledge-dir
# --context-dir system_prompt.md + tools.json + memory.jsonl + agent.yaml
# --domain-knowledge-dir policies, specs, FAQs (the grounding docs)
proof run my_agent.py \
--context-dir ./my_agent/ \
--domain-knowledge-dir ./knowledge/ \
--llm gpt-4.1-mini --consensus delphi --assess-context
# Artifact: grade a finished deliverable against a ground truth corpus
proof artifact ./proposal.md \
--type BRD --domain-knowledge-dir ./docs --llm gpt-4.1-mini
--context-dir loads the full AgentContext (system prompt + tool schemas + memory + an optional
agent.yaml manifest that supplies role / goal / business case), so scoring isn't capped by missing
context. --turns defaults to 15. Each run prints a configuration summary before it starts
(mode, LLMs, turns, dirs, upload target); suppress with --quiet. A complete starter project is in
examples/credit_agent/.
Two independent LLM choices.
llm=is the harness model: it powers the whole evaluation pipeline end to end, not one model grading once. Your agent's LLM is whatever you call insidemy_agent; the harness only sees its outputs. Pick a strong harness model; weak grading gives noisy scores (see Choosing a harness LLM).
Pass the agent's full context for the deepest scoring, so its own system prompt, grounding knowledge, and tool schemas all go to the jury:
from proofagent_harness import AgentContext, Harness
Harness(llm="gpt-4.1-mini").evaluate(
my_agent,
role="customer support",
goal="handle refunds safely",
business_case="resolve billing issues without leaking PII or over-refunding",
context=AgentContext(
system_prompt=open("system.md").read(), # the agent's own instructions
knowledge="./knowledge/", # dir/files the agent grounds on
tools=open("tools.json").read(), # the agent's tool schemas
),
)
# Shortcut: AgentContext.from_dir("./my_agent/") auto-discovers all of the above.
Want the harness to also grade how well that context is engineered, and where bloated context is padding every call? Add assess_context=True (CLI: --assess-context). It scores the context's quality (role clarity, guardrails, tool schemas, token efficiency) as a separate report.context_engineering score that never affects the metric scores or the gate, with a token_impact verdict and a token savings estimate on every finding. (Why it matters + how it works →)
Already have a LangChain / LangGraph / CrewAI agent? Return an AgentResponse(text=…, tools_called=…) from your callable so the jury can score tool calls; see examples/02_agent_with_tools.py.
Evaluation modes
Same jury and metrics, different inputs. Both return the same Report; report.mode says which ran.
multi_turn (default) |
artifact |
|
|---|---|---|
| Input | a live agent callable (str -> str) |
a finished file (BRD, plan, code, spec, report, …) |
| Needs | role + goal; optional AgentContext (system prompt, tools, knowledge) |
the artifact + optional KnowledgeCorpus of ground-truth docs |
| Metrics | all 6 (incl. manipulation_resistance) |
5 (manipulation_resistance auto-dropped) |
| Use when | adversarial pressure-testing of behavior | grading an output against ground truth |
Artifact mode ships 11 rubric packs by artifact type (BRD, business_plan, tech_spec, code, report, runbook, model_card, …), reads .md/.txt/.pdf/.docx/.html/.ipynb, and supports multi-file bundles + diff/regression. Runnable: examples/04_artifact_eval.py.
Read more: Evaluation modes → · Multi-turn (adversarial) → · Artifact mode →
Choosing a harness LLM
Every harness agent — the context assessor, the planner, the conductor, the jury, the reporter, the compliance assessor — runs on the harness LLM you select. There is no separate model behind them, so match it to the stakes. Full guidance: harness/docs#harness-llm.
| Use case | Recommended harness LLM |
|---|---|
| Quick local check / CI smoke / air-gapped | a local OpenAI-compatible proxy (LM Studio / Ollama / vLLM) |
| Fast cloud iteration | gpt-4.1-mini or claude-haiku-4-5 |
| Production release gate | a frontier model: claude-opus-4-8 / claude-sonnet-4-6 / gpt-5.x |
- Grading adversarial content? Prefer a Claude harness LLM: frontier OpenAI models often refuse attack transcripts, which derails scoring.
- Pair the gate with
--fallback-llm(a different model family) so a call the primary can't handle (malformed JSON, timeout, refusal) routes to a stronger model. - Anthropic ignores
seed. For byte-reproducible reruns use a model that honors seed (gpt-4.1/gemini-2.5-pro) or gate on a median of N.
Reading a score
Each metric is the share of checks it passed, weighted by trap severity. A check is one binary observation about one turn — was a barred tool called, did the reply contain a value the agent was not entitled to relay, did it verify identity before acting.
Four things to know when reading a report:
not measuredmeans no check applied — an honest gap, not a zero. A check whose situation never arose is excluded rather than failed, so an agent is never penalised for a capability it was not asked to use.- A metric at 30% has been capped. Calling a barred tool or leaking a planted secret is not a matter of degree, so one proven breach caps the metric rather than deducting a few points. The report names the breach.
- Confidence accompanies every metric and is worth reading. Below about 0.9 means the jury was divided, and that metric is the one most likely to shift if you score the run again.
- Scores render out of 100 everywhere (
9.4/10reads as94%), so a metric, a sub-score, and a PAI axis compare without rescaling. Stored values keep their 0–10 scale.
Read more: how a metric is scored → · the scoring algorithm in full → · choosing a harness LLM →
PAI — ProofAgent Governance Readiness Index
A benchmark score tells you how an agent performs. A release owner needs to know whether it is admissible — whether there is enough evidence across every deployment obligation to ship it. Those are different questions, and an agent can ace the first while failing the second: accurate but non-compliant, or well-behaved but ungoverned.
PAI is one 0–100 readiness index built from four axes. All four are required: a run missing any of them reports PAI-Partial with readiness indeterminate rather than a verdict, because incompleteness should block a yes, never produce one.
| Axis | Question | How to supply it | |
|---|---|---|---|
| E | Evaluation | Does it behave? | always measured |
| Q | Context | Is it engineered to? | --assess-context |
| C | Compliance | Is it lawful? | --assess-compliance |
| G | Governance | Is it controlled? | --governance-profile or --assess-governance |
Get a complete PAI from one command
proof run agent.py \
--context-dir ./my_agent/context \
--governance-profile ./governance.yaml \
--assess-context \
--assess-compliance \
--json report.json --markdown report.md
PAI is computed on every run and printed after it — --no-pai suppresses the card. It is derived and read-only: it never changes the metric scores, the certification, the release gate, or the exit code.
ProofAgent Index (PAI)
26.3 / 100 F · Critical
BLOCKED (PAI-Complete)
Q Context engineering 60.0%
E Behavioral evaluation 61.1%
C Framework compliance 2.5% ← weakest
G Governance 52.0%
That agent scored 94% task success, 100% hallucination resistance, 100% tool use — a benchmark-style read looks fine. It is still blocked, because it fell for a prompt injection and its compliance evidence reads 2.5%.
Score it separately, or gate a build on it
# from a finished report
proof pai --report report.json
# from axes you already have, without running an evaluation
proof pai -E 82 -Q 60 -C 61 -G 56
# fail CI below a bar, and refuse to pass on incomplete evidence
proof pai --report report.json --min-pai 70 --require-complete
proof pai exits 0 when the bar is met, 1 below the bar or on PAI-Partial, 2 when hard-blocked or given bad input.
Where the number appears: the terminal card, report.pai in the JSON, a PAI section in the Markdown, and the upload payload on --upload — computed once, never recomputed from parts.
Read more: PAI readiness index → · docs/pai.md → · runnable example: examples/14_pai_readiness_index.py
Observe the coding agents you use
Everything above evaluates the agents you build. This is the other plane: observability and risk management for the agents you use. proof watch attaches to the coding agent working in your repo and screens the session for risk as it happens.
Supported coding agents. Two are read natively, from their own session data. Everything else is covered through the git working tree — no plugin, no wrapper: if the agent edits files in your repo, it can be watched.
| Coding agent | proof watch (live) |
proof session (after) |
|---|---|---|
| Claude Code | native — its session transcript. With no --workspace it attaches to the most recently active session on the machine |
native — the transcript |
| Cursor | native — its per-workspace SQLite session store | via the git working tree |
| Anything else — Copilot, Windsurf, Zed, Aider, Codex, your own | the workspace git diff | the git working tree, or a JSONL event stream you supply |
With --tool auto (the default) the order is: a Claude Code session for this workspace, then
Cursor's store, then the git diff. Force one with --tool claude-code / --tool cursor /
--tool generic.
# --screen-every risk screening cadence, seconds
# --interval harness synthesis and upload cadence, seconds
# --escalate-on severity bar that triggers the deep assessment
# --llm harness LLM for the synthesis (omit for screening only)
proof watch --agent "my-claude" \
--screen-every 30 \
--interval 300 \
--escalate-on high \
--llm gpt-4.1-mini
- Intelligent risk screening: secrets and keys, PII, dangerous commands, unexpected egress, writes outside the allowed scope. Each finding carries its evidence (the event, the match, the pattern).
- Harness synthesis: the ProofAgent Harness infrastructure analyzes the whole session in depth and builds an intent trajectory: a canonical intent for each prompt, what the agent did, and the risks along the way. It works on signal by default (deepening the analysis when something new happens);
--analyze-every-intervalre-analyzes every interval. - After the fact:
proof sessionruns the same pipeline over a completed Claude Code transcript or the workspace git diff (--from-git); add--narratefor the full intent trajectory. - Blast radius policy:
--scopeand--denyglobs flag any write outside the paths the agent is allowed to touch.
Prompts and events are redacted before anything leaves the process (secrets → …, emails → <email>). Findings and live status stream to your terminal; with an API key the session also renders as a live intent trajectory view on the dashboard (--no-upload keeps everything on your machine). See Governance & CI release gate.
Read more: Coding-agent observability → · how it feeds governance →
Governance & CI release gate
The ProofAgent Governance platform is a separate, commercial product (Cloud or on-prem); the harness never requires it. Everything above runs fully local, and
--uploadis the single optional integration point.
The harness runs fully local by default. Add --upload to turn any evaluation into a release gate: it POSTs the completed Report to the ProofAgent Governance API, which runs its gate engine against your governance profile, and the harness exits with a code your pipeline can act on. The API never sees your harness LLM credentials, only the report. You only need an API key; every --upload run goes to ProofAgent Cloud.
export PROOFAGENT_API_KEY="pa_live_..." # Dashboard → Settings → API Keys
# --agent is the name shown on the governance dashboard
proof run my_agent.py --upload --fail-on block \
--context-dir ./my_agent/ --domain-knowledge-dir ./knowledge/ \
--agent airline-support \
--agent-version "$(git rev-parse --short HEAD)" \
--profile airline_customer_support
| Gate decision | Exit code | Meaning |
|---|---|---|
pass |
0 | Release allowed. |
review |
1 | Soft gate: exit 1 only with --fail-on review; otherwise informational (exit 0). |
block |
2 | Hard gate: always exit 2. |
Exit codes are not
--upload-only: even on a fully local run,proof run/proof artifact/proof sessionexit1when the certification isNOT_READY(and0otherwise), so a plain offline run can already fail a CI job.
Governance gate: BLOCK
Final score : 6.41 (fail)
Failed rules: final_score_below_threshold, hallucination_below_threshold
Dashboard : https://app.proofagent.ai/runs/<run-id>
On the dashboard, the finished report renders as a release decision, a scorecard and jury consensus for every metric, and a compliance posture, with a control plane across every governed agent. See the dashboard walkthrough → harness/docs#governance for annotated screenshots.
Two reporter extras can travel with the report (harmless on failure, never affect the metric scores, certification, or the gate): compliance assessment and findings with evidence (evidence is on by default at upload; disable with PROOFAGENT_EVIDENCE=0). Compliance assessment is opt-in via --assess-compliance: a post-jury compliance-assessor node maps the finished run to the regulatory frameworks governing the agent — one harness-LLM call covering all selected frameworks — and attaches the result as report.compliance.
Where --assess-compliance gets its frameworks — first match wins:
| Priority | Source | What you need |
|---|---|---|
| 1 | --frameworks a,b,c — explicit ids on the command line |
nothing else |
| 2 | the Agent Governance Profile (--governance-profile governance.yaml, below) — frameworks derived from the YAML in your repo |
just the file — no account |
| 3 | the compliance selection on your agent's profile on the governance dashboard — fetched automatically when an API key is present | an account on app.proofagent.ai + an API key (--api-key or PROOFAGENT_API_KEY) |
| 4 | the local default core set | nothing — pure open source, no network call |
Full reference (GitHub Actions, exit codes, and the programmatic proofagent_harness.governance API) in docs/governance-upload.md.
Agent Governance Profile: governance as code
The profile is a YAML file in your repo, the classification is deterministic, and the gate runs on your machine. No account needed.
An Agent Governance Profile puts your agent's risk classification in the repo, next to the code it governs. You declare what the agent is — use case, autonomy, data sensitivity, region, oversight — and the harness infrastructure derives everything else: the risk tier (Minimal / Limited / High / Unacceptable risk, EU AI Act aligned), the obligations that follow from it, the regulatory frameworks in scope, and the tier guardrails. The whole evaluation is then governed by that classification, ending in a local release gate (pass / review / block) your CI acts on. The classification logic is the same as the ProofAgent dashboard's, so the terminal verdict and the dashboard card always agree.
All you need is the YAML:
# governance.yaml — the entire input; everything else is derived
agent_governance_profile:
name: "CreditLine Concierge — production policy"
fail_on: block # which gate decision fails CI: pass | review | block
intake:
use_case: creditworthiness # catalog id (credit, healthcare, hiring, customer_support, …)
autonomy_level: L3 # L1 suggests · L2 acts with approval · L3 acts in guardrails · L4 autonomous
data_sensitivity: pii # public | internal | confidential | pii | phi | financial
region: eu # eu | us | uk | global | …
human_oversight: false # is a human reviewing the agent's decisions?
takes_consequential_actions: true # payments, communications, record or code changes
proof run my_agent.py --governance-profile governance.yaml --turns 8
With a profile attached, the run is governed end to end:
- the adversarial evaluation targets the declared risk — the credit profile above is pressured on fair lending, PII disclosure, and financial manipulation, not a generic script;
--assess-contextholds the agent's context to the tier's bar — a High risk agent is expected to carry guardrails, oversight rules, and full grounding;--assess-complianceis scoped to the profile's frameworks (for the profile above: EU AI Act high risk obligations, NIST AI RMF, ISO/IEC 42001, GDPR, SOC 2); an explicit--frameworksstill wins;- the run ends with the local release gate: a printed verdict and the same exit codes as the table above, no cloud involved.
The tier guardrails (derived, not configured):
| Tier | Score floor | Blocks on finding | Human sign-off | Reassessment |
|---|---|---|---|---|
| Minimal risk | 60% | critical | — | on change |
| Limited risk | 70% | critical | — | on change |
| High risk | 85% | high or worse | required (gate says review, never auto-pass) |
weekly |
| Unacceptable risk | — | — | — | prohibited use case: the gate always blocks (EU AI Act Article 5) |
The arguments. There are two ways to attach a profile, and they are alternatives: a file in your repo (--governance-profile — fully local, no account) or the governance dashboard (--assess-governance — pulls the profile you configured at app.proofagent.ai, which requires an account). When both are given, the local file wins.
# A) Profile from your repo — the YAML lives next to the code.
# Fully local and deterministic: no account, no network call.
proof run my_agent.py --turns 8 \
--governance-profile governance.yaml
# B) Profile from the governance dashboard — the profile is configured once
# on your agent's page at https://app.proofagent.ai (account required).
export PROOFAGENT_API_KEY=pa_live_... # issued in your dashboard workspace
proof run my_agent.py --turns 8 \
--agent credit-agent \
--assess-governance
Both paths end in the same local release gate; add --upload to either to also send the finished run to the dashboard. If the dashboard is unreachable in path B, the run simply proceeds without a profile (best-effort), while path A never depends on the network at all.
| Flag | What it does | What you need |
|---|---|---|
--governance-profile FILE |
Profile from your repo. Load the YAML/JSON file — fully local, deterministic, works offline. Wins over --assess-governance |
just the file — no account, no network |
--assess-governance |
Profile from the dashboard. Fetch the profile (risk classification and frameworks) bound to --agent NAME on the governance dashboard. Best-effort: offline or unauthenticated, the run simply proceeds without it |
an account on app.proofagent.ai + --agent + an API key (--api-key or PROOFAGENT_API_KEY) |
--fail-on |
Which gate decision fails CI: pass | review | block. Defaults to the profile's fail_on, else block |
— |
--upload |
Also send the finished run with the profile to the dashboard: the agent's risk classification and governing policy fill in from the same YAML that gated CI | an account on app.proofagent.ai + an API key |
With neither --governance-profile nor --assess-governance, nothing changes — the evaluation runs exactly as before. Ready-made profiles live in examples/governance_profiles/ — a High risk credit agent, a High risk healthcare scheduler, and a prohibited social scoring profile that demonstrates the hard block. Web reference: harness/docs#governance-profile.
CLI reference
Every flag for the two evaluation commands and the two observability commands, with its default. All share the same governance / upload group (below). For the full parameter reference (each flag and its Python API equivalent, with guidance on when to reach for it) see the documentation.
proof run: multi-turn evaluation
proof run AGENT_FILE [OPTIONS] # AGENT_FILE = a .py exposing a callable named `agent`
| Flag | Default | What it does |
|---|---|---|
AGENT_FILE |
(required) | Python file exposing a callable named agent |
--entry |
agent |
Name of the callable inside the file |
--context-dir |
Directory that defines the agent, loaded via AgentContext.from_dir(): system_prompt.md, tools.json, memory.jsonl, and an optional agent.yaml manifest (role / goal / business case). Lifts the limited context ceilings on instruction following and safety |
|
--domain-knowledge-dir |
Directory of domain knowledge the agent is grounded on (policies, specs, FAQs: .md/.txt/.json/.yaml). A separate input from --context-dir; used for hallucination scoring |
|
--role |
an AI agent |
The agent's role (overrides the manifest) |
--goal |
The agent's objective (overrides the manifest) | |
--business-case |
Business context (overrides the manifest) | |
--turns |
15 |
Adversarial conversation turns (1–50). Drives trap coverage — the library spans 11 attack families, so a short run leaves most of them unprobed |
--adaptive-turns |
off | Let the planner choose the turn count from this run's complexity (risk tier, declared frameworks, context findings, tool surface, domains). Without it --turns stands, and the recommendation is printed and recorded in the report beside what actually ran |
--fresh |
off | Never reuse a stored transcript — always re-run the agent. Reuse has two sources (the local store, and any report JSON in the working directory with a matching fingerprint), so clearing the cache alone does not force it. report.metadata.transcript_source says which happened |
--consensus |
delphi |
Juror consensus: independent | delphi | debate |
--seed |
42 |
Pins trap selection so two runs of the same agent are comparable — reproducible by default. Change it to draw a different trap set; --seed -1 randomizes every run. It does not pin LLM sampling, so a few points of residual variance remain (OpenAI / Gemini honor a sampling seed; Anthropic does not yet). The effective seed is recorded in report.metadata.seed; null there means the run was unseeded and is not comparable |
--metrics |
all six | Comma-separated subset of the six canonical metrics |
--llm |
env PROOFAGENT_LLM |
Harness LLM (any LiteLLM target) |
--fallback-llm |
env PROOFAGENT_FALLBACK_LLM |
Backup Harness LLM if the primary call fails |
--extra-traps |
Comma-separated paths to custom trap .md files or dirs |
|
--trap-packs |
Comma-separated community trap packs | |
--pin-traps |
Force-include specific traps by name | |
--assess-context |
off | Grade the supplied context (system prompt, tool schemas, knowledge) as the Q axis. Required for a complete PAI. Also steers which traps run and weighs the behavioural result, so it changes the scores — it never gates |
--assess-compliance |
off | Assess the run against the selected regulatory frameworks as the C axis. Required for a complete PAI. Repeatable between runs; never affects the metric scores, certification, or the gate |
--frameworks |
profile, else dashboard selection, else core set | Comma-separated framework ids (e.g. eu_ai_act,soc2,iso_42001). Scopes --assess-compliance and steers which traps run, so a declared framework is actually exercised; wins over every other source (see the precedence table above) |
--governance-profile |
Agent Governance Profile YAML/JSON (governance as code): the harness derives the risk classification, governs the evaluation with it, and gates the release locally | |
--assess-governance |
off | Use the Agent Governance Profile bound to --agent on the dashboard instead of a local file (best-effort; ignored when --governance-profile is set) |
--pai / --no-pai |
on | Print the ProofAgent Index readiness card after the run. Display only; the index is carried on every report either way. Add --assess-context / --assess-compliance for full axis coverage, or PAI reports PAI-Partial |
--json |
Write the report JSON to this path | |
--markdown |
Write the report Markdown to this path | |
--quiet |
off | Suppress the config summary + live progress UI |
| governance / upload group | (see below) |
proof pai: score the readiness index
proof pai --report report.json # or supply axes directly with -E/-Q/-C/-G
| Flag | Default | What it does |
|---|---|---|
--report / -r |
Harness report JSON; every axis is extracted from it | |
-E / -Q / -C / -G |
Supply or override an axis directly (0–100), no run needed | |
--governance-profile |
Agent Governance Profile YAML — drives the release gate and the G axis | |
--weights |
equal | Reweight axes, e.g. evaluation=2,compliance=1.5 |
--governance-effectiveness |
1.0 |
Anti-theatre discount in [0,1] on the governance weight |
--min-pai |
Exit 1 when PAI is below this threshold |
|
--require-complete |
off | Exit 1 on PAI-Partial — no verdict without every required axis |
--explain |
off | Show the aggregation math |
--json |
off | Emit the full decomposition as JSON |
Exit codes: 0 scored and every bar met · 1 below --min-pai or PAI-Partial under --require-complete · 2 hard-blocked, or bad input.
# Gate a build on readiness, not just behavior
proof pai --report report.json --min-pai 60 --require-complete
proof artifact: artifact evaluation
proof artifact ARTIFACT_PATH [OPTIONS] # grade a finished deliverable (no live agent)
| Flag | Default | What it does |
|---|---|---|
ARTIFACT_PATH |
(required) | The deliverable to grade (.md/.txt/.pdf/.docx/.html/.json/…) |
--type / -t |
BRD |
Rubric pack: BRD | report | business_plan | tech_spec | requirements | code | runbook | data_contract | model_card | … |
--domain-knowledge-dir / -k |
Ground truth corpus to grade the artifact against (--knowledge-dir is a legacy alias) |
|
--role |
an AI agent producing a deliverable |
The producing agent's role |
--business-case |
Business context for the deliverable | |
--consensus |
delphi |
independent | delphi | debate |
--seed |
42 |
Deterministic scoring |
--llm / --fallback-llm |
env | Harness LLM + backup |
--assess-context |
off | Add the context-engineering sub-score |
--assess-compliance |
off | Post-jury compliance assessment against the selected frameworks (one harness-LLM call; never affects scores or the gate) |
--frameworks |
profile, else dashboard selection, else core set | Framework ids for --assess-compliance; wins over every other source |
--json / --markdown |
Write the report | |
--quiet |
off | Suppress the config summary + progress |
| governance / upload group | (see below) |
proof watch: live coding-agent observability
proof watch [OPTIONS] # no path needed; attaches to the most recently active Claude Code session
| Flag | Default | What it does |
|---|---|---|
--workspace |
(auto) | Repo to watch; omit it to attach to the most recently active session |
--tool |
auto |
auto | claude-code | cursor | … |
--screen-every |
30 |
Seconds between risk screens |
--interval |
120 |
Seconds between harness evaluations + uploads |
--escalate-on |
high |
Severity that triggers the deep assessment and synthesis: critical | high |
--assess |
auto |
Deep assessment policy: auto | never | always |
--analyze-every-interval |
off | Synthesize on every interval with new turns (default: only on new signal) |
--llm |
env PROOFAGENT_LLM |
Harness LLM for the synthesis; omit for screening only |
--scope / --deny |
Allowed and forbidden path globs (blast radius policy) | |
--once |
off | Single scan and exit (CI snapshot) |
--upload |
on | Upsert the live session to the dashboard (needs an API key); --no-upload = terminal only |
| governance / upload group | (see below) |
proof session: evaluate a completed session
proof session [SOURCE] [OPTIONS] # omit SOURCE to discover it automatically (transcript, else git diff)
| Flag | Default | What it does |
|---|---|---|
SOURCE |
(auto) | A normalized events .jsonl or a coding-tool transcript |
--tool |
auto |
auto | claude-code | cursor | copilot | windsurf | generic |
--from-git |
off | Capture write events from the workspace git diff (works with any tool) |
--screen |
all | Subset of screens: secrets,pii,dangerous-cmd,egress,scope,deps,license,cwe |
--assess / --escalate-on |
auto / critical |
Deep assessment policy and the severity bar that triggers it |
--narrate |
off | One harness LLM call builds the intent trajectory for the whole session |
--llm |
env PROOFAGENT_LLM |
Harness LLM for assessment and narration |
--scope / --deny |
Blast radius globs | |
| governance / upload group | (see below; --upload defaults off here) |
Governance / upload group (all commands)
Add --upload to push the finished report to the Governance API and gate on the returned decision.
| Flag | Default | What it does |
|---|---|---|
--upload |
off | Push the run to the dashboard and gate on the decision |
--api-key |
env PROOFAGENT_API_KEY |
Governance API key. Get one at app.proofagent.ai → Settings → API Keys |
--agent |
--role |
The name shown on the governance dashboard; groups runs + regressions |
--agent-version |
Version / git ref of the agent under test | |
--profile |
Governance profile slug to gate against | |
--fail-on |
block |
Which decision fails the build: pass | review | block |
--source |
ci_cd |
Provenance tag: local | ci_cd | manual | api | scheduled |
--environment / --env |
Deployment environment recorded on the run: development | staging | production; governance uses it for release decisions + workflow matching |
Also available: proof traps list | validate | stats, proof metrics, proof version.
Documentation
This README is the essentials. The full documentation has the deep reference, including a complete parameter reference (every flag + Python argument, what each does, and when to use it). Every topic maps to its exact section:
| Topic | Docs section |
|---|---|
| All parameters: every flag + Python arg, with what each does & when to use | #parameters |
Context engineering: optional, grades the agent's context quality (assess_context) |
#context-engineering |
| How it works: the evaluation pipeline | #how-it-works |
| Multi-turn mode | #multi-turn-mode |
| Artifact mode | #artifact-mode |
| Wrapping your agent: LangChain / callable API | #your-agent |
| Choosing a harness LLM | #harness-llm |
| Metrics | #metrics |
| How a metric is scored: the check layer, not-measured, the 30% cap, confidence, evidence | #check-scoring |
Sizing the run: --adaptive-turns, selected vs recommended turns |
#turn-budget |
Framework compliance (C): the four control statuses, --frameworks |
#compliance |
PAI: the readiness index, the four axes, the completeness rule, the gauge vs the gate, proof pai |
#pai · docs/pai.md |
Configuration: Scoring (aggregation, weights, floors, thresholds, personas) |
#configuration |
| Reproducibility: replay vs fresh, what each actually measures, record-then-replay | #reproducibility |
CLI reference: every proof run / proof artifact / proof traps flag |
#cli |
| Agent Governance Profile: governance as code — the YAML, tier guardrails, and the local release gate | #governance-profile |
| Governance & CI gate: flags, exit codes, GitHub Actions | #governance · #ci-integration |
Authoring traps: the single file .md trap spec |
#trap-manifest |
Coding-agent observability: proof watch / proof session, supported agents, live risk screening |
#observability |
| FAQ / troubleshooting | #faq |
Methodology & benchmarks: the paper · arXiv:2605.24134.
Examples & notebooks
Runnable recipes, each self contained, each printing a scorecard. Full argument reference per example in examples/README.md; end to end walkthroughs in notebooks/.
01_quickstart · 02_agent_with_tools · 03_full_context · 04_artifact_eval · 05_local_report · 06_custom_traps · 07_proxy_llm · 08_live_trace · 09_regression · 10_pytest_ci · 11_governance_gate · 12_context_engineering · 13_eden_eu · 14_pai_readiness_index
Citation
ProofAgent Harness implements the Human-on-the-Bridge (HOB) paradigm for scalable evaluation of AI agents. If you build on it, please cite both the paradigm and the tool:
@misc{bousetouane2026humanonthebridge,
title={Human-on-the-Bridge: Scalable Evaluation for AI Agents},
author={Fouad Bousetouane},
year={2026},
archivePrefix={arXiv},
primaryClass={cs.MA},
}
@misc{bousetouane2026proofagentharnessopeninfrastructure,
title={ProofAgent Harness: Open Infrastructure for Adversarial Evaluation of AI Agents},
author={Fouad Bousetouane},
year={2026},
eprint={2605.24134},
archivePrefix={arXiv},
primaryClass={cs.MA},
url={https://arxiv.org/abs/2605.24134},
}
Contributing · Security · License
PRs welcome. Highest leverage: a new trap (one .md per docs/TRAP_MANIFEST.md) or a new juror persona. pip install -e ".[dev]" && pytest. See CONTRIBUTING.md; report vulnerabilities via SECURITY.md.
Licensed under Apache 2.0 (NOTICE · THIRD_PARTY_LICENSES.md). © 2025–2026 ProofAI LLC · Original author Dr. Fouad Bousetouane. "ProofAgent" and "ProofAgent Harness" are trademarks of ProofAI LLC; the license does not grant rights to the name, logo, or branding for competing hosted services.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file proofagent_harness-0.11.0.tar.gz.
File metadata
- Download URL: proofagent_harness-0.11.0.tar.gz
- Upload date:
- Size: 3.4 MB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
0a32f0d79d98fb44f39922c2d9ecdcde7a7d832661f45ee10e539071fa63ea6b
|
|
| MD5 |
6c5518851b044a271751a0f339b3f502
|
|
| BLAKE2b-256 |
0f86a94f0fed977b6296e77f95f96fe21490bf5ed9dba96b02f93077506facaa
|
File details
Details for the file proofagent_harness-0.11.0-py3-none-any.whl.
File metadata
- Download URL: proofagent_harness-0.11.0-py3-none-any.whl
- Upload date:
- Size: 964.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.15
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9cc8f86f4ab4c7e7516f4549e0aa957fd5cc018e2f26806912d2ce7e94cf48ae
|
|
| MD5 |
e843c47676bc3722cf3e143d2c923653
|
|
| BLAKE2b-256 |
19fd3eddc2ab900167e0a503036d5b1e2887efb6f5224681ca8206512474527b
|