Skip to main content

arcezia

Runtime safety verification for autonomous AI agents — official Python SDK.

All verification runs in Arcezia's secure cloud. The SDK makes HTTPS calls to api.arcezia.com and returns typed result objects. Zero inference on the client.

Install

pip install arcezia

Framework extras:

pip install "arcezia[langchain]"    # LangChain + LangGraph
pip install "arcezia[openai]"       # OpenAI Agents SDK
pip install "arcezia[anthropic]"    # Anthropic (Claude) SDK
pip install "arcezia[autogen]"      # AutoGen (legacy + modern)
pip install "arcezia[llamaindex]"   # LlamaIndex
pip install "arcezia[all]"          # everything

Quick start

import arcezia

az = arcezia.Arcezia(api_key="ar_live_...", task="clean up test records")

cert = az.verify(
    action_type="execute_sql",
    action_description="DELETE FROM analytics_staging WHERE date < '2024-01-01'",
    domain="database_ops",
)

if cert.degraded:
    # Arcezia could not be reached, so nothing was actually verified.
    # The default on_error="fail_closed" raises before you get here; check this
    # explicitly if you set on_error="review" or "fail_open".
    raise RuntimeError("Not verified — Arcezia unreachable")

# Gate on ALLOW positively — never on "not blocked". A verdict can be
# review (insufficient evidence, human confirmation required), which is
# neither allow nor block; treating it as runnable executes an action the
# engine explicitly declined to clear. cert.allow is True only for a
# grounded ALLOW.
if not cert.allow:
    raise RuntimeError(f"Not allowed ({'review' if cert.review else 'blocked'}): {cert.summary}")

db.execute(sql)  # only reached when the verdict is ALLOW

Reaching ALLOW on a fresh key

Re-run on 2026-09-16 on a free-tier key, a free test key and an enterprise key — identical verdicts on all three. A read is held until you say what the engine cannot see; a write additionally needs a person's approval.

az = arcezia.Arcezia(api_key="ar_live_...", task="report on last year's events")
az.start_session(capability_envelope={          # what this agent may do at most
    "allowed_domains": ["database_ops"],
    "allowed_action_types": ["execute_sql"],
})
cert = az.verify(action_type="execute_sql", domain="database_ops",
                 action_description="SELECT COUNT(*) FROM events")
cert.verdict   # REVIEW — cert.missing starts with the three facts a statement cannot show:
               # action_direction_is_outbound, action_crosses_trust_boundary, action_involves_sensitive_data

Declare those three for execute_sql once, as the account admin (never "*"):

curl -X POST https://api.arcezia.com/v1/declarations \
  -H "Authorization: Bearer $ARCEZIA_API_KEY" -H "Content-Type: application/json" \
  -d '{"declared_absent": {"action_direction_is_outbound": ["execute_sql"],
                           "action_crosses_trust_boundary": ["execute_sql"],
                           "action_involves_sensitive_data": ["execute_sql"]}}'
cert = az.verify(action_type="execute_sql", domain="database_ops",
                 action_description="SELECT COUNT(*) FROM events")
cert.verdict     # ALLOW — cert.credential is set

cert = az.verify(action_type="execute_sql", domain="database_ops",
                 action_description="UPDATE customers SET tier='pro' WHERE id = 42")
cert.verdict     # REVIEW — a write needs an approval the agent cannot give itself
                 # (cert.missing names user_explicit_authorization, action_scope_is_mass, ...)

# Declare action_scope_is_mass and action_is_destructive for execute_sql the same
# way, attach the approval your backend minted, and the same write clears:
az.authorize(signed_user_token).authorize_production(signed_production_token)
cert = az.verify(action_type="execute_sql", domain="database_ops",
                 action_description="UPDATE customers SET tier='pro' WHERE id = 42")
cert.verdict     # ALLOW

# A declaration never silences what the engine detects: with the same five
# declarations and the same approvals, "DELETE FROM customers WHERE id = 42"
# is still held — a destructive statement needs a verified backup your own
# system confirms — and an agent that claims an approval it does not have is
# BLOCK with cert.fabrication_detected == True.

The framework adapters below do this for you: a degraded certificate always raises ArceziaUnavailableError and the tool never executes.

When Arcezia is unreachable

One setting, on_error, decides this — and it applies to every method that makes a network call, not only verify().

method fail_closed (default) review fail_open
verify raises ArceziaUnavailableError synthetic REVIEW cert synthetic ALLOW cert
verify_chain raises ArceziaChainResult(overall_verdict="REVIEW_REQUIRED", degraded=True) overall_verdict="SAFE", degraded=True
verify_outcome raises synthetic REVIEW result synthetic ALLOW result
start_session raises pending, retried on next call pending, retried on next call
authorize, authorize_production raises and keeps the token pending pending, retried pending, retried
usage raises raises raises
audit_subject raises raises raises

Three properties worth quoting in a security review:

  1. Under the default, an outage never becomes an ALLOW. There is exactly one place in the client that chooses between raising and returning a degraded value, and under fail_closed it always raises.
  2. A degraded verdict is identifiable and carries no credential. cert.degraded — and result.degraded on a chain, the same field name on all three result types — is set only by local construction; it is stripped from anything parsed off the wire, so a response cannot claim it and a synthetic ALLOW cannot be replayed as a verified one. Note that a degraded chain result's overall_verdict is the string "SAFE" under fail_open: use result.safe, which is False on a degraded result, rather than comparing the string.
  3. A human approval is never silently dropped. A failed POST /v1/authorize leaves the token pending and re-sends it before the next verdict is asked for; under the default it also raises, so the person who clicked Approve finds out.

A deterministic 4xx — including a WAF's HTML 403 — is an answer, not an outage. It raises ArceziaAPIError under all three settings, fail_open included: an edge-blocked deployment fails loudly rather than running every tool unverified while looking healthy.

on_error belongs to the client. Adapters that take it (DispatchGuard) build their client with it; passing it alongside an az that disagrees raises rather than being ignored.

Framework integrations

LangChain / LangGraph

from arcezia.integrations.langchain import ArceziaToolkit

toolkit = ArceziaToolkit(az)
safe_tools = toolkit.wrap(tools)                    # classic AgentExecutor
safe_tools = toolkit.wrap_for_langgraph(tools)      # LangGraph / tool-calling

OpenAI function calling

from arcezia.integrations.openai import ArceziaGuard

guard = ArceziaGuard(az)
result = guard.execute_tool_call(
    tool_call=response.choices[0].message.tool_calls[0],
    tool_implementations={"execute_sql": db.execute},
)
# or wrap a single function:
safe_execute = guard.wrap_function("execute_sql", db.execute)

CrewAI

from arcezia.integrations.openai import ArceziaCrewTool

class SafeSQLTool(ArceziaCrewTool):
    az = your_arcezia_client
    domain = "database_ops"
    name = "execute_sql"
    description = "Execute SQL"

    def _run(self, sql: str) -> str:
        return db.execute(sql)

Anthropic (Claude tool_use)

from arcezia.integrations.anthropic import ArceziaAnthropicGuard

guard = ArceziaAnthropicGuard(az)
safe_uses, blocked = guard.filter_tool_uses(message.content)

AutoGen

from arcezia.integrations.autogen import ArceziaAutoGenGuard

guard = ArceziaAutoGenGuard(az)
safe_fn = guard.wrap("execute_sql", db.execute, "database_ops")   # name first
safe_map = guard.wrap_many([                                      # list of tuples
    ("execute_sql", db.execute, "database_ops"),
    ("send_data", exporter.send, "agent_action"),
])

LlamaIndex

from arcezia.integrations.llamaindex import ArceziaLlamaToolkit
safe_tools = ArceziaLlamaToolkit(az).wrap(tools)

Any framework (Pydantic AI, smolagents, Google ADK, Strands, …)

from arcezia import guard_callable
safe_fn = guard_callable(run_sql, az)

Claude Code CLI hook (gated at the harness level — every tool call)

arcezia-hook install      # merges a PreToolUse hook into ~/.claude/settings.json
export ARCEZIA_API_KEY=ar_live_...
export TASK="refactor auth module"

install never overwrites your settings: a file that does not parse as a JSON object raises and names the path, and every write copies the original to settings.json.bak-<timestamp> first. The hook always exits 0 and always prints a decision — a crash denies rather than passing the tool through, because the harness reads a non-zero exit as a non-blocking error. WebFetch and WebSearch are verified as outbound actions, not treated as reads.

Generic dispatch-loop agents (OpenCLAW, AutoAgent, …)

from arcezia.integrations.openclaw import DispatchGuard
guard = DispatchGuard(api_key="ar_live_...", task="...")
result = guard.dispatch("write_file", {"path": "/etc/app.conf", "content": "..."})

n8n workflows

from arcezia.integrations.n8n import workflow_template, save_template
save_template("arcezia_gate.json")  # import into n8n

The template's human-approval path needs two things from you before it enforces anything: a signing key registered at POST /v1/account/token_key, and an n8n HTTP-header credential named "Arcezia Approval Resume Auth" for the Wait node's resume URL. The approval token is minted by your backend after a person approves; the workflow supplies no default for it and stops the run when the resume carries none.

Verdicts

Verdict Meaning
cert.allow Safe to execute — all required evidence is grounded
cert.block Execution blocked — violated constraint or fabrication detected
cert.review Insufficient evidence — human confirmation required

Two scores travel with every certificate — they measure different things:

Field Meaning
cert.precondition_score [0,1] severity-weighted fraction of required preconditions satisfied
cert.trust_score [0,1] fraction of evidence that is externally grounded, not agent-claimed

When a verdict is not ALLOW, two more fields tell you why — and they answer different questions:

Field Meaning
cert.missing facts you can act on. Ground these and re-verify. Legitimately empty when nothing is caller-groundable.
cert.unresolved every ungrounded fact, including ones held by a rule rather than directly required. Read this when missing is empty but the verdict still is not ALLOW.
cert.denied_authority_axes axes you declared False in the capability envelope. An action that crosses one cannot reach ALLOW, and no token lifts it — widening means signing a new envelope. Three-state: a list is what the server reported; None means the server did not report it, which is never the same as "nothing was denied". Read it with cert.denied_axes_or_unknown(), which returns (axes, reported).
cert.fabrication_detected three-state as well: True (the server found fabricated evidence), False (it looked and found none), None (it did not report). cert.is_clean() is the fail-closed reading — it is False on None, because an absent accusation is not a clearance. cert.allow / .block / .review are unchanged: they gate on verdict, which is the decision and is always present.
cert.semantic_block a cross-step danger pattern fired for this session — e.g. a sensitive read earlier and an outbound send now. cert.chain_patterns names which.
cert.chain_status three-state: "SEMANTIC_BLOCK" (the cross-step scan ran and fired), "CLEAR" (it ran and found nothing), or None (it did not run — there was no session to scan across, or it raised). cert.chain_status_reported answers "did the scan run"; cert.is_clean() does not require it, because a sessionless single verify legitimately has no cross-step context. If your deployment always runs sessions and a missing scan should stop the action, write if not (cert.is_clean() and cert.chain_status_reported): halt().

denied_authority_axes is the most common reason a correctly-wired integration stays stuck: declaring "irreversible": False and then verifying a DELETE denies the very axis the action needs. If every fact is grounded and the verdict still is not ALLOW, read it first.

A dangerous sequence can have a safe-looking step

cert.verdict describes this action. A sequence can be dangerous while every step in it is unobjectionable alone — read customer records, then send data to an external host. When that happens the verdict stays ALLOW and cert.semantic_block is True.

cert.allow accounts for this, so the gate below is correct as written and you do not need a second check:

if not cert.allow:              # False on a cross-step block, even when verdict == "ALLOW"
    raise RuntimeError(cert.summary)
run_tool(...)

If you gate on cert.verdict == "ALLOW" instead, you will miss it. Gate on cert.allow.

Declaring authority — how an action reaches ALLOW

Arcezia never infers what you permit; a principal declares it. Without that declaration the scope of an action is unresolved, and an unresolved action is never allowed — so a fresh session returns REVIEW even for a harmless read. That is the design, not a misconfiguration: Arcezia does not allow what it cannot positively verify.

Declare authority once, when the session opens:

az = arcezia.Arcezia(task="read analytics for the weekly report")
az.start_session(capability_envelope={
    "max_scope": "batch",              # single_record | batch | limited | mass
    "structural_authority": {
        "sensitive_data":          True,   # may touch credentials/PII
        "outbound":                False,  # may send data out
        "persistent_mutation":     False,  # may change stored state
        "mass_scope":              False,  # may act on many records at once
        "trust_boundary_crossing": False,  # may call external principals
        "irreversible":            False,  # may take unrecoverable actions
    },
})

cert = az.verify(action_type="execute_sql",
                 action_description="SELECT COUNT(*) FROM events",
                 domain="database_ops")
# → ALLOW, with a signed credential, once every fact the action depends on is
#   established. If it comes back REVIEW, cert.missing names what is left.

Scope and authority are not the only facts an action depends on. Some cannot be seen in the action itself — whether a query sends data anywhere, reaches outside your organisation, or touches sensitive data. When nothing in the action settles one of these, it is left unresolved and the action is held, never assumed safe. You settle it by declaring, per action type, which of those risks your deployment never has (POST /v1/declarations, admin role). Declare a risk absent only if it is true of every call of that type — a declaration is read as "not present" for everything the engine does not detect, and it never overrides something the engine does detect.

Those six axes are the complete set, and the names are exact. The SDK rejects an unrecognised axis at start_session with a ValueError (v1.0.1+), because a silently dropped axis would leave you believing you had granted or denied something you had not. Over raw HTTP the server accepts the session but grants nothing for the unknown axis and reports it back as ignored_authority_keys in the response — never a silent grant either way. Two are easy to get wrong: it is persistent_mutation (not mutation) and trust_boundary_crossing (not trust_crossing).

The envelope is a ceiling, not a permission slip. Declaring outbound: False and then attempting an outbound action does not produce ALLOW — the action contradicts the authority you signed, so it is blocked, and no runtime approval token can lift it. Widening authority is your act: sign a new envelope. Declaring an axis True does not force ALLOW either; it only removes that axis as a blocker, and every other check still applies.

Declare all six axes. An axis you omit is not a ceiling — it is an open question, and a signed human token (az.authorize(...)) can answer it for the session. That is the intended escalation path for work nobody pre-authorized, but it means one authorize() call covers every axis you left unspecified. Only an axis you declared False is a hard limit.

The four levels

Each level is useful on its own and assumes the one below it. Every framework adapter implements Level 1 for you; Levels 2–4 are reached through the adapter's .az property — the same client, no private access.

Level What you get How
1 — Drop-in gating Every tool call verified before it runs toolkit.wrap(tools)
2 — Chain verification Verify the whole plan, not just each step toolkit.az.verify_chain(...)
3 — Grounded evidence Arcezia asks your systems for facts instead of trusting the agent register a probe webhook
4 — Custom domains Your own constraint domains and compliance packs POST /v1/domains

Level 2 — verify the plan before running any of it

result = toolkit.az.verify_chain({
    "steps": [
        {"step_id": "s1", "action_type": "execute_sql", "domain": "database_ops",
         "action_description": "SELECT ssn, name FROM customers"},
        {"step_id": "s2", "action_type": "send_email", "domain": "email_ops",
         "action_description": "email the list to external-analytics@gmail.com"},
    ]
}, stop_on_block=True)
# → overall_verdict "SEMANTIC_BLOCK", blocked_at "s2",
#   semantic_triggers [{"pattern_name": "structural_exfiltration", ...}]

# An ArceziaChainResult: .overall_verdict, .blocked_at, .steps,
# .semantic_triggers, .human_summary, .degraded, .safe — plus .raw for
# anything else the server sent (final_state, session_state_updated).
# There is no top-level "verdict" — per-step verdicts live under .steps.
if not result.safe:
    # blocked_at names the step only when execution was actually stopped.
    # On REVIEW_REQUIRED nothing was blocked, so it is None — find the step
    # that needs attention in .steps instead.
    step = result.blocked_at or next(
        (s["id"] for s in result.steps if s["verdict"] != "ALLOW"), None
    )
    abort(step)

The request field is step_id; the response's steps[] echo it as id.

verify_chain returned a plain dict before 1.0.5. Indexing still works for every documented key — result["overall_verdict"], result["blocked_at"], result["steps"], result["summary"], result["final_state"] — and is deprecated. Prefer result.safe over result["overall_verdict"] != "SAFE": the string reads "SAFE" on a degraded result too, and .safe does not.

Other behaviour changes in 1.0.5 (the fail-closed direction, deliberately):

  • A response that omits fabrication_detected now parses as None — not reported — rather than False. cert.allow / cert.block / cert.review are unchanged; they read verdict, which is always present. But every framework adapter now gates on cert.is_clean(), which treats not reported as not clean. No current server omits the field, so no live deployment is affected; a much older self-built server would now be refused at the adapter rather than executed against.
  • Constructing an adapter with a client set to on_error="fail_open" emits one warning explaining that the adapter still refuses a degraded (synthetic) certificate. The behaviour is unchanged; the warning exists so the contradiction cannot be hit silently.

overall_verdict is one of:

Value Meaning blocked_at
SAFE every step cleared null
BLOCKED a single step was blocked on its own merits the step id
SEMANTIC_BLOCK the steps are individually fine but compose into harm — check semantic_triggers (e.g. structural_exfiltration, credential_exfiltration, recon_then_exfil) the step id
REVIEW_REQUIRED a step needs evidence or human approval null — nothing was blocked

Describe the artefact, not just the operation. Arcezia grounds its verdicts on concrete referents in the description — file paths, URLs, recipient addresses, credential and PII field names. "SELECT ssn, name FROM customers" names PII, so the read grounds as sensitive access and the chain above composes into SEMANTIC_BLOCK. "SELECT email, name FROM customers" names only column identifiers, so the same chain returns REVIEW_REQUIRED instead: still not SAFE, still not executable, but held for a human rather than positively identified as exfiltration.

The rule this reflects: a vague description degrades a verdict toward review — never toward approval. Arcezia never allows what it could not verify, so imprecision costs you review latency, not safety. The framework adapters get this right automatically because they pass the real tool arguments; it is worth attention only when you hand-build chain manifests.

Chain steps do not inherit the session's capability envelope, so action_within_task_scope stays unresolved and a chain will not reach SAFE on the envelope alone. Ground it per step with an evidence dict (the key is evidence — agent_evidence is ignored on chain steps):

{"id": "s1", "action_type": "execute_sql", "domain": "database_ops",
 "action_description": "SELECT ssn, name FROM customers",
 "evidence": {"action_within_task_scope": True}}

id and step_id are accepted interchangeably. Note that state_mutations may only add danger, never remove it: asserting a danger flag True is accepted, asserting it False is rejected, and flags the engine derives for itself (the g_* world-state namespace) are not caller-writable at all.

Audit after execution — did reality match the prediction?

toolkit.az.verify_outcome(
    action_type="execute_sql",
    action_description="DELETE FROM orders WHERE test = true",
    outcome={"rows_affected": 50000},      # what ACTUALLY happened
    expected={"rows_affected": 1},         # what you intended
)

Level 3 — ground the evidence. Register a probe webhook so evidence is GROUNDED rather than CLAIMED. Human intent can never be produced by a model, so ground it explicitly:

toolkit.az.authorize(user_token)              # user_explicit_authorization
toolkit.az.authorize_production(prod_token)   # production_explicit_authorization

These ground different constraints. Actions touching production generally need both — authorize() alone will leave production_explicit_authorization unresolved and the action stays in REVIEW.

Integrating over raw HTTP (n8n, curl, another language)? Two things the SDK handles for you: the API is behind a WAF that rejects the default library agent strings (e.g. Python-urllib/*), so send an explicit User-Agent of your own; and the precondition score is on the wire as precondition_score (with dc_score kept as a legacy alias for older consumers) — the SDK exposes it as cert.precondition_score.

Full guide: arcezia.com/docs

Enforcing at the resource

An ALLOW carries a single-use credential (cert.credential). The strongest pattern is an endpoint that refuses work without one: a gate that was skipped then has nothing to present, so the action cannot succeed. Placement of a check can be forgotten; a missing token cannot be.

Validate it from the resource before executing:

answer = az.validate_credential(cert)          # pass the certificate, not the token
if not answer["ok"]:
    refuse(answer["error"])                     # e.g. "action_digest_mismatch"

Passing the certificate is what makes the check strict. It sends cert.action_digest — the sha256 of the action the verdict was actually about — alongside the token, so the answer is "this credential was issued for this action". With only action_type, a credential minted for a single-row SELECT authorises a table-dropping statement of the same type in the same session.

A refusal comes back as {"ok": False, "error": …}; only a transport failure raises, and a raise means not validated — there is no degraded fallback here. Over raw HTTP the same call is POST /v1/validate_credential with token and action_digest; the n8n template forwards both as X-Arcezia-Credential and X-Arcezia-Action-Digest.

Development mode

Use an ar_test_ key for local development — infrastructure constraints (backup APIs, capability envelopes, CI gates) are relaxed so you are not blocked by production infra that does not exist on your laptop:

az = arcezia.Arcezia(api_key="ar_test_...", task="...")   # dev mode by default

Development mode is only available on ar_test_ keys and is re-checked server-side. Live ar_live_ keys are always pinned to production and cannot point at localhost.

Metadata

Release files for arcezia 1.0.6

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for arcezia 1.0.6
File Size Uploaded
arcezia-1.0.6.tar.gz 107.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for arcezia 1.0.6
File Interpreter ABI Platform
arcezia-1.0.6-py3-none-any.whl Python 3 none any Details

Total release size: 218.3 kB

Release files / arcezia-1.0.6.tar.gz

Download URL arcezia-1.0.6.tar.gz
Size 107.6 kB
Tags Source
SHA-256 checksum
How to use checksums
69410eb189277722fc6bb63ddeaf53aa8c70f7ca3cdf7c7f6fa26e4bf59cd0c0
BLAKE2b-256 checksum
How to use checksums
c71acd6d6ceb0b61a56d6bfa56b59f5d4aa805445867fbd5f1fbaae0179e87cd
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release files / arcezia-1.0.6-py3-none-any.whl

Download URL arcezia-1.0.6-py3-none-any.whl
Size 110.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
f38efb0941926361324ce2c45af5b7262ee4025c4db4bafcd03c46c9eee9b669
BLAKE2b-256 checksum
How to use checksums
1b4ff09bd22df1ec0c3a1d5a3209334d72d395bb8f291385bef42d18693351d3
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.11.9

Release history Release notifications | RSS feed

1.0.7

2 release files

This release

1.0.6 This release

2 release files

1.0.5

2 release files

1.0.4

2 release files

1.0.3

2 release files

1.0.2

2 release files

1.0.1

2 release files

1.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page