Zero-trust security runtime for AI agents: policy-mediated tool calls, tamper-evident audit, network and identity guards, adaptive security.
Project description
Warden
An explainable zero-trust security gateway for AI agents.
Warden sits between an AI agent and the tools it can call, and enforces security policy on every single tool call in real time. It exists because agents are increasingly handed real capabilities — file systems, shells, APIs, network access — with no firewall between the model and those tools. Warden is that firewall.
Every tool call flows through the proxy and receives one of four decisions:
- allow — forward it unchanged
- redact — strip secrets/PII from arguments and responses, then forward
- escalate — pause and require a human yes/no
- deny — block it and return a safe error
The guiding principle is deny by default: a tool or path not explicitly permitted by policy does not run.
Why
AI agents can be steered — by a malicious web page they scrape, a poisoned document they read, or a crafted user prompt — into exfiltrating data, escaping their workspace, or taking destructive actions. Most guardrails today try to make the model safer. Warden takes the complementary approach used everywhere else in security: put a policy-enforcing checkpoint between the untrusted actor and the sensitive resource, and log everything, tamper-evidently.
What's built (v1 → v7)
An installable runtime (pip install warden-security), proven by a 518-test synthetic attack, fuzz, and performance suite:
- Path canonicalization — blocks
../../etc/passwd, absolute-path escapes, and symlink escapes; everything is confined to a workspace root. - Argument parameterization — no shell strings, ever;
shell=Falseargument vectors make command injection impossible by construction. (No shell tool is enabled in v1 by design; the guard exists so the safe path is the only path available when one is.) - Secret & PII screening, both directions — credentials in tool arguments block the call outright; PII is flagged as risk and surfaced at the human-approval gate; tool responses pass through the redactor before they reach the agent's context.
- Indirect prompt-injection inspection — scans data returned from tools for hidden instructions before it reaches the agent's context, live in the transport (annotate / escalate / deny per policy).
- Real MCP transport — Warden physically between the client and a real MCP server (stdio JSON-RPC), with an execution watchdog, graceful drain, and generic errors to the agent (rule ids stay in the audit log, for the human).
- Tool-definition pinning — the trust boundary covers what a tool is, not just what it does. Tool schemas are canonicalized and SHA-256 hashed into a persistent, versioned registry at approval time; a server that swaps a tool's definition later (MCP rug-pull / tool poisoning) trips drift detection and is denied (rule PIN-001) until a human re-approves the new schema. Every register / approve / drift / reapprove event is on the audit chain.
- Egress allowlist — network destinations outside
egress.allowed_hostsare denied; the exfiltration half of the injection kill chain dies here. - Human approval gate — escalated actions pause for a yes/no on the terminal, with the full explainable decision shown; every failure mode (no terminal, timeout, error) resolves to DENY.
- Fail-closed pipeline — any internal failure becomes an audited DENY,
proven by fault-injection tests; monitor mode (
mode: monitor) computes and audits every decision without enforcing, for rollout and weight tuning. - Unicode hardening — NFKC, zero-width/bidi stripping, and homoglyph folding run before every inspector, so lookalike characters cannot evade detection.
- Tamper-evident audit log — hash-chained record of every decision,
independently verifiable with
warden verify. - Audit telemetry —
warden statsderives operational metrics straight from the forensic log (verdict counts, highest-risk tools, rule frequency, watchdog/injection/traversal counters): one source of truth, read-only. - Optional Presidio backend — opt-in richer PII detection behind the same detector interface; regex + entropy stay the always-on default, and a policy that enables presidio without the dependency present is rejected at startup rather than silently running weaker detection.
- v2 model-security layer — expanded adversarial detection (role
confusion, jailbreak scaffolds, hidden Unicode, markup smuggling, context-
window abuse), JSON-Schema validation of tool calls (rule SCHEMA-001), and
a measured detection posture: every classifier scored against a
labeled attack corpus with miss rates published per class
(
docs/DETECTION_POSTURE.md). All defense-in-depth on top of v1 enforcement, never a replacement for it. - v3 network-security layer — one ordered battery for every outbound URL (scheme, DNS sinkhole, global allowlist, per-tool scopes that narrow and never widen, domain reputation, SSRF resolve-then-validate with DNS-rebinding attribution), reused verbatim for every redirect hop; a download guard (executable magics, zip bombs judged without inflating them, nested/encrypted archives, base64-dressed binaries); token-bucket rate limiting; and canary tokens — the only detector permitted to claim certainty, because its false-positive cost is structurally zero. The whole subsystem runs against injectable resolvers: zero network I/O in its tests.
- v4 identity & trust layer — agent RBAC (the invoking user's role INTERSECTED with the deployment's scope, deny-by-default for unknown identities), HMAC-signed per-session capability tokens verified against canonical targets (single-use by default, replay refused, all revoked at the key when the session closes), per-capability approval policies with audit-chain history and escalation chains that cannot approval-shop, secure sessions that revoke keys BEFORE wiping workspaces, and signed, hash-chained, versioned agent memory with rollback pinning — tampering, truncation, and head-deletion all detected, reads verify before they return.
- v5 runtime-containment layer — the downstream MCP server runs inside a Warden-provisioned sandbox: the docker→gvisor→wasmtime isolation ladder with required-or-stronger selection (never a silent downgrade), an unbreachable floor (network none, read-only rootfs, cap-drop ALL, no-new-privileges), verified-destruction ephemeral workspaces, resource quotas validated at load with a host-held wall clock, and a process monitor for fork breaches, zombies, overstay, and unexpected executables. Rendered argv is the tested artifact — no container runtime needed to prove the posture.
- v6 adaptive-security layer — the payoff of monitor mode. Per-agent behavioral baselines learn each agent's normal and flag deviation (unfamiliar tool, unseen egress class, risk spike) as escalate hints, never autonomous denies; learning is gated, denied calls never teach "normal", and a suspect profile freezes. A user→agent→tool→file→network trust graph finds read-then-exfiltrate paths, privilege bridges, and blast radius no per-call guard can see. A read-only replay engine spends the recorded audit corpus to answer "what would this policy have done to the traffic we actually saw?" before rollout. Tighten-only adaptive controls — context floors, sticky human-clear-only quarantine, and intent verification against a stated goal — can only ever add caution, never lower a verdict below the static floor.
- v7 platform layer — Warden as a product, not a repo.
pip install warden-securitywith the runtime still exactly stdlib + PyYAML (enforced by package metadata); signed, attested releases (SHA256SUMS + Sigstore provenance, exact-pin lock file); framework adapters that run the same Normalize → Policy → Approve → Execute → Audit pipeline inside OpenAI Agents SDK, LangChain, AutoGen, and CrewAI processes with zero framework dependencies (MCP is served by the transport itself); a localhost-only, token-gated dashboard — audit chain, telemetry, sessions, live policy, and candidate-policy replay through the v6 engine — doubling as the desktop app; a VS Code extension (preview) for live policy decisions while developing; and signed, hash-pinned policy bundles with deny-by-default install as the shareable-policy format. - Tiered policy engine — reads auto-allow, writes/deletes escalate to a human, dangerous tools denied; deny by default.
See docs/THREAT_MODEL.md for the full defense-to-threat
mapping.
Warden and the landscape
Prior art exists and is worth naming. Guardrails libraries and safety classifiers (PromptGuard, Llama Guard) work at the model layer — they try to detect bad content. MCP gateways and API proxies exist for routing and observability. Warden's position is different on three axes: it is local-first (your machine, your policy, no cloud dependency), it is deny-by-default runtime enforcement rather than detection-only (a missed detection still hits the policy wall), and every decision is explainable and audit-chained (rule id, risk score, reason — verifiable after the fact). Detection layers are welcome on top; they are defense-in-depth, never the control. Warden assumes detection will sometimes fail and is built so that failing detection does not mean failing open.
Install
pip install warden-security # the runtime is stdlib + PyYAML, nothing else
Optional capability is opt-in via extras: [sign] (Ed25519 policy-bundle
signing), [pii] (Presidio-backed PII detection), [desktop] (native
dashboard window), [dev]. The distribution is warden-security (the name
warden on PyPI belongs to an unrelated project); the import package and the
command are both warden. Releases ship with SHA256SUMS and Sigstore
provenance attestations (gh attestation verify --repo Mecapixel/warden <file>).
Run it
warden init # starter policy + workspace
warden inspect read_file '{"path": "a.txt"}' # one explained decision
warden run -- <your-mcp-server-command> # mediate a live MCP server
warden verify # audit-chain integrity
warden stats # operational telemetry
warden dashboard # localhost web dashboard (token-gated)
warden desktop # the same dashboard in a native window
warden bundle pack|sign|verify|install # shareable, signed policy bundles
Inside an agent process (OpenAI Agents SDK, LangChain, AutoGen, CrewAI), the same pipeline guards plain-Python tool callables:
from warden.adapters import WardenGate, guard_autogen_map
gate = WardenGate("warden.policy.yaml", "audit.db")
safe_tools = guard_autogen_map(gate, {"read_file": read_file}) # DENY never runs
With no policy file present, Warden runs on a built-in deny-all default — unconfigured Warden is a wall, not a hole.
Run the tests
pip install -r requirements.lock
python -m pytest
Every test is a synthetic, benign attack payload — the structure of a real attack with no destructive function. No real malware or live exploits exist in this repository, by design.
Honest status
Built and tested (518 passing tests): the full v1 pipeline — request
normalization, path canonicalization, safe-exec guarding, secret/PII detection
and response redaction, inbound-injection heuristics, risk scoring, the
explainable Decision object, the hash-chained audit log, real MCP stdio
transport with watchdog and fail-closed guarantees, the human approval gate,
monitor mode — plus v1.5 registry hardening (tool-definition pinning with
drift detection, Hypothesis fuzzing that found and fixed two relay-crash
vectors), v2 model security (measured detection posture with published
per-class miss rates), v3 network security (the full outbound battery, whose
first test run found and fixed an escalate-masks-deny ordering flaw), v4
identity & trust (whose review found and fixed a self-referential hash bug
and a head-deletion rollback bypass in memory verification), v5 runtime
containment, v6 adaptive security, and v7 platform (packaging verified by
installing the built wheel into a clean venv and exercising the CLI through
it; framework adapters; signed policy bundles; the token-gated localhost
dashboard). Every phase is summarized in the What's built section above and
documented release by release in CHANGELOG.md.
python demo.py exercises the decision core end to end;
tests/test_transport.py proves the full story over real pipes.
Open, stated plainly: no shell-style tool is enabled by design (the
safe-exec guard waits, armed). Detection is heuristic-plus-measured, not
ML — every classifier is scored against a labeled corpus with miss rates
published in docs/DETECTION_POSTURE.md, and
detection is defense-in-depth either way: the policy layer is the control
that prevents harm. Containment renders and audits hardened sandbox specs;
running them end to end against live Docker/gVisor/Wasmtime backends is
deployment work, not library work, and is stated as such. Measured overhead
is published in
docs/PERFORMANCE.md: ≈ 2 ms per mediated tool call
end to end, deny path effectively free. The VS Code
extension ships as source in integrations/vscode/ (preview: run from
source or package with vsce; not yet on the Marketplace). The policy
marketplace exists as the format and tooling — signed, hash-pinned .wpb
bundles with deny-by-default install — not as a hosted registry; publishing
one is a service decision, not library work. Roadmap v1 through v7 are
complete.
License
AGPL-3.0 — the safeguards stay open and cannot be quietly stripped out.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file warden_security-7.0.0.tar.gz.
File metadata
- Download URL: warden_security-7.0.0.tar.gz
- Upload date:
- Size: 181.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
706eddc43e66316fffe90cb25ef244e80b18c9424d83bae90785d8de2e848bc6
|
|
| MD5 |
6fa82937f72f8a2c9034284e03957b7b
|
|
| BLAKE2b-256 |
d9488ce5b4e53f5860472f011e4aef70816b499dfc9f9c092130834ab7b0b0d6
|
Provenance
The following attestation bundles were made for warden_security-7.0.0.tar.gz:
Publisher:
publish.yml on Mecapixel/Warden
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
warden_security-7.0.0.tar.gz -
Subject digest:
706eddc43e66316fffe90cb25ef244e80b18c9424d83bae90785d8de2e848bc6 - Sigstore transparency entry: 2152322081
- Sigstore integration time:
-
Permalink:
Mecapixel/Warden@15363fa33f235e7439a3fe8089eee0705f0af556 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Mecapixel
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@15363fa33f235e7439a3fe8089eee0705f0af556 -
Trigger Event:
workflow_dispatch
-
Statement type:
File details
Details for the file warden_security-7.0.0-py3-none-any.whl.
File metadata
- Download URL: warden_security-7.0.0-py3-none-any.whl
- Upload date:
- Size: 149.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
81fe4c1c5be84c87f9e1da7ecb1bcd03fde3550902856bbecec87bdd8b606df7
|
|
| MD5 |
77f0b0a324045af2b628b02cb74b4760
|
|
| BLAKE2b-256 |
6d1e06105ea2b22f7b027c3742bffc5d303378f6884d534a3a51081d42a1f294
|
Provenance
The following attestation bundles were made for warden_security-7.0.0-py3-none-any.whl:
Publisher:
publish.yml on Mecapixel/Warden
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
warden_security-7.0.0-py3-none-any.whl -
Subject digest:
81fe4c1c5be84c87f9e1da7ecb1bcd03fde3550902856bbecec87bdd8b606df7 - Sigstore transparency entry: 2152322199
- Sigstore integration time:
-
Permalink:
Mecapixel/Warden@15363fa33f235e7439a3fe8089eee0705f0af556 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Mecapixel
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@15363fa33f235e7439a3fe8089eee0705f0af556 -
Trigger Event:
workflow_dispatch
-
Statement type: