Runtime firewall for AI agents - blocks catastrophic tool calls and self-heals the run.
Project description
๐ก๏ธ AgentX: The Action Firewall for AI Agents
This package (agentx-security-sdk) is MIT licensed. The gateway and control plane are separate, closed-source products; see Licensing below.
An agent that can act will eventually act badly: a DROP TABLE arriving through a prompt-injected field, a secret read, an SSRF to a cloud metadata endpoint, an installer piped straight into a shell. Most agent frameworks already give you a place to stop a tool call before it runs. What they don't give you is the decision about which calls are worth stopping.
AgentX is that decision, and it starts with zero keys. The Shield is a deterministic floor that runs inside your process, with no LLM call and no network hop. It hard-blocks the never-legitimate call (DROP TABLE, SSRF, secret reads, supply-chain RCE, destructive shell and cloud teardown) and escalates to a human the ones that can be legitimate (large transfers, external publishes, runaway spend, bulk deletes). No API key, no signup. The only outbound call it ever makes is the dependency-reputation check against the public npm and PyPI registries.
A block is not a dead end. Every block hands your agent a coaching message naming a safe path to try instead, so the run continues rather than ending. Add a Gemini key and the gateway's reasoning layer writes that message to fit the specific task rather than from a template, catches what a keyword floor cannot see, and runs the retry for you.
๐ก๏ธ Block the catastrophic deterministically. ๐ง Coach the recoverable when you add a key.
โก 1. Quickstart
AgentX needs no changes to your agent logic, your tools, or your payload schemas. It reads your function signatures at runtime and works out what to inspect on its own.
Step 1: Install the SDK
pip install agentx-security-sdk
Step 1b: See it work in 10 seconds (no key, no gateway)
agentx demo
Runs a canned agent that tries a DROP TABLE and watches the in-process shield block it offline, the fastest way to confirm the install works before you wire it into your own agent.
Hit a block you're proud of? Turn it into a postable card:
agentx share
Renders your most recent catch as a clean, screenshot-able receipt + a ready-to-post draft and link. Privacy-safe by construction: it uses the policy class and your own tool name only, never the query or payload (the local ledger never stores one).
Step 2: Decorate Sensitive Tool Operations
Attach the @agentx_protect decorator over any high-risk system tool. The SDK automatically serializes parameters and enforces the evaluation wedge:
# โ
MODERN REFLECTIVE IMPORTS (No boilerplate functions required)
from agentx_sdk.decorators import agentx_protect
@agentx_protect(agent_id="demo_frictionless_agent")
def dispatch_crm_update(client_id: str, profile_notes: str, db_session=None):
"""
AgentX automatically inspects string elements, ignores connection objects
like 'db_session', and evaluates intents out-of-prompt natively in RAM.
"""
print(f"Updating records for {client_id}")
Step 2b: Handle the Block
Decorating is half the job. Your code also has to react when AgentX blocks a call. You never parse the message text. Use is_block() and read the structured fields:
from agentx_sdk import agentx_protect, is_block
result = dispatch_crm_update(client_id="CLI-99401", profile_notes=untrusted)
if is_block(result):
print(f"Blocked by policy: {result.policy}")
llm.send(result.challenge) # feed the safe-path challenge back to your agent to self-correct
else:
use(result) # not blocked โ the real return value
For strictly-typed tools (e.g. LangChain / Pydantic tools that validate a -> dict return), AgentX raises instead of returning, so the framework doesn't crash. Catch it and feed the same challenge back:
from agentx_sdk import AgentXSecurityBlock
try:
data = fetch_user(uid) # -> dict
except AgentXSecurityBlock as block:
llm.send(block.challenge)
A circuit-breaker trip (runaway loop) is not a policy block. It raises
AgentXCircuitBreakerTrippedandis_block()returnsFalsefor it. Catch it separately to abort the run.
Step 2c: Or protect an MCP server with zero code
Don't own the tool's Python? Running a non-Python agent? Wrap any MCP server with agentx-mcp and every tools/call is screened by the same keyless Shield before it runs. No decorator, no key, no code change. Just one line in your mcp.json (Claude Code, Cursor, or any MCP client):
{
"mcpServers": {
"filesystem": {
"command": "agentx-mcp",
"args": ["npx", "-y", "@modelcontextprotocol/server-filesystem", "/data"]
}
}
}
agentx-mcp <real server command> spawns the real server and relays the protocol untouched, intercepting only tool calls. A blocked call comes back to the agent as a coaching tool error it can self-correct on, so the run keeps going and the dangerous call never reaches the server.
Beyond the calls your agent makes, agentx-mcp also watches the server's advertised tools for a bait-and-switch: a malicious or compromised server can advertise a benign tool at install (you approve it once), then silently rewrite that tool's description or schema on a later run to steer your agent. The proxy fingerprints each tool the first time it sees it and warns you when an already-approved tool's definition later changes (the NSA's 2026 MCP guidance names this exact attack). It runs in advisory mode by default (a loud warning, never breaks a run); set AGENTX_MCP_TOOL_PINNING=block to also gate calls to a changed tool until you re-verify it, or off to disable.
No Python in your stack? You don't need a persistent install. Run it on demand with uvx (or pipx run), which fetches the package into a throwaway environment, so your mcp.json stays one line:
{
"mcpServers": {
"filesystem": {
"command": "uvx",
"args": ["agentx-mcp", "npx", "-y", "@modelcontextprotocol/server-filesystem", "/data"]
}
}
}
(pipx run agentx-mcp <real server command> works the same way.)
Step 3: See your first block: no key, no gateway, no signup
This is the point of AgentX: the deterministic Shield runs inside the SDK. The catastrophic call is blocked in-process, with zero keys and nothing else running, no LLM, no gateway.
If you installed from PyPI, agentx demo (Step 1b) is this, and it needs nothing else. If you have the repo checked out, the same thing with the source visible:
python examples/08_frictionless_agent_protection.py # requires the repo, not the pip install
The script wraps an ordinary CRM function, then feeds it a prompt-injected DROP TABLE users;. The Layer-0 keyword shield intercepts it before it executes, with no .env, no API key and no gateway required. You'll see the block and a clean session summary:
๐น๏ธ Console Session Output Logs:
========================================================================
๐ค AGENTX DEMO 08: ZERO-CONFIGURATION ENTERPRISE TOOL PROTECTION
========================================================================
Scenario: An engineer wraps an existing corporate function with AgentX.
The SDK automatically parses inputs via Python signature reflection.
๐ --- Agent Execution Step ---
Agent Attempting Call: dispatch_crm_update(client_id='CLI-99401', profile_notes='...')
Injected Input Payload: 'Customer requested normal update. Retain account historical state; DROP TABLE users;'
๐ก๏ธ [AgentX SDK] Intercepting tool call to 'dispatch_crm_update' and active_stats = 0...
โก [LOCAL KEYWORD SHIELD] Fast-path intercept engaged on policy 'Mass Destructive Intent' (offline, no LLM judge).
๐ [LOCAL BLOCK] Policy 'Mass Destructive Intent' matched a blocked intent locally.
๐ [AGENTX SHIELD] Request Intercepted & Blocked (deterministic floor โ no key, no LLM)!
-> The reflection engine successfully captured the nested SQL payload.
-> The SQLAlchemy session context object was safely ignored.
๐ Enterprise data assets protected via zero-configuration injection monitoring.
========================================================================
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ก๏ธ AgentX Session Summary (Trace: bbfda7a1-8e1b-41dd-8c59-b784a3bbdf6d)
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โฑ๏ธ Uptime: 0.27 seconds
๐ ๏ธ Tools Monitored: 1
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ Intercepts: 1 | Cumulative: 1
๐ฅ Critical Blocks: 1 | Cumulative: 1
๐จ Human Escalations: 0 | Cumulative: 0
๐ Self-Corrections: 0 | Cumulative: 0
๐ Recovery Rate: 0.0% (keyless block โ add a key for recovery, Step 4)
๐ฐ Tokens Saved: ~1500
โณ Time Saved: ~5 min
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
Step 3b: Customize the coaching (keyless, by name)
Every block above hands your agent a coaching message: what went wrong, plus a safe path to try instead. That wording is what drives whether the agent recovers, so you can override it for any built-in floor policy, by name, keyless. No UUID, no gateway.
List the policies you can customize and the coaching each one ships with:
agentx policies
Then override one by its name, inline or in your editor:
agentx customize "Mass Destructive Intent" --text "Our house rule: never DROP; snapshot then soft-delete first."
agentx customize "Network Sandbox (SSRF)" --edit # opens $EDITOR seeded with the current coaching
Your customized coaching is saved to ./.agentx/overrides.json (commit it to share with your team) and applies keyless on both surfaces: the @agentx_protect decorator and the agentx-mcp proxy. Add --safe-path "..." to set the concrete safe alternative distinctly from the coaching. Validate the store anytime, so a hand-edit typo is loud instead of a silent disable:
agentx policies --check
Step 4: (Upgrade) Recover: turn hard blocks into recoverable challenges
Everything above is keyless Shield: it blocks the blatant calls and coaches your agent to self-correct. The reasoning axis (orthogonal to AGENTX_MODE) is Recover: the gateway judge catches what the keyword floor can't see, writes the task-fitting challenge when your policy carries none, and runs the coach-and-retry for you, plus Discovery for novel-intent classification. It takes two things, and they go together:
1. A GEMINI_API_KEY, the cheap part. Copy .env.example to .env and set it:
# Reasoning axis โ OPTIONAL. Unset = Shield (floor-only, zero LLM, zero keys).
GEMINI_API_KEY=your_gemini_key_here # from aistudio.google.com
# AGENTX_REASONING=off # force Shield even when a key is present
2. The gateway running with that key (Step 5), because the judge that writes the task-fitting challenge lives in the gateway, not the SDK. The gateway is closed-source, so you run a prebuilt private image. Drop your email at agentx-core.com/gateway and you get a one-paste docker command back, no wait. This is the wedge worth the two minutes.
The recovery demo needs the key (it makes a real LLM call to re-plan after a block; run it keyless and it points you to example 08 instead):
python examples/01_self_healing_agent.py
Run it with the key but without the gateway and it still completes, but fail-open: the Layer-0 shield applies a static challenge and the deep semantic checks are skipped. The demo says so itself (
โ ๏ธ RECOVERED, DEGRADED โฆ start the gateway to verify). The verified, judge-written recovery, the actual Recover tier, is the gateway path in Step 5.
Step 5: (Upgrade) Run the gateway: the reasoning judge (Recover), telemetry, control plane, HITL
The SDK protects you on its own (keyless Shield). The gateway is what unlocks Recover from Step 4: with a GEMINI_API_KEY set, the judge here writes the verified, task-fitting challenge (the keyless path only ever gets a static one). Run the gateway and the Next.js console to add the data plane (the reasoning judge, deep AST evaluation, the policy lifecycle) and the control plane (the ROI dashboard, Chain-of-Thought review, human-in-the-loop approvals). This is the AGENTX_MODE data axis, where data lives and whether it syncs:
# local isolated; local SQLite store; no sync; no keys required
# linked local-authoritative; explicit pull/push; no auto-sync
# cloud control plane authoritative; continuous sync + upload + HITL/SOC
AGENTX_MODE=local
NEXT_PUBLIC_AGENTX_MODE=local # the UI's build-time copy; keep it equal
# AGENTX_API_KEY=agentx_sk_your_key_here # required for `cloud` (and remote `linked`)
Connecting to the cloud is one variable. You don't need to set all three. An
AGENTX_API_KEY is only ever meaningful for cloud upload, so just setting it (with
no AGENTX_MODE/CONTROL_PLANE_URL) puts the gateway in cloud mode against the
public plane, and it says so loudly at boot (๐ AGENTX_API_KEY detected โ CLOUD
mode โฆ). Set AGENTX_MODE=local to override and stay isolated, or linked (with
a CONTROL_PLANE_URL) for explicit pull/push without auto-upload.
Boot the data-plane wedge and the dashboard:
docker compose up -d
OR
cd backend
uvicorn gateway:app --host "0.0.0.0" --port=8000
Both commands above build the gateway from source (internal / source-access). The gateway is closed-source, so you run a prebuilt private image instead,
pip install agentx-security-sdkdoes not ship it. Drop your email at agentx-core.com/gateway and you get a one-pastedockercommand back immediately. No sales form, no wait. The starter kit indeploy/partner/then runs the image (no source needed).
The reasoning engine then listens on http://localhost:8000 and the dashboard on http://localhost:3000. In local/linked mode the gateway owns the policy lifecycle from a local SQLite store (seeded on first boot, editable via the Policy & Discovery tabs); in cloud mode it mirrors policies from your Supabase-backed Control Plane.
Safety posture (optional). By default AgentX fails open, if the gateway is unreachable, tool calls still execute (with a loud warning and an audit tally in the session summary), and the in-process keyword shield still blocks deterministic threats like DROP TABLE. For high-stakes or irreversible actions, set AGENTX_FAIL_MODE=closed to instead block any call the engine can't verify until it recovers.
Enforcement level: try it before you enforce it (optional). Set AGENTX_ENFORCEMENT=audit to run the SAME detection non-blocking: AgentX records what it would have blocked and lets every call proceed, so you can run it in staging for a week and see exactly what it would have caught (and anything it would have caught wrongly) with zero risk. agentx insights then shows the would-have-blocked report, per policy. Flip to AGENTX_ENFORCEMENT=enforce (the default) when the catches look right. Keep one dangerous tool hard-blocked even while auditing with a per-tool override: @agentx_protect(..., enforcement="enforce"). Audit is a distinct posture from the fail-mode above; the circuit breaker and fail-closed still stop even in audit.
๐ Control Plane Telemetry
When your agent script finishes or exits, AgentX writes its telemetry to the local .agentx.db. In cloud mode that also syncs to your control plane; in local and linked mode it stays on the machine. With the dashboard running, open http://localhost:3000/dashboard to see the summary. The numbers below are illustrative, not measured results:
+-----------------------------------------------------------------------------------+
| EXECUTIVE COMMAND CONSOLE |
+-----------------------------------------------------------------------------------+
| [Catastrophic Actions] [Autonomous Recovery] [Runs Protected] [Time Saved] |
| 92 36.3% 37 12.3 hrs |
| ๐ Irreversible/exfil ๐ Self-corrected รท ๐ก๏ธ Runs that โฑ๏ธ ~20 min/run |
| stopped pre-exec challenged loops self-corrected reclaimed |
+-----------------------------------------------------------------------------------+
Executive ROI Mappings (all computed per session, grouped by trace_id):
- Catastrophic Actions Blocked (hero): A pure count of distinct sessions whose intercepted action fell in an irreversible / exfiltrative class,
failure_mode โ {DESTRUCTIVE_ACTION, PII_EXFILTRATION, NETWORK_TRAVERSAL, SECRETS_LEAK}, with a policy-name keyword fallback. Every incident in the ledger is a pre-execution interception, so this is harm averted, not harm survived. - Autonomous Recovery Rate: Of the sessions that entered the challenge loop, the share whose terminal status is
COMPLIED(the agent self-corrected). Counted per session sorecovered โ challenged, bounded โค100% by construction. HITL-approved sessions are excluded; only autonomous self-correction counts. - Agent Runs Protected: Sessions the agent self-corrected after a block (terminal
COMPLIED). - Engineering Time Saved: ~20 min of manual triage credited per protected run, valued at $75/hr.
Operator dashboard metrics are scoped to production agents only, demo, simulation, blind-eval, and test/probe traffic are excluded so benchmarks never inflate an operator's numbers (they showcase the engine on the public landing page instead).
๐ง The 5 Pillars of Agentic Security
AgentX is built on a "Reasoning Engine" architecture that treats AI agents as autonomous employees rather than static scripts:
- Cognitive Interception: We intercept tool calls to compare the agent's stated intent (Chain of Thought) against its actual deterministic action.
- Socratic Nudging: Instead of crashing the agent, we issue a Socratic Challenge to guide them to a safe, desired end-goal.
- Shared Immunity Network (roadmap): Novel zero-day signatures discovered on one node are designed to graduate into the deterministic floor and propagate to other Edge nodes for O(1) interception. The local Discovery โ Promote โ live-in-3s loop works today; cross-node global distribution is a Day-100 capability and is not yet active, we don't claim it until it is.
- Circuit Breakers: If an agent enters an infinite hallucination loop, AgentX hard-locks the runtime after 3 strikes to prevent massive LLM token billing overages.
- Human-in-the-Loop (HITL): If an agent pulls the "Andon Cord" (requests help), the system suspends the execution thread (
202 Accepted) and parks it in the SOC Sandbox for human approval.
๐ The 4 Shields (Defense-in-Depth)
-
The Inbound Shield (Prompt Injection): Sanitizes inbound user text to prevent cognitive hijacking ("Ignore previous instructions") before the agent reads it.
-
The Logic Shield (Database Guard): Uses AST parsing and Gemini to catch destructive queries (DROP, DELETE) and nudges the agent to write safer SQL.
-
The Network Shield (SSRF Guard): Prevents agents from acting as confused deputies to hit cloud metadata IPs (e.g., 169.254.169.254).
-
The Egress Shield (DLP/PII Scrubber): Dynamically masks PII and API keys on the wire, maintaining clean audit logs without triggering SOC alert fatigue.
๐ Local Telemetry & Agent Health
AgentX ships with a built-in, privacy-first SQLite time-series event log (.agentx.db). It tracks every interception locally. When your agent script finishes or crashes, AgentX automatically prints a comprehensive Session Summary and Lifetime ROI dashboard:
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ก๏ธ AgentX Session Summary
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โฑ๏ธ Uptime: 9.17 seconds
๐ ๏ธ Tools Monitored: 2
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ Intercepts: 1 | Cumulative: 5
๐ฅ Critical Blocks: 1 | Cumulative: 5
๐ฐ Tokens Saved: ~1500 | Cumulative: ~7500
โณ Time Saved: ~5m | Cumulative: ~25m
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ฉบ AGENT HEALTH INSIGHT
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ ๏ธ Top Offender: 'Database Isolation'
๐ ๏ธ Tip: Consider refining your agent's system prompt to avoid this.
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
๐ฆ Try the other Developer Demos
These live in the examples/ folder of the repo, so they need a checkout rather than the pip install:
- 01_self_healing_agent.py: Watch AgentX catch a hallucination and coach the agent to self-correct (Saving tokens and uptime).
- 02_cognitive_intent_block.py: Watch AgentX catch malicious intent even when the raw syntax is perfectly safe.
- 04_circuit_breaker_demo.py: AgentX catches and prevents an infinite apology loop, saving time and tokens.
- 06_hitl_escalation.py: See how an agent safely pauses execution and pings a SOC analyst for approval using a 202 Accepted queue.
- 09_budget_ceiling_demo.py: Watch AgentX meter a runaway agent's cumulative spend and halt the session the moment it crosses your budget ceiling, a deterministic, gateway-side escalation (no LLM judge). Needs the gateway running + a key.
- 10_self_correction_coaching.py: The Recover tier, where the gateway judge does the coaching for you. AgentX blocks a dangerous action and returns a task-fitting challenge that names a safe path, so your agent finishes the job instead of dead-ending. Puts the AgentX challenge side by side with a bare refusal, then watches the agent recover on it. Needs the gateway running + your Gemini key (the keyless Shield already coaches in-band; Recover's judge catches more and runs the retry for you).
- And many more...
๐น๏ธ Human-in-the-Loop (HITL) & Control Plane
Sometimes, an agent needs to drop a table for a valid business reason.
AgentX ships a Next.js Control Plane dashboard. If an agent requests an escalation, the SDK pauses local execution and polls the gateway. A human reviewer clicks "Approve" or "Deny" in the UI, and the Python execution loop resumes on its own.
Running the dashboard needs the repo, not the pip install:
cd ui
npm install
npm run dev
๐๏ธ The Architecture (Split-Plane)
Three pieces. Only the first one is required, and it is the only one that ships in the pip package.
- The Edge SDK (
agentx_sdk): The Python package that instruments your tools and carries the keyless deterministic floor. Runs in your process. This is whatpip installgives you. - The Data Plane (Reasoning Engine): A FastAPI service that evaluates what the floor waves through, via a Layer 0 โ Layer 1 โ Layer 2 funnel: the deterministic floor first, then AST evaluation, then the LLM judge. Closed source, run as a prebuilt image.
- The Control Plane (Dashboard): A Next.js app where a human reviews intercepted traffic, reads chains of thought, and approves or denies parked requests.
- Where state lives: Mode-dependent. In
cloud, Supabase is authoritative and both planes sync through it. Inlocalandlinked, the gateway's own SQLite stores (.agentx/incidents.db,.agentx/policies.db) are authoritative and nothing syncs. See What leaves your machine. - The Evaluator: Google Gemini 2.5 Flash or Pro, configurable by environment variable. It reads the agent's stated intent and evaluates it against your policies. Only used on the gateway path; the keyless floor never calls it.
โจ Key Features & Built-in Policies
- Automated Socratic Self-Healing: Intercepts dangerous tool calls and challenges the agent to revise its strategy.
- Fast Pass Heuristic Traps: Instantly intercepts structurally dangerous queries (e.g.,
DROP TABLE, or an unscoped /WHERE 1=1massDELETE/UPDATE) with minimal latency. - Zero-Knowledge Intent Extraction: Prevents malicious prompt injection by translating raw agent logic into a strict schema before policy evaluation.
- Dynamic Policies: In
cloud, enforces isolation rules via a Supabase-backed Control Plane that syncs to edge caches in ~3 seconds. Inlocal/linked, the gateway owns the policy lifecycle locally, create/edit/toggle/delete and AI-drafted promotions from the Policy & Discovery tabs are armed live (re-embedded into the in-RAM vector index) with no restart.
๐ What leaves your machine
Keyless, in local mode, the answer is almost nothing. Your queries, payloads and agent chain-of-thought stay on the machine and are never uploaded unless you explicitly push them. Two things do leave, and both are narrow and named:
- The dependency-reputation check. Sends package names only to the public npm and PyPI registries, to catch slopsquats.
- An anonymous daily usage pulse. Version, OS and block counts. Never your code, queries, chain-of-thought or identity. It is on by default, prints a one-time notice before the first pulse, and turns off with
AGENTX_TELEMETRY=off. See.env.example.
Contributing to the shared corpus is separate and off by default. It happens only when you push it (agentx push, or agentx sync which pulls policies and pushes in one step), and only after you opt in with AGENTX_CONTRIBUTE. Left unset on a terminal, it asks once and saves your answer; left unset in CI or a script, it stays off and never blocks. What it sends is abstract: which policy fired and when, never a query, payload, chain-of-thought or identifier. It also needs the gateway, because the gateway is what reduces your local records to abstract signal before anything leaves the machine.
๐ Licensing
This package is MIT. The agentx_sdk/ edge client, published to PyPI as agentx-security-sdk, is licensed under the MIT License. The agentx-mcp launcher is MIT as well. Use them freely, including commercially.
The gateway and control plane are not. The Reasoning Engine and Control Plane are closed source and proprietary, all rights reserved. No public open-source or source-available license is granted for them. Source access for evaluation is available to qualified customers and partners under written agreement.
๐ Roadmap & Milestones
โ Trust Boundary Shift: Moved neuro-symbolic evaluation entirely into the Data Plane container to eliminate agent runtime bypasses. (Completed)
โ Hard Split-Plane: Telemetry is stripped of payload and chain-of-thought content at the edge, so what crosses the plane boundary is counts and classes rather than your data. (Completed)
โ Zero-Config Reflection Engine: Eliminated manual query and CoT boilerplate writing using dynamic signature parameters compilation hooks. (Completed)
โ Local Keyword Shield (Layer 0): Deterministic, dependency-free keyword/intent pre-filter in the SDK that intercepts obvious threats offline, in-process, with zero gateway/LLM calls. Scans the action payload only; chain-of-thought intent is deferred to the gateway's LLM judge. (Completed)
โ Judge Verdict Memoization: Bounded in-memory cache on the Data Plane that reuses prior LLM verdicts for identical (payload + reasoning + policy set), eliminating repeat Gemini calls during agent retry loops. (Completed)
โ Catastrophic-Action Hero Metric: Reframed the Executive ROI strip to lead with severity-filtered "Catastrophic Actions Blocked" (irreversible / exfiltration intents stopped pre-execution), with per-session metric accounting that bounds Recovery Rate โค100% by construction across the dashboard, the Supabase summary view, and the SDK. (Completed)
โ
Detection-vs-Recovery Eval Harness (eval/): Independent instruments that measure the engine honestly, blind_agent_eval.py (end-to-end detection recall via a blind LLM agent + independent oracle), probe_judge.py (isolates the reasoning layer's marginal recall over the deterministic floor), and recovery_eval.py (A/B marginal-recovery lift of the Socratic challenge vs a bare 403). (Completed)
โ
Incident-Persistence Hardening & Fail-Mode Switch: Restored the CHALLENGEDโCOMPLIED persistence pipeline (gateway-pinned UUID receipts, /v1/incident Layer-0 registration, COMPLIED PATCH gated on a real 200) and added AGENTX_FAIL_MODE=open|closed. (Completed)
โ
Deterministic Floor, Hard-Block + HITL-Escalation + Loop-Abort Tiers: The zero-LLM core that runs before (and without) the judge, so the engine fully protects keyless. Hard-block tier DENIES never-legitimate actions (destructive DDL/DML incl. ALTER โฆ DROP COLUMN, cluster/cloud teardown, SSRF, secret reads + egress exfiltration, filesystem whole-scope deletes + path-boundary escapes, remote-pipe-to-shell installs, and bidi-override / Unicode-Tags carrier payloads, Trojan-Source / invisible-instruction smuggling). HITL-escalation tier returns 202 ESCALATED โ human SOC for consequence-gated actions that can be legitimate: High-Value Transfer Approval (AGENTX_TRANSFER_ESCALATION_THRESHOLD), External Publication Approval, Comms Bulk-Deletion Approval, and Budget Ceiling Approval (cumulative session token/$ spend vs AGENTX_SESSION_TOKEN_CEILING / AGENTX_SESSION_COST_CEILING_USD; report real usage with agentx.record_spend(...) or rely on the built-in volume estimate). Loop-abort tier terminates runaway loops (strike-count breaker + the detect_no_progress_loop no-progress repeat breaker, AGENTX_LOOP_REPEAT_CEILING). Every detector has a fires-in-anger test asserting attribution + zero LLM calls. See AGENT_FAILURE_CATALOG.md for the per-incident coverage state. (Completed)
โฌ Downloadable Vector Seeds (agentx compile): Real pre-compiled fastembed vectors for offline semantic matching, scoped to air-gapped deployments. (Future)
โฌ Containerized Multi-Region Edge Cluster: Standardize container blueprints for automated high-availability deployments onto AWS ECS and Render clusters. (Future)
๐ค Support & design partners
- Docs: agentx-core.com/docs
- Community and support: the Discord invite is on agentx-core.com, which always has the current link.
- Issues in the SDK: the MIT client is mirrored at github.com/vdalal/agentx-security-sdk, which is where SDK issues belong. This monorepo is private, so an issue opened here will not reach anyone.
If you are running agents that write to systems you care about and want to compare notes on what they try, we are actively looking for design partners.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file agentx_security_sdk-0.4.26.tar.gz.
File metadata
- Download URL: agentx_security_sdk-0.4.26.tar.gz
- Upload date:
- Size: 249.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6204039d25aeb14d4e3c6e05801b9e1321a6950ddea273c5e3d431bf21017049
|
|
| MD5 |
214687c4e62e52b90d3a0e8fe57b44e6
|
|
| BLAKE2b-256 |
e3bcfbd7036e00505e9973b895c122dc82b7740d2629cd7e4e9153d0ff40dd1f
|
File details
Details for the file agentx_security_sdk-0.4.26-py3-none-any.whl.
File metadata
- Download URL: agentx_security_sdk-0.4.26-py3-none-any.whl
- Upload date:
- Size: 232.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.1
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a5d54c6082ddfdcdc7f1d5d67952fac91f6338e677e4fb1f920f439556121096
|
|
| MD5 |
d052ef87e43790aff0919fd19d9807ba
|
|
| BLAKE2b-256 |
edd68be198f7467c2d7dab3422682da7ac7a43f799d1cfecf93709a42f3abe56
|