Skip to main content

CodeShield: Autonomous Code Execution Engine

CI Python License: MIT Code Style: Ruff

Deterministic, Isolated, and Self-Healing Python Execution Runtime for AI Agents.

This open-source engine executes Python code generated by LLMs inside a disposable, isolated sandbox, validates it statically with the Python ast module, and recovers from runtime errors through a deterministic self-healing loop backed by local heuristics and plug-and-play LLM-guided patch generation (OpenAI, Anthropic Claude, DeepSeek, Ollama, Google Gemini).

Comparison

Feature Vanilla subprocess Docker Container CodeShield (This Engine)
Startup Overhead ~5 – 10 ms ~1,500 – 3,000 ms Sub-second (~30–250 ms via uv)
Isolation Mechanism None (Host Process) Container Namespaces / cgroups Ephemeral Virtualenv (tempfile + uv)
AST Security Gate ❌ None ❌ None ✅ Static AST inspection (os.system, eval)
Silent Failure Detection ❌ None ❌ None ✅ Regex scanning for empty DataFrames/NaNs
Self-Healing Loop ❌ None ❌ None ✅ 3-Tier Traceback Diagnosis + LLM Patch (Any Provider)

Architecture

[LLM Generated Code]
        │
        ▼
[AST Static Gate] ──(Syntax/Security Violation)──► [Validation Error Report]
        │ (Passed)
        ▼
[uv Isolated Sandbox] ──(Runtime Error/Silent Failure)──► [Traceback Classifier]
        │                                                          │
        │ (Clean Execution: exit 0)                                ▼
        ▼                                               [LLM Self-Healing (Any Provider) / Local Heuristic]
[Verified Output (JSON)] ◄──(AST Validated Patch)─────────────────┘

Key Features

1. Ephemeral Sandboxing with Dual-Mode Backend

  • Primary: uv venv for ultra-fast environment creation and package installation.
  • Fallback: native python -m venv + pip when uv is unavailable, so the engine works out of the box on any machine.
  • Each execution lands in its own temporary workspace that is destroyed after use.

2. Deterministic AST Security Gates

The engine parses every snippet with the standard ast module and rejects:

  • SyntaxErrors before execution.
  • Bare except: / except Exception: / except BaseException: handlers.
  • Calls to dangerous parametrizable functions: eval(), exec(), compile().
  • Calls to system/subprocess primitives: os.system(), subprocess.call(), subprocess.run(), subprocess.Popen().

3. Silent Failure Detection

Even when a process exits with 0, the engine flags suspicious output patterns such as:

  • empty DataFrame
  • all NaN
  • Traceback
  • Pipeline failed
  • Fatal Error

4. Model-Agnostic Self-Healing Loop

AST Validation ──► Sandbox Execution ──► Traceback Classification ──► Patch ──► Re-run
                                      (3 attempts max)
  • Local heuristic fallback: handles NameError, ImportError, ModuleNotFoundError by injecting safe imports or placeholder definitions.
  • LLM-guided healing: when an LLM is configured (built-in Gemini Flash by default, or any custom provider via patch_generator), it asks the model for a corrected version of the code, validates it with the AST gate, and re-executes the patched snippet.

Quickstart

Installation

# Install from PyPI
pip install codeshield-runtime

# Install with all extras (LLM + Dev tools)
pip install "codeshield-runtime[llm,dev]"

# Or clone for development
git clone https://github.com/AlgorithmicMind/codeshield.git
cd codeshield

# With uv (recommended)
uv venv
uv pip install -e ".[test,lint,llm,dev]"

# Or with pip
python -m venv .venv
.venv\Scripts\activate  # Windows
pip install -e ".[test,lint,llm,dev]"

Offline Usage (No API Key)

from codeshield.loop import SelfHealingEngine

engine = SelfHealingEngine(use_llm=False)
with engine:
    result, diagnosis = engine.run("print('hello world')")
    print(result.stdout)

Model-Agnostic Self-Healing (Plug-and-Play)

CodeShield is not locked into a single LLM. Pass any Python callable as the patch_generator to use OpenAI, Anthropic Claude, DeepSeek, Ollama, LiteLLM or your own service:

from codeshield.loop import SelfHealingEngine


def custom_openai_patcher(code: str, diagnosis) -> str:
    # Any LLM call (OpenAI, Anthropic, DeepSeek, Ollama, LiteLLM)
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {
                "role": "user",
                "content": f"Fix this code:\n{code}\nError: {diagnosis.message}",
            }
        ],
    )
    return response.choices[0].message.content


engine = SelfHealingEngine(patch_generator=custom_openai_patcher)

Zero-Config Self-Healing with Gemini Flash

For the built-in zero-config experience, create a .env file from .env.example:

GEMINI_API_KEY=your_key_here
GEMINI_MODEL=gemini-3.7-flash
from dotenv import load_dotenv
from codeshield.loop import SelfHealingEngine

load_dotenv()

engine = SelfHealingEngine()
with engine:
    result, diagnosis = engine.run('print("Result: " + 42)')
    print(result.stdout)  # Result: 42

Run the included demo:

python demo.py

CLI Usage

Execute any Python file directly from the terminal with the built-in CLI:

python -m codeshield run script.py
python -m codeshield run script.py --timeout 30
python -m codeshield run script.py --llm        # try LLM self-healing if configured
python -m codeshield run script.py --no-llm     # force local fallback

🤖 Agent Tool Integration (LangChain, CrewAI, OpenAI, Gen AI)

from codeshield import create_code_execution_tool

# Pass the tool directly to your agent
tools = [create_code_execution_tool()]

create_code_execution_tool() returns a ready-to-register execute_python_code(code: str) -> str function. It runs the provided Python in a self-healing sandbox and returns either the stdout or a structured error report with error_type and stderr.


Verified Examples

The examples/ folder contains ready-to-run recipes that have been executed and verified:

  • 01_basic_sandboxing.py: isolated execution with timing measurements.
  • 02_security_gatekeeper.py: AST rejection of unsafe code.
  • 03_llm_healing_workflow.py: self-healing workflow with an LLM or local fallback.
  • 04_agent_tool_dropin.py: end-to-end agentic tool-calling workflow with dynamic code generation.
python examples/01_basic_sandboxing.py
python examples/02_security_gatekeeper.py
python examples/03_llm_healing_workflow.py
python examples/04_agent_tool_dropin.py

Running Tests & Lint

The suite currently has 50 tests with >82% code coverage on src/codeshield.

ruff check src tests examples
pytest tests -v --cov=src/codeshield

Enterprise Architecture & Custom Deployments

This repository ships the core execution and healing engine. For production multi-tenant deployments, the enterprise extension adds:

  • Multi-tenant orchestrator with queue-based job scheduling.
  • PostgreSQL state persistence for execution history, audit trails and replay.
  • Automated billing and token governance (cost caps per tenant, per-execution budgets).
  • Prometheus/Grafana observability, RBAC, and signed artifact provenance.
  • SLA-backed support and custom agentic architecture consulting.

Want the production-grade version or a tailored integration for your platform?

We offer enterprise licensing, dedicated onboarding and custom agentic-architecture consulting.


License

This project is licensed under the MIT License. See LICENSE for details.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

codeshield_runtime-0.1.1.tar.gz (25.4 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

codeshield_runtime-0.1.1-py3-none-any.whl (24.5 kB view details)

Uploaded Python 3

File details

Details for the file codeshield_runtime-0.1.1.tar.gz.

File metadata

  • Download URL: codeshield_runtime-0.1.1.tar.gz
  • Upload date:
  • Size: 25.4 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.4

File hashes

Hashes for codeshield_runtime-0.1.1.tar.gz
Algorithm Hash digest
SHA256 bba35734e1d5efdaa3a0cce6c296a64956f7d8dc654d8efe3438f99e6ab98577
MD5 bb2613976cfc655dbab9b7c56614e394
BLAKE2b-256 5b8242ed2f6d37be5d2beda9f407c0a9df5e2ee71ba3c1df0ba006fedf8aa29f

See more details on using hashes here.

File details

Details for the file codeshield_runtime-0.1.1-py3-none-any.whl.

File metadata

File hashes

Hashes for codeshield_runtime-0.1.1-py3-none-any.whl
Algorithm Hash digest
SHA256 78b9126039e83e55a9ec2ed08dbe46205345d2c9947a9c4b38a164cc45fd24d1
MD5 6a9c6de8ed9d13df3a3ec014adea9a8e
BLAKE2b-256 e5909343b240f41333a2207c2f4b7d1c5399e822c73a62b352bcd5eb14cee100

See more details on using hashes here.

Release history Release notifications | RSS feed

0.1.3

2 files

0.1.2

2 files

This release

0.1.1 This release

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page