Skip to main content

🧠 Agent Sandbox Runtime

The Self-Correcting AI Agent with Swarm Intelligence

An open-source, production-grade AI agent platform that writes code, executes it safely, learns from failures, and self-corrects until it works.

MIT License CI Python 3.11+ Benchmark PRs Welcome Docker LangGraph


🎬 See it in action

Swarm Intelligence Activating Parallel Code Generation
Swarm Init Code Gen
Generated Solution Mission Accomplished 🏆
Solution Result

📺 Video Demo

Watch Demo


📖 Documentation · 🚀 Quick Start · 🏗️ Architecture · 🤝 Contributing


⚡ One-Click Deploy

Deploy on Railway Deploy to Render


🎯 Why This Exists

Most AI coding assistants generate code and hope it works. Agent Sandbox Runtime takes a fundamentally different approach:

You describe what you want → Agent writes code → Executes in Docker sandbox → 
If it fails → Analyzes the error → Rewrites with improvements → Repeats until success

This is Reflexion - the same self-improvement loop that makes humans good at coding. Combined with Swarm Intelligence (5 specialist AI agents reviewing each solution), you get code that actually works.

Real-world problems this solves:

  • 🔄 "The AI gave me broken code" — Self-correction fixes bugs automatically
  • 🔒 "I can't run untrusted code" — Docker isolation makes it safe
  • 🐌 "AI suggestions are slow" — Groq inference at 743ms average
  • 💸 "AI APIs are expensive" — Free tier models supported (Ollama, OpenRouter)

🏗️ System Architecture

The Reflexion Loop

This is the core innovation. Instead of generating code once, we generate → test → improve:

                    ┌─────────────────────────────────────────────────┐
                    │           REFLEXION LOOP (LangGraph)            │
                    │                                                 │
     Your Task ───► │  ┌──────────┐    ┌─────────┐    ┌─────────┐   │
                    │  │ GENERATE │───►│ EXECUTE │───►│ SUCCESS │───┼──► Result
                    │  │  (LLM)   │    │(Docker) │    │    ?    │   │
                    │  └──────────┘    └─────────┘    └────┬────┘   │
                    │       ▲                              │        │
                    │       │         ┌───────────┐        │ No     │
                    │       │         │  CRITIQUE │◄───────┘        │
                    │       │         │  (LLM)    │                 │
                    │       │         └─────┬─────┘                 │
                    │       │               │                       │
                    │       │         ┌─────▼─────┐                 │
                    │       └─────────┤   RETRY   │                 │
                    │                 │ (≤3 times)│                 │
                    │                 └───────────┘                 │
                    └─────────────────────────────────────────────────┘

Component Overview

Component Purpose Technology
Orchestrator Manages the reflexion loop state machine LangGraph
Generator Produces Python code from natural language LLM (6 providers)
Sandbox Executes code in isolated Docker containers Docker SDK
Critic Analyzes failures and suggests improvements LLM
Swarm Multi-agent code review (Architect, Coder, Critic, Optimizer, Security) Async LLM calls

Data Flow (Peer-to-Peer)

┌─────────────┐     ┌─────────────┐     ┌─────────────┐
│   CLI/API   │────►│   Runtime   │────►│ Orchestrator│
│   (Input)   │     │  (Entry)    │     │ (LangGraph) │
└─────────────┘     └─────────────┘     └──────┬──────┘
                                               │
                    ┌──────────────────────────┼──────────────────────────┐
                    │                          ▼                          │
                    │  ┌─────────────┐   ┌─────────────┐   ┌───────────┐ │
                    │  │  Generator  │◄─►│   Critic    │◄─►│  Sandbox  │ │
                    │  │   (LLM)     │   │   (LLM)     │   │  (Docker) │ │
                    │  └──────┬──────┘   └─────────────┘   └───────────┘ │
                    │         │                                          │
                    │         ▼                                          │
                    │  ┌─────────────────────────────────────┐          │
                    │  │         SWARM INTELLIGENCE          │          │
                    │  │  ┌────────┐ ┌──────┐ ┌───────────┐  │          │
                    │  │  │Architect│ │Critic│ │ Security  │  │          │
                    │  │  └────────┘ └──────┘ └───────────┘  │          │
                    │  │  ┌────────┐ ┌──────────┐            │          │
                    │  │  │ Coder  │ │Optimizer │            │          │
                    │  │  └────────┘ └──────────┘            │          │
                    │  └─────────────────────────────────────┘          │
                    │                    NODE POOL                       │
                    └────────────────────────────────────────────────────┘

✨ Features

Feature Description
🔄 Self-Correction Loop Automatically detects and fixes bugs through iterative refinement
🐝 Swarm Intelligence 5 specialist agents (Architect, Coder, Critic, Optimizer, Security) collaborate
🔒 Docker Sandbox Code runs in isolated containers with memory/CPU limits, no network by default
🔌 6 LLM Providers Groq, OpenRouter, Anthropic, Google Gemini, OpenAI, Ollama (local)
Fast Inference Groq's LPU delivers ~743ms average response time
📊 Structured Output Pydantic-validated JSON responses from LLMs
🌐 API & CLI FastAPI server + command-line interface

🏆 Benchmark Results

Metric Value
Total Tests 12
Passed 11/12
Success Rate 92%
Rating 🔥 GOD TIER
Avg Response 743ms

Charts

Success by Difficulty Response Time
Success Time

vs Competitors

Tool Success Speed Self-Correct Sandbox Cost
Agent Sandbox 92% 743ms Free
GPT-4 Code Interpreter 87% 3.2s $0.03/1K
Claude 3.5 Sonnet 89% 2.1s $0.015/1K
Devin 85% 45s $500/mo
Cursor 78% 2.8s $20/mo

🚀 Quick Start

Option 1: One-Click Deploy

Click the Railway or Render button above ☝️

Option 2: Docker

docker run -e GROQ_API_KEY=your_key ghcr.io/ixchio/agent-sandbox-runtime

Option 3: Local Installation

# Clone the repository
git clone https://github.com/ixchio/agent-sandbox-runtime.git
cd agent-sandbox-runtime

# Create virtual environment
python -m venv .venv
source .venv/bin/activate  # On Windows: .venv\Scripts\activate

# Install dependencies
pip install -e .

# Configure environment
cp .env.example .env
# Edit .env and add your GROQ_API_KEY (get free key at https://console.groq.com)

# Run your first task
agent-sandbox run "Calculate fibonacci(10)"

Option 4: API Server

# Start the API server
agent-sandbox serve

# POST a request
curl -X POST http://localhost:8000/execute \
  -H "Content-Type: application/json" \
  -d '{"task": "Write a function to check if a number is prime"}'

⚙️ Configuration

Environment Variables

Variable Required Default Description
LLM_PROVIDER No groq Provider: groq, openrouter, anthropic, google, ollama, openai
GROQ_API_KEY Yes* - Get free key
OPENROUTER_API_KEY Yes* - Get key
ANTHROPIC_API_KEY Yes* - Get key
GOOGLE_API_KEY Yes* - Get key
OPENAI_API_KEY Yes* - Get key
SANDBOX_TIMEOUT_SECONDS No 5.0 Max execution time per run
SANDBOX_MEMORY_LIMIT_MB No 256 Container memory limit
MAX_REFLEXION_ATTEMPTS No 3 Max retry attempts
API_PORT No 8000 Server port

*Only one provider API key is required

Recommended Models by Provider

Provider Model Best For
Groq llama-3.3-70b-versatile Speed + Quality
OpenRouter qwen/qwen-2.5-coder-32b-instruct:free Free tier
Anthropic claude-3-5-sonnet-20241022 Complex reasoning
Google gemini-1.5-flash Fast + cheap
Ollama qwen2.5-coder:7b Local/private
OpenAI gpt-4o-mini Balanced

📂 Project Structure

agent-sandbox-runtime/
├── src/agent_sandbox/
│   ├── api/              # FastAPI endpoints
│   ├── cli.py            # Command-line interface
│   ├── config.py         # Settings & environment
│   ├── orchestrator/     # LangGraph workflow
│   │   ├── graph.py      # Main state machine
│   │   ├── nodes/        # Generate, Execute, Critique, Retry
│   │   └── state.py      # Workflow state model
│   ├── providers/        # LLM provider adapters
│   ├── sandbox/          # Docker execution engine
│   │   ├── manager.py    # Container lifecycle
│   │   ├── executor.py   # Code execution
│   │   └── models.py     # Request/Response types
│   ├── swarm/            # Multi-agent intelligence
│   └── runtime.py        # Main entry point
├── docs/                 # Documentation
├── tests/                # Test suite
├── Dockerfile            # Container build
├── docker-compose.yml    # Local development stack
└── pyproject.toml        # Dependencies & config

📚 Documentation

Document Description
Architecture System design & component breakdown
How It Works Deep dive into the reflexion loop
Capabilities What problems this solves
API Reference Endpoint documentation
Contributing How to contribute

🤝 Contributing

We welcome contributions! See CONTRIBUTING.md for:

  • 🔧 Development setup
  • 📝 Code style guidelines
  • 🧪 Testing requirements
  • 📬 Pull request process
  • 💡 Feature request guidelines

Quick Contribution Steps

# Fork & clone
git clone https://github.com/YOUR_USERNAME/agent-sandbox-runtime.git

# Create branch
git checkout -b feature/your-feature

# Install dev dependencies
pip install -e ".[dev]"

# Make changes, run tests
pytest tests/unit/ -v

# Submit PR

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.


Built with 💜 by the open-source community

⭐ Star us on GitHub · 🐛 Report Bug · 💡 Request Feature

Release files for agent-sandbox-runtime 1.0.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for agent-sandbox-runtime 1.0.0
File Size Uploaded
agent_sandbox_runtime-1.0.0.tar.gz 831.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for agent-sandbox-runtime 1.0.0
File Interpreter ABI Platform
agent_sandbox_runtime-1.0.0-py3-none-any.whl Python 3 none any Details

Total release size: 902.9 kB

Release files / agent_sandbox_runtime-1.0.0.tar.gz

Download URL agent_sandbox_runtime-1.0.0.tar.gz
Size 831.0 kB
Tags Source
SHA-256 checksum
How to use checksums
c4b17940393762badd6fd709dc818ed33fe3944c409fe0ce4e9cf2262037fc2d
BLAKE2b-256 checksum
How to use checksums
6b51a68080dccab7c7b982ab10d2c5632ffabee6fb90669c79c44923a43bca17
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Dec 18, 2025.

Transparency log

Release files / agent_sandbox_runtime-1.0.0-py3-none-any.whl

Download URL agent_sandbox_runtime-1.0.0-py3-none-any.whl
Size 71.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4c7c9ead84f62396321b2e4ea937411c8a3211a39ea437e171b553b60975dd41
BLAKE2b-256 checksum
How to use checksums
024f90ddaeb5b4345ea1c1c4acea417e35e131d0e2f8f045a95938124bd3abf6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/6.1.0 CPython/3.13.7

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Dec 18, 2025.

Transparency log

Release history Release notifications | RSS feed

This release

1.0.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page