Skip to main content

Box Agent

A general-purpose AI agent with sandboxed code execution, sub-agent parallelism, and multi-provider LLM support.

PyPI Downloads Python License Release

English | 中文


Get started in 30 seconds:

uv tool install box-agent   # or: pip install box-agent (Python 3.10+)
box-agent setup              # interactive config wizard
box-agent                    # start chatting

Or run a one-shot task:

box-agent --task "Analyze sales.csv — show top 10 products by revenue with a bar chart"

Why Box Agent?

Most agent frameworks are either too simple (no sandbox, no tools) or too complex (massive dependencies, rigid architecture). Box Agent hits the sweet spot:

Feature Box Agent Open Interpreter Aider
Sandboxed code execution Jupyter kernel in isolated venv Runs in host Python N/A
Sub-agent parallelism Multiple sub-agents run concurrently No No
Multi-provider LLM Anthropic, OpenAI, DeepSeek, SiliconFlow, any API OpenAI + a few others OpenAI + Anthropic
MCP tool integration Native No No
ACP protocol (embed in apps) Full support No No
Standalone binary PyInstaller runtime, no Python needed No No
Context compression Staged automatic compaction + LLM summary Manual Git-based

Key Features

Sub-Agent Parallelism

Delegate isolated work to sub-agents with explicit tools, Skills, inputs, constraints, and hard step/tool-call budgets. For many known local text files, one batch_files child reads the files concurrently and performs one tool-free synthesis call; heterogeneous work can use bounded general_loop children. The parent remains responsible for conflict handling, the final deliverable, and verification.

You: "Analyze data1.csv, data2.csv, and data3.csv separately, then give me a combined summary"

┌─ Sub-Agent 1 ──────┐  ┌─ Sub-Agent 2 ──────┐  ┌─ Sub-Agent 3 ──────┐
│ Read data1.csv      │  │ Read data2.csv      │  │ Read data3.csv      │
│ Run statistics      │  │ Run statistics      │  │ Run statistics      │
│ Generate charts     │  │ Generate charts     │  │ Generate charts     │
│ → Summary: ...      │  │ → Summary: ...      │  │ → Summary: ...      │
└─────────────────────┘  └─────────────────────┘  └─────────────────────┘
                              ↓ parallel ↓
                    ┌─ Parent Agent ──────────┐
                    │ Combines 3 summaries    │
                    │ Produces final report   │
                    └─────────────────────────┘

New-style delegation is deny-by-default (read_only: true, network: false, external_side_effect: false). See the sub-agent delegation contract for schemas, limits, compatibility behavior, and host diagnostics.

Sandboxed Code Execution

Python runs in an isolated Jupyter kernel with pre-installed data science packages (pandas, numpy, matplotlib, scikit-learn, openpyxl, xlrd). Generated files (charts, CSVs, PDFs) are automatically detected and surfaced as structured artifacts.

Multi-Provider LLM

One config, any provider:

# Anthropic
api_base: "https://api.anthropic.com"
provider: "anthropic"
model: "claude-sonnet-4-20250514"

# DeepSeek
api_base: "https://api.deepseek.com"
provider: "openai"
model: "deepseek-chat"

# Any OpenAI-compatible endpoint
api_base: "https://your-api.example.com/v1"
provider: "openai"
model: "your-model"

Staged Context Compression

  • Layer 0 — Large content: Generated-artifact reads are compacted immediately. Large write/edit arguments remain intact for one subsequent model turn, then become structured placeholders so the model can confirm success without repeatedly paying the full context cost.
  • Layer 1 — Micro-compact: Every step, older tool results are replaced with short placeholders. Zero cost, no LLM call.
  • Layer 2 — Auto-summary: When tokens exceed the derived threshold (about 104k tokens for user-configured endpoints by default), an LLM call summarizes the conversation. Original data is preserved in logs.
  • Self-healing guard: Internal history placeholders are rejected if a later model turn tries to reuse them as executable file/code arguments; Box-Agent requests one clean regeneration instead of writing the placeholder to disk.

More

  • MCP Tools: Connect to any MCP server — web search, knowledge graphs, databases
  • Claude Skills: 32 built-in skills for documents (DOCX, PDF, PPTX, XLSX), canvas design, Obsidian, web app testing, and more
  • ACP Protocol: Embed Box Agent in Electron apps, Zed Editor, or any ACP-compatible host via JSON-RPC over stdio
  • Standalone Runtime: PyInstaller binary bundles Python + all dependencies. No external Python needed — download and run
  • Cross-session Memory: Persistent memory lets the agent retain key information across conversations
  • Safety Layer: Dangerous command detection, workspace scope control, auto-backup before file modifications. Interactive permission negotiation for out-of-workspace access (CLI prompts user, ACP sends reverse RPC to host)
  • Planning Snapshots: Structured plan tool for rendering objective, scope, steps, verification, and risks in host UIs
  • Task Tracking: Built-in todo tool for multi-step task decomposition and progress tracking

Demos

Task Execution

The agent creates a webpage and opens it in the browser.

Demo: Task Execution

Claude Skill — PDF Generation

The agent uses a skill to create a professional document.

Demo: Claude Skill

Web Search via MCP

The agent searches the web and summarizes results.

Demo: Web Search

Installation

Requires Python 3.10+. If your system Python is older (e.g. 3.9), use uv tool install — it manages Python automatically.

Quick Start (uv, recommended)

uv handles Python version management for you — no need to upgrade your system Python:

# Install uv (if not already)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install box-agent (auto-downloads Python 3.10+ if needed)
uv tool install box-agent
box-agent setup    # interactive config wizard
box-agent          # start chatting

# Upgrade later
uv tool upgrade box-agent

Quick Start (pip)

If you already have Python 3.10+:

pip install box-agent
box-agent setup
box-agent

From Source

git clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
uv sync
uv run python -m box_agent.cli

Contributor Quickstart

If you are joining the project as a collaborator, start here before changing code:

git clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
git submodule update --init --recursive   # needed for bundled skills
uv sync
uv run python -m box_agent.cli --help
uv run pytest tests/test_core.py -q

Read these files first:

  • AGENTS.md — repo-local engineering rules and verification expectations.
  • CONTRIBUTING.md — contribution flow, PR checklist, and commit style.
  • docs/REVIEW_GUIDE.md — maintainer review order, blockers, and proof requirements.
  • docs/DEVELOPMENT_GUIDE.md — deeper architecture and development notes.
  • docs/INTEGRATION.md — ACP/runtime integration details for host apps.

Project map:

Area Where to start
Agent execution loop box_agent/core.py, box_agent/agent.py, box_agent/events.py
CLI and config box_agent/cli.py, box_agent/config.py, box_agent/config/
LLM providers box_agent/llm/
Built-in tools box_agent/tools/
ACP server/runtime embedding box_agent/acp/, box_agent/build_runtime_cli.py
Skills box_agent/skills/, box_agent/tools/skill_loader.py
Tests tests/test_<area>.py

Common development loop:

# Run the smallest relevant test while iterating
uv run pytest tests/test_bash_tool.py -q

# Run the broader suite before handing off
uv run pytest tests/ -q

# Catch whitespace/patch formatting issues
git diff --check

Use focused tests for the area you touched: tools in tests/test_*_tool.py, LLM behavior in tests/test_llm*.py / tests/test_error_messages.py, ACP in tests/test_acp*.py, memory in tests/test_memory*.py, and runtime packaging in tests/test_build_runtime.py / tests/test_cli_runtime.py. Tests that need real provider credentials are skipped unless the required API keys are present.

When a change affects the standalone runtime used by a host app, source changes are not enough: rebuild the runtime, install it into the host, restart the running ACP process, then probe the installed runtime. For local packaging:

uv run box-agent-build-runtime

Configuration

After running box-agent setup, your config lives at ~/.box-agent/config/config.yaml:

api_key: "your-api-key"
api_base: "https://api.anthropic.com"
model: "claude-sonnet-4-20250514"
provider: "anthropic" # "anthropic" or "openai"
max_steps: 200
max_parallel_tools: 8
parallel_tool_timeout_seconds: 900
sub_agent_batch_synthesis_timeout_seconds: 300 # 0 disables the extra batch synthesis cap
goal_autopilot_enabled: true
goal_autopilot_max_turns: 3
goal_autopilot_max_seconds: 14400
goal_autopilot_no_progress_turns: 2
box-agent config                    # show current config summary
box-agent config --get model        # print one config value
box-agent config --set max_steps 300
box-agent config --set goal_autopilot_max_turns 5
box-agent config --json             # machine-readable config summary
box-agent config --edit             # open in editor
box-agent doctor                    # check environment & API connectivity
box-agent doctor --json             # machine-readable health check

CLI Usage

# Interactive mode
box-agent
box-agent --workspace /path/to/project
box-agent --no-sandbox           # disable Jupyter sandbox

# Non-interactive (CI/CD, scripts)
box-agent --task "analyze data.csv and create a report"
box-agent --task "analyze data.csv" --json          # append execution summary JSON
box-agent --task "local file task" --no-verify-api  # skip startup API probe
box-agent --task "create a PPT" --force-plan-start  # publish a plan before work
box-agent --task "create a PPT" --no-completion-gate
box-agent --goal "Ship CLI parity" --task "finish tests"
box-agent --goal "Ship CLI parity" --task "finish tests" --no-goal-autopilot
box-agent --deep-think --task "review this repo"    # enable thinking mode when supported

# Subcommands
box-agent setup              # config wizard
box-agent config             # show/edit config
box-agent doctor             # health check
box-agent log                # open log directory
box-agent goal status        # show persistent workspace goal
box-agent goal complete --evidence "tests passed"
box-agent install-browser   # install Chromium for Playwright MCP (~200MB)
box-agent install-node      # install managed Node.js runtime for skills (macOS)

Browser automation (optional)

Box-Agent ships with a disabled @playwright/mcp entry. To enable browser tools locally:

box-agent install-browser   # downloads Chromium and flips the entry to enabled

Requires Node.js ≥ 18 on PATH. Chromium lands in ~/.box-agent/browsers/ (shared by CLI and ACP runtime) and mcpServers.playwright.disabled in ~/.box-agent/config/mcp.json is set to false.

ACP embedders: no env-var plumbing required — box-agent-acp defaults PLAYWRIGHT_BROWSERS_PATH to the same ~/.box-agent/browsers/ path. To point at a different cache, export PLAYWRIGHT_BROWSERS_PATH=<your path> before spawning box-agent-acp (our setdefault won't override it).

In-session commands: /help, /clear, /clear_all, /history, /stats, /sandbox_status, /log, /goal, /memory review, /exit

Use /goal <objective> or --goal "<objective>" to keep a durable workspace objective attached to later turns. The CLI persists it under ~/.box-agent/goals/; later turns include that goal until you run /goal pause, /goal resume, /goal block <reason>, /goal complete <evidence>, or /goal clear. Scripted runs can manage it with box-agent goal ....

In non-interactive --task mode and ACP sessions, active goals also use bounded autopilot: when a turn ends naturally but the goal is still active, Box-Agent automatically continues in the same session until the model marks the goal complete, marks it blocked, the user cancels, goal_autopilot_max_turns / goal_autopilot_max_seconds is reached, or goal_autopilot_no_progress_turns consecutive automatic continuations make no recorded goal progress. Use --no-goal-autopilot for one CLI run, or set goal_autopilot_enabled: false in config.

ACP & Editor Integration

Box Agent supports the Agent Communication Protocol for embedding in editors and apps.

Zed Editor — add to settings.json:

{
  "agent_servers": {
    "box-agent": {
      "command": "/path/to/box-agent-acp"
    }
  }
}

Standalone Runtime — for Electron apps and other hosts:

# Download pre-built binary (latest release; omit the tag to always get the newest)
gh release download --repo Raccoon-Office/Box-Agent --pattern "box-agent-runtime-*.tar.gz"

# Or build from source (current platform)
uv run box-agent-build-runtime

# Build macOS Intel/x64 runtime from Apple Silicon
# Requires a separate x86_64 venv because PyInstaller cannot bundle arm64 wheels into an x64 binary.
# One-time setup:
#   arch -x86_64 /bin/bash -c 'curl -LsSf https://astral.sh/uv/install.sh | INSTALLER_NO_MODIFY_PATH=1 UV_INSTALL_DIR="$HOME/.local/bin-x64" sh'
#   UV_PROJECT_ENVIRONMENT=.venv-x64 arch -x86_64 ~/.local/bin-x64/uv sync
# Build:
UV_PROJECT_ENVIRONMENT=.venv-x64 BOX_AGENT_RUNTIME_TARGET=darwin-x64 arch -x86_64 ~/.local/bin-x64/uv run box-agent-build-runtime

The runtime communicates via JSON-RPC over stdio. stdout = protocol only, stderr = diagnostics. macOS runtime archives include Box-Agent's pinned Node.js runtime for skills under box-agent-runtime/runtimes/node/; npm cache/prefix state remains in ~/.box-agent/runtimes/node/sandbox/.

Testing

uv run pytest tests/ -v          # all tests
uv run pytest tests/test_core.py -v   # core + context compression
uv run pytest --cov              # with coverage

Troubleshooting

SSL Certificate Error: pip install --upgrade certifi or set verify=False for testing.

Module Not Found: Make sure you're in the project directory: cd Box-Agent && uv run python -m box_agent.cli

Contributing

Issues and PRs welcome! See Contributing Guide.

License

MIT

Links


If this project helps you, give it a ⭐!

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

box_agent-0.8.79.tar.gz (11.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

box_agent-0.8.79-py3-none-any.whl (11.9 MB view details)

Uploaded Python 3

File details

Details for the file box_agent-0.8.79.tar.gz.

File metadata

  • Download URL: box_agent-0.8.79.tar.gz
  • Upload date:
  • Size: 11.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for box_agent-0.8.79.tar.gz
Algorithm Hash digest
SHA256 ee0ca09273ece392395fd3fd4227688b214ec32ec3b79af1fc7894b9b48389dc
MD5 7e1799a25a31b15b21220d41d4f8bde8
BLAKE2b-256 0262eda79738d5970dbb64fdfae9c1c49a980429563296d5373cd42775f6a879

See more details on using hashes here.

File details

Details for the file box_agent-0.8.79-py3-none-any.whl.

File metadata

  • Download URL: box_agent-0.8.79-py3-none-any.whl
  • Upload date:
  • Size: 11.9 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.13.12

File hashes

Hashes for box_agent-0.8.79-py3-none-any.whl
Algorithm Hash digest
SHA256 51b4b5313866cac5564491698254bda5ccf9702626e85c5ec67559045e162e31
MD5 e092c2fa054fb8b028cf214983b8b0c3
BLAKE2b-256 05edfdd38be3b60e62e409121d3e19ce3f1fd3ede037168630ef49f083a5b81d

See more details on using hashes here.

Release history Release notifications | RSS feed

0.9.7

2 files

0.9.6

2 files

This release

0.8.79 This release

2 files

0.8.71

2 files

0.8.70

2 files

0.8.68

2 files

0.8.67

2 files

0.8.66

2 files

0.8.65

2 files

0.8.64

2 files

0.8.63

2 files

0.8.62

2 files

0.8.60

2 files

0.8.59

2 files

0.8.58

2 files

0.8.57

2 files

0.8.56

2 files

0.8.55

2 files

0.8.54

2 files

0.8.53

2 files

0.8.52

2 files

0.8.51

2 files

0.8.50

2 files

0.8.47

2 files

0.8.46

2 files

0.8.45

2 files

0.8.44

2 files

0.8.43

2 files

0.8.42

2 files

0.8.41

2 files

0.8.40

2 files

0.8.39

2 files

0.8.37

2 files

0.8.36

2 files

0.8.33

2 files

0.8.30

2 files

0.8.29

2 files

0.8.28

2 files

0.8.27

2 files

0.8.26

2 files

0.8.25

2 files

0.8.24

2 files

0.8.23

2 files

0.8.22

2 files

0.8.21

2 files

0.8.20

2 files

0.8.19

2 files

0.8.18

2 files

0.8.17

2 files

0.8.16

2 files

0.8.15

2 files

0.8.14

2 files

0.8.13

2 files

0.8.12

2 files

0.8.11

2 files

0.8.10

2 files

0.8.9

2 files

0.8.8

2 files

0.8.7

2 files

0.8.6

2 files

0.8.5

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.9

2 files

0.7.8

2 files

0.7.7

2 files

0.7.6

2 files

0.7.5

2 files

0.7.4

2 files

0.7.3

2 files

0.7.2

2 files

0.7.1

2 files

0.7.0

2 files

0.6.9

2 files

0.6.8

2 files

0.6.7

2 files

0.6.6

2 files

0.6.5

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.16

2 files

0.4.15

2 files

0.4.14

2 files

0.4.13

2 files

0.4.12

2 files

0.4.11

2 files

0.4.10

2 files

0.4.9

2 files

0.4.8

2 files

0.4.7

2 files

0.4.6

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page