Skip to main content

Box Agent

A general-purpose AI agent with sandboxed code execution, sub-agent parallelism, and multi-provider LLM support.

PyPI Downloads Python License Release

English | 中文


Get started in 30 seconds:

uv tool install box-agent   # or: pip install box-agent (Python 3.10+)
box-agent setup              # interactive config wizard
box-agent                    # start chatting

Or run a one-shot task:

box-agent --task "Analyze sales.csv — show top 10 products by revenue with a bar chart"

Why Box Agent?

Most agent frameworks are either too simple (no sandbox, no tools) or too complex (massive dependencies, rigid architecture). Box Agent hits the sweet spot:

Feature Box Agent Open Interpreter Aider
Sandboxed code execution Jupyter kernel in isolated venv Runs in host Python N/A
Sub-agent parallelism Multiple sub-agents run concurrently No No
Multi-provider LLM Anthropic, OpenAI, DeepSeek, SiliconFlow, any API OpenAI + a few others OpenAI + Anthropic
MCP tool integration Native No No
ACP protocol (embed in apps) Full support No No
Standalone binary PyInstaller runtime, no Python needed No No
Context compression Staged automatic compaction + LLM summary Manual Git-based

Key Features

Sub-Agent Parallelism

Delegate isolated work through a flat task contract with optional tools, Skills, files, write scope, and hard step/tool-call budgets. Omitted tools resolve only to trusted local readers; explicit capabilities still pass a fail-closed runtime policy. Passing known local text paths in files selects the bounded, completeness-checked batch fast path automatically when read_file is the only resolved tool; requesting additional tools keeps the general child loop. The parent remains responsible for conflict handling, the final deliverable, and verification.

You: "Analyze data1.csv, data2.csv, and data3.csv separately, then give me a combined summary"

┌─ Sub-Agent 1 ──────┐  ┌─ Sub-Agent 2 ──────┐  ┌─ Sub-Agent 3 ──────┐
│ Read data1.csv      │  │ Read data2.csv      │  │ Read data3.csv      │
│ Run statistics      │  │ Run statistics      │  │ Run statistics      │
│ Generate charts     │  │ Generate charts     │  │ Generate charts     │
│ → Summary: ...      │  │ → Summary: ...      │  │ → Summary: ...      │
└─────────────────────┘  └─────────────────────┘  └─────────────────────┘
                              ↓ parallel ↓
                    ┌─ Parent Agent ──────────┐
                    │ Combines 3 summaries    │
                    │ Produces final report   │
                    └─────────────────────────┘

Child policy is derived by the runtime: process tools, external side effects, and unknown MCP tools fail closed; path writes require an exact scope. See the sub-agent delegation contract for schemas, limits, compatibility behavior, and host diagnostics.

Sandboxed Code Execution

Python runs in an isolated Jupyter kernel with pre-installed data science packages (pandas, numpy, matplotlib, scikit-learn, openpyxl, xlrd). Generated files (charts, CSVs, PDFs) are automatically detected and surfaced as structured artifacts.

Multi-Provider LLM

One config, any provider:

# Anthropic
api_base: "https://api.anthropic.com"
provider: "anthropic"
model: "claude-sonnet-4-20250514"

# DeepSeek
api_base: "https://api.deepseek.com"
provider: "openai"
model: "deepseek-chat"

# Any OpenAI-compatible endpoint
api_base: "https://your-api.example.com/v1"
provider: "openai"
model: "your-model"

Staged Context Compression

  • Oversized tool results: Individual results are persisted immediately when needed; fresh parallel results also share a 50k-character pre-request budget. The model receives a stable preview while full text remains on disk. Read results are exempt and stay bounded by Read's own line/character controls.
  • Usage-aware auto-summary: The next request is estimated from the latest real API usage plus subsequent messages. When it reaches the model-derived safety threshold, older history is summarized into a user message while bounded recent messages and todo, plan, and skill state are restored.
  • Tool-call arguments: Write/edit arguments remain verbatim until a whole-history summary replaces their turn; they are not independently compacted.
  • Legacy safety guard: Internal history placeholders from older or externally supplied sessions are rejected if a model tries to reuse them as executable file/code arguments; Box-Agent requests one clean regeneration instead of writing the placeholder to disk.

More

  • MCP Tools: Connect to any MCP server — web search, knowledge graphs, databases
  • Claude Skills: 32 built-in skills for documents (DOCX, PDF, PPTX, XLSX), canvas design, Obsidian, web app testing, and more
  • ACP Protocol: Embed Box Agent in Electron apps, Zed Editor, or any ACP-compatible host via JSON-RPC over stdio
  • Standalone Runtime: PyInstaller binary bundles Python + all dependencies. No external Python needed — download and run
  • Cross-session Memory: Persistent memory lets the agent retain key information across conversations
  • Safety Layer: Dangerous command detection, workspace scope control, auto-backup before file modifications. Interactive permission negotiation for out-of-workspace access (CLI prompts user, ACP sends reverse RPC to host)
  • Planning Snapshots: Structured plan tool for rendering objective, scope, steps, verification, and risks in host UIs
  • Task Tracking: Built-in todo tool for multi-step task decomposition and progress tracking

Demos

Task Execution

The agent creates a webpage and opens it in the browser.

Demo: Task Execution

Claude Skill — PDF Generation

The agent uses a skill to create a professional document.

Demo: Claude Skill

Web Search via MCP

The agent searches the web and summarizes results.

Demo: Web Search

Installation

Requires Python 3.10+. If your system Python is older (e.g. 3.9), use uv tool install — it manages Python automatically.

Quick Start (uv, recommended)

uv handles Python version management for you — no need to upgrade your system Python:

# Install uv (if not already)
curl -LsSf https://astral.sh/uv/install.sh | sh

# Install box-agent (auto-downloads Python 3.10+ if needed)
uv tool install box-agent
box-agent setup    # interactive config wizard
box-agent          # start chatting

# Upgrade later
uv tool upgrade box-agent

Quick Start (pip)

If you already have Python 3.10+:

pip install box-agent
box-agent setup
box-agent

From Source

git clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
uv sync
uv run python -m box_agent.cli

Contributor Quickstart

If you are joining the project as a collaborator, start here before changing code:

git clone https://github.com/Raccoon-Office/Box-Agent.git
cd Box-Agent
git submodule update --init --recursive   # needed for bundled skills
uv sync
uv run python -m box_agent.cli --help
uv run pytest tests/test_core.py -q

Read these files first:

  • AGENTS.md — repo-local engineering rules and verification expectations.
  • CONTRIBUTING.md — contribution flow, PR checklist, and commit style.
  • docs/REVIEW_GUIDE.md — maintainer review order, blockers, and proof requirements.
  • docs/DEVELOPMENT_GUIDE.md — deeper architecture and development notes.
  • docs/INTEGRATION.md — ACP/runtime integration details for host apps.

Project map:

Area Where to start
Agent execution loop box_agent/core.py, box_agent/agent.py, box_agent/events.py
CLI and config box_agent/cli.py, box_agent/config.py, box_agent/config/
LLM providers box_agent/llm/
Built-in tools box_agent/tools/
ACP server/runtime embedding box_agent/acp/, box_agent/build_runtime_cli.py
Skills box_agent/skills/, box_agent/tools/skill_loader.py
Tests tests/test_<area>.py

Common development loop:

# Run the smallest relevant test while iterating
uv run pytest tests/test_bash_tool.py -q

# Run the broader suite before handing off
uv run pytest tests/ -q

# Catch whitespace/patch formatting issues
git diff --check

Use focused tests for the area you touched: tools in tests/test_*_tool.py, LLM behavior in tests/test_llm*.py / tests/test_error_messages.py, ACP in tests/test_acp*.py, memory in tests/test_memory*.py, and runtime packaging in tests/test_build_runtime.py / tests/test_cli_runtime.py. Tests that need real provider credentials are skipped unless the required API keys are present.

When a change affects the standalone runtime used by a host app, source changes are not enough: rebuild the runtime, install it into the host, restart the running ACP process, then probe the installed runtime. For local packaging:

uv run box-agent-build-runtime

Build a versioned runtime and install the resulting archive into the usual officev3 checkout in one command:

uv run box-agent-build-runtime --version 0.8.82 --install-officev3

Pass an explicit checkout path after --install-officev3, or set BOX_AGENT_OFFICEV3_DIR, when officev3 is stored elsewhere.

Configuration

After running box-agent setup, your config lives at ~/.box-agent/config/config.yaml:

api_key: "your-api-key"
api_base: "https://api.anthropic.com"
model: "claude-sonnet-4-20250514"
provider: "anthropic" # "anthropic" or "openai"
max_steps: 300
max_parallel_tools: 8
parallel_tool_timeout_seconds: 900
provider_stale_seconds: 300
sub_agent_token_limit: 50000
sub_agent_batch_synthesis_timeout_seconds: 600 # 0 disables the extra batch synthesis cap
goal_autopilot_enabled: true
goal_autopilot_max_turns: 3
goal_autopilot_max_seconds: 14400
goal_autopilot_no_progress_turns: 2

Tool limits are omitted by default so runtime upgrades can supply updated defaults from box_agent/config.py. Add only deliberate overrides under tool_limits:; inspect the current effective values with box-agent config --json.

box-agent config                    # show current config summary
box-agent config --get model        # print one config value
box-agent config --set max_steps 300
box-agent config --set goal_autopilot_max_turns 5
box-agent config --set tool_limits.external_skill.max_tool_calls 160
box-agent config --set tool_limits.external_skill.max_delegated_tool_calls 512
box-agent config --set tool_limits.completion.deadline_seconds 1800
box-agent config --set tools.bash_default_timeout_seconds 300
box-agent config --json             # machine-readable config summary
box-agent config --edit             # open in editor
box-agent doctor                    # check environment & API connectivity
box-agent doctor --json             # machine-readable health check

CLI Usage

# Interactive mode
box-agent
box-agent --workspace /path/to/project
box-agent --no-sandbox           # disable Jupyter sandbox

# Non-interactive (CI/CD, scripts)
box-agent --task "analyze data.csv and create a report"
box-agent --task "analyze data.csv" --json          # append execution summary JSON
box-agent --task "local file task" --no-verify-api  # skip startup API probe
box-agent --task "create a PPT" --force-plan-start  # publish a plan before work
box-agent --task "create a PPT" --no-completion-gate
box-agent --goal "Ship CLI parity" --task "finish tests"
box-agent --goal "Ship CLI parity" --task "finish tests" --no-goal-autopilot
box-agent --deep-think --task "review this repo"    # enable thinking mode when supported

# Subcommands
box-agent setup              # config wizard
box-agent config             # show/edit config
box-agent doctor             # health check
box-agent log                # open log directory
box-agent trace-viewer       # open the offline Agent Trace diagnostics page
box-agent goal status        # show persistent workspace goal
box-agent goal complete --evidence "tests passed"
box-agent install-browser   # install Chromium for Playwright MCP (~200MB)
box-agent install-node      # install managed Node.js runtime for skills (macOS)

Agent Trace diagnostics

Run box-agent trace-viewer to open the packaged, offline developer viewer. Open ~/.box-agent/log/sessions/ for a newest-first overview of every trace, then select one run to inspect per-turn metrics, LLM/tool waterfalls, raw events, and the complete system → user → assistant/tool → final-response chain. You can still open one .jsonl file directly.

If an embedded browser does not expose the native file picker, run the loopback-only service and enter the trace directory path in the page:

uv run python -m box_agent.trace_viewer.server --port 8766

The offline page reads files in the browser. Service mode reads only top-level .jsonl files from the directory you enter, checks their metadata once per second, and refreshes the ledger when files are added or changed; trace bodies are transferred over 127.0.0.1 only when that metadata changes. The service rejects requests whose Host or Origin is not its exact loopback authority, preventing a rebinding site from reading local traces. Neither mode makes external network requests. Chromium and Edge can keep following appended records after you grant a file handle; drag/drop and ordinary file inputs load a snapshot. Session traces may contain prompts, tool arguments, outputs, and business data—handle them as sensitive diagnostic artifacts.

Browser automation (optional)

Box-Agent ships with a disabled @playwright/mcp entry. To enable browser tools locally:

box-agent install-browser   # downloads Chromium and flips the entry to enabled

Requires Node.js ≥ 18 on PATH. Chromium lands in ~/.box-agent/browsers/ (shared by CLI and ACP runtime) and mcpServers.playwright.disabled in ~/.box-agent/config/mcp.json is set to false.

ACP embedders: no env-var plumbing required — box-agent-acp defaults PLAYWRIGHT_BROWSERS_PATH to the same ~/.box-agent/browsers/ path. To point at a different cache, export PLAYWRIGHT_BROWSERS_PATH=<your path> before spawning box-agent-acp (our setdefault won't override it).

In-session commands: /help, /clear, /clear_all, /history, /stats, /sandbox_status, /log, /goal, /memory review, /exit

ACP session traces keep their existing ~/.box-agent/log/sessions/<session-id>.jsonl name and box-agent-session-trace/v1 record format. Retention removes only whole, inactive session files: files older than 7 days are eligible, and the directory has a soft 512 MiB cap. The current append target, files modified within 24 hours, and the newest two sessions are protected. Cleanup runs best-effort at most once every 6 hours; cleanup failures never interrupt agent execution. Operators can override the defaults with BOX_AGENT_SESSION_TRACE_RETENTION_DAYS, BOX_AGENT_SESSION_TRACE_MAX_TOTAL_BYTES, and BOX_AGENT_SESSION_TRACE_CLEANUP_INTERVAL_SECONDS, or disable cleanup with BOX_AGENT_SESSION_TRACE_RETENTION_ENABLED=0.

Use /goal <objective> or --goal "<objective>" to keep a durable workspace objective attached to later turns. The CLI persists it under ~/.box-agent/goals/; later turns include that goal until you run /goal pause, /goal resume, /goal block <reason>, /goal complete <evidence>, or /goal clear. Scripted runs can manage it with box-agent goal ....

In non-interactive --task mode and ACP sessions, active goals also use bounded autopilot: when a turn ends naturally but the goal is still active, Box-Agent automatically continues in the same session until the model marks the goal complete, marks it blocked, the user cancels, goal_autopilot_max_turns / goal_autopilot_max_seconds is reached, or goal_autopilot_no_progress_turns consecutive automatic continuations make no recorded goal progress. Use --no-goal-autopilot for one CLI run, or set goal_autopilot_enabled: false in config.

ACP & Editor Integration

Box Agent supports the Agent Communication Protocol for embedding in editors and apps.

Zed Editor — add to settings.json:

{
  "agent_servers": {
    "box-agent": {
      "command": "/path/to/box-agent-acp"
    }
  }
}

Standalone Runtime — for Electron apps and other hosts:

# Download pre-built binary (latest release; omit the tag to always get the newest)
gh release download --repo Raccoon-Office/Box-Agent --pattern "box-agent-runtime-*.tar.gz"

# Or build from source (current platform)
uv run box-agent-build-runtime

# Build macOS Intel/x64 runtime from Apple Silicon
# Requires a separate x86_64 venv because PyInstaller cannot bundle arm64 wheels into an x64 binary.
# One-time setup:
#   arch -x86_64 /bin/bash -c 'curl -LsSf https://astral.sh/uv/install.sh | INSTALLER_NO_MODIFY_PATH=1 UV_INSTALL_DIR="$HOME/.local/bin-x64" sh'
#   UV_PROJECT_ENVIRONMENT=.venv-x64 arch -x86_64 ~/.local/bin-x64/uv sync
# Build:
UV_PROJECT_ENVIRONMENT=.venv-x64 BOX_AGENT_RUNTIME_TARGET=darwin-x64 arch -x86_64 ~/.local/bin-x64/uv run box-agent-build-runtime

The runtime communicates via JSON-RPC over stdio. stdout = protocol only, stderr = diagnostics. macOS runtime archives include Box-Agent's pinned Node.js runtime for skills under box-agent-runtime/runtimes/node/; npm cache/prefix state remains in ~/.box-agent/runtimes/node/sandbox/.

Testing

uv run pytest tests/ -v          # all tests
uv run pytest tests/test_core.py -v   # core + context compression
uv run pytest --cov              # with coverage

Troubleshooting

SSL Certificate Error: pip install --upgrade certifi or set verify=False for testing.

Module Not Found: Make sure you're in the project directory: cd Box-Agent && uv run python -m box_agent.cli

Contributing

Issues and PRs welcome! See Contributing Guide.

License

MIT

Links


If this project helps you, give it a ⭐!

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

box_agent-0.9.7.tar.gz (13.9 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

box_agent-0.9.7-py3-none-any.whl (13.7 MB view details)

Uploaded Python 3

File details

Details for the file box_agent-0.9.7.tar.gz.

File metadata

  • Download URL: box_agent-0.9.7.tar.gz
  • Upload date:
  • Size: 13.9 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.20

File hashes

Hashes for box_agent-0.9.7.tar.gz
Algorithm Hash digest
SHA256 e666b1533f0dee5e08fec07e68a9a2943c6e29d9dbf3891c0e69f15f098b00ee
MD5 3adebf76020a213e65fd75c76a91628a
BLAKE2b-256 2804a6b015723a7c891b6964a8efdf84bbd0b7a6b04c1c8c1a000c8bf5407526

See more details on using hashes here.

File details

Details for the file box_agent-0.9.7-py3-none-any.whl.

File metadata

  • Download URL: box_agent-0.9.7-py3-none-any.whl
  • Upload date:
  • Size: 13.7 MB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.10.20

File hashes

Hashes for box_agent-0.9.7-py3-none-any.whl
Algorithm Hash digest
SHA256 a80a628f492ba1ec7634ab4cd99aaf52c9755c41756cbc7ec1645dd7e81672a4
MD5 d217f49fa70ca72c5000310db0e302f9
BLAKE2b-256 628601e53c98ffabcd40f9d87e1f4a8909c6c03896437f70dcb084dccd9e9d57

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.9.7 This release

2 files

0.9.6

2 files

0.8.79

2 files

0.8.71

2 files

0.8.70

2 files

0.8.68

2 files

0.8.67

2 files

0.8.66

2 files

0.8.65

2 files

0.8.64

2 files

0.8.63

2 files

0.8.62

2 files

0.8.60

2 files

0.8.59

2 files

0.8.58

2 files

0.8.57

2 files

0.8.56

2 files

0.8.55

2 files

0.8.54

2 files

0.8.53

2 files

0.8.52

2 files

0.8.51

2 files

0.8.50

2 files

0.8.47

2 files

0.8.46

2 files

0.8.45

2 files

0.8.44

2 files

0.8.43

2 files

0.8.42

2 files

0.8.41

2 files

0.8.40

2 files

0.8.39

2 files

0.8.37

2 files

0.8.36

2 files

0.8.33

2 files

0.8.30

2 files

0.8.29

2 files

0.8.28

2 files

0.8.27

2 files

0.8.26

2 files

0.8.25

2 files

0.8.24

2 files

0.8.23

2 files

0.8.22

2 files

0.8.21

2 files

0.8.20

2 files

0.8.19

2 files

0.8.18

2 files

0.8.17

2 files

0.8.16

2 files

0.8.15

2 files

0.8.14

2 files

0.8.13

2 files

0.8.12

2 files

0.8.11

2 files

0.8.10

2 files

0.8.9

2 files

0.8.8

2 files

0.8.7

2 files

0.8.6

2 files

0.8.5

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

0.7.9

2 files

0.7.8

2 files

0.7.7

2 files

0.7.6

2 files

0.7.5

2 files

0.7.4

2 files

0.7.3

2 files

0.7.2

2 files

0.7.1

2 files

0.7.0

2 files

0.6.9

2 files

0.6.8

2 files

0.6.7

2 files

0.6.6

2 files

0.6.5

2 files

0.6.4

2 files

0.6.3

2 files

0.6.2

2 files

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.16

2 files

0.4.15

2 files

0.4.14

2 files

0.4.13

2 files

0.4.12

2 files

0.4.11

2 files

0.4.10

2 files

0.4.9

2 files

0.4.8

2 files

0.4.7

2 files

0.4.6

2 files

0.4.4

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.2

2 files

0.3.1

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page