Sandbox Environment Manager MCP server
Project description
Sandbox Environment Manager MCP
An MCP (Model Context Protocol) server that provides persistent sandbox environment management for AI agents. Manages Docker containers and SSH machines as execution targets, with shell-based command execution and full file operation capabilities.
Designed as a replacement for Hermes Agent's built-in terminal/file/code_execution tools, adding persistent environment management that the built-in tools lack.
Features
- Compact MCP surface: 7 exposed tools, with progressive management discovery via
sandbox_env - Dual transport: stdio (Hermes child process) or HTTP (independent service)
- Multi-backend: Docker containers (SDK, works with remote daemons) + SSH remote machines
- Persistent machines: Docker containers survive MCP restart; discover with
docker_ps - Shell-based execution: dual-marker confirmation, read for long-running commands
- Full file operations: read, write (atomic), patch (fuzzy match), search (ripgrep/glob)
- In-process linters: Python
ast, JSON, optional YAML/TOML pre-write validation - Safety advisories: non-blocking warnings for sensitive paths (
.ssh,.aws,.env*) - Audit logging: JSON-line stream of all tool invocations (content hashed)
Quick Start
Install
pip install .
pip install -e ".[dev]" # + test/lint tools
# Run unit tests (integration tests are skipped by default)
pytest tests/ -v
# Run integration tests (requires a running Docker daemon)
pytest tests/ -m integration -v
Run
sandbox-mcp has two transports:
sandbox-mcp-http— standalone HTTP service. Start it from a shell:sandbox-mcp-http # Then connect any MCP client to http://127.0.0.1:8010/mcpsandbox-mcp(stdio) — launched by an MCP host as a child process. Don't run this from a shell directly; configure it in the host (see Register with Hermes below).
CLI flags
| flag | applies to | purpose |
|---|---|---|
--config PATH / -c PATH |
both | path to a TOML config file |
--host ADDR / -H ADDR |
sandbox-mcp-http |
HTTP bind address |
--port N / -p N |
sandbox-mcp-http |
HTTP port |
# standalone HTTP server (default: streamable-http on /mcp)
sandbox-mcp-http -c /etc/sandbox-mcp/prod.toml --port 9000
# stdio (passed via the MCP host's config; not run from a shell)
# see Register with Hermes below for an example
Precedence (highest first): CLI flag → env var → config file → built-in default.
Configuration
sandbox-mcp reads config in this priority order (highest wins):
- CLI flags (see above)
- Environment variables —
SANDBOX_MCP_*(e.g.SANDBOX_MCP_SERVER_PORT) - Config file —
~/.sandbox-mcp/config.tomlby default; overridden by--config PATH/SANDBOX_MCP_CONFIG - Built-in defaults (declared in
src/sandbox_mcp/config.py)
To customize, copy config/config.example.toml from the
repo root to ~/.sandbox-mcp/config.toml and edit what you need.
Leaving it in place means all defaults are used.
Config sections:
[server] # HTTP server
host = "0.0.0.0"
port = 8010
transport = "streamable-http"
[storage] # persistent workspace directory (MUST be an absolute HOST path)
work_home = "/var/lib/sandbox-mcp" # no '~' — passed verbatim to the Docker daemon
[audit] # SQLite audit log (one row per tool call)
log_path = "~/.sandbox-mcp/audit.db"
# "" = stderr (sandbox_audit_query hidden); file = query tool enabled
[docker] # container defaults
default_image = "debian:stable-slim"
restart_policy_name = "on-failure"
restart_max_retry_count = 3
[ssh]
connect_timeout = 10
socket_dir_prefix = "sandbox-mcp-ssh-"
tmpfile_pattern = ".sandbox-mcp-tmp.XXXXXX"
# Default SSH target for [default_machine] backend = "ssh":
default_host = ""
default_user = ""
default_port = 22
default_key = ""
[shell]
default_max_output = 50000
head_size = 5120
tail_size = 46080
[files]
max_file_size = 51200
default_read_limit = 500
max_read_limit = 2000
default_search_limit = 50
[default_machine] # opt-in: provision a default machine at startup
enabled = false # false = lazy (agent creates its first machine)
backend = "docker" # "docker" or "ssh"
name = "admin" # defaults to "admin" -> triggers admin mount layout
purpose = "" # backend params live in [docker] / [ssh], not here
Every value can also be overridden via env var (uppercased, dots → underscores), e.g.:
SANDBOX_MCP_SERVER_PORT=9000 sandbox-mcp-http
SANDBOX_MCP_DOCKER_CONTAINER_NAME_PREFIX="box-" sandbox-mcp
SANDBOX_MCP_AUDIT_LOG_PATH=/var/log/sandbox-mcp/audit.db sandbox-mcp
The work_home directory is created automatically. When docker_run is called,
a subdirectory work_home/<machine-name>/ is created and bind-mounted to
/workspace inside the container - the agent works in /workspace without
ever seeing a host path.
Default machine at startup
By default sandbox-mcp is lazy: the agent creates its first machine on
demand with docker_run / ssh_connect, and there is no default machine
until it does. Set [default_machine] enabled = true to instead provision a
default machine at startup so the agent can call sandbox_shell_exec /
sandbox_file_* immediately:
[default_machine]
enabled = true
backend = "docker" # or "ssh"
name = "dev"
# Docker needs nothing else here -- the image comes from [docker] default_image.
# For SSH, set the target under [ssh] (backend params live in their own
# section, not under [default_machine]):
# [ssh]
# default_host = "10.0.0.5"
# default_user = "ubuntu"
# default_port = 22
# default_key = "/home/ubuntu/.ssh/id_ed25519"
Behaviour:
- The
docker_psreconciliation pass runs first, so on a restart the surviving default container is re-adopted (not re-created). - Provisioning failure is fatal (fail-closed): the operator opted in, so a missing default machine would surprise the agent at first use. The server refuses to start with a clear error instead.
namedefaults to"admin". Whenadmin_machineis enabled andname = "admin",DockerBackend.createdetects the match and applies the admin mount layout automatically. Pick any other name (e.g.name = "dev") to opt out of admin and provision a normal peer as the default.docker_runis now idempotent within a session: if a container with the requested name already exists, it reattaches (starting it if stopped) rather than erroring on the name conflict. The response carries anote("reattached to existing container ...") so the agent knows it was a reuse.docker_runno longer overrides the image's CMD. The container runs with whatever CMD/ENTRYPOINT the image author chose —postgres:16actually starts postgres,redis:7actually starts redis, etc. To run a generic image as an exec-only sandbox, build your own image withCMD sleep infinity, or any long-lived shell, anddocker_runagainst that.docker_runaccepts an optionalshellparameter (default"bash") for the binary used bydocker execinto this machine. Set it to e.g./bin/shfor alpine / distroless / busybox images that don't ship bash.start()/ reattach catch fast-crashing containers (CMD exits within milliseconds of start) by reloading state immediately after the start request returns. A container that crashes a moment later may briefly report"running"before transitioning to"exited"— for robust liveness, polldocker_inspectyourself with an appropriate delay.
Register with Hermes
Stdio transport (the sandbox-mcp command):
Add to ~/.hermes/config.yaml:
mcp_servers:
sandbox:
command: sandbox-mcp
# Optional: pass CLI flags to the server.
args:
- --config
- /etc/sandbox-mcp/prod.toml
# Disable built-in tools (optional, to avoid duplicate schemas)
agent:
disabled_toolsets:
- terminal
- file
- code_execution
Hermes spawns sandbox-mcp as a child process and pipes JSON-RPC over
its stdin/stdout. The server has no GUI; it just waits for requests.
HTTP transport (the sandbox-mcp-http command):
mcp_servers:
sandbox:
url: "http://localhost:8010/mcp"
headers:
Authorization: "Bearer <your-token>"
agent:
disabled_toolsets:
- terminal
- file
- code_execution
Hermes connects to the HTTP MCP endpoint (/mcp, the current MCP spec
"Streamable HTTP" transport). Useful when the MCP server runs on a
different machine or is managed as a systemd service.
Tools
| Tool | Purpose |
|---|---|
sandbox_shell_exec |
Execute a shell command (wait or non-blocking) |
sandbox_shell_read |
Read new output from a shell |
sandbox_file_read |
Read a text file with line numbers |
sandbox_file_write |
Write a file (auto mkdir, syntax check, atomic) |
sandbox_file_patch |
Targeted edit with fuzzy match |
sandbox_file_search |
Ripgrep content search + glob file search |
sandbox_env |
Progressive discovery: default_set, shell_*, docker_*, ssh_* |
sandbox_audit_query |
Read the audit log (filtered, paginated) — only when [audit] log_path is set |
sandbox_env Actions
sandbox_env advertises only help and status by default. Call action=help to discover the full action set, or action=docker_help / action=ssh_help for backend-specific actions:
| namespace | actions |
|---|---|
| Discovery | help, status |
| General | machine_list, default_set |
| Shell | shell_new, shell_list, shell_remove |
| Docker | docker_run, docker_build, docker_commit, docker_stop, docker_start, docker_remove, docker_restart, docker_ps, docker_images, docker_image_history, docker_inspect, docker_logs, docker_diff, docker_stats |
| SSH | ssh_connect, ssh_disconnect, ssh_reconnect, ssh_remove |
docker_run is idempotent: if a container with the same name already exists
(e.g. after an MCP restart), it reattaches instead of failing.
Container networking
All containers created by docker_run join a shared user-defined bridge
network (sandbox-mcp by default). This means containers can reach each
other by the same name you passed to docker_run (DNS-resolvable):
sandbox_env(action="docker_run", name="db", image="postgres:16")
sandbox_env(action="docker_run", name="dev", image="debian:stable-slim")
# Inside the "dev" container: psql -h db
# ^ DNS resolves to the "db" container's IP
sandbox_env(action="docker_run", name="web", image="nginx:latest")
# Inside the "dev" container: curl http://web
# ^ DNS resolves to the "web" container's IP
The network name is configured via [docker] auto_network (default
"sandbox-mcp"). Set it to an empty string to opt out:
[docker]
auto_network = ""
The network is created lazily on the first docker_run call, so no
startup dependency exists.
docker_build Usage
The agent never touches the host filesystem. docker_build only
accepts file mode:
sandbox_file_write(path="/workspace/Dockerfile",
content="FROM debian:stable-slim\nRUN apt install -y python3\n")
sandbox_env(action="docker_build",
machine="dev",
image_tag="myapp:v1")
# Defaults: dockerfile=/workspace/Dockerfile, context_dir=/workspace
# sandbox-mcp translates the container path to work_home/<machine>/ on the host
Sandbox boundary: dockerfile and context_dir must live under
/workspace/. Host paths are rejected — the agent cannot reach files
outside its assigned work_home/<machine>/. And only that subtree is
bind-mounted to the host, so a file at e.g. /etc/foo inside the
container has no host-side counterpart for the docker daemon to read;
the build will fail with "context not a directory" even though the
file is visible to the agent's shell_exec.
Why no inline
dockerfile_content? An inline Dockerfile would skip the sandbox's file-write audit trail AND be fed verbatim to the docker daemon, whose build steps execute with full host kernel capabilities (e.g. BuildKit--mount=type=bind,source=/,...). The agent has to commit its Dockerfile to disk viasandbox_file_writefirst, which keeps every line auditable and the build context underwork_home.
Inspecting images and containers
# Container view: state, cmd, entrypoint, mounts, labels, restart policy.
# Env values are deliberately omitted (use shell_exec env/pwd/whoami).
sandbox_env(action="docker_inspect", machine="dev")
# Image view: identity, tags, size, cmd/entrypoint, env KEYS (values redacted),
# exposed ports, declared volumes. Pass any image ref — name:tag, short id, full id.
sandbox_env(action="docker_inspect", machine="python:3.12", kind="image")
sandbox_env(action="docker_inspect", machine="sha256:abc123def456", kind="image")
# Layer-by-layer build history for a single image (mirrors `docker history`).
# Use this when you need ONE image's provenance; use docker_images to enumerate.
sandbox_env(action="docker_image_history", image="python:3.12")
# Returns: {image, layers: [{id (12-hex), created, created_by, size_bytes, tags}], total_size_bytes, layer_count}
docker_inspect (container view), docker_logs, docker_diff, docker_stats,
docker_restart all operate on a managed machine (docker_run-created
container). docker_inspect with kind="image" is the exception: it takes an
image ref directly and never touches the registry. docker_logs and
docker_diff are container-only — images have no log stream and no overlay
filesystem to diff against; for image provenance, use docker_image_history.
docker_run Sandbox Boundary
The agent cannot smuggle arbitrary host paths into a sandboxed container:
- The Docker SDK's raw
volumes=[]is not accepted. Attempts to passvolumes=["/:/host", "/etc:/host-etc"]are silently dropped. - The auto-mounted bind set is fixed: the per-machine workspace
(
work_home/<name>→/workspace, rw) and the inter-container share dir (see below). There is no per-run mount parameter — agents cannot reach arbitrary host paths via sandbox-mcp. - The agent can run any image and
docker execany command inside the container, but cannot mount arbitrary host paths, cannot read host/etc,/root, etc. from inside.
Inter-container share directory
Every docker_run automatically bind-mounts work_home/<share_subdir>/
(default _share/) into the container at /share/. The
mount spec is fixed at two bind mounts, regardless of how many peer
containers exist:
- The whole share root is mounted read-only at
/share/. - The container's own subdirectory
work_home/_share/<machine>/is overlaid read-write at/share/<machine>/— the agent can drop its own output, but the ro parent mount still blocks writes to any peer subdirectory (kernel-enforced mount flag).
Convention:
# inside the "dev" container:
echo "build output" > /share/dev/result.txt # self rw
cat /share/alice/notes.md # peer ro (via parent ro)
ls /share/ # discover peers
New peers appear automatically. Because the parent mount covers
the whole _share/ tree, the kernel evaluates its contents at
access time — a peer subdir created on the host after the container
starts is visible to it on the next ls. No container recreation
needed.
Disable by setting [storage] share_subdir = "" (env:
SANDBOX_MCP_STORAGE_SHARE_SUBDIR).
This is a deliberate first line of defense: the sandbox's
file-write boundary extends into docker_run. Caveats still apply:
the container shares the host kernel, so kernel-capability exploits
(unshare, kernel CVEs) are not stopped by this. For stronger
isolation, deploy with rootless docker or gVisor (runsc).
Admin machine (cross-machine ops)
The admin machine is a special container identified purely by name. Two configs collaborate:
[docker] admin_machine(adminby default; empty disables) — the feature flag + name. When non-empty,DockerBackend.createdetects matching names and applies the god-mode mount layout below. When empty, no name triggers it.[default_machine] enabled = true+name = "admin"(the default) — drives actual container creation via the existing default-machine machinery. Setnameto anything else (e.g."dev") to opt out of admin and provision a normal peer as default.
Default mount layout (peer container):
| Container mount | Source (host) | Mode |
|---|---|---|
/workspace |
work_home/<name>/ |
rw |
/share |
work_home/_share/ |
ro |
/share/<self> |
work_home/_share/<self>/ |
rw |
Admin mount layout (when name == admin_machine):
| Container mount | Source (host) | Mode | Purpose |
|---|---|---|---|
/workspace |
work_home/admin/ |
rw | own scratch |
/host |
work_home/ (whole tree) |
rw | global view: every peer + share |
Share bindings (/share/*) are skipped because the
global /host mount already covers work_home/_share/.
Convention:
# inside the "admin" container:
ls /workspace/ # admin's own scratch
ls /host/ # ALL workspaces + _share + admin/
ls /host/dev/ # peer's workspace (read-only by convention)
rm -rf /host/dev/build # clean up a peer's stale build
cp /workspace/notes.txt /host/alice/ # deliver to a peer
cat /host/_share/bob/log.txt # read peer's share output
Why two mounts: agents should default to /workspace for their own
work. Targeting /host/<peer>/... makes the cross-machine intent
explicit and visible in the agent's command history. Both paths
overlap on work_home/admin/ (same inodes), so writes through either
path land on the same data.
WARNING — god-mode container. /host is rw and covers every peer's
workspace. Operations there are irreversible and can affect
concurrent peer work. Use sparingly — prefer peers cleaning their
own /workspace whenever possible.
Provisioning pattern — to make admin the initial default:
[docker]
admin_machine = "admin" # feature flag (default ON, name = "admin")
[default_machine]
enabled = true # opt in to auto-create
name = "admin" # default — drives the admin container creation
default_machine calls DockerBackend.create("admin", ...), which
detects the name match and applies the god-mode mount layout. No
special-casing in the server startup path.
Upgrading an existing deployment: if a container named admin
already exists as a peer (with the per-machine work_home/admin/
mount), remove it first: docker_remove admin (or
docker rm admin on the host). The server will recreate it with the
admin mount on the next start. No auto-migration — silent failure
otherwise.
Disable by setting [docker] admin_machine = "" (env:
SANDBOX_MCP_DOCKER_ADMIN_MACHINE). With it empty, the name admin
behaves like any peer (own mount + share, no /host).
Connecting to a Remote Docker Daemon
By default sandbox-mcp talks to the docker daemon at
unix:///var/run/docker.sock (or wherever $DOCKER_HOST points). To
point at a remote daemon, set [docker] host in config.toml (env
override: SANDBOX_MCP_DOCKER_HOST):
# Remote daemon over TLS (recommended for non-local daemons).
[docker]
host = "tcp://docker.internal:2376"
tls_verify = true
cert_path = "/etc/sandbox-mcp/docker-certs"
# Or ride your existing SSH trust — uses paramiko, no cert needed.
# host = "ssh://deploy@docker-prod.internal"
# Or a custom socket path when bind-mounted into a container.
# host = "unix:///var/run/docker.sock"
URL scheme (unix:// / tcp:// / ssh://) selects transport. See
config/config.example.toml for all options.
HTTP authentication
The HTTP transport (sandbox-mcp-http) requires a bearer token
on every request. Tokens are stored in a file, one per line:
~/.sandbox-mcp/auth_tokens # default path
The file must be mode 0600 before sandbox-mcp will start.
World/group readable files are rejected (fail-closed):
chmod 600 ~/.sandbox-mcp/auth_tokens
The path can be changed via the config file:
[server]
auth_tokens_file = "/etc/sandbox-mcp/auth_tokens"
Or via an env var (overrides everything):
SANDBOX_MCP_SERVER_AUTH_TOKENS_FILE=/run/secrets/auth_tokens sandbox-mcp-http
When you connect an MCP client, include the token in the
Authorization header:
# default streamable-http transport
curl -X POST -H "Authorization: Bearer <your-token>" \
-H "Content-Type: application/json" \
-d '{"jsonrpc":"2.0","id":1,"method":"ping"}' \
http://127.0.0.1:8010/mcp
Auto-generating a dev token
Set auto_generate_if_empty = true in the config file or export
SANDBOX_MCP_SERVER_AUTO_GENERATE_IF_EMPTY=true. If the token file
is missing or empty, an ephemeral token is generated at startup and
printed to stderr:
[sandbox-mcp-http] WARNING: no tokens found at ~/.sandbox-mcp/auth_tokens.
Generated ephemeral token (capture now, will not be shown again):
XKTUv1Gjv2...33-chars-long
Pass it as: Authorization: Bearer <token>
Capture this token and use it for the session. The server will not regenerate it on restart without the file present.
Limitations
- SSH backend uses key auth only. Password authentication is not supported in the initial release.
- No PTY / interactive stdin. Commands run non-interactively. Commands that expect a TTY (vim, ssh password prompts) are not supported.
- State is in-memory. Shell sessions are lost on server restart; re-create with
shell_new. Containers survive restart and can be reattached viadocker_runor inspected viadocker_ps. - Shell auto-restarts on death. If the agent runs
exit(or bash dies for any other reason), the nextshell_exectransparently runs in a fresh bash. The response includes abash_pidfield — track it across calls; if it changes, all in-shell state (exports, cwd, background jobs) is gone. Persist any state you need across restarts via files, not env vars.- The first response after restart also carries a one-shot
previous_shellsnapshot (previous_bash_pid,last_command,exit_reason,exit_code); act on it then — it won't appear again. (These fields are deliberately not surfaced in the tool description — agents discover them from the response itself, which keeps the per-session tool-list context cost down.)
- The first response after restart also carries a one-shot
- No built-in session isolation. Multiple agents connecting to the same server share the same machine/shell registry. This matches Hermes's own MCP behavior.
Architecture Overview
Agent (LLM)
│
▼
MCP Client (Hermes Gateway | any MCP host)
│ JSON-RPC over stdio │ or │ HTTP (/mcp)
▼ ▼
sandbox-mcp sandbox-mcp-http
│ (stdio transport) │ (streamable-http, port 8010)
│ │
└──────────┬───────────────────┘
│
▼
Application Layer
┌──────────────────────┐
│ 7 MCP tools │
│ sandbox_env dispatcher│
│ ShellSession / ShellReg│
│ MachineRegistry │
│ FileOperations │
│ AuditLogger / Safety │
└──────────┬───────────┘
│
┌───────┴───────┐
▼ ▼
Docker SDK SSH (subprocess)
(put_archive, (ControlMaster,
exec_run, exec_oneoff,
exec socket) stdin pipe)
Design
See docs/design-spec-v2.md for the current design specification. See docs/implementation-plan.md for the TDD implementation plan.
License
This project is dual-licensed:
- Open-source use — GNU Affero General Public License v3.0 (AGPL-3.0-only). You are free to use, modify, and distribute this software under the terms of the AGPLv3, including the requirement that modified versions serving users over a network must also provide their source code.
- Commercial use — if you wish to use this software in a closed-source or proprietary context without the AGPLv3 obligations (notably the source-disclosure obligation of Section 13), a separate commercial license is available. Email 1606272735@qq.com with a description of your intended use; terms (fee, term, seat count) are negotiated per-engagement.
Contributing
By submitting a contribution you agree to the project's contribution terms: your contribution is licensed under AGPLv3, and the maintainer is permitted to relicense it under the commercial terms above. See CONTRIBUTING.md for the development workflow, code style, and sign-off instructions.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sandbox_env_mcp-0.2.1.tar.gz.
File metadata
- Download URL: sandbox_env_mcp-0.2.1.tar.gz
- Upload date:
- Size: 144.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.7 {"installer":{"name":"uv","version":"0.11.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
f544251ba8cec870877e21c0ca6a199daa95901ee47465053288f5e6e5b64330
|
|
| MD5 |
bc58f552ffcdfd9c3cd08a84613e9c69
|
|
| BLAKE2b-256 |
8e06cc92565280527fd031579d74dfe506f7ceff6ee0e055d0afdcae7b9c0b10
|
File details
Details for the file sandbox_env_mcp-0.2.1-py3-none-any.whl.
File metadata
- Download URL: sandbox_env_mcp-0.2.1-py3-none-any.whl
- Upload date:
- Size: 100.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.7 {"installer":{"name":"uv","version":"0.11.7","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Debian GNU/Linux","version":"13","id":"trixie","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
96491ccf2cd7961736e7da8014bd85d6dcb48964aa9b0feb1503dab49bb6d1bb
|
|
| MD5 |
f1f1f0fc2373eca05314ab7a9124bcee
|
|
| BLAKE2b-256 |
505db60d80b4910b811068a516b44b5b4bac20ae38b5da85374b1c088fd37172
|