Skip to main content

maf-sandbox-codeact

PyPI Python License

Experimental. This package is early-stage (pre-1.0, Development Status :: 4 - Beta) — its API may change or be removed in a future release without notice. Importing it emits a one-time MafSandboxCodeactExperimentalWarning; suppress it with warnings.filterwarnings("ignore", category=maf_sandbox_codeact.MafSandboxCodeactExperimentalWarning) once you've read the notice.

This package is not affiliated with, endorsed by, or a product of Microsoft — it is a third-party reference implementation of microsoft/agent-framework#7568 for Microsoft Agent Framework.

CodeAct as a Microsoft Agent Framework tool: the agent gets one tool, execute_code; the model writes a short Python program; the program runs inside a sandbox and the tool returns what it printed. Computing an answer beats reasoning about what the computation would produce — and the code that does it runs somewhere the host is not.

app  ->  maf_sandbox  ->  a backend (maf-sandbox-acas, maf-sandbox-wslc, ...)  ->  this workload

This package is a sandbox kind in the sense of maf-sandbox's protocol. It contains no Azure import, no backend import and no sandbox lifecycle code; it asks a SandboxRouter for a sandbox and gets back write_file and exec, so the same tool runs unchanged against ACA Sandboxes, a WSL container or an in-process fake. Tests enforce both boundaries.

Quickstart

pip install maf-sandbox-codeact
from maf_sandbox_codeact import make_codeact_tools

tools = make_codeact_tools(router, "data-analyst", context,
                           image="mcr.microsoft.com/devcontainers/python:3.13-bookworm")

Pass router=None — or a router with no backend — and you get [] back: an unconfigured host attaches no tool rather than one that fails when called. A backend that cannot exec, or cannot take files in, is refused right there with SandboxCapabilityNotSupported, before the model is shown a capability it does not have.

router and context are the host's, and this snippet shows neither being built. samples/03_acas_codeact and samples/04_wslc_codeact are the whole wiring as runnable programs — the same agent on a microVM-isolated Azure backend and on a container on your own machine.

What the model gets

One tool, execute_code. The program is written to a directory of its own and run as the argv ["python3", ".../program.py"], and the result is its stdout, its stderr when it wrote any, and its exit code when that was not zero. There is no REPL echo, so a program that computes without printing returns a sentence saying so.

Every call gets a fresh directory, and that is load-bearing rather than hygiene. acquire is get-or-create, so the same sandbox serves every call in a conversation. Without a per-call directory a file deleted from the workspace between rounds would still be there for the next program to read as current, and last round's output file would be collected as this round's — a stale answer presented as a live one, in a kind whose whole job is transforming files.

Two further channels exist and neither is on by default. Wire neither and this is the stdout-only kind it has always been.

Files in

Pass a workspace_store and the tool grows a files parameter:

tools = make_codeact_tools(router, "data-analyst", context,
                           workspace_store=store, image=...)

Each named file is read from the store and written into the program's working directory under its own name, so data/sales.csv is what the program opens. The caller's listing is the authority: only a name present in WorkspaceContext.list_files is ever shared, so a name the model invented — or read out of a file it was given — has nowhere to go. A name outside the listing comes back as a refusal naming the near misses; a name that traverses comes back as a refusal that echoes nothing.

Files out

Produced files never come back as bytes. They go to a host-supplied OutputSink, and the model gets the reference the sink returned. Two ways to name them, and the host picks one:

from maf_sandbox_codeact import CodeactOutputs, make_codeact_tools

tools = make_codeact_tools(router, "data-analyst", context,
                           output_sink=sink, outputs=CodeactOutputs.DECLARED, image=...)
  • DECLARED adds an outputs parameter: the model says what its program will write before it runs. Names are validated and capped up front, and one declared but not written is reported back by name rather than dropped. Prefer this.
  • MANIFEST has the program write outputs.json listing what it produced — for a program whose output names it can only know once it has read its input. The names are then the guest's rather than the model's, settled after the fact. The manifest is itself a file the collection moved, so it takes one slot of files_out.max_files and its bytes count against the ceilings; a cap below 2 leaves no room for an artifact and is refused at attach.

Either way the kind requires FILES_OUT and never FILES_LIST: it collects literal paths and never enumerates a directory, so it runs on every backend that serves the pull surface at all rather than only on the one with the richest file API. files_out.max_files is what bounds how many artifacts a single call may produce, and files_in bounds what one call may share in — count, per-file bytes and total. Both are enforced by this kind, because no backend's write_file or read_file knows the workload's caps.

No media type is ever taken from the guest. Artifact.media_type is None on both roads: the kind does not know what a model-written program produced, and a value read out of outputs.json would be the guest telling the host how to handle its own bytes — which a sink may act on to choose inline rendering. A host that wants to decide by extension has Artifact.name and its own policy.

Where files land is the host's decision, never this kind's. That is the point of the sink, and it matters more here than for any other kind: these bytes were authored by model-written code. A host that points the sink at the same store the agent's own file tools write to has given that code an unapproved file_access_write, and one that lets it overwrite has given it a way to influence a different tool on the next call. Point it somewhere the agent cannot otherwise reach.

Threat model

The source is never a command line. Model-written code reaches the interpreter as file content and the command is a fixed two-element argv — a sequence, not a shell string — so there is no command line for the source to be part of, nothing to quote, and nothing to escape. That is the security-relevant decision in this package and it is pinned by a test.

Egress is closed. SandboxSpec.egress_allow is empty, stated as a property of the workload rather than of configuration: the program computes, it does not fetch. A backend that cannot confine egress at all is refused at attach.

Nothing is dispatchable from inside. There is no host-tool registry in this version, and that emptiness is the security story rather than a missing feature: the program cannot open a socket and cannot call a host function, so it initiates nothing. The output sink does not change that — the kind calls it host-side, after the program has exited, and nothing inside the sandbox can reach it. A host wanting a hard stop denies FILES_OUT.

A workspace store is ingress, and it is the host's own. "Nothing can get in" describes what the program can initiate, not what the host puts there. With workspace_store wired, caller-selected files are written into the sandbox before the program runs — deliberately, and constrained to the caller's listing, so the model cannot widen the set. What that content is remains the host's to know: a workspace file may itself carry text from somewhere untrusted, and a program that parses it is running on input the sandbox did not vet. Wire no store and this paragraph does not apply.

The tool declares no source_integrity. The library's default is "trusted", which is right for a workload whose result is a compiler's own diagnostics and wrong for this one: what comes back is whatever a model-written print(...) chose to emit. Undeclared, MAF's information-flow tracker applies its untrusted default and the result taints the conversation — the fail-safe direction, and the honest one.

Isolation is the host's call, and a store changes what that call is about. This kind does not raise SandboxSpec.min_isolation, so the router's floor governs — MICROVM unless the host opted down. A kind that ran code influenced by untrusted external content would pin the floor itself, and this one cannot know whether it is one: with no store, the program's only input is source the model wrote, and opting down to CONTAINER weighs model-written code against a shared kernel. With a store, the program also reads whatever those files contain, so the floor should be chosen against the provenance of the workspace, not against this kind's defaults. Only the host knows that.

What this version is not

Host-tool dispatch and the RUN_CODE road served by an embedded-interpreter backend are absent on purpose. The design that governs them — capabilities declared by backends and required by specs, and what HOST_TOOLS would have to carry before it ships — is docs/design/two-axis-sandbox-policy.md; the file channels above are specified in docs/design/files-out.md.

There is also no way to delete a file from a sandbox, which is why staleness is answered by a fresh directory per call rather than by cleaning the old one. A long conversation therefore accumulates one directory per call until the sandbox is disposed — and a program that walks upwards can still open them. The fresh directory removes staleness from the namespace, so a program reading data.csv gets this call's or nothing; it does not put earlier rounds out of reach. Everything reachable that way belongs to the same conversation and the same agent, since that is what a sandbox is keyed by.


Maintained by SOKOLAI BV.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

maf_sandbox_codeact-0.2.3.tar.gz (21.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

maf_sandbox_codeact-0.2.3-py3-none-any.whl (21.8 kB view details)

Uploaded Python 3

File details

Details for the file maf_sandbox_codeact-0.2.3.tar.gz.

File metadata

  • Download URL: maf_sandbox_codeact-0.2.3.tar.gz
  • Upload date:
  • Size: 21.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for maf_sandbox_codeact-0.2.3.tar.gz
Algorithm Hash digest
SHA256 cbc828d35a6fc53056759ffeaaad36a4773bc187e8693f5398d9e40c34a304ea
MD5 e64880fb23aa7736245ac3db1302c6fb
BLAKE2b-256 b1de331df3e7261a0adc35cb301ccc005927bf5d045a6db487e82ebdf0824818

See more details on using hashes here.

Provenance

The following attestation bundles were made for maf_sandbox_codeact-0.2.3.tar.gz:

Publisher: publish-packages.yml on sokolaidev/maf-extensions

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file maf_sandbox_codeact-0.2.3-py3-none-any.whl.

File metadata

File hashes

Hashes for maf_sandbox_codeact-0.2.3-py3-none-any.whl
Algorithm Hash digest
SHA256 c193969172c86b3178d822a955aa5feb66dc7752eb5e5e538199e01ad93b2e2f
MD5 3c7d74c8ede8e9210d3b9f73a1d8da44
BLAKE2b-256 ad15717bfc5eb932d319edb79508090529ec9da30236ca84e27271abad9fd389

See more details on using hashes here.

Provenance

The following attestation bundles were made for maf_sandbox_codeact-0.2.3-py3-none-any.whl:

Publisher: publish-packages.yml on sokolaidev/maf-extensions

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.7.6

2 files

0.7.5

2 files

0.7.4

2 files

0.7.3

2 files

0.7.2

2 files

0.7.1

2 files

0.7.0

2 files

0.6.1

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.3

2 files

0.4.2

2 files

0.4.1

2 files

0.4.0

2 files

0.3.0

2 files

This release

0.2.3 This release

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page