mini-agent
mini-agent is one small agent loop for three environment families: software
engineering, web research, and computer use. Provider codecs, tools, benchmark
loaders, graders, storage, and multi-agent scheduling stay outside that loop.
The design follows the useful constraint in mini-swe-agent: keep the agent boring and make everything else replaceable. This project is alpha software. It is infrastructure for controlled experiments, not a claim that one normalized prompt reproduces every provider's proprietary harness.
Quickstart
pip install mini-agent-cmu
(The distribution is mini-agent-cmu — PyPI blocks the bare name as too
similar to an unrelated existing project — but the import stays
mini_agent and the command stays mini-agent.)
Run a research agent over a local JSONL corpus — no VM, no benchmark checkout, no search API key; retrieval is deterministic in-process BM25:
export OPENAI_API_KEY=...
mini-agent run --environment web --web-backend jsonl \
--corpus examples/corpus.jsonl \
--task 'Which document explains when the wind turbine furls?' \
--model openai/MODEL
Each corpus line is one JSON object with docid (or id) and text (or
contents); see examples/corpus.jsonl. The run
prints its output directory on start — tail -f <output>/trace.jsonl streams
every model and tool event live.
Or drive the loop as a library with no key at all
(examples/library_quickstart.py is the
complete runnable version):
import asyncio
from mini_agent import MiniAgent, RunContext, ScriptedModel, ModelResponse
from mini_agent.models import build_model # also exported as mini_agent.build_model
model = ScriptedModel([ModelResponse("done")]) # or build_model("openai/MODEL")
agent = MiniAgent(
model=model,
environment=my_environment, # any BaseEnvironment subclass
system_prompt="Solve the task with the available tools.",
max_steps=64,
context=RunContext(),
)
result = asyncio.run(agent.run("the task"))
See docs/library.md for custom environments, budgets, and accounting.
Design contract
MiniAgentowns only linear history, model calls, declared tool calls, and stopping.- A SWE environment exposes one
bashtool. A web environment exposes onebrowsertool; general runs enablesearchandopen, while fixed BrowseComp-Plus is search-only. A computer environment exposes one batchedcomputertool using native screenshot pixels. - OpenAI Responses, OpenAI-compatible Chat Completions, and Anthropic Messages
are thin maintained codecs. Meta deployments require an explicit endpoint and
a
--protocolchoice; protocol- or model-specific behavior belongs downstream unless it is truly universal. - Every maintained run path uses the same concurrency-safe budget ledger and streaming trace recorder.
- Multi-agent mode adds one
agenttool withspawn,send, blocking or non-blockinginbox,wait, descendantstop, and explicitadoptactions. Every worker is still an ordinaryMiniAgent. - Evaluation control data, expected answers, qrels, and verifiers never enter an agent environment or prompt.
Prefer delegation on by default? pip install multi-mini-agent adds a
multi-mini-agent console script — the identical CLI with --multi-agent
implied for profile/run/eval (a thin companion distribution from
packaging/multi-mini-agent, not a second
codebase).
The recursive scheduler has no topology depth constant. It is still finite: global and per-agent budgets, active-agent and total-agent limits, mailbox limits, tool limits, timeouts, and environment leases all apply. It is a minimal inference topology; it does not claim to reproduce any policy-trained multi-agent method.
Install
Python 3.10 through 3.13 is supported. The upstream-locked web-fixed extra
supports Python 3.10 through 3.12 because Pyjnius 1.6.1 does not publish a
Python 3.13 wheel.
python -m pip install -e .
# Only install the extras needed on this machine.
python -m pip install -e '.[web-live]'
python -m pip install -e '.[web-fixed]'
SWE-bench grading requires an exact editable upstream checkout because the tag's built wheel omits tracked runtime resource files used by the official harness:
git clone https://github.com/SWE-bench/SWE-bench.git /src/SWE-bench
git -C /src/SWE-bench checkout 726c5461e2ef52d83cf1ea2107870a8bb3328d57
python -m pip install -e /src/SWE-bench
Upstream did not publish version 4.1.0 to PyPI, so a plain
swebench==4.1.0 requirement is not installable. Grading verifies the installed
editable source and its required resource inventory before and after use; a
wheel built from the tag is intentionally rejected as incomplete.
Computer benchmarks use their pinned upstream checkouts and environment
dependencies; there is no synthetic catch-all computer extra.
The fixed-web extra installs only the pinned Rust tokenizer, its Hub downloader,
and the JNI bridge. It reads the same Qwen tokenizer.json used upstream without
installing a general model framework. The 185 MB Anserini fat JAR and Lucene
index are explicit, hashed benchmark assets rather than an excuse to install
Pyserini's unused dense-retrieval and server stack.
Models use provider/model syntax:
openai/MODELusesOPENAI_API_KEYand OpenAI Responses by default.anthropic/MODELusesANTHROPIC_API_KEYand Anthropic Messages.meta/MODELusesMODEL_API_KEY; it requires--base-urlbecause mini-agent does not guess a Meta deployment's hostname or protocol support. Without--protocolit remains the opt-in Responses-compatible adapter.
--protocol chat-completions selects the OpenAI-compatible Chat Completions
adapter for openai/ and meta/ models. It resends the full wire transcript
each turn (no previous_response_id) and never sets sampling parameters, so
the server's defaults always apply. --provider-header NAME=VALUE attaches
non-secret static headers (for example a deployment-required session id) to
every provider call; credential-looking header names are rejected.
Transcript-replay codecs (Chat Completions and Anthropic Messages) keep only
the newest --max-history-images screenshots when replaying history (default
4); older image blocks become a fixed text placeholder, declared as a
translation loss and recorded in run manifests. This is what makes long
computer-use runs feasible on those paths — a 64-step run replays at most K
images per call instead of every prior screenshot. Pass
--max-history-images unlimited to restore full replay. The Responses path
keeps continuation server-side and rejects the option.
Transient provider failures (408/429/5xx and transport errors) are retried
with jittered backoff, honoring Retry-After; --provider-retries N bounds
the attempts (0 disables) and --provider-timeout overrides the 300-second
per-request default. Retries never extend a run past its wall-clock budget,
and retry counts appear in trace events.
Override a trusted compatible endpoint with --base-url and its credential name
with --api-key-env. Put credentials only in environment variables, never in
--provider-body or --provider-header.
For a live evaluation intended to be reproducible, pass the exact model value
the provider promises to return as --expected-provider-model (and
--grader-expected-provider-model for BrowseComp). The adapter hashes every
observed response model, rejects an unexpected snapshot, and rejects a snapshot
change within one agent run. Without this option a requested alias is recorded
honestly as an alias, not treated as an immutable model-card identity.
The OpenAI Responses adapter chains tool turns with previous_response_id.
Accordingly, it rejects {"store": false}: Zero Data Retention requires a
different downstream adapter that replays every response item, including encrypted
reasoning items, instead of pretending this continuation mode is compatible.
CLI
The public surface is intentionally small:
mini-agent profile
mini-agent run --environment swe|web|computer
mini-agent eval --benchmark swebench|browsecomp|browsecomp-plus|osworld-v1|osworld-v2|cua-speed-run
mini-agent grade --benchmark swebench|browsecomp-plus
mini-agent doctor
Resolve the maintained baseline without making a model call:
mini-agent profile \
--application web \
--profile default \
--model openai/MODEL
The provider-neutral, versioned agent contract is available as canonical JSON:
mini-agent profile \
--application web \
--model openai/MODEL \
--multi-agent \
--format agent-spec
Use --format translation-report for an explicit field-level loss report. The
report includes the selected provider codec's declared losses — the
OpenAI-protocol codecs drop the tool-result error flag and relocate tool-result
images into a synthetic user message, and every codec restricts tools to the
generic function kind — so exact is only reported when the codec declares
none. Even then the claim is deliberately scoped to declared fields; it is
never a claim of behavioral, policy-training, timing, tool, or benchmark
equivalence.
Downstream Python adapters can load that document with AgentSpecV1.from_json
and call spec.bind(model=..., model_id=..., environment=..., environment_id=...). Binding checks the explicit model/domain identities,
runtime tool names, communication action enum, step limit, prompt, and shared
budget before constructing the same MiniAgent; endpoint credentials and
benchmark assets stay in their respective downstream constructors.
Run a SWE task. Single-agent mode edits the selected workspace directly:
mini-agent run \
--environment swe \
--workspace /path/to/repository \
--task 'Fix the failing tests and verify the change.' \
--model openai/MODEL \
--home /path/to/durable/mini-agent \
--scratch /path/to/local-scratch/mini-agent
With --multi-agent, every SWE worker gets an independent private repository
copy. The root's selected state is exported as patch.diff; the source workspace
is not changed.
Run live web research with SerpAPI, or use JSONL BM25 for a small deterministic fixture:
export SERPAPI_API_KEY=...
mini-agent run \
--environment web \
--web-backend serpapi \
--page-reader http \
--task 'Answer the question and cite the returned references.' \
--model openai/MODEL
mini-agent run \
--environment web \
--web-backend jsonl \
--corpus /path/to/corpus.jsonl \
--task 'Find the relevant document.' \
--model openai/MODEL
Direct computer use expects a trusted cua-speed-run-compatible gateway. Direct multi-agent computer use is rejected because it has no machine-lease factory:
mini-agent run \
--environment computer \
--env-url http://127.0.0.1:8000 \
--task 'Complete the visible desktop task.' \
--model anthropic/MODEL
Non-loopback gateways must use HTTPS and a bearer token supplied only through an
environment variable, for example --env-token-env CUA_TASK_TOKEN.
Use --max-model-calls, token limits, --max-cost-usd, wall time, tool limits,
and agent limits for every paid run. A cost cap requires explicit input and output
prices; unknown or incomplete usage fails closed.
Evaluations
eval defaults to one task as a canary. Use --all deliberately. Generation is
separate from official grading.
# SWE-bench generation in independent persistent containers.
mini-agent eval \
--benchmark swebench \
--dataset /data/swebench-verified.jsonl \
--runtime docker \
--model openai/MODEL \
--output /path/to/durable/mini-agent/runs/swe-canary \
--scratch /path/to/local-scratch/mini-agent-swe
# Live BrowseComp plus an independently configured private-answer grader.
mini-agent eval \
--benchmark browsecomp \
--dataset /data/browsecomp.csv \
--model openai/MODEL \
--grader-model openai/GRADER_MODEL \
--output /path/to/durable/mini-agent/runs/browsecomp-canary
# Fixed-corpus BrowseComp-Plus generation.
mini-agent eval \
--benchmark browsecomp-plus \
--dataset /data/browsecomp-plus/queries.tsv \
--index /data/browsecomp-plus/index \
--anserini-jar /data/browsecomp-plus/anserini-1.1.1-fatjar.jar \
--snippet-tokenizer Qwen/Qwen3-0.6B \
--snippet-tokenizer-revision COMMIT \
--model openai/MODEL
# OSWorld lifecycle and hidden evaluation from an exact pinned checkout.
# On a Docker-less KVM host, run the official container image with Apptainer.
mini-agent eval \
--benchmark osworld-v1 \
--checkout /src/OSWorld \
--provider-name docker \
--runtime apptainer \
--path-to-vm /assets/osworld/Ubuntu.qcow2 \
--osworld-apptainer-image /assets/osworld/osworld-docker.sif \
--model openai/MODEL
# Local cua-speed-run environment plane and checker.
mini-agent eval \
--benchmark cua-speed-run \
--checkout /src/cua-speed-run \
--benchmark-path /data/cua-benchmark \
--backend gym-anything-qemu-apptainer \
--qemu-cache /assets/gym-anything/qemu \
--model openai/MODEL
Every evaluation writes an immutable manifest fingerprint, one shared redacted
JSONL trace, atomic task results with commit markers, accounting, timing, and
hash-bound benchmark-specific artifacts. Artifact bytes and their atomic rename
are synced before a commit marker is treated as durable. --resume accepts the
exact same task data, configuration, and limits, and restores already-spent
accounting.
If a crash leaves an uncommitted task after a model or tool operation started,
resume refuses to guess its spend or side effects; start a new output rather
than silently repeating it. The start record is synced to durable storage before
the provider or environment request begins.
run/
manifest.json
trace.jsonl
summary.json
instances/<hashed-task-id>/result.json
instances/<hashed-task-id>/completed.json
predictions.jsonl # SWE-bench
official_runs/ # BrowseComp-Plus
Official graders are explicit commands:
mini-agent grade \
--benchmark swebench \
--evaluation /path/to/evaluation \
--dataset /data/SWE-bench_Verified.jsonl \
--run-id mini-agent-canary
mini-agent grade \
--benchmark browsecomp-plus \
--evaluation /path/to/evaluation \
--checkout /src/BrowseComp-Plus \
--ground-truth /data/ground_truth.jsonl \
--qrel-evidence /data/qrel_evidence.txt \
--judge-model /models/Qwen3-32B-materialized
Grading accepts only a fingerprint-valid matching evaluation manifest and local
content-hashed inputs. It snapshots predictions/runs and hidden answer data into
a private 0700 grade directory before invoking upstream. SWE-bench requires
the current Python to contain exactly swebench==4.1.0. Image verification uses
the same Docker SDK, explicit DOCKER_HOST, and allowlisted environment as the
upstream grader; the generation manifest's runtime command is inert provenance.
BrowseComp-Plus checks
the pinned checkout and lockfile plus the lock's direct grader versions
(numpy==1.26.4, tqdm==4.67.1, vllm==0.9.0.1). Its judge model must be a
materialized, symlink-free local directory, not a mutable model name. The grade
manifest, captured stdout/stderr, return code, and hash inventory of official
outputs are private evidence; do not publish the grade directory.
See benchmark fidelity for exact pins, prerequisites, and intentional differences from upstream model harnesses.
Storage and preflight
Keep durable evidence and disposable machine state separate:
export MINI_AGENT_HOME=/path/to/durable/mini-agent
export MINI_AGENT_SCRATCH=/path/to/local-scratch/mini-agent
mini-agent doctor --target storage
mini-agent doctor --target swebench --runtime docker
mini-agent doctor --target web --web-mode live --page-reader http
mini-agent doctor \
--target computer \
--checkout /src/OSWorld \
--osworld-version v1 \
--runtime apptainer \
--path-to-vm /assets/osworld/Ubuntu.qcow2 \
--osworld-apptainer-image /assets/osworld/osworld-docker.sif
doctor is non-paid and machine-readable. It reports prerequisites; it does not
claim that inference, a VM launch, or official grading passed.
See machine images for exact upstream pins, acquisition commands, hashes from the validated node, and the optional adjacent provenance sidecar contract.
Environment variables
| Variable | Read by | Purpose |
|---|---|---|
OPENAI_API_KEY / ANTHROPIC_API_KEY / MODEL_API_KEY |
provider codecs | default credentials for openai/, anthropic/, and meta/ models; override the name with --api-key-env |
MINI_AGENT_HOME |
storage | durable root for runs, assets, and caches (default ~/.local/share/mini-agent) |
MINI_AGENT_SCRATCH |
storage | disposable node-local work area (default $MINI_AGENT_HOME/work) |
SERPAPI_API_KEY |
live web backend | SerpAPI credential for --web-backend serpapi |
GYM_ANYTHING_QEMU_CACHE |
cua-speed-run | upstream QEMU image cache directory (also --qemu-cache) |
GYM_ANYTHING_QEMU_WORK_DIR |
cua-speed-run | set by the CLI to the run's scratch QEMU work area |
OSWORLD_QEMU_BASE_IMAGE |
cua-speed-run | optional base-image override honored by the upstream runner |
CS_ALLOW_QEMU_TCG |
cua-speed-run | opt-in to software-emulated QEMU when KVM is unavailable (slow) |
GYM_ANYTHING_QEMU_CONTAINER |
cua-speed-run (upstream) | runtime container reference; pin a SIF path or @sha256: digest for comparable runs |
DOCKER_HOST |
SWE-bench / OSWorld docker paths | explicit Docker engine endpoint; required for official SWE-bench grading |
What results mean
The adapters preserve task loading, environment isolation, output schemas, and hidden evaluator boundaries where documented. The generic tool topology and baseline prompt are intentionally not the official provider-specific leaderboard harnesses. In particular:
- BrowseComp-Plus uses one search-only
browseraction rather than the upstream provider-specific search name, while retaining fixed retrieval and official run/grader artifacts. Full-documentopenis intentionally absent because it is opt-in, not the pinned upstream runner's default. - OSWorld uses a batched native-pixel
computeraction protocol adapted onto the official reset/step/evaluate lifecycle. - cua-speed-run executes the environment plane and checker in process; it is not the upstream two-file submission/executor protocol.
- Multi-agent mode is a bounded scheduler, not best-of-N, a trained recursive policy, or a simulation of proprietary team runtimes.
A passing smoke test is internal QA, not evidence that a configuration is a release winner. Current machine evidence and blockers live in validation.
Development
PYTHONPATH=src python3 -m unittest discover -s tests -v
PYTHONPATH=src python3 -m compileall -q src tests
python3 -m ruff check src tests
python3 -m mypy src/mini_agent
PYTHONPATH=src python3 -m mini_agent profile \
--application web --profile default --model openai/test-model
See architecture, the pinned reference ledger, contributing, and security.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mini_agent_cmu-0.5.0.tar.gz.
File metadata
- Download URL: mini_agent_cmu-0.5.0.tar.gz
- Upload date:
- Size: 288.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c18c9ba92cb92399de6650306fcebe45ad87d4a586aef63adad68362c7ffde9b
|
|
| MD5 |
0b8b5115da8a534b43dfb5c1abf896a4
|
|
| BLAKE2b-256 |
b11419c45688f47f082fbd073b0f885d5940af183ebc5ab21b13f6d9270acbef
|
Provenance
The following attestation bundles were made for mini_agent_cmu-0.5.0.tar.gz:
Publisher:
release.yml on ljang0/mini-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
mini_agent_cmu-0.5.0.tar.gz -
Subject digest:
c18c9ba92cb92399de6650306fcebe45ad87d4a586aef63adad68362c7ffde9b - Sigstore transparency entry: 2441872981
- Sigstore integration time:
-
Permalink:
ljang0/mini-agent@0fdb0d7c3cb7bf9808fa5952e3d9d86e3df6880e -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/ljang0
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@0fdb0d7c3cb7bf9808fa5952e3d9d86e3df6880e -
Trigger Event:
release
-
Statement type:
File details
Details for the file mini_agent_cmu-0.5.0-py3-none-any.whl.
File metadata
- Download URL: mini_agent_cmu-0.5.0-py3-none-any.whl
- Upload date:
- Size: 185.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
41e844d1e6841c546ba65368d939f4d2475f404f036d63efd5465cd23e46edbe
|
|
| MD5 |
be601ed7375eb18553d5e54644e1ae90
|
|
| BLAKE2b-256 |
750f29fa69b6e2564b899ee4bb74a6f3c9137c458b944ce1c05f9d7b776a4f8a
|
Provenance
The following attestation bundles were made for mini_agent_cmu-0.5.0-py3-none-any.whl:
Publisher:
release.yml on ljang0/mini-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
mini_agent_cmu-0.5.0-py3-none-any.whl -
Subject digest:
41e844d1e6841c546ba65368d939f4d2475f404f036d63efd5465cd23e46edbe - Sigstore transparency entry: 2441873074
- Sigstore integration time:
-
Permalink:
ljang0/mini-agent@0fdb0d7c3cb7bf9808fa5952e3d9d86e3df6880e -
Branch / Tag:
refs/tags/v0.5.0 - Owner: https://github.com/ljang0
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@0fdb0d7c3cb7bf9808fa5952e3d9d86e3df6880e -
Trigger Event:
release
-
Statement type: