Skip to main content

autocontext Python package

This package is the Python control plane for autocontext: scenario runs, solve, simulations, investigations, MCP/HTTP surfaces, persistent knowledge, training-data export, and local training hooks.

Use it when you want the full harness in Python, a CLI installed with uv/pip, or the MCP/HTTP server that coding agents can call.

Generation output that changes prompts or harness behavior is staged as an immutable ContextBundle. It becomes active only after matched screening, adaptive confirmation, and held-out evaluation; a successful strategy gate by itself does not activate a context edit. The artifact layout and Python API are documented in context bundles. The autocontext.analytics.context_attribution API joins controlled trials to those immutable digests, plans bounded re-ablation, and returns non-destructive prompt-selection decisions. See ablation-backed attribution.

Install

pip install autocontext
# or, for an isolated CLI tool:
uv tool install autocontext

Optional extras:

pip install 'autocontext[browser]'          # Chrome/CDP capture
pip install 'autocontext[primeintellect]'   # PrimeIntellect sandbox backend
pip install 'autocontext[mcp]'              # MCP server dependencies

The CLI entrypoint is autoctx. Provider env vars are listed in the repo-level .env.example.

Autoresearch checkpoint selection uses a minimum-effect gate by default and supports adaptive matched-trial confirmation through the Python API. Raw trials and stopping rationale are persisted separately from deployment promotion; see trainer-local statistical confirmation.

Long-running campaign operators can opt into a separately routed, frozen CampaignAuditor that reviews a sanitized evidence packet without mutation authority. Live checkpoints and the pre-promotion gate are idempotent, and production calls require a cancellable transport by default. Reviews are cached, bounded, and advisory; see the read-only campaign auditor.

Campaign plans can be executed through the optional, provider-neutral CampaignScheduler, with durable leases, resource/capability matching, comparable lanes, bounded retries and budgets, capability-driven warm reuse, worker cancellation, late-usage accounting, and a restart-safe service loop. See the campaign scheduler.

Code and research scenarios can opt into a process-backed ResearchWorkspace with explicit file, import, network, and host-bridge grants. Only approved trusted_local workspaces may import packages or run allow-listed subprocesses; isolated_sandbox requires a complete OS backend and never falls back. The shipped Docker backend provides a pinned, read-only, deny-network container with resource limits, opaque host-side credential brokering, transactional state, and verified cleanup; allowlisted egress needs a stronger deployment backend. The live queued-task path selects it with AUTOCONTEXT_WORKSPACE_INTERPRETER_BACKEND=docker, candidate execution, and explicit capability-approval settings. The existing restricted interpreter remains the default; see capability-scoped research workspaces.

Remote tasks use a provider-neutral request/result contract for resources, artifacts, lifecycle, events, usage, and cleanup. Prime Intellect implements that contract as an optional adapter and no longer owns scenario scoring logic; see remote execution sessions. Accelerators remain opt-in. Configure both the requested type/count and an operator-verified capability allowlist; image, region, required telemetry, resolved hardware, and attributable usage then remain identity-bound through campaign scheduling and external-evaluation accounting. Unsupported requests fail before provider creation and never fall back to local CPU execution. The shipped Prime runtime also commits a durable pre-dispatch claim and full typed result under runs/external-evaluations/; restart replays a committed result or surfaces an unresolved, possibly billable claim instead of dispatching the paid request again.

Run from a checkout

cd autocontext
uv venv
source .venv/bin/activate
uv sync --group dev

AUTOCONTEXT_AGENT_PROVIDER=deterministic \
uv run autoctx solve "improve customer-support replies for billing disputes" --iterations 3

Use a real provider by changing AUTOCONTEXT_AGENT_PROVIDER and setting its credential:

AUTOCONTEXT_AGENT_PROVIDER=anthropic \
ANTHROPIC_API_KEY=... \
uv run autoctx solve "improve customer-support replies for billing disputes" --iterations 3

Pi and local CLI providers avoid API-key plumbing when those tools are already authenticated:

AUTOCONTEXT_AGENT_PROVIDER=pi AUTOCONTEXT_PI_COMMAND=pi uv run autoctx solve "..." --iterations 3
AUTOCONTEXT_AGENT_PROVIDER=claude-cli AUTOCONTEXT_CLAUDE_MODEL=sonnet uv run autoctx solve "..." --iterations 3
AUTOCONTEXT_AGENT_PROVIDER=codex AUTOCONTEXT_CODEX_MODEL=o4-mini uv run autoctx solve "..." --iterations 3

Ollama uses llama3.1 by default. Set one local model override to fill every otherwise-unset role and tier slot, regardless of whether role routing is off or automatic:

AUTOCONTEXT_AGENT_PROVIDER=ollama \
AUTOCONTEXT_LOCAL_MODEL=qwen3:32b \
AUTOCONTEXT_PROVIDER_CAPABILITY=frontier \
uv run autoctx solve "..." --iterations 3

An explicit AUTOCONTEXT_MODEL_<ROLE> or AUTOCONTEXT_TIER_<TIER>_MODEL still takes precedence. AUTOCONTEXT_LOCAL_MODEL does not alter Anthropic's shipped per-role defaults.

With AUTOCONTEXT_ROLE_ROUTING=auto, unset OpenAI and OpenAI-compatible tier slots resolve to GPT-5.6 Sol/Terra/Luna for frontier/mid-tier/fast roles; OpenRouter resolves the same tiers to Claude Opus/Sonnet/Haiku. A generic OpenAI-compatible endpoint does not necessarily serve those ids, so set AUTOCONTEXT_LOCAL_MODEL when one gateway model should serve every role.

For endpoint-aware routing and cost estimates, set AUTOCONTEXT_PROVIDER_HOSTING to local or remote and, for a local endpoint, set AUTOCONTEXT_PROVIDER_CAPABILITY to fast, mid_tier, or frontier. Empty hosting retains conservative transport inference. A role-specific endpoint uses the corresponding AUTOCONTEXT_<ROLE>_PROVIDER_HOSTING and AUTOCONTEXT_<ROLE>_PROVIDER_CAPABILITY declarations instead of the default endpoint's declarations.

OpenAI-compatible role generation requests schema-constrained output by default. If that changes output quality for a backend, set AUTOCONTEXT_CONSTRAINED_OUTPUT=false to omit schemas from every role request and use the existing Markdown parsers instead. The setting also applies to roles with dedicated provider overrides.

Build-time teacher trace collection supports native deep_think tool loops on Anthropic and OpenAI-compatible providers. The collector requires a structured tool stream by default and keeps it separate from the final answer. Local and CLI/runtime providers report the capability as unsupported; visible-preamble fallback is available only through the collector's explicit require_thinking_stream=False option. For GPT 5.6+ models, reasoning_effort selects the external numeric prompt budget while native reasoning is requested off; a compatible gateway that rejects none is clamped to its lowest advertised level. Thinking payloads may contain sensitive prompt-derived data and should be redacted before persistence or export.

Common commands

Command Purpose
uv run autoctx solve "..." --iterations 3 Generate and run a scenario from a plain-language goal
uv run autoctx run <scenario> --iterations 3 Improve an existing scenario
uv run autoctx status <run-id> --json / watch <run-id> --ndjson Read one run snapshot or stream snapshots
uv run autoctx show <run-id> --best --json Inspect the best generation
uv run autoctx simulate --description "..." Create/replay/compare modeled-world simulations
uv run autoctx investigate --description "..." Run synthetic or iterative investigations
uv run autoctx list / status <run_id> / show <run_id> Inspect runs
uv run autoctx replay <run_id> --generation 1 Replay a generation before accepting knowledge
uv run autoctx queue add --task-prompt "..." --rubric "..." Queue evaluation/improvement work
uv run autoctx scenario create --family workflow --name support --description "..." Create a reusable scenario through a family-specific pipeline
uv run autoctx serve --host 127.0.0.1 --port 8000 Start the local HTTP API
uv run autoctx worker --poll-interval 5 --concurrency 2 Process queued tasks beside the API server
uv run autoctx serve mcp Expose the MCP tool surface
uv run autoctx export-training-data --scenario <name> --all-runs --output data.jsonl Build a training corpus (quarantined scores excluded by default; --include-quarantined to keep them)
uv run autoctx train --scenario <name> --data data.jsonl --time-budget 300 Run the local training hook
uv run autoctx epoch list [--scenario <name>] List evaluator-epoch registry records (candidate/active)
uv run autoctx epoch approve <scenario> <epoch_id> --charter ambient-charter.yaml Approve a candidate evaluator epoch and clear its quarantine
uv run autoctx hermes inspect --json Inspect Hermes Curator state

Saved custom scenarios under knowledge/_custom_scenarios/ can be rerun and benchmarked by name after their spec.json is persisted.

HTTP, MCP, and agents

The HTTP/WebSocket control plane is local-only unless explicitly secured. To bind beyond loopback, set a unique AUTOCONTEXT_SERVER_TOKEN of at least 32 characters and send it as an Authorization: Bearer value. Browser WebSocket clients use the autocontext.bearer.<base64url-token> subprotocol; query-string credentials are rejected. This token is a single-tenant containment control, not scoped multi-user authorization. Use TLS and an authenticated reverse proxy for any network-visible deployment.

uv sync --group dev --extra mcp
uv run autoctx serve mcp

Python runtime-backed run and solve calls append provider prompts/responses to run-scoped runtime-session logs. The same logs are readable through the cockpit HTTP API and MCP tools.

Detailed setup moved out of this README:

Contract probes

Contract probes turn observed harness traces into executable checks:

uv run autoctx probes check --suite contract-probes.json
uv run autoctx probes check --suite contract-probes.json --json
uv run autoctx probes extract --trace harness-trace.json --output contract-probes.json

Probe suites are strict JSON: unknown keys fail validation and required observation fields must be present. Pipe stdin with --suite - when another tool generates the suite.

Production traces

Wrap an existing Anthropic/OpenAI client once, then persist emitted traces through a sink:

from anthropic import Anthropic
from autocontext.integrations.anthropic import FileSink, instrument_client

sink = FileSink("./traces/anthropic.jsonl")
client = instrument_client(
    Anthropic(),
    sink=sink,
    app_id="billing-bot",
    environment_tag="prod",
)

For lower-level emit APIs, use autocontext.production_traces.build_trace and write_jsonl. Architecture notes are in ../docs/analytics.md and ../docs/opentelemetry-bridge.md.

Training

uv run autoctx export-training-data \
  --scenario support_triage --all-runs \
  --output training/support_triage.jsonl
uv run autoctx train \
  --scenario support_triage \
  --data training/support_triage.jsonl \
  --time-budget 300

For MLX/CUDA setup and case studies, use:

Repository layout

autocontext/
├── src/autocontext/       # Python package
├── tests/                 # pytest suite
├── docs/                  # package-specific docs
├── demo_data/             # small bundled examples
├── migrations/            # SQLite migrations
└── pyproject.toml

Development

uv run ruff check .
uv run mypy src
uv run pytest

Keep this README concise. Add deep reference prose to docs/ or the repo-level docs index instead.

Architect cadence and search efficiency

Python runs the architect on generations divisible by AUTOCONTEXT_ARCHITECT_EVERY_N_GENS (default 3). Other generations retain a skipped architect execution with zero token usage and no tool or harness changes. Set the value to 1 for an architect call every generation. This applies to direct, pipeline, and RLM generation paths.

Direct and pipeline execution overlap the independent architect with the analyst and coach; the coach still receives the current analyst's findings. RLM sessions remain serial. Strategy search reuses unchanged knowledge-file contents while checking file revisions and current database summaries on every query.

Metadata

Release files for autocontext 0.18.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for autocontext 0.18.0
File Size Uploaded
autocontext-0.18.0.tar.gz 3.8 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for autocontext 0.18.0
File Interpreter ABI Platform
autocontext-0.18.0-py3-none-any.whl Python 3 none any Details

Total release size: 6.0 MB

Release files / autocontext-0.18.0.tar.gz

Download URL autocontext-0.18.0.tar.gz
Size 3.8 MB
Tags Source
SHA-256 checksum
How to use checksums
ed0a5988f66c3076c48b5ac38a572f800fbd959a06c49bee6e823ed89c3873a0
BLAKE2b-256 checksum
How to use checksums
f793480386d54def4a6796bdf32f8b3b1b650c3599984b0d90e5cd67516628e6
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release files / autocontext-0.18.0-py3-none-any.whl

Download URL autocontext-0.18.0-py3-none-any.whl
Size 2.2 MB
Tags Python 3
SHA-256 checksum
How to use checksums
0ff06060d55f7e85441129be3cff2e285ad0a4759782cc1ca5bc4649d31aec69
BLAKE2b-256 checksum
How to use checksums
3b24ab01ec8b2f805659f8c3c8e4e33ba1f51725852dc3de2fd2aadcc815375e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 16, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.18.0 This release

2 release files

0.17.0

2 release files

0.16.1

2 release files

0.15.1

2 release files

0.15.0

2 release files

0.14.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.1

2 release files

0.5.0

2 release files

0.4.9

2 release files

0.4.8

2 release files

0.4.7

2 release files

0.4.6

2 release files

0.4.5

2 release files

0.4.4

2 release files

0.4.3

2 release files

0.4.2

2 release files

0.4.1

2 release files

0.4.0

2 release files

0.3.7

2 release files

0.3.6

2 release files

0.3.5

2 release files

0.3.4

2 release files

0.3.3

2 release files

0.3.2

2 release files

0.3.1

2 release files

0.0.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page