autocontext Python package
This package is the Python control plane for autocontext: scenario runs, solve, simulations, investigations, MCP/HTTP surfaces, persistent knowledge, training-data export, and local training hooks.
Use it when you want the full harness in Python, a CLI installed with uv/pip, or the MCP/HTTP server that coding agents can call.
Generation output that changes prompts or harness behavior is staged as an
immutable ContextBundle. It becomes active only after matched screening,
adaptive confirmation, and held-out evaluation; a successful strategy gate by
itself does not activate a context edit. The artifact layout and Python API are
documented in context bundles.
The autocontext.analytics.context_attribution API joins controlled trials to
those immutable digests, plans bounded re-ablation, and returns non-destructive
prompt-selection decisions. See ablation-backed attribution.
Install
pip install autocontext
# or, for an isolated CLI tool:
uv tool install autocontext
Optional extras:
pip install 'autocontext[browser]' # Chrome/CDP capture
pip install 'autocontext[primeintellect]' # PrimeIntellect sandbox backend
pip install 'autocontext[mcp]' # MCP server dependencies
The CLI entrypoint is autoctx. Provider env vars are listed in the repo-level .env.example.
Autoresearch checkpoint selection uses a minimum-effect gate by default and supports adaptive matched-trial confirmation through the Python API. Raw trials and stopping rationale are persisted separately from deployment promotion; see trainer-local statistical confirmation.
Long-running campaign operators can opt into a separately routed, frozen
CampaignAuditor that reviews a sanitized evidence packet without mutation
authority. Live checkpoints and the pre-promotion gate are idempotent, and
production calls require a cancellable transport by default. Reviews are
cached, bounded, and advisory; see the
read-only campaign auditor.
Campaign plans can be executed through the optional, provider-neutral
CampaignScheduler, with durable leases, resource/capability matching,
comparable lanes, bounded retries and budgets, capability-driven warm reuse,
worker cancellation, late-usage accounting, and a restart-safe service loop.
See the campaign scheduler.
Code and research scenarios can opt into a process-backed ResearchWorkspace
with explicit file, import, network, and host-bridge grants. Only approved
trusted_local workspaces may import packages or run allow-listed subprocesses;
isolated_sandbox requires a complete OS backend and never falls back. The
shipped Docker backend provides a pinned, read-only, deny-network container
with resource limits, opaque host-side credential brokering, transactional state, and
verified cleanup; allowlisted egress needs a stronger deployment backend. The
live queued-task path selects it with
AUTOCONTEXT_WORKSPACE_INTERPRETER_BACKEND=docker, candidate execution, and
explicit capability-approval settings. The existing restricted interpreter
remains the default; see
capability-scoped research workspaces.
Remote tasks use a provider-neutral request/result contract for resources,
artifacts, lifecycle, events, usage, and cleanup. Prime Intellect implements
that contract as an optional adapter and no longer owns scenario scoring logic;
see remote execution sessions.
Accelerators remain opt-in. Configure both the requested type/count and an
operator-verified capability allowlist; image, region, required telemetry,
resolved hardware, and attributable usage then remain identity-bound through
campaign scheduling and external-evaluation accounting. Unsupported requests
fail before provider creation and never fall back to local CPU execution.
The shipped Prime runtime also commits a durable pre-dispatch claim and full
typed result under runs/external-evaluations/; restart replays a committed
result or surfaces an unresolved, possibly billable claim instead of dispatching
the paid request again.
Run from a checkout
cd autocontext
uv venv
source .venv/bin/activate
uv sync --group dev
AUTOCONTEXT_AGENT_PROVIDER=deterministic \
uv run autoctx solve "improve customer-support replies for billing disputes" --iterations 3
Use a real provider by changing AUTOCONTEXT_AGENT_PROVIDER and setting its credential:
AUTOCONTEXT_AGENT_PROVIDER=anthropic \
ANTHROPIC_API_KEY=... \
uv run autoctx solve "improve customer-support replies for billing disputes" --iterations 3
Pi and local CLI providers avoid API-key plumbing when those tools are already authenticated:
AUTOCONTEXT_AGENT_PROVIDER=pi AUTOCONTEXT_PI_COMMAND=pi uv run autoctx solve "..." --iterations 3
AUTOCONTEXT_AGENT_PROVIDER=claude-cli AUTOCONTEXT_CLAUDE_MODEL=sonnet uv run autoctx solve "..." --iterations 3
AUTOCONTEXT_AGENT_PROVIDER=codex AUTOCONTEXT_CODEX_MODEL=o4-mini uv run autoctx solve "..." --iterations 3
Ollama uses llama3.1 by default. Set one local model override to fill every
otherwise-unset role and tier slot, regardless of whether role routing is off
or automatic:
AUTOCONTEXT_AGENT_PROVIDER=ollama \
AUTOCONTEXT_LOCAL_MODEL=qwen3:32b \
AUTOCONTEXT_PROVIDER_CAPABILITY=frontier \
uv run autoctx solve "..." --iterations 3
An explicit AUTOCONTEXT_MODEL_<ROLE> or AUTOCONTEXT_TIER_<TIER>_MODEL
still takes precedence. AUTOCONTEXT_LOCAL_MODEL does not alter Anthropic's
shipped per-role defaults.
With AUTOCONTEXT_ROLE_ROUTING=auto, unset OpenAI and OpenAI-compatible tier
slots resolve to GPT-5.6 Sol/Terra/Luna for frontier/mid-tier/fast roles;
OpenRouter resolves the same tiers to Claude Opus/Sonnet/Haiku. A generic
OpenAI-compatible endpoint does not necessarily serve those ids, so set
AUTOCONTEXT_LOCAL_MODEL when one gateway model should serve every role.
For endpoint-aware routing and cost estimates, set
AUTOCONTEXT_PROVIDER_HOSTING to local or remote and, for a local
endpoint, set AUTOCONTEXT_PROVIDER_CAPABILITY to fast, mid_tier, or
frontier. Empty hosting retains conservative transport inference. A
role-specific endpoint uses the corresponding
AUTOCONTEXT_<ROLE>_PROVIDER_HOSTING and
AUTOCONTEXT_<ROLE>_PROVIDER_CAPABILITY declarations instead of the default
endpoint's declarations.
OpenAI-compatible role generation requests schema-constrained output by
default. If that changes output quality for a backend, set
AUTOCONTEXT_CONSTRAINED_OUTPUT=false to omit schemas from every role request
and use the existing Markdown parsers instead. The setting also applies to
roles with dedicated provider overrides.
Build-time teacher trace collection supports native deep_think tool loops on
Anthropic and OpenAI-compatible providers. The collector requires a structured
tool stream by default and keeps it separate from the final answer. Local and
CLI/runtime providers report the capability as unsupported; visible-preamble
fallback is available only through the collector's explicit
require_thinking_stream=False option. For GPT 5.6+ models,
reasoning_effort selects the external numeric prompt budget while native
reasoning is requested off; a compatible gateway that rejects none is clamped
to its lowest advertised level. Thinking payloads may contain sensitive
prompt-derived data and should be redacted before persistence or export.
Common commands
| Command | Purpose |
|---|---|
uv run autoctx solve "..." --iterations 3 |
Generate and run a scenario from a plain-language goal |
uv run autoctx run <scenario> --iterations 3 |
Improve an existing scenario |
uv run autoctx status <run-id> --json / watch <run-id> --ndjson |
Read one run snapshot or stream snapshots |
uv run autoctx show <run-id> --best --json |
Inspect the best generation |
uv run autoctx simulate --description "..." |
Create/replay/compare modeled-world simulations |
uv run autoctx investigate --description "..." |
Run synthetic or iterative investigations |
uv run autoctx list / status <run_id> / show <run_id> |
Inspect runs |
uv run autoctx replay <run_id> --generation 1 |
Replay a generation before accepting knowledge |
uv run autoctx queue add --task-prompt "..." --rubric "..." |
Queue evaluation/improvement work |
uv run autoctx scenario create --family workflow --name support --description "..." |
Create a reusable scenario through a family-specific pipeline |
uv run autoctx serve --host 127.0.0.1 --port 8000 |
Start the local HTTP API |
uv run autoctx worker --poll-interval 5 --concurrency 2 |
Process queued tasks beside the API server |
uv run autoctx serve mcp |
Expose the MCP tool surface |
uv run autoctx export-training-data --scenario <name> --all-runs --output data.jsonl |
Build a training corpus (quarantined scores excluded by default; --include-quarantined to keep them) |
uv run autoctx train --scenario <name> --data data.jsonl --time-budget 300 |
Run the local training hook |
uv run autoctx epoch list [--scenario <name>] |
List evaluator-epoch registry records (candidate/active) |
uv run autoctx epoch approve <scenario> <epoch_id> --charter ambient-charter.yaml |
Approve a candidate evaluator epoch and clear its quarantine |
uv run autoctx hermes inspect --json |
Inspect Hermes Curator state |
Saved custom scenarios under knowledge/_custom_scenarios/ can be rerun and benchmarked by name after their spec.json is persisted.
HTTP, MCP, and agents
uv sync --group dev --extra mcp
uv run autoctx serve mcp
Python runtime-backed run and solve calls append provider prompts/responses to run-scoped runtime-session logs. The same logs are readable through the cockpit HTTP API and MCP tools.
Detailed setup moved out of this README:
- External agents and provider routing: docs/agent-integration.md
- Self-hosted models, end to end: docs/self-hosted-models.md
- Persistent worker trust boundaries: docs/persistent-host.md
- Sandbox/executor notes: docs/sandbox.md
- Extension hooks: docs/extensions.md
- Correctness-first external kernel evolution, including durable autonomous provider campaigns, resume, and fail-closed multi-workload transfer studies: docs/kernel-evolution.md
Contract probes
Contract probes turn observed harness traces into executable checks:
uv run autoctx probes check --suite contract-probes.json
uv run autoctx probes check --suite contract-probes.json --json
uv run autoctx probes extract --trace harness-trace.json --output contract-probes.json
Probe suites are strict JSON: unknown keys fail validation and required observation fields must be present. Pipe stdin with --suite - when another tool generates the suite.
Production traces
Wrap an existing Anthropic/OpenAI client once, then persist emitted traces through a sink:
from anthropic import Anthropic
from autocontext.integrations.anthropic import FileSink, instrument_client
sink = FileSink("./traces/anthropic.jsonl")
client = instrument_client(
Anthropic(),
sink=sink,
app_id="billing-bot",
environment_tag="prod",
)
For lower-level emit APIs, use autocontext.production_traces.build_trace
and write_jsonl. Architecture notes are in
../docs/analytics.md and
../docs/opentelemetry-bridge.md.
Training
uv run autoctx export-training-data \
--scenario support_triage --all-runs \
--output training/support_triage.jsonl
uv run autoctx train \
--scenario support_triage \
--data training/support_triage.jsonl \
--time-budget 300
For MLX/CUDA setup and case studies, use:
Repository layout
autocontext/
├── src/autocontext/ # Python package
├── tests/ # pytest suite
├── docs/ # package-specific docs
├── demo_data/ # small bundled examples
├── migrations/ # SQLite migrations
└── pyproject.toml
Development
uv run ruff check .
uv run mypy src
uv run pytest
Keep this README concise. Add deep reference prose to docs/ or the repo-level
docs index instead.
Metadata
Release files for autocontext 0.17.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| autocontext-0.17.0.tar.gz | 3.4 MB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| autocontext-0.17.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 5.5 MB
Release files / autocontext-0.17.0.tar.gz
| Download URL | autocontext-0.17.0.tar.gz |
|---|---|
| Size | 3.4 MB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5555df967ef5d8dcbae68565c8c03d8c7d22cbef0ba494b8a96ad09691822fc9
|
|
BLAKE2b-256 checksum How to use checksums |
f201e2b23e98373c9051cf49aa52340a1032bf587792fc487658f10bbc5a2b13
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 26, 2026.
Transparency logRelease files / autocontext-0.17.0-py3-none-any.whl
| Download URL | autocontext-0.17.0-py3-none-any.whl |
|---|---|
| Size | 2.1 MB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c29faa916f50fac8aacc95e37862d35836e43d839d791224136377977728189c
|
|
BLAKE2b-256 checksum How to use checksums |
5d4214842a620877895919abe7115f6150575fdab6fb15fdae211be940f34bbf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Aug 26, 2026.
Transparency log