Skip to main content

pxx 2.0 - local-first AI coding agent runtime

CI PyPI Python License: MIT

pxx is an async, event-sourced coding-agent runtime that runs against your own inference (Ollama, vLLM, or any OpenAI-compatible endpoint) - no cloud dependency, no telemetry, no API keys required. It pairs a native tool-calling agent loop with persistent cross-session memory, deterministic safety gates, and MCP interop, and can delegate to aider as an optional edit engine.

New here? → Hands-on tutorial Every claim this project makes is evidence-gated - dated records, reproducible procedures, and explicit non-claims live in docs/RECEIPTS.md.

pxx 2.0 is a ground-up rewrite of pxx 1.x. The 1.x control-plane semantics (fail-closed gates, scope limits, bounded loops, audit) are preserved; the execution layer (os.execv into aider, argv scanning, sidecar services) is replaced by an async runtime where pxx owns the agent loop. See DESIGN.md for the architecture contract and docs/MIGRATION.md for 1.x → 2.0 changes.

Why pxx

  • Offline-capable: all inference stays on your machine or your LAN.
  • pxx owns the runtime: every model/tool event flows through pxx's event bus. Backends cannot bypass policy - scope, permissions, budgets, and hooks are enforced by the host, never by the model.
  • Persistent memory: observations from previous sessions (files changed, tool outcomes, your own remember notes) are stored in a local SQLite database with hybrid BM25 + vector search and injected deterministically at session start.
  • Fail-closed safety: read-only by default. Writes require edit/run mode, stay inside a canonicalized scope (symlinks resolved), shell commands are gated in every write-capable mode (a hook, sandbox, or explicit opt-in - even under unattended run), and every run ends with a machine-readable terminal code in a hash-chained audit log.
  • Interop: consumes MCP servers as tools, and exposes its own memory as an MCP server for other agents (Claude Code, goose, opencode, …).

Where pxx fits

This is a fit guide, not a sales pitch - pxx isn't competing on raw capability. It exists because one distinction is architectural, not marketing: a hosted SaaS coding agent cannot be self-hosted. Your source and your inference run on the vendor's infrastructure by construction. Where that's permitted you have many good options. Where it isn't - regulated, air-gapped, IP-sensitive, sovereignty-required - that entire category is off the table before capabilities are even discussed, because those environments demand a higher standard than "the vendor is trustworthy": the code and the inference must stay inside controls the operator owns, and every action must be auditable.

One honest caveat up front: running locally and air-gapped is not unique to pxx. Local agents such as aider and opencode run entirely on your own hardware too, so locality is the floor here, not the differentiator. What pxx adds sits above the model, in the points below, and it is defense in depth inside the isolation you build, never a substitute for it.

pxx is built to that sovereign standard:

  • Self-hosted, always (the floor, not the pitch). Code and inference stay on hardware you own: local, on-prem, or fully air-gapped; your models, your network, your key management, your audit boundary. No vendor in the loop, no telemetry, no data egress. Other local agents do this too; the difference is the two points below.
  • Host-enforced, not model-trusted. Scope, permissions, budgets, and hooks are enforced by the host process, so a jailbroken or confused model still can't write outside the files you allowed. This is an action-level boundary and defense in depth inside your own isolation (network, sandbox, air-gap), not a replacement for it. And because the gate is our code, the honest control is not to trust it: read it (MIT) and verify the hash-chained, tamper-evident audit of what it actually did (pxx audit verify).
  • Open, inspectable, and yours to change. MIT, under the same community model that made Linux and the tools you already rely on - not "source-available," not open-core with the real logic hidden behind a service. You don't just get to read every gate, tool, and decision (no hidden functions, no telemetry) - you have the full open-source freedoms to run, study, modify, and redistribute it. For a regulated shop that's the deepest control there is: fork it, harden the enforcement to your compliance regime, add your own gates, and own that fork - no vendor lock-in, and no dependence on a vendor's roadmap or continued existence. (MIT goes even further than Linux's copyleft: adapt it freely, even inside proprietary internal tooling, with no strings.) SaaS grants you none of these rights; pxx grants all of them.
  • Evidence-gated - a step beyond typical open source. Open source hands you the source; pxx also keeps the receipts. Every capability claim maps to a dated, reproducible record in docs/RECEIPTS.md with an explicit statement of what is not claimed - evidence, not just availability.

The through-line is verify, don't trust: no vendor to trust (there isn't one), no binary to trust (read it), no claims to trust (check the receipts). You should not have to trust software you did not write, so this is built to be read and its record verified.

Use pxx when self-hosting, sovereignty, and provable audit are requirements - when the work simply cannot leave your perimeter and "trust the vendor" isn't a sufficient control.

pxx is deliberately not trying to be the tool for every environment. If your environment permits SaaS and you want maximum capability on hard or greenfield work, a hosted frontier agent will do more; if you want turn-by-turn pair editing, a local assistant such as aider (which pxx can delegate to) may fit better. pxx has no browser yet and leans on smaller local models - it trades breadth for sovereignty and auditability, on purpose. It's early (2.x) and solo-maintained; the receipts, not the prose, are the source of truth.

Install

pip install pxx-orchestrator          # the command is `pxx`
# optional extras:
pip install "pxx-orchestrator[aider]"   # aider delegation backend (Python < 3.13)
pip install "pxx-orchestrator[server]"  # headless HTTP API (pxx serve)

Prerequisites: Python 3.11+, and a reachable model endpoint - Ollama by default (ollama pull qwen2.5-coder:7b).

Backends

pxx runs tasks on one of two execution backends, selected per run:

  • native (the tool-calling agent loop this README describes): pxx drives the model directly through its own tool surface and gates.
  • aider (delegation): edits are handed to aider as the edit engine, still under pxx's scope/budget/hook gates.

Default selection: aider when the aider binary is on PATH, else native (run/loop always default to native). Force one with --backend native|aider on any run verb.

Tool calling. The native backend needs an endpoint that accepts tool calls (tools in the chat-completions request). Ollama supports tool calling out of the box. vLLM must be launched with --enable-auto-tool-choice --tool-call-parser <parser> - without those flags every native round fails with HTTP 400 ("auto" tool choice requires …). pxx doctor probes for this under a realistic agent context and reports if a model accepts tools but answers in prose (some small models tool-call on a toy probe yet degrade under a real loop prompt on constrained hardware). ask/edit (and run with --backend aider) can sidestep via the aider backend; pxx loop is native-only and cannot.

Quick start

pxx doctor                          # check your setup

pxx ask -m "Explain main.py"        # read-only (default): no writes, no shell
pxx edit -m "Add error handling to main.py"   # writes allowed, in scope
pxx edit --commit -m "…"            # + commit the change on COMPLETED (opt-in)
pxx run  -m "Add tests for utils.py"          # unattended, budget-capped
pxx loop -m "Fix the failing tests" --scope src  # bounded edit→test→review loop
pxx chat                            # interactive session

Edits land uncommitted by default so you can review them first - pass --commit (or PXX_AUTO_COMMIT=1 / auto_commit = true) to have a COMPLETED run commit its work. The pxx-pre/<ts> safety tag always points at the pre-session HEAD, so undo is git reset --hard <tag> either way.

Permission modes: ask (read-only) → plan (plan only) → edit (writes in scope, shell via hooks) → auto (unattended, budgets enforced). Every run is bounded: max rounds/tokens/cost/wall-clock/diff-lines, all configurable.

New to pxx? The hands-on tutorial takes you from install to building (and safely undoing) a small tool in ~25 minutes - the fastest way to get the mental model: read-only by default, scoped edits, and the safety tag that nets your work.

Memory

pxx memory add "we use ruff, not black" --tags conventions
pxx memory search "linting"
pxx memory list

Memory is hybrid-retrieved (FTS5 BM25 0.4 + embedding cosine 0.6). Embeddings come from a local Ollama model when reachable, else a deterministic hash embedder - search always works offline. TTL'd observations archive to JSONL monthly. Memory is context, never policy.

Memory is project-scoped by working directory: search/list see only the current directory's project (its directory name) - run them from the directory the memory was added in. Keyword search matches whole tokens exactly (no stemming): searching round will not match rounding.

Expose it to other agents over MCP:

pxx mcp            # stdio MCP server: memory_search / memory_add / memory_list

Configuration

Layered, highest precedence wins: CLI flags → PXX_* env → ./pxx.toml (or .pxx/config.toml) → ~/.config/pxx/config.toml → defaults. Unknown keys are rejected (fail-closed, no silent typos). Example pxx.toml:

model = "qwen2.5-coder:14b"
provider = "ollama"
permission = "edit"
scope = ["src", "tests"]
test_command = "pytest -q"

[budgets]
max_rounds = 20
max_cost_usd = 2.0

[[fallback_models]]
model = "served-model"
provider = "vllm"
base_url = "http://gpu-box:8000"

[[hooks]]
event = "PreToolUse"
command = "/usr/local/bin/my-gate"   # exit 0 allow / 2 deny - deterministic

[[mcp_servers]]
name = "filesystem"
command = ["npx", "-y", "@modelcontextprotocol/server-filesystem", "."]

1.x PXX_OLLAMA_BASE / PXX_OLLAMA_MODEL env vars and ~/.config/pxx/env still work.

Headless API

pxx serve --port 8400     # FastAPI: POST /v1/sessions, SSE event stream,
                          # cancel, memory proxy. Loopback-only by default.

Operator commands

Beyond the everyday verbs (ask/edit/plan/run/loop/chat):

Safety & release

pxx check [--all-files]   # secret/PII scan - staged files, or all tracked files
pxx upgrade               # upgrade the pxx install in place
pxx review [--staged|--since SHA]  # read-only review of the current diff (exit 2 on REVISE)
pxx doctor                # diagnose setup (endpoints, backend, memory, config)
pxx audit verify <path>   # verify a hash-chained audit log

Run evidence (every run is recorded with an immutable agent manifest)

pxx runs list|show|export         # recorded runs, per-agent projections
pxx runs resume <run-id>          # resume a run from its checkpoint
pxx agents list|show              # agent versions + success rates (drift quarantine)
pxx verify [run-id]               # verification packet for a run (gates fired)
pxx metrics summary|failures|memory-impact|export|compare

Evaluation & improvement (the self-improvement platform)

pxx eval run|self-check|report [--partition held-out]
pxx calibrate                     # reviewer calibration (recall/fp/agreement)
pxx improve analyze|clusters|proposals|cycle|status|daemon|pause|resume
pxx improve triage list|qualify|reject   # durable human verdicts on proposals
pxx improve evaluate-candidate <id>   # held-out, both arms
pxx improve readiness|auto-promote|principles
pxx propose                       # create a constrained improvement candidate
pxx compare <baseline> <candidate>    # promotion verdict (held-out, multi-metric)
pxx promote <candidate-id>        # human-gated promotion (needs a real scorecard)
pxx agent activate|rollback|history|channels|canary
pxx goal -m "<goal>"              # goal -> task DAG -> isolated per-node loops

Legibility (docs/workflow contracts)

pxx workflow validate             # validate this repo's WORKFLOW.md
pxx context audit                 # docs present + trust mirrors in sync
pxx docs check                    # every documented verb exists

Every verb self-describes: append --help (e.g. pxx check --help).

Safety model (short version)

  • Edit-capable sessions (edit/run/loop/goal, in a git repo) tie a safety net before anything can write: uncommitted work is stashed (--include-untracked, message carries the run id) and HEAD is tagged pxx-pre/<ts>. Restore with git reset --hard <tag> + git stash pop - pop is your move, never pxx's. Disable with safety_net = false.
  • Paths are canonicalized with symlinks resolved before any gate decision - model output never defines the trust boundary.
  • Hooks are deterministic gates (like Claude Code's PreToolUse): they cannot be overridden by model judgment.
  • The audit log (~/.local/state/pxx/audit/YYYY-MM-DD.jsonl) is hash-chained and metadata-only - no prompts, file contents, or secrets. Verify with pxx audit verify <path>.
  • Bounded loops stop on: round cap, diff cap, budget, scope violation, non-monotonic test progress (NO_TEST_PROGRESS), a detected oscillation (LOOP_DETECTED), or a blocking review verdict.

Upgrading

With 2.0 on PyPI:

  • uv tool: uv tool upgrade pxx-orchestrator
  • pipx: pipx upgrade pxx-orchestrator
  • pip: pip install -U pxx-orchestrator
  • from source: git pull && uv sync --extra dev --extra server
  • in-place: pxx upgrade - upgrades the pxx install (detects uv tool / pipx / pip automatically)

Settings, memory, and audit state carry forward - 2.0 migrates them on first run (see docs/MIGRATION.md).

Development

git clone https://github.com/cdnwetzel/pxx && cd pxx   # the 2.0 tree (branch v2)
uv sync --extra dev --extra server
uv run pytest          # 870+ tests, no network/Ollama/aider required
uv run ruff check

2.0 lives on cdnwetzel/pxx (this repo); the 1.x line continues on its v1.x branch. The public history is a curated series - the full development history stays private.

Pull requests are reviewed automatically by CodeRabbit (config in .coderabbit.yaml) in addition to human review.

License

MIT - see LICENSE.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pxx_orchestrator-2.4.1.tar.gz (385.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pxx_orchestrator-2.4.1-py3-none-any.whl (262.9 kB view details)

Uploaded Python 3

File details

Details for the file pxx_orchestrator-2.4.1.tar.gz.

File metadata

  • Download URL: pxx_orchestrator-2.4.1.tar.gz
  • Upload date:
  • Size: 385.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for pxx_orchestrator-2.4.1.tar.gz
Algorithm Hash digest
SHA256 1e916c102fc6be709d0b4f845b931b4dc5ed3238f5490d2368bcca8e7324add4
MD5 aadc8366d244c98b63ec564c96560cfd
BLAKE2b-256 e1b0025972054fb93b65610c020d63a65b0e86650701fdbc325dde8c009c0bc0

See more details on using hashes here.

Provenance

The following attestation bundles were made for pxx_orchestrator-2.4.1.tar.gz:

Publisher: release.yml on cdnwetzel/pxx

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pxx_orchestrator-2.4.1-py3-none-any.whl.

File metadata

File hashes

Hashes for pxx_orchestrator-2.4.1-py3-none-any.whl
Algorithm Hash digest
SHA256 9311dc9419739be99b06da4416d111a99aeab2f288df9090a8be5dddb924cef5
MD5 b599cabab04c35f4bbbb01cde091353c
BLAKE2b-256 0afa2da328f051e54a036c6332f4eab6261682810787cb62da504e42e6d84d2e

See more details on using hashes here.

Provenance

The following attestation bundles were made for pxx_orchestrator-2.4.1-py3-none-any.whl:

Publisher: release.yml on cdnwetzel/pxx

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page