sovereign-agent
The eight architectural decisions every serious agent system converges on — implemented as a library you can use, and a tutorial you can read.
Debug by cat. Crash-recover by ls. Teach by reading the same code that runs in production.
Repository identity. The canonical home is
zeroemployeeorg/sovereign-agent. Oldersovereignagents/...URLs still resolve by GitHub redirect but are not canonical; new links should usezeroemployeeorg.
pip install sovereign-agent
Reusable actions are ZeoCore capabilities (@capability). Session control
(complete_task, handoff_to_structured) is a Sovereign runtime command.
make_session_callable_surface merges both for the model. @register_tool
still works during the compatibility window and is deprecated.
from pydantic import BaseModel
from zeo_core.contracts import CapabilityResult, EffectKind
from zeo_core.tools import ToolContext, capability
from sovereign_agent import run_task
class WeatherQuery(BaseModel):
city: str
@capability(
id="demo.weather.get@1.0.0",
description="Look up current weather for a city.",
effects={EffectKind.READ},
)
def get_weather(request: WeatherQuery, ctx: ToolContext) -> CapabilityResult:
return CapabilityResult.ok(
data={"city": request.city, "temperature": 18, "condition": "rainy"},
msg="ok",
)
result = run_task("What's the weather in Edinburgh?")
print(result.summary)
# → "Weather in Edinburgh: 18°C, rainy."
The agent ran a planner, called into the callable surface, wrote a trace, and
saved every artifact to sessions/sess_<id>/. Inspect it with
ls -R sessions/sess_<id>. That's not a metaphor — it's how you debug this
system.
run_task still accepts deprecated @register_tool helpers. New reusable
actions should be ZeoCore capabilities. See v0.5 Unit 4
for session filesystem capabilities (read_file / write_file / list_files)
and runtime commands.
What sovereign-agent is, exactly
Two things in one repository:
A production library you pip install and use to build agents.
A build-from-scratch curriculum that reconstructs the library in five chapters with tests, so you learn by implementing.
The chapters rebuild the v0.2 substrate and re-export the production primitives
they teach. A CI check (tools/verify_chapter_drift.py) prevents those mapped
primitives from drifting. v0.3 features are taught in numbered unit documents;
the chapters intentionally do not claim whole-package parity. See the
teaching-surface decision.
sovereign-agent is not trying to be the next Claude Code or OpenHands. It's the thing you read to understand why they and every other production agent system converged on the same eight architectural decisions.
The eight decisions (this is the product)
Every production agent system I've read the internals of — Claude Code, OpenHands, Aider, SWE-agent, Devin, Cognition's public work — has independently arrived at the same eight decisions. sovereign-agent is those decisions, made explicit, with code you can run.
- Sessions are directories.
sessions/sess_<12hex>/contains everything — memory, IPC, state, logs, artifacts. No database. No shared tables. - State is forward-only. A session never transitions backwards. Retries are new sessions, linked to the old one.
- Tickets for every operation. Append-only audit trail. Tickets are to your agent what commits are to a git repo.
- Manifests verify by SHA-256. Detect accidental edits, disk corruption, and tampering. Not cryptographic security — this is about catching the mistakes you actually make.
- Atomic rename for IPC. Two halves of the agent communicate by writing files. No brokers, no Kafka, no Redis. POSIX's
rename()is the only IPC primitive you need until proven otherwise. - Lock at the session level. Not finer (deadlocks), not coarser (unscalable). One serialization boundary per session gives you multi-file consistency for free.
- Parse JSON defensively. The LLM is lying. Write a parser that handles what the model actually produces, not what the prompt instructed it to produce.
- Register tools explicitly. Prompts are advisory; the registry is physics. When the model persistently reaches for a tool you don't want used, remove the tool — don't write the third negative instruction.
Each decision removes a class of bugs rather than handling them. That's why the decisions age well.
Full walk-through in docs/architecture.md. Note that
the architecture doc enumerates the same system from an implementation angle and
so numbers its decisions differently; the two lists overlap but are not a 1:1
mapping, and neither is a renumbering of the other.
The ninth thing nobody converges on, but should: dataflow integrity
The framework can guarantee that tools were called, tickets were written, manifests verified, state advanced. That is necessary but not sufficient. You also need to verify the LLM used its tool outputs.
Every scenario in this repo ships with a dataflow integrity audit. The research-assistant scenario checks that every arXiv ID cited in the report was actually returned by web_lookup — not fabricated from training data. The code-reviewer scenario checks that the review references findings the analyzer actually returned.
This is not a hypothetical concern:
One morning I had a framework with 148 passing tests and three clean scenarios. I ran the code reviewer against a real LLM for the first time. It produced a perfectly-formatted review of code that did not exist — named
add,multiply,divide. The framework reported ✓ success, ✓ manifest, ✓ complete. Every structural guarantee held. The output was pure fiction.
(148 was the test count at v0.1.0; see Status for the current number.)
"It ran" is not "it worked." The library gives you the first; your scenario has to verify the second. Every example in this repo demonstrates the pattern.
Install
pip install sovereign-agent # core
pip install "sovereign-agent[all]" # + optional extras (voice, observability)
Development tooling lives in the PEP 735 dev dependency group, not an extra —
so it is uv sync --group dev (or pip install -e . --group dev), not
pip install "sovereign-agent[dev]".
Requires Python 3.13+. zeocore>=0.5,<0.6 is a required dependency.
Docker is not available. There is a docker extra, but it only installs the
Docker SDK; DockerWorker is a stub that raises NotImplementedError on
run_session(). The supported isolated backend is worker_backend='subprocess'
(Landlock on Linux ≥ 5.13, sandbox-exec on macOS). See
v0.3 non-goals.
The two halves
The decision-8 principle — prompts are advisory, registries are physics — generalizes. Not every constraint is a tool registry. Some constraints are rules. sovereign-agent ships a loop half (ReAct-style LLM reasoning) and a structured half (deterministic Python) that communicate by atomic-rename IPC.
from sovereign_agent import LoopHalf, StructuredHalf, Rule
loop = LoopHalf(planner=planner, executor=executor)
structured = StructuredHalf(rules=[
Rule(name="commit_under_cap",
condition=lambda d: d["deposit"] <= 300,
action=commit_booking),
Rule(name="escalate_over_cap",
condition=lambda d: d["deposit"] > 300,
escalate_if=lambda d: True),
])
The loop decides what to try. The structured half decides what's allowed. No amount of prompt engineering can bypass a rule — the business constraint lives in Python where it belongs.
This isn't new. It's how every financial, regulatory, or high-stakes agent system eventually ends up structured. The name varies ("policy layer," "guardrails," "rules engine") but the pattern is the same. sovereign-agent makes it the default.
What shipped in v0.2.0 (the published release)
Five capabilities that came out of running sovereign-agent against real LLMs and hitting real failures:
Parallel tool dispatch. Tools marked parallel_safe=True run concurrently. Writes and handoffs serialize automatically. Five 0.3s calls finish in 0.31s instead of 1.5s.
Worker isolation without Docker. Linux ≥5.13 gets kernel-level Landlock. macOS gets sandbox-exec. No container runtime needed. Tool compromise can't escape the session directory.
Session resume. Resume any terminal session as a new child. Parent context auto-prepends to the child's SESSION.md. Chains preserve ancestry. Forward-only state is preserved — resumes are new sessions, never edits of the old.
Verifier protocol. Rule conditions accept callables, scikit-learn classifiers, or LLM judges. Same audit trail, different backend. Decision 6 generalized.
Human-in-the-loop approval. Tools return requires_human_approval=True. The executor exits cleanly, writes the request to disk, and resumes when a human decides — seconds, hours, or days later. Nothing in memory.
Each is ~200-400 lines of code with tests. See examples/ for end-to-end scenarios that use each one.
What shipped in v0.3.0
The v0.3 package adds:
- Channel adapters —
ChannelAdapterprotocol plus a CLI adapter and an inbound router. OnlyCHANNEL_REGISTRYis exported in__all__; the adapter types are importable but not yet part of the semver surface. - Plugin registries — a generic
Pluginprotocol andRegistry[T]. - Worker backend integration — orchestrator dispatch routes through
WorkerBackendviamake_worker_backend(). - Native CLI providers — Codex CLI and Claude Code are probed, capability
gated, parsed into normalized events, and executed through
WorkerBackendwithout Sandcastle. - Liveness monitor — stalled-session detection and heartbeat
(
LivenessMonitor, importable but not in__all__).
These additions are part of the v0.3 public contract where listed in
sovereign_agent.__all__. See
docs/v0.3-non-goals.md for what v0.3 will not do,
and docs/branch-consolidation-2026-08-22.md
for how these branches landed.
The architecture in one picture
┌─────────────────────────────────────────────────────────────┐
│ LOOP HALF Planner → Executor → Tools │
│ (reasoning) ReAct-style, free-form LLM │
│ │
│ ▼ handoff via ipc/handoff_to_structured.json │
│ │
│ STRUCTURED HALF Deterministic rules, classifiers, │
│ (constraints) LLM judges — whatever you configure. │
│ Binds what's allowed. │
└─────────────────────────────────────────────────────────────┘
│
▼
┌─────────────────────────────────────────────────────────────┐
│ sessions/sess_a382a2149fc1/ │
│ ├── session.json # state machine, forward-only │
│ ├── SESSION.md # the system prompt │
│ ├── workspace/ # tool outputs, agent artifacts │
│ ├── memory/ # persistent facts across runs │
│ ├── ipc/ # atomic-rename message passing │
│ ├── tickets/ # every operation recorded │
│ └── logs/trace.jsonl # every event, every tool call │
└─────────────────────────────────────────────────────────────┘
Debugging by cat
This is the visceral version of "sessions are directories." Last week I debugged five separate failures across three scenarios. Every one was diagnosed from three lines of JSONL:
$ cat sessions/sess_5ab10359c72a/logs/trace.jsonl
{"event_type": "executor.tool_called", "payload":
{"tool": "read_file", "arguments": {"path": "workspace/source.py"},
"success": false, "summary": "file not found"}}
{"event_type": "executor.tool_called", "payload":
{"tool": "list_files", "arguments": {"path": "workspace"},
"success": true, "summary": "0 entries"}}
{"event_type": "executor.tool_called", "payload":
{"tool": "complete_task", "arguments":
{"result": {"status": "failed",
"reason": "No source file found in workspace"}}}}
The LLM never tried the analyzer tool. It looked for files that weren't there, gave up, marked the session complete. Three lines and I knew the fix (decision 8: remove the directory-listing tool from the registry; the model reaches for it reflexively).
No SELECT queries. No vendor viewer. No SDK. cat.
Three surfaces, one codebase
sovereign-agent/
├── src/sovereign_agent/ # the library you pip install (src layout)
├── chapters/ # 5 tutorial chapters (minitorch-style, fill in the TODOs)
├── examples/ # 8 reference scenarios (research, code review, HITL, etc.)
├── docs/ # architecture, API stability, deployment
└── tests/ # 500 collected tests — library + chapters + examples
- If you want to ship an agent today → read
src/sovereign_agent/and pick scenarios fromexamples/ - If you want to understand how it works → read the chapters. Each one rebuilds a piece of the library. Your tests pass when you're done.
- If you want the architecture argument →
docs/architecture.md
The chapter/library drift check (tools/verify_chapter_drift.py) runs in CI for
the v0.2 primitives those chapters teach. The explicit v0.3 teaching scope is
documented in docs/teaching-surface.md.
Lineage
sovereign-agent stands on three lineages:
Teaching artifacts with real libraries — the pattern:
fastai(Jeremy Howard) — the library + course pattern. sovereign-agent's biggest debt.minitorch(Sasha Rush) — rebuild-the-framework pedagogy. Chapters work like this.LLMs-from-scratch(Sebastian Raschka) — book + code as the same artifact. Reading order matters.nanoGPT(Andrej Karpathy) — small, readable, no magic.
Production agent systems that converged on the same architecture:
- NanoClaw (Gavriel Cohen) — TypeScript reference implementation. Read
src/group-queue.ts; two hours of reading and you'll see where sovereign-agent's patterns come from. - Claude Code — session-as-directory, sub-agent isolation.
- OpenHands — closest OSS cousin architecturally.
- Aider — per-repo state in
.aider/, same pattern simpler scope. - SWE-agent, Devin — per-task sandboxes.
Papers that shaped the module design:
- ReAct — the executor loop.
- Reflexion — memory patterns.
- MemGPT — hierarchical memory.
- Voyager — procedural memory; file-based skills.
- SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering — "give the agent a typewriter, not a console."
See CREDITS.md for the full list.
Quick-start with a real LLM
# Configure your LLM endpoint (Nebius, OpenAI, or any OpenAI-compatible provider)
cp .env.example .env
# edit: set NEBIUS_KEY=sk-... and optionally swap models
# Verify everything is wired up — Python, uv, .env, models, imports, CI
make doctor
# Run a real scenario end-to-end
make example-research-real
A real-LLM run prints every step, every tool call, and finishes with a dataflow integrity audit. Illustrative shape of example-research-real output (model names and byte counts depend on your .env and the run):
▶ research-assistant (real LLM)
planner: MiniMaxAI/MiniMax-M2.5
executor: Qwen/Qwen3-235B-A22B-Instruct-2507
✓ plan produced: 1 subgoal, loop half
✓ web_lookup("retrieval augmented generation") → 2 results
✓ write_file(report.md) — 287 bytes
✓ complete_task
=== Dataflow integrity audit ===
web_lookup calls: 1, successful hits: 2, unique papers returned: 2
✓ all 2 arXiv ID(s) came from web_lookup
titles from web_lookup referenced in report: 2/2
If the model fabricates a paper the audit catches it and flags ✗ with the fabricated ID. This pattern is what you want to copy into your own scenarios.
Launch-checklist-as-Makefile
The Makefile is also the documentation for the release workflow:
make doctor # tabular status — Python, uv, .env, deps, imports, CI
make preflight # lint + drift + pytest collection + demo importability
make test # full suite (500 collected: 497 pass, 3 opt-in/platform skips)
make ci-real-estimate # cost preview for a full ci-real run (no API calls)
make ci-real # run every -real scenario against a live LLM
make pre-publish # audit for secrets, PII, forbidden files before public push
make ready-to-ship # deterministic CI + audit + wheel/sdist clean-install proof
make help groups every target by category. make doctor output is tabular; paste it into an issue and a maintainer has full diagnostic context.
Where things live
sovereign-agent is a well-behaved Python library. It never writes to your CWD when you import sovereign_agent. Different entry points write to different places by design:
| Entry point | Artifacts go to | Why |
|---|---|---|
sovereign-agent run <task> (production) |
./sessions/ (your CWD) |
Your deployment, your call |
python -m chapters.<n>.demo |
$XDG_DATA_HOME/sovereign-agent/demos/ |
Persists for inspection; outside your git tree |
python -m examples.<n>.run (offline) |
tempdir, auto-cleans | Deterministic dev runs |
python -m examples.<n>.run --real |
$XDG_DATA_HOME/sovereign-agent/examples/ |
Real-LLM runs burn tokens; keep artifacts |
Override via SOVEREIGN_AGENT_DATA_DIR=<path>.
Status
v0.5.1. The package exports 161 symbols in sovereign_agent.__all__
(the 152-symbol v0.4 surface plus nine capability-adapter names). Python 3.13
is the floor; zeocore>=0.5,<0.6 is required. Git tag v0.5.0 is the
capability-migration snapshot and is not moved. A git tag is not a public
release — announce v0.5 as published only after this version is on PyPI.
See docs/API.md, docs/roadmap.md, and
docs/v0.4-operator-guide.md.
- ✅ Framework: sessions, tickets, IPC, parallelism, isolation, resume, verifiers, HITL
- ✅ 522 tests collected — 519 pass, 3 opt-in/platform skips
- ✅ 8 reference scenarios, all with dataflow integrity checks
- ✅ 5 tutorial chapters, drift-checked against production code in CI
- 🚧 Voice pipeline, observability backends (Evidently/OTel) — skeletons, not implementations
- ✅ Channels, native Codex/Claude CLI providers, plugin registries, worker lifecycle, governed repository execution, durable seat registry and relay
- ❌ Docker worker backend — stub only;
DockerWorker.run_session()raisesNotImplementedError - ❌ Vector-DB memory backends — not started, and a v0.3 non-goal
Nothing in this repository is load-bearing for a production deployment you have not read end to end yourself. The alpha label is not modesty.
What sovereign-agent is not
It's not trying to replace Claude Code for daily coding or LangGraph for orchestration-heavy workflows. It's not a framework I'm trying to grow into the next big thing. It's not abandoned; it's not vibe-coded; it's not a thin wrapper over LangChain.
What it is: a substrate for teaching the eight architectural decisions, plus a library that implements them cleanly enough that you can use it for real work. fastai for agents.
If you want an agent you can own, audit, reproduce, teach, and — crucially — understand at the bottom of the stack, this is probably the smallest codebase in the world that gives you all five.
For what this series will not attempt through v0.7, see
docs/non-goals.md and
docs/v0.3-non-goals.md. Those are commitments, not
moods. The public sequence is docs/roadmap.md.
This is a work repo, not a corpus
This repository holds code. It does not hold the planning and reporting chain for the work done on it. There is no Statement of Work, no SOW directory, and no filing chain in this tree, and there never should be:
work_repo— this repository. Code that genuinely collides, so changes land on branches and get merged through pull requests. The public sequence isdocs/roadmap.md; binding authorization stays in the corpus.sow_repo— a separate corpus repository, elsewhere. Where work is scoped and reported.
They are different fields for a reason. If you came looking for a root
SOW.md, it is not missing — it was never supposed to be here. Read the code,
the architecture doc, and the
CHANGELOG; those are the artifacts this repository owes you.
Learn more
- 📖
docs/architecture.md— the architectural decisions in detail - 🧭
docs/roadmap.md— post-v0.5 sequence (v0.6 correctness, v0.7 fleet) - 🚫
docs/non-goals.md— durable refusals through v0.7 - 🧭
chapters/— rebuild the v0.2 substrate yourself in 5 runnable chapters - 🧪
examples/— 8 reference scenarios, each with a dataflow integrity audit - 📋
docs/API.md— semver contract for the public symbols - 🚫
docs/v0.3-non-goals.md— historical v0.3 refusals that still hold - 🌿
docs/branch-consolidation-2026-08-22.md— how the pre-v0.3 branches landed - 📝
CHANGELOG.md— what shipped and when
Contributing
Pull requests, issues, and architectural criticism are welcome. There is no
CONTRIBUTING.md yet; the contribution contract is make verify — apply, then
verify, then commit.
git clone https://github.com/zeroemployeeorg/sovereign-agent
cd sovereign-agent
make first-run # install, preflight, sanity check
make test # full deterministic suite
make demo-ch5 # see a working agent end-to-end
Credits
NanoClaw (Gavriel Cohen) is the TypeScript reference implementation whose patterns — session-as-directory, group-queue serialization, tickets, filesystem IPC — sovereign-agent ports and extends in Python.
sovereign-agent is also built by reading and comparing production agent systems (Claude Code, OpenHands, Aider) and the foundational papers (ReAct, Reflexion, SWE-agent, Voyager, MemGPT). The pedagogical format is modelled on nanoGPT (Karpathy), minitorch (Rush), LLMs-from-scratch (Raschka), and fastai (Howard).
See CREDITS.md for full attributions.
License
Apache 2.0. Use commercially, modify, fork — just keep the notice.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sovereign_agent-0.5.1.tar.gz.
File metadata
- Download URL: sovereign_agent-0.5.1.tar.gz
- Upload date:
- Size: 275.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
994949bfce12522c9cbb12756d301f559a6dde8510fff89aa0a0a23fe3b37659
|
|
| MD5 |
2909bd145864d9d8b958a2dac82f385d
|
|
| BLAKE2b-256 |
856b6a5b2e88fd8ac8165495ed0beab4145b325fe40cc10f7d32ac8e7166ef2f
|
Provenance
The following attestation bundles were made for sovereign_agent-0.5.1.tar.gz:
Publisher:
publish.yml on zeroemployeeorg/sovereign-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sovereign_agent-0.5.1.tar.gz -
Subject digest:
994949bfce12522c9cbb12756d301f559a6dde8510fff89aa0a0a23fe3b37659 - Sigstore transparency entry: 2573041285
- Sigstore integration time:
-
Permalink:
zeroemployeeorg/sovereign-agent@221ece545c7af3cfea9b57f0a0b126f1d6a1c10c -
Branch / Tag:
refs/tags/v0.5.1 - Owner: https://github.com/zeroemployeeorg
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@221ece545c7af3cfea9b57f0a0b126f1d6a1c10c -
Trigger Event:
push
-
Statement type:
File details
Details for the file sovereign_agent-0.5.1-py3-none-any.whl.
File metadata
- Download URL: sovereign_agent-0.5.1-py3-none-any.whl
- Upload date:
- Size: 309.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e8a8cbef1d9dfa42de77424347f92a7bd93a9fbcad72fd2abbd010a5f0fff85f
|
|
| MD5 |
2c07911a913d56194da24f4513590611
|
|
| BLAKE2b-256 |
381c5f55c7c18a9843b3118b8ced2ed781dac4c2cf49088cff9acb5a2359679c
|
Provenance
The following attestation bundles were made for sovereign_agent-0.5.1-py3-none-any.whl:
Publisher:
publish.yml on zeroemployeeorg/sovereign-agent
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sovereign_agent-0.5.1-py3-none-any.whl -
Subject digest:
e8a8cbef1d9dfa42de77424347f92a7bd93a9fbcad72fd2abbd010a5f0fff85f - Sigstore transparency entry: 2573041840
- Sigstore integration time:
-
Permalink:
zeroemployeeorg/sovereign-agent@221ece545c7af3cfea9b57f0a0b126f1d6a1c10c -
Branch / Tag:
refs/tags/v0.5.1 - Owner: https://github.com/zeroemployeeorg
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@221ece545c7af3cfea9b57f0a0b126f1d6a1c10c -
Trigger Event:
push
-
Statement type: