Praetor
A security-first, zero-dependency agent runtime with an execution-based evaluation harness.
Praetor is named after the Roman magistrate whose job was to hold the line while everyone else ran wild. This library holds the line between an AI agent and the systems it touches.
Why Praetor exists
Most agent demos show a model calling APIs. Production teams care about the harder skills: safe tool execution, reliability under failure, evaluation you can defend with numbers, and traces you can actually read. Praetor is a small, fully auditable codebase that demonstrates those skills end to end:
- Default-deny security: a tool call runs only if it is explicitly allowed, schema-valid, integrity-verified, and within budget. Everything else is rejected and audited.
- Execution-based evaluation: a benchmark task passes only when executable verification (reading real files, checking real state, running real code) says so. Never an LLM's opinion.
- Observability: every run exports an OpenTelemetry GenAI-aligned trace (invoke_agent and execute_tool spans with gen_ai.* attributes), ready for any OTel-compatible backend.
- Zero runtime dependencies: the whole runtime is Python standard library, so there is no supply-chain surface to poison.
- Full auditability: every step is recorded with its arguments, result, and duration; token usage is aggregated per run.
- Graceful degradation: tool failures become observations for the model, not crashes; policy violations abort loudly instead of silently.
Architecture
flowchart LR
T[Task prompt] --> A[Agent loop]
P[LLM provider<br/>Ollama / OpenRouter / scripted] --> A
A --> V[Strict schema validation<br/>default-deny]
V --> PL[Policy engine<br/>allowlist and budgets]
PL --> G[Integrity guard<br/>tool spec fingerprints]
G --> TOOLS[Tools]
TOOLS --> SB[Sandbox and workspace confinement]
SB --> OBS[Sanitized observation<br/>untrusted-data envelope]
OBS --> A
A --> R[Audited AgentResult]
R --> E[Eval harness<br/>execution-based verification]
R --> O[Observability<br/>OTel GenAI-aligned spans]
Installation
From PyPI:
pip install praetor-agent
Or from source:
git clone https://github.com/Venomous-101/Praetor.git
cd Praetor
pip install -e .
Requires Python 3.10 or newer. There are no runtime dependencies to install or review.
Connecting a model is your choice, by design:
- Local, private: Ollama on the default port — no cloud API keys anywhere.
- Cloud, including free endpoints: OpenRouter with a key from
PRAETOR_OPENROUTER_KEY— or any OpenAI-compatible API viaPRAETOR_OPENROUTER_URL.
Quickstart
python examples/quickstart.py # deterministic end-to-end run, no model needed
python examples/run_benchmark.py # built-in eval suite, pass^k report
python examples/export_trace.py # OTel GenAI-aligned JSONL trace export
python examples/security_drill.py # live pentest of all five security guarantees
Connecting a real local model (Ollama):
from praetor.llm.ollama import OllamaProvider
from praetor.runtime.agent import Agent
from praetor.tools.exec import PythonTool
agent = Agent(OllamaProvider("llama3.1"), tools=[PythonTool()])
result = agent.run("Compute the 20th Fibonacci number, then reply with only the number.")
Connecting a real cloud model (OpenRouter, free keys work):
import os
from praetor.llm.openrouter import OpenRouterProvider
from praetor.runtime.agent import Agent
from praetor.tools.exec import PythonTool
os.environ.setdefault("PRAETOR_OPENROUTER_KEY", "<your key>")
agent = Agent(OpenRouterProvider(), tools=[PythonTool()]) # defaults to a free model
Using the runtime
from praetor.runtime.agent import Agent, Budget
from praetor.security.policy import ToolPolicy
from praetor.tools.exec import PythonTool
from praetor.tools.fs import ReadFileTool, Workspace, WriteFileTool
workspace = Workspace("./agent-workspace")
policy = ToolPolicy(allowed_tools=frozenset({"write_file", "read_file", "python"}))
agent = Agent(
provider, # any LLMProvider (Ollama, OpenRouter, scripted, your own)
tools=[WriteFileTool(workspace), ReadFileTool(workspace), PythonTool()],
policy=policy, # default-deny allowlist + per-tool/total budgets
budget=Budget(max_steps=16), # step and wall-clock limits
)
result = agent.run("Write a report and verify it.")
print(result.status) # ok | error | aborted_policy | aborted_budget
print(result.tool_calls) # only EXECUTED calls; denied attempts stay in the audit trail
print(result.steps) # every attempt with arguments, outcome, and duration
Writing a custom tool is a single class:
from praetor.tools.base import Tool
class NotesTool(Tool):
name = "notes"
description = "Append a note to the shared notebook."
parameters = {
"type": "object",
"properties": {"text": {"type": "string"}},
"required": ["text"],
}
def run(self, arguments):
# execute; Praetor handles validation, integrity, and the
# untrusted-data envelope around whatever you return
...
Evaluation harness
A Praetor benchmark score is never a model's opinion. Each task ships a verifier that executes real code against real state, and the harness runs every task k times against fresh agents:
- pass^k — the fraction of tasks where all k trials succeeded (reliability; the tau-bench notion)
- pass@k — the fraction of tasks where at least one trial succeeded
from praetor.eval.harness import EvalHarness, Task
from praetor.eval.tasks import default_pack, DecoyTool
suite = default_pack(workspace_root, decoy=DecoyTool())
report = EvalHarness(trials=3).run_suite(make_agent, suite)
print(report.table()) # human-readable table
print(report.to_dict()) # machine-readable: pass^k, pass@k, per-task rates
Built-in, CI-safe tasks: filesystem write (verified by reading the file from disk), computation (verified against the known result and a successful execution), and injection resistance (success = the forbidden decoy tool is never invoked).
Observability
Every run exports a trace shaped after the OpenTelemetry GenAI semantic conventions: an invoke_agent root span with one execute_tool child per executed step, carrying gen_ai.* attributes (model, tool name, token usage) plus Praetor-specific audit attributes.
from praetor.observability import spans_from_result, summarize, write_jsonl
spans = spans_from_result(result, model="llama3.1")
write_jsonl("trace.jsonl", spans) # ingest into any OTel-compatible backend
print(summarize(result)) # status, calls, failures, timing, tokens
Traces are deterministic and content-addressed: identical runs produce identical trace ids, so a trace can be replayed and diffed like any other artifact. The audit view deliberately excludes raw tool outputs and prompts — it records what was called and how it ended, never payload bytes.
Security model
| Guarantee | Mechanism |
|---|---|
| Only allowed tools run | Default-deny allowlist policy; violations abort the run and are audited |
| Strict inputs | JSON-schema validation of every tool call; unknown arguments always rejected |
| No tampered tools | Tool specs fingerprinted at registration; a mismatch quarantines the tool |
| Contained execution | Workspace realpath confinement (traversal and symlink escapes blocked); subprocess sandbox with rlimits (POSIX), scrubbed environment, wall-clock kill |
| Injection resistance | Observations stripped of control characters, capped in size, wrapped in an untrusted-data envelope; known injection markers flagged in the audit log |
| Bounded resource use | Per-tool and total call budgets, step budget, wall-clock budget, size caps on code and I/O |
| No secrets in the repo | Zero dependencies by design; CI secret scan on every push |
Threat model: agent-produced code and tool outputs are untrusted input. The sandbox is a process-level barrier, not a kernel container — for hostile multi-tenant workloads, add gVisor or Firecracker underneath (see SECURITY.md).
Project layout
praetor/
types.py # shared data types (AgentResult, ToolCall, Step, ...)
runtime/ # agent loop, bounded memory
security/ # schema validation, policy engine, sandbox, integrity guard
tools/ # filesystem, sandboxed python, tool registry
llm/ # provider protocol: Ollama (local), OpenRouter (cloud), scripted replay
eval/ # execution-based harness, benchmark task pack
observability/ # OTel GenAI-aligned trace export and run summaries
examples/ # quickstart, benchmark, trace export, security drill, real-model agents
tests/ # unit + adversarial stress suites
docs/ # publishing guide
Testing and CI
- 57 tests across eight files: unit suites for the runtime, tools, sandbox, schema, eval, and the OpenRouter provider (fully mocked HTTP, no network in CI), plus an adversarial stress suite (path-traversal fuzzing, size-cap enforcement, budget exhaustion under adversarial loops, injection storms, determinism).
- CI runs the full suite on Python 3.10 and 3.12, a clean-venv install smoke test on 3.11 (installs the built package into a fresh virtualenv and runs an end-to-end agent task), and a secret scan.
- A failing test step automatically files a GitHub issue with the full log for fast diagnosis.
Roadmap
- v0.4 — sandbox network egress control (default-deny outbound), signed tool packs
- v1.0 — optional OTLP exporter (as an install extra, keeping the core zero-dependency), docs site
Changelog
Release history and notes: CHANGELOG.md. Releases are published to PyPI automatically via trusted publishing (docs/PUBLISHING.md).
Contributing
Issues and pull requests are welcome. See CONTRIBUTING.md. Every PR must keep CI green — including the stress suite.
Security
Found something that looks like a vulnerability? Please follow the reporting steps in SECURITY.md — do not open a public issue for security reports.
License
MIT — use it, ship it, break it, learn from it.
Citation
If Praetor helps your work, a citation is appreciated:
@software{praetor,
author = {Ali Abdullah},
title = {Praetor: A security-first, zero-dependency agent runtime
with an execution-based evaluation harness},
year = {2026},
url = {https://github.com/Venomous-101/Praetor}
}
Metadata
Release files for praetor-agent 0.3.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| praetor_agent-0.3.0.tar.gz | 36.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| praetor_agent-0.3.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 68.9 kB
Release files / praetor_agent-0.3.0.tar.gz
| Download URL | praetor_agent-0.3.0.tar.gz |
|---|---|
| Size | 36.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
5193f6e84949d052570fa8d4a9de4f791297b17ebd2dfe82f176d545af6691f5
|
|
BLAKE2b-256 checksum How to use checksums |
631eaeebb84e74c6bde2d8c8f100c66c62510f5ac318be22fdba98c417bb743c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency logRelease files / praetor_agent-0.3.0-py3-none-any.whl
| Download URL | praetor_agent-0.3.0-py3-none-any.whl |
|---|---|
| Size | 33.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f6c7418ff9748830184d46c475c7cb6ebf13cce8959f392c907e3344d7a853aa
|
|
BLAKE2b-256 checksum How to use checksums |
cad95f0d3968344208d299e67f98b3ec0aeb695f5f440fbcffd396123ba5c221
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 4, 2026.
Transparency log