Skip to main content

PlannerCritic Engine

License: MIT Python 3.11+ PyPI Ruff Type checked Coverage Contributor Covenant OpenSSF Best Practices Field Test

Hierarchical task planning with an independent LLM critic. A planner decomposes a goal into a structured plan; a critic audits every subtask; the plan is revised until approval — or escalated to a human.

[!NOTE] Status: v0.2.2 released · PyPI · pip install planner-critic License: MIT


Why

Planning is the weakest part of agent systems. Most agents act too early, skip hard subproblems, and start executing a plan that was never reviewed. A single-pass decomposition embeds silent assumptions, and the first sign of trouble arrives mid-execution — after state has already diverged.

Single-pass planning fails silently on multi-step goals, and a model "reviewing" its own plan is agreement with extra steps. There is no draft to review, no independent reviewer to catch the gap, and no structured escalation.

PlannerCritic Engine closes this gap by treating planning as a first-class, productized artifact rather than a hidden chain-of-thought side effect.


What It Is

The Draft → Critique → Revise → Escalate Loop

 Goal + Constraints → PLANNER → typed plan → CRITIC → findings
                         ↑  │                        │
                         │  └────── revise ←──────────┘
                         │                             │
                         └──────── approved plan ──────┘
                                      │
                                   ESCALATE (if no convergence)
  • Draft — a planner LLM decomposes a goal into a structured, typed plan: tasks, dependencies, ordering, verification steps, and rollback points.
  • Critique — a separate critic LLM audits every subtask against six heuristic families: feasibility, risk, missing steps, unsafe sequencing, unverified dependencies, weak rollback — producing severity-graded findings.
  • Revise — the planner revises in response, in a bounded loop with a revision budget and convergence detection, preserving draft history.
  • Escalate — if the loop cannot converge, a human gets a minimal, precise question about exactly what is blocking approval.

The plan is a persisted, versioned artifact — you can diff revisions, see which critiques drove which changes, and trace whether a failed run was a planning failure or an execution failure.

Key Features

Feature Description Added
Risk tolerance balanced (findings are advisory warnings) or strict (zero tolerance, fail-closed) v0.1.0
Deterministic gates 7 injection-immune gates — ordering, branch-sanity, rollback, verification, preconditions, branch-tasks, high-risk completeness v0.1.0
Escalation management Human-in-the-loop with override, patch, and restart decisions v0.1.0
Convergence detection Early termination when the planner stops making progress — saves LLM calls v0.1.0
Provider registry Pluggable LLM providers (OpenRouter, OpenAI, oMLX, Ollama) via TOML config v0.1.0
StructuredEnforcer Retry mechanism for LLM JSON output — fail-closed after 3 retries v0.1.0
Plan versioning Every revision is a persisted artifact with diff support v0.1.0
Deterministic auto-repair Topological re-ordering + precondition closure — fixes ordering/dependency defects without LLM cost (#130, #131) v0.2.0
Oscillation detection Detects structural cycling and auto-converges (#152) v0.2.0
Domain Pack framework Domain-specific gate packs (SecOps, Supply Chain, FinOps, Data Eng) with plancritic init --template (#139, #140–143) v0.2.0
Policy-as-Code engine OPA/Rego + CEL policy evaluation — deterministic gates for custom compliance (#129, #156) v0.2.0
Security oracle SWE-bench-derived security corpus validates gates against human ground truth (#123–127) v0.2.0
Enterprise safety Dynamic posture, run budgets, state locking, precondition ledger, blast-radius quotas, secret/PII redaction (#149–151, #158, #159) v0.2.0
Developer surfaces plancritic check, diagnose, domains, policy, templates CLI + @guardrail decorator + seed Rego library (#137, #153, #162) v0.2.0
CI/CD integrations GitHub Action, GitLab CI template, AutoGen adapter, webhook notifier (#128, #134, #161) v0.2.0
Probe system Health probes for pre-execution precondition validation (DB, deploy, env, HTTP) v0.2.0
Drift observability Finding-drift detection — track how findings change across revisions (#181) v0.2.0
pytest plugin pytest-planner-critic — use gate assertions in your test suite (#156) v0.2.0
Frozen acceptance contract Bound pre-run; post-bind mutation creates new version + audit trail (#215) v0.2.1
Rollback credibility gate Detects unreachable / self-dependent / inconsistent-state / post-consumed rollbacks (#216) v0.2.1
Verification ordering gate Catches vacuous verifications — consumer runs before verified mutation (#219) v0.2.1
Histogram cycling detection Period-2 reshuffling stall signal — A→B→A→B pattern detection (#217) v0.2.1
Live-critic boundary evaluator Repeated-trial label-flip, evidence-drift, family-migration, underclaim metrics (#218) v0.2.1
Failure-mode register Intentional vs. needs-evidence assumptions register (14 rows) (#220) v0.2.1
Runtime precondition verification On by default; coverage honesty, fail-closed path (#244) v0.2.2
Typed rollback contracts Declarative restored-state + restoration-evidence on RollbackStep (#245) v0.2.2
Requirement-traceability gate Every plan step traces back to an acceptance criterion (#255) v0.2.2
Machine-actionable findings Edge targeting, observed-state, evidence refs, schema versioning (#243) v0.2.2
Decision-context capture Model metadata + unsupported-evidence frequency in boundary runner (#242) v0.2.2
Indirect-injection defense Capability-scoped state transitions + tool-result provenance (#249, #258) v0.2.2
Benign-twin control 11 benign twins measure injection isolation vs gate strictness (#260) v0.2.2
Compositional injection traps Individually feasible steps harmful in combination (#256) v0.2.2
Escalation audit trail Actor field, explain output, identity plumbing (#261) v0.2.2
Cost-vs-rigor guardrails Immutable gates config prevents skipping deterministic gates (#262) v0.2.2
Critic satisfaction signals Strict mode approves when critic explicitly endorses (#254) v0.2.2
Adaptive revision cap Strict goals cap at 1 revision to save LLM calls (#251) v0.2.2
Failure-origin taxonomy 51 bugs classified by first-detectable layer (#264) v0.2.2

What It Is Not

  • It does not execute the plan — an existing runner consumes the approved plan.
  • It does not guarantee plan correctness — it reduces risk, it cannot eliminate it.
  • It does not replace execution engines or agent frameworks (LangGraph, CrewAI, etc.).

Quick Start

Install

pip install planner-critic

Requires Python 3.11+.

Configure an LLM provider

Create a config file (or set an env var for the API key):

# plancritic.toml
[roles]
planner = "local"
critic = "local"

[providers.local]
transport = "openai-compatible"
base_url = "https://openrouter.ai/api/v1"      # or your provider
model = "openai/gpt-4o-mini"
api_key = "${OPENROUTER_API_KEY}"               # or set in your shell
max_tokens = 16384
timeout_s = 300.0

Run your first plan

plancritic plan path/to/goal.json --config plancritic.toml
plancritic demo       # run the bundled demo scenario
plancritic quickstart # create and run a sample goal

Requires an LLM provider (OpenRouter, OpenAI, or a local model). See the User Guide for a full walkthrough.


CLI

plancritic plan <goal.json>                # Plan a goal
plancritic critique <plan.json>            # Critique a plan
plancritic check --plan <plan.json>        # Quality check a plan
plancritic diagnose <plan-id>              # Diagnose plan issues
plancritic domains list                    # List available domain packs
plancritic policy check <plan.json>        # Evaluate Rego/CEL policies
plancritic templates list                  # List scaffold templates
plancritic field-test run --goals <dir>    # Run field test
plancritic eval --regression               # Security oracle regression
plancritic findings list                   # List plan findings
plancritic lessons                         # List learned lesson codes
plancritic plan replay <plan-id>           # Replay plan history
plancritic escalate list                   # List escalations
plancritic demo                            # Run demo scenario
plancritic quickstart                      # Quickstart demo
plancritic init [--dir]                    # Scaffold config + store
plancritic providers add/list/rm           # Manage LLM providers
plancritic migrate <old> <new>             # Migrate schema
plancritic serve                           # Start HTTP server
plancritic quota show                      # Show blast-radius quotas

See API Reference for full CLI docs, HTTP endpoints, and MCP tools.


Field Test

170 goals across 40 domains, all run against a real LLM (gpt-4o-mini via OpenRouter):

Metric v0.2.2 Result v0.2.1 Result
Balanced goals approved 73/73 (100%) 73/73 (100%)
Strict goals escalated 96/97 (99%) ⚠ 1 transient error 96/97 (99%) ⚠ 1 transient error
Adversarial goals escalated 11/11 (100%) ⚠ 3 adversarial-policy added 8/8 (100%)
True failures 0 0
Deterministic gate passes 170/170 (100%) 170/170 (100%)
Deterministic tests 1342 pass 1294 pass
Benchmarks 5 3
Scorecard A PASS PASS

Full results: field-test-results-0.2.2.md | v0.2.1 results


Documentation

Doc Path Contents
Field Test Results v0.2.2 results BLUF, regression diff, per-goal data, scorecards, blocker analysis
Field Test Results v0.2.1 results v0.2.1 results for reference
Field Test Results v0.2.0 results v0.2.0 results for reference
Field Test Results v0.1.0 results v0.1.0 results for reference
Field Test Plan plan Corpus, invariant assertions, execution guide
Release Notes v0.2.2 release-notes What's new, M1-M5 changes, field test results
Release Notes v0.2.1 release-notes What's new, hardening, field test results
Release Notes v0.2.0 release-notes v0.2.0 release notes for reference
Release Notes v0.1.0 release-notes v0.1.0 release notes for reference
Architecture architecture Component diagram, module map, data flow
API Reference api CLI cheat-sheet, HTTP endpoints, MCP tools
Design Decisions decisions DD-01..N decision records
Domain Pack Design domain-packs Domain pack protocol, pack format, engine integration
Policy Engine Design policy-engine OPA/Rego/CEL integration
Enterprise Safety Design enterprise-safety Posture, budgets, state, ledger, quotas, redaction
Developer Surfaces Design developer-surfaces CLI commands, decorator, seed Rego
Integration Surfaces Design integration CI runners, AutoGen, notifier
Security security Security policy, OWASP, OpenSSF
WBS Index (v0.2.2) wbs Milestone overview, dependency graph, M1-M6
WBS Index (v0.2.1) wbs v0.2.1 WBS for reference
WBS Index (v0.2.0) wbs v0.2.0 WBS for reference

Project Layout

planner-critic-engine/
├── docs/                    Documentation
│   ├── architecture/          System architecture and spec
│   ├── design/                PRD, design spec, design decisions
│   ├── field-test/            Field test plan + results (170 goals, 40 domains)
│   ├── reference/             API reference, quickstart, release notes
│   └── wbs/                   Work breakdown structure (M1–M10)
├── src/planner_critic/       Engine source
│   ├── adapters/               AutoGen, CrewAI, LangGraph, OpenAI Agents, PydanticAI adapters
│   ├── cli/                    21 CLI commands
│   ├── critique/               LLM critic with severity guardrail
│   ├── domains/                Domain-specific gate packs (SecOps, Supply Chain, FinOps, Data Eng)
│   ├── eval/                   Security oracle, injection harness, regression, label migration
│   ├── gates/                  7 deterministic gates
│   ├── llm/                    Provider registry, transport, logging
│   ├── loop/                   Plan revision loop, auto-repair, convergence, oscillation
│   ├── probe/                  Health probes (DB, deploy, env, HTTP)
│   ├── schema/                 Goal and plan schemas
│   ├── server/                 HTTP + MCP servers
│   └── store/                  SQLite plan store with versioning
├── tests/                    Test suite (field test + 90 deterministic tests)
├── .github/                  Issue templates, PR template, CI workflows
├── CHANGELOG.md               Release history
├── CONTRIBUTING.md            How to contribute
├── SECURITY.md                Security policy + OWASP + OpenSSF
└── pyproject.toml             Package metadata

Known Gaps (v0.3.0)

The authoritative register of known failure modes and load-bearing assumptions — each classified as an intentional trade-off or a claim still needing evidence — lives in docs/reference/failure-modes.md. Summary:

  • Planner capability gap — strict mode = escalation for non-trivial plans (intentional posture; see register F-01).
  • Local model support for planner — Qwen3-4B JSON valid; critic role still too weak (F-09).
  • TUI / studio / IDE surfaces — deferred to v0.3.0 (F-11).
  • Backstage developer portal plugin — deferred to v0.3.0 (F-12).
  • Adaptive revision cap — detect strict goals and reduce cap to 1, saving LLM calls (F-10).

See CHANGELOG.md for full details.


License

MIT — see LICENSE.

Metadata

Release files for planner-critic 0.2.2

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for planner-critic 0.2.2
File Size Uploaded
planner_critic-0.2.2.tar.gz 5.1 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for planner-critic 0.2.2
File Interpreter ABI Platform
planner_critic-0.2.2-py3-none-any.whl Python 3 none any Details

Total release size: 5.3 MB

Release files / planner_critic-0.2.2.tar.gz

Download URL planner_critic-0.2.2.tar.gz
Size 5.1 MB
Tags Source
SHA-256 checksum
How to use checksums
5d0b05db6d9301a20b7121c777df2e902483a9a4fc2c2502fe3e023038169249
BLAKE2b-256 checksum
How to use checksums
264189be8c413e1b7d458dcc5f6463423e437d581669728cd35c19a792be2698
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.5

Release files / planner_critic-0.2.2-py3-none-any.whl

Download URL planner_critic-0.2.2-py3-none-any.whl
Size 265.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
10971f3c8bc8751fcd217fa500b2b3ac633558defbfd4b12e58805e289af614d
BLAKE2b-256 checksum
How to use checksums
ef959bd135e42a7a9108840a0e0ecbc49e5c297956cf489cfaa10eaf67569e7d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.14.5

Release history Release notifications | RSS feed

0.2.3

2 release files

This release

0.2.2 This release

2 release files

0.2.1

2 release files

0.2.0

1 release file

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page