Skip to main content

PandaProbe Harness

PandaProbe Harness evaluates any developer-owned, PandaProbe-instrumented task agent and maintains a shared learned-rules workspace with a separate, package-owned repair agent.

PyPI CI Python License

Ownership model

The developer owns the task agent, model, framework, prompts, domain tools, execution loop, and environment. PandaProbe owns task instrumentation and evaluation, trajectory detection, diagnostic notices, the managed repair-agent loop, workspace administration, validation, and read-only rule delivery.

Task and repair activity are isolated:

developer task agent → task trace → evaluation/gate → notice
                                                ↓
                                  PandaProbe managed repair
                                                ↓
                                     candidate rule
                                                ↓
developer task agent ── read-only tools ──→ shared workspace

The task agent never reads notices, inspects diagnostic traces, acknowledges notices, writes or retires rules, or controls validation. Its harness tools are exactly harness_rules_read, harness_rules_search, harness_rules_list, and harness_rule_status, all read-only.

Learned rules are read on demand, never injected

system_context() never includes rule bodies or an expanded rule-file index and performs no implicit rule read or search. It contains only a stable note that learned rules are available through those four tools. The task agent decides whether and when to list scopes, search, or read one. Workspace rules.md holds SKILL-style task-facing instructions and the compact generated scope index; each entry carries a bounded description and active/provisional counts, never rule text.

Scope: where a rule is filed

Managed repair decides, from the failure evidence it already has, whether a rule is broadly applicable or belongs to a specific context. That decision is part of the existing repair call — no extra model round.

  • global is the default: broadly reusable rules, not tied to one task, workflow, application, tool, or domain.
  • A concise contextual name — an application, workflow, or domain drawn from the evidence — is preferred whenever the rule really belongs to that context. The catalog is open, and a new name simply creates its rules/<scope>.md.
  • scoped is the fallback: the rule is specific, but no meaningful stable name could be determined.

Scope naming is generic and not tied to any benchmark or integration; a host's own label for itself is rejected as a scope. Hosts may pass bounded RuleScopeHint metadata and a short task_summary to inform the decision, but neither dictates it. PandaProbe normalizes names only for filename safety, owns the resulting path, and imposes no prefix convention.

Validation decides promotion and retirement

A new rule enters as a provisional candidate and only validation promotes or retires it — never the repair agent. Replay is the strong path; a cheap forward trial over live sessions decides candidates replay cannot reach, so every candidate reaches a verdict. Call settle_validation() at a phase boundary before snapshotting or reporting a ruleset.

Installation

pip install pandaprobe-harness

Managed repair uses PandaProbe's official LiteLLM wrapper. The same normalized path accepts LiteLLM identifiers for OpenAI, Anthropic API, Anthropic on Bedrock, and Gemini on Vertex AI:

openai/...
anthropic/...
bedrock/anthropic....
vertex_ai/...

No potentially billable model is selected by default. Configure one explicitly with repair_model= or HARNESS_REPAIR_MODEL. Provider credentials and cloud settings follow LiteLLM conventions.

Generic task-loop integration

from pandaprobe_harness import Harness, HarnessConfig, RuleScopeHint
from pandaprobe_harness.agent_tools.native import as_anthropic_tools

harness = Harness.create(
    HarnessConfig(
        repair_model="openai/gpt-...",
        repair_timeout_s=60,
        repair_max_turns=6,
        repair_max_tokens=4096,
        repair_temperature=None,
        repair_reasoning_effort="none",  # current OpenAI reasoning models + tools
        trace_repair_agent=False,
        domain_policy="Describe authorized domain behavior here.",
    )
)

async def one_turn(session_id: str, user_input: str) -> str:
    context = harness.system_context(session_id, task_hint=user_input)
    rule_specs, rule_dispatch = as_anthropic_tools(harness.task_tools)

    # You still construct and run your own agent with your own domain tools.
    answer = await my_agent_step(
        system_prompt=context + MY_PROMPT,
        tools=[*my_domain_tools, *rule_specs],
        tool_dispatch=rule_dispatch,
        user_input=user_input,
    )

    # Flush/export task tracing first, then register the completed task turn.
    harness.on_turn_end({
        "session_id": session_id,
        "turn_index": next_index(),
        "end_state": end_state(),
        "rule_scope_hints": [
            RuleScopeHint(
                key="payments",
                description="Payment authorization and transaction workflows.",
            ).to_json()
        ],
    })
    settlement = await harness.settle(session_id)
    return answer

The host must settle before starting the next task turn when same-session repair is desired. Settlement waits for task evaluation, notice persistence, and one bounded managed repair attempt. It does not synchronously wait for domain replay validation. A candidate is immediately discoverable through list/read after settlement, but is never inserted into the next prompt; validation may promote or retire it later.

Before a task turn, the host may attach the stable capability preamble and four read-only tools. The harness injects no learned content and makes no automatic rule-tool call. After the turn, tracing is flushed, evaluation/gating run, related notices are grouped into one bounded repair episode, and managed repair may add at most one provisional candidate or resolve without one. Successful settlement atomically refreshes the index/scope artifacts before the next task turn can query them; timeout or failure leaves every notice recoverable.

settlement.repair exposes status, repair/task session IDs, repair episode and notice IDs, recommended/selected scope, considered/existing/candidate rule IDs, suppression reason, model turns/tool calls, normalized usage when available, and error category. Repair failure or timeout never fails the developer task.

Framework turn detectors remain available through Harness.for_langgraph(), for_langchain(), for_deepagents(), for_crewai(), for_claude_agent_sdk(), and for_openai_agents().

Configuration

Managed repair fields and environment equivalents are:

Field Environment
repair_model HARNESS_REPAIR_MODEL
repair_timeout_s HARNESS_REPAIR_TIMEOUT_S
repair_max_turns HARNESS_REPAIR_MAX_TURNS
repair_max_tokens HARNESS_REPAIR_MAX_TOKENS
repair_temperature HARNESS_REPAIR_TEMPERATURE
repair_reasoning_effort HARNESS_REPAIR_REASONING_EFFORT
trace_repair_agent HARNESS_TRACE_REPAIR_AGENT
domain_policy HARNESS_DOMAIN_POLICY

observe_only=True remains non-mutating and does not require a repair model. Managed repair requires rule_validation=True, so a repair-authored rule can never skip the candidate lifecycle. Task tracing is unchanged when repair tracing is disabled. When enabled, the PandaProbe SDK records repair completions under repair-<task-session>-<episode-id> with repair-role metadata; exact-session task trace discovery excludes them.

Each enabled repair run exports one trace named pandaprobe. Its harness CHAIN span contains repeated repair-agent and tools AGENT spans; each repair-agent contains the official wrapper's litellm-chat LLM span, and each tools span contains one TOOL child per restricted workspace operation. No second tool-only trace is created.

Offline examples and operator tools

uv run python examples/misc/offline_self_heal.py
uv run python examples/misc/closed_loop_self_heal.py
uv run python examples/misc/calibration_demo.py

The examples use deterministic completion fakes and require no credentials. Operator-only CLIs remain pandaprobe-harness-eval for regression replay and pandaprobe-harness-calibrate for threshold calibration. The former task-administration companion CLI has been removed.

Development

uv run pytest -q
uv run ruff check .
uv run mypy

See CHANGELOG.md for migration notes and CONTRIBUTING.md for project invariants.

License

MIT © Chirpz AI

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pandaprobe_harness-0.9.0.tar.gz (1.0 MB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pandaprobe_harness-0.9.0-py3-none-any.whl (157.2 kB view details)

Uploaded Python 3

File details

Details for the file pandaprobe_harness-0.9.0.tar.gz.

File metadata

  • Download URL: pandaprobe_harness-0.9.0.tar.gz
  • Upload date:
  • Size: 1.0 MB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for pandaprobe_harness-0.9.0.tar.gz
Algorithm Hash digest
SHA256 67ddd684fe357d94bf147134e5870b0d40b6039177e74d5666538cdc51ae5977
MD5 d2ddbbda036b1ed0ad9f020d2a71e91a
BLAKE2b-256 c4c224ecae7acaaefbddb86ef4fc41611acb028b8a2c5f25aac257c57905cc03

See more details on using hashes here.

Provenance

The following attestation bundles were made for pandaprobe_harness-0.9.0.tar.gz:

Publisher: release.yml on chirpz-ai/pandaprobe-harness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pandaprobe_harness-0.9.0-py3-none-any.whl.

File metadata

File hashes

Hashes for pandaprobe_harness-0.9.0-py3-none-any.whl
Algorithm Hash digest
SHA256 6b43d1d7ccef88e5af8af607d843e3f68ac2b37f6cde4de5bae5eae1f4644c86
MD5 d8b90582a292f291b1fccd0664e96e19
BLAKE2b-256 ce6d429a3eb4973f94ba1af88ff135e778daae75004297bbb2d31b4512eb726f

See more details on using hashes here.

Provenance

The following attestation bundles were made for pandaprobe_harness-0.9.0-py3-none-any.whl:

Publisher: release.yml on chirpz-ai/pandaprobe-harness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.9.0 This release

2 files

0.8.0

2 files

0.7.0

2 files

0.6.3

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page