Skip to main content

Synthetic-prompt evaluation and persona-audit harness for web and API AI systems.

Project description

PromptPolygraph

Synthetic-prompt evaluation and persona-audit harness for web and API AI systems.

PromptPolygraph pushes thousands of synthetic prompts through any web/API/LLM system, scores the responses against a pluggable rubric (plus cheap deterministic assertions and an optional multi-judge ensemble), reacts to them through a panel of personas, traces low scores with a forensic audit, and renders the result as a markdown / docx / pdf / html report.

It is local-first and cloud-agnostic: one process, an async concurrency pool, and a SQLite store — runs on a laptop, a CI runner, or a container. No work queue, no database server, nothing to deploy. It also ships a deployable API + worker (Postgres-backed, Dockerized) for multi-user, team-scale use — see Service mode.

Status: the local harness (CLI) and the service layer are both built and tested. Runs fully offline in --mock mode; the Docker image and the postgres + api + worker compose stack are verified end to end.

Why

Evaluating an AI system well means more than a single accuracy number. PromptPolygraph combines:

  • Synthetic corpora — a fixed deterministic set for clean run-over-run baselining, or LLM-generated varied / adversarial probes, with dials for quantity, categories, difficulty, and seed.
  • Adapters — one small class (async def query(case) -> Response) per system under test. Ships an HTTP/REST adapter, an LLM-chat adapter, and an in-process callable adapter.
  • Layered scoring — deterministic assertions (contains / regex / json-schema / latency / semantic similarity / custom Python), with weights, per-test thresholds, and named + derived (F1-style) metrics, then an LLM rubric scorer with a multi-judge ensemble + inter-judge agreement.
  • Persona panel — a pool of distinct individuals react to the real responses (trust, usefulness, clarity, would-return), reconciled against the rubric so you optimize real value, not just the number.
  • Forensic audit — per-category agents trace low scores to failure modes and emit a concrete suggested fix (the file:line to change + a diff) for systems whose source you provide.
  • Comparison & trends — comparability-gated N-run comparison, per-dimension trends, and regression detection vs a pinned or rolling baseline, with statistical significance (a two-sample test, Benjamini-Hochberg-corrected across dimensions) so noise is not flagged as a regression.
  • Statistical rigor — proportions (red-team ASR, pass/assertion rates) carry confidence intervals, continuous aggregates carry seeded bootstrap intervals, and the gate can decline to fail inside the noise band — so the numbers a report stakes a claim on are defensible.
  • Visuals & reports — a local dashboard (score heatmaps, compare matrix, trend lines, persona radar, root-cause→fix view) and presentation-grade markdown / docx / pdf / html reports with inline charts.

See CHANGELOG.md for what's new in each release.

Install

pip install -e ".[dev]"            # core + test deps
pip install -e ".[dev,llm]"        # + OpenAI-compatible adapter
pip install -e ".[dev,redteam]"    # + OSS-grounded red-team sources (garak, PyRIT, DeepTeam, datasets)

Requires Python 3.10+. The red-team engine and Arena work with no extras (the built-in catalog probe source needs no dependencies); the [redteam] extra only lights up the external OSS sources.

Quickstart

# Offline end-to-end demo against the bundled example (no API key needed):
# runs a fixed corpus through a built-in demo target, scores it, runs the persona
# + forensic audit, and writes md/html/docx/pdf reports under polygraph_out/.
polygraph all --config examples/everyday_assistant/config.yaml --mock --format md,html,docx,pdf

With an ANTHROPIC_API_KEY set and a real adapter configured, drop --mock.

Bundled example packs (each a self-contained config + corpus + rubric + personas):

  • examples/everyday_assistant/ — the default: a neutral general-purpose assistant (general knowledge, how-tos, reasoning, refusal, safety, edge input).
  • examples/support_bot/ — a customer-support assistant for a fictional SaaS.
  • examples/clinical_trials/ — a domain pack for a clinical-trial-protocol assistant (personas for medical writer, clinical scientist, data manager, medical director, regulatory, biostatistician, clinical ops, pharmacovigilance).

Copy one as a starting point, or generate a fresh tailored pack with polygraph tune (see below).

Runbook

1. Point an adapter at your system

The adapter is the only thing that changes per target. Set it in config.yaml:

# HTTP / REST
adapter:
  type: http
  options:
    url: https://api.example.com/chat
    method: POST
    headers: { Authorization: "Bearer ${API_KEY}" }   # ${VAR} reads the env
    body_template: { message: "{{prompt}}" }           # {{prompt}} {{category}} {{id}}
    response_path: "reply.text"                         # JMESPath into the JSON response

# LLM chat (OpenAI / Anthropic / OpenAI-compatible; openai needs the [llm] extra)
adapter:
  type: llm
  options: { provider: anthropic, model: claude-opus-4-8, system: "You are ...", max_tokens: 512 }

For anything exotic, implement async def query(self, case) -> Response and pass an in-process callable with --callable mymodule:my_fn.

Run on local / open models (Ollama, vLLM, LM Studio)

The grader/judge, audit agents, and corpus generation can run against any OpenAI-compatible endpoint — not just Anthropic — so the whole eval can run 100% locally on open models, no API key:

llm:
  provider: ollama          # anthropic (default) | openai | ollama | vllm | lmstudio | openai-compatible
  base_url: http://localhost:11434/v1   # default for ollama; set for other servers
model: llama3.1             # the model NAME for the judge/audit/generation

A local provider (ollama, …) needs no key and runs live; Anthropic/OpenAI fall back to --mock when their key is absent. To also point the system under test at a local model: adapter: { type: llm, options: { provider: ollama, model: llama3.1 } }.

2. Choose a corpus mode (dials)

polygraph run --config c.yaml --mode fixed                       # deterministic set; stable ids for baselining
polygraph run --config c.yaml --mode varied      --count 1000    # fresh LLM-generated each run
polygraph run --config c.yaml --mode adversarial --count 500 --difficulty aggressive
polygraph run --config c.yaml --mode hybrid       --count 800    # fixed core + generated supplement

Dials: --count, --per-category, --categories a,b,c, --difficulty {mild|standard|aggressive}, --concurrency, plus rps/timeout/retries/resume in the config. Thousands of prompts is just the concurrency pool; runs are durable and --resume-able.

3. Score

polygraph analyze --config c.yaml --run <run_id> --judges 3 --ci

Deterministic assertions (declared per probe: contains / regex / json-schema / latency …) run first; then an LLM rubric scorer. --judges N adds an ensemble with inter-judge agreement. --ci exits non-zero if the gate fails (per-dimension threshold, cumulative across categories). The rubric is a YAML pack — edit rubric.yaml (dimensions, anchors, per-category applicability, threshold) for your domain.

4. Audit (persona panel + forensic)

polygraph audit --config c.yaml --run <run_id>

A panel of personas reacts to the real responses (trust / usefulness / clarity / would-return) and is reconciled against the rubric so you optimize real value, not just the number; per-category forensic agents trace low scores to the highest-leverage fixes.

Manage personas:

polygraph personas list
polygraph personas new "a grumpy retiree who hates phone trees" --out persona.yaml
polygraph personas generate --domain "developer tools" --count 8 --out panel.yaml

Select a panel via personas_path: (a YAML file), audit.persona_pool: N (sample the 13-persona library), or audit.personas: [ids].

5. Report, compare, trend & regressions

polygraph report  --config c.yaml --run <run_id> --format md,docx,pdf,html --baseline <prior_run_id>
polygraph compare --config c.yaml --run-a <id> --run-b <id>        # A/B win/loss/tie
polygraph compare --config c.yaml --runs <id1>,<id2>,<id3>         # N-run comparison report
polygraph trend   --config c.yaml --project <name> --window 30     # per-dimension trend over runs
polygraph regressions --config c.yaml --run <id> --against rolling:20   # or --against <baseline_id>

--baseline adds Δ-vs-prior columns + regression flags. Comparison is comparability-gated (runs are only score-compared when their corpus + rubric fingerprints match). trend plots per-dimension means + slope over the recent comparable runs; regressions flags dimensions that dropped vs a pinned or rolling baseline. PDF uses LibreOffice headless when available and is skipped gracefully otherwise.

One-shot

polygraph all --config c.yaml --format md,html        # run -> analyze -> audit -> report

Browse results in a local dashboard

A zero-dependency web UI (stdlib only — no server to deploy) over the runs the CLI has produced:

polygraph dashboard --out-dir polygraph_out          # opens http://127.0.0.1:8765

Browse runs, drill into category/dimension scores, per-case responses + assertions, the persona panel and forensic audit, and open the rendered reports. (For a multi-user, Postgres-backed team server with a job queue + API, see Service mode.)

Export the generated prompts as a reusable dataset

PromptPolygraph generates synthetic corpora — export them as a portable dataset to reuse anywhere:

polygraph export --run <run_id> --out prompts.json                 # full cases from a stored run
polygraph export --corpus mycorpus/ --out prompts.jsonl --format jsonl
polygraph export --corpus mycorpus/ --out prompts.csv --format csv --prompts-only

Formats: json / jsonl / csv; --prompts-only emits just {prompt, category} for a clean prompt corpus.

Customize reports (templates & branding)

Reports render from Jinja2 templates you can override and brand:

report:
  template: default            # built-in: default | minimal
  template_dir: my_templates/  # optional: your own report.html.j2 / report.md.j2 (overrides built-in)
  branding: { title: "Acme Eval", accent: "#0aa", logo: "https://…/logo.png" }

Tailor it to your domain

Set a domain (a one-line description of the system under test) and the generated prompts — and, with audit.tailor_personas, the persona panel — are specific to it instead of generic:

polygraph all --config c.yaml --mode adversarial --count 500 \
  --domain "Clinical trial protocol digitalization assistant"

Or scaffold a whole tailored project (rubric + persona panel + starter corpus + config) for a domain in one command, then refine the files and point the adapter at your system:

polygraph tune --domain "Clinical trial protocol digitalization assistant" --out projects/ctpd
polygraph all  --config projects/ctpd/config.yaml --mock

Bootstrap a golden set with a subject-matter expert

polygraph tune auto-scaffolds; polygraph elicit builds an expert-validated golden corpus through a guided, human-in-the-loop walkthrough:

polygraph elicit interview --domain "Clinical trial protocol digitalization assistant" --out brief.yaml
#   (or: polygraph elicit init --domain "..." --out brief.yaml   — fill the brief by hand/async)
polygraph elicit build    --brief brief.yaml --out projects/ctpd     # draft probes + review sheet + rubric + personas
#   the SME edits projects/ctpd/review.yaml (set `decision: reject` to drop a probe, or edit any field)
polygraph elicit finalize --review projects/ctpd/review.yaml --out projects/ctpd   # accepted probes -> golden corpus
polygraph all --config projects/ctpd/config.yaml --mock

The interview asks the expert what the system does, what good/bad answers look like per category, the failure modes and must-refuse cases, and real example queries; probes are drafted grounded in those answers, then gated by the expert's review — which is what makes the set golden. Personas are drawn from the expert roles named in the brief.

A ready-made clinical-trials pack ships under examples/clinical_trials/ (personas for medical writer, clinical scientist, data manager, medical director, regulatory, biostatistician, clinical ops, and pharmacovigilance; a protocol-focused corpus and rubric):

polygraph all --config examples/clinical_trials/config.yaml --mock --format md,html

Red-team & Arena

An authorized adversarial red team for a system you own — a roster of attacker agents pressurizes the target, a breach judge rules each exchange, and findings roll up into a severity-ranked vulnerability report mapped to the OWASP Top-10 for LLM Apps and MITRE ATLAS.

# Offline demo against the built-in target (zero tokens, deterministic):
polygraph redteam --profile all_frontier --mock --format md,html,json

# Ground the attack in OSS tooling + datasets, judge with a safety classifier:
polygraph redteam --profile deep \
  --sources catalog,garak,pyrit,deepteam,dataset:advbench \
  --guard --format md,html,json

polygraph redteam profiles      # list the preconfigured teams
  • Preconfigured teamsall_frontier (default; every strategy on a frontier model, each agent a distinct persona + temperature), multi_frontier (mixed vendors), mixed (local open + frontier), local_swarm (100% offline), pressure, quick, and deep (adaptive smart multi-turn).
  • OSS sources — attackers seed + mutate from a catalog of known techniques. --sources folds in optional garak, PyRIT, DeepTeam, and on-demand datasets (AdvBench / HarmBench / JailbreakBench); each runs through the same target → judge loop. Install the [redteam] extra.
  • Smart multi-turn — PAIR iterative refinement, Crescendo escalation, and TAP candidate trees, per-attacker mode.
  • Converters — base64 / rot13 / leetspeak / payload-split / roleplay-wrap / many-shot transforms, per-attacker converter.
  • Judges — an LLM reviewer (default) or --guard for a Llama-Guard-style safety classifier on a local model.
  • Arena — the dashboard (/redteam, SSE) or service (/ws/redteam, WebSocket) streams the run live, persists it for replay (Live/Replay toggle), and offers a per-attacker drill-down: the multi-turn timeline plus a root cause that — when redteam.code_path points at a local checkout and a model is available — traces the breach to file:line rungs of the target source (local model by default; air-gap and consent controls for any remote provider). Without a clone it shows the finding summary (control, OWASP/ATLAS, mitigation).

The report leads with the attack success rate (ASR), an OWASP coverage table (tested vs breached), and OWASP/ATLAS tags on every finding. See docs/REDTEAM_LOCAL.md for running attacker/judge models locally (Ollama / HuggingFace, Mac Studio / NVIDIA DGX Spark / cloud).

Continuous integration (CI gate)

PromptPolygraph drops into a pipeline as a gate — validate the config, score a run, fail on a regression, and emit reports the CI renders natively. Generate a starter pipeline:

polygraph scaffold-ci github      # or: gitlab | jenkins | precommit
polygraph validate-config --config config.yaml           # fail fast on a bad config (precise error path)
polygraph analyze --config config.yaml --run <id> --ci \
    --baseline rolling:5 --github-annotations --pr-comment pr.md
polygraph report  --config config.yaml --run <id> --format junit,sarif

--ci exits non-zero on gate failure. --baseline (a run id / rolling:N / HEAD) adds a statistically-tested regression check. --github-annotations emits ::error/::warning workflow commands + a job summary; --pr-comment writes a markdown summary. JUnit XML lands in the CI's test tab; SARIF 2.1.0 renders findings inline on the PR via GitHub/GitLab code scanning (a code-grounded red-team trace carries the file:line). polygraph schema writes JSON Schemas for editor autocomplete. See docs/CI.md.

Validation & audit (toward 1.0)

For institutions that need to defend the evaluation itself, PromptPolygraph ships an evidence layer:

polygraph validate --out evidence/   # IQ/OQ/PQ evidence bundle (install, components, reproducibility)
polygraph references --check         # the OWASP/MITRE-ATLAS technique mapping is pinned (CI-enforced)
polygraph calibrate --min-f1 0.7     # score the breach judge vs labeled ground truth (precision/recall/F1/κ)
polygraph manifest --run <id>        # provenance: tool/deps/fingerprints/sources that produced a result
polygraph bundle --run <id>          # seal artifacts into a tamper-evident .tar.gz (checksum manifest)
polygraph verify run.polygraph.tar.gz   # re-hash + refuse on any tampered/missing file (HMAC-signable)

These make a run's numbers defensible: provenance-tracked probes, a calibrated judge, reproducible runs, confidence intervals, and a self-checking evidence bundle. See docs/VALIDATION.md.

Architecture

One local process: Corpus → Runner (async pool, timeout/retry/resume) → Adapter → target → Response → SQLite + cache → Analyze (assertions + LLM ensemble) → Gate → Audit (persona + forensic) → Report. No work queue, no database server. The store and dispatch sit behind interfaces, so a service/worker deployment can wrap the same engine without touching it.

The red-team engine reuses the same adapter + report spine: Profile (attacker roster) → [LLM attackers ⨯ multi-turn modes] + [OSS sources] → Adapter → target → Breach judge (LLM reviewer | Llama-Guard) → Vulnerability report (ASR + OWASP/ATLAS), with a live event stream (emit) that the CLI logs and the dashboard/service render as the Arena.

Service mode (multi-user, scalable)

The same engine runs as a deployable API + worker for a team:

pip install -e ".[service]"
polygraph-server      # FastAPI: trigger runs, poll status, fetch reports, compare, manage personas
polygraph-worker      # claims queued jobs and executes the pipeline (scale horizontally)
# or one container that does both (in-process worker) for local use

Or the whole stack in containers (Postgres + API + a separate worker):

docker compose up --build           # api on :8080, worker scales with --scale worker=N
curl -XPOST localhost:8080/api/runs -H 'X-API-Key: change-me-dev-key' \
     -H 'Content-Type: application/json' -d '{"config_name":"support_bot"}'

The default Docker image includes LibreOffice, so PDF reports render in-container; build with --build-arg INCLUDE_PDF=false for a slimmer image (html/md/docx still work).

  • Store: SQLite locally, Postgres in production (same code, just POLYGRAPH_DATABASE_URL). A durable job queue lets many workers pull work safely (FOR UPDATE SKIP LOCKED on Postgres).
  • API (/api, API-key auth): create runs, poll status + live progress, fetch summary/report (html & markdown re-render from the DB), browse cases, A/B compare, manage personas. Plus a dashboard at /.
  • Scheduling & CI: cron schedules enqueue recurring runs; per-run or global webhooks POST a summary on completion; --ci-style gating is available via the run verdict.
  • Deploy: one Docker image, two roles, on AWS (App Runner / ECS Fargate / EKS) or GCP (Cloud Run).

See docs/SERVICE.md for the operator runbook and deploy/README.md for cloud deployment.

Development

pip install -e ".[dev,service]"
pytest -q                                   # full suite
python -m promptpolygraph --help            # same as the `polygraph` CLI
python -m build                             # build wheel + sdist

CI (.github/workflows/ci.yml) runs the suite on Python 3.11/3.12, smoke-tests the default example offline, and verifies the built wheel ships the persona data.

License

Apache-2.0. See LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

promptpolygraph-0.7.0.tar.gz (451.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

promptpolygraph-0.7.0-py3-none-any.whl (381.6 kB view details)

Uploaded Python 3

File details

Details for the file promptpolygraph-0.7.0.tar.gz.

File metadata

  • Download URL: promptpolygraph-0.7.0.tar.gz
  • Upload date:
  • Size: 451.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for promptpolygraph-0.7.0.tar.gz
Algorithm Hash digest
SHA256 bc6835bd8ec9e220dfdb7b015c4d583ccd954f625f906f6490228847e6ff217c
MD5 197bfb369198f386a54a4b9442cd8865
BLAKE2b-256 f3c23f5362308e5aad03c7efac79f9cd76397f61a1b453de638f3a74a9dd0a8e

See more details on using hashes here.

Provenance

The following attestation bundles were made for promptpolygraph-0.7.0.tar.gz:

Publisher: publish.yml on justinkastil/promptpolygraph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file promptpolygraph-0.7.0-py3-none-any.whl.

File metadata

  • Download URL: promptpolygraph-0.7.0-py3-none-any.whl
  • Upload date:
  • Size: 381.6 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for promptpolygraph-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ffe38d6f0c51ac2b6142ca925910afded7345841bdf92f77012f435cca472b07
MD5 74d20404534bfffb5b779872a27050bb
BLAKE2b-256 8fd2e8285af717edc9ab4224f0fd40104ffbd88bfe165c98f364bb05ee0b46af

See more details on using hashes here.

Provenance

The following attestation bundles were made for promptpolygraph-0.7.0-py3-none-any.whl:

Publisher: publish.yml on justinkastil/promptpolygraph

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page