Skip to main content

promptdrift

Know what your prompt change actually did.

Prompt regression testing with CI gating. Snapshot-diff first, flake-aware by design, Python native.

CI Python License Code style: ruff

中文文档

PyPI name note: the distribution is published as promptdrift-py because PyPI's name-confusion policy blocks the bare name (someone registered the empty project prompt-drift, edit distance 1). The install command, CLI, and Python imports are all still promptdrift.

Teams that ship LLM features edit prompts constantly, and every edit can quietly break behavior that used to work. promptdrift catches that: it re-runs the same cases against your prompt, compares pass rates against a baseline snapshot you commit to git, and fails CI when behavior got statistically worse. Noise never trips the gate.

Quick tour, no API key needed

pipx install promptdrift-py
promptdrift demo

The demo runs a complete loop offline on a scripted mock provider:

  1. run executes a suite (3 samples per case)
  2. approve records the behavior as the baseline
  3. the mock answer gets "improved" and loses its structured steps
  4. diff catches the drift and exits 1, the way CI would
step 4/4 — diff: catch the drift
┌──────────────┬─────────────────────┬────────┬───────┬────────────────┐
│ case         │ assertion           │ before │ after │ verdict        │
├──────────────┼─────────────────────┼────────┼───────┼────────────────┤
│ refund-steps │ contains-申请退款    │ 3/3    │ 0/3   │ 🔴 regressed   │
│              │ regex-3-5-个工作日   │ 3/3    │ 0/3   │ 🔴 regressed   │
│ invoice-json │ is_json-yes         │ 3/3    │ 3/3   │ ✅ stable      │
└──────────────┴─────────────────────┴────────┴───────┴────────────────┘
❌ gate FAILED — 1 regressed case(s)

Why plain assertions aren't enough

LLM output is non-deterministic. A CI check like "output must contain X" fails randomly, gets retried, and eventually gets disabled. The verdict engine handles the noise instead:

Verdict Meaning Gate behavior
🔴 regressed pass rate fell beyond sampling noise (Wilson score intervals separate) fails regression (default)
⚠️ unstable samples disagree where the baseline was deterministic fails only flaky
🟢 improved pass rate rose beyond noise never fails
✅ stable within noise passes
🆕 new / ➖ removed case or assertion added or removed since the baseline fails any-fail when imperfect

Concretely: at samples: 3, a 3/3 → 2/3 drop is ⚠️ unstable, a warning. At samples: 10, 10/10 → 3/10 is 🔴 regressed. Raise the sample count to sharpen the statistics; run promptdrift cost first to see what that costs.

How it fits together

promptest.yaml ──► Runner ──► Run ──► Differ ◄── Snapshot (.promptest/baselines/*.snap.yaml)
                       │                          ▲
                       └── provider ──► sqlite cache

Baselines are YAML files you commit. When a PR touches one, that diff is the behavior change under review, the same way jest snapshots work. A config fingerprint (model, sampling params, sample count, judge model) protects comparability: swap the model and the old baseline goes stale instead of being silently mis-compared. Editing prompt text, vars, or mock answers is never staleness; that is exactly the change the diff should measure.

One provider adapter covers every OpenAI-compatible endpoint (OpenAI, GLM, DeepSeek, Qwen, Moonshot, vLLM, Ollama), configured with base_url.

Suite file

$schema: https://raw.githubusercontent.com/ritaprieto900/promptdrift/main/schema/promptest.schema.json
suite: support-agent
provider:
  openai_compat:
    model: glm-4.7
    base_url: https://open.bigmodel.cn/api/paas/v4
    api_key_env: ZHIPUAI_API_KEY
    pricing: {prompt_per_1m: 1.0, completion_per_1m: 8.0}
judge:                          # optional: grades `judge` assertions
  openai_compat:
    model: glm-4.7-flash
    base_url: https://open.bigmodel.cn/api/paas/v4
    api_key_env: ZHIPUAI_API_KEY
    temperature: 0.1
samples: 3
cases:
  - id: refund-steps
    messages:
      - role: user
        content: "订单有点问题,我想申请{{topic}}"
    vars:
      topic: 退款
    assertions:
      - contains: "申请退款"
      - regex: "3-5 个工作日"
      - not_contains: "抱歉"
      - latency_under: 3000
      - judge:
          rubric: |
            回复是否包含明确的退款步骤(入口、操作、时限)?
            5 = 步骤完整且语气友好;1 = 没有回答问题。
          min_score: 4

Deterministic assertions: equals, contains, not_contains, regex, is_json, json_schema, latency_under, completion_tokens_under. The judge assertion sends each sample to the judge model with your rubric and expects a 1-5 score; min_score sets the pass bar. Judges are only human (well, only models), so their scores feed the same multi-sample statistics as everything else: a wavering judge shows up as ⚠️ unstable, not as a blocked merge.

Commands

Command Purpose
promptdrift init scaffold a starter suite (mock provider, runs offline)
promptdrift run execute once; exit 1 if any assertion failed
promptdrift approve record current behavior as the baseline
promptdrift diff run and compare against the baseline; this is the CI gate
promptdrift show inspect the latest run or the baseline
promptdrift cost estimate calls and cost (when pricing is configured)
promptdrift demo the offline end-to-end walkthrough
promptdrift schema JSON Schema for suite files (editor completion)

diff options that matter in CI: --fail-on regression|flaky|any-fail, --require-baseline, and --md report.md to write a PR-ready markdown report.

Exit codes: 0 pass, 1 gate or assertion failure, 2 config or runtime error.

# GitHub Actions, minimal version (see docs/ci.md for the full one)
- run: uv tool install promptdrift-py
- run: promptdrift diff --require-baseline --md pr-report.md
  env:
    OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}

How it compares

promptdrift promptfoo DeepEval
Language Python Node/TypeScript Python
Core primitive baseline snapshot diff eval runs + web viewer pytest metrics
Non-determinism multi-sample + interval statistics per-run assertions metric thresholds
Review flow snapshot diff in the PR open the web UI read pytest output

promptfoo is the broader eval platform and DeepEval is strong on metric research. This project stays narrow: gate prompt changes in CI, with a verdict you can trust.

Roadmap

  • selective re-runs of changed cases only; embedding-similarity assertion
  • composite GitHub Action + sticky PR comments once the PyPI release lands
  • multi-suite projects, cross-model comparison mode

Development

git clone https://github.com/ritaprieto900/promptdrift
cd promptdrift
uv sync
uv run pytest              # 150+ tests, fully offline
uv run ruff check .
uv run pyright

The whole test suite runs on the mock provider, so CI never touches a real endpoint. CONTRIBUTING.md has the architecture map and the design invariants.

License

Apache-2.0

Metadata

Release files for promptdrift-py 0.1.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for promptdrift-py 0.1.1
File Size Uploaded
promptdrift_py-0.1.1.tar.gz 174.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for promptdrift-py 0.1.1
File Interpreter ABI Platform
promptdrift_py-0.1.1-py3-none-any.whl Python 3 none any Details

Total release size: 226.3 kB

Release files / promptdrift_py-0.1.1.tar.gz

Download URL promptdrift_py-0.1.1.tar.gz
Size 174.0 kB
Tags Source
SHA-256 checksum
How to use checksums
83b301345dae27ee2221209a4f1e4bb428b527e4d8d9ba5e679482afea353b0a
BLAKE2b-256 checksum
How to use checksums
da6d0eabb4b01b229ec68a14a1380f284d22283e189d3cc99e56c88ca1cb2979
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.20 {"installer":{"name":"uv","version":"0.12.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release files / promptdrift_py-0.1.1-py3-none-any.whl

Download URL promptdrift_py-0.1.1-py3-none-any.whl
Size 52.3 kB
Tags Python 3
SHA-256 checksum
How to use checksums
bcc04687eaa5aa81af9c6b4734c0edbdfe49a671c104971b988ba9fa984e4469
BLAKE2b-256 checksum
How to use checksums
7b3cdb7e0ca2279d314634454ceb8dc5b1c80cfcb8cb75125940c21f6b2ab303
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via uv/0.12.20 {"installer":{"name":"uv","version":"0.12.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

Release history Release notifications | RSS feed

This release

0.1.1 This release

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page