promptdrift
Know what your prompt change actually did.
Prompt regression testing with CI gating. Snapshot-diff first, flake-aware by design, Python native.
Teams that ship LLM features edit prompts constantly, and every edit can quietly break behavior that used to work. promptdrift catches that: it re-runs the same cases against your prompt, compares pass rates against a baseline snapshot you commit to git, and fails CI when behavior got statistically worse. Noise never trips the gate.
Quick tour, no API key needed
pipx install promptdrift
promptdrift demo
The demo runs a complete loop offline on a scripted mock provider:
runexecutes a suite (3 samples per case)approverecords the behavior as the baseline- the mock answer gets "improved" and loses its structured steps
diffcatches the drift and exits 1, the way CI would
step 4/4 — diff: catch the drift
┌──────────────┬─────────────────────┬────────┬───────┬────────────────┐
│ case │ assertion │ before │ after │ verdict │
├──────────────┼─────────────────────┼────────┼───────┼────────────────┤
│ refund-steps │ contains-申请退款 │ 3/3 │ 0/3 │ 🔴 regressed │
│ │ regex-3-5-个工作日 │ 3/3 │ 0/3 │ 🔴 regressed │
│ invoice-json │ is_json-yes │ 3/3 │ 3/3 │ ✅ stable │
└──────────────┴─────────────────────┴────────┴───────┴────────────────┘
❌ gate FAILED — 1 regressed case(s)
Why plain assertions aren't enough
LLM output is non-deterministic. A CI check like "output must contain X" fails randomly, gets retried, and eventually gets disabled. The verdict engine handles the noise instead:
| Verdict | Meaning | Gate behavior |
|---|---|---|
🔴 regressed |
pass rate fell beyond sampling noise (Wilson score intervals separate) | fails regression (default) |
⚠️ unstable |
samples disagree where the baseline was deterministic | fails only flaky |
🟢 improved |
pass rate rose beyond noise | never fails |
✅ stable |
within noise | passes |
🆕 new / ➖ removed |
case or assertion added or removed since the baseline | fails any-fail when imperfect |
Concretely: at samples: 3, a 3/3 → 2/3 drop is ⚠️ unstable, a warning. At samples: 10,
10/10 → 3/10 is 🔴 regressed. Raise the sample count to sharpen the statistics; run
promptdrift cost first to see what that costs.
How it fits together
promptest.yaml ──► Runner ──► Run ──► Differ ◄── Snapshot (.promptest/baselines/*.snap.yaml)
│ ▲
└── provider ──► sqlite cache
Baselines are YAML files you commit. When a PR touches one, that diff is the behavior change under review, the same way jest snapshots work. A config fingerprint (model, sampling params, sample count, judge model) protects comparability: swap the model and the old baseline goes stale instead of being silently mis-compared. Editing prompt text, vars, or mock answers is never staleness; that is exactly the change the diff should measure.
One provider adapter covers every OpenAI-compatible endpoint (OpenAI, GLM, DeepSeek, Qwen,
Moonshot, vLLM, Ollama), configured with base_url.
Suite file
$schema: https://raw.githubusercontent.com/ritaprieto900/promptdrift/main/schema/promptest.schema.json
suite: support-agent
provider:
openai_compat:
model: glm-4.7
base_url: https://open.bigmodel.cn/api/paas/v4
api_key_env: ZHIPUAI_API_KEY
pricing: {prompt_per_1m: 1.0, completion_per_1m: 8.0}
judge: # optional: grades `judge` assertions
openai_compat:
model: glm-4.7-flash
base_url: https://open.bigmodel.cn/api/paas/v4
api_key_env: ZHIPUAI_API_KEY
temperature: 0.1
samples: 3
cases:
- id: refund-steps
messages:
- role: user
content: "订单有点问题,我想申请{{topic}}"
vars:
topic: 退款
assertions:
- contains: "申请退款"
- regex: "3-5 个工作日"
- not_contains: "抱歉"
- latency_under: 3000
- judge:
rubric: |
回复是否包含明确的退款步骤(入口、操作、时限)?
5 = 步骤完整且语气友好;1 = 没有回答问题。
min_score: 4
Deterministic assertions: equals, contains, not_contains, regex, is_json,
json_schema, latency_under, completion_tokens_under. The judge assertion sends each
sample to the judge model with your rubric and expects a 1-5 score; min_score sets the
pass bar. Judges are only human (well, only models), so their scores feed the same
multi-sample statistics as everything else: a wavering judge shows up as ⚠️ unstable, not as
a blocked merge.
Commands
| Command | Purpose |
|---|---|
promptdrift init |
scaffold a starter suite (mock provider, runs offline) |
promptdrift run |
execute once; exit 1 if any assertion failed |
promptdrift approve |
record current behavior as the baseline |
promptdrift diff |
run and compare against the baseline; this is the CI gate |
promptdrift show |
inspect the latest run or the baseline |
promptdrift cost |
estimate calls and cost (when pricing is configured) |
promptdrift demo |
the offline end-to-end walkthrough |
promptdrift schema |
JSON Schema for suite files (editor completion) |
diff options that matter in CI: --fail-on regression|flaky|any-fail,
--require-baseline, and --md report.md to write a PR-ready markdown report.
Exit codes: 0 pass, 1 gate or assertion failure, 2 config or runtime error.
# GitHub Actions, minimal version (see docs/ci.md for the full one)
- run: uv tool install git+https://github.com/ritaprieto900/promptdrift
- run: promptdrift diff --require-baseline --md pr-report.md
env:
OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}
How it compares
| promptdrift | promptfoo | DeepEval | |
|---|---|---|---|
| Language | Python | Node/TypeScript | Python |
| Core primitive | baseline snapshot diff | eval runs + web viewer | pytest metrics |
| Non-determinism | multi-sample + interval statistics | per-run assertions | metric thresholds |
| Review flow | snapshot diff in the PR | open the web UI | read pytest output |
promptfoo is the broader eval platform and DeepEval is strong on metric research. This project stays narrow: gate prompt changes in CI, with a verdict you can trust.
Roadmap
- selective re-runs of changed cases only; embedding-similarity assertion
- composite GitHub Action + sticky PR comments once the PyPI release lands
- multi-suite projects, cross-model comparison mode
Development
git clone https://github.com/ritaprieto900/promptdrift
cd promptdrift
uv sync
uv run pytest # 150+ tests, fully offline
uv run ruff check .
uv run pyright
The whole test suite runs on the mock provider, so CI never touches a real endpoint. CONTRIBUTING.md has the architecture map and the design invariants.
License
Metadata
Release files for promptdrift-py 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| promptdrift_py-0.1.0.tar.gz | 173.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| promptdrift_py-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 225.7 kB
Release files / promptdrift_py-0.1.0.tar.gz
| Download URL | promptdrift_py-0.1.0.tar.gz |
|---|---|
| Size | 173.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
15d3e4df0795fcb8efe018cd8406f5a8ebbd5146cd00dbd07191d23b6c3f03f1
|
|
BLAKE2b-256 checksum How to use checksums |
076774b791834f39bfc119d5186d56e7bce7ea5d5d76b04ae89de53fa18bc076
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.20 {"installer":{"name":"uv","version":"0.12.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|
Release files / promptdrift_py-0.1.0-py3-none-any.whl
| Download URL | promptdrift_py-0.1.0-py3-none-any.whl |
|---|---|
| Size | 52.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7114b9a6d54d62639045531d392635165a91dc2aee40053340862732e9e6fd4d
|
|
BLAKE2b-256 checksum How to use checksums |
efdd37a682b4341ed0e5c84ed1f5ebf987fb18818da858266fae357a465f6b6b
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
uv/0.12.20 {"installer":{"name":"uv","version":"0.12.20","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":null,"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
|