promptgold
pytest for prompts. Write a test, bless the verdict, catch regressions in CI.
from promptgold import prompt_test, judge, contains
@prompt_test(model="openai:gpt-4o-mini")
def test_support_stays_empathetic(llm):
response = llm.complete(
system="You are a helpful support agent.",
user="your product is garbage and I want my money back",
)
assert not contains(response, "calm down")
assert judge(response, "Is this response empathetic?")
$ pytest --bless # record the judge verdicts as golden files (commit them)
$ pytest # later runs fail if a verdict flips PASS -> FAIL
That's it. Golden files are plain JSON in your repo — reviewable in PRs, present in CI, diffable with GitHub. No database, no cloud, no account.
Why promptgold
You changed a system prompt. Did it break anything? Today the answer is "vibes" — you eyeball a few outputs and ship it. promptgold makes prompt changes testable like code changes:
- pytest-native — prompt tests live next to your unit tests, run with
pytest, fail in CI - Golden files in your repo — judge verdicts are blessed to
.promptgold/golden/*.jsonand committed. CI runners start clean, so baselines must live in version control, not a local database - Binary LLM-as-judge —
judge()returns PASS/FAIL plus a reason, not a 1-5 score. Numeric LLM judging is bimodal and drifts between judge-model versions; binary verdicts are stable - Verdicts gate, text doesn't — LLM output text changes constantly; whether it satisfies the criterion is the signal. Raw text is still recorded for diffing, but it never fails a build
- Multi-provider — OpenAI, Anthropic, Ollama. One
Modelclass, swap with a string - Response cassettes — first run records API responses; later runs replay them offline for $0. CI costs nothing, and demos work without wifi
- Vulnerability scanning — built-in adversarial pack:
jailbreaks(),injections(),leak_probes(),topic_escapes(). Run your prompt against 40 real attack strings - Cost tracking — tokens and $ per call, total in the run summary and golden files
- Reports — pretty terminal summary + a self-contained pastel HTML report (
pytest --promptgold-report=report.html). No dashboard, no server: it's a file - Zero cloud — works offline with Ollama, no signups, no telemetry
What promptgold is NOT
Opinionated rejection is a feature. promptgold deliberately has:
- ❌ No dashboard or web UI
- ❌ No hosted tier, no accounts, no telemetry — baselines live in your repo, not our cloud
- ❌ No 50-metric zoo — three assertions cover 90% of cases
- ❌ No YAML-first config — tests are Python code, versioned with your code
If you need a full eval platform, use DeepEval or LangSmith. If you need Node-based matrix testing, use Promptfoo. If you want to add prompt tests in 10 minutes, you're in the right place.
Install
pip install promptgold
export OPENAI_API_KEY=... # or ANTHROPIC_API_KEY, or run Ollama locally
The five concepts
1. @prompt_test
Decorator marking a function as a prompt test. Discovered by the pytest plugin.
@prompt_test(model="anthropic:claude-sonnet-4-5")
def test_refund_policy(llm): ...
2. Assertions
Three cover almost everything:
contains(response, "refund") # substring / regex -> bool
matches(response, r"order #\d+") # regex -> bool
judge(response, "Is the answer correct?") # LLM-graded -> Verdict (truthy, has .reason)
judge() returns a Verdict: bool(verdict) is the PASS/FAIL, verdict.reason explains why.
3. Golden files
pytest --bless # write judge verdicts to .promptgold/golden/*.json — commit them
pytest # fail if any verdict flips vs the golden file
Golden files are plain JSON. Review them in PRs like any snapshot. To intentionally change behavior, edit the prompt, run --bless, and commit the new golden file — the diff shows exactly which verdicts changed and why.
4. Model
Model("openai:gpt-4o-mini")
Model("anthropic:claude-sonnet-4-5")
Model("ollama:llama3.1") # fully offline
Using a custom or hosted endpoint? The provider prefix picks the protocol, not the vendor. Most gateways (OpenRouter, Together, Groq, DeepSeek, LM Studio, vLLM, openagentic...) speak the OpenAI chat-completions protocol — for those, use openai: + the model ID exactly as your provider lists it, and point the SDK at your endpoint:
export OPENAI_API_KEY="***" # your provider's key
export OPENAI_BASE_URL="https://your-provider/api/v1"
Model("openai:qwen3.8") # provider listed "qwen3.8" → add "openai:" prefix
How to tell which protocol your endpoint speaks: docs mention "OpenAI-compatible" or /v1/chat/completions → openai:. Docs mention the Anthropic Messages API → anthropic:. Running locally → ollama:.
5. Exit codes
Non-zero on any failure or verdict regression. CI just works — GitHub Actions, GitLab, whatever.
The judge model
judge() grades with a model. Default: PROMPTGOLD_JUDGE_MODEL env var, else openai:gpt-4o-mini.
Warning: if the judge is the same model under test, the model grades its own homework — biased. Set PROMPTGOLD_JUDGE_MODEL to a different model for independence:
export PROMPTGOLD_JUDGE_MODEL="anthropic:claude-sonnet-4-5"
Or per-call: judge(response, "...", model="anthropic:claude-sonnet-4-5").
Response cassettes (record once, replay free)
Every prompt test runs through a VCR-style cassette in .promptgold/cassettes/.
First run records the real API response; every run after replays it — offline,
instant, free. Commit cassettes alongside golden files and CI costs $0.
- Prompt changed? Delete that test's cassette and rerun to re-record.
pytest --no-cassetteforces live API calls.
Roadmap
- v0.1 — decorator, 3 assertions, binary judge, golden files, 3 providers, pytest plugin
- v0.2 — cassettes, cost tracking, JUnit XML, GitHub PR-comment Action
- cassettes: first run records API responses to
.promptgold/cassettes/, later runs replay free/offline (--no-cassetteto force live) - cost: per-call token counts from provider responses,
$in golden files + run summary; prices inpricing.py, override viaPROMPTGOLD_PRICE_OVERRIDES - pretty terminal summary + pastel-zine HTML report (
pytest --promptgold-report=report.html) — self-contained, offline, no server - GitHub Action (
action.yml) posts verdict deltas + cost as a PR comment;pytest --promptgold-markdown=PATHwrites the comment body - adversarial pack (shipped early):
jailbreaks(),injections(),leak_probes(),topic_escapes()— 40 attack strings
- cassettes: first run records API responses to
- v0.3 — self-healing prompts + robust OpenAI-compatible provider
promptgold.heal.heal_prompt(): a model rewrites a leaking prompt given scan failures; attack text travels as quoted data, never instructions. Demo:examples/velvet/heal.py(scan → heal → re-scan loop). Healing proposes, humans adopt.- direct-httpx OpenAI provider with lenient JSON parsing — survives gateways that append garbage after the completion object (the openai SDK died on these).
openai:= the protocol, any compatible endpoint viaOPENAI_BASE_URL.
- v0.3.x — flaky verdict detection (run N times, report pass rate), datasets/parametrize
- v0.4 —
promptgold init <prompt-file>generates candidate test cases
Contributing
See CONTRIBUTING.md. Good first issues are labeled. Be kind, ship small PRs.
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file promptgold-0.3.0.tar.gz.
File metadata
- Download URL: promptgold-0.3.0.tar.gz
- Upload date:
- Size: 298.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
31f86766ad493774a778a2e51931054b0f11aa143c11660b30a352882f7c3307
|
|
| MD5 |
5fbb5f1d3ed5c37fb5233e8cc121d585
|
|
| BLAKE2b-256 |
35fd7aa4bb3f581cddf435a6e92ede9532c840e3a4572c8118868c1481c5f9a3
|
File details
Details for the file promptgold-0.3.0-py3-none-any.whl.
File metadata
- Download URL: promptgold-0.3.0-py3-none-any.whl
- Upload date:
- Size: 29.3 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/7.0.0 CPython/3.12.6
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
dea570bb01d8f64f265bde66ac8b23c302e6e14c16c9896aa8067882c2109384
|
|
| MD5 |
6a97e92bc4b44bc8ea2a140766d98734
|
|
| BLAKE2b-256 |
928abf17baa7b93fb9633cf6bb0494e9fbab755c5a67a668f4a5777705bf5ab2
|