Skip to main content

Decorator-based integration testing for LLM prompts

Project description

prompt-tester

Decorator-based integration testing for LLM prompts. Wrap a test function with @prompt_run, and it sends your prompt to a real model, then hands you the response to assert on — with an LLM judge for the assertions that plain string checks can't express.

When to use this

prompt-tester runs integration tests against real models with real API calls. It is not a mocking or unit-testing framework. It fits prompts that are part of a larger system:

  • Custom tools / function calling — verify the model calls the right tool with the right arguments under realistic conditions.
  • MCP-connected agents — test prompts that drive an agentic loop against your live MCP server, not a stub.
  • RAG pipelines — assert that retrieved context is actually used in the final response.
  • Multi-step workflows — validate that each prompt stage produces output fit for the next.

Because model output is non-deterministic, prompt-tester can run a prompt N times and assert a minimum pass rate — see Multi-run testing.

Installation

pip install prompt-tester            # core
pip install prompt-tester[anthropic] # + Anthropic SDK
pip install prompt-tester[google]    # + Google GenAI SDK
pip install prompt-tester[all]       # both providers

Requires Python 3.11+.

Quick start

from pathlib import Path
import prompt_tester
from prompt_tester import prompt_run, Model, Provider

prompt_tester.configure(
    judge_model    = Model.GEMINI_2_5_FLASH,
    judge_provider = Provider.GOOGLE,
)

@prompt_run(
    target_prompt    = "Summarise in one sentence: {text}",
    subject_model    = Model.GEMINI_3_1_FLASH_LITE,
    subject_provider = Provider.GOOGLE,
    target_prompt_vars = {"text": "Alice leads Project Phoenix. Budget: $2M. Deadline: Q3."},
)
def test_summary_mentions_alice(run):
    assert run.output                       # raw response text
    verdict = run.ask("Is Alice mentioned in the output?")
    assert verdict.passed, verdict.reasoning

The decorator injects a PromptRun as the first argument. Assert on run.output directly, or use run.ask(...) to have the judge answer a yes/no question about the response.

Configuration

Call configure() once before your tests run — at the top of a test module or in your test setup. Both parameters are required; there is no default judge. If configure() has not been called, @prompt_run raises ConfigurationError with setup instructions when a test runs.

import prompt_tester
from prompt_tester import Model, Provider

prompt_tester.configure(
    judge_model    = Model.GEMINI_2_5_FLASH,
    judge_provider = Provider.GOOGLE,
)
Function Description
configure(judge_model, judge_provider) Set the judge model and provider. Both required. Accepts Model/Provider enums or plain strings.
reset() Clear configuration and cached provider/judge instances.

API keys

ANTHROPIC_API_KEY=<your key>   # for Anthropic models
GOOGLE_API_KEY=<your key>      # for Gemini models

Keys are loaded automatically from a .env file in your project root (via python-dotenv), or from the environment directly.

@prompt_run

Runs a prompt and injects the result into your test function as a PromptRun. Assert on anything — raw output, token counts, cost, or judge verdicts. With runs > 1, each run executes in its own thread and the decorator asserts the pass rate once all runs finish (see Multi-run testing).

Parameter Type Default Description
target_prompt str required Prompt text. Use {key} placeholders for target_prompt_vars.
subject_model Model | str required Model to run the prompt against.
subject_provider Provider | str required Provider for subject_model: Provider.ANTHROPIC or Provider.GOOGLE.
target_prompt_vars dict {} Values substituted into {key} placeholders before the call. An unmatched {placeholder} raises a UserWarning.
max_tokens int 2048 Maximum output tokens for the prompt call.
runs int 1 Number of times to run the prompt. When > 1, each run executes in its own thread.
pass_threshold float 1.0 Fraction of runs that must pass (only meaningful with runs > 1). Uses math.ceil, so runs=5, pass_threshold=0.8 requires 4 passes.
enable_cache bool False Attach an ephemeral 5-minute cache to the judge system prompt. With runs > 1, the first run executes solo to prime the cache before the rest run in parallel. See Caching.
run_fn callable | None None Custom executor for tool use, MCP, or agentic loops. See run_fn.
@prompt_run(
    target_prompt    = PROMPT,
    subject_model    = Model.GEMINI_3_1_FLASH_LITE,
    subject_provider = Provider.GOOGLE,
    target_prompt_vars = {"text": INPUT},
)
def test_compactor_is_concise(run):
    # Raw output
    assert len(run.output) < len(run.target_prompt_vars["text"]) * 0.6

    # Judge verdict
    alice = run.ask("Is Alice mentioned in the output?")
    assert alice.passed, alice.reasoning

    # API metadata
    assert run.cost_usd is not None
    assert run.stop_reason in ("end_turn", "STOP")

PromptRun

The object injected into your test function. Access the response and metadata as fields; call the ask* methods to invoke the judge.

Methods

# One question — one API call
verdict = run.ask("Is Alice mentioned?")
assert verdict.passed, verdict.reasoning

# Multiple questions in one API call (cheapest)
alice, budget = run.ask_all([
    "Is Alice mentioned?",
    "Is the $2M budget mentioned?",
])

# Multiple questions, one API call each, fired concurrently
alice, budget = run.ask_parallel([
    "Is Alice mentioned?",
    "Is the $2M budget mentioned?",
])
Method API calls Wall-clock Cost Cross-contamination Returns
ask(q) 1 none JudgeVerdict
ask_all(qs) 1 for all fastest lowest low list[JudgeVerdict]
ask_parallel(qs) 1 per question, concurrent ~1 call higher (N calls) none list[JudgeVerdict]
  • ask_all — many independent checks where cost matters and slight cross-contamination between questions is acceptable.
  • ask_parallel — questions that must be fully isolated, without the latency of sequential ask calls.
  • ask — a single question, or when you need to branch on the result before asking the next.

ask_all and ask_parallel return [] immediately for an empty question list. If any ask_parallel call raises, the pool waits for in-flight calls to land, then raises a RuntimeError listing the failures.

Fields

Input

Field Type Description
prompt str Rendered prompt text — target_prompt_vars already substituted.
target_prompt_vars dict[str, Any] Key/value pairs substituted into the prompt.

Model response

Field Type Description
output str The model's response text.
stop_reason str | None Why generation stopped. Anthropic: "end_turn", "max_tokens", "stop_sequence". Gemini: "STOP", "MAX_TOKENS", "SAFETY", "RECITATION", "OTHER".
safety_filtered bool True if the provider blocked the response (output will be empty).

Model identity

Field Type Description
model str The subject_model value you passed.
provider str The subject_provider value you passed.
model_used str | None Actual model ID reported by the provider (may differ if an alias resolves).
model_version str | None Provider version string. Populated by Gemini; None for Anthropic.
request_id str | None Provider request ID for log correlation.

Token usage & cost

Field Type Description
input_tokens int Tokens in the prompt.
output_tokens int Tokens in the response.
cached_input_tokens int Input tokens served from cache. 0 when unused.
cache_creation_tokens int Tokens written to cache (Anthropic only).
thoughts_tokens int Reasoning tokens (Gemini thinking models only).
cost_usd float | None Total cost in USD. None if the model is not in the pricing table.
run.to_dict()   # all fields as a plain dict — safe to log or serialise

Multi-run testing

A single test run is a point-in-time sample, not a reliable signal — the model may have passed by chance or failed on noise. Set runs > 1 and a pass_threshold to require a minimum pass rate:

@prompt_run(
    target_prompt    = PROMPT,
    subject_model    = Model.GEMINI_3_1_FLASH_LITE,
    subject_provider = Provider.GOOGLE,
    target_prompt_vars = {"text": INPUT},
    runs             = 5,
    pass_threshold   = 0.8,   # at least 4 of 5 runs must pass
)
def test_compactor(run):
    alice, budget = run.ask_all([
        "Is Alice mentioned in the output?",
        "Is the $2M budget mentioned?",
    ])
    assert alice.passed,  alice.reasoning
    assert budget.passed, budget.reasoning

Execution model:

  • Each run gets its own thread. With enable_cache=False (default), all runs fire at once, so wall-clock time is roughly one run's worth.
  • With enable_cache=True, run 1 executes solo first to prime the judge cache; the remaining runs - 1 then fire in parallel as cache hits.
  • The required pass count is ceil(runs * pass_threshold) — with 5 runs and 0.8, you need 4.
  • Failed runs are printed with their run number ([run 3/5] FAILED: ...) before the final assertion, so you can see exactly which failed and why.

Tuning: raise pass_threshold toward 1.0 for hard requirements; lower it for prompts with known variance. Start at 0.8 and tighten once you have a baseline.

pytest.mark.parametrize alternative

For pytest to report each run as a separate test item, use parametrize with runs=1 (the default):

@pytest.mark.parametrize("_", range(5))
@prompt_run(
    target_prompt    = PROMPT,
    subject_model    = Model.GEMINI_3_1_FLASH_LITE,
    subject_provider = Provider.GOOGLE,
    target_prompt_vars = {"text": INPUT},
)
def test_compactor_always_concise(run, _):
    assert len(run.output) < len(run.target_prompt_vars["text"]) * 0.6

This gives individual pass/fail per run but no pass-rate control — one failure fails the whole group. Use it when each run has different inputs, or when you want strict all-or-nothing behaviour.

Caching

Caching is opt-in through enable_cache=True on @prompt_run, and it applies to the judge calls automatically — no extra wiring.

When enabled, Anthropic judge calls attach an ephemeral (5-minute) cache to the shared judge system prompt via cache_control. Repeated run.ask() / run.ask_all() calls within the window then hit the cache instead of re-reading the system text. To make those hits land in a multi-run test, run 1 executes solo first to prime the cache, and only then do the remaining runs fire in parallel. The cached_input_tokens and cache_creation_tokens fields reflect actual cache activity, and cost is computed accordingly (reads at 10% of the input rate, cache creation at 125%).

Anthropic only activates caching above a minimum token threshold — 4096 tokens for Opus and Haiku 4.5, 2048 for Sonnet 4.6, 1024 for older Sonnets. Below the threshold the cache_control block is silently ignored. The built-in judge system prompt is roughly 2,000–2,700 tokens, so use Sonnet 4.6 as the Anthropic judge to reliably cross the threshold; Opus would need a larger prompt to hit its 4096-token minimum.

Gemini caches server-side automatically and ignores this flag — its cached_input_tokens populate regardless.

Subject-model calls are cached only when a custom run_fn passes a cached system prompt itself.

run_fn

@prompt_run sends a single prompt and records the response. If your prompt drives an agentic loop — calling tools, querying an MCP server, or taking multiple model turns before a final answer — pass a run_fn to replace the built-in provider call entirely.

(prompt: str, model: str, max_tokens: int) -> CompletionResult

run_fn receives the rendered prompt and owns the full loop, including tool round-trips. It must return a CompletionResult whose text field holds the final output; the judge then evaluates that output via run.ask() as normal. CompletionResult requires only text, input_tokens, and output_tokens — all other fields (stop_reason, model_used, request_id, …) are optional and surface on the PromptRun for metadata and assertions.

When run_fn is set, subject_provider is metadata only — recorded on the PromptRun but not used to pick an SDK. Your run_fn makes the actual API calls.

The goal is not merely "did a tool get called" — it is asserting on the quality of the final answer the whole agentic process produced. Tool-call checks are one assertion among many.

import anthropic
from prompt_tester import prompt_run, Model, Provider, CompletionResult

MCP_TOOLS = [...]   # tool schemas from your MCP server

def run_with_mcp(prompt: str, model: str, max_tokens: int) -> CompletionResult:
    client   = anthropic.Anthropic()
    messages = [{"role": "user", "content": prompt}]
    total_input = total_output = 0

    while True:
        response = client.messages.create(
            model=model, max_tokens=max_tokens, tools=MCP_TOOLS, messages=messages,
        )
        total_input  += response.usage.input_tokens
        total_output += response.usage.output_tokens

        if response.stop_reason != "tool_use":
            final_text = next(b.text for b in response.content if hasattr(b, "text"))
            return CompletionResult(
                text          = final_text,
                input_tokens  = total_input,
                output_tokens = total_output,
                stop_reason   = response.stop_reason,
                model_used    = response.model,
                request_id    = response.id,
            )

        # Execute tool calls and feed results back
        tool_results = []
        for block in response.content:
            if block.type == "tool_use":
                result = mcp_client.call_tool(block.name, block.input)  # noqa: F821
                tool_results.append({
                    "type": "tool_result", "tool_use_id": block.id, "content": result.content,
                })
        messages.append({"role": "assistant", "content": response.content})
        messages.append({"role": "user",      "content": tool_results})


@prompt_run(
    target_prompt    = "Use the search tool to find the capital of France, then summarise it.",
    subject_model    = Model.CLAUDE_SONNET_4_6,
    subject_provider = Provider.ANTHROPIC,
    run_fn           = run_with_mcp,
)
def test_agent_answers_correctly(run):
    # The judge evaluates the final answer, after all tool calls have completed.
    correct, cited = run.ask_all([
        "Does the response state that Paris is the capital of France?",
        "Does the response mention a source or search result?",
    ])
    assert correct.passed, correct.reasoning
    assert cited.passed,   cited.reasoning
    assert run.stop_reason == "end_turn"
    assert run.input_tokens > 0

Cost tracking

Token counts and USD cost are on every PromptRun and every JudgeVerdict:

run.input_tokens, run.output_tokens, run.cost_usd    # the prompt call
verdict.judge_input_tokens, verdict.judge_cost_usd   # a judge call

Prices live in prompt_tester/config.py in the _PRICES table, keyed by Model. cost_usd is None for any model not in the table. To add one, add a Model member and a matching _PRICES entry.

Judge verdicts

The judge answers a yes/no question about the response. It classifies each question into one of seven types and applies type-specific rules, so you can phrase assertions naturally:

Type Purpose Example
A. Factual presence Something is present "Does the response mention Paris?"
B. Avoidance Something is absent "Does the response avoid mentioning competitor pricing?"
C. Format / validity Output has a format property "Is the output valid JSON?"
D. Structure / count Counts or measures hold "Does the response have exactly 3 bullet points?"
E. Tone / style A holistic trait applies "Is the tone professional?"
F. Comparison / preference A preference is expressed "Does it recommend Python over JavaScript?"
G. Multi-part Every listed part is covered "Does the response cover both setup and teardown?"

For avoidance (type B), the logic is inverted: a "yes" verdict means the forbidden content was not found. Numbers, dates, abbreviations, plurals, and similar variants are treated as equivalent where sensible (5k = 5,000, 2024-01-15 = January 15, 2024, API = its expansion).

JudgeVerdict fields

Field Type Description
question str The question you asked.
answer str "yes" or "no".
passed bool True when the answer is "yes".
snippet str | None Verbatim excerpt the judge cited as evidence.
reasoning str | None The judge's stated reasoning — most useful as a failure message.
judge_model str Model ID of the judge.
judge_provider str | None Provider of the judge.
judge_input_tokens int Input tokens for this verdict.
judge_output_tokens int Output tokens for this verdict.
judge_cost_usd float | None Cost for this verdict.
verdict.to_dict()   # all fields as a plain dict

Models reference

Model and Provider enums with their pricing (USD per million tokens). Plain model-ID strings work anywhere an enum does.

Model Provider Model ID Input Output
CLAUDE_HAIKU_4_5 ANTHROPIC claude-haiku-4-5-20251001 $0.80 $4.00
CLAUDE_SONNET_4_6 ANTHROPIC claude-sonnet-4-6 $3.00 $15.00
CLAUDE_OPUS_4_7 ANTHROPIC claude-opus-4-7 $15.00 $75.00
GEMINI_2_5_FLASH GOOGLE gemini-2.5-flash $0.15 $0.60
GEMINI_2_5_PRO GOOGLE gemini-2.5-pro $1.25 $10.00
GEMINI_2_0_FLASH GOOGLE gemini-2.0-flash $0.10 $0.40
GEMINI_2_0_FLASH_LITE GOOGLE gemini-2.0-flash-lite $0.075 $0.30
GEMINI_3_1_FLASH_LITE GOOGLE gemini-3.1-flash-lite $0.25 $1.50

Architecture

prompt_tester/
  __init__.py   Public API: prompt_run, render_prompt, PromptRun, JudgeVerdict,
                CompletionResult, Model, Provider, configure, reset
  config.py     Model/Provider enums, pricing table, configure() / reset()
  decorator.py  @prompt_run — runs the prompt, injects PromptRun into the test
  models.py     Pydantic types: PromptRun, JudgeVerdict, UsageMetrics
  judge.py      Judge.evaluate() / evaluate_all() — LLM-as-judge
  llm/providers/
    base.py        LLMProvider ABC + CompletionResult (unified return type)
    anthropic.py   Anthropic SDK wrapper
    gemini.py      Google GenAI SDK wrapper

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

prompt_tester-0.5.0.tar.gz (41.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

prompt_tester-0.5.0-py3-none-any.whl (29.0 kB view details)

Uploaded Python 3

File details

Details for the file prompt_tester-0.5.0.tar.gz.

File metadata

  • Download URL: prompt_tester-0.5.0.tar.gz
  • Upload date:
  • Size: 41.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.4

File hashes

Hashes for prompt_tester-0.5.0.tar.gz
Algorithm Hash digest
SHA256 a6e3945e83e9dbd74287c5328976584ddc55c7db24d83f63b990f6a1fa53030f
MD5 3b3440fa8e5847bf31fc99a06a4e1d04
BLAKE2b-256 ba9fd3a045609663337c2d9ab521c3523bee4287006cb8fbd6d10b506f903fff

See more details on using hashes here.

File details

Details for the file prompt_tester-0.5.0-py3-none-any.whl.

File metadata

  • Download URL: prompt_tester-0.5.0-py3-none-any.whl
  • Upload date:
  • Size: 29.0 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.4

File hashes

Hashes for prompt_tester-0.5.0-py3-none-any.whl
Algorithm Hash digest
SHA256 34804f25d50a48a293de97abd83fe4a255ee2f45baddf85b46167c04f1661d6e
MD5 cde3e2dfc888b4e31b05dbd4bd87a698
BLAKE2b-256 d9faf8e7d55d09165835e55f25661d99cd99aa2014713e4a1a22297465c1bc4f

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page