LLM Preflight
Last reviewed: 2026-08-31 · As of: v2.9.0
Know whether an AI-generated LLM integration is safe before it reaches production. LLM Preflight is a local contract preflight for LLM integration changes: model, prompt, structured-output, and provider-call changes. It runs a small cross-provider preflight and compares validated output, response speed, tokens, and estimated cost.
Try it in 60 seconds
Create and run a deterministic local benchmark—no API key or network request:
python3 -m pip install llm-preflight
llm-preflight init
llm-preflight benchmark.json --no-save
From a source checkout:
python3 -m llm_preflight init
python3 -m llm_preflight benchmark.json --no-save
init never overwrites an existing config. It creates a mock benchmark so
you can see the report and exit behavior before making a paid request.
Its result is intentionally inconclusive (exit code 3): a local mock
validates configuration and output handling, but cannot approve a live model.
Choose your path
- Validate a change. Compare an approved model, prompt, schema, or provider route with a candidate using the model-change guide.
- Review a new model. Discover metadata, deliberately probe a route, then prepare a bounded candidate smoke with the model-catalogue guide.
- Automate an established contract. Add the no-spend GitHub Marketplace Action or use CI and JSON output.
Safety boundary
flowchart LR
A[Integration change] --> B[No-spend validation\ndoctor, pricing, dry run]
B --> C{Human reviews\nevidence and cost bound}
C -->|Explicit approval| D[Bounded paid smoke]
C -->|No approval or missing evidence| E[Inconclusive: fix or stop]
D --> F[Local evidence for\nproduction approval]
LLM Preflight is local evidence, not production approval. It is not a hosted evaluation platform, tracing system, RAG framework, or public leaderboard. Its results apply to your account, network, prompts, and validation rules.
[!WARNING] Live benchmarks make paid API requests. Start with the no-key demo, preview the plan before a live run, and keep limits and repetitions small.
What is new in 2.9.0
- Use a safe GitHub Action. The Marketplace Action runs doctor, pricing, and dry-run checks without provider traffic by default. A paid smoke needs an explicit workflow input and credentials supplied by the caller. See the Action guide.
- Report drift safely. Provider-breakage and pricing-drift issue forms ask for redacted reproduction metadata, never credentials or private prompts.
- Choose the right tool. Read when to use LLM Preflight for its boundary with evaluation, observability, and provider tools.
For earlier releases, see the changelog.
Purpose
Mission: make every LLM integration change evidence-based before production.
Vision: AI-assisted software delivery where an agent can validate its LLM changes as routinely as it runs tests, while people retain control of spend and production approval.
Positioning: LLM Preflight is the fast, local, cross-provider contract preflight for AI-powered application changes. It is not a general evaluation, observability, or autonomous-deployment platform.
It is built for engineers and coding agents working on AI features: teams that need to check a real application contract against live model APIs before a model ID, prompt, parser, tool definition, or provider option ships. Read the north star and the AI implementation testing guide for the intended workflow and boundaries.
Common jobs
-
Switch a model or provider. Run the bounded migration check, then add the contract test your feature needs.
-
Check a prompt, schema, parser, or tool change. Define an explicit output contract before the smoke.
-
Review a newly discovered model. Refresh metadata, then prepare—not run— a bounded candidate plan:
llm-preflight catalog refresh benchmarks/watch.json llm-preflight catalog prepare benchmarks/watch.json \ --against benchmarks/approved.json --output benchmarks/candidates.json llm-preflight benchmarks/candidates.json --migration-check --dry-run
Only explicitly approved, fully evidenced models proceed to paid work; see the model catalogue guide.
-
Investigate a provider or price change. Run
--doctor,--pricing-check, and a dry-run; report a suspected regression through the redacted issue forms. -
Automate a known contract. Use the no-spend GitHub Action or the CI guide with a saved baseline and
--ci.
It measures deterministic test validity, end-to-end latency (p50/p95), time to first token, throughput when the stream is incremental and usage is available, token totals, and estimated cost. Result files retain request metadata and per-request observations for reproducibility.
"Deterministic" describes the validator, not the model: every response is checked against explicit structural rules — a regular expression, a JSON shape, an exact routing label — so the same response always produces the same verdict. The tool does not score semantic quality; that is your task-specific evaluation, and it stays out of scope on purpose.
What a live run reports
Real output from a cross-provider run (2026-08-14, one short support prompt, three repetitions per model, total spend $0.051867):
| Model | Success | Latency p50 | Latency p95 | TTFT p50 | Tokens/s p50 | Cost |
|---|---|---|---|---|---|---|
| gpt-5.6-luna | 100% | 1.525s | 1.654s | 0.977s | 150.9 | $0.000350 |
| claude-opus-5 | 100% | 6.052s | 6.991s | 2.088s | 67.0 | $0.022805 |
| claude-sonnet-5 | 100% | 4.068s | 4.277s | 1.928s | 90.2 | $0.006092 |
| gemini-3.7-flash | 100% | 2.244s | 3.060s | 2.119s | 3273.0 | $0.005689 |
| grok-4.6 | 100% | 9.734s | 9.955s | 7.889s | 53.5 | $0.003096 |
| deepseek/deepseek-v4-pro-0813 | 100% | 6.883s | 10.460s | 5.003s | 116.5 | $0.000891 |
Tokens/s reads n/a when a provider delivers the response as a terminal
burst instead of an incremental stream — the observable window measures
transport, not generation, so no rate is reported. Cost reads n/a when
pricing for the model is unknown.
The report ends with an executive summary:
- Fastest: gpt-5.6-luna — 1.560s mean latency.
- Cheapest: gpt-5.6-luna — $0.000350 total.
- Best value: gpt-5.6-luna — 100% composite score.
- Recommended: gpt-5.6-luna — passed every selected test and led the
qualified value ranking.
- Total spent: $0.051867 including warmups.
Numbers like these are evidence for one environment at one time, not a leaderboard. Latency depends on your network and region; run the preflight from the host that will serve production traffic.
The same comparison can be driven interactively — pick models and tests at the terminal, read the cost ceiling before anything is sent, watch each request report its own cost, and end on the decision. This capture is a real two-model paid run that cost $0.005404 (config, details):
First live run
Python 3.10+ is required. There are no third-party runtime dependencies:
pip install llm-preflight installs this package and nothing else, and the
CLI runs on the Python standard library alone. Development tools (pytest,
ruff, mypy) are optional extras that never reach a production install.
cp benchmark.example.json benchmark.json
cp .env.example .env.production
# Edit benchmark.json and add only the provider keys you use.
python3 -m llm_preflight benchmark.json --dry-run
python3 -m llm_preflight benchmark.json
The CLI reads .env.production beside the config without overriding environment
variables already set by your shell. Use --no-env-file or --env-file PATH
when needed. Runs print a terminal report and, unless --no-save is used,
write JSON and Markdown results under results/.
Install the command globally in a virtual environment if preferred:
python3 -m pip install llm-preflight
llm-preflight --init
Run --doctor and --dry-run before the final command. They make no generation
requests; the final command is the paid work.
Change a model safely
This is the core workflow. Put your approved model and candidate model in one config, then run the small response-and-contract preflight:
llm-preflight benchmark.json --migration-check --dry-run
llm-preflight benchmark.json --migration-check
It sends three short representative cases to each selected model, once each. It answers: did the API work, did each response meet the basic contract, and how quickly did the provider start and finish responding? It is a cheap compatibility check, not a statistical performance conclusion.
When that passes, run the task-specific checks that match your application—for
example exact-routing-check or structured-output-check—before approving a
switch.
Use custom contract tests to express the outputs your
own feature must preserve.
Using a coding agent
Give an agent the same evidence you would use yourself: a reviewed config, an explicit output contract, and a dry run before paid work. Start with the recommended five-check suite:
# No generation request: inspect credentials, model selection, and paid-work plan.
llm-preflight benchmark.json --doctor --json
llm-preflight benchmark.json --tests agent-smoke --smoke --dry-run --json
# Paid run, only after reviewing the plan.
llm-preflight benchmark.json --tests agent-smoke --smoke --json --no-save
An agent should not infer model IDs, weaken a validator to turn a failure into a pass, or approve a model without an explicit instruction. The compact LLM and coding-agent guide covers commands, result JSON, exit codes, and automation guardrails. The AI implementation testing guide shows how to make this validation an agent's default testing step.
MCP for coding agents
Use the local stdio MCP server when an agent needs the preflight evidence without shell parsing or arbitrary command execution:
{
"mcpServers": {
"llm-preflight": {
"command": "llm-preflight-mcp",
"args": ["--workspace", "/absolute/path/to/repository"]
}
}
}
It exposes only four tools: validate a config, prepare a dry-run plan, run an explicitly confirmed preflight, and compare saved baselines. The first, second, and fourth tools never contact providers or load credentials. A live run still needs an explicit paid-run confirmation. See the MCP server guide for tool semantics, workspace boundaries, and the safe agent workflow.
Useful commands once you know your path
# Inspect configuration, credentials, and model selection without generation.
# --doctor provides pricing advisory; use --pricing-check as the fail-closed coverage gate.
llm-preflight benchmark.json --doctor
llm-preflight benchmark.json --pricing-check
llm-preflight benchmark.json --dry-run
# Run a reduced live benchmark.
llm-preflight benchmark.json --smoke
# Run a single ad hoc prompt.
llm-preflight --quick "Return only valid JSON with a status field." \
--models openai:gpt-5.4-mini
For advanced discovery, interactive runs, CI, baselines, replay, and stop modes, see workflows. For models, environment files, custom prompts, and provider-specific options, see configuration.
What makes a comparison useful
- Keep prompts, system instructions, temperature, and output limits fixed.
- Validate outputs: a fast malformed response is a failed result.
- Run from the same host; network distance and provider load affect latency.
- Treat single-user latency and load testing as separate experiments.
- Prefer dated model IDs over moving aliases.
The CLI distinguishes API FAIL (transport, credentials, provider, or request
failure) from API OK / TEST FAIL (a response that fails your validator).
Recommendations only consider models that pass every selected test.
How it compares
Several good tools live near this space. Use them when their job is your job:
- promptfoo, deepeval — full evaluation suites: scored quality metrics, red-teaming, large ongoing test matrices in CI. Use them to grade prompt and model quality over time.
- Braintrust, LangSmith — hosted platforms: tracing, dashboards, team collaboration, production observability.
llm(Simon Willison) — a general multi-provider CLI for running prompts, not a comparison harness.
LLM Preflight does one narrower job: the local go/no-go check in the moment before an LLM integration change. Your prompt, candidate models, structural validation, latency, and cost — one command, one report, no hosted service, no telemetry, and no vendor between you and the verdict.
Documentation
Start at the documentation homepage, then choose the path that matches your work:
- Start safely: safe demo and model change.
- Validate a change: output contracts, model catalogue, and pricing and safety.
- Automate: CI and JSON output, coding agents, and MCP.
- Look up details: CLI reference, configuration, result JSON, and troubleshooting.
- Understand the product: north star, product decisions, and AI implementation testing.
- Contributing — development setup and the TDD workflow.
- Security — reporting vulnerabilities.
Contributing and license
Contributions are welcome; see CONTRIBUTING.md. Released under the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file llm_preflight-2.9.0.tar.gz.
File metadata
- Download URL: llm_preflight-2.9.0.tar.gz
- Upload date:
- Size: 93.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c1147251d01b128f631bba1aa90eb231c9b71dec7daee87601bdfa9ca0d43f28
|
|
| MD5 |
8350243217ea98924f90596db40b7dd1
|
|
| BLAKE2b-256 |
5bdbd5b1e0d45a9ff9d0a8d226e3a7ebd86cad72372c5ec51c7e7fea3e758aed
|
Provenance
The following attestation bundles were made for llm_preflight-2.9.0.tar.gz:
Publisher:
release.yml on feronovak/llm-preflight
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_preflight-2.9.0.tar.gz -
Subject digest:
c1147251d01b128f631bba1aa90eb231c9b71dec7daee87601bdfa9ca0d43f28 - Sigstore transparency entry: 2661532868
- Sigstore integration time:
-
Permalink:
feronovak/llm-preflight@ecd4fc78e921bfdab0e1077815bb2692e5531413 -
Branch / Tag:
refs/tags/v2.9.0 - Owner: https://github.com/feronovak
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@ecd4fc78e921bfdab0e1077815bb2692e5531413 -
Trigger Event:
release
-
Statement type:
File details
Details for the file llm_preflight-2.9.0-py3-none-any.whl.
File metadata
- Download URL: llm_preflight-2.9.0-py3-none-any.whl
- Upload date:
- Size: 83.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b91061ff1a530f867e0bac275b874d426f3b802507b3bb6e717e735c3376a42e
|
|
| MD5 |
ef97216919a8e7f9cbdc98c971939f48
|
|
| BLAKE2b-256 |
ef45292f7b90ca3796c12182991df6e5d472888a311b48e4d3dde5dcd7940d2f
|
Provenance
The following attestation bundles were made for llm_preflight-2.9.0-py3-none-any.whl:
Publisher:
release.yml on feronovak/llm-preflight
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
llm_preflight-2.9.0-py3-none-any.whl -
Subject digest:
b91061ff1a530f867e0bac275b874d426f3b802507b3bb6e717e735c3376a42e - Sigstore transparency entry: 2661533003
- Sigstore integration time:
-
Permalink:
feronovak/llm-preflight@ecd4fc78e921bfdab0e1077815bb2692e5531413 -
Branch / Tag:
refs/tags/v2.9.0 - Owner: https://github.com/feronovak
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@ecd4fc78e921bfdab0e1077815bb2692e5531413 -
Trigger Event:
release
-
Statement type: