Skip to main content

wiretruth

wiretruth replays real LLM provider responses through LLM SDKs and checks whether each SDK reports what actually happened.

Does your LLM SDK tell you the truth?

CI PyPI License: MIT

Version 0.1 tests one library, any-llm, on the OpenAI Chat Completions API with three truncation scenarios. More providers and libraries follow, and a public results matrix arrives in v0.4. Until then, the grades are in results/.

Why

A model hits its output token limit halfway through a JSON object. The provider says so plainly: "finish_reason": "length". What your code sees depends on the SDK in between. If the SDK raises a generic JSON parse error, nothing tells you that a higher limit would fix it. If it returns the partial object as a success, your pipeline stores an incomplete record.

wiretruth has this case recorded as truncation.structured-output on openai-chat. Replayed through any-llm:

Library What the caller gets Grade
any-llm 1.33.0 Raises openai.LengthFinishReasonError. Its completion has finish_reason: "length" and the token usage. ✓ PASS

wiretruth runs the same recorded edge cases through every library it supports and grades each one against what the provider sent.

What it checks

Bug class What goes wrong for users Scenarios
Silent truncation Agents accept partial output or retry without raising the limit 3
Finish reason mapping Control flow branches on the wrong reason planned
Usage accounting Cost and quota tracking is off (cached and reasoning tokens) planned
Stream integrity Corrupted or unfinished text is treated as final planned
Tool call fidelity Tools run with broken or merged arguments planned
Error classification Retries and backoff misbehave planned
Request fidelity The library sends options the caller didn't set planned

The scenarios are in suite/scenarios/, and the bug classes in suite/bug-classes.json.

Grades

Grade Meaning
✓ PASS Reports what happened: meets every MUST requirement of the scenario.
~ LOSSY Doesn't contradict what happened, but the fact is only in raw metadata, or missing.
✗ WRONG Reports something that contradicts what happened.
! CRASH Fails with an internal error, such as a KeyError in its own parser.
⧗ HANG Doesn't return within 20 s on a response that ended within 1 s.
– UNSUPPORTED The library doesn't support this provider or feature (documented).
? HARNESS A problem on our side, not the library's.

Each scenario lists its requirements. MUST requirements set the grade. SHOULD requirements, such as the exact partial text in the truncation scenarios, add a note when they aren't met. How grades are computed: docs/grading.md.

Quickstart

You need Python 3.11 or newer and uv, which builds each library's own environment.

git clone https://github.com/kartsan03/wiretruth && cd wiretruth
uv run wiretruth run --lib any-llm
scenario                      openai-chat
truncation.structured-output  ✓ PASS
truncation.text               ✓ PASS
truncation.text-stream        ✓ PASS

any-llm 1.33.0: 3 cells, 3 PASS
stderr: .wiretruth/runs/20261010T193444.312067Z

No API keys are needed. Everything runs against recorded responses on 127.0.0.1. The first run downloads any-llm into .wiretruth/envs/; later runs work offline. uv run wiretruth run --shim shims/reference runs the built-in reference client, which must pass every cell. The scenarios, fixtures and shims live in this repository, so wiretruth runs from a clone.

Add your library

A shim is a small program. It reads a call spec as JSON on stdin, calls your library the way its docs show, and prints what the library reported as JSON on stdout. The any-llm shim is 150 lines. The protocol is in shims/README.md; to propose a library, open a new-library issue.

  • Open Responses: a spec and compliance tests for servers that implement the Responses API. wiretruth tests clients across several native APIs.
  • LLMConform: conformance tests for LLM gateways and APIs, run live.
  • llmprobe and conform: conformance for inference engines (vLLM, llama.cpp, Ollama, …).
  • aimock: a mock server for testing AI apps, with record and replay. Use it for your app's tests; wiretruth audits the SDK layer itself.
  • The design borrows from JSONTestSuite, Wycheproof, toml-test and compat-table.

Author note

I've contributed fixes to libraries that wiretruth tests or will test: mozilla-ai/any-llm#1389 and TanStack/ai#1548. Grades are computed by the code in this repository. If you think a grade is wrong, please open a grade dispute.

Contributing

Scenario ideas, new libraries and grade disputes are all welcome. See CONTRIBUTING.md.

Citing

See CITATION.cff.

License

MIT. Fixtures contain model outputs recorded from provider APIs; docs/recording.md explains how they're recorded.

Metadata

Release files for wiretruth 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for wiretruth 0.1.0
File Size Uploaded
wiretruth-0.1.0.tar.gz 66.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for wiretruth 0.1.0
File Interpreter ABI Platform
wiretruth-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 139.6 kB

Release files / wiretruth-0.1.0.tar.gz

Download URL wiretruth-0.1.0.tar.gz
Size 66.3 kB
Tags Source
SHA-256 checksum
How to use checksums
40bd9d350c934c508b531a2a581c4b0016451f44db2cc941c68cd6c5c78e4216
BLAKE2b-256 checksum
How to use checksums
6146a82a288e1c8fd5f9153dbf130e4a02161b0d4d7a5371a2a432d2c8736dbb
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release files / wiretruth-0.1.0-py3-none-any.whl

Download URL wiretruth-0.1.0-py3-none-any.whl
Size 73.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
dd1d70a0091a7ced893fde5135f6065259bd49ad1c61feb92f1b5680d89b14a5
BLAKE2b-256 checksum
How to use checksums
a15704832f8bb2289f5b4c7f3eb21be1c217a6cb1b082144a3423db4ea81f129
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page