wiretruth
wiretruth replays real LLM provider responses through LLM SDKs and checks whether each SDK reports what actually happened.
Does your LLM SDK tell you the truth?
Version 0.1 tests one library, any-llm, on the OpenAI Chat Completions API with three truncation scenarios. More
providers and libraries follow, and a public results matrix arrives in v0.4. Until then, the grades are in
results/.
Why
A model hits its output token limit halfway through a JSON object. The provider says so plainly:
"finish_reason": "length". What your code sees depends on the SDK in between. If the SDK raises a generic JSON
parse error, nothing tells you that a higher limit would fix it. If it returns the partial object as a success,
your pipeline stores an incomplete record.
wiretruth has this case recorded as truncation.structured-output on openai-chat. Replayed through any-llm:
| Library | What the caller gets | Grade |
|---|---|---|
| any-llm 1.33.0 | Raises openai.LengthFinishReasonError. Its completion has finish_reason: "length" and the token usage. |
✓ PASS |
wiretruth runs the same recorded edge cases through every library it supports and grades each one against what the provider sent.
What it checks
| Bug class | What goes wrong for users | Scenarios |
|---|---|---|
| Silent truncation | Agents accept partial output or retry without raising the limit | 3 |
| Finish reason mapping | Control flow branches on the wrong reason | planned |
| Usage accounting | Cost and quota tracking is off (cached and reasoning tokens) | planned |
| Stream integrity | Corrupted or unfinished text is treated as final | planned |
| Tool call fidelity | Tools run with broken or merged arguments | planned |
| Error classification | Retries and backoff misbehave | planned |
| Request fidelity | The library sends options the caller didn't set | planned |
The scenarios are in
suite/scenarios/, and the bug classes in
suite/bug-classes.json.
Grades
| Grade | Meaning | |
|---|---|---|
| ✓ | PASS | Reports what happened: meets every MUST requirement of the scenario. |
| ~ | LOSSY | Doesn't contradict what happened, but the fact is only in raw metadata, or missing. |
| ✗ | WRONG | Reports something that contradicts what happened. |
| ! | CRASH | Fails with an internal error, such as a KeyError in its own parser. |
| ⧗ | HANG | Doesn't return within 20 s on a response that ended within 1 s. |
| – | UNSUPPORTED | The library doesn't support this provider or feature (documented). |
| ? | HARNESS | A problem on our side, not the library's. |
Each scenario lists its requirements. MUST requirements set the grade. SHOULD requirements, such as the exact partial text in the truncation scenarios, add a note when they aren't met. How grades are computed: docs/grading.md.
Quickstart
You need Python 3.11 or newer and uv, which builds each library's own environment.
git clone https://github.com/kartsan03/wiretruth && cd wiretruth
uv run wiretruth run --lib any-llm
scenario openai-chat
truncation.structured-output ✓ PASS
truncation.text ✓ PASS
truncation.text-stream ✓ PASS
any-llm 1.33.0: 3 cells, 3 PASS
stderr: .wiretruth/runs/20261010T193444.312067Z
No API keys are needed. Everything runs against recorded responses on 127.0.0.1. The first run downloads any-llm
into .wiretruth/envs/; later runs work offline. uv run wiretruth run --shim shims/reference runs the
built-in reference client, which must pass every cell. The scenarios, fixtures and shims live in this repository,
so wiretruth runs from a clone.
Add your library
A shim is a small program. It reads a call spec as JSON on stdin, calls your library the way its docs show, and prints what the library reported as JSON on stdout. The any-llm shim is 150 lines. The protocol is in shims/README.md; to propose a library, open a new-library issue.
Related work
- Open Responses: a spec and compliance tests for servers that implement the Responses API. wiretruth tests clients across several native APIs.
- LLMConform: conformance tests for LLM gateways and APIs, run live.
- llmprobe and conform: conformance for inference engines (vLLM, llama.cpp, Ollama, …).
- aimock: a mock server for testing AI apps, with record and replay. Use it for your app's tests; wiretruth audits the SDK layer itself.
- The design borrows from JSONTestSuite, Wycheproof, toml-test and compat-table.
Author note
I've contributed fixes to libraries that wiretruth tests or will test: mozilla-ai/any-llm#1389 and TanStack/ai#1548. Grades are computed by the code in this repository. If you think a grade is wrong, please open a grade dispute.
Contributing
Scenario ideas, new libraries and grade disputes are all welcome. See CONTRIBUTING.md.
Citing
See CITATION.cff.
License
MIT. Fixtures contain model outputs recorded from provider APIs; docs/recording.md explains how they're recorded.
Metadata
Release files for wiretruth 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| wiretruth-0.1.0.tar.gz | 66.3 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| wiretruth-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 139.6 kB
Release files / wiretruth-0.1.0.tar.gz
| Download URL | wiretruth-0.1.0.tar.gz |
|---|---|
| Size | 66.3 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
40bd9d350c934c508b531a2a581c4b0016451f44db2cc941c68cd6c5c78e4216
|
|
BLAKE2b-256 checksum How to use checksums |
6146a82a288e1c8fd5f9153dbf130e4a02161b0d4d7a5371a2a432d2c8736dbb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.
Transparency logRelease files / wiretruth-0.1.0-py3-none-any.whl
| Download URL | wiretruth-0.1.0-py3-none-any.whl |
|---|---|
| Size | 73.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
dd1d70a0091a7ced893fde5135f6065259bd49ad1c61feb92f1b5680d89b14a5
|
|
BLAKE2b-256 checksum How to use checksums |
a15704832f8bb2289f5b4c7f3eb21be1c217a6cb1b082144a3423db4ea81f129
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.
Transparency log