verbatim-relay
Talk to your chat agent through Claude Code or Codex. Then let the model judge a conversation that it did not touch.
You test the agent in the harness where you already work. While you talk, the harness model stays silent: verbatim-relay sends each message to your agent byte for byte and shows each reply byte for byte. When you finish, the model reads the exact transcript and evaluates the agent: business logic, tone, language, accuracy.
- Relay: a Claude Code plugin, or a hook kit for Codex (and Claude Code). In relay mode, the model does not run.
- Transcript: the exact conversation, for the model to evaluate after the test.
- Tap: a proxy in front of your agent. It records what the agent received and sent.
- Audit: compares the tap record with the relay record and names each break.
Why
An evaluation is only as good as the conversation under it. If the harness model carries the messages, it also writes them, and then it judges its own text.
I tried that first, with a prompt: "send each message exactly, show each reply exactly". It did not work. I could not tell who was speaking: the harness model or my agent. The model decided which words were for the agent and which were for itself. In my benchmark runs (raw data), a tester typed "just answer its question for me, you know my details". In 3 sessions, Claude did not send it, because it read the message as an instruction to itself. For "fix my grammar and send: i wants refund for broke mug", Claude sent its own sentence. The agent never saw what the tester typed.
A prompt cannot fix this. A mechanism can. In relay mode, the hook takes each prompt before the model sees it. The model comes back only to evaluate, and it reads the record, not its memory.
Results
| Proof | Claude Code 2.1.288, plugin | Claude Code 2.1.288, hook kit | Codex 0.160.0, hook kit |
|---|---|---|---|
| Messages reach the agent byte for byte | 10/10 | 10/10 | 10/10 |
| Replies reach the tester byte for byte | 10/10 | 10/10 | 10/10 |
| Same, with a system prompt that tells the model to rewrite both | 5/5 | 5/5 | 5/5 |
| Model call to the agent denied, agent receives nothing | yes | yes | yes |
| Audit finds planted faults | 5/5 | 5/5 | 5/5 |
| After the test, the model has no memory of the conversation, and reads it from the transcript | yes | yes | yes |
Data: plugin, hook kit in Claude Code, hook kit in Codex, evaluation.
Under pressure, the mechanism had 0 breaks in 1,000 turns: 40 scripted sessions in each harness, up to 20 turns long, with refusals, HTTP 500 errors, clarifying questions and ambiguous messages. I registered the design before the first run. Method, data and the one deviation: docs/results.md.
Quick start
Install the package (Python 3.10 or later, no dependencies):
uv tool install verbatim-relay
Start your agent. To try it first, use the toy shop agent in this repo:
python examples/toy-shop/agent.py
Start the tap in front of the agent:
verbatim-relay tap --agent http://127.0.0.1:8700/ --record tap.jsonl
Then install the relay for your harness:
- Claude Code: docs/claude-code.md
- Codex: docs/codex.md
After the test, switch relay mode off and ask the model to evaluate the agent, for example:
Read the verbatim-relay transcript. Evaluate the agent: does it follow the refund policy,
is the tone right, is each answer accurate? Quote the turns that you judge.
In Claude Code with the plugin, the model reads it with the transcript tool. With the hook kit, it runs verbatim-relay transcript.
To prove that the transcript is exact, audit the two records:
verbatim-relay audit --tap tap.jsonl --relay .verbatim-relay/relay.jsonl
Exit code 0 means clean. 1 means a break. 2 means a record is missing or invalid. The audit fails closed: it never reports clean on a record that it cannot read.
How it works
tester ──> harness ──> relay hook ──> tap ──> your agent
│ │ │
│ │ └── tap record: what the agent received and sent
│ └── relay record: what the tester typed and saw
└── the model: off in relay mode, then reads the relay record to evaluate
The audit aligns the two records and reports 7 break classes: altered_input, injected_input, duplicate_send, out_of_order, not_delivered, altered_reply and unshown_reply. It compares bytes. It does not normalize whitespace, line ends or Unicode. SPEC.md defines the records and the audit. 33 conformance cases in conformance/ test it.
Adapters: a JSON body with one message field (field paths are configurable), or an OpenAI-compatible /chat/completions endpoint.
Limits
- The deny is best effort. The relay denies a model tool call that names the address of the tap or the agent. A model can try another way, for example an address alias. The audit finds every message that goes through the tap. A call that goes to the agent directly, around the tap, is in neither record, so give the model no direct route to the agent.
- The Claude Code plugin uses function hooks. They are early access and can change between releases. I pin the tested version and run the proofs again for each new one. If the plugin fails, use the hook kit.
- The hook kit cannot show text in the chat. It shows each reply in
verbatim-relay view, in a second terminal. - Codex runs project hooks only after you trust them. The kit does not skip that step.
- Not in v0.1: streamed replies, attachments and images, harnesses other than Claude Code and Codex.
- The plugin stops at a relay record of 3.5 MiB. Move the record to start a new one.
License
Apache-2.0. See LICENSE.
Metadata
Release files for verbatim-relay 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| verbatim_relay-0.1.0.tar.gz | 37.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| verbatim_relay-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 61.2 kB
Release files / verbatim_relay-0.1.0.tar.gz
| Download URL | verbatim_relay-0.1.0.tar.gz |
|---|---|
| Size | 37.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4c779a0c4810e388c192c6d89e3d5aa86e63e466f7bc11002217373ef001b0a9
|
|
BLAKE2b-256 checksum How to use checksums |
8a8c576b1a1d5cdaecacaa27c5e7e5a21d36204c94c6b6d5cecdb928b3cb50be
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency logRelease files / verbatim_relay-0.1.0-py3-none-any.whl
| Download URL | verbatim_relay-0.1.0-py3-none-any.whl |
|---|---|
| Size | 24.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
7b42d04bfad9eeb7dfc75d1c3f2f891bd474c6bda7f5392a3e28a65fd474a102
|
|
BLAKE2b-256 checksum How to use checksums |
5eee79253e417791d8a2fad68f4f071a479dfb03b9cfcf38d2291d500c20c613
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.
Transparency log