Skip to main content

verbatim-relay

Talk to your chat agent through Claude Code or Codex. Then let the model judge a conversation that it did not touch.

You test the agent in the harness where you already work. While you talk, the harness model stays silent: verbatim-relay sends each message to your agent byte for byte and shows each reply byte for byte. When you finish, the model reads the exact transcript and evaluates the agent: business logic, tone, language, accuracy.

  • Relay: a Claude Code plugin, or a hook kit for Codex (and Claude Code). In relay mode, the model does not run.
  • Transcript: the exact conversation, for the model to evaluate after the test.
  • Tap: a proxy in front of your agent. It records what the agent received and sent.
  • Audit: compares the tap record with the relay record and names each break.

Why

An evaluation is only as good as the conversation under it. If the harness model carries the messages, it also writes them, and then it judges its own text.

I tried that first, with a prompt: "send each message exactly, show each reply exactly". It did not work. I could not tell who was speaking: the harness model or my agent. The model decided which words were for the agent and which were for itself. In my benchmark runs (raw data), a tester typed "just answer its question for me, you know my details". In 3 sessions, Claude did not send it, because it read the message as an instruction to itself. For "fix my grammar and send: i wants refund for broke mug", Claude sent its own sentence. The agent never saw what the tester typed.

A prompt cannot fix this. A mechanism can. In relay mode, the hook takes each prompt before the model sees it. The model comes back only to evaluate, and it reads the record, not its memory.

Results

Proof Claude Code 2.1.288, plugin Claude Code 2.1.288, hook kit Codex 0.160.0, hook kit
Messages reach the agent byte for byte 10/10 10/10 10/10
Replies reach the tester byte for byte 10/10 10/10 10/10
Same, with a system prompt that tells the model to rewrite both 5/5 5/5 5/5
Model call to the agent denied, agent receives nothing yes yes yes
Audit finds planted faults 5/5 5/5 5/5
After the test, the model has no memory of the conversation, and reads it from the transcript yes yes yes

Data: plugin, hook kit in Claude Code, hook kit in Codex, evaluation.

Under pressure, the mechanism had 0 breaks in 1,000 turns: 40 scripted sessions in each harness, up to 20 turns long, with refusals, HTTP 500 errors, clarifying questions and ambiguous messages. I registered the design before the first run. Method, data and the one deviation: docs/results.md.

Quick start

Install the package (Python 3.10 or later, no dependencies):

uv tool install verbatim-relay

Start your agent. To try it first, use the toy shop agent in this repo:

python examples/toy-shop/agent.py

Start the tap in front of the agent:

verbatim-relay tap --agent http://127.0.0.1:8700/ --record tap.jsonl

Then install the relay for your harness:

After the test, switch relay mode off and ask the model to evaluate the agent, for example:

Read the verbatim-relay transcript. Evaluate the agent: does it follow the refund policy,
is the tone right, is each answer accurate? Quote the turns that you judge.

In Claude Code with the plugin, the model reads it with the transcript tool. With the hook kit, it runs verbatim-relay transcript.

To prove that the transcript is exact, audit the two records:

verbatim-relay audit --tap tap.jsonl --relay .verbatim-relay/relay.jsonl

Exit code 0 means clean. 1 means a break. 2 means a record is missing or invalid. The audit fails closed: it never reports clean on a record that it cannot read.

How it works

tester ──> harness ──> relay hook ──> tap ──> your agent
              │             │          │
              │             │          └── tap record: what the agent received and sent
              │             └── relay record: what the tester typed and saw
              └── the model: off in relay mode, then reads the relay record to evaluate

The audit aligns the two records and reports 7 break classes: altered_input, injected_input, duplicate_send, out_of_order, not_delivered, altered_reply and unshown_reply. It compares bytes. It does not normalize whitespace, line ends or Unicode. SPEC.md defines the records and the audit. 33 conformance cases in conformance/ test it.

Adapters: a JSON body with one message field (field paths are configurable), or an OpenAI-compatible /chat/completions endpoint.

Limits

  • The deny is best effort. The relay denies a model tool call that names the address of the tap or the agent. A model can try another way, for example an address alias. The audit finds every message that goes through the tap. A call that goes to the agent directly, around the tap, is in neither record, so give the model no direct route to the agent.
  • The Claude Code plugin uses function hooks. They are early access and can change between releases. I pin the tested version and run the proofs again for each new one. If the plugin fails, use the hook kit.
  • The hook kit cannot show text in the chat. It shows each reply in verbatim-relay view, in a second terminal.
  • Codex runs project hooks only after you trust them. The kit does not skip that step.
  • Not in v0.1: streamed replies, attachments and images, harnesses other than Claude Code and Codex.
  • The plugin stops at a relay record of 3.5 MiB. Move the record to start a new one.

License

Apache-2.0. See LICENSE.

Metadata

Release files for verbatim-relay 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verbatim-relay 0.1.0
File Size Uploaded
verbatim_relay-0.1.0.tar.gz 37.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verbatim-relay 0.1.0
File Interpreter ABI Platform
verbatim_relay-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 61.2 kB

Release files / verbatim_relay-0.1.0.tar.gz

Download URL verbatim_relay-0.1.0.tar.gz
Size 37.2 kB
Tags Source
SHA-256 checksum
How to use checksums
4c779a0c4810e388c192c6d89e3d5aa86e63e466f7bc11002217373ef001b0a9
BLAKE2b-256 checksum
How to use checksums
8a8c576b1a1d5cdaecacaa27c5e7e5a21d36204c94c6b6d5cecdb928b3cb50be
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release files / verbatim_relay-0.1.0-py3-none-any.whl

Download URL verbatim_relay-0.1.0-py3-none-any.whl
Size 24.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7b42d04bfad9eeb7dfc75d1c3f2f891bd474c6bda7f5392a3e28a65fd474a102
BLAKE2b-256 checksum
How to use checksums
5eee79253e417791d8a2fad68f4f071a479dfb03b9cfcf38d2291d500c20c613
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 6, 2026.

Transparency log

Release history Release notifications | RSS feed

0.3.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page