Skip to main content

Callwitness

Record every tool call an AI agent makes. Block nothing.

tests

A transparent MCP proxy. It sits between an agent and its tools, forwards every byte unchanged, and writes down what happened.

No dependencies. Python 3.8+. MIT.


The thing it shows you

Same tool. Same permission. Two very different actions:

      51B  send_email   ops@acme.com
    20085B send_email   exfil.example.net, drop@unknown.example

An allowlist cannot tell those apart — the agent is permitted to send email in both cases. The difference is how much is leaving and where it is going, and those are the two signals Callwitness records on every call.

Why it blocks nothing

Because it should be installable in production on a Tuesday afternoon.

Callwitness cannot corrupt what an agent sends or receives: it relays every message whether or not it can parse it, and every write to storage is wrapped so a recorder bug can't reach the stream. That property is tested, not asserted — see tests/test_passthrough.py, which asserts the proxied output is byte-identical to running the server directly.

It cannot delay, either. Observation runs on its own thread behind a bounded queue, so the relay only ever does a non-blocking hand-off. A deliberately half-second-slow observer moves the gap between two forwarded messages by 0.05ms — it used to move it by 4.2 seconds. The queue drops rather than growing without limit under load, and counts what it dropped: unrecorded data nobody can see is worse than data that was never collected.

It matters because the security industry is currently writing rules against agent failures nobody has measured. Enforcement without data is guessing with extra steps. Collect first.

Install

pip install -e .

Use

Wrap any stdio MCP server:

callwitness run --echo -- npx -y @modelcontextprotocol/server-filesystem /data

Or let it wrap the servers you already have. It finds your client's config, shows you exactly what would change, and writes nothing until you say so:

$ callwitness install

/Users/you/Library/Application Support/Claude/claude_desktop_config.json
  filesystem
    - npx -y @modelcontextprotocol/server-filesystem /data
    + callwitness run --label filesystem -- npx -y @modelcontextprotocol/server-filesystem /data
  git
    - uvx mcp-server-git --repository /repo
    + callwitness run --label git -- uvx mcp-server-git --repository /repo
  remote-api  SKIPPED: remote server -- needs `callwitness proxy --upstream
              https://mcp.acme.com/mcp --port <port>` and a port you choose

2 servers would be wrapped. Nothing has been changed.

--apply writes it, after a timestamped backup. callwitness uninstall --apply puts everything back. Running install twice does nothing the second time.

Dry-run is the default because this edits a file you did not write and a broken MCP config means a broken agent — the one outcome this whole tool promises not to cause. Anything it does not recognise is skipped and named rather than guessed at.

Knows about Claude Desktop, Cursor, Windsurf, Claude Code, and project-local .mcp.json / .vscode/mcp.json. If yours lives elsewhere: callwitness install --config /path/to/mcp.json.

Remote servers

Production agents mostly talk to remote MCP servers over Streamable HTTP. Put Callwitness in front of one and point the client at the local address instead:

callwitness proxy --upstream https://mcp.example.com/mcp --port 8100 --echo
{
  "mcpServers": {
    "example": { "url": "http://127.0.0.1:8100/mcp" }
  }
}

POST, the SSE response stream, the server-initiated GET stream and session teardown are all relayed verbatim, headers included, so the Mcp-Session-Id handshake works without Callwitness understanding it. Both transports share one recorder (CallTracker), so a row looks the same whichever produced it.

The agent behaves exactly as before. Then look at what it did:

callwitness stats            # per-tool volume, errors, latency, destinations
callwitness tail -n 20       # the most recent calls
callwitness verify           # check nothing has been altered since it was written
callwitness export out.jsonl # everything, for analysis

Then let it write the rules

The next tier was going to be YAML you write by hand. But a person typing max_payload: 8KB for send_email is guessing at a number they have no way to know — which is the thing this project says the industry is doing wrong. Enforcement without data is guessing with extra steps, and a rule language is not data. So the rules come out of the observation tier instead:

$ callwitness suggest --since 14d

  send_email     max_payload            3.5KB    # p99 observed 2.3KB over n=400; 1.5x headroom
  send_email     destinations_emails    3 allowed # 3 distinct emails covering 100% of traffic over n=400
  send_email     rate_limit_per_hour    21       # 10.0/hour average over 39.9 hours; 2x headroom
? fetch_url      destinations_hosts     --       # 99 distinct hosts across 150 calls -- too varied for an
                                                 #   allowlist; this reads as a general-purpose fetcher
? delete_record  insufficient_data      --       # only 6 calls observed; 30 needed before a threshold
                                                 #   means anything

--format yaml emits the same thing as a policy draft, every rule commented with the evidence it rests on.

Note what it refuses to do. A tool below 30 calls gets no threshold, because a p99 over n=6 is an anecdote. A tool whose destinations are too varied is flagged for a human rather than handed an allowlist that would fire constantly. And a destination that was never seen is not a destination that is forbidden — it may simply not have happened yet, and the output says so rather than letting you forget it. Every line is a hypothesis with its evidence attached, not a finding.

Baseline poisoning. If the bad thing already happened while Callwitness was watching, it is in the distribution, and a plain percentile quietly raises the ceiling to permit it. The demo above showed exactly that: a 29KB exfiltration produced a 43KB proposed ceiling — one that would have allowed the very call this tool exists to catch.

So ceilings come from the bulk of a distribution, not all of it. Calls far above the median set no limit; they are named, with timestamps and destinations, and handed to a person:

  send_email  max_payload   4.3KB  # p99 of the bulk is 2.9KB; 1.5x headroom.
                                   #   EXCLUDES 1 call above 11.4KB
! send_email  tail_review   1      # 1 call more than 8x the 1.4KB median. A rare
                                   #   enormous call is the most interesting thing
                                   #   here, so it sets no limit until you have
                                   #   looked at it: 28.8KB at 2026-09-09T18:09
                                   #   -> exfil.example.net

The reference is the median, because it is the one statistic a single enormous call cannot move — which is the point when that call may be the attack. If more than 10% of traffic sits above the threshold it is not a tail, it is the shape, and nothing is excluded; misdescribing the distribution is a different failure, and just as wrong.

This is not a solution to baseline poisoning. Nothing that learns from unlabelled traffic has one. It is a refusal to hide it.

Try it without an agent

No agent, no API key, no network, no Node:

python examples/demo.py                     # throwaway run, nothing kept
python examples/demo.py --keep --repeat 40  # record into your own store
callwitness suggest                             # then let it propose rules

The plain run uses a temporary directory so trying the tool doesn't pollute anyone's data — but the obvious next thing to type is callwitness stats, and "No data yet" is a bad first hour. --keep records into ~/.callwitness, and --repeat sends enough varied traffic that suggest has a distribution to work from rather than an anecdote.

Privacy

Flag Effect
(default) Credentials in argument values are redacted before storage
--no-redact Stores argument values verbatim, credentials included
--no-args Stores argument shape only ({"to": "<str:20>"}), never values
--max-arg-bytes N Caps stored bytes; the true size is still recorded
--home DIR Where data lives (default ~/.callwitness)

Redaction is on by default. Tool arguments routinely carry API keys, bearer tokens and connection strings, and without this every install would be a plaintext credential store that didn't exist before Callwitness was installed. Known key formats, credentials inside URLs, sensitively-named parameters and high-entropy tokens are replaced with <redacted:reason> on the write path — never on read, because by then the plaintext is already on disk. The true pre-redaction byte count is still recorded, so the volume signal survives.

Destinations survive redaction on purpose: postgres://admin:hunter2@db.internal stores as postgres://admin:<redacted:url_password>@db.internal. The host is the signal; the password is not.

--no-args still records destinations — hosts, emails, IPs — because destinations are the signal. That's deliberate, it's tested, and you should say it out loud to anyone you ask to run this.

Everything stays on the machine that ran it. Nothing is transmitted anywhere.

Evidence, not just a log

An append-only file is trivially editable by anyone with filesystem access — including a compromised agent running as the same user. A record that can be silently rewritten is a convenience, not evidence.

So every call commits to the one before it. Editing, deleting, reordering or inserting a record breaks the chain from that point, and callwitness verify says where:

$ callwitness verify
BROKEN  filesystem  642 records, breaks at seq 118
                    content does not match its hash: this record was edited
                    after it was written

Exit code 1 on a break, so it works in a cron job without anyone parsing text.

It is tamper-evident, not tamper-proof, and the tool says so out loud. Someone who can write to the file can also recompute every hash after a change and produce a chain that verifies — nothing local can stop that, because the verifier and the attacker read the same file. What defeats it is an anchor the operator does not control, so verify prints the head hash and tells you to store it somewhere the machine cannot reach. That is a deployment decision, and inventing one for you would be worse than naming the gap.

Records written before chaining existed are reported as predating it, not as tampering. A verifier that cries wolf on an upgraded install is worse than no verifier.

What gets stored

calls — one row per tool call: tool, arguments, args_bytes, whether they were truncated, signals, duration_ms, is_error, result_bytes, and a result preview.

signals splits two things a naive scan conflates:

{
  "destinations":   {"emails": ["archive@unknown-host.example"],
                     "hosts":  ["exfil.example.net"]},
  "content_counts": {"emails": 400}
}

Destinations are entities found in routing fields — to, url, webhook, attach_url and so on. Content counts are how many entities appear in the payload. Scanning the whole blob for email addresses would report four hundred customer emails from inside a message body as "destinations" and bury the one address the message is actually addressed to. These are different signals and they compose: 29KB addressed to an unknown host, containing 400 email addresses is a shape worth stopping. 29KB containing 400 addresses, sent to the CRM you always use is a Tuesday.

eventsinitialize and tools/list, so you know which tools were exposed.

sessions — one row per wrapped process, with the exit code.

SQLite at ~/.callwitness/callwitness.db, plus an append-only calls.jsonl.

Design rule

The recorder must never corrupt the protocol stream, and must never delay it. Every message is forwarded first, then handed to a background queue; parsing happens on another thread, inside a try. If recording throws, traffic still flows. If recording is slow, traffic still moves.

If you contribute, keep it that way. test_a_broken_recorder_never_raises and tests/test_hardening.py are there to make sure you do.

The experiment

experiments/ runs the measurement this tool exists to make possible: 15 tasks x 4 injection channels, an agent with real tools and real side effects, all traffic recorded through Callwitness itself.

python experiments/run.py --driver scripted --out runs/pilot --fresh --repeats 4
python experiments/analyze_runs.py runs/pilot

That validates the pipeline with no API key and no network. Swap --driver llm with a Groq free-tier key for the real thing. Protocol and design are in experiments/README.md.

Where this is going

  1. Now — observe. Record every call, block nothing.
  2. Next — deterministic policy: the rules callwitness suggest proposes, evaluated inline, sub-millisecond, fail-open by default. The generator ships first on purpose; an engine that enforces numbers nobody could justify is the problem, not the product.
  3. Then — context: an LLM judge, but only on calls the deterministic tier flags. Payload volume × destination reputation first.

Scope: MCP tool calls over stdio and Streamable HTTP. The deprecated two-endpoint HTTP+SSE transport is not covered. Direct API calls made inside agent code need an SDK wrapper, and that is deliberately not in v1.

Where it came from

Out of an experiment on chain-of-thought faithfulness (cot-hint-verbalization), which turned up a measurement problem: "hint verbalisation rate" reads 100% on the reasoning trace and 12% on the user-facing answer, for the same responses. Same data, same model, an order of magnitude apart depending only on where you look.

A field whose headline metric moves by 10x depending on the instrument does not need another opinion about agent risk. It needs somebody to start writing down what actually happens.

Tests

pip install -e ".[dev]"
pytest

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

callwitness-0.1.0.tar.gz (76.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

callwitness-0.1.0-py3-none-any.whl (47.1 kB view details)

Uploaded Python 3

File details

Details for the file callwitness-0.1.0.tar.gz.

File metadata

  • Download URL: callwitness-0.1.0.tar.gz
  • Upload date:
  • Size: 76.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for callwitness-0.1.0.tar.gz
Algorithm Hash digest
SHA256 03adeac8ee9ce82c6283d1c663903677b3a4c0fb9045334ef990b6a95feacc14
MD5 0429834c1e136769aceeb13f56ea6f76
BLAKE2b-256 fe51c7a6896d06d698b223d962d5a6c219fa2645f6d5979593b10903d8d83789

See more details on using hashes here.

Provenance

The following attestation bundles were made for callwitness-0.1.0.tar.gz:

Publisher: release.yml on AditiChaudharyy14/callwitness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file callwitness-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: callwitness-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 47.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for callwitness-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ed17d5d4b04eb031dcb261a48e2f0f9fbe5fbc40c790208db0dac9686764ba6a
MD5 83342a67ce206242c9d0c237dc8f5912
BLAKE2b-256 bbbaf250b580e037e1d31e09516b6266b05eba160fe105687be8fe3ee82b4937

See more details on using hashes here.

Provenance

The following attestation bundles were made for callwitness-0.1.0-py3-none-any.whl:

Publisher: release.yml on AditiChaudharyy14/callwitness

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page