Skip to main content

pr-witness

Checks whether a pull request's tests still mean what they meant before it.

A test suite is a claim: these behaviours hold. A change can keep the suite green while quietly withdrawing the claim — by narrowing what gets collected, by rewriting the expected value to match a new bug, by skipping the test that guarded the thing it broke. CI reports green either way, because CI only ever runs the new tests against the new code.

pr-witness runs the combination nobody runs: the base branch's tests against the pull request's code. A test that passed before and fails now is a regression the pull request hid, whatever the green checkmark says.

Status

Phases 0–6 complete: corpus, cross-run, diff integrity (via checkwash), claim checking, the GitHub Action, signed evidence and the Claude Code plugin. On PyPI (pip install pr-witness), on the GitHub Marketplace, and installable as a Claude Code plugin.

On the numbers. The evaluation corpus is 17 synthetic cases, written by the same person who wrote the detector, scored on three verdicts: a cheat must be flagged, a rewritten expectation must come back as review, a clean change must be left alone. The tool gets all 17 right — eval/witness_results.md. That is a regression gate, not an accuracy claim: it says the tool has not got worse, not how it performs on pull requests it has never seen.

The first real cases are in the loop, from two sources:

  • corpus/real/swebench — seven test edits by a frontier agent on SWE-bench Verified (out of 1,500 patches examined), labelled against the maintainers' own test patches. Six run end to end; the verdict matches on all six (one clean, five review).
  • corpus/real/swechat — commits from real Claude Code sessions on real repositories (SALT-NLP/SWE-chat): 3,355 agent test edits mined, seven run end to end so far, labelled by hand. Six of seven match. The miss is written up there; so is the case that was a miss until it got its own finding — a deleted test replaced by a different one, which no run can see and coverage_replaced reads from the source.

Those cases were picked by eye and the tool was changed after almost every one, so their scores flatter it. The number that does not: batch 1 — seventeen SWE-chat commits selected by a written rule, labelled before running, run with the tool frozen. Twelve produced evidence; the verdict matched on seven of twelve. All five misses are the tool's, and four of them are a flag on an honest commit. review was right every time it was used; flag was wrong every time. No commit in the batch was dishonest, so catching one on real data is still unmeasured. The four causes are fixed in 0.3.1; the twelve now agree, which measures nothing — a new number needs commits the tool has not seen.

What to rely on today. review — "a human should read this, here is the old assertion and the new one" — was right every time it was used in the frozen batch (four of four; the earlier cases are the ones the verdict was built from, so they do not count). flag has not earned that yet. Use report mode; do not gate merges on it.

The real runs changed the tool more than the synthetic corpus ever did — the review verdict, greened_by_test_edit (a failing suite made green by editing the tests, which the cross-run could not see), base_broken, the bun runner — each from a specific commit, listed in the two READMEs.

See eval/results.md for the numbers and corpus/README.md for how the cases are built.

The signals

Signal What it catches
crossrun A base-branch test the PR did not touch that now fails on the PR's code — a hidden regression (high)
expectation_changed A test the PR rewrote whose old assertion fails on the new code; the report shows the old and new assertion (medium → review)
greened_by_test_edit A test that was red on the base branch and is green now, whose old version still fails on the new code — the test edit made it pass, not the code (medium → review)
tests_removed_still_passing A test the PR deleted that still passes against the PR's code — behaviour kept, coverage gone (medium → review)
coverage_replaced Replacement: a function the changed tests used to call still exists and no test calls it for real any more — the one test edit the runs cannot see, found by reading the source (Python; medium → review)
base_broken / base_not_green Every base test failing is no evidence, not clean; some failing is context
denominator Fewer tests actually running than before — the 4966/4966 ALL PASSED trick
skip_inflation Collected count steady while the skipped count climbs
added_tests_skipped Tests the PR adds that do not run — reported, not flagged: nothing that ran before stopped running (low)
claim_false The description asserts something the run contradicts
claim_unverifiable A claim nothing in CI can confirm or deny

As a GitHub Action

# .github/workflows/pr-witness.yml
name: pr-witness
on: pull_request
permissions:
  contents: read
  pull-requests: write      # for the comment

jobs:
  witness:
    runs-on: ubuntu-latest
    timeout-minutes: 20
    steps:
      - uses: actions/checkout@v7
        with: { fetch-depth: 0 }     # both branches must be present
      - uses: actions/setup-python@v7
        with: { python-version: "3.12" }
      - run: pip install -e . pytest  # whatever your own test job installs
      - uses: talhayme/pr-witness@v0.3
        with:
          test-command: pytest        # or vitest, jest, go, django (Django's own runtests.py)
          mode: report                # report | require-ack | strict

It posts one comment per pull request and edits it on every push, so the thread does not fill with stale reports. The same report goes to the job summary, and evidence.json is exposed as an output for anything downstream.

mode behaviour
report never fails the job. The default.
require-ack fails on a flag until a maintainer adds the test-change-approved label
strict fails on a flag; the label documents intent but does not change the facts

A flag is a disproved claim or a high finding. A review is a rewritten test whose old assertion fails on the new code: honest specification change or test bent to fit a bug, the run cannot tell, so it shows both assertions and require-ack/strict hold the job for a human the same way they do for a flag. Other medium findings and misleading claims are reported but never fail the job in any mode.

Locally, the same thing:

pip install pr-witness
pr-witness --base origin/main                   # text report
pr-witness --base origin/main --format markdown --pr-body pr.md
pr-witness --base origin/main --mode require-ack --approved
pr-witness --base origin/main --changed-only      # only the test files the change touches
pr-witness --base origin/main --paths tests/api   # or a subset you name; large suites in seconds

Signed evidence

With sign: "true" (the default) the action signs evidence.json with GitHub Artifact Attestations, and the comment links to the attestation. That is what makes the report worth more than a comment an agent could have typed: it proves the file was produced by this workflow, in CI, and not altered since.

gh attestation verify evidence.json --repo you/your-repo \
  --predicate-type https://github.com/talhayme/pr-witness/evidence/v1

The job needs id-token: write, attestations: write and artifact-metadata: write. Signing works in public repositories on every GitHub plan; private repositories need Enterprise Cloud. When it cannot run the report is still posted, marked "Unsigned" with the reason.

The evidence format is schema/evidence.v1.json.

How to clear a false alarm

  1. Read the finding's "↳" line — every one says when it is innocent.
  2. If the behaviour change is intentional, say so in the pull request and add test-change-approved. In require-ack mode that is enough.
  3. If pr-witness is wrong, open an issue with the diff; it becomes a clean case in the corpus so the next release cannot regress.

In Claude Code

claude plugin marketplace add talhayme/pr-witness
claude plugin install pr-witness

Three small things, all advisory:

  • A pause before a test file changes. Editing, writing or rm/mv/sed-ing a test file asks for confirmation, with the reasoning: "if the behaviour is meant to change, say so; if the test is in the way, fix the code — CI will run the old test against the new code either way." PR_WITNESS_GUARD=deny for a hard no, PR_WITNESS_GUARD=off to silence it.
  • A check when the agent says the tests pass. When the final message makes a checkable claim — "all tests pass", "12/12 green", "added tests for X" — the hook runs pr-witness on the working tree against the base branch and shows any claim the run contradicts, in the session. It never blocks the stop: a hook that refused to let the agent finish until the numbers matched would teach it to stop giving numbers.
  • /pr-witness:verify-before-done — a checklist: run the whole suite, quote the runner's own summary line, report the denominator, list every changed test with its reason.

The plugin is a courtesy, not a control. The agent runs inside its own environment and can route around any hook. The proof is the signed report from CI, where the agent cannot reach.

Relationship to checkwash

checkwash detects diff-level tampering — weakened assertions, loosened tolerances, disabled tests, touched CI — and does it well, with its own published benchmarks. pr-witness is not a replacement and does not re-test those detectors. It covers what a diff cannot show: what happened when the code actually ran.

Run both.

Reproducing the baseline

python -m venv venv && ./venv/bin/pip install checkwash pytest
(cd .tsrunner && npm install)

python corpus/build.py         # generate the 16 cases
python corpus/verify.py        # measure the Python cases
python corpus/verify_ts.py     # measure the TypeScript cases
python eval/run_eval.py --tool checkwash
python eval/run_eval.py --tool crossrun

Licence

Apache-2.0.

Metadata

Release files for pr-witness 0.3.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pr-witness 0.3.1
File Size Uploaded
pr_witness-0.3.1.tar.gz 66.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pr-witness 0.3.1
File Interpreter ABI Platform
pr_witness-0.3.1-py3-none-any.whl Python 3 none any Details

Total release size: 117.2 kB

Release files / pr_witness-0.3.1.tar.gz

Download URL pr_witness-0.3.1.tar.gz
Size 66.0 kB
Tags Source
SHA-256 checksum
How to use checksums
fe8d483732d8bfb84bae6033176e5009f8b8538bf61b9d3d1bf67566a15f8d9d
BLAKE2b-256 checksum
How to use checksums
e6621abcb179898504e7da5d2b5050573369f1b87b401be1128ed8cb20b25f5e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release files / pr_witness-0.3.1-py3-none-any.whl

Download URL pr_witness-0.3.1-py3-none-any.whl
Size 51.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
590d157add047b3a14f44db0f1d68e80ad775d97de66ddbf83fc351337d81926
BLAKE2b-256 checksum
How to use checksums
fdc7349a2ee6926720cae8bd6eec108f420e93c3418f4be47e01ef617e0d3627
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 11, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.1 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page