Skip to main content

TestGuard

CI npm PyPI node license deps

Proves that a test suite actually defends the claims a project makes — by injecting the faults those claims say cannot happen, and reporting every fault the tests fail to detect.

Not a test generator. A claim verifier. Test generation is what happens after a claim turns out to be unfalsifiable.

Third tool following the Guard pattern, alongside docguard-cli (docs ↔ code) and websec-validator (attack surface ↔ code). All three run one loop:

declare what must be true → try mechanically to falsify it → freeze a baseline → gate only the delta → brief the agent before it writes code.

Why

Coverage cannot tell a test that pins correct behaviour from one that pins a defect. An agent that writes both the code and its tests encodes whatever it believed — including its bugs — and the suite goes green.

Measured on a real, entirely AI-authored production codebase with ~4,900 disciplined tests (no snapshots, 0.4% zero-assertion): 8 of 9 real historical bugs were invisible to the suite, worst case 2,451 tests green on known-broken code. The largest gap was a compliance-critical path with 100% coverage, where the one assertion that mattered used expect.objectContaining({...}) and omitted the field carrying the data.

A second, independent run on a different AI-authored codebase (63 test files, 458 tests, 24 hand-written security claims, 39 faults): 21 of 39 faults survived a fully green suite — 9 of them critical. Super-admin gating, membership checks, cookie flags and the whole authorization callback could be disabled without a single test noticing. One test file had re-implemented the authorization logic inside the test and asserted against the copy: fifteen green tests, zero detection. After wrapper-level tests were written against the survivors, 39/39 were killed.

Install

How Command
npx (no install) npx testguard-cli probe
npm npm i -D testguard-cli then npx testguard probe
pip pip install testguard-cli then testguard probe (needs Node ≥ 20)
Homebrew brew tap raccioly/tap && brew install testguard
GitHub Action uses: raccioly/testguard@v0.4.0 — see action.yml
pre-commit repo: https://github.com/raccioly/testguard, hooks testguard-claims, testguard-probe

Projects that set min-release-age in .npmrc cannot see a version published less than that many days ago (ENOVERSIONS); install that one with npm i -D testguard-cli --min-release-age=0.

How it works

npx testguard-cli init        # install the agent layer: skill, session-start hook, AGENTS.md section
npx testguard-cli status --json   # where the project is and the ONE next action — the machine entry point
npx testguard-cli claims      # what does this project claim, and is every claim probeable?
npx testguard-cli probe       # try to falsify each claim; report what the tests missed
npx testguard-cli baseline    # freeze today's unproven findings; from now on only new ones gate
npx testguard-cli brief       # tell the agent where the suite is blind, before it writes
npx testguard-cli scaffold src/x.ts   # propose faults for a file, as a draft to keep or drop
  1. Claims live in testguard.claims.json (editors validate it against "$schema": "./node_modules/testguard-cli/spec/schemas/claims.schema.json"): a statement, where it comes from, which tests supposedly defend it, and one or more faults — each a deterministic source change that would make the statement false. Every claim and every fault records who produced it. testguard claims validates the file and reports drift against @claim <ID> annotations in source. Test files are deliberately not scanned — a claim asserted by a test is the authorship trap the tool exists for — and annotation ids must contain a hyphen so prose is never mistaken for one.

  2. Probe confirms the defenders are green N times unmodified, applies each fault in a scratch git worktree (your tree is never touched), runs the defenders N times, re-runs survivors against the whole suite with N-run attribution, restores, and classifies. Verdicts are a closed set:

    Verdict Meaning
    killed a test body rejected the behaviour, N/N — the only pass
    SURVIVED the defenders stayed green while the claim was false
    NOCOVER no test file defends the claim at all
    UNVERIFIABLE the fault's anchor is missing or ambiguous — loud, never a skip
    TIMEOUT the defenders hung; a hang is not a detection
    FAULT-INVALID the replacement does not load — a bad fault, not a finding
    FLAKY-DEFENDER the defenders are not reliably green, or disagreed across runs

    Never a single score. Findings are ranked by severity, claim provenance and blast radius (relative imports, tsconfig path aliases and package.json#imports are resolved; bare package names are not), and written to .testguard/evidence.json — validated against the spec before it is written.

    Worktree mode probes a commit. If a defender or target file has uncommitted changes, probe refuses and says so — otherwise your new tests would be silently absent and the same survivors would come back with no hint why. --include-dirty snapshots the working tree (tracked edits and new files) into a throwaway commit and probes that; your tree, HEAD and index are never touched. Every summary names the commit probed.

    A claim with no defendedBy has its defenders discovered: the test files that import the fault's target, by relative path or resolved alias. NOCOVER then means exactly "no test file imports this source".

    Runners: vitest and jest (--runner auto picks the first that resolves; both read the same jest-compatible JSON report). Anything else goes through --runner-cmd. Each runner is proven against its own copy of the known-answer fixture.

    Practical loop: first pass --no-escalate (escalation re-runs the whole suite N times per survivor); iterate on one claim with --claim <ID> and either --include-dirty or --in-place (only fault target files must be clean there; test files may be dirty), optionally --confirm 1 for a fast provisional signal — verdicts print with a ?, evidence goes to evidence-provisional.json, and baseline refuses it; final pass with defaults. By default the stream shows only unproven faults plus a killed count — --verbose shows every fault. A custom runner (pnpm --filter, a specific config) goes in --runner-cmd "<cmd> {files} … {out}"; if the scratch worktree cannot see your node_modules, pass --node-modules <dir>.

  3. Baseline freezes every non-passing fingerprint. Later probes suppress what was already known and exit non-zero only on what is new. Claims whose source and defenders are unchanged reuse their prior verdict, so a probe in CI costs only what changed.

  4. Brief turns evidence plus baseline into a ranked, capped ## TEST BLINDSPOT CONTEXT block, printed and also written to .testguard/brief.json (--text prints only). Wire it into an agent's session start — for Claude Code, in .claude/settings.json:

    { "hooks": { "SessionStart": [ { "hooks": [
      { "type": "command", "command": "npx testguard-cli brief --text" }
    ] } ] } }
    

    --text prints only, and exits 0 silently when there is no evidence yet, so the hook can never break a session.

Built for agents to run

TestGuard is meant to be driven by an AI agent, not typed by a person. Three things make that safe:

  • One source of truth. testguard status --json computes state and the one next action from the claims file, the evidence, the baseline and the working tree. Every human rendering — the CLI text, the session-start brief, the skill — derives from it, so they cannot disagree. Every command accepts --json.
  • An installable operating loop. testguard init writes .claude/skills/testguard/SKILL.md (state → action, verdict → the only acceptable fix, the two-gate rule for any test the agent writes), the brief --text session-start hook, an AGENTS.md section and the .gitignore lines. Idempotent.
  • Gaming is visible. The cheapest way to make a survivor disappear is to weaken its fault, not to write a test. Evidence records every fault's content hash; status lists any fault edited after it survived, with its previous verdict, and makes reviewing that edit the next action. Editing a claim is allowed — claims can be wrong — but it is never invisible.

Authoring faults mechanically

Writing faults by hand means reading the code to find exact anchors. Two field reports found that ~80% of hand-written faults are one of five shapes, so scaffold proposes them for you:

npx testguard-cli scaffold src/auth.ts          # → .testguard/scaffold-auth.json (a draft, never your claims file)
npx testguard-cli scaffold src/auth.ts --claim AUTH-ADMIN   # every proposal under one claim; copies it if it exists
Shape What it proposes
condition-forced if (<guard>) {if (false) { — a guard is a !… condition or one whose body returns, throws or 4xx-es
statement-deleted a single-line guard (if (…) return …;) or a state change (x = …;) removed
return-altered return <check>; (===, .includes(, &&, …) → return true;
literal-changed httpOnly/secure flipped, sameSitenone, a cost/rounds → 1, a ttl/tolerance/window/limit ×1000
call-removed a bare verify…() / validate…() / check…() / authorize…() call removed

Every proposal's find is the exact line with expectHits/occurrence computed from the file, so it is verifiable by construction; provenance is producer: derived; defendedBy is prefilled from the tests that import the module; proposals are grouped under a preceding @claim <ID> annotation or by enclosing function. Statements are TODO: placeholders — a proposal becomes a claim only when a human states what it defends. Deterministic heuristics, no AST, no LLM; a proposal the tool cannot anchor is never emitted.

Commit .testguard/baseline.json; ignore evidence.json and brief.json. The baseline is the frozen contract; the other two are regenerated per run.

The fault model is the auditable artifact. You never reach 100% of correctness; you reach 100% of stated claims verified, and the statement of claims is what an assessor reads. A claims file is code — its replace strings run under your test runner — so review it like code.

Try it

The repository ships a known-answer fixture with a real blind spot:

git clone <this repo> && cd testguard && npm install
npm test                                  # includes probing the fixture end to end

fixtures/known-answer/ is a tiny project whose audit-row test asserts with expect.objectContaining({...}) and omits the content key. Swap the redacted text for the raw input and the test stays green. probe reports it as SURVIVED; the fixture's README walks through every verdict.

Status

v0.4. Seven commands, vitest and jest runners, hand-authored faults plus a mechanical scaffold, and an agent operating layer (status, init). The contract spine — six JSON Schemas shared with the other Guard tools — is under spec/. One exact-pinned runtime dependency (ajv, for schema validation); Node ≥ 20.

Not yet: test generation (the two-gate acceptance loop), runners beyond vitest and jest, AST-aware producers, and calibration of fault classes against real escaped bugs. Each is designed for; none is claimed.

Licence

MIT.

Release files for testguard-cli 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for testguard-cli 0.4.0
File Size Uploaded
testguard_cli-0.4.0.tar.gz 157.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for testguard-cli 0.4.0
File Interpreter ABI Platform
testguard_cli-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 166.8 kB

Release files / testguard_cli-0.4.0.tar.gz

Download URL testguard_cli-0.4.0.tar.gz
Size 157.3 kB
Tags Source
SHA-256 checksum
How to use checksums
89bb1c0aab25ee4d153a19f8f83c9fcab34cc22c8739f1db0eb73534df70f769
BLAKE2b-256 checksum
How to use checksums
140112f7a7ca360ddf2f09bd29a891885c39d71c6ba10c1491b8f03d68524dc7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release files / testguard_cli-0.4.0-py3-none-any.whl

Download URL testguard_cli-0.4.0-py3-none-any.whl
Size 9.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7ffe64ab9fd21aab9bed5282c2ce3afc1932c1e41367a5de00e3dcdd63405ffe
BLAKE2b-256 checksum
How to use checksums
9bbb96816dafe128efca1ef4f211ae0cb06e7eb287c45c40fcc96567985a8c34
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.

Transparency log

Release history Release notifications | RSS feed

0.14.0

2 release files

0.13.0

2 release files

0.12.0

2 release files

0.11.0

2 release files

0.10.4

2 release files

0.10.3

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.4

2 release files

0.9.3

2 release files

0.9.2

2 release files

0.9.1

2 release files

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

This release

0.4.0 This release

2 release files

0.3.1

2 release files

0.3.0

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page