TestGuard
Proves that a test suite actually defends the claims a project makes — by injecting the faults those claims say cannot happen, and reporting every fault the tests fail to detect.
Not a test generator. A claim verifier. Test generation is what happens after a claim turns out to be unfalsifiable.
Third tool following the Guard pattern, alongside
docguard-cli (docs ↔ code) and
websec-validator (attack
surface ↔ code). All three run one loop:
declare what must be true → try mechanically to falsify it → freeze a baseline → gate only the delta → brief the agent before it writes code.
Why
Coverage cannot tell a test that pins correct behaviour from one that pins a defect. An agent that writes both the code and its tests encodes whatever it believed — including its bugs — and the suite goes green.
Measured on a real, entirely AI-authored production codebase with ~4,900
disciplined tests (no snapshots, 0.4% zero-assertion): 8 of 9 real
historical bugs were invisible to the suite, worst case 2,451 tests green
on known-broken code. The largest gap was a compliance-critical path with
100% coverage, where the one assertion that mattered used
expect.objectContaining({...}) and omitted the field carrying the data.
A second, independent run on a different AI-authored codebase (63 test files, 458 tests, 24 hand-written security claims, 39 faults): 21 of 39 faults survived a fully green suite — 9 of them critical. Super-admin gating, membership checks, cookie flags and the whole authorization callback could be disabled without a single test noticing. One test file had re-implemented the authorization logic inside the test and asserted against the copy: fifteen green tests, zero detection. After wrapper-level tests were written against the survivors, 39/39 were killed.
Install
| How | Command |
|---|---|
| npx (no install) | npx testguard-cli probe |
| npm | npm i -D testguard-cli then npx testguard probe |
| pip | pip install testguard-cli then testguard probe (needs Node ≥ 20) |
| Homebrew | brew tap raccioly/tap && brew install testguard |
| GitHub Action | uses: raccioly/testguard@v0.5.0 — see action.yml |
| pre-commit | repo: https://github.com/raccioly/testguard, hooks testguard-claims, testguard-probe |
Projects that set min-release-age in .npmrc cannot see a version published
less than that many days ago (ENOVERSIONS); install that one with
npm i -D testguard-cli --min-release-age=0.
How it works
npx testguard-cli init # install the agent layer: skill, session-start hook, AGENTS.md section
npx testguard-cli status --json # where the project is and the ONE next action — the machine entry point
npx testguard-cli claims # what does this project claim, and is every claim probeable?
npx testguard-cli probe # try to falsify each claim; report what the tests missed
npx testguard-cli baseline # freeze today's unproven findings; from now on only new ones gate
npx testguard-cli brief # tell the agent where the suite is blind, before it writes
npx testguard-cli gate --changed origin/main # fail when a changed source file carries no claim at all
npx testguard-cli scaffold src/x.ts # propose faults for a file, as a draft to keep or drop
-
Claims live in
testguard.claims.json(editors validate it against"$schema": "./node_modules/testguard-cli/spec/schemas/claims.schema.json"): a statement, where it comes from, which tests supposedly defend it, and one or more faults — each a deterministic source change that would make the statement false. Every claim and every fault records who produced it.testguard claimsvalidates the file and reports drift against@claim <ID>annotations in source. Test files are deliberately not scanned — a claim asserted by a test is the authorship trap the tool exists for — and annotation ids must contain a hyphen so prose is never mistaken for one. -
Probe confirms the defenders are green N times unmodified, applies each fault in a scratch git worktree (your tree is never touched), runs the defenders N times, re-runs survivors against the whole suite with N-run attribution, restores, and classifies. Verdicts are a closed set:
Verdict Meaning killeda test body rejected the behaviour, N/N — the only pass SURVIVEDthe defenders stayed green while the claim was false NOCOVERno test file defends the claim at all UNVERIFIABLEthe fault's anchor is missing or ambiguous — loud, never a skip TIMEOUTthe defenders hung; a hang is not a detection FAULT-INVALIDthe replacement does not load — a bad fault, not a finding FLAKY-DEFENDERthe defenders are not reliably green, or disagreed across runs Never a single score. Findings are ranked by severity, claim provenance and blast radius (relative imports,
tsconfigpath aliases andpackage.json#importsare resolved; bare package names are not), and written to.testguard/evidence.json— validated against the spec before it is written.Worktree mode probes a commit. If a defender or target file has uncommitted changes,
proberefuses and says so — otherwise your new tests would be silently absent and the same survivors would come back with no hint why.--include-dirtysnapshots the working tree (tracked edits and new files) into a throwaway commit and probes that; your tree, HEAD and index are never touched. Every summary names the commit probed.A claim with no
defendedByhas its defenders discovered: the test files that import the fault's target, by relative path or resolved alias.NOCOVERthen means exactly "no test file imports this source".Runners: vitest and jest (
--runner autopicks the first that resolves; both read the same jest-compatible JSON report). Anything else goes through--runner-cmd. Each runner is proven against its own copy of the known-answer fixture.Practical loop: first pass
--no-escalate(escalation re-runs the whole suite N times per survivor); iterate on one claim with--claim <ID>and either--include-dirtyor--in-place(only fault target files must be clean there; test files may be dirty), optionally--confirm 1for a fast provisional signal — verdicts print with a?, evidence goes toevidence-provisional.json, andbaselinerefuses it; final pass with defaults. By default the stream shows only unproven faults plus a killed count —--verboseshows every fault. A custom runner (pnpm --filter, a specific config) goes in--runner-cmd "<cmd> {files} … {out}"; if the scratch worktree cannot see yournode_modules, pass--node-modules <dir>. -
Baseline freezes every non-passing fingerprint. Later probes suppress what was already known and exit non-zero only on what is new. Claims whose source and defenders are unchanged reuse their prior verdict, so a probe in CI costs only what changed.
-
Brief turns evidence plus baseline into a ranked, capped
## TEST BLINDSPOT CONTEXTblock, printed and also written to.testguard/brief.json(--textprints only). Wire it into an agent's session start — for Claude Code, in.claude/settings.json:{ "hooks": { "SessionStart": [ { "hooks": [ { "type": "command", "command": "npx testguard-cli brief --text" } ] } ] } }
--textprints only, and exits 0 silently when there is no evidence yet, so the hook can never break a session.
Every change needs a claim
probe can only verify claims that exist. Every escaped defect in the field
reports so far was a claim gap: the feature shipped green with zero
claims, and a verifier with no claim about a feature is silent about it by
construction. gate closes that hole on the delta:
npx testguard-cli gate --changed origin/main # in a PR: the files changed since the base branch
npx testguard-cli gate --changed HEAD --include-dirty # before a commit: the working tree, staged or not
Every changed source file must carry a fault, resolve as a defender of a
claim (test files), or be excused by an unexpired path entry in
testguard.ignore.json — with a reason a reviewer will accept. One
unclaimed file exits 1; there is no percentage. Every reliance on an ignore
entry is printed, so a reviewer sees why the gate passed; an expired entry
excuses nothing. Non-source files and documented never-claimed patterns
(*.d.ts, *.config.*, fixtures, mocks; --explain lists them) are
excluded and said so; --strict fails a change that evaluated nothing.
With a reference known, status --changed <ref> reports unclaimed-changes
before any evidence state and makes the claim the next action; the brief
lists the unclaimed files first. The claim is written before more code.
In CI the base is detected — GitHub Actions (GITHUB_BASE_REF) and GitLab
merge request pipelines (CI_MERGE_REQUEST_DIFF_BASE_SHA, then
CI_MERGE_REQUEST_TARGET_BRANCH_NAME); TESTGUARD_CHANGED_REF overrides both.
The base must exist locally: GitHub — actions/checkout with fetch-depth: 0;
GitLab — the diff base sha needs nothing extra on a merge request pipeline,
the branch name needs GIT_DEPTH: 0 or a git fetch origin <target>. A
detected base that does not resolve is a warning for status and brief
(they keep working) and an error for gate (its whole job is the measurement).
# GitHub Actions
- uses: actions/checkout@v4
with: { fetch-depth: 0 }
- uses: raccioly/testguard@v0.5.0
with: { command: gate }
# GitLab CI — or include: remote: the template in packaging/gitlab/
testguard:gate:
image: node:22
rules: [{ if: $CI_PIPELINE_SOURCE == "merge_request_event" }]
script: [npx -y testguard-cli gate .]
Built for agents to run
TestGuard is meant to be driven by an AI agent, not typed by a person. Three things make that safe:
- One source of truth.
testguard status --jsoncomputesstateand the onenextaction from the claims file, the evidence, the baseline and the working tree. Every human rendering — the CLI text, the session-start brief, the skill — derives from it, so they cannot disagree. Every command accepts--json. - An installable operating loop.
testguard initwrites.claude/skills/testguard/SKILL.md(state → action, verdict → the only acceptable fix, the two-gate rule for any test the agent writes), thebrief --textsession-start hook, anAGENTS.mdsection and the.gitignorelines. Idempotent. - Gaming is visible. The cheapest way to make a survivor disappear is to
weaken its fault, not to write a test. Evidence records every fault's
content hash;
statuslists any fault edited after it survived, with its previous verdict, and makes reviewing that edit the next action. Editing a claim is allowed — claims can be wrong — but it is never invisible.
Authoring faults mechanically
Writing faults by hand means reading the code to find exact anchors. Two
field reports found that ~80% of hand-written faults are one of five shapes,
so scaffold proposes them for you:
npx testguard-cli scaffold src/auth.ts # → .testguard/scaffold-auth.json (a draft, never your claims file)
npx testguard-cli scaffold src/auth.ts --claim AUTH-ADMIN # every proposal under one claim; copies it if it exists
| Shape | What it proposes |
|---|---|
condition-forced |
if (<guard>) { → if (false) { — a guard is a !… condition or one whose body returns, throws or 4xx-es |
statement-deleted |
a single-line guard (if (…) return …;) or a state change (x = …;) removed |
return-altered |
return <check>; (===, .includes(, &&, …) → return true; |
literal-changed |
httpOnly/secure flipped, sameSite → none, a cost/rounds → 1, a ttl/tolerance/window/limit ×1000 |
call-removed |
a bare verify…() / validate…() / check…() / authorize…() call removed |
Every proposal's find is the exact line with expectHits/occurrence
computed from the file, so it is verifiable by construction; provenance is
producer: derived; defendedBy is prefilled from the tests that import
the module; proposals are grouped under a preceding @claim <ID> annotation
or by enclosing function. Statements are TODO: placeholders — a proposal
becomes a claim only when a human states what it defends. Deterministic
heuristics, no AST, no LLM; a proposal the tool cannot anchor is never
emitted.
Commit .testguard/baseline.json; ignore evidence.json and brief.json.
The baseline is the frozen contract; the other two are regenerated per run.
The fault model is the auditable artifact. You never reach 100% of
correctness; you reach 100% of stated claims verified, and the statement
of claims is what an assessor reads. A claims file is code — its replace
strings run under your test runner — so review it like code.
Try it
The repository ships a known-answer fixture with a real blind spot:
git clone <this repo> && cd testguard && npm install
npm test # includes probing the fixture end to end
fixtures/known-answer/ is a tiny project whose audit-row test asserts with
expect.objectContaining({...}) and omits the content key. Swap the
redacted text for the raw input and the test stays green. probe reports it
as SURVIVED; the fixture's README walks
through every verdict.
Status
v0.4. Eight commands, vitest and jest runners, hand-authored faults plus
a mechanical scaffold, an agent operating layer (status, init) and a
change gate (gate). The contract
spine — eight JSON Schemas shared with the other Guard tools — is under
spec/. One exact-pinned runtime dependency (ajv, for schema validation); Node ≥ 20.
Not yet: test generation (the two-gate acceptance loop), runners beyond vitest and jest, AST-aware producers, and calibration of fault classes against real escaped bugs. Each is designed for; none is claimed.
Licence
MIT.
Release files for testguard-cli 0.5.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| testguard_cli-0.5.0.tar.gz | 178.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| testguard_cli-0.5.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 188.5 kB
Release files / testguard_cli-0.5.0.tar.gz
| Download URL | testguard_cli-0.5.0.tar.gz |
|---|---|
| Size | 178.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
8724aac9580478442126194de53f2ec3c516ee783eb0dbe5dcf6eaeaf2f08a96
|
|
BLAKE2b-256 checksum How to use checksums |
2eb5705e40fc4e6d06a41480c6974d0ed7477d6de6b376da09baef020deb1d04
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency logRelease files / testguard_cli-0.5.0-py3-none-any.whl
| Download URL | testguard_cli-0.5.0-py3-none-any.whl |
|---|---|
| Size | 10.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
88b5cb6dd3c3f4ca733615e6f1ccfeec77f2b71276bf121b55c29a777369e5d4
|
|
BLAKE2b-256 checksum How to use checksums |
ca413a0a8c839c44b0aa6335dc98e618f7df6c9053ecd0910561f2ff151ec62c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 17, 2026.
Transparency log