Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

MateProbe by Mate4B

Test whether your Python validator catches invalid AI outputs and preserves valid ones.

MateProbe audits your existing validator with paired input variants: known faults and valid controls. It reports missed faults, rejections for the wrong reason, incomplete evidence, and execution errors. You supply the cases and policy; the audit runs without an LLM judge.

Here, mutations are changes to the samples passed to the validator. The audit does not rewrite its Python source. See how to test AI output validators for the workflow and how it fits with ordinary pytest and source mutation testing.

Alpha 0.1.0a4. Python 3.11+. Core runtime has zero third-party dependencies and makes no model or network calls. The optional pytest plugin adds a fixture and JSON reports.

The library checks structured state invariants and explicitly labelled lexical heuristics. It does not certify arbitrary prose as truthful, meaningful, or good writing. A passing declaration check only establishes consistency of the supplied declarations with supplied authoritative state.

Documentation · Agent integration guide · Published API · Pydantic recipe · Documentation index for agents

Try the one-file audit: a refund claim without a matching receipt, a retained prose survivor, and a pytest regression.

Previously Narrative Contracts. Existing users: see the migration guide.

Audit an existing validator

You do not need to change your validator to start. Wrap its existing result and supply the failures and valid variations you care about. The APIs below are available in 0.1.0a4; older 0.1.0a2 wheels do not include them.

Pytest can express every individual assertion. This library supplies paired baseline/variant execution, targeted finding attribution, valid controls, honest error accounting, obligation inventories, and CI reports. You still define the domain obligations and justify the cases.

For example, a test expects customer_mismatch, but the validator rejects with malformed_input. A generic assert not validate(sample) passes; the audit reports unattributed rejection, not successful detection of the customer error.

python examples/audit_existing_validator.py --output /tmp/refund-audit

The refund demo detects 3 of 9 authored faults before the fix and 8 of 9 afterward, preserving both valid controls. The prose-only contradiction still passes and an untested idempotency obligation stays visible. The prose challenge remains inside the nine-fault denominator. This is an illustrative audit of validators, not measured production accuracy. See scope and limits.

Start with the adoption levels: a boolean validator can expose accepted bad cases and rejected valid controls; finding IDs enable targeted detection, scopes refine attribution, and completeness/evidence explain unknowns. The external-validator trials adapt JSON Schema and Pydantic without adding runtime dependencies to the core.

Reports lead with known gaps, incomplete evidence, and untested obligations. Opt-in provenance records the library version, corpus digest, and observed or caller-supplied Git metadata. These are audit results for supplied cases, not a percentage of total agent coverage.

The corpus becomes a regression suite for your validation policy. Follow the survivor-to-regression guide to keep a discovered gap as a permanent pytest check and share a sanitized integration report.

When to use this

  • Your application owns authoritative state and needs to check explicit output declarations against it: refund status, account balances, workflow outcomes, or narrative branches.
  • You need repeatable pytest failures and reports for specified constraints, without an inference call during evaluation.
  • You want to audit a validator using faulty variants and valid controls, retaining missed faults and false rejections as evidence.

Start with the published alpha quickstart. Schema validation and state checks solve different problems; the Pydantic recipe shows both.

When this is insufficient

  • Checking whether unrestricted prose is truthful or entails the supplied declarations.
  • Discovering authoritative facts, executing state transitions, or enforcing runtime permissions.
  • Measuring writing quality, broad semantic equivalence, or production accuracy from a mutation score alone.

An LLM judge may evaluate open-ended properties outside these predicates. The two approaches can coexist; this library does not claim to replace every judge or guardrail system. See choosing an evaluation method.

Install the current alpha

Install the pinned alpha from PyPI with:

python -m pip install mateprobe==0.1.0a4 pytest-mateprobe==0.1.0a4

Install only mateprobe==0.1.0a4 if you do not need the pytest integration.

Install from this checkout

git clone https://github.com/Mate4b/mateprobe.git
cd mateprobe
python -m venv .venv
. .venv/bin/activate
pip install -e '.[dev]' -e ./packages/pytest-mateprobe
pytest

Both packages use the MIT license. Versioned wheels and sdists are distributed through GitHub Releases. To install the pinned alpha without a checkout:

python -m pip install \
  https://github.com/Mate4b/mateprobe/releases/download/v0.1.0a4/mateprobe-0.1.0a4-py3-none-any.whl \
  https://github.com/Mate4b/mateprobe/releases/download/v0.1.0a4/pytest_mateprobe-0.1.0a4-py3-none-any.whl

The core wheel can also be installed alone. Installation downloads packages; evaluation itself never calls a model. Model collection is a separate, opt-in benchmark script.

Start with the five-minute offline demo: a detected declaration error, a preserved valid variation, and a prose contradiction that passes.

State-conditioned checks

from mateprobe import (
    Claim,
    Context,
    DeclaredClaimsConsistent,
    Document,
    MinimumTokens,
    RequiredFact,
    Surface,
    evaluate,
)

context = Context(
    {
        "current": {"offer.status": "pending"},
        "accept": {"offer.status": "accepted"},
        "reject": {"offer.status": "rejected"},
    }
)
document = Document(
    (
        Surface(
            "outcome",
            "You sign the agreement and arrange a meeting with your new team.",
            state_ref="accept",
            claims=(Claim("offer.status", "accepted"),),
        ),
    )
)
contracts = (
    RequiredFact("offer-was-pending", "offer.status", "pending"),
    DeclaredClaimsConsistent("outcome-facts", ("outcome",)),
    MinimumTokens("lexical-floor", ("outcome",), minimum=8, minimum_unique=5),
)
report = evaluate(document, context, contracts)
report.assert_accepted()
print(report.to_dict())

A claim under reject cannot use facts from accept. Missing snapshots or facts produce undetermined, not success. Rule exceptions produce error. Strict policy blocks both; heuristic violations can be configured as advisory with Policy(block_heuristics=False).

Pytest

Install the second package for automatic pytest11 discovery:

def test_outcome(mateprobe):
    mateprobe.check(document, context, contracts)
pytest --mateprobe-report=contract-results.json

mateprobe.audit(cases, contracts, detection=0.9, preservation=0.95) checks a mutation campaign. Both denominators must exist; an empty suite cannot claim perfect performance. See mutation examples. Distributed pytest report merging is not yet supported; use a serial run with --mateprobe-report.

What is included

Contract Kind Actual guarantee / limitation
RequiredFact Invariant Type-sensitive equality of a declared required fact in a named snapshot; does not inspect prose.
DeclaredClaimsConsistent Invariant Declared claims match their own branch snapshot; undeclared assertions remain unchecked.
StateChanged Invariant At least one key in an explicit state projection changes; excludes unrelated bookkeeping.
MinimumTokens Heuristic Minimum word-token count and lexical diversity; not information content.
LexicalRestatement Heuristic High source-token overlap with too few novel tokens; no character-length exemption.
ForbiddenPattern Heuristic Matches configured regex on Unicode-normalized text; no negation or quotation reasoning.
SettledPremise Heuristic Configured phrase recognizers conditioned on closed premise IDs; unsupported IDs are explicit.
NoRepeatedText Heuristic Repeated normalized token sequence in supplied history; no semantic paraphrase detection.

All reports carry input/configuration digests, rule versions, evidence, scope and result status. No timestamps contaminate deterministic reports. See architecture and contract semantics.

Audit the evaluator

The mutation runner records baseline acceptance, attributable detections, survivors, valid-variant regressions and exclusions. A hit must match rule ID + finding code + scope; an unrelated rejection or exception does not count as a detection. Semantic validity of a mutation is caller-supplied and requires provenance. No-op, equivalent and unreviewed cases are reported as exclusions.

python examples/mutation_audit.py
python benchmarks/run.py

The included benchmark has author-constructed state/text variations and deliberate scope challenges. It is a feasibility artifact, not evidence of accuracy on natural LLM outputs. Original synthetic results, expanded mutation campaign, and real-model pilot are separate evidence streams. The follow-up 384-case mutation audit of captured Gemma/Qwen outputs separates declaration, lexical, schema, control and out-of-scope prose evidence. All 64 deliberately inserted prose-only contradictions pass; no semantic accuracy is claimed. Read the technical report and frozen release protocol. No human labels are required to use or reproduce the alpha; without them, natural-output acceptance must not be called semantic accuracy.

JSON and CLI

mateprobe examples/valid.json --output report.json

Exit status: 0 accepted, 1 rejected, 2 invalid input/configuration/I/O. Configuration accepts only built-in contract types and known fields. It never evaluates Python expressions. Library extensions use the Contract protocol, not untrusted imports from JSON.

LifeCard integration

mateprobe.adapters.lifecard_document maps a card to stable surface paths and an individual state reference for every outcome. Supply post-state snapshots calculated by the trusted engine, not by the generating LLM. The adapter does not modify LifeCard or execute effects. See example and migration guide.

Independent support workflow

The customer-support example computes refund eligibility with trusted Python state transitions, checks branch selection and declarations, and demonstrates valid paraphrases and failures that remain outside the prose guarantee. It is independent of LifeCard; this is an executable second integration, not evidence of broad domain generalization.

Reproduce or challenge the results

python benchmarks/expanded_mutations.py --output /tmp/expanded-campaign
python benchmarks/natural.py replay --input benchmarks/natural-results --output /tmp/natural-replay

Replay requires no model, network or API key. Output directories must be new. Submit counterexamples with a minimal bundle and evidence; see contribution guidance. Publication does not imply that independent reviewers have validated the method.

Development

pytest
ruff check .
ruff format --check .
mypy
python -m build --no-isolation
python -m build --no-isolation packages/pytest-mateprobe
python benchmarks/run.py

Roadmap and release gates · Contributing · Prior art

Metadata

Release files for mateprobe 0.1.0a4

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mateprobe 0.1.0a4
File Size Uploaded
mateprobe-0.1.0a4.tar.gz 677.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mateprobe 0.1.0a4
File Interpreter ABI Platform
mateprobe-0.1.0a4-py3-none-any.whl Python 3 none any Details

Total release size: 708.0 kB

Release files / mateprobe-0.1.0a4.tar.gz

Download URL mateprobe-0.1.0a4.tar.gz
Size 677.9 kB
Tags Source
SHA-256 checksum
How to use checksums
b7af4085f13cf96ebca572d844ff8fe1a68afacc5666783dc456a9ea16238924
BLAKE2b-256 checksum
How to use checksums
ec6e0b26a9dbb06956294d3b188716691d366693a343dd74b1eb2ad37de4c1a8
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release files / mateprobe-0.1.0a4-py3-none-any.whl

Download URL mateprobe-0.1.0a4-py3-none-any.whl
Size 30.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
462419a35c7f3cdecfde6609748331aaaaf77aeaca7836bb2db89a01ea79f1ca
BLAKE2b-256 checksum
How to use checksums
30f92ce921c06ac65a8a8f6e425511caf0f1cce82589d6359612d22fab893b7f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 29, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0a4 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page