Skip to main content

Mendmark

Mutation testing for agent evals.

Find the broken tool calls and coordination failures your passing tests still accept.

Tests Security PyPI Python 3.10–3.13 MIT license

Quick start · Harnesses · Golden set · Multi-agent · How it works · Run a pilot

Mendmark changes one part of a passing agent trace, reruns existing evaluators, and identifies killed faults and surviving blind spots.

Your agent tests can all pass and still miss a broken tool call. Mendmark tests the tests themselves: it plants one controlled fault, runs the same evaluators again, and turns every survivor into a concrete blind spot to fix.

Killed means the eval noticed the planted fault. Survived means the damaged case still passed. A critical survivor can fail the release gate.

Quick start

Mendmark requires Python 3.10 or newer.

pip install 'mendmark-evals[deepeval]'

git clone https://github.com/danielgaskins/mendmark.git
cd mendmark

mendmark audit examples/order_agent_suite.py \
  --output mendmark-report.json \
  --write-baseline
Mendmark agent-eval audit
Cases: 1
Mutations: 19  Killed: 19  Survived: 0  Errors: 0
Mutation kill rate: 100.0%
New tools: lookup_order, refund_order
Gate: PASS

Run the same command in CI without --write-baseline. Mendmark compares the current tool schemas and mutation results with the last accepted baseline.

For a native parallel multi-agent audit:

mendmark audit-json examples/multi_agent_suite.json \
  --evaluator-command "python3 examples/multi_agent_evaluator.py"

Equip an agent harness

Mendmark has dependency-light adapters for LangChain/LangGraph, CrewAI, and the OpenAI Agents SDK. In an existing agent repository:

python -m pip install 'mendmark-evals==0.6.1'
mendmark equip --framework auto --agent auto

The command detects bounded dependency files and creates a reviewed capture guide, offline evaluator, and inactive CI template under .mendmark/. It does not edit application code, upload a trace, overwrite existing work, enable CI, or accept a baseline.

Want Codex or Claude Code to perform the integration? Install its native, repo-scoped skill (use all to install both):

mendmark equip --framework auto --agent codex       # invoke with $mendmark
mendmark equip --framework auto --agent claude-code # invoke with /mendmark

For any other repository-capable agent, print a portable self-equip prompt:

mendmark equip --agent generic --print-agent-prompt

See the agent harness integration guide for the direct Python APIs, explicit trace-approval boundary, framework compatibility, and multi-agent guidance.

See the blind spot in two minutes

▶ Watch the narrated weak-eval demonstration (original v1 fault inventory)

A refund-agent test checks only the final sentence. Mendmark finds that 15 of 19 faults escape—including a wrong refund amount and a duplicated refund. A complete evaluator checks the calls, arguments, results, and response, killing all 19.

Agent Eval Golden Set

The Mendmark Agent Eval Golden Set is the canonical, versioned benchmark for the mutation engine.

24
reviewable cases
13
tool contracts
39
tool calls
263
pinned mutations
10
domains
Evaluator profile Killed Survived Kill rate
Complete trace and outcome 263 0 100.000%
Trace only 215 48 81.749%
Response only 87 176 33.080%

The response-only profile leaves 162 critical tool-behavior mutations undetected. The complete profile catches every mutation in the golden set.

python benchmarks/benchmark_golden_set.py

Review the manifest, case suite, evaluator profiles, methodology, and reference performance directly. The benchmark is deterministic, offline, and makes no model calls.

Native coordination behavior has its own reviewable Multi-Agent Golden Set v2: 6 workflows, 17 agent declarations, 41 causal events, 30 operators, and 294/294 mutations killed by the complete reference evaluator.

python benchmarks/benchmark_multi_agent_golden_set_v2.py

For contrast, the v2 output-only evaluator detects just 23/294 and leaves 271 survivors, 181 critical. Mendmark groups those blind spots by category and pinpoints the affected agent, event, and tool using privacy-safe identifiers. See the multi-agent guide for both commands.

Built for real CI

🔎 Expose evaluator blind spots
Wrong arguments, corrupted results, reordered calls, duplicate side effects, false recovery, and damaged responses.
🧰 Fit the existing stack
DeepEval, Rubric, or any local evaluator through a validated JSON subprocess protocol.
🚦 Gate tool rollouts
Per-tool coverage, schema-change detection, accepted baselines, mutation budgets, JUnit, and SARIF.
🔐 Keep case content local
Reports omit prompts, answers, tool arguments, and tool outputs; artifacts can be signed with Cosign.

One engine for single-agent and multi-agent systems

Single-agent suites use a simple ordered tool trace. Multi-agent suites add a causal event graph with agent identities, delegation targets, explicit tool permissions, returned results, shared-state events, and dependencies. Existing 1.0 suites remain unchanged; native graphs use the validated 2.0 JSON contract.

Mendmark mutates both layers. It can break an individual tool call, route work to the wrong specialist, omit handoff context, drop or misattribute a result, violate an agent's tool permissions, remove a causal dependency, or insert a delegation loop. Independent parallel branches are not forced into an arbitrary wall-clock order.

The included three-agent reference suite has 9 events across parallel billing and risk branches. Its complete evaluator kills all 64 currently applicable mutations. The broader v2 golden set covers six topologies and kills 294/294. See the multi-agent guide and reviewable JSON suite.

See an eval fail the test

The repository also includes a deliberately weak evaluator. It checks whether the final sentence is correct and ignores the tool trace.

mendmark audit examples/order_agent_weak_suite.py \
  --output /tmp/mendmark-weak-report.json

The original refund case passes. Mendmark then changes the refund amount, removes required calls, and duplicates the side effect. Many of those faults survive because the final sentence never changed. The command exits with a failed gate and names each blind spot.

Run the complete order_agent_suite.py next. Its tool-trace evaluator checks the ordered calls, arguments, and results, so the same planted faults are caught. This before-and-after pair is the shortest demonstration of what Mendmark measures.

Define a suite

A suite is a trusted local Python file. It exports three things:

from deepeval.metrics import BaseMetric
from deepeval.test_case import LLMTestCase, ToolCall

TOOLS = [
    {
        "name": "refund_order",
        "input_schema": {
            "type": "object",
            "properties": {
                "order_id": {"type": "string"},
                "amount": {"type": "number"},
            },
            "required": ["order_id", "amount"],
        },
        "side_effecting": True,
    }
]

MENDMARK_POLICY = {
    "minimum_kill_rate": 0.9,
    "fail_on_critical_survivor": True,
    "fail_on_untested_tools": True,
    "fail_on_tool_contract_issues": True,
    "fail_on_regression": True,
}


def get_metrics():
    # Return your DeepEval metrics here. The complete example includes a
    # deterministic tool-trace metric that runs without an API key.
    return [MyToolMetric()]


def get_cases():
    refund = ToolCall(
        name="refund_order",
        input_parameters={"order_id": "104", "amount": 29.99},
        output={"status": "accepted"},
    )
    return [
        LLMTestCase(
            name="refund-order",
            input="Refund order 104 in full.",
            actual_output="The refund was accepted.",
            expected_output="The refund was accepted.",
            tools_called=[refund],
            expected_tools=[refund],
        )
    ]

get_metrics() must return fresh metric instances on every call. Metric names must be unique. Mendmark reruns those metrics against the original case and each mutated copy.

See the complete example and the mutation audit guide.

Use JSON instead of DeepEval

Teams can export cases and traces as JSON and connect any language or eval framework through a local stdin/stdout command:

mendmark audit-json examples/order_agent_suite.json \
  --evaluator-command "python3 examples/json_evaluator.py" \
  --junit /tmp/mendmark.xml \
  --sarif /tmp/mendmark.sarif

The command runs locally and receives the original and mutated cases in one batch. Mendmark strictly validates its metric results. See the JSON adapter and protocol.

Custom domain faults can be loaded from a suite, trusted Python file, installed entry point, or module attribute. See custom mutation plugins.

The repository also includes a tested Rubric integration that runs Rubric metrics through the same JSON protocol.

CI provenance, budgets, and signatures

Reports automatically record the Mendmark version, adapter, canonical policy digest, and supported GitHub/GitLab CI metadata. Explicit versions can be added with --source-commit, --source-ref, --suite-version, and --policy-version. Use --maximum-mutants to stop before evaluator work when a suite exceeds its approved cost ceiling.

Mendmark delegates signatures to Sigstore Cosign:

mendmark sign mendmark-report.json --bundle mendmark-report.sigstore.json
mendmark verify-signature mendmark-report.json \
  --bundle mendmark-report.sigstore.json \
  --certificate-identity "EXPECTED_OIDC_IDENTITY" \
  --certificate-oidc-issuer "EXPECTED_OIDC_ISSUER"

See artifact signing, the compatibility policy, the engine benchmark, the user assurance contracts, and the packaged report and baseline JSON Schemas.

Tool rollout checks

Mendmark hashes each declared tool's name, schema, description, and side-effect flag. The baseline lets it answer four concrete questions in a pull request:

  1. Was a tool added?
  2. Did its contract change?
  3. Does at least one eval exercise it?
  4. Do those evals catch faults in its calls and results?

Mendmark also checks required arguments and basic JSON Schema types in the actual and expected traces. Reports identify the case, tool, field, and problem without storing the argument value.

This makes a tool launch visible before it reaches production. It does not prove the tool is safe. It shows whether the team's current evals can recognize the failures Mendmark introduced.

Security and privacy

The suite file is executable Python. Only run suites from code you trust.

Mendmark's JSON report stores case IDs, operator names, severities, metric names, statuses, and tool schema digests. It does not store prompts, expected answers, tool arguments, or tool outputs. Teams can run the engine inside their own CI boundary and publish only the report.

See SECURITY.md for private vulnerability reporting, the threat model and deployment checklist for the complete trusted-code boundary, and SUPPORT.md for version and support expectations.

Design-partner pilot

Teams with a real tool-using agent can follow the pilot guide and open a privacy-safe Mendmark pilot request. Do not include prompts, traces, payloads, credentials, or customer data in a public issue.

Completed pilots use the machine-validated design-partner evidence contract to record time-to-value, mutation realism, equivalent faults, discovered/remediated blind spots, runtime, cost, and CI retention without storing customer content. Until that external utility gate passes, golden-set results are engine evidence—not a claim of universal agent safety or customer validation.

Existing ML integrity pack

Mendmark started as a benchmark for coding agents that repair ML pipelines. That work remains available through mendmark prepare, mendmark grade, and the MendmarkIntegrityMetric DeepEval adapter. It checks failures such as label leakage, train-serve skew, invalid metric aggregation, and broken reproducibility.

The ML pack is now one specialized use of the broader idea. An evaluator should be tested against known bad outcomes before its score is trusted.

See the DeepEval guide and the ML evaluation card.

Current boundary

Version 0.6 is a local, open-source engine. It does not yet provide a hosted dashboard, team accounts, remote trace ingestion, or a secrets service. The planned control plane is described in the product design.

Release history is maintained in CHANGELOG.md.

Release files for mendmark-evals 0.6.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mendmark-evals 0.6.1
File Size Uploaded
mendmark_evals-0.6.1.tar.gz 157.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mendmark-evals 0.6.1
File Interpreter ABI Platform
mendmark_evals-0.6.1-py3-none-any.whl Python 3 none any Details

Total release size: 252.8 kB

Release files / mendmark_evals-0.6.1.tar.gz

Download URL mendmark_evals-0.6.1.tar.gz
Size 157.1 kB
Tags Source
SHA-256 checksum
How to use checksums
566db92ebca76cfa91e300eaf7b03256c249a7c7a2db1313b152fef8440fe000
BLAKE2b-256 checksum
How to use checksums
68c2753a22b7d9098e4239201e52d8a2a09ab9322596cc3417570456679207d7
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release files / mendmark_evals-0.6.1-py3-none-any.whl

Download URL mendmark_evals-0.6.1-py3-none-any.whl
Size 95.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
4c14e20839fae7754135a8c27eed200d0031c1b1f9042857dc190d16770f0e08
BLAKE2b-256 checksum
How to use checksums
549724cc5e8b421839acfcb8c5f8b3d6dc58c77f07ab20e9f00b12cfa18a471c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Aug 11, 2026.

Transparency log

Release history Release notifications | RSS feed

0.7.1

2 release files

0.7.0

2 release files

This release

0.6.1 This release

2 release files

0.6.0

2 release files

0.4.2

2 release files

0.4.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page