Skip to main content

Mylonite

A safe model is not the same thing as a safe app. A top-tier model can shrug off every generic prompt-injection you throw at it and still hand an attacker a win — because the weakness is in how your app is wired, not in how the model behaves. Mylonite checks whether your app's own safeguards are what stop an attack, writes a test for every weakness it finds, and wires that test into CI so a future model upgrade cannot quietly remove the protection.

CI PyPI GitHub release License

For teams shipping MCP or agentic apps who want CI-enforced regression coverage on the AI layer.

Point Mylonite at any MCP (Model Context Protocol) app. Mylonite's own agent reads your system prompt and tool descriptions and drives your tools itself — attacking the part an ordinary code scanner cannot see — finds weaknesses specific to your app, and turns each one into a pytest test that gates your CI. The API key you configure pays for that agent: the planner driving the attack, the customiser crafting payloads, and the judge scoring the result — not for a call your own app's agent made.

Mylonite deliberately does not review your ordinary application code. That is what SAST/DAST tools are for.

How it works

The idea is simple: run the same attack twice.

  1. Once against your app as it is.
  2. Once with the safeguard in place.

A weakness is only reported if the attack succeeds the first time and is stopped the second — repeated several times over, so that a model having a good or bad day cannot decide the outcome. (In the codebase and the docs this is the differential, or control-efficacy check.)

That answers a question a one-off scan cannot: is your safeguard doing the work, or is the model's current good behaviour doing it? Only one of those survives a model upgrade.

Two levels of proof

Mylonite always tells you which one you got.

What plays the "safeguard" role What a kept finding tells you How to get it
Your own safeguard, switched off and on Your implementation is doing the work. The stronger result. Declare control_env in your target.yaml
A standard safeguard Mylonite applies at the boundary The attack is real, and this kind of safeguard closes it — but what stopped it was Mylonite's stand-in, not your code The default; works on any app, no setup

The second is genuinely useful, and it is what runs against most real apps. It is a weaker statement than the first, and no Mylonite output will dress it up as the stronger one.

How it has been tested

Mylonite ships an independent verification harness that scores it against material it did not write — runnable third-party MCP servers and published academic benchmarks. The harness is in this repository and you can run it yourself.

Results against third-party targets

  • A kept finding on a third-party MCP email server. The attack succeeded 5 times out of 5 against the server as shipped, and 0 times out of 5 with the safeguard in place. Scope of that run: the server needed two fixes before it would start at all (a launch wrapper, and a one-line bug in its own send_email), the weakness only appears when the app's system prompt tells the agent to send without asking, and the safeguard was Mylonite's boundary stand-in rather than a second build of the server. Full detail in the capability matrix.
  • No false alarms on a third-party server with nothing wrong with it (Enkrypt's echo_mcp). Running the same comparison against a hardened third-party server is still outstanding — see verification.
  • A real weakness found in a published vulnerable-MCP corpus (MCPSecBench). When that finding was re-run to confirm it, it failed to reproduce (0 of 3 runs), so Mylonite discarded it instead of shipping a flaky test.
  • The judge checked against real third-party examples — AgentDojo transcripts from models that genuinely fell for attacks, not examples we wrote ourselves.
  • The safety rails hold under test. A check that could not be run is never reported as a pass; a score with nothing to measure is labelled as such; work outside the tool's scope is marked rather than graded.

Results that came back negative

Published for the same reason the positive ones are.

  • DVMCP recall is unmeasured, not a result. The published 0.9.0 figure read 0/8, but the harness that produced it had two defects that could each force a miss regardless of what the scan actually found: the scorer folded any untested challenge into "missed" instead of reporting it separately (fixed, issue #136), and the documented campaign workflow saves a report with no per-attempt weakness class, so the found count reads 0 whatever the scan did (still open). Treat 0/8 as unmeasured until Layer 1 is re-run.
  • On InjecAgent (100 cases per split, using a local llama3.2:3b) the judge scored F1 1.000 on the direct-harm split and F1 0.833 at 0.714 recall on the data-stealing split in 0.10.0 (0.9.0 measured 0.400 at 0.25 recall). That recall rests on only 7 attacks that succeeded, so we record it as unresolved at this sample size, not as an improvement. The gap between the splits is the finding, so both are published.
  • Judge agreement of F1 0.41 against AgentDojo's own labels. Mylonite's judge asks "did harm actually happen?"; AgentDojo asks "was the exact goal achieved?". Some of that gap is a genuine difference in question, which we have not resolved.

Current limits

  • Against a single-build app, only the weaker statement is available. With no control_env to switch, the safeguard is Mylonite's stand-in rather than your code, and no output will claim otherwise.
  • Finding nothing is the normal outcome for a well-built app on a robust model. See below.
  • The evidence rests largely on one model — Claude Haiku 4.5 — at small, deliberately cost-capped sample sizes.
  • The published figures were measured between 25 June and 14 September 2026. The benchmark results carry the version they were measured against (verification/results/0.9.0/ and verification/results/0.10.0/); the run logs in the capability matrix do not. Figures have not been re-measured for every release since, so read them as a floor rather than a current reading. Per-release re-measurement is planned.
  • Run transcripts are not published. The harness and its scorers are, so you can produce your own numbers; you cannot yet audit ours.
  • No third party has built a plugin against the extension points yet.

Full scorecard with caveats: docs/verification.md. Everything that limits the tool's reach is collected in docs/limitations.md.

Finding nothing is also a result

Worth knowing before you run it. Against a well-built app on a robust model, Mylonite will often correctly find nothing — that is the tool working, not failing. Proving a safeguard carries the security requires a weakness that actually lands, which in practice means a design flaw (an action with real consequences and no approval step, an unrestricted outbound request) or an app configured to act on its own.

That is why Try it starts with the bundled practice app rather than your code. That app is deliberately insecure, so it finds something every time and you can watch the machinery work before pointing it somewhere the honest answer may be "nothing".

Where this sits

Static scanners read your tool descriptions and flag whatever looks risky, leaving you to judge which flags matter. Model-eval harnesses swap models and score which behaves best. Mylonite's purpose is the step neither of those takes: run the attack against your app, then hold the model constant and switch only your safeguard. The result is evidence about your safeguard, not about how a description reads or how a model scored today.

Project status

Beta, and essentially a single maintainer — one outside contribution to date, the rest of the history from the maintainer and Dependabot. Over 2,300 tests, with CI (ruff, mypy, pytest, pre-commit) enforced on every pull request. The extension points are versioned public API, but nobody outside the project has built against them yet. If you are weighing this as a dependency in a security pipeline, pin a version — and read Known limitations first.

Install

pip install mylonite                      # the CLI, from PyPI
pip install "mylonite[demo]"              # ...plus the bundled practice app

Python 3.11–3.14.

The [demo] extra installs the bundled practice app, and you need it for any reference:... command — demo and scan reference:... alike.

Scanning your own app needs a model: an API key for a hosted provider, or no key at all for one you host yourself (Ollama, vLLM, or a LiteLLM proxy — see self-hosted models). scan --scaffold and report never need a model at all, and demo replays recorded responses rather than calling one.

Try it

No API key, no install, one command (needs uv):

uvx --from "mylonite[demo]" mylonite demo

On every pull request, CI builds Mylonite from source and runs this command against that build on Python 3.14, on Linux and Windows, in an 80-column terminal.

With the [demo] extra already installed (see Install), the same demo is:

mylonite demo

Either way, that runs the comparison against the bundled practice app — deliberately insecure, runs in-process, opens no network ports — and prints what got through on the unguarded build next to what was stopped on the guarded one. Same attacks, two builds, different outcomes. That contrast is the point of the tool.

demo replays model responses recorded against those bundled apps, so it is offline and gives the same answer every time. The scan engine, the adapters and the comparison are all the real ones; only the model's replies are pre-recorded, and the output tells you which model produced them and when. Every demo mode — replay and --live alike — turns off the per-seed customiser and the LLM-judge fallback, so a verdict is decided purely by deterministic predicates; scan itself runs each payload once by default, and the repeat-run consensus belongs to validate (five iterations by default), which gate runs for you. Treat the numbers as a demonstration of the machinery rather than a fresh measurement of today's model — mylonite demo --live is the fresh measurement, and it does call a model (by default one you host locally). Where a cell could not be decided either way the table says so rather than showing it as a pass, and if a recording is ever missing or out of date the command fails and explains why instead of reporting a clean result it did not earn.

Then, with a model configured, the real thing:

mylonite scan reference:vulnerable   # finds the weaknesses built into it
mylonite scan reference:guarded      # same attacks, comes up clean

See the practice app for what is built into it and why.

Then point it at your own app

The first step is free — no API key, no model call, no spend.

# Inspect a server and write a starter target.yaml
mylonite scan --command "python" --arg "my_server.py" --scaffold app.yaml --scope my-app

--scaffold connects to your server, lists its tools, says which weakness classes apply to it, and flags the tools whose actions have real consequences. It works from the names, descriptions and schemas your server advertises, matching them against keyword patterns, so treat everything it suggests as a hint to confirm rather than a verdict.

Proving which weaknesses actually land, and whether a safeguard closes them, is the scan itself. By default the safeguard side is Mylonite's own stand-in guard at the tool boundary — real evidence that this kind of guard closes the attack, not yet that your implementation does; declare control_env in the target file to measure your own guard instead (see docs/concepts.md). That needs a model:

mylonite scan --target-file app.yaml --authorize my-app

Expect this to find less than the practice app did — often nothing. See Finding nothing is also a result above, and docs/limitations.md for where the tool's reach genuinely ends.

From a scan to a pull request that gates CI

mylonite gate runs the whole sequence — find a weakness, write a test for it, confirm the test is meaningful, and optionally open a pull request that makes CI depend on it:

mylonite gate reference:vulnerable                                   # find -> test -> confirm
mylonite gate --target-file app.yaml --authorize my-app --open-pr    # ...and open the PR

gate does not touch your repository unless you ask it to. By default it writes its files under .mylonite/gate/ — the test, the weakness record, the confirmation report — and prints the exact git and gh commands so you can commit and open the PR yourself. Add --open-pr to have it create the branch, commit and open the PR; add --workflows to also write two CI templates (a cheap per-PR gate and a nightly discovery run).

The pull request carries the finding, its OWASP/ASI/ATLAS/NIST tags, the supporting evidence, and a recommended fix that names the actual tool and argument the attack used. Full guide: docs/ci-gating.md. Behind a corporate network, see docs/enterprise-networking.md.

Commands

Command What it does Needs a model?
mylonite demo Replays the unguarded-vs-guarded comparison on the bundled practice app, offline. No (--live does)
mylonite scan The weakness-finding loop. --scaffold inspects a server and writes a starter target.yaml. Yes (except --scaffold)
mylonite generate Writes the pytest test from a confirmed weakness. No
mylonite validate Confirms a test is meaningful by running the comparison. --fast makes it cheaper. Yes
mylonite gate End to end: scan → generate → validate → optionally open a gating PR. Yes
mylonite report Terminal summary, SARIF 2.1.0, or a JSON bundle — each carrying the supporting evidence and compliance tags. No
mylonite plugins Lists installed plugins across all five extension points. No
mylonite version Prints the installed version. No

Two more commands, check (a structural, no-LLM pre-check) and ablate (scores each safeguard as load-bearing, redundant, security theatre, no-attack, or inconclusive), exist but are hidden and experimental — see docs/experimental.md.

--fast trades thoroughness for cost, and what it skips depends on the target: against your own app it skips the safeguard comparison itself, leaving a weaker check; against the bundled practice app the comparison is not optional, so it reduces the robustness checks instead.

Exit codes are a documented contract (0 success · 1 structural findings present, the experimental check --enforce · 2 configuration · 3 budget · 4 provider · 5 not confirmed · 6 generate failed · 7 validate failed · 8 PR step failed). A scan that finds something exits 0. Budget exhaustion always wins: a run that finds something AND runs out of --max-llm-calls exits 3, not 0 — the findings are still written to disk and (for gate) still turned into a test and gated, so nothing is lost. Full details in the CLI reference.

Remote MCP transport (SSE / streamable-HTTP), the versioned extension points, and entry-point plugins are covered in the architecture guide.

Compliance metadata

Every test and every finding carries tags from four frameworks: OWASP LLM Top 10 2025, OWASP ASI 2026, MITRE ATLAS, and NIST AI RMF. They ride into the pytest markers, the SARIF output and the JSON bundle, so a finding traces back to the control catalogue your auditors already use. See docs/standards-mapping.md.

Documentation

Full docs site: abidemialade.github.io/mylonite (or mkdocs serve from a checkout). Highlights:

Responsible use

Mylonite reproduces working attacks against AI agents. Use it only against targets you control or are contractually authorised to test. Every command that drives a real target — scan, gate, validate and ablate — refuses to run without an explicit --authorize flag naming that target: the value must match the target's declared scope, or its family name where no scope is declared. The bundled insecure practice app runs in-process and opens no network ports.

Full policy: SECURITY.md.

Contributing

Bug reports, adapter requests and attack-pattern submissions are welcome — see CONTRIBUTING.md for development setup, how to write a plugin, and the pull-request conventions. The five extension points (attack modules, test generators, validators, target adapters, compliance mappers) are versioned public API with reference implementations in this repository.

License

Apache License 2.0. See LICENSE and NOTICE.

Metadata

Release files for mylonite 0.11.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mylonite 0.11.0
File Size Uploaded
mylonite-0.11.0.tar.gz 810.2 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mylonite 0.11.0
File Interpreter ABI Platform
mylonite-0.11.0-py3-none-any.whl Python 3 none any Details

Total release size: 1.6 MB

Release files / mylonite-0.11.0.tar.gz

Download URL mylonite-0.11.0.tar.gz
Size 810.2 kB
Tags Source
SHA-256 checksum
How to use checksums
41fd26291b1c5c12f5cc97e90e7082c822cffc4258dd4558c8eb144487eb9309
BLAKE2b-256 checksum
How to use checksums
0ee84960de0ca2673e12c1f2e873dc24cda8f71908d4f78da89a7cc505d815bf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release files / mylonite-0.11.0-py3-none-any.whl

Download URL mylonite-0.11.0-py3-none-any.whl
Size 775.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
48f834f088628e354b43466379265cd2432e3551afabab4f3d1cddc58a4428f5
BLAKE2b-256 checksum
How to use checksums
f1f5fb25cf41fb3dc07c4b2781938b2fe9831d3752a96a93fef497e43b4d522c
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.11.0 This release

2 release files

0.10.4

2 release files

0.10.3

2 release files

0.10.2

2 release files

0.10.1

2 release files

0.10.0

2 release files

0.9.0

2 release files

0.8.6

2 release files

0.8.5

2 release files

0.8.4

2 release files

0.8.3

2 release files

0.8.2

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.8

2 release files

0.7.7

2 release files

0.7.6

2 release files

0.7.5

2 release files

0.7.4

2 release files

0.7.3

2 release files

0.7.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page