Mylonite
A safe model is not the same thing as a safe app. A top-tier model can shrug off every generic prompt-injection you throw at it and still hand an attacker a win — because the weakness is in how your app is wired, not in how the model behaves. Mylonite checks whether your app's own safeguards are what stop an attack, writes a test for every weakness it finds, and wires that test into CI so a future model upgrade cannot quietly remove the protection.
For teams shipping MCP or agentic apps who want CI-enforced regression coverage on the AI layer.
Point Mylonite at any MCP (Model Context Protocol) app. Mylonite's own agent reads your
system prompt and tool descriptions and drives your tools itself — attacking the part an
ordinary code scanner cannot see — finds weaknesses specific to your app, and turns each
one into a pytest test that gates your CI. The API key you configure pays for that
agent: the planner driving the attack, the customiser crafting payloads, and the judge
scoring the result — not for a call your own app's agent made.
Mylonite deliberately does not review your ordinary application code. That is what SAST/DAST tools are for.
How it works
The idea is simple: run the same attack twice.
- Once against your app as it is.
- Once with the safeguard in place.
A weakness is only reported if the attack succeeds the first time and is stopped the second — repeated several times over, so that a model having a good or bad day cannot decide the outcome. (In the codebase and the docs this is the differential, or control-efficacy check.)
That answers a question a one-off scan cannot: is your safeguard doing the work, or is the model's current good behaviour doing it? Only one of those survives a model upgrade.
Two levels of proof
Mylonite always tells you which one you got.
| What plays the "safeguard" role | What a kept finding tells you | How to get it |
|---|---|---|
| Your own safeguard, switched off and on | Your implementation is doing the work. The stronger result. | Declare control_env in your target.yaml |
| A standard safeguard Mylonite applies at the boundary | The attack is real, and this kind of safeguard closes it — but what stopped it was Mylonite's stand-in, not your code | The default; works on any app, no setup |
The second is genuinely useful, and it is what runs against most real apps. It is a weaker statement than the first, and no Mylonite output will dress it up as the stronger one.
How it has been tested
Mylonite ships an independent verification harness that scores it against material it did not write — runnable third-party MCP servers and published academic benchmarks. The harness is in this repository and you can run it yourself.
Results against third-party targets
- A kept finding on a third-party MCP email server. The attack succeeded 5 times out
of 5 against the server as shipped, and 0 times out of 5 with the safeguard in
place. Scope of that run: the server needed two fixes before it would start at all (a
launch wrapper, and a one-line bug in its own
send_email), the weakness only appears when the app's system prompt tells the agent to send without asking, and the safeguard was Mylonite's boundary stand-in rather than a second build of the server. Full detail in the capability matrix. - No false alarms on a third-party server with nothing wrong with it (Enkrypt's
echo_mcp). Running the same comparison against a hardened third-party server is still outstanding — see verification. - A real weakness found in a published vulnerable-MCP corpus (MCPSecBench). When that finding was re-run to confirm it, it failed to reproduce (0 of 3 runs), so Mylonite discarded it instead of shipping a flaky test.
- The judge checked against real third-party examples — AgentDojo transcripts from models that genuinely fell for attacks, not examples we wrote ourselves.
- The safety rails hold under test. A check that could not be run is never reported as a pass; a score with nothing to measure is labelled as such; work outside the tool's scope is marked rather than graded.
Results that came back negative
Published for the same reason the positive ones are.
- DVMCP recall is unmeasured, not a result. The published 0.9.0 figure read 0/8, but the harness that produced it had two defects that could each force a miss regardless of what the scan actually found: the scorer folded any untested challenge into "missed" instead of reporting it separately (fixed, issue #136), and the documented campaign workflow saves a report with no per-attempt weakness class, so the found count reads 0 whatever the scan did (still open). Treat 0/8 as unmeasured until Layer 1 is re-run.
- On InjecAgent (100 cases per split, using a local
llama3.2:3b) the judge scored F1 1.000 on the direct-harm split and F1 0.833 at 0.714 recall on the data-stealing split in 0.10.0 (0.9.0 measured 0.400 at 0.25 recall). That recall rests on only 7 attacks that succeeded, so we record it as unresolved at this sample size, not as an improvement. The gap between the splits is the finding, so both are published. - Judge agreement of F1 0.41 against AgentDojo's own labels. Mylonite's judge asks "did harm actually happen?"; AgentDojo asks "was the exact goal achieved?". Some of that gap is a genuine difference in question, which we have not resolved.
Current limits
- Against a single-build app, only the weaker statement is available. With no
control_envto switch, the safeguard is Mylonite's stand-in rather than your code, and no output will claim otherwise. - Finding nothing is the normal outcome for a well-built app on a robust model. See below.
- The evidence rests largely on one model — Claude Haiku 4.5 — at small, deliberately cost-capped sample sizes.
- The published figures were measured between 25 June and 14 September 2026. The
benchmark results carry the version they were measured against
(
verification/results/0.9.0/andverification/results/0.10.0/); the run logs in the capability matrix do not. Figures have not been re-measured for every release since, so read them as a floor rather than a current reading. Per-release re-measurement is planned. - Run transcripts are not published. The harness and its scorers are, so you can produce your own numbers; you cannot yet audit ours.
- No third party has built a plugin against the extension points yet.
Full scorecard with caveats: docs/verification.md. Everything that limits the tool's reach is collected in docs/limitations.md.
Finding nothing is also a result
Worth knowing before you run it. Against a well-built app on a robust model, Mylonite will often correctly find nothing — that is the tool working, not failing. Proving a safeguard carries the security requires a weakness that actually lands, which in practice means a design flaw (an action with real consequences and no approval step, an unrestricted outbound request) or an app configured to act on its own.
That is why Try it starts with the bundled practice app rather than your code. That app is deliberately insecure, so it finds something every time and you can watch the machinery work before pointing it somewhere the honest answer may be "nothing".
Where this sits
Static scanners read your tool descriptions and flag whatever looks risky, leaving you to judge which flags matter. Model-eval harnesses swap models and score which behaves best. Mylonite's purpose is the step neither of those takes: run the attack against your app, then hold the model constant and switch only your safeguard. The result is evidence about your safeguard, not about how a description reads or how a model scored today.
Project status
Beta, and essentially a single maintainer — one outside contribution to date, the rest of the history from the maintainer and Dependabot. Over 2,300 tests, with CI (ruff, mypy, pytest, pre-commit) enforced on every pull request. The extension points are versioned public API, but nobody outside the project has built against them yet. If you are weighing this as a dependency in a security pipeline, pin a version — and read Known limitations first.
Install
pip install mylonite # the CLI, from PyPI
pip install "mylonite[demo]" # ...plus the bundled practice app
Python 3.11–3.14.
The [demo] extra installs the bundled practice app, and you need it for any
reference:... command — demo and scan reference:... alike.
Scanning your own app needs a model: an API key for a hosted provider, or no key at all for
one you host yourself (Ollama, vLLM, or a LiteLLM proxy — see
self-hosted models). scan --scaffold and report
never need a model at all, and demo replays recorded responses rather than calling one.
Try it
No API key, no install, one command (needs uv):
uvx --from "mylonite[demo]" mylonite demo
On every pull request, CI builds Mylonite from source and runs this command against that build on Python 3.14, on Linux and Windows, in an 80-column terminal.
With the [demo] extra already installed (see Install), the same demo is:
mylonite demo
Either way, that runs the comparison against the bundled practice app — deliberately insecure, runs in-process, opens no network ports — and prints what got through on the unguarded build next to what was stopped on the guarded one. Same attacks, two builds, different outcomes. That contrast is the point of the tool.
demo replays model responses recorded against those bundled apps, so it is offline and
gives the same answer every time. The scan engine, the adapters and the comparison are all
the real ones; only the model's replies are pre-recorded, and the output tells you which
model produced them and when. Every demo mode — replay and --live alike — turns off the
per-seed customiser and the LLM-judge fallback, so a verdict is decided purely by
deterministic predicates; scan itself runs each payload once by default, and the
repeat-run consensus belongs to validate (five iterations by default), which gate runs
for you. Treat the numbers as a demonstration of the machinery rather than a fresh
measurement of today's model — mylonite demo --live is the fresh measurement, and it does
call a model (by default one you host locally). Where a cell could not be
decided either way the table says so rather than showing it as a pass, and if a recording is
ever missing or out of date the command fails and explains why instead of reporting a clean
result it did not earn.
Then, with a model configured, the real thing:
mylonite scan reference:vulnerable # finds the weaknesses built into it
mylonite scan reference:guarded # same attacks, comes up clean
See the practice app for what is built into it and why.
Then point it at your own app
The first step is free — no API key, no model call, no spend.
# Inspect a server and write a starter target.yaml
mylonite scan --command "python" --arg "my_server.py" --scaffold app.yaml --scope my-app
--scaffold connects to your server, lists its tools, says which weakness classes apply to
it, and flags the tools whose actions have real consequences. It works from the names,
descriptions and schemas your server advertises, matching them against keyword patterns, so
treat everything it suggests as a hint to confirm rather than a verdict.
Proving which weaknesses actually land, and whether a safeguard closes them, is the scan
itself. By default the safeguard side is Mylonite's own stand-in guard at the tool boundary
— real evidence that this kind of guard closes the attack, not yet that your
implementation does; declare control_env in the target file to measure your own guard
instead (see docs/concepts.md).
That needs a model:
mylonite scan --target-file app.yaml --authorize my-app
Expect this to find less than the practice app did — often nothing. See Finding nothing is also a result above, and docs/limitations.md for where the tool's reach genuinely ends.
From a scan to a pull request that gates CI
mylonite gate runs the whole sequence — find a weakness, write a test for it, confirm the
test is meaningful, and optionally open a pull request that makes CI depend on it:
mylonite gate reference:vulnerable # find -> test -> confirm
mylonite gate --target-file app.yaml --authorize my-app --open-pr # ...and open the PR
gate does not touch your repository unless you ask it to. By default it writes its
files under .mylonite/gate/ — the test, the weakness record, the confirmation report — and
prints the exact git and gh commands so you can commit and open the PR yourself. Add
--open-pr to have it create the branch, commit and open the PR; add --workflows to also
write two CI templates (a cheap per-PR gate and a nightly discovery run).
The pull request carries the finding, its OWASP/ASI/ATLAS/NIST tags, the supporting evidence, and a recommended fix that names the actual tool and argument the attack used. Full guide: docs/ci-gating.md. Behind a corporate network, see docs/enterprise-networking.md.
Commands
| Command | What it does | Needs a model? |
|---|---|---|
mylonite demo |
Replays the unguarded-vs-guarded comparison on the bundled practice app, offline. | No (--live does) |
mylonite scan |
The weakness-finding loop. --scaffold inspects a server and writes a starter target.yaml. |
Yes (except --scaffold) |
mylonite generate |
Writes the pytest test from a confirmed weakness. |
No |
mylonite validate |
Confirms a test is meaningful by running the comparison. --fast makes it cheaper. |
Yes |
mylonite gate |
End to end: scan → generate → validate → optionally open a gating PR. | Yes |
mylonite report |
Terminal summary, SARIF 2.1.0, or a JSON bundle — each carrying the supporting evidence and compliance tags. | No |
mylonite plugins |
Lists installed plugins across all five extension points. | No |
mylonite version |
Prints the installed version. | No |
Two more commands, check (a structural, no-LLM pre-check) and ablate (scores each
safeguard as load-bearing, redundant, security theatre, no-attack, or inconclusive), exist
but are hidden and experimental — see docs/experimental.md.
--fast trades thoroughness for cost, and what it skips depends on the target: against
your own app it skips the safeguard comparison itself, leaving a weaker check; against the
bundled practice app the comparison is not optional, so it reduces the robustness checks
instead.
Exit codes are a documented contract (0 success · 1 structural findings present,
the experimental check --enforce · 2 configuration · 3 budget · 4 provider ·
5 not confirmed · 6 generate failed · 7 validate failed · 8 PR step failed). A scan that finds
something exits 0. Budget exhaustion always wins: a run that finds something AND runs
out of --max-llm-calls exits 3, not 0 — the findings are still written to disk and
(for gate) still turned into a test and gated, so nothing is lost. Full details in the
CLI reference.
Remote MCP transport (SSE / streamable-HTTP), the versioned extension points, and entry-point plugins are covered in the architecture guide.
Compliance metadata
Every test and every finding carries tags from four frameworks: OWASP LLM Top 10 2025, OWASP ASI 2026, MITRE ATLAS, and NIST AI RMF. They ride into the pytest markers, the SARIF output and the JSON bundle, so a finding traces back to the control catalogue your auditors already use. See docs/standards-mapping.md.
Documentation
Full docs site: abidemialade.github.io/mylonite
(or mkdocs serve from a checkout). Highlights:
- Quickstart · Test your own app — install and point it at your MCP server.
- Weakness classes · Attack modes — what is tested, and how the attacks work.
- The validation engine — how the safeguard comparison works.
- Independent verification — the full scorecard against material Mylonite did not write.
- Known limitations — where the tool's reach ends, in one place.
- Reading the results · CLI reference · target.yaml.
- CI gating · Re-validate on a new model — keep the gate proving your safeguard as models change.
- Architecture · Plugin authoring · Threat model.
- ROADMAP.md · CONTRIBUTING.md · GOVERNANCE.md · SECURITY.md.
Responsible use
Mylonite reproduces working attacks against AI agents. Use it only against targets you
control or are contractually authorised to test. Every command that drives a real target —
scan, gate, validate and ablate — refuses to run without an explicit --authorize
flag naming that target: the value must match the target's declared scope, or its family
name where no scope is declared. The bundled insecure practice app runs in-process and opens
no network ports.
Full policy: SECURITY.md.
Contributing
Bug reports, adapter requests and attack-pattern submissions are welcome — see CONTRIBUTING.md for development setup, how to write a plugin, and the pull-request conventions. The five extension points (attack modules, test generators, validators, target adapters, compliance mappers) are versioned public API with reference implementations in this repository.
License
Metadata
Release files for mylonite 0.11.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| mylonite-0.11.0.tar.gz | 810.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| mylonite-0.11.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 1.6 MB
Release files / mylonite-0.11.0.tar.gz
| Download URL | mylonite-0.11.0.tar.gz |
|---|---|
| Size | 810.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
41fd26291b1c5c12f5cc97e90e7082c822cffc4258dd4558c8eb144487eb9309
|
|
BLAKE2b-256 checksum How to use checksums |
0ee84960de0ca2673e12c1f2e873dc24cda8f71908d4f78da89a7cc505d815bf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency logRelease files / mylonite-0.11.0-py3-none-any.whl
| Download URL | mylonite-0.11.0-py3-none-any.whl |
|---|---|
| Size | 775.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
48f834f088628e354b43466379265cd2432e3551afabab4f3d1cddc58a4428f5
|
|
BLAKE2b-256 checksum How to use checksums |
f1f5fb25cf41fb3dc07c4b2781938b2fe9831d3752a96a93fef497e43b4d522c
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 2, 2026.
Transparency log