gaslight
A Lighthouse score for MCP agents. Point it at your agent, get a graded security report — and everything it confirms, it proves.
gaslight is a black-box security tester for AI agents built on the Model Context Protocol. You point it at a running MCP server; it connects like any other client, runs a suite of real probes against the tools it finds, and hands you a graded, screenshot-ready report. It never reads your source, your database, or an API key.
Think Lighthouse, but for agent security: five scores out of 100, a letter grade, and a picture of how far a breach could travel — so you can see at a glance whether your agent is solid or leaking, and exactly where.
The one thing that sets it apart from every "AI security" scanner that grades one model's output with another model: gaslight never trusts an opinion. A finding is only marked CONFIRMED when something physically happened — a unique token you planted arrived at a listener gaslight controls, real out-of-bounds data came back, or a call that should have been refused went through. If it says confirmed, it's real. Go fix it.
$ gaslight -- npx -y some-mcp-server
🛡 These are safe probes, not real break-ins — testing your guards, not stealing your data.
🧪 Network Egress Abuse (SSRF) — whether a URL tool can be aimed at addresses it should never reach
🔥 CONFIRMED — fetch_url reached http://127.0.0.1:<sink> (your canary arrived). No destination check.
🧪 Claim Integrity — whether a tool that promises to be read-only keeps that promise
🔥 CONFIRMED — create_invoice says "stages for approval; does not issue",
but the record it created shows status "issued". Its own read tools contradict its own description.
Network 42 Filesystem 100 Leakage 88 Authorization 60 Integrity 55
Grade: F · blast radius: THIS MACHINE ▓ YOUR NETWORK ▓ DATA LEAVING ░
Quick start
pip install gaslight
gaslight -- npx -y your-mcp-server # point it at your agent
That's it — no API key, no config. gaslight connects like any MCP client,
attacks the tools it finds, prints a graded report to your terminal, and writes
a shareable HTML report next to it. Add --json for machine-readable output, or
--baseline tools.json to catch a tool that changes on you in CI. What the
grades mean is under How to read a report; to drive it
from an AI assistant, see AGENTS.md.
The score
Every finding rolls up into five gauges, each 0–100, plus one letter grade:
| Gauge | The question it answers |
|---|---|
| Network | Can the agent be steered into reaching out, or sending your data, over the wire? |
| Filesystem | Do file tools stay in their lane, or can they be walked out of scope? |
| Leakage | Do secrets, keys or tokens slip out during ordinary use? |
| Authorization | Are destructive, consequential or costly actions actually gated? |
| Integrity | Do the agent's instructions and its tools' own promises hold under pressure? |
A breach never scores green — a confirmed exploit caps the metric and drops the grade. And the blast radius only lights up where an attack physically succeeded: working guards keep the lit area small, so the picture rewards real defense instead of just counting capabilities.
What it tests
Every agent, whatever its domain, is built on the same few primitives: tools that move data out, read resources, take consequential actions, and make claims about themselves. gaslight aims at the primitives, so the same tests apply to a hospital agent and a bank agent alike.
17 attack modules, all deterministic, all proven physically — including indirect prompt injection → exfiltration, SSRF, path traversal, code execution, confused-deputy, tool-metadata poisoning, memory poisoning, destructive-action authorization, verbose-error disclosure, denial-of-wallet (unbounded, costly calls), and the headline: claim integrity — it takes the safety promise a tool's description makes (the promise the driving model trusts) and checks whether it's still true, using the target's own read tools as the verification channel.
Plus a zero-call static surface pass (schema hygiene, hidden instructions in descriptions, tool-shadowing / homoglyph names) that flags red flags without firing a shot.
Coverage, mapped to a standard
gaslight is anchored to the OWASP MCP Top 10, not an ad-hoc list:
- 7 of 10 fully covered — secret exposure, tool poisoning, command injection, prompt injection, weak authorization, context over-sharing, and shadow-servers/rug-pulls.
- 2 out of scope by design — supply-chain/dependency tampering and internal audit logging. A black-box behavioral tool can't see those, and says so rather than faking a pass.
- 1 partial — privilege scope-creep (stateful over time).
Stating what it doesn't test is the difference between an honest benchmark and a scanner that over-credits itself.
The optional LLM layer
The whole deterministic core needs no model and no key — it works out of the box. An LLM is an optional lens that makes probes smarter and the report richer, and there is one hard rule:
The LLM may aim an attack and explain a result. It may never decide whether the attack succeeded.
Every CONFIRMED still comes from a canary reaching our sink — never a model's
opinion. Turn it on with a key, or with --llm ollama for a free local
model (nothing leaves your machine — the right default for a security tool).
No model configured? You still get the full deterministic run. Every run states
plainly whether the LLM layer is on, off, and where its boundary is.
Extra modes
- Doctor mode — if the target won't start (built for an older SDK, needs a credential, wrong Node/Python), gaslight reads its startup output and gives you one plain sentence to fix it, instead of a wall of someone else's stack trace.
- Rug-pull guard —
gaslight --baseline tools.jsonrecords your tools on first run and flags any that changed since (a description quietly rewritten into a poisoned payload, a new dangerous parameter). Drop it in CI to catch a tool that turns malicious after you approved it.
How to read a report
- CONFIRMED — physically proven. Real. Fix it.
- not tested — no tool of the shape this attack needs, or a claim with no black-box way to verify it. Honest about what it couldn't reach, never silently green.
- A clean run means these attack vectors, tested, found nothing — a floor, not a certificate. It's the first security check to run on your agent, not the last.
Safety
Probes are harmless by construction: they only ever touch a sink gaslight
controls, an ordinary system file, or a synthetic canary record — never real
data. On the default --safe, an action with a real irreversible or external
effect is never triggered, and a soft, description-only signal never fires one.
When you point it at a target that needs a credential, use a throwaway/test
one via --env KEY=VALUE — never production.
Install
pip install gaslight # (pre-launch)
gaslight --help
Point it at any stdio MCP server:
gaslight -- npx -y some-mcp-server
gaslight --url https://my-agent.example/mcp # remote (HTTP+SSE)
gaslight --json -- python my_server.py # machine-readable, for CI
Use it from an AI agent (or CI)
gaslight is built to be driven by an AI assistant — just tell yours "use gaslight to check my agent." It ships an AGENTS.md that tells the AI how to run it, how to read the result, and — crucially — the safety contract, so it can run without any worry about touching your code:
--jsonemits a structured report on stdout (findings, grade, metrics); human text goes to stderr, so stdout stays clean JSON to parse.- Exit codes:
0clean ·1a CONFIRMED finding ·2the target couldn't start (with a plain-language reason). - The boundary: black-box (never reads your source),
--safeby default (never fires a destructive action), probes aimed at a local sink it controls, no API key, nothing leaves your machine. SeeAGENTS.mdfor the full contract.
Roadmap
v1 tests MCP-based agents — one target, through the MCP boundary. The next frontier is the agentic / multi-agent layer: an adaptive red-team agent that reasons about your agent and holds a real multi-turn conversation, plus agent-to-agent (A2A) attacks between agents that talk to each other. The physical-proof rule holds there too — the attacker gets smarter, the verdict stays deterministic.
Status
Pre-launch. 17 attack modules, validated against deliberately-vulnerable
benchmarks and 50+ real, independently-built community MCP servers, with matched
vulnerable/hardened fixtures and zero false positives on the controlled set. See
docs/ for the design specs and docs/TESTING_STRATEGY.md for the validation
record.
License
Apache-2.0 — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file gaslight-0.1.0.tar.gz.
File metadata
- Download URL: gaslight-0.1.0.tar.gz
- Upload date:
- Size: 402.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
68dd9b29ba531b7aea09a522dd55865c1a1f57287e8354dcd3f5d63d2ef35fe0
|
|
| MD5 |
34eb1c3c6791ab3e563afecdf3976e9f
|
|
| BLAKE2b-256 |
edb52f1f14f987ad8b8d98f65d5a018640bd8e8d59602ed0441e2370f724197f
|
File details
Details for the file gaslight-0.1.0-py3-none-any.whl.
File metadata
- Download URL: gaslight-0.1.0-py3-none-any.whl
- Upload date:
- Size: 130.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
uv/0.12.5 {"installer":{"name":"uv","version":"0.12.5","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"macOS","version":null,"id":null,"libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
12e6511f73084b2a3c2d9dae3a26ce314eb208b77d194294f5081fd0b11f6f62
|
|
| MD5 |
7472249b08f3e055f24ab2eb90fae45a
|
|
| BLAKE2b-256 |
cbad9835a31ef7b5e03c343bca8148ce2cc8ee1ffaa24402819267cbb2814fe7
|