English · Español
An MCP firewall that stops prompt-injection data theft by tracking where data came from,
not by guessing what attacks look like.
In One Piece, kairoseki (seastone) cancels Devil Fruit powers. Kairoseki does the same for your agent's most dangerous power: reading something an attacker wrote, and then quietly sending your data somewhere.
uv tool install git+https://github.com/AnthonyRiveraI/kairoseki
kairoseki scan # what can a single prompt injection do with your MCP setup?
kairoseki wrap # put every MCP server in your Claude / Cursor / VS Code config behind Kairoseki
kairoseki attack # replay 9 real-world attacks against your setup and get a grade
Why
An agent is exploitable by design when it has all three legs of the lethal trifecta:
- access to private data (your files, repos, inbox),
- exposure to untrusted content (a web page, a GitHub issue, an email), and
- a way to send data out (HTTP, email, a comment, a pull request).
Plug a filesystem server and a fetch server into Claude Code, Cursor or Claude Desktop and you have all three. This keeps happening in the real world:
| Incident | What happened |
|---|---|
| GitHub MCP exploit (May 2025) | A malicious public issue made an agent copy private repo data into a public pull request |
| Tool poisoning (Apr 2025) | Hidden instructions in a tool description stole ~/.cursor/mcp.json |
| MCPoison, CVE-2025-54136 (Jul 2025) | An approved MCP config was silently swapped later (rug pull) |
| Comment and Control (Apr 2026) | Injections in PR titles and comments made Claude Code, Gemini CLI and Copilot agents leak their own secrets |
Most defenses scan text for attack patterns, and attackers just rephrase. Kairoseki breaks the trifecta instead. The idea is inspired by CaMeL (Google DeepMind): track taint across the whole session and step in exactly when untrusted content, private data and an outbound channel meet.
What it does
| 🔗 Session taint across servers | Every kairoseki run in one agent session shares state. The web page comes from the fetch server, the secret from filesystem, the leak goes through github: Kairoseki still sees one chain. |
| 🧪 Secret fingerprints | Secrets seen in tool output or in a server's environment are fingerprinted (never stored). If one shows up in a later tool call, even base64, hex, URL-encoded or reversed, the call is denied. |
| 🫥 Redaction | API keys, tokens and private keys are replaced with [REDACTED:kind] before they reach the model. The model can't leak what it never saw. |
| ☠️ Poisoned tool detection | Tool descriptions and schemas with injection text, ANSI escapes or invisible Unicode are neutralized before the model reads them, and the tool is blocked. |
| 📌 Rug-pull pins | Tool definitions are pinned on first use. If a server changes one later, that tool is blocked until you re-approve it. |
| ✋ Approvals that fit your client | When the trifecta closes, Kairoseki asks you: an in-client prompt (MCP elicitation, in both protocol eras), or a one-time kairoseki approve K-1A2B3C from any terminal. |
| ⚔️ Attack lab and badge | kairoseki attack replays real attacks against a fully hijacked agent and checks, on the attacker's side, whether the canary leaked. |
| 🔍 Scanner | kairoseki scan labels every tool in your config and tells you if you already have the lethal trifecta. |
It is a transparent stdio proxy that speaks raw JSON-RPC, so it works with any MCP server and client, in
both the handshake era (initialize, 2024-11-05 → 2025-11-25) and the modern era (server/discover,
2026-07-28). Two dependencies: pyyaml and rich.
Quickstart
0. Prerequisites
| You need | Why | How to get it |
|---|---|---|
| uv (recommended) or pipx | Installs Kairoseki as an isolated command-line tool | macOS / Linux: curl -LsSf https://astral.sh/uv/install.sh | shWindows (PowerShell): powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex" |
| Python 3.10+ | Kairoseki is written in Python | uv downloads a suitable Python automatically if you don't have one. With pipx, install Python yourself. |
| Git | To install straight from GitHub | git-scm.com |
| An MCP client | Something to protect | Claude Code, Claude Desktop, Cursor, VS Code, Windsurf... |
| Node.js (optional) | Only if your MCP servers start with npx |
nodejs.org |
1. Install
Kairoseki is not on PyPI yet, so install it from GitHub:
uv tool install git+https://github.com/AnthonyRiveraI/kairoseki
# or with pipx:
pipx install git+https://github.com/AnthonyRiveraI/kairoseki
Then make sure the kairoseki command is on your PATH, and open a new terminal:
uv tool update-shell # or: pipx ensurepath
kairoseki --version # should print: kairoseki 0.1.2
kairoseki: command not found, or your MCP client can't start it? The tool lives in~/.local/bin(%USERPROFILE%\.local\binon Windows). Runuv tool update-shell, then fully restart your terminal and your MCP client so they pick up the newPATH.kairoseki wrapalso writes the absolute path into your config, which avoids the problem entirely.
To update later: uv tool upgrade kairoseki (or pipx upgrade kairoseki).
Windows: close your MCP client (or disable its Kairoseki-wrapped servers) before upgrading. While a client is running
kairoseki.exe, Windows locks the file anduv tool upgradefails with os error 32.
2. See your exposure
kairoseki scan
It reads the MCP configs of Claude Desktop, Claude Code, Cursor, Windsurf and VS Code, starts each server, and
labels every tool as private, untrusted, sink and/or destructive. For Claude Code that includes project
servers (.mcp.json), user servers and the current project's local servers in ~/.claude.json; add
--all-projects to include every project's local servers.
3. Wrap your servers
kairoseki wrap # all detected configs (a .kairoseki.bak backup is written first)
kairoseki wrap --undo # restore
kairoseki status # which servers are protected, and what each live session has seen
Or wrap a single server by hand. Every client uses the same pattern, kairoseki run --name <name> -- <original command>:
{
"mcpServers": {
"fetch": {
"command": "kairoseki",
"args": ["run", "--name", "fetch", "--", "uvx", "mcp-server-fetch"]
}
}
}
With Claude Code:
claude mcp add fetch -- kairoseki run --name fetch -- uvx mcp-server-fetch
Restart your client and check that the server connects (with Claude Code: claude mcp get fetch). That's it.
Testing with
@modelcontextprotocol/server-filesystem? It replaces the directories you pass on its command line with the client's roots. Claude Code sends the project directory, so the server will serve that folder, not the one in your config.
4. Try to break it
kairoseki attack # grade your current policy
kairoseki attack --badge kairoseki.svg # and get a README badge
How decisions are made
flowchart LR
A[tool call] --> B{denied by policy,<br/>poisoned or rug-pulled?}
B -- yes --> X[⛔ deny]
B -- no --> C{arguments contain a<br/>secret seen this session?}
C -- yes --> X
C -- no --> D{session saw untrusted content<br/>AND private data<br/>AND this tool is a sink?}
D -- yes --> Q[✋ ask the user]
D -- no --> E{untrusted content seen AND private data<br/>flows into an unclassified tool?}
E -- yes --> Q
E -- no --> OK[✅ forward to the server]
OK --> R[result: sanitize, redact,<br/>fingerprint secrets, update taint]
| Mode | Behaviour |
|---|---|
monitor |
Never blocks. Logs what would have happened (redaction still applies). Good for your first week. |
balanced (default) |
Denies exfiltration, poisoned tools and rug pulls. Asks before the lethal trifecta closes. |
strict |
Also asks before any sink or destructive tool once untrusted content entered the session. |
When Kairoseki asks:
- In-client prompt. If your client supports MCP elicitation, you get an "Allow this call once?" form.
Kairoseki sends
elicitation/createin the handshake era and aninput_requiredresult (SEP-2322) in the 2026-07-28 era. - Terminal. Otherwise the agent gets a clear refusal with an id. Run
kairoseki approve K-1A2B3C, then ask the agent to retry. Approvals are single-use, bound to the exact arguments, and expire after 10 minutes.
Learn more in docs/how-it-works.md.
Policy
kairoseki init # writes a commented kairoseki.yaml
mode: balanced
redact:
secrets: true
pii: false
servers:
github:
tools:
create_or_update_file: [sink, destructive] # explicit labels replace the heuristics
allow: [search_repositories] # never ask (redaction still applies)
deny: [delete_repository] # always block
Kairoseki looks for --policy, then $KAIROSEKI_POLICY, then ./kairoseki.yaml, then ~/.kairoseki/kairoseki.yaml.
Other commands
kairoseki log # recent decisions, redactions and detections
kairoseki status # protected vs unprotected servers, and live sessions (--all for ended ones)
kairoseki session --explain # which session this process joins, and why
kairoseki approve # list pending approvals
kairoseki pins list # servers whose tools changed since you pinned them
kairoseki pins approve github
How it compares
| Kairoseki | Pattern scanners / guardrail models | mcp-context-protector | Enterprise MCP gateways | |
|---|---|---|---|---|
| Blocks rephrased or novel injections (data-flow based) | ✅ | ❌ | ❌ | ➖ |
| Taint shared across servers in one session | ✅ | ❌ | ❌ | ➖ |
| Detects encoded secret exfiltration | ✅ | ➖ | ❌ | ➖ |
| Rug-pull pinning | ✅ | ❌ | ✅ | ✅ |
| Poisoned descriptions, ANSI, invisible Unicode | ✅ | ✅ | ✅ | ➖ |
| Runs locally, no service or API key | ✅ | ➖ | ✅ | ❌ |
| Measurable attack lab with a grade | ✅ | ❌ | ❌ | ❌ |
➖ = depends on the product. Kairoseki borrows trust-on-first-use pinning and ANSI sanitization from Trail of Bits' excellent mcp-context-protector. The two are complementary.
Limitations (please read)
- It only sees MCP traffic. Built-in client tools (for example Claude Code's own
BashorWebFetch) never go through MCP. Pair Kairoseki with your client's permission rules. - Labels are heuristics. Tool names and descriptions are read as verb + object (
get_issue,send_email). They can be wrong, and the policy lets you fix them. Server annotations can only add risk, never remove it. - Taint is per session and coarse on purpose. Once untrusted content is in the context, Kairoseki assumes it may
have influenced everything after it. That is what makes it robust, and it is why
strictmode asks more often. - Secret fingerprints are exact matching. They catch a secret copied whole, base64/hex/URL-encoded, reversed, or
split into pieces of 12+ characters, but not one interleaved character by character or run through a custom
cipher. The lethal-trifecta rule is the safety net that does not need to recognize the data, so be careful with
monitormode and policyallow:entries, which turn it off. - stdio servers only in v0.1. Streamable HTTP servers are on the roadmap.
- It is not a sandbox. A malicious server binary can still do anything your user account can. Kairoseki protects against malicious content, not malicious code.
Roadmap
- Publish on PyPI (
pipx install kairoseki) - Streamable HTTP transport
- 🐌 Den Den Mushi: Kairoseki calls your phone to approve risky actions
- OpenTelemetry export of decisions
- More attack scenarios. Propose one!
Contributing
The most valuable contribution is a new attack scenario: a real write-up turned into a replayable test. See CONTRIBUTING.md. Found a bypass? Please report it privately.
License
MIT © Anthony Rivera
Metadata
Release files for kairoseki 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| kairoseki-0.1.2.tar.gz | 232.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| kairoseki-0.1.2-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 292.3 kB
Release files / kairoseki-0.1.2.tar.gz
| Download URL | kairoseki-0.1.2.tar.gz |
|---|---|
| Size | 232.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
305dc3a9321f1aea9f531c3708b3f911c963500d01a42407a2e5021c17b3766a
|
|
BLAKE2b-256 checksum How to use checksums |
2b18580110572a4264a34c9deab37e09f528b98911d3707c39c5f551055bdb83
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / kairoseki-0.1.2-py3-none-any.whl
| Download URL | kairoseki-0.1.2-py3-none-any.whl |
|---|---|
| Size | 60.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
87396b3d9b5b1b76795ea7ec1baa813b9129ed0c9d258acb7d35b9ec283d9e67
|
|
BLAKE2b-256 checksum How to use checksums |
bc6292dcb34c75860eaa9cdec6a0908e9f2928d081aff9e8e4426925c27e2237
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log