Skip to main content

mcp-policy-gateway

CI Python 3.11+ License: Apache 2.0

An authorising reverse proxy for MCP servers. It sits between an agent and the tools it can reach, and decides — per call, on the actual arguments — whether the call happens.

MCP servers are being wired directly into agent loops with nothing between "the model decided to call this tool" and "the tool ran". Those are two different events. This puts something in the gap.

  agent ──▶ mcp-policy-gateway ──▶ [ hass ] [ github ] [ filesystem ] ...
                   │
                   ├── policy.yaml   per-tool and per-argument rules, per token
                   ├── audit.jsonl   every decision, hash-chained
                   └── dry-run       log what would break before it breaks

The demo

An assistant is asked to summarise a web page. The page contains instructions the user cannot see. The assistant follows them.

$ uv run demo/prompt_injection.py

Round 1 — agent connected directly to the MCP server, as almost every MCP setup is wired today:

  agent calls fetch_url(url='https://smart-home-weekly.example/morning-routines')
    page fetched — it contains hidden instructions
    the agent now believes it has three maintenance steps to run

  agent calls ha_remove_entity(entity_id='camera.driveway')
    EXECUTED Entity camera.driveway permanently removed.

  agent calls ha_call_service(domain='lock', service='unlock', entity_id='lock.front_door')
    EXECUTED Called lock.unlock on lock.front_door.

  agent calls ha_restart()
    EXECUTED Home Assistant is restarting.

Round 2 — same attack, same agent, same server, gateway in between:

  tools visible to the agent: 3
    fetch_url, ha_call_service, ha_get_overview

  agent calls ha_remove_entity(entity_id='camera.driveway')
    not in the tool list the agent was given; calling it by name anyway
    BLOCKED  Denied by gateway policy: the call to 'ha_remove_entity' was not
             permitted (destructive operations require a human at a keyboard).

  agent calls ha_call_service(domain='lock', service='unlock', entity_id='lock.front_door')
    BLOCKED  Denied by gateway policy: the call to 'ha_call_service' was not
             permitted (physical security domains are out of scope for an assistant).

  agent calls ha_restart()
    BLOCKED  Denied by gateway policy: the call to 'ha_restart' was not
             permitted (destructive operations require a human at a keyboard).
                            no gateway    with gateway
  entities remaining                 5               6
  entities deleted                   1               0
  front door                  unlocked          locked
  restarted                       True           False

The model was compromised in both rounds. It made exactly the same three calls. What changed is that in round two those calls had to get past something that was not reading the web page.

On the agent in this demo. It is scripted, not a language model. Whether a given model falls for a given payload depends on the model, the phrasing and the day; building the demo on that would make it a benchmark of one model's gullibility rather than a test of this software, and it would not run in CI. The claim here is about what happens after the model is convinced. The gateway's guarantee does not depend on the model resisting — which is the whole reason to have a gateway rather than a better system prompt.


Install

Not on PyPI yet. From a checkout:

git clone https://github.com/Spuddy10345/mcp-policy-gateway
cd mcp-policy-gateway
uv sync --extra dev
uv run mcp-policy-gateway --help

Or install the built wheel as a standalone tool:

uv build && uv tool install dist/*.whl

Use

Write a policy:

version: 1

upstreams:
  hass:
    transport: stdio
    command: uvx
    args: [hass-mcp]
    env:
      HA_URL: ${HA_URL}
      HA_TOKEN: ${HA_TOKEN}

policies:
  general-assistant:
    default: deny                # a policy that fails open is not a policy
    rules:
      - name: read-the-house
        effect: allow
        tools: [ha_get_*, ha_search]

      - name: safe-domain-actuation
        effect: allow
        tools: [ha_call_service]
        when:
          args.domain:
            in: [light, climate, switch, media_player]
        reason: actuation limited to non-safety-critical domains

      - name: never-unlock
        effect: deny
        tools: [ha_call_service]
        when:
          args.domain:
            in: [lock, alarm_control_panel]
        reason: physical security is out of scope for an assistant

Check it before you trust it:

$ mcp-policy-gateway validate -c policy.yaml --strict
$ mcp-policy-gateway explain -c policy.yaml --policy general-assistant \
    --tool ha_call_service --args '{"domain": "lock", "service": "unlock"}'

DENY  ha_call_service  (policy 'general-assistant', upstream 'hass')
  matched:   never-unlock
  reason:    physical security is out of scope for an assistant
  listed:    yes
  trace:
    [safe-domain-actuation] FAIL args.domain='lock' not in ['light', 'climate', ...]
    [never-unlock] PASS args.domain satisfied

Then point your MCP client at the gateway instead of the server:

Preparing PyPi Release{
  "mcpServers": {
    "hass": {
      "command": "mcp-policy-gateway",
      "args": ["run", "-c", "/path/to/policy.yaml", "--policy", "general-assistant"]
    }
  }
}

With one upstream and default settings the tool names are unchanged, so this is a drop-in swap for the server it fronts.


What it does

Argument-level rules. ha_call_service can turn on a lamp or unlock your front door depending on one string. Allow-listing by tool name cannot tell those apart, which is why rules reach into arguments:

when:
  args.domain: {in: [light, climate]}
  args.entity_id: {matches: "(light|climate)\\..+"}
  args.targets[*].id: {not_matches: "lock\\..*"}    # every element must pass

Capability-scoped tokens. Each caller gets a token bound to one named policy — readonly, general-assistant, automation-editor — not a blanket credential. Tokens live in the config as SHA-256 digests, so the config can be committed with no secret in it.

Denied tools are never advertised. A tool your policy can never permit is left out of tools/list entirely, so an injected model is not told it exists. Tools that are conditionally allowed stay listed, because the model needs them for the permitted case. Hiding is blast-radius reduction, not the control — tools/call is enforced independently, and a client that calls a hidden tool by name is still refused.

Rate limits. Token buckets, per caller or shared, consumed only by calls that are actually allowed. A rule naming several buckets debits all or none.

Dry-run. --dry-run logs every decision and enforces none of them, so you can point a new policy at real traffic and read what it would have broken before it breaks it. Available per-policy too, so one policy can be under test while the rest stay enforced.

A tamper-evident audit log. Every decision is one JSON line: who called what, with which arguments (redacted), which rule fired, and whether it ran. Records are hash-chained, so editing or deleting one breaks the chain:

$ mcp-policy-gateway audit verify audit.jsonl
line 2: content does not match its hash (stored a7895c2e372b...,
        recomputed ac50a6edb9f2...) — record was edited
FAILED: 1 problem(s) across 4 records

A linter for your policy. The failure mode of a policy language is a rule that parses, starts, and enforces less than its author believed. validate catches the fail-open shapes: unreachable rules, allow rules with no constraints, regexes that narrow nothing, optional where it weakens a check.

$ mcp-policy-gateway validate -c policy.yaml
warn   policy 'assistant' no-secrets: is unreachable: reads already matches
       everything this rule would, unconditionally. Rules are evaluated in
       order, first match wins.

Design decisions worth knowing about

Rules are first-match-wins, top to bottom. No priorities, no implicit "deny overrides allow" pass. A reviewer reads the rules in order and knows what happens — which matters more than expressive power in a file whose whole job is to be reviewed.

matches is a full match, never a substring search. An allow rule written light\..* does not also permit evil/light.bypass. ^ and $ are the two characters people forget, so the language does not depend on them.

A missing argument fails a constraint. when: {args.domain: {in: [light]}} against a call with no domain at all is not satisfied. Vacuous truth in an allow rule is a bypass. Say optional: true if you meant it — and the linter will mention that you did.

Wildcards are universally quantified. args.entity_id[*] requires every element to pass, so a batch call cannot smuggle lock.front_door in alongside two lamps.

Denials come back as tool results, not protocol errors. The agent reads "policy denied this, tell the user" and can act on it. A JSON-RPC error would surface to most clients as a transport failure and get retried.

Unknown config keys are rejected. deney: silently ignored would leave a rule allowing everything it was meant to block.


Non-goals

  • Not a sandbox. An allowed call runs with whatever authority the upstream server has. This decides whether a call happens, not what it can do once it does.
  • Not a prompt-injection detector. No classifier, no scanning of tool arguments for "suspicious" content. It enforces a policy you wrote. That is the point: it works against attacks nobody has thought of yet.
  • Not an LLM guardrail product. It never sees the conversation, only the calls.
  • Not a secrets manager. Upstream credentials come from your environment.
  • Resources and prompts are not proxied yet — only tools. They are not exposed rather than passed through unchecked, so the omission fails closed.
  • Truncation of the audit log is not detectable by hash-chaining alone. Chaining proves nothing was changed within what remains; catching a lopped-off tail needs an external anchor. See THREAT_MODEL.md.

Documentation

Transports

stdio streamable HTTP
Upstreams yes yes
Serving clients yes yes
Identity fixed at launch (--policy / --token) bearer token per request

Under stdio the client launches the gateway, so there is exactly one caller and the OS process boundary is the authentication. Over HTTP one process serves many callers, each with their own token and policy, and identity is read from the request being served.

Development

uv sync --extra dev
uv run pytest              # 248 tests
uv run ruff check .
uv run pyright src

License

Apache 2.0.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mcp_policy_gateway-0.1.0.tar.gz (88.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mcp_policy_gateway-0.1.0-py3-none-any.whl (53.7 kB view details)

Uploaded Python 3

File details

Details for the file mcp_policy_gateway-0.1.0.tar.gz.

File metadata

  • Download URL: mcp_policy_gateway-0.1.0.tar.gz
  • Upload date:
  • Size: 88.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mcp_policy_gateway-0.1.0.tar.gz
Algorithm Hash digest
SHA256 78d8002b2c6c625197f154286784c66603b5aaaa5f9eddc1e61cc8b8e21ef547
MD5 a2f3eaa907922050e1c39d249b5c6596
BLAKE2b-256 90864732f15313b73fcf205cc25f0530f8b42f7c79800c61056a27060963e6c7

See more details on using hashes here.

Provenance

The following attestation bundles were made for mcp_policy_gateway-0.1.0.tar.gz:

Publisher: publish.yml on Spuddy10345/mcp-policy-gateway

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mcp_policy_gateway-0.1.0-py3-none-any.whl.

File metadata

File hashes

Hashes for mcp_policy_gateway-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 4f8c162f83b8fba83e5dc4d502997f35b0779491fd3c8f14d19395edf3894bef
MD5 0cf9f3eb271a6fd0146f8c52c5e69237
BLAKE2b-256 0f1f1b277ddde894f3d065d92e67b200e8c40fd9767cc7bc6702f78ee5387f63

See more details on using hashes here.

Provenance

The following attestation bundles were made for mcp_policy_gateway-0.1.0-py3-none-any.whl:

Publisher: publish.yml on Spuddy10345/mcp-policy-gateway

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page