Skip to main content

MCPSentinel

+----------------------------------------------------------------+
|                          MCPSENTINEL                           |
|       Security review for Model Context Protocol servers        |
|                     Read-only by default                       |
+----------------------------------------------------------------+

Discover MCP metadata. Triage suspicious intent. Review changes before you trust them.

CI PyPI Python License MCP Registry GitHub Action

MCPSentinel is a precision-first security scanner for Model Context Protocol servers. It treats a static rule hit as a candidate, then applies semantic intent analysis before reporting it. This keeps the fast coverage of pattern matching without making every normal-looking fetch or delete tool a noisy vulnerability.

Start in 60 seconds

python -m pip install mcp-guardian-scan
mcpsentinel                    # safe, no-write onboarding
mcpsentinel scan http://localhost:8000/mcp

The first command you see is deliberately friendly and branded, while scan output stays free of decorative text when you select json or sarif for automation:

$ mcpsentinel
+----------------------------------------------------------------+
|                          MCPSENTINEL                           |
|       Security review for Model Context Protocol servers        |
|                     Read-only by default                       |
+----------------------------------------------------------------+

Welcome to MCPSentinel 0.7.0

1. Run your first offline scan:
   mcpsentinel scan http://localhost:8000/mcp

2. Save CI-friendly output and fail on high-severity findings:
   mcpsentinel scan http://localhost:8000/mcp --format sarif --output results.sarif --fail-on high
I want to… Start here
inspect one local or remote server Scan a server
add a review gate to CI GitHub Action
expose scanning to an AI client MCP-native scanner
run it in a container Container image
understand scope and limits What MCPSentinel can—and cannot—tell you

The review loop

discover metadata  ->  static candidates  ->  semantic triage  ->  human review
                                                                        |
                                                                        v
                                                        explicitly approve baseline

What you get

  • MCP discovery over stdio and Streamable HTTP
  • configurable static pattern rules for tool, prompt, and resource descriptors, including tool poisoning, shadowing, cross-server, and OAuth confused-deputy signals
  • semantic triage: offline heuristic by default, optional OpenAI structured-output judge with bounded fallback
  • explicit baseline approval and rug-pull definition diffs
  • terminal, JSON, SARIF, and self-contained HTML risk reports
  • allow/deny policy configuration
  • explicit, Docker-sandboxed owned-tool validation with no network egress
  • GitHub Action and MCP-native scanner interfaces

Static scans are metadata-only. Dynamic invocation is a separate opt-in path described below and never runs from the GitHub Action or MCP-native server.

Install and onboard

Install the published package, then use the MCPSentinel CLI:

python -m pip install mcp-guardian-scan
mcpsentinel

Running mcpsentinel with no command starts a short, no-write terminal onboarding guide. It explains the read-only scan model, gives a copy-pasteable first scan, and keeps OpenAI optional. Use mcpsentinel onboard (or the alias mcpsentinel init) to show it again, or tailor the suggested command without contacting a server:

mcpsentinel onboard --target https://mcp.example.com/mcp
mcpsentinel onboard --target "python -m example_mcp_server" --transport stdio

The onboarding flow never asks for, stores, or transmits an API key. Use mcpsentinel --help or mcpsentinel scan --help for the complete reference.

For development from source:

python -m venv .venv
. .venv/bin/activate
python -m pip install -e '.[dev]'

Scan a server

For a Streamable HTTP server:

mcpsentinel scan http://localhost:8000/mcp

For a stdio server, quote its command as the target:

mcpsentinel scan "python -m example_mcp_server" --transport stdio

Or keep the executable and arguments separate. Arguments beginning with a dash need the --arg=value form:

mcpsentinel scan python --transport stdio --arg=-m --arg=example_mcp_server

Useful options:

# Machine-readable report and CI failure gate
mcpsentinel scan http://localhost:8000/mcp --format sarif --output results.sarif --fail-on high

# Visual portfolio-ready report
mcpsentinel scan http://localhost:8000/mcp --format html --output risk-report.html

# Use OpenAI's structured-output semantic judge (OPENAI_API_KEY is required)
mcpsentinel scan http://localhost:8000/mcp --judge openai --judge-model gpt-4o-mini

# Use a repository-local directory for reviewed baseline snapshots
mcpsentinel scan http://localhost:8000/mcp --baseline-dir .mcpsentinel/baselines

# Create or replace a baseline only after reviewing the report
mcpsentinel scan http://localhost:8000/mcp --baseline-dir .mcpsentinel/baselines --approve-baseline

Baseline snapshots are kept in ~/.mcpsentinel/baselines by default, but are never updated by an ordinary scan. A changed, added, or removed descriptor is surfaced as an MCP-B001 rug-pull review finding while the prior approved snapshot is preserved. Review the report, then use --approve-baseline deliberately to create or replace the snapshot. This prevents an unattended scan from silently accepting a rug-pull change.

The first scan reports that no approved baseline exists. That is an onboarding state, not a vulnerability finding. Establish a baseline only from a server version and environment you trust.

The risk score is a capped 0–100 weighted sum of severity and semantic confidence. It is a prioritization signal, not a claim that the server is safe or unsafe in isolation.

Semantic judges

--judge heuristic is the default and is fully offline. --judge openai requires OPENAI_API_KEY; --judge auto opts into using OpenAI when that key is present, otherwise it uses the heuristic. The OpenAI judge uses the Python SDK's Responses structured-output API, so an API response cannot bypass the scanner's expected verdict schema. Results are cached by descriptor hash plus a versioned judge/prompt identity in the baseline directory to avoid repeat API charges without retaining verdicts after judging methodology changes.

Each OpenAI judgement uses a 30-second client deadline and at most two SDK retries. Before an OpenAI request, MCPSentinel redacts common API keys, bearer credentials, private keys, and secret-valued JSON fields. Prompts are capped at 12,000 characters with field-aware head-and-tail excerpts, so a long descriptor cannot simply hide all final evidence behind filler. Candidate assessment uses a bounded concurrency of four requests. Redaction is defense-in-depth, not a guarantee that arbitrary sensitive metadata is safe to send. Choose heuristic when metadata must remain local.

If --judge auto encounters an OpenAI outage or malformed response, the scan completes with the offline heuristic and emits a visible report note; a fallback verdict is not cached as an OpenAI verdict. --judge openai remains strict and fails rather than silently changing the configured provider.

The semantic threshold defaults to 0.70. Candidate findings below it are withheld from the report; lower it only when you prefer recall over precision.

Custom static rules

Pass --rules path/to/rules.json to add rule objects to the built-in rules. Each rule has this shape:

{
  "id": "ORG001",
  "title": "Example organization policy",
  "category": "tool_poisoning",
  "severity": "high",
  "description": "Why this candidate deserves semantic review.",
  "patterns": ["(?i)example pattern"],
  "fields": ["description", "schema"]
}

Supported categories are prompt_injection, tool_poisoning, tool_shadowing, ssrf, secret_exfiltration, command_execution, destructive_operation, cross_server_attack, oauth_confused_deputy, and rug_pull.

Before regex evaluation, the scanner applies Unicode NFKC normalization, removes format controls such as zero-width characters, and collapses whitespace in an analysis-only view. Original descriptor text remains unchanged in reports and baselines. It intentionally does not rewrite cross-script homoglyphs because that would risk misrepresenting legitimate metadata; use the benchmark to track those coverage gaps before claiming support for them.

Policy configuration

--policy path/to/policy.json supplies organization-specific allow/deny controls. An allow selector suppresses matching static candidates; a deny selector emits a policy-enforced finding without relying on the semantic judge. Selectors can be rule IDs or objects scoped to a tool-name regex.

{
  "allow": [{"rule_id": "MCP003", "subject_pattern": "^controlled_fetch$"}],
  "deny": ["MCP002"],
  "semantic_threshold": 0.75
}

See examples/policy.json for a complete file. Keep policy files under source control and review changes as security-sensitive configuration.

Dynamic Docker validation

Dynamic testing is intentionally opt-in and limited to a server you own or a local test fixture. It requires an explicit acknowledgement, a pre-built local image, an explicit high-confidence tool name, and JSON arguments. The runner creates a fresh Docker container with no network, no host mounts, a read-only root filesystem, dropped capabilities, an unprivileged user, resource limits, and a call timeout. It never forwards the scan process environment into the container.

mcpsentinel scan "python -m my_server" --transport stdio \
  --dynamic --i-own-this-target \
  --dynamic-image my-mcp-server:test \
  --dynamic-entrypoint "python -m my_server" \
  --dynamic-invoke 'unsafe_tool={"fixture": true}'

The dynamic server image must already exist locally; MCPSentinel uses --pull=never. Every explicit tool invocation receives its own fresh container/session, so state from one selected tool cannot affect another. Docker is not needed for normal metadata scans. A dynamic response is retained only as a SHA-256 digest and content-type summary.

For each owned-target invocation, MCPSentinel also records the Docker process count immediately before and after the call, plus the number of copy-on-write filesystem changes reported by Docker. It never retains process arguments or filesystem paths. An additional process still running after the call produces MCP-D002; it is a review signal for background work, not evidence of a host escape. Credential-like response material produces MCP-D001 without writing response text to disk. This bounded telemetry does not trace syscalls, inspect arbitrary environment reads, or prove that no network connection was attempted—the container's --network=none boundary remains the network control.

The repository includes a deliberately local-only Docker fixture to verify this boundary end to end. It is excluded from the normal test suite because it needs a running Docker daemon and builds an image:

MCPSENTINEL_RUN_DOCKER_TESTS=1 pytest tests/test_dynamic_docker_e2e.py

GitHub Action

The repository root is a composite GitHub Action. It installs MCPSentinel, restores a scoped baseline cache, emits SARIF, and fails at the selected severity. It does not enable dynamic testing. Reference a release tag from another repository; pinning a full commit SHA is recommended for stricter supply-chain controls.

- uses: actions/checkout@3d3c42e5aac5ba805825da76410c181273ba90b1 # v7.0.1
- uses: actions/setup-python@5fda3b95a4ea91299a34e894583c3862153e4b97 # v7.0.0
  with:
    python-version: "3.12"
- uses: gentaArnezzi/MCPSentinel@v0.7.0
  id: mcpsentinel
  with:
    target: https://mcp.example.com/mcp
    transport: http
    fail-on: high
    policy: .mcpsentinel/policy.json
- uses: github/codeql-action/upload-sarif@d6317709a54fd87078d323eeb0e48ec331c8e621 # v3
  with:
    sarif_file: ${{ steps.mcpsentinel.outputs.sarif }}

Action scans preserve an approved baseline by default. Use approve-baseline: "true" only in a reviewed workflow on a protected branch, after the scan's output is accepted. Do not enable it for pull requests from contributors.

- uses: gentaArnezzi/MCPSentinel@v0.7.0
  if: github.event_name == 'push' && github.ref == 'refs/heads/main'
  with:
    target: https://mcp.example.com/mcp
    transport: http
    approve-baseline: "true"

Set OPENAI_API_KEY in the workflow only when choosing judge: openai or auto; heuristic remains the default. For example, expose a GitHub Actions secret only to the scan step with env: OPENAI_API_KEY: ${{ secrets.OPENAI_API_KEY }}.

MCP-native scanner

Run mcpsentinel-mcp to expose the scanner as the MCP tool scan_mcp_server over stdio. It is intentionally more constrained than the CLI: dynamic execution is unavailable, target configuration is operator-controlled, HTTP targets must be explicitly allowlisted, and stdio targets are disabled unless the operator enables them.

export MCPSENTINEL_ALLOWED_HOSTS="mcp.example.com,localhost"
mcpsentinel-mcp

The allowlist accepts either host or an exact host:port. HTTP redirects are refused, each discovery session has a 30-second deadline, and private or reserved addresses are denied by default. For public MCP-native HTTP scans, the validated DNS address set is pinned to the HTTP transport while the original hostname remains the HTTP Host and TLS SNI name; this prevents a second DNS lookup from changing a validated public hostname into an internal destination. For a deliberately trusted local network, set MCPSENTINEL_ALLOW_PRIVATE_HTTP_TARGETS=true alongside its exact allowlist entry.

Optional operator settings are MCPSENTINEL_MCP_BASELINE_DIR, MCPSENTINEL_RULES_PATH, MCPSENTINEL_POLICY_PATH, MCPSENTINEL_MCP_JUDGE, and MCPSENTINEL_MCP_JUDGE_MODEL. Set MCPSENTINEL_ALLOW_STDIO_TARGETS=true only in a trusted local environment. The MCP caller cannot choose arbitrary policy files, baseline paths, or approve a baseline. For a deliberate one-time approval, an operator can set MCPSENTINEL_MCP_APPROVE_BASELINE=true, execute the reviewed scan, then remove the variable.

Registry publication

MCPSentinel is published to PyPI as mcp-guardian-scan and to the official MCP Registry. The PyPI package has a different name because mcpsentinel was unavailable; the product name, import package, and CLI stay MCPSentinel and mcpsentinel.

The concrete registry/server.json is kept version-locked with the package. The release workflow builds and audits the artifact, publishes it to PyPI through trusted publishing, then submits matching Registry metadata through GitHub OIDC. See registry/README.md for release details and the official package-type documentation.

Container image

Every non-prerelease GitHub Release publishes a versioned image and latest to GitHub Container Registry:

docker pull ghcr.io/gentaarnezzi/mcpsentinel:0.7.0
docker run --rm ghcr.io/gentaarnezzi/mcpsentinel:0.7.0 scan https://mcp.example.com/mcp --transport http

The first GHCR package may need its visibility set to Public in GitHub Packages by the repository owner. For local development, build the scanner image directly:

docker build -t mcpsentinel:local .
docker run --rm mcpsentinel:local scan https://mcp.example.com/mcp --transport http

The image intentionally has no Docker socket and cannot run the dynamic layer. Run dynamic validation from a trusted host with Docker configured.

Dataset

datasets/vulnerable_by_design holds controlled descriptor-level ground truth for regression tests across every default static rule. It expands deterministically to 200 synthetic descriptors: 35 hand-curated controls and 165 template-generated variants. It includes safe hard negatives, Unicode/zero-width evasion, non-English controls, metadata/schema variants, and intentionally uncovered controls; it contains no live third-party targets or runnable destructive payloads. The labelling protocol documents the provenance and review contract.

Run a reproducible accuracy and timing measurement with the offline judge:

mcpsentinel benchmark datasets/vulnerable_by_design/manifest.json --format json --output benchmark.json

The benchmark measures both raw static candidates and semantic findings against the dataset's expected reportable rules. It reports precision, recall, false-positive rate, F1, confusion-matrix counts, stage timings, per-category breakdowns, and the count for each provenance type. Its JSON and terminal reports include the source-manifest SHA-256 and scanner version for traceability. Ten bounded-fetch controls intentionally count as static false positives but semantic true negatives, so regressions in noise suppression are visible in CI or release review.

On the bundled 200-case corpus with the offline heuristic and default threshold (0.70), static candidates measure precision 0.932, recall 0.965, F1 0.948, and false-positive rate 0.006 (TP=136, FP=10, TN=1649, FN=5). Semantic triage measures precision 1.000, recall 0.965, F1 0.982, and false-positive rate 0.000 (TP=136, FP=0, TN=1659, FN=5). The five deliberate misses—four non-English prompt-injection controls and one metadata-placement destructive-operation control—remain visible rather than being excluded. The SSRF category shows why both stages are reported: static precision is 0.545 while semantic precision is 1.000 on its controlled cases.

This is a reproducible regression signal—not a claim about public MCP-server accuracy, recall, real-world false-positive rate, or superiority over another scanner. The 165 generated variants are useful coverage controls, not 165 independent real-world observations.

Curated public metadata v2

datasets/curated_public_metadata_v2 adds 428 literal tool descriptors from source-pinned, permissively licensed MCP implementations: 329 from AWS Labs' Apache-2.0 repository and 99 from GitHub's MIT-licensed MCP server. Every case records repository, full commit SHA, license, source path, line, and source-file SHA-256. The extractor only reads local checkouts and never contacts or invokes an upstream MCP server.

This is a negative-control benchmark: ordinary documented tool metadata is expected to produce no unbounded-risk finding. A tool that can perform a scoped cloud deletion or write operation is not automatically a vulnerability, so the corpus does not label source projects as insecure. At the 0.7.0 release configuration, the frozen heuristic produces zero candidates and a false-positive rate of 0.000 across 3,852 descriptor/rule negative pairs. Because it has no labelled positives, precision, recall, and F1 correctly display as n/a, not 1.000.

mcpsentinel benchmark datasets/curated_public_metadata_v2/manifest.json \
  --judge heuristic --format json --output benchmark-v2.json

The v2 corpus has one maintainer review and is explicitly marked independent-review-pending. It strengthens public-metadata false-positive evidence; it does not establish public-server recall, real-world vulnerability prevalence, or superiority over another scanner.

Authorized metadata positive v3

datasets/authorized_positive_metadata_v3 adds 16 literal, intentionally malicious metadata fixtures from Cisco's Apache-2.0 licensed MCP Scanner evaluation corpus. Cisco's first-party scenario labels cover prompt injection and unauthorized code execution; MCPSentinel maps them into 18 rule/case pairs. Every case pins a source path, function line, file digest, and full commit. The extractor only reads a local checkout.

mcpsentinel benchmark datasets/authorized_positive_metadata_v3/manifest.json \
  --judge heuristic --format json --output benchmark-v3.json

V3 is a calibration regression control, not a held-out accuracy study: its labels informed the narrow metadata rules added in 0.7.0. At that frozen configuration it reports all 18 labelled pairs while the 428-case v2 public negative control remains at zero candidates. This is useful evidence that the refinement did not create a false-positive in those exact public snapshots; it is not proof of real-world recall. One maintainer has reviewed the v3 mapping; see the independent-review protocol before citing it beyond regression coverage.

What MCPSentinel can—and cannot—tell you

MCPSentinel is useful as a preflight signal for three workflows: an individual developer deciding whether to inspect an MCP server more deeply, a maintainer self-auditing metadata before release, and a security team adding a non-blocking or reviewed CI gate.

It discovers advertised MCP metadata; it does not read a server's source code, prove authorization boundaries, or guarantee that runtime behavior matches an honest description. A clean report is not proof that a server is safe. Dynamic validation is intentionally narrower still: it can only invoke explicitly named, high-confidence tools from an image you own, with arguments you supply. Its process and filesystem counters are bounded review evidence, not full behavioral instrumentation. It is not a safe way to probe arbitrary public servers.

The default scanner is read-only. It never calls a discovered tool, follows HTTP redirects, or enables dynamic execution from the GitHub Action or MCP-native server. Use the result as evidence for review and combine it with source review, dependency review, permissions/egress controls, and normal incident response processes.

Development

pytest
ruff check .

The project is intentionally dependency-light: mcp handles protocol discovery, while the core rule engine, snapshot store, and report writers use the standard library.

Security

See SECURITY.md for vulnerability reporting and supported-version information.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

mcp_guardian_scan-0.7.0.tar.gz (90.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

mcp_guardian_scan-0.7.0-py3-none-any.whl (51.4 kB view details)

Uploaded Python 3

File details

Details for the file mcp_guardian_scan-0.7.0.tar.gz.

File metadata

  • Download URL: mcp_guardian_scan-0.7.0.tar.gz
  • Upload date:
  • Size: 90.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for mcp_guardian_scan-0.7.0.tar.gz
Algorithm Hash digest
SHA256 b15150701e6309914f86399779169afdf89bd24770af825b1dc136070ad80734
MD5 12ec6c8936144e3a5f16b263cc17c903
BLAKE2b-256 fadbbc15cffee5070e6168ec798d0524af08205c7a4bd1bd9f5a324db6e07275

See more details on using hashes here.

Provenance

The following attestation bundles were made for mcp_guardian_scan-0.7.0.tar.gz:

Publisher: publish.yml on gentaArnezzi/MCPSentinel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file mcp_guardian_scan-0.7.0-py3-none-any.whl.

File metadata

File hashes

Hashes for mcp_guardian_scan-0.7.0-py3-none-any.whl
Algorithm Hash digest
SHA256 f79d2e1d6ff1eecb0115e6d1494ff7cb5cb6e89213e123b6ee76c5998df57157
MD5 c53b370095e74c85b420bd9bf8283560
BLAKE2b-256 b3778cb9aa2702c564e7218f46834c64174370e382708967a28035aaeeea6011

See more details on using hashes here.

Provenance

The following attestation bundles were made for mcp_guardian_scan-0.7.0-py3-none-any.whl:

Publisher: publish.yml on gentaArnezzi/MCPSentinel

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.8.7

2 files

0.8.6

2 files

0.8.5

2 files

0.8.4

2 files

0.8.3

2 files

0.8.2

2 files

0.8.1

2 files

0.8.0

2 files

This release

0.7.0 This release

2 files

0.6.0

2 files

0.5.1

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.2

2 files

0.2.1

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page