Strac MCP DLP — PII/PHI/PCI & secret redaction for AI agents (MCP)
An open-source Model Context Protocol server that gives any AI agent — Claude, Cursor, VS Code, your own — a way to find and strip sensitive data before it reaches the model.
Point your agent at it, and redact_text turns
Please onboard the user with SSN 123-45-6789 and email jane@acme.com
into
Please onboard the user with SSN [REDACTED] and email [REDACTED]
...along with a structured list of what was found. It works the same on images, PDFs and scanned documents.
This is a thin client over the Strac DLP API. Bring your own API key; all detection and redaction happens server-side at Strac.
Why
Agents are now wired into inboxes, ticketing systems, CRMs, databases and file stores. Every MCP tool call is a chance for an SSN, a card number, a patient record or an AWS key to be pulled into a prompt — and from there into a model provider's logs, a vector store, a Slack summary or a support ticket.
Filtering that data after the model has seen it is too late. This server puts the check in front of the model: detect first, redact, then let the agent reason over text that no longer carries the sensitive values.
This server is one door into Strac
Redacting a string you hand it is the smallest thing Strac does. The product is coverage: connect Strac to the SaaS and cloud apps where your sensitive data already lives, and it discovers, classifies, redacts and remediates it there — continuously, under your policies, with an audit trail — rather than waiting for someone to paste it into a prompt.
Slack, Google Workspace, Microsoft 365, Salesforce, Zendesk, Box, Dropbox, Jira, Confluence, GitHub, Notion, OneDrive, SharePoint, Snowflake, Databricks, BigQuery, Postgres, MongoDB, AWS, Azure, GCP, browsers and endpoints — 60+ integrations, agentless, no code to write.
That matters most when those systems are reached over MCP. When an agent — Claude Code, Claude Desktop, Cursor, GitHub Copilot, OpenAI Codex — pulls a Salesforce record or a Drive file through an MCP connector, Strac redacts the sensitive data inline, before the agent receives it. That is the difference between asking an agent to redact and enforcing it whether or not it asks.
strac.io/mcp-integrations is that product. This repo is its developer-facing sliver: the same detection engine, reachable from any MCP client, for when you want to sanitise a string or a file yourself.
Every MCP invocation your agents make — tools called, files read, the identity behind the prompt — captured and inspected.
60-second quickstart
1. Install
pip install strac-mcp-dlp
2. Request an API key
Request a key, or email hello@strac.io.
Keys are prefixed sk_live_ (production) or sk_test_ (sandbox); the server picks the matching endpoint automatically.
3. Add it to your MCP client
Claude Desktop — ~/Library/Application Support/Claude/claude_desktop_config.json on macOS, %APPDATA%\Claude\claude_desktop_config.json on Windows:
{
"mcpServers": {
"strac-dlp": {
"command": "strac-mcp-dlp",
"env": {
"STRAC_API_KEY": "sk_live_your_key_here"
}
}
}
}
Claude Code:
claude mcp add strac-dlp --env STRAC_API_KEY=sk_live_your_key_here -- strac-mcp-dlp
Cursor — .cursor/mcp.json in your project, or ~/.cursor/mcp.json globally: use the same block as Claude Desktop.
4. Restart your client and ask it to redact something
Use Strac to redact this before I paste it into the ticket: "Customer Jane Doe, SSN 123-45-6789, card 4111 1111 1111 1111."
No install at all, if you have uv:
{
"mcpServers": {
"strac-dlp": {
"command": "uvx",
"args": ["strac-mcp-dlp"],
"env": { "STRAC_API_KEY": "sk_live_your_key_here" }
}
}
}
See examples/ for ready-to-copy config files.
Tools
-
redact_text
- Redact PII, PHI, PCI and secrets out of a block of text, returning the sanitised string plus what was found
- Inputs:
text(string): the text to redactredact_field_mode(string, optional):REDACTED(default),BLANK,MASK_SEVEN_XorTOKEN_LINK_PLAINTEXTinclude_matched_text(boolean, optional): also return the raw sensitive values. Off by default
- Call it before putting untrusted or user-supplied text into a prompt, a log line, a ticket or any downstream system
redact_text(text="SSN 123-45-6789")→"SSN [REDACTED]"
-
detect_sensitive_data
- Report which sensitive data types are present in text, without changing it
- Inputs:
text(string): the text to scaninclude_matched_text(boolean, optional): also return the raw values. Off by default
- Use it to decide whether text is safe to send onward
detect_sensitive_data(text="call me at jane@acme.com")→EMAIL
-
detect_file
- Scan a local file — image, PDF, scan, config or source — using Strac OCR and classifiers
- Inputs:
path(string): path to the file to scaninclude_matched_text(boolean, optional): also return the raw values. Off by default
- The file is never modified
detect_file(path="./w2.pdf")→TAX_ID_NUMBER, NAME, ADDRESS
-
redact_file
- Write a redacted copy of a local file to disk
- Inputs:
path(string): path to the file to redactoutput_path(string, optional): where to write the copy. Defaults to a.redactedsuffix beside the original. Pointing it at the source is refusedoverwrite(boolean, optional): allow replacing an existing destination. Off by default
- The original is never modified
redact_file(path="./w2.pdf")→./w2.redacted.pdf
-
detokenize
- Resolve Strac vault tokens (
tkn_…) back to their original values, for authorised callers - Inputs:
token_ids(string[]): the token identifiers to resolve. Maximum 10 per call
- Requires an IP-allowlisted server-to-server key in live mode
detokenize(token_ids=["tkn_abc"])→"111-22-3333"
- Resolve Strac vault tokens (
Redaction styles
redact_text takes a redact_field_mode:
| Mode | Result |
|---|---|
REDACTED (default) |
[REDACTED] |
MASK_SEVEN_X |
XXXXXXX |
BLANK |
removed entirely |
TOKEN_LINK_PLAINTEXT |
<sensitive_data: https://…/detokenize/tkn_…> — a link to the value in the Strac vault, so an authorised human can still retrieve it |
Redaction that doesn't leak
By default these tools return the types and positions of what they found, not the values:
{
"redacted_text": "Please onboard the user with SSN [REDACTED]",
"detection_count": 1,
"data_element_types": ["TAX_ID_NUMBER"],
"detections": [{ "type": "TAX_ID_NUMBER", "begin_index": 33, "end_index": 44, "length": 11 }]
}
Returning the matched text alongside the redacted text would hand the model exactly the data you just removed. Pass include_matched_text=true when a caller genuinely needs the raw values.
What gets detected
Strac ships 191 built-in data elements across 10 categories, plus custom elements you define with regex or your own trained model:
| Category | Elements | Examples |
|---|---|---|
| Identification | 126 | SSN/TIN, passports, driver licences and national IDs across ~60 countries — Aadhaar, PAN, PESEL, BSN, Fiscal Code, IRD — plus NPI and DEA registration numbers |
| Secrets | 30 | AWS access and secret keys, GitHub and GitLab tokens, Slack tokens, GCP credentials, Azure storage and service-principal keys, private keys, JDBC and MongoDB connection strings, seed phrases |
| Financial Account | 10 | Card number and tail, CVV, expiry, bank account and routing numbers, IBAN, SWIFT |
| Advertisement Identifiers | 7 | Apple IDFA and IDFV, Google GAID, Roku, Amazon Fire OS, Huawei OAID |
| Contact | 6 | Name, address, email, phone, date of birth, age |
| Device Tracking | 5 | IP address, MAC address, IMEI, webpage URL, date/time |
| Asset | 3 | Source code, VIN, vehicle licence plate |
| Document Properties | 2 | Invoice, password-protected document |
| Content Moderation | 1 | Offensive content |
| Intellectual Property | 1 | Chemical/molecular structure |
Every element named individually: Strac Catalog of Sensitive Data Elements.
Detection runs on text and, via OCR, on PDFs, JPEGs, PNGs, DOCX, XLSX, screenshots and .msg email files — which is what detect_file and redact_file reach.
One caveat worth setting expectations on: the type values these MCP tools return depend on which endpoint answered, and the two use different vocabularies for the same element. A US Social Security Number comes back as TAX_ID_NUMBER from redact_text and as SOCIAL_SECURITY_NUMBER from detect_sensitive_data — which additionally reports SSN under reported_element_types. Both vocabularies are surfaced as returned rather than normalised, so nothing is invented on your behalf. Which elements are detected at all depends on what is enabled for your account; the full catalog above is what runs across your connected apps, where policies, remediation and audit live.
Configuration
| Variable | Required | Default |
|---|---|---|
STRAC_API_KEY |
yes | — |
STRAC_API_BASE |
no | https://api.live.tokenidvault.com for sk_live_ keys, https://api.test.tokenidvault.com for sk_test_ keys |
STRAC_API_TIMEOUT |
no | 60 seconds |
Without a key, every tool returns: Set STRAC_API_KEY — request one at https://www.strac.io/mcp-integrations.
The server speaks stdio by default. strac-mcp-dlp --transport streamable-http serves HTTP instead.
How it works
Your MCP client spawns this server locally; it forwards each tool call to the Strac API over HTTPS with your X-Api-Key, and returns the result. Nothing is classified or redacted on your machine, and this repository contains no detection models.
Your data element definitions, custom policies, remediation rules and audit trail live in your Strac account, not in this repo — which is why the same key that powers these five tools also governs every connected app.
MCP client ──stdio──▶ strac-mcp-dlp ──HTTPS──▶ Strac DLP API
(Claude, Cursor, (this repo, (classifiers, OCR,
your agent) ~600 lines) vault, policies, audit)
Limits inherited from the API: 4 MB for inline content, 10 MB for document uploads, 10 tokens per detokenize call.
Development
git clone https://github.com/strac-io/strac-mcp-dlp
cd strac-mcp-dlp
python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/python -m pytest
The test suite mocks the Strac API with respx, so it runs without a key. To try the server by hand:
STRAC_API_KEY=sk_test_… .venv/bin/python -m strac_mcp_dlp
Or point the MCP Inspector at it:
STRAC_API_KEY=sk_test_… npx @modelcontextprotocol/inspector strac-mcp-dlp
FAQ
Do I need a Strac account?
Yes. This is a thin client over the Strac DLP API and every call is authenticated — there is no anonymous mode. Request a key at strac.io/mcp-integrations or email hello@strac.io.
Does my data stay on my machine?
No. Text and files you pass to these tools are sent to the Strac API over HTTPS for classification. Nothing is classified locally — this repository contains no detection models, no classifier and no vault. detect_file sends files under 4 MB inline without storing them; larger files, and anything passed to redact_file, are uploaded to your Strac document vault.
Which MCP clients does it work with?
Any client that speaks stdio — Claude Desktop, Claude Code, Cursor, VS Code and others. --transport streamable-http is available if you need to run it as a service rather than a subprocess.
Does it detect secrets, or only PII?
Both. Alongside personal, health and payment data, Strac classifies AWS access and secret keys, GitHub and GitLab tokens, Slack tokens, GCP credentials, Azure storage and service-principal keys, private keys, and JDBC and MongoDB connection strings. See the full catalog.
Why doesn't redact_text return the values it found?
Because handing them back would undo the redaction. If the tool returned "123-45-6789" in its detections alongside the redacted string, the model would end up with exactly the data you just removed. You get types and positions by default; pass include_matched_text=true when a caller genuinely needs the raw values.
Will redact_file modify my original?
No. It writes a copy, refuses an output_path that resolves to the source file (including via symlink), and will not replace an existing file unless you pass overwrite=true.
What happens if the Strac API returns something unexpected?
The tool errors. It never reports a clean scan it could not verify — a malformed or unparseable response raises rather than returning "no sensitive data found", because a silent false negative is the one failure a DLP tool cannot have.
How is this different from the Strac platform?
This repo is on-demand tooling your agent chooses to call. The platform sits across your SaaS, cloud and database connectors and enforces policy whether or not the agent asks — redacting inline before the agent receives the data, with remediation and an audit trail. See strac.io/mcp-integrations.
Links
- Strac MCP integrations — the full MCP DLP gateway, connectors, policies and audit
- Strac catalog of sensitive data elements — every data element, named
- Strac API reference
- MCP DLP: protecting data across Model Context Protocol
- Strac integrations — the full SaaS and cloud connector list
Security
Never commit an API key. sk_live_ keys are server-side credentials; keep them in your MCP client's env block or your secret manager, not in source control. To report a vulnerability, email security@strac.io.
License
MIT — see LICENSE.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file strac_mcp_dlp-0.1.1.tar.gz.
File metadata
- Download URL: strac_mcp_dlp-0.1.1.tar.gz
- Upload date:
- Size: 24.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
87e494fbc9fc0f2b08f94e0086d6210b68667204ef9d34be3928599da3786e62
|
|
| MD5 |
f7cad2444e6d9c5e2b7d5c7b3f82cc0c
|
|
| BLAKE2b-256 |
2cab5336d0b421ee3c680f8efc91e34a129edc63e0870faf8b39e5d124015599
|
Provenance
The following attestation bundles were made for strac_mcp_dlp-0.1.1.tar.gz:
Publisher:
publish.yml on strac-io/strac-mcp-dlp
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
strac_mcp_dlp-0.1.1.tar.gz -
Subject digest:
87e494fbc9fc0f2b08f94e0086d6210b68667204ef9d34be3928599da3786e62 - Sigstore transparency entry: 2704729315
- Sigstore integration time:
-
Permalink:
strac-io/strac-mcp-dlp@75b97ca649c9ea0ac982b08234dbf8816a3a725f -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/strac-io
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@75b97ca649c9ea0ac982b08234dbf8816a3a725f -
Trigger Event:
release
-
Statement type:
File details
Details for the file strac_mcp_dlp-0.1.1-py3-none-any.whl.
File metadata
- Download URL: strac_mcp_dlp-0.1.1-py3-none-any.whl
- Upload date:
- Size: 20.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
2dfd4f40665d6f73842c23c65d3cee94ea7234bd01a107fbe171e729ec802a55
|
|
| MD5 |
969d7bf606fa14e3047853f1f32fccc2
|
|
| BLAKE2b-256 |
c71d9a1faf04da8f9b23603dfd8594c8be4d95e89b7b6a4ba89c851cd466d8bb
|
Provenance
The following attestation bundles were made for strac_mcp_dlp-0.1.1-py3-none-any.whl:
Publisher:
publish.yml on strac-io/strac-mcp-dlp
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
strac_mcp_dlp-0.1.1-py3-none-any.whl -
Subject digest:
2dfd4f40665d6f73842c23c65d3cee94ea7234bd01a107fbe171e729ec802a55 - Sigstore transparency entry: 2704729393
- Sigstore integration time:
-
Permalink:
strac-io/strac-mcp-dlp@75b97ca649c9ea0ac982b08234dbf8816a3a725f -
Branch / Tag:
refs/tags/v0.1.1 - Owner: https://github.com/strac-io
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
publish.yml@75b97ca649c9ea0ac982b08234dbf8816a3a725f -
Trigger Event:
release
-
Statement type: