Skip to main content

pyveil: PII and secret redaction for Python AI agents

Stop sensitive data before it reaches an LLM, tool, MCP resource, memory store, log, or trace.

PyPI version Tests Python 3.8 to 3.14 Zero core dependencies Typed package Synthetic evaluation: 39 cases passing MIT license

Documentation · Guides · Evaluation · PyPI · Cookbook · Detection reference · Support · Security

pyveil is local, dependency-free redaction middleware for LLM applications and AI agents. It replaces high-confidence PII and credentials with deterministic, scoped HMAC placeholders before data crosses an application boundary.

Raw application context Context sent to the model
Email alice@example.com Email [EMAIL:a13f7c91b0d2]
api_key: sk-proj-... api_key: [API_KEY:38ded98a17e7]
Authorization: Bearer ... [AUTH_HEADER:4fe2926b7d20]

No network calls. No reversible vault. No raw values in findings by default.

Try It

pip install pyveil
pyveil demo
# or: python -m pyveil demo

Or run the synthetic demo in an isolated environment:

uvx pyveil demo
before: Email alice@example.com, call 010-1234-5678, and use API key sk-proj-...
after:  Email [EMAIL:...], call [PHONE:...], and use API key [API_KEY:...]
found:  API_KEY, EMAIL, PHONE

Protect An LLM Call

Put pyveil immediately before the provider call. The same code works with OpenAI, Azure OpenAI, Anthropic, Gemini, LiteLLM, or an internal gateway.

from pyveil import Channel, Veil

veil = Veil.high(
    secret=b"tenant-or-run-secret",
    scope="tenant/session",
)

messages = [
    {"role": "user", "content": "Email alice@example.com about my account."},
]

safe = veil.redact_data(messages, channel=Channel.PROMPT_INPUT)
response = call_llm(safe.data)  # Your provider SDK call

The provider receives the same list and dictionary shape, with sensitive values replaced before serialization or transmission.

pyveil redacts synthetic PII and secrets before an AI agent boundary

Why pyveil

Need What pyveil provides
Keep data local Standard-library core, zero required dependencies, zero network calls
Preserve references Stable [TYPE:12hex] placeholders from HMAC-SHA256
Isolate tenants and runs Caller-defined scope changes placeholders across boundaries
Redact real agent payloads Recursive dictionaries, lists, tuples, and JSON strings
Cover more than prompts Policy channels for tools, MCP, memory, logs, traces, input, and output
Stop credentials in tool calls Auth headers, private keys, API keys, JWTs, and tokens block by default
Cover app-specific data Exact known-value and trusted custom-regex rules
Audit without leaking Findings contain type, rule, path, placeholder, and fingerprint, not raw values
Verify supported behavior Public 39-case synthetic regression corpus, evaluator, and CI gate

Reproducible Evidence

The repository ships a public synthetic detector corpus and a standard-library evaluator:

python evaluation/evaluate.py --check

For corpus v1, pyveil 0.2.1 matches all 36 expected findings across 39 cases (33 positive, 6 negative), with no corpus false positives, false negatives, labeled-value leaks, or non-empty Finding.raw values.

These numbers describe documented supported shapes only. They are not a real-world PII recall benchmark and do not cover unknown names, addresses, languages, documents, or images. Read the methodology and limits.

Known Names And Domain IDs

Regex cannot discover arbitrary names or addresses. When your application already knows a value is sensitive, teach that value to pyveil without adding an NER model:

from pyveil import CustomRule, Veil

rules = [
    CustomRule.exact("PERSON", ["Alice Kim", "Hong Gildong"]),
    CustomRule("CUSTOMER_ID", r"\bCUS-[A-Z0-9]{8}\b", rule_id="customer_id"),
]

veil = Veil.high(
    secret=b"tenant-secret",
    scope="tenant/session",
    rules=rules,
)

result = veil.redact_text("Alice Kim owns CUS-A1B2C3D4.")
print(result.text)
# [PERSON:...] owns [CUSTOMER_ID:...].

Custom patterns are trusted application code. Keep them narrow and test them against realistic positive and negative samples.

Agent Boundaries

Classic masking often stops at text input. Agents move data across several surfaces, so channels are first-class policy inputs:

Channel Redact before
prompt.input User, RAG, or application context reaches a model
prompt.output Model output is displayed or chained
tool.call.arguments A model-controlled tool executes
tool.call.result Tool output returns to a model
mcp.resource.content MCP resource content enters context
memory.write Text is embedded or persisted
trace.span.attributes Attributes leave through telemetry
log.record Records reach handlers or external sinks
user / retrieval / tool / resource data
                    |
                  pyveil
                    |
model / tool / MCP / memory / trace / log boundary

Detection Board

Core detection is intentionally conservative and high precision:

Type Examples HIGH output
EMAIL Email addresses [EMAIL:12hex]
PHONE Korean, separated international, and compact E.164 phone shapes [PHONE:12hex]
CREDIT_CARD Card numbers that pass Luhn validation [CREDIT_CARD:12hex]
JWT Compact JSON Web Tokens [JWT:12hex]
AUTH_HEADER Bearer and Basic authorization headers [AUTH_HEADER:12hex]
PRIVATE_KEY PEM private-key blocks [PRIVATE_KEY:12hex]
API_KEY OpenAI, GitHub, Slack, Google, and AWS-style keys [API_KEY:12hex]
URL_QUERY_SECRET Token, key, secret, code, and auth query values [URL_QUERY_SECRET:12hex]
KV_SECRET Password, cookie, secret, and token key-value pairs [KV_SECRET:12hex]
Custom Known values and application regex rules [YOUR_TYPE:12hex]

See the full redaction reference for examples, validation rules, LOW masking output, and limitations.

HIGH And LOW

Use HIGH at model, agent, tool, MCP, memory, trace, and external log boundaries. It produces stable HMAC placeholders such as [EMAIL:a13f7c91b0d2].

Use LOW only for human-facing previews where preserving shape is useful:

alice@example.com  -> al***@e******.com
010-1234-5678      -> 010-****-5678
4242 4242 4242 4242 -> **** **** **** 4242

Credential-like values remain aggressively hidden in both levels.

Policy

The default high policy redacts supported findings and blocks credentials in tool.call.arguments:

from pyveil import Action, Channel, Entity, Policy, Veil

policy = Policy.default_high().override(
    Channel.PROMPT_INPUT,
    Entity.EMAIL,
    Action.PASS,
)

veil = Veil.high(secret=b"tenant-secret", policy=policy)

When both policy and level are supplied, the explicit policy decides channel levels and actions. Build one Veil per tenant, session, or run and reuse it.

CLI

Use stdin for shell pipelines, files for preflight checks, and JSON output for structured automation:

export PYVEIL_SECRET="tenant-or-run-secret"
export PYVEIL_SCOPE="tenant/session"

printf 'Email alice@example.com' | pyveil redact -
pyveil redact request.json --channel tool.call.result --format json
pyveil scan prompt.txt --format json
pyveil init
pyveil test-config pyveil.yaml

scan emits finding metadata without raw sensitive values. JSON-shaped input is parsed and traversed structurally.

Integration Recipes

Stack or boundary Copy-paste example
Any LLM provider Provider-neutral client wrapper
OpenAI Agents SDK Input guardrail example
Azure OpenAI Use the provider-neutral wrapper; only call_llm changes
LiteLLM Proxy filter example
FastAPI Request middleware example
MCP Server wrapper and integration guide
Python logging Logging filter example
Agent memory Before-embedding example

The cookbook covers prompts, tools, MCP, memory, logging, tracing, JSON, and CLI workflows.

Pick The Right Tool

Choose pyveil when you want a small local boundary filter for structured PII, credentials, known values, and domain identifiers across agent context flows.

Choose Presidio, GLiNER, or another NER-backed system when you need broad semantic discovery of unknown people, organizations, locations, or addresses. Choose an enterprise DLP product when you need managed policy, document/image coverage, incident workflows, or compliance reporting.

These tools can be layered. pyveil does not claim perfect recall. See the full pyveil vs Presidio, NER, guardrails, and DLP decision guide.

Safety Contract

  • Raw sensitive values are not stored in Finding objects by default.
  • Placeholders use HMAC-SHA256 with a caller-provided secret and scope.
  • Credential-like values can be blocked before model-controlled tools execute.
  • max_input_chars bounds work performed on text and structured payloads.
  • The core makes no network calls and has no required third-party dependency.
  • pyveil has no reversible vault or unmasking API.

pyveil is not a compliance guarantee, enterprise DLP system, secret-scanning replacement, or prompt-injection firewall. Read the threat model, known limitations, and security policy before production use.

Guides

Development

uv run --extra dev ruff check .
uv run --extra dev mypy pyveil tests
uv run --extra test pytest
uv run --extra test python evaluation/evaluate.py --check

CI runs the test suite on Python 3.8 through 3.14. The core remains typed, dependency-free, and MIT licensed.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pyveil-0.2.1.tar.gz (935.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

pyveil-0.2.1-py3-none-any.whl (26.4 kB view details)

Uploaded Python 3

File details

Details for the file pyveil-0.2.1.tar.gz.

File metadata

  • Download URL: pyveil-0.2.1.tar.gz
  • Upload date:
  • Size: 935.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pyveil-0.2.1.tar.gz
Algorithm Hash digest
SHA256 35ee73781ec7c31edf85928aa378443394e1e795c1ac573b1ba7e441aa9bbc0e
MD5 6f0c2728a5d93840c700c9fa0d392925
BLAKE2b-256 18c367101c169911c98107429905d57c04858657c02cb28062ee8c0e633a2a6f

See more details on using hashes here.

Provenance

The following attestation bundles were made for pyveil-0.2.1.tar.gz:

Publisher: release.yml on hyeonsangjeon/pyveil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file pyveil-0.2.1-py3-none-any.whl.

File metadata

  • Download URL: pyveil-0.2.1-py3-none-any.whl
  • Upload date:
  • Size: 26.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for pyveil-0.2.1-py3-none-any.whl
Algorithm Hash digest
SHA256 3907b124ad39a3e8fccbc0ee3db968653a0b3d812aec4231f2097e17c93ba508
MD5 ca6d5877bdfa8fe91bff457fa6474475
BLAKE2b-256 a697ebf2d20d81ea904e9af109770f83c9263a6df37f80193bcaa5369c8137b3

See more details on using hashes here.

Provenance

The following attestation bundles were made for pyveil-0.2.1-py3-none-any.whl:

Publisher: release.yml on hyeonsangjeon/pyveil

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.2.5

2 files

0.2.4

2 files

0.2.3

2 files

0.2.2

2 files

This release

0.2.1 This release

2 files

0.2.0

2 files

0.1.2

2 files

0.1.1

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page