Skip to main content

Pseudonymize

PyPI PyPI downloads Python CI License Typed

Typed, dependency-free, local-first PII pseudonymization for Python applications and LLM payloads.

Email paolo@example.com from 192.0.2.10.
                ↓
Email <EMAIL_1> from <IP_ADDRESS_1>.

Pseudonymize detects structured sensitive values locally and transforms them into numbered, generic, deterministic, or redacted tokens. The base package has no runtime dependencies, performs no telemetry or model downloads, and denies remote-capable backends by default.

Why Pseudonymize

  • Local by default. The standard-library core makes no network calls.
  • Useful identity semantics. Repeated normalized values share an alias inside an explicit processing scope.
  • Safe observability. Detailed reports expose types, offsets, provenance, and counts without copying matched values.
  • Small installation. The wheel is typed and has zero base runtime dependencies. Document, OCR, ML, HTML, and remote HTTP support remain optional.
  • Explicit extension points. Detection and format handling are separate, so custom backends and adapters do not replace the core policy and transformation logic.
  • Designed for LLM boundaries. Nested payload processing preserves structure and lets policies include or exclude paths such as messages.*.content.

What works today

Capability Status Notes
Strings Shipped Pseudonymization, redaction, batch processing, and safe reports
Nested Python data Shipped Dictionaries, lists, tuples, JSON scalars, and path policies
Structured detection Shipped Email, phone, IP, IBAN, payment card, Italian fiscal and VAT identifiers, URL credentials, and common secrets
Document representation Shipped Immutable blocks, typed locations, sanitized metadata, and inspection
Generic file orchestration Shipped Built-in or caller-provided adapters and atomic safe-copy output
TXT, Markdown, log, JSON, JSONL, and CSV files Shipped Explicit format or recognized suffix; no content guessing
Names, organizations, and locations (PII) Shipped Requires the optional [ml] extra and a caller-provided local ONNX model

Quality Benchmarks

The engine is strictly gated on detection accuracy against the ai4privacy/pii-masking-openpii-1.5m dataset (validation split, 1000 randomly sampled rows).

Current Baseline (1.20.0, 1000 rows, strict boundaries and type matching, ONNX backend enabled):

  • Precision: 0.8587
  • Recall: 0.8016
  • F1 Score: 0.8292

Each detection is paired with at most one annotation and must agree with it on entity type. Figures published before the scoring was strictly corrected are not comparable; see docs/benchmarks.md.

Installation

python -m pip install pseudonymize

Python 3.11 through 3.14 is supported.

Optional Local ML for PII

The strict and exclusive goal of this package is PII pseudonymization. The optional ML backend is provided solely to identify unstructured personal data (names, organizations, locations) for pseudonymization, and must never be treated as a general-purpose NLP tool. To detect and pseudonymize this data, you can integrate a local ONNX model by installing the ml extra:

python -m pip install pseudonymize[ml]

This installs onnxruntime, tokenizers, and numpy allowing you to configure the LocalONNXPIIBackend to run a lightweight local quantized DistilBERT model. No models are downloaded implicitly; you must provide your own paths to your downloaded model.onnx, tokenizer.json, and config.json.

Optional Document & OCR Support

To extract and safely redact complex file formats, install the relevant extras:

# For PDF support (via PyMuPDF)
python -m pip install pseudonymize[pdf]

# For Microsoft Office support (.docx, .xlsx, .pptx)
python -m pip install pseudonymize[office]

# For optical character recognition (scanned PDFs and images)
python -m pip install pseudonymize[ocr]

Quickstart

Transform text

from pseudonymize import pseudonymize, redact

safe = pseudonymize("Email paolo@example.com")
hidden = redact("Email paolo@example.com")

assert safe == "Email <EMAIL_1>"
assert hidden == "Email [REDACTED]"

Numbered aliases are the default. Numbering starts from one for each convenience call.

Get a safe report

from pseudonymize import Pseudonymizer

result = Pseudonymizer().process_with_report("Email paolo@example.com from 192.0.2.10.")

assert result.output == "Email <EMAIL_1> from <IP_ADDRESS_1>."
assert result.statistics.detections_found == 2
assert result.detections[0].backend == "rules"
assert "paolo@example.com" not in repr(result)

Reports include entity type, block identifier, typed source location, relative offsets, confidence, detector, backend provenance, and an optional replacement token. They never include the matched value.

Process an LLM payload

from pseudonymize import Policy, Pseudonymizer

payload = {
    "model": "example-model",
    "messages": [
        {"role": "user", "content": "Email paolo@example.com"},
        {"role": "user", "content": "Use paolo@example.com again"},
    ],
    "temperature": 0.2,
}

result = Pseudonymizer(policy=Policy.llm()).process_data_with_report(payload)

assert result.output["messages"][0]["content"] == "Email <EMAIL_1>"
assert result.output["messages"][1]["content"] == "Use <EMAIL_1> again"
assert result.output["model"] == "example-model"

The input is not mutated. Dictionary keys and non-string values are preserved.

Stream LLM responses

Real-time WebSocket chunks from OpenAI or Anthropic can be processed seamlessly without risking split-entity leakage across chunks. Both process_stream and process_stream_async are available:

import asyncio
from pseudonymize import Pseudonymizer


async def handle_stream(socket):
    engine = Pseudonymizer()

    # Process the async stream chunk-by-chunk. Overlapping contexts are automatically
    # managed behind the scenes so that chunks splitting "john.doe" and "@example.com"
    # are safely recombined and redacted before yielding to the user.
    async for safe_chunk in engine.process_stream_async(socket):
        print(safe_chunk, end="", flush=True)

Transformation modes

Mode Example Identity behavior
numbered <EMAIL_1> Stable inside one explicit scope
generic <EMAIL> Does not distinguish values of the same type
deterministic <EMAIL_K8M42PX7D3Q> Stable for the same key, namespace, type, and normalized value
redacted [REDACTED] Removes type and identity distinction
from pseudonymize import Pseudonymizer

scope = Pseudonymizer().new_scope()

assert scope.process("paolo@example.com").text == "<EMAIL_1>"
assert scope.process("maria@example.com and paolo@example.com").text == ("<EMAIL_2> and <EMAIL_1>")

Deterministic mode uses HMAC-SHA256 and requires a key of at least 32 bytes:

engine = Pseudonymizer(
    mode="deterministic",
    key=b"a-32-byte-or-longer-secret-key...",
    namespace="customer-42",
)

Different tenants should use different keys or namespaces. The package never generates, stores, or transmits a key silently.

Documents and files

Document contains immutable ContentBlock values. Each block has a stable identifier, text, a typed source location, and immutable JSON-scalar metadata. Detection offsets remain relative to the block text.

process_document() returns a transformed document. inspect_document() returns detections without transformed output.

process_file() selects a built-in adapter from an explicit format or a recognized suffix. It never overwrites the source, defaults to <stem>.safe<suffix>, refuses an existing destination unless overwrite=True, and publishes rendered bytes atomically:

from pseudonymize import Pseudonymizer

result = Pseudonymizer().process_file("requests.json")

assert result.output.name == "requests.safe.json"
assert result.statistics.replacements_applied >= 0

TXT, Markdown, log, JSON, JSONL, and strict comma-separated CSV files are dependency-free. inspect_file() reports detections without writing output. Pass format="json" to override an unknown suffix or use input_adapter and output_adapter for a custom format.

JSON, JSONL, and CSV outputs preserve data semantics but normalize insignificant whitespace, quoting, and record endings. UTF-8 is strict by default, an existing UTF-8 BOM is preserved, and an explicit codec can be supplied with encoding=.

The CLI exposes the same workflow:

pseudonymize file requests.json
pseudonymize inspect-file requests.json

Detection backends

The dependency-free RulesBackend handles structured values. A custom backend receives one ContentBlock and the active Policy, then returns relative Detection offsets. Backends declare supported entity types, provenance, remote capability, and remote-processing consent.

CompositeBackend merges leaf results through the same deterministic overlap resolver used by the core. Malformed and out-of-range detections fail with sanitized exceptions.

No capitalization heuristic is used for names. PERSON, ORGANIZATION, and LOCATION are public entity types; optional local ML can detect them.

Network policy

NetworkPolicy.DENY is the default. A remote-capable backend is called only when both conditions are true:

  1. The active policy is ALLOW_CONFIGURED with the backend allowlisted, or ALLOW_ALL.
  2. The backend explicitly sets allow_remote_processing=True.

An API key alone never enables network access. The optional remote extra includes HTTPRemoteBackend; when explicitly enabled, it sends each configured content block to the caller-selected HTTPS endpoint without following redirects. Callers are responsible for bounding content block sizes and configuring transport timeouts appropriate for the target provider.

Reversible mappings

Mappings are opt-in and available only in numbered and deterministic modes:

result = Pseudonymizer().process(
    "paolo@example.com",
    include_mapping=True,
)

assert result.restore("Reply to <EMAIL_1>.") == "Reply to paolo@example.com."

Mappings contain sensitive source values. They are hidden from repr, never persisted by the package, and must be protected separately by the application.

Compatibility

The dependency-free core API is stable throughout the 0.1 release line:

  • documented public names, constructors, methods, CLI behavior, and token formats are preserved;
  • patch releases preserve documented behavior and token formats;
  • additive capabilities may appear in minor releases;
  • a breaking core redesign requires an incompatible release;
  • security fixes may reject unsafe inputs and are called out in release notes.

See the compatibility policy before upgrading the package in an application.

Development

git clone https://github.com/ma2za/pseudonymize.git
cd pseudonymize
uv sync --all-extras --all-groups
uv run ruff format --check .
uv run ruff check .
uv run mypy
uv run pytest
uv run mkdocs build --strict

Tests must use synthetic values only. New capabilities require positive, negative, boundary, Unicode, adversarial, and cross-feature coverage where relevant.

Read CONTRIBUTING.md, SUPPORT.md, and SECURITY.md before opening an issue or pull request.

Project status

Release files for pseudonymize 1.26.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for pseudonymize 1.26.0
File Size Uploaded
pseudonymize-1.26.0.tar.gz 437.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for pseudonymize 1.26.0
File Interpreter ABI Platform
pseudonymize-1.26.0-py3-none-any.whl Python 3 none any Details

Total release size: 529.9 kB

Release files / pseudonymize-1.26.0.tar.gz

Download URL pseudonymize-1.26.0.tar.gz
Size 437.4 kB
Tags Source
SHA-256 checksum
How to use checksums
d6eae279001f345d556084863aa71a5e0e45244e0fd0b37f446e48cdc64260d2
BLAKE2b-256 checksum
How to use checksums
3e063b933e52cdd5fe0f32b6b6e6034a58bd5125af933a28f37f13669a6f9fd4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release files / pseudonymize-1.26.0-py3-none-any.whl

Download URL pseudonymize-1.26.0-py3-none-any.whl
Size 92.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
7f4f8759d6b17c414efea9ad2de221ca78c9acc830db4f663f3445afd11f5ccb
BLAKE2b-256 checksum
How to use checksums
ee35714e44e512453feb3fb54441d55a6547ca0487378509a9965eda23a86e7f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log
Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page