Skip to main content
Pre-release

This release is a pre-release and may not be stable for production use.

DataFog Python

DataFog is a Python library for detecting and redacting personally identifiable information (PII).

It provides:

  • Fast structured PII detection via regex
  • An offline PII firewall for AI agents: a Claude Code hook and a LiteLLM gateway guardrail (introduced in 4.6)
  • Optional NER support via spaCy and GLiNER
  • A simple agent-oriented API for LLM applications
  • Backward-compatible DataFog and TextService classes

Agent & Gateway Firewall

DataFog includes two ready-made enforcement points that catch PII at the moment it would leave your machine — offline, in microseconds, with matched values never echoed into logs or transcripts:

  • Claude Code hook (datafog-hook): gates agent tool calls (shell commands, web requests, file writes, MCP tools) and warns the model when prompts or tool results carry PII. ~70–90ms per invocation including process startup. Easiest install is the Claude Code plugin:

    /plugin marketplace add DataFog/datafog-claude-plugin
    /plugin install datafog@datafog
    

    Manual hook setup and limitations: examples/claude_code_hook/.

  • LiteLLM guardrail (DataFogGuardrail): redacts or blocks PII in requests and responses at the gateway, for any LiteLLM-proxied provider. In-process (~40µs per message scanned; a request clears the guardrail in well under a millisecond), no sidecar service. Setup: examples/litellm_guardrail/.

Both default to the high-precision entity set (EMAIL, PHONE, CREDIT_CARD, SSN); noisier types are opt-in. Known-safe values can be exempted with an allowlist: scan(text, allowlist=[...]) for exact values, allowlist_patterns=[...] for full-match regexes (e.g. ^\d{10}$ to stop unix timestamps matching as phone numbers) — available in both adapters and the API. Presidio-style entity names (EMAIL_ADDRESS, PHONE_NUMBER, US_SSN) are accepted as aliases for easy migration.

Every performance number above is reproducible with one command — methodology, pinned payloads, and comparisons against Presidio and spaCy NER live in benchmarks/.

Installation

# Core install (regex engine)
pip install datafog

# Add spaCy support
pip install datafog[nlp]

# Add GLiNER + spaCy support
pip install datafog[nlp-advanced]

# Add local OCR support
pip install datafog[ocr]

# Add Spark/distributed support
pip install datafog[distributed]

# Everything
pip install datafog[all]

Python 3.13 support is certified for the core SDK, CLI, nlp, nlp-advanced, and ocr install profiles. Donut OCR still requires a model that is available locally before runtime use. distributed and all remain optional, heavier profiles and are not part of the lightweight core path.

Python 3.14 support is certified for the core SDK and CLI. Optional profiles remain dependent on upstream Python 3.14 package availability until they are covered by the corresponding CI install-profile checks.

Quick Start

import datafog

text = "Contact john@example.com or call (555) 123-4567"
clean = datafog.sanitize(text, engine="regex")
print(clean)
# Contact [EMAIL_1] or call [PHONE_1]

For LLM Applications

import datafog

# 1) Scan prompt text before sending to an LLM
prompt = "My SSN is 123-45-6789"
scan_result = datafog.scan_prompt(prompt, engine="regex")
if scan_result.entities:
    print(f"Detected {len(scan_result.entities)} PII entities")

# 2) Redact model output before returning it
output = "Email me at jane.doe@example.com"
safe_result = datafog.filter_output(output, engine="regex")
print(safe_result.redacted_text)
# Email me at [EMAIL_1]

# 3) One-liner redaction
print(datafog.sanitize("Card: 4111-1111-1111-1111", engine="regex"))
# Card: [CREDIT_CARD_1]

German Structured PII

German structured PII is country-specific and opt-in. Use explicit locale selection or entity-type filtering when you want German VAT IDs, German IBANs, tax IDs, postal codes, passports, or residence permits.

import datafog

text = "Steuer-ID 12345678901 liegt vor."

print(datafog.scan(text, engine="regex").entities)
# []

print(datafog.scan(text, engine="regex", locales=["de"]).entities)
# [Entity(type='DE_TAX_ID', text='12345678901', ...)]

Guardrails

import datafog

# Reusable guardrail object
guard = datafog.create_guardrail(engine="regex", on_detect="redact")

@guard
def call_llm() -> str:
    return "Send to admin@example.com"

print(call_llm())
# Send to [EMAIL_1]

Engines

Use the engine that matches your accuracy and dependency constraints:

  • regex:
    • Fastest and always available.
    • Best for default structured entities: EMAIL, PHONE, SSN, CREDIT_CARD, IP_ADDRESS, DATE, ZIP_CODE (DOB and ZIP are accepted as input aliases).
    • Use locales=["de"] for German structured IDs such as DE_VAT_ID, DE_IBAN, DE_TAX_ID, DE_POSTAL_CODE, and passport or residence permit numbers.
  • spacy:
    • Requires pip install datafog[nlp].
    • Useful for unstructured entities like person and organization names.
  • gliner:
    • Requires pip install datafog[nlp-advanced].
    • Stronger NER coverage than regex for unstructured text.
  • smart:
    • Cascades regex with optional NER engines.
    • If optional deps are missing, it degrades gracefully and warns.

Optional OCR And Spark Surfaces

The 4.x line keeps the main package story centered on lightweight text PII screening. OCR and Spark remain supported optional surfaces for users who already rely on them, but they are not required for the core import, default scan/redact helpers, or guardrail helpers.

  • OCR:
    • Install datafog[ocr] for local image OCR helpers.
    • URL-based image downloading also needs datafog[web,ocr].
    • Tesseract usage requires the system tesseract binary.
    • Python 3.13 is validated for the OCR install profile, Pillow, pytesseract, and system Tesseract smoke checks.
    • Donut OCR requires datafog[nlp-advanced,ocr] and a model already available locally.
  • Spark:
    • Install datafog[distributed] for SparkService.
    • Spark PII UDF helpers also require datafog[nlp] and an installed spaCy model.
    • A Java runtime is required by PySpark.

OCR and Spark are not deprecated. Their broader API and packaging overhaul is deferred; the 4.x goal is to keep them explicit, documented, and isolated from the lightweight core path.

Backward-Compatible APIs

The existing public API remains available.

DataFog class

from datafog import DataFog

result = DataFog().scan_text("Email john@example.com")
print(result["EMAIL"])

TextService class

from datafog.services import TextService

service = TextService(engine="regex")
result = service.annotate_text_sync("Call (555) 123-4567")
print(result["PHONE"])

CLI

# Scan text
datafog scan-text "john@example.com"

# Redact text
datafog redact-text "john@example.com"

# Replace text with pseudonyms
datafog replace-text "john@example.com"

# Hash detected entities
datafog hash-text "john@example.com"

# Enable German regex identifiers
datafog redact-text "Steuer-ID 12345678901" --locale de

Telemetry

DataFog telemetry is disabled by default.

To opt in:

export DATAFOG_TELEMETRY=1

To force telemetry off:

export DATAFOG_NO_TELEMETRY=1
# or
export DO_NOT_TRACK=1

Telemetry does not include input text or detected PII values.

Development

git clone https://github.com/datafog/datafog-python
cd datafog-python
python -m venv .venv
source .venv/bin/activate  # Windows: .venv\Scripts\activate
pip install -e ".[all,dev]"
pip install -r requirements-dev.txt
pytest tests/

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datafog-4.8.0b4.tar.gz (102.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datafog-4.8.0b4-py3-none-any.whl (76.4 kB view details)

Uploaded Python 3

File details

Details for the file datafog-4.8.0b4.tar.gz.

File metadata

  • Download URL: datafog-4.8.0b4.tar.gz
  • Upload date:
  • Size: 102.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.15

File hashes

Hashes for datafog-4.8.0b4.tar.gz
Algorithm Hash digest
SHA256 5cc50271d41aad74badeac3606a756b38d89b3b9c61ea8ef4ccff9f939172c2c
MD5 7a6e588e034f81df925d102e803d4877
BLAKE2b-256 addac9210192fe1f692dcf2db52abb68f8aec7e272ef3e6a6f125f90441e76ab

See more details on using hashes here.

File details

Details for the file datafog-4.8.0b4-py3-none-any.whl.

File metadata

  • Download URL: datafog-4.8.0b4-py3-none-any.whl
  • Upload date:
  • Size: 76.4 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.11.15

File hashes

Hashes for datafog-4.8.0b4-py3-none-any.whl
Algorithm Hash digest
SHA256 fd5d76895aa16a0fe527403a7060349c6db975a098f1e8d620cef0812a9c7ed5
MD5 ca0e767ebed0a2ffb8a7073e08b1f107
BLAKE2b-256 c9e6bb7b60c05bb314d8955c1b3a64525572c159d791ef49124c0c7988625ce5

See more details on using hashes here.

Release history Release notifications | RSS feed

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page