Skip to main content

Detect and redact text + visual PII in document images and plain text.

Project description

Eka-PII-redaction

Detect and redact PII in document images — both text PII (names, addresses, IDs, dates, phone/email, …) and visual entities (signatures, stamps/seals, QR/barcodes, face photos, fingerprints, logos) — behind a single class.

  • One install, one Hugging Face repo. The package internally pulls the model weights it needs from one HF repo; you only ever reference one repo id.
  • Text + visual in one call, or text-only if you don't need the detector.
  • CPU or GPU — auto-selects CUDA if available, else CPU.
  • Choose what to redact — exclude any categories; everything is on by default.
  • Returns a structured entity list (text, bbox, category, …) and/or a redacted image.

Install

From PyPI:

pip install eka-pii-redaction          # core library
pip install "eka-pii-redaction[server]"  # + FastAPI service

From source:

git clone https://github.com/eka-care/Eka-PII-redactors.git
cd Eka-PII-redactors
pip install -e .            # core library
pip install -e ".[server]"  # + FastAPI service

System dependency: Tesseract OCR — required for the image modality (it OCRs the document before the text-in-image classifier runs on the words); not needed for the text-only modality.

# Debian/Ubuntu
sudo apt-get install -y tesseract-ocr
# macOS
brew install tesseract

(The provided Docker image installs everything for you — see below.)

Quickstart

from eka_pii_redaction import ImagePIIRedactor

# Loads both models from one HF repo; GPU if available, else CPU.
redactor = ImagePIIRedactor(
    "ekacare/pii-redactors",
    detect_visual=True,          # set False to skip visual entities (QR codes,
                                  # face photos, signatures, etc.) — text PII only
    # device="cpu",              # force CPU (default: auto)
    # exclude_entities=["logo", "brandname"],  # never redact these
)

# 1) Get the structured entity list
for e in redactor.detect("page.jpg"):
    print(e.kind, e.category, e.bbox, e.text, e.score)

# 2) Get a redacted image (PIL.Image)
redactor.redact("page.jpg").save("redacted.png")
redactor.redact("page.jpg", mode="blur").save("blurred.png")  # solid|blur|pixelate

detect() / redact() accept a file path, raw bytes, or a PIL.Image.

Text-only PII (plain strings)

from eka_pii_redaction import TextPIIRedactor

r = TextPIIRedactor("ekacare/pii-redactors")     # GPU if available, else CPU

# 1) Character-span entities
for s in r.detect("John Doe, DOB 1990-01-01, john@x.com"):
    print(s.category, s.l1, s.start, s.end, s.text, s.score)

# 2) Redacted string (mask supports {category} / {l1} placeholders)
r.redact("Call John at john@x.com", mask="[REDACTED]")    # -> "Call [REDACTED] at [REDACTED]"
r.redact("Call John at john@x.com", mask="[{category}]")  # -> "Call [primary_subject_name] at [email]"

TextPIIRedactor runs a lightweight multilingual token classifier on raw text — no OCR, no image — and returns TextPIISpan(category, start, end, l1, text, score) with character offsets.

API

ImagePIIRedactor

ImagePIIRedactor(hf_repo, *, detect_visual=True, device=None,
                  exclude_entities=None, visual_score_threshold=0.25,
                  ocr_lang=None, cache_dir=None)
arg meaning
hf_repo HF repo id or a local dir with text_model/ + visual_model/best.pt.
detect_visual If False, visual entities (QR codes, face photos, signatures, etc.) are not downloaded or loaded — text PII only.
device "cuda" / "cpu". None → auto (CUDA if available).
exclude_entities Categories to never detect/redact. Default: none excluded (all on).
visual_score_threshold Visual-entity confidence cutoff.
ocr_lang Tesseract language/script, e.g. "eng", "eng+Devanagari".

detect

detect(image, *, exclude_entities=None, ocr_lang=None) -> list[PIIEntity]

Each PIIEntity has:

field description
category fine category, e.g. primary_subject_name, signature
kind "text" or "visual"
bbox [x0, y0, x1, y1] in original-image pixels
l1 coarse group: person/location/contact/uid/… or biometric_visual
text OCR text (text entities) or None (visual)
score confidence in [0,1]

redact

redact(image, *, mode="solid", color=(0,0,0), exclude_entities=None,
       ocr_lang=None, pad=2) -> PIL.Image

Returns a redacted copy. modesolid | blur | pixelate.

list_entities

ImagePIIRedactor.list_entities() -> {"text": [...], "visual": [...]}

All selectable categories.

Categories

  • Text (47): person (name, age, gender, occupation, …), location (address, city, state, postcode, …), date_time, contact (phone, email, web_url, fax), uid (aadhaar, pan, passport, mrn/uhid, abha, insurance policy, bank/iban/upi, …), device_net, credential, and brandname.
  • Visual (6): signature, seal_stamp, qr_barcode, face_photo, fingerprint_thumb_impression, logo.

ImagePIIRedactor.list_entities() returns the exact set.

Run as a container

The image bundles torch+CUDA, Tesseract, the FastAPI service, and the React demo UI (web/, built at image-build time and served by the same process at /). Same image runs on GPU or CPU. This is also what runs on the ekacare/pii-redactor-demo HF Space.

docker build -t eka-pii-redaction .

# GPU
docker run --gpus all -p 7860:7860 \
  -e EKA_PII_HF_REPO=ekacare/pii-redactors \
  -v $HOME/.cache/huggingface:/root/.cache/huggingface \
  eka-pii-redaction

# CPU
docker run -e EKA_PII_DEVICE=cpu -p 7860:7860 eka-pii-redaction

Open http://localhost:7860 for the UI. API endpoints: GET /health, GET /entities, GET /entities-text, POST /detect / POST /redact (multipart file, optional exclude query param, mode/color form fields for /redact), and POST /detect-text / POST /redact-text (JSON {"text": ..., "exclude": [...]}). Env: EKA_PII_HF_REPO, EKA_PII_DETECT_VISUAL, EKA_PII_DEVICE, EKA_PII_EXCLUDE.

curl -F file=@page.jpg http://localhost:7860/detect
curl -F file=@page.jpg -F mode=blur http://localhost:7860/redact -o redacted.png

Structure (by modality)

The library is organized by modality, so additional models slot in cleanly:

eka_pii_redaction/
  taxonomy.py, entities.py          # shared
  image/                            # IMAGE modality (implemented)
    redactor.py  -> ImagePIIRedactor
    layoutlmv3.py -> text-PII-in-image detector
    yolo11m.py    -> visual-entity detector
  text/                             # TEXT modality (implemented)
    redactor.py  -> TextPIIRedactor  (PII inside plain-text strings, no image)
    minilm.py    -> text-PII token classifier (char-span detector)

The single model repo mirrors this:

<hf_repo>/
  image/ layoutlmv3/   yolo/best.pt
  text/  minilm/                    # multilingual text-PII model

Both modalities share the category taxonomy (eka_pii_redaction.taxonomy); the text model also detects mac_address (device_net).

Deploying the demo to HF Spaces

The Space (ekacare/pii-redactor-demo) runs this same repo's Docker image — see the "Run as a container" section above. To (re)deploy:

  1. Create the Space once, as private (matches the model's current visibility — flip both to public together later):
    huggingface-cli repo create pii-redactor-demo --organization ekacare \
      --type space --space_sdk docker --private
    
    (or via the HF UI: New Space → owner ekacare → SDK Docker → Private.)
  2. Add the model's read token as a Space secret: Space → Settings → Repository secrets → add HF_TOKEN. The server already reads it via huggingface_hub's standard auth — no code change needed.
  3. Deploy: ./scripts/push_space.sh. It pushes the current commit to the Space's git remote, with .space-metadata.yaml's front matter prepended to README.md for that push only — the Space needs that front matter to render its card (title/sdk/app_port/...), but GitHub and PyPI don't know to strip HF-specific front matter, so it's kept out of the tracked README.md and only injected at deploy time. Auth: set HF_TOKEN, or it falls back to your cached hf auth login token.
  4. Watch the build under the Space's "Logs" tab. The base image (pytorch/pytorch:...-cudnn9-runtime) is large, so the first build can take a while; subsequent pushes reuse Docker layer caching.
  5. Once it shows Running, open the Space URL and click through both tabs.

Publishing the model weights

The trained checkpoints are assembled into the single HF repo with:

python scripts/build_hf_repo.py \
  --layoutlmv3   .../checkpoints/base_v3_combined_4ep/final \
  --yolo-weights .../checkpoints/visual/visual_yolo11m/weights/best.pt \
  --out /tmp/eka-pii-hf --push --repo-id ekacare/pii-redactors

How it works (image modality)

  1. Text-in-image: Tesseract OCR (via the processor) → words + boxes → token classifier → per-word BIO labels → merged spans.
  2. Visual: a detector over the page → boxes + categories.
  3. Redact: fill / blur / pixelate every selected entity's box.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

eka_pii_redaction-0.1.3.tar.gz (24.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

eka_pii_redaction-0.1.3-py3-none-any.whl (22.1 kB view details)

Uploaded Python 3

File details

Details for the file eka_pii_redaction-0.1.3.tar.gz.

File metadata

  • Download URL: eka_pii_redaction-0.1.3.tar.gz
  • Upload date:
  • Size: 24.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/7.0.0 CPython/3.12.11

File hashes

Hashes for eka_pii_redaction-0.1.3.tar.gz
Algorithm Hash digest
SHA256 c34f8380cba1a33268ed7622d90d53fc74674209ec45df72013b329904a3e0b5
MD5 f9b37b75ec3c641ea9c140a37fe22b9b
BLAKE2b-256 cce8d7712cd741a87834b332cad94ed8f19eda919f30a5eeaac9abf4805ca224

See more details on using hashes here.

File details

Details for the file eka_pii_redaction-0.1.3-py3-none-any.whl.

File metadata

File hashes

Hashes for eka_pii_redaction-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 0dc60fe1d07e83a568fe742398242925a98c8711a0c34830eb7f05fba11bb3b4
MD5 2ab1ae6a60ba6d871d78e258c108fe8e
BLAKE2b-256 9c9586e873547be96f553e514d6f0d806ba0e95a8ff8aed54a672e9f8a62c697

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page