Skip to main content

Eka-PII-redaction

Detect PII in document images and plain text — both text PII (names, addresses, IDs, dates, phone/email, …) and visual entities (signatures, stamps/seals, QR/barcodes, face photos, fingerprints, logos) — then redact, de-identify, or anonymize it, behind a single class per modality.

  • Three modes, one detector. Redact destroys the value (mask/black-out — nothing kept). De-identify replaces each entity with a consistent pseudonym (Person_1) and returns the entity→pseudonym mapping so an authorized key-holder can re-link later. Anonymize is one-way: ages become 10-year buckets, dates keep only the year, fine geography collapses, and names/IDs become unnumbered tokens — no mapping exists anywhere.
  • One install, one Hugging Face repo. The package internally pulls the model weights it needs from one HF repo; you only ever reference one repo id.
  • Text + visual in one call, or text-only if you don't need the detector.
  • CPU or GPU — auto-selects CUDA if available, else CPU.
  • Choose what to process — exclude any categories; everything is on by default.
  • Returns a structured entity list (text, bbox, category, …) and/or a transformed image/string.

Note: anonymization here is best-effort removal/generalization of detected identifiers. It is not a k-anonymity guarantee and not, by itself, a compliance determination.

Install

From PyPI:

pip install eka-pii-redaction          # core library
pip install "eka-pii-redaction[server]"  # + FastAPI service

From source:

git clone https://github.com/eka-care/Eka-PII-redactors.git
cd Eka-PII-redactors
pip install -e .            # core library
pip install -e ".[server]"  # + FastAPI service

System dependency: Tesseract OCR — required for the image modality (it OCRs the document before the text-in-image classifier runs on the words); not needed for the text-only modality.

# Debian/Ubuntu
sudo apt-get install -y tesseract-ocr
# macOS
brew install tesseract

(The provided Docker image installs everything for you — see below.)

Quickstart

detect() is the core primitive — it finds every PII entity with its location, category, and confidence. The transforms (redact / anonymize / de-identify) are consumers of its output: run detection once, feed the same result into whichever transform you need.

Detect — images

from eka_pii_redaction import ImagePIIRedactor

# Loads the models from one HF repo; GPU if available, else CPU.
redactor = ImagePIIRedactor(
    "ekacare/pii-redactors",
    detect_visual=True,          # set False to skip visual entities (QR codes,
                                  # face photos, signatures, etc.) — text PII only
    # device="cpu",              # force CPU (default: auto)
    # exclude_entities=["logo", "brandname"],  # never detect these
)

entities = redactor.detect("page.jpg")
for e in entities:
    print(e.kind, e.category, e.bbox, e.text, e.score)

Detect — plain text

from eka_pii_redaction import TextPIIRedactor

r = TextPIIRedactor("ekacare/pii-redactors")     # GPU if available, else CPU

spans = r.detect("John Doe, DOB 1990-01-01, john@x.com")
for s in spans:
    print(s.category, s.l1, s.start, s.end, s.text, s.score)

TextPIIRedactor runs a lightweight multilingual token classifier on raw text — no OCR, no image — and returns TextPIISpan(category, start, end, l1, text, score) with character offsets. detect() (image) accepts a file path, raw bytes, or a PIL.Image.

Feed detections into transforms

Every transform accepts a prior detect() result (entities= for images, spans= for text) — detection runs once, transforms reuse it. Called without it, each transform just detects internally.

# --- Image: one detection, three outputs ---
entities = redactor.detect("page.jpg")

redactor.redact("page.jpg", mode="blur", entities=entities).save("redacted.png")
redactor.anonymize("page.jpg", entities=entities).save("anonymized.png")
deid = redactor.deidentify("page.jpg", entities=entities)   # ImageDeidResult
deid.image.save("deidentified.png")      # pseudonyms rendered in place;
                                          # faces/signatures become placeholders
deid.mapping.to_dict()                   # store securely to re-link later

# --- Text: same pattern ---
text = "Mr. John Doe, 45 yrs, DOB 12-03-1979, Indiranagar, Karnataka."
spans = r.detect(text)

r.redact(text, mask="[{category}]", spans=spans)
r.anonymize(text, spans=spans)
# -> "[PERSON], 40–49 yrs, DOB 1979, [LOCATION], Karnataka."  (no mapping exists)
result = r.deidentify(text, spans=spans)
result.text      # -> "Person_1, Age_1 yrs, DOB Date_1, City_1, State_1."
result.mapping   # entity -> pseudonym map; pass mapping=result.mapping on the
                 # next page of the same record to keep numbering consistent

(Example outputs are illustrative — exact spans depend on the model's tagging of the input.) De-identification consistency is per exact surface form ("John Doe" and a later bare "John" get different pseudonyms).

API

ImagePIIRedactor

ImagePIIRedactor(hf_repo, *, detect_visual=True, device=None,
                  exclude_entities=None, visual_score_threshold=0.25,
                  ocr_lang=None, cache_dir=None)
arg meaning
hf_repo HF repo id or a local dir with text_model/ + visual_model/best.pt.
detect_visual If False, visual entities (QR codes, face photos, signatures, etc.) are not downloaded or loaded — text PII only.
device "cuda" / "cpu". None → auto (CUDA if available).
exclude_entities Categories to never detect/redact. Default: none excluded (all on).
visual_score_threshold Visual-entity confidence cutoff.
ocr_lang Tesseract language/script, e.g. "eng", "eng+Devanagari".

detect

detect(image, *, exclude_entities=None, ocr_lang=None) -> list[PIIEntity]

Each PIIEntity has:

field description
category fine category, e.g. primary_subject_name, signature
kind "text" or "visual"
bbox [x0, y0, x1, y1] in original-image pixels
l1 coarse group: person/location/contact/uid/… or biometric_visual
text OCR text (text entities) or None (visual)
score confidence in [0,1]

redact

redact(image, *, mode="solid", color=(0,0,0), exclude_entities=None,
       ocr_lang=None, pad=2, entities=None) -> PIL.Image

Returns a redacted copy. mode ∈ solid | blur | pixelate. All transforms (redact / anonymize / deidentify, both modalities) accept a prior detect() result — entities= for images, spans= for text — so detection runs once.

list_entities

ImagePIIRedactor.list_entities() -> {"text": [...], "visual": [...]}

All selectable categories.

Categories

  • Text (47): person (name, age, gender, occupation, …), location (address, city, state, postcode, …), date_time, contact (phone, email, web_url, fax), uid (aadhaar, pan, passport, mrn/uhid, abha, insurance policy, bank/iban/upi, …), device_net, credential, and brandname.
  • Visual (6): signature, seal_stamp, qr_barcode, face_photo, fingerprint_thumb_impression, logo.

ImagePIIRedactor.list_entities() returns the exact set.

Run as a container

The image bundles torch+CUDA, Tesseract, the FastAPI service, and the React demo UI (web/, built at image-build time and served by the same process at /). Same image runs on GPU or CPU. This is also what runs on the ekacare/pii-redactor-demo HF Space.

docker build -t eka-pii-redaction .

# GPU
docker run --gpus all -p 7860:7860 \
  -e EKA_PII_HF_REPO=ekacare/pii-redactors \
  -v $HOME/.cache/huggingface:/root/.cache/huggingface \
  eka-pii-redaction

# CPU
docker run -e EKA_PII_DEVICE=cpu -p 7860:7860 eka-pii-redaction

Open http://localhost:7860 for the UI. API endpoints: GET /health, GET /entities, GET /entities-text, POST /detect / POST /redact / POST /anonymize (multipart file, optional exclude query param, mode/color form fields for /redact), POST /deidentify (multipart file → JSON {"image": <base64 png>, "mapping": ...}), and POST /detect-text / POST /redact-text / POST /deidentify-text / POST /anonymize-text (JSON {"text": ..., "exclude": [...]}; /deidentify-text also accepts "mapping" from a prior call and returns the updated one). Env: EKA_PII_HF_REPO, EKA_PII_DETECT_VISUAL, EKA_PII_DEVICE, EKA_PII_EXCLUDE.

curl -F file=@page.jpg http://localhost:7860/detect
curl -F file=@page.jpg -F mode=blur http://localhost:7860/redact -o redacted.png

Structure (by modality)

The library is organized by modality, so additional models slot in cleanly:

eka_pii_redaction/
  taxonomy.py, entities.py          # shared
  image/                            # IMAGE modality (implemented)
    redactor.py  -> ImagePIIRedactor
    layoutlmv3.py -> text-PII-in-image detector
    yolo11m.py    -> visual-entity detector
  text/                             # TEXT modality (implemented)
    redactor.py  -> TextPIIRedactor  (PII inside plain-text strings, no image)
    minilm.py    -> text-PII token classifier (char-span detector)

The single model repo mirrors this:

<hf_repo>/
  image/ layoutlmv3/   yolo/best.pt
  text/  minilm/                    # multilingual text-PII model

Both modalities share the category taxonomy (eka_pii_redaction.taxonomy); the text model also detects mac_address (device_net).

Deploying the demo to HF Spaces

The Space (ekacare/pii-redactor-demo) runs this same repo's Docker image — see the "Run as a container" section above. To (re)deploy:

  1. Create the Space once, as private (matches the model's current visibility — flip both to public together later):
    huggingface-cli repo create pii-redactor-demo --organization ekacare \
      --type space --space_sdk docker --private
    
    (or via the HF UI: New Space → owner ekacare → SDK Docker → Private.)
  2. Add the model's read token as a Space secret: Space → Settings → Repository secrets → add HF_TOKEN. The server already reads it via huggingface_hub's standard auth — no code change needed.
  3. Deploy: ./scripts/push_space.sh. It pushes the current commit to the Space's git remote, with .space-metadata.yaml's front matter prepended to README.md for that push only — the Space needs that front matter to render its card (title/sdk/app_port/...), but GitHub and PyPI don't know to strip HF-specific front matter, so it's kept out of the tracked README.md and only injected at deploy time. Auth: set HF_TOKEN, or it falls back to your cached hf auth login token.
  4. Watch the build under the Space's "Logs" tab. The base image (pytorch/pytorch:...-cudnn9-runtime) is large, so the first build can take a while; subsequent pushes reuse Docker layer caching.
  5. Once it shows Running, open the Space URL and click through both tabs.

Publishing the model weights

The trained checkpoints are assembled into the single HF repo with:

python scripts/build_hf_repo.py \
  --layoutlmv3   .../checkpoints/base_v3_combined_4ep/final \
  --yolo-weights .../checkpoints/visual/visual_yolo11m/weights/best.pt \
  --out /tmp/eka-pii-hf --push --repo-id ekacare/pii-redactors

How it works (image modality)

  1. Text-in-image: Tesseract OCR (via the processor) → words + boxes → token classifier → per-word BIO labels → merged spans.
  2. Visual: a detector over the page → boxes + categories.
  3. Redact: fill / blur / pixelate every selected entity's box.

Release files for eka-pii-redaction 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for eka-pii-redaction 0.3.0
File Size Uploaded
eka_pii_redaction-0.3.0.tar.gz 40.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for eka-pii-redaction 0.3.0
File Interpreter ABI Platform
eka_pii_redaction-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 76.5 kB

Release files / eka_pii_redaction-0.3.0.tar.gz

Download URL eka_pii_redaction-0.3.0.tar.gz
Size 40.6 kB
Tags Source
SHA-256 checksum
How to use checksums
3e4fe7d8954633cca2f7bd0ad5fc1e81270e9e2d00e49ebfd967bc7b2bb97323
BLAKE2b-256 checksum
How to use checksums
333a21590a22b2d4e5ac81d5b70441ef50b04b536d6582f1d359c5470b4190a0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.11

Release files / eka_pii_redaction-0.3.0-py3-none-any.whl

Download URL eka_pii_redaction-0.3.0-py3-none-any.whl
Size 35.9 kB
Tags Python 3
SHA-256 checksum
How to use checksums
92325a7a61b93b4415da5efe5c8e1922b5e65304392ea529b700e2b41f2a46e3
BLAKE2b-256 checksum
How to use checksums
b0ce05577bb5c52c2f12137bb8af94f64e3bb8b84794190c693942e45b9ad6fc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.12.11

Release history Release notifications | RSS feed

0.7.0

2 release files

0.6.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.6

2 release files

0.2.5

2 release files

0.2.4

2 release files

0.2.3

2 release files

0.2.2

2 release files

0.2.1

2 release files

0.2.0

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page