Detect, redact, de-identify, or anonymize PII in document images and plain text.
Project description
Eka-PII-redaction
Detect PII in document images and plain text — both text PII (names, addresses, IDs, dates, phone/email, …) and visual entities (signatures, stamps/seals, QR/barcodes, face photos, fingerprints, logos) — then redact, de-identify, or anonymize it, behind a single class per modality.
- Three modes, one detector.
Redact destroys the value (mask/black-out — nothing kept).
De-identify replaces each entity with a consistent pseudonym
(
Person_1) and returns the entity→pseudonym mapping so an authorized key-holder can re-link later. Anonymize is one-way: ages become 10-year buckets, dates keep only the year, fine geography collapses, and names/IDs become unnumbered tokens — no mapping exists anywhere. - One install, one Hugging Face repo. The package internally pulls the model weights it needs from one HF repo; you only ever reference one repo id.
- Text + visual in one call, or text-only if you don't need the detector.
- CPU or GPU — auto-selects CUDA if available, else CPU.
- Choose what to process — exclude any categories; everything is on by default.
- Returns a structured entity list (
text,bbox,category, …) and/or a transformed image/string.
Note: anonymization here is best-effort removal/generalization of detected identifiers. It is not a k-anonymity guarantee and not, by itself, a compliance determination.
Install
From PyPI:
pip install eka-pii-redaction # core library
pip install "eka-pii-redaction[server]" # + FastAPI service
From source:
git clone https://github.com/eka-care/Eka-PII-redactors.git
cd Eka-PII-redactors
pip install -e . # core library
pip install -e ".[server]" # + FastAPI service
System dependency: Tesseract OCR — required for the image modality (it OCRs the document before the text-in-image classifier runs on the words); not needed for the text-only modality.
# Debian/Ubuntu
sudo apt-get install -y tesseract-ocr
# macOS
brew install tesseract
(The provided Docker image installs everything for you — see below.)
Quickstart
from eka_pii_redaction import ImagePIIRedactor
# Loads both models from one HF repo; GPU if available, else CPU.
redactor = ImagePIIRedactor(
"ekacare/pii-redactors",
detect_visual=True, # set False to skip visual entities (QR codes,
# face photos, signatures, etc.) — text PII only
# device="cpu", # force CPU (default: auto)
# exclude_entities=["logo", "brandname"], # never redact these
)
# 1) Get the structured entity list
for e in redactor.detect("page.jpg"):
print(e.kind, e.category, e.bbox, e.text, e.score)
# 2) Get a redacted image (PIL.Image)
redactor.redact("page.jpg").save("redacted.png")
redactor.redact("page.jpg", mode="blur").save("blurred.png") # solid|blur|pixelate
detect() / redact() accept a file path, raw bytes, or a PIL.Image.
Text-only PII (plain strings)
from eka_pii_redaction import TextPIIRedactor
r = TextPIIRedactor("ekacare/pii-redactors") # GPU if available, else CPU
# 1) Character-span entities
for s in r.detect("John Doe, DOB 1990-01-01, john@x.com"):
print(s.category, s.l1, s.start, s.end, s.text, s.score)
# 2) Redacted string (mask supports {category} / {l1} placeholders)
r.redact("Call John at john@x.com", mask="[REDACTED]") # -> "Call [REDACTED] at [REDACTED]"
r.redact("Call John at john@x.com", mask="[{category}]") # -> "Call [primary_subject_name] at [email]"
TextPIIRedactor runs a lightweight multilingual token classifier on raw text —
no OCR, no image — and returns TextPIISpan(category, start, end, l1, text, score)
with character offsets.
De-identify and anonymize
Both modalities support all three modes:
# --- Text ---
r = TextPIIRedactor("ekacare/pii-redactors")
result = r.deidentify("John Doe met Asha. Call John at +91 98765 43210.")
result.text
# -> "Person_1 met Person_2. Call Person_3 at Phone_1."
result.mapping.to_dict() # entity -> pseudonym map; store it securely to re-link.
# Pass mapping=result.mapping on the next page/document
# of the same record to keep numbering consistent.
r.anonymize("Mr. John Doe, 45 yrs, DOB 12-03-1979, Indiranagar, Karnataka.")
# -> "[PERSON], 40–49 yrs, DOB 1979, [LOCATION], Karnataka." (no mapping exists)
# --- Image ---
redactor = ImagePIIRedactor("ekacare/pii-redactors")
deid = redactor.deidentify("page.jpg") # ImageDeidResult
deid.image.save("deidentified.png") # pseudonyms rendered in place;
# faces/signatures become placeholders
deid.mapping.to_dict()
redactor.anonymize("page.jpg").save("anonymized.png")
# generalized values rendered in place; faces/signatures filled solid black
De-identification consistency is per exact surface form ("John Doe" and a
later bare "John" get different pseudonyms).
API
ImagePIIRedactor
ImagePIIRedactor(hf_repo, *, detect_visual=True, device=None,
exclude_entities=None, visual_score_threshold=0.25,
ocr_lang=None, cache_dir=None)
| arg | meaning |
|---|---|
hf_repo |
HF repo id or a local dir with text_model/ + visual_model/best.pt. |
detect_visual |
If False, visual entities (QR codes, face photos, signatures, etc.) are not downloaded or loaded — text PII only. |
device |
"cuda" / "cpu". None → auto (CUDA if available). |
exclude_entities |
Categories to never detect/redact. Default: none excluded (all on). |
visual_score_threshold |
Visual-entity confidence cutoff. |
ocr_lang |
Tesseract language/script, e.g. "eng", "eng+Devanagari". |
detect
detect(image, *, exclude_entities=None, ocr_lang=None) -> list[PIIEntity]
Each PIIEntity has:
| field | description |
|---|---|
category |
fine category, e.g. primary_subject_name, signature |
kind |
"text" or "visual" |
bbox |
[x0, y0, x1, y1] in original-image pixels |
l1 |
coarse group: person/location/contact/uid/… or biometric_visual |
text |
OCR text (text entities) or None (visual) |
score |
confidence in [0,1] |
redact
redact(image, *, mode="solid", color=(0,0,0), exclude_entities=None,
ocr_lang=None, pad=2) -> PIL.Image
Returns a redacted copy. mode ∈ solid | blur | pixelate.
list_entities
ImagePIIRedactor.list_entities() -> {"text": [...], "visual": [...]}
All selectable categories.
Categories
- Text (47): person (name, age, gender, occupation, …), location (address,
city, state, postcode, …), date_time, contact (phone, email, web_url, fax),
uid (aadhaar, pan, passport, mrn/uhid, abha, insurance policy, bank/iban/upi, …),
device_net, credential, and
brandname. - Visual (6):
signature,seal_stamp,qr_barcode,face_photo,fingerprint_thumb_impression,logo.
ImagePIIRedactor.list_entities() returns the exact set.
Run as a container
The image bundles torch+CUDA, Tesseract, the FastAPI service, and the React
demo UI (web/, built at image-build time and served by the same process at
/). Same image runs on GPU or CPU. This is also what runs on the
ekacare/pii-redactor-demo
HF Space.
docker build -t eka-pii-redaction .
# GPU
docker run --gpus all -p 7860:7860 \
-e EKA_PII_HF_REPO=ekacare/pii-redactors \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
eka-pii-redaction
# CPU
docker run -e EKA_PII_DEVICE=cpu -p 7860:7860 eka-pii-redaction
Open http://localhost:7860 for the UI. API endpoints: GET /health,
GET /entities, GET /entities-text,
POST /detect / POST /redact / POST /anonymize (multipart file,
optional exclude query param, mode/color form fields for /redact),
POST /deidentify (multipart file → JSON {"image": <base64 png>, "mapping": ...}), and POST /detect-text / POST /redact-text /
POST /deidentify-text / POST /anonymize-text
(JSON {"text": ..., "exclude": [...]}; /deidentify-text also accepts
"mapping" from a prior call and returns the updated one).
Env: EKA_PII_HF_REPO, EKA_PII_DETECT_VISUAL, EKA_PII_DEVICE, EKA_PII_EXCLUDE.
curl -F file=@page.jpg http://localhost:7860/detect
curl -F file=@page.jpg -F mode=blur http://localhost:7860/redact -o redacted.png
Structure (by modality)
The library is organized by modality, so additional models slot in cleanly:
eka_pii_redaction/
taxonomy.py, entities.py # shared
image/ # IMAGE modality (implemented)
redactor.py -> ImagePIIRedactor
layoutlmv3.py -> text-PII-in-image detector
yolo11m.py -> visual-entity detector
text/ # TEXT modality (implemented)
redactor.py -> TextPIIRedactor (PII inside plain-text strings, no image)
minilm.py -> text-PII token classifier (char-span detector)
The single model repo mirrors this:
<hf_repo>/
image/ layoutlmv3/ yolo/best.pt
text/ minilm/ # multilingual text-PII model
Both modalities share the category taxonomy (eka_pii_redaction.taxonomy); the
text model also detects mac_address (device_net).
Deploying the demo to HF Spaces
The Space (ekacare/pii-redactor-demo) runs this same repo's Docker image — see
the "Run as a container" section above. To (re)deploy:
- Create the Space once, as private (matches the model's current
visibility — flip both to public together later):
huggingface-cli repo create pii-redactor-demo --organization ekacare \ --type space --space_sdk docker --private
(or via the HF UI: New Space → ownerekacare→ SDKDocker→ Private.) - Add the model's read token as a Space secret: Space → Settings →
Repository secrets → add
HF_TOKEN. The server already reads it viahuggingface_hub's standard auth — no code change needed. - Deploy:
./scripts/push_space.sh. It pushes the current commit to the Space's git remote, with.space-metadata.yaml's front matter prepended toREADME.mdfor that push only — the Space needs that front matter to render its card (title/sdk/app_port/...), but GitHub and PyPI don't know to strip HF-specific front matter, so it's kept out of the trackedREADME.mdand only injected at deploy time. Auth: setHF_TOKEN, or it falls back to your cachedhf auth logintoken. - Watch the build under the Space's "Logs" tab. The base image
(
pytorch/pytorch:...-cudnn9-runtime) is large, so the first build can take a while; subsequent pushes reuse Docker layer caching. - Once it shows Running, open the Space URL and click through both tabs.
Publishing the model weights
The trained checkpoints are assembled into the single HF repo with:
python scripts/build_hf_repo.py \
--layoutlmv3 .../checkpoints/base_v3_combined_4ep/final \
--yolo-weights .../checkpoints/visual/visual_yolo11m/weights/best.pt \
--out /tmp/eka-pii-hf --push --repo-id ekacare/pii-redactors
How it works (image modality)
- Text-in-image: Tesseract OCR (via the processor) → words + boxes → token classifier → per-word BIO labels → merged spans.
- Visual: a detector over the page → boxes + categories.
- Redact: fill / blur / pixelate every selected entity's box.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file eka_pii_redaction-0.2.2.tar.gz.
File metadata
- Download URL: eka_pii_redaction-0.2.2.tar.gz
- Upload date:
- Size: 36.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
d0aba1516104fa13d22cd6a84d22ce06f313239757adb84a06bc226c6917e10a
|
|
| MD5 |
6fc269251f68438c8cfcb1051c2cd90a
|
|
| BLAKE2b-256 |
e7d5e762f96c03e2d3f4c14e06e49bdd897163420a7e5c941094f466449d26c8
|
File details
Details for the file eka_pii_redaction-0.2.2-py3-none-any.whl.
File metadata
- Download URL: eka_pii_redaction-0.2.2-py3-none-any.whl
- Upload date:
- Size: 32.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/7.0.0 CPython/3.12.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c66653a3268bfb448c8ddb820cbb808d4a8b8d21225f343f571e63810a1826cc
|
|
| MD5 |
9da4605b4d3ae6ed0ab4df11dc17a7c3
|
|
| BLAKE2b-256 |
ae5fa767187cde6c4ae95ee4630ef8da3923a3c3f0028014b6d8ecb7ec07c892
|