Skip to main content

Local Redact

CI

A local command-line redaction tool that removes personally identifiable information (PII), DevOps secrets, API keys, tokens, credentials, and other sensitive values from text files and images. It is built on top of Microsoft Presidio.

The Python distribution is named local-redactor (on PyPI); the installed command is redact.

The goal of this project is to provide a simple command such as:

redact sensitive-file.txt
redact screenshot.png

which creates sanitized copies that are safer to share in public forums, GitHub issues, support tickets, AI tools, documentation, and other external systems.

The project is designed to run locally so that the original sensitive data does not need to be uploaded to a third-party redaction service.

It runs on Windows, macOS, and Linux. The tool installs as a proper Python package with a redact console command, so the same command works identically on every operating system. Only the system prerequisites (Tesseract OCR, the spaCy model) differ per OS.


Quick Start

1. Install Tesseract OCR (required for image redaction)

OS Command
macOS brew install tesseract
Ubuntu / Debian sudo apt install -y tesseract-ocr
Windows winget install -e --id tesseract-ocr.tesseract

Skip this step if you only need text redaction.

2. Install Local Redact

pip install local-redactor

3. Download the spaCy language model

python -m spacy download en_core_web_lg

This is a one-time ~400 MB download that provides the NLP engine for detecting names, emails, phone numbers, and other PII.

4. Redact

redact sensitive-file.txt          # -> sensitive-file.redacted.txt
redact screenshot.png              # -> screenshot.redacted.png

The original file is never modified. Common options:

redact file.log --show-detections  # see what was detected and confidence scores
redact file.log -o clean.log       # custom output filename
redact file.log --force            # overwrite existing output

See Prerequisites and Installation below for virtual-environment setup, per-OS details, and development instructions.


Image Redaction Example

Before After
Unredacted DevOps secrets screenshot Redacted DevOps secrets screenshot

Why This Project Exists

Engineers regularly need to share:

  • Application logs
  • Terminal output
  • Configuration files
  • JSON/YAML
  • Error messages
  • Screenshots
  • Debugging information
  • Infrastructure configuration
  • Support diagnostics

These files can unintentionally contain sensitive information such as:

  • Names
  • Email addresses
  • Phone numbers
  • IP addresses
  • URLs
  • Account identifiers
  • Authentication tokens
  • API keys
  • Cloud credentials
  • Private keys
  • Connection strings

Manually finding and redacting every sensitive value is slow and error-prone.

This project provides a local automated redaction layer before information is shared externally.


Current Status

The current version supports text-based files using:

  • Presidio Analyzer
  • Presidio Anonymizer
  • spaCy
  • en_core_web_lg

It also supports PNG/JPG/JPEG image redaction using:

  • Presidio Image Redactor
  • Tesseract OCR
  • pytesseract

DevOps-specific secret detection is implemented with custom recognizers for common infrastructure credentials, API keys, tokens, private keys, and connection strings.

The project ships as an installable Python package (local-redactor) with a redact console entry point.


Architecture

Text Redaction

Input file
    |
    v
Presidio Analyzer
    |
    +-- Pattern recognizers
    |
    +-- spaCy NLP model
    |
    v
Detected entities
    |
    v
Presidio Anonymizer
    |
    v
Redacted output file

Example:

user-pii.txt
        |
        v
Presidio Analyzer
        |
        v
PERSON
EMAIL_ADDRESS
PHONE_NUMBER
        |
        v
Presidio Anonymizer
        |
        v
user-pii.redacted.txt

Image Redaction

Screenshot
    |
    v
Tesseract OCR
    |
    v
Extracted text + coordinates
    |
    v
Presidio Analyzer
    |
    v
Sensitive entities
    |
    v
Presidio Image Redactor
    |
    v
Opaque redaction boxes
    |
    v
screenshot.redacted.png

Prerequisites

Two things are required on every operating system and cannot be installed by pip, so they are installed per-OS below:

  1. Python 3.9 or newer
  2. Tesseract OCR (only needed for image redaction)

A third requirement, the spaCy language model, is installed with the same command on every OS after the Python package is installed (see Installation).

Pick your operating system:


Windows prerequisites

Python

Install from python.org or with winget:

winget install -e --id Python.Python.3.12
python --version

Tesseract OCR

Tesseract is a native system dependency and is not installed by pip.

winget install -e --id tesseract-ocr.tesseract

A typical installation location is:

C:\Program Files\Tesseract-OCR

Verify:

tesseract --version
tesseract --list-langs

At minimum, the English language model (eng) should be listed.

If Windows cannot find tesseract after installation, either add C:\Program Files\Tesseract-OCR to your PATH, or point the redactor at the binary directly with the TESSERACT_CMD environment variable:

$env:TESSERACT_CMD = "C:\Program Files\Tesseract-OCR\tesseract.exe"

macOS prerequisites

Python

macOS ships with Python 3, but a dedicated install via Homebrew is recommended:

brew install python
python3 --version

Tesseract OCR

brew install tesseract
tesseract --version
tesseract --list-langs

Homebrew places tesseract on your PATH automatically. If you installed it somewhere non-standard, point the redactor at it explicitly:

export TESSERACT_CMD="/opt/homebrew/bin/tesseract"

Linux (Ubuntu/Debian) prerequisites

Python

sudo apt update
sudo apt install -y python3 python3-venv python3-pip
python3 --version

Tesseract OCR

sudo apt install -y tesseract-ocr
tesseract --version
tesseract --list-langs

For other distributions, use the equivalent package (for example sudo dnf install tesseract on Fedora). If tesseract is not on your PATH, set TESSERACT_CMD to its full path:

export TESSERACT_CMD="/usr/bin/tesseract"

Installation

The steps below are the same on every OS once the prerequisites are in place. Windows users can run the same commands in PowerShell (adjusting only the virtual-environment activation line, noted below).

1. Clone the repository

git clone https://github.com/mumehta/local-redact.git
cd local-redact

2. Create and activate a virtual environment

A dedicated virtual environment keeps Presidio, spaCy, OpenCV, OCR libraries, and their dependencies isolated from system-wide Python packages.

Create it:

python3 -m venv .venv

Activate it:

macOS / Linux

source .venv/bin/activate

Windows (PowerShell)

.\.venv\Scripts\Activate.ps1

Your prompt should now be prefixed with (.venv).

3. Upgrade pip

python -m pip install --upgrade pip

4. Install the package

Install the project (and its pinned dependencies) into the virtual environment. This also creates the redact command.

python -m pip install .

For development (editable install plus test dependencies):

python -m pip install -e ".[dev]"

5. Install the spaCy English model

This command is identical on every OS:

python -m spacy download en_core_web_lg

Verify:

python -m spacy validate

en_core_web_lg should be listed as compatible.

6. Verify the installation

redact --help

You should see the CLI usage. Optionally verify the underlying engines:

python -c "from presidio_analyzer import AnalyzerEngine; a=AnalyzerEngine(); print([r.entity_type for r in a.analyze(text='My name is John Smith and my email is john@example.com', language='en')])"
python -c "import pytesseract; print(pytesseract.get_tesseract_version())"

The redact Command

Installing the package with pip install . (or pip install -e .) creates a real redact executable on every operating system:

  • On Windows, pip generates redact.exe in the environment's Scripts directory.
  • On macOS/Linux, pip generates a redact executable in the environment's bin directory.

No hand-written wrapper script is required. This replaces the older Windows-only redact.cmd approach (see Legacy Windows wrapper if you still want a global command that does not require activating the virtual environment).

pipx installs the command into an isolated environment and puts redact on your PATH, so you never have to activate a virtual environment to use it. This works the same on Windows, macOS, and Linux.

Install the published release from PyPI:

pipx install local-redactor

Or install from a local clone (for unreleased changes):

pipx install .

Then, from anywhere:

redact application.log
redact screenshot.png

You still need Tesseract and the spaCy model installed as described in Prerequisites and Installation. When using pipx, the spaCy model must be downloaded into the pipx-managed environment for local-redactor (pipx isolates each app, so a model installed elsewhere is not visible to it):

pipx runpip local-redactor -- python -m spacy download en_core_web_lg

Using the Redactor

With the virtual environment activated (or after pipx install), run:

redact ./example.txt

The tool creates:

example.redacted.txt

The original file is left unchanged.

Display detected entities

redact ./example.txt --show-detections

Example:

Detections:

PERSON               score=0.85 position=11:21
EMAIL_ADDRESS        score=1.00 position=38:54
PHONE_NUMBER         score=0.75 position=71:86

This is useful when testing detection accuracy.

Specify an output file

redact ./example.txt -o ./safe-to-share.txt

Overwrite an existing redacted file

By default, the tool will not overwrite an existing output file. Use --force when intentional overwriting is required:

redact ./example.txt --force

Image redaction

redact screenshot.png

produces:

screenshot.redacted.png

Supported image formats: .png, .jpg, .jpeg.

If Tesseract is not installed or not on your PATH, the tool prints a clear error with the correct install command for your OS. You can also point it at a specific Tesseract binary with the TESSERACT_CMD environment variable.


Supported Text Files

The current implementation supports text-based formats including:

.txt
.log
.json
.yaml
.yml
.env
.conf
.config
.ini
.xml
.csv
.md

Files are expected to contain UTF-8 text.

Structured formats such as JSON, YAML, XML, .env, and .ini are redacted as text. The tool preserves useful structure in many common cases, but it does not parse and reserialize these formats yet, so review generated output before using it as machine-readable configuration.


Example

Input:

My name is John Smith.
My email address is john@example.com.
My phone number is +61 412 345 678.

Run:

redact example.txt

Output (example.redacted.txt):

My name is <PERSON>.
My email address is <EMAIL_ADDRESS>.
My phone number is <PHONE_NUMBER>.

Dependency Management

Dependencies are declared in pyproject.toml:

  • Runtime dependencies live under [project].dependencies.
  • Development/test dependencies live under [project.optional-dependencies].dev and are installed with pip install -e ".[dev]".

Two supporting files remain for convenience:

requirements.txt

The direct runtime dependencies, mirroring pyproject.toml, for environments that prefer a plain requirements file:

presidio-analyzer==2.2.364
presidio-anonymizer==2.2.364
presidio-image-redactor==0.0.60

requirements-lock.txt

Captures the complete known-working Python environment, including transitive dependencies. Regenerate it with:

python -m pip freeze > requirements-lock.txt

Use it when an exact environment needs to be reproduced:

python -m pip install -r requirements-lock.txt

Development And Tests

Install the project with development dependencies:

python -m pip install -e ".[dev]"

Run the automated tests:

python -m pytest

The tests use synthetic PII and credential-shaped fixtures only. Image tests use fakes/mocks and do not require Tesseract to be installed.


System Dependencies

Some dependencies cannot be represented in pyproject.toml and are installed per-OS (see Prerequisites):

Tesseract OCR 5.x

The spaCy language model is also installed separately (same command on all operating systems):

python -m spacy download en_core_web_lg

Security Considerations

Automated redaction is not a security guarantee

The output of this tool should not automatically be assumed safe for public disclosure.

PII and secret detection systems can produce:

  • False positives
  • False negatives
  • Incorrect entity boundaries
  • OCR errors
  • Unrecognized credential formats

Review highly sensitive output before publishing it.

DevOps Secrets

Standard Presidio recognizers are primarily designed for PII.

Infrastructure and DevOps material may contain secrets that are not detected by default, including:

AWS_ACCESS_KEY_ID
AWS_SECRET_ACCESS_KEY

ghp_...
github_pat_...

Authorization: Bearer ...

JWT tokens

client_secret=...

password=...

N8N_ENCRYPTION_KEY=...

privkey:...

nlpriv:...

Kubernetes secrets

database connection strings

SSH/private keys

OAuth credentials

cloud provider credentials

Custom recognizers cover many of these patterns, but manually inspect infrastructure-related output before sharing it publicly.

Current Limitations

  • Files are processed one at a time; directory/batch redaction is not implemented yet.
  • Text files must be UTF-8 encoded.
  • JSON, YAML, XML, .env, .ini, and similar files are processed as text, not parsed as structured data.
  • OCR accuracy depends on screenshot quality, font size, contrast, and layout.
  • Detection is best-effort and can miss unfamiliar token formats or redact too much context.

Legacy Windows wrapper

Before the project was packaged with a console entry point, Windows users exposed a global redact command with a small .cmd wrapper that called the virtual environment's Python directly. This is no longer necessary — use pip install . or pipx install . instead, which produce a redact command on every OS.

The wrapper is documented here only for historical reference. If you still want a wrapper that invokes the project without activating the virtual environment, create C:\Users\<username>\bin\redact.cmd:

@echo off
"C:\path\to\local-redact\.venv\Scripts\python.exe" -m presidio_redactor.cli %*

and add C:\Users\<username>\bin to your user PATH.


Repository Structure

local-redact/
|
+-- .github/
|   +-- workflows/
|       +-- ci.yml               # cross-OS test matrix
|       +-- release.yml          # build + PyPI trusted publishing on v* tags
|
+-- .gitignore
+-- LICENSE                      # MIT
+-- README.md
+-- pyproject.toml               # packaging, deps, `redact` entry point
+-- requirements.txt
+-- requirements-lock.txt
+-- requirements-dev.txt
|
+-- src/
|   +-- presidio_redactor/
|       +-- __init__.py
|       +-- cli.py               # argparse + main(); the `redact` command
|       +-- text.py              # text redaction pipeline
|       +-- image.py             # image redaction + Tesseract resolution
|       +-- recognizers/
|           +-- __init__.py
|           +-- devops.py        # custom DevOps/secret recognizers
|
+-- tests/
|   +-- fixtures/
|   +-- test_devops_recognizers.py
|   +-- test_image_redaction.py
|   +-- test_text_redaction.py
|   +-- test_tesseract_resolution.py
|
+-- user-pii.txt                 # synthetic local example
+-- user-pii.redacted.txt        # synthetic redacted example
|
+-- .venv/                       # ignored by Git

Planned Features

Future development includes:

  • Content-based file type detection
  • Broader OAuth token detection
  • Broader GCP/Azure credential detection
  • GitLab token detection
  • Format-aware Kubernetes Secret parsing
  • Batch directory redaction
  • Dry-run mode
  • Configurable entity selection
  • Confidence thresholds
  • Windows Explorer "Redact before sharing" integration

Development Roadmap

The recommended implementation order is:

1. Text PII redaction                 DONE
        |
2. Global redact command              DONE
        |
3. Tesseract installation             DONE
        |
4. Image redaction                    DONE
        |
5. DevOps secret recognizers          DONE
        |
6. Automated synthetic tests          DONE
        |
7. Cross-platform support             DONE
        |
8. Python CLI packaging               DONE
        |
9. CI + PyPI release automation       DONE
        |
10. Batch redaction
        |
11. Windows Explorer integration

Git Safety

Test-data notice: All names, email addresses, phone numbers, passwords, tokens, credentials, connection strings, and other sensitive-looking values committed in tests/fixtures/, user-pii.txt, and user-pii.redacted.txt are synthetic dummy data. They are deliberately shaped like real PII and secrets to exercise the redaction pipeline. They are not valid credentials and are not associated with real accounts. Automated secret scanners may still flag these fixtures because their formats intentionally resemble real credentials.

Never commit real sensitive data simply to test the redactor.

Avoid committing:

.env
real application logs
credentials
private keys
access tokens
unredacted screenshots
production configuration
customer information
personal information

Use synthetic test data instead.

Do not rely on .gitignore as a security boundary. A file that has already been committed remains in Git history even if it is subsequently added to .gitignore.


Privacy Model

The primary design principle of this project is:

Sensitive source material should remain local wherever possible.

No data leaves your machine at runtime

Unlike an online redaction service, the entire redaction pipeline runs locally. When you invoke redact, your files are read, analyzed, and written on your own machine. No file content, detected entities, or redaction results are transmitted over the network. The tool does not collect telemetry, usage data, or analytics of any kind.

Network activity during setup only

Network access occurs only during initial installation and model download:

  • pip install / pipx install fetches Python packages from PyPI.
  • python -m spacy download en_core_web_lg downloads the NLP model (~400 MB) from GitHub/spaCy's CDN.

Once installed, the tool operates fully offline. If your environment requires air-gapped operation, you can pre-download the wheel and spaCy model, transfer them via removable media, and install from local files.

Third-party dependencies

This project relies on open-source libraries (Microsoft Presidio, spaCy, Pillow, pytesseract, and their transitive dependencies). Their privacy behavior is governed by their respective projects. None of these libraries are known to transmit user data at runtime, but this project does not control their code.

Image redaction uses opaque pixel replacement

Sensitive regions in images are overwritten with solid-colored boxes, not blurred. This means the original pixel data is destroyed in the output file and cannot be recovered, unlike Gaussian blur which can sometimes be reversed.

Redacted output still requires review

Redaction removes detected sensitive values, but surrounding context may still reveal operational information (system names, endpoints, usernames, timestamps, log structure). Users remain responsible for reviewing redacted output before publishing or transmitting it, particularly for high-sensitivity material.


License

This project is licensed under the MIT License.

The key dependencies and their licenses:

Dependency License
Microsoft Presidio MIT
spaCy MIT
Tesseract OCR Apache 2.0
pytesseract Apache 2.0
Pillow HPND (MIT-like)

All dependency licenses are compatible with the MIT License.


Acknowledgements

This project builds on:

  • Microsoft Presidio
  • spaCy
  • Tesseract OCR
  • pytesseract

These projects provide the underlying PII detection, natural-language processing, and optical character recognition capabilities.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

local_redactor-0.3.2.tar.gz (25.9 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

local_redactor-0.3.2-py3-none-any.whl (18.2 kB view details)

Uploaded Python 3

File details

Details for the file local_redactor-0.3.2.tar.gz.

File metadata

  • Download URL: local_redactor-0.3.2.tar.gz
  • Upload date:
  • Size: 25.9 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for local_redactor-0.3.2.tar.gz
Algorithm Hash digest
SHA256 742f3e21f940168d7f6cc2de95cf2d655c028eb09daac371db53ada36af2b59c
MD5 1c7676a25d3a187360199509a3a14a6d
BLAKE2b-256 98e34d248375c236cebee4c4d439602cec73d96d0d34eeb231d807e9d2193144

See more details on using hashes here.

Provenance

The following attestation bundles were made for local_redactor-0.3.2.tar.gz:

Publisher: release.yml on mumehta/local-redact

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file local_redactor-0.3.2-py3-none-any.whl.

File metadata

  • Download URL: local_redactor-0.3.2-py3-none-any.whl
  • Upload date:
  • Size: 18.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for local_redactor-0.3.2-py3-none-any.whl
Algorithm Hash digest
SHA256 0994394260c5ac56a38b6928b38d5f43ced4b4b5a79f2245957e9847feb0a24e
MD5 0a34b9f56e8696d4b48cc1e3f8788ec5
BLAKE2b-256 bbc4c48b098fa1fc15880bf07025bc999bb37c2433eda9a66e22a7e5a52c0ec2

See more details on using hashes here.

Provenance

The following attestation bundles were made for local_redactor-0.3.2-py3-none-any.whl:

Publisher: release.yml on mumehta/local-redact

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.3.2 This release

2 files

0.3.1

2 files

0.3.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page