Skip to main content

Unmasker

What a human sees in a document, against what a machine reads out of it.

Local, read-only detection of hidden and failed-redaction content in documents.

PyPI Python 3.10+ 25 detectors Runtime dependencies Local and read-only Network requests CI License

Why Unmasker? · What it finds · Install · Usage · Examples · How it decides · JSON


A black rectangle drawn over text is not a redaction. The text is still in the file and every parser reads it. unmasker reports each place the two layers disagree — and says nothing beyond what it can show.

$ unmasker leaked.pdf

  unmasker  leaked.pdf                                          4 findings
  ────────────────────────────────────────────────────────────────────────

  ● 22 characters under a black shape at x 117.5-268.7, y           page 1
    684.2-698.5; the rest of the line still reads "Name:"
  │ human sees     ██████████████████████
  │ machine reads  Wanda Testowa-Przyklad

  ● 21 characters under a black shape at x 114.8-259.3, y           page 1
    664.5-678.7; the rest of the line still reads "Email:"
  │ human sees     █████████████████████
  │ machine reads  w.testowa@example.org

Every finding names the two readings, where to look, and how the tool knows. There are no verdicts: it does not say a document was manipulated, because that is the reader's conclusion to draw and a tool that draws it has to be trusted blindly.

Why Unmasker?

The person redacting a document sees a black bar and believes the job is done. It keeps happening in court filings and government releases, because the tool that draws the bar does not remove what is under it.

The same reading serves two more uses. A PDF fed to a retrieval pipeline can carry instructions a human reviewer will never see. And a leaked or altered document can be checked: what did the tracked changes hold, whose name is in the metadata, does the text layer agree with the picture.

What it finds

On the page

Detector What it reports
covered-text text under a filled shape, reported per character — a bar dragged too short is reported as covering exactly what it covers
text-under-image text under a picture, kept separate because a scan of a printed page looks the same and usually agrees with itself
invisible-text text a render mode never paints, or an opacity that paints nothing — color: transparent is one CSS declaration and changes no render mode at all
low-contrast-text text in the colour of what is behind it, whether that is a shape or the bare paper
off-page-text text outside the visible page — a crop box smaller than the media box is how a "cropped" file keeps what was cropped off

In the characters

Works on anything that yields text, so it covers DOCX, HTML, Markdown and source as well as PDF.

Detector What it reports
zero-width zero-width spaces, joiners, soft hyphens, word joiners
bidi-control direction overrides — a filename written invoice⟨U+202E⟩gpj.exe in the file reads as invoiceexe.jpg on screen, and is an executable
tag-characters plane-14 tag characters, decoded; the channel of choice for hiding instructions in text meant for a model
mixed-script a single word spanning two scripts — Cyrillic а inside a Latin domain

In a spreadsheet

A row, a column or a whole sheet carries an attribute saying not to draw it, and every value in it stays in the file exactly as typed. Someone selects three columns, right-clicks, chooses Hide, and sends the workbook out believing the numbers are gone.

Detector What it reports
hidden-sheet a sheet the workbook carries and never shows — and it says so louder when the sheet is marked veryHidden, which the application offers no way to undo
hidden-rows rows nobody sees, collapsed into one finding per block, because hiding rows 10 to 40 is one act by one person
hidden-columns the same by column, addressed the way the person who hid it saw it: column D, not an index
filtered-rows rows a filter is holding back rather than a person having hidden them — a weaker claim, and reported as one
changed-cell what change tracking took out of a cell and left in the file, with who changed it and when — the one finding here whose both columns carry text, because the current value is sitting in the cell

A spreadsheet stores a date as 45366 and a price as 240000. Dates are rendered, because a date cell holds a count of days and the conversion is exact. Everything else is quoted as the file stores it, with a note naming the format the sheet applies — a number formatter that is nearly right quotes a figure that is nearly right, which is worse than an exact quotation and a sentence of context.

In a presentation

Detector What it reports
hidden-slide a slide the deck skips when it is shown, quoted in full — the one that was cut before the meeting and never deleted
speaker-notes a note that was never on the screen, which is what notes are for and why the candid line ends up in one

In a photograph

A picture has no text layer, so the question inverts: not what does this document say that it does not show, but what does this file show that the picture does not.

Detector What it reports
stale-thumbnail the preview in the EXIF is a different shape from the picture, so it was not made from it — cropping a photograph does not regenerate the preview, and ImageMagick carries the old one through unasked

With --ocr the same file is asked the stronger question: what is legible in the preview and absent from the picture. On the specimen that is the witness name the crop removed.

In the container

Detector What it reports
deleted-text what a tracked deletion took off the page and left in the file, with who deleted it and when
comment comments — in Word, in OpenDocument, and in a PDF where they are annotations hanging off the page that no text extraction reports
revision-history one line naming who edited the file and when, never one finding per change
undisclosed-metadata a name, a client, a codename or a classification the document's own text never shows
metadata-path a filesystem path, which leaks a directory structure and usually an account
metadata-conflict the file contradicting itself — a PDF states its metadata twice, and a tool that clears one copy and not the other leaves exactly that

By reading the page back

--ocr renders each page and reads the picture with an OCR engine. It needs ghostscript and tesseract, and costs seconds a page, which is why it is not the default.

Detector What it reports
unrendered-text words the file holds that the page does not show — found without knowing how they are hidden
unextractable-text words the page shows that the file does not hold; the only finding here where the two columns swap

Every other detector knows a trick, and each was written after a producer was caught doing something particular. This one knows nothing, so a method nobody has thought of still fails it. On the specimens the two approaches name the same words on every file that hides text, and both stay silent on every control — which is not something either could arrange for the other.

Installation

pipx install unmasker

or uv tool install unmasker. Either puts unmasker on the PATH in an environment of its own, which is what you want for a command-line tool: it never shares a site-packages with anything you are investigating.

From a checkout, to work on it:

python3 -m venv .venv
.venv/bin/pip install -e .
.venv/bin/unmasker document.pdf

Python 3.10 or later, and one runtime dependency — pypdf, which is pure Python, BSD-licensed and has no dependencies of its own. It was checked rather than assumed, and any second one has to earn its place the same way — in writing.

Reads PDF, DOCX, ODT, XLSX, ODS, PPTX, ODP, JPEG and any text file. Local, read-only, no network, and it never writes to the file it is given.

Usage

unmasker <file>   [options]
unmasker <folder> [options]

Point it at a folder and it surveys the lot: how much was read, how much could not be, which kinds of finding turned up and in how many files, and which files to open next. The screen triages and --json carries every finding, so there is no --full — the archive already exists.

It never ranks the files. Sorting the worst documents to the top would be the judgement this tool leaves to its reader, so the tally counts files per kind and the list is in path order.

Option Purpose
--json one object on stdout, for a pipeline that wants to sort or filter
--html one self-contained page on stdout — redirect it into a file and send it
--md Markdown on stdout, for a wiki, a ticket or a pull request
--ocr render each page and read the picture back (needs ghostscript, tesseract)
--width N wrap at N columns instead of measuring the terminal
--version print the version and exit
-h, --help the full option list
Code Meaning
0 read, searched, nothing found
1 read, searched, findings exist
2 could not be read

[!IMPORTANT] Three exit statuses rather than two, because a file that could not be read is not a file that came back clean. A pipeline that cannot tell those apart will eventually wave through the one document it should have stopped.

Practical examples

One document, reported for a person

unmasker tests/specimens/pdf/libreoffice-writer-black-bars.pdf
  unmasker  libreoffice-writer-black-bars.pdf                     4 findings
  ──────────────────────────────────────────────────────────────────────────

  ● 22 characters under a black shape at x 117.5-268.7, y             page 1
    684.2-698.5; the rest of the line still reads "Name:"
  │ human sees     ██████████████████████
  │ machine reads  Wanda Testowa-Przyklad

  ● 15 characters under a black shape at x 123.1-228.1, y             page 1
    644.7-658.9; the rest of the line still reads "Telephone:"
  │ human sees     ███████████████
  │ machine reads  +48 601 000 000

  notes                                                               1 note
  ──────────────────────────────────────────────────────────────────────────
    the file says it was made by Creator Writer; Producer LibreOffice 24.2,
    and dates itself CreationDate 2026-08-31T20:25:50Z

  ──────────────────────────────────────────────────────────────────────────
  searched 1 page of 1. 4 findings in 1 kind.

A workbook whose columns were hidden rather than removed

unmasker tests/specimens/xlsx/libreoffice-calc-hidden-columns.xlsx
  hidden-sheet                                                     1 finding
  ──────────────────────────────────────────────────────────────────────────
  ● sheet "Workings" is marked hidden, and holds 1 value still    whole file
    in the file
  │ human sees     nothing on the page
  │ machine reads  Reserve set at 240,000. Kowalski came in 12% under; the
  │                others were told nothing.

  hidden-columns                                                   1 finding
  ──────────────────────────────────────────────────────────────────────────
  ● column D of sheet "Evaluation" is hidden, and holds 5 values  whole file
    still in the file

A folder, surveyed: which file to open next

unmasker tests/specimens/docx
  unmasker  tests/specimens/docx                                3 of 6 files
  ──────────────────────────────────────────────────────────────────────────
    read      6 files, 15 findings
    not read  0 files

  what was found                                                     9 kinds
  ──────────────────────────────────────────────────────────────────────────
    zero-width            1 file
    bidi-control          1 file
    tag-characters        1 file
    mixed-script          1 file
    undisclosed-metadata  1 file
    metadata-path         1 file
    deleted-text          1 file
    comment               1 file
    revision-history      1 file

  files that hide something                                          3 files
  ──────────────────────────────────────────────────────────────────────────
    libreoffice-writer-hidden-characters.docx  zero-width, bidi-control,
                                               tag-characters, mixed-script
    libreoffice-writer-metadata-leak.docx      undisclosed-metadata,
                                               metadata-path
    libreoffice-writer-tracked-changes.docx    deleted-text, comment,
                                               revision-history

  ──────────────────────────────────────────────────────────────────────────
  searched 6 files. unmasker <file> for the detail, --json for all of it.

A page somebody can be sent

unmasker ~/cases/kowalski --html > report.html

One file, no external anything, no JavaScript, and print rules for the day it goes into a case file. It carries the full detail rather than the survey's summary — a browser has search and a scrollbar where a terminal has neither.

When the technique is unknown

unmasker scan.pdf --ocr

Report safety

Everything this tool quotes came out of a document somebody else wrote. That makes every report an untrusted document too, and the two shareable formats are escaped accordingly.

--html writes to stdout and is redirected, rather than taking an --out option, because this tool never writes anything. Every value on the page is escaped: a PDF whose metadata reads <img src=x onerror=…> would otherwise put a live handler into the report of itself.

[!WARNING] --md is the more dangerous of the two to get wrong, not the safer. An HTML renderer handed <script> prints it; a Markdown renderer runs it, because passing raw HTML through is what Markdown does.

So quoted evidence goes in a fenced block — with a fence grown longer than any run of backticks inside it — prose is escaped, and a | never reaches a table cell unescaped.

How it decides what to say

Colour encodes how the tool knows, never how bad it is. Three classes, learned once:

  • direct — the bytes are in the file and were read out; nothing is inferred
  • circumstantial — consistent with hiding, and with innocent explanations too. A word spanning two scripts may be an attack or may be how somebody writes; OCR failing to read text looks exactly like text not being there
  • self-reported — the file's own account of itself. A name in a document is whatever the application that wrote it was configured to say

There are no scores. 55 reads as a probability and never was one; a word can be argued with by the person reading the report, which is the point.

Different questions are never ranked against each other. A page can have a bar over its text and an invisible character and stale metadata. Those are three findings, not one winner.

[!NOTE] "Nothing found" has two meanings, and the report keeps them apart. Searched and it is not there is one. There was nothing this tool could search is the other — a page with no text layer, or text sitting on a picture whose colour at that point is not in the file. The second comes with a note saying so, and --json carries it as "searched": false.

Nothing is truncated. A value too long for the line wraps. An ellipsis sends the reader to fetch the value another way, which defeats having read the report.

JSON and automation

--json writes one object on stdout. The exit status still gates, so a pipeline needs no second mode:

unmasker leaked.pdf --json > findings.json || echo "findings exist"

Two shapes, named in the document itself:

Field Meaning
schema unmasker.scan/1 for one file, unmasker.survey/1 for a folder. The two carry different keys and this is what tells them apart
version which build wrote the document — a different question from which shape it is
searched false means there was nothing to search, not that the search came back empty
{
  "tool": "unmasker",
  "schema": "unmasker.scan/1",
  "version": "0.1.0",
  "file": "leaked.pdf",
  "kind": "pdf",
  "searched": true,
  "remarks": ["…"],
  "findings": [
    {
      "detector": "covered-text",
      "basis": "direct",
      "summary": "22 characters under a black shape at x 117.5-268.7, y 684.2-698.5; the rest of the line still reads \"Name:\"",
      "human_sees": "██████████████████████",
      "machine_reads": "Wanda Testowa-Przyklad",
      "location": { "page": 1 },
      "codepoints": ["U+0057", "U+0061", "…"]
    }
  ]
}

The field order is deliberate and is the order a consumer sees. codepoints carries every codepoint behind the finding rather than a sample, for the same reason nothing on the screen is truncated. A folder survey replaces file with root and nests the same per-file objects under files.

The /1 is what lets the shape change later without silently breaking anything built against it.

How it is tested

Every detector fires on a committed specimen that a real producer wrote — LibreOffice, headless Chrome, Ghostscript, tesseract, exiftool, ImageMagick. There are 32 of them and they are the test suite; each has a .md beside it recording which tool made it, what a human sees when it is opened, and what is actually inside.

That discipline is not decoration. The first specimen disproved the design: LibreOffice draws a redaction bar as a polygon filled with f* and emits no re at all, so the rectangle-based detector the plan called for would have found nothing on the archetypal case — and a fixture built from the PDF specification would have hidden that behind a green test suite.

Where a specimen needs a measurement — where a bar goes, how wide a glyph is — it comes from poppler, not from this project's own code. A fixture measured with the tool under test proves only that the tool agrees with itself.

Findings are also mutation-tested: each rule is broken on purpose and the suite has to notice. That has repeatedly found behaviour the docstrings claimed and nothing checked.

The detector tables above are held against the source by tests/test_documented_detectors.py: every slug the code can emit has to appear in a table, every slug in a table has to be one the code emits, and the badge has to agree with both. This document went stale once — it claimed 22 detectors against 25 — and a number beside a list does not keep itself true.

716 tests.

Limits

  • No verdicts. It will not tell you a document was manipulated.
  • No writing. It never modifies the file it is given.
  • No network.
  • Not every container. Legacy OLE2 (.doc, .xls, .ppt) is unread, and so is XMP outside a PDF. The sibling project filetrail reads several of those and answers a different question with them — where a file came from.
  • Not every producer. Word and Acrobat are not on the machine this was built on, and two producers never agree about everything. tests/specimens/README.md keeps the current list of what is untested and why.

Development

git clone https://github.com/osint-shifu/unmasker
cd unmasker
python3 -m venv .venv
.venv/bin/pip install -e ".[dev]"
.venv/bin/pytest
.venv/bin/ruff check .

ruff check ., and not ruff format --check .. This project lints and does not auto-format: the report layout, the specimens' XML and the docstrings are placed deliberately, and a formatter run at publication time would rewrite a dozen files nobody asked it to. CI asserts the standard the project actually holds.

  • CONTRIBUTING.md — the rules this project works to, each one named after the failure that produced it
  • SECURITY.md — what to do about a finding in this tool itself
  • CHANGELOG.md — what changed, and when
  • tests/specimens/README.md — the specimens, what each proves, and the gaps that are named rather than hidden
  • the sibling project's design notes — the design language this report follows

License

Apache License 2.0.


Unmasker

What a human sees in a document, against what a machine reads out of it.

Made by osint-shifu

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

unmasker-0.1.0.tar.gz (700.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

unmasker-0.1.0-py3-none-any.whl (146.2 kB view details)

Uploaded Python 3

File details

Details for the file unmasker-0.1.0.tar.gz.

File metadata

  • Download URL: unmasker-0.1.0.tar.gz
  • Upload date:
  • Size: 700.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for unmasker-0.1.0.tar.gz
Algorithm Hash digest
SHA256 013c2c519c2b66df061c52b238675ceb66f34460c1565c0fbe4c5ee2018f274c
MD5 0af1c4e315bc8cfa932e878c422c4e4e
BLAKE2b-256 9b1cc9a384b273ee70657519aa54bab4f210dd475d7a61bfbc8033f0ffa3d261

See more details on using hashes here.

Provenance

The following attestation bundles were made for unmasker-0.1.0.tar.gz:

Publisher: release.yml on osint-shifu/unmasker

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file unmasker-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: unmasker-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 146.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for unmasker-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 43efb3792be53b852c4eb1096718923a8fb93c2730d02063ddaf561c803cbceb
MD5 28e3db7230d20a98fbe21074a4296244
BLAKE2b-256 d10bb3ba781ec456357b5e8716dc0df9810391a18d7391743fdaa1ec3d1b8199

See more details on using hashes here.

Provenance

The following attestation bundles were made for unmasker-0.1.0-py3-none-any.whl:

Publisher: release.yml on osint-shifu/unmasker

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.3.1

2 files

0.3.0

2 files

0.2.1

2 files

0.2.0

2 files

0.1.12

2 files

0.1.11

2 files

0.1.10

2 files

0.1.9

2 files

0.1.8

2 files

0.1.7

2 files

0.1.6

2 files

0.1.5

2 files

0.1.4

2 files

0.1.3

2 files

0.1.2

2 files

0.1.1

2 files

This release

0.1.0 This release

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page