Watermark Cleaner
Remove AI text artifacts, hidden unicode characters and file metadata before you publish.
Watermark Cleaner is a deterministic, offline, zero-dependency tool for content you own. It cleans the mechanical traces that mark text as machine generated, removes the citation markers and tracking parameters that assistants leave in copied text, flags the stylistic phrases that read as AI writing (English and German), and strips identifying metadata from images and office documents without touching a single pixel. It ships as a Python CLI and a Node CLI that share one rule set and are tested to behave identically.
Try it in your browser, no install: christianstrunk.com/tools/ai-watermark-cleaner
See it
The problem is invisible by definition. Escaped, it looks like this:
before "Smart quotes “work” — and a hidden watermark"
after "Smart quotes \"work\", and a hidden watermark"
$ watermark-cleaner check post.md
post.md (1 error, 5 would fix)
[block] ai phrases present (rewrite required, not auto-fixed) x1 (in today's fast-paced world)
[would fix] removed zero width space (U+200B) x2
[would fix] straightened smart quotes x2
[would fix] replaced em/en dash per policy x1
1 files scanned | 1 would change | 5 would fix | 0 warnings | 1 blocking
$ echo $?
1
check reports and never changes anything. fix cleans in place and writes a .bak backup next to every changed file. The backup keeps the first original: running fix again after editing does not overwrite an existing .bak.
Install
| Ecosystem | One-shot | Install |
|---|---|---|
| Python | uvx watermark-cleaner check . or pipx run watermark-cleaner check . |
pipx install watermark-cleaner |
| Node | npx watermark-cleaner check . |
npm install -g watermark-cleaner |
| pre-commit | see below | .pre-commit-config.yaml snippet |
| From source | pip install . or npm install -g ./node |
Both packages install a single command, watermark-cleaner, available in every folder, like git.
Use
watermark-cleaner check . # report only, changes nothing, exits non zero if blocked
watermark-cleaner fix . # clean in place, writes a .bak backup per changed file
watermark-cleaner check post.md
watermark-cleaner fix ./content --no-voice
watermark-cleaner fix ./assets --aggressive # also replace homoglyphs and strip variation selectors
watermark-cleaner check . --json
watermark-cleaner check . --strict # exit 1 when anything would change, for ci
pbpaste | watermark-cleaner fix - | pbcopy # clean the clipboard (macos); report goes to stderr
check is safe to run over an entire folder to see what is inside without touching anything. Exit codes: check returns 1 when a file contains AI phrases or AI sentence shapes, so it works as a publish gate. check --strict also returns 1 when anything would change, which turns it into a CI lint for hidden characters and metadata. fix returns 0 after cleaning what it can; pass --strict to make fix return 1 when blocking findings remain that need a human rewrite.
A single - instead of a path reads standard input. fix - writes the cleaned text to stdout and the report to stderr, so it drops into a pipe: copy text from a chat, run it through the cleaner, paste the result.
Markdown structure is protected: typography, voice, entity and artifact rules never touch fenced code blocks, inline code (single or double backticks) or YAML frontmatter, so code examples in documentation survive cleaning. Invisible characters are still removed inside code, because hidden characters in code are exactly the Trojan Source attack this tool defends against. Disable with "protect_code": false if you want full-document typography.
What it removes, and what it deliberately keeps
| Category | Codepoints | Action |
|---|---|---|
| Zero width space, word joiner, BOM | U+200B, U+2060, U+FEFF | removed |
| Soft hyphen, mongolian vowel separator, combining grapheme joiner | U+00AD, U+180E, U+034F | removed |
| Invisible math operators | U+2061 to U+2064 | removed |
| Deprecated format controls, interlinear annotation characters, hangul fillers | U+206A to U+206F, U+FFF9 to U+FFFB, U+3164, U+FFA0 | removed |
| Unicode tag characters (hidden payloads, tag space) | U+E0000 to U+E007F | removed, except inside well-formed subdivision flag emoji |
| Private use characters (ChatGPT citation delimiters U+E200 to U+E203 and everything else vendors hide there) | U+E000 to U+F8FF, planes 15 and 16 | removed |
| Noncharacters | U+FDD0 to U+FDEF, U+FFFE, U+FFFF | removed |
| Hangul fillers | U+115F, U+1160, U+3164, U+FFA0 | removed |
| Variation selectors without a base that supports them (letters, spaces, selector chains) | U+FE00 to U+FE0F, U+E0100 to U+E01EF, U+180B to U+180D | removed; after emoji, symbols, digits and CJK they stay |
HTML entities of all of the above (​, ‌, ­, ) |
decoded, then the same rules apply | removed |
Assistant copy artifacts (citeturn0search0, [cite: 1], 【1†source】, utm_source=chatgpt.com) |
per rules/artifacts.json |
removed |
| Exotic spaces and blanks (nbsp, thin space, ideographic space, line and paragraph separator, braille blank) | U+00A0, U+2000 to U+200A, U+202F, U+205F, U+3000, U+2028, U+2029, U+2800 | replaced with a normal space |
| Smart quotes, ellipsis, bullets, dashes, typographic hyphens | U+2018, U+2010, U+2011 and friends | normalized to ascii |
| Homoglyphs (cyrillic і inside a latin word and friends) | per rulebook | warned, replaced only with --aggressive |
Deliberately preserved, because removing them breaks legitimate text:
- Subdivision flag emoji (England, Scotland, Wales). They are a black flag followed by tag characters and a cancel tag; the well-formed sequence stays, a tag character anywhere else is removed.
- Zero width joiner (U+200D) next to an emoji or inside a script that requires it (Indic scripts, Arabic). Between latin letters it is a watermark and gets removed.
- Zero width non-joiner (U+200C) when adjacent to a script that requires it (Persian, Arabic, Indic scripts and others). Between latin letters it is a watermark and gets removed.
- Bidi marks and bidi controls in documents that contain right-to-left text. In pure left-to-right documents they are removed, which is the Trojan Source defense.
- Non breaking spaces next to a digit (
12 000,10 %,5 kg,§ 5,Nr. 5), after an ordinal (5. Mai), inside spaced abbreviations (z. B.,d. h.,i. d. R.) and in French punctuation (Il dit : bonjour !,« Oui »). Everywhere else, between two words or between two sentences, a non breaking space is a watermark and becomes a normal space. Disable the guard withkeep_nbsp_in_numbers: false. - A spaced dash between two numbers (
10 – 20 Uhr,1990 — 2000) becomes a plain hyphen, never the comma fromdash_policy, because a range is not an enumeration. - Variation selectors that belong to something:
❤️, keycaps like1️⃣,©️, CJK ideographic variations and Mongolian free variation selectors. A selector after a latin letter, a space or another selector carries no meaning and is removed as a hidden payload. - The combining grapheme joiner (U+034F) next to Hebrew and the braille blank (U+2800) inside braille text.
- Genuine Cyrillic and Greek text. Homoglyph detection works per word: a word that mixes latin letters with look-alike letters is flagged, a word written entirely in Cyrillic or Greek is normal text and stays untouched even with
--aggressive. A word that consists only of look-alike letters (СОРЕspelled in Cyrillic) is flagged only when the document contains no other Cyrillic or Greek text.
Coverage
| Channel | Handled | Notes |
|---|---|---|
| Invisible/format Unicode (ZWSP, bidi, tag chars, private use, variation selector payloads, homoglyphs) | Yes | Deterministic, verifiable |
| HTML entity forms of invisible characters | Yes | Decoded, then the character rules apply |
| Assistant copy artifacts (citation markers, tracking parameters) | Yes | Deterministic, from rules/artifacts.json |
| Typography (smart quotes, dashes, ellipsis) | Yes | Deterministic, verifiable |
| AI filler phrases and sentence shapes | English and German | Filler removed (English only), shapes flagged |
| Image metadata: EXIF, XMP, C2PA, comments | JPEG, PNG, WebP, SVG | Lossless, container-level |
| Image metadata: GIF, TIFF, HEIC, HEIF, AVIF | No | Reported, file left untouched |
| Document metadata: DOCX, PPTX, XLSX, ODT, ODP, ODS | Yes | Lossless, container-level (see below) |
| Document metadata: PDF | Reported only | check lists author, creator, producer and XMP presence; the file is never rewritten, see Non-goals |
| Statistical/token-level text watermarks (SynthID-Text in Gemini and, since August 2026, Claude) | Best-effort only, opt-in | Via DeepL rewrite, not a deterministic strip |
| Pixel-domain image watermarks | No | Out of scope, see below |
| Training-time backdoors | No | Out of scope, not a watermarking concern this tool addresses |
Image metadata, losslessly
watermark-cleaner fix strips EXIF, XMP, C2PA and comment segments from JPEG, PNG and WebP at the container level. Pixels are never re-encoded, so there is no quality loss and the operation is verifiable with a byte diff. ICC color profiles are kept by default, because removing them visibly shifts colors in the browser; strip them too with "strip_icc": true or --aggressive. SVG metadata, XMP blocks and XML comments are removed as text; an SVG that is not valid UTF-8 (older Illustrator exports in ISO-8859-1, for example) is skipped with a warning instead of being re-encoded. GIF and TIFF are never modified; the tool tells you it cannot strip them losslessly and leaves them alone. Files that fail to parse are left untouched.
Document metadata, losslessly
DOCX, PPTX, XLSX, ODT, ODP and ODS are ZIP containers. watermark-cleaner fix removes the people and the software from them, container-level and lossless. For DOCX that means the author and last-editor fields in docProps/core.xml, the application name, application version, template name and total editing time in docProps/app.xml, the hidden people list in word/people.xml, and the author names and initials on every comment and tracked change in the document body, headers, footers and notes. Comments and tracked changes themselves stay in place, only the name on them is blanked. PPTX and XLSX get the same core and app treatment plus the names and initials of comment authors (ppt/commentAuthors.xml, ppt/authors.xml, xl/comments*.xml, xl/persons/person.xml). For ODT, ODP and ODS it means creator, initial creator, generator, editing duration and printed-by in meta.xml plus the creator on every annotation and tracked change in content.xml. Every part that carries no such field, including styles and embedded images, is copied through byte-for-byte, verifiable with a byte diff exactly like the image path. Rewritten parts are stored uncompressed so both CLIs produce identical bytes; a document with heavy tracked changes can therefore grow by a few hundred kilobytes. Title, subject, keywords, description and edit dates are left alone; those are content you likely want, not a machine fingerprint. Files that fail to parse as a well-formed ZIP are left untouched. PDF is scanned read-only: check tells you when author, creator, producer, XMP or a C2PA manifest is present, and both commands leave the file alone.
Use as a publish gate
With the pre-commit framework:
repos:
- repo: https://github.com/pixelstrunk/watermark-cleaner
rev: v0.3.0
hooks:
- id: watermark-cleaner-fix
watermark-cleaner-fix cleans the mechanical layers on every commit and never blocks on style. Add id: watermark-cleaner-check if you also want commits blocked on AI phrases. A plain git hook and a deploy gate script live in integrations/.
Coding agents can drive the CLI through the skill in skills/watermark-cleaner/, which enforces an inspect-first workflow.
All writes are safe by construction: files are replaced atomically, writes through symlinks are refused, and files larger than max_file_bytes (default 256 MiB) are skipped instead of loaded into memory. A file or folder the tool is not allowed to read or write is reported as a warning and skipped; the run continues with the next file instead of aborting.
The voice layer, honestly
Cleaning AI text falls into buckets, and this tool is precise about which bucket each action lives in.
| Layer | What | Result |
|---|---|---|
| Characters | Invisible and format characters, exotic spaces, homoglyphs, NFC normalization | Fixed, 100% verifiable |
| Typography | Smart quotes, ellipsis and bullet glyphs, em and en dashes | Fixed, 100% verifiable |
| File metadata | EXIF, XMP and C2PA in images, metadata and comments in SVG, author/app fields in DOCX and ODT | Stripped losslessly, 100% verifiable |
| Voice | AI filler phrases and AI sentence shapes from a portable rulebook | Filler removed, shapes flagged and blocked |
The voice layer auto-deletes only phrases that are pure filler ("without further ado"). When a deletion opens a sentence, the sentence is repaired: leftover spaces go away and the next word is capitalized. Phrases that carry an object ("let's explore the API") are never cut mid-sentence; they are flagged as errors for a human to rewrite. The rulebook covers English and German. German phrases are only ever flagged, never deleted, because removing a German clause opener changes the word order of what follows. Chat residue such as "Certainly! Here's" or "I hope this helps" blocks outright; it is the clearest sign that text was pasted from an assistant.
Non-goals, stated plainly
- It does not remove statistical text watermarks. Google has embedded SynthID-Text in Gemini output since 2024, and Anthropic announced in August 2026 that all Claude models released after 2 August 2026 carry the same kind of watermark; OpenAI has not deployed one for text. These watermarks live in word choice, add no hidden characters, survive light editing and disappear only under a complete rewrite. No character-level tool can detect or remove them, and no public detector exists to verify removal. The April 2025 reports of ChatGPT inserting narrow no-break spaces were a training artifact that OpenAI removed within days, not a watermark; this tool removes those spaces anyway.
- It does not auto-rewrite AI sentence shapes. Rewriting a sentence needs judgment, so the tool detects and blocks those shapes instead of replacing them and producing nonsense.
- It does not certify that text will pass any AI detector, and it is not a way to misrepresent authorship.
- It does not rewrite PDF files. Lossless PDF editing needs a full object parser, and an incremental update would leave the old metadata in the file for anyone who looks.
checkreports what a PDF carries; to strip it, useexiftool -all= file.pdforqpdf, or export to a supported format before running the cleaner. - It does not detect or remove pixel-domain image watermarks (SynthID-class, StegaStamp, Tree-Ring and similar signals embedded in the pixels themselves). That is a fundamentally different, model-heavy problem; this tool only ever touches metadata containers and text, never re-encodes pixels.
- It does not address training-time backdoors or any watermarking mechanism baked into a model's weights. Out of scope for a client-side content cleaner.
Configure
Drop a watermark-cleaner.config.json in a project root. Any key overrides the default.
{
"voice": true,
"replace_homoglyphs": false,
"keep_nbsp_in_numbers": true,
"protect_code": true,
"strip_icc": false,
"dash_policy": { "spaced_replacement": " - ", "unspaced_replacement": "-" },
"custom_banned_phrases": ["our innovative product", "leverage synergies"],
"ignore_phrases": ["delve", "not-only-but-also"],
"text_extensions": [".md", ".mdx", ".txt"],
"exclude": ["node_modules", ".git", "dist"]
}
By default a spaced em dash ("fast — slow") becomes a comma ("fast, slow") and an unspaced one becomes a hyphen. That is an opinionated policy; dash_policy lets you change both replacements per project.
custom_banned_phrases adds your own blocking phrases on top of the rulebook. ignore_phrases switches off any built-in phrase, lexicon word or sentence shape id that does not fit your domain.
Optional DeepL rewrite
export DEEPL_API_KEY=your-key
watermark-cleaner rewrite post.md
watermark-cleaner rewrite post.md --source-lang DE --pivot-lang EN
watermark-cleaner rewrite post.md --write
The document language is auto-detected by DeepL when --source-lang is not given.
This back translates the text (source to pivot and back) using a model that is not the one that wrote the text, then runs the deterministic cleaner on the result. It changes word choice, which is what degrades a statistical watermark, but it can shift meaning and it sends your text to DeepL. It is off by default and opt in. Do not use it on confidential content. It cannot certify that any vendor detector will fail.
Use as a library
The Node package has two entries. watermark-cleaner is for Node and loads the rules from disk, watermark-cleaner/browser has no file system access and ships the rules bundled, so it works in a web page or a worker. Both export the same functions. The web version at christianstrunk.com/tools/ai-watermark-cleaner runs on exactly this browser entry.
import { cleanText, classifyCharacters, findPhrases, stripJpeg, RULES } from "watermark-cleaner/browser";
const { text, report } = cleanText(input);
report.findings[0].by_rule; // { "chatgpt-citation-token": 2, "utm_source-tracking": 1 }
classifyCharacters("a\u200bb"); // [{ index: 1, char: "\u200b", code: 8203, action: "remove", name: "zero width space" }, ...]
findPhrases("Without further ado, we delve in."); // [{ start: 0, end: 19, kind: "filler", id: "without further ado", severity: "fixed", text: "Without further ado" }, ...]
stripJpeg(bytes); // { cleaned: Uint8Array, stripped: 2, orientation: 6, blocks: [{ kind: "exif", label: "APP1 (EXIF)", bytes: 84 }, ...] } or null when the file does not parse
cleanText(text, config?, rules?) takes a partial config merged over the CLI defaults. Deep imports of src/ and rules_data/ are not part of the public API; use RULES (browser) or loadRules() (Node) instead.
In Python, from watermark_cleaner import clean_text returns (text, report) with the same findings and by_rule counts.
Rules are one source
All character lists, phrase lists and copy-artifact patterns live in rules/*.json. The Python and Node packages read the same files, scripts/sync-rules.sh copies them into each package, and CI fails when they drift or when the two CLIs produce different output. Edit the rules once, both tools follow. Contributions to the rulebook are the easiest way to help; see CONTRIBUTING.md.
Scope and intent
This tool is for hygiene, privacy and transparency on content you own: see what is hidden in your text, and publish without machine fingerprints you did not choose to include.
License
Metadata
Release files for watermark-cleaner 0.4.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| watermark_cleaner-0.4.0.tar.gz | 62.6 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| watermark_cleaner-0.4.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 107.5 kB
Release files / watermark_cleaner-0.4.0.tar.gz
| Download URL | watermark_cleaner-0.4.0.tar.gz |
|---|---|
| Size | 62.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
091352a91ef900eeb542ab2ced6492b75c86d31f09e3c7fd2c6776ad7d40697e
|
|
BLAKE2b-256 checksum How to use checksums |
97f23c293b11f869271b38f29bc5f943670c5870c027f711bdfd5a33a2c87924
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency logRelease files / watermark_cleaner-0.4.0-py3-none-any.whl
| Download URL | watermark_cleaner-0.4.0-py3-none-any.whl |
|---|---|
| Size | 44.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
42a716035764939558b42fd06c10fdb09895e93b94cb439c4dcd934c794890a6
|
|
BLAKE2b-256 checksum How to use checksums |
883dd6cdcc0110038c2f2533e85b566feafd1e4737f2f21136e1eb9cdea99419
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 5, 2026.
Transparency log