Skip to main content

cringe-filter

Strip the AI tells out of prose: the em dashes, the "it's not X, it's Y", the bolded preambles and the rest. A linter, a scorer and a prompt builder, all measured rather than guessed: 570k words Sterling Baird wrote on GitHub, contrasted against 1.35M words Claude wrote in the same threads, plus his Discussions posts, LinkedIn posts and comments, direct-message counts, sent-mail counts, first-author manuscripts and tutorial pages. The rates and intervals are measured; the thresholds, severities and budgets are declared in code and listed in this file, and a corpus rebuild changes the rules without anyone editing them.

The bundled profile is his voice, so score and prompt pull toward how he writes. The lint tells are about Claude's prose and apply to anyone's.

No dependencies. Python 3.9 or later. The optional rewrite command needs the anthropic package.

Install

pip install cringe-filter
uvx cringe-filter contexts            # or run it without installing

# Drop SKILL.md into a project's .claude/skills/cringe-filter/ so a coding
# agent applies the filter to its own drafts (see below)

The corpus and the pipeline that compiles it into cringe_filter/data/profile.json live in a private repository, because the corpus includes mail and messages. Paths below under scripts/voice/ and docs/voice/ refer to that repository.

Contexts

Every command takes a context, because the same person writes very differently in each. cringe-filter contexts prints the table. Aliases are accepted (dm, docs, bug-report, pr, grant), and --url infers the context from where the text is going:

Context What it is His median Budget
github Reply in his own repos 27 words 110
discussion GitHub Discussions post 51 297
third-party Issue or comment in someone else's repo 30 126
email Sent mail 57 170
message DM, chat reply, comment on a post 22 65
linkedin LinkedIn post 54 283
tutorial Docs and teaching pages 357 1500
paper Scientific prose, his papers before 2023 359 per section 6000
proposal Grant and research proposals 359 per section 6000

Commands

lint: the cringe filter. Deterministic, zero tokens, exit code 1 on any error.

cringe-filter lint --context github draft.md
cringe-filter lint --url https://github.com/pytorch/pytorch/issues/1 draft.md
git diff --name-only | grep '\.md$' | xargs cringe-filter lint -c tutorial --quiet
# only what a revision added
cringe-filter lint -c paper --against draft-v1.tex draft-v2.tex

--against OLD reports the findings the new version has that the old one did not (same rule on the same text, wherever it moved). In the lab's repos, paragraphs an agent revised after a reviewer's comment came back with more em dashes (261 to 329) and semicolons (371 to 470) than they had, so a review round can reopen what the last one closed.

Each finding says what backs it. [15.2x, CI 9.7-23.9] means Claude used the pattern 15 times as often as Sterling did in that register, with the family-wise bootstrap interval. [preventive, no corpus support] means the pattern never separated the two authors there; those never rise above info. A rule the corpus contradicts (he uses "ensure", "leverage", "comprehensive" and "streamline" more than Claude does) is switched off, per register, from the data. Suppress a real false positive inline:

<!-- cringe-filter: disable-line em-dash -->
<!-- cringe-filter: disable-next-line bold-run, md-header -->
<!-- cringe-filter: disable-file -->

LaTeX is read as the prose a reader of the PDF sees (latex.py): the preamble, comments, math, tables, footnotes and \cite/\ref/\label keys are masked, --- counts as an em dash and -- as an en dash, and findings keep their line numbers. A .tex file given without a context is linted as paper, and rewrite checks that every citation key, reference key and inline math span survives. In LaTeX the suppression comment is % cringe-filter: disable-line em-dash.

score: the voice filter as a measurement. Each feature is scored by the Poisson log-likelihood ratio of Claude's rate against his rate in the chosen context, and the sum is one style score with the features that produced it. The features overlap and the rates were measured on the same corpus, so the score is a ranking of what to fix, not a calibrated probability; the JSON carries calibrated: false and the review in docs/voice/research/ says what a calibrated version needs.

cringe-filter score --context email draft.txt
cringe-filter score -c github --text "Not sure this is right. Could you check?"
context: github (GitHub reply in his own repos)   words: 143   sentences: 7, longest 31
style score: +3.41 log-odds, reads like Claude (a ranking of what to fix, not a calibrated probability)
length: 143 words; his median here is 27, p90 110, budget 110  [over budget]
Burrows' Delta over the 150 most frequent words: 1.31 to his register, 1.12 to Claude (closer to Claude)
feature                       n  yours/1k   his/1k  Claude/1k  log-odds
bold runs                     4     28.0     2.93      21.47     +2.13
hedges (might, maybe, ...)    0      0.0    10.58       0.33     +1.46

The Delta line is Burrows' Delta over the 150 most frequent words of the pooled corpus (mostly function words, by construction, not a curated function-word list), z-scored against the per-document spread: the draft's distance to the register's centroid and to Claude's. Function words carry authorship in the stylometry literature, and this is the standard implementation of that idea. It is a second signal, though not an independent one: its centroids come from the same corpus as the tells.

The structure rows (marked per 1000 sentences or paragraphs) measure how the sentences are built rather than which words they use: how often the subject is I, we or you, how often a sentence opens on "The", a number, a conjunction or a lead-in like "Also,", modal and "to" verbs per sentence, colons, dashes, semicolons and parentheses inside a sentence, and the share of one-sentence paragraphs. structure.py counts them without a parser; each was checked against a dependency parse of the corpus (docs/voice/research/structure.md). Registers whose text never reaches the build (email, tutorials, manuscripts) have no structure rows.

It is not a detector. It says where a draft sits between two measured writers and which features put it there, which is what you need to fix it.

prompt: the voice filter as a prompt, for any model. Two real passages first, matched to the target length and free of anything the rules forbid; then five plain constraints derived from the register's measurements; then the measured tells in plain words; then a few notes from the register card. No rates or intervals reach the model, because numeric constraints are what models follow worst; --evidence appends them for a person.

cringe-filter prompt --context linkedin draft.txt        # full prompt
cringe-filter prompt --context linkedin --system-only    # the filter alone
cringe-filter prompt -c github --evidence draft.md       # plus the numbers
cringe-filter prompt -c github --json draft.md           # {system, user, evidence}
cringe-filter prompt -c github --structure draft.md      # with the sentence-structure priority

The prompt carries one priority written from the register's structure rates: make a person the subject, open few sentences on "The" or a number, carry plans in verbs, ask where the draft leaves a choice open, one idea per sentence, one point per paragraph. Each clause appears only where the register's own rates put it on his side by a clear margin. It is on by default: on 36 fresh held-out comments it moved every grammar judge toward him, at a cost of about 1.5 points of the source's content words. --no-structure leaves it out, and build_prompt(..., structure=1) rebuilds the wording that test used (docs/voice/research/structure.md).

rewrite: apply the prompt with Claude, lint the result, and make one revision pass against lint errors and measured warnings (bold, headers, tables and arrows are warnings, and they are the dominant Claude features), and against any number, link, code span or path the rewrite dropped (preserve.py). The length budget is reported but does not trigger a revision by itself: in the held-out rewrite experiment every length-driven revision cut about 6% of the source's content words without moving any classifier toward him. Never more than two model calls: repeated self-revision degrades text that was already fine.

pip install 'cringe-filter[rewrite]'      # or: pip install anthropic
export ANTHROPIC_API_KEY=...
cringe-filter rewrite --context email draft.txt
cringe-filter rewrite -c github --model claude-sonnet-5 --effort low draft.md
cringe-filter rewrite -c github --dry-run draft.md      # print the prompt
cringe-filter rewrite -c paper --minimal draft.tex      # edit, do not rewrite

--minimal (also on prompt) is for a draft headed to review. It asks for the smallest edit a reviewer would make: a nine-line checklist of what the lab's reviewers asked agents to change (claim only what was done, plain words, defined abbreviations, concrete detail, no process talk, no em dashes or "X, not Y"), the draft's own lint findings, and no revision pass. On 74 agent drafts that people later corrected, the full rewrite moved the text away from the version the reviewer wrote, and the minimal edit did not (docs/voice/research/corrections.md).

Defaults to claude-opus-5 with adaptive thinking at medium effort, and requests the server-side refusal fallback so a declined request is re-run on a fallback model inside the same call (--no-fallback to turn that off). Output goes to stdout; the pass count, before-and-after score and any remaining findings go to stderr.

audit turns the lint findings into an edit spec: each sentence that needs work, its line, and why, plus the flagged sentences to leave alone. Hand the spec and the draft to any frontier model with prompt --minimal. The "X, not Y" family and the "X is what did Y" cleft are judgment calls (his own uses are mostly instructions), so --model puts each one to a model as a single question: score from 0 to 100 how likely it is that Y was set up only to be knocked down. Hits under --threshold (default 30) move to "leave as written". Any OpenAI-compatible endpoint works: Ollama on localhost by default, so a private draft stays on the machine, or llama.cpp's llama-server, LM Studio or vLLM through --endpoint. A frontier model asked for this score separated Claude's contrasts from his; the 1.5B and 3B models that fit on two CPU cores did worse (docs/voice/research/local-models.md). Without --model nothing is dropped and the editor decides.

cringe-filter audit -c paper draft.tex > spec.md
ollama pull qwen2.5:14b
cringe-filter audit -c github --model qwen2.5:14b draft.md

instructions <context>... writes path-scoped instruction files for GitHub Copilot, one per context, into .github/instructions/ (--dir to change it). Copilot's cloud agent and its code review apply each file to the paths its applyTo globs match: **/*.tex,**/*.bib for paper, folders named for a proposal or grant for proposal, Markdown, RST and notebooks for tutorial. --apply-to replaces the globs and --exclude-agent code-review keeps review from reading a file. Replies are not files, so for github and the other reply contexts put prompt -c <context> --system-only in .github/copilot-instructions.md.

cringe-filter instructions paper proposal tutorial
cringe-filter instructions paper --apply-to "manuscript/**"

contexts and profile -c <context> print what the bundle knows.

Inside a coding agent

An agent that drafts prose is itself the model, so it does not need rewrite. SKILL.md in this folder is a Claude Code skill that runs lint and prompt and has the agent apply the filter to its own draft. Copy it to .claude/skills/cringe-filter/SKILL.md in any repository that has this package on its path.

What is in the bundle

profile.json carries, per context: document and sentence length statistics, per-1000-word rates for 31 surface markers, the register's own rate for each of 104 candidate tells, the Claude-over-register ratio for each with its family-wise interval, the register card, and which exemplar banks to read. Plus Claude's reference rates, the corpus-wide verdicts, and the Holm-surviving bigrams on each side, limited in the published build to function-word pairs so project vocabulary stays out. The exemplar banks are real passages from public repositories and public LinkedIn posts only; sent mail and direct messages never leave the machine that harvested them, and only their counts are here.

The private repository rebuilds it after the pipeline runs (scripts/voice/build_voicekit.py) and opens a pull request here with the new profile. Tests:

python -m unittest discover -s tests

A profile can also live outside the package. Point CRINGE_FILTER_PROFILE at it, and an exemplars/ folder beside it replaces the packaged banks of the same name:

export CRINGE_FILTER_PROFILE=~/private/voice/profile.json   # + ~/private/voice/exemplars/*.md

python scripts/voice/privacy_audit.py reports what the package would reveal if it were published: whether every quoted passage is public, any private repository name, email or phone number, which numbers come from private text, and which word lists carry project vocabulary.

Limits

  • The reference for most contexts is Claude's GitHub prose, because that is the only place both authors wrote about the same work. Scoring an email against it asks "does this read like the model or like his email", which is the useful question, but the two are not the same genre.
  • paper is scored against Claude's own scientific prose instead: the manuscripts claude[bot] wrote in the lab's repos (references in the profile). With one of his papers and one of Claude's repositories held out at a time, score -c paper reads 253 of 291 chunks of his earlier papers as his and 17 as Claude's, and 66 of 79 chunks of Claude's manuscripts as Claude's and 3 as his. Before, it read every one of Claude's drafts of his sections as his.
  • proposal is scored against the coding agents' proposals in the lab's public and private repos, 15k words, most of them the Copilot agent running Claude Opus models. Semicolons, mid-sentence colons, significance words and "we" separate them from his proposals. Held out, it reads 32 of 40 chunks of his as his and 38 of 49 of the agents' as Claude's, against 33 and 34 with Claude's manuscripts as the reference. His side rests on two proposals, 13k words, so treat the score as a hint and the card as the rules (docs/voice/research/technical-writing.md).
  • Small registers carry little signal for rare phrases. LinkedIn is 89 posts; the linter falls back to the corpus-wide verdict when a register has none of its own for a phrase.
  • The scorer treats features as independent, which they are not; read the log-odds as a ranking of what to fix, not as a probability with guarantees.
  • Detector evasion is not the objective, and a low score is not proof of anything. The one-question test still applies: does this sound pretentious to someone who already knows what they are doing?

Metadata

Release files for cringe-filter 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for cringe-filter 0.1.0
File Size Uploaded
cringe_filter-0.1.0.tar.gz 147.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for cringe-filter 0.1.0
File Interpreter ABI Platform
cringe_filter-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 296.5 kB

Release files / cringe_filter-0.1.0.tar.gz

Download URL cringe_filter-0.1.0.tar.gz
Size 147.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a1b0ff1a3734f6cff1f10ce1bc2edcdea733ec04f306e7a749cc9795c3d313be
BLAKE2b-256 checksum
How to use checksums
d4167b18d4acc10967331e5f1dbe6678e67019cb4d4805c13f74d4d4d26c86a5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release files / cringe_filter-0.1.0-py3-none-any.whl

Download URL cringe_filter-0.1.0-py3-none-any.whl
Size 149.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
32ab051906390cca6f33b2a633d17e3a9e477f8e93cd550e40f6bf89dae6c9b9
BLAKE2b-256 checksum
How to use checksums
21b62645802bfe86bf88565f8996431543dae0ff30851098e57ee193e4c838c0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 26, 2026.

Transparency log

Release history Release notifications | RSS feed

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page