styleprofile
Compare a draft with a writer's usual style, then see which habits make it different. With LLM drafts of the same briefs as a contrast set, also measure resemblance to those drafts. Scores describe this comparison; they do not prove who wrote a text.
Install
Python 3.11+; CI tests 3.11–3.14. While a PyPI release is unavailable, use the verified GitHub installation below. Choose pip in an activated virtual environment, or uv:
python -m pip install "styleprofile[syntax] @ git+https://github.com/wangjohn/styleprofile.git"
styleprofile setup
uv tool install --python 3.11 "styleprofile[syntax] @ git+https://github.com/wangjohn/styleprofile.git"
styleprofile setup
setup installs spaCy's English model once. Make sure uv's tool executable directory is
on your PATH. For virtual-environment activation, Windows, and a standard-library-only
install, see installation details.
Try it in 60 seconds
styleprofile demo
Expected verdict: Overall: close. The command copies bundled, original samples into
styleprofile-demo/, builds a profile and scores a draft; it needs no checkout. The draft
contains two off-voice paragraphs that its whole-document verdict can hide. The samples
are deliberately small, so the command also warns about a thin reference. See the
sample guide for the experimental paragraph demonstration.
Your own texts
Use one genre per profile. In these recipes, replace the example paths with your own:
book.md is a long manuscript, posts/ holds independent essays, draft.md and drafts/
are new writing. Inputs are UTF-8 Markdown, text, HTML or JSONL. Profiles are summaries,
not copies of the text; keep both your source texts and generated reports out of version control.
One long file. Build from a book or archive, then score a new draft. The file is split at suitable chapter headings or issue boundaries; automatic stand-in sections are less reliable when it has no such divisions.
styleprofile build book.md
styleprofile score draft.md book.profile.json
A folder of essays or posts. Build once, then reuse the profile. Without -o, the
output is <folder-name>.profile.json in the current directory.
styleprofile build posts/
styleprofile score drafts/ posts.profile.json
Tweets, comments or emails. Use one JSON object per line with a text field and a thread or conversation ID. Group independent conversations and pool short records:
styleprofile build comments.jsonl --text-field text --group-field thread --pool
styleprofile score new-comments.jsonl comments.profile.json --pool
Records in different groups are never pooled together. Without --pool at score time,
each record is judged separately, and a record under 75 words gets no verdict. See
grouping and pooling for data requirements and pitfalls.
One-shot checks. Build the reference in memory and score immediately:
styleprofile score draft.md --against posts/
LLM-likeness. First make model drafts from the writer's own briefs, keeping topics,
genre and lengths similar and using several models. Save those drafts in llm-drafts/:
styleprofile build posts/ --contrast llm-drafts/ -o writer.json
styleprofile score draft.md writer.json
The contrast is optional: without it, you get Delta but no LLM-likeness. The weights know only the drafts they learned from. The contrast recipe includes a copyable prompt and advice on counts and independent held-out checks.
How much text you need
| Feature | Minimum |
|---|---|
| A useful Delta reference | 15+ windows, 20,000+ words, from independent documents |
| Calibrated verdicts on short texts | 3+ documents in the reference |
| Paragraph drift checks | 10+ documents, about 200 paragraphs; build with --by-paragraph |
| A scorable text | 75 prose words, and a reference calibrated at that length |
The first row is a recommended size, not a hard cutoff: the small demo still returns a verdict with a warning. Three documents alone do not guarantee a sound reference. Aim for 20 or more substantial documents for paragraph checks; check whether the profile actually has thresholds. Short drafts get wider ranges, and texts below the shortest calibrated length can remain “too short to judge” even above 75 words.
Reading the output
Delta measures distance from the writer's reference in the writer's own standard deviations, with each area (sentence shape, vocabulary, punctuation, and so on) counting equally. Lower is closer. The verdict compares it with the writer's own held-out range.
| Delta | Meaning |
|---|---|
| close | within the writer's usual held-out range |
| somewhat different | up to 1.5 times that range |
| clearly different | up to twice that range |
| very different | further still |
LLM-likeness measures deviations in the directions the contrast drafts differ, weighted by how well those habits separate the two sets. It appears only with a contrast set.
| LLM-likeness | Meaning |
|---|---|
| like the reference | within the writer's own range |
| a few LLM traits | above it, but less than halfway to the LLM drafts |
| leans LLM | more than halfway to the LLM drafts |
| like the LLM drafts | at or beyond the drafts' typical score |
Both scores use the reference range at the draft's length. Under 75 words there is no verdict; numbers and traits are indicative only. Chunks too short to judge are left out of the headline. “By area” divides each area's Delta by the top of its own held-out range: close up to 1x, somewhat different to 1.5x, clearly different to 2x, very different above. Bars use a log scale, full at 32x; biggest-difference arrows show direction, one per standard deviation.
Human output uses short notes by default. --verbose on build, score or evaluate restores
full explanations; --all shows every metric. JSON always keeps full warnings. Use
styleprofile show writer.json for a saved report and styleprofile metrics for metric definitions.
The command reference covers formats,
splitting, cache, worker processes and evaluate; Method
explains the mathematics, reliability checks and resolution floors.
What a verdict can't tell you
A whole-document average can hide one or two paragraphs in another voice: the demo draft reads “close” despite its two LLM-style paragraphs. English is the supported language. A human can drift from their reference, and a model can be prompted toward it; resemblance is not proof of authorship.
Paragraph checks are experimental and off by default. To try them, build with
--by-paragraph, then score with the same flag. On synthetic writer texts with new topics,
a paragraph falsely drifted in up to about a third of documents. These are historical
measurements, not a promise about your writing:
| Writer heldouts with false paragraph drift | A | B | C | D |
|---|---|---|---|---|
| single documents, without / with spaCy | 4% / 2% | 0% / 6% | 3% / 0% | 0% / 33% |
| 4–10 documents joined, without / with spaCy | 0% / 0% | 0% / 0% | 3% / 0% | 0% / 30% |
A uses reference topics; B–D use unseen topics, D the widest shift. Runs of two or three LLM blocks were found 76–100% of the time; single blocks 57–74%. Short paragraphs under 30 words were often missed (17–48% found), as was one paragraph in a long document. “No paragraph drifts” does not mean a text is clean. The synthetic corpora share sample material and are optimistic; expect more uncertainty on real new topics. Treat a drift as a lead to read. See the full research limits.
Troubleshooting
Match the message below; identifiers name existing library StyleProfileError.code or
Note.code values. Pip and uv installation errors have no styleprofile code.
| What happened | What to do |
|---|---|
| “No matching distribution” | Use Python 3.11+ and the GitHub install above while the PyPI release is pending. |
| “Externally managed environment” (PEP 668) | Use an activated virtual environment or uv tool install; don't replace system Python packages. |
spaCy/model versions differ (code=setup_spacy_version, code=syntax_model_mismatch) |
Run styleprofile setup in the same environment, then rebuild the profile. If spaCy itself is outside the required range, reinstall the syntax extra there. |
Syntax metrics left out (code=no_syntax, code=syntax_unavailable) |
Install the syntax extra and run styleprofile setup, or choose --no-syntax. Scoring a surface-only reference stays surface-only; rebuild to add syntax. |
Cache unavailable (code=cache_unavailable) |
Check styleprofile cache; choose a writable absolute XDG_CACHE_HOME, or use --no-cache. The run continues without the cache. |
“Rebuild your profile” (code=outdated) |
Rebuild a reference from its original texts, rescore a saved score, or reevaluate an evaluation. Different report majors are incompatible. |
Newer minor report (code=newer_report_version) |
Matching majors still load; upgrade if you need the added fields. Missing paragraph calibration requires rebuilding with --by-paragraph. |
| Windows console symbols look different | Symbols fall back to ASCII when the stream cannot encode them. UTF-8 inputs are still required. Use a UTF-8 terminal or python -X utf8 -m styleprofile demo for redirected output. |
No usable prose (code=no_chunks) |
Convert unsupported formats and check that the input contains prose, not only code or lists. |
CLI build/evaluate caching is on by default and stores counts that can reveal wording; treat the cache like your texts. Library calls leave it off unless you opt in. See cache and workers, report compatibility, and the changelog for details.
Use it in CI
Start by reading ordinary scores before setting a failure threshold. Each document is
judged separately; one drifting document can fail a batch. --fail-above clearly checks
whole-document Delta and --fail-likeness leans checks resemblance to the supplied contrast:
styleprofile score -q --fail-above clearly --fail-likeness leans drafts/ writer.json
A threshold hit returns 3, which is an expected gate failure. Whole-document scores
can mask a few off-voice chunks; --fail-flagged N catches N chunks that individually read
clearly different or lean LLM. These occur by chance too: about 0–0.14% of writer chunks
in the measured corpora, with at least one in 15–27% of runs of 200 chunks. Choose N for
the document length; 1 suits a few windows, while long documents may need 2 or more.
For pre-commit, save this as .pre-commit-config.yaml. The installed
command and writer.json must be available in the repository:
repos:
- repo: local
hooks:
- id: styleprofile
name: styleprofile
entry: styleprofile score -q --fail-above clearly --fail-flagged 1 -r writer.json
language: system
types: [markdown]
require_serial: true
| Exit status | Meaning |
|---|---|
| 0 | scored; no document reached a requested fail level |
| 1 | an error, such as a missing file or unreadable profile |
| 2 | invalid command-line usage |
| 3 | a document reached a fail level, or could not be compared when a fail flag was given |
A text too short to judge never fails a run. An incomparable document fails when a fail
flag is given. JSON records requested thresholds under fail and failures under failed;
see automation details.
Using it from Python
import styleprofile as sp
profile = sp.build("posts/", contrast="llm-drafts/")
result = profile.score_text("A draft to check against the writer.")
print(result.verdict, result.delta)
The tiny placeholder draft abstains; supply at least 75 prose words for a verdict. For the string-first API, settings, saving/loading and notes, see the library guide.
Development, benchmark gates, snapshots and release checks are in CONTRIBUTING. The package began as the stylometry module of GoodProse.
License
MIT
Metadata
Release files for styleprofile 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| styleprofile-0.2.0.tar.gz | 476.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| styleprofile-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 750.7 kB
Release files / styleprofile-0.2.0.tar.gz
| Download URL | styleprofile-0.2.0.tar.gz |
|---|---|
| Size | 476.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
64918d815c992e90fa2c1b79c9ccbe303335a4ebd470b7e618ed0c3c097b12dc
|
|
BLAKE2b-256 checksum How to use checksums |
72d7001aa25d1abd90bbe2a32384354266b221fe492375628d7d04b3e2524224
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency logRelease files / styleprofile-0.2.0-py3-none-any.whl
| Download URL | styleprofile-0.2.0-py3-none-any.whl |
|---|---|
| Size | 274.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f0b0bf651faf736c7c9b75ce8a0a6160788f0d46b9015b706553d793b941f125
|
|
BLAKE2b-256 checksum How to use checksums |
094a5bc2d4989cf7da9cd4a913533dba125ca02c1f59400be8bcf7ed7e489cd0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 9, 2026.
Transparency log