Remove boilerplate from scraped markdown before it reaches an LLM. Subtractive-only, auditable, zero dependencies.
Project description
Winnow
Remove boilerplate from scraped markdown before it reaches an LLM.
Scraping and extract APIs like Tavily hand you "LLM-ready" markdown that still carries navigation, footers, cookie banners, promos, and link farms — 10–90% of the tokens depending on the page. The same goes for your own scraper or loader pipeline: if it produces markdown (or HTML converted to markdown), Winnow slots in right after it — any markdown in, leaner markdown out, with a receipt for every removed block.
- Subtractive-only. Winnow deletes blocks; it never rewrites a word. Zero hallucination risk by construction.
- Recall-first. Dropping real content is data loss; keeping boilerplate just costs tokens. When uncertain, Winnow keeps.
- Auditable. Every removed block comes back with a reason code and score.
- Template memory. Feed Winnow multiple pages from one site and it learns the site's template — blocks repeated across pages are boilerplate, near-certainly.
- Zero runtime dependencies. The core install adds nothing to your tree.
Install
pip install winnow-md
Optional extras:
pip install "winnow-md[model]" # learned block-sequence scorer (numpy + model2vec)
pip install "winnow-md[tokens]" # exact token counts via tiktoken
pip install "winnow-md[model,tokens]"
Python 3.9+. The package installs as winnow-md and imports as winnow. The
core is pure Python with no dependencies; if the [model] extra isn't
installed, the learned scorer is silently skipped and the heuristics run alone.
Quickstart
import winnow
# One page
res = winnow.clean(markdown_text)
print(res.markdown) # cleaned markdown
print(res.stats) # tokens before/after, reduction %
for r in res.removed: # the receipt
print(r.reasons, r.text[:60])
# A crawl — template memory kicks in across pages of the same domain
w = winnow.Winnow(aggressiveness=0.5)
results = w.clean_many(pages) # list of markdown strings (or (md, url) tuples)
# Streaming with a persistent per-domain template store
w = winnow.Winnow(store="winnow.db")
res = w.clean(md, url="https://example.com/post/1")
winnow clean page.md # cleaned markdown to stdout
winnow clean ./crawl/ --report out.html # batch + filterable HTML audit report
winnow clean ./crawl/ --url-mode strip # also strip URL bodies from kept links
Jina Reader output (Title: / URL Source: preamble) is auto-detected and the
source URL is used for template memory.
Benchmark
Five generations of independently-labeled, adversarially-arbitrated exam batches (each fetched fresh, dual-labeled blind, disputes refereed) — ~12,000 hand-adjudicated blocks across 21 domains:
| Exam batch | Content recall | Junk recall | Token cut |
|---|---|---|---|
| batch 5 (newest, still converging) | 0.960 | 0.61 | −42% |
| batch 4 | 0.979 | 0.61 | −49% |
| batch 3 | 0.992 | 0.57 | −35% |
| batch 2 | 0.997 | 0.58 | −29% |
| batch 1 | 1.000 | 0.69 | −41% |
Content recall is the fraction of real content kept — the number that must never slip. Junk recall is the fraction of boilerplate actually removed; what it misses costs tokens, never correctness. Each batch was a fresh exam nothing had been tuned on when first scored, then became training data — the newest batch is always the honest one.
Add --url-mode strip for roughly 15 additional points of token cut with zero
text loss. With the [model] extra installed, a learned block-sequence scorer
raises junk recall further; it is capped so that it can never delete a block
on its own.
The benchmark harness, labeling pipeline, and mutation self-test live in
bench/ — see ARCHITECTURE.md.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file winnow_md-0.1.1.tar.gz.
File metadata
- Download URL: winnow_md-0.1.1.tar.gz
- Upload date:
- Size: 361.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
72d1d68fdee4ebf91101f62f41600e0aed316c175b597fc25a848b9ee3d721de
|
|
| MD5 |
e55a98b5012c37ac54375f0e92ec2961
|
|
| BLAKE2b-256 |
cdcf8e2246c98ad6953d36606040f6f47a8f67580592e762fdba3e910a2162ad
|
File details
Details for the file winnow_md-0.1.1-py3-none-any.whl.
File metadata
- Download URL: winnow_md-0.1.1-py3-none-any.whl
- Upload date:
- Size: 354.5 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
e22c1ead4acf76d79407b6c3af2ea27613a1028f976bd067cc774716b65ea2f1
|
|
| MD5 |
7d7c720b0e816523efe25607be6f5f48
|
|
| BLAKE2b-256 |
9aa30e16a0c2ae9a7ed6f1ad30d530405979cab07f01bd75ef0260ec0e898789
|