Skip to main content

Remove boilerplate from scraped markdown before it reaches an LLM. Subtractive-only, auditable, zero dependencies.

Project description

Winnow

PyPI Python License: MIT

Remove boilerplate from scraped markdown before it reaches an LLM.

Scraping and extract APIs like Tavily hand you "LLM-ready" markdown that still carries navigation, footers, cookie banners, promos, and link farms — 10–90% of the tokens depending on the page. The same goes for your own scraper or loader pipeline: if it produces markdown (or HTML converted to markdown), Winnow slots in right after it — any markdown in, leaner markdown out, with a receipt for every removed block.

  • Subtractive-only. Winnow deletes blocks; it never rewrites a word. Zero hallucination risk by construction.
  • Recall-first. Dropping real content is data loss; keeping boilerplate just costs tokens. When uncertain, Winnow keeps.
  • Auditable. Every removed block comes back with a reason code and score.
  • Template memory. Feed Winnow multiple pages from one site and it learns the site's template — blocks repeated across pages are boilerplate, near-certainly.
  • Zero runtime dependencies. The core install adds nothing to your tree.

Install

pip install winnow-md

Optional extras:

pip install "winnow-md[model]"    # learned block-sequence scorer (numpy + model2vec)
pip install "winnow-md[tokens]"   # exact token counts via tiktoken
pip install "winnow-md[model,tokens]"

Python 3.9+. The package installs as winnow-md and imports as winnow. The core is pure Python with no dependencies; if the [model] extra isn't installed, the learned scorer is silently skipped and the heuristics run alone.

Quickstart

import winnow

# One page
res = winnow.clean(markdown_text)
print(res.markdown)          # cleaned markdown
print(res.stats)             # tokens before/after, reduction %
for r in res.removed:        # the receipt
    print(r.reasons, r.text[:60])

# A crawl — template memory kicks in across pages of the same domain
w = winnow.Winnow(aggressiveness=0.5)
results = w.clean_many(pages)          # list of markdown strings (or (md, url) tuples)

# Streaming with a persistent per-domain template store
w = winnow.Winnow(store="winnow.db")
res = w.clean(md, url="https://example.com/post/1")
winnow clean page.md                      # cleaned markdown to stdout
winnow clean ./crawl/ --report out.html   # batch + filterable HTML audit report
winnow clean ./crawl/ --url-mode strip    # also strip URL bodies from kept links

Jina Reader output (Title: / URL Source: preamble) is auto-detected and the source URL is used for template memory.

Benchmark

Five generations of independently-labeled, adversarially-arbitrated exam batches (each fetched fresh, dual-labeled blind, disputes refereed) — ~12,000 hand-adjudicated blocks across 21 domains:

Exam batch Content recall Junk recall Token cut
batch 5 (newest, still converging) 0.960 0.61 −42%
batch 4 0.979 0.61 −49%
batch 3 0.992 0.57 −35%
batch 2 0.997 0.58 −29%
batch 1 1.000 0.69 −41%

Content recall is the fraction of real content kept — the number that must never slip. Junk recall is the fraction of boilerplate actually removed; what it misses costs tokens, never correctness. Each batch was a fresh exam nothing had been tuned on when first scored, then became training data — the newest batch is always the honest one.

Add --url-mode strip for roughly 15 additional points of token cut with zero text loss. With the [model] extra installed, a learned block-sequence scorer raises junk recall further; it is capped so that it can never delete a block on its own.

The benchmark harness, labeling pipeline, and mutation self-test live in bench/ — see ARCHITECTURE.md.

Contributing

Contributions are welcome — bug fixes, new signals, docs, tests, or ideas. One especially easy and useful report: a page Winnow handles badly, since most improvements so far came from being shown a real page it got wrong. See CONTRIBUTING.md for setup, the three design rules any change must respect, and how to run the benchmark gates before opening a PR.

License

MIT — see LICENSE.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

winnow_md-0.1.2.tar.gz (362.8 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

winnow_md-0.1.2-py3-none-any.whl (355.5 kB view details)

Uploaded Python 3

File details

Details for the file winnow_md-0.1.2.tar.gz.

File metadata

  • Download URL: winnow_md-0.1.2.tar.gz
  • Upload date:
  • Size: 362.8 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for winnow_md-0.1.2.tar.gz
Algorithm Hash digest
SHA256 fbe18ce5453e961eb5bb03161cf15705f8bae89d3dbcba5c60348c58cab8ab7b
MD5 0944e81543c22582067c21b4cf3cf803
BLAKE2b-256 301c8c083e44f0becb25cdfcf567224ebfca0c4e645a31d1143db7399c31ec2c

See more details on using hashes here.

File details

Details for the file winnow_md-0.1.2-py3-none-any.whl.

File metadata

  • Download URL: winnow_md-0.1.2-py3-none-any.whl
  • Upload date:
  • Size: 355.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: uv/0.11.18 {"installer":{"name":"uv","version":"0.11.18","subcommand":["publish"]},"python":null,"implementation":{"name":null,"version":null},"distro":{"name":"Ubuntu","version":"24.04","id":"noble","libc":null},"system":{"name":null,"release":null},"cpu":null,"openssl_version":null,"setuptools_version":null,"rustc_version":null,"ci":null}

File hashes

Hashes for winnow_md-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 063b4a0f54f1819fe3c322ddfd2bce17123b56d2f6152937864c355c60aa78f5
MD5 3d0efe20ec5047e42960f1331f2ade9a
BLAKE2b-256 897229672502ed49126740318bf85dedd24de02404e340f6b39cf1aa3a03eb7b

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page