Skip to main content

Python Dependencies Tests

Features

  • Tree-based parsing into an editable document tree
  • CSS selectors: combinators, attributes, pseudo-classes, selector lists
  • Extractors: metadata, links, tables, text, images, forms, headings, scripts, stylesheets
  • Sanitization with a removal report
  • Minify, pretty-print, Markdown conversion
  • JSON / CSV / Markdown export
  • CLI tool
  • Standard library only at runtime

Installation

pip install webreader

Quickstart

import webreader

html = '<h1 class="title">Hello</h1><a href="/x">link</a>'
doc = webreader.parse(html)

webreader.select(doc, 'h1.title')   # [{'tag': ..., 'attrs': ..., 'text': ..., 'html': ...}]
doc.find('a').get_attr('href')       # '/x'

webreader.minify(html)              # collapsed whitespace
webreader.pretty(html)              # re-indented, readable markup
webreader.to_markdown(html)         # convert to Markdown

clean, report = webreader.sanitize(html)

Bytes input is decoded using detected encoding: BOM sniffing first, then <meta charset> in the first 2 KB, then the configured encoding fallback.

In non-strict mode, malformed HTML does not raise; the warnings are available on the document:

doc = webreader.parse('<div><p>x</p>')   # unclosed <div>
doc.errors                                 # ['unclosed tag(s) at end of input: <div>']

Extraction

webreader.extract_metadata(html)    # title, OG/Twitter, canonical, charset, JSON-LD
webreader.extract_links(html)       # links and asset references
webreader.extract_tables(html)      # matrices with colspan/rowspan expanded
webreader.extract_text(html)        # readable text without boilerplate
webreader.extract_images(html)      # src, alt, dimensions, srcset candidates
webreader.extract_forms(html)       # form fields with types and defaults
webreader.extract_headings(html)    # h1-h6 outline + word count + reading time
webreader.extract_scripts(html)     # inline and external script inventory
webreader.extract_stylesheets(html) # link[rel~=stylesheet] and inline styles

CSS selectors

Supported syntax: type selectors, *, #id, .class, attribute selectors ([attr], [attr=value] and the ^=, $=, *=, ~=, |= operators), the pseudo-classes :first-child, :last-child, :nth-child(an+b), :not(compound) and :contains(text), compounds such as div.content#main[href="x"], the combinators whitespace (descendant), > (child), + (adjacent sibling) and ~ (general sibling), and comma-separated selector lists.

webreader.select(doc, 'a[href^="https://"]')
webreader.select(doc, 'ul > li:nth-child(2n+1):not(.skip)')
webreader.select(doc, 'h2 + p, blockquote p:contains("note")')

Invalid or unsupported selectors raise SelectorSyntaxError.

Tree editing

Nodes support a mutation API, and the edited tree can be re-serialized:

doc = webreader.parse('<div><a href="/x">link</a></div>')
link = doc.find('a')
link.set_attr('rel', 'noopener').remove_attr('class')

new = webreader.Node('p')
new.children.append('added text')
doc.find('div').append_child(new)

link.unwrap()        # replace the <a> with its inner text
doc.to_html()        # '<div>link<p>added text</p></div>'

Node also provides remove() and replace_with(*nodes).

Markdown conversion

to_markdown renders headings, paragraphs, links, images, emphasis/strong/strikethrough, inline and fenced code (with language detection from class="language-..."), nested lists, blockquotes, horizontal rules and tables.

webreader.to_markdown('<h1>T</h1><p>Hi <b>bold</b></p>')
# '# T\n\nHi **bold**'

Sanitization

sanitize() returns (clean_html, report):

  • tags outside the whitelist are unwrapped, inner content kept
  • script, style, noscript, template are dropped entirely
  • attributes filtered through a safe-list, on* handlers removed
  • javascript:, vbscript: and non-image data: URLs blocked
  • the report lists removed tags/attributes, blocked URLs and detected threats

CLI

python -m webreader parse -f input.html
python -m webreader select -f input.html -s 'a[href]'
python -m webreader links -f input.html --base-url https://example.com
python -m webreader meta -f input.html
python -m webreader tables -f input.html --format csv
python -m webreader text -f input.html
python -m webreader images -f input.html
python -m webreader forms -f input.html
python -m webreader minify -f input.html --out min.html
python -m webreader pretty -f input.html --out pretty.html
python -m webreader md -f input.html --out page.md
python -m webreader sanitize -f input.html --out clean.html --allowed p,a,ul

Use -f - to read from stdin. Every output command accepts --out to write to a file instead of stdout.

Default settings

Key Default Purpose
strip_whitespace True collapse whitespace in text nodes
encoding utf-8 fallback when bytes input has no detectable encoding
max_depth 100 maximum element nesting
preserve_comments False keep comments in the tree
strict_mode False raise on malformed HTML
remove_comments True strip comments when minifying
allowed_tags p, br, b, i, a, ul, ol, li, div, span sanitize whitelist
max_file_size_mb 10 reject oversized input
cache_size 50 parse cache capacity

Exceptions

HTMLPaferError (base), ConfigLoadError, SanitizationError, SelectorSyntaxError.

Metadata

Release files for webreader 2.3.8

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for webreader 2.3.8
File Size Uploaded
webreader-2.3.8.tar.gz 30.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for webreader 2.3.8
File Interpreter ABI Platform
webreader-2.3.8-py3-none-any.whl Python 3 none any Details

Total release size: 60.7 kB

Release files / webreader-2.3.8.tar.gz

Download URL webreader-2.3.8.tar.gz
Size 30.5 kB
Tags Source
SHA-256 checksum
How to use checksums
975caf19329653d7f5e2d36e498349009d182f074017d468d17ad6346689d8bd
BLAKE2b-256 checksum
How to use checksums
9d76eafaab0dfd8158ab637ba9945dd83fc2e270d1421301ef63d19e1135aaf5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.0 CPython/3.13.7

Release files / webreader-2.3.8-py3-none-any.whl

Download URL webreader-2.3.8-py3-none-any.whl
Size 30.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
02496f5dfe6776e2edfc6a40e22e834ca38538f88e81b62501ecc6d89a2ae986
BLAKE2b-256 checksum
How to use checksums
abd902d42951edc66c59ed2977f8b2ad47be1957285d30bf5c91829d7136ba28
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.0.0 CPython/3.13.7

Release history Release notifications | RSS feed

This release

2.3.8 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page