Features
- Tree-based parsing into an editable document tree
- CSS selectors: combinators, attributes, pseudo-classes, selector lists
- Extractors: metadata, links, tables, text, images, forms, headings, scripts, stylesheets
- Sanitization with a removal report
- Minify, pretty-print, Markdown conversion
- JSON / CSV / Markdown export
- CLI tool
- Standard library only at runtime
Installation
pip install webreader
Quickstart
import webreader
html = '<h1 class="title">Hello</h1><a href="/x">link</a>'
doc = webreader.parse(html)
webreader.select(doc, 'h1.title') # [{'tag': ..., 'attrs': ..., 'text': ..., 'html': ...}]
doc.find('a').get_attr('href') # '/x'
webreader.minify(html) # collapsed whitespace
webreader.pretty(html) # re-indented, readable markup
webreader.to_markdown(html) # convert to Markdown
clean, report = webreader.sanitize(html)
Bytes input is decoded using detected encoding: BOM sniffing first,
then <meta charset> in the first 2 KB, then the configured
encoding fallback.
In non-strict mode, malformed HTML does not raise; the warnings are available on the document:
doc = webreader.parse('<div><p>x</p>') # unclosed <div>
doc.errors # ['unclosed tag(s) at end of input: <div>']
Extraction
webreader.extract_metadata(html) # title, OG/Twitter, canonical, charset, JSON-LD
webreader.extract_links(html) # links and asset references
webreader.extract_tables(html) # matrices with colspan/rowspan expanded
webreader.extract_text(html) # readable text without boilerplate
webreader.extract_images(html) # src, alt, dimensions, srcset candidates
webreader.extract_forms(html) # form fields with types and defaults
webreader.extract_headings(html) # h1-h6 outline + word count + reading time
webreader.extract_scripts(html) # inline and external script inventory
webreader.extract_stylesheets(html) # link[rel~=stylesheet] and inline styles
CSS selectors
Supported syntax: type selectors, *, #id, .class, attribute
selectors ([attr], [attr=value] and the ^=, $=, *=, ~=,
|= operators), the pseudo-classes :first-child, :last-child,
:nth-child(an+b), :not(compound) and :contains(text), compounds
such as div.content#main[href="x"], the combinators whitespace
(descendant), > (child), + (adjacent sibling) and ~ (general
sibling), and comma-separated selector lists.
webreader.select(doc, 'a[href^="https://"]')
webreader.select(doc, 'ul > li:nth-child(2n+1):not(.skip)')
webreader.select(doc, 'h2 + p, blockquote p:contains("note")')
Invalid or unsupported selectors raise SelectorSyntaxError.
Tree editing
Nodes support a mutation API, and the edited tree can be re-serialized:
doc = webreader.parse('<div><a href="/x">link</a></div>')
link = doc.find('a')
link.set_attr('rel', 'noopener').remove_attr('class')
new = webreader.Node('p')
new.children.append('added text')
doc.find('div').append_child(new)
link.unwrap() # replace the <a> with its inner text
doc.to_html() # '<div>link<p>added text</p></div>'
Node also provides remove() and replace_with(*nodes).
Markdown conversion
to_markdown renders headings, paragraphs, links, images,
emphasis/strong/strikethrough, inline and fenced code (with language
detection from class="language-..."), nested lists, blockquotes,
horizontal rules and tables.
webreader.to_markdown('<h1>T</h1><p>Hi <b>bold</b></p>')
# '# T\n\nHi **bold**'
Sanitization
sanitize() returns (clean_html, report):
- tags outside the whitelist are unwrapped, inner content kept
script,style,noscript,templateare dropped entirely- attributes filtered through a safe-list,
on*handlers removed javascript:,vbscript:and non-imagedata:URLs blocked- the report lists removed tags/attributes, blocked URLs and detected threats
CLI
python -m webreader parse -f input.html
python -m webreader select -f input.html -s 'a[href]'
python -m webreader links -f input.html --base-url https://example.com
python -m webreader meta -f input.html
python -m webreader tables -f input.html --format csv
python -m webreader text -f input.html
python -m webreader images -f input.html
python -m webreader forms -f input.html
python -m webreader minify -f input.html --out min.html
python -m webreader pretty -f input.html --out pretty.html
python -m webreader md -f input.html --out page.md
python -m webreader sanitize -f input.html --out clean.html --allowed p,a,ul
Use -f - to read from stdin. Every output command accepts --out
to write to a file instead of stdout.
Default settings
| Key | Default | Purpose |
|---|---|---|
strip_whitespace |
True |
collapse whitespace in text nodes |
encoding |
utf-8 |
fallback when bytes input has no detectable encoding |
max_depth |
100 |
maximum element nesting |
preserve_comments |
False |
keep comments in the tree |
strict_mode |
False |
raise on malformed HTML |
remove_comments |
True |
strip comments when minifying |
allowed_tags |
p, br, b, i, a, ul, ol, li, div, span |
sanitize whitelist |
max_file_size_mb |
10 |
reject oversized input |
cache_size |
50 |
parse cache capacity |
Exceptions
HTMLPaferError (base), ConfigLoadError, SanitizationError,
SelectorSyntaxError.
Metadata
Release files for webreader 2.3.8
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| webreader-2.3.8.tar.gz | 30.5 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| webreader-2.3.8-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 60.7 kB
Release files / webreader-2.3.8.tar.gz
| Download URL | webreader-2.3.8.tar.gz |
|---|---|
| Size | 30.5 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
975caf19329653d7f5e2d36e498349009d182f074017d468d17ad6346689d8bd
|
|
BLAKE2b-256 checksum How to use checksums |
9d76eafaab0dfd8158ab637ba9945dd83fc2e270d1421301ef63d19e1135aaf5
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.0 CPython/3.13.7
|
Release files / webreader-2.3.8-py3-none-any.whl
| Download URL | webreader-2.3.8-py3-none-any.whl |
|---|---|
| Size | 30.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
02496f5dfe6776e2edfc6a40e22e834ca38538f88e81b62501ecc6d89a2ae986
|
|
BLAKE2b-256 checksum How to use checksums |
abd902d42951edc66c59ed2977f8b2ad47be1957285d30bf5c91829d7136ba28
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.0.0 CPython/3.13.7
|