python-readability
Given an HTML document, extract and clean up the main body text and title.
This is a Python port of a Ruby port of arc90's Readability project.
Installation
It's easy using pip, just run:
$ pip install readability-lxml
As an alternative, you may also use conda to install, just run:
$ conda install -c conda-forge readability-lxml
Usage
>>> import requests
>>> from readability import Document
>>> response = requests.get('http://example.com')
>>> doc = Document(response.content)
>>> print(doc.title())
Example Domain
>>> print(doc.summary())
<html><body><div><body id="readabilityBody">
<div>
<h1>Example Domain</h1>
<p>This domain is established to be used for illustrative examples in documents. You may
use this domain in examples without prior coordination or asking for permission.</p>
<p><a href="http://www.iana.org/domains/example">More information...</a></p>
</div>
</body>
</div></body></html>
The command-line interface accepts a local HTML file or a URL:
python -m readability -u https://example.com
Run the real-page extraction benchmark with the default minimum F1 of 0.95:
make benchmark
BENCHMARK_MIN_SCORE=0.97 make benchmark
The corpus methodology, version reports, and ten-engine comparison are
documented in
docs/quality/.
Security
Document.summary() removes common active content while extracting an article,
but readability-lxml is not a security boundary or a general-purpose HTML
sanitizer. Applications that render untrusted output must apply a dedicated
allowlist sanitizer and an appropriate Content Security Policy.
Change Log
- 0.9
- Expanded the declared Python range through 3.14, added
lxml6 support, and updated packaging metadata for Markdown documentation and currentcssselectreleases. - Added
python -m readabilityexecution for local files and URLs. - Fixed bytes input and encoding detection.
- Fixed
get_clean_html()before another API call has initialized the document. - Fixed shortened-title selection and CJK title length handling.
- Fixed XPath annotations incorrectly affecting the ruthless-parser retry length.
- Preserved inline links and formatting elements when converting misused
<div>elements into paragraphs. - Removed inline
display: none,visibility: hidden, HTMLhidden, and<noscript>fallback content before scoring. - Preserved content containers whose class or ID also contains an unlikely-candidate term such as
sidebar. - Preserved code blocks and semantic
<main>or<article>containers during unlikely-candidate filtering. - Recovered editorial leads, heading preambles, split article segments, and substantial list-based articles.
- Removed trailing linked calls to action without weakening general link-density filtering.
- Restricted retained video iframes to exact HTTP(S) YouTube and Vimeo hosts, preventing lookalike-host and userinfo URL bypasses, and removed active
srcdoccontent. - Added 181 reproducible fixtures: 166 quality fixtures split into base, Dragnet, GitHub user-issue, and complete 130-page Mozilla Readability corpora, plus 15 manually curated real-page regression fixtures from
jcharum/lxml-readability. - Replaced the unmaintained nose test runner with pytest in local, tox, and GitHub Actions workflows.
- Updated development and release targets for portable module execution, PEP 517 builds, version synchronization, isolated artifact checks, and current-version uploads.
- Improved the 166-page benchmark from precision 0.971, recall 0.884, and F1 0.926 in 0.8.4.1 to precision 0.991, recall 0.956, and F1 0.973.
- Corrected the README usage examples.
- Fixes GitHub issues #14, #108, #119, #130, #143, #146, #153, #158, #159, #163, #170, #176, #182, and #194. Release tracking issue #196 can be closed after 0.9 is published to PyPI.
- Expanded the declared Python range through 3.14, added
- 0.8.4 Better CJK support, thanks @cdhigh
- 0.8.3.1 Support for python 3.8 - 3.13
- 0.8.3 We can now save all images via keep_all_images=True (default is to save 1 main image), thanks @botlabsDev
- 0.8.2 Added article author(s) (thanks @mattblaha)
- 0.8.1 Fixed processing of non-ascii HTMLs via regexps.
- 0.8 Replaced XHTML output with HTML5 output in summary() call.
- 0.7.1 Support for Python 3.7 . Fixed a slowdown when processing documents with lots of spaces.
- 0.7 Improved HTML5 tags handling. Fixed stripping unwanted HTML nodes (only first matching node was removed before).
- 0.6 Finally a release which supports Python versions 2.6, 2.7, 3.3 - 3.6
- 0.5 Preparing a release to support Python versions 2.6, 2.7, 3.3 and 3.4
- 0.4 Added Videos loading and allowed more images per paragraph
- 0.3 Added Document.encoding, positive_keywords and negative_keywords
Licensing
This code is under the Apache License 2.0 license.
Thanks to
- Latest readability.js
- Ruby port by starrhorne and iterationlabs
- Python port by gfxmonk
- Decruft effort to move to lxml
- "BR to P" fix from readability.js which improves quality for smaller texts
- Github users contributions.
Release files for readability-lxml 0.9
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| readability_lxml-0.9.tar.gz | 25.2 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| readability_lxml-0.9-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 51.6 kB
Release files / readability_lxml-0.9.tar.gz
| Download URL | readability_lxml-0.9.tar.gz |
|---|---|
| Size | 25.2 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
f7a5f88ee194ed6c5aa36d14593fbfb20c5d7bb8f3f5bc57734288d45a5abe26
|
|
BLAKE2b-256 checksum How to use checksums |
e9fbe7c40afabd660f121fca9f0993e39dc042974c80a6c04ec4b27d7fbd7323
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|
Release files / readability_lxml-0.9-py3-none-any.whl
| Download URL | readability_lxml-0.9-py3-none-any.whl |
|---|---|
| Size | 26.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f43cf9ee74a15996581ae9ba1968d224d54344f6d134738bd0a5d5ea4b3ef55f
|
|
BLAKE2b-256 checksum How to use checksums |
1c2c0555d2c8d6c99152258d888f709ec1d46d64ae928a16313f760a2b264c41
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.14.2
|