Skip to main content

PyPI version

python-readability

Given an HTML document, extract and clean up the main body text and title.

This is a Python port of a Ruby port of arc90's Readability project.

Installation

It's easy using pip, just run:

$ pip install readability-lxml

As an alternative, you may also use conda to install, just run:

$ conda install -c conda-forge readability-lxml

Usage

>>> import requests
>>> from readability import Document

>>> response = requests.get('http://example.com')
>>> doc = Document(response.content)
>>> print(doc.title())
Example Domain

>>> print(doc.summary())
<html><body><div><body id="readabilityBody">
<div>
    <h1>Example Domain</h1>
<p>This domain is established to be used for illustrative examples in documents. You may
use this domain in examples without prior coordination or asking for permission.</p>
    <p><a href="http://www.iana.org/domains/example">More information...</a></p>
</div>
</body>
</div></body></html>

The command-line interface accepts a local HTML file or a URL:

python -m readability -u https://example.com

Run the real-page extraction benchmark with the default minimum F1 of 0.95:

make benchmark
BENCHMARK_MIN_SCORE=0.97 make benchmark

The corpus methodology, version reports, and ten-engine comparison are documented in docs/quality/.

Security

Document.summary() removes common active content while extracting an article, but readability-lxml is not a security boundary or a general-purpose HTML sanitizer. Applications that render untrusted output must apply a dedicated allowlist sanitizer and an appropriate Content Security Policy.

Change Log

  • 0.9
    • Expanded the declared Python range through 3.14, added lxml 6 support, and updated packaging metadata for Markdown documentation and current cssselect releases.
    • Added python -m readability execution for local files and URLs.
    • Fixed bytes input and encoding detection.
    • Fixed get_clean_html() before another API call has initialized the document.
    • Fixed shortened-title selection and CJK title length handling.
    • Fixed XPath annotations incorrectly affecting the ruthless-parser retry length.
    • Preserved inline links and formatting elements when converting misused <div> elements into paragraphs.
    • Removed inline display: none, visibility: hidden, HTML hidden, and <noscript> fallback content before scoring.
    • Preserved content containers whose class or ID also contains an unlikely-candidate term such as sidebar.
    • Preserved code blocks and semantic <main> or <article> containers during unlikely-candidate filtering.
    • Recovered editorial leads, heading preambles, split article segments, and substantial list-based articles.
    • Removed trailing linked calls to action without weakening general link-density filtering.
    • Restricted retained video iframes to exact HTTP(S) YouTube and Vimeo hosts, preventing lookalike-host and userinfo URL bypasses, and removed active srcdoc content.
    • Added 181 reproducible fixtures: 166 quality fixtures split into base, Dragnet, GitHub user-issue, and complete 130-page Mozilla Readability corpora, plus 15 manually curated real-page regression fixtures from jcharum/lxml-readability.
    • Replaced the unmaintained nose test runner with pytest in local, tox, and GitHub Actions workflows.
    • Updated development and release targets for portable module execution, PEP 517 builds, version synchronization, isolated artifact checks, and current-version uploads.
    • Improved the 166-page benchmark from precision 0.971, recall 0.884, and F1 0.926 in 0.8.4.1 to precision 0.991, recall 0.956, and F1 0.973.
    • Corrected the README usage examples.
    • Fixes GitHub issues #14, #108, #119, #130, #143, #146, #153, #158, #159, #163, #170, #176, #182, and #194. Release tracking issue #196 can be closed after 0.9 is published to PyPI.
  • 0.8.4 Better CJK support, thanks @cdhigh
  • 0.8.3.1 Support for python 3.8 - 3.13
  • 0.8.3 We can now save all images via keep_all_images=True (default is to save 1 main image), thanks @botlabsDev
  • 0.8.2 Added article author(s) (thanks @mattblaha)
  • 0.8.1 Fixed processing of non-ascii HTMLs via regexps.
  • 0.8 Replaced XHTML output with HTML5 output in summary() call.
  • 0.7.1 Support for Python 3.7 . Fixed a slowdown when processing documents with lots of spaces.
  • 0.7 Improved HTML5 tags handling. Fixed stripping unwanted HTML nodes (only first matching node was removed before).
  • 0.6 Finally a release which supports Python versions 2.6, 2.7, 3.3 - 3.6
  • 0.5 Preparing a release to support Python versions 2.6, 2.7, 3.3 and 3.4
  • 0.4 Added Videos loading and allowed more images per paragraph
  • 0.3 Added Document.encoding, positive_keywords and negative_keywords

Licensing

This code is under the Apache License 2.0 license.

Thanks to

  • Latest readability.js
  • Ruby port by starrhorne and iterationlabs
  • Python port by gfxmonk
  • Decruft effort to move to lxml
  • "BR to P" fix from readability.js which improves quality for smaller texts
  • Github users contributions.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

readability_lxml-0.9.tar.gz (25.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

readability_lxml-0.9-py3-none-any.whl (26.5 kB view details)

Uploaded Python 3

File details

Details for the file readability_lxml-0.9.tar.gz.

File metadata

  • Download URL: readability_lxml-0.9.tar.gz
  • Upload date:
  • Size: 25.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for readability_lxml-0.9.tar.gz
Algorithm Hash digest
SHA256 f7a5f88ee194ed6c5aa36d14593fbfb20c5d7bb8f3f5bc57734288d45a5abe26
MD5 1a7c3d735a6652467738f70e27681769
BLAKE2b-256 e9fbe7c40afabd660f121fca9f0993e39dc042974c80a6c04ec4b27d7fbd7323

See more details on using hashes here.

File details

Details for the file readability_lxml-0.9-py3-none-any.whl.

File metadata

  • Download URL: readability_lxml-0.9-py3-none-any.whl
  • Upload date:
  • Size: 26.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.14.2

File hashes

Hashes for readability_lxml-0.9-py3-none-any.whl
Algorithm Hash digest
SHA256 f43cf9ee74a15996581ae9ba1968d224d54344f6d134738bd0a5d5ea4b3ef55f
MD5 3bf26c55897909e14fce17b15999e0ed
BLAKE2b-256 1c2c0555d2c8d6c99152258d888f709ec1d46d64ae928a16313f760a2b264c41

See more details on using hashes here.

Release history Release notifications | RSS feed

This release

0.9 This release

2 files

0.8.4.1

2 files

0.8.4

2 files

0.8.1

2 files

0.7.1

1 file

0.7

1 file

0.6.2

1 file

0.6.1

1 file

0.6.0.5

1 file

0.6.0.4

1 file

0.6.0.3

1 file

0.5.1

1 file

0.5

1 file

0.3.0.6

1 file

0.3.0.5

1 file

0.3.0.3

2 files

0.3.0.2

2 files

0.3.0.1

2 files

0.3

2 files

0.2.6.1

1 file

0.2.6

1 file

0.2.5

1 file

0.2.3

1 file

0.2.2

1 file

0.2.1

1 file

0.2

1 file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page