This code is under the Apache License 2.0. http://www.apache.org/licenses/LICENSE-2.0
This is a python port of a ruby port of arc90’s readability project
http://lab.arc90.com/experiments/readability/
In few words, Given a html document, it pulls out the main body text and cleans it up. It also can clean up title based on latest readability.js code.
- Based on:
Latest readability.js ( https://github.com/MHordecki/readability-redux/blob/master/readability/readability.js )
Ruby port by starrhorne and iterationlabs
Python port by gfxmonk ( https://github.com/gfxmonk/python-readability , based on BeautifulSoup )
Decruft effort to move to lxml ( http://www.minvolai.com/blog/decruft-arc90s-readability-in-python/ )
“BR to P” fix from readability.js which improves quality for smaller texts.
Github users contributions.
Installation:
easy_install readability-xml or pip install readability-xml
Usage:
from readability.readability import Document import urllib html = urllib.urlopen(url).read() readable_article = Document(html).summary() readable_title = Document(html).short_title()
Command-line usage:
python -m readability.readability -u http://pypi.python.org/pypi/readability-lxml
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file readability-lxml-0.2.2.zip.
File metadata
- Download URL: readability-lxml-0.2.2.zip
- Upload date:
- Size: 13.6 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
41259a306fda066b2305fdb3573ef704b1527d96ee646039f37372aa90e63e0b
|
|
| MD5 |
40a583ee57485c02b5feb0eb6f05fff2
|
|
| BLAKE2b-256 |
063e68426e5a09c72688f99642f852ca71c220129c0a93becc6f9588bbeb3b68
|