fast python port of arc90's readability tool
Project description
This code is under the Apache License 2.0. http://www.apache.org/licenses/LICENSE-2.0
This is a python port of a ruby port of arc90’s readability project
http://lab.arc90.com/experiments/readability/
In few words, Given a html document, it pulls out the main body text and cleans it up. It also can clean up title based on latest readability.js code.
- Based on:
Latest readability.js ( https://github.com/MHordecki/readability-redux/blob/master/readability/readability.js )
Ruby port by starrhorne and iterationlabs
Python port by gfxmonk ( https://github.com/gfxmonk/python-readability , based on BeautifulSoup )
Decruft effort to move to lxml ( http://www.minvolai.com/blog/decruft-arc90s-readability-in-python/ )
“BR to P” fix from readability.js which improves quality for smaller texts.
Github users contributions.
Installation:
easy_install readability-lxml or pip install readability-lxml
Usage:
from readability.readability import Document import urllib html = urllib.urlopen(url).read() readable_article = Document(html).summary() readable_title = Document(html).short_title()
Command-line usage:
python -m readability.readability -u http://pypi.python.org/pypi/readability-lxml
Document() kwarg options:
attributes:
debug: output debug messages
min_text_length:
retry_length:
url: will allow adjusting links to be absolute
Updates
0.2.5 Update setup.py for uploading .tar.gz to pypi
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
File details
Details for the file readability-lxml-0.2.5.zip
.
File metadata
- Download URL: readability-lxml-0.2.5.zip
- Upload date:
- Size: 15.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
File hashes
Algorithm | Hash digest | |
---|---|---|
SHA256 | fa013226a4f2f01d160fe173200d4a1984d931f5d0a38c4868b9598d41a18beb |
|
MD5 | a7615ee319f26320ce2414b5bd414541 |
|
BLAKE2b-256 | bafd80f04050a06a46426402cd98ba5596831ef8c0e08ace714661e3def48e8e |