readable-content
Collects actual content of any article, blog, news, etc.
Installation
pip install readable-content
Usage
After installing you need to do just add following two variables in settings.py of your Scrapy project
from readable_content.parser import ContentParser
parser = ContentParser("https://ideas.ted.com/how-do-animals-learn-how-to-be-well-animals-through-a-shared-culture/")
content = parser.get_content()
print(readable_content)
In case the website does not allow getting the content and throws 4XX or 3XX or any other error codes, we can first get the HTML using other techniques like using requests, using user-agent, applying proxies on your own, etc. Then the html content can be passed as following:
parser = ContentParser("https://ideas.ted.com/how-do-animals-learn-how-to-be-well-animals-through-a-shared-culture/", html_content)
Here html_content variable is string representation of the HTML.
Thank you!
Release files for readable-content 0.1.2
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| readable-content-0.1.2.tar.gz | 4.6 kB | Details |
Release files / readable-content-0.1.2.tar.gz
| Download URL | readable-content-0.1.2.tar.gz |
|---|---|
| Size | 4.6 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
94e6c41c95db814473a190170c861d7a2a20ce04b245c2ecbd20da9eef1f8f93
|
|
BLAKE2b-256 checksum How to use checksums |
9464f47d78e2ff5c7174795b492843e73fe033d825264fb658199a7afbea6295
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/3.1.1 pkginfo/1.5.0.1 requests/2.21.0 setuptools/45.3.0 requests-toolbelt/0.9.1 tqdm/4.46.0 CPython/3.7.7
|