Newspaper-Scraper
The all-in-one Python package for seamless newspaper article indexing, scraping, and processing – supports public and premium content!
Intro
While tools like newspaper3k and goose3 can be used for extracting articles from news websites, they need a dedicated article url for older articles and do not support paywall content. This package aims to solve these issues by providing a unified interface for indexing, extracting and processing articles from newspapers.
- Indexing: Index articles from a newspaper website using the beautifulsoup package for public articles and selenium for paywall content.
- Extraction: Extract article content using the goose3 package.
- Processing: Process articles for nlp features using the spaCy package.
The indexing functionality is based on a dedicated file for each newspaper. A few newspapers are already supported, but it is easy to add new ones.
Supported Newspapers
| Logo | Newspaper | Country | Time span | Number of articles |
|---|---|---|---|---|
| Der Spiegel | Germany | Since 2000 | tbd | |
| Die Welt | Germany | Since 2000 | tbd | |
| Bild | Germany | Since 2006 | tbd | |
| Die Zeit | Germany | Since 1946 | tbd | |
| Handelsblatt | Germany | Since 2003 | tbd | |
| Der Tagesspiegel | Germany | Since 2000 | tbd | |
| Süddeutsche Zeitung | Germany | Since 2001 | tbd |
Setup
It is recommended to install the package in an dedicated Python environment.
To install the package via pip, run the following command:
pip install newspaper-scraper
To also include the nlp extraction functionality (via spaCy), run the following command:
pip install newspaper-scraper[nlp]
Usage
To index, extract and process all public and premium articles from Der Spiegel, published in August 2021, run the following code:
import newspaper_scraper as nps
from credentials import username, password
with nps.Spiegel(db_file='articles.db') as news:
news.index_articles_by_date_range('2021-08-01', '2021-08-31')
news.scrape_public_articles()
news.scrape_premium_articles(username=username, password=password)
news.nlp()
This will create a sqlite database file called articles.db in the current working directory. The database contains the following tables:
tblArticlesIndexed: Contains all indexed articles with their scraping/ processing status and whether they are public or premium content.tblArticlesScraped: Contains metadata for all parsed articles, provided by goose3.tblArticlesProcessed: Contains nlp features of the cleaned article text, provided by spaCy.
Metadata
Release files for newspaper-scraper 0.2.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| newspaper_scraper-0.2.1.tar.gz | 21.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| newspaper_scraper-0.2.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.2 kB
Release files / newspaper_scraper-0.2.1.tar.gz
| Download URL | newspaper_scraper-0.2.1.tar.gz |
|---|---|
| Size | 21.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
4fd8ada2a5e714b2b09b5872f92032b6136a0c918d33d3e5f5e6e54c306392f4
|
|
BLAKE2b-256 checksum How to use checksums |
e12d592a27f65aba43a1efd4e2861e1418f112aac5ad7241dd0d6b82ceaa0627
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.11.2
|
Release files / newspaper_scraper-0.2.1-py3-none-any.whl
| Download URL | newspaper_scraper-0.2.1-py3-none-any.whl |
|---|---|
| Size | 31.5 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
a2d956af0272b67ee0bab7f3630d97767a04d9d4c8f2cf5fb0be31f163bffc39
|
|
BLAKE2b-256 checksum How to use checksums |
186ffa7e82dbee757fd9322e4e09fbb1d99b968ea8c9f5786067253e4d7d9c67
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/4.0.2 CPython/3.11.2
|