Skip to main content

Quickly build a news/web corpus with specifc topics or terms automatically from Google News or by specifying article links in a file. This module automatically extracts the body and title from each article and saves the result to either flatfiles or sqlite database.

Project description

News Corpus Builder

A simple module that can be used to quickly build a corpus from news articles. The generated corpus can be stored in a sqlite database or as flat files.

See http://skillachie.github.io/news-corpus-builder/ for installation and usage

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

news-corpus-builder-0.1.5.zip (6.2 kB view details)

Uploaded Source

File details

Details for the file news-corpus-builder-0.1.5.zip.

File metadata

File hashes

Hashes for news-corpus-builder-0.1.5.zip
Algorithm Hash digest
SHA256 1494f5e793f4bd1d7f04b869463167c2db0f09b98061c386708bd0f78b687895
MD5 c620a67dbc812d6fa403b3ed27c1fd90
BLAKE2b-256 5de143b98c242ce342d30fd574751513dbb6f75137503893e30813abec278ae5

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page