Skip to main content

Scrapers that helps journalists at Kristeligt Dagblad

Project description

scrapers-for-journalists

Scraper(s) to help the journalists retrieve data or monitor sites for potential leads for stories.

Using the scrapers

pip install scrapers_for_journalists==0.1.2

And then import a scraper, e.g. from domstoldk.retrive import DomStolScrape

Every file in utils/can be imported in your scrapers, as it is added as a package in pyproject.toml. For example, you can import the BaseScraper with generic utilities like: from base import BaseScraper.

Description of current scrapers

domstol.dk

This scrapers retrieves information about current court cases ("retslister") in Danish "byretter" (Currently, Højesteret etc. are not included). Civil cases and tvangsauktioner are filtered away. Relevance of the cases are estimated based on keywords and "gerningskoder" (types of crimes) from the Danish Police.

To run it manually, use:

poetry run python domstoldk/retrieve.py --outfile test.xlsx

afdøde.dk

This scraper retrieves information about public "dødsannoncer" from afdøde.dk.

To run it manually, use:

poetry run python afdoededk/retrieve.py --outfile test.xlsx --from_date "2024-11-01" --to_date "2024-11-02"

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapers_for_journalists-0.1.2.tar.gz (27.2 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapers_for_journalists-0.1.2-py3-none-any.whl (27.5 kB view details)

Uploaded Python 3

File details

Details for the file scrapers_for_journalists-0.1.2.tar.gz.

File metadata

  • Download URL: scrapers_for_journalists-0.1.2.tar.gz
  • Upload date:
  • Size: 27.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/1.8.4 CPython/3.12.3 Linux/6.8.0-48-generic

File hashes

Hashes for scrapers_for_journalists-0.1.2.tar.gz
Algorithm Hash digest
SHA256 e8dad62ac7033737e68091b3b6ca34835587cd8bab706f270c946c85e0b57458
MD5 5cedf400438d30fbfc26e3f3b8fb40cf
BLAKE2b-256 c9921fa0a7547bed8eb86a37f6d67f32471bab12d435fd4f1ba4abaa4f01b6d2

See more details on using hashes here.

File details

Details for the file scrapers_for_journalists-0.1.2-py3-none-any.whl.

File metadata

File hashes

Hashes for scrapers_for_journalists-0.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 aa06d6d968419b86db1fa23ea70e6d5cb56b66da1f2fd0c0fd39ae232fdc9708
MD5 3969dfc5c687fc57b35e065c8badb9be
BLAKE2b-256 2667a58ac3876739dc9a0a9894aa07ff6b45e0bf3d5450fe835c2a4a14523985

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page