Skip to main content

Scrapers that helps journalists at Kristeligt Dagblad

Project description

scrapers-for-journalists

Scraper(s) to help the journalists retrieve data or monitor sites for potential leads for stories.

Using the scrapers

pip install scrapers_for_journalists

And then import a scraper, e.g. from domstoldk.retrive import DomStolScrape

Every file in utils/can be imported in your scrapers, as it is added as a package in pyproject.toml. For example, you can import the BaseScraper with generic utilities like: from base import BaseScraper.

Description of current scrapers

domstol.dk

This scrapers retrieves information about current court cases ("retslister") in Danish "byretter" (Currently, Højesteret etc. are not included). Civil cases and tvangsauktioner are filtered away. Relevance of the cases are estimated based on keywords and "gerningskoder" (types of crimes) from the Danish Police.

To run it manually, use:

poetry run python domstoldk/retrieve.py --outfile test.xlsx

afdøde.dk

This scraper retrieves information about public "dødsannoncer" from afdøde.dk.

To run it manually, use:

poetry run python afdoededk/retrieve.py --outfile test.xlsx --from_date "2024-11-01" --to_date "2024-11-02"

politi.dk

This scraper retrieves police reports ("døgnrapporter") from politi.dk.

To run it manually, use:

poetry run python doegnrapporter/retrieve.py --outfile test.xlsx --from_date "2024-11-01" --to_date "2024-11-01"

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

scrapers_for_journalists-0.1.7.tar.gz (27.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

scrapers_for_journalists-0.1.7-py3-none-any.whl (28.7 kB view details)

Uploaded Python 3

File details

Details for the file scrapers_for_journalists-0.1.7.tar.gz.

File metadata

  • Download URL: scrapers_for_journalists-0.1.7.tar.gz
  • Upload date:
  • Size: 27.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: poetry/2.1.1 CPython/3.12.3 Linux/6.8.0-51-generic

File hashes

Hashes for scrapers_for_journalists-0.1.7.tar.gz
Algorithm Hash digest
SHA256 c157ee970bf6bf6111e52c0f1ae0716a9b8e83a9847527cce36d3784130eb985
MD5 758915fbe6e69d261bd1198e037ab195
BLAKE2b-256 7dfdbb074ea06da2aebffb5edbb8d65a5bfcf67314a5cb0e14addaf4bda8926d

See more details on using hashes here.

File details

Details for the file scrapers_for_journalists-0.1.7-py3-none-any.whl.

File metadata

File hashes

Hashes for scrapers_for_journalists-0.1.7-py3-none-any.whl
Algorithm Hash digest
SHA256 fd9c36d0b5e1936cf532902e37cffaeb5249c9bba553fa5cd40243540203b4cf
MD5 74a9b8ada34526524f0118015b39516e
BLAKE2b-256 debe40f80280418088214745a30234f3573d1a6f4bd1eb511eb23f91cd5cc078

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page