Skip to main content

An agenda scraper framework for municipalities

Project description

Engage Scraper

Installation

pip i engage-scraper

About

The Engage Scraper is a standalone library that can be included in any service. The purpose of the scraper is to catalog a municipality's council meeting agendas in a usable format for such things as the engage-client and engage-backend.

To extend this library for your municipality, override the methods of the base class from the scraper_core/ directory and put it in scraper_logics/, prefacing it with your municipality name. For an example see the Santa Monica, CA example in the scraper_logics/ directory. The Santa Monica example makes use of htmlutils.py because it requires HTML scraping for its sources. Feel free to make PRs with new utilities (for example, PDF scraping, RSS scraping, JSON parsing, etc.). The Santa Monica example also uses SQLAlchemy for its models and that is what is preferred for use in the dbutils.py, however you can use anything. ORMs are preferred rather than vanilla psycopg2 or the like.

To use the postgres dbutils.py make sure to set these 5 environment variables (check dev.env and see docker-compose usage below):

  • POSTGRES_HOST optional a host or hostname that is resolvable. Defaults to localhost
  • POSTGRES_USER required
  • POSTGRES_PASSWORD required
  • POSTGRES_PORT optional defaults to 5432
  • POSTGRES_DB required The database used for cataloging your municipality's agendas.

An example of using the Santa Monica scraper library

from engage_scraper.scraper_logics import santamonica_scraper_logic

scraper = santamonica_scraper_logic.SantaMonicaScraper(committee="Santa Monica City Council")
scraper.get_available_agendas()
scraper.scrape()

For SantaMonicaScraper instantiation

For twitter utils used in SantaMonicaScraer

To use the santa monica logic, you must create an App on twitter (will work to make this optional). Following making an app, please use the structure dev.env file to insert the appropriate parameters. But make sure not to make changes to the repository's file. Copy the file up one directory and edit it there. Following the edit, use the docker-compose.yml for testing. You can add examples to examples/ and run them from the script in scripts/ using the docker container.

For the SantaMonicaScraper class the init has these options

  • tz_string="America/Los_Angeles" # defaulted string
  • years=["2019"] # defaulted array of strings of years
  • committee="Santa Monica City Council" # defaulted string of council name

The exposed API methods for scraper are

  • .get_available_agendas() # To get available agendas, no arguments
  • .scrape() # To process agendas and store contents

Feel free to expose more

  • Write wrappers for internal functions if you want to expose them
  • Write extra functions to handle more complex municipality-specific tasks

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

engage_scraper-0.0.32.tar.gz (13.5 kB view details)

Uploaded Source

Built Distribution

engage_scraper-0.0.32-py3-none-any.whl (16.5 kB view details)

Uploaded Python 3

File details

Details for the file engage_scraper-0.0.32.tar.gz.

File metadata

  • Download URL: engage_scraper-0.0.32.tar.gz
  • Upload date:
  • Size: 13.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.42.1 CPython/3.6.9

File hashes

Hashes for engage_scraper-0.0.32.tar.gz
Algorithm Hash digest
SHA256 4618d6b45cd73fcc404151cfe83a86344212b6f2f168507a1cb2c766d73ab385
MD5 b5948c358203f2d69ec6d86d45d89c1e
BLAKE2b-256 e54501429136dc4b4da07ebd8b213d927a864e6150142f54c411db31c03f2cdf

See more details on using hashes here.

File details

Details for the file engage_scraper-0.0.32-py3-none-any.whl.

File metadata

  • Download URL: engage_scraper-0.0.32-py3-none-any.whl
  • Upload date:
  • Size: 16.5 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/3.1.1 pkginfo/1.5.0.1 requests/2.22.0 setuptools/45.2.0 requests-toolbelt/0.9.1 tqdm/4.42.1 CPython/3.6.9

File hashes

Hashes for engage_scraper-0.0.32-py3-none-any.whl
Algorithm Hash digest
SHA256 2b8c850f7844c6d723a0ba3adada1ca435c5266aaea72ceb6c0f098fc52e0c90
MD5 822463cd1e8f55d5bce63bc57ff2c513
BLAKE2b-256 0885392d25e0e7d81cda7aa4307bfeb25cd00c26dc4d1fde0422fa7e45130dad

See more details on using hashes here.

Supported by

AWS AWS Cloud computing and Security Sponsor Datadog Datadog Monitoring Fastly Fastly CDN Google Google Download Analytics Microsoft Microsoft PSF Sponsor Pingdom Pingdom Monitoring Sentry Sentry Error logging StatusPage StatusPage Status page