Skip to main content

Playbooks for data. Open, process and save table based data.

Project description

Data Playbook

:book: Playbooks for data. Open, process and save table based data.

Automate repetitive tasks on table based data. Include various input and output tasks. Can be extended with custom modules.

Install: pip install dataplaybook

Use: dataplaybook playbook.yaml

Playbook structure

The playbook.yaml file allows you to load additional modules (containing tasks) and specify the tasks to execute in sequence, with all their parameters.

The tasks to perform typically follow the the structure of read, process, write.

Example yaml: (please note yaml is case sensitive)

modules: [list, of, modules]

tasks:
  - task: *name
    tables: # List of tables used by this task
    target: # Target table name of this function
    debug*: True/False # default: False
    # task specific properties, refer to each task

Tasks

Tasks are implemented as simple Python functions and the modules can be found in the dataplaybook/tasks folder.

Default tasks

  • drop
  • extend
  • filter
  • fuzzy_match (pip install fuzzywuzzy)
  • print
  • replace
  • unique
  • vlookup

Module io_xlsx (loaded by default)

  • read_excel
  • write_excel

Module io_misc (loaded by default)

  • read_tab_delim
  • read_text_regex
  • wget
  • write_csv

Module io_mongo (uses pymongo)

  • read_mongo
  • write_mongo
  • columns_to_list
  • list_to_columns

Module io_pdf (requires pdftotext)

  • read_pdf_pages
  • read_pdf_files

Module io_xml

Module ietf

Module gis

Module fnb

Install development version

  1. Clone the repo
  2. pip install <path> -e

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Filename, size & hash SHA256 hash help File type Python version Upload date
dataplaybook-0.2.4.tar.gz (22.2 kB) Copy SHA256 hash SHA256 Source None Jul 12, 2018

Supported by

Elastic Elastic Search Pingdom Pingdom Monitoring Google Google BigQuery Sentry Sentry Error logging AWS AWS Cloud computing DataDog DataDog Monitoring Fastly Fastly CDN DigiCert DigiCert EV certificate StatusPage StatusPage Status page