Piper
Piper is a python package designed to simplify data wrangling tasks with pandas. It provides a set of wrapper functions or 'verbs' that provide a simpler interface to standard Pandas functions.
Piper functions accept and receive pandas dataframe objects. They can be used as standalone functions but are more powerful when used together in a Jupyter notebook cell to form a data pipeline. This is achieved by linking the functions using the '>>' link operator within a cell using %%piper magic command.
So, instead of the traditional pandas method of calling a method associated with an object, in this case showing the 'head' (first 5 rows) of the dataframe:
df.head()
Piper passes the result of the dataframe object as the first parameter to the next function in the pipeline that, in turn, both accepts and returns dataframe objects. So the equivalent of above with piper is:
%%piper
df >> head()
This 'chaining' or linking of functions provides a rapid, easy to use/remember approach to exploring, cleaning or building a data pipeline from csv, xml, excel, databases etc. Custom functions that accept and return dataframe objects can be linked together using this kind of syntax:
%%piper
read_oracle_database()
>> validate_data()
>> cleanup_data()
>> generate_summary()
>> write_to_target_system()
The concept is based on the approach used in the R language tidyverse and magrittr packages. The main functions are:
- select()
- assign()
- relocate()
- where()
- group_by()
- summarise()
- order_by()
For other piper functionality, please see the Goals and Features section.
Alternatives
For a comprehensive alternative, please check out Michael Chow's siuba package.
Table of contents
Installation
To install the package, enter the following:
pip install dpiper
Basic use
Example #1 - A dataframe consisting of two columns A and B.
import pandas as pd
import numpy as np
np.random.seed(42)
df = pd.DataFrame({'A': np.random.randint(10, 1000, 10),
'B': np.random.randint(10, 1000, 10)})
df.head()
| A | B | |
|---|---|---|
| 0 | 112 | 476 |
| 1 | 445 | 224 |
| 2 | 870 | 340 |
| 3 | 280 | 468 |
| 4 | 116 | 97 |
Let's create two further calculated columns and filter the 'D' column values.
df['C'] = df['A'] + df['B']
df['D'] = df['C'] < 1000
df[df['D'] == False]
| A | B | C | D | |
|---|---|---|---|---|
| 2 | 870 | 340 | 1210 | False |
| 8 | 624 | 673 | 1297 | False |
The equivalent in piper would be:
%%piper
df
>> assign(C = lambda x: x.A + x.B,
D = lambda x: x.C < 1000)
>> where("~D")
Example #2 Suppose you need the following function to trim columnar text data.
def trim_columns(df):
''' Trim blanks for given dataframe '''
str_cols = df.select_dtypes(include='object').columns
for col in str_cols:
df[col] = df[col].str.strip()
return df
Standard Pandas can combine the new function into a pipeline along with other transformation/filtering tasks by using the .pipe method:
import pandas as pd
from piper.factory import get_sample_data
df = get_sample_data()
# Select all columns EXCEPT 'dates'
subset_cols = ['order_dates', 'regions', 'countries', 'values_1', 'values_2']
criteria1 = ~df['countries'].isin(['Italy', 'Portugal'])
criteria2 = df['values_1'] > 40
criteria3 = df['values_2'] < 25
df2 = (df[subset_cols][criteria1 & criteria2 & criteria3]
.pipe(trim_columns)
.sort_values('countries', ascending=False))
df2.head()
Result:
| dates | order_dates | countries | ids | values_1 | values_2 |
|---|---|---|---|---|---|
| 2020-03-03 | 2020-03-09 | Sweden | E | 194 | 20 |
| 2020-05-02 | 2020-05-08 | Sweden | D | 322 | 14 |
| 2020-01-20 | 2020-01-26 | Spain | A | 183 | 20 |
| 2020-02-01 | 2020-02-07 | Norway | D | 344 | 21 |
| 2020-05-06 | 2020-05-12 | Norway | B | 135 | 21 |
The equivalent in piper would be to import the piper magic function, and the required 'verbs'.
from piper import piper
from piper.verbs import head, select, where, group_by, summarise, order_by
Using the %%piper magic function, piper verbs can be combined with standard python functions like trim_columns() using the linking symbol '>>' to form a data pipeline.
%%piper
get_sample_data()
>> trim_columns()
>> select('-dates')
>> where(""" ~countries.isin(['Italy', 'Portugal']) &
values_1 > 40 &
values_2 < 25 """)
>> order_by('countries', ascending=False)
>> head(5)
--info option If you specify this option, you see the equivalent pandas 'piped' version below the cell.
%%piper --info
get_sample_data()
>> trim_columns()
>> select('-dates')
>> where(""" ~countries.isin(['Italy', 'Portugal']) &
values_1 > 40 &
values_2 < 25 """)
>> order_by('countries', ascending=False)
>> head(5)
gives:
(get_sample_data()
.pipe(select, '-dates')
.pipe(where, """ ~countries.isin(['Italy', 'Portugal']) &values_1 > 40 &values_2 < 25 """)
.pipe(order_by, 'countries', ascending=False)
.pipe(head, 5))
Documentation
Further examples are available in these jupyter notebooks:
Goals and Features
- Enhance working with Excel files through the WorkBook class
- Exporting high quality formatted Excel Workbooks using xlsxwriter
- Provide access to databases with support for SQL based scripting and connections.
To-do list
- TBD
Contact
This is very much a personal library, in that its highly opinionated, flawed and probably of no use to anyone else :). However if it helps anyone else in their endeavours, that would be fantastic to hear about.
If you'd like to contact me, I'm miketarpey@gmx.net.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file dpiper-0.1.0.tar.gz.
File metadata
- Download URL: dpiper-0.1.0.tar.gz
- Upload date:
- Size: 76.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/3.3.0 pkginfo/1.7.0 requests/2.23.0 setuptools/44.0.0 requests-toolbelt/0.9.1 tqdm/4.42.1 CPython/3.8.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a8909cabfc71b7bd255b4850bb75135c2ba05871d9c47c3a87d5bf49fb9a27be
|
|
| MD5 |
13d0f267d3f7349b8373970863399603
|
|
| BLAKE2b-256 |
9fa530871eb5812fdac9df501f70c441db54a9d59f102756f1dcff91a066561d
|
File details
Details for the file dpiper-0.1.0-py3-none-any.whl.
File metadata
- Download URL: dpiper-0.1.0-py3-none-any.whl
- Upload date:
- Size: 84.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via:
twine/3.3.0 pkginfo/1.7.0 requests/2.23.0 setuptools/44.0.0 requests-toolbelt/0.9.1 tqdm/4.42.1 CPython/3.8.5
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
35beb8132a192ad55d2e4eac5c5e8ee7218595119bd1cc4054bd3f36d24fb025
|
|
| MD5 |
c38932e1c9fecd9cbe6824c53e09a88e
|
|
| BLAKE2b-256 |
b63b5e52e72e41935e1afb93732cecfb5a6628d9554b71175e9c90f502e71ca6
|