pdpipe

Easy pipelines for pandas.

These details have not been verified by PyPI

Project links

Homepage

Project description

Easy pipelines for pandas DataFrames.

>>> df = pd.DataFrame(
        data=[[4, 165, 'USA'], [2, 180, 'UK'], [2, 170, 'Greece']],
        index=['Dana', 'Jack', 'Nick'],
        columns=['Medals', 'Height', 'Born']
    )
>>> pipeline = pdp.ColDrop('Medals').Binarize('Born')
>>> pipeline(df)
            Height  Born_UK  Born_USA
    Dana     165        0         1
    Jack     180        1         0
    Nick     170        0         0

1 Installation

Install pdpipe with:

pip install pdpipe

Some pipeline stages require scikit-learn; they will simply not be loaded if scikit-learn is not found on the system, and pdpipe will issue a warning. To use them you must also install scikit-learn.

2 Features

Pure Python.
Compatible with Python 3.5+.
A simple interface.
Informative prints and errors on pipeline application.
Chaining pipeline stages constructor calls for easy, one-liners pipelines.
Pipeline arithmetics.

2.1 Design Decisions

Data science-oriented naming (rather than statistics).
A functional approach: Pipelines never change input DataFrames. Nothing is done “in place”.
Opinionated operations: Help novices avoid mistake by default appliance of good practices; e.g., binarizing (creating dummy variables) a column will drop one of the resulting columns by default, to avoid the dummy variable trap (perfect multicollinearity).
Machine learning-oriented: The target use case is transforming tabular data into a vectorized dataset on which a machine learning model will be trained; e.g., column transformations will drop the source columns to avoid strong linear dependence.

3 Use

3.1 Pipeline Stages

3.1.1 Creating Pipeline Stages

You can create stages with the following syntax:

import pdpipe as pdp
drop_name = pdp.ColDrop("Name")

All pipeline stages have a predefined precondition function that returns True for dataframes to which the stage can be applied. By default, pipeline stages raise an exception if a DataFrame not meeting their precondition is piped through. This behaviour can be set per-stage by assigning exraise with a bool in the constructor call. If exraise is set to False the input DataFrame is instead returned without change:

drop_name = pdp.ColDrop("Name", exraise=False)

3.1.2 Applying Pipelines Stages

You can apply a pipeline stage to a DataFrame using its apply method:

res_df = pdp.ColDrop("Name").apply(df)

Pipeline stages are also callables, making the following syntax equivalent:

drop_name = pdp.ColDrop("Name")
res_df = drop_name(df)

The initialized exception behaviour of a pipeline stage can be overridden on a per-application basis:

drop_name = pdp.ColDrop("Name", exraise=False)
res_df = drop_name(df, exraise=True)

Additionally, to have an explanation message print after the precondition is checked but before the application of the pipeline stage, pass verbose=True:

res_df = drop_name(df, verbose=True)

3.1.3 Extending PipelineStage

To use other stages than the built-in ones (see Types of Pipeline Stages) you can extend the PipelineStage class. The constructor must pass the PipelineStage constructor the exmsg, appmsg and desc keyword arguments to set the exception message, application message and description for the pipeline stage, respectively. Additionally, the _prec and _op abstract methods must be implemented to define the precondition and the effect of the new pipeline stage, respectively.

3.1.4 Ad-Hoc Pipeline Stages

To create a custom pipeline stage without creating a proper new class, you can instantiate the AdHocStage class which takes a function in its op constructor parameter to define the stage’s operation, and the optional prec parameter to define a precondition (an always-true function is the default).

3.2 Pipelines

3.2.1 Creating Pipelines

Pipelines can be created by supplying a list of pipeline stages:

pipeline = pdp.Pipeline([pdp.ColDrop("Name"), pdp.Binarize("Label")])

3.2.2 Pipeline Arithmetics

Alternatively, you can create pipelines by adding pipeline stages together:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")

Or even by adding pipelines together or pipelines to pipeline stages:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
pipeline += pdp.ApplyToRows("Job", {"Part": True, "Full":True, "No": False})
pipeline += pdp.Pipeline([pdp.ColRename({"Job": "Employed"})])

3.2.3 Pipeline Chaining

Pipeline stages can also be chained to other stages to create pipelines:

pipeline = pdp.ColDrop("Name").Binarize("Label").ValDrop([-1], "Children")

3.2.4 Pipeline Slicing

Pipelines are Python Sequence objects, and as such can be sliced using Python’s slicing notation, just like lists:

pipeline = pdp.ColDrop("Name").Binarize("Label").ValDrop([-1], "Children").ApplyByCols("height", math.ceil)
result_df = pipeline[1:2](df)

3.2.5 Applying Pipelines

Pipelines are pipeline stages themselves, and can be applied to a DataFrame using the same syntax, applying each of the stages making them up, in order:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
res_df = pipeline(df)

Assigning the exraise parameter to a pipeline apply call with a bool sets or unsets exception raising on failed preconditions for all contained stages:

pipeline = pdp.ColDrop("Name") + pdp.Binarize("Label")
res_df = pipeline.apply(df, exraise=False)

Additionally, passing verbose=True to a pipeline apply call will apply all pipeline stages verbosely:

res_df = pipeline.apply(df, verbose=True)

4 Types of Pipeline Stages

4.1 Basic Stages

AdHocStage - Define custom pipeline stages on the fly.
ColDrop - Drop columns by name.
ValDrop - Drop rows by by their value in specific or all columns.
ValKeep - Keep rows by by their value in specific or all columns.
ColRename - Rename columns.

4.2 Column Generation

Bin - Convert a continuous valued column to categoric data using binning.
Binarize - Convert a categorical column to the several binary columns corresponding to it.
ApplyToRows - Generate columns by applying a function to each row.
ApplyByCols - Generate columns by applying an element-wise function to columns.

4.3 Scikit-learn-dependent Stages

Encode - Encode a categorical column to corresponding number values.

5 Contributing

Package author and current maintainer is Shay Palachy (shay.palachy@gmail.com); You are more than welcome to approach him for help. Contributions are very welcomed, especially since this package is very much in its infancy and many other pipeline stages can be added.

5.1 Installing for development

Clone:

git clone git@github.com:shaypal5/pdpipe.git

Install in development mode with test dependencies:

cd pdpipe
pip install -e ".[test]"

5.2 Running the tests

To run the tests, use:

python -m pytest --cov=pdpipe

5.3 Adding documentation

This project is documented using the numpy docstring conventions, which were chosen as they are perhaps the most widely-spread conventions that are both supported by common tools such as Sphinx and result in human-readable docstrings (in my personal opinion, of course). When documenting code you add to this project, please follow these conventions.

6 Credits

Created by Shay Palachy (shay.palachy@gmail.com).

Project details

These details have not been verified by PyPI

Project links

Homepage

Release history Release notifications | RSS feed

0.3.2

Sep 19, 2022

0.3.1

Aug 9, 2022

0.3.0

Jul 4, 2022

0.2.8

Jun 23, 2022

0.2.7

Jun 22, 2022

0.2.6

Jun 22, 2022

0.2.5

Jun 7, 2022

0.2.4

May 12, 2022

0.2.3

Mar 13, 2022

0.2.2

Mar 10, 2022

0.2.1

Feb 23, 2022

0.2.0

Feb 14, 2022

0.1.6

Jan 30, 2022

0.1.5

Jan 29, 2022

0.1.4

Jan 29, 2022

0.1.3

Jan 26, 2022

0.1.2

Jan 23, 2022

0.1.0

Jan 23, 2022

0.0.72

Jan 19, 2022

0.0.71

Dec 26, 2021

0.0.70

Dec 19, 2021

0.0.69

Dec 10, 2021

0.0.68

Dec 8, 2021

0.0.67

Nov 15, 2021

0.0.66

Nov 8, 2021

0.0.65

Nov 8, 2021

0.0.64

Nov 8, 2021

0.0.63

Nov 3, 2021

0.0.62

Oct 27, 2021

0.0.61

Oct 27, 2021

0.0.60

Sep 29, 2021

0.0.59

Aug 30, 2021

0.0.58

Aug 28, 2021

0.0.57

Aug 25, 2021

0.0.56

Aug 18, 2021

0.0.55

Aug 18, 2021

0.0.54

Aug 18, 2021

0.0.53

Nov 9, 2020

0.0.52

Oct 30, 2020

0.0.51

Oct 1, 2020

0.0.50

Aug 27, 2020

0.0.49

May 5, 2020

0.0.48

May 5, 2020

0.0.46

Feb 26, 2020

0.0.45

Feb 24, 2020

0.0.44

Feb 24, 2020

0.0.43

Feb 17, 2020

0.0.42

Feb 5, 2020

0.0.41

Feb 3, 2020

0.0.40

Feb 3, 2020

0.0.39

Jan 26, 2020

0.0.38

Jan 20, 2020

0.0.37

Jan 7, 2020

0.0.35

Dec 21, 2019

0.0.33

Dec 7, 2019

0.0.32

Dec 3, 2019

0.0.31

Jun 27, 2019

0.0.30

Jun 14, 2019

0.0.29

Jun 14, 2019

0.0.27

May 28, 2018

0.0.26

May 28, 2018

0.0.25

May 9, 2018

0.0.24

May 2, 2018

0.0.23

Apr 22, 2018

0.0.22

Apr 16, 2018

0.0.21

Apr 16, 2018

0.0.20

Apr 8, 2018

0.0.19

Apr 8, 2018

0.0.18

Mar 20, 2018

0.0.17

Mar 12, 2018

0.0.16

Mar 11, 2018

0.0.15

Mar 7, 2018

0.0.14

Mar 7, 2018

0.0.13

Mar 3, 2018

0.0.12

Feb 12, 2018

0.0.11

Feb 12, 2018

0.0.10

Feb 12, 2018

0.0.9

Feb 5, 2018

0.0.8

Feb 5, 2018

0.0.7

Jan 30, 2018

0.0.6

Jan 14, 2018

This version

0.0.5

May 24, 2017

0.0.4

May 5, 2017

0.0.3

May 5, 2017

0.0.2

Mar 17, 2017

0.0.1

Mar 16, 2017

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

pdpipe-0.0.5.tar.gz (30.4 kB view hashes)

Uploaded May 24, 2017 Source

Hashes for pdpipe-0.0.5.tar.gz

Hashes for pdpipe-0.0.5.tar.gz
Algorithm	Hash digest
SHA256	`7a0d63a5a718f06c33d385d027b757b17dad67064502f80a6b98ed0e028dd0f0`
MD5	`c9266c1647954021c405584c2b9de308`
BLAKE2b-256	`5aa17026f88bc9229a432b842f1057d1014b459b7a508f4d48af7fc00f860570`