Skip to main content

PreProcess1

Super Easy Way of PreProcessing your Data!

Good News!

Preprocess1 simplifies the preprocessing steps that are some time essential for ML/modelling, such as imputations, one hot encoding. There are over 20 preprocessing steps available. A summary of the options is below:

  • Auto infer data types
  • Impute (simple or with surrogate columns)
  • Ordinal Encoder
  • Drop categorical variables that have zero variance or near-zero variance
  • Club categorical variables levels together as a new level (other_infrequent) that are rare / at the bottom 5% of the variable
    distribution
  • Club unseen levels in test dataset with most/least frequent levels in train dataset
  • Reduce high cardinality in categorical features using clustering or counts
  • Generate sub-features from time feature such as 'month','weekday',is_month_end','is_month_start' & 'hour'
  • Group features by calculating min, max, mean, median & sd of similar features
  • Make nonlinear features (polynomial, sin, cos & tan)
  • Scales & Power Transform (zscore,minmax,yeo-johnson,quantile,maxabs,robust) , including option to transform target variable
  • Apply binning to variables when numeric features are provided as a list
  • Detect & remove outliers using isolation forest, KNN and PCA
  • Apply clusters to segment entire data
  • One Hot / Dummy encoding
  • Remove special characters from column names such as commas, square brackets etc. to make it compatible with Jason dependent models
  • Feature Selection through Random Forest, LightGBM and Pearson Correlation
  • Fix multicollinearity
  • Feature Interaction (DFS), multiply, divided, add and subtract features
  • Apply dimension reduction techniques such as pca_liner, pca_kernal, incremental or Tsne. except for pca_liner, all other methods only take the number of components (as integer) i.e no variance explanation method available

You can install the library as

pip install preprocess1
from preprocess1 import toolkit as t

Although one can use the methods individually (by calling the respective class) , such as:

binn = t.Binning(['feature_tobin'])
binned_data = binn.fit_transform(training_data)
binned_new_data = binn.transform(test_data)

However, there is more power to it. We have made pre-built complete pipelines to deploy all sorts of preprocessing transformers. Path1 is for supervised ML, and Path2 is for unsupervised ML problems. Below is how you use it:

# apply the path to the training dataset while clubbing rare categorical levels & scaling numerical features
# Imputation & One Hot Encoding is automatically applied
data_training_transformed = t.Preprocess_Path_One(training_data, 'target_column', club_rare_levels = True, scale_data= True)
# apply the pipeline to the test data set
data_test_transformed = pipe.fit_transform(test_data)

You can find more information under the docstring of each class/function. Enjoy coding! Please share your ideas, suggestions and critique with me.

License

Copyright 2019-2020 Fahad Akbar fahad.akbar@gmail.com

Metadata

Release files for preprocess1 0.1.42

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for preprocess1 0.1.42
File Size Uploaded
preprocess1-0.1.42.tar.gz 29.6 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for preprocess1 0.1.42
File Interpreter ABI Platform
preprocess1-0.1.42-py3-none-any.whl Python 3 none any Details

Total release size: 70.8 kB

Release files / preprocess1-0.1.42.tar.gz

Download URL preprocess1-0.1.42.tar.gz
Size 29.6 kB
Tags Source
SHA-256 checksum
How to use checksums
22045d7174552832d98f7cab0862f7a5b07831f26673a2e762da27f4fa92ac03
BLAKE2b-256 checksum
How to use checksums
0c12e44e25dd08b1a77e229b0f29ed1ff58c2649c2017c0f62b0f1c9461b4850
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.2.0 pkginfo/1.5.0.1 requests/2.24.0 setuptools/47.1.0 requests-toolbelt/0.9.1 tqdm/4.47.0 CPython/3.8.4rc1

Release files / preprocess1-0.1.42-py3-none-any.whl

Download URL preprocess1-0.1.42-py3-none-any.whl
Size 41.2 kB
Tags Python 3
SHA-256 checksum
How to use checksums
54ae56d1c3884432e1da8c2959c4172ed8ffd5cf0719e313a2a556e329e7f94e
BLAKE2b-256 checksum
How to use checksums
210c2b9cec0516e91a0c53b4d34c67512b893686e7a3bb2591f0765d5fdf1899
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.2.0 pkginfo/1.5.0.1 requests/2.24.0 setuptools/47.1.0 requests-toolbelt/0.9.1 tqdm/4.47.0 CPython/3.8.4rc1

Release history Release notifications | RSS feed

This release

0.1.42 This release

2 release files

0.1.38

2 release files

0.1.37

2 release files

0.1.36

2 release files

0.1.35

2 release files

0.1.30

2 release files

0.1.29

2 release files

0.1.28

2 release files

0.1.27

2 release files

0.1.26

2 release files

0.1.25

2 release files

0.1.24

2 release files

0.1.23

2 release files

0.1.22

2 release files

0.1.21

2 release files

0.1.20

2 release files

0.1.19

2 release files

0.1.18

2 release files

0.1.17

2 release files

0.1.16

2 release files

0.1.15

2 release files

0.1.14

2 release files

0.1.13

2 release files

0.1.12

2 release files

0.1.11

2 release files

0.1.10

2 release files

0.1.9

2 release files

0.1.8

2 release files

0.1.7

2 release files

0.1.6

2 release files

0.1.5

2 release files

0.1.4

2 release files

0.1.3

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

0.0.9

2 release files

0.0.5

2 release files

0.0.4

2 release files

0.0.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page