Skip to main content

Data Preprocessing fns

Project description

Installation

pip install dsfns

FNS Package

Function Descriptions

  1. outl_iqr(df, columns) Identifies and handles outliers in the specified columns using the Interquartile Range (IQR) method. Parameters: o df: DataFrame — The input data in which outliers will be detected. o columns: list — List of column names in which outliers need to be identified.

  2. outl_winsor(df, column, capping_method='iqr') Applies Winsorization to cap outliers in the specified column using either IQR or other capping methods. Parameters: o df: DataFrame — The input data to apply Winsorization. o column: str — The name of the column to apply the Winsorization to. o capping_method: str, default 'iqr' — Method used to define the outlier thresholds (options: 'iqr' , std, 'quantiles' or 'mad').

  3. outl_clip(df, columns) Clips extreme values to a predefined threshold in the specified columns, effectively handling outliers. Parameters: o df: DataFrame — The input data to clip outliers from. o columns: list — List of columns in which to clip the outliers.

  4. miss_repl(df, columns, type='mean') Replaces missing values in the specified columns using a chosen method. Parameters: o df: DataFrame — The input data in which missing values will be replaced. o columns: list — List of column names where missing values need to be replaced. o type: str, default 'mean' — The method used for replacement ('mean', 'median', or mode)

  5. miss_all(df) Identifies and returns all rows in the DataFrame that contain missing values with mean for numeric columns and mode (with index[0]) for object. Parameters: o df: DataFrame — The input data to check for missing values.

  6. norm(df) Normalizes a given value (or a set of values) to specific scale [0, 1]. Parameters: o df: data to be normalized.

  7. outlierColumns(df) Returns a list of columns that contain outliers based on IQR. Parameters: o df: DataFrame — The input data to check for outliers.

VERSION 1.3

  1. outlierCount(df, columns) Counts the number of outliers in the specified columns. Parameters: o df: DataFrame — The input data to count outliers in. o columns: list — List of columns to check for outliers.

  2. highFrequency(df, perc=0.5) Identifies and returns columns where more than the given percentage (default 70%) of values are identical, typically used to detect low-variance or high-frequency columns. Parameters: o df: DataFrame — The input data to identify high-frequency columns. o perc: float, default 0.7 — The percentage threshold for identifying high-frequency columns.

  3. miss_impute(df, columns, strategy='mean') The miss_impute function is designed to handle missing values in a pandas DataFrame by applying various imputation strategies. It utilizes the SimpleImputer from scikit-learn to efficiently fill missing values in specified columns. Parameters: o df: DataFrame — The input data to identify high-frequency columns. o columns (list): A list of column names in the DataFrame where missing values need to be imputed. o strategy (str, default='mean'):The imputation strategy to be applied. Options include:

    'mean': Replace missing values with the mean of the column. 'median': Replace missing values with the median of the column. 'mode' or 'most_frequent': Replace missing values with the most frequently occurring value in the column.

  4. Encoding_Label(df) Encodes categorical columns into numeric labels for compatibility with machine learning algorithms.

    Parameters: df: A Pandas DataFrame containing the dataset. Functionality: Identifies columns with data types object or category. Uses LabelEncoder from scikit-learn to transform these categorical columns into numerical labels.

  5. Scaler(df, method='minmax') Purpose: Scales numerical data for better performance during machine learning model training.

    Parameters: df: A Pandas DataFrame containing numeric data. method: Specifies the scaling technique to use. Options are: 'minmax' (default): Rescales data to a range of 0 to 1. 'standard': Standardizes data to have a mean of 0 and a standard deviation of 1. 'robust': Scales data using the median and interquartile range, making it robust to outliers.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

dsfns-1.4.tar.gz (3.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

dsfns-1.4-py3-none-any.whl (3.8 kB view details)

Uploaded Python 3

File details

Details for the file dsfns-1.4.tar.gz.

File metadata

  • Download URL: dsfns-1.4.tar.gz
  • Upload date:
  • Size: 3.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.12.3

File hashes

Hashes for dsfns-1.4.tar.gz
Algorithm Hash digest
SHA256 9d8d7e715b281c6c5e187bb1aeb1280e689a39e95c774474278b006d26f7e01e
MD5 f6ee5ee2b930cb826d9f8018602b4807
BLAKE2b-256 2962ffc0b135517c5ecc83e6f6ba17686f12524b6f01597d6d8b45817dde3592

See more details on using hashes here.

File details

Details for the file dsfns-1.4-py3-none-any.whl.

File metadata

  • Download URL: dsfns-1.4-py3-none-any.whl
  • Upload date:
  • Size: 3.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.0.1 CPython/3.12.3

File hashes

Hashes for dsfns-1.4-py3-none-any.whl
Algorithm Hash digest
SHA256 215afcf106574fcb04c6b2a00058f64c9bb13bb22b3978ff8454cd27f5460b04
MD5 616f5da86995da2de2a4927e15839617
BLAKE2b-256 10972237509730990bb5505866e97171e9da23561830b9143b291a28dcc48dd4

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page