A data preprocessing library for machine learning.
Project description
DatasetPreprocessor
DatasetPreprocessor is a Python library for preprocessing datasets for machine learning. It provides functionalities to handle missing values, encode categorical features, normalize numerical features, select important features, and more.
Features
- Read Data: Load data from a CSV file.
- Handle Missing Values: Fill or drop missing values.
- Encode Categorical Features: Encode categorical features using label encoding.
- Normalize Numerical Features: Normalize numerical features using standard scaling or min-max scaling.
- Select Features: Perform feature selection based on correlation with the target variable.
- Drop Useless Columns: Drop common useless columns from the dataset.
- Save and Load Preprocessed Data: Save preprocessed data to a CSV file and load it back.
- Generate Report: Generate a report summarizing the preprocessing steps.
Installation
You can install DatasetPreprocessor using pip:
pip install DatasetPreprocessor
Usage
Here is an example of how to use the DatasetPreprocessor library:
from DataPreprocessor import DatasetPreprocessor
# Initialize the preprocessor
preprocessor = DatasetPreprocessor()
# Read data
data = preprocessor.read_data("data.csv")
# Preprocess data
processed_data = preprocessor.preprocess_data(
file_path="data.csv",
target_column="target",
handle_missing="fill",
encode_categorical=True,
normalize_numerical="standard",
feature_selection_method="correlation",
n_features_to_select=5,
variance_ratio_threshold=0.1
)
# Save preprocessed data
preprocessor.save_preprocessed_data(processed_data, "processed_data.csv")
# Load preprocessed data
loaded_data = preprocessor.load_preprocessed_data("processed_data.csv")
# Print the first few rows of the preprocessed data
print(loaded_data.head())
# Generate a report
original_shape = data.shape
report = preprocessor.get_report(processed_data, original_shape)
print(report)
Methods
read_data(file_path)
Reads data from a CSV file.
-
Args:
file_path(str): Path to the CSV file.
-
Returns:
pandas.DataFrame: The loaded data.
handle_missing_values(data, strategy="fill")
Handles missing values in the dataset.
-
Args:
data(pandas.DataFrame): The input data.strategy(str, optional): The strategy to handle missing values. Options are'fill'(default) and'drop'.
-
Returns:
pandas.DataFrame: The data with missing values handled.
encode_categorical_features(data)
Encodes categorical features using label encoding.
-
Args:
data(pandas.DataFrame): The input data.
-
Returns:
pandas.DataFrame: The data with categorical features encoded.
normalize_numerical_features(data, norm_type="standard", variance_ratio_threshold=0.1)
Normalizes numerical features with high variance relative to other numerical features.
-
Args:
data(pandas.DataFrame): The input data.norm_type(str, optional): The normalization type. Options are'standard'(default) and'minmax'.variance_ratio_threshold(float, optional): The threshold for the ratio of feature variance to the maximum variance among numerical features. Default is0.1.
-
Returns:
pandas.DataFrame: The data with numerical features normalized.
select_features(data, target_column, n_features_to_select=None)
Performs feature selection based on correlation with the target variable.
-
Args:
data(pandas.DataFrame): The input data.target_column(str): The name of the target column.n_features_to_select(int, optional): The number of features to select. IfNone, all features will be kept.
-
Returns:
pandas.DataFrame: The data with selected features.
drop_useless_columns(data, useless_columns=None)
Drops useless columns from the dataset.
-
Args:
data(pandas.DataFrame): The input data.useless_columns(list, optional): A list of column names to drop. IfNone, a default list of common useless columns will be used.
-
Returns:
pandas.DataFrame: The data with useless columns dropped.
preprocess_data(file_path, target_column, handle_missing="fill", encode_categorical=True, normalize_numerical="standard", feature_selection_method="correlation", n_features_to_select=None, variance_ratio_threshold=0.1)
Preprocesses the dataset for machine learning.
-
Args:
file_path(str): Path to the CSV file.target_column(str): The name of the target column.handle_missing(str, optional): Strategy to handle missing values. Default is'fill'.encode_categorical(bool, optional): Whether to encode categorical features. Default isTrue.normalize_numerical(str, optional): The normalization type for numerical features. Default is'standard'.feature_selection_method(str, optional): The feature selection method to use. Options are'correlation'(default).n_features_to_select(int, optional): The number of features to select. IfNone, all features will be kept.variance_ratio_threshold(float, optional): The threshold for the ratio of feature variance to the maximum variance among numerical features. Default is0.1.
-
Returns:
pandas.DataFrame: The preprocessed data.
save_preprocessed_data(data, file_path)
Saves the preprocessed data to a file.
- Args:
data(pandas.DataFrame): The preprocessed data.file_path(str): The path to save the preprocessed data.
load_preprocessed_data(file_path)
Loads the preprocessed data from a file.
-
Args:
file_path(str): The path to load the preprocessed data from.
-
Returns:
pandas.DataFrame: The loaded preprocessed data.
get_report(data, original_shape)
Generates a report summarizing the preprocessing steps applied to the dataset.
-
Args:
data(pandas.DataFrame): The preprocessed data.original_shape(tuple): The original shape of the dataset before preprocessing.
-
Returns:
str: The report summarizing the preprocessing steps.
License
This project is licensed under the MIT License. See the LICENSE file for details.
Contributing
Contributions are welcome! Please open an issue or submit a pull request on GitHub.
Author
Abhishek Nair - abhishek,naiir@gmail.com Priyanshi Furiya - furiyapriyanshi@gmail.com
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file DatasetPreprocessor-0.1.1.tar.gz.
File metadata
- Download URL: DatasetPreprocessor-0.1.1.tar.gz
- Upload date:
- Size: 5.2 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.1.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
a07c168c644f5cadfa11985ae1694884bebd763f0102165f3e37390f05b9b567
|
|
| MD5 |
d89c1cde83bec578e2e801518ad8ec77
|
|
| BLAKE2b-256 |
fb65b87280e677b4927711c595f360679d5b00919ccf8e58463852882111dd9e
|
File details
Details for the file DatasetPreprocessor-0.1.1-py3-none-any.whl.
File metadata
- Download URL: DatasetPreprocessor-0.1.1-py3-none-any.whl
- Upload date:
- Size: 6.0 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/5.1.0 CPython/3.11.2
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
93f4b7647b8adbf9c9f2da0a8994a306aefd30c036a572259bea6846eb3915dd
|
|
| MD5 |
5090826f7a80154950eaad36149cc30d
|
|
| BLAKE2b-256 |
bb10320b04f8a600b90a2349b96b591dc129d9fec62893a825c2a7e973b211b3
|