A Python package for automatic statistical analysis and imputation
Project description
AutoStats
AutoStats is a Python library designed to simplify the process of cleaning, imputing, and analyzing datasets with minimal coding effort. It provides tools for generating exploratory reports, handling missing data, and optimizing imputation methods, making it ideal for data scientists and analysts.
🚀 Features
Report Module
- Auto Report: Automatically generates an initial exploratory report from your dataset, categorizing columns and visualizing data distributions.
- Manual Report: Allows users to specify categorical, continuous, and discrete columns for a more customized report.
Impute Module
- Data Preprocessing: Automatically preprocesses datasets by handling missing values, encoding categorical variables, and identifying column types (categorical, continuous, discrete).
- Imputation Methods:
- KNN Imputation: Uses K-Nearest Neighbors to fill missing values.
- MICE Imputation: Implements Multiple Imputation by Chained Equations.
- MissForest Imputation: Uses Random Forests to impute missing values.
- MIDAS Imputation: Leverages deep learning for advanced imputation.
- Hyperparameter Optimization: Automatically tunes imputation methods using Optuna for the best performance.
- Best Method Selection: Evaluates multiple imputation methods and selects the best-performing one for each column.
📦 Installation
To install AutoStats, ensure you have Python 3.8 or higher and run the following command:
pip install AutoStats
You can also view the project on PyPI.
Usage
Auto Report
To generate an automated exploratory data analysis report:
from AutoStats.report import auto_report
import pandas as pd
# Load your dataset
df = pd.read_csv("your_dataset.csv")
# Generate the report
auto_report(df, tresh=10, output_file="auto_report.pdf", df_name="Your Dataset")
Manual Report
To create a report with manually specified column types:
from AutoStats.report import manual_report
import pandas as pd
# Load your dataset
df = pd.read_csv("your_dataset.csv")
# Specify column types
categorical_cols = ['col1', 'col2']
continuous_cols = ['col3', 'col4']
discrete_cols = ['col5']
# Generate the report
manual_report(df, categorical_cols, continuous_cols, discrete_cols, output_file="manual_report.pdf", df_name="Your Dataset")
Data Imputation
To run the complete missing data imputation pipeline:
from AutoStats.impute import run_full_pipeline
import pandas as pd
# Load your dataset (with missing values)
df = pd.read_csv("your_dataset.csv")
# Run the imputation pipeline
best_imputed_df, summary_table = run_full_pipeline(df, simulate=True, build=True)
# The pipeline returns the imputed dataframe and a summary of the best methods used.
print(best_imputed_df.head())
print(summary_table)
How the imputation pipeline works:
simulate=True: This mode is for datasets that already have missing values. The pipeline will identify the missing data patterns and find the best imputation method for each column. This is the most common use case.simulate=False: This mode is for evaluating the imputation methods on a complete dataset. You must specify amissingness_value(e.g.,missingness_value=0.10for 10%). The pipeline will artificially introduce missing values into your complete dataset and then impute them, allowing you to assess the performance of the different methods.
For a detailed technical explanation of the imputation module, please refer to the Imputation Technical Report.
📜 License
This project is licensed under the MIT License - see the LICENSE file for details.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file autostats-0.1.3.tar.gz.
File metadata
- Download URL: autostats-0.1.3.tar.gz
- Upload date:
- Size: 19.0 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8e0c515dcf863757be0e028b330ee2da50450208f3943f1398c6351b09a43cc4
|
|
| MD5 |
853447c36345b605b722be62ecfac1c6
|
|
| BLAKE2b-256 |
1d19d785c506ddc59119b8af736f48475b578a8694ae5db4d5e9dd46fbeea81b
|
File details
Details for the file autostats-0.1.3-py3-none-any.whl.
File metadata
- Download URL: autostats-0.1.3-py3-none-any.whl
- Upload date:
- Size: 19.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.10.11
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
406ad501d13ec43a559a46d1fcc13add7ea166e699ea3c23c2571238c217b73e
|
|
| MD5 |
916606d87e38096e913e203c0d2d61e4
|
|
| BLAKE2b-256 |
e375c4685d8621bc3525c43d919e9f40a30c6bcfc40914f6283a89c6aa175b03
|