Skip to main content

A Python package for automatic statistical analysis and imputation

Project description

AutoStats

PyPI version Python versions PyPI - License

AutoStats is a Python library designed to simplify the process of cleaning, imputing, and analyzing datasets with minimal coding effort. It provides tools for generating exploratory reports, handling missing data, and optimizing imputation methods, making it ideal for data scientists and analysts.


🚀 Features

Report Module

  • Auto Report: Automatically generates an initial exploratory report from your dataset, categorizing columns and visualizing data distributions.
  • Manual Report: Allows users to specify categorical, continuous, and discrete columns for a more customized report.

Impute Module

  • Data Preprocessing: Automatically preprocesses datasets by handling missing values, encoding categorical variables, and identifying column types (categorical, continuous, discrete).
  • Imputation Methods:
    • KNN Imputation: Uses K-Nearest Neighbors to fill missing values.
    • MICE Imputation: Implements Multiple Imputation by Chained Equations.
    • MissForest Imputation: Uses Random Forests to impute missing values.
    • MIDAS Imputation: Leverages deep learning for advanced imputation.
  • Hyperparameter Optimization: Automatically tunes imputation methods using Optuna for the best performance.
  • Best Method Selection: Evaluates multiple imputation methods and selects the best-performing one for each column.

📦 Installation

To install AutoStats, ensure you have Python 3.8 or higher and run the following command:

pip install AutoStats

You can also view the project on PyPI.


Usage

Auto Report

To generate an automated exploratory data analysis report:

from AutoStats.report import auto_report
import pandas as pd

# Load your dataset
df = pd.read_csv("your_dataset.csv")

# Generate the report
auto_report(df, tresh=10, output_file="auto_report.pdf", df_name="Your Dataset")

Manual Report

To create a report with manually specified column types:

from AutoStats.report import manual_report
import pandas as pd

# Load your dataset
df = pd.read_csv("your_dataset.csv")

# Specify column types
categorical_cols = ['col1', 'col2']
continuous_cols = ['col3', 'col4']
discrete_cols = ['col5']

# Generate the report
manual_report(df, categorical_cols, continuous_cols, discrete_cols, output_file="manual_report.pdf", df_name="Your Dataset")

Data Imputation

To run the complete missing data imputation pipeline:

from AutoStats.impute import run_full_pipeline
import pandas as pd

# Load your dataset (with missing values)
df = pd.read_csv("your_dataset.csv")

# Run the imputation pipeline
best_imputed_df, summary_table = run_full_pipeline(df, simulate=True, build=True)

# The pipeline returns the imputed dataframe and a summary of the best methods used.
print(best_imputed_df.head())
print(summary_table)

How the imputation pipeline works:

  • simulate=True: This mode is for datasets that already have missing values. The pipeline will identify the missing data patterns and find the best imputation method for each column. This is the most common use case.
  • simulate=False: This mode is for evaluating the imputation methods on a complete dataset. You must specify a missingness_value (e.g., missingness_value=0.10 for 10%). The pipeline will artificially introduce missing values into your complete dataset and then impute them, allowing you to assess the performance of the different methods.

For a detailed technical explanation of the imputation module, please refer to the Imputation Technical Report.


📜 License

This project is licensed under the MIT License - see the LICENSE file for details.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

autostats-0.1.3.tar.gz (19.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

autostats-0.1.3-py3-none-any.whl (19.2 kB view details)

Uploaded Python 3

File details

Details for the file autostats-0.1.3.tar.gz.

File metadata

  • Download URL: autostats-0.1.3.tar.gz
  • Upload date:
  • Size: 19.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.11

File hashes

Hashes for autostats-0.1.3.tar.gz
Algorithm Hash digest
SHA256 8e0c515dcf863757be0e028b330ee2da50450208f3943f1398c6351b09a43cc4
MD5 853447c36345b605b722be62ecfac1c6
BLAKE2b-256 1d19d785c506ddc59119b8af736f48475b578a8694ae5db4d5e9dd46fbeea81b

See more details on using hashes here.

File details

Details for the file autostats-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: autostats-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 19.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.1.0 CPython/3.10.11

File hashes

Hashes for autostats-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 406ad501d13ec43a559a46d1fcc13add7ea166e699ea3c23c2571238c217b73e
MD5 916606d87e38096e913e203c0d2d61e4
BLAKE2b-256 e375c4685d8621bc3525c43d919e9f40a30c6bcfc40914f6283a89c6aa175b03

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page