Skip to main content

Automated data cleaning and preparation library for analysts before EDA.

Project description

data_cleaner_lib

A comprehensive Python library for automated data cleaning and preparation before Exploratory Data Analysis (EDA).

data_cleaner_lib provides a simple yet powerful pipeline that detects common data quality issues and fixes them with minimal code. It is designed for data analysts, data scientists, and ML practitioners who want a fast and reliable way to prepare datasets.


Key Features

Automatic Column Detection

Detects semantic column types automatically:

  • Numeric
  • Text
  • Email
  • Datetime
  • Categorical

Smart Data Type Conversion

Automatically converts mixed or string data to proper types:

  • "25" → Int64
  • "45.5" → float
  • "2024-01-01" → datetime

Missing Value Handling

Handles missing data using appropriate strategies:

Column Type Strategy
Numeric Median
Text Mode
Categorical Mode
Datetime Forward / Backward fill

Duplicate Removal

Detects and removes duplicate rows automatically.


Outlier Detection

Supports statistical outlier detection:

  • IQR method
  • Z-score method

Extreme values are capped automatically.


Text Standardization

Cleans textual data by:

  • Lowercasing
  • Removing extra spaces
  • Removing special characters
  • Preserving valid email formats

Data Quality Score

Generates a dataset quality score (0–100) based on:

  • Missing values
  • Duplicate rows

Example:

Quality Score: 87.5

Cleaning Recommendations

Automatically suggests potential cleaning actions:

Example:

age: 1 potential outliers detected → consider outlier treatment

EDA-Ready Summary

Generates a structured dataset summary including:

  • Column types
  • Missing values
  • Unique values
  • Numeric statistics

Example:

column detected_type dtype missing% unique min max
age numeric float64 0 3 25 42

One-Line Cleaning Pipeline

Prepare a dataset for analysis with a single command.

pipeline.prepare_for_eda()

Installation

Install using pip:

pip install data_cleaner_lib

Requirements:

  • Python ≥ 3.8
  • pandas
  • numpy

Quick Example

import pandas as pd
from data_cleaner_lib import CleanPipeline


df = pd.DataFrame({
    "name": ["John", "Jane", "Giri", None],
    "email": [" John@email.com ", "JANE@email.COM", "GIri@Gmail.com", None],
    "age": [25, 30, None, 47],
    "salary": [50000, 60000.00, None, None],
    "weight": ["45.5", "90", 68.76, "100"]
})

pipeline = CleanPipeline(df)

cleaned_df = pipeline.prepare_for_eda()

print(cleaned_df)

Output:

   name           email   age  salary  weight
0  john  john@email.com  25.0   50000   45.50
1  jane  jane@email.com  30.0   60000   90.00
2  giri  giri@gmail.com  30.0   55000   68.76
3  giri  john@email.com  42.5   55000  100.00

Generate EDA Summary

summary = pipeline.eda_summary()
print(summary)

Example Output:

column detected_type dtype missing_percent unique_values
name text object 0 3
age numeric float64 0 3

Get Quality Score

score = pipeline.quality_score()
print(score)

Cleaning Recommendations

pipeline.recommend_cleaning()

Public API (Available Functions)

After installing and importing the library, the following main functions are available through the CleanPipeline class.

Example import:

from data_cleaner_lib import CleanPipeline

Core Pipeline

Function Purpose
prepare_for_eda() Runs the full automated cleaning pipeline
detect_types() Detect semantic column types
profile() Generate dataset profiling report

Cleaning Operations

Function Purpose
fix_dtypes() Convert mixed data into proper datatypes
handle_missing() Fill missing values using smart strategies
remove_duplicates() Detect and remove duplicate rows
handle_outliers() Detect and cap outliers using statistical methods
clean_text() Normalize and clean text columns

Data Intelligence

Function Purpose
quality_score() Calculate dataset quality score (0–100)
recommend_cleaning() Generate automatic cleaning recommendations

Reporting

Function Purpose
eda_summary() Produce an EDA-ready dataset summary
report() Generate cleaning operation report

Example Usage

import pandas as pd
from data_cleaner_lib import CleanPipeline

pipeline = CleanPipeline(df)

pipeline.detect_types()
pipeline.profile()
pipeline.fix_dtypes()
pipeline.handle_missing()
pipeline.remove_duplicates()
pipeline.handle_outliers()
pipeline.clean_text()

summary = pipeline.eda_summary()

Project Structure

data_cleaner_lib/

src/
  data_cleaner_lib/

    detection/
    profiling/
    cleaning/
    reporting/
    quality/
    recommendations/

examples/
tests/

Running Tests

pytest

Roadmap

Future improvements:

  • Schema validation
  • Rule-based cleaning engine
  • Config-driven pipelines
  • Advanced anomaly detection

License

MIT License


Contributing

Pull requests and improvements are welcome.

If you find a bug or want a feature, open an issue.


Author

Giri V

Developed as a data cleaning framework for analysts preparing datasets before EDA and machine learning workflows.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

data_cleaner_lib-2.tar.gz (12.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

data_cleaner_lib-2-py3-none-any.whl (14.2 kB view details)

Uploaded Python 3

File details

Details for the file data_cleaner_lib-2.tar.gz.

File metadata

  • Download URL: data_cleaner_lib-2.tar.gz
  • Upload date:
  • Size: 12.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.1

File hashes

Hashes for data_cleaner_lib-2.tar.gz
Algorithm Hash digest
SHA256 4112f8c6db50a0c2d864da260eccb4d29bcb104393e6e8222edc47ce472051c7
MD5 17113373ac09255bb032f981cdf32278
BLAKE2b-256 58591b6b36937909964904b016b5e33677084b72fd2cd3245334b3863b67a0bf

See more details on using hashes here.

File details

Details for the file data_cleaner_lib-2-py3-none-any.whl.

File metadata

  • Download URL: data_cleaner_lib-2-py3-none-any.whl
  • Upload date:
  • Size: 14.2 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.1

File hashes

Hashes for data_cleaner_lib-2-py3-none-any.whl
Algorithm Hash digest
SHA256 dfe02c1122c385efc9177937038ca89c239dd89b60fb76662b8e9de8ebd81f91
MD5 a9e4ded7f1f1dc6456ecbf524121a9e9
BLAKE2b-256 7a220dd2fac8c28756c98f911bd67abfcb2906ba96d21fd505d59b433d9ff913

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page