A comprehensive data cleaning library
Project description
CleanFusion: A Comprehensive Data Cleaning Library
Overview
CleanFusion is a powerful Python library designed to streamline and automate data cleaning tasks. Built with flexibility and ease-of-use in mind, it provides a comprehensive suite of tools for data assessment, preprocessing, and transformation. Whether you're dealing with missing values, outliers, inconsistent text data, or need to extract information from various file formats, CleanFusion offers an integrated solution for all your data cleaning needs.
Authors
- Himanshu Chopade, Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune
- Hriday Thaker, Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune
- Aryan Bachute, Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune
- Gautam Rajhans, Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune
Mentor
- Dr. Aditi Sharma (Mentor), Department of Computer Science and Engineering, Symbiosis Institute of Technology, Pune
Features
Core Functionality
- Data Assessment: Comprehensive analysis of data quality issues including missing values, outliers, data type inconsistencies, duplicates, and distribution analysis
- Missing Value Handling: Multiple strategies for imputing missing data (mean, median, mode, KNN, constant values)
- Outlier Detection and Treatment: Z-score and IQR-based methods with configurable thresholds and handling strategies
- Decision Engine: Intelligent recommendations for data cleaning strategies based on data characteristics
- Categorical Data Encoding: Various encoding techniques (label, one-hot, ordinal, binary, target encoding)
Text Processing
- Text Cleaning: Lowercasing, punctuation removal, stopword removal, lemmatization, and negation handling
- Text Vectorization: Convert text to numerical features using TF-IDF, count vectorization, or BERT embeddings
File Handling
- Multiple Format Support: Process CSV, TXT, DOCX, and PDF files
- Seamless Conversion: Extract and clean text from various document formats
Installation
pip install cleanfusion
Requirements
- Python 3.7+
- pandas
- numpy
- scikit-learn
- nltk
- PyPDF2
- pdfplumber (optional, for enhanced PDF extraction)
- sentence-transformers (optional, for BERT embeddings)
Quick Start
Command Line Interface
CleanFusion provides a convenient command-line interface for common data cleaning tasks:
cleanfusion --help
Replace file names and column names as needed for your use case.
# Assess data quality
cleanfusion assess data.csv --output assessment_report.txt
# Get cleaning recommendations
cleanfusion recommend data.csv --output recommendations.txt
# Clean a data file
cleanfusion clean data.csv --numerical median --categorical most_frequent --output cleaned_data.csv
# Clean text data
cleanfusion text document.txt --lowercase --remove-punctuation --remove-stopwords --output cleaned_text.txt
# Vectorize text data
cleanfusion vectorize document.txt --method tfidf --output vectors.csv
# Encode categorical data (one-hot encoding for specific columns)
cleanfusion encode data.csv --method onehot --columns Category,Region --output encoded_data.csv
# Label encoding
cleanfusion encode data.csv --method label --output encoded_data.csv
# Target encoding (with a target variable)
cleanfusion encode data.csv --method target --target TargetColumn --output target_encoded.csv
# Clean a DOCX file
cleanfusion clean document.docx --output cleaned_document.docx
# Clean a PDF file
cleanfusion clean document.pdf --output extracted_text.txt
Tip:
For details and all options for any command, run:
cleanfusion <command> --help
Python API
# Data assessment
from cleanfusion.core.data_assessment import DataAssessment
assessment = DataAssessment()
results = assessment.assess(df)
report = assessment.generate_report(output_path="report.txt")
# Data cleaning
from cleanfusion.core.data_preprocessor import DataPreprocessor
preprocessor = DataPreprocessor(
numerical_strategy="median",
categorical_strategy="most_frequent",
outlier_threshold=2.5
)
cleaned_df = preprocessor.transform(df)
# Text processing
from cleanfusion.text.text_cleaner import TextCleaner
from cleanfusion.text.text_vectorizer import TextVectorizer
cleaner = TextCleaner(remove_stopwords=True, lemmatize=True)
cleaned_text = cleaner.clean_text(text)
vectorizer = TextVectorizer(method="tfidf")
vectors = vectorizer.vectorize(texts)
# Categorical encoding
from cleanfusion.text.categorical_encoder import CategoricalEncoder
encoder = CategoricalEncoder(method="onehot")
encoded_df = encoder.fit_transform(df, columns=["Category", "Region"])
Detailed Usage
Data Assessment
The DataAssessment class provides comprehensive analysis of data quality:
from cleanfusion.core.data_assessment import DataAssessment
assessment = DataAssessment()
results = assessment.assess(df)
# Access specific assessment results
missing_values = results["missing_values"]
outliers = results["outliers"]
correlations = results["correlations"]
distributions = results["distributions"]
# Generate a comprehensive report
report = assessment.generate_report(output_path="assessment_report.txt")
Decision Engine
The DecisionEngine provides intelligent recommendations for data cleaning:
from cleanfusion.core.decision_engine import DecisionEngine
engine = DecisionEngine()
recommendations = engine.analyze(df)
# Get specific recommendations
missing_value_recs = recommendations["missing_values"]
outlier_recs = recommendations["outliers"]
feature_selection_recs = recommendations["feature_selection"]
text_cleaning_recs = recommendations["text_cleaning"]
# Generate a human-readable report
report = engine.get_recommendation_report()
Missing Value Handling
from cleanfusion.core.missing_value_handler import MissingValueHandler
handler = MissingValueHandler()
# Handle missing values in numerical columns
df["numeric_column"] = handler.handle_numerical_missing(df["numeric_column"], method="median")
# Handle missing values in categorical columns
df["category_column"] = handler.handle_categorical_missing(df["category_column"], method="most_frequent")
# Drop rows or columns with too many missing values
df = handler.drop_missing(df, threshold=0.5)
Outlier Detection and Handling
from cleanfusion.core.outlier_handler import OutlierHandler
handler = OutlierHandler(df)
# Detect outliers
z_outliers = handler.detect_outliers_zscore("column_name", threshold=3.0)
iqr_outliers = handler.detect_outliers_iqr("column_name", multiplier=1.5)
# Handle outliers
df = handler.handle_outliers_zscore("column_name", method="clip")
df = handler.handle_outliers_iqr("column_name", method="remove")
# Get summary of detected outliers
summary = handler.get_outlier_summary()
Text Processing
from cleanfusion.text.text_cleaner import TextCleaner
cleaner = TextCleaner(
remove_stopwords=True,
lemmatize=True,
handle_negation=True,
preserve_structure=False
)
cleaned_text = cleaner.clean_text(
text,
lowercase=True,
remove_punctuation=True
)
Text Vectorization
from cleanfusion.text.text_vectorizer import TextVectorizer
vectorizer = TextVectorizer(method="tfidf") # Options: "tfidf", "count", "bert"
vectors = vectorizer.vectorize(texts, return_array=True)
# Get feature names
feature_names = vectorizer.get_features()
Categorical Encoding
from cleanfusion.text.categorical_encoder import CategoricalEncoder
# Label encoding
encoder = CategoricalEncoder(method="label")
encoded_df = encoder.fit_transform(df, columns=["Category", "Region"])
# One-hot encoding
encoder = CategoricalEncoder(method="onehot")
encoded_df = encoder.fit_transform(df, columns=["Category"])
# Ordinal encoding with custom ordering
ordinal_mappings = {"Size": ["Small", "Medium", "Large"]}
encoder = CategoricalEncoder(method="ordinal")
encoded_df = encoder.fit_transform(df, columns=["Size"], ordinal_mappings=ordinal_mappings)
# Target encoding
encoder = CategoricalEncoder(method="target")
encoded_df = encoder.fit_transform(df, columns=["Category"], target_column="Target")
File Handling
# CSV files
from cleanfusion.file_handlers.csv_handler import CSVHandler
handler = CSVHandler()
df = handler.read_file("data.csv")
handler.write_file(df, "output.csv")
# Text files
from cleanfusion.file_handlers.txt_handler import TXTHandler
handler = TXTHandler()
text = handler.read_file("document.txt")
handler.write_file(text, "output.txt")
# PDF files
from cleanfusion.file_handlers.pdf_handler import PDFHandler
handler = PDFHandler()
text = handler.read_file("document.pdf")
handler.write_file(text, "extracted_text.txt")
# DOCX files
from cleanfusion.file_handlers.docx_handler import DOCXHandler
handler = DOCXHandler()
text = handler.read_file("document.docx")
handler.write_file(text, "extracted_text.txt")
Contributing
Contributions are welcome! Please feel free to submit a Pull Request.
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgments
- The scikit-learn team for their excellent machine learning library
- The NLTK project for natural language processing tools
- All contributors who have helped improve this library
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file cleanfusion-0.1.9.tar.gz.
File metadata
- Download URL: cleanfusion-0.1.9.tar.gz
- Upload date:
- Size: 30.9 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
c649f997ec9e0a56931096705f0bc5be2240092f6a79242735444141de085beb
|
|
| MD5 |
c22e98c7d21fa10f406659e14424a2c9
|
|
| BLAKE2b-256 |
08692b6bedfbc0da8e34470e0375244b91fa6143643a769644465ae2e75f2c0b
|
File details
Details for the file cleanfusion-0.1.9-py3-none-any.whl.
File metadata
- Download URL: cleanfusion-0.1.9-py3-none-any.whl
- Upload date:
- Size: 34.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.1.0 CPython/3.12.10
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
723a76797c1a3ea7b3c75fc880b67d48c5c861c16bdab01f3471e25d74a3fdc8
|
|
| MD5 |
b409b9eee13db969612b65652dcf5329
|
|
| BLAKE2b-256 |
06b5a4ca75b057f681602d335b6f78e54097f8eff6d93ecbdc564fd32a382843
|