SUCLEPY — Smart Universal Cleaner Library for Python
Project description
🧹 SUCLEPY — Smart Universal Cleaner Library for Python
✨ Overview
SUCLEPY v1.0.1 is a Smart Universal Cleaner Library for Python designed to make data cleaning automatic, fast, and reliable.
It handles missing values, duplicates, invalid emails, date formatting, string normalization, and generates detailed cleaning reports.
Perfect for data scientists, analysts, and developers working with messy datasets.
Compatible with Python 3.11+.
🗂 Version History
v1.0.1
- String columns missing values filled with
"abcd" - Dates automatically formatted to
DD-MM-YYYY - Missing or invalid dates filled with
"00-00-0000" - Emails missing entirely filled with
"abc@gmail.com" - Emails with missing domain automatically corrected
- Minor improvements to cleaning report generation and performance
v1.0.0
- Automatic Cleaning: Dataset cleaned intelligently using default strategies
- Duplicate Removal: Repeated rows removed to avoid redundancy
- Missing Value Handling: Numeric missing values handled via mean/median; categorical via mode; rows can be dropped
- Email Validation: Detects invalid email addresses
- Date Parsing: Converts various date formats to standardized format
- Text Normalization: Capitalizes and strips unnecessary spaces
- CSV Export: Saves cleaned data easily
- Cleaning Report: Generates a clear and printable cleaning report
⚙️ Installation
pip install suclepy
🧩 Library Structure
SUCLEPY ke modules aur folder structure:
suclepy/
├── __init__.py
├── core/
│ ├── __init__.py
│ ├── cleaner.py
│ ├── analyzer.py
│ ├── role_infer.py
│ ├── validator.py
│ └── reporter.py
└── utils.py
🧑💻 Usage Example
import pandas as pd
import suclepy as sp
# Create a sample dataset with 10 rows
df = pd.DataFrame({
"Name": ["Subodh", None, "", "Amit", "Riya", None, "Riya", "Mohan", "", "Geeta"],
"Age": [21, None, 22, 25, 20, 23, 20, None, 24, 22],
"Join_Date": [
"2024/05/10", None, "12-05-2024", "2024-05-11", "May 12, 2024",
"13/05/2024", "", "2024-05-14", None, "2024/05/15"
],
"Email": [
"subodh@", None, "amit@example.com", "amit@", "", "riya@gmail.com",
None, "mohan@", "geeta@example.com", ""
],
"City": [None, "Jaipur", "", "Delhi", None, "Mumbai", "Delhi", "", "Jaipur", "Pune"]
})
# Clean the dataset
report = sp.auto_clean(df)
# Print summary and cleaned data
print("=== Cleaning Summary ===")
print(report.summary())
print("\n=== Cleaned Data ===")
print(report.head(10)) # Show all 10 rows
# Save cleaned data to CSV
report.to_csv("cleaned_dataset.csv")
=== Cleaning Summary ===
SUCLEPY CLEANING REPORT
------------------------------
Total Rows (before): 10
Rows After Cleaning: 10
Duplicates Removed: 0
Missing Values Filled: 10
Invalid Emails Fixed: 7
Missing Strings Filled: 5
Dates Standardized: 7
Default Date Used: 3
Status: SUCCESS ✅
=== Cleaned Data ===
Name Age Join_Date Email City
0 Subodh 21.0 05-10-2024 subodh@gmail.com abcd
1 abcd 22.0 00-00-0000 abc@gmail.com Jaipur
2 abcd 22.0 12-05-2024 amit@example.com abcd
3 Amit 25.0 05-11-2024 amit@gmail.com Delhi
4 Riya 20.0 12-05-2024 abc@gmail.com abcd
5 abcd 23.0 13-05-2024 riya@gmail.com Mumbai
6 Riya 20.0 00-00-0000 abc@gmail.com Delhi
7 Mohan 22.0 14-05-2024 mohan@gmail.com abcd
8 abcd 24.0 00-00-0000 geeta@example.com Jaipur
9 Geeta 22.0 15-05-2024 abc@gmail.com Pune
File saved as cleaned_dataset.csv
🔧 Configuration Options
import suclepy as sp
sp.config({
"fill_missing_strategy": "mean", # Strategy for numeric missing values
"fill_string": "abcd", # Fill string missing values
"fill_date": "00-00-0000", # Fill invalid/missing dates
"validate_email": True, # Validate emails
"auto_correct_email": True, # Auto-fix incomplete emails
"drop_duplicates": True, # Remove duplicate rows
"date_format": "DD-MM-YYYY" # Standard date format
})
📝 Cleaning Rules / Behavior
- String Columns: missing (
None,NaN,"") →"abcd" - Dates: any format → standardized to
DD-MM-YYYY, missing →"00-00-0000" - Emails: missing →
"abc@gmail.com", incomplete domains auto-corrected - Numeric Columns: missing →
meanor configured strategy - Duplicates: automatically removed
- Text Normalization: capitalized and stripped extra spaces
❓ FAQ / Notes
Q1: What happens if an email has no @ or domain?
A: Automatically corrected or filled with "abc@gmail.com"
Q2: How are missing dates handled?
A: Any invalid/missing date is replaced with "00-00-0000"
Q3: Can I customize fill values?
A: Yes, see Configuration Options above
Q4: Compatible Python versions?
A: Python 3.11+
📚 Documentation & Resources
- GitHub Repository: https://github.com/subodhkryadav/suclepy
- LinkedIn: Subodh Kumar Yadav
📝 License
This project is licensed under the MIT License.
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file suclepy-1.0.1.tar.gz.
File metadata
- Download URL: suclepy-1.0.1.tar.gz
- Upload date:
- Size: 12.5 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
23dc8b2fea0b2c3eaf56e1cea4044ef67823e69a014ef2bb765972b688458875
|
|
| MD5 |
4fd7aee04afadb87933eff9f52531f62
|
|
| BLAKE2b-256 |
34300f89d1bb92beca3b46a05509fcf8a76375962a30d0e3dc825d645ba3633e
|
File details
Details for the file suclepy-1.0.1-py3-none-any.whl.
File metadata
- Download URL: suclepy-1.0.1-py3-none-any.whl
- Upload date:
- Size: 13.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.11.9
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
52388b9276c5c4a880ad4a77bf1bdfc3409e9f1f3aeae782d963555095efad2b
|
|
| MD5 |
7039b057164019b9bc59c691ffa58b37
|
|
| BLAKE2b-256 |
22268ed7df060d365ea6442aea174b68c5043e069a01a8e84a95fb8a1f0af068
|