sanitizepy
An open-source Python engine for automated tabular data quality inspection, explainable cleaning, and preprocessing.
Overview
sanitizepy provides modular, high-performance data engineering components built on top of pandas, numpy, scipy, rich, and pydantic. The library is designed around a transparent Detect → Explain → Recommend → Preview → Apply → Validate → Audit workflow.
sanitizepy separates responsibilities into dedicated subsystems:
- Core Engine & High-Level API: Centralized
Cleanerentry point supporting.inspect(),.plan(), and.clean(..., dry_run=True). - Dataset Health & Inspection: Read-only dataset analysis covering completeness, uniqueness, consistency, validity, datatypes, memory consumption, and statistical distributions with a composite Dataset Health Score (0–100).
- Explainable Recommendations & Planning: Rule-based issue detection with human-readable explanations (
WHAT,WHY,SEVERITY,EVIDENCE,RECOMMENDATION) and previewableCleaningPlaninstances. - Cleaning Engine: Safe, deterministic dataset transformations with dry-run support, before/after impact metrics, and detailed audit trails.
- Preprocessing & Feature Engineering: Stateful fit/transform operations for interactions, ratio features, polynomial terms, logarithmic transformations, and datetime extraction.
- Rule Engine: Quality validation framework with built-in rules, severity levels, and category classifications.
- Report Engine: Structured report generation, rendering (Text, JSON), and exporting (String, File).
- Pipeline Engine: Execution workflow orchestration with step timing and metadata tracking.
Capabilities
| Area | Component | Key Functionality |
|---|---|---|
| Core & High-Level API | Cleaner, CleanerConfig |
Unified entry point for .inspect(), .plan(), and .clean(..., dry_run=True) |
| Dataset Health & Inspection | Cleaner.inspect(), MissingValueInspector, DuplicateInspector, DatatypeInspector, MemoryInspector, StatisticsInspector |
Dataset Health Score (0-100), severity scoring, memory estimation, distribution stats |
| Recommendations & Planning | CleaningPlan, IssueDetector |
Human-readable recommendations, issue severity classification (critical, warning, info), previewable execution plan |
| Cleaning Engine | CleaningEngine, DropMissingRows, DropMissingColumns, FillMissing, DropDuplicates, DropColumns |
Deterministic cleaning operations with dry_run support, OperationResult, and immutable audit log |
| Preprocessing | FeatureEngineeringEngine, ColumnInteraction, RatioFeature, PolynomialFeature, LogFeature, DatetimeFeatures |
Stateful fit/transform feature generation preserving dataset indices |
| Rules | RuleEngine, RuleRegistry, Rule, register_builtin_rules |
Data quality rules, severity levels (info, warning, error, critical), custom rules |
| Reporting | ReportEngine, TextRenderer, JSONRenderer, StringExporter, FileExporter |
Structured immutable reports with multi-format rendering and exporting |
| Pipeline | PipelineEngine, CallableStep, TransformStep |
Sequenced workflow execution with step duration and row/column metrics |
Requirements
- Python:
>=3.11 - Core Dependencies:
numpy >= 1.24.0pandas >= 2.0.0scipy >= 1.10.0rich >= 13.0.0pydantic >= 2.0.0
Installation
Standard User Installation
Install sanitizepy using pip:
pip install sanitizepy
Or via python -m pip:
python -m pip install sanitizepy
Developer / Contributor Installation
For local development or contributing to the codebase, clone the repository and perform an editable installation with development dependencies:
git clone https://github.com/tahahssn/sanitizepy.git
cd sanitizepy
pip install -e .[dev]
Quick Start — High-Level API
The recommended entry point is the Cleaner class or the module-level convenience functions inspect(), plan(), and clean().
import pandas as pd
from sanitizepy import Cleaner
# Load your dataset
df = pd.read_csv("your_data.csv")
# 1. Inspect — Understand what's wrong
c = Cleaner()
report = c.inspect(df)
report.show() # Rich terminal health report
print(f"Health Score: {report.health_score}/100")
print(f"Critical Issues: {len(report.critical_issues)}")
print(f"Recommendations: {len(report.recommendations)}")
# 2. Plan — Generate a previewable cleaning plan
plan = c.plan(report)
plan.show() # Tabular plan preview
# Optional: disable or enable specific steps
plan.disable(2) # Disable step #2
plan.enable(2) # Re-enable step #2
# 3. Clean — Execute with dry-run or for real
# Dry run: see what WOULD happen without changing data
dry_result = c.clean(df, plan=plan, dry_run=True)
print(dry_result.summary())
# Apply for real
result = c.clean(df, plan=plan, dry_run=False)
cleaned_df = result.data
print(result.summary()) # Human-readable summary
print(result.audit_log) # JSON-serializable audit trail
Convenience Functions
from sanitizepy import inspect, plan, clean
report = inspect(df)
cleaning_plan = plan(report)
result = clean(df, cleaning_plan=cleaning_plan, dry_run=True)
Advanced Usage — Direct Engine Access
For granular control, use the individual engines directly:
from sanitizepy.cleaning import CleaningEngine, DropDuplicates, FillMissing
from sanitizepy.inspection import MissingValueInspector
# Read-only inspection
inspector = MissingValueInspector()
inspection_result = inspector.inspect(df)
# Manual cleaning engine
engine = CleaningEngine([
FillMissing(value=0.0, subset=["numeric_column"]),
DropDuplicates(keep="first"),
])
# Run with full result tracking
result = engine.run_with_result(df, dry_run=False)
print(result.summary())
print(result.audit_log)
Architecture & Design
sanitizepy adopts the following transparent workflow:
Detect → Explain → Recommend → Preview → Apply → Validate → Audit
┌───────────────────────┐
│ Dataset │
└───────────┬───────────┘
│
┌─────────────┴─────────────┐
│ Dataset Profiler / │
│ Issue Detector │ (Read-Only)
└─────────────┬─────────────┘
│
┌─────────────┴─────────────┐
│ Recommendation Engine │ (Explainable)
└─────────────┬─────────────┘
│
┌─────────────┴─────────────┐
│ Cleaning Plan │ (Previewable)
└─────────────┬─────────────┘
│
user approves
│
┌─────────────┴─────────────┐
│ Transformation Engine │ (Deterministic)
└─────────────┬─────────────┘
│
┌─────────────┴─────────────┐
│ Validation / Audit │ (Auditable)
└───────────────────────────┘
Documentation
Detailed documentation is available in the docs/ directory:
- Installation Guide: Requirements, virtual environments, installation commands, verification, and upgrade procedures.
- Quick Start Guide: Step-by-step examples for inspection, cleaning, feature engineering, rules, reporting, and pipelines.
- API Reference: Complete technical API documentation for classes, functions, dataclasses, models, and exceptions.
Development & Testing
To run the project test suite or linting tools:
Run Tests
pytest
Code Formatting & Linting
black --check src tests
ruff check src tests
mypy src
License
sanitizepy is distributed under the terms of the MIT License.
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file sanitizepy-0.1.0.tar.gz.
File metadata
- Download URL: sanitizepy-0.1.0.tar.gz
- Upload date:
- Size: 59.8 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
7a879e74f7937e2310ef4804a010b485e711c10bb017f324fe113ed0cd61d226
|
|
| MD5 |
29c7db2e2ef783eccba39109bbf0fea9
|
|
| BLAKE2b-256 |
738c0f50afb4482877982d999abad2ba22449c177347704fc8902a65e0b7c877
|
Provenance
The following attestation bundles were made for sanitizepy-0.1.0.tar.gz:
Publisher:
release.yml on tahahssn/sanitizepy
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sanitizepy-0.1.0.tar.gz -
Subject digest:
7a879e74f7937e2310ef4804a010b485e711c10bb017f324fe113ed0cd61d226 - Sigstore transparency entry: 2638418039
- Sigstore integration time:
-
Permalink:
tahahssn/sanitizepy@d82af254db8e18fec8b16de34c0ad936de9704ea -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/tahahssn
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d82af254db8e18fec8b16de34c0ad936de9704ea -
Trigger Event:
release
-
Statement type:
File details
Details for the file sanitizepy-0.1.0-py3-none-any.whl.
File metadata
- Download URL: sanitizepy-0.1.0-py3-none-any.whl
- Upload date:
- Size: 59.2 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4592e98baf6fee6e0d29b25fd7dbcdaa97b909c9a681c0b2c4d9bfe4599f6dbe
|
|
| MD5 |
2cacb082ea2502056a0c9d0e92b00937
|
|
| BLAKE2b-256 |
e20596fe0039e5ea460185bc5f4456961fd1336b04bfaddcbf956ea60d6143b9
|
Provenance
The following attestation bundles were made for sanitizepy-0.1.0-py3-none-any.whl:
Publisher:
release.yml on tahahssn/sanitizepy
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
sanitizepy-0.1.0-py3-none-any.whl -
Subject digest:
4592e98baf6fee6e0d29b25fd7dbcdaa97b909c9a681c0b2c4d9bfe4599f6dbe - Sigstore transparency entry: 2638418141
- Sigstore integration time:
-
Permalink:
tahahssn/sanitizepy@d82af254db8e18fec8b16de34c0ad936de9704ea -
Branch / Tag:
refs/tags/v0.1.0 - Owner: https://github.com/tahahssn
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@d82af254db8e18fec8b16de34c0ad936de9704ea -
Trigger Event:
release
-
Statement type: