Skip to main content

Sanitify

Intelligent data quality analysis and ML-assisted data cleaning for production Python workflows.

Sanitify is a production-grade Python library built to systematically analyze, validate, score, and improve structured datasets before they enter analytics or machine learning pipelines.

It provides structured profiling, rule-based quality validation, explainable quality scoring, deterministic cleaning utilities, and ML-assisted fix suggestions — all through a single public entry point:

from sanitify import DataCleaner

🚀 Why Sanitify?

Modern data systems fail more often due to poor data quality than model limitations.

Sanitify helps teams:

  • Detect structural issues early
  • Quantify dataset health
  • Enforce quality standards
  • Apply deterministic cleaning safely
  • Receive ML-assisted improvement suggestions
  • Maintain transparency and reproducibility

Designed for:

  • Data Engineers
  • ML Engineers
  • Analytics Teams
  • Data Platform Teams
  • Startups building internal data tooling

🧠 Design Philosophy

Sanitify follows strict engineering principles:

  • Single public API — DataCleaner
  • No hidden mutations — data is never altered silently
  • ML never auto-applies fixes — human-in-the-loop by design
  • Deterministic-first approach — rules before models
  • Explainable scoring logic
  • Modular architecture
  • Production-ready src layout
  • Test-covered implementation

📦 Core Features (v1.0)

  • Structured dataset profiling (schema-aware)
  • Scalable sampling for large datasets
  • Column-level metadata extraction
  • Rule-based quality validation engine
  • Explainable weighted quality scoring
  • Deterministic cleaning operations
  • ML-assisted fix suggestions (confidence-based)
  • Structured JSON report export
  • Clean, modular architecture

⚙️ Installation

Development Install

pip install -e ".[dev]"

Future PyPI Install

pip install sanitify

🔍 Quick Example

import pandas as pd
from sanitify import DataCleaner

df = pd.read_csv("data.csv")

dc = DataCleaner(df)

# 1. Profile dataset
profile = dc.profile()

# 2. Detect quality issues
issues = dc.check_quality()

# 3. Compute explainable quality score
score = dc.quality_score()

print(score)

📊 Example Output

Quality Issues

[
  {
    "column": "age",
    "rule": "high_missing",
    "severity": "medium",
    "metric": 0.42,
    "threshold": 0.3
  }
]

Quality Score

{
  "score": 72,
  "max_score": 100,
  "penalties": [
    {"rule": "high_missing", "deduction": 20}
  ]
}

Fully transparent. Fully explainable.


🏗 Architecture Overview

Sanitify uses a modular, production-oriented structure:

sanitify/
    datacleaner.py   # Public API
    core/            # Profiling, rules, scoring
    cleaning/        # Deterministic cleaning
    ai/              # ML suggestions
    report/          # Structured exports
    utils/           # Validation & helpers

Architectural Principles

  • src/ layout for clean packaging
  • Clear separation of concerns
  • Rule engine abstraction
  • Configurable scoring engine
  • No visualization inside core
  • No notebook dependencies
  • Stable output schemas

🧮 Quality Engine

Sanitify includes a rule-based validation system:

Built-in rules (v1):

  • High missing rate detection
  • Constant column detection
  • High cardinality detection
  • High duplicate rate detection

The engine is extensible and designed for future plugin support.


🔐 ML-Assisted Suggestions

Sanitify supports ML-driven fix recommendations.

Key guarantees:

  • Suggestions are confidence-scored
  • Fixes are never auto-applied
  • Users must explicitly approve changes
  • Deterministic cleaning remains primary

🛣 Roadmap

  • Structured profiling engine
  • Rule-based validation engine
  • Weighted scoring engine
  • Deterministic cleaning utilities
  • ML suggestion engine
  • Report exporters (JSON/YAML)
  • Streamlit demo application
  • PyPI release

🧪 Development

Run tests:

pytest

Run with coverage:

pytest --cov=sanitify --cov-report=term-missing

Sanitify follows:

  • Test-driven development
  • Modular design
  • Public API stability
  • Semantic commit conventions

🤝 Contributing

Contributions are welcome.

Before submitting a PR:

  • Ensure tests pass
  • Maintain coverage
  • Follow existing architecture patterns
  • Avoid breaking public API
  • Keep changes modular

📜 License

MIT License


🔮 Vision

Sanitify aims to become a lightweight but powerful data quality foundation layer for modern Python data stacks — sitting between raw ingestion and analytics/ML pipelines.

Transparent. Deterministic. Extensible. Production-ready. "# trigger CI"

Release files for sanitify 0.1.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for sanitify 0.1.3
File Size Uploaded
sanitify-0.1.3.tar.gz 13.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for sanitify 0.1.3
File Interpreter ABI Platform
sanitify-0.1.3-py3-none-any.whl Python 3 none any Details

Total release size: 27.0 kB

Release files / sanitify-0.1.3.tar.gz

Download URL sanitify-0.1.3.tar.gz
Size 13.8 kB
Tags Source
SHA-256 checksum
How to use checksums
bf45efe08030f3fd381059d4ca1c71d1ed194d1cf770ab9aeb2435257b1677bd
BLAKE2b-256 checksum
How to use checksums
61b7df62da4902cabeb107f7ad8aa8df2acedaba0a5b3ca0ed7dbd06102582ed
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release files / sanitify-0.1.3-py3-none-any.whl

Download URL sanitify-0.1.3-py3-none-any.whl
Size 13.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
3930100a975074b1243dae348e6021dad9dc0447674c3037a034f9d6e5069965
BLAKE2b-256 checksum
How to use checksums
4ef8a2bd8a453852be7321a986c1668c5e80e5c0b6f34c07a2e803753fe5a7d4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.12.10

Release history Release notifications | RSS feed

This release

0.1.3 This release

2 release files

0.1.2

2 release files

0.1.1

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page