Skip to main content
freshdata logo

freshdata

The explainable cleaning layer for pandas — decision-preserving data hygiene.

One call turns a messy CSV, Excel, or SQL export into analysis- and ML-ready data, and tells you exactly what it changed and why.

PyPI Version Python Versions License: MIT CI Docs Coverage

Documentation · Quickstart · API Reference · Changelog

What is freshdata?

freshdata is an automated data-cleaning library for Python. A rule-based decision engine profiles every column — missing ratio, dtype, skewness, cardinality, inferred role — and chooses the right action per column. Every decision carries a rationale, a risk level, and a confidence score, so nothing happens silently and nothing is left unexplained.

It fills the gap between tools that only describe data and tools that only validate it: freshdata makes the cleaning decision and shows its work, producing reproducible, auditable, ML-ready output with an audit trail you can hand to a reviewer.

Key features

  • One-call cleaning — fd.clean(df) handles missing values, outliers, duplicate detection (removal is opt-in), dtype repair, and messy column names.
  • Per-column decision engine — infers each column's role and applies explicit, documented rules instead of one blunt global strategy.
  • Explainable by design — every action carries a rationale, risk level, and confidence score; if a NaN survives, the report says why.
  • Safe defaults — never imputes an identifier, modifies a target column, or removes outliers blindly.
  • pandas-first, scalable when needed — pandas + NumPy core; pass a Polars frame and get one back, with optional Polars/DuckDB/Spark execution backends for larger-than-memory data.
  • CLI included — clean, plan, apply-plan, profile, learn, and trust subcommands for scripting and CI pipelines without writing Python.
  • Typed and tested — fully type-hinted (py.typed), vectorized, with a 93% coverage gate enforced in CI.

Installation

pip install freshdata-cleaner

The PyPI distribution is freshdata-cleaner; the import name is freshdata.

Requires Python >= 3.9 and pandas >= 1.5. The core install depends only on pandas and NumPy; everything else is an optional extra:

pip install "freshdata-cleaner[ml,polars]"
Extra Adds
ml KNN/model-based imputation
polars Polars DataFrame support
duckdb Out-of-core execution via DuckDB
spark Out-of-core execution via PySpark
viz Interactive HTML report rendering
privacy PII detection and anonymization
enterprise Compliance reporting, orchestration hooks, quality-ops exporters
all Everything above

See the installation guide for the full list of extras (domain packs, format parsers, streaming, entity resolution, and more).

Quickstart

import pandas as pd
import freshdata as fd

df = pd.read_csv("messy_export.csv")

cleaned = fd.clean(df)                               # one line
cleaned, report = fd.clean(df, return_report=True)   # ... with a full audit trail
print(report.summary())
freshdata clean report
  rows:    525 -> 525
  columns: 7 -> 6 (-1)
  missing: 421 -> 0 cell(s)
  actions (7):
    - [drop_duplicates] detected 25 duplicate row(s) (4.8%), none removed
    - [missing] 'age': filled 12 missing value(s) with median (39.6846)
    - [outliers] 'amount': flagged 15 outlier(s) in new column 'amount_outlier'

Duplicate rows are reported but kept by default; pass drop_duplicates=True to remove them.

The same operation is available from the command line:

freshdata clean messy_export.csv -o clean.csv --report audit.json

See the quickstart guide for strategies, reports, and CLI usage.

Beyond core cleaning

Optional layers, all off by default and covered in the documentation:

  • Repair plans — suggest a reviewable plan, then apply exactly the approved actions.
  • Context policies — compile plain-English cleaning rules into an enforceable policy.
  • Streaming — micro-batch and time-series-aware cleaning with bounded memory.
  • Privacy — PII detection, masking, and jurisdiction-aware anonymization policies.
  • Plugins — extend the engine with your own experts, validators, and backends.
  • AI Copilot (experimental) — deterministic, offline dataset analysis that returns an explainable cleaning plan and copy-ready freshdata code; no API key required.

Documentation and examples

Contributing

Contributions are welcome — standard GitHub flow: fork, branch, add tests, open a pull request. CI runs ruff, mypy, and the fast pytest lane on every PR. See CONTRIBUTING.md for setup and guidelines, and CODE_OF_CONDUCT.md for community standards.

New here? Good places to start:

License

MIT — see LICENSE.

Metadata

Release files for freshdata-cleaner 2.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for freshdata-cleaner 2.1.0
File Size Uploaded
freshdata_cleaner-2.1.0.tar.gz 2.7 MB Details

Built distribution (wheel)

Table of built distributions (wheels) for freshdata-cleaner 2.1.0
File Interpreter ABI Platform
freshdata_cleaner-2.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 3.6 MB

Release files / freshdata_cleaner-2.1.0.tar.gz

Download URL freshdata_cleaner-2.1.0.tar.gz
Size 2.7 MB
Tags Source
SHA-256 checksum
How to use checksums
bf445031373fde1a8e17c317c244db781826af184da32d8326f74e156c5af94d
BLAKE2b-256 checksum
How to use checksums
6e051a0437b5bb417b4117bfe583dddde2c0e4cfe1c9d258a1eac55f85d2aebf
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release files / freshdata_cleaner-2.1.0-py3-none-any.whl

Download URL freshdata_cleaner-2.1.0-py3-none-any.whl
Size 815.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
b232b4224cffdceda385584ac61d9d3ddc3076ab437fe636e1f4f0043b1cda2d
BLAKE2b-256 checksum
How to use checksums
77ac6d56eaea0cb67a0a396589399e12792f73ede8129df884068537d9d51e39
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 15, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2.1.0 This release

2 release files

2.0.0

2 release files

1.1.1

2 release files

1.1.0

2 release files

1.0.1

2 release files

1.0.0

2 release files

0.5.0

2 release files

0.4.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page