Skip to main content

datadiff-engine

Tests Python License

A Python library and CLI for comparing datasets, detecting schema changes, profiling columns, and identifying simple numeric data drift.

Built for data engineers who want a quick answer to:

What changed between yesterday's dataset and today's dataset?


Why datadiff-engine?

Data pipelines frequently produce datasets that look valid but have unexpected changes:

  • Row counts suddenly increase or decrease
  • Columns are added or removed
  • Data types change
  • Null rates increase
  • Unique values change
  • Numeric distributions shift

datadiff-engine provides a structured comparison so these changes can be inspected programmatically or directly from the command line.


Features

  • Compare row counts
  • Detect added columns
  • Detect removed columns
  • Detect data-type changes
  • Compare null rates
  • Compare unique-value counts
  • Generate numeric statistics
  • Detect simple numeric drift
  • Analyze categorical columns
  • Configure null-rate severity thresholds
  • Ignore selected columns for numeric drift analysis
  • Python API
  • Command-line interface
  • JSON output
  • CSV support
  • Parquet support

Installation

pip install datadiff-engine

Python usage

import pandas as pd

from datadiff_engine import compare

old = pd.DataFrame({
    "customer_id": [1, 2, 3],
    "amount": [100, 200, 300],
})

new = pd.DataFrame({
    "customer_id": [1, 2, 3, 4],
    "amount": [100, 200, None, 500],
    "country": ["IN", "IN", "US", "IN"],
})

diff = compare(old, new)

print(diff)

Example output:

DATASET DIFF
========================================

ROWS
  3 → 4
  Change: +1 (+33.33%)

SCHEMA
  + Added:   ['country']
  - Removed: none
  ~ Changed: amount (int64 → float64)

COLUMN CHANGES
----------------------------------------
amount
  Data type: int64 → float64
  Null rate: 0.00% → 25.00% (+25.00 pp)
  Severity:  CRITICAL

  Statistics
    Mean:    200.00 → 266.67
    Median:  200.00 → 200.00
    Min:     100.00 → 100.00
    Max:     300.00 → 500.00
    P95:     290.00 → 470.00

Column-level analysis

Access details for an individual column:

amount = diff.column("amount")

print(amount.name)
print(amount.old_dtype)
print(amount.new_dtype)

print(amount.null_rate_old)
print(amount.null_rate_new)
print(amount.null_rate_change)

print(amount.severity)

For numeric columns, statistics are available:

print(amount.old_stats.mean)
print(amount.new_stats.mean)

print(amount.old_stats.median)
print(amount.new_stats.median)

print(amount.old_stats.min)
print(amount.new_stats.max)

print(amount.old_stats.p95)
print(amount.new_stats.p95)

Numeric drift information is also available:

if amount.drift:
    print(amount.drift.mean_change_pct)
    print(amount.drift.median_change_pct)
    print(amount.drift.p95_change_pct)
    print(amount.drift.has_drift)

Custom severity thresholds

The default null-rate thresholds can be customized:

from datadiff_engine import DiffConfig, compare

config = DiffConfig(
    warning_null_rate=0.10,
    critical_null_rate=0.50,
)

diff = compare(
    old,
    new,
    config=config,
)

This allows different projects or pipelines to define their own thresholds.


Ignoring columns for numeric drift

Some numeric columns represent identifiers rather than measurements.

For example:

customer_id
order_id
account_id

These columns may not be appropriate for numeric drift analysis.

You can exclude them:

from datadiff_engine import DiffConfig, compare

config = DiffConfig(
    ignored_columns={"customer_id"},
)

diff = compare(
    old,
    new,
    config=config,
)

The column is still part of the comparison, but numeric drift analysis is skipped for the ignored column.


JSON output

The comparison can be converted into a machine-readable dictionary:

result = diff.to_dict()

print(result)

This is useful when integrating the library into:

  • Data pipelines
  • Airflow tasks
  • CI/CD pipelines
  • Logging systems
  • APIs
  • Monitoring tools

Example:

{
    "rows": {
        "old": 3,
        "new": 4,
        "change": 1,
        "change_pct": 33.33
    },
    "schema": {
        "added": ["country"],
        "removed": [],
        "changed_types": {
            "amount": ["int64", "float64"]
        }
    }
}

Command-line interface

datadiff-engine also provides a CLI.

Compare two CSV files:

datadiff old.csv new.csv

Compare Parquet files:

datadiff old.parquet new.parquet

Output JSON:

datadiff old.csv new.csv --json

Ignore columns for numeric drift:

datadiff old.csv new.csv --ignore customer_id

Show help:

datadiff --help

Example

Suppose yesterday's dataset contains:

customer_id,amount,country
1,100,IN
2,200,IN
3,300,US

and today's dataset contains:

customer_id,amount,country
1,100,IN
2,200,US
3,,US
4,500,BD

Running:

datadiff old.csv new.csv

can identify:

Rows: 3 → 4

amount
  Data type: int64 → float64
  Null rate: 0.00% → 25.00%

country
  Unique values: 2 → 3

customer_id
  Unique values: 3 → 4

Drift methodology

The current release uses a lightweight heuristic for numeric drift.

It compares relative changes in selected summary statistics:

  • Mean
  • Median
  • P95

A configurable percentage threshold is then used to determine whether the tracked statistics changed beyond the configured level.

This is intentionally simple and lightweight.

It is not intended to replace formal statistical distribution-drift tests.


Supported input formats

Python API

pandas.DataFrame

CLI

CSV
Parquet

Development

Clone the repository:

git clone https://github.com/rohesen/datadiff-engine.git
cd datadiff-engine

Create a virtual environment:

python -m venv .venv

Activate it on Windows:

.venv\Scripts\activate

Install the project with development dependencies:

pip install -e ".[dev]"

Run the test suite:

pytest

Expected result for the current development checkpoint:

14 passed

Building the package

Install the build tools:

pip install build twine

Build the package:

python -m build

This creates distribution files inside:

dist/

Validate them:

python -m twine check dist/*

Project structure

datadiff-engine/
│
├── src/
│   └── datadiff_engine/
│       ├── __init__.py
│       ├── compare.py
│       └── cli.py
│
├── tests/
│   └── test_compare.py
│
├── demo.py
├── README.md
├── LICENSE
├── pyproject.toml
└── .gitignore

Roadmap

Future versions may include:

  • More advanced statistical drift methods
  • Better categorical drift analysis
  • Datetime-aware profiling
  • Polars support
  • DuckDB integration
  • Rich terminal output
  • Markdown reports
  • Additional CI/CD integrations
  • More configurable analysis rules

Contributing

Contributions, ideas, bug reports, and improvements are welcome.

Before submitting a change, please run:

pytest

License

MIT License.

See LICENSE for details.


Author

Rohesen Rajkamal Maurya

Built as an open-source data-engineering utility for comparing datasets and understanding changes in pipeline outputs.

Metadata

Release files for datadiff-engine 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for datadiff-engine 0.1.0
File Size Uploaded
datadiff_engine-0.1.0.tar.gz 13.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for datadiff-engine 0.1.0
File Interpreter ABI Platform
datadiff_engine-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 23.5 kB

Release files / datadiff_engine-0.1.0.tar.gz

Download URL datadiff_engine-0.1.0.tar.gz
Size 13.4 kB
Tags Source
SHA-256 checksum
How to use checksums
97b11aa51a89028255bd8a11b14329617c638a42355e2ad3e5fbbcf623def52b
BLAKE2b-256 checksum
How to use checksums
79b96b052087d6d37f477e35a67520cd2147a5716bf9f985c3382dccbc2d5944
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release files / datadiff_engine-0.1.0-py3-none-any.whl

Download URL datadiff_engine-0.1.0-py3-none-any.whl
Size 10.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
2847c902a0d3e7abfc2e98ecda2dcc20cc1fd271db49d92f27632d4b5ea4e337
BLAKE2b-256 checksum
How to use checksums
87bf3dc73fcb69e5e058f5601f0238ada345fa9d20dce9dc407618be5e190732
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page