datadiff-engine
A Python library and CLI for comparing datasets, detecting schema changes, profiling columns, and identifying simple numeric data drift.
Built for data engineers who want a quick answer to:
What changed between yesterday's dataset and today's dataset?
Why datadiff-engine?
Data pipelines frequently produce datasets that look valid but have unexpected changes:
- Row counts suddenly increase or decrease
- Columns are added or removed
- Data types change
- Null rates increase
- Unique values change
- Numeric distributions shift
datadiff-engine provides a structured comparison so these changes can be inspected programmatically or directly from the command line.
Features
- Compare row counts
- Detect added columns
- Detect removed columns
- Detect data-type changes
- Compare null rates
- Compare unique-value counts
- Generate numeric statistics
- Detect simple numeric drift
- Analyze categorical columns
- Configure null-rate severity thresholds
- Ignore selected columns for numeric drift analysis
- Python API
- Command-line interface
- JSON output
- CSV support
- Parquet support
Installation
pip install datadiff-engine
Python usage
import pandas as pd
from datadiff_engine import compare
old = pd.DataFrame({
"customer_id": [1, 2, 3],
"amount": [100, 200, 300],
})
new = pd.DataFrame({
"customer_id": [1, 2, 3, 4],
"amount": [100, 200, None, 500],
"country": ["IN", "IN", "US", "IN"],
})
diff = compare(old, new)
print(diff)
Example output:
DATASET DIFF
========================================
ROWS
3 → 4
Change: +1 (+33.33%)
SCHEMA
+ Added: ['country']
- Removed: none
~ Changed: amount (int64 → float64)
COLUMN CHANGES
----------------------------------------
amount
Data type: int64 → float64
Null rate: 0.00% → 25.00% (+25.00 pp)
Severity: CRITICAL
Statistics
Mean: 200.00 → 266.67
Median: 200.00 → 200.00
Min: 100.00 → 100.00
Max: 300.00 → 500.00
P95: 290.00 → 470.00
Column-level analysis
Access details for an individual column:
amount = diff.column("amount")
print(amount.name)
print(amount.old_dtype)
print(amount.new_dtype)
print(amount.null_rate_old)
print(amount.null_rate_new)
print(amount.null_rate_change)
print(amount.severity)
For numeric columns, statistics are available:
print(amount.old_stats.mean)
print(amount.new_stats.mean)
print(amount.old_stats.median)
print(amount.new_stats.median)
print(amount.old_stats.min)
print(amount.new_stats.max)
print(amount.old_stats.p95)
print(amount.new_stats.p95)
Numeric drift information is also available:
if amount.drift:
print(amount.drift.mean_change_pct)
print(amount.drift.median_change_pct)
print(amount.drift.p95_change_pct)
print(amount.drift.has_drift)
Custom severity thresholds
The default null-rate thresholds can be customized:
from datadiff_engine import DiffConfig, compare
config = DiffConfig(
warning_null_rate=0.10,
critical_null_rate=0.50,
)
diff = compare(
old,
new,
config=config,
)
This allows different projects or pipelines to define their own thresholds.
Ignoring columns for numeric drift
Some numeric columns represent identifiers rather than measurements.
For example:
customer_id
order_id
account_id
These columns may not be appropriate for numeric drift analysis.
You can exclude them:
from datadiff_engine import DiffConfig, compare
config = DiffConfig(
ignored_columns={"customer_id"},
)
diff = compare(
old,
new,
config=config,
)
The column is still part of the comparison, but numeric drift analysis is skipped for the ignored column.
JSON output
The comparison can be converted into a machine-readable dictionary:
result = diff.to_dict()
print(result)
This is useful when integrating the library into:
- Data pipelines
- Airflow tasks
- CI/CD pipelines
- Logging systems
- APIs
- Monitoring tools
Example:
{
"rows": {
"old": 3,
"new": 4,
"change": 1,
"change_pct": 33.33
},
"schema": {
"added": ["country"],
"removed": [],
"changed_types": {
"amount": ["int64", "float64"]
}
}
}
Command-line interface
datadiff-engine also provides a CLI.
Compare two CSV files:
datadiff old.csv new.csv
Compare Parquet files:
datadiff old.parquet new.parquet
Output JSON:
datadiff old.csv new.csv --json
Ignore columns for numeric drift:
datadiff old.csv new.csv --ignore customer_id
Show help:
datadiff --help
Example
Suppose yesterday's dataset contains:
customer_id,amount,country
1,100,IN
2,200,IN
3,300,US
and today's dataset contains:
customer_id,amount,country
1,100,IN
2,200,US
3,,US
4,500,BD
Running:
datadiff old.csv new.csv
can identify:
Rows: 3 → 4
amount
Data type: int64 → float64
Null rate: 0.00% → 25.00%
country
Unique values: 2 → 3
customer_id
Unique values: 3 → 4
Drift methodology
The current release uses a lightweight heuristic for numeric drift.
It compares relative changes in selected summary statistics:
- Mean
- Median
- P95
A configurable percentage threshold is then used to determine whether the tracked statistics changed beyond the configured level.
This is intentionally simple and lightweight.
It is not intended to replace formal statistical distribution-drift tests.
Supported input formats
Python API
pandas.DataFrame
CLI
CSV
Parquet
Development
Clone the repository:
git clone https://github.com/rohesen/datadiff-engine.git
cd datadiff-engine
Create a virtual environment:
python -m venv .venv
Activate it on Windows:
.venv\Scripts\activate
Install the project with development dependencies:
pip install -e ".[dev]"
Run the test suite:
pytest
Expected result for the current development checkpoint:
14 passed
Building the package
Install the build tools:
pip install build twine
Build the package:
python -m build
This creates distribution files inside:
dist/
Validate them:
python -m twine check dist/*
Project structure
datadiff-engine/
│
├── src/
│ └── datadiff_engine/
│ ├── __init__.py
│ ├── compare.py
│ └── cli.py
│
├── tests/
│ └── test_compare.py
│
├── demo.py
├── README.md
├── LICENSE
├── pyproject.toml
└── .gitignore
Roadmap
Future versions may include:
- More advanced statistical drift methods
- Better categorical drift analysis
- Datetime-aware profiling
- Polars support
- DuckDB integration
- Rich terminal output
- Markdown reports
- Additional CI/CD integrations
- More configurable analysis rules
Contributing
Contributions, ideas, bug reports, and improvements are welcome.
Before submitting a change, please run:
pytest
License
MIT License.
See LICENSE for details.
Author
Rohesen Rajkamal Maurya
Built as an open-source data-engineering utility for comparing datasets and understanding changes in pipeline outputs.
Metadata
Release files for datadiff-engine 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| datadiff_engine-0.1.0.tar.gz | 13.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| datadiff_engine-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 23.5 kB
Release files / datadiff_engine-0.1.0.tar.gz
| Download URL | datadiff_engine-0.1.0.tar.gz |
|---|---|
| Size | 13.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
97b11aa51a89028255bd8a11b14329617c638a42355e2ad3e5fbbcf623def52b
|
|
BLAKE2b-256 checksum How to use checksums |
79b96b052087d6d37f477e35a67520cd2147a5716bf9f985c3382dccbc2d5944
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency logRelease files / datadiff_engine-0.1.0-py3-none-any.whl
| Download URL | datadiff_engine-0.1.0-py3-none-any.whl |
|---|---|
| Size | 10.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
2847c902a0d3e7abfc2e98ecda2dcc20cc1fd271db49d92f27632d4b5ea4e337
|
|
BLAKE2b-256 checksum How to use checksums |
87bf3dc73fcb69e5e058f5601f0238ada345fa9d20dce9dc407618be5e190732
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Oct 3, 2026.
Transparency log