Skip to main content

leakcheck

tests PyPI Python

Catch data leakage before it ruins your model.

Machine-learning data leakage, not security leaks: leakcheck finds problems in train/test splits that make a model look better than it really is.

A model that scores 99% in testing and fails in production usually has a leak: test rows that were also in training, the same patient on both sides of the split, a feature that secretly contains the answer, or a time series split in the wrong order. leakcheck is a linter for your train/test split. Run it before you train.

$ leakcheck train.csv test.csv --target price_nok

leakcheck  train: 1,600 rows × 7 cols   test: 500 rows × 7 cols

  ✗ 60 test rows (12.0%) also appear in train
  ✗ 62 test rows (12.4%) are near-duplicates of train rows (≥75% of values match) — about 30 more than the 32 expected by chance
  ✗ Column 'final_price_eur' has correlation 1.000 with target 'price_nok'
  ⚠ Column 'listing_id' looks like an ID (unique per row) — drop it before training
  ⚠ 100.0% of test rows on 'listed_date' are not after the latest train date (2024-05-14)

3 problem(s), 2 warning(s) — likely leakage

Install

pip install leakcheck-ml

The package is leakcheck-ml on PyPI; the command and the import are both leakcheck.

Usage

leakcheck train.csv test.csv --target label             # basic
leakcheck train.parquet test.parquet -t label           # parquet (pip install "leakcheck[parquet]")
leakcheck train.csv test.csv -t label --time-col date   # enforce a time-based split
leakcheck train.csv test.csv -t label -g patient_id     # enforce a group split (no patient on both sides)
leakcheck train.csv test.csv -t label --similarity 0.9 # stricter near-duplicate matching
leakcheck train.csv test.csv -t label --json            # machine-readable output
leakcheck train.csv test.csv -t label --strict          # fail on warnings too

The exit code is 1 when problems are found, so you can drop it into CI and block a pipeline that would train on leaky data.

What it checks

Check What it catches How
Duplicates Test rows the model already saw in training Each row is hashed to 64 bits; overlap is a set lookup, so it runs in O(n + m) instead of comparing every pair of rows
Near-duplicates Copies that were slightly changed: re-rounded numbers, new IDs, different capitalisation or spacing MinHash + locality-sensitive hashing find candidate pairs without comparing every row to every other row; candidates are then scored exactly. Numbers are bucketed on two offset grids so tiny changes don't break a match. Only flagged when the count is significantly above a chance baseline (train matched against itself, z ≥ 3), because similar rows are normal in low-dimensional data
Group leakage The same patient, user or customer in both train and test, so the model learns to recognise the person instead of the pattern Checks which test rows belong to groups already seen in train. Pass --group-col, or it auto-detects entity-like columns (user_id, patient…) whose values repeat. Row numbers that restart in each file are ignored because they don't repeat
Target leakage Features that are the target in disguise (e.g. price in another currency, a status set after the outcome) Pearson correlation for numeric features; predictive purity for categorical ones, meaning how much better than the majority-class baseline a feature predicts the target on its own
ID columns Unique-per-row identifiers the model can memorise Uniqueness + dtype + naming heuristics
Temporal leakage Randomly split time series, where the model trains on the future Compares the test dates against the latest training date; date columns are detected automatically

Try it

python examples/make_demo.py       # two demo datasets with planted leaks
leakcheck examples/train.csv examples/test.csv --target price_nok
leakcheck examples/patients_train.csv examples/patients_test.csv --target diagnosis
  ✗ 176 of 183 'patient_id' groups (263 test rows, 93.9%) also appear in train — split by group instead of by row

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest

Roadmap

  • Near-duplicate detection with MinHash + LSH (rows that are almost identical)
  • Group leakage (same patient/user in both train and test)
  • GitHub Action
  • Publish to PyPI
  • HTML report

License

MIT

Release files for leakcheck-ml 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for leakcheck-ml 0.3.0
File Size Uploaded
leakcheck_ml-0.3.0.tar.gz 17.3 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for leakcheck-ml 0.3.0
File Interpreter ABI Platform
leakcheck_ml-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 31.2 kB

Release files / leakcheck_ml-0.3.0.tar.gz

Download URL leakcheck_ml-0.3.0.tar.gz
Size 17.3 kB
Tags Source
SHA-256 checksum
How to use checksums
6ef33a41a914ff053855a22bce714661ba86045265cc44af0ee1505adc748a27
BLAKE2b-256 checksum
How to use checksums
0713dfde0686792ccdc4502a1d7999ac85790c18831b79d0a7046049f4e6c4c2
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release files / leakcheck_ml-0.3.0-py3-none-any.whl

Download URL leakcheck_ml-0.3.0-py3-none-any.whl
Size 14.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
1b212c72723be243aab175e483fbdddbe01d5337fc1bb911d3be2ca0f7b7fdf9
BLAKE2b-256 checksum
How to use checksums
e8407a832122cce0b86fb75958372e7bbe0128b0dcac00fe6a0f88491ca272ce
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 24, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.3.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page