Skip to main content

PreventLeak

Measure and prevent data leakage in machine learning evaluation.

Most tools tell you whether leakage is present. PreventLeak also tells you how much of your reported score is leakage and what the honest number is, with a calibrated confidence interval, and gives you leak-safe splits, folds, and cleaned data to remove it. It works on any tabular feature matrix and any scikit-learn compatible model, and it covers the data-resident leakage (duplicate, group, temporal) that code-static analyzers cannot see.

  • Measure. For every leakage channel, the optimism gap, the corrected (leak-free) score, a 95% confidence interval, and a significance test.
  • Prevent. Leak-safe train/test splitting, cross-validation, duplicate removal, and a pipeline that fits every step on the training rows only.
  • Model-flexible. Pass one model, several, or none (a documented default panel).
  • Any modality. Tabular, and anything represented as features (text, image, time series).

Table of contents


Installation

pip install preventleak

Requires Python >= 3.9 and numpy, pandas, scikit-learn, scipy, matplotlib (installed automatically).


Quickstart

Measure the leakage optimism gap

from preventleak import PreventLeak
from sklearn.ensemble import RandomForestClassifier

report = PreventLeak(RandomForestClassifier()).audit(
    X, y,
    channels=["scaling", "feature_selection", "mean_encoding"],
    cat_cols=[3],          # high-cardinality categorical -> target-encoding leakage
    group=patient_id,      # optional entity key -> group leakage
    time=timestamp,        # optional time order -> temporal leakage
)

print(report.summary())    # per-channel reported, corrected, gap, 95% CI, p
report.plot(save="audit.png")
records = report.to_dict() # machine-readable list of dicts

Example summary() output:

PreventLeak audit
channel                reported  Delta_hat         gap 95% CI corrected        p
mean_encoding[3]         0.7966     0.0089 [-0.0251, 0.0430]    0.7877    0.012
feature_selection        0.7744     0.0021 [-0.0365, 0.0408]    0.7723    0.695
scaling                  0.7943     0.0000 [-0.0345, 0.0346]    0.7943    0.746
-> largest optimism from 'mean_encoding[3]': reported 0.7966, honest 0.7877 ...

Prevent it

from preventleak import safe_split, safe_cv, clean, LeakSafePipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier

train_idx, test_idx = safe_split(X, y, group=patient_id, time=timestamp)

for tr, te in safe_cv(X, y, n_splits=5, group=patient_id):
    ...

X_clean, y_clean, kept = clean(X, y)   # drop cross-split duplicate rows

model = LeakSafePipeline(
    steps=[("scale", StandardScaler())],
    estimator=RandomForestClassifier(),
).fit(X[train_idx], y[train_idx])      # every step fit on train only

Concepts

The optimism gap

Let S_leaky be the score a practitioner reports under a protocol that contains a leak, and S_clean the score of the same model under a leak-free protocol that performs every data-dependent step inside the training portion of each fold. The leakage-induced optimism gap is

Delta = S_leaky - S_clean

and the corrected score is S_clean = S_leaky - Delta. PreventLeak estimates Delta for your specific dataset and pipeline, not its average over a benchmark, by running your model under both protocols on the same cross-validation folds and taking the paired difference.

Leakage channels

A channel defines how the leaky and clean arms differ. Six are supported.

channel leaky arm clean arm discovered
scaling standardizer fit on all rows fit inside the fold from pipeline
feature_selection features ranked on all rows ranked inside the fold from pipeline
mean_encoding category means on all rows means inside the fold from pipeline
duplicate duplicate rows straddle the split duplicates kept together automatic
group same entity on both sides entity kept on one side from group= key
temporal train on past and future train only on the past from time= order

Duplicate contamination is discovered without supervision; group and temporal channels are activated when you pass a key or a time order; preprocessing channels are activated for the steps you list.

The confidence interval

Cross-validation folds share training rows, so the naive variance of the fold scores is too small. PreventLeak uses the Nadeau-Bengio corrected resampled-t variance, which inflates the naive variance by the test-to-train ratio rho = 1/(k-1) for k-fold, so the interval reflects the true uncertainty of a cross-validation estimate. Significance of the gap is a Wilcoxon signed-rank test on the paired fold scores.

An honest caveat: cross-validation pessimism

The corrected score is a leak-free cross-validation estimate, so it carries the usual cross-validation training-size pessimism: it estimates the performance of a model trained on (k-1)/k of the data, which is slightly below a model trained on all of it. This effect is a property of cross-validation, not of leakage; it shrinks as the fold count rises, and PreventLeak does not count it as leakage.


API reference

PreventLeak(model=None, k_folds=5, repeats=3, seed0=0, models=None)

Audit a dataset for leakage-induced optimism.

  • model - a single scikit-learn compatible estimator. Leave None to use a documented default panel (RandomForest(200), ExtraTrees(200), HistGradientBoosting).
  • models - a dict {name: estimator} or a list of estimators to audit several at once; the report then carries a model column and plot() facets per model.
  • k_folds, repeats - the repeated stratified k-fold design (default 5 x 3).
  • seed0 - base random seed.

.audit(X, y, channels=None, cat_cols=None, feature_k=10, group=None, time=None, detect_duplicates=True) -> AuditReport

  • X, y - array-like features and binary target.
  • channels - any of "scaling", "feature_selection", "mean_encoding".
  • cat_cols - column indices for mean_encoding (one channel per column).
  • feature_k - number of features kept by feature_selection.
  • group - per-row entity key; activates the group channel.
  • time - per-row order or timestamp; activates the temporal channel.
  • detect_duplicates - auto-discover duplicate contamination (default True).

AuditReport

Returned by audit. Fields and methods:

  • .results - list of ChannelResult (channel, detected, gap, note, model).
  • .summary() -> str - a formatted per-channel table.
  • .to_dict() -> list[dict] - one record per channel with reported, corrected, gap, ci_low, ci_high, p_value, model.
  • .worst() -> ChannelResult - the channel with the largest gap.
  • .plot(ax=None, alpha=0.05, title=..., save=None) - a forest plot of the per-channel gaps with 95% intervals; filled marks are significant. With several models it becomes small multiples.

Each gap is a GapEstimate with reported, corrected, delta_hat, ci_low, ci_high, wilcoxon_p, n_pairs.

estimate_gap

estimate_gap(X, y, model, leak, groups=None, k_folds=5, repeats=5,
             seed0=0, rng_seed=0, metric="auc") -> GapEstimate

Low-level estimator for a single LeakSpec. metric is one of "auc", "acc", "f1", "prauc", "mcc".

LeakSpec

LeakSpec(kind, k=10, cat_col=-1, time=None)

Specifies one channel. kind is one of the six channel names; k is the feature count for feature_selection; cat_col is the column for mean_encoding; time is the per-row order for temporal.

safe_split

safe_split(X, y=None, test_size=0.25, group=None, time=None, dedup=True,
           stratify=True, random_state=0, decimals=6) -> (train_idx, test_idx)

A leak-free train/test split. A time order gives a forward split (train earlier, test later); otherwise a group key or detected duplicates give a group-aware split that never places rows of one group on both sides; otherwise a stratified split.

safe_cv

safe_cv(X, y=None, n_splits=5, group=None, time=None, dedup=True,
        random_state=0, decimals=6)  # yields (train_idx, test_idx)

Leak-free fold iterator: forward-chaining folds for time, group-disjoint folds for a group key or detected duplicates, otherwise stratified folds.

clean

clean(X, y=None, dedup=True, decimals=6)
# -> (X_clean, y_clean, keep_mask)  or  (X_clean, keep_mask) when y is None

Drop exact or near-duplicate rows, keeping the first occurrence of each.

LeakSafePipeline

LeakSafePipeline(steps=None, estimator=None, dedup=True, decimals=6)

A thin wrapper that deduplicates the training rows and fits every transformer and the estimator on the training data only. steps is a list of (name, transformer); use with safe_cv / safe_split for a fully leak-free evaluation. Provides fit, predict, predict_proba.

auto_groups and detect_cross_split_duplicates

auto_groups(X, decimals=6) -> (group_ids, n_duplicate_rows)
detect_cross_split_duplicates(X_train, X_val, decimals=6) -> boolean_mask

auto_groups assigns a shared id to exact/near-duplicate rows; the second flags validation rows whose match appears in the training portion.


Command-line interface

Works on any CSV. Categorical columns are encoded automatically; the highest-cardinality one is audited for target-encoding leakage.

# audit, save a plot and a JSON report
preventleak audit data.csv --target label --plot audit.png --json audit.json

# one or more models (shortcuts rf/hgb/et/logreg, or any dotted import path)
preventleak audit data.csv --target label --model rf logreg sklearn.ensemble.GradientBoostingClassifier

# with an entity key and a time column
preventleak audit data.csv --target label --group patient_id --time visit_date

# remove duplicate rows, or write a leak-safe split
preventleak clean data.csv --target label -o clean.csv
preventleak split data.csv --target label --time ts --out-prefix split

audit options: --model (one or more), --folds, --repeats, --feature-k, --group, --time, --plot, --json.


Reproducibility

The estimator is deterministic given the seeds (seed0, rng_seed). The default model panel is fixed (RandomForest(200), ExtraTrees(200), HistGradientBoosting) so a no-model audit is reproducible. Duplicate discovery rounds features to decimals places before hashing.


Limitations

  • The corrected score inherits cross-validation training-size pessimism (see the caveat above); it is reported separately and is not counted as leakage.
  • Group and temporal channels require you to supply an entity key or a time order; inferring them from features alone is not always possible.
  • Preprocessing channels are audited for the steps you declare.
  • Targets are treated as binary; multi-class support is on the roadmap.

Citation

If you use PreventLeak, please cite:

M. A. Bouke, A. Abdullah, N. I. Udzir, N. Samian, M. Othman. Quantifying the Optimism Gap from Data Leakage in Machine Learning Evaluation.


License

MIT.

Metadata

Release files for preventleak 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for preventleak 0.1.0
File Size Uploaded
preventleak-0.1.0.tar.gz 23.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for preventleak 0.1.0
File Interpreter ABI Platform
preventleak-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 44.4 kB

Release files / preventleak-0.1.0.tar.gz

Download URL preventleak-0.1.0.tar.gz
Size 23.4 kB
Tags Source
SHA-256 checksum
How to use checksums
a72cd0854e8da7ad0e08e284e1c4422bf781d458760238190d9d254f146b3abf
BLAKE2b-256 checksum
How to use checksums
794abfeb3ebebc6d1dfdd60f2089faf2e7daf65d1f2150525cfcced91e1b36ba
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release files / preventleak-0.1.0-py3-none-any.whl

Download URL preventleak-0.1.0-py3-none-any.whl
Size 21.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
549e5a2f653e09d6154be66dfda6a972433b7024abb43a43a96f1b04868d4cb7
BLAKE2b-256 checksum
How to use checksums
1b4ebb6f9986905e075498abcc9dd8494006a0bd4926b689da760cfa35992238
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.14.0

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page