PreventLeak
Measure and prevent data leakage in machine learning evaluation.
Most tools tell you whether leakage is present. PreventLeak also tells you how much of your reported score is leakage and what the honest number is, with a calibrated confidence interval, and gives you leak-safe splits, folds, and cleaned data to remove it. It works on any tabular feature matrix and any scikit-learn compatible model, and it covers the data-resident leakage (duplicate, group, temporal) that code-static analyzers cannot see.
- Measure. For every leakage channel, the optimism gap, the corrected (leak-free) score, a 95% confidence interval, and a significance test.
- Prevent. Leak-safe train/test splitting, cross-validation, duplicate removal, and a pipeline that fits every step on the training rows only.
- Model-flexible. Pass one model, several, or none (a documented default panel).
- Any modality. Tabular, and anything represented as features (text, image, time series).
Table of contents
- Installation
- Quickstart
- Concepts
- API reference
- Command-line interface
- Reproducibility
- Limitations
- Citation
- License
Installation
pip install preventleak
Requires Python >= 3.9 and numpy, pandas, scikit-learn, scipy, matplotlib
(installed automatically).
Quickstart
Measure the leakage optimism gap
from preventleak import PreventLeak
from sklearn.ensemble import RandomForestClassifier
report = PreventLeak(RandomForestClassifier()).audit(
X, y,
channels=["scaling", "feature_selection", "mean_encoding"],
cat_cols=[3], # high-cardinality categorical -> target-encoding leakage
group=patient_id, # optional entity key -> group leakage
time=timestamp, # optional time order -> temporal leakage
)
print(report.summary()) # per-channel reported, corrected, gap, 95% CI, p
report.plot(save="audit.png")
records = report.to_dict() # machine-readable list of dicts
Example summary() output:
PreventLeak audit
channel reported Delta_hat gap 95% CI corrected p
mean_encoding[3] 0.7966 0.0089 [-0.0251, 0.0430] 0.7877 0.012
feature_selection 0.7744 0.0021 [-0.0365, 0.0408] 0.7723 0.695
scaling 0.7943 0.0000 [-0.0345, 0.0346] 0.7943 0.746
-> largest optimism from 'mean_encoding[3]': reported 0.7966, honest 0.7877 ...
Prevent it
from preventleak import safe_split, safe_cv, clean, LeakSafePipeline
from sklearn.preprocessing import StandardScaler
from sklearn.ensemble import RandomForestClassifier
train_idx, test_idx = safe_split(X, y, group=patient_id, time=timestamp)
for tr, te in safe_cv(X, y, n_splits=5, group=patient_id):
...
X_clean, y_clean, kept = clean(X, y) # drop cross-split duplicate rows
model = LeakSafePipeline(
steps=[("scale", StandardScaler())],
estimator=RandomForestClassifier(),
).fit(X[train_idx], y[train_idx]) # every step fit on train only
Concepts
The optimism gap
Let S_leaky be the score a practitioner reports under a protocol that contains a
leak, and S_clean the score of the same model under a leak-free protocol that
performs every data-dependent step inside the training portion of each fold. The
leakage-induced optimism gap is
Delta = S_leaky - S_clean
and the corrected score is S_clean = S_leaky - Delta. PreventLeak estimates
Delta for your specific dataset and pipeline, not its average over a benchmark, by
running your model under both protocols on the same cross-validation folds and
taking the paired difference.
Leakage channels
A channel defines how the leaky and clean arms differ. Six are supported.
| channel | leaky arm | clean arm | discovered |
|---|---|---|---|
scaling |
standardizer fit on all rows | fit inside the fold | from pipeline |
feature_selection |
features ranked on all rows | ranked inside the fold | from pipeline |
mean_encoding |
category means on all rows | means inside the fold | from pipeline |
duplicate |
duplicate rows straddle the split | duplicates kept together | automatic |
group |
same entity on both sides | entity kept on one side | from group= key |
temporal |
train on past and future | train only on the past | from time= order |
Duplicate contamination is discovered without supervision; group and temporal channels are activated when you pass a key or a time order; preprocessing channels are activated for the steps you list.
The confidence interval
Cross-validation folds share training rows, so the naive variance of the fold scores
is too small. PreventLeak uses the Nadeau-Bengio corrected resampled-t variance,
which inflates the naive variance by the test-to-train ratio rho = 1/(k-1) for
k-fold, so the interval reflects the true uncertainty of a cross-validation estimate.
Significance of the gap is a Wilcoxon signed-rank test on the paired fold scores.
An honest caveat: cross-validation pessimism
The corrected score is a leak-free cross-validation estimate, so it carries the usual
cross-validation training-size pessimism: it estimates the performance of a model
trained on (k-1)/k of the data, which is slightly below a model trained on all of
it. This effect is a property of cross-validation, not of leakage; it shrinks as the
fold count rises, and PreventLeak does not count it as leakage.
API reference
PreventLeak(model=None, k_folds=5, repeats=3, seed0=0, models=None)
Audit a dataset for leakage-induced optimism.
- model - a single scikit-learn compatible estimator. Leave
Noneto use a documented default panel (RandomForest(200), ExtraTrees(200), HistGradientBoosting). - models - a dict
{name: estimator}or a list of estimators to audit several at once; the report then carries a model column andplot()facets per model. - k_folds, repeats - the repeated stratified k-fold design (default 5 x 3).
- seed0 - base random seed.
.audit(X, y, channels=None, cat_cols=None, feature_k=10, group=None, time=None, detect_duplicates=True) -> AuditReport
- X, y - array-like features and binary target.
- channels - any of
"scaling","feature_selection","mean_encoding". - cat_cols - column indices for
mean_encoding(one channel per column). - feature_k - number of features kept by
feature_selection. - group - per-row entity key; activates the
groupchannel. - time - per-row order or timestamp; activates the
temporalchannel. - detect_duplicates - auto-discover duplicate contamination (default
True).
AuditReport
Returned by audit. Fields and methods:
.results- list ofChannelResult(channel,detected,gap,note,model)..summary() -> str- a formatted per-channel table..to_dict() -> list[dict]- one record per channel withreported,corrected,gap,ci_low,ci_high,p_value,model..worst() -> ChannelResult- the channel with the largest gap..plot(ax=None, alpha=0.05, title=..., save=None)- a forest plot of the per-channel gaps with 95% intervals; filled marks are significant. With several models it becomes small multiples.
Each gap is a GapEstimate with reported, corrected, delta_hat, ci_low,
ci_high, wilcoxon_p, n_pairs.
estimate_gap
estimate_gap(X, y, model, leak, groups=None, k_folds=5, repeats=5,
seed0=0, rng_seed=0, metric="auc") -> GapEstimate
Low-level estimator for a single LeakSpec. metric is one of "auc", "acc",
"f1", "prauc", "mcc".
LeakSpec
LeakSpec(kind, k=10, cat_col=-1, time=None)
Specifies one channel. kind is one of the six channel names; k is the feature
count for feature_selection; cat_col is the column for mean_encoding; time
is the per-row order for temporal.
safe_split
safe_split(X, y=None, test_size=0.25, group=None, time=None, dedup=True,
stratify=True, random_state=0, decimals=6) -> (train_idx, test_idx)
A leak-free train/test split. A time order gives a forward split (train earlier,
test later); otherwise a group key or detected duplicates give a group-aware split
that never places rows of one group on both sides; otherwise a stratified split.
safe_cv
safe_cv(X, y=None, n_splits=5, group=None, time=None, dedup=True,
random_state=0, decimals=6) # yields (train_idx, test_idx)
Leak-free fold iterator: forward-chaining folds for time, group-disjoint folds for
a group key or detected duplicates, otherwise stratified folds.
clean
clean(X, y=None, dedup=True, decimals=6)
# -> (X_clean, y_clean, keep_mask) or (X_clean, keep_mask) when y is None
Drop exact or near-duplicate rows, keeping the first occurrence of each.
LeakSafePipeline
LeakSafePipeline(steps=None, estimator=None, dedup=True, decimals=6)
A thin wrapper that deduplicates the training rows and fits every transformer and the
estimator on the training data only. steps is a list of (name, transformer);
use with safe_cv / safe_split for a fully leak-free evaluation. Provides fit,
predict, predict_proba.
auto_groups and detect_cross_split_duplicates
auto_groups(X, decimals=6) -> (group_ids, n_duplicate_rows)
detect_cross_split_duplicates(X_train, X_val, decimals=6) -> boolean_mask
auto_groups assigns a shared id to exact/near-duplicate rows; the second flags
validation rows whose match appears in the training portion.
Command-line interface
Works on any CSV. Categorical columns are encoded automatically; the highest-cardinality one is audited for target-encoding leakage.
# audit, save a plot and a JSON report
preventleak audit data.csv --target label --plot audit.png --json audit.json
# one or more models (shortcuts rf/hgb/et/logreg, or any dotted import path)
preventleak audit data.csv --target label --model rf logreg sklearn.ensemble.GradientBoostingClassifier
# with an entity key and a time column
preventleak audit data.csv --target label --group patient_id --time visit_date
# remove duplicate rows, or write a leak-safe split
preventleak clean data.csv --target label -o clean.csv
preventleak split data.csv --target label --time ts --out-prefix split
audit options: --model (one or more), --folds, --repeats, --feature-k,
--group, --time, --plot, --json.
Reproducibility
The estimator is deterministic given the seeds (seed0, rng_seed). The default
model panel is fixed (RandomForest(200), ExtraTrees(200), HistGradientBoosting) so a
no-model audit is reproducible. Duplicate discovery rounds features to decimals
places before hashing.
Limitations
- The corrected score inherits cross-validation training-size pessimism (see the caveat above); it is reported separately and is not counted as leakage.
- Group and temporal channels require you to supply an entity key or a time order; inferring them from features alone is not always possible.
- Preprocessing channels are audited for the steps you declare.
- Targets are treated as binary; multi-class support is on the roadmap.
Citation
If you use PreventLeak, please cite:
M. A. Bouke, A. Abdullah, N. I. Udzir, N. Samian, M. Othman. Quantifying the Optimism Gap from Data Leakage in Machine Learning Evaluation.
License
MIT.
Metadata
Release files for preventleak 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| preventleak-0.1.0.tar.gz | 23.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| preventleak-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 44.4 kB
Release files / preventleak-0.1.0.tar.gz
| Download URL | preventleak-0.1.0.tar.gz |
|---|---|
| Size | 23.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
a72cd0854e8da7ad0e08e284e1c4422bf781d458760238190d9d254f146b3abf
|
|
BLAKE2b-256 checksum How to use checksums |
794abfeb3ebebc6d1dfdd60f2089faf2e7daf65d1f2150525cfcced91e1b36ba
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|
Release files / preventleak-0.1.0-py3-none-any.whl
| Download URL | preventleak-0.1.0-py3-none-any.whl |
|---|---|
| Size | 21.0 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
549e5a2f653e09d6154be66dfda6a972433b7024abb43a43a96f1b04868d4cb7
|
|
BLAKE2b-256 checksum How to use checksums |
1b4ebb6f9986905e075498abcc9dd8494006a0bd4926b689da760cfa35992238
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.14.0
|