Building a Python Library for Detecting and Preventing Data Leakage
Project description
mleakdetect: Data Leakage Detection Toolkit
A Python toolkit for detecting and measuring data leakage in machine learning failure prediction tasks, implementing methodologies from the ICPE 2025 paper "Quantifying Data Leakage in Failure Prediction Tasks".
Table of Contents
- Overview
- Features
- Installation
- Quick Start
- Core Modules
- Usage Examples
- Testing
- Project Structure
- License
Overview
Data leakage is a critical issue in machine learning that occurs when information from the test set inadvertently influences the training process. This toolkit provides:
- Quantitative measurement of data leakage using the temporal leakage metric (L̃) from the ICPE 2025 paper
- Temporal and group-based splitting strategies to prevent leakage
- Duplicate detection (exact and near-duplicates) that can cause contamination
- Identifier overlap analysis between train and test sets
- Automated analysis pipeline with comprehensive reporting
The implementation is designed to be modular, extensible, and easy to integrate into existing ML pipelines.
Features
Core Functionality
- Leakage Measurement: Compute normalized leakage scores using exponential decay-based temporal similarity
- Temporal Data Splitting: Split time-series data chronologically with configurable time gaps
- Group-based Splitting: Prevent group contamination by keeping entities (e.g., devices, users) separate
- Duplicate Detection: Find and remove exact and near-duplicate rows
- Overlap Analysis: Identify shared identifiers between training and test sets
- Automated Reporting: Generate comprehensive analysis reports
Design Principles
- Modular Architecture: Each component can be used independently
- NumPy-style Documentation: Complete docstrings following scientific Python standards
- Comprehensive Testing: 48 unit tests covering all core functionality
- pip-installable: Standard Python packaging with pyproject.toml
Prerequisites
- Install Python 3.12. We recommend to use a virtual environment (e.g., with conda) to avoid conflicts with other Python projects.
- pip install mleakdetect
Installation
This repository uses source/ as the package root (it contains pyproject.toml).
Run the following commands inside the source/ directory.
Option 1: Install via pip (recommended)
If you just want to use the toolkit, install the latest released version:
pip install mleakdetect
Option 2: Install from source (for development)
If you want to run the latest code or modify the package locally:
git clone https://github.com/song_min_0111/mleakdetect.git
cd mleakdetect
pip install .
Basic Workflow
import pandas as pd
import mleakdetect as mld
def main() -> None:
df = pd.read_csv("examples/Bitcoin 2024.csv")
result = mld.analyze_dataset(
df=df,
split_mode="temporal",
split_type="train_test",
time_column="Date",
id_columns=None,
alpha=1.0,
similarity_threshold=0.99,
remove_exact_dups=True,
remove_near_dups=False,
)
mld.print_leakage_report(result)
mld.save_report_to_file(result, "leakage_report.md")
print("Saved: leakage_report.md")
Core Modules
1. Leakage Measurement (mleakdetect.measure)
Implements the temporal leakage metric L̃ from the ICPE 2025 paper:
from mleakdetect.core.measure import compute_data_leakage
leakage = compute_data_leakage(
train_df=train,
test_df=test,
time_column='date',
alpha=1.0, # Weight for backward temporal leakage
group_column=None # Optional: for group-aware measurement
)
Interpretation:
L̃ = 0.0: No leakage (clean split)L̃ = 1.0: Maximum leakage (train == test)L̃ ∈ (0, 1): Partial leakage
2. Data Splitting (mleakdetect.split)
Temporal Split
Preserves chronological order to prevent temporal leakage:
from mleakdetect.core.split import temporal_split
train, test = temporal_split(
df,
time_column='timestamp',
train_ratio=0.7,
gap='7D' # Time gap between train and test
)
Group-based Split
Keeps groups (e.g., devices, users) separate across splits:
from mleakdetect.core.split import group_split
train, test = group_split(
df,
group_column='device_id',
train_ratio=0.7,
random_state=42
)
3. Duplicate Detection (mleakdetect.duplicates)
Exact Duplicates
from mleakdetect.core.duplicates import find_exact_duplicates, remove_exact_duplicates
# Find duplicates
duplicates, summary = find_exact_duplicates(df)
print(f"Found {summary.exact_duplicate_rows} duplicates")
# Remove duplicates
clean_df = remove_exact_duplicates(df, keep='first')
Near-duplicates
from mleakdetect.core.duplicates import find_near_duplicates
# Find near-duplicates using cosine similarity
pairs, summary = find_near_duplicates(
df,
threshold=0.95, # Similarity threshold
subset=['feature1', 'feature2'] # Columns to compare
)
4. Identifier Overlap Analysis (mleakdetect.overlap)
from mleakdetect.core.overlap import compute_id_overlap
overlap_df = compute_id_overlap(
train_df=train,
test_df=test,
id_columns=['device_id', 'user_id']
)
print(overlap_df)
# column unique_train unique_test overlap_count overlap_percentage_test
# 0 device_id 70 30 15 50.0
Usage Examples
Example 1: Bitcoin Price Dataset (Single Time Series)
import pandas as pd
from mleakdetect import temporal_split, compute_data_leakage
# Load Bitcoin price data
df = pd.read_csv('Bitcoin 2024.csv')
df['date'] = pd.to_datetime(df['date'])
# gernerate report
result = analyze_dataset(
df,
split_mode="temporal",
time_column="Date",
id_columns=None,
alpha=1.0,
similarity_threshold=0.99,
)
# Display report in notebook
print_leakage_report(result)
save_report_to_file(result, "bitcoin_leakage_report.md")
Example 2: HDD Failure Prediction (Multi-group)
from mleakdetect import group_split, compute_data_leakage
# Load HDD failure data (Backblaze dataset)
df = pd.read_parquet('hdd_data.parquet')
df['date'] = pd.to_datetime(df['date'])
# gernerate report
result = analyze_dataset(
df,
split_mode="group",
time_column=date_col,
group_column=group_col,
id_columns=["serial_number"],
alpha=1.0,
similarity_threshold=0.999,
)
print_leakage_report(result)
save_report_to_file(result, "hdd_leakage_report.md")
Testing
The package includes comprehensive unit tests covering all core functionality.
Run all tests
# Simple
pytest
Test Coverage
Current test coverage: 48 tests, 100% pass rate
- Core leakage measurement (6 tests)
- Temporal and group splitting (8 tests)
- Duplicate detection (8 tests)
- Identifier overlap analysis (8 tests)
- Input validation (18 tests)
Project Structure
source/
├── mleakdetect/ # Main package
│ ├── __init__.py # Package initialization & exports
│ ├── core/ # Core functionality modules
│ │ ├── __init__.py
│ │ ├── measure.py # Leakage measurement
│ │ ├── split.py # Data splitting utilities
│ │ ├── duplicates.py # Duplicate detection
│ │ ├── overlap.py # ID overlap analysis
│ │ ├── analysis.py # High-level orchestrator
│ │ └── report.py # Report generation
│ └── utils/ # Utility modules
│ ├── __init__.py
│ └── validation.py # Input validation
├── tests/ # Unit tests
│ ├── __init__.py
│ ├── test_measure.py
│ ├── test_split.py
│ ├── test_duplicates.py
│ ├── test_overlap.py
│ └── test_validation.py
├── examples/ # Usage examples
│ ├── example_notebook_bitcoin.ipynb
│ └── example_notebook_hdd.ipynb
├── pyproject.toml # Package configuration
├── README.md # This file
└── LICENSE # MIT License
API Reference
Main Functions
| Function | Module | Description |
|---|---|---|
compute_data_leakage() |
measure |
Compute normalized leakage score |
temporal_split() |
split |
Split data chronologically |
group_split() |
split |
Split data by groups |
find_exact_duplicates() |
duplicates |
Find exact duplicate rows |
find_near_duplicates() |
duplicates |
Find near-duplicate rows |
remove_exact_duplicates() |
duplicates |
Remove exact-duplicate rows |
remove_near_duplicates() |
duplicates |
Remove near-duplicate rows |
compute_id_overlap() |
overlap |
Analyze ID overlap |
analyze_dataset() |
analysis |
Full analysis pipeline |
print_leakage_report() |
report |
Print analysis result |
Parameters Guide
Alpha Parameter (α)
Controls the weight of backward temporal leakage:
- α = 0: Only future leakage counts (strict temporal-only task)
- α = 1: Past and future equally weighted (general failure prediction)
- α > 1: Emphasize group contamination over temporal leakage
Time Gap
Specifies the time buffer between train and test:
gap='0D' # No gap
gap='1D' # 1 day
gap='7D' # 1 week
gap='1W' # 1 week (alternative)
gap='12h' # 12 hours
gap='30min' # 30 minutes
License
This project is licensed under the MIT License - see the LICENSE file for details.
Acknowledgments
- Based on methodologies from the ICPE 2025 paper "Quantifying Data Leakage in Failure Prediction Tasks"
- Developed at Wurzburg University in Germany
- Supervisor: Daniel Grillmeyer
Contact
For questions or issues, please:
- Open an issue on GitHub
Changelog
Version 0.1.0 (2025-02-16)
- Initial release
- Core leakage measurement implementation
- Temporal and group-based splitting
- Duplicate detection (exact and near)
- ID overlap analysis
- Automated analysis pipeline
- Comprehensive test suite (48 tests)
- Complete documentation
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file mleakdetect-0.1.0.tar.gz.
File metadata
- Download URL: mleakdetect-0.1.0.tar.gz
- Upload date:
- Size: 32.4 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
8caba2b96b6e8a01482930be5575354b230bd541dbd2de5106d692c139840b77
|
|
| MD5 |
578e395f8d05b3a1d0d7556fa6e21db1
|
|
| BLAKE2b-256 |
06c62cdf21084d82b9a35899e6b561e0dca46556e8fac9cd7a68a96f7a934925
|
File details
Details for the file mleakdetect-0.1.0-py3-none-any.whl.
File metadata
- Download URL: mleakdetect-0.1.0-py3-none-any.whl
- Upload date:
- Size: 35.7 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: twine/6.2.0 CPython/3.12.3
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
56a9c3503545f19bc7f104dce15ed06e92681ccc196685240adf601e75308166
|
|
| MD5 |
bcb374127c89f350e233731c9acb6b19
|
|
| BLAKE2b-256 |
4572c9aa759d45a7188a7052112b6cf1809acf6ac7538420a9121fef7c47bab9
|