Verascan
Data Contamination & Leakage Detection for AI / ML Workflows
Detect exact, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark integrity, RAG validation, and synthetic data auditing.
🎯 Overview
Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.
Verascan provides a multi-tier contamination detection pipeline:
- Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
- Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
- Semantic match — Dense embedding similarity search (
sentence-transformers+ FAISS) for paraphrased content.
✨ Features
- Multi-Tier Detection: Run exact, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
- Cross-Method Deduplication: Matches identified by exact hashing are automatically excluded from fuzzy/semantic passes to prevent double-counting.
- Multi-Format Ingestion: Natively accepts
pandas.DataFrame,JSONL,CSV, Hugging Facedatasets.Dataset, and Pythonlist[str]. - Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring live search, method filtering, and word-level diffs.
- Cleaned Eval Export: Drop contaminated eval rows and write a reusable CSV/JSONL set (
cleaned_eval(),to_cleaned(), CLI--output-cleaned). - CI/CD Integration: CLI includes
--fail-aboveto fail builds if contamination exceeds an allowed threshold. - Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (
sentence-transformers,faiss-cpu) are optional extras. - Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting
stderr.
🔬 Detection Engines
| Method | Algorithm | Complexity / Speed | Best For |
|---|---|---|---|
exact |
SHA-256 Content Hashing (normalised) | $O(N + M)$ • Microseconds | Verbatim duplicates, casing/whitespace variations |
fuzzy |
MinHash + LSH (datasketch) |
$O(N + M)$ • Milliseconds | Minor edits, word insertions/deletions, truncations |
semantic |
Dense Vector Cosine Similarity (FAISS) | $O(M \cdot d)$ • Seconds | Paraphrased sentences, reworded questions, synonyms |
📦 Installation
# Core installation (exact + fuzzy matching)
pip install verascan
# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"
# With Hugging Face datasets support
pip install "verascan[hf]"
# Complete installation with all optional extras
pip install "verascan[all]"
🚀 Quickstart
Python API
import verascan
# Run contamination audit across training and evaluation splits
report = verascan.check(
train="data/train.jsonl",
eval="data/eval.jsonl",
methods=["exact", "fuzzy", "semantic"],
threshold=0.85,
)
# Print terminal summary
report.summary()
# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs : {report.total_matches}")
# Query high-confidence matches
for match in report.flagged(min_score=0.90):
print(
f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
)
print(f" Eval : {match.eval_text}")
print(f" Train: {match.train_text}")
# Export interactive HTML and machine-readable JSON reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")
# Export a decontaminated eval set (contaminated rows removed)
cleaned = report.cleaned_eval() # list[str] or pandas.DataFrame
report.to_cleaned("eval_cleaned.jsonl") # .jsonl, .csv, or .json
Terminal Output
===============================================
Verascan Contamination Report
===============================================
Train size : 50,000
Eval size : 1,000
Methods : exact, fuzzy, semantic
Threshold : 0.85
-----------------------------------------------
Total matches : 14
Contaminated : 12 / 1,000 eval samples (1.2%)
Exact matches : 4
Fuzzy matches : 7
Semantic hits : 3
===============================================
📂 Supported Input Formats
Verascan normalises inputs into clean text sequences automatically:
import pandas as pd
import verascan
# 1. Plain String Lists
report = verascan.check(
train=["The quick brown fox.", "Artificial intelligence."],
eval=["The quick brown fox."],
)
# 2. File Paths (CSV or JSONL)
report = verascan.check(
train="data/train.jsonl",
eval="data/eval.csv",
column="text",
)
# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")
# 4. Hugging Face Datasets
from datasets import load_dataset
train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")
💻 CLI Usage
The verascan command-line interface enables automated checks in terminal workflows and CI/CD pipelines:
# Basic contamination check
verascan check --train data/train.jsonl --eval data/eval.jsonl
# Specify custom column, methods, and threshold
verascan check \
--train data/train.csv \
--eval data/eval.csv \
--methods exact,fuzzy \
--threshold 0.80 \
--column text \
--output report.html \
--output-cleaned eval_cleaned.jsonl
# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
--train data/train.jsonl \
--eval data/eval.jsonl \
--fail-above 0.01
📊 Interactive HTML Reports
The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):
- Health Status Banner: Visual indicator (
Clean,Low Risk,High Risk) with contamination percentage and progress meter. - Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
- Live Search & Filtering: Instant client-side search across text samples and index numbers.
- Word-Level Diffs: Color-coded
<del>and<ins>tags illustrating textual overlap. - Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.
⚙️ ContaminationReport API
report = verascan.check(train, eval)
# Properties
report.contamination_rate # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches # int: Total flagged pairs
report.exact_count # int: Exact duplicate count
report.fuzzy_count # int: Fuzzy / near-duplicate count
report.semantic_count # int: Semantic match count
report.train_size # int: Size of training corpus
report.eval_size # int: Size of evaluation corpus
# Methods
report.flagged(min_score=0.9) # Returns list of MatchRecord objects >= min_score
report.summary() # Prints ASCII summary to stdout
report.to_dict() # Serialises report to a Python dict
report.to_json("report.json") # Exports JSON file
report.to_html("report.html") # Exports self-contained interactive HTML report
report.cleaned_eval() # Eval examples with contaminated rows removed
report.contaminated_eval() # Eval examples that were flagged
report.to_cleaned("eval_clean.jsonl") # Write cleaned eval as CSV / JSONL / JSON
report.to_contaminated("eval_flagged.jsonl") # Write flagged eval rows
MatchRecord Structure
Each match in report.matches contains:
eval_index: int— Index of the sample in the evaluation dataset.train_index: int— Index of the sample in the training dataset.eval_text: str— Evaluation sample text.train_text: str— Matching training sample text.score: float— Similarity metric (1.0for exact matches, Jaccard for fuzzy, cosine for semantic).method: str— Engine that produced the match ("exact","fuzzy","semantic").
⚠️ Limitations
- Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with
methods=["exact", "fuzzy"]. - Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
- Cross-Lingual Matching: The default semantic model (
all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model viamodel_name="paraphrase-multilingual-MiniLM-L12-v2".
🛠️ Development
# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan
# Install development dependencies
pip install -e ".[all,dev]"
# Run test suite
pytest
# Code formatting and linting
ruff check .
ruff format --check .
# Type checking
mypy src/
📄 License
Distributed under the MIT License.
Metadata
Release files for verascan 0.2.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| verascan-0.2.0.tar.gz | 105.8 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| verascan-0.2.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 133.0 kB
Release files / verascan-0.2.0.tar.gz
| Download URL | verascan-0.2.0.tar.gz |
|---|---|
| Size | 105.8 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
1ba9366863529f3adbf6e5858e2ad91a3f15387e05629244452505f67ee129af
|
|
BLAKE2b-256 checksum How to use checksums |
2002915ab9735b3e6e46875640578b7b296b7b786325f40f9ddbee2ffa74b839
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|
Release files / verascan-0.2.0-py3-none-any.whl
| Download URL | verascan-0.2.0-py3-none-any.whl |
|---|---|
| Size | 27.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
f2acb2e90ef014b70a8943cdd66f1b304fe0109b507e6d97579411f2526c6fc9
|
|
BLAKE2b-256 checksum How to use checksums |
95db41356d04cb4cab19d5e68929d55530f18ca651fd5df5f30db3b4796fc2eb
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.5
|