Verascan
Data Contamination & Leakage Detection for AI / ML Workflows
Detect exact, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark validation, RAG pipelines, and synthetic data auditing.
Overview
Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.
Verascan provides a multi-tier contamination detection pipeline:
- Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
- Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
- Semantic match — Dense embedding similarity search (
sentence-transformers+ FAISS) for paraphrased content.
Features
- Multi-Tier Detection: Run exact, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
- Cross-Method Deduplication: Matches identified by exact hashing are automatically excluded from fuzzy/semantic passes to prevent double-counting.
- Multi-Format Ingestion: Natively accepts
pandas.DataFrame,JSONL,CSV, Hugging Facedatasets.Dataset, and Pythonlist[str]. - Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring search, method filtering, and word-level diffs.
- CI/CD Integration: CLI includes
--fail-aboveto fail builds if contamination exceeds an allowed threshold. - Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (
sentence-transformers,faiss-cpu) are optional extras. - Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting
stderr.
Detection Engines
| Method | Algorithm | Complexity / Speed | Best For |
|---|---|---|---|
exact |
SHA-256 Content Hashing (normalised) | $O(N + M)$ • Microseconds | Verbatim duplicates, casing/whitespace variations |
fuzzy |
MinHash + LSH (datasketch) |
$O(N + M)$ • Milliseconds | Minor edits, word insertions/deletions, truncations |
semantic |
Dense Vector Cosine Similarity (FAISS) | $O(M \cdot d)$ • Seconds | Paraphrased sentences, reworded questions, synonyms |
Installation
# Core installation (exact + fuzzy matching)
pip install verascan
# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"
# With Hugging Face datasets support
pip install "verascan[hf]"
# Complete installation with all optional extras
pip install "verascan[all]"
Quickstart
Python API
import verascan
# Run contamination audit across training and evaluation splits
report = verascan.check(
train="data/train.jsonl",
eval="data/eval.jsonl",
methods=["exact", "fuzzy"],
threshold=0.85,
)
# Print terminal summary
report.summary()
# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs : {report.total_matches}")
# Query high-confidence matches
for match in report.flagged(min_score=0.90):
print(
f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
)
print(f" Eval : {match.eval_text}")
print(f" Train: {match.train_text}")
# Export reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")
Terminal Output
===============================================
Verascan Contamination Report
===============================================
Train size : 50,000
Eval size : 1,000
Methods : exact, fuzzy
Threshold : 0.85
-----------------------------------------------
Total matches : 14
Contaminated : 12 / 1,000 eval samples (1.2%)
Exact matches : 4
Fuzzy matches : 10
===============================================
Supported Input Formats
Verascan normalises inputs into clean text sequences automatically:
import pandas as pd
import verascan
# 1. Plain String Lists
report = verascan.check(
train=["The quick brown fox.", "Artificial intelligence."],
eval=["The quick brown fox."],
)
# 2. File Paths (CSV or JSONL)
report = verascan.check(
train="data/train.jsonl",
eval="data/eval.csv",
column="text",
)
# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")
# 4. Hugging Face Datasets
from datasets import load_dataset
train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")
CLI Usage
The verascan command-line interface enables automated checks in terminal workflows and CI/CD pipelines:
# Basic contamination check
verascan check --train train.jsonl --eval eval.jsonl
# Specify custom column, methods, and threshold
verascan check \
--train data/train.csv \
--eval data/eval.csv \
--methods exact,fuzzy \
--threshold 0.80 \
--column instruction \
--output report.html
# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
--train train.jsonl \
--eval eval.jsonl \
--fail-above 0.01
Interactive HTML Reports
The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):
- Health Status Banner: Visual indicator (
Clean,Low Risk,High Risk) with contamination percentage and progress meter. - Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
- Live Search & Filtering: Instant client-side search across text samples and index numbers.
- Word-Level Diffs: Color-coded
<del>and<ins>tags illustrating textual overlap. - Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.
ContaminationReport API
report = verascan.check(train, eval)
# Properties
report.contamination_rate # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches # int: Total flagged pairs
report.exact_count # int: Exact duplicate count
report.fuzzy_count # int: Fuzzy / near-duplicate count
report.semantic_count # int: Semantic match count
report.train_size # int: Size of training corpus
report.eval_size # int: Size of evaluation corpus
# Methods
report.flagged(min_score=0.9) # Returns list of MatchRecord objects >= min_score
report.summary() # Prints ASCII summary to stdout
report.to_dict() # Serialises report to a Python dict
report.to_json("report.json") # Exports JSON file
report.to_html("report.html") # Exports self-contained interactive HTML report
MatchRecord Structure
Each match in report.matches contains:
eval_index: int— Index of the sample in the evaluation dataset.train_index: int— Index of the sample in the training dataset.eval_text: str— Evaluation sample text.train_text: str— Matching training sample text.score: float— Similarity metric (1.0for exact matches, Jaccard for fuzzy, cosine for semantic).method: str— Engine that produced the match ("exact","fuzzy","semantic").
Limitations
- Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with
methods=["exact", "fuzzy"]. - Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
- Cross-Lingual Matching: The default semantic model (
all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model viamodel_name="paraphrase-multilingual-MiniLM-L12-v2".
Development
# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan
# Install development dependencies
pip install -e ".[all,dev]"
# Run test suite
pytest
# Code formatting and linting
ruff check .
ruff format --check .
# Type checking
mypy src/
License
Distributed under the MIT License.
Metadata
Release files for verascan 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| verascan-0.1.0.tar.gz | 28.9 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| verascan-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 53.7 kB
Release files / verascan-0.1.0.tar.gz
| Download URL | verascan-0.1.0.tar.gz |
|---|---|
| Size | 28.9 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
e0f1b23924d17609f26cd5352855c30797bba59508432cc749c44038cca0f4e0
|
|
BLAKE2b-256 checksum How to use checksums |
28d67eff1d1367deb1b3aab3b841c5711130bcfd597dfea4c6985b7a33c8c8e0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|
Release files / verascan-0.1.0-py3-none-any.whl
| Download URL | verascan-0.1.0-py3-none-any.whl |
|---|---|
| Size | 24.8 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6f45764488a4307ee88c3133c7f2b918e06ea4b13847239187db6578ba3d8c11
|
|
BLAKE2b-256 checksum How to use checksums |
f142d5d3b0679a07179828b80a1cbd52bc8f8c4cb9c95a9b0b2065744307704d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.11.9
|