Skip to main content

Verascan

Data Contamination & Leakage Detection for AI / ML Workflows

PyPI Version Python Versions License: MIT Tests Type Checking Code Style


Detect exact, n-gram, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark integrity, RAG validation, and synthetic data auditing.


Verascan Interactive HTML Report Preview


🎯 Overview

Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.

Verascan provides an end-to-end contamination detection and prevention pipeline:

  1. Prevention (verascan.split) — Partition raw datasets into train and eval splits with a mathematical guarantee of zero exact, fuzzy, or semantic leakage.
  2. Audit (verascan.check) — Multi-tier contamination scanner across existing splits:
    • Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
    • N-gram overlap — GPT-3-style word 13-gram collisions (Brown et al., 2020) with frequency filtering.
    • Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
    • Semantic match — Dense embedding similarity search (sentence-transformers + FAISS) for paraphrased content.
  3. Remediation (report.cleaned_eval / to_cleaned) — Automatically drop contaminated rows and export a pristine benchmark.

✨ Features

  • Leak-Free Dataset Splitting: Proactively partition raw datasets into train and eval splits with 0.0% residual leakage (verascan.split() and CLI verascan split).
  • Multi-Tier Detection: Run exact, ngram, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
  • Cross-Method Deduplication: Matches identified by earlier methods are automatically excluded from later passes to prevent double-counting.
  • Multi-Format Ingestion: Natively accepts pandas.DataFrame, JSONL, CSV, Hugging Face datasets.Dataset, and Python list[str].
  • Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring live search, method filtering, and word-level diffs.
  • Cleaned Eval Export: Drop contaminated eval rows and write a reusable CSV/JSONL benchmark (cleaned_eval(), to_cleaned(), CLI --output-cleaned).
  • CI/CD Integration: CLI includes --fail-above to fail builds if contamination exceeds an allowed threshold.
  • Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (sentence-transformers, faiss-cpu) are optional extras.
  • Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting stderr.

🔬 Detection Engines

Method Algorithm Complexity / Speed Best For
exact SHA-256 Content Hashing (normalised) $O(N + M)$ • Microseconds Verbatim duplicates, casing/whitespace variations
ngram Word n-gram overlap (GPT-3 / Brown et al.) $O(N + M)$ • Milliseconds Long shared phrases; classic 13-gram contamination
fuzzy MinHash + LSH (datasketch) $O(N + M)$ • Milliseconds Minor edits, word insertions/deletions, truncations
semantic Dense Vector Cosine Similarity (FAISS) $O(M \cdot d)$ • Seconds Paraphrased sentences, reworded questions, synonyms

📦 Installation

# Core installation (exact + fuzzy matching + split utility)
pip install verascan

# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"

# With Hugging Face datasets support
pip install "verascan[hf]"

# Complete installation with all optional extras
pip install "verascan[all]"

🛡️ Leak-Free Splitting (verascan.split)

Instead of auditing datasets after training, prevent leakage upfront with verascan.split(). It partitions your dataset and iteratively purges candidate evaluation examples that are too similar to training samples:

Python API

import verascan

# Create a leak-free train and eval split
train, eval_set = verascan.split(
    data="data/raw_dataset.jsonl",
    eval_size=0.2,
    methods=["exact", "fuzzy"],
    seed=42,
    output_train="data/train.jsonl",
    output_eval="data/eval.jsonl",
)

print(f"Train samples: {len(train):,}")
print(f"Eval samples : {len(eval_set):,}")

CLI Command

verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl \
  --methods exact,fuzzy \
  --threshold 0.85

🚀 Contamination Audit (verascan.check)

Python API

import verascan

# Run contamination audit across training and evaluation splits
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.jsonl",
    methods=["exact", "ngram", "fuzzy", "semantic"],
    threshold=0.85,
)

# Print terminal summary
report.summary()

# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs     : {report.total_matches}")

# Query high-confidence matches
for match in report.flagged(min_score=0.90):
    print(
        f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
    )
    print(f"  Eval : {match.eval_text}")
    print(f"  Train: {match.train_text}")

# Export interactive HTML and machine-readable JSON reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")

# Export a decontaminated eval set (contaminated rows removed)
cleaned = report.cleaned_eval()  # list[str] or pandas.DataFrame
report.to_cleaned("eval_cleaned.jsonl")  # .jsonl, .csv, or .json

Terminal Output

===============================================
  Verascan Contamination Report
===============================================
  Train size      : 50,000
  Eval size       : 1,000
  Methods         : exact, fuzzy, semantic
  Threshold       : 0.85
-----------------------------------------------
  Total matches   : 14
  Contaminated    : 12 / 1,000 eval samples (1.2%)
    Exact matches : 4
    Fuzzy matches : 7
    Semantic hits : 3
===============================================

📂 Supported Input Formats

Verascan normalises inputs into clean text sequences automatically:

import pandas as pd
import verascan

# 1. Plain String Lists
report = verascan.check(
    train=["The quick brown fox.", "Artificial intelligence."],
    eval=["The quick brown fox."],
)

# 2. File Paths (CSV or JSONL)
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.csv",
    column="text",
)

# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")

# 4. Hugging Face Datasets
from datasets import load_dataset

train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")

💻 CLI Usage

The verascan CLI enables automated auditing and dataset splitting in terminal workflows and CI/CD pipelines:

1. Split Datasets

# Partition dataset into leak-free splits
verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl

2. Audit Contamination

# Basic contamination check
verascan check --train data/train.jsonl --eval data/eval.jsonl

# Custom columns, methods, threshold, HTML report, and cleaned output
verascan check \
  --train data/train.csv \
  --eval data/eval.csv \
  --methods exact,fuzzy \
  --threshold 0.80 \
  --column text \
  --output report.html \
  --output-cleaned eval_cleaned.jsonl

# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
  --train data/train.jsonl \
  --eval data/eval.jsonl \
  --fail-above 0.01

📊 Interactive HTML Reports

The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):

  • Health Status Banner: Visual indicator (Clean, Low Risk, High Risk) with contamination percentage and progress meter.
  • Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
  • Live Search & Filtering: Instant client-side search across text samples and index numbers.
  • Word-Level Diffs: Color-coded <del> and <ins> tags illustrating textual overlap.
  • Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.

⚙️ Python API Reference

verascan.split(data, ...)

train, eval = verascan.split(
    data,  # str | pd.DataFrame | list[str] | Dataset
    eval_size=0.2,  # float ratio (0-1) or int row count
    methods=["exact", "fuzzy"],  # exact, fuzzy, semantic
    threshold=0.85,  # similarity threshold (0-1)
    column="text",  # text column for tabular inputs
    seed=42,  # random seed for reproducible partition
    move_to="train",  # "train" (preserve data) or "drop" (discard)
    output_train="train.jsonl",  # optional path (.csv, .jsonl, .json)
    output_eval="eval.jsonl",  # optional path (.csv, .jsonl, .json)
)

verascan.check(train, eval, ...)

report = verascan.check(train, eval)

# Properties
report.contamination_rate  # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches  # int: Total flagged pairs
report.exact_count  # int: Exact duplicate count
report.ngram_count  # int: N-gram overlap match count
report.fuzzy_count  # int: Fuzzy / near-duplicate count
report.semantic_count  # int: Semantic match count
report.train_size  # int: Size of training corpus
report.eval_size  # int: Size of evaluation corpus

# Methods
report.flagged(min_score=0.9)  # Returns list of MatchRecord objects >= min_score
report.summary()  # Prints ASCII summary to stdout
report.to_dict()  # Serialises report to a Python dict
report.to_json("report.json")  # Exports JSON file
report.to_html("report.html")  # Exports self-contained interactive HTML report
report.cleaned_eval()  # Eval examples with contaminated rows removed
report.contaminated_eval()  # Eval examples that were flagged
report.to_cleaned("eval_clean.jsonl")  # Write cleaned eval as CSV / JSONL / JSON
report.to_contaminated("eval_flagged.jsonl")  # Write flagged eval rows

MatchRecord Structure

Each match in report.matches contains:

  • eval_index: int — Index of the sample in the evaluation dataset.
  • train_index: int — Index of the sample in the training dataset.
  • eval_text: str — Evaluation sample text.
  • train_text: str — Matching training sample text.
  • score: float — Similarity metric (1.0 for exact matches, n-gram overlap ratio for ngram, Jaccard for fuzzy, cosine for semantic).
  • method: str — Engine that produced the match ("exact", "ngram", "fuzzy", "semantic").

⚠️ Limitations

  • Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with methods=["exact", "fuzzy"].
  • Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
  • Word N-Gram Length: The ngram method needs at least ngram_n words (default 13) after cleaning; shorter eval texts produce no n-gram matches.
  • Cross-Lingual Matching: The default semantic model (all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model via model_name="paraphrase-multilingual-MiniLM-L12-v2".

🛠️ Development

# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan

# Install development dependencies
pip install -e ".[all,dev]"

# Run test suite
pytest

# Code formatting and linting
ruff check .
ruff format --check .

# Type checking
mypy src/

📄 License

Distributed under the MIT License.

Metadata

Release files for verascan 0.3.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verascan 0.3.0
File Size Uploaded
verascan-0.3.0.tar.gz 116.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verascan 0.3.0
File Interpreter ABI Platform
verascan-0.3.0-py3-none-any.whl Python 3 none any Details

Total release size: 150.6 kB

Release files / verascan-0.3.0.tar.gz

Download URL verascan-0.3.0.tar.gz
Size 116.1 kB
Tags Source
SHA-256 checksum
How to use checksums
0a788dd3999bb5ae5059f2482f578995bd41118fcc30e7ac9ec45f429de01bb9
BLAKE2b-256 checksum
How to use checksums
b3dad7a18002898bd629ce88e80a333394f755758bafd6c2985cc51751fcd9a4
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / verascan-0.3.0-py3-none-any.whl

Download URL verascan-0.3.0-py3-none-any.whl
Size 34.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
efc0a2e5ed737303d2799770672aec12e79a44cdc8030c4a18639f1a62251f5f
BLAKE2b-256 checksum
How to use checksums
47010445ba64adbfd8da2e6e9b19e34dd9953545a59a4d251d83459b0053336e
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.4.1

2 release files

0.4.0

2 release files

This release

0.3.0 This release

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page