Skip to main content

Verascan

Data Contamination & Leakage Detection for AI / ML Workflows

PyPI Version CI Python Versions License: MIT Tests Type Checking Code Style


Detect exact, n-gram, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark integrity, RAG validation, and synthetic data auditing.


Verascan Interactive HTML Report Preview


🎯 Overview

Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.

Verascan provides an end-to-end contamination detection and prevention pipeline:

  1. Prevention (verascan.split) — Partition raw datasets into train and eval splits with a mathematical guarantee of zero exact, fuzzy, or semantic leakage.
  2. Benchmark Audit (verascan.audit) — Audit training data against popular public benchmarks (mmlu, gsm8k, humaneval).
  3. Contamination Audit (verascan.check) — Multi-tier contamination scanner across existing splits:
    • Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
    • N-gram overlap — GPT-3-style word 13-gram collisions (Brown et al., 2020) with frequency filtering.
    • Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
    • Semantic match — Dense embedding similarity search (sentence-transformers + FAISS) for paraphrased content.
  4. Remediation (report.cleaned_eval / to_cleaned) — Automatically drop contaminated rows and export a pristine benchmark.

✨ Features

  • Leak-Free Dataset Splitting: Proactively partition raw datasets into train and eval splits with 0.0% residual leakage (verascan.split() and CLI verascan split).
  • Multi-Tier Detection: Run exact, ngram, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
  • Cross-Method Deduplication: Matches identified by earlier methods are automatically excluded from later passes to prevent double-counting.
  • Multi-Format Ingestion: Natively accepts pandas.DataFrame, JSONL, CSV, Hugging Face datasets.Dataset, and Python list[str].
  • Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring live search, method filtering, and word-level diffs.
  • Cleaned Eval Export: Drop contaminated eval rows and write a reusable CSV/JSONL benchmark (cleaned_eval(), to_cleaned(), CLI --output-cleaned).
  • CI/CD Integration: CLI includes --fail-above to fail builds if contamination exceeds an allowed threshold.
  • Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (sentence-transformers, faiss-cpu) are optional extras.
  • Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting stderr.

🔬 Detection Engines

Method Algorithm Complexity / Speed Best For
exact SHA-256 Content Hashing (normalised) $O(N + M)$ • Microseconds Verbatim duplicates, casing/whitespace variations
ngram Word n-gram overlap (GPT-3 / Brown et al.) $O(N + M)$ • Milliseconds Long shared phrases; classic 13-gram contamination
fuzzy MinHash + LSH (datasketch) $O(N + M)$ • Milliseconds Minor edits, word insertions/deletions, truncations
semantic Dense Vector Cosine Similarity (FAISS) $O(M \cdot d)$ • Seconds Paraphrased sentences, reworded questions, synonyms

📦 Installation

# Core installation (exact + fuzzy matching + split utility)
pip install verascan

# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"

# With Hugging Face datasets support
pip install "verascan[hf]"

# Complete installation with all optional extras
pip install "verascan[all]"

🛡️ Leak-Free Splitting (verascan.split)

Instead of auditing datasets after training, prevent leakage upfront with verascan.split(). It partitions your dataset and iteratively purges candidate evaluation examples that are too similar to training samples:

Python API

import verascan

# Create a leak-free train and eval split
train, eval_set = verascan.split(
    data="data/raw_dataset.jsonl",
    eval_size=0.2,
    methods=["exact", "fuzzy"],
    seed=42,
    output_train="data/train.jsonl",
    output_eval="data/eval.jsonl",
)

print(f"Train samples: {len(train):,}")
print(f"Eval samples : {len(eval_set):,}")

CLI Command

verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl \
  --methods exact,fuzzy \
  --threshold 0.85

🏛️ Benchmark Auditing (verascan.audit)

Audit your training data directly against popular public evaluation benchmarks (mmlu, gsm8k, humaneval) to detect pre-training contamination before benchmark evaluation:

import verascan

report = verascan.audit(
    train="data/train.jsonl",
    benchmarks=["mmlu", "gsm8k", "humaneval"],
    methods=["exact", "fuzzy"],
    threshold=0.85,
)

report.summary()
report.to_html("audit_report.html")

🚀 Contamination Audit (verascan.check)

Python API

import verascan

# Run contamination audit across training and evaluation splits
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.jsonl",
    methods=["exact", "ngram", "fuzzy", "semantic"],
    threshold=0.85,
)

# Print terminal summary
report.summary()

# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs     : {report.total_matches}")

# Query high-confidence matches
for match in report.flagged(min_score=0.90):
    print(
        f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
    )
    print(f"  Eval : {match.eval_text}")
    print(f"  Train: {match.train_text}")

# Export interactive HTML and machine-readable JSON reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")

# Export a decontaminated eval set (contaminated rows removed)
cleaned = report.cleaned_eval()  # list[str] or pandas.DataFrame
report.to_cleaned("eval_cleaned.jsonl")  # .jsonl, .csv, or .json

Terminal Output

===============================================
  Verascan Contamination Report
===============================================
  Train size      : 50,000
  Eval size       : 1,000
  Methods         : exact, fuzzy, semantic
  Threshold       : 0.85
-----------------------------------------------
  Total matches   : 14
  Contaminated    : 12 / 1,000 eval samples (1.2%)
    Exact matches : 4
    Fuzzy matches : 7
    Semantic hits : 3
===============================================

📂 Supported Input Formats

Verascan normalises inputs into clean text sequences automatically:

import pandas as pd
import verascan

# 1. Plain String Lists
report = verascan.check(
    train=["The quick brown fox.", "Artificial intelligence."],
    eval=["The quick brown fox."],
)

# 2. File Paths (CSV or JSONL)
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.csv",
    column="text",
)

# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")

# 4. Hugging Face Datasets
from datasets import load_dataset

train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")

💻 CLI Usage

The verascan CLI enables automated auditing and dataset splitting in terminal workflows and CI/CD pipelines:

1. Split Datasets

# Partition dataset into leak-free splits
verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl

2. Audit Contamination

# Basic contamination check
verascan check --train data/train.jsonl --eval data/eval.jsonl

# Custom columns, methods, threshold, HTML report, and cleaned output
verascan check \
  --train data/train.csv \
  --eval data/eval.csv \
  --methods exact,fuzzy \
  --threshold 0.80 \
  --column text \
  --output report.html \
  --output-cleaned eval_cleaned.jsonl

# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
  --train data/train.jsonl \
  --eval data/eval.jsonl \
  --fail-above 0.01

3. Benchmark Auditing

# Fast offline audit with synthetic stand-ins
verascan audit --train train.jsonl --benchmarks mmlu,gsm8k --synthetic

# Full benchmark audit with interactive HTML report
verascan audit \
  --train data/train.jsonl \
  --benchmarks mmlu,gsm8k,humaneval \
  --output audit_report.html

📊 Interactive HTML Reports

The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):

  • Health Status Banner: Visual indicator (Clean, Low Risk, High Risk) with contamination percentage and progress meter.
  • Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
  • Live Search & Filtering: Instant client-side search across text samples and index numbers.
  • Word-Level Diffs: Color-coded <del> and <ins> tags illustrating textual overlap.
  • Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.

⚙️ Python API Reference

verascan.split(data, ...)

train, eval = verascan.split(
    data,  # str | pd.DataFrame | list[str] | Dataset
    eval_size=0.2,  # float ratio (0-1) or int row count
    methods=["exact", "fuzzy"],  # exact, fuzzy, semantic
    threshold=0.85,  # similarity threshold (0-1)
    column="text",  # text column for tabular inputs
    seed=42,  # random seed for reproducible partition
    move_to="train",  # "train" (preserve data) or "drop" (discard)
    output_train="train.jsonl",  # optional path (.csv, .jsonl, .json)
    output_eval="eval.jsonl",  # optional path (.csv, .jsonl, .json)
)

verascan.audit(train, ...)

report = verascan.audit(
    train="train.jsonl",  # path, DataFrame, list[str], or Dataset
    benchmarks=["mmlu", "gsm8k", "humaneval"],  # preset names
    methods=["exact", "fuzzy"],  # exact, ngram, fuzzy, semantic
    threshold=0.85,  # similarity cutoff
    synthetic=False,  # set True for local offline testing
)

report.benchmark_counts  # dict: e.g. {"mmlu": 1, "gsm8k": 0, "humaneval": 2}
report.by_benchmark()  # dict: matches grouped by benchmark name

verascan.check(train, eval, ...)

report = verascan.check(train, eval)

# Properties
report.contamination_rate  # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches  # int: Total flagged pairs
report.exact_count  # int: Exact duplicate count
report.ngram_count  # int: N-gram overlap match count
report.fuzzy_count  # int: Fuzzy / near-duplicate count
report.semantic_count  # int: Semantic match count
report.train_size  # int: Size of training corpus
report.eval_size  # int: Size of evaluation corpus

# Methods
report.flagged(min_score=0.9)  # Returns list of MatchRecord objects >= min_score
report.summary()  # Prints ASCII summary to stdout
report.to_dict()  # Serialises report to a Python dict
report.to_json("report.json")  # Exports JSON file
report.to_html("report.html")  # Exports self-contained interactive HTML report
report.cleaned_eval()  # Eval examples with contaminated rows removed
report.contaminated_eval()  # Eval examples that were flagged
report.to_cleaned("eval_clean.jsonl")  # Write cleaned eval as CSV / JSONL / JSON
report.to_contaminated("eval_flagged.jsonl")  # Write flagged eval rows

MatchRecord Structure

Each match in report.matches contains:

  • eval_index: int — Index of the sample in the evaluation dataset.
  • train_index: int — Index of the sample in the training dataset.
  • eval_text: str — Evaluation sample text.
  • train_text: str — Matching training sample text.
  • score: float — Similarity metric (1.0 for exact matches, n-gram overlap ratio for ngram, Jaccard for fuzzy, cosine for semantic).
  • method: str — Engine that produced the match ("exact", "ngram", "fuzzy", "semantic").

⚠️ Limitations

  • Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with methods=["exact", "fuzzy"].
  • Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
  • Word N-Gram Length: The ngram method needs at least ngram_n words (default 13) after cleaning; shorter eval texts produce no n-gram matches.
  • Cross-Lingual Matching: The default semantic model (all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model via model_name="paraphrase-multilingual-MiniLM-L12-v2".

🛠️ Development

# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan

# Install development dependencies
pip install -e ".[all,dev]"

# Run test suite
pytest

# Code formatting and linting
ruff check .
ruff format --check .

# Type checking
mypy src/

📄 License

Distributed under the MIT License.

Metadata

Release files for verascan 0.4.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verascan 0.4.1
File Size Uploaded
verascan-0.4.1.tar.gz 128.1 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verascan 0.4.1
File Interpreter ABI Platform
verascan-0.4.1-py3-none-any.whl Python 3 none any Details

Total release size: 170.5 kB

Release files / verascan-0.4.1.tar.gz

Download URL verascan-0.4.1.tar.gz
Size 128.1 kB
Tags Source
SHA-256 checksum
How to use checksums
9411f1a9ae7ff14264a339d0d4179da86b17f92f405f3df1ae452f8aa6dcfaf6
BLAKE2b-256 checksum
How to use checksums
cb20df1659c686ce4ceedceec6e052cb5c268cb6d50ebef81eee7334b66b509f
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / verascan-0.4.1-py3-none-any.whl

Download URL verascan-0.4.1-py3-none-any.whl
Size 42.4 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ce1c99edf3c8c4239f53d6b303e8395a45df3a80eaae88a40f2ff8a44a2171e5
BLAKE2b-256 checksum
How to use checksums
b80bf099e1aa2362bc56d43d062139e037b3e04a820fb2553fc4f2a405ddbde0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

This release

0.4.1 This release

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page