Skip to main content

Verascan

Data Contamination & Leakage Detection for AI / ML Workflows

PyPI Version CI Python Versions License: MIT Tests Type Checking Code Style


Detect exact, n-gram, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark integrity, RAG validation, and synthetic data auditing.


Verascan Interactive HTML Report Preview


🎯 Overview

Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.

Verascan provides an end-to-end contamination detection and prevention pipeline:

  1. Prevention (verascan.split) — Partition raw datasets into train and eval splits with a mathematical guarantee of zero exact, fuzzy, or semantic leakage.
  2. Benchmark Audit (verascan.audit) — Audit training data against popular public benchmarks (mmlu, gsm8k, humaneval).
  3. Contamination Audit (verascan.check) — Multi-tier contamination scanner across existing splits:
    • Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
    • N-gram overlap — GPT-3-style word 13-gram collisions (Brown et al., 2020) with frequency filtering.
    • Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
    • Semantic match — Dense embedding similarity search (sentence-transformers + FAISS) for paraphrased content.
  4. Remediation (report.cleaned_eval / to_cleaned) — Automatically drop contaminated rows and export a pristine benchmark.

✨ Features

  • Leak-Free Dataset Splitting: Proactively partition raw datasets into train and eval splits with 0.0% residual leakage (verascan.split() and CLI verascan split).
  • Multi-Tier Detection: Run exact, ngram, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
  • Cross-Method Deduplication: Matches identified by earlier methods are automatically excluded from later passes to prevent double-counting.
  • Multi-Format Ingestion: Natively accepts pandas.DataFrame, JSONL, CSV, Hugging Face datasets.Dataset, and Python list[str].
  • Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring live search, method filtering, and word-level diffs.
  • Cleaned Eval Export: Drop contaminated eval rows and write a reusable CSV/JSONL benchmark (cleaned_eval(), to_cleaned(), CLI --output-cleaned).
  • CI/CD Integration: CLI includes --fail-above to fail builds if contamination exceeds an allowed threshold.
  • Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (sentence-transformers, faiss-cpu) are optional extras.
  • Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting stderr.

🔬 Detection Engines

Method Algorithm Complexity / Speed Best For
exact SHA-256 Content Hashing (normalised) $O(N + M)$ • Microseconds Verbatim duplicates, casing/whitespace variations
ngram Word n-gram overlap (GPT-3 / Brown et al.) $O(N + M)$ • Milliseconds Long shared phrases; classic 13-gram contamination
fuzzy MinHash + LSH (datasketch) $O(N + M)$ • Milliseconds Minor edits, word insertions/deletions, truncations
semantic Dense Vector Cosine Similarity (FAISS) $O(M \cdot d)$ • Seconds Paraphrased sentences, reworded questions, synonyms

📦 Installation

# Core installation (exact + fuzzy matching + split utility)
pip install verascan

# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"

# With Hugging Face datasets support
pip install "verascan[hf]"

# Complete installation with all optional extras
pip install "verascan[all]"

🛡️ Leak-Free Splitting (verascan.split)

Instead of auditing datasets after training, prevent leakage upfront with verascan.split(). It partitions your dataset and iteratively purges candidate evaluation examples that are too similar to training samples:

Python API

import verascan

# Create a leak-free train and eval split
train, eval_set = verascan.split(
    data="data/raw_dataset.jsonl",
    eval_size=0.2,
    methods=["exact", "fuzzy"],
    seed=42,
    output_train="data/train.jsonl",
    output_eval="data/eval.jsonl",
)

print(f"Train samples: {len(train):,}")
print(f"Eval samples : {len(eval_set):,}")

CLI Command

verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl \
  --methods exact,fuzzy \
  --threshold 0.85

🏛️ Benchmark Auditing (verascan.audit)

Audit your training data directly against popular public evaluation benchmarks (mmlu, gsm8k, humaneval) to detect pre-training contamination before benchmark evaluation:

import verascan

report = verascan.audit(
    train="data/train.jsonl",
    benchmarks=["mmlu", "gsm8k", "humaneval"],
    methods=["exact", "fuzzy"],
    threshold=0.85,
)

report.summary()
report.to_html("audit_report.html")

🚀 Contamination Audit (verascan.check)

Python API

import verascan

# Run contamination audit across training and evaluation splits
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.jsonl",
    methods=["exact", "ngram", "fuzzy", "semantic"],
    threshold=0.85,
)

# Print terminal summary
report.summary()

# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs     : {report.total_matches}")

# Query high-confidence matches
for match in report.flagged(min_score=0.90):
    print(
        f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
    )
    print(f"  Eval : {match.eval_text}")
    print(f"  Train: {match.train_text}")

# Export interactive HTML and machine-readable JSON reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")

# Export a decontaminated eval set (contaminated rows removed)
cleaned = report.cleaned_eval()  # list[str] or pandas.DataFrame
report.to_cleaned("eval_cleaned.jsonl")  # .jsonl, .csv, or .json

Terminal Output

===============================================
  Verascan Contamination Report
===============================================
  Train size      : 50,000
  Eval size       : 1,000
  Methods         : exact, fuzzy, semantic
  Threshold       : 0.85
-----------------------------------------------
  Total matches   : 14
  Contaminated    : 12 / 1,000 eval samples (1.2%)
    Exact matches : 4
    Fuzzy matches : 7
    Semantic hits : 3
===============================================

📂 Supported Input Formats

Verascan normalises inputs into clean text sequences automatically:

import pandas as pd
import verascan

# 1. Plain String Lists
report = verascan.check(
    train=["The quick brown fox.", "Artificial intelligence."],
    eval=["The quick brown fox."],
)

# 2. File Paths (CSV or JSONL)
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.csv",
    column="text",
)

# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")

# 4. Hugging Face Datasets
from datasets import load_dataset

train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")

💻 CLI Usage

The verascan CLI enables automated auditing and dataset splitting in terminal workflows and CI/CD pipelines:

1. Split Datasets

# Partition dataset into leak-free splits
verascan split \
  --input data/raw_dataset.jsonl \
  --eval-size 0.2 \
  --output-train data/train.jsonl \
  --output-eval data/eval.jsonl

2. Audit Contamination

# Basic contamination check
verascan check --train data/train.jsonl --eval data/eval.jsonl

# Custom columns, methods, threshold, HTML report, and cleaned output
verascan check \
  --train data/train.csv \
  --eval data/eval.csv \
  --methods exact,fuzzy \
  --threshold 0.80 \
  --column text \
  --output report.html \
  --output-cleaned eval_cleaned.jsonl

# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
  --train data/train.jsonl \
  --eval data/eval.jsonl \
  --fail-above 0.01

3. Benchmark Auditing

# Fast offline audit with synthetic stand-ins
verascan audit --train train.jsonl --benchmarks mmlu,gsm8k --synthetic

# Full benchmark audit with interactive HTML report
verascan audit \
  --train data/train.jsonl \
  --benchmarks mmlu,gsm8k,humaneval \
  --output audit_report.html

📊 Interactive HTML Reports

The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):

  • Health Status Banner: Visual indicator (Clean, Low Risk, High Risk) with contamination percentage and progress meter.
  • Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
  • Live Search & Filtering: Instant client-side search across text samples and index numbers.
  • Word-Level Diffs: Color-coded <del> and <ins> tags illustrating textual overlap.
  • Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.

⚙️ Python API Reference

verascan.split(data, ...)

train, eval = verascan.split(
    data,  # str | pd.DataFrame | list[str] | Dataset
    eval_size=0.2,  # float ratio (0-1) or int row count
    methods=["exact", "fuzzy"],  # exact, fuzzy, semantic
    threshold=0.85,  # similarity threshold (0-1)
    column="text",  # text column for tabular inputs
    seed=42,  # random seed for reproducible partition
    move_to="train",  # "train" (preserve data) or "drop" (discard)
    output_train="train.jsonl",  # optional path (.csv, .jsonl, .json)
    output_eval="eval.jsonl",  # optional path (.csv, .jsonl, .json)
)

verascan.audit(train, ...)

report = verascan.audit(
    train="train.jsonl",  # path, DataFrame, list[str], or Dataset
    benchmarks=["mmlu", "gsm8k", "humaneval"],  # preset names
    methods=["exact", "fuzzy"],  # exact, ngram, fuzzy, semantic
    threshold=0.85,  # similarity cutoff
    synthetic=False,  # set True for local offline testing
)

report.benchmark_counts  # dict: e.g. {"mmlu": 1, "gsm8k": 0, "humaneval": 2}
report.by_benchmark()  # dict: matches grouped by benchmark name

verascan.check(train, eval, ...)

report = verascan.check(train, eval)

# Properties
report.contamination_rate  # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches  # int: Total flagged pairs
report.exact_count  # int: Exact duplicate count
report.ngram_count  # int: N-gram overlap match count
report.fuzzy_count  # int: Fuzzy / near-duplicate count
report.semantic_count  # int: Semantic match count
report.train_size  # int: Size of training corpus
report.eval_size  # int: Size of evaluation corpus

# Methods
report.flagged(min_score=0.9)  # Returns list of MatchRecord objects >= min_score
report.summary()  # Prints ASCII summary to stdout
report.to_dict()  # Serialises report to a Python dict
report.to_json("report.json")  # Exports JSON file
report.to_html("report.html")  # Exports self-contained interactive HTML report
report.cleaned_eval()  # Eval examples with contaminated rows removed
report.contaminated_eval()  # Eval examples that were flagged
report.to_cleaned("eval_clean.jsonl")  # Write cleaned eval as CSV / JSONL / JSON
report.to_contaminated("eval_flagged.jsonl")  # Write flagged eval rows

MatchRecord Structure

Each match in report.matches contains:

  • eval_index: int — Index of the sample in the evaluation dataset.
  • train_index: int — Index of the sample in the training dataset.
  • eval_text: str — Evaluation sample text.
  • train_text: str — Matching training sample text.
  • score: float — Similarity metric (1.0 for exact matches, n-gram overlap ratio for ngram, Jaccard for fuzzy, cosine for semantic).
  • method: str — Engine that produced the match ("exact", "ngram", "fuzzy", "semantic").

⚠️ Limitations

  • Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with methods=["exact", "fuzzy"].
  • Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
  • Word N-Gram Length: The ngram method needs at least ngram_n words (default 13) after cleaning; shorter eval texts produce no n-gram matches.
  • Cross-Lingual Matching: The default semantic model (all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model via model_name="paraphrase-multilingual-MiniLM-L12-v2".

🛠️ Development

# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan

# Install development dependencies
pip install -e ".[all,dev]"

# Run test suite
pytest

# Code formatting and linting
ruff check .
ruff format --check .

# Type checking
mypy src/

📄 License

Distributed under the MIT License.

Metadata

Release files for verascan 0.4.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verascan 0.4.0
File Size Uploaded
verascan-0.4.0.tar.gz 126.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verascan 0.4.0
File Interpreter ABI Platform
verascan-0.4.0-py3-none-any.whl Python 3 none any Details

Total release size: 168.0 kB

Release files / verascan-0.4.0.tar.gz

Download URL verascan-0.4.0.tar.gz
Size 126.0 kB
Tags Source
SHA-256 checksum
How to use checksums
5e17c0de369e7b7cc39cb8041aa4e8e0aec8e96db9cad8fadde69fdc30cc338e
BLAKE2b-256 checksum
How to use checksums
31381d2b26f31fec6dcda3aff5924d8525439996d6c34b333bd056b50ebdc148
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / verascan-0.4.0-py3-none-any.whl

Download URL verascan-0.4.0-py3-none-any.whl
Size 42.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
68ba1f9fc0a359b4bc25dc959997755223ad4bb9e20285994d85725ce110c59f
BLAKE2b-256 checksum
How to use checksums
9e0cda9315ee0f532d6a9888dd1f849d45a0458b8c5c7a241cea470f23590eb5
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.4.1

2 release files

This release

0.4.0 This release

2 release files

0.3.0

2 release files

0.2.0

2 release files

0.1.0

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page