Skip to main content

Verascan

Data Contamination & Leakage Detection for AI / ML Workflows

PyPI Version Python Versions License: MIT Tests Type Checking Code Style


Detect exact, fuzzy, and semantic data leakage between training and evaluation datasets.
Built for LLM fine-tuning, benchmark validation, RAG pipelines, and synthetic data auditing.


Overview

Data contamination occurs when evaluation or benchmark examples leak into a model's training data. This compromises evaluation validity, inflates benchmark scores, and masks real-world model degradation.

Verascan provides a multi-tier contamination detection pipeline:

  1. Exact match — $O(N)$ hash-based verbatim duplicate detection with normalisation.
  2. Fuzzy match — MinHash + Locality-Sensitive Hashing (LSH) for near-duplicates and minor edits.
  3. Semantic match — Dense embedding similarity search (sentence-transformers + FAISS) for paraphrased content.

Features

  • Multi-Tier Detection: Run exact, fuzzy, and semantic algorithms independently or in a cascaded pipeline.
  • Cross-Method Deduplication: Matches identified by exact hashing are automatically excluded from fuzzy/semantic passes to prevent double-counting.
  • Multi-Format Ingestion: Natively accepts pandas.DataFrame, JSONL, CSV, Hugging Face datasets.Dataset, and Python list[str].
  • Interactive HTML Reports: Generates self-contained, offline-ready HTML reports featuring search, method filtering, and word-level diffs.
  • CI/CD Integration: CLI includes --fail-above to fail builds if contamination exceeds an allowed threshold.
  • Lightweight Core: Installs cleanly with minimal dependencies; heavy ML dependencies (sentence-transformers, faiss-cpu) are optional extras.
  • Noise-Free Execution: Built-in log suppression prevents noisy C++/oneDNN and framework deprecation logs from polluting stderr.

Detection Engines

Method Algorithm Complexity / Speed Best For
exact SHA-256 Content Hashing (normalised) $O(N + M)$ • Microseconds Verbatim duplicates, casing/whitespace variations
fuzzy MinHash + LSH (datasketch) $O(N + M)$ • Milliseconds Minor edits, word insertions/deletions, truncations
semantic Dense Vector Cosine Similarity (FAISS) $O(M \cdot d)$ • Seconds Paraphrased sentences, reworded questions, synonyms

Installation

# Core installation (exact + fuzzy matching)
pip install verascan

# With semantic similarity matching (sentence-transformers + FAISS)
pip install "verascan[semantic]"

# With Hugging Face datasets support
pip install "verascan[hf]"

# Complete installation with all optional extras
pip install "verascan[all]"

Quickstart

Python API

import verascan

# Run contamination audit across training and evaluation splits
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.jsonl",
    methods=["exact", "fuzzy"],
    threshold=0.85,
)

# Print terminal summary
report.summary()

# Inspect metrics
print(f"Contamination Rate: {report.contamination_rate:.1%}")
print(f"Flagged Pairs     : {report.total_matches}")

# Query high-confidence matches
for match in report.flagged(min_score=0.90):
    print(
        f"[{match.method}] Eval #{match.eval_index} <-> Train #{match.train_index} (Score: {match.score:.3f})"
    )
    print(f"  Eval : {match.eval_text}")
    print(f"  Train: {match.train_text}")

# Export reports
report.to_html("contamination_report.html")
report.to_json("contamination_report.json")

Terminal Output

===============================================
  Verascan Contamination Report
===============================================
  Train size      : 50,000
  Eval size       : 1,000
  Methods         : exact, fuzzy
  Threshold       : 0.85
-----------------------------------------------
  Total matches   : 14
  Contaminated    : 12 / 1,000 eval samples (1.2%)
    Exact matches : 4
    Fuzzy matches : 10
===============================================

Supported Input Formats

Verascan normalises inputs into clean text sequences automatically:

import pandas as pd
import verascan

# 1. Plain String Lists
report = verascan.check(
    train=["The quick brown fox.", "Artificial intelligence."],
    eval=["The quick brown fox."],
)

# 2. File Paths (CSV or JSONL)
report = verascan.check(
    train="data/train.jsonl",
    eval="data/eval.csv",
    column="text",
)

# 3. Pandas DataFrames
train_df = pd.DataFrame({"prompt": ["Translate to French...", "Summarize..."]})
eval_df = pd.DataFrame({"prompt": ["Translate to French..."]})
report = verascan.check(train=train_df, eval=eval_df, column="prompt")

# 4. Hugging Face Datasets
from datasets import load_dataset

train_ds = load_dataset("imdb", split="train")
eval_ds = load_dataset("imdb", split="test")
report = verascan.check(train=train_ds, eval=eval_ds, column="text")

CLI Usage

The verascan command-line interface enables automated checks in terminal workflows and CI/CD pipelines:

# Basic contamination check
verascan check --train train.jsonl --eval eval.jsonl

# Specify custom column, methods, and threshold
verascan check \
  --train data/train.csv \
  --eval data/eval.csv \
  --methods exact,fuzzy \
  --threshold 0.80 \
  --column instruction \
  --output report.html

# CI/CD Gate: Fail build if contamination rate exceeds 1%
verascan check \
  --train train.jsonl \
  --eval eval.jsonl \
  --fail-above 0.01

Interactive HTML Reports

The HTML report generated via report.to_html("report.html") is 100% self-contained (no external fonts, CDNs, or scripts required):

  • Health Status Banner: Visual indicator (Clean, Low Risk, High Risk) with contamination percentage and progress meter.
  • Method Breakdown: Color-coded badges for exact (Rose), fuzzy (Amber), and semantic (Indigo) detections.
  • Live Search & Filtering: Instant client-side search across text samples and index numbers.
  • Word-Level Diffs: Color-coded <del> and <ins> tags illustrating textual overlap.
  • Responsive Layout: Designed for seamless viewing across desktop monitors and mobile devices.

ContaminationReport API

report = verascan.check(train, eval)

# Properties
report.contamination_rate  # float: Fraction of eval examples found in train (0.0 to 1.0)
report.total_matches  # int: Total flagged pairs
report.exact_count  # int: Exact duplicate count
report.fuzzy_count  # int: Fuzzy / near-duplicate count
report.semantic_count  # int: Semantic match count
report.train_size  # int: Size of training corpus
report.eval_size  # int: Size of evaluation corpus

# Methods
report.flagged(min_score=0.9)  # Returns list of MatchRecord objects >= min_score
report.summary()  # Prints ASCII summary to stdout
report.to_dict()  # Serialises report to a Python dict
report.to_json("report.json")  # Exports JSON file
report.to_html("report.html")  # Exports self-contained interactive HTML report

MatchRecord Structure

Each match in report.matches contains:

  • eval_index: int — Index of the sample in the evaluation dataset.
  • train_index: int — Index of the sample in the training dataset.
  • eval_text: str — Evaluation sample text.
  • train_text: str — Matching training sample text.
  • score: float — Similarity metric (1.0 for exact matches, Jaccard for fuzzy, cosine for semantic).
  • method: str — Engine that produced the match ("exact", "fuzzy", "semantic").

Limitations

  • Large-Scale Semantic Search: While FAISS provides fast approximate search, semantic matching encodes all samples using transformer models, which is compute-intensive on CPU for corpora with millions of rows. For very large datasets, start with methods=["exact", "fuzzy"].
  • Character N-Gram Sensitivity: Fuzzy matching relies on character 5-grams by default. Very short texts (fewer than 5 characters) fall back to exact matching.
  • Cross-Lingual Matching: The default semantic model (all-MiniLM-L6-v2) is optimized for English text. For multilingual evaluation datasets, specify a multilingual model via model_name="paraphrase-multilingual-MiniLM-L12-v2".

Development

# Clone repository
git clone https://github.com/balamuruganpg/verascan.git
cd verascan

# Install development dependencies
pip install -e ".[all,dev]"

# Run test suite
pytest

# Code formatting and linting
ruff check .
ruff format --check .

# Type checking
mypy src/

License

Distributed under the MIT License.

Metadata

Release files for verascan 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for verascan 0.1.0
File Size Uploaded
verascan-0.1.0.tar.gz 28.9 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for verascan 0.1.0
File Interpreter ABI Platform
verascan-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 53.7 kB

Release files / verascan-0.1.0.tar.gz

Download URL verascan-0.1.0.tar.gz
Size 28.9 kB
Tags Source
SHA-256 checksum
How to use checksums
e0f1b23924d17609f26cd5352855c30797bba59508432cc749c44038cca0f4e0
BLAKE2b-256 checksum
How to use checksums
28d67eff1d1367deb1b3aab3b841c5711130bcfd597dfea4c6985b7a33c8c8e0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release files / verascan-0.1.0-py3-none-any.whl

Download URL verascan-0.1.0-py3-none-any.whl
Size 24.8 kB
Tags Python 3
SHA-256 checksum
How to use checksums
6f45764488a4307ee88c3133c7f2b918e06ea4b13847239187db6578ba3d8c11
BLAKE2b-256 checksum
How to use checksums
f142d5d3b0679a07179828b80a1cbd52bc8f8c4cb9c95a9b0b2065744307704d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/7.0.0 CPython/3.11.9

Release history Release notifications | RSS feed

0.4.1

2 release files

0.4.0

2 release files

0.3.0

2 release files

0.2.0

2 release files

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page