Skip to main content

RupuData

Language: English | Español

Follow the path of your data.

RupuData answers an uncomfortable question: does my training data overlap with data I use to evaluate my model?

It detects dataset duplicates, train/eval overlap, and benchmark overlap signals — locally, with deterministic, machine-readable audit reports. Technical signals, not legal certification.

Install

Requires Python 3.9+.

pip install rupudata
rupudata --help

Quick start

pip install rupudata

rupudata scan dataset.jsonl
rupudata compare train.jsonl eval.jsonl
rupudata benchmark-check train.jsonl --benchmark gsm8k

Three commands. One goal: understand whether data overlaps where it shouldn't.

Real-world example (GSM8K)

Against the full GSM8K test split as --reference (1,319 questions), a 7,476-row training file with three injected test questions reports:

rupudata benchmark-check leak.jsonl \
  --benchmark gsm8k \
  --reference gsm8k_test.jsonl
RupuData v0.6.5

Dataset:       7,476 records
Reference:     GSM8K test — 1,319 records (user_reference)

Exact matches:       3
Normalized matches:  3
Near matches:        0

Status: OVERLAP_DETECTED

Evidence
  dataset_row  reference_row  field
  0            0              question
  1            1              question
  2            2              question

3 exact GSM8K test-set overlaps detected under the configured matching methodology.

RupuData reports textual overlap. It does not determine why the overlap exists or certify that a model or dataset is contaminated.

On the clean GSM8K train split (7,473 records) alone:

rupudata scan gsm8k_train.jsonl

→ no exact duplicates; only 2 lexical near-duplicate pairs flagged (Jaccard ≥ 0.85, MinHash/LSH).

7,476 training records
        │
        ▼
     RupuData
        │
        ├── scan: 0 exact duplicates, 2 near-dupe pairs
        │
        ▼
   GSM8K test (1,319) via --reference
        │
        ▼
   3 exact overlaps → OVERLAP_DETECTED

Benchmark reference: demo sample vs real audit

Warning: Without --reference, default gsm8k uses a tiny packaged sample. A NO_OVERLAP_DETECTED against that sample does not mean your dataset is free of GSM8K. For an actual audit, pass the real export with --reference.

rupudata benchmark-check train.jsonl --benchmark gsm8k --reference /path/to/gsm8k.jsonl

What each command answers

scan
  → Do I have duplicates?

compare
  → Do train and eval share data?

benchmark-check
  → Does my dataset contain known benchmark text?

Formats & reports

  • Input: JSONL and Parquet
  • Output: terminal summary + JSON audit contract (input → configuration → method → result)
  • Defaults: rupudata-report.json, rupudata-compare.json, rupudata-benchmark.json
rupudata compare a.jsonl b.jsonl -o reports/compare.json

Compare full records by default. For a single text column:

rupudata compare train.jsonl eval.jsonl --text-field text

Try the packaged examples

git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install rupudata

rupudata compare examples/train.jsonl examples/eval.jsonl
rupudata benchmark-check examples/train_with_gsm8k_overlap.jsonl --benchmark gsm8k
rupudata scan examples/near_dupes.jsonl --near-duplicate-threshold 0.85

Status

0.6.5 — contextual benchmark notes + real-world GSM8K README demo.

Not in this release (on purpose): semantic / paraphrase matching, CI fail gates, streaming multi-GB scans, provenance/license detectors.

Next: driven by real usage. Likely first bridge: exit codes for pipelines (--fail-on-overlap). Matching model stays stable unless users need a change.

What RupuData does not do

  • determine legal ownership or certify copyright compliance
  • guarantee a dataset is legally safe
  • detect every form of benchmark contamination
  • claim semantic understanding of near-duplicates
  • replace specialized license scanners or large-scale frameworks

Technical audit contract

For readers who need to reproduce findings exactly. JSON reports are shaped so a third party can audit the method:

input → configuration → method → result

Matching model (source → unit → spec)

method.unit Text source Specs
full_record whole record record_exact_v1 / record_normalized_v1
field_text explicit method.field (--text-field) text_exact_v1 / text_normalized_v1
extracted_text text_extraction (e.g. first_non_empty) text_exact_v1 / text_normalized_v1

text_* specs define how a plain text value is hashed. They do not define where the text came from — the command declares the source.

Matching vocabulary by command

Command Field Spec / unit
scan exact_duplicates record_normalized_v1 (unit=full_record)
compare exact_overlap / normalized_overlap record_* or text_* per method.unit
benchmark-check exact / normalized text_* with unit=extracted_text

Compare match evidence: exact lists row pairs only; normalized always includes also_exact, and adds differing_fields (full record) or difference (field text) only when also_exact is false.

What “normalized” means (fingerprint + scan exact duplicates)

Policy id: record_normalized_v1.

Step Behavior
Fields All record fields
Keys Sorted recursively
Strings str.strip() only
Internal spaces Kept
Case / Unicode No lowercasing, no NFC/NFKC
Hash SHA-256 over compact sorted JSON

Fingerprint = SHA-256 over the sorted multiset of per-record hashes → rupu: + 16 hex chars.

Near-duplicates use a different text prep (near_text_v1: lowercase + collapse whitespace). Lexical similarity only — not paraphrases.

rupudata scan dataset.parquet --skip-near-duplicates

Development

git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install -e ".[dev]"
pytest
python -m build

License

Apache License 2.0

Release files for rupudata 0.6.5

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rupudata 0.6.5
File Size Uploaded
rupudata-0.6.5.tar.gz 37.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rupudata 0.6.5
File Interpreter ABI Platform
rupudata-0.6.5-py3-none-any.whl Python 3 none any Details

Total release size: 70.7 kB

Release files / rupudata-0.6.5.tar.gz

Download URL rupudata-0.6.5.tar.gz
Size 37.0 kB
Tags Source
SHA-256 checksum
How to use checksums
7e835a471c001b2bdbcf3b026f140aa36fe1cd4a98ab2ac278f152b943e9e655
BLAKE2b-256 checksum
How to use checksums
a1cb1cb76357b7743f47f23271c92f051abc8fe5d644ffd01539f06c14d1bd36
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release files / rupudata-0.6.5-py3-none-any.whl

Download URL rupudata-0.6.5-py3-none-any.whl
Size 33.7 kB
Tags Python 3
SHA-256 checksum
How to use checksums
ffc734eee7afdd0a2c5391b4ff0024bbdb5f9fd0d0e4b56ea891543d76f9c6d2
BLAKE2b-256 checksum
How to use checksums
1e64cd3454d4f41a4f69fbb2e04ad016cf1e0f0ff2236280ae901f851eb73a6d
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release history Release notifications | RSS feed

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.6

2 release files

This release

0.6.5 This release

2 release files

0.6.4

2 release files

0.6.3

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page