Skip to main content

RupuData

Language: English | Español

Follow the path of your data.

RupuData detects dataset duplicates, train/eval overlap, and benchmark overlap signals — locally, with deterministic, machine-readable audit reports.

Compare datasets, scan for duplicates, and check whether training data overlaps known evaluation benchmarks. Technical signals, not legal certification.

Install

Requires Python 3.9+.

pip install rupudata
rupudata --help

Quick start

pip install rupudata

rupudata scan dataset.jsonl
rupudata compare train.jsonl eval.jsonl
rupudata benchmark-check train.jsonl --benchmark gsm8k

Three commands. One goal: understand whether data overlaps where it shouldn't.

Example output (compare)

From the repo examples/:

RupuData v0.6.3

Exact overlap:       1
Normalized overlap:  2

Evidence (sample)
  dataset_a_row  dataset_b_row
  0              0

Benchmark reference: demo sample vs real audit

Status is OVERLAP_DETECTED or NO_OVERLAP_DETECTED under the matching methodology — not a claim that a model is contaminated.

Warning: The packaged GSM8K reference is a tiny demo sample, not the full benchmark. A NO_OVERLAP_DETECTED result against the sample does not mean your dataset is free of GSM8K. For an actual audit, provide the benchmark export with --reference.

rupudata benchmark-check train.jsonl --benchmark gsm8k --reference /path/to/gsm8k.jsonl

What each command answers

scan
  → Do I have duplicates?

compare
  → Do train and eval share data?

benchmark-check
  → Does my dataset contain known benchmark text?

Formats & reports

  • Input: JSONL and Parquet
  • Output: terminal summary + JSON audit contract (input → configuration → method → result)
  • Defaults: rupudata-report.json, rupudata-compare.json, rupudata-benchmark.json
rupudata compare a.jsonl b.jsonl -o reports/compare.json

Compare full records by default. For a single text column:

rupudata compare train.jsonl eval.jsonl --text-field text

Try the packaged examples

Clone the repo (or download examples/ from GitHub):

git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install rupudata

rupudata compare examples/train.jsonl examples/eval.jsonl
rupudata benchmark-check examples/train_with_gsm8k_overlap.jsonl --benchmark gsm8k
rupudata scan examples/near_dupes.jsonl --near-duplicate-threshold 0.85

Status

0.6.3 — product release: install from PyPI, user-facing docs, CI.

Not in this release (on purpose): semantic / paraphrase matching, CI fail gates, streaming multi-GB scans, provenance/license detectors.

Next: driven by real usage. Likely first bridge: exit codes for pipelines (--fail-on-overlap). Matching model stays stable unless users need a change.

What RupuData does not do

  • determine legal ownership or certify copyright compliance
  • guarantee a dataset is legally safe
  • detect every form of benchmark contamination
  • claim semantic understanding of near-duplicates
  • replace specialized license scanners or large-scale frameworks

Technical audit contract

For readers who need to reproduce findings exactly. JSON reports are shaped so a third party can audit the method:

input → configuration → method → result

Matching model (source → unit → spec)

method.unit Text source Specs
full_record whole record record_exact_v1 / record_normalized_v1
field_text explicit method.field (--text-field) text_exact_v1 / text_normalized_v1
extracted_text text_extraction (e.g. first_non_empty) text_exact_v1 / text_normalized_v1

text_* specs define how a plain text value is hashed. They do not define where the text came from — the command declares the source.

Matching vocabulary by command

Command Field Spec / unit
scan exact_duplicates record_normalized_v1 (unit=full_record)
compare exact_overlap / normalized_overlap record_* or text_* per method.unit
benchmark-check exact / normalized text_* with unit=extracted_text

Compare match evidence: exact lists row pairs only; normalized always includes also_exact, and adds differing_fields (full record) or difference (field text) only when also_exact is false.

What “normalized” means (fingerprint + scan exact duplicates)

Policy id: record_normalized_v1.

Step Behavior
Fields All record fields
Keys Sorted recursively
Strings str.strip() only
Internal spaces Kept
Case / Unicode No lowercasing, no NFC/NFKC
Hash SHA-256 over compact sorted JSON

Fingerprint = SHA-256 over the sorted multiset of per-record hashes → rupu: + 16 hex chars.

Near-duplicates use a different text prep (near_text_v1: lowercase + collapse whitespace). Lexical similarity only — not paraphrases.

rupudata scan dataset.parquet --skip-near-duplicates

Development

git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install -e ".[dev]"
pytest
python -m build

License

Apache License 2.0

Release files for rupudata 0.6.3

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for rupudata 0.6.3
File Size Uploaded
rupudata-0.6.3.tar.gz 36.0 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for rupudata 0.6.3
File Interpreter ABI Platform
rupudata-0.6.3-py3-none-any.whl Python 3 none any Details

Total release size: 69.1 kB

Release files / rupudata-0.6.3.tar.gz

Download URL rupudata-0.6.3.tar.gz
Size 36.0 kB
Tags Source
SHA-256 checksum
How to use checksums
9526549af5a98c7a18f67927a084926573230a22c4cc00013b8084a8c077f8e7
BLAKE2b-256 checksum
How to use checksums
96e40ccf6d647bfedfd9fb77dd1710ac2cafa287f6029dd9998e9ade36440643
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release files / rupudata-0.6.3-py3-none-any.whl

Download URL rupudata-0.6.3-py3-none-any.whl
Size 33.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
98ea57bc2584d754c4f6a0e99ca97013c75213489e3af2141555e9be2035de28
BLAKE2b-256 checksum
How to use checksums
b0236547281c0335172dbbcc62c524a664508a771338142843437c756f130524
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/6.2.0 CPython/3.9.6

Release history Release notifications | RSS feed

0.9.0

2 release files

0.8.1

2 release files

0.8.0

2 release files

0.7.0

2 release files

0.6.6

2 release files

0.6.5

2 release files

0.6.4

2 release files

This release

0.6.3 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page