RupuData
Follow the path of your data.
RupuData answers an uncomfortable question: does my training data overlap with data I use to evaluate my model?
It detects dataset duplicates, train/eval overlap, and benchmark overlap signals — locally, with deterministic, machine-readable audit reports. Technical signals, not legal certification.
Install
Requires Python 3.9+.
pip install rupudata
rupudata --help
Quick start
pip install rupudata
rupudata scan dataset.jsonl
rupudata compare train.jsonl eval.jsonl
rupudata benchmark-check train.jsonl --benchmark gsm8k
Three commands. One goal: understand whether data overlaps where it shouldn't.
Real-world example (GSM8K)
Against the full GSM8K test split as --reference (1,319 questions), a 7,476-row training file with three injected test questions reports:
rupudata benchmark-check leak.jsonl \
--benchmark gsm8k \
--reference gsm8k_test.jsonl
RupuData v0.6.5
Dataset: 7,476 records
Reference: GSM8K test — 1,319 records (user_reference)
Exact matches: 3
Normalized matches: 3
Near matches: 0
Status: OVERLAP_DETECTED
Evidence
dataset_row reference_row field
0 0 question
1 1 question
2 2 question
3 exact GSM8K test-set overlaps detected under the configured matching methodology.
RupuData reports textual overlap. It does not determine why the overlap exists or certify that a model or dataset is contaminated.
On the clean GSM8K train split (7,473 records) alone:
rupudata scan gsm8k_train.jsonl
→ no exact duplicates; only 2 lexical near-duplicate pairs flagged (Jaccard ≥ 0.85, MinHash/LSH).
7,476 training records
│
▼
RupuData
│
├── scan: 0 exact duplicates, 2 near-dupe pairs
│
▼
GSM8K test (1,319) via --reference
│
▼
3 exact overlaps → OVERLAP_DETECTED
Benchmark reference: demo sample vs real audit
Warning: Without
--reference, defaultgsm8kuses a tiny packaged sample. ANO_OVERLAP_DETECTEDagainst that sample does not mean your dataset is free of GSM8K. For an actual audit, pass the real export with--reference.
rupudata benchmark-check train.jsonl --benchmark gsm8k --reference /path/to/gsm8k.jsonl
What each command answers
scan
→ Do I have duplicates?
compare
→ Do train and eval share data?
benchmark-check
→ Does my dataset contain known benchmark text?
Formats & reports
- Input: JSONL and Parquet
- Output: terminal summary + JSON audit contract (
input → configuration → method → result) - Defaults:
rupudata-report.json,rupudata-compare.json,rupudata-benchmark.json
rupudata compare a.jsonl b.jsonl -o reports/compare.json
Compare full records by default. For a single text column:
rupudata compare train.jsonl eval.jsonl --text-field text
Try the packaged examples
git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install rupudata
rupudata compare examples/train.jsonl examples/eval.jsonl
rupudata benchmark-check examples/train_with_gsm8k_overlap.jsonl --benchmark gsm8k
rupudata scan examples/near_dupes.jsonl --near-duplicate-threshold 0.85
Status
0.6.5 — contextual benchmark notes + real-world GSM8K README demo.
Not in this release (on purpose): semantic / paraphrase matching, CI fail gates, streaming multi-GB scans, provenance/license detectors.
Next: driven by real usage. Likely first bridge: exit codes for pipelines (--fail-on-overlap). Matching model stays stable unless users need a change.
What RupuData does not do
- determine legal ownership or certify copyright compliance
- guarantee a dataset is legally safe
- detect every form of benchmark contamination
- claim semantic understanding of near-duplicates
- replace specialized license scanners or large-scale frameworks
Technical audit contract
For readers who need to reproduce findings exactly. JSON reports are shaped so a third party can audit the method:
input → configuration → method → result
Matching model (source → unit → spec)
method.unit |
Text source | Specs |
|---|---|---|
full_record |
whole record | record_exact_v1 / record_normalized_v1 |
field_text |
explicit method.field (--text-field) |
text_exact_v1 / text_normalized_v1 |
extracted_text |
text_extraction (e.g. first_non_empty) |
text_exact_v1 / text_normalized_v1 |
text_* specs define how a plain text value is hashed. They do not define where the text came from — the command declares the source.
Matching vocabulary by command
| Command | Field | Spec / unit |
|---|---|---|
scan |
exact_duplicates |
record_normalized_v1 (unit=full_record) |
compare |
exact_overlap / normalized_overlap |
record_* or text_* per method.unit |
benchmark-check |
exact / normalized |
text_* with unit=extracted_text |
Compare match evidence: exact lists row pairs only; normalized always includes also_exact, and adds differing_fields (full record) or difference (field text) only when also_exact is false.
What “normalized” means (fingerprint + scan exact duplicates)
Policy id: record_normalized_v1.
| Step | Behavior |
|---|---|
| Fields | All record fields |
| Keys | Sorted recursively |
| Strings | str.strip() only |
| Internal spaces | Kept |
| Case / Unicode | No lowercasing, no NFC/NFKC |
| Hash | SHA-256 over compact sorted JSON |
Fingerprint = SHA-256 over the sorted multiset of per-record hashes → rupu: + 16 hex chars.
Near-duplicates use a different text prep (near_text_v1: lowercase + collapse whitespace). Lexical similarity only — not paraphrases.
rupudata scan dataset.parquet --skip-near-duplicates
Development
git clone https://github.com/EmanuelCorreaAR/rupudata.git
cd rupudata
pip install -e ".[dev]"
pytest
python -m build
License
Apache License 2.0
Release files for rupudata 0.6.5
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rupudata-0.6.5.tar.gz | 37.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rupudata-0.6.5-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 70.7 kB
Release files / rupudata-0.6.5.tar.gz
| Download URL | rupudata-0.6.5.tar.gz |
|---|---|
| Size | 37.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
7e835a471c001b2bdbcf3b026f140aa36fe1cd4a98ab2ac278f152b943e9e655
|
|
BLAKE2b-256 checksum How to use checksums |
a1cb1cb76357b7743f47f23271c92f051abc8fe5d644ffd01539f06c14d1bd36
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|
Release files / rupudata-0.6.5-py3-none-any.whl
| Download URL | rupudata-0.6.5-py3-none-any.whl |
|---|---|
| Size | 33.7 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
ffc734eee7afdd0a2c5391b4ff0024bbdb5f9fd0d0e4b56ea891543d76f9c6d2
|
|
BLAKE2b-256 checksum How to use checksums |
1e64cd3454d4f41a4f69fbb2e04ad016cf1e0f0ff2236280ae901f851eb73a6d
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.6
|