Skip to main content

Reference-free auditing of synthetic datasets for generator artifacts before machine learning

Project description

SynthAudit logo

CI Docs License: Apache-2.0 Python 3.11 | 3.12 Linted with Ruff Release

SynthAudit is a reference-free auditor for synthetic and simulated tabular datasets. Point it at a released CSV, with no access to the real source data and no access to the generator, and it recovers the generator's fingerprints before you train anything: exact equations between columns, labels that are thresholds or rules over shipped features, functional dependencies, balance constraints, schedule leakage, duplicated and constant columns. It then classifies every column, separates benchmark leakage from legitimate physical structure, and scores the release with a decomposable Benchmark Trustworthiness Index (BTI).

Why this matters: many widely used synthetic benchmarks quietly ship their own answer key, and models "solve" the generator rather than the task. Findings from the audits in reports/, each produced by one command:

Dataset What SynthAudit found (from the file alone) Grade
Electrical Grid Stability (UCI) stabf == 'unstable' iff stab > 0, the label is the sign of a shipped column, plus the power balance p4 = -(p1+p2+p3) F
AI4I 2020 (UCI) Machine failure = rules(TWF, HDF, PWF, OSF) at fidelity 0.9991, and the 9 rule-violating rows are the dataset's independently documented generator bugs F
Synthea medications TOTALCOST = BASE_COST x DISPENSES, exact on all 42,989 rows F
NSL-KDD Label is clean (the famous deduplication worked), but 610 rows of the official test file appear verbatim in the training file B
BATADAL Dozens of hydraulic identities, all correctly classified as physics rather than leakage; attack windows mandate temporal splits A
Tennessee Eastman 19 structural constraints, zero target leakage: determinism is not punished, shipping the answer is B

The methodology, the index, formal guarantees, a planted-artifact validation testbed (full recall, zero false positives), a leave-one-module-out ablation, and twelve cross-domain case studies are described in the accompanying paper (see Citation).

Installation

pip install synthaudit            # core  (PyPI release pending; see below)
pip install "synthaudit[causal]"  # + causal-learn for the causal structure scan

Until the PyPI release lands, install from GitHub:

pip install "synthaudit[causal] @ git+https://github.com/rvndye/synthaudit.git"

Quick start

import pandas as pd
from synthaudit import Audit

df = pd.read_csv("datasets/grid_stability.csv")

audit = Audit(df, target="stabf", name="grid_stability",
              metadata={"generator_described": True,
                        "generator_code_available": False,
                        "seed_reported": False,
                        "artifacts_disclosed": True})
results = audit.run()

audit.generate_report("grid_report.html")   # self-contained HTML report
audit.export_pdf("grid_scorecard.pdf")      # one-page scorecard
audit.export_json("grid_audit.json")        # machine-readable findings
print(audit.bti, results["scoring"]["grade"])

Or from the command line, including as a CI gate for your data pipeline:

synthaudit audit datasets/grid_stability.csv --target stabf --html report.html
synthaudit audit train.csv --target label --fail-below C   # nonzero exit if grade < C
synthaudit selftest                                        # planted-artifact self-validation

SynthAudit HTML report for the Grid Stability dataset

What it detects

Artifact class Example Detector
Linear identities cost = power x tariff Gram-based minimal OLS with backward elimination
Balance constraints p4 = -(p1+p2+p3) same miner, symmetric constraint collapse
Power laws TOTALCOST = BASE_COST x DISPENSES log-space mining
Regime-affine equations eff = base(type) - 0.015*age fixed-effects demeaning
Rule-derived labels AI4I's failure flag shallow-tree fidelity with imbalance guards
Threshold / sign labels stabf = sign(stab) exact separation scan
Functional dependencies CODE -> DESCRIPTION g3 mining with non-degeneracy guard
Chained derivations duplicate of x re-expresses every identity of x iterative peeling
Residual determinism nonlinear derived columns out-of-sample GBM sweep (skill-banded)
Schedule leakage attacks in contiguous windows target autocorrelation screen
Contamination train/test row overlap hash join across supplied splits

Every recovered relation carries a disposition: target_leakage (the posed task is answerable by arithmetic), structural_constraint (legitimate physics or accounting among non-target columns), or redundancy. Conservation laws are not crimes; shipping the label's source column is.

The Benchmark Trustworthiness Index

Five measured pillars, each in [0, 1] with an explicit estimator: Label integrity, Feature integrity, difficulty Headroom, sampling Realism, Information density, plus optional metadata-driven Transparency. Aggregation is a weighted geometric mean, which is provably non-compensatory: a dataset whose label is exactly recoverable is capped in the F band no matter how clean everything else is, and a release disclosing nothing about its generator caps near C. The pillar vector and the recovered evidence ship with every score; the scalar alone is never the deliverable. Grades map to actions: A/B use with the recommended feature view, C repair and re-audit, D/F the shipped task measures generator recovery.

API overview

Object / function Purpose
Audit(data, target, name, metadata, test_data, seed, modules, identity_kwargs) orchestrates the seven audit modules
Audit.run() returns the full results tree (profile, identity, determinism, causal, leakage, taxonomy, scoring, recommendations)
Audit.generate_report / export_html / export_pdf / export_json / export_notebook deliverables
Audit.bti the scalar index (pillar vector lives in results["scoring"])
synthaudit.make_planted(n, seed, extended) the planted-artifact validation testbed
synthaudit.score_detection(results, truth) recall and negative-control scoring against planted ground truth

Full reference: documentation site (installation, tutorial, theory, architecture, BTI, artifact taxonomy, API, developer guide, FAQ).

Reproducing the paper's results

make selftest          # planted-artifact validation (recall + negative controls)
make audit-example     # audits the three bundled datasets into reports/
bash scripts/download_datasets.sh   # fetches the remaining public cohort datasets
python scripts/audit_cohort.py      # regenerates reports/ and benchmarks/ tables

Bundled small datasets (with attribution in datasets/README.md): AI4I 2020 and Electrical Grid Stability (UCI, CC BY 4.0) and one Synthea table (Apache-2.0). Audit outputs for all twelve cohort datasets are in reports/; ablation and stability tables in benchmarks/.

Citation

If you use SynthAudit in your research, please cite the software release:

@software{erwin2026synthaudit,
  author  = {Erwin, Randy},
  title   = {SynthAudit: Reference-Free Auditing of Synthetic Datasets
             for Generator Artifacts Before Machine Learning},
  year    = {2026},
  version = {0.1.0},
  url     = {https://github.com/rvndye/synthaudit}
}

The methodology paper is under submission (JMLR); the entry above will be updated on acceptance, and CITATION.cff carries the canonical metadata.

Contributing

Contributions are welcome: new artifact detectors, dataset audits, docs, and bug reports alike. Start with CONTRIBUTING.md and the developer guide. Found a generator artifact in a public benchmark using SynthAudit? Open an issue with the Audit finding template; confirmed findings are collected in the documentation.

Roadmap

Planned directions, roughly ordered: temporal auditing (lagged identities, dynamics mining, split-protocol certification), relational multi-table audits (cross-table dependency graphs), an opt-in symbolic-regression backend for out-of-class equations, latent-regime and masked-category scans, a multi-probe headroom panel, and auditors for LLM-generated benchmarks. See the open issues for current status.

License

Apache License 2.0. Bundled datasets keep their original licenses, documented in datasets/README.md.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

synthaudit-0.1.0.tar.gz (47.1 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

synthaudit-0.1.0-py3-none-any.whl (47.1 kB view details)

Uploaded Python 3

File details

Details for the file synthaudit-0.1.0.tar.gz.

File metadata

  • Download URL: synthaudit-0.1.0.tar.gz
  • Upload date:
  • Size: 47.1 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for synthaudit-0.1.0.tar.gz
Algorithm Hash digest
SHA256 96b046ffbb27ccd13fd78b59abf7fab7eb046e93fce2f3df0cd90c2de044ac78
MD5 b3d06afd75623db4ca78efaededadeab
BLAKE2b-256 8dd4a36b4209040fd3ac11f0bca034342d6b269073e61d5836aab412cf74af79

See more details on using hashes here.

File details

Details for the file synthaudit-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: synthaudit-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 47.1 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.9.6

File hashes

Hashes for synthaudit-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 ae9a357c127d79d6e94dc1b9e6acd8206ef7accae7cf8024dd0899a6f092034b
MD5 ab2a68374136fb09255afc8b973cc7a0
BLAKE2b-256 6937c06b70ceb4707309f09e9f169624d762b11344be811a9bb92cae711dfbf7

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page