Skip to main content
synthbench

PyPI Python CI Docs License: MIT

synthbench generates synthetic datasets for benchmarking. You choose how complicated the signal is, layer noise or missing data on top, and get back a dataset that carries a record of how it was made. One integer seed reproduces the whole thing.

The point is knowing the answer in advance. On real data you can measure that your model scored 0.83 and have no idea whether 0.85 was available or whether you already hit the ceiling. Here the ceiling is in the metadata.

Installation

pip install synthbench

Parquet serialization needs pip install "synthbench[io]", and RandomNeuralDGP needs pip install "synthbench[neural]" for PyTorch. Everything else works from the base install.

Basic usage

from synthbench import BenchPipeline, LinearDGP, MissingDataCorruptor

pipeline = BenchPipeline(
    LinearDGP(complexity="medium", task_type="classification"),
    corruptors=[MissingDataCorruptor(proportion=0.1, mechanism="mar")],
)
result = pipeline.run(n_samples=500, n_features=10, random_state=42)

print(result.X.shape)  # (500, 10)
print(result.metadata["bayes_error_analytic"])  # exact error floor
print(result.metadata["effective_rank"])  # feature space dimensionality

Data-generating processes

Each one takes a complexity parameter and records which features actually carry signal.

DGP Signal
LinearDGP Linear combination, sparsity and noise set by complexity
PolynomialDGP Polynomial terms and interactions
TreeDGP Axis-aligned splits, depth set by complexity
FriedmanDGP The Friedman 1/2/3 benchmark functions
AdditiveDGP Sum of univariate functions, GAM-style
SparseDGP Explicit number of informative features via k
GeometricDGP Moons, circles, spirals
RandomNeuralDGP A randomly initialised MLP (needs [neural])

Corruptors

MeasurementNoise, Outlier, MissingData, Collinearity, and Categorical transform the feature matrix; LabelNoiseCorruptor goes in the separate label_corruptors= argument and touches only y. Feature corruptors always run in a fixed order regardless of how you list them, so two pipelines with the same components produce the same data. Each one records how much of each feature's information it destroyed, in metadata["effective_feature_importances"].

Severity is a preset: "low", "medium", or "high". Pass the underlying parameter (proportion, noise_level, n_bins) when you want a specific value instead.

More than two classes

Pass n_classes to any DGP except FriedmanDGP, with class_weight="balanced" or an explicit list of priors.

from synthbench import BenchPipeline, LinearDGP

result = BenchPipeline(
    LinearDGP(
        task_type="classification", n_classes=4, class_weight=[0.4, 0.3, 0.2, 0.1]
    )
).run(n_samples=1000, n_features=10, random_state=0)

How the classes are produced differs by DGP, and metadata["label_mechanism"] says which you got. LinearDGP, SparseDGP, TreeDGP, and RandomNeuralDGP emit one latent score per class and draw through a softmax, so the classes are unordered. PolynomialDGP and AdditiveDGP are built on a fixed term basis with only one signal to work with, so theirs is cut into ordered bins — fine for genuinely ordinal targets, misleading as a general multiclass benchmark. FriedmanDGP refuses: its functions are specific published formulas, and a multi-output variant would not be the thing anyone cites.

How hard is the dataset?

Two numbers, and the difference between them matters:

  • bayes_error_analytic is exact. The DGP draws each label from a known probability, so the irreducible error is computable rather than estimated. No model can do better. It reflects the labels only, so feature corruption does not change it.
  • bayes_error is an empirical 1-NN leave-one-out error on the corrupted features. It moves when you corrupt things, which makes it useful for comparing severity levels against each other, but it is biased upwards in more than a few dimensions. It is not a bound.

For a balanced LinearDGP classification task the analytic floor is about 0.325, while the 1-NN estimate reads about 0.456 at n=500, p=10. Use the first when you need a floor and the second when you need a relative signal.

Ablations

severity_sweep and difficulty_sweep vary one axis and return a dict keyed by level. experiment_grid crosses sample size, complexity, and severity. Seeds come from a nested SeedSequence hierarchy, so cells are independent of each other but reproducible across runs.

from synthbench import LinearDGP, OutlierCorruptor, experiment_grid

grid = experiment_grid(
    LinearDGP,
    OutlierCorruptor,
    n_samples_list=[200, 500, 1000],
    complexities=["low", "medium", "high"],
    severities=["low", "medium", "high"],
    n_features=10,
    random_state=0,
    task_type="classification",
)

print(grid[(500, "high", "medium")].metadata["bayes_error_analytic"])

Suites and serialization

BenchSuite("easy-classification").run() generates a curated collection in one call and returns it keyed by label — handy as a shared baseline, since the suite name alone pins the data. BenchSuite.from_dict and BenchSuite.from_json take your own specs.

to_parquet / to_csv round-trip the data and the full metadata, and BenchPipeline.from_metadata rebuilds the pipeline from that metadata so a saved dataset can be regenerated rather than shipped.

Docs

Full reference and runnable notebooks at JanTeichertKluge.github.io/synth-bench.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

synthbench-0.2.0.tar.gz (576.0 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

synthbench-0.2.0-py3-none-any.whl (67.7 kB view details)

Uploaded Python 3

File details

Details for the file synthbench-0.2.0.tar.gz.

File metadata

  • Download URL: synthbench-0.2.0.tar.gz
  • Upload date:
  • Size: 576.0 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for synthbench-0.2.0.tar.gz
Algorithm Hash digest
SHA256 144d9683a014e1c7c20feac40e2d90e1f821784d8bed7d3774f7bdca8b34192a
MD5 55c1a75a0647a5f11a48f349cafb8703
BLAKE2b-256 72496364a82c30e24701a0f325cdde1386da196ffc9c554c15de1aa36d0ec74e

See more details on using hashes here.

Provenance

The following attestation bundles were made for synthbench-0.2.0.tar.gz:

Publisher: publish.yml on JanTeichertKluge/synth-bench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file synthbench-0.2.0-py3-none-any.whl.

File metadata

  • Download URL: synthbench-0.2.0-py3-none-any.whl
  • Upload date:
  • Size: 67.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for synthbench-0.2.0-py3-none-any.whl
Algorithm Hash digest
SHA256 3b485eac24db02524a1a8edac0b214c7f2fec2d3796b1cfe02583007a44b94d4
MD5 7f2caab3587f8e08c0a77ff942fb8789
BLAKE2b-256 fb3c3ab8e1df60a486a23091cc657557be60fdbbfeff0df9b140c1c382b0158c

See more details on using hashes here.

Provenance

The following attestation bundles were made for synthbench-0.2.0-py3-none-any.whl:

Publisher: publish.yml on JanTeichertKluge/synth-bench

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

This release

0.2.0 This release

2 files

0.1.0

2 files

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page