Skip to main content

A user-friendly synthetic data generator with automatic topological sorting

Project description

DataSandbox: Synthetic Data Generator

A lightweight, rule-based engine for generating realistic synthetic datasets. It allows you to define complex relationships between variables using a straightforward mathematical schema, ensuring your generated data maintains logical consistency and statistical correlations.


1. Why is this needed?

In modern data science and software development, getting access to high-quality, realistic data is often difficult due to privacy regulations, data scarcity, or security concerns. This tool is essential for:

  • Protecting Privacy: Generate mock datasets that mirror real user behavior without exposing sensitive or personally identifiable information.
  • Testing Machine Learning Models: Create controlled, synthetic environments with known correlations to evaluate how well your models capture specific patterns.
  • Stress-Testing Applications: Rapidly produce thousands or millions of rows of relational data to validate database performance and system scalability.
  • Automating Mock Data: Replace tedious manual mock-data creation with an automated, reproducible, and mathematically sound architecture.

2. What tasks does it perform?

  • Dependency Resolution: Automatically calculates the correct execution order for your columns. If one column depends on another, the engine guarantees the parent data is generated first.
  • Distribution Generation: Seamlessly produces normal, uniform, integer, or weighted categorical distributions based on your statistical parameters.
  • Linear Combinations: Calculates continuous variables based on the weighted sum of parent variables, while injecting realistic statistical noise.
  • Conditional Branching: Routes data generation through entirely different logic paths depending on the discrete category of a parent column.
  • DataFrame Export: Compiles all generated relationships directly into a clean, ready-to-use tabular format.

3. User Guide: Understanding Mathematical Categories

The core philosophy of this tool is that every column in your dataset is defined by a specific mathematical rule. You organize these rules into a schema, and the engine handles the rest.

Here are the three mathematical categories you can apply to your columns, what they represent, and how they function:

A. The Independent Rule

What it means: This is a base-level generator. A column with this rule relies on absolutely nothing else in the dataset; it is mathematically independent. When to use it: Use this for foundational data points that kick off your logic tree, such as Age, Gender, Base Economy, or starting geographic zones. How to apply it: You define the core statistical method you want to use (such as generating random integers, floats, or picking from a list of choices) and provide the necessary statistical boundaries, like minimums, maximums, or probability weights.

B. The Conditional Rule

What it means: This is a branching generator. The output of a conditional column changes entirely depending on the specific value generated by a single parent column. When to use it: This is perfect for cascading categorical logic. For example, a person's expected education level or housing status might have completely different probability distributions depending on which city tier they live in. How to apply it: You specify which parent column to watch, and then provide a map of conditions. For every possible value the parent column might output, you assign a brand new Independent rule to dictate how the child column should behave in that specific scenario.

C. The Linear Rule

What it means: This is a continuous correlation generator. It creates numerical values by applying mathematical weights to one or more parent columns, adding a starting baseline, and injecting random standard deviation to make the data look organic. When to use it: Use this when numerical values naturally scale based on other numbers in your dataset. For instance, a salary typically increases as years of experience increase, or a monthly spend increases based on a combination of salary and credit score. How to apply it: You define exactly which parent columns influence the result and by what multiplier. Then, you provide a base bias (the starting value before any additions) and define the noise level to ensure the resulting correlation isn't perfectly artificial.

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

datasandbox-0.1.0.tar.gz (4.6 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

datasandbox-0.1.0-py3-none-any.whl (4.8 kB view details)

Uploaded Python 3

File details

Details for the file datasandbox-0.1.0.tar.gz.

File metadata

  • Download URL: datasandbox-0.1.0.tar.gz
  • Upload date:
  • Size: 4.6 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for datasandbox-0.1.0.tar.gz
Algorithm Hash digest
SHA256 7feeffef6c516c484cf47622ed17f7c05bf7abbd4dc8521f56e0fca42490b6db
MD5 218f856b9106ca7901fa13dfe11fe6cf
BLAKE2b-256 8eaae0eca03753a31af6a887bc555ca4c724797e7e1c87808ff7b64efceecf08

See more details on using hashes here.

File details

Details for the file datasandbox-0.1.0-py3-none-any.whl.

File metadata

  • Download URL: datasandbox-0.1.0-py3-none-any.whl
  • Upload date:
  • Size: 4.8 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? No
  • Uploaded via: twine/6.2.0 CPython/3.12.3

File hashes

Hashes for datasandbox-0.1.0-py3-none-any.whl
Algorithm Hash digest
SHA256 6d171b9f54e769130f2f891a383746d07814efd20d1f33f848e041f3232157f4
MD5 b2c458057fcb0d9a40f211f42e540ddc
BLAKE2b-256 0a3e0708ba305222db64306802a5de22199756d3ca4ad9feda78a42c6b6644c3

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page