Template-driven synthetic data generation engine.
Project description
pysynthgen
Template-driven synthetic data generation engine.
pysynthgen takes a validated JSON template describing a dataset and yields synthetic
rows as an Iterator[dict], writing them to a configurable sink (JSON, CSV, Parquet,
or Avro).
Installation
pip install pysynthgen
That gives you the engine and the JSON and CSV sinks, which need nothing beyond the standard library. Requires Python 3.10+.
The Parquet and Avro sinks depend on pyarrow and fastavro, which are not
installed by default — they are heavy, and most runs don't need them. Pick the extra
for the formats you want:
pip install "pysynthgen[parquet]" # + pyarrow, enables the parquet sink
pip install "pysynthgen[avro]" # + fastavro, enables the avro sink
pip install "pysynthgen[all]" # both
The quotes matter in zsh, which would otherwise treat the brackets as a glob.
You only pay for what you install: import pysynthgen never imports pyarrow or
fastavro, and each sink imports its backend lazily when you build it. Asking for a
format whose extra is missing raises an error naming the extra to install.
Design
- Engine is standalone and dependency-light. It takes a validated template and returns an iterator of rows; the heavy format dependencies are optional extras.
- Templates are validated with Pydantic, not raw dict parsing. Each field type
is its own model in a discriminated union keyed on
type, so validation is type-specific and mistakes are caught up front. - Generation is seeded and reproducible. A single
seedat the template level drives all randomness (numeric, Faker, regex), so a run is byte-for-byte repeatable. Output depends only on the template and its seed — never on how you consume it, soiter_rows()anditer_batches(n)agree for anyn. - Generation is vectorized. Rows are built a column at a time: each field's
values are drawn for a whole chunk of rows in one numpy call, then transposed into
row dicts. Field types that can't vectorize (
faker,regex) fall back to a per-row draw. - Streaming = chunked generation. The engine exposes
iter_rows()anditer_batches(batch_size); the consumer (a sink or a file writer) decides how to persist batches. The internal generation chunk is fixed, so peak memory stays flat however largerow_countgets. - Relationships are out of scope for now but the schema is shaped to grow into
multi-table generation without a rewrite (a future
entities: {name: template}wrapper, and cross-tablereferencefields).
Template format
{
"row_count": 100000,
"seed": 42,
"fields": [
{"name": "user_id", "type": "uuid"},
{"name": "signup_date", "type": "date", "start": "2023-01-01", "end": "2026-01-01"},
{"name": "age", "type": "int", "distribution": "normal", "mean": 35, "stddev": 10, "min": 18},
{"name": "country", "type": "category", "values": ["US", "NL", "DE"], "weights": [0.5, 0.3, 0.2]},
{"name": "email", "type": "faker", "provider": "email"},
{"name": "sku", "type": "regex", "pattern": "[A-Z]{3}-\\d{4}"},
{"name": "referrer_id", "type": "reference", "field": "user_id", "null_probability": 0.7}
],
"constraints": [
{"type": "unique", "fields": ["user_id"]}
]
}
Field types
type |
Produces | Key params |
|---|---|---|
uuid |
UUIDv4 string | — |
date |
date between bounds | start, end |
datetime |
datetime between bounds | start, end |
int |
integer | distribution (uniform/normal), min/max or mean/stddev |
float |
float | same as int |
category |
value from a set | values, optional weights (sum to 1) |
faker |
Faker output | provider (e.g. email), optional max_length |
regex |
string matching a pattern | pattern, optional max_length |
reference |
copy of an earlier field in the same row | field |
String-producing fields (faker, regex) accept an optional max_length that
truncates the generated value. For regex, truncation may break the pattern match.
Every field also accepts null_probability (0–1): the chance of emitting null
for that field on a given row.
Constraints
unique— the listed field(s) form a unique key across generated rows.
Generating templates with AI
If you use Claude Code, the bundled skill in
.claude/skills/pysynthgen-template/ writes
templates for you: describe the dataset in words, or point it at a sample file (CSV,
JSON, Parquet, Avro) and it infers the schema. It ships a profiler you can also run
directly:
python .claude/skills/pysynthgen-template/profile_sample.py sample.csv --rows 2000
Usage
from pysynthgen import load_and_validate_template, SynthEngine, build_sink
spec = load_and_validate_template("template.json")
engine = SynthEngine(spec)
# iterate rows directly
for row in engine.iter_rows():
...
# or write the whole dataset to a sink in one call (batches internally)
sink = build_sink("parquet", "out.parquet")
path = sink.write(engine.iter_rows()) # -> "out.parquet"
# for finer control, drive the batches yourself
sink = build_sink("parquet", "out.parquet")
for batch in engine.iter_batches(batch_size=10_000):
sink.write_batch(batch)
path = sink.finalize()
Command line:
python -m pysynthgen template.json # validate + echo the normalized spec
python -m pysynthgen template.json --rows 20 # print 20 sample rows as JSON
python -m pysynthgen template.json --out data.parquet # generate full dataset to a file
python -m pysynthgen template.json --out data --format avro
Sinks
A sink consumes rows in batches and writes one output artifact. All sinks
implement write_batch(rows) / finalize() -> path, and build_sink(format, path)
selects one by name.
| format | extension | dependency |
|---|---|---|
json |
.json |
stdlib (streamed JSON array) |
csv |
.csv |
stdlib (delimiter/quotechar configurable) |
parquet |
.parquet, .pq |
pyarrow — pysynthgen[parquet] |
avro |
.avro |
fastavro — pysynthgen[avro] |
Notes: parquet/avro infer their schema from the first batch with all columns
nullable; avro writes naive datetimes as UTC so output is deterministic across
machines. Install both format deps with pysynthgen[all].
Benchmarks
benchmarks/bench_sinks.py streams a wide dataset into each sink format and reports
throughput, output size, and peak resident memory sampled while writing. Because
the engine yields rows lazily and sinks write in batches, memory stays flat as row
count grows — a multi-GB file is written with tens of MB of RSS.
python benchmarks/bench_sinks.py --rows 200000 # quick pass
python benchmarks/bench_sinks.py --rows 10000000 # full ~10M-row target (slow, multi-GB)
python benchmarks/bench_sinks.py --rows 1000000 --formats parquet avro --keep
benchmarks/bench_generators.py measures the generator draw itself, comparing a
per-row draw against the column-at-a-time draw the engine uses and checking the two
agree value-for-value:
python benchmarks/bench_generators.py --rows 200000
Development
This repo uses uv:
uv venv
uv pip install -e ".[dev]"
uv run pytest # tests
uv run ruff check . # lint
uv run mypy src/pysynthgen benchmarks/ # types (strict)
Contributing & releases
Development happens on short-lived branches merged into main via squash-merge, with
Conventional Commit PR titles. Merging to
main runs an automated pipeline that computes the next semantic version, tags a
release, and publishes to PyPI. See CONTRIBUTING.md.
Status
- ✅ Template schema + loader (Pydantic, validated)
- ✅ Generation engine — generators +
SynthEngine - ✅ File sinks — json, csv, parquet, avro
- ⬜ More sinks (DB table, Kafka)
- ⬜ Multi-table / linked-entity generation
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file pysynthgen-0.1.3.tar.gz.
File metadata
- Download URL: pysynthgen-0.1.3.tar.gz
- Upload date:
- Size: 41.7 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
27d67685a8aa9a1c8c9594f5276eed12c280e8f46f11bb81656d6211b448900a
|
|
| MD5 |
c643c8a7ea4935cc70c72104193c0b14
|
|
| BLAKE2b-256 |
97b63d56c275e4dac82923e3a85530ab127d5d9a43585a1138e1d8497a91bb0b
|
Provenance
The following attestation bundles were made for pysynthgen-0.1.3.tar.gz:
Publisher:
release.yml on mohsensafari/pysynthgen
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pysynthgen-0.1.3.tar.gz -
Subject digest:
27d67685a8aa9a1c8c9594f5276eed12c280e8f46f11bb81656d6211b448900a - Sigstore transparency entry: 2173219836
- Sigstore integration time:
-
Permalink:
mohsensafari/pysynthgen@87092d2285b74d553a5bc97dff1805b18ad1577a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/mohsensafari
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@87092d2285b74d553a5bc97dff1805b18ad1577a -
Trigger Event:
push
-
Statement type:
File details
Details for the file pysynthgen-0.1.3-py3-none-any.whl.
File metadata
- Download URL: pysynthgen-0.1.3-py3-none-any.whl
- Upload date:
- Size: 23.6 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b3b145eb31f5b70e11ae19e0ac135859b9a78e2a67d1cf47c96eb71939725ec7
|
|
| MD5 |
10d3f06ef502adb94f019cbcc0d4c491
|
|
| BLAKE2b-256 |
c338f45fb95cedab2544ac7dabea51670d98d9bef21eb3e0454211f32068faf6
|
Provenance
The following attestation bundles were made for pysynthgen-0.1.3-py3-none-any.whl:
Publisher:
release.yml on mohsensafari/pysynthgen
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
pysynthgen-0.1.3-py3-none-any.whl -
Subject digest:
b3b145eb31f5b70e11ae19e0ac135859b9a78e2a67d1cf47c96eb71939725ec7 - Sigstore transparency entry: 2173219918
- Sigstore integration time:
-
Permalink:
mohsensafari/pysynthgen@87092d2285b74d553a5bc97dff1805b18ad1577a -
Branch / Tag:
refs/heads/main - Owner: https://github.com/mohsensafari
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@87092d2285b74d553a5bc97dff1805b18ad1577a -
Trigger Event:
push
-
Statement type: