django-data-shape
A realistically shaped test database from Django models.
Declare the shape of your data -- cardinality, value skew, foreign-key fan-out as
a distribution with a long tail, and where related rows physically sit -- then
load it by COPY and ANALYZE it, so the query planner makes the same choices
it will make in production.
It exists because a plan over ten rows is a lie, and because the loop it replaces is not merely smaller: uniform fan-out makes the planner always right, and generating children parent-by-parent clusters them perfectly, which flatters every index scan. A test database can be wrong in the flattering direction, and usually is.
Install
pip install 'django-data-shape[postgres]'
The quotes are not decoration: zsh globs the brackets and reports
no matches found without them.
Use
import datetime
from django_data_shape import Sequential, Shape, Skew, Table, Uniform, build
from myapp.models import Order
shape = Shape(
Table(
Order,
rows=1_000_000,
status=Skew({"complete": 0.98, "pending": 0.015, "cancelled": 0.005}),
total=Uniform(0, 500, places=2),
created_at=Sequential(
datetime.datetime(2020, 1, 1, tzinfo=datetime.timezone.utc),
datetime.timedelta(seconds=3),
),
),
seed=1234,
)
build(shape)
build() generates the rows, loads them with COPY, moves the identity sequence
past the keys it assigned, and runs ANALYZE so the planner can see the shape.
It raises on any backend that is not PostgreSQL rather than degrading quietly --
unless you say require_statistics=False, which asks for rows and cardinality
instead of a database the planner can reason about, and is what the growth
harness below is built on.
Relations
from django_data_shape import Constant, FanOut, Shape, Table, Zipf, build
build(
Shape(
Table(Company, rows=50, name=Constant("acme")),
Table(
Order,
rows=2_000_000,
# A distribution, not a number: giving every parent ten children is
# the one shape in which the planner is never wrong, because its
# n_distinct average is then the truth.
company=FanOut(Zipf(1.2), childless=0.35),
status=Constant("complete"),
),
)
)
The parents can be rows this package built or rows your own code did -- their real keys are read, not assumed, so the ORM can own the small tables while this owns the large ones.
Derivations
A distribution says what a column looks like across rows. A derivation says what one column is given the others -- and the four faces are one mechanism, differing only in where the inputs are read from.
from django_data_shape import After, Aligned, Derived, FanOut, Given, Table, Zipf
Table(
Ticket,
rows=2_000_000,
account=FanOut(Zipf(1.2)),
# the parent row: this ticket's own account, read across the fan-out
opened_at=After("account.signed_up_at", within=timedelta(days=365)),
severity=Given("account.plan", {"free": mostly_low, "enterprise": mostly_high}),
# a shared rank: the big tickets are big in both columns at once
quantity=Aligned("size", Uniform(1, 100, places=0)),
unit_price=Aligned("size", Uniform(1, 500, places=2)),
# this row
total=Derived("quantity", "unit_price", compute=operator.mul),
)
Derived is the mechanism and takes scope= directly, so a correlation nobody
shipped a face for is still declarable. Column order is the COPY column list
and says nothing about dependencies, so derivations get a computation order of
their own and a cycle among them is refused by name.
This package may call your code, and your code may not call the database. Generation runs under a wrapper on the connection being built, so a query raises rather than quietly costing a round trip per row. A per-row creation hook is the thing this package replaces, and a hook whose body may query is a hook whose body will.
Projections
Some tables are copied rather than distributed. An Event is created from a
Template, and its EventSession rows mirror that template's TemplateSession
rows -- so the child count is determined, and correlated with the template.
Shape(
Table(Template, rows=500, name=Constant("t")),
Table(TemplateSession, rows=4_000, template=FanOut(Zipf()), title=Constant("s")),
Table(Event, rows=200_000, template=FanOut(Zipf()), name=Constant("e")),
Projection(EventSession, per=Event, copying=TemplateSession),
)
One INSERT ... SELECT, derived from the model graph, with raw SQL as the escape
hatch. There is no rows=: the count comes from the join, and comes back in the
BuildResult. It is what a creation service collapses into at scale -- one event
from a template is a service call, a million is one statement -- and it
reproduces a correlation a FanOut on the child would destroy.
Invariants
A company has many projects, at most one of which may be ACTIVE. With 50,000
companies and 2,000,000 projects that is exactly 50,000 active rows -- 2.5%,
derived from the fan-out rather than chosen. Declare Skew({"ACTIVE": 0.1, ...})
beside the fan-out and you have asked for 200,000 of them in a schema that
permits 50,000.
Shape(
Table(Company, rows=50_000, name=Constant("acme")),
Table(
Project,
rows=2_000_000,
company=FanOut(Zipf(1.2)),
created_at=Sequential(start, timedelta(minutes=1)),
status=PerParent("company", last="ACTIVE", rest="COMPLETE"),
),
invariants=[
Invariant("no company has two active projects", sql=ONE_ACTIVE_PER_COMPANY),
],
)
Three nets. PerParent generates it right -- the last row of each group is
the special one, computed from the fan-out partition in O(1) so rows still stream
into COPY interleaved. Declared invariants check it as SQL after the
load, failing and rolling back the build, which is the only net that covers rules
the database does not state. And the schema itself refuses, because the rows go
into the real migrated tables.
Because that last message is a unique index failing at row 700,000, the
arithmetic is also done statically off Model._meta.constraints when the shape
is declared:
one_active_project_per_company permits at most 50000 rows with status='ACTIVE',
one per (company); Project.status is filled by Skew({'ACTIVE': 0.1, ...}), which
asks for 200000 of them.
A constraint must be satisfiable by construction within one group, or it is declared as an invariant and checked, not generated. Scheduling is NP-hard and is refused.
Statistics, and building once
ANALYZE runs at the end of every build, because rows the planner cannot see are
worse than no rows. How much of a column it records is decided by that column's
statistics target, so a declaration wider than the target is one PostgreSQL
would build and then not see -- and this refuses rather than producing it:
Shape(
Table(
Event,
rows=2_000_000,
kind=Skew(weights), # 150 event types
statistics={"kind": 300}, # the planner keeps 100 unless asked
)
)
The target is declared, never inferred: it is a property of the column rather than of the distribution, and a package choosing one for you would be deciding how the planner sees your data on evidence your declaration does not contain. What the distributions are read for is the refusal.
Building that database costs about seventeen seconds. Copying it costs about two hundred milliseconds, statistics included:
from django_data_shape import clone_database, template_database
template = template_database(shape) # builds the first time, finds it after
clone_database(template, "test_myapp", replace=True)
The template is named after a content hash of the declaration, the schema, the
relevant settings and this package's version, so a stale one is never asked for.
A shape holding a Derived or a KeyFunction is refused rather than hashed --
there is no honest digest of a callable, and every way of guessing one agrees
while the data has changed. See
Statistics and reuse.
From pytest
# conftest.py
from django_data_shape import Constant, Shape, Table
from django_data_shape.fixtures import scale_fixture, shape_fixture
orders = shape_fixture(Shape(Table(Order, rows=100_000, status=Constant("complete"))))
world = scale_fixture(Shape(Table(Order, rows=100, status=Constant("complete"))))
orders is one world built once for the whole session, composed with
pytest-django rather than replacing it. world is the scale protocol: make
the world be at factor F, then let the caller run its block, which is what a
query count asserted to be O(1) rather than O(N) needs.
def test_the_dashboard_does_not_grow(world, django_assert_num_queries):
for factor in (1, 10):
with world(factor):
with django_assert_num_queries(3):
dashboard()
A factor varies the declaration rather than subsetting one larger build, and
pip install 'django-data-shape[pytest]' is what these two need. The growth
harness works on any backend Django supports, because a query count is an ORM
property and means the same everywhere; the session world skips with a stated
reason where a shaped database cannot exist, because a plan over it is the
thing it exists to make honest.
What it expects, and what it refuses
A declaration that cannot describe a database raises before a row is generated, naming the field. In particular:
- PostgreSQL and psycopg 3. Rows stream into
COPY FROM STDIN, which psycopg 2 cannot do without materialising them first. Both are refused by name rather than degraded around. PostgreSQL is required for the statistics half only:build(shape, require_statistics=False)loads rows on any backend and claims nothing about a plan. psycopg 2 is refused either way, because the vendor picks the route and not the caller. - A key type it can assign. Integer keys count from one and UUID keys are
derived from the seed; anything else is refused rather than guessed, and
keys=KeyFunction(...)declares one. - Empty tables. Keys start at 1 on every build, so
build()checks first and raises rather than colliding partway through. - A callable model default such as
default=uuid4must be declared as a distribution:uuid4varies per row anddictdoes not, and nothing on the field distinguishes them. - A derivation that queries the database. The generation pass is guarded, so the rule is a refusal rather than advice. Read what you need before the build and close over it.
Status
Early. Single tables, the model graph, the pytest surface, the derivation mechanism, projections, statistics targets, template-database reuse and business invariants. Many-to-many edges come next.
Full documentation: https://artui.github.io/django-data-shape/
License
MIT
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file django_data_shape-0.10.0.tar.gz.
File metadata
- Download URL: django_data_shape-0.10.0.tar.gz
- Upload date:
- Size: 327.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
b87476e01a009a2784222b01461568e7bcc029f5821ba05a7b1514af0745a45e
|
|
| MD5 |
50afdcd3567e8559b2e031c45762f038
|
|
| BLAKE2b-256 |
3cfd9740a4430bd3c43c5ef0d4f6f4b73038a389fbac99bbd416302f11f0328f
|
Provenance
The following attestation bundles were made for django_data_shape-0.10.0.tar.gz:
Publisher:
release.yml on Artui/django-data-shape
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
django_data_shape-0.10.0.tar.gz -
Subject digest:
b87476e01a009a2784222b01461568e7bcc029f5821ba05a7b1514af0745a45e - Sigstore transparency entry: 2697328388
- Sigstore integration time:
-
Permalink:
Artui/django-data-shape@6b525dd3c70db8b55690a74e1a013abbe49beb15 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Artui
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@6b525dd3c70db8b55690a74e1a013abbe49beb15 -
Trigger Event:
push
-
Statement type:
File details
Details for the file django_data_shape-0.10.0-py3-none-any.whl.
File metadata
- Download URL: django_data_shape-0.10.0-py3-none-any.whl
- Upload date:
- Size: 140.1 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via:
twine/7.0.0 CPython/3.13.14
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
9de3e0a6f12c61e40f201af1abc14818fb0f7570a542d59c927686886afe9a88
|
|
| MD5 |
0412a6824c098016deec73e130a436b4
|
|
| BLAKE2b-256 |
5cab4f80b86c144db3558c90f8139fc8ae7ce61d2658dec7675b277930d0740c
|
Provenance
The following attestation bundles were made for django_data_shape-0.10.0-py3-none-any.whl:
Publisher:
release.yml on Artui/django-data-shape
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
django_data_shape-0.10.0-py3-none-any.whl -
Subject digest:
9de3e0a6f12c61e40f201af1abc14818fb0f7570a542d59c927686886afe9a88 - Sigstore transparency entry: 2697328447
- Sigstore integration time:
-
Permalink:
Artui/django-data-shape@6b525dd3c70db8b55690a74e1a013abbe49beb15 -
Branch / Tag:
refs/heads/main - Owner: https://github.com/Artui
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
release.yml@6b525dd3c70db8b55690a74e1a013abbe49beb15 -
Trigger Event:
push
-
Statement type: