Skip to main content

django-data-shape

CI PyPI Python versions Django versions Docs Coverage Ruff License

A realistically shaped test database from Django models.

Declare the shape of your data -- cardinality, value skew, foreign-key fan-out as a distribution with a long tail, and where related rows physically sit -- then load it by COPY and ANALYZE it, so the query planner makes the same choices it will make in production.

It exists because a plan over ten rows is a lie, and because the loop it replaces is not merely smaller: uniform fan-out makes the planner always right, and generating children parent-by-parent clusters them perfectly, which flatters every index scan. A test database can be wrong in the flattering direction, and usually is.

Install

pip install 'django-data-shape[postgres]'

The quotes are not decoration: zsh globs the brackets and reports no matches found without them.

Use

import datetime

from django_data_shape import Sequential, Shape, Skew, Table, Uniform, build

from myapp.models import Order

shape = Shape(
    Table(
        Order,
        rows=1_000_000,
        status=Skew({"complete": 0.98, "pending": 0.015, "cancelled": 0.005}),
        total=Uniform(0, 500, places=2),
        created_at=Sequential(
            datetime.datetime(2020, 1, 1, tzinfo=datetime.timezone.utc),
            datetime.timedelta(seconds=3),
        ),
    ),
    seed=1234,
)

build(shape)

build() generates the rows, loads them with COPY, moves the identity sequence past the keys it assigned, and runs ANALYZE so the planner can see the shape. It raises on any backend that is not PostgreSQL rather than degrading quietly -- unless you say require_statistics=False, which asks for rows and cardinality instead of a database the planner can reason about, and is what the growth harness below is built on.

Relations

from django_data_shape import Constant, FanOut, Shape, Table, Zipf, build

build(
    Shape(
        Table(Company, rows=50, name=Constant("acme")),
        Table(
            Order,
            rows=2_000_000,
            # A distribution, not a number: giving every parent ten children is
            # the one shape in which the planner is never wrong, because its
            # n_distinct average is then the truth.
            company=FanOut(Zipf(1.2), childless=0.35),
            status=Constant("complete"),
        ),
    )
)

The parents can be rows this package built or rows your own code did -- their real keys are read, not assumed, so the ORM can own the small tables while this owns the large ones.

Derivations

A distribution says what a column looks like across rows. A derivation says what one column is given the others -- and the four faces are one mechanism, differing only in where the inputs are read from.

from django_data_shape import After, Aligned, Derived, FanOut, Given, Table, Zipf

Table(
    Ticket,
    rows=2_000_000,
    account=FanOut(Zipf(1.2)),
    # the parent row: this ticket's own account, read across the fan-out
    opened_at=After("account.signed_up_at", within=timedelta(days=365)),
    severity=Given("account.plan", {"free": mostly_low, "enterprise": mostly_high}),
    # a shared rank: the big tickets are big in both columns at once
    quantity=Aligned("size", Uniform(1, 100, places=0)),
    unit_price=Aligned("size", Uniform(1, 500, places=2)),
    # this row
    total=Derived("quantity", "unit_price", compute=operator.mul),
)

Derived is the mechanism and takes scope= directly, so a correlation nobody shipped a face for is still declarable. Column order is the COPY column list and says nothing about dependencies, so derivations get a computation order of their own and a cycle among them is refused by name.

This package may call your code, and your code may not call the database. Generation runs under a wrapper on the connection being built, so a query raises rather than quietly costing a round trip per row. A per-row creation hook is the thing this package replaces, and a hook whose body may query is a hook whose body will.

Projections

Some tables are copied rather than distributed. An Event is created from a Template, and its EventSession rows mirror that template's TemplateSession rows -- so the child count is determined, and correlated with the template.

Shape(
    Table(Template, rows=500, name=Constant("t")),
    Table(TemplateSession, rows=4_000, template=FanOut(Zipf()), title=Constant("s")),
    Table(Event, rows=200_000, template=FanOut(Zipf()), name=Constant("e")),
    Projection(EventSession, per=Event, copying=TemplateSession),
)

One INSERT ... SELECT, derived from the model graph, with raw SQL as the escape hatch. There is no rows=: the count comes from the join, and comes back in the BuildResult. It is what a creation service collapses into at scale -- one event from a template is a service call, a million is one statement -- and it reproduces a correlation a FanOut on the child would destroy.

Invariants

A company has many projects, at most one of which may be ACTIVE. With 50,000 companies and 2,000,000 projects that is exactly 50,000 active rows -- 2.5%, derived from the fan-out rather than chosen. Declare Skew({"ACTIVE": 0.1, ...}) beside the fan-out and you have asked for 200,000 of them in a schema that permits 50,000.

Shape(
    Table(Company, rows=50_000, name=Constant("acme")),
    Table(
        Project,
        rows=2_000_000,
        company=FanOut(Zipf(1.2)),
        created_at=Sequential(start, timedelta(minutes=1)),
        status=PerParent("company", last="ACTIVE", rest="COMPLETE"),
    ),
    invariants=[
        Invariant("no company has two active projects", sql=ONE_ACTIVE_PER_COMPANY),
    ],
)

Three nets. PerParent generates it right -- the last row of each group is the special one, computed from the fan-out partition in O(1) so rows still stream into COPY interleaved. Declared invariants check it as SQL after the load, failing and rolling back the build, which is the only net that covers rules the database does not state. And the schema itself refuses, because the rows go into the real migrated tables.

Because that last message is a unique index failing at row 700,000, the arithmetic is also done statically off Model._meta.constraints when the shape is declared:

one_active_project_per_company permits at most 50000 rows with status='ACTIVE',
one per (company); Project.status is filled by Skew({'ACTIVE': 0.1, ...}), which
asks for 200000 of them.

A constraint must be satisfiable by construction within one group, or it is declared as an invariant and checked, not generated. Scheduling is NP-hard and is refused.

Statistics, and building once

ANALYZE runs at the end of every build, because rows the planner cannot see are worse than no rows. How much of a column it records is decided by that column's statistics target, so a declaration wider than the target is one PostgreSQL would build and then not see -- and this refuses rather than producing it:

Shape(
    Table(
        Event,
        rows=2_000_000,
        kind=Skew(weights),  # 150 event types
        statistics={"kind": 300},  # the planner keeps 100 unless asked
    )
)

The target is declared, never inferred: it is a property of the column rather than of the distribution, and a package choosing one for you would be deciding how the planner sees your data on evidence your declaration does not contain. What the distributions are read for is the refusal.

Building that database costs about seventeen seconds. Copying it costs about two hundred milliseconds, statistics included:

from django_data_shape import clone_database, template_database

template = template_database(shape)  # builds the first time, finds it after
clone_database(template, "test_myapp", replace=True)

The template is named after a content hash of the declaration, the schema, the relevant settings and this package's version, so a stale one is never asked for. A shape holding a Derived or a KeyFunction is refused rather than hashed -- there is no honest digest of a callable, and every way of guessing one agrees while the data has changed. See Statistics and reuse.

From pytest

# conftest.py
from django_data_shape import Constant, Shape, Table
from django_data_shape.fixtures import scale_fixture, shape_fixture

orders = shape_fixture(Shape(Table(Order, rows=100_000, status=Constant("complete"))))
world = scale_fixture(Shape(Table(Order, rows=100, status=Constant("complete"))))

orders is one world built once for the whole session, composed with pytest-django rather than replacing it. world is the scale protocol: make the world be at factor F, then let the caller run its block, which is what a query count asserted to be O(1) rather than O(N) needs.

def test_the_dashboard_does_not_grow(world, django_assert_num_queries):
    for factor in (1, 10):
        with world(factor):
            with django_assert_num_queries(3):
                dashboard()

A factor varies the declaration rather than subsetting one larger build, and pip install 'django-data-shape[pytest]' is what these two need. The growth harness works on any backend Django supports, because a query count is an ORM property and means the same everywhere; the session world skips with a stated reason where a shaped database cannot exist, because a plan over it is the thing it exists to make honest.

What it expects, and what it refuses

A declaration that cannot describe a database raises before a row is generated, naming the field. In particular:

  • PostgreSQL and psycopg 3. Rows stream into COPY FROM STDIN, which psycopg 2 cannot do without materialising them first. Both are refused by name rather than degraded around. PostgreSQL is required for the statistics half only: build(shape, require_statistics=False) loads rows on any backend and claims nothing about a plan. psycopg 2 is refused either way, because the vendor picks the route and not the caller.
  • A key type it can assign. Integer keys count from one and UUID keys are derived from the seed; anything else is refused rather than guessed, and keys=KeyFunction(...) declares one.
  • Empty tables. Keys start at 1 on every build, so build() checks first and raises rather than colliding partway through.
  • A callable model default such as default=uuid4 must be declared as a distribution: uuid4 varies per row and dict does not, and nothing on the field distinguishes them.
  • A derivation that queries the database. The generation pass is guarded, so the rule is a refusal rather than advice. Read what you need before the build and close over it.
  • One table per declaration. A model using multi-table inheritance puts one logical row in two tables sharing a key, and this package owns each table's keys, so it can write either half and has nothing to pair them with. It is refused by name. Abstract bases and proxies are ordinary single-table models and are fine.
  • A uniqueness it can keep. A multi-column UniqueConstraint over two fan-outs -- the through table of a many-to-many -- fits comfortably and still cannot be built: two partitions of the same rows are computed without either seeing the other, so a collision is a matter of the seed. It is refused at declaration time, pointing at the Projection with your own sql= that builds a deduplicated edge table today.

Status

Early. Single tables, the model graph, the pytest surface, the derivation mechanism, projections, statistics targets, template-database reuse and business invariants. Many-to-many edges come next.

Full documentation: https://artui.github.io/django-data-shape/

License

MIT

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

django_data_shape-0.16.0.tar.gz (410.3 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

django_data_shape-0.16.0-py3-none-any.whl (178.2 kB view details)

Uploaded Python 3

File details

Details for the file django_data_shape-0.16.0.tar.gz.

File metadata

  • Download URL: django_data_shape-0.16.0.tar.gz
  • Upload date:
  • Size: 410.3 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/7.0.0 CPython/3.13.14

File hashes

Hashes for django_data_shape-0.16.0.tar.gz
Algorithm Hash digest
SHA256 6046496093315f44ff36361f25e7b6a032d03c5d3df502f534955e6995beb89d
MD5 a683864da6eeed1f2297edc065e7a1a4
BLAKE2b-256 f7fa5f42eea9b4c38a0f0fc864d173efd25f2209113a766451f46ae0e3648a62

See more details on using hashes here.

Provenance

The following attestation bundles were made for django_data_shape-0.16.0.tar.gz:

Publisher: release.yml on Artui/django-data-shape

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file django_data_shape-0.16.0-py3-none-any.whl.

File metadata

File hashes

Hashes for django_data_shape-0.16.0-py3-none-any.whl
Algorithm Hash digest
SHA256 7fa0510a7e174fa012523411d3fca3386e6e425bd7e81ade282c26bb04139f13
MD5 8700507de12e3949f3f72af4ec270f71
BLAKE2b-256 65de1a341f8cb404221bb70d97dec488a21dc52eb5d5178d21c38a05e71b3daa

See more details on using hashes here.

Provenance

The following attestation bundles were made for django_data_shape-0.16.0-py3-none-any.whl:

Publisher: release.yml on Artui/django-data-shape

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

0.21.0

2 files

0.20.0

2 files

0.19.0

2 files

0.18.1

2 files

0.18.0

2 files

0.17.1

2 files

0.17.0

2 files

This release

0.16.0 This release

2 files

0.15.0

2 files

0.14.0

2 files

0.13.0

2 files

0.12.0

2 files

0.11.0

2 files

0.10.0

2 files

0.9.0

2 files

0.8.0

2 files

0.7.0

2 files

0.6.0

2 files

0.5.0

2 files

0.4.0

2 files

0.3.0

2 files

0.2.0

2 files

0.1.1

2 files

0.1.0

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page