Skip to main content

mitol-django-benchmark

Prove a Django or DRF performance change actually made an endpoint faster.

A performance PR that says "this should be faster" is a guess. This package turns it into a measurement: seed a production-shaped throwaway database, run the same request on two git refs against identical rows, and attribute the difference query by query.

Two failure modes make a benchmark worse than none, and most of this package exists to avoid them:

  1. Measuring the harness. Instrumentation that scales with the thing you changed will manufacture a result. An N+1 profiler hooks every ORM fetch, so it costs more in whichever arm hydrates more objects — which is the arm you are trying to show is slower. The harness refuses to run with one active.
  2. Measuring the wrong shape. Factory defaults are nothing like production: blank rich-text columns, one related row where production has dozens. A seed that is off structurally produces a number with no bearing on the endpoint you care about.

Related: mitol-drf-lint and the drf-api-performance skill are about writing fast endpoints. This is about proving one got faster.

Installation

uv add --dev "mitol-django-benchmark[drf,factories,django,postgres]"

click comes with it rather than as an extra: the runner drives every step by invoking python -m mitol.benchmark step <name> inside the application environment, so the CLI has to work wherever the package is installed.

Extra Gives you
drf DRF's APIClient, so authentication does not cost a login round-trip per call
factories factory-boy, for seed steps of kind factory
django OpenTelemetry Django instrumentation — request spans
postgres OpenTelemetry psycopg instrumentation — query spans, which is where attribution comes from
postgres2 the psycopg2 equivalent, for a project still on it

Registering the app in INSTALLED_APPS is optional. The ol-benchmark console script starts Django itself; add mitol.benchmark.apps.BenchmarkApp only if you want the manage.py ol_benchmark fallback for a project whose bootstrap the console script cannot reproduce.

The three configuration layers

Three files, because they have three different owners and lifetimes. Nothing is repeated between them, and a file that declares a section belonging to another layer is rejected rather than quietly ignored.

File Committed? Owner Holds
benchmarks/benchmark.toml yes the project [django], [database], [defaults.measure]
benchmarks/<name>.toml yes the benchmark author [benchmark], [knobs], [seed], [target], [auth], [trace], [calibration]
benchmarks/benchmark.local.toml no — gitignore it the developer [backend], and any connection strings their machine needs

The project and local layers are discovered by searching at and above the benchmark file's directory. The local layer is genuinely optional: the defaults are the local backend against localhost Postgres, so a fresh clone runs without one.

Precedence, lowest to highest: project → benchmark → local → environment → CLI flags.

Every string in [django], [database] and [backend] is expanded with ${VAR} and ${VAR:-default}. That is what lets a committed file work on every machine — and what a k3d/Tilt cluster needs, where the DSN is assigned by the cluster rather than known to the repository. An unset variable with no default is an error naming the variable, not an empty string.

Getting started

ol-benchmark init --project                      # benchmarks/benchmark.toml
ol-benchmark init --local                        # benchmarks/benchmark.local.toml
ol-benchmark init --benchmark library-list       # benchmarks/library_list.toml
ol-benchmark validate benchmarks/library_list.toml
ol-benchmark run benchmarks/library_list.toml --base-ref main

init --project also adds this package's artifacts to the project's .gitignore — .bench/, the local configuration layer, and traces/ plus *.trace.json. The exports are the reason it happens at adoption time rather than being left as advice: a raw OTel trace carries statement literals, query strings and user identifiers, and the first request for one comes a few steps later. The entries are appended behind a marker comment, so re-running adds nothing.

ol-benchmark -h lists the subcommands; each takes -h of its own. Exit codes are part of the contract, because the backends run these as subprocesses: 2 is a configuration problem, 1 is everything else that is known to go wrong — including a void comparison.

Writing the seed

Every shape dimension is a knob, so the same benchmark can be re-run at another shape without editing anything, and the report can state the exact shape the number was measured at:

[knobs]
rows = 25
nested_per_row = 10
blob_bytes = 2048

Override at run time with --knob rows=100 or BENCH_ROWS=100.

Steps run in order and can refer to each other:

[[seed.step]]
name = "authors"
factory = "libraries.factories:AuthorFactory"
count = 50

[[seed.step]]
name = "books"
factory = "libraries.factories:BookFactory"
count = "$knob:rows"
kwargs = { author = "$cycle:authors", title = "Book {index}" }
m2m = { topics = { source = "topics", per = "$knob:nested_per_row", strategy = "cycle" } }
Step kind Needs Does
factory factory factory-boy, one create() per row (or one bulk_create with bulk = true)
model model plain model instances, no factory-boy needed
fixture fixtures loaddata of committed JSON
sql statements raw SQL
hook hook calls pkg.mod:fn(context) — the escape hatch for fan-out a declarative format cannot express

Tokens, resolved anywhere in kwargs, m2m, count, target params and auth:

Token Means
$knob:NAME a shape knob
$index the 0-based index of the row being built
$blob:N | $blob:KNOB filler text of roughly N bytes
$ref:STEP, $ref:STEP[2] one object from an earlier step
$cycle:STEP that step's objects, cycled by $index
$sample:STEP:5 5 of them, from the seeded RNG (deterministic)
$ids:KEY a scalar from [seed.export] — how the request reaches seeded rows
$$ a literal $

Plain strings are also run through str.format with index and every knob, so "Book {index}" works without a token.

Structure matters more than sizing. How rows fan out across joins dominates how wide they are. Putting everything under one tenant is the classic error: it makes the filter match everything and turns an unchanged query pathological. The seed result reports many-to-many pair totals for exactly this reason — join multiplicity is invisible in an API response.

Backends

kind Runs steps with Use when
local subprocesses of the harness the application runs in the environment you are typing in
compose docker compose run --rm --no-deps -T <service> the repository has its own compose stack, already up
k8s kubectl exec a local-dev cluster (k3d + Tilt or similar)

The k8s backend requires context and never inherits your current kubectl context: a developer's current context is routinely a deployed environment, and this harness runs DROP DATABASE. It also refuses a context whose name contains ci, qa, prod, applications, residential, data or operations, because being made to name the cluster does not help if the name you type is applications-qa. Override the markers with [backend].context_denylist in your own local config if one of them is a false positive on a cluster you really do benchmark against. It also waits for a push-based sync to finish after a ref switch, and the runner separately verifies the switched-to files actually arrived before measuring.

Comparing against production

A/B-ing two refs tells you whether a change helped. It does not tell you whether your seed is shaped like reality — and a delta measured on the wrong shape is a number about your seed. For that you need production.

Exported OTel traces are never committed. They carry statement literals, query strings, user and tenant identifiers and absolute timestamps. What gets committed is the aggregate:

ol-benchmark baseline benchmarks/library_list.toml traces/*.json

Pass as many exported traces as you have — the medians improve with the sample, exactly as trace_repeats does locally. Files may mix the two OTLP export layouts, a trace split across files is one request, and a trace present in two files is counted once.

That writes benchmarks/library_list.baseline.json: per-query medians keyed by the labels you wrote in [[trace.classify]], plus row-count floors and the sampled trace IDs. Nothing else. The safety property is structural rather than a scrubbing pass —

the only free-text field in the baseline is drawn from a closed set that is already committed in the same repository.

A production query matching none of your classifiers is labelled "unclassified" rather than by its SQL, so there is no code path — and no flag — that can put statement text into the file. The command warns when that happens so you can add a rule. Use --stdout to look before anything is written, and read the file before you commit it.

The trace IDs are the one deliberate exception. A trace ID is a pointer, not data: it resolves only for someone who already has access to your tracing backend. It is what lets a reviewer open the exact request a number came from. It does record that a request existed, so drop the field if your team is not comfortable publishing that.

Two things to get right when exporting:

  • Include the server/root span. A trace filtered down to database spans makes the final query's gap read as zero, and that gap is where serialization lives. The command warns if it spots this.
  • One endpoint per baseline. A glob that sweeps in unrelated traces averages their queries into nonsense; the command warns when the traces cover more than one route.

Then reference it, and say which query the change is meant to move:

[[trace.classify]]
label = "topics prefetch"
pattern = "book_topics"
targeted = true            # excluded from the drift check; it should differ

[calibration]
baseline = "library_list.baseline.json"
drift_factor = 5

The attribution table gains a production column, and any untargeted query more than drift_factor away from production is reported under "Seed drift from production". That is the falsification check: a seed parameter that makes an unchanged query wildly slower than production is wrong, however good the story behind it was. Drift is a warning about the seed and never changes the verdict — the verdict is about whether the two arms are comparable to each other.

Local faster than production is expected and says so. Local slower on a query the change does not touch is the real alarm.

What you get

Written to .bench/out/<benchmark>/ by default:

File Contents
comparison.json the primary artifact — verdict, deltas, per-query attribution, calibration
report.md the same, for a human
config.resolved.json every setting in force and which layer it came from, credentials redacted
seed.json per-step row counts, m2m pair totals, exports, floor warnings
base.json, branch.json the wall-clock results, with the conditions they were measured under
trace-base.json, trace-branch.json every span: duration and the gap to the next one
agg-base.json, agg-branch.json median per logical query across the traced repeats

benchmarks/<name>.baseline.json is the one output that is committed — it is production evidence distilled to numbers, and a colleague needs it to re-run the drift check from a fresh clone.

Verdicts

comparison.json never reports a bare number. It reports one of:

  • void — the arms returned different responses. They did not do the same work, so no delta may be quoted. Find out why.
  • inconclusive — the difference is inside the run-to-run spread, or min and median disagree about which arm is faster. This is a real answer.
  • ok — the difference exceeds the spread and both statistics agree.

Reading the trace

A database span wraps cursor.execute and nothing else, so row fetch, model instantiation and serialization all live in the gaps between spans. That is why every span carries gap_ms and the aggregation reports sql and gap separately: a trace that shows only span durations will tell you the queries are fast while the request is slow.

The traced numbers are for attribution, never for the headline. Instrumentation is not free, and it is normally inert in an application with no OTLP endpoint — which is what keeps the wall-clock pass clean. Set [trace].otlp_endpoint to additionally ship the spans to a real collector.

What the number means

Local Postgres has no network round-trip, a warm cache and no contention. Production has all three, and they penalise larger result sets disproportionately. What you measure here is a floor, not an estimate. Every report says so; do not delete it when you quote the number.

Also state, every time:

  • the seed shape, next to the numbers;
  • which knobs were guesses, and that they are identical in both arms, so they move the baseline rather than the delta;
  • any production signal the harness failed to reproduce.

Pitfalls the harness cannot see

Pitfall Why it ruins the result
Tuning the seed against the query you changed Circular. Calibrate only on observables the change does not touch
Factory defaults as "realistic" Blank rich-text fields, one related row where production has dozens
One traced request Per-query gaps are far too noisy; the default is a median of seven
Quoting a traced total as the result The trace is for attribution; the uninstrumented pass is for the number
Extrapolating local ms to production ms Different hardware, cache state and network

See AGENTS.md in this package for the full calibration procedure.

Metadata

Release files for mitol-django-benchmark 2026.9.30

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for mitol-django-benchmark 2026.9.30
File Size Uploaded
mitol_django_benchmark-2026.9.30.tar.gz 79.8 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for mitol-django-benchmark 2026.9.30
File Interpreter ABI Platform
mitol_django_benchmark-2026.9.30-py3-none-any.whl Python 3 none any Details

Total release size: 183.3 kB

Release files / mitol_django_benchmark-2026.9.30.tar.gz

Download URL mitol_django_benchmark-2026.9.30.tar.gz
Size 79.8 kB
Tags Source
SHA-256 checksum
How to use checksums
360d6ee12ab6edb9356f5c4668bca92a2809ec4b4f43ba51e4e3bfc5acb489e2
BLAKE2b-256 checksum
How to use checksums
07cf2b766430a25bcaa96e0f56e9c7446ede4ae8ed19d18829fa33756caeef8b
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / mitol_django_benchmark-2026.9.30-py3-none-any.whl

Download URL mitol_django_benchmark-2026.9.30-py3-none-any.whl
Size 103.5 kB
Tags Python 3
SHA-256 checksum
How to use checksums
946e24adee19ace02366fc4ece03eeb33a186fe916ea94c3e9ad45714640d237
BLAKE2b-256 checksum
How to use checksums
635ecc8c4591050f6ffb4919c3a3086bea29b14c7011f61e7ff03124a915a8fc
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

2026.9.30 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page