Skip to main content

QBMP — Quoins, Bricks, Mortar & Pointing

  • Declarative synthetic dataset generation through mathematical modelling.
  • Declare a model — inputs with ranges or categories, outputs with rule functions, and the weights that bind them — and QBMP fills every output's declared range with no gaps, then publishes a datasheet certifying exactly how well it did.
  • Check out the example code in repo ( https://github.com/Palani-SN/QBMP ) for reference.

Installation

  python -m pip install QBMP

Usage

  • A model is a class that subclasses QBMP and declares two things: Inputs (what varies) and Outputs (what gets computed), wired together by a weights dict per output. Each output also names a rule method, decorated with @rule(...), that turns a row of input values into that output's value.
  • Sample code as shown below declares a two-output slice of a real-estate pricing model (refer Demo.py under EXAMPLES/ for the full seven-output version).
from typing import ClassVar
from QBMP.engine import QBMP, rule


class Real_Estate_Pricing_Model(QBMP):
    model: ClassVar[dict] = {
        "Outputs": {
            "price_lakhs": {
                # weights order the inputs quoins -> bricks -> mortar
                "weights": {
                    "area_sqft": 100,
                    "property_type": 55,
                    "distance_km": 43,
                    "age_years": 20,
                    "floor": 12,
                    "bedrooms": 4,
                },
                "range": (7, 418),
                "engine": "populate_price_lakhs",
            },
            "price_band": {
                # a COMBINATIONAL output - categories, not a range
                "weights": {"area_sqft": 100, "property_type": 55},
                "categories": ["Budget", "Mid", "Premium", "Luxury"],
                "engine": "populate_price_band",
            },
        },
        "Inputs": {
            "area_sqft": {"range": (500, 3500), "default": 1500},
            "property_type": {
                "categories": ["Studio", "Apartment", "Rowhouse", "Villa"],
                "default": "Apartment",
            },
            "bedrooms": {"range": (1, 5), "default": 2},
            "age_years": {"range": (0, 50), "default": 10},
            "distance_km": {"range": (1, 30), "default": 10},
            "floor": {"range": (0, 20), "default": 3},
        },
    }

    @rule("price_lakhs")
    def populate_price_lakhs(
        self, area_sqft, bedrooms, age_years, distance_km, floor, property_type
    ):
        rate = 8000 - 150 * distance_km - 40 * age_years + 60 * floor + 100 * bedrooms
        return round(rate * area_sqft / 1e5, 2)

    @rule("price_band")
    def populate_price_band(
        self, area_sqft, bedrooms, age_years, distance_km, floor, property_type
    ):
        price = self.populate_price_lakhs(
            area_sqft=area_sqft,
            bedrooms=bedrooms,
            age_years=age_years,
            distance_km=distance_km,
            floor=floor,
            property_type=property_type,
        )
        return (
            "Budget"
            if price < 80
            else "Mid"
            if price < 180
            else "Premium"
            if price < 300
            else "Luxury"
        )


if __name__ == "__main__":
    data_set = Real_Estate_Pricing_Model(seed=24)
    data_set.save(
        min_rows=1000,
        dataset="prices",
        format="csv",
        max_bins=4096,
        outputs=["price_lakhs", "price_band"],
    )
  • save() is the single entrypoint: it generates rows, qualifies them, checks coherence, and publishes a folder named for the dataset.
prices/
    index.html      the datasheet - self-contained, drag-to-zoom
    prices.csv      the dataset
  • Console output on a run looks like this — a preview of the rows, the coverage report per output, and where the folder landed:

  • And the published datasheet (index.html) looks like this:

  • outputs takes a name, a list, or None for every output declared in the model. format is "csv" or "xlsx" (the latter needs openpyxl). min_rows is a floor, not an exact count, and max_bins is the resolution mortar aims for — both explained below.

Design

In masonry, quoins, bricks and mortar are laid in that order. Quoins are the large dressed cornerstones set first — few in number, widely spaced, fixing the geometry of the whole wall. Bricks fill the field between them. Mortar closes whatever gap is left, at the finest grain of all. Pointing then goes back over the joints and finesses them.

That is exactly what filling an output's range needs, and the courses map onto three passes plus a fourth on the roadmap:

graph LR
    L1["<b>Quoins</b><br/>landmarks at equidistant points<br/>across the declared range"]
    L2["<b>Bricks</b><br/>widen each landmark along<br/>the next input, tiling the gap"]
    L3["<b>Mortar</b><br/>find the bins still empty and<br/>solve a row into each one"]
    L4["<b>Pointing</b> <i>(roadmap)</i><br/>purge over-full bins,<br/>populate the thin ones"]

    L1 --> L2 --> L3 -.-> L4

    style L1 fill:#3b6ea5,color:#fff,stroke:none
    style L2 fill:#6f96bd,color:#fff,stroke:none
    style L3 fill:#3f7d58,color:#fff,stroke:none
    style L4 fill:none,stroke:#999,stroke-dasharray:4 3,color:#888

An input's weight decides which course it belongs to: heavy inputs (most leverage on the output) place the landmarks, light inputs perturb the value just enough to fill gaps. Quoins and bricks are blind — they subdivide on a schedule without checking where the gaps are. Mortar is targeted — it bins the output, finds the empty bins, and solves a row into each one specifically.

Every input and output is exactly one of two kinds:

Continuational Combinational
declared with "range": (min, max) "categories": [...]
values are swept and solved enumerated
coverage means the span is spanned, no gaps every declared option appears

A combinational output can never drive the sampling hierarchy — there is no span to place landmarks across — so it just rides the rows the continuational outputs produce.

Principles

  • Coverage is chosen over uniformity. These genuinely compete; QBMP optimises for reaching every corner of the declared range rather than for a flat histogram. Rebalancing toward uniformity is future work ("pointing").
  • The coherence invariant. A row is only ever produced by choosing inputs and running the rule engines — output values are never written, interpolated, or carried across passes. Every save() re-derives each output from its own row and reports the worst disagreement (0.0 when everything checks out).
  • Declared ranges are not silently corrected. If a model's declared range is wider than it can actually produce, the unreachable bins are reported and greyed on the datasheet rather than hidden — a declaration exceeding reality is worth seeing.
  • min_rows is a floor, not a target. Row totals are products of per-pass counts, so the sampler lands on the closest reachable count at or above what was asked, not on the exact number.
  • The datasheet is self-contained. Inline CSS, SVG and script, no network requests — it opens from a file:// path on any machine, with drag-to-zoom on every ladder.

API at a glance

@rule(output_name) decorator binding a method as an output's rule engine; completes its kwargs from declared defaults
QBMP(seed) validates the model wiring, binds every engine, raises early on a bad wire-up
save(min_rows, dataset, format, outputs, max_bins, ...) the single entrypoint: generate, qualify, check coherence, publish a folder — returns the DataFrame, with self.report / self.mix / self.drift left on the instance

Metadata

Release files for QBMP 0.0.1

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for QBMP 0.0.1
File Size Uploaded
qbmp-0.0.1.tar.gz 41.8 kB Details

Release files / qbmp-0.0.1.tar.gz

Download URL qbmp-0.0.1.tar.gz
Size 41.8 kB
Tags Source
SHA-256 checksum
How to use checksums
a1b0d0c8dbcf4049848cead84daf6a914a6296b52187417aa72c34ff3e009e17
BLAKE2b-256 checksum
How to use checksums
263381b6c3895de0b41615bb2f680e1c8ba41cbca975d903dd45c556fa727240
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
No
Uploaded via twine/3.4.2 importlib_metadata/9.0.1 pkginfo/1.13 requests/2.34.2 requests-toolbelt/1.0.0 tqdm/4.70.0 CPython/3.14.7

Release history Release notifications | RSS feed

0.0.3

1 release file

0.0.2

1 release file

This release

0.0.1 This release

1 release file

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page