Skip to main content

Aggregation-bias and subgroup-composition sensitivity audits in Python

Project description

updatesupport

PyPI Version

Check whether a reported aggregate is stable to subgroup recomposition.

Reports often publish a number by coarse categories. updatesupport asks:

If those public category counts stayed fixed, could the aggregate still move because the finer subgroup mix inside them changed?

This is a practical audit for aggregation bias, Simpson's paradox risk, ecological-fallacy risk, subgroup composition sensitivity, and coarsened fairness or disparity reports. It quantifies hidden-composition ambiguity: how much an aggregate rate, effect, or risk metric could move under a declared recomposition stress test.

Thirty-second example. In the Folktables ACSIncome demo, a report grouped by age band, education band, and sex observes an income-threshold rate of 12.37%. Holding those public group proportions fixed, retained subgroup composition can move the compatible rate to:

11.79% to 13.44%

That 1.65 percentage-point width is not sampling error. It is the amount of aggregation ambiguity left by the public categories. The best one-column refinement in that example is occupation major group.

Here, "hidden" means retained but not publicly reported, not unobserved or unknowable. The ambiguity interval is always relative to the retained refinement you choose and the admissible shift class Q you declare. It is not an absolute bound on every possible omitted variable or population shift.

It is useful when a dashboard, model-review pack, policy table, or causal analysis reports a number by coarse public categories and you need to know:

Are these categories enough to support this reported estimate?

Why It Exists

Many reports publish estimates by coarse buckets:

age band x education band x sex
product x region x FICO band x LTV band
segment x channel

Those buckets may contain retained subgroups with different target rates or effects. A confidence interval answers sampling uncertainty. A model metric answers model fit. updatesupport answers a different review question:

Holding the reported public mix fixed, how far could the aggregate move under plausible hidden subgroup shifts?

The output is a review-ready Markdown audit with:

  • claim-level pass/fail/inconclusive verdicts
  • observed estimate
  • hidden-composition stress interval
  • transport ambiguity, the interval width
  • estimator-uncertainty-aware interval when hidden-cell standard errors are supplied
  • public adequacy flag
  • public cells driving the ambiguity
  • refinement recommendations
  • sensitivity checks
  • public-representation frontier search for small stable reporting designs

This is not a replacement for causal identification, model validation, or sampling uncertainty. It is an audit layer for the reporting representation.

The project does not claim new mathematics. The core calculation is a partial-identification / sensitivity-analysis problem with roots in ecological inference, Frechet-style bounds, and distributionally robust optimization. The value is the packaged workflow: compile a retained fine table, solve the stress test, explain the ambiguity, rank refinements, run sensitivity checks, and emit review-ready artifacts. See docs/positioning-and-lineage.md.

Why the Math Is Sound

updatesupport reduces the review question to a finite optimization problem. Given retained fine cells D, public buckets pi(D), supplied retained-cell target values h, and an explicit admissible distribution class Q, it computes:

inf/sup  sum_d h(d) q(d)
subject to q in Q and fixed public law pi#q

For saturated reweighting this has a closed-form public-fiber range formula. For linear Q it is solved as a linear program. For convex transport budgets such as TV, chi-square, KL, and Wasserstein, it is solved as a convex CVXPY problem with a linear objective.

The interval is statistically meaningful as a partial-identification or sensitivity interval conditional on the retained support, target values, and chosen Q. It is not a confidence interval. If retained-cell target standard errors are supplied, reports can add an estimator-uncertainty-aware outer interval, but causal, survey-design, and broader model uncertainty still belong to the upstream statistical workflow. With CVXPY-compatible Q sets, reports can also include an SOCP confidence-core diagnostic showing whether all admissible composition-specific confidence bands have a common overlap.

The core solver target is fixed after compilation. Most tabular reports compile to the linear plug-in aggregate sum_d h(d) q(d); RatioTarget covers supported fixed ratio targets; MomentTransformTarget covers affine moment transforms, one-sided convex/concave CVXPY endpoints, and conservative monotone moment bounds; and ProcedureTarget handles representation-dependent reporting procedures by recompiling the target for each representation before solving. Other nonlinear targets that depend directly on q still need an explicit reformulation or a dedicated target-functional backend.

See docs/mathematical-statistical-soundness.md for the full assumptions, formulas, backend checks, and limitations.

Quickstart

pip install updatesupport

or, in a uv project:

uv add updatesupport
import updatesupport as us

audit = us.claim(
    "Income-threshold rate is stable enough to report",
    public=["AGE_BAND", "EDU_BAND", "SEX"],
    hidden=["AGE_BAND", "EDU_BAND", "SEX", "OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    target="income_over_threshold",
    weight="sample_weight",
    candidate_refinements=["OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    ambiguity_limit=0.015,
    min_cell_weight=25,
    q_presets=["saturated"],
).audit(rows_or_frame)

print(audit.to_markdown())

The claim audit asks:

If we only report by AGE_BAND x EDU_BAND x SEX, how much could the aggregate income-threshold rate move if occupation, work hours, and race composition changed inside those public cells?

The stress interval is not a confidence interval. It is a partial-identification / sensitivity interval for hidden composition. The audit returns one review artifact with the verdict, interval, witness, limitations, and claim-centered refinement recommendations.

Front-Door Demo

Start with the Folktables ACSIncome case if you want the shortest path to the idea. It is close to fairness and disparity-audit workflows because it asks whether a coarse demographic public report survives retained subgroup composition changes.

In that demo, the observed target rate is 12.37%. When the public mix by age band, education band, and sex is held fixed, hidden subgroup composition can move the target rate from:

11.79% to 13.44%

The transport ambiguity is 1.65 percentage points. In plain English:

The coarse public categories almost determine the aggregate income-threshold rate in this sample, but not quite. Hidden occupational and household differences still matter.

See the full analyst interpretation in docs/folktables-acs-income-interpretation.md.

The finance plugin is useful as a domain showcase and model-risk vocabulary layer, but the ACS/Folktables workflow is the best first demo for most Python users interested in aggregation bias, ecological fallacy, Simpson's paradox, or coarsened subgroup reporting.

Main Workflows

1. Validate a Claim

Use us.claim(...) when you want one review artifact that certifies the claim, breaks it with a hidden-composition witness, or proposes a stable repair. ClaimSpec is the underlying dataclass when you want to instantiate or serialize the spec directly, and ClaimAudit is the report object returned by .audit(...).

claim = us.claim(
    estimate_name="Income-threshold target rate",
    public=["AGE_BAND", "EDU_BAND", "SEX"],
    hidden=["AGE_BAND", "EDU_BAND", "SEX", "OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    target="income_over_threshold",
    weight="sample_weight",
    q_presets=[us.q_tv_budget(0.10), us.q_chi_square_budget(0.25)],
    candidate_refinements=["OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    ambiguity_limit=0.015,
    bucket_budget=40,
    statistical_interval=(0.119, 0.128),
)

verdict = claim.audit(rows_or_frame)
print(verdict.to_markdown())

The auditor separates the reported estimate, statistical uncertainty, hidden-composition ambiguity, public-refinement repair, counterexample witness, and limitations. verdict.recommend_refinements() returns claim-centered refinement rows: whether a candidate actually repairs the claim, whether it satisfies the ambiguity limit, and how much ambiguity it removes. See docs/reporting-claims.md.

For model-assisted plausibility checks, fit a nonparametric public/hidden joint distribution and run the claim across sampled hidden compositions:

joint = us.fit_joint_distribution(
    rows_or_frame,
    public=claim.public,
    hidden=claim.hidden,
    target=claim.target,
    weight=claim.weight,
)

verdict = us.audit_claim(
    rows_or_frame,
    claim,
    joint_model=joint,
    joint_draws=500,
    joint_seed=123,
)

Use hidden_composition_uncertainty(...) when you want the posterior/bootstrap draw summary as its own report. By default it preserves the observed public bucket mix and resamples hidden composition inside each public bucket.

See docs/model-assisted-joint-analysis.md.

2. Inspect The Evidence Behind A Claim

Most users should start with us.claim(...). Use public_descent_report(...) when you explicitly want only the lower-level partial-ID evidence without a claim verdict, decision rule, repair search, or claim-centered recommendation language.

report = us.public_descent_report(
    rows_or_frame,
    public=["segment"],
    hidden=["segment", "region", "tenure_band"],
    target="outcome_rate",
    candidate_refinements=["region", "tenure_band"],
    q=us.q_bounded_shift(0.5),
)

Use AuditSpec for serializable lower-level audit configurations that need to be reviewed or rerun:

spec = us.AuditSpec(
    public=["segment"],
    hidden=["segment", "region", "tenure_band"],
    target="outcome_rate",
    candidate_refinements=["region", "tenure_band"],
    q={"name": "bounded_shift", "radius": 0.5},
)

run = spec.run(rows_or_frame)
print(run.to_markdown())

See docs/audit-specs.md.

Reports also expose structured JSON and tables:

run.to_json()
run.to_tables()
run.to_dataframes()  # requires pandas

See docs/structured-exports.md.

Reports include pre-solve data diagnostics for sparse cells, dropped mass, singleton public fibers, constant-target fibers, and skipped refinement candidates. See docs/data-diagnostics.md.

When you need to inspect the actual lower-vs-upper endpoint worlds behind an ambiguity result, use report.witness_report() or us.witness_report(...) to see which hidden cells move while the public distribution stays fixed.

3. Add Robustness And Repair Intelligence

Claim audits already embed the primary report, optional certificate search, and claim-centered refinement recommendations. Use the lower-level robustness tools when you need to inspect those pieces separately:

  • sensitivity_report(...) checks Q presets, sparse-cell thresholds, or alternate hidden-state definitions.
  • certify_public_representation(...) finds the smallest evaluated public representation satisfying an ambiguity limit.
  • public_representation_frontier(...) exposes the Pareto frontier behind that certificate.
  • breakdown_point(...) finds the stress radius where a decision or ambiguity claim stops passing.
sensitivity = us.sensitivity_report(
    rows_or_frame,
    public=["AGE_BAND", "SEX"],
    hidden=["AGE_BAND", "SEX", "OCC_MAJOR", "WKHP_BAND"],
    target="income_over_threshold",
    min_cell_weights=[1, 10, 25],
    q_presets=["saturated", us.q_bounded_shift(0.5), "observed"],
)

See docs/transport-presets.md for guidance on Q presets and docs/representation-adequacy.md for interpretation rules.

4. Audit Causal or Uplift Reports

Use a causal inference library such as EconML, DoWhy, DoubleML, or CausalML to estimate effects first. Then pass row-level or subgroup-level effects into updatesupport.

suite = us.causal_reporting_stability(
    df,
    public=["AGE_BAND", "SEX"],
    hidden=["AGE_BAND", "SEX", "OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    effect="tau_hat",
    effect_standard_error="tau_se",
    weight="sample_weight",
    candidate_refinements=["OCC_MAJOR", "WKHP_BAND", "RAC1P"],
    q=us.q_bounded_shift(0.5),
    statistical_estimate=ate_hat,
    statistical_interval=(ci_low, ci_high),
    statistical_method="causal estimator bootstrap",
)

This keeps causal estimate, statistical uncertainty, hidden-composition ambiguity, estimator-uncertainty-aware adjusted ambiguity, and public refinement recommendations separate.

Try the tutorial notebook: DoWhy downstream reporting audit.

For causal/model-review stress tests based on hidden balance drift, use q=us.q_covariate_balance(epsilon, moments) to bound ||standardized_hidden_moment_shift||_2.

See docs/causal-library-integration.md.

5. Financial Model-Risk Plugin

Financial model-risk use cases live in the separate plugin package:

pip install "updatesupport[finance]"
# or
uv add "updatesupport[finance]"

It adds finance-oriented metrics such as expected loss and default rate without cluttering the core package.

See packages/updatesupport-finance/README.md.

Install Extras

pip install "updatesupport[cvxpy]"      # TV, chi-square, KL, Wasserstein, custom CVXPY Q
pip install "updatesupport[scip]"       # CVXPY plus PySCIPOpt for solver="SCIP" and MIP presets
pip install "updatesupport[examples]"   # Folktables, pandas, plotting examples
pip install "updatesupport[causal]"     # causal-effect examples
pip install "updatesupport[dowhy]"      # CausalRefutation conversion
pip install "updatesupport[finance]"    # financial model-risk plugin package

The same extras work with uv add:

uv add "updatesupport[cvxpy]"
uv add "updatesupport[scip]"
uv add "updatesupport[examples]"
uv add "updatesupport[causal]"
uv add "updatesupport[dowhy]"
uv add "updatesupport[finance]"

The causal docs cover EconML, DoWhy, DoubleML, and CausalML handoffs.

For common examples and optional integrations:

pip install "updatesupport[examples,causal,dowhy,cvxpy,finance]"
# or
uv add "updatesupport[examples,causal,dowhy,cvxpy,finance]"

updatesupport[scip] is intentionally separate because solver wheels and platform requirements are heavier than the default CVXPY workflow. It enables SCIP-backed mixed-integer presets such as q_fiber_support_floor(...).

Examples

Run the no-download Folktables smoke demo from a source checkout:

uv run python examples/folktables_acs.py --synthetic

Run the no-download AI / ML evaluation stability demo:

uv run python examples/ml_eval_stability.py

Run the no-download product experimentation / A/B testing stability demo:

uv run python examples/product_experiment_stability.py

Run the real Folktables ACSIncome example:

uv run --extra examples python examples/folktables_acs.py \
  --task income \
  --states CA \
  --year 2018 \
  --sample 50000 \
  --min-cell-weight 25

Generate the benchmark gallery:

uv run --extra examples --extra causal python examples/benchmark_gallery.py

Run the finance plugin example:

uv run --package updatesupport-finance python \
  packages/updatesupport-finance/examples/model_risk_portfolio.py

See docs/benchmark-gallery.md for reproducible case studies.

Source Development

git clone https://github.com/nahuaque/updatesupport.git
cd updatesupport
uv sync --all-packages --group dev --extra cvxpy --extra examples
uv run pytest

Documentation

Start Here

Positioning And Foundations

Claim Audits And Reports

Refinement And Robustness

Integrations And Case Studies

Project

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

updatesupport-0.1.3.tar.gz (474.7 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

updatesupport-0.1.3-py3-none-any.whl (160.7 kB view details)

Uploaded Python 3

File details

Details for the file updatesupport-0.1.3.tar.gz.

File metadata

  • Download URL: updatesupport-0.1.3.tar.gz
  • Upload date:
  • Size: 474.7 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for updatesupport-0.1.3.tar.gz
Algorithm Hash digest
SHA256 4c2e51ec92b0db5bcc72f1698359e580385ad676c1c311c861eb164b1468d8fb
MD5 137049b27cb28539cdb96a8482371804
BLAKE2b-256 257aa25f2d2b63b38a2b320c7e924770e53463c0720666c7fead730d755e81d7

See more details on using hashes here.

File details

Details for the file updatesupport-0.1.3-py3-none-any.whl.

File metadata

  • Download URL: updatesupport-0.1.3-py3-none-any.whl
  • Upload date:
  • Size: 160.7 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for updatesupport-0.1.3-py3-none-any.whl
Algorithm Hash digest
SHA256 b20c3f38793cddab6150f7aee795ad2bf3c70ece33d93cc16b9f09683fdb6610
MD5 1805a3caea94c3ca75ba0cc0305cee9f
BLAKE2b-256 c4bd6df606de35dc4f0c330038554c8cc404a32db5ab08b73943f74d08def1f6

See more details on using hashes here.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page