Skip to main content

population-resemblance

Python tools for monitoring changes in categorical populations with the Population Resemblance Statistic (PRS), sample-size-aware thresholds, PSI and discrete KS benchmarks, and Monte Carlo simulation.

Documentation · Getting started · API reference · Contributing

What it does

Compare current category counts against a fixed reference distribution, assess whether the shift is acceptable under a chosen tolerance, and identify which categories contribute to the discrepancy. Applications include model-output classes, customer segments, and other discrete populations.

The core statistic is

$$ \mathrm{PRS} = \sum_{j=1}^{B} \frac{(\hat p_j - p_{0j})^2}{p_{0j}}. $$

The package implements the $\delta$-resemblance framework of Potgieter et al., including three decision regions. PSI and discrete KS are benchmarks with different decision rules. Named categories, temporal reports, reusable monitors, calibration diagnostics, and JSON export support recurring monitoring workflows.

Installation

Status: preparing the first public release, 0.1.0; not yet published to PyPI. Use a source installation with Python 3.12, 3.13, or 3.14:

git clone https://github.com/DiogoRibeiro7/population-resemblance.git
cd population-resemblance
python -m venv .venv

Activate the environment with source .venv/bin/activate on macOS/Linux or .\.venv\Scripts\Activate.ps1 in Windows PowerShell, then install:

python -m pip install .

NumPy and SciPy are installed automatically. Contributors should use Poetry 2.5.x (2.5.1 matches CI); see the development setup. Track publication in the release notes.

Quick start

Run this in the installed environment. The count-based API derives the sample size from the data; the counts and reference probabilities must use the same category order.

from population_resemblance import assess_population_counts

result = assess_population_counts(
    counts=[400, 350, 250],
    reference=[0.50, 0.30, 0.20],
)

print(f"PRS: {result.statistic:.6f}")
print(f"Region: {result.region} ({result.label})")
print(
    f"Thresholds: {result.critical_values.lower:.6f}, "
    f"{result.critical_values.upper:.6f}"
)
PRS: 0.040833
Region: R3 (fully discrepant)
Thresholds: 0.000710, 0.007793

Here the PRS exceeds the upper threshold, placing the sample in R3 under the default calibration. Investigate the population change; this classification alone does not establish that a predictive model has failed.

Region Label Interpretation
R1 acceptable PRS is at or below the lower threshold; continue monitoring.
R2 partially discrepant PRS is between the thresholds; increase monitoring.
R3 fully discrepant PRS exceeds the upper threshold; investigate the discrepancy.

PRS is a discrepancy statistic, not a p-value. Supply delta= for an explicit category-probability tolerance, or use the default that varies with sample size. See calibration choices before interpreting decisions in your application.

Assumptions and limitations

  • PRS uses a fixed reference with strictly positive category probabilities. Reference counts are treated conditionally as fixed probabilities; this is not a two-sample test.
  • PRS thresholds use an asymptotic approximation. Sparse categories and small samples need calibration checks.
  • Discrete KS depends on category order. Choose a meaningful, fixed order before comparing periods; arbitrary ordering of nominal labels changes the benchmark.
  • PSI follows the source paper's convention of omitting zero observed-probability contributions. It does not automatically smooth them.
  • Temporal reports assess each period separately, without adjustment for repeated testing or dependence between periods.

Read the input and interpretation guidance for zero counts, category alignment, and sparse data. Two-sample extensions remain research.

More examples

Task Documentation
Named categories and reusable monitors Getting started
Temporal monitoring, diagnostics, and export User guide
Calibration and sensitivity Calibration choices
PRS/PSI comparisons and operating-characteristic curves Simulation
Complete runnable workflows Examples
Reproducing published numerical results Reproducibility

Citation and license

This is an independent implementation and extension of:

Potgieter, C. J., Van Zyl, C., Schutte, W. D., & Lombard, F. (2026). The population resemblance statistic: a chi-square measure of fit for banking. Annals of Operations Research, 361, 413–435. doi:10.1007/s10479-025-07024-6

For research use, cite both the methodology and this software using CITATION.cff. Distributed under the MIT license.

Metadata

Release files for population-resemblance 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for population-resemblance 0.1.0
File Size Uploaded
population_resemblance-0.1.0.tar.gz 24.4 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for population-resemblance 0.1.0
File Interpreter ABI Platform
population_resemblance-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 61.5 kB

Release files / population_resemblance-0.1.0.tar.gz

Download URL population_resemblance-0.1.0.tar.gz
Size 24.4 kB
Tags Source
SHA-256 checksum
How to use checksums
87ffc509364486f454629595fdcbacfc41b9d589bcaf34165341a7a62c27a2ed
BLAKE2b-256 checksum
How to use checksums
c64ebd0e41533751dcbb1120da25fa8d80fc1c887f539a35233868702e426ba1
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release files / population_resemblance-0.1.0-py3-none-any.whl

Download URL population_resemblance-0.1.0-py3-none-any.whl
Size 37.1 kB
Tags Python 3
SHA-256 checksum
How to use checksums
c50a53bcf4e98adb91e44fe6562d162c220cadad9343878bd1563f00170851bf
BLAKE2b-256 checksum
How to use checksums
b9ad981b7c31c049742645fd709ef98c401c657058af07b44269e568a76c54f0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.

Transparency log

Release history Release notifications | RSS feed

This release

0.1.0 This release

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page