population-resemblance
Python tools for monitoring changes in categorical populations with the Population Resemblance Statistic (PRS), sample-size-aware thresholds, PSI and discrete KS benchmarks, and Monte Carlo simulation.
Documentation · Getting started · API reference · Contributing
What it does
Compare current category counts against a fixed reference distribution, assess whether the shift is acceptable under a chosen tolerance, and identify which categories contribute to the discrepancy. Applications include model-output classes, customer segments, and other discrete populations.
The core statistic is
$$ \mathrm{PRS} = \sum_{j=1}^{B} \frac{(\hat p_j - p_{0j})^2}{p_{0j}}. $$
The package implements the $\delta$-resemblance framework of Potgieter et al., including three decision regions. PSI and discrete KS are benchmarks with different decision rules. Named categories, temporal reports, reusable monitors, calibration diagnostics, and JSON export support recurring monitoring workflows.
Installation
Status: preparing the first public release, 0.1.0; not yet published to PyPI. Use a source installation with Python 3.12, 3.13, or 3.14:
git clone https://github.com/DiogoRibeiro7/population-resemblance.git
cd population-resemblance
python -m venv .venv
Activate the environment with source .venv/bin/activate on macOS/Linux or
.\.venv\Scripts\Activate.ps1 in Windows PowerShell, then install:
python -m pip install .
NumPy and SciPy are installed automatically. Contributors should use Poetry 2.5.x (2.5.1 matches CI); see the development setup. Track publication in the release notes.
Quick start
Run this in the installed environment. The count-based API derives the sample size from the data; the counts and reference probabilities must use the same category order.
from population_resemblance import assess_population_counts
result = assess_population_counts(
counts=[400, 350, 250],
reference=[0.50, 0.30, 0.20],
)
print(f"PRS: {result.statistic:.6f}")
print(f"Region: {result.region} ({result.label})")
print(
f"Thresholds: {result.critical_values.lower:.6f}, "
f"{result.critical_values.upper:.6f}"
)
PRS: 0.040833
Region: R3 (fully discrepant)
Thresholds: 0.000710, 0.007793
Here the PRS exceeds the upper threshold, placing the sample in R3 under the default calibration. Investigate the population change; this classification alone does not establish that a predictive model has failed.
| Region | Label | Interpretation |
|---|---|---|
| R1 | acceptable | PRS is at or below the lower threshold; continue monitoring. |
| R2 | partially discrepant | PRS is between the thresholds; increase monitoring. |
| R3 | fully discrepant | PRS exceeds the upper threshold; investigate the discrepancy. |
PRS is a discrepancy statistic, not a p-value. Supply delta= for an explicit
category-probability tolerance, or use the default that varies with sample size.
See calibration choices
before interpreting decisions in your application.
Assumptions and limitations
- PRS uses a fixed reference with strictly positive category probabilities. Reference counts are treated conditionally as fixed probabilities; this is not a two-sample test.
- PRS thresholds use an asymptotic approximation. Sparse categories and small samples need calibration checks.
- Discrete KS depends on category order. Choose a meaningful, fixed order before comparing periods; arbitrary ordering of nominal labels changes the benchmark.
- PSI follows the source paper's convention of omitting zero observed-probability contributions. It does not automatically smooth them.
- Temporal reports assess each period separately, without adjustment for repeated testing or dependence between periods.
Read the input and interpretation guidance for zero counts, category alignment, and sparse data. Two-sample extensions remain research.
More examples
| Task | Documentation |
|---|---|
| Named categories and reusable monitors | Getting started |
| Temporal monitoring, diagnostics, and export | User guide |
| Calibration and sensitivity | Calibration choices |
| PRS/PSI comparisons and operating-characteristic curves | Simulation |
| Complete runnable workflows | Examples |
| Reproducing published numerical results | Reproducibility |
Citation and license
This is an independent implementation and extension of:
Potgieter, C. J., Van Zyl, C., Schutte, W. D., & Lombard, F. (2026). The population resemblance statistic: a chi-square measure of fit for banking. Annals of Operations Research, 361, 413–435. doi:10.1007/s10479-025-07024-6
For research use, cite both the methodology and this software using CITATION.cff. Distributed under the MIT license.
Metadata
Release files for population-resemblance 0.1.0
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| population_resemblance-0.1.0.tar.gz | 24.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| population_resemblance-0.1.0-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 61.5 kB
Release files / population_resemblance-0.1.0.tar.gz
| Download URL | population_resemblance-0.1.0.tar.gz |
|---|---|
| Size | 24.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
87ffc509364486f454629595fdcbacfc41b9d589bcaf34165341a7a62c27a2ed
|
|
BLAKE2b-256 checksum How to use checksums |
c64ebd0e41533751dcbb1120da25fa8d80fc1c887f539a35233868702e426ba1
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency logRelease files / population_resemblance-0.1.0-py3-none-any.whl
| Download URL | population_resemblance-0.1.0-py3-none-any.whl |
|---|---|
| Size | 37.1 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c50a53bcf4e98adb91e44fe6562d162c220cadad9343878bd1563f00170851bf
|
|
BLAKE2b-256 checksum How to use checksums |
b9ad981b7c31c049742645fd709ef98c401c657058af07b44269e568a76c54f0
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 30, 2026.
Transparency log