Skip to main content

riskval

A regulatory validation battery for credit- and market-risk models.

Python has no maintained, tested implementation of the statistical tests that bank supervisors actually apply to internal models. Model risk management teams rewrite them in Excel or SAS every year. riskval is that battery, in one package, with every number checked against a published reference.

pip install riskval

Quickstart

>>> import numpy as np
>>> from riskval import validate
>>>
>>> rng = np.random.default_rng(0)
>>> pd_true = rng.uniform(0.01, 0.30, size=4000)   # the true default rates
>>> y = rng.binomial(1, pd_true)                   # who actually defaulted
>>>
>>> report = validate(y, pd_true, seed=0)
>>> report.rejected
[]

Now hand the same battery a model that ranks obligors perfectly but prices them at half the risk:

>>> report = validate(y, pd_true / 2, seed=0)
>>> report.rejected
['Hosmer-Lemeshow', 'Spiegelhalter Z']

The ranking is untouched, so every discrimination measure is bit-for-bit identical to the good model's:

>>> good = validate(y, pd_true, seed=0)
>>> round(report.metrics["gini"], 6) == round(good.metrics["gini"], 6)
True

That gap is the whole point. A model can look excellent on AUC and still fail every test a supervisor runs. And it is not free: put the misestimated PD through the Basel IRB formula and read off the capital shortfall.

>>> from riskval import rwa_deviation
>>> out = rwa_deviation(pd_true=0.03, pd_hat=0.01, lgd=0.45, ead=1_000_000)
>>> out["direction"]
'understated'
>>> round(out["rwa_relative_deviation"], 4)
-0.2812

A grade priced at 1% that actually defaults at 3% carries 28% too little capital.

What is in it

Module Question Contents
riskval.calibration Is the level of the PDs right, and is their spread? Exact binomial test, Jeffreys test, Hosmer–Lemeshow, Spiegelhalter's Z, Cox calibration slope and joint test, calibration curve table, grade_level_report
riskval.discrimination Is the ordering right? AUC with bootstrap CI, Gini, Kolmogorov–Smirnov, Somers' D
riskval.backtest Does the VaR model hold up? Kupiec POF, Christoffersen independence and conditional coverage, Basel traffic light, capital multiplier add-on
riskval.capital What does being wrong cost? Basel IRB risk weights for four asset classes, RWA, rwa_deviation
riskval.report All of it at once validate, ValidationReport

Dependencies: numpy, scipy, pandas. Nothing else, and nothing proprietary — no Bloomberg, no Refinitiv, no vendor data.

The promise: every number traces to a published reference

Nobody installs a package for a formula they could write in twenty lines. They install it so they do not have to defend those twenty lines. So each test is checked in the test suite against a value someone else published:

Test Reference
Exact binomial BCBS (1996) Table 1: cumulative binomial probabilities for 250 days at 1%, all nine rows. Plus the tail in exact rational arithmetic via fractions.Fraction.
Jeffreys The beta–binomial identity, and quadrature of the posterior density independent of scipy.special.betainc
Hosmer–Lemeshow Closed-form arithmetic done by hand, plus a reproduction of Hosmer and Lemeshow's own simulation establishing the null distribution
Spiegelhalter Z Algebraic equivalence of the compact form and the standardised Brier score, plus a Monte Carlo check that the null really is N(0,1)
Cox calibration slope A saturated two-point design, whose intercept and slope have a closed form, matched to 1e-10; statsmodels for the coefficients, standard errors and log-likelihood; and recovery of a known distortion (predictions pushed to c times their true log-odds must return a slope of 1/c)
AUC, Gini, Somers' D Hanley and McNeil (1982) Table 1, whose published area of 0.893 is reproduced as the exact rational 1321/1479; cross-checked against sklearn and scipy.stats.mannwhitneyu, and against brute-force pair counting
Kolmogorov–Smirnov scipy.stats.ks_2samp
Kupiec POF Kupiec (1995) Table 1: the 95% non-rejection regions, all fifteen cells, bound by bound
Christoffersen Longhand re-derivation of both likelihood ratios, plus Monte Carlo null distributions
Basel IRB risk weights BCBS (2006) Annex 3 "Illustrative IRB Risk Weights": all 19 PD rows × 4 asset classes, and the formula rebuilt from the regulatory text
Basel traffic light BCBS (1996) Tables 1 and 2

If a function has no reference value, it does not go in the package.

Four places where the conventional implementation is wrong

These are the reasons to use riskval rather than a snippet off the internet.

1. The Hosmer–Lemeshow degrees of freedom depend on where the predictions came from. The familiar G − 2 is correct when the predictions come from a logistic model fitted to the very data under test — two degrees of freedom go on the fitted intercept and slope. In a validation exercise the predictions are exogenous: a model estimated on a development sample, or a vendor model, scored on fresh data. Nothing is estimated from the test sample, and the null distribution is χ²*G*. riskval verifies both regimes by simulation:

Predictions Mean statistic (G = 10) KS vs χ²₁₀ KS vs χ²₈
Exogenous 9.86 p = 0.58 p = 9 × 10⁻¹⁸
Fitted in-sample 8.04 p = 2 × 10⁻¹⁵ p = 0.99

Using G − 2 on exogenous predictions overstates the evidence against the model. validate() therefore defaults to G, and hosmer_lemeshow(..., df=...) lets you say which you mean.

2. The Jeffreys test is what the ECB prescribes, and it materially outperforms the exact binomial test on small grades. The binomial test is conservative because the binomial is discrete. For a grade of 100 obligors at a 2% PD, at a nominal 5% level:

Test Realised size Rejects from
Exact binomial 1.5% 6 defaults
Jeffreys 5.0% 5 defaults

3. A one-sided supervisory test does not answer "is this model calibrated". Pillar 1 exists to catch insufficient capital, so the binomial test a supervisor specifies is one-sided: is the PD understated? A model that errs on the conservative side clears it while being badly miscalibrated. That is not a prudential problem, but it is an economic one — capital held against losses that will not happen, credit refused to borrowers who would have repaid.

grade_level_report therefore reports all three directions, and ValidationReport.summary() shows the split:

>>> import numpy as np, pandas as pd
>>> from riskval import validate
>>> rng = np.random.default_rng(1)
>>> pd_true = rng.uniform(0.005, 0.25, size=6000)
>>> y = rng.binomial(1, pd_true)
>>> grades = np.asarray(pd.cut(pd_true, bins=5, labels=list("ABCDE")))
>>>
>>> report = validate(y, pd_true * 1.6, grades=grades, n_boot=50, seed=0)
>>> frame = report.grade_report
>>> frame.attrs["rejection_rate_understated"]   # what a supervisor tests
0.0
>>> frame.attrs["rejection_rate_twosided"]      # whether it is calibrated
1.0

Every PD inflated by 60%, and the supervisory test flags nothing.

4. The 1.06 IRB scaling factor. Basel II and the EU CRR multiply IRB risk weights by 1.06; the finalised Basel III framework removed it. Implementations disagree and rarely say which they use. riskval defaults to 1.0, says so, and takes scaling_factor=1.06 for CRR figures.

riskval also documents where its tests stop working. Simulated under the true model at the Basel 250-day window and 99% coverage, a nominal 5% Kupiec test rejects correct models 9.1% of the time, while Christoffersen's independence test rejects only 1.3%. Both figures are pinned in the test suite and stated in the docstrings.

The gap nobody tests: dispersion

Every test above asks whether the PD level is right — for the portfolio, or grade by grade. None of them asks whether the predictions are spread right, and that is what costs capital, because the Basel IRB risk weight is concave in PD. Mean-preserving spread in the PDs therefore destroys risk-weighted assets: an over-dispersed model reports less capital than the realised defaults require, while clearing the level tests.

calibration_slope is the Cox (1958) recalibration slope, and it is in validate() for exactly this reason. Take a portfolio and widen the spread of its predictions on the logit scale, leaving the level alone:

spread factor Cox slope one-sided supervisory rejection capital deviation
1.0 0.985 0.2 −0.3%
1.3 0.758 0.3 −3.4%
1.6 0.616 0.4 −7.2%
1.8 0.547 0.4 −9.9%
2.2 0.448 0.4 −15.8%

The capital shortfall doubles from −7.2% to −15.8% and the one-sided supervisory rejection rate does not move at all. The slope tracks the damage; the level test does not. Both columns are pinned by tests in tests/test_calibration.py.

The complement holds too, which is why both belong in the battery: halve every PD and the level tests reject while the slope correctly reports 1.0 — a level failure is what the supervisory test is for.

Use a materiality band, not a point null

A point null on 10 000 obligors rejects a slope of 0.97 — statistically real, economically irrelevant. materiality=m tests |b − 1| ≤ m instead, a minimal-effect test in the sense of Wellek (2010). Scored against the realised capital deviation on 402 credit models:

trigger rule detects a >10% capital shortfall false alarm
slope test rejects at 5% (point null) 99.5% 76.9%
grade-level one-sided rejection rate > 20% 31.6% 38.5%
grade-level two-sided rate > 50% 60.9% 53.8%
materiality=0.15, i.e. the interval clears [0.85, 1.15] 97.7% 26.9%

The rank correlation with the capital deviation is +0.80 for the slope and +0.02 (p = 0.72) for the grade-level one-sided rate — the instrument in use today is uncorrelated with the damage. Caveat: the false-alarm column rests on 26 cells whose capital was within 2% of correct, so read it as indicative.

res = calibration_slope(y_true, y_prob, materiality=0.15)
res.statistic          # the slope
res.detail["band"]     # (0.85, 1.15)
res.reject             # materially over- or under-dispersed

What it does not do

By design:

  • No model training. riskval evaluates models; it does not build them.
  • No plotting. It returns tables; you plot them.
  • No data downloading. It is not a data source.
  • No proprietary dependencies.

Development

python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest --cov=riskval --cov-report=term-missing
ruff check src tests

Test coverage is 100% and the floor is 85%. The suite includes the README's own examples, so the quickstart above cannot go stale.

Citing

If riskval contributed to published work, please cite it. See CITATION.cff.

License

MIT. See LICENSE.

Release files for riskval 0.1.0

For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.

Source distribution (sdist)

Source distribution for riskval 0.1.0
File Size Uploaded
riskval-0.1.0.tar.gz 81.5 kB Details

Built distribution (wheel)

Table of built distributions (wheels) for riskval 0.1.0
File Interpreter ABI Platform
riskval-0.1.0-py3-none-any.whl Python 3 none any Details

Total release size: 128.6 kB

Release files / riskval-0.1.0.tar.gz

Download URL riskval-0.1.0.tar.gz
Size 81.5 kB
Tags Source
SHA-256 checksum
How to use checksums
577e226e4202f72c925858615b36a39dd223c40b537f2c1b5a2d1d8e1c26fb1d
BLAKE2b-256 checksum
How to use checksums
13c0bf493cbae546e991046cebc942a166c9a8df26599e0fb7f345fabdc48030
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.

Transparency log

Release files / riskval-0.1.0-py3-none-any.whl

Download URL riskval-0.1.0-py3-none-any.whl
Size 47.0 kB
Tags Python 3
SHA-256 checksum
How to use checksums
a8a937435ed4eb40b94494964dcaafc58065ca7b7b8697d17c83e44d127b9f6c
BLAKE2b-256 checksum
How to use checksums
e59f7df6558fdbaad38a8b855369e79c6657178374ab7ab2d9f7e512d67789d0
Upload date
Uploaded using Trusted Publishing?
What is trusted publishing?
Yes
Uploaded via twine/7.0.0 CPython/3.13.14

Provenance

Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.

PyPI Publish Attestation

PyPI verified that this artifact, at this checksum, originated from the publisher listed below.

Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.

Transparency log

Release history Release notifications | RSS feed

0.1.1

2 release files

This release

0.1.0 This release

2 release files

0.0.1

2 release files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page