riskval
A regulatory validation battery for credit- and market-risk models.
Python has no maintained, tested implementation of the statistical tests that
bank supervisors actually apply to internal models. Model risk management teams
rewrite them in Excel or SAS every year. riskval is that battery, in one
package, with every number checked against a published reference.
pip install riskval
Quickstart
>>> import numpy as np
>>> from riskval import validate
>>>
>>> rng = np.random.default_rng(0)
>>> pd_true = rng.uniform(0.01, 0.30, size=4000) # the true default rates
>>> y = rng.binomial(1, pd_true) # who actually defaulted
>>>
>>> report = validate(y, pd_true, seed=0)
>>> report.rejected
[]
Now hand the same battery a model that ranks obligors perfectly but prices them at half the risk:
>>> report = validate(y, pd_true / 2, seed=0)
>>> report.rejected
['Hosmer-Lemeshow', 'Spiegelhalter Z']
The ranking is untouched, so every discrimination measure is bit-for-bit identical to the good model's:
>>> good = validate(y, pd_true, seed=0)
>>> round(report.metrics["gini"], 6) == round(good.metrics["gini"], 6)
True
That gap is the whole point. A model can look excellent on AUC and still fail every test a supervisor runs. And it is not free: put the misestimated PD through the Basel IRB formula and read off the capital shortfall.
>>> from riskval import rwa_deviation
>>> out = rwa_deviation(pd_true=0.03, pd_hat=0.01, lgd=0.45, ead=1_000_000)
>>> out["direction"]
'understated'
>>> round(out["rwa_relative_deviation"], 4)
-0.2812
A grade priced at 1% that actually defaults at 3% carries 28% too little capital.
What is in it
| Module | Question | Contents |
|---|---|---|
riskval.calibration |
Is the level of the PDs right, and is their spread? | Exact binomial test, Jeffreys test, Hosmer–Lemeshow, Spiegelhalter's Z, Cox calibration slope and joint test, calibration curve table, grade_level_report |
riskval.discrimination |
Is the ordering right? | AUC with bootstrap CI, Gini, Kolmogorov–Smirnov, Somers' D |
riskval.backtest |
Does the VaR model hold up? | Kupiec POF, Christoffersen independence and conditional coverage, Basel traffic light, capital multiplier add-on |
riskval.capital |
What does being wrong cost? | Basel IRB risk weights for four asset classes, RWA, rwa_deviation |
riskval.report |
All of it at once | validate, ValidationReport |
Dependencies: numpy, scipy, pandas. Nothing else, and nothing proprietary
— no Bloomberg, no Refinitiv, no vendor data.
The promise: every number traces to a published reference
Nobody installs a package for a formula they could write in twenty lines. They install it so they do not have to defend those twenty lines. So each test is checked in the test suite against a value someone else published:
| Test | Reference |
|---|---|
| Exact binomial | BCBS (1996) Table 1: cumulative binomial probabilities for 250 days at 1%, all nine rows. Plus the tail in exact rational arithmetic via fractions.Fraction. |
| Jeffreys | The beta–binomial identity, and quadrature of the posterior density independent of scipy.special.betainc |
| Hosmer–Lemeshow | Closed-form arithmetic done by hand, plus a reproduction of Hosmer and Lemeshow's own simulation establishing the null distribution |
| Spiegelhalter Z | Algebraic equivalence of the compact form and the standardised Brier score, plus a Monte Carlo check that the null really is N(0,1) |
| Cox calibration slope | A saturated two-point design, whose intercept and slope have a closed form, matched to 1e-10; statsmodels for the coefficients, standard errors and log-likelihood; and recovery of a known distortion (predictions pushed to c times their true log-odds must return a slope of 1/c) |
| AUC, Gini, Somers' D | Hanley and McNeil (1982) Table 1, whose published area of 0.893 is reproduced as the exact rational 1321/1479; cross-checked against sklearn and scipy.stats.mannwhitneyu, and against brute-force pair counting |
| Kolmogorov–Smirnov | scipy.stats.ks_2samp |
| Kupiec POF | Kupiec (1995) Table 1: the 95% non-rejection regions, all fifteen cells, bound by bound |
| Christoffersen | Longhand re-derivation of both likelihood ratios, plus Monte Carlo null distributions |
| Basel IRB risk weights | BCBS (2006) Annex 3 "Illustrative IRB Risk Weights": all 19 PD rows × 4 asset classes, and the formula rebuilt from the regulatory text |
| Basel traffic light | BCBS (1996) Tables 1 and 2 |
If a function has no reference value, it does not go in the package.
Four places where the conventional implementation is wrong
These are the reasons to use riskval rather than a snippet off the internet.
1. The Hosmer–Lemeshow degrees of freedom depend on where the predictions
came from. The familiar G − 2 is correct when the predictions come from a
logistic model fitted to the very data under test — two degrees of freedom go
on the fitted intercept and slope. In a validation exercise the predictions
are exogenous: a model estimated on a development sample, or a vendor model,
scored on fresh data. Nothing is estimated from the test sample, and the null
distribution is χ²G. riskval verifies both regimes by simulation:
| Predictions | Mean statistic (G = 10) | KS vs χ²₁₀ | KS vs χ²₈ |
|---|---|---|---|
| Exogenous | 9.86 | p = 0.58 | p = 9 × 10⁻¹⁸ |
| Fitted in-sample | 8.04 | p = 2 × 10⁻¹⁵ | p = 0.99 |
Using G − 2 on exogenous predictions overstates the evidence against the
model. validate() therefore defaults to G, and
hosmer_lemeshow(..., df=...) lets you say which you mean.
2. The Jeffreys test is what the ECB prescribes, and it materially outperforms the exact binomial test on small grades. The binomial test is conservative because the binomial is discrete. For a grade of 100 obligors at a 2% PD, at a nominal 5% level:
| Test | Realised size | Rejects from |
|---|---|---|
| Exact binomial | 1.5% | 6 defaults |
| Jeffreys | 5.0% | 5 defaults |
3. A one-sided supervisory test does not answer "is this model calibrated". Pillar 1 exists to catch insufficient capital, so the binomial test a supervisor specifies is one-sided: is the PD understated? A model that errs on the conservative side clears it while being badly miscalibrated.
And a conservative level error is not the same as surplus prudence. Which way the miscalibration moves capital is not settled by the level: the IRB risk weight is concave in PD, so the capital number responds to the spread of the predictions, not their level — see the dispersion section below.
grade_level_report therefore reports all three directions, and
ValidationReport.summary() shows the split:
>>> import numpy as np, pandas as pd
>>> from riskval import validate
>>> rng = np.random.default_rng(1)
>>> pd_true = rng.uniform(0.005, 0.25, size=6000)
>>> y = rng.binomial(1, pd_true)
>>> grades = np.asarray(pd.cut(pd_true, bins=5, labels=list("ABCDE")))
>>>
>>> report = validate(y, pd_true * 1.6, grades=grades, n_boot=50, seed=0)
>>> frame = report.grade_report
>>> frame.attrs["rejection_rate_understated"] # what a supervisor tests
0.0
>>> frame.attrs["rejection_rate_twosided"] # whether it is calibrated
1.0
Every PD inflated by 60%, and the supervisory test flags nothing.
4. The 1.06 IRB scaling factor. Basel II and the EU CRR multiply IRB risk
weights by 1.06; the finalised Basel III framework removed it. Implementations
disagree and rarely say which they use. riskval defaults to 1.0, says so, and
takes scaling_factor=1.06 for CRR figures.
riskval also documents where its tests stop working. Simulated under the
true model at the Basel 250-day window and 99% coverage, a nominal 5% Kupiec
test rejects correct models 9.1% of the time, while Christoffersen's
independence test rejects only 1.3%. Both figures are pinned in the test suite
and stated in the docstrings.
The gap nobody tests: dispersion
Every test above asks whether the PD level is right — for the portfolio, or grade by grade. None of them asks whether the predictions are spread right, and that is what costs capital, because the Basel IRB risk weight is concave in PD. Mean-preserving spread in the PDs therefore destroys risk-weighted assets: an over-dispersed model reports less capital than the realised defaults require, while clearing the level tests.
calibration_slope is the Cox (1958) recalibration slope, and it is in
validate() for exactly this reason. Take a portfolio and widen the spread of
its predictions on the logit scale, leaving the level alone:
| spread factor | Cox slope | one-sided supervisory rejection | capital deviation |
|---|---|---|---|
| 1.0 | 0.985 | 0.2 | −0.3% |
| 1.3 | 0.758 | 0.3 | −3.4% |
| 1.6 | 0.616 | 0.4 | −7.2% |
| 1.8 | 0.547 | 0.4 | −9.9% |
| 2.2 | 0.448 | 0.4 | −15.8% |
The capital shortfall doubles from −7.2% to −15.8% and the one-sided
supervisory rejection rate does not move at all. The slope tracks the damage;
the level test does not. Both columns are pinned by tests in
tests/test_calibration.py.
The complement holds too, which is why both belong in the battery: halve every PD and the level tests reject while the slope correctly reports 1.0 — a level failure is what the supervisory test is for.
Use a materiality band, not a point null
A point null on 10 000 obligors rejects a slope of 0.97 — statistically real,
economically irrelevant. materiality=m tests |b − 1| ≤ m instead, a
minimal-effect test in the sense of Wellek (2010). Scored against the realised
capital deviation on 402 credit models:
| trigger rule | detects a >10% capital shortfall | false alarm |
|---|---|---|
| slope test rejects at 5% (point null) | 99.5% | 76.9% |
| grade-level one-sided rejection rate > 20% | 31.6% | 38.5% |
| grade-level two-sided rate > 50% | 60.9% | 53.8% |
materiality=0.15, i.e. the interval clears [0.85, 1.15] |
97.7% | 26.9% |
The rank correlation with the capital deviation is +0.80 for the slope and +0.02 (p = 0.72) for the grade-level one-sided rate — the instrument in use today is uncorrelated with the damage. Caveat: the false-alarm column rests on 26 cells whose capital was within 2% of correct, so read it as indicative.
res = calibration_slope(y_true, y_prob, materiality=0.15)
res.statistic # the slope
res.detail["band"] # (0.85, 1.15)
res.reject # materially over- or under-dispersed
What it does not do
By design:
- No model training.
riskvalevaluates models; it does not build them. - No plotting. It returns tables; you plot them.
- No data downloading. It is not a data source.
- No proprietary dependencies.
Development
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev]"
pytest --cov=riskval --cov-report=term-missing
ruff check src tests
Test coverage is 100% and the floor is 85%. The suite includes the README's own examples, so the quickstart above cannot go stale.
Citing
If riskval contributed to published work, please cite it. See CITATION.cff.
License
MIT. See LICENSE.
Release files for riskval 0.1.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| riskval-0.1.1.tar.gz | 82.0 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| riskval-0.1.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 129.2 kB
Release files / riskval-0.1.1.tar.gz
| Download URL | riskval-0.1.1.tar.gz |
|---|---|
| Size | 82.0 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
bcd8af1fa095d10875b7afb64d00c16c106302402e57f981c398ef257120fffd
|
|
BLAKE2b-256 checksum How to use checksums |
4c602203346c833463b10d2cb9bc1f3cbfe839b25955dca55f59eaa83b235f30
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency logRelease files / riskval-0.1.1-py3-none-any.whl
| Download URL | riskval-0.1.1-py3-none-any.whl |
|---|---|
| Size | 47.2 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
632b094aa5f964e0cb723b3ad288d1233ef2ef490ef0c17353491134a1f81838
|
|
BLAKE2b-256 checksum How to use checksums |
fd91fe6d9067e114562a8fcdc72ae21ea1a55e29bac3a3e3e5bd3ee6645fbcbf
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
Yes |
| Uploaded via |
twine/7.0.0 CPython/3.13.14
|
Provenance
Provenance describes where a file came from. On PyPI, provenance is shared via attestations, which provide a verifiable record of the build or publishing details. View details, limitations and caveats.
PyPI Publish Attestation
PyPI verified that this artifact, at this checksum, originated from the publisher listed below.
Signed by GitHub Actions, verified by PyPI on Sep 7, 2026.
Transparency log