shap-recommender
Exclusion / non-linearity / interaction recommendations from saved SHAP attribution files, and application of them to a design matrix.
One shared idea runs through two of the three rules. A feature's own contribution is represented flexibly (indicator columns when it takes few values, a restricted cubic spline when it is continuous), and a model built on that flexible basis is compared against a straight line in the feature. That comparison answers two different questions:
- non-linearity -- does the attribution deviate from a linear function of the feature? This directly tests the linear-trend assumption, rather than relying on a raw correlation coefficient, which conflates "no effect" with "non-linear effect".
- interaction -- does the attribution vary among subjects who share the same feature value? Under additivity the attribution is a deterministic function of the feature, so residual dispersion implies effect modification.
The interaction screen routes each candidate partner to the statistic that is valid for its type:
- Binary partners go to a within-stratum contrast. The stratifying
feature is split at a pre-specified cut point (never on the attribution
itself, which would condition on the candidate modifier), and
E[phi_y | y=1] - E[phi_y | y=0]is compared across strata. For a binary partner that two-point contrast equalsg(1) - g(0)exactly, so it is free of the within-stratum distribution ofy. - Continuous and ordinal partners go to a product-term regression, with
no stratification at all: the stratifying feature's own attribution is
regressed on a flexible basis in
x, the partners, and the productspartner * (x - mean x), and the product coefficient is the modification signal. A within-stratum OLS slope isCov(y, h(y)) / Var(y)under additivity, which moves with the distribution ofyinside the stratum wheneverhis non-linear, so comparing that slope across strata can declare an interaction where none exists.
Both use bootstrap standard errors to report approximate two-sided Wald
p-values and confidence intervals (n_boot_screen resamples), and both are flagged against
min_abs_effect applied to the effect scaled by the typical attribution
magnitude (phi_scale) rather than by the estimate's own magnitude, since
a ratio to the estimate itself blows up whenever it sits near zero. The
test table records which statistic was used for each pair in a method
column.
The stratifying feature and each continuous or ordinal partner are standardized before the product-term regression. This makes interaction effect sizes comparable across features measured in different units.
Install
pip install shap-recommender
Expected input files
For each of the two dataset tags, Recommender.load(tag) expects two
tab-separated files in res_dir:
shap_values_<tag>.tsv-- SHAP values, one row per subject, one column per feature, first column = row index.sel_data_<tag>.tsv-- the corresponding feature values (design matrix), same row index.
Command-line use
shap-recommender \
--res-dir ./shap_results \
--tags cohort_low cohort_high \
--nonlinear-candidates age bmi creatinine \
--out ./recommendations
The two tags are screened separately. Exclusion recommendations are retained only when a feature is negligible in both cohorts; non-linearity and interaction recommendations use the union of results across cohorts, with a single multiplicity correction over pooled interaction tests.
This writes cohort-specific exclusion, sensitivity, non-linearity, and
attribution-pattern tables, plus interaction_tests.tsv,
attribution_patterns.tsv, and recommendations.json to --out. Run
shap-recommender --help for all options (thresholds, bootstrap count, spline
degrees of freedom, a --cutpoints JSON file for pre-specified stratification
cut points, etc).
Library use
from shap_recommender import Recommender
rec = Recommender(res_dir="./shap_results")
recommendations = rec.generate(
candidates_nonlinear=["age", "bmi", "creatinine"],
low_tag="cohort_low",
high_tag="cohort_high",
out="./recommendations",
)
# apply the recommendations to a design matrix
X_train_adj, X_test_adj = Recommender.apply(
X_train, X_test, recommendations, variant="all",
)
Recommender.apply(..., variant=...) accepts "baseline", "exclusion",
"nonlinear", "interaction", or "all", so each rule's effect on
downstream model performance can be evaluated separately.
A feature flagged non-linear is expanded one of two ways, chosen by how
many distinct values it takes in the training data (n_bins, default 10):
a low-cardinality (ordinal) feature is expanded into per-level indicator
columns (<feature>_lvl<value>, one per level after a reference level),
since a quadratic in the level code cannot represent an arbitrary
threshold effect; a feature with more distinct values than n_bins is
centred on its training mean and given a <feature>_quad column instead.
An interaction partner that is itself such an expanded ordinal feature
attaches to every one of its indicator columns, rather than being dropped
for a missing main effect.
Validating the interaction rule
Recommender.null_sim() runs a small simulation under an additive null
(no true interaction with the feature being tested) and reports the type-I
error rate of the within-stratum contrast used by stratified_screen,
compared against naively partitioning on the attribution itself:
from shap_recommender import Recommender
Recommender(res_dir=".").null_sim()
License
MIT
Release files for shap-recommender 0.6.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| shap_recommender-0.6.1.tar.gz | 21.7 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| shap_recommender-0.6.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 39.7 kB
Release files / shap_recommender-0.6.1.tar.gz
| Download URL | shap_recommender-0.6.1.tar.gz |
|---|---|
| Size | 21.7 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
12df60990fd4de2eb3f6d0ad3b7e398faa31d4650d5cb400d3e322bfa5075547
|
|
BLAKE2b-256 checksum How to use checksums |
3b5231b501e56d5dbcd391c733e93769b019dff7c2de9640c91c2b53e1905b14
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.9
|
Release files / shap_recommender-0.6.1-py3-none-any.whl
| Download URL | shap_recommender-0.6.1-py3-none-any.whl |
|---|---|
| Size | 17.9 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
c5b59fdc6c51c818a7f4144de0eae2c798d50a3b314a15650463d4a589fab34b
|
|
BLAKE2b-256 checksum How to use checksums |
21247f5877128fc27ee6c282f9fad61dba2ce46a0514afef06d903a74e5e651e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/7.0.0 CPython/3.13.9
|