English | 中文
RuleLift · Credit-Risk Rule Mining / Scorecard / Strategy Evaluation
From a labeled loan dataset to a risk-control strategy you can review, deploy, and monitor — RuleLift collapses the four stages Explore → Model → Decide → Monitor into one shared vocabulary, one family of result objects, and three isomorphic entry layers, so rule mining and scorecard building become "one line per step, what-you-see-is-what-you-get".
Version v2.10 | MIT License | Requires pandas / numpy / scikit-learn / scipy
Table of Contents
- 1. What is it
- 2. Install, import & namespace
- 3. Full-workflow guide (the core chapter)
- 3.0 The big picture
- 3.1 Step 0: prepare data, declare column roles
- 3.2 Explore I: profile — batch variable check
- 3.3 Explore II: inspect — bin detail & tuning cut points
- 3.4 Mine: single / cross / tree
- 3.5 Model: Scorecard (with bin/WOE export)
- 3.6 Decide: Strategy — evaluate / simulate / compare / reason codes
- 3.7 Monitor: live-rule effectiveness
- 3.8 One-shot: Pipeline
- 3.9 Ship it: save / load / package
- 3.10 Strategy tuning: A-type & D-type
- 4. Using it on your own big data (wide tables)
- 5. Data: the exploration context
- 6. Scorecard: modeling in depth
- 7. Strategy: the decision context
- 8. RuleSet / Rule
- 9. The uniform result-object protocol
- 10. Pipeline / top-level functions / Agent
- 11. Performance & capacity (big data)
- 12. Offline install
- License & info
1. What is it
RuleLift is a credit risk-control toolkit that helps you turn a labeled loan dataset (good/bad flag) into a deployable rule strategy or scorecard. It wires the four everyday tasks into one smooth pipeline, each essentially "one line, one step":
- Analyze variables — which features separate good from bad customers (
profile/inspect); - Mine rules — automatically find "which threshold isolates the high-risk crowd" (
mine/cross); - Build a scorecard — turn strong features into an interpretable, better-performing score table (
Scorecard); - Decide & monitor — evaluate a rule's business impact and track it after go-live (
Strategy/monitor).
No jumping between scattered scripts: metric definitions are unified, results are comparable, and statistical gating is built in.
A key design idea: evaluating live rules needs no A/B (traffic-split) test and no waiting for the performance labels of rejected customers. RuleLift estimates a rule's bad rate from the credit-rating distribution of the customers it hits, then contrasts it with the observable actuals — so even deployed rules stay measurable.
The four core objects
| Object | Stage | Responsibility |
|---|---|---|
Data |
Explore | variable analysis, bin detail, single/cross/tree rule mining |
Scorecard |
Model | fit logistic regression on WOE of (custom or auto) bins → reliable scorecard + score rules |
Strategy |
Decide / Monitor | evaluate candidate rules' business impact; evaluate live-strategy effectiveness |
RuleSet |
Exchange | a serializable/filterable/mergeable set of rules passed between objects |
They connect linearly: Data (explore) → Scorecard (model) → Strategy (decide) → monitor,
bridged by to_scorecard() / to_strategy(). Pipeline provides a one-shot batch entry.
Key terms (plain language)
| Term | One-line meaning | How to read |
|---|---|---|
| good / bad | target=1 = bad (overdue/default), =0 = good. Bad rate = share of bad |
lower is better |
| IV | a variable's overall power to separate good from bad | >0.02 usable, 0.1–0.3 medium, >0.3 strong |
| KS / AUC | how far a variable/model pulls good and bad apart | higher is better (KS>0.3, AUC>0.7 good) |
| WOE | a bin's "bad concentration" vs the whole | positive = bin skews bad; ideally monotonic in the variable |
| lift | bad rate of the hit crowd ÷ overall bad rate | lift=3 → 3× the base bad rate; higher = more precise rule |
| coverage / recall | how many people a rule hits / how many bads it catches | very low coverage rarely matters |
| PSI | distribution drift of a variable/score between two periods | <0.1 stable, 0.1–0.25 watch, >0.25 drifting |
| OOT (out-of-time) | a later time slice used to check whether a rule/model still holds "in the future" | train vs OOT metrics should be close |
| cutoff | a score/value threshold; crossing it means reject | direction='<=' = below threshold = high risk = reject |
Capability overview
RuleLift
├── Rule intelligence — evaluate live rules without A/B testing
├── Auto rule mining — single-feature / cross-feature / tree-model
├── Deep variable analysis— IV / KS / AUC / PSI in one pass
├── Scorecard modeling — binning → WOE → logistic regression → reliability guards
├── Strategy & monitoring — simulate / compare / adverse-action reason codes / live monitor
└── One-shot pipeline — automated end-to-end mining
Want a runnable full example right away? Run
python rulelift/examples/full_guide.py— it runs every feature in this doc with assertions and prints real output; its comments are a line-by-line tutorial.
2. Install, import & namespace
pip install rulelift
For air-gapped environments see 12. Offline install.
v2.10 top-level namespace
After import rulelift as rl, rl.<Tab> surfaces only the tasks and first-class types you call
directly (~26):
| Category | Members |
|---|---|
| Top-level functions (run once) | profile / inspect / mine / cross / evaluate / compare / monitor / scorecard |
| Context classes | Data (explore) / Strategy (decide) / Scorecard (model) / Pipeline (one-shot) |
| Rule model | Rule / RuleSet / Condition / parse_rule_string |
| Result objects | ProfileResult / BinDetail / EvalResult / ImpactTable / CompareResult / CrossResult / ScorecardResult |
| Other | RuleLiftConfig / load_example_data / evaluate_rule_on_mask |
Internal parts (individual miner classes, VariableAnalyzer, low-level plot_*, config normalizers,
nested configs…) were moved into submodules. They are still reachable two ways:
from rulelift.mining import SingleFeatureRuleMiner # recommended: submodule import
rl.SingleFeatureRuleMiner # compatible: lazy-forwarded + one DeprecationWarning
Two usage styles: top-level functions vs the Data context
The same feature has two equivalent forms; pick by scenario.
① Top-level functions — one line per feature (fastest start). Best for "run once" or scripts: pass data, target column, and other params as keywords; no object needed.
import rulelift as rl
df = rl.load_example_data('feas_target')
rl.profile(df, target='ISBAD') # variable check
rl.mine(df, target='ISBAD', method='single', top_n=10) # mine rules (method only here)
rl.evaluate('ALI_FQZSCORE <= 700', df, target='ISBAD') # evaluate one rule
rl.scorecard(df, target='ISBAD', custom_bins={'ALI_FQZSCORE': [560, 700]}) # build a scorecard
The eight top-level functions mirror their context methods exactly (name, params, return):
| Top-level function | Equivalent to | What it does |
|---|---|---|
rl.profile |
Data(...).profile() |
variable check (IV/KS/missing…) |
rl.inspect |
Data(...).inspect() |
per-bin detail of a variable |
rl.mine |
Data(...).mine() |
mine rules (single/cross/tree) |
rl.cross |
Data(...).cross() |
two-feature cross matrix |
rl.evaluate |
Strategy(...).evaluate() |
evaluate a rule's business effect |
rl.compare |
challenger.compare(champion) |
champion / challenger contrast |
rl.monitor |
Strategy(...).monitor() |
live-strategy effectiveness |
rl.scorecard |
Data(...).to_scorecard() |
scorecard modeling |
② Data context — reuse one dataset across many calls (recommended for iteration). Declare column
roles (target/date/amount…) once; then profile / inspect / mine / cross / to_scorecard / to_strategy
all reuse them, no repeated params:
data = rl.Data(df, target='ISBAD', date_col='CREATE_TIME',
amount_col='AMOUNT', exclude_cols=['ID']) # declare once
data.profile() # call repeatedly; roles are reused automatically
data.inspect('ALI_FQZSCORE')
data.mine('single', top_n=10)
sc = data.to_scorecard(custom_bins={'ALI_FQZSCORE': [560, 700]}) # bridge straight to a scorecard
In one line: run once → top-level function; iterate / chain the whole flow → build
Datafirst.
Bundled example data
The main examples use datasets shipped with the package — no external files, fully reproducible:
import rulelift as rl
df = rl.load_example_data('feas_target') # 499×8 feature table, bad rate 0.307
hit = rl.load_example_data('hit_rule_info') # 998-row rule-hit log, for monitor()
feas_target roles: target=ISBAD, date=CREATE_TIME, amount=AMOUNT, ovd_bal=OVD_BAL; key
features ALI_FQZSCORE (IV 0.936), BAIDU_FQZSCORE, `NUMBER OF LOAN APPLICATIONS TO PBOC`
(name has spaces → wrap in backticks inside rule expressions).
3. Full-workflow guide (the core chapter)
This chapter walks the full path from raw data to a deployable strategy on feas_target. Every
step shows reproducible output; run it top to bottom and you know the whole toolkit.
3.0 The big picture
┌────────────────────────────────────────────────────────────┐
raw df ───▶│ Explore Data │
│ profile() batch check → find strong variables │
│ inspect() per-bin detail → set cut points, check mono │
│ mine() mine candidate rules (single / cross / tree) │
└───────────────┬───────────────────────────┬────────────────┘
│ to_scorecard() │ to_strategy() / RuleSet
▼ ▼
┌────────────────────────┐ ┌───────────────────────────────┐
│ Model Scorecard │ │ Decide Strategy │
│ fit → report │ │ evaluate one/many rules │
│ export / export_bins │ │ simulate before/after │
│ predict / score │ │ compare champion vs chall │
│ to_strategy() ────────┼──▶│ reason_codes │
└────────────────────────┘ └───────────────┬───────────────┘
│ package() / save()
▼
┌────────────────────────────────────────────────────────────┐
│ Monitor Strategy.monitor() live effectiveness (est vs actual)│
└────────────────────────────────────────────────────────────┘
Three entry styles, pick one (same name/params/result — see §2 two usage styles):
- Context objects (iterate): methods of
Data / Scorecard / Strategy; - Top-level functions (run once):
rl.profile / rl.mine / rl.evaluate / ...; - One-shot workflow (batch/report):
Pipeline(...).run_all().
Below we use the context-object style (declaring roles once makes the whole flow cleanest).
3.1 Step 0: prepare data, declare column roles
Column roles are declared once when building Data; later methods need not repeat them:
import rulelift as rl
df = rl.load_example_data('feas_target') # 499×8, bad rate 0.307
data = rl.Data(
df,
target='ISBAD', # binary label (1 = bad)
date_col='CREATE_TIME', # time column (for OOT split and PSI)
oot_split_date='2024-09-01', # rows after this date become out-of-time (OOT)
amount_col='AMOUNT', # loan amount (for amount-based false-reject / interception)
ovd_bal_col='OVD_BAL', # overdue balance (optional, for loss metrics)
exclude_cols=['ID'], # non-feature columns to exclude
) # n_bins defaults to 10; the bin count affects IV ranking
Dataautomatically moves target/date/amount/ovd_bal out of the feature universe to prevent leakage, and copies the frame once on construction (narrow columns first for wide tables, see §4). Escape hatches:data.analyzer(underlyingVariableAnalyzer),data.miner(...)(underlying miner).
3.2 Explore I: profile — batch variable check
What this does: a "health report" per variable (power IV, KS, missing rate, stability PSI…), sorted by IV. The first thing to do after loading data — see which variables are useful.
prof = data.profile() # returns ProfileResult, sorted by IV
prof.top(5, by='iv') # returns a DataFrame
Key parameters (shared by data.profile(...) and top-level rl.profile(df, target=..., ...)):
| Param | Default | Meaning |
|---|---|---|
features |
None = all |
check only these columns. Narrow the subset first on wide tables, else per-feature binning is slow |
parallel |
'auto' |
multiprocessing ('auto' decides by feature count / True / False) |
include_exact_ks |
False |
True also computes exact KS (slower); default uses approximate KS |
batch_size |
50 |
features per batch when parallel |
profile() computes per variable: iv / ks / auc / gini / corr_with_target / psi (if OOT declared)
plus count / nunique / unique_rate / mean / std / min / max / median / missing_rate / single_value_rate. On feas_target the top-IV variables:
| variable | iv |
|---|---|
| ALI_FQZSCORE | 0.936 |
| NUMBER OF LOAN APPLICATIONS TO PBOC | 0.834 |
| BAIDU_FQZSCORE | 0.460 |
The result follows the uniform protocol — filter and export directly:
prof.query('iv > 0.1 and missing_rate < 0.2') # boolean filter (returns a DataFrame)
prof.to_excel('variables.xlsx') # export
prof['iv'] # direct column access
Metric interpretation cheat sheet (rule of thumb):
| Metric | Strong | Medium | Weak / unstable |
|---|---|---|---|
| IV | >0.3 | 0.1–0.3 | <0.1 (input threshold ~0.02) |
| KS | >0.3 | 0.2–0.3 | <0.2 |
| AUC | >0.7 | — | — |
| PSI | <0.1 stable | 0.1–0.25 watch | >0.25 unstable |
3.3 Explore II: inspect — bin detail & tuning cut points
What this does: cut a variable into bins and read bad rate, share, and WOE per bin — to judge monotonicity (bad rate moving one direction with the variable) and whether your cut points are good. This is the key prep before building a scorecard.
Key parameters:
| Param | Default | Meaning |
|---|---|---|
features |
required | variable name(s) (string or list) |
custom_bins |
None |
custom cut points. Numeric: {'f':[560,700]}; categorical groups: {'f':[['a','b'],['c']]}. Else auto-bins by n_bins |
n_bins |
global config | number of auto bins when no custom_bins. Defaults to equal-frequency when set ("as many bins as you asked"); columns with heavy duplicates produce the max bins the data allows |
binning_method |
auto | 'quantile' (equal-frequency) / 'chi2'. Auto = equal-frequency when n_bins is set; pass 'chi2' to make n_bins only an upper bound (bins get merged) |
plot |
True |
draw "bin-share bars + per-bin bad-rate line"; pass False for table only |
det = data.inspect('ALI_FQZSCORE', custom_bins={'ALI_FQZSCORE': [560, 700]}, plot=False)
det.to_frame()[['bin_range', 'total', 'bad', 'badrate', 'woe']]
bin_range total bad badrate woe
(-inf, 560] 5 4 0.800 2.202
(560, 700] 92 57 0.620 1.304
(700, inf] 402 92 0.229 -0.399
WOE 2.202 → 1.304 → −0.399 is monotonically decreasing and bad rate 0.80 → 0.62 → 0.23 separates
cleanly — good cut points. If an adjacent bin's WOE "turns back" (non-monotonic), adjust and re-check.
inspect shares the same chi2 core as the scorecard, so the bins you see here are the bins the
scorecard will use — WYSIWYG.
3.4 Mine: single / cross / tree
What this does: automatically "find thresholds that isolate high-risk crowds", producing a batch
of candidate rules (with lift, coverage, statistical significance). One mine, switch via method:
method |
mines what | when to use | own params |
|---|---|---|---|
'single' |
single-variable cut, e.g. score <= 500 |
most common, most interpretable | by_feature (top_n per feature) |
'cross' |
two-variable combo, e.g. A low AND B low |
variables weak alone, strong combined | top_features (pool size), max_pairs |
'tree' |
multi-variable tree paths (conditions AND-ed) | non-linear, multi-factor combos | algorithm (see below) |
method='tree' algorithm options:
| algorithm | meaning | note |
|---|---|---|
'dt' |
single decision tree | most interpretable, good for starting |
'rf' |
random forest | multi-tree vote, more stable (default) |
'gbdt' |
gradient boosting | stronger fit |
'chi2' |
chi-square tree | splits by chi-square |
'isf' |
isolation forest | anomaly-detection style |
Shared params: top_n (take top N), min_lift (keep lift ≥), sort_by, features,
custom_bins, min_hit_count (default 30, hit-count gate), min_coverage (default 1%, coverage
gate) — the last two decide the result's passes_gate.
# ① single-feature: search the best cut per feature, top_n each
rs = data.mine('single', features=['ALI_FQZSCORE', 'BAIDU_FQZSCORE',
'NUMBER OF LOAN APPLICATIONS TO PBOC'],
top_n=3, by_feature=True)
rs.to_frame().sort_values('lift', ascending=False).head()
# ② two-feature cross
rc = data.mine('cross', features=[...], top_n=5, max_pairs=6)
# ③ decision-tree paths (algorithm dt/rf/gbdt/chi2/isf)
rt = data.mine('tree', algorithm='dt', top_n=5)
Strongest single rule on feas_target:
rule lift badrate coverage lift_ci_lower passes_gate
(`NUMBER OF LOAN APPLICATIONS TO PBOC` >= 12) 2.718 0.833 0.124 ... True
mine returns a RuleSet; each rule carries rule_id / rule_description / features / lift / badrate / coverage / recall / hit_bad_rate / n_samples / hit_count / lift_ci_lower / passes_gate / passes_stat / source. Statistical gating is explicit:
strong = rs.query('passes_gate and lift_ci_lower > 1 and n_samples >= 30')
top1_expr = rs.top(1, by='lift')[0].to_expr() # parseable expression of the strongest rule
Cross matrix (heat map of the bad rate for two-variable combos)
mine('cross') mines cross rules; rl.cross / data.cross produce a cross matrix — bin
each variable, then read bad rate/lift/count per cell, to see "which cell is most dangerous".
res = rl.cross(df, features=['ALI_FQZSCORE', 'BAIDU_FQZSCORE'], target='ISBAD',
n_bins=3, min_samples=20)
res.to_frame() # tidy long table: one row per cell
res.matrix('lift') # wide matrix for a single pair
res.plot('badrate') # heat map
Key parameters:
| Param | Default | Meaning |
|---|---|---|
features |
— | ≥2 features → all C(n,2) pairs; or feature1 + feature2 for a single pair |
n_bins |
5 |
bins per variable (variables with fewer uniques stay unbinned) |
binning_method |
'quantile' |
'quantile' (equal-freq) / 'chi2' (bin edges fit the bad rate) |
min_samples |
0 |
suppress low-count cells: cells below it get NaN rates (avoids misleading small-sample lift) |
metrics |
None = all |
keep a subset, e.g. ('badrate','lift') |
save_path |
None |
dump to Excel (long table + one matrix sheet per pair/metric) |
3.5 Model: Scorecard (with bin/WOE export)
What this does: turn a few strong variables into a "score table" — bin each variable, compute WOE, fit logistic regression, and scale to a score (higher = safer). Produces readable score rules plus six reliability guards (monotonicity / OOT / PSI / coefficient sign / collinearity / sample size) that raise explicit warnings rather than passing silently.
A scorecard is one straight line: custom/auto bins → WOE → logistic regression → scaling → six
guards. Recommended via the Data bridge (inherits OOT split & column roles, prevents leakage):
Key parameters (data.to_scorecard(...) / rl.scorecard(...) / rl.Scorecard(...)):
| Param | Default | Meaning |
|---|---|---|
custom_bins |
None |
custom cut points (tune with inspect first, then pass as-is); else auto-bin |
select |
True |
auto-select features (by IV / collinearity / sign) |
iv_threshold |
0.02 |
min IV to enter the model when select=True |
corr_threshold |
0.7 |
collinearity cutoff (drop one of a too-correlated pair) |
max_features |
None |
cap on features entering the model |
monotonic |
True |
monotonize WOE of auto bins (custom cut points kept as-is) |
enforce_sign |
True |
force coefficient sign to match WOE (kills uninterpretable "wrong-sign" features) |
pdo/rate/base_odds/base_score |
60/2/35/750 |
score scaling, defaults aligned with toad |
oot |
None |
pass an explicit OOT DataFrame (when not going through Data) |
sc = data.to_scorecard(
custom_bins={'ALI_FQZSCORE': [560, 700], # cut points tuned via inspect, passed as-is
'BAIDU_FQZSCORE': [400, 470]},
select=True, monotonic=True, pdo=60, base_score=750,
)
Reliability report — train/OOT KS·AUC + score PSI + six guards, warns explicitly if any fails:
rep = sc.report()
# feas_target: 3 features in model; train KS=0.5533 AUC=0.847;
# OOT KS=0.0968 PSI=0.2884; verdict='warn' (overfit/PSI warnings on OOT)
Export score rules — four forms:
print(sc.export(readable=True)) # ① readable IF-THEN
# base = 481.87 (base_placement=spread)
# ALI_FQZSCORE:
# IF ALI_FQZSCORE in (-inf, 560] -> +12.02
# IF ALI_FQZSCORE in (560, 700] -> +10.59
# IF ALI_FQZSCORE in (700, inf] -> +210.60
# ...
sc.export(to_frame=True) # ② ScorecardResult table (feature/bin/bin_range/woe/coef/points)
sc.export(to_dict=True) # ③ score-point dict {feature: {bin_label: points}}
sc.export(to_json='card.json') # ④ pure JSON (no pickle)
Score / probability / per-record reason:
sc.predict({'ALI_FQZSCORE': 600, 'BAIDU_FQZSCORE': 420,
'NUMBER OF LOAN APPLICATIONS TO PBOC': 5}) # -> 366.0 (higher = safer)
sc.predict_proba(df) # -> P(bad), inverse of score via scaling
sc.score(df) # a copy with a 'SCORECARD' column added
Export the auto-binning / WOE process (new in v2.10)
The scorecard bins and computes WOE internally; that binning process can be exported — to audit
"how it binned and each bin's WOE/bad rate", and to feed the auto bins back into custom_bins for
tweaking.
# ① bin result dict — same shape as custom_bins input, re-feedable
sc.bins_
# {'ALI_FQZSCORE': [560.0, 700.0],
# 'BAIDU_FQZSCORE': [400.0, 470.0],
# 'NUMBER OF LOAN APPLICATIONS TO PBOC': [7.0, 10.0]} # numeric: inner cut points (no ±inf)
# categorical: {'feat': [['A','B'], ['C']]}
# re-feed: start from auto bins, tweak by hand, rebuild
bins = sc.bins_
bins['NUMBER OF LOAN APPLICATIONS TO PBOC'] = [5, 12]
sc2 = rl.Scorecard(target='ISBAD', custom_bins=bins).fit(df)
# ② binning-process long table (per-bin counts + WOE/IV, for audit)
sc.export_bins(to_frame=True).to_frame()
# feature bin_range count badrate woe iv
# ALI_FQZSCORE (-inf, 560] 4 1.000000 1.8345 1.0435
# ALI_FQZSCORE (560, 700] 71 0.802817 1.8522 1.0435
# BAIDU_FQZSCORE (-inf, 400] 95 0.600000 0.8537 0.5609
# ... missing goes to its own bin labeled 'Missing'; per-feature counts sum to the fit sample size
# ③ dict / JSON
sc.export_bins(to_dict=True) # same as sc.bins_
sc.export_bins(to_json='bins.json') # pure JSON
sc.export_bins(to_frame=True, all_features=True) # default: model features only; all_features = every candidate
sc.export(to_dict=True)exports the score-point table (for scoring/review);sc.export_bins(...)exports the cut points & WOE process (for audit and re-feeding). They complement each other.
3.6 Decide: Strategy — evaluate / simulate / compare / reason codes
A strategy = baseline + candidate rules + cutoffs, combined with OR (hit any ⇒ reject).
strat = (rl.Strategy(baseline='ALI_FQZSCORE <= 500', target='ISBAD', amount_col='AMOUNT')
.add('BAIDU_FQZSCORE <= 400')) # add a candidate rule (also .add(ruleset))
strat.simulate(df).to_frame() # before/after business impact
scenario approval_rate approved_bad_rate reject_rate bad_intercepted recall_intercepted
baseline 1.000 0.307 0.000 0.0 0.000
proposed 0.711 0.251 0.289 64.0 0.418
marginal(proposed-baseline) 0.711 0.251 0.289 64.0 0.418
Reading: adding BAIDU_FQZSCORE <= 400 drops approval to 71.1%, cuts the approved crowd's bad rate
30.7% → 25.1%, intercepts 64 more bads (recall 41.8%). simulate columns: scenario / n_total / approval_rate / reject_rate / approved_bad_rate / rejected_bad_rate / bad_intercepted / recall_intercepted / false_reject / reject_amount_ratio.
ev = strat.evaluate('ALI_FQZSCORE <= 700', df) # single-rule offline eval → EvalResult
ev.overall # reject_rate / lift / lift_ci_lower / passes_gate ...
champ = rl.Strategy(baseline='ALI_FQZSCORE <= 500', target='ISBAD')
chal = champ.copy().add('BAIDU_FQZSCORE <= 400')
chal.compare(champ, df) # CompareResult: two strategies side by side
chal.reason_codes(df.head(100), top=2, lang='en') # per-record reject reasons (adverse action)
strat.add_cutoff('ALI_FQZSCORE', 450, direction='<=') # score cutoff (low value = high risk = reject)
3.7 Monitor: live-rule effectiveness
After go-live, rejected customers have no observed performance, so you cannot directly compute a bad
rate. monitor contrasts a rating-based estimated bad rate with the (partially observable)
actual bad rate to gauge real effectiveness. It consumes a rule-hit log (each row: user ×
hit rule × hit date × rating / rating bad rate):
hit = rl.load_example_data('hit_rule_info') # 998-row example hit log
res = rl.Strategy(target='USER_TARGET').monitor(
hit, rule_col='RULE', user_id_col='USER_ID', date_col='HIT_DATE',
user_level_badrate_col='USER_LEVEL_BADRATE')
res.to_frame() # per rule: estimated_badrate_pred / estimated_lift_pred / hit_bad_rate / lift /
# passes_gate / hit-rate drift hit_rate_change_rate ...
res.correlation # rule-to-rule hit correlation matrix
res.strategy_gain # pairwise strategy-combination gain matrix
3.8 One-shot: Pipeline
When you don't want to step through manually, Pipeline.run_all() does "variable analysis → grouping
→ three mining types → statistical checks → rule portfolio → (optional) report" in one call, and can
bridge to a strategy:
pipe = rl.Pipeline(df, target_col='ISBAD', date_col='CREATE_TIME',
amount_col='AMOUNT', preset='balanced') # fast / balanced / thorough
out = pipe.run_all(methods=('single', 'cross', 'tree'))
out['ruleset'] # unified RuleSet
imp = pipe.to_strategy(baseline='ALI_FQZSCORE <= 400').simulate(df) # bridge
run_all(oot=oot_df)does OOT validation at the rule layer only when you pass an explicit OOT frame;oot_split_dateonly drives profile PSI. Categorical grouping columns (e.g. frompd.qcut) must go intoexclude_cols, else mining crashes withCannot setitem on a Categorical.
3.9 Ship it: save / load / package
strat.save('strategy.json'); s2 = rl.Strategy.load('strategy.json') # pure JSON
pkg = strat.package(market='PK', version='v2.10') # deploy bundle (rules/cutoffs/meta)
sc.save('scorecard.json'); rl.Scorecard.load('scorecard.json') # scorecard persistence (versioned)
3.10 Strategy tuning: A-type & D-type
Two complementary directions for tuning a live strategy:
- A-type (Ascending — raise approval): without letting the bad rate rise, let more good
customers through — essentially "loosen rules". Analyze the cumulative interception curve
(
from rulelift.metrics import add_cumulative_metrics): once the remaining crowd's bad rate drops below the overall, the "tail" rules stop being worth it — loosen/remove them. - D-type (Descending — cut the bad rate): add effective new rules to intercept more bads —
essentially "add rules". Use mining to find fresh high-risk signals, then confirm their marginal
value with
Strategy.simulate()(themarginalrow) orcompare()before shipping.
actual bad actual good
strategy rejects reject bad ✓ false reject ✗
strategy approves let bad pass ✗ approve good ✓
↑ ↑
D-type: fewer A-type: fewer
passed bads false rejects
Usual rhythm: first D-type to push the bad rate to target, then A-type to lift approval without worsening it. Always confirm rules generalize (train vs OOT lift decay) before loosening.
4. Using it on your own big data (wide tables)
Two things save a lot of pain on production data.
① Narrow columns before building Data on very wide tables. Data/Pipeline copy the frame on
construction; thousands of columns can OOM. Right way: read, immediately pick "role columns +
candidate features", del the original, gc, then continue with the small table:
import pandas as pd, gc
df = pd.read_pickle('your_big_data.pkl') # peak memory here
roles = ['label', 'repay_datetime', 'amount']
feats = [c for c in df.select_dtypes('number').columns
if c not in roles and df[c].notna().mean() > 0.6][:100]
sub = df[roles + feats].copy()
del df; gc.collect()
data = rl.Data(sub, target='label', date_col='repay_datetime', amount_col='amount')
You can also run data.profile(features=...) on small feature batches to screen high-IV ones first.
② Check the label first. Make sure the target is a clean 0/1: drop "not-yet-matured / unknown" samples (sometimes coded 0.5, -1…) or the bad rate and every metric will be distorted. If the time column is only populated for part of the sample (e.g. only bad customers have a repay date), don't split OOT by date — use a random split.
The simplest way to verify the toolkit on your data: swap
examples/full_guide.py's dataset for your narrowed subset, adjust the column roles, and run it — everything should stay green.
5. Data: the exploration context
data = rl.Data(df, target, *, date_col=None, oot_split_date=None, amount_col=None,
ovd_bal_col=None, exclude_cols=None, include_cols=None, n_bins=10,
config=None, **kwargs)
| Method | Purpose | Returns |
|---|---|---|
profile(features=None, ...) |
batch variable check (IV/KS/AUC/PSI/missing…) | ProfileResult |
inspect(features, *, custom_bins=None, n_bins=None, plot=True) |
per-bin detail + plot | BinDetail |
mine(method='single', *, top_n, by_feature, features, algorithm, max_pairs, ...) |
three mining types | RuleSet |
cross(features, *, n_bins, min_samples) |
two-feature cross matrix | CrossResult |
to_scorecard(**kw) |
bridge to scorecard | Scorecard |
to_strategy(baseline=None) |
bridge to strategy | Strategy |
Escape hatches: data.analyzer (VariableAnalyzer, full-precision analysis + select_features()
with collinearity removal), data.miner(method=...) (underlying miner).
6. Scorecard: modeling in depth
Three equivalent entries:
sc = data.to_scorecard(custom_bins=..., select=True, monotonic=True, pdo=60, base_score=750) # ① Data bridge (recommended)
sc = rl.scorecard(df, target='ISBAD', date_col=..., oot_split_date=..., custom_bins=...) # ② top-level, one shot
sc = rl.Scorecard(target='ISBAD', custom_bins=..., select=True).fit(df, oot=oot_df) # ③ construct + fit
Six reliability guards (all in report(), warn on failure): monotonic WOE, two-segment OOT
check, score PSI, coefficient sign, collinearity removal, sample sufficiency. Scaling defaults align
with toad (pdo=60, rate=2, base_odds=35, base_score=750); with bin_closed='left' + base_placement='spread' the points match toad.ScoreCard.export digit-for-digit.
Export interfaces:
| Interface | Output |
|---|---|
sc.export(readable=True) |
readable IF-THEN text |
sc.export(to_frame=True) |
ScorecardResult table (feature/bin/bin_range/woe/coef/points) |
sc.export(to_dict=True) |
score-point dict {feature: {bin_label: points}} |
sc.export(to_json='card.json') |
pure JSON scorecard |
sc.bins_ / sc.export_bins(...) |
bin/WOE process: custom_bins-shaped dict (re-feedable) / long table / JSON (§3.5) |
sc.predict(X) / predict_proba(X) / score(df) |
score / probability / add score column |
sc.to_strategy(cutoff=, direction=) |
back to the strategy loop |
7. Strategy: the decision context
strat = rl.Strategy(baseline=None, target=..., *, amount_col=None, ovd_bal_col=None, config=None)
| Method | Purpose |
|---|---|
add(rule_or_ruleset) / add_cutoff(feature, threshold, direction='<=') |
add candidate rule / cutoff |
evaluate(rule(s), data, *, schema='business', by=None) |
offline eval of one/many rules → EvalResult |
simulate(data) |
before/after business impact → ImpactTable |
compare(other, data) |
champion vs challenger → CompareResult |
reason_codes(data, *, top=3, lang='en') |
per-record reject reasons → DataFrame |
monitor(hit_log, *, rule_col, user_id_col, date_col, user_level_badrate_col, ...) |
live effectiveness → EvalResult |
copy() / save(path) / load(path) / package(market=, version=) |
copy / persist / deploy bundle |
schema='business' (default, business column names) vs schema='ml' (modeling names). Multi-rule
evaluate uses the same business names as single-rule (reject_rate / lift / false_reject_rate
/ lift_ci_lower / passes_gate).
8. RuleSet / Rule
rs = data.mine('single', features=[...], top_n=5)
rs.to_frame() # all rules + metrics (DataFrame passthrough)
rs.query('lift_ci_lower > 1 and n_samples >= 30') # statistical filter (RuleSet.query returns RuleSet)
rs.top(3, by='lift') # top 3 (returns RuleSet)
rs[0].to_expr() # single rule → parseable expression '(feat >= x)'
rs.to_json() / RuleSet.from_json(js) # serialize
rs.to_excel('rules.xlsx')
rs_a + rs_b # merge; dedup available
Rule.to_mask(df) returns all-False for missing columns (no raise — intentional); to_expr() output
is parenthesized and backtick-quotes spaced column names. Parse rule strings with
rl.parse_rule_string("NUMBER OF LOAN >= 12 and ALI_FQZSCORE <= 700") (supports and / or / comparisons / in).
9. The uniform result-object protocol
All result objects (ProfileResult / BinDetail / EvalResult / ImpactTable / CompareResult / CrossResult / ScorecardResult) share one protocol and behave like a DataFrame:
| Method | Purpose |
|---|---|
.to_frame() |
the underlying DataFrame |
.top(n, by=) |
top n by a column (returns a DataFrame) |
.query(expr) |
boolean filter (on the underlying DataFrame, returns a DataFrame) |
.to_excel(path) |
export to Excel |
.plot(...) |
visualize |
res['col'] / res.col / for r in res |
DataFrame passthrough (since v2.8, no .to_frame() needed) |
prof['iv'] # direct column access
prof.query('iv > 0.1') # boolean filter → DataFrame
prof.top(5, by='ks')['variable'] # top 5 (DataFrame), then chain with pandas
Note:
query/ passthrough delegate to the underlying DataFrame and return a DataFrame (not a result object), so don't chain.top()/.to_frame()after them;RuleSet.query()is the exception — it returns aRuleSet(keeps rule semantics, can.top()).
10. Pipeline / top-level functions / Agent
Pipeline (one-shot six steps): analyze → select → mine → validate → portfolio → report, see
§3.8. preset='fast'|'balanced'|'thorough'; iv_ranges controls the
per-mining-type feature-selection range (default None = all).
Top-level functions (8, one line per feature): rl.profile / inspect / mine / cross / evaluate / compare / monitor / scorecard — full table & usage in
§2 two usage styles.
Agent / MCP (LLM integration surface, 4 JSON-schema-validated tools):
from rulelift.agent import register_dataset, get_tool_schemas, run_tool
ref = register_dataset(df, 'feas')
run_tool('mine', dataset_ref=ref, target='ISBAD', method='single', top_n=5) # {summary, data}
Tools: profile / mine / evaluate / simulate, output contract {summary, data}. Self-implemented
stdio JSON-RPC MCP server: python -m rulelift.agent.mcp.
11. Performance & capacity (big data)
- Narrow columns first: thousands of columns straight into
Data/Pipelinewill OOM (the constructor copies the frame). Pick role + candidate columns first, see §4. - Faster profile: on wide tables,
profileuses incremental recompute + alokyprocess pool — 10k+ columns can drop from hours to minutes, values unchanged. - Truncate:
Pipeline'sselect_max_featurescaps features by IV;include_colsis a whitelist. - Parallel:
parallel='auto'decides serial/parallel by data size; job resolution viautils.resolve_n_jobs. - Memory engineering (from the library's design): per-batch
gc.collect(),np.digitizeinstead ofpd.cut, binning-result caching, and real-time memory monitoring with auto-degradation to avoid OOM.
12. Offline install
For risk-control intranets without internet access.
Option A — download deps on a networked machine, install offline:
pip download rulelift -d ./packages/ # on a networked machine
# transfer ./packages to the offline host, then:
cd ./packages/ && pip install *.whl --no-index --find-links=.
Option B — run from source: download the source from GitHub, manually install pandas / numpy / scikit-learn / matplotlib / seaborn, then add the source path:
import sys
sys.path.append('/path/to/rulelift-master')
import rulelift as rl
License & info
MIT License.
- GitHub: https://github.com/aialgorithm/rulelift
- PyPI: https://pypi.org/project/rulelift/
- Issues / PRs welcome: https://github.com/aialgorithm/rulelift/issues
Runnable end-to-end example with assertions:
python rulelift/examples/full_guide.py. A methodology walkthrough (Chinese) that threads these features along the real strategy lifecycle:rulelift/examples/信贷风控策略全流程.md.
Metadata
Release files for rulelift 2.10.1
For a detailed explanation of source distributions (sdists) and built distributions (wheels), please see the package formats documentation.
Source distribution (sdist)
| File | Size | Uploaded | |
|---|---|---|---|
| rulelift-2.10.1.tar.gz | 383.4 kB | Details |
Built distribution (wheel)
| File | Interpreter | ABI | Platform | Reset |
|---|---|---|---|---|
| rulelift-2.10.1-py3-none-any.whl | Python 3 | none | any | Details |
Total release size: 673.8 kB
Release files / rulelift-2.10.1.tar.gz
| Download URL | rulelift-2.10.1.tar.gz |
|---|---|
| Size | 383.4 kB |
| Tags | Source |
|
SHA-256 checksum How to use checksums |
481cad3f855000aa7fbd9df058cef34f672b74d7f42654b1868429f7595d9f9c
|
|
BLAKE2b-256 checksum How to use checksums |
e4ed41ac59a85d05e0ad01da5eb32be60b81f3243df7badecca02e2f51d8b49e
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.18
|
Release files / rulelift-2.10.1-py3-none-any.whl
| Download URL | rulelift-2.10.1-py3-none-any.whl |
|---|---|
| Size | 290.4 kB |
| Tags | Python 3 |
|
SHA-256 checksum How to use checksums |
6e7c2ccdca4c8ab14ee84bfe8cd1a0cb276f27c8f0a4479c1e167a21af9afe3a
|
|
BLAKE2b-256 checksum How to use checksums |
867975e7bb280f50a4b6a7b60a9fd41f6858722fc4fd2b651446fafe5fae7625
|
| Upload date | |
|
Uploaded using Trusted Publishing? What is trusted publishing? |
No |
| Uploaded via |
twine/6.2.0 CPython/3.9.18
|