Feature selection and model monitoring toolkit for credit and risk modeling.
Project description
Model Track CR
Read this in other languages: English, Português
model-track-cr is a Python library designed to structure, standardize, and operationalize the full statistical and machine learning modeling workflow, with a strong focus on credit, risk, and supervised modeling use cases.
Rather than offering isolated utilities, the library provides a cohesive set of tools that cover the most critical stages of real-world modeling:
- variable diagnostics and exploration
- binning and categorization
- Weight of Evidence (WOE) and Information Value (IV)
- temporal stability analysis
- foundations for post-model monitoring
All components are designed to work together, with explicit APIs, reproducibility in mind, and strong test coverage.
🎯 Project Philosophy
The model-track-cr project was created to address a common problem in applied modeling:
modeling workflows are often fragmented across notebooks, ad-hoc scripts, and hard-to-reproduce logic.
This library is built around the following principles:
- Workflow-first: tools make sense primarily when used together
- Pandas-first: native integration with DataFrames
- Clear responsibilities: each module has a well-defined scope
- Reproducibility: modeling decisions are explicit and traceable
- Technical governance: stability and monitoring are first-class concerns
- Strict TDD: tests act as living documentation
🧭 High-Level Modeling Workflow
A typical modeling flow using this library follows these steps:
- Initial data diagnostics
- Binning / categorization
- WOE and IV computation
- Temporal stability analysis
- Post-model monitoring support
Each step is supported by a dedicated module that integrates naturally with the others.
🏗️ Project Architecture
model-track-cr is organized in pandas-first modules. Several classes inherit from BaseTransformer and expose fit / transform, but method signatures are DataFrame-oriented (for example, explicit column names and optional columns lists). That keeps the API explicit for modeling work; if you need a strict scikit-learn Pipeline, plan for thin wrapper steps that adapt arguments and return types.
classDiagram
namespace model_track_base {
class BaseTransformer {
<<abstract>>
+fit(df, target) *
+transform(df) *
}
}
namespace binning {
class TreeBinner {
+bins
+fit(df, column, target)
+transform(df, column)
}
}
namespace woe {
class WoeCalculator {
+mapping_: dict
+fit(df, target, columns)
+transform(df, columns)
}
class WoeStability {
+date_col: str
+calc: WoeCalculator
+calculate_stability_matrix(df, feature_col, target_col)
+generate_view(matrix, title, ax)
}
class CategoryMapper {
+mapping_dict_: dict
+auto_group(matrix, min_groups, is_ordered)
}
}
BaseTransformer <|-- TreeBinner
BaseTransformer <|-- WoeCalculator
WoeStability *-- WoeCalculator : composes (self.calc)
WoeStability ..> CategoryMapper : stability matrices
Key components
- BaseTransformer: Abstract
fit/transform/fit_transformcontract shared across several transformers. - TreeBinner: Supervised binning via
sklearn.tree.DecisionTreeClassifiersplit thresholds. - WoeCalculator: WoE mappings with Laplace smoothing;
fit/transformoperate on one or more categorical columns. - WoeStability (
model_track.woe): Per-period WoE matrices and plotting helpers; composes aWoeCalculator. - CategoryMapper (
model_track.woe): Grouping suggestions over stability matrices (inversions / SSE cost). - ProjectContext (
model_track): Serializable bag for bins, WoE maps, selected features, and free-form metadata (intended to grow with the library).
🧩 Core Modules Overview
📊 1. Diagnostics and early statistics (preprocessing, stats)
Use preprocessing for table-level audits and summaries, and stats for IV / association style metrics and selection helpers.
Example (variable audit summary):
from model_track.preprocessing import DataAuditor
auditor = DataAuditor(target="target")
summary = auditor.get_summary(df)
Example (IV-driven feature screening lives under stats):
from model_track.stats import StatisticalSelector
selector = StatisticalSelector()
selector.fit(df, target="target", features=["feat_a", "feat_b"])
df_selected = selector.transform(df)
This step typically informs:
- which variables to bin
- how to handle missing values
- potential stability risks
🪜 2. Binning & categorization (binning)
Responsible for turning continuous inputs into ordered categories for WoE and stability workflows.
Exported binners: supervised tree-based binning (TreeBinner), unsupervised quantile binning (QuantileBinner), and BinApplier to consistently apply saved bins to new datasets.
Example:
from model_track.binning import TreeBinner
binner = TreeBinner(max_depth=2)
binner.fit(df, column="income", target="target")
df["income_cat"] = binner.transform(df, column="income")
🧮 3. WoE and IV (woe, stats)
Implements classic supervised modeling metrics:
- WoE mappings and transformed columns via
WoeCalculator - IV and related helpers via
model_track.stats(for examplecompute_iv) - Per-period WoE paths via
WoeStability(see section 4)
Example:
from model_track.woe import WoeCalculator
calc = WoeCalculator()
calc.fit(df, target="target", columns=["income_cat"])
df_woe = calc.transform(df, columns=["income_cat"])
These outputs are typically used for:
- feature selection
- model interpretability
- direct input into linear models
📈 4. Temporal WoE stability (model_track.woe)
Tools to assess whether category-level WoE stays consistent over time (e.g. by vintage or calendar period).
Example:
from model_track.woe import WoeStability
ws = WoeStability(date_col="period")
matrix = ws.calculate_stability_matrix(
df=df,
feature_col="income_cat",
target_col="target",
)
ws.generate_view(matrix, title="Income_cat WoE stability")
This stage is essential for:
- validating model robustness
- supporting production decisions
- monitoring deployed models
🚀 Installation
Install in user mode:
pip install model-track-cr
# Or with heavy ML dependencies (LightGBM, etc.):
pip install "model-track-cr[tuning]"
The lib is available on https://pypi.org/project/model-track-cr/
Install in development mode:
git clone https://github.com/Cristiano2132/model-track-cr.git
cd model-track-cr
pip install -e .
# Or using Poetry:
poetry install
🧪 Testing and code quality
Run tests: make test
Run tests with coverage: make cov
HTML coverage report: htmlcov/index.html
The project follows strict Test-Driven Development (TDD) and is backed by a robust Testing Pyramid Strategy. Every new feature must be accompanied by automated tests.
For testing tiers (unit, statistical, integration, benchmarks), see the Testing Strategy Guide. Contributor environment notes live in AGENTS.md.
Workflow and contributions
This project uses Git Flow, GitHub Issues, and Pull Requests. Day-to-day commands, CI expectations, and local setup are summarized in AGENTS.md.
Documentation in this repository
Right now, the main narrative lives in this README (EN / PT-BR) plus the testing strategy linked above. Docstrings in src/model_track are the API reference. Deeper per-module guides and an end-to-end modeling walkthrough can be added as the surface area grows.
📓 Example Notebooks
| Notebook | Description |
|---|---|
notebooks/multiclass_example.ipynb |
End-to-end multiclass pipeline: Wine dataset → binning → MulticlassSelector → OvRWoeAdapter → LightGBM → MulticlassEvaluator → ProjectContext. |
When should you use this library?
This project is a good fit if you want to:
- standardize modeling workflows across teams
- move from fragile notebooks to reproducible pipelines
- evaluate stability before and after deployment
- improve technical governance of models
- document modeling decisions clearly
Technical roadmap
- Quantile-based binning and helpers to apply saved bins consistently across tables
- Automated PSI and additional drift metrics
- Feature selection informed by stability matrices
- Clearer optional adapters for scikit-learn
Pipelineusers - Richer monitoring and visualization utilities
License
MIT
Project details
Release history Release notifications | RSS feed
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file model_track_cr-1.1.0.tar.gz.
File metadata
- Download URL: model_track_cr-1.1.0.tar.gz
- Upload date:
- Size: 39.1 kB
- Tags: Source
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.4.0 CPython/3.13.13 Linux/6.17.0-1010-azure
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
4f235d230a3e4a773ca8981ecc57dc1f4b1155421611ea7c52ca5bb100ef4f57
|
|
| MD5 |
424fa4d61c48822e6bc22ec56d17b2e9
|
|
| BLAKE2b-256 |
4d8abed11a239ab0766288ec911861be42df6c2fd7e3babd5e30073f108b6b2d
|
File details
Details for the file model_track_cr-1.1.0-py3-none-any.whl.
File metadata
- Download URL: model_track_cr-1.1.0-py3-none-any.whl
- Upload date:
- Size: 52.4 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? No
- Uploaded via: poetry/2.4.0 CPython/3.13.13 Linux/6.17.0-1010-azure
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
6ec00da01115ee2d2d39fb76e542089066a1a2be1c8a0983164addb3038d4ba7
|
|
| MD5 |
047c6f5476973ea24ab029a4c525a4ef
|
|
| BLAKE2b-256 |
50be2902acca9f9611d4830a21a955ba41b34438dbdc7d7ebb88795fa3dc1fdc
|