Skip to main content

High-performance interpretable rule-based ML — HUG-IML classifier, adaptive binning, EBM-style plots, pattern pruning, and benchmark runner (IEEE Access 2024).

Project description

hugiml-core

High-performance interpretable rule-based ML infrastructure built on the HUG-IML algorithm published in IEEE Access (2024).

CI PyPI Docs Python License DOI

HUGIML: interpretable tabular ML through compact human-readable patterns

HUGIML learns human-readable High Utility Gain patterns and uses those patterns as the model representation itself. Instead of explaining a black-box after training, the learned model is already composed of inspectable intervals, categories, supports, utilities, and coefficients.

glucose=[157.1,177.3)                coef= +1.4077   support=0.067
bmi=[31.8,39.1)                      coef= +1.0839   support=0.200
duration=[24,48)                     coef= +0.84     support=0.28
checking_status=no_checking          coef= +1.12     support=0.39

Where HUGIML fits


Table of Contents

  1. What Is HUG-IML?
  2. Installation
  3. Quick Start
  4. Feature Modes
  5. Execution Modes
  6. Hyperparameter Search
  7. Governance Studio Dashboard
  8. LLM Assistant
  9. Augmented Pair Features
  10. Adaptive Binning
  11. Missing Value Handling
  12. Model Explanation and Visualisations
  13. Native Mining Pruning Controls
  14. Pattern Pruning
  15. Interpretability Metrics
  16. Multiclass, Imbalanced Data, High-Cardinality
  17. Drift Detection & Monitoring
  18. Calibration
  19. Serialisation
  20. Governance & Model Cards
  21. Benchmark Suite
  22. Validation Highlights
  23. Inference Server
  24. CI / CD
  25. Repository Structure
  26. License
  27. Citation

What Is HUG-IML?

The High Utility Gain Interpretable Machine Learning (HUG-IML) framework extracts High Utility Gain patterns from labelled tabular data, transforms the input into a binary pattern-presence matrix, and fits an interpretable downstream classifier (logistic regression by default) on that matrix.

The resulting patterns are human-readable and serve as the primary source of model explanations, making the system suitable for regulated domains such as credit scoring, healthcare, and risk management.

Key reference:

Krishnamoorthy, S. (2024). Interpretable Classifier Models for Decision Support Using High Utility Gain Patterns. IEEE Access, 12, 126088–126107. DOI: 10.1109/ACCESS.2024.3455563


Installation

# Core
pip install hugiml-core

# With profile plots
pip install "hugiml-core[plots]"

# With Streamlit UIs: Governance Studio dashboard and HUGIML LLM Assistant
pip install "hugiml-core[dashboard]"

# With benchmark comparison suite
pip install "hugiml-core[benchmarks]"

# With imbalanced-data helpers
pip install "hugiml-core[imbalanced]"

# With SHAP interoperability
pip install "hugiml-core[explainability]"

# With MLflow integration
pip install "hugiml-core[mlflow]"

# Everything
pip install "hugiml-core[all]"

Build from source requires a C++17 compiler. The recommended source-build path uses the build requirements declared in pyproject.toml so that pybind11 is installed before compilation:

git clone https://github.com/srikumar2050/hugiml-core.git
cd hugiml-core
python -m pip install -e ".[dev]"
python scripts/build_batched.py --inplace

The helper enables conservative native build batching by default (HUGIML_BUILD_BATCH_SIZE=4, HUGIML_BUILD_JOBS=2) to avoid memory spikes in constrained environments. Direct python setup.py build_ext --inplace still works when local build requirements are already installed. Avoid --no-build-isolation unless pybind11, setuptools, and wheel are already available in the active environment.

Recommended local validation uses deterministic pytest batches instead of one long pytest process:

python scripts/run_pytest_batches.py
ruff check .

make build, make test, make lint, and make validate provide the same default workflows when make is available.


Quick Start

HUGIMLClassifier is the primary public class name. HUGIMLClassifierNative remains available as a backward-compatible alias for existing code.

Note on prepareXy: prepareXy performs schema and type preparation only — it detects integer, float, and categorical columns and encodes the target. Discretisation, HUG pattern mining, and downstream classifier fitting occur inside fit() on the training data supplied to that call.

Path A — prepareXy

import pandas as pd
from sklearn.model_selection import train_test_split
from hugiml import HUGIMLClassifier

clf = HUGIMLClassifier(adaptive_binning=True, L=1, G=5e-3, topK=100)

X_enc, y_enc = clf.prepareXy(X_df, y)   # schema/type prep — no model fitting

X_tr, X_te, y_tr, y_te = train_test_split(
    X_enc, y_enc, stratify=y_enc, random_state=42
)

clf.fit(X_tr, y_tr)                     # mining + downstream fit on train only
proba = clf.predict_proba(X_te)

print(clf.get_hug_features())
print(clf.feature_importances())
print(clf.model_summary())

Path B — explicit allCols for CV and production pipelines

from hugiml import HUGIMLClassifier

clf = HUGIMLClassifier(
    allCols=[int_col_names, float_col_names, cat_col_names],
    origColumns=X.columns.tolist(),
    B=-1,
    adaptive_binning=True,
    b_candidates=[2, 3, 5, 7, 10, 15],
    L=1,
    G=1e-5,
    topK=150,
)

clf.fit(X_train, y_train)

pred = clf.predict(X_test)
proba = clf.predict_proba(X_test)

Feature Modes

HUGIML can use the mined binary pattern matrix in three downstream feature modes. The default remains pattern-only behavior, so existing code keeps the same high-interpretability semantics unless feature_mode is set explicitly.

feature_mode Downstream estimator input When to use
"patterns_only" HUGIML binary pattern matrix only Standard HUGIML; best when the mined pattern space itself captures the decision boundary.
"original_plus_patterns" Original features plus all mined binary patterns Useful when original features contain strong marginal signal and HUGIML patterns add supervised nonlinear refinements.
"original_plus_interactions" Original features plus only L > 1 mined patterns Useful when original features should handle marginal effects and HUGIML should contribute interaction/compound-region features only.

The recommended tuning grid and configuration choices are described in Hyperparameter Search. Start there for first-pass model selection, then select a representation based on interpretability and runtime needs.

from hugiml import HUGIMLClassifier

# Backward-compatible default: pattern matrix only
clf = HUGIMLClassifier(B=-1, L=2, G=1e-2, topK=150,
                              adaptive_binning=True, feature_mode="patterns_only")

# Hybrid: original features + all binary HUGIML patterns
clf_hybrid = HUGIMLClassifier(B=-1, L=2, G=1e-2, topK=150,
                                    adaptive_binning=True, feature_mode="original_plus_patterns")

# Hybrid: original features + higher-order/interaction patterns only
clf_interactions = HUGIMLClassifier(B=-1, L=2, G=1e-2, topK=150,
                                          adaptive_binning=True,
                                          feature_mode="original_plus_interactions")

transform(X) always returns the HUGIML binary pattern matrix, regardless of feature_mode. The feature mode only changes the matrix passed to the downstream estimator inside fit(), predict(), predict_proba(), and score().

For hybrid modes, HUGIML standardizes numeric original features internally before concatenating them with the sparse binary pattern matrix and any active augmented-pair columns. feature_importances(), model_summary(), and get_model_composition() report the downstream feature representation, while get_hug_features() and get_pattern_info() remain pattern-only APIs.


Downstream Solver Support

When base_estimator is not supplied, HUGIML now exposes the built-in downstream linear classifier choice through lr_solver. The default remains "auto", which preserves the historical behavior: binary problems use LogisticRegression(solver="liblinear"), and multiclass problems use LogisticRegression(solver="lbfgs").

lr_solver Downstream estimator When to use
"auto" LogisticRegression with the historical binary/multiclass solver choice Recommended default for most datasets and for backward-compatible results.
"saga" LogisticRegression(solver="saga") Useful for larger or sparse downstream matrices when you still want logistic-regression coefficients and probability estimates.
"sgd" SGDClassifier(loss="log_loss") Useful for very large downstream matrices where stochastic optimization can reduce memory pressure or wall-clock time. Validate accuracy because SGD can be more sensitive to scaling and convergence settings.

All built-in solver choices keep deterministic defaults aligned with the existing classifier path: random_state=0 and max_iter=500. If you need complete control over solver-specific hyperparameters, pass a fully configured base_estimator; that continues to override lr_solver.

from hugiml import HUGIMLClassifier

# Historical default
clf_default = HUGIMLClassifier(lr_solver="auto")

# LogisticRegression through saga
clf_saga = HUGIMLClassifier(lr_solver="saga", feature_mode="original_plus_patterns")

# Logistic loss through SGDClassifier
clf_sgd = HUGIMLClassifier(lr_solver="sgd", feature_mode="original_plus_patterns")

The versioned .hugiml serializer records lr_solver in the classifier initialization state and natively round-trips both LogisticRegression and the built-in SGDClassifier downstream estimator.

Execution Modes

HUGIML supports two execution modes:

execution_mode Purpose Behavior
"audit" Default mode for development, validation, governance, and regulated review Keeps the complete training and traceability artifacts needed by audit, governance, and dashboard APIs.
"production" Lean mode for deployment after validation Keeps prediction, probability scoring, save, and load behavior, while dropping training/audit-heavy artifacts to reduce retained memory.
from hugiml import HUGIMLClassifier

# Full traceability; this is the default.
audit_model = HUGIMLClassifier(execution_mode="audit")
audit_model.fit(X_train, y_train)

# Lean retained state for deployment.
prod_model = HUGIMLClassifier(execution_mode="production")
prod_model.fit(X_train, y_train)
prod_model.save_model("model.hugiml")
loaded = HUGIMLClassifier.load_model("model.hugiml")

In production mode, audit-oriented methods return a clear guidance result or raise a clear message asking you to refit with execution_mode="audit" when complete traceability is required.


Hyperparameter Search

HUGIML provides a fast cached tuning path for adaptive-binning grids. When adaptive_binning=True, the binning and transaction construction work is reused across eligible candidates, so compact grids can be evaluated without rebuilding the same mining inputs repeatedly.

Recommended named parameter grids

HUGIML tuning reads the recommended grids from hugiml.hyperparameter_configs. Use the default "performance" grid for a compact first pass, then switch to "interpretability" when the final representation should remain pattern-only.

from hugiml import HUGIMLClassifier

performance_grid = HUGIMLClassifier.default_param_grid()
interpretability_grid = HUGIMLClassifier.default_param_grid("interpretability")

# Equivalent default performance grid:
performance_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [1, 2],
    "topK": [50, 100],
    "feature_mode": ["original_plus_patterns"],
    "G": [0.01, 0.001],
}

# Equivalent interpretability grid:
interpretability_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [1, 2],
    "topK": [50, 100],
    "feature_mode": ["patterns_only"],
    "G": [0.01, 0.001],
    "interaction_relaxed_mining": [True],
    "augmented_pair_transforms": [False],
}
Grid Recommended use Main values
performance First-pass predictive tuning feature_mode=["original_plus_patterns"], L=[1,2], topK=[50,100], G=[0.01,0.001]
interpretability Pattern-only representation review feature_mode=["patterns_only"], interaction_relaxed_mining=True, augmented_pair_transforms=False

Both grids keep B=[-1] and adaptive_binning=[True], so each numerical feature chooses a supervised bin count. Do not enable interaction_relaxed_mining=True and augmented_pair_transforms=True in the same L >= 2 candidate.

Use focused follow-up grids when you want to explore interaction-relaxed mining or augmented-pair transforms.

tune() — cross-validated search with automatic fast path

result = HUGIMLClassifier.tune(
    X, y,
    param_grid="performance",
    cv=5,
    shuffle=True,
    random_state=42,
    scoring="roc_auc",
    refit=True,
)

print(result.best_params_)
print(f"CV score: {result.best_score_:.4f}")
print(f"Fast path used: {result.fast_path_used_}")

best_model = result.best_estimator_

A custom grid is supplied via param_grid. For the cached adaptive-binning path, keep the varying dimensions compact and centered on mining or representation choices such as G, L, topK, and feature_mode. Fixed values such as B=-1 and adaptive_binning=True may be included for clarity.

custom_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "G": [1e-2, 5e-3],
    "L": [1, 2],
    "topK": [50, 100],
    "feature_mode": ["patterns_only", "original_plus_patterns"],
}

result = HUGIMLClassifier.tune(
    X, y,
    param_grid=custom_grid,
    cv=3,
    scoring="roc_auc",
    refit=True,
)

Choosing the model configuration

After the default grid identifies a useful budget range, choose one of these focused configurations based on the representation you want.

Option Feature mode Interaction path Extra downstream pair columns? Interpretability Runtime profile Good default when...
Pure HUG patterns patterns_only Standard L=1 or L=2 mining No Very high Lowest to moderate You want the simplest pattern-only model.
Patterns + interaction-relaxed mining patterns_only interaction_relaxed_mining=True No Very high Higher than augmented pairs You want interaction evidence to affect HUG pattern discovery without adding a new feature family.
Patterns + augmented pairs patterns_only augmented_pair_transforms=True Yes High Often faster than relaxed mining You want selected pair evidence with better runtime control.
Originals + patterns original_plus_patterns Standard L=1 or L=2 mining No High Moderate Original variables have strong marginal signal and patterns add readable refinements.
Originals + patterns + relaxed mining original_plus_patterns interaction_relaxed_mining=True No High Higher than augmented pairs You want original features plus survivor-led HUG patterns, but no pair-operator columns.
Originals + patterns + augmented pairs original_plus_patterns augmented_pair_transforms=True Yes Moderate Moderate to higher You want the highest representation capacity among the recommended options.

A survivor is a source feature that remains after interaction-information screening. It may not be one of the strongest features by itself, but it has useful pairwise or synergy evidence with another feature. In interaction-relaxed mining, these survivor source features are allowed to participate in native HUG pattern mining. A survivor is not automatically a final model feature; it is a candidate source that can help form mined patterns.

interaction_relaxed_mining=True relaxes the usual entry path for interaction-useful source features. Instead of adding product, difference, or sum columns to the downstream estimator, it lets a small survivor pool enter the native mining step, so the final representation remains HUG patterns plus any original features selected by feature_mode.

Use these focused follow-up grids:

# Pattern-only with interaction-relaxed mining.
patterns_relaxed_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [2],
    "G": [1e-2, 5e-3],
    "topK": [50, 100],
    "feature_mode": ["patterns_only"],
    "augmented_pair_transforms": [False],
    "interaction_relaxed_mining": [True],
    "interaction_relaxed_feature_size": [8, 12],
}

# Pattern-only with augmented pair features.
patterns_augmented_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [2],
    "G": [1e-2, 5e-3],
    "topK": [50, 100],
    "feature_mode": ["patterns_only"],
    "augmented_pair_transforms": [True],
    "augmented_pair_mode": ["interaction_information"],
    "aug_feature_size": [8, 12],
}

# Originals plus patterns with interaction-relaxed mining.
originals_relaxed_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [2],
    "G": [1e-2, 5e-3],
    "topK": [50, 100],
    "feature_mode": ["original_plus_patterns"],
    "augmented_pair_transforms": [False],
    "interaction_relaxed_mining": [True],
    "interaction_relaxed_feature_size": [8, 12],
}

# Originals plus patterns with augmented pair features.
originals_augmented_grid = {
    "B": [-1],
    "adaptive_binning": [True],
    "L": [2],
    "G": [1e-2, 5e-3],
    "topK": [50, 100],
    "feature_mode": ["original_plus_patterns"],
    "augmented_pair_transforms": [True],
    "augmented_pair_mode": ["interaction_information"],
    "aug_feature_size": [8, 12],
}

fast_grid_tune() — single-split cached path for custom CV loops

tune_result = HUGIMLClassifier.fast_grid_tune(
    X_train, y_train,
    X_val,   y_val,
    param_grid="performance",
    scoring="roc_auc",
    refit_full=False,
)

print(tune_result["best_params"])
print(f"Validation score: {tune_result['best_score']:.4f}")

Governance Studio Dashboard

The HUGIML Governance Studio is an interactive Streamlit dashboard for preparing model runs, comparing candidate models, reviewing model evidence, and producing governance-ready summaries. It keeps the existing Workbench/Governance layout and exposes evidence views for adaptive binning, interaction-relaxed mining, augmented pairs, feature families, pattern coverage, monitoring, and validation review.

Installation

pip install "hugiml-core[dashboard]"

The dashboard extra includes the UI and plotting dependencies used by the Governance Studio experience.

Launch

# Installed console script
hugiml-dashboard

# Pass Streamlit or dashboard arguments after the separator
hugiml-dashboard -- --cv 5 --random-state 42

# Source-tree development
python -m streamlit run src/hugiml/dashboard/app.py

When installed, hugiml-dashboard starts the packaged Streamlit app automatically, so you do not need to know the source file location.

What is included

Area What it supports
Workbench Demo data or uploaded tabular data, target and column-role setup, candidate run configuration, model comparison, and drill-down review
Governance Evidence summaries, validation results, representation review, adaptive-binning and augmented-pair evidence, feature-family review, pattern coverage, case-level explanations, data quality checks, policy review, monitoring signals, and model-card-oriented outputs

Evidence views

View What it shows
Overview Dataset summary, active configuration, validation score, feature mode, and top evidence
Validation Cross-validation metrics, fold-level results, and calibration-oriented review
Representation Audit Original features, HUG patterns, augmented pairs, binary indicators, feature-family provenance, and complexity budget
Pattern Inventory Pattern table with coefficients, support, utility, information gain, review filters, and population coverage
Case Review Row-level predictions, probabilities, active pattern evidence, and explanation details
Data Quality & Policy Missingness review, sensitive/proxy column checks, and policy-oriented notes
Configuration Comparison Side-by-side comparison across HUGIML settings and optional baseline models
Representation Pruning Interactive removal of original features or representation columns with re-evaluation
Monitoring PSI and KL-divergence drift signals across fitted training baselines and review data

Data sources

  • Demo datasets — built-in examples for dashboard exploration without uploading data.
  • Upload — CSV, TSV, Excel (.xlsx/.xls), or Parquet files. The sidebar lets you choose the target, ID, protected/sensitive, date, numeric, categorical, and excluded columns before fitting.

Binary indicators

Numeric two-value columns are treated as categorical indicators during HUGIML preparation, so encoded flags remain visible as discrete evidence in the dashboard instead of being shown as numeric intervals.

Demo preview


LLM Assistant

The HUGIML LLM Assistant is a chat-first Streamlit interface for working with HUGIML models and documentation from one place. It is designed for model-review workflows where a user asks natural-language questions, sees a structured answer, and then continues with follow-up questions in the same review session.

Unlike a general chatbot, the assistant grounds its responses in two HUGIML-specific sources:

Query type Grounding source Typical output
Model/run questions The fitted classifier, run artifacts, pattern inventory, validation metrics, pruning state, and governance metadata Decision-maker summaries, key model drivers, risk/audit checks, and next actions
API/documentation questions The local Sphinx/API documentation index, with fallback to package README and public API docstrings Condensed API guidance, usage examples, parameter notes, and caveats
Mixed questions Both model artifacts and documentation Practical guidance that explains what the current model shows and which API or governance action applies

Launch

# Installed console script
hugiml-llm

# Source-tree development
python -m streamlit run src/hugiml/llm/ui_app.py

The UI uses a single Q&A-style view. The session appears as a chronological transcript, the follow-up input stays inline below the latest answer, and quick model/run details are available in expandable panes below the chat.

Local model policy

The assistant can run in deterministic mode, or use a supported local Ollama model as a writer/synthesis layer over retrieved HUGIML evidence. The default policy favours small models that work on modest local machines, while still exposing larger configured models when they are installed and memory is available.

Role Model Purpose
Default qwen3:1.7b Primary local writer for grounded API and model-review answers
Light mode gemma3:1b Lower-memory local response generation
Fallback llama3.2:1b Retry model before falling back to deterministic routing
Deterministic router built in Always available when no supported local model is selected or available

Additional configured Ollama models remain visible in the selector when available:

Model Profile
llama3.2:3b Minimum larger LLM
qwen3:4b Balanced local LLM
gemma3:4b Balanced alternative
qwen3:8b Stronger local LLM
gemma3:12b Large-context local LLM

The selector uses available system memory as a safety check. These thresholds are available-RAM requirements, not model file sizes:

Model Role / profile Minimum available RAM
qwen3:1.7b Default local LLM 5.0 GB
gemma3:1b Light mode 3.5 GB
llama3.2:1b Fallback before deterministic routing 3.5 GB
llama3.2:3b Minimum larger LLM 6.0 GB
qwen3:4b Balanced local LLM 10.0 GB
gemma3:4b Balanced alternative 10.0 GB
qwen3:8b Stronger local LLM 16.0 GB
gemma3:12b Large-context local LLM 32.0 GB

Supported tiny models are shown explicitly. Other extra small Ollama models that are not part of the configured policy are omitted from the selector, while deterministic routing remains available without Ollama.

Install the preferred Ollama models before launching the assistant:

ollama pull qwen3:1.7b
ollama pull gemma3:1b
ollama pull llama3.2:1b

Optional larger models can also be installed and selected when the local machine has enough available memory:

ollama pull llama3.2:3b
ollama pull qwen3:4b
ollama pull gemma3:4b
ollama pull qwen3:8b
ollama pull gemma3:12b

Only configured supported models are listed in the UI. Extra experimental or unsupported small Ollama models that may be installed locally are omitted from the selector.

Documentation-aware answers

For API questions such as “what hyperparameters should I tune?”, “how does pruning work?”, or “what governance artifacts are created?”, the deterministic runner builds a lightweight local index over the Sphinx documentation and package docs. It retrieves the relevant sections internally, then presents a concise, structured response with the API details needed to act.

Model-review answers

For run-specific questions such as “summarize findings”, “what changed after pruning?”, or “is this model ready for governance review?”, the assistant analyzes the active classifier outputs and produces a structured decision summary covering performance, main evidence drivers, interpretability constraints, audit concerns, and recommended next steps.

Static examples

These GitHub Pages examples show the intended Q&A-style review flow:


Augmented Pair Features

For interaction-oriented models, HUGIML can add native augmented-pair features to the downstream estimator. These are continuous product or absolute-difference transforms built from informative numeric features, for example:

glucose * bmi
abs(age - duration)

They are active when L > 1, adaptive_binning=True, and augmented_pair_transforms=True (the default). They are appended only to the downstream estimator; the mined HUG pattern matrix and transform(X) remain pattern-space APIs.

The default augmented_pair_mode="interaction_information" scores candidate source columns using pair context before building product, absolute-difference, sum, and signed-difference features. Set augmented_pair_mode="marginal_ig" to use the v1.1.11 marginal-information-gain source selection behavior. aug_feature_size controls how many source columns are retained in interaction-information mode; ii_partner_size optionally bounds partner search; max_pair_features controls the source budget for marginal-IG mode.

clf = HUGIMLClassifier(
    B=-1,
    adaptive_binning=True,
    L=2,
    topK=50,
    G=1e-2,
    feature_mode="original_plus_patterns",
    augmented_pair_transforms=True,
    augmented_pair_mode="interaction_information",
    aug_feature_size=10,
    topk_budget_strict=True,
)
clf.fit(X_train, y_train)

print(clf.get_model_composition())
print(clf.explain_augmented_pair_effects())

For selected pair features, HUGIML reports the raw formula, standardized formula, observed-row coverage, missing-pair policy, and raw-scale coefficient interpretation.


Adaptive Binning

The global B parameter controls how many quantile bins each numerical feature is discretised into. Adaptive binning selects the optimal bin count per feature via supervised information-gain search and elbow stopping. For larger datasets, adaptive_binning_sample_frac can choose bin counts from a deterministic stratified row sample, then apply the selected bin edges to the full training data.

from hugiml.adaptive import HUGIMLAdaptive

clf = HUGIMLAdaptive(b_candidates=[3, 5, 7, 10, 15], L=2, G=1e-2)

X_enc, y_enc = clf.prepareXy(X_df, y)
clf.fit(X_tr, y_tr)

print(clf.per_feature_b_)
clf.plot_bin_profiles()
clf.ig_heatmap()

Alternatively, enable adaptive binning directly on HUGIMLClassifier:

from hugiml import HUGIMLClassifier

clf = HUGIMLClassifier(
    adaptive_binning=True,
    b_candidates=[3, 5, 7, 10],
    min_marginal_gain_ratio=0.02,
    adaptive_binning_sample_frac=0.20,  # optional for large adaptive-binning runs
)

How it works: for each numerical feature, HUGIML evaluates information gain at candidate B values and stops when the marginal gain falls below min_marginal_gain_ratio × current_IG. This prevents blindly selecting the maximum bin count. Set adaptive_binning_sample_frac=False for full-data bin selection, or a float in (0, 1] to use a stratified sample for the selection step.


Missing Value Handling

HUGIML treats NaN and Inf values as not observed — no imputation and no special parameter are required.

How it works: numerical columns are pre-binned at fit time. Non-finite cells become np.nan in the label array, and the C++ transaction builder skips them. The corresponding item is absent from the transaction. Patterns requiring that feature do not fire for that row.

import numpy as np
from hugiml import HUGIMLClassifier

X_train.iloc[5, 2] = np.nan

clf = HUGIMLClassifier(B=5, L=2, G=1e-4)
clf.fit(X_train, y_train)

X_test.iloc[0, 0] = np.nan
proba = clf.predict_proba(X_test)       # scored using available feature items

Mining Patterns About Missingness

To mine patterns that involve missingness (e.g., Glucose_MISSING=1 AND HeartRate=[110,140]), add binary missingness indicators as preprocessing features:

def add_missingness_indicators(X, threshold=0.05):
    X_aug = X.copy()
    for col in X.columns:
        if X[col].isna().mean() > threshold:
            X_aug[f"{col}__MISSING"] = X[col].isna().astype(int)
    return X_aug

X_with_indicators = add_missingness_indicators(X_raw)
clf = HUGIMLClassifier(B=7, L=2, G=1e-4)
clf.fit(X_with_indicators, y)

The Governance Studio Data Quality & Policy view shows feature-level missingness rates alongside sensitive column review.


Model Explanation and Visualisations

Interactive Plotly dashboard

from hugiml.plots import HUGPlotter

plotter = HUGPlotter(clf)

plotter.plot_dashboard(
    X_test,
    dataset_name="My Dataset",
    feature_names_for_profile=["age", "income", "glucose"],
    output_path="hugiml_dashboard.html",
)

plotter.plot_marginal_bin_profile("glucose", X=X_test).show()
plotter.plot_top_patterns(top_n=20).show()
plotter.plot_feature_importance(top_n=15).show()
plotter.plot_active_patterns(X_test, sample_idx=0).show()

Each profile panel shows the learned bin/pattern behavior for a feature: utility or coefficient-like contribution per bin, with support overlay where available.

Existing example dashboards:

Public tabular benchmark classification Feature shape profiles — public tabular benchmark

Credit risk scoring Feature shape profiles — credit risk

The static benchmark dashboard is reproducible from the repository source with experiments/benchmark/benchmark_dashboard.py; see Benchmark Suite for the exact rerun and assemble commands.

Profile visualisations

plotter.plot_marginal_bin_profile("age", X=X_test).show()  # EBM-style 1-D shape function
plotter.plot_feature_combinations("age").show()             # Feature-combination view
plotter.plot_top_patterns(top_n=20).show()                  # Top patterns by importance
plotter.plot_active_patterns(X_test, sample_idx=0).show()   # Local explanation for one sample

Native Mining Pruning Controls

This section covers native HUIM search pruning, which is different from the user-facing Pattern Pruning workflow below. Native pruning controls how the C++ miner avoids unnecessary candidate work during fit() while preserving the same public model outputs.

Pruning path When it is active What it does User-facing controls
LIU Active for compound-pattern mining (L > 1, including the L=2 hot path). Bounded classifier mining uses exact candidate evidence before raising the utility floor. Raises the utility threshold from locally strong candidate sequences so low-utility branches can be skipped earlier. Tune L, G, and topK; there is no separate public LIU switch.
LA Active during generic utility-list child construction when the current branch uses the ordinary utility-ranked path. Stops building a child utility list once the remaining upper bound can no longer pass the current utility floor. Tune topK and G; relaxed-root interaction branches bypass this utility-floor shortcut where needed.
EUCS Considered for L > 1 after the admitted item set is known. It is skipped for small, dense, or very wide pair spaces where the cache would not pay off. Builds a pair co-occurrence utility cache and skips pair intersections whose pair-level utility cannot enter the retained set. Environment variables below.

EUCS is enabled by default for eligible L > 1 native mining paths, but it has safety gates so small or dense workloads continue without the extra cache. The relevant environment variables are:

Variable Default Meaning
HUGIML_EUCS_ENABLE or HUGIML_EUCS_ENABLED enabled Set to 0, false, no, off, disable, or disabled to disable EUCS. Set to 1, true, yes, on, enable, or enabled to enable it. Invalid values keep the default.
HUGIML_EUCS_MIN_ITEMS 32 EUCS is skipped when the admitted item universe is this size or smaller.
HUGIML_EUCS_MAX_CELLS 6000000 Maximum pair-cache cells allowed before EUCS is skipped.
HUGIML_EUCS_MAX_DENSITY 0.20 Maximum observed active-item density allowed before EUCS is skipped.

Typical users should leave these settings at their defaults and tune model-level parameters first: L for maximum pattern length, G for the information-gain gate, and topK for the retained pattern budget. EUCS controls are mainly useful when benchmarking native mining behavior or diagnosing a workload whose pair space is unusually sparse or dense.


Pattern Pruning

In regulated domains, analysts often need to remove patterns that reference protected attributes, have high PSI, or are operationally invalid. HUGIML provides a controlled editing workflow with a JSON audit trail.

from hugiml.pruning import PatternEditor

editor = PatternEditor(clf, operator_name="risk-team")

print(editor.list_patterns().head(10))

editor.remove([3, 7], reason="references protected attribute 'gender'")
editor.remove_by_keyword("postcode", reason="high PSI — unstable feature")
editor.remove_low_support(min_support=0.01, reason="noise patterns")

editor.refit(X_tr, y_tr)
editor.calibrate(X_cal, y_cal, method="isotonic")

new_clf = editor.finalize()
print(editor.audit_report())

The Representation Pruning view in the Governance Studio provides an interactive version of this workflow without writing code.


Interpretability Metrics

from hugiml.metrics import compute_all_metrics

m = compute_all_metrics(clf, X_test)
print(m)

Example output:

InterpretabilityMetrics
==========================================
n_patterns              : 87
avg_pattern_length       : 1.34
coverage                 : 0.9812
mean_active_patterns     : 6.21
overlap_rate             : 0.0714
explanation_sparsity     : 0.0230

top-k cumulative |coef|:
top- 1 : 8.4%
top- 5 : 31.2%
top-10 : 54.7%

Multiclass, Imbalanced Data, High-Cardinality

Multiclass Classification

from hugiml.multiclass import MulticlassHUGReport

report = MulticlassHUGReport(clf)
print(report.importances_for_class(class_label=2, top_n=10))
print(report.summary())

Imbalanced Data Handling

from hugiml.multiclass import make_imbalanced_pipeline

clf_bal = make_imbalanced_pipeline(clf_proto, strategy="smote")
clf_bal.fit(X_tr, y_tr)

High-Cardinality Categorical Reduction

When categorical features have hundreds or thousands of unique values (ZIP codes, ICD-10 diagnoses, merchant IDs), grouping rare categories prevents combinatorial explosion in pattern mining:

def reduce_high_cardinality(X, y, threshold=50, min_frequency=0.01):
    """Group rare categories (<min_frequency) as '__OTHER__' for high-cardinality columns."""
    X_reduced = X.copy()
    for col in X.select_dtypes(include=["object", "category"]).columns:
        if X[col].nunique() <= threshold:
            continue
        value_counts = X[col].value_counts()
        min_count = len(X) * min_frequency
        rare_categories = value_counts[value_counts < min_count].index
        X_reduced[col] = X[col].apply(
            lambda x: "__OTHER__" if x in rare_categories else x
        )
    return X_reduced

X_reduced = reduce_high_cardinality(X_raw, y, threshold=50, min_frequency=0.01)
clf = HUGIMLClassifier(B=7, L=2, G=1e-4)
clf.fit(X_reduced, y)

# Or use built-in target encoding:
from hugiml.multiclass import encode_high_cardinality, apply_encoding
X_enc, enc_map = encode_high_cardinality(X_tr, y_tr, threshold=20, method="target_mean")
X_te_enc = apply_encoding(X_te, enc_map)

Note: Learn category groupings on training data only, then apply the same mapping to test/production data.


Drift Detection & Monitoring

clf.enable_monitoring(window_size=1000)

clf.predict_proba(X_new)

print(clf.monitor.report())

report = clf.detect_drift(X_new, current_labels=y_new)
print(report)

The Monitoring view in the Governance Studio shows PSI and KL-divergence drift signals per feature from the fitted model's training baseline.


Calibration

from hugiml.calibration import evaluate_calibration

result = evaluate_calibration(y_te.values, proba[:, 1])

print(f"ECE: {result.ece:.4f}")
print(f"Brier: {result.brier_score:.4f}")

Serialisation

from hugiml.serialization import save_model, load_model, generate_sbom

save_model(clf, "model.hugiml")
clf2 = load_model("model.hugiml")

sbom = generate_sbom(clf)

Governance & Model Cards

from hugiml.governance import generate_model_card

card = generate_model_card(
    clf,
    model_id="credit-scorer-v1.0.0",
    intended_use="Credit risk assessment for SME lending.",
    training_data_description="German Credit dataset, 1000 samples",
)

print(card.to_markdown())
card.save("model_card.json")

Model cards should include top positive/negative patterns, missing-value behavior, calibration metrics, drift-monitoring plan, and any pattern-pruning audit trail.

The Governance Studio dashboard provides interactive governance evidence views that complement programmatic model cards with visual audit artifacts.


Benchmark Suite

HUGIML includes two reproducible benchmark workflows:

  1. Package benchmark runner for quick CV-style comparisons from the installed package.
  2. Experiment dashboard runners in experiments/ for regenerating the published static benchmark and scalability dashboards.

The package-level runner is useful for ad hoc benchmark checks:

# Run full CV comparison
python -m hugiml.benchmarks.runner

# Specific datasets
python -m hugiml.benchmarks.runner --datasets german_credit pima adult

# Save results
python -m hugiml.benchmarks.runner --output benchmarks/results/

Or use the installed console script:

hugiml-bench --datasets german_credit --output results/

Reproduce the benchmark analysis dashboard

The public benchmark analysis dashboard is generated from experiments/benchmark/benchmark_dashboard.py. This script defines the 50-dataset panel, model grids, preprocessing policy, checkpointing, result aggregation, and static HTML assembly used for the dashboard.

From the repository root:

# Full fresh run; writes checkpoint, CSV summaries, and revised HTML
python experiments/benchmark/benchmark_dashboard.py --fresh

# Resume a partially completed run from checkpoint
python experiments/benchmark/benchmark_dashboard.py --resume

# Rebuild only the HTML/CSV summaries from an existing checkpoint
python experiments/benchmark/benchmark_dashboard.py --assemble

Default outputs are written under:

experiments/benchmark/results/

The dashboard runner is deterministic for a fixed code version and dependency environment: dataset generation, train/validation/test splits, row subsampling, and model seeds are all controlled by the script. The generated artifacts include details.csv, summary_by_scope.csv, scope_tests.csv, overall.csv, and hugiml_benchmark_analysis_dashboard_revised.html.

Scalability dashboard

For runtime and memory scaling evidence, see the static scalability dashboard:

The dashboard summarizes measured fit time, prediction latency, memory delta, pattern counts, and test AUC against XGBoost and LightGBM. It covers sample-size scaling, feature-count scaling, and parameter sweeps over B, G, topK, L, and adaptive binning. HUGIML retains many training and test artifacts to support governance and audit requirements.

The scalability dashboard is reproducible from experiments/scalability/scalability_dashboard.py:

# Full scalability run with checkpointing
python experiments/scalability/scalability_dashboard.py --fresh

# Resume a partially completed scalability run
python experiments/scalability/scalability_dashboard.py --resume

# Rebuild only the static dashboard from an existing checkpoint
python experiments/scalability/scalability_dashboard.py --assemble

# Assemble with a privacy-sanitized reproducibility/SBOM manifest embedded in Methodology
python experiments/scalability/scalability_dashboard.py --assemble --include-sbom

Default outputs are written under the scalability results directory configured by the script and include the JSON checkpoint, flat CSV export, and hugiml_scalability_dashboard.html. With --include-sbom, assembly also writes scalability_reproducibility_sbom.json and embeds the same privacy-sanitized manifest as a collapsed block under the Methodology tab.

Worked notebooks in notebooks/ are organized as 12 self-contained folders:

Folder Notebook Brief description
00_quickstart nb00_pattern_explanation_walkthrough.ipynb Quick end-to-end walkthrough of fitting HUGIML, extracting patterns, and reading pattern-level explanations.
01_benchmark_baselines nb01_benchmark_baselines.ipynb Benchmark comparison across HUGIML and common tabular baselines such as XGBoost, LightGBM, Random Forest, and logistic regression.
02_hug_vs_ebm nb02_hug_vs_ebm.ipynb Side-by-side comparison of HUGIML pattern profiles and EBM-style additive shape functions.
03_modeling_special_cases nb03_modeling_special_cases.ipynb Practical modeling cases including multiclass targets, imbalance, high-cardinality categoricals, adaptive binning, and pruning workflows.
04_credit_risk nb04_credit_risk.ipynb Credit-risk governance example using German Credit-style data, scorecard-style features, and auditable risk patterns.
05_aml nb05_aml.ipynb Anti-money-laundering example focused on suspicious transaction pattern discovery and model review artifacts.
06_mobile_money nb06_mobile_money_fraud.ipynb Mobile-money fraud example showing compact transaction-risk patterns and operational fraud-review signals.
07_basel_ca nb07_basel_ca.ipynb Basel capital-adequacy oriented example for regulated risk analytics and explainable model validation.
08_clinical nb08_healthcare_breast_cancer.ipynb Clinical classification example using breast-cancer features to demonstrate interpretable healthcare pattern explanations.
09_insurance nb09_insurance_underwriting.ipynb Insurance underwriting example with risk-selection patterns and model-card-friendly feature narratives.
10_medicare nb10_medicare_program_integrity.ipynb Medicare program-integrity example for suspicious provider/claim behavior and audit-ready pattern summaries.
11_workforce_analytics nb11_workforce_attrition.ipynb Workforce attrition analytics example showing HR risk patterns, explanation tables, and governance-oriented summaries.

Validation Highlights

The finance panels use German Credit / HELOC-style risk features such as loan duration, credit amount, checking status, and repayment-risk signals. The healthcare panels use Pima diabetes-style features such as glucose, BMI, pregnancies, pedigree, and age.

HUGIML vs EBM shape profiles

HUGIML native shape profiles compared with EBM shape functions

EBM is excellent for smooth effect inspection; HUGIML is strong when the explanation needs to be reviewed as a set of readable thresholds and pattern contributions.

Real-world and synthetic benchmarks

Real-world credit risk benchmark comparing HUGIML, LR, XGBoost, LightGBM, Random Forest, and EBM

Synthetic non-monotonic benchmark comparing HUGIML, LR, XGBoost, LightGBM, Random Forest, and EBM

Native missing-value handling

Native missing-value schemes in HUGIML, XGBoost, LightGBM, and EBM

Model Native missing-value behavior What to monitor
HUGIML Missing numerical values are absent from the transaction. Patterns requiring that feature item do not fire. Missingness rate and activation frequency of top patterns.
XGBoost Each split learns a default route for missing values. Whether default-route behavior changes under deployment shift.
LightGBM Histogram splits learn how missing values are routed. Missing-value routing and feature missingness drift.
EBM Missing values can be modeled as a separate bin/effect. Size and sign of each missing-bin effect.

Adaptive binning

Adaptive binning benchmark against fixed bin counts

Adaptive binning is a safe default when you do not want to tune B; fixed B=5 is a useful fast baseline. For larger adaptive workflows, the sampling option reduces bin-selection memory while preserving full-data training after edges are selected; in Governance Studio this option is exposed from the Workbench Advanced configuration path.

Pattern explanations

HUGIML pattern explanations on finance and healthcare datasets

Model-card-ready artifacts

Model-card-ready HUGIML explanations

Observed benchmark results

Benchmark comparison

Model AUC (mean±std) Fit time/fold Complexity budget Remarks
HUG B=3 0.9907 ± 0.0031 0.32 s topK patterns topK is an explicit cap; actual mined patterns can be lower.
HUG B=5 0.9909 ± 0.0028 0.34 s topK patterns More bins per feature.
HUG adaptive 0.9954 ± 0.0022 1.20 s topK patterns Per-feature B increases fit time.
EBM 0.9940 ± 0.0025 11.0 s Additive terms + interactions Reference interpretable baseline.
XGBoost 0.9882 ± 0.0040 0.12 s Trees × leaves High-performing ensemble; not directly pattern-interpretable.
LightGBM 0.9921 ± 0.0028 0.07 s Leaves × trees Fast histogram boosting.

Complexity budget

topK is the feature-selection budget K. It caps each selected feature family before the final estimator is built, unless topk_budget_strict=True is used to apply one global cap. The effective downstream width D can be lower than these limits when fewer valid features are mined or selected.

Configuration Downstream feature budget when topK = K
patterns_only, L = 1 Up to K HUG pattern features.
patterns_only, L > 1, interaction_relaxed_mining=True Up to K HUG pattern features. The mining search may admit up to interaction_relaxed_feature_size interaction-information survivor source columns, but no extra downstream feature family is added.
patterns_only, L > 1, augmented pairs enabled Up to K HUG pattern features + up to K augmented-pair features, so D ≤ 2K.
original_plus_patterns, L = 1 Up to K selected original features + up to K HUG pattern features, so D ≤ 2K.
original_plus_patterns, L > 1, interaction_relaxed_mining=True Up to K selected original features + up to K HUG pattern features, so D ≤ 2K. Survivor-led mining affects which patterns are available, not the number of downstream feature families.
original_plus_patterns, L > 1, augmented pairs enabled Up to K selected original features + up to K HUG pattern features + up to K augmented-pair features, so D ≤ 3K.
original_plus_interactions Original features are capped at K; retained interaction/pattern features are also bounded by the HUG pattern budget. With augmented pairs enabled, the same additional K augmented-pair cap applies.
topk_budget_strict=True HUGIML first avoids oversized family blocks, then applies one global TopK selection across the constructed original, pattern, and augmented-pair candidates, so final D ≤ K.

Feature-family budgets

topK defines the per-family selection budget used by HUGIML when constructing downstream representations. A configuration may include one, two, or three selected feature families:

  • HUG pattern features
  • selected original input features
  • augmented-pair features, when enabled for higher-order configurations

Interaction-relaxed mining changes the native search path but does not add a separate downstream feature family; its budget is interaction_relaxed_feature_size, which controls survivor-source admission before pattern mining.

Each active family can contribute up to topK downstream columns before strict global selection. Therefore, the maximum downstream width is the number of active selected families multiplied by topK:

  • one active family: up to topK columns
  • two active families: up to 2 × topK columns
  • three active families: up to 3 × topK columns

For example, with topK=150, original_plus_patterns at L=1 can retain up to 150 selected original columns and up to 150 HUG pattern columns, for a maximum downstream width of 300. With L>1 and interaction_relaxed_mining=True, the same downstream width bound remains 300; the relaxed path affects pattern discovery rather than adding feature columns. With L>1 and augmented-pair transforms enabled, the same configuration can retain up to 150 selected original columns, 150 HUG pattern columns, and 150 augmented-pair columns, for a maximum downstream width of 450. When topk_budget_strict=True, HUGIML applies one final global TopK selection across the constructed downstream candidates, so the final downstream width is capped at topK.

With strict budgeting enabled, HUGIML applies the TopK budget during feature construction rather than after building a full expanded matrix. This keeps the practical downstream width bounded and avoids large intermediate matrices. In hybrid modes, original features are scored and preselected before prediction-time preparation, so prediction prepares only the retained original columns.

Missing value robustness

Missing value benchmark


Capabilities Summary

Capability Details
HUG pattern mining C++ accelerated via pybind11; optional OpenMP parallelism
scikit-learn API Full BaseEstimator / ClassifierMixin compliance
Mixed feature types Integer, float, categorical — auto-detected or explicitly supplied
Feature modes Pattern-only, original-plus-patterns, original-plus-interactions, augmented-pair downstream features
Fast hyperparameter search Cached adaptive-binning grid; mining runs once per unique (G, L, topK) group
Governance Studio Multi-view Streamlit dashboard with audit evidence views and upload support
Profile visualisations EBM-style 1-D/2-D HUG profiles, active-pattern explanations, coefficient-support views (Plotly)
Interpretability metrics Pattern count, coverage, overlap, sparsity, top-k cumulative contribution
Adaptive binning Per-feature supervised B selection with optional stratified sampling — addresses the B-sensitivity trap
Pattern pruning Regulated remove/refit/calibrate workflow with full JSON audit trail
Multiclass & imbalance Multiclass report, SMOTE/class-weight pipeline, high-cardinality encoding
Benchmark suite Reproducible CV comparison and dashboard regeneration via experiments/benchmark/benchmark_dashboard.py
Scalability dashboard Static runtime, latency, memory, n-scaling, p-scaling, and parameter-sweep evidence reproducible via experiments/scalability/scalability_dashboard.py
Calibration ECE, MCE, Brier score, reliability diagram data
Drift detection PSI + symmetric KL divergence + label drift
Monitoring Thread-safe PredictionMonitor, latency tracking
Governance Model cards (JSON + Markdown), audit artifacts, SBOM
Observability OpenTelemetry tracing, Prometheus metrics (both optional)
Secure serialisation Allowlist-based _RestrictedUnpickler, versioned schema
Deployment FastAPI inference server, Docker image, Kubernetes manifests
CI/CD GitHub Actions: lint → coverage → native tests → wheels → PyPI

Inference Server

A FastAPI-based inference server is included for containerised deployments.

docker build -t hugiml-core:latest -f docker/Dockerfile .

docker run -p 8080:8080 -v /path/to/models:/models hugiml-core:latest

curl -s -X POST http://localhost:8080/predict \
  -H "Content-Type: application/json" \
  -d '{"instances": [{"age": 35, "savings": "moderate"}]}'

Kubernetes manifests are in kubernetes/deployment.yaml.


CI / CD

Workflow Trigger What it does
ci.yml Every push / PR Lint, type-check, coverage gate, native tests, sanitizer build, benchmark regression, wheel build
release.yml Git tag v*.*.* Build platform wheels, generate SBOM, publish to PyPI, create GitHub release

Repository Structure

hugiml-core/
├── native/                      C++ extension sources
├── src/
│   └── hugiml/
│       ├── classifier.py        HUGIMLClassifier / HUGIMLClassifierNative
│       ├── calibration.py       ECE, Brier, reliability diagrams
│       ├── explainability.py    SHAP bridge, feature lineage, stability
│       ├── governance.py        Model cards, audit artifacts
│       ├── monitoring.py        PredictionMonitor, DriftDetector
│       ├── serialization.py     save/load, SBOM, restricted unpickler
│       ├── telemetry.py         OpenTelemetry, Prometheus
│       ├── exceptions.py        Exception hierarchy
│       ├── metrics.py           Interpretability-complexity metrics
│       ├── plots.py             EBM-style profile visualisations
│       ├── pruning.py           Pattern editor + audit trail
│       ├── adaptive.py          Per-feature adaptive binning
│       ├── multiclass.py        Multiclass / imbalanced / encoding
│       ├── dashboard/           Governance Studio Streamlit application
│       │   ├── app.py           Entry point (hugiml-dashboard console script)
│       │   ├── runner.py        Model training and scoring helpers
│       │   ├── components/      Individual evidence-view renderers
│       │   └── ...
│       ├── llm/                 LLM Assistant package and runtime
│       │   ├── cli.py           hugiml-llm console entry point
│       │   ├── ui_app.py        Streamlit chat interface
│       │   ├── orchestrator.py  Evidence routing and answer assembly
│       │   ├── docs_index.py    Local documentation index
│       │   ├── assets/          Packaged configs and demo datasets
│       │   └── ...
│       └── benchmarks/          CV comparison suite
├── LLM/                         Static assistant assets and GitHub Pages examples
│   ├── config/                  Assistant model policy config
│   ├── datasets/                Built-in and user dataset folders
│   ├── examples/                Static demo pages and source data
│   ├── prompts/                 Assistant prompt templates
│   └── ui/                      Standalone chat UI helpers
├── notebooks/                   Worked examples (12 domain folders)
├── tests/                       Pytest suite
│   ├── dashboard/               Dashboard component tests
│   └── llm/                     Optional LLM Assistant tests
├── benchmarks/                  Micro-benchmarks and regression gate
├── experiments/                 Reproducible dashboard-generation workflows
│   ├── benchmark/               Benchmark analysis runner, checkpointing, CSV summaries, HTML assembly
│   └── scalability/             Scalability runner, checkpointing, flat exports, HTML assembly
├── docker/                      Dockerfile + FastAPI inference server
├── kubernetes/                  Deployment manifests
├── scripts/                     Build and utility scripts
├── docs/                        Sphinx documentation, LLM assistant docs, and model-card templates
├── .github/workflows/           CI/CD pipelines
├── pyproject.toml
└── setup.py

License

Apache License 2.0 — see LICENSE.


Citation

If you use hugiml-core in research or commercial work, please cite:

@article{krishnamoorthy2026interpretability,
  title        = {Interpretability Myopia: Governance Fitness in Financial Risk Models},
  author       = {Krishnamoorthy, Srikumar},
  journal      = {SSRN Electronic Journal},
  year         = {2026},
  doi          = {10.2139/ssrn.6821418},
  url          = {https://dx.doi.org/10.2139/ssrn.6821418},
  keywords     = {Interpretable machine learning, analytics, financial risk governance, deployment evaluation, regulatory compliance, model risk management}
}

@article{krishnamoorthy2024hugIML,
  author  = {Krishnamoorthy, Srikumar},
  title   = {Interpretable Classifier Models for Decision Support Using High Utility Gain Patterns},
  journal = {IEEE Access},
  volume  = {12},
  pages   = {126088--126107},
  year    = {2024},
  doi     = {10.1109/ACCESS.2024.3455563}
}

Project details


Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hugiml_core-1.1.17.tar.gz (688.2 kB view details)

Uploaded Source

Built Distributions

If you're not sure about the file name format, learn more about wheel file names.

hugiml_core-1.1.17-cp313-cp313-win_amd64.whl (820.7 kB view details)

Uploaded CPython 3.13Windows x86-64

hugiml_core-1.1.17-cp313-cp313-manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.13manylinux: glibc 2.28+ x86-64

hugiml_core-1.1.17-cp313-cp313-macosx_15_0_x86_64.whl (1.8 MB view details)

Uploaded CPython 3.13macOS 15.0+ x86-64

hugiml_core-1.1.17-cp313-cp313-macosx_15_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.13macOS 15.0+ ARM64

hugiml_core-1.1.17-cp312-cp312-win_amd64.whl (820.7 kB view details)

Uploaded CPython 3.12Windows x86-64

hugiml_core-1.1.17-cp312-cp312-manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.12manylinux: glibc 2.28+ x86-64

hugiml_core-1.1.17-cp312-cp312-macosx_15_0_x86_64.whl (1.8 MB view details)

Uploaded CPython 3.12macOS 15.0+ x86-64

hugiml_core-1.1.17-cp312-cp312-macosx_15_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.12macOS 15.0+ ARM64

hugiml_core-1.1.17-cp311-cp311-win_amd64.whl (818.0 kB view details)

Uploaded CPython 3.11Windows x86-64

hugiml_core-1.1.17-cp311-cp311-manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.11manylinux: glibc 2.28+ x86-64

hugiml_core-1.1.17-cp311-cp311-macosx_15_0_x86_64.whl (1.8 MB view details)

Uploaded CPython 3.11macOS 15.0+ x86-64

hugiml_core-1.1.17-cp311-cp311-macosx_15_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.11macOS 15.0+ ARM64

hugiml_core-1.1.17-cp310-cp310-win_amd64.whl (816.8 kB view details)

Uploaded CPython 3.10Windows x86-64

hugiml_core-1.1.17-cp310-cp310-manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.10manylinux: glibc 2.28+ x86-64

hugiml_core-1.1.17-cp310-cp310-macosx_15_0_x86_64.whl (1.8 MB view details)

Uploaded CPython 3.10macOS 15.0+ x86-64

hugiml_core-1.1.17-cp310-cp310-macosx_15_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.10macOS 15.0+ ARM64

hugiml_core-1.1.17-cp39-cp39-win_amd64.whl (816.9 kB view details)

Uploaded CPython 3.9Windows x86-64

hugiml_core-1.1.17-cp39-cp39-manylinux_2_28_x86_64.whl (1.5 MB view details)

Uploaded CPython 3.9manylinux: glibc 2.28+ x86-64

hugiml_core-1.1.17-cp39-cp39-macosx_15_0_x86_64.whl (1.8 MB view details)

Uploaded CPython 3.9macOS 15.0+ x86-64

hugiml_core-1.1.17-cp39-cp39-macosx_15_0_arm64.whl (1.7 MB view details)

Uploaded CPython 3.9macOS 15.0+ ARM64

File details

Details for the file hugiml_core-1.1.17.tar.gz.

File metadata

  • Download URL: hugiml_core-1.1.17.tar.gz
  • Upload date:
  • Size: 688.2 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for hugiml_core-1.1.17.tar.gz
Algorithm Hash digest
SHA256 38fefed5130fb6e4d8a31152600bb46e5333b97795cdda5e54ff0efcbbae64fd
MD5 c9a5027dfa7274055dc59d718ad55725
BLAKE2b-256 e1a95965a1240e96af810ddc678e8b9434e2ab991d0d59385638ac8d5bd1617d

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17.tar.gz:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp313-cp313-win_amd64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp313-cp313-win_amd64.whl
Algorithm Hash digest
SHA256 2268b476c2f8ae805a58b9984aca7ccacccd866beab711809733e1a8d633f86b
MD5 95d78a1ab83e51265fd2e9116ac8f29b
BLAKE2b-256 40504cfb79ba3b90adb70861c7d188e96f56abb232ff58d8568a6474a71e89b0

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp313-cp313-win_amd64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp313-cp313-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp313-cp313-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 b2db45e9eaf2e77a7b95070f832baf3b0c8321a6e2ce261c8ddd50f16b5d6e3d
MD5 fbdfca3982bfac56d1786142af55dae0
BLAKE2b-256 7cec56566589787ac3d1f7541d7f1403e82ba42a63ab03422245d54c80fb1802

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp313-cp313-manylinux_2_28_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp313-cp313-macosx_15_0_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp313-cp313-macosx_15_0_x86_64.whl
Algorithm Hash digest
SHA256 47e428f6c58972c0030f3f671f8721fde89d3791e8efb3e171af0f8a77d15cec
MD5 8c522e8b5a042711cfc26a44fe1fcc1c
BLAKE2b-256 d8e6263c11696ee933e43f291c8ad9e48b8733bf8712f049378c598aa88cac72

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp313-cp313-macosx_15_0_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp313-cp313-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp313-cp313-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 018f3536edba964e60a045e0cfec66a5dedce1a0a349e6a193a6d8e35c6a7d10
MD5 6cfa44f8da50819cfe5edbb39853732b
BLAKE2b-256 edb8c40d9d27021582cc0a2e54e045e4d018efb6147a3bd70a2007f33eaed94f

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp313-cp313-macosx_15_0_arm64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp312-cp312-win_amd64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp312-cp312-win_amd64.whl
Algorithm Hash digest
SHA256 6613aadab14ed369b2edf9b189fb491c254f11ef87aacb596015fbf1b7a52142
MD5 d22cce3dffe5c076eeb431ea11b01b79
BLAKE2b-256 d95574906aa55053e2842e54f95fb3adf29e0e36503ee34bf1bc6ae360ed6458

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp312-cp312-win_amd64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp312-cp312-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp312-cp312-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 a271439314e0a95b23ee6f423dd6722c33ff0a0f14973a9bf66d74f56c9af8be
MD5 ee69976dc70d1ee790ad58cba76b392e
BLAKE2b-256 0689aad75e5a4e638c0896b08ae04d8b3268cccdcc989ae75ab1dd4073b355b7

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp312-cp312-manylinux_2_28_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp312-cp312-macosx_15_0_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp312-cp312-macosx_15_0_x86_64.whl
Algorithm Hash digest
SHA256 5ca2fb6094e5a8476ca0892296f2055af9d7ce00b76f2b807735f7a141761056
MD5 35ee24858e355c217b7970ad20797e3a
BLAKE2b-256 8318065613b1946a53769116eeda24eb35dddacedffaf0f3f3155019f2dd2daf

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp312-cp312-macosx_15_0_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp312-cp312-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp312-cp312-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 d14ef653535c04d107eb3c40ea64df4f25d8a041f131182b29117f6fe4409b9b
MD5 9d512a2052df0cff3d1079b6724ba87d
BLAKE2b-256 cd10a4ee47ce9b1fbafa69a9841548d0cb9e8c25c965a363cf8eef68712b7b34

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp312-cp312-macosx_15_0_arm64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp311-cp311-win_amd64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp311-cp311-win_amd64.whl
Algorithm Hash digest
SHA256 f33a8c2b31d5591608882dfce8fed96487969a86329cf010f574ecc7bcb5c422
MD5 c505c7723840be4fe302bd699185419a
BLAKE2b-256 c36518ee2e056a3d709d78235c8361668e38c74413f3d5b4eca6e6f8b537a76b

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp311-cp311-win_amd64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp311-cp311-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp311-cp311-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 0afe67b0119d7895fa3e47bb406866bc00af6b3c8cb65e392a4d12e411b5543c
MD5 8b2e8cb5045bca7d7d0433db93ef679b
BLAKE2b-256 2a6ec642293f11dbd310a8aaec573dd8ee81f08f373a84d36f7c7e4c1457c0d0

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp311-cp311-manylinux_2_28_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp311-cp311-macosx_15_0_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp311-cp311-macosx_15_0_x86_64.whl
Algorithm Hash digest
SHA256 e35454dae1a34c5446ffb911da37d8e67f6ee8e1ddfcc30a158778aca4ce2ee0
MD5 a5e093e1234d8dcf92fe367721c782b5
BLAKE2b-256 a90541b5c7e4896f10fbf0c870187c49038b9fc577eb72b365c6396881231712

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp311-cp311-macosx_15_0_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp311-cp311-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp311-cp311-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 f7daa01763ce92b4fc52388f8d07d08dc0b07e96b9740dd05f32677467d33804
MD5 f69026390441091a78f123879a53b913
BLAKE2b-256 7a307a3778ae0c3b46a716f0c1bbd73c3d2e59facc68b441cc11de49c6eaa779

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp311-cp311-macosx_15_0_arm64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp310-cp310-win_amd64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp310-cp310-win_amd64.whl
Algorithm Hash digest
SHA256 9688ebdd1c016bee23df2ab367ed28564b2f149f91a7611b2825586d3d86eba8
MD5 ffc78875914d82763d085261973ae318
BLAKE2b-256 d4c5c8800a7e41db42d24a1feb9aa51f628ebd8975c9ab17d22f712bf4f219e8

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp310-cp310-win_amd64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp310-cp310-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp310-cp310-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 7c145c7e220fc81ede615e81fe80627221d3854b8e8dbe7b36379beea33b538c
MD5 8dc92f13214ca73999626fc170117b69
BLAKE2b-256 7cb86e51b07405450a56b00a7e913f45d1ba6d7a54cdbbd338a5b199c97e0bb1

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp310-cp310-manylinux_2_28_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp310-cp310-macosx_15_0_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp310-cp310-macosx_15_0_x86_64.whl
Algorithm Hash digest
SHA256 1daba73056fe75369993e8ddfd7115c68c9df7494826f7c6798b016178e4bd6b
MD5 96bcc0caaed4cc71222a4392196037b6
BLAKE2b-256 c14ec58d9a2da297553998e4855d052821d817d1cc079395a9ef1192c0779181

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp310-cp310-macosx_15_0_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp310-cp310-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp310-cp310-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 5a193ea4fe1daec478aa4cbbcc81972cb010995a9b2f5fd7fa0fc1778c414f16
MD5 d0860bebf9bac3c29f72eeaaab6d7dcb
BLAKE2b-256 397d9dbde96cd11389d1f99dfa4148984406779d484cb61262749234003bef67

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp310-cp310-macosx_15_0_arm64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp39-cp39-win_amd64.whl.

File metadata

  • Download URL: hugiml_core-1.1.17-cp39-cp39-win_amd64.whl
  • Upload date:
  • Size: 816.9 kB
  • Tags: CPython 3.9, Windows x86-64
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.12

File hashes

Hashes for hugiml_core-1.1.17-cp39-cp39-win_amd64.whl
Algorithm Hash digest
SHA256 0acdb13d5b1467a71649a17022073c20ad6314409a68c81ed654b484205df70e
MD5 2574cffc69743c0e60f1b1a4cec13b4f
BLAKE2b-256 7d546a8e410228947818824d4c328b785410dea0db8796efaf8877499440ec85

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp39-cp39-win_amd64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp39-cp39-manylinux_2_28_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp39-cp39-manylinux_2_28_x86_64.whl
Algorithm Hash digest
SHA256 ceb5cdeeea194cba33eb0cee38048b73443816393b644d21fb8a280bc1af6680
MD5 05e211bb95ee1bf0995a2277150a49d1
BLAKE2b-256 41eb44fa1e511479c41fe67d435a8dd5d48fee0e0e81a3f3e41c42d8fd959e17

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp39-cp39-manylinux_2_28_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp39-cp39-macosx_15_0_x86_64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp39-cp39-macosx_15_0_x86_64.whl
Algorithm Hash digest
SHA256 873d45f3a009848d6a24a671cdbe869ecf5884d3b5daaefb5d1dd3a5fce2602d
MD5 cff3f048c27ce02159c4eea2c824aee8
BLAKE2b-256 cebb2f834701fdf9f1edbd92b7f6bb1757d42109373adfc99e0836ee37eedd4f

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp39-cp39-macosx_15_0_x86_64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hugiml_core-1.1.17-cp39-cp39-macosx_15_0_arm64.whl.

File metadata

File hashes

Hashes for hugiml_core-1.1.17-cp39-cp39-macosx_15_0_arm64.whl
Algorithm Hash digest
SHA256 6a381467a7002f9ae91ff7ac1f9d0dd93d0f902cbf65fa5bd1283138f49a8026
MD5 ecc9d66bb3ce2f5638ee11bc3d532a84
BLAKE2b-256 ddcc32ca6ab9e734e54a7fe05073751114a683a5b75723029393fd00fe72482c

See more details on using hashes here.

Provenance

The following attestation bundles were made for hugiml_core-1.1.17-cp39-cp39-macosx_15_0_arm64.whl:

Publisher: release.yml on srikumar2050/hugiml-core

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Supported by

AWS Cloud computing and Security Sponsor Datadog Monitoring Depot Continuous Integration Fastly CDN Google Download Analytics Pingdom Monitoring Sentry Error logging StatusPage Status page