Skip to main content

Hierarchical Genetic Programming Library

CI codecov PyPI version Python versions License Docs

A Python library for explainable rule-based classification. It evolves human-readable boolean rule trees via hierarchical genetic programming, with automatic binarization and parallel benchmarking.

Full documentation: https://fii-optim-lab.github.io/hgp-lib/

What it does

hgp_lib evolves boolean rules that classify tabular data. A rule is a tree of logical operators (And, Or) over literals, for example And(age < 50, Or(income >= 30k, employed)). Rules are readable, so a trained classifier can be inspected and explained.

The method is genetic programming. A population of candidate rules is scored against the data, the best rules are selected, and crossover and mutation produce the next generation. Over many epochs the population converges toward rules with high fitness. Hierarchical GP extends this with child populations that evolve on sampled subsets of features, then combine into larger rules.

Boolean GP operates on boolean data. Numeric and categorical columns are binarized first, so a numeric feature becomes a set of boolean bins. See Data Preparation for details.

The model is a single boolean rule, so it is readable on its own and needs no separate explanation. See Theory for how the search works and Interpretability for why this matters.

Installation

pip install hgp-lib
# or
pip install 'hgp-lib[dev]'

Quickstart

BooleanRuleClassifier is the fastest way to train an interpretable rule end to end. It binarizes the raw data for you, evolves a rule, and applies the same binarization when predicting. The example below is fully runnable on the scikit-learn breast_cancer dataset.

from sklearn.datasets import load_breast_cancer
from sklearn.model_selection import train_test_split

from hgp_lib import BooleanRuleClassifier
from hgp_lib.configs import BooleanGPConfig, TrainerConfig
from hgp_lib.utils.metrics import fast_f1_score

X, y = load_breast_cancer(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.2, stratify=y, random_state=0
)
X_train, X_val, y_train, y_val = train_test_split(
    X_train, y_train, test_size=0.25, stratify=y_train, random_state=0
)

config = TrainerConfig(
    gp_config=BooleanGPConfig(score_fn=fast_f1_score), num_epochs=1000, val_every=100
)
clf = BooleanRuleClassifier(config)  # StandardBinarizer by default; pass binarizer=... to customize
clf.fit(X_train, y_train, X_val, y_val)  # validation data is binarized internally too

predictions = clf.predict(X_test)  # raw data is binarized internally
print(clf.format_rule())           # the evolved rule as plain logic

Validation data is optional; when given, it is binarized with the same fitted binarizer and used to track a validation score during training. clf.format_rule() prints the rule with the binarized column names, so the model reads as plain logic. To binarize and train manually with GPTrainer, see Training; the Data Preparation guide shows how to avoid leaking data between splits.

Benchmarking

GPBenchmarker runs multiple independent experiments and aggregates the results. Each run takes a stratified train/test split, performs k-fold cross-validation on the training set, and evaluates the best rule on the held-out test set. Runs execute in parallel by default.

The benchmarker binarizes data internally, per fold, so you pass a raw pandas.DataFrame and skip manual binarization.

import numpy as np
from sklearn.datasets import load_breast_cancer
from hgp_lib.configs import BenchmarkerConfig, BooleanGPConfig, TrainerConfig
from hgp_lib.benchmarkers import GPBenchmarker
from hgp_lib.utils.metrics import fast_f1_score

X, y = load_breast_cancer(return_X_y=True, as_frame=True)

gp_config = BooleanGPConfig(score_fn=fast_f1_score)
trainer_config = TrainerConfig(gp_config=gp_config, num_epochs=1000, val_every=100)
config = BenchmarkerConfig(
    data=X,
    labels=y.to_numpy(),
    trainer_config=trainer_config,
    num_runs=30,
    n_folds=5,
    test_size=0.2,
    n_jobs=-1,
)
benchmarker = GPBenchmarker(config)
result = benchmarker.fit()

test_scores = result.test_scores
print(f"Test score: {np.mean(test_scores):.4f} ± {np.std(test_scores):.4f}")

# Human-readable best rule
print(result.best_rule.to_str(result.best_run.feature_names))

# sklearn-style predict on raw data (binarized internally with the best run's binarizer)
predictions = benchmarker.predict(X)

See Benchmarking for scorer optimization, custom binarizers, and the aggregated result fields.

Customizing the algorithm

The population, mutation, and crossover behavior is configured through factories passed to BooleanGPConfig. The default factories cover the common case. To use custom initialization strategies or mutations, subclass a factory and override its construction hook.

from hgp_lib.populations import PopulationGeneratorFactory

factory = PopulationGeneratorFactory(population_size=100)

The Configuring HGP guide covers the built-in factories and hierarchical GP. The Extending HGP guide covers custom strategies, mutations, and low-level use of BooleanGP directly.

Documentation

Contributing

See CONTRIBUTING.md.

Download files

Download the file for your platform. If you're not sure which to choose, learn more about installing packages.

Source Distribution

hgp_lib-1.1.2.tar.gz (114.5 kB view details)

Uploaded Source

Built Distribution

If you're not sure about the file name format, learn more about wheel file names.

hgp_lib-1.1.2-py3-none-any.whl (97.9 kB view details)

Uploaded Python 3

File details

Details for the file hgp_lib-1.1.2.tar.gz.

File metadata

  • Download URL: hgp_lib-1.1.2.tar.gz
  • Upload date:
  • Size: 114.5 kB
  • Tags: Source
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for hgp_lib-1.1.2.tar.gz
Algorithm Hash digest
SHA256 3c4f361ac8f0ceca457d3d6380ad73ee96a46d3b84324458a2fd4b9fe3f94a6c
MD5 cf4439c28eec981250588cfb7e5e1f16
BLAKE2b-256 efcb620cc54dc8929cfb581ddd0a7983a4dbb09dd152b4c5dcd2bc769fc0b8fc

See more details on using hashes here.

Provenance

The following attestation bundles were made for hgp_lib-1.1.2.tar.gz:

Publisher: python-publish.yml on fii-optim-lab/hgp-lib

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

File details

Details for the file hgp_lib-1.1.2-py3-none-any.whl.

File metadata

  • Download URL: hgp_lib-1.1.2-py3-none-any.whl
  • Upload date:
  • Size: 97.9 kB
  • Tags: Python 3
  • Uploaded using Trusted Publishing? Yes
  • Uploaded via: twine/6.1.0 CPython/3.13.14

File hashes

Hashes for hgp_lib-1.1.2-py3-none-any.whl
Algorithm Hash digest
SHA256 cbaf0063e9322c0ffc370637cd25b2951b36c422f462cfdc2d4bbfeb3f2ae8a4
MD5 cc24e65c887814b7cdcd17c78f8236b1
BLAKE2b-256 7d5c3e631a6b94b52c0863b01f6108d257e3bf6090c8b212c903e68a48617607

See more details on using hashes here.

Provenance

The following attestation bundles were made for hgp_lib-1.1.2-py3-none-any.whl:

Publisher: python-publish.yml on fii-optim-lab/hgp-lib

Attestations: Values shown here reflect the state when the release was signed and may no longer be current.

Release history Release notifications | RSS feed

1.2.2

2 files

1.2.1

2 files

1.2.0

2 files

This release

1.1.2 This release

2 files

1.1.1

2 files

1.1.0

2 files

1.0.1

2 files

1.0.0

2 files

0.0.1

2 files

Anthropic, PBC Visionary sponsor Bloomberg Visionary sponsor Hudson River Trading Visionary sponsor Meta Visionary sponsor NVIDIA Visionary sponsor Microsoft Sustainability sponsor Depot Continuous Integration AWS Cloud computing and Security Sponsor Datadog Monitoring Fastly CDN Google Download Analytics Sentry Error logging StatusPage Status page