hgp-lib: Hierarchical Genetic Programming Library for Generating Boolean Rules from Tabular Data
Project description
Hierarchical Genetic Programming Library
A Python library for explainable rule-based classification. It evolves human-readable boolean rule trees via hierarchical genetic programming, with automatic binarization and parallel benchmarking.
Full documentation: https://fii-optim-lab.github.io/hgp-lib/
What it does
hgp_lib evolves boolean rules that classify tabular data.
A rule is a tree of logical operators (And, Or) over literals, for example And(age < 50, Or(income >= 30k, employed)).
Rules are readable, so a trained classifier can be inspected and explained.
The method is genetic programming. A population of candidate rules is scored against the data, the best rules are selected, and crossover and mutation produce the next generation. Over many epochs the population converges toward rules with high fitness. Hierarchical GP extends this with child populations that evolve on sampled subsets of features, then combine into larger rules.
Boolean GP operates on boolean data. Numeric and categorical columns are binarized first, so a numeric feature becomes a set of boolean bins. See Data Preparation for details.
The model is a single boolean rule, so it is readable on its own and needs no separate explanation. See Theory for how the search works and Interpretability for why this matters.
Installation
pip install hgp-lib
Quickstart
Binarize the data, train a rule with GPTrainer, then use it to predict and print it as plain logic.
from hgp_lib.preprocessing import StandardBinarizer
from hgp_lib.configs import BooleanGPConfig, TrainerConfig
from hgp_lib.trainers import GPTrainer
from hgp_lib.utils.metrics import fast_f1_score
binarizer = StandardBinarizer(num_bins=5)
train_bin = binarizer.fit_transform(train_data, train_labels)
test_bin = binarizer.transform(test_data)
gp = BooleanGPConfig(
score_fn=fast_f1_score,
train_data=train_bin.to_numpy(),
train_labels=train_labels,
)
trainer = GPTrainer(TrainerConfig(gp_config=gp, num_epochs=1000))
history = trainer.fit()
rule = history.global_best_rule
predictions = trainer.predict(test_bin.to_numpy())
# Equivalent notation
# predictions = rule.evaluate(test_bin.to_numpy())
column_names = dict(enumerate(train_bin.columns))
print(rule.to_str(column_names))
The column_names map turns literal indices back into the binarized column names, so the printed rule reads as plain logic.
The Data Preparation guide shows how to use StandardBinarizer without leaking data between splits.
Benchmarking
GPBenchmarker runs multiple independent experiments and aggregates the results.
Each run takes a stratified train/test split, performs k-fold cross-validation on the training set, and evaluates the best rule on the held-out test set.
Runs execute in parallel by default.
The benchmarker binarizes data internally, per fold, so you pass a raw pandas.DataFrame and skip manual binarization.
import numpy as np
import pandas as pd
from hgp_lib.configs import BenchmarkerConfig, BooleanGPConfig, TrainerConfig
from hgp_lib.benchmarkers import GPBenchmarker
data = pd.DataFrame(...) # raw features (bool / categorical / numeric)
labels = np.array(...) # 1-D target array
gp_config = BooleanGPConfig(score_fn=score_fn)
trainer_config = TrainerConfig(gp_config=gp_config, num_epochs=1000, val_every=100)
config = BenchmarkerConfig(
data=data,
labels=labels,
trainer_config=trainer_config,
num_runs=30,
n_folds=5,
test_size=0.2,
n_jobs=-1,
)
benchmarker = GPBenchmarker(config)
result = benchmarker.fit()
test_scores = result.test_scores
print(f"Test score: {np.mean(test_scores):.4f} ± {np.std(test_scores):.4f}")
# Human-readable best rule
print(result.best_rule.to_str(result.best_run.feature_names))
# sklearn-style predict on raw data (binarized internally with the best run's binarizer)
predictions = benchmarker.predict(data)
See Benchmarking for scorer optimization, custom binarizers, and the aggregated result fields.
Customizing the algorithm
The population, mutation, and crossover behavior is configured through factories passed to BooleanGPConfig.
The default factories cover the common case.
To use custom initialization strategies or mutations, subclass a factory and override its construction hook.
from hgp_lib.populations import PopulationGeneratorFactory
factory = PopulationGeneratorFactory(population_size=100)
The Configuring HGP guide covers the built-in factories and hierarchical GP.
The Extending HGP guide covers custom strategies, mutations, and low-level use of BooleanGP directly.
Documentation
- Getting Started
- Theory
- Interpretability
- Data Preparation
- Training
- Benchmarking
- Configuring HGP
- Extending HGP
- Rule Trees
- Experiments
- API Reference
Contributing
See CONTRIBUTING.md.
Project details
Download files
Download the file for your platform. If you're not sure which to choose, learn more about installing packages.
Source Distribution
Built Distribution
Filter files by name, interpreter, ABI, and platform.
If you're not sure about the file name format, learn more about wheel file names.
Copy a direct link to the current filters
File details
Details for the file hgp_lib-1.0.0.tar.gz.
File metadata
- Download URL: hgp_lib-1.0.0.tar.gz
- Upload date:
- Size: 108.3 kB
- Tags: Source
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
bbd82dac7caceda477ad119d2425dc2adeb4a5de538df70141b6586e0fbc029d
|
|
| MD5 |
86475fdf8025a2401bef302b57e3a534
|
|
| BLAKE2b-256 |
4283dd2e10a702568476a052fbab9b54dad720bed808b35180158955350fab33
|
Provenance
The following attestation bundles were made for hgp_lib-1.0.0.tar.gz:
Publisher:
python-publish.yml on fii-optim-lab/hgp-lib
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hgp_lib-1.0.0.tar.gz -
Subject digest:
bbd82dac7caceda477ad119d2425dc2adeb4a5de538df70141b6586e0fbc029d - Sigstore transparency entry: 2144463499
- Sigstore integration time:
-
Permalink:
fii-optim-lab/hgp-lib@da31aea2ba7b10af7e36250518ef54cf4457c7ca -
Branch / Tag:
refs/tags/1.0.0 - Owner: https://github.com/fii-optim-lab
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@da31aea2ba7b10af7e36250518ef54cf4457c7ca -
Trigger Event:
release
-
Statement type:
File details
Details for the file hgp_lib-1.0.0-py3-none-any.whl.
File metadata
- Download URL: hgp_lib-1.0.0-py3-none-any.whl
- Upload date:
- Size: 93.8 kB
- Tags: Python 3
- Uploaded using Trusted Publishing? Yes
- Uploaded via: twine/6.1.0 CPython/3.13.12
File hashes
| Algorithm | Hash digest | |
|---|---|---|
| SHA256 |
afdc5c314a045a7e1edb3853c005a9849330cf1d1003bd099dee079ce42a3f7f
|
|
| MD5 |
35e130a47016182285eb9bcf932dfc57
|
|
| BLAKE2b-256 |
7f2b6eacf1f2e68a25b40ed7689def1781612bbde6c926bc9bd8391b8efdf7c5
|
Provenance
The following attestation bundles were made for hgp_lib-1.0.0-py3-none-any.whl:
Publisher:
python-publish.yml on fii-optim-lab/hgp-lib
-
Statement:
-
Statement type:
https://in-toto.io/Statement/v1 -
Predicate type:
https://docs.pypi.org/attestations/publish/v1 -
Subject name:
hgp_lib-1.0.0-py3-none-any.whl -
Subject digest:
afdc5c314a045a7e1edb3853c005a9849330cf1d1003bd099dee079ce42a3f7f - Sigstore transparency entry: 2144463522
- Sigstore integration time:
-
Permalink:
fii-optim-lab/hgp-lib@da31aea2ba7b10af7e36250518ef54cf4457c7ca -
Branch / Tag:
refs/tags/1.0.0 - Owner: https://github.com/fii-optim-lab
-
Access:
public
-
Token Issuer:
https://token.actions.githubusercontent.com -
Runner Environment:
github-hosted -
Publication workflow:
python-publish.yml@da31aea2ba7b10af7e36250518ef54cf4457c7ca -
Trigger Event:
release
-
Statement type: